Wheat early-stage disease monitoring model construction method and system

By domain adaptation and parameter optimization of the spectral reflectivity data of wheat samples, a high-precision model suitable for early disease monitoring of wheat gibberellia was constructed, which solved the problem of model performance degradation caused by the difference in source domain and target domain data, and achieved efficient monitoring of early diseases.

CN120408560AActive Publication Date: 2025-08-01ANHUI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510801660.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-08-01
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In the prior art, due to the large difference in data distribution between the source domain and the target domain, the early disease prediction model of wheat gibberellia has negative migration and significantly reduced performance.

Method used

By obtaining the spectral reflectivity data of wheat ear samples, dividing them into source domain and target domain data sets, sensitive band range screening is performed, feature mapping and ALA algorithm optimization parameters are used to combine with Ridge regression algorithm to build a disease monitoring model.

Benefits of technology

It improves the accuracy and mobility of the early disease monitoring model of wheat gibberellia, and solves the problems of poor generalization ability and low accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408560A_ABST
    Figure CN120408560A_ABST
Patent Text Reader

Abstract

The invention provides a wheat early-stage disease monitoring model construction method and system, and relates to the technical field of crop disease remote sensing monitoring, and the wheat early-stage disease monitoring model construction method comprises the steps: obtaining the spectral reflectivity data of a wheat ear sample and the disease severity of the spectral reflectivity data; dividing the spectral reflectivity data into a source domain data set and a target domain data set according to disease severity; wherein the source domain data set comprises first spectral reflectivity data corresponding to health and a first infection degree, and the target domain data set comprises second spectral reflectivity data corresponding to health and a second infection degree; and respectively carrying out sensitive wave band range screening on the source domain data set and the target domain data set. According to the model construction method and system provided by the invention, the precision of the early disease monitoring model can be improved to an ideal level through feature migration and intelligent dynamic parameter adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing monitoring of crop diseases, and particularly to a method and system for constructing a monitoring model for early wheat diseases. Background Art

[0002] The symptoms of wheat scab are particularly obvious in the later stage, that is, the stage of severe infection of wheat. Therefore, currently, researchers usually choose to conduct monitoring research for a period of time after the outbreak of wheat scab (usually in the later filling stage) in order to capture more significant specific characteristics, thereby improving the accuracy of the monitoring model. However, in fact, once wheat scab breaks out severely, it is very difficult to control and will cause irreversible serious harm to the grains.

[0003] However, there is relatively little research on the early monitoring of wheat scab at present. Although certain progress has been made in current disease monitoring models and methods, due to their weak generalization ability and low transferability, these models and methods still have many limitations in the application of early wheat scab diseases. For example, when the data distribution difference between the source domain and the target domain is large, the model may exhibit negative transfer phenomenon, resulting in a significant decline in the performance of predicting early wheat scab diseases. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method and system for constructing a monitoring model for early wheat diseases, which is used to solve the problem that in the prior art, due to the large difference in data distribution between the source domain and the target domain, the model may exhibit negative transfer phenomenon, resulting in a significant decline in the performance of predicting early wheat scab diseases.

[0005] To achieve the above purpose and other related purposes, the present invention provides a method for constructing a monitoring model for early wheat diseases, including: obtaining the spectral reflectance data of wheat ear samples and the disease severity of the spectral reflectance data; dividing the spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity; wherein, the source domain dataset includes the first spectral reflectance data corresponding to healthy and the first disease degree, and the target domain dataset includes the second spectral reflectance data corresponding to healthy and the second disease degree; respectively screening the sensitive band ranges of the source domain dataset and the target domain dataset to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset; constructing a source domain disease monitoring model according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance and the disease severity label.

[0006] In an embodiment of the present invention, the sensitive band ranges of the source domain dataset and the target domain dataset are respectively screened to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset, including: establishing a first partial least squares regression model according to the source domain dataset within the set band range; predicting the first disease severity prediction value through the first partial least squares regression model according to the first spectral reflectance data; obtaining the first maximum determination coefficient according to the actual value of the first disease severity corresponding to the first spectral reflectance data and the first disease severity prediction value; establishing a second partial least squares regression model according to the target domain dataset within the set band range; predicting the second disease severity prediction value through the second partial least squares regression model according to the second spectral reflectance data; obtaining the second maximum determination coefficient according to the actual value of the second disease severity corresponding to the second spectral reflectance data and the second disease severity prediction value; respectively screening the sensitive band ranges of the source domain dataset and the target domain dataset according to the first maximum determination coefficient and the second maximum determination coefficient to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

[0007] In an embodiment of the present invention, respectively screening the sensitive band ranges of the source domain dataset and the target domain dataset according to the first maximum determination coefficient and the second maximum determination coefficient to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance, including: performing sensitive band range screening within the set band range according to the first maximum determination coefficient and the second maximum determination coefficient to obtain the target sensitive band range; wherein the target sensitive band range is the same band range of the source domain dataset and the target domain dataset; selecting the sensitive wavelength reflectance according to the target sensitive band range to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

[0008] In an embodiment of the present invention, constructing a source domain disease monitoring model according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance and the disease severity label, including: performing standardization processing on the first sensitive wavelength reflectance and the second sensitive wavelength reflectance; mapping the standardized first sensitive wavelength reflectance and the second sensitive wavelength reflectance to the same high-dimensional reproducing kernel Hilbert space for feature space alignment, and according to the disease severity label, to obtain the target source domain dataset after feature space alignment; optimizing the regularization parameter of the regression model through the target source domain dataset to construct the source domain disease monitoring model.

[0009] In one embodiment of the present invention, the normalization processing of the first sensitive wavelength reflectance and the second sensitive wavelength reflectance includes: according to the first mean value of all the first sensitive wavelength reflectances and the first standard deviation of all the first sensitive wavelength reflectances, the first sensitive wavelength reflectance is normalized by the z-score method to obtain the normalized first sensitive wavelength reflectance; according to the second mean value of all the second sensitive wavelength reflectances and the second standard deviation of all the second sensitive wavelength reflectances, the second sensitive wavelength reflectance is normalized by the z-score method to obtain the normalized second sensitive wavelength reflectance.

[0010] In one embodiment of the present invention, the normalized first sensitive wavelength reflectance and the second sensitive wavelength reflectance are mapped to the same high-dimensional reproducing kernel Hilbert space for feature space alignment, and according to the disease severity label, the target source domain dataset after feature space alignment is obtained, including: merging the normalized first sensitive wavelength reflectance and the second sensitive wavelength reflectance row by row to obtain a joint matrix; obtaining the dimension information of the joint matrix according to the joint matrix; obtaining a symmetric matrix according to the dimension information and a preset weight coefficient; obtaining a centralized matrix according to the dimension information, the identity matrix and the all-ones matrix; obtaining a kernel matrix according to the normalized first sensitive wavelength reflectance, the second sensitive wavelength reflectance, a preset linear kernel function type and matrix transpose; obtaining an optimal projection matrix according to the symmetric matrix, the kernel matrix, the centralized matrix, the identity matrix and a regularization parameter; obtaining a target joint matrix of the joint matrix mapped to the reproducing kernel Hilbert space according to the optimal projection matrix, matrix transpose and kernel matrix; according to the disease severity label, the target joint matrix is segmented and reconstructed for the source domain and the target domain to obtain the target source domain dataset after feature space alignment.

[0011] In one embodiment of the present invention, the regularization parameter of the regression model is optimized by the target source domain dataset to construct a source domain disease monitoring model, including: obtaining a regression model according to the feature matrix, the disease severity label and the error of the source domain training data in the target source domain dataset; wherein, the regression model includes a regression parameter vector to be solved; obtaining a loss function of the regression model according to the regression model, the L2 regularization term and the number of samples of the source domain training data; wherein, the loss function includes a regularization parameter; dynamically adjusting the regularization parameter according to the initial search range of the regularization parameter, a random number generation function and an energy factor to construct a source domain disease monitoring model.

[0012] In an embodiment of the present invention, the regularization parameter is dynamically adjusted according to the initial search range of the regularization parameter, the random number generation function, and the energy factor to construct a source domain disease monitoring model, including: generating an initial candidate value of the regularization parameter according to the initial search range of the regularization parameter; obtaining the regression model loss value according to the initial candidate value and the loss function; when the energy factor is greater than the first set value and the generated value of the random number generation function is less than the second set value, performing migration perturbation global exploration according to the preset number of iterations to obtain the first regularization parameter; when the energy factor is greater than the first set value and the generated value of the random number generation function is greater than or equal to the second set value, performing local fine search according to the preset number of iterations to obtain the second regularization parameter; when the energy factor is less than the first set value and the generated value of the random number generation function is less than the third set value, performing spiral path density search according to the preset number of iterations to obtain the third regularization parameter; when the energy factor is less than the first set value and the generated value of the random number generation function is greater than or equal to the third set value, performing Levy flight hybrid search according to the preset number of iterations to obtain the fourth regularization parameter; obtaining the target regularization parameter corresponding to the minimum regression model loss value at the preset number of iterations according to the first regularization parameter, the second regularization parameter, the third regularization parameter, the fourth regularization parameter, and the loss function of the regression model; constructing a source domain disease monitoring model according to the target regularization parameter.

[0013] In an embodiment of the present invention, it further includes: inputting the source domain validation data in the target source domain dataset into the source domain disease monitoring model to obtain a model prediction result; fitting the model prediction result with the disease severity label to obtain the first coefficient of determination; inputting the source domain validation data in the target source domain dataset into the comparison model to obtain the second coefficient of determination; wherein, the comparison model includes a random forest model and a support vector regression model; comparing the first coefficient of determination and the second coefficient of determination; when the first coefficient of determination is greater than the second coefficient of determination, the source domain disease monitoring model is effective; when the first coefficient of determination is less than the second coefficient of determination, the source domain disease monitoring model is invalid.

[0014] To achieve the above and other related objectives, the present invention further provides a system for constructing a wheat early disease monitoring model, including: an acquisition unit, configured to acquire the spectral reflectance data of wheat ear samples and the disease severity of the spectral reflectance data; a division unit, configured to divide the spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity; wherein, the source domain dataset includes the first spectral reflectance data corresponding to health and the first disease degree, and the target domain dataset includes the second spectral reflectance data corresponding to health and the second disease degree; a selection unit, configured to respectively screen the sensitive band ranges of the source domain dataset and the target domain dataset to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset; and a modeling unit, configured to construct a source domain disease monitoring model according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance, and the disease severity label.

[0015] To achieve the above and other related objectives, the present invention further provides an electronic device, which includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the aforementioned method for constructing a wheat early disease monitoring model.

[0016] To achieve the above and other related objectives, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of the computer, the computer executes the aforementioned method for constructing a wheat early disease monitoring model.

[0017] As described above, a method and system for constructing a wheat early disease monitoring model of the present invention have the following beneficial effects: by dividing the spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity, then selecting the sensitive wavelength reflectance, and then using the TCA method to map the sample features of the source domain and the target domain with different disease degrees to a high-dimensional space to achieve the alignment of the feature distributions of the source domain and the target domain. At the same time, the ALA algorithm can be used to achieve intelligent parameter optimization, and the Ridge regression algorithm is combined to construct a source domain disease monitoring model. This method combines the domain adaptation ability of TCA, the adaptive optimization of ALA, and the regularization advantage of ridge regression, and is applicable to the modeling scenario of early disease monitoring of wheat scab, that is, it improves the early disease monitoring ability of the source domain disease monitoring model and solves the technical problems such as poor model generalization ability, low transferability, and low accuracy in the early monitoring scenario of wheat scab by traditional methods. Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of the method for constructing a wheat early disease monitoring model provided by an embodiment of the present invention.

[0019] Figure 2It shows a healthy sample diagram provided by an embodiment of the present invention.

[0020] Figure 3 It shows a relationship curve diagram between the number of healthy samples and the disease severity provided by an embodiment of the present invention.

[0021] Figure 4 It shows a mild sample diagram provided by an embodiment of the present invention.

[0022] Figure 5 It shows a relationship curve diagram between the number of mild samples and the disease severity provided by an embodiment of the present invention.

[0023] Figure 6 It shows a severe sample diagram provided by an embodiment of the present invention.

[0024] Figure 7 It shows a relationship curve diagram between the number of severe samples and the disease severity provided by an embodiment of the present invention.

[0025] Figure 8 It shows a PLSR fitting result diagram of the wavelength spectral reflectance within any wavelength band range of 350 - 1050 nm in the source domain and the disease severity provided by an embodiment of the present invention.

[0026] Figure 9 It shows a PLSR fitting result diagram of the wavelength spectral reflectance within any wavelength band range of 350 - 1050 nm in the target domain and the disease severity provided by an embodiment of the present invention.

[0027] Figure 10 It shows a VIP weight distribution diagram of the sensitive wavelengths in the source domain provided by an embodiment of the present invention.

[0028] Figure 11 It shows a VIP weight distribution diagram of the sensitive wavelengths in the target domain provided by an embodiment of the present invention.

[0029] Figure 12 It shows a source domain monitoring result diagram based on the TCA - ALA - Ridge regression model provided by an embodiment of the present invention.

[0030] Figure 13 It shows a source domain monitoring result diagram based on the SVR regression model provided by an embodiment of the present invention.

[0031] Figure 14 It shows a source domain monitoring result diagram based on the RF regression model provided by an embodiment of the present invention.

[0032] Figure 15 It shows a structural block diagram of a wheat early disease monitoring system provided by an embodiment of the present invention.

[0033] Figure 16 Shown is a schematic structural diagram of an electronic device according to an embodiment of the present invention.

[0034] Element number description

[0035] Electronic device 1; Wheat early disease monitoring model construction system 11; Memory 12; Processor 13; Acquisition unit 111; Division unit 112; Selection unit 113; Modeling unit 114. Specific implementation manners

[0036] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] It should be noted that the drawings provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0038] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0039] Please refer to Figure 1, the present invention provides a method for constructing a wheat early disease monitoring model. By using the spectral reflectance data of wheat ear samples and the disease severity of the spectral reflectance data, it is possible to divide the data into a source domain dataset and a target domain dataset according to the disease severity, so as to improve the monitoring accuracy of the source domain dataset to the monitoring accuracy of the target domain dataset; after dividing the source domain dataset and the target domain dataset, further screen the sensitive band range to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset, and perform model training based on the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset, so as to realize the construction of the source domain disease monitoring model, and can improve the monitoring accuracy of the source domain disease monitoring model to the level of the monitoring accuracy of the target domain disease monitoring model, that is, improve the accuracy of the disease early monitoring model to an ideal level, and solve the technical problems such as poor model generalization ability, low transferability, and low accuracy in the traditional method in the early monitoring scenario of wheat scab.

[0040] Figure 1 FIG. shows a flowchart of a method for constructing a wheat early disease monitoring model in an exemplary embodiment of the present application, which is applied to a wheat early disease monitoring model construction system, and includes steps S10 - step S40. The following will be combined with Figure 1 to elaborate the technical solution of the present application in detail. It should be noted that the sensitive wavelength reflectance can also be directly referred to as the sensitive wavelength in the following content.

[0041] First, execute step S10 to obtain the spectral reflectance data of wheat ear samples and the disease severity of the spectral reflectance data.

[0042] Before obtaining the spectral reflectance data of wheat ear samples and the disease severity of the spectral reflectance data, by first obtaining the diseased data of wheat ears under different disease infection degrees, the disease severity of each sample is calculated and statistically analyzed. Specifically, when calculating the disease severity of each sample, each sample can correspond to each wheat ear. Then, it is determined based on the proportion of the number of diseased grains to the total number of grains on each wheat ear. The disease severity can be set with a value range of 0 - 100% according to the severity level, and the smaller the ratio, the lighter the infection degree. At the same time, each wheat ear is also measured for the spectral reflectance data of the wheat ear sample through a spectral data measurement system, such as using the SOC710E imaging spectrometer (Surface Optics Corporation, San Diego, USA), in a completely enclosed black box environment to reduce the interference of external light on data collection. Moreover, during the measurement of the spectral reflectance data, parameters such as the height, focal length, and exposure of the spectrometer can be adjusted, and hyperspectral imaging data of all samples are taken with an exposure time of 18 ms; and the threshold method is used to segment the wheat ear samples from the background and morphological operations are used to eliminate details, and the average spectral reflectance of all pixels within the region of interest (ROI) of each sample is used as the final spectral data of each sample, that is, as the spectral reflectance data of each wheat ear sample. Then, the spectral reflectance data of the corresponding wheat ear samples and the disease severity of the spectral reflectance data are obtained through the wheat early disease monitoring model construction system.

[0043] Next, step S20 is executed to divide the spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity; among them, the source domain dataset includes the first spectral reflectance data corresponding to healthy and the first disease infection degree, and the target domain dataset includes the second spectral reflectance data corresponding to healthy and the second disease infection degree.

[0044] When the wheat early disease monitoring model construction system divides the source domain dataset and the target domain dataset, it mainly uses the spectral reflectance data as sample features and the disease severity corresponding to the spectral reflectance data as labels to divide the spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity. Moreover, when dividing, in order to ensure that the monitoring accuracy of the source domain dataset is improved to the monitoring accuracy of the target domain dataset, the disease severity labels of the source domain dataset can include healthy and the first disease infection degree, and the disease severity labels of the target domain dataset can include healthy and the second disease infection degree. Thus, the source domain dataset includes the first spectral reflectance data corresponding to healthy and the first disease infection degree, and the target domain dataset includes the second spectral reflectance data corresponding to healthy and the second disease infection degree. As Figures 2 - 7 shown by the samples and the distribution of disease severity under different disease infection degrees, it can be seen that Figures 2 - 3The wheat ear sample data obtained is a healthy sample (infection rate is 0); Figures 4 - 5 The wheat ear sample data obtained is a mild sample (0 < infection rate < 15%); Figures 6 - 7 The wheat ear sample data obtained is a severe sample (infection rate > 15%). Through the above classification method, the monitoring accuracy of the source domain dataset can be effectively improved to that of the target domain dataset, so as to provide a technical method guidance with high precision for the early disease monitoring of Fusarium head blight of wheat.

[0045] Next, perform step S30 to screen the sensitive band ranges of the source domain dataset and the target domain dataset respectively, so as to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset.

[0046] After the source domain dataset and the target domain dataset are divided, the sensitive band range is further screened through the wheat early disease monitoring model construction system, so that the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset can be obtained.

[0047] In step S30, screening the sensitive band ranges of the source domain dataset and the target domain dataset respectively to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset may further include:

[0048] Establish a first partial least squares regression model according to the source domain dataset with a set band range;

[0049] Predict the first disease severity prediction value through the first partial least squares regression model according to the first spectral reflectance data;

[0050] Obtain the first maximum determination coefficient according to the first actual disease severity value corresponding to the first spectral reflectance data and the first disease severity prediction value;

[0051] Establish a second partial least squares regression model according to the target domain dataset with a set band range;

[0052] Predict the second disease severity prediction value through the second partial least squares regression model according to the second spectral reflectance data;

[0053] Obtain the second maximum determination coefficient according to the second actual disease severity value corresponding to the second spectral reflectance data and the second disease severity prediction value;

[0054] Screen the sensitive band ranges of the source domain dataset and the target domain dataset respectively according to the first maximum determination coefficient and the second maximum determination coefficient, so as to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

[0055] In the process of extracting the first sensitive wavelength reflectance and the second sensitive wavelength reflectance through the wheat early disease monitoring model construction system, according to the set band range, based on the source domain dataset and the target domain dataset within this set band range, a first partial least squares regression model and a second partial least squares regression model are respectively established; of course, according to specific circumstances, it can also be other band ranges. Then, according to the first spectral reflectance data of the source domain dataset, the first disease severity prediction value can be predicted through the first partial least squares regression model, and then according to the second spectral reflectance data in the target domain dataset, the second disease severity prediction value can be predicted through the second partial least squares regression model. Then, based on the actual value of the first disease severity corresponding to the first spectral reflectance data, that is, the disease severity label corresponding to the first spectral reflectance data and the first disease severity prediction value, the first maximum determination coefficient can be determined, and based on the actual value of the second disease severity corresponding to the second spectral reflectance data, that is, the disease severity label corresponding to the second spectral reflectance data and the second disease severity prediction value, the second maximum determination coefficient can be determined. Then, based on the first maximum determination coefficient and the second maximum determination coefficient, the sensitive band range of the source domain dataset and the target domain dataset is determined. The sensitive band range can be the same within the source domain dataset and the target domain dataset, so that according to this same sensitive band range, the corresponding first sensitive wavelength reflectance in the source domain dataset and the second sensitive wavelength reflectance in the target domain dataset can be determined, so as to ensure that the sensitive wavelength reflectances in the datasets of the source domain and the target domain can effectively reflect the physiological state differences of wheat, and provide spectral characteristic parameters with biochemical interpretability for the subsequent construction of the TCA-ALA-Ridge model (that is, the source domain disease monitoring model).

[0056] Specifically, in the process of extracting features of the first sensitive wavelength reflectance and the second sensitive wavelength reflectance, for example, within the band range of 350 - 1050 nm, a partial least squares regression (PLSR) model of the spectral reflectance data in any interval and the disease verification degree can be established, such as the first partial least squares regression model corresponding to the source domain dataset and the second partial least squares regression model corresponding to the target domain dataset. The PLSR algorithm corresponding to the partial least squares regression (PLSR) model can handle the multicollinearity problem between the wavelength reflectances of the sample spectra at different wavelengths. It projects the original variables into a new space through projection, thereby generating new latent variables that are orthogonal and uncorrelated to each other. Among them, the partial least squares regression (PLSR) model (which can be the first partial least squares regression model or the second partial least squares regression model) can be expressed as:

[0057] X = T × P t + E

[0058] Y = T × Q t + F

[0059] Wherein, X and Y respectively represent the wavelength reflectance of the sample and the corresponding disease severity. The sample includes the sample in the source domain dataset or the sample in the target domain dataset. E and F respectively represent the projection residuals of X and Y. T represents the common latent variable matrix connecting X and Y, which is essentially a low-dimensional comprehensive feature matrix extracted from the high-dimensional spectral data X and directly associated with the prediction target Y. t represents a single latent variable (a single column vector of T), and P and Q respectively represent the weight matrices of X and Y, which are used to reveal the contribution of each band to the latent variable and establish the quantitative relationship between the latent variable and the disease severity.

[0060] Then, the maximum determination coefficient R is determined through the maximum determination coefficient calculation formula 2 :

[0061]

[0062] Wherein, y i is the actual value of the disease severity of each sample. The sample includes the sample in the source domain dataset or the sample in the target domain dataset. y' i is the predicted value of the disease severity after being processed by the PLSR algorithm, that is, the prediction target Y obtained after the above PLSR algorithm is processed. is the actual average value of the disease severity of all samples, n is the number of samples, and the maximum determination coefficient R 2 includes the first maximum determination coefficient and the second maximum determination coefficient.

[0063] Furthermore, according to the first maximum determination coefficient and the second maximum determination coefficient, the sensitive band ranges of the source domain dataset and the target domain dataset are respectively screened to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance, which may include:

[0064] According to the first maximum determination coefficient and the second maximum determination coefficient, the sensitive band range is screened within the set band range to obtain the target sensitive band range; wherein, the target sensitive band range is the same band range of the source domain dataset and the target domain dataset;

[0065] According to the target sensitive band range, the sensitive wavelength reflectance is selected to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

[0066] After obtaining the first maximum determination coefficient of the source domain dataset and the second maximum determination coefficient of the target domain dataset, the wheat early disease monitoring model construction system can screen the sensitive band ranges according to the source domain dataset and the target domain dataset, and obtain the target sensitive band range that is the same for both the source domain dataset and the target domain dataset. Then, using this target sensitive band range, the sensitive wavelength reflectances are selected in the source domain dataset and the target domain dataset respectively, so as to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

[0067] For example, when screening the sensitive band ranges for the source domain and the target domain, the sensitive wavelengths and fitting results of the source domain and the target domain can be as shown in the following table:

[0068]

[0069] As shown in the above table and Figure 8 and Figure 9 the PLSR fitting results of the wavelength spectral reflectance and the disease severity in any band range of 350 - 1050nm for the source domain and the target domain are given. It can be seen that in any interval range of the set band range of 350 - 1050nm, the PLSR fitting results of the spectral reflectance and the disease severity of the source domain and target domain samples are respectively shown. And overall, the maximum determination coefficient R 2 precision of the target domain is higher than that of the source domain. The maximum determination coefficient R 2 precision in most band intervals exceeds 0.8. The bands with the source domain precision exceeding 0.8 are mainly concentrated around 400 - 700nm and 700 - 1000nm. The maximum determination coefficient R 2 precision of the source domain reaches a maximum of 0.83. The corresponding sensitive band range at this time is between 648 - 681nm. The maximum determination coefficient R 2 precision of the target domain reaches a maximum of 0.88. The corresponding sensitive band range at this time is also between 648 - 681nm. Therefore, the sensitive band range corresponding to the maximum determination coefficient R 2 precision of the source domain and the target domain can be selected as 648 - 681nm. The first 14 sensitive wavelengths within the sensitive band range can be selected as the subsequent model input variables, that is, as the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

[0070] In addition, after determining the first sensitive wavelength reflectance and the second sensitive wavelength reflectance, the relative contribution of each wavelength to the PLSR model can be further evaluated by using Variable Importance in Projection (VIP). Of course, other methods can also be used for contribution evaluation, such as Significance Multivariate Correlation (SMC), Selectivity Ratio (SR), Regression Coefficients (RC), etc. In Figure 10 and Figure 11 In the example of the VIP weight distribution of the sensitive wavelengths in the source domain and the target domain shown, it can be seen that through the Variable Importance in Projection (VIP) analysis, the spectral feature distributions of the first 14 key sensitive wavelengths in the source domain and the target domain show significant correlation (p < 0.01). In the spectral range of 648 - 681 nm, the VIP scores of both domains exceed 0.25, and the relative contribution degrees of the wavelength points at 655 and 648 nm reach the peak values (VIP of the source domain = 0.66, VIP of the target domain = 0.67). This wavelength range highly coincides with the absorption peaks of plant chlorophyll a / b (640 - 660 nm), indicating that the extracted sensitive wavelengths can effectively characterize the physiological state differences of wheat and provide spectral feature parameters with biochemical interpretability for the subsequent construction of the TCA-ALA-Ridge model.

[0071] Next, step S40 is executed: Construct a source domain disease monitoring model based on the first sensitive wavelength reflectance, the second sensitive wavelength reflectance, and the disease severity label.

[0072] After determining the first sensitive wavelength reflectance of the source domain dataset and the second sensitive wavelength reflectance of the target domain dataset, further, the disease severity label can be combined to construct the source domain disease monitoring model, and the accurate prediction of the disease severity of early diseases can be achieved through the constructed source domain disease monitoring model.

[0073] In step S40, constructing the source domain disease monitoring model based on the first sensitive wavelength reflectance, the second sensitive wavelength reflectance, and the disease severity label includes:

[0074] Perform standardization processing on the first sensitive wavelength reflectance and the second sensitive wavelength reflectance;

[0075] Map the standardized first sensitive wavelength reflectance and the second sensitive wavelength reflectance to the same high-dimensional reproducing kernel Hilbert space for feature space alignment, and obtain the target source domain dataset according to the disease severity label;

[0076] Optimize the regularization parameters of the regression model using the target source domain dataset to construct a source domain disease monitoring model.

[0077] In the process of constructing the source domain disease monitoring model, first perform standardized preprocessing on the reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength to ensure that the data can be on the same scale. Then, map the standardized reflectance of the first sensitive wavelength and the standardized reflectance of the second sensitive wavelength to the same high-dimensional reproducing kernel Hilbert space for feature space alignment to reduce the distribution difference between the reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength. After feature space alignment, combine the disease severity labels to obtain the target source domain dataset after feature space alignment. Subsequently, use the target source domain dataset to optimize the regularization parameters of the regression model during model training to achieve the construction of the source domain disease monitoring model.

[0078] Among them, the standardized processing of the reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength includes:

[0079] According to the first mean value of all the reflectances of the first sensitive wavelength and the first standard deviation of all the reflectances of the first sensitive wavelength, perform standardized processing on the reflectance of the first sensitive wavelength by the z-score method to obtain the standardized reflectance of the first sensitive wavelength;

[0080] According to the second mean value of all the reflectances of the second sensitive wavelength and the second standard deviation of all the reflectances of the second sensitive wavelength, perform standardized processing on the reflectance of the second sensitive wavelength by the z-score method to obtain the standardized reflectance of the second sensitive wavelength.

[0081] In the process of performing standardized preprocessing on the reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength, the first mean value and the first standard deviation of the reflectance of the first sensitive wavelength, and the second mean value and the second standard deviation of the reflectance of the second sensitive wavelength can be used. The reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength can be respectively standardized by using the z-score method, so that the standardized reflectance of the first sensitive wavelength and the standardized reflectance of the second sensitive wavelength can be obtained. By standardizing the sensitive wavelength, it can be ensured that the reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength can be on the same scale. The z-score method is a basic and powerful data standardization tool, and its core value lies in eliminating the incomparability between data and providing a unified benchmark for subsequent analysis; of course, it can also be achieved by other methods to perform standardized processing on the reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength.

[0082] Specifically, in the process of data standardization preprocessing, the source domain dataset and the target domain dataset can be set first, which can be respectively expressed as:

[0083] data1 = [x s1 , x s2 , x s3 , … x sns , y1]

[0084] data2 = [x t1 , x t2 , x t3 , … x tnt , y2]

[0085] Among them, data1 represents the source domain dataset, data2 represents the target domain dataset, X1 = [x s1 , x s2 , x s3 , … x sns ∈ R ns×d represents the source domain sensitive wavelength, X2 = [x t1 , x t2 , x t3 , … x tnt ∈ R nt×d represents the target domain sensitive wavelength, and y1 and y2 respectively represent the disease severity labels of the source domain and the target domain.

[0086] Then, the data of the source domain sensitive wavelength X1 and the target domain sensitive wavelength X2 can be read respectively in software such as matlab, and the z-score method can be used to standardize the data. The corresponding formula for the standardization process can be expressed as:

[0087]

[0088] Among them, X src and X tar respectively represent the sensitive wavelength data after z-score processing, that is, the first sensitive wavelength reflectivity and the second sensitive wavelength reflectivity after standardization processing. μ1 and μ2 respectively represent the means of the original sensitive wavelength reflectivities of all samples of X1 and X2, and σ1 and σ2 respectively represent the standard deviations of the original sensitive wavelength reflectivities of all samples of X1 and X2.

[0089] In addition, map the first sensitive wavelength reflectivity and the second sensitive wavelength reflectivity after standardization processing to the same high-dimensional reproducing kernel Hilbert space for feature space alignment, and obtain the target source domain dataset after feature space alignment according to the disease severity label, including:

[0090] Merge the first sensitive wavelength reflectance and the second sensitive wavelength reflectance after standardization by rows to obtain a joint matrix;

[0091] Obtain the dimension information of the joint matrix according to the joint matrix;

[0092] Obtain a symmetric matrix according to the dimension information and a preset weight coefficient;

[0093] Obtain a centering matrix according to the dimension information, the identity matrix, and the all-ones matrix;

[0094] Obtain a kernel matrix according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance after standardization, a preset linear kernel function type, and matrix transpose;

[0095] Obtain an optimal projection matrix according to the symmetric matrix, the kernel matrix, the centering matrix, the identity matrix, and a regularization parameter;

[0096] Obtain a target joint matrix that maps the joint matrix to the reproducing kernel Hilbert space according to the optimal projection matrix, matrix transpose, and kernel matrix;

[0097] According to the disease severity label, perform source domain and target domain segmentation and reconstruction on the target joint matrix to obtain a target source domain dataset after feature space alignment.

[0098] In the process of aligning the feature spaces of the first sensitive wavelength reflectance and the second sensitive wavelength reflectance after standardization in the high-dimensional reproducing kernel Hilbert space and establishing a target source domain dataset after feature space alignment, it is necessary to first merge the first sensitive wavelength reflectance and the second sensitive wavelength reflectance after standardization by rows to obtain a joint matrix. Then, according to the dimension information of the joint matrix, a preset weight coefficient, the identity matrix, the all-ones matrix, a preset linear kernel function type, matrix transpose, the identity matrix, and a regularization parameter, a symmetric matrix, a centering matrix, a kernel matrix, and an optimal projection matrix can be obtained respectively. Then, using the optimal projection matrix, matrix transpose, and kernel matrix, a target joint matrix that maps the joint matrix to the reproducing kernel Hilbert space can be obtained. Then, perform source domain and target domain segmentation and reconstruction on the target joint matrix. Combining the disease severity label, a target source domain dataset after feature space alignment can be obtained, realizing the feature space alignment of the first sensitive wavelength reflectance and the second sensitive wavelength reflectance after standardization to reduce the distribution difference between the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

[0099] Specifically, the transfer component analysis (TCA) method can be adopted to process the first sensitive wavelength reflectance X after standardization src and the second sensitive wavelength reflectance X tarMap to a common high-dimensional subspace, such as a reproducing kernel Hilbert space, to reduce the distribution difference between the two and provide a more consistent feature space for model construction. Among them, when preprocessing the reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength, methods such as the improved TCA method (such as TCADS), Joint Distribution Adaption (JDA), and Deep Domain Adaptation Networks (DDAN) can also be used for substitution.

[0100] When processing by the Transfer Component Analysis (TCA) method, first merge the preprocessed reflectance of the first sensitive wavelength and the reflectance of the second sensitive wavelength by row to obtain a joint matrix. Specifically, first set the TCA parameters, select the "linear" kernel function type, set the regularization parameter to 1.0, and set the kernel parameter to 2.0. Let X src and X tar be merged by row to form a new joint matrix X. And according to the new joint matrix X, calculate the dimension information of X, and the formula is expressed as:

[0101] X = [X src ; X tar ∈ R n×d

[0102] n = n s + n t

[0103] where n s and n t are the number of samples in the source domain and the target domain respectively, and d is the feature dimension.

[0104] Then, according to the dimension information and the preset weight coefficient, obtain a symmetric matrix. Construct a symmetric matrix L ∈ R n×n to reflect the distribution difference between cross-domain samples, and perform Frobenius norm normalization on the L matrix to enhance numerical stability. The calculation formula of the symmetric matrix L is:

[0105]

[0106] Immediately afterwards, according to the dimension information, the identity matrix, and the all-ones matrix, obtain a centering matrix. By constructing the centering matrix H, it can be used to eliminate the mean shift of X and retain the relative relationship between samples. Among them, the calculation formula of the centering matrix H is:

[0107]

[0108] Among them, \(l\) is an \(n\times n\) identity matrix, and \(P\) is an \(n\times n\) all-ones matrix.

[0109] Furthermore, according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance after standardization processing, the preset linear kernel function type, and matrix transpose, a kernel matrix is obtained. That is, the kernel matrix \(K\) can be calculated according to the preset linear kernel function type, implicitly mapping the data to a high-dimensional reproducing kernel Hilbert space (RKHS). For the linear kernel function "linear", \(K\) i,j = \(X\) i T \(X\) j , so the kernel matrix \(K\) can be expressed as:

[0110]

[0111] Among them, \(K\) src,src , \(K\) src,tar , \(K\) tar,src and \(K\) tar,tar are cross-domain data of the embedded Hilbert space defined on the source domain and the target domain, and \(T\) represents matrix transpose.

[0112] Next, an optimal projection matrix can be obtained according to the symmetric matrix, the kernel matrix, the centering matrix, the identity matrix, and the regularization parameter. That is, based on the settings of the symmetric matrix \(L\), the centering matrix \(H\), and the kernel matrix \(K\), eigenvalue decomposition and matrix projection decomposition are performed to solve the optimal projection matrix \(A\), and the formula can be expressed as:

[0113] \(A=(KLK + \lambda l)\) -1 \(KHK\)

[0114] Among them, \(\lambda\) is the regularization parameter, and \(l\) is the identity matrix.

[0115] Subsequently, according to the optimal projection matrix, matrix transpose, and kernel matrix, the target joint matrix of the joint matrix mapped to the reproducing kernel Hilbert space is obtained. That is, data mapping and output are performed to realize the mapping calculation of the joint matrix \(X\) to the subspace, and the formula can be expressed as:

[0116] \(X\) new = \(A\) T \(K\)

[0117] Among them, \(X\) new represents the joint matrix formed by mapping \(X\) into the new subspace, which is the target joint matrix.

[0118] Then, according to the disease severity label, the target joint matrix is segmented and reconstructed for the source domain and the target domain to obtain the target source domain dataset after feature space alignment. That is, \(X\) newThe source domain and target domain data are reconstructed separately, and finally new source domain and target domain feature parameters are obtained to achieve feature distribution alignment, that is, new source domain and target domain sensitive wavelengths are obtained. The formula can be expressed as:

[0119] X src_new = X new [:, 1:n s T

[0120] X tar_new = X new [:, n s+1 :n] T

[0121] Among them, X src_new and X tar_new respectively represent the new source domain and target domain sensitive wavelengths obtained after TCA processing.

[0122] Furthermore, based on the new source domain sensitive wavelength X src_new and the target domain sensitive wavelength X tar_new obtained after TCA processing, the source domain and target domain data are redefined as:

[0123] data 1_new = [X src_new , y1]

[0124] data 2_new = [X src_new , y2]

[0125] Among them, data 1_new = [X src_new , y1] represents the target source domain dataset after feature space alignment.

[0126] Immediately afterwards, the regularization parameter of the regression model is optimized through the target source domain dataset to construct a source domain disease monitoring model, which may further include:

[0127] Obtain a regression model according to the feature matrix, disease severity label and error of the source domain training data in the target source domain dataset; among them, the regression model includes a regression parameter vector to be solved;

[0128] Obtain the loss function of the regression model according to the regression model, L2 regularization term, and the number of samples of the source domain training data; among them, the loss function includes a regularization parameter;

[0129] Dynamically adjust the regularization parameter according to the initial search range of the regularization parameter, random number generation function, and energy factor to construct a source domain disease monitoring model.

[0130] ​In the process of constructing the source domain disease monitoring model, model training can be carried out based on the target source domain dataset after feature space alignment. Moreover, in the process of model training, the Artificial Lemming Algorithm (ALA) can be applied to optimize the regularization parameter of the regression model (for example, it can be a Ridge regression model) to ensure the stability and accuracy of the model in a complex data environment. In the process of optimizing the regularization parameter, the Artificial Lemming Algorithm (ALA) can be used. Of course, other algorithms can also be used, such as Particle Swarm Optimization (PSO), Genetic Algorithm (GA), Differential Evolution (DE), etc. Specifically, an initial regression model, which can be a Ridge regression model, or of course, other types of regression models, can be established by using the feature matrix, disease severity label, and error of the source domain training data in the target source domain dataset. Then, according to the regression model, L2 regularization term, and the number of samples of the source domain training data, the loss function of the regression model is further determined. And according to the regularization parameter to be optimized in the loss function, the initial search range of the regularization parameter, the random number generation function, and the energy factor, the dynamic adjustment of the regularization parameter is realized to obtain the optimized regularization parameter to construct the source domain disease monitoring model, so as to ensure the stability and accuracy of the source domain disease monitoring model in a complex data environment.

[0131] The Ridge regression algorithm is an improvement of the least squares estimation algorithm and is a biased estimation. It abandons the unbiasedness and partial accuracy of the least squares algorithm and seeks a more practical regression process. In the Ridge regression model, the larger the value of the regularization parameter λ, the greater the bias of the model. Therefore, determining the appropriate value of λ is the key to the Ridge regression method. The ALA algorithm can simulate the natural behaviors of lemmings such as migration, burrowing, foraging, and avoiding predators. Through mathematical modeling, the lemming behavior is transformed into four cooperative search operators, and an energy decreasing mechanism is introduced to dynamically regulate the selection probability of the operators, so as to gradually transition from global coarse-grained search to local refined optimization in the iterative process.

[0132] When constructing the Ridge regression model, for example, 2 / 3 can be randomly selected from the target source domain dataset as the source domain training data to construct the Ridge regression model of the source domain. Then, when using the Artificial Lemming Algorithm (ALA) to optimize the regularization parameter λ of the Ridge regression model, first set the feature matrix of the source domain training data as X ∈ Rm×d , the target vector y ∈ R m , then the Ridge regression model of the source domain can be defined as:

[0133] y = Xw + ε

[0134] where X and y represent the sensitive wavelengths and disease severity labels of the training data respectively, w is the regression parameter vector to be solved, m is the number of samples of the training data, and ε represents the error.

[0135] Then, since the goal of Ridge regression is to find w and make Xw infinitely close to y, Ridge regression introduces an L2 regularization term on the basis of ordinary least squares (OLS), and the loss function is:

[0136]

[0137] where min w means to find the weight vector w that minimizes the loss of the objective function. is the L2 regularization term, and λ is used to constrain the magnitude of the weight vector w to prevent the model from overfitting and underfitting.

[0138] Finally, according to the set initial search range of λ as [10 -5 , 10 3 , the λ parameter is dynamically adjusted by the ALA algorithm to construct the source domain disease monitoring model with the optimized λ parameter.

[0139] Specifically, according to the initial search range of the regularization parameter, the random number generation function, and the energy factor, the regularization parameter is dynamically adjusted to construct the source domain disease monitoring model, which may further include:

[0140] Generate the initial candidate values of the regularization parameter according to the initial search range of the regularization parameter;

[0141] Obtain the regression model loss value according to the initial candidate value and the loss function;

[0142] When the energy factor is greater than the first set value and the generated value of the random number generation function is less than the second set value, global exploration of migration perturbation is performed according to the preset number of iterations to obtain the first regularization parameter;

[0143] When the energy factor is greater than the first set value and the generated value of the random number generation function is greater than or equal to the second set value, local fine search is performed according to the preset number of iterations to obtain the second regularization parameter;

[0144] When the energy factor is less than the first set value and the generated value of the random number generation function is less than the third set value, spiral path density search is performed according to the preset number of iterations to obtain the third regularization parameter;

[0145] When the energy factor is less than the first set value and the generated value of the random number generation function is greater than or equal to the third set value, perform Levy flight hybrid search according to the preset number of iterations to obtain the fourth regularization parameter;

[0146] According to the first regularization parameter, the second regularization parameter, the third regularization parameter, the fourth regularization parameter, and the loss function of the regression model, obtain the target regularization parameter corresponding to the minimum regression model loss value under the preset number of iterations;

[0147] Construct a source domain disease monitoring model according to the target regularization parameter.

[0148] When dynamically adjusting the regularization parameter, the initial candidate value of the regularization parameter obtained by using the initial search range of the regularization parameter and the loss function are used to obtain the regression model loss value. Then, according to the size relationship between the energy factor and the first set value and the relationship between the generated value of the random number generation function and the second set value or the third set value, different regularization parameters can be obtained, namely the first regularization parameter, the second regularization parameter, the third regularization parameter, and the fourth regularization parameter. Then, by using the first regularization parameter, the second regularization parameter, the third regularization parameter, the fourth regularization parameter, and the loss function of the regression model, the target regularization parameter corresponding to the minimum regression model loss value under the preset number of iterations can be further obtained, and the source domain disease monitoring model is constructed by using the target regularization parameter, thus ensuring that the constructed source domain disease monitoring model can still maintain stability and accuracy in a complex data environment.

[0149] When dynamically adjusting the regularization parameter, the ALA algorithm can be used to dynamically adjust the λ parameter. When dynamically adjusting the λ parameter by the ALA algorithm, the parameters can be initialized first; that is, set the population size and set the maximum number of iterations T max to be 100, and initialize the population position to generate the initial candidate value of λ.

[0150] Then, take the derivative of the J(w) loss function and set the derivative to 0, and solve for the current w through the closed-form solution:

[0151] w = (X T X + λP) -1 X T y

[0152] where P is an m×m identity matrix.

[0153] Then, calculate the mean squared error EMS to measure the deviation between the predicted value and the actual value:

[0154]

[0155] Calculate the model loss at this moment based on the mean square error:

[0156]

[0157] Subsequently, the current λ parameter is dynamically updated by combining the rand() function and the energy factor E. Among them, rand is a function that can randomly generate numbers uniformly distributed in the interval [0, 1], and the energy factor where the angle decay term θ = 2arctan(1 - t / T max) Decreases dynamically with the number of iterations.

[0158] Moreover, when E > 1 and rand() < 0.3, perform migratory perturbation global exploration of λ. By simulating the migratory behavior of mouse populations, combining Gaussian noise and random direction perturbation to explore potential high-quality λ regions globally, the formula is:

[0159]

[0160] where F ∈ {-1, 1} is the random direction flag, BM ∼ N(0, 1) is the Brownian motion perturbation, the random vector R ∈ [-1, 1] is the mixing weight, t represents the number of iterations, i represents the population individuals, represents the position vector of the i-th individual at the t-th iteration, represents the updated position vector of the i-th individual at the (t + 1)-th iteration, represents the position of the optimal individual in the population at the t-th iteration, represents the average position vector of the population at the t-th iteration. And in the above formula, the position vector of each individual corresponds to a candidate value of λ, and based on the current optimal Perturbation to explore new λ values.

[0161] When E > 1 and rand() ≥ 0.3, perform local fine search of λ. Based on the dynamic sine coefficient, adjust the local search step size L, and finely search near the optimal solution by controlling L to avoid skipping high-quality λ regions due to too large a step size. The formula for L is:

[0162]

[0163] where L represents the search step size. At this time, the calculation formula for λ is:

[0164]

[0165] When E < 1 and rand() < 0.5, perform spiral path density search of λ. Generate a spiral path around the current optimal λ, conduct density-controlled circular exploration, and use the random phase angle to control the search density:

[0166] spird = radius × (sin(2 × π × rand) + cos(2 × π × rand))

[0167] Among them, sprial is the coordinate value of a certain point on the spiral path, and radius is the radius of the spiral, which is used to control the search range. radius can be defined as:

[0168]

[0169] Among them, Dim represents the dimension of the optimization problem, which directly affects the complexity of the search space, and z best,j (t) represents the global optimal position of the population in the j-th dimension at the t-th iteration, and z i,j (t) represents the current position of the i-th particle in the j-th dimension at the t-th iteration. At this time, the calculation formula of λ is:

[0170]

[0171] When E < 1 and rand() ≥ 0.5, execute the Levy flight hybrid search for λ. Introduce Levy Flight to achieve a hybrid search of long jump steps (global) and short step lengths (local):

[0172]

[0173] Among them, levy_step is the random step length based on Levy flight, which follows the heavy-tailed distribution levy ~ t -λ (1 < λ ≤ 3), sign is a function that returns -1, 0, and 1, and its numerical value is determined by rand(). G is a dynamic coefficient that balances exploration and exploitation. Among them, the dynamic coefficient G can be defined as:

[0174]

[0175] Among them, T max is the maximum number of iterations.

[0176] Based on the above, the model loss calculation will be repeatedly executed for each search result of λ, and the calculated model loss will be compared with the previous model loss. When the maximum number of iterations T max is reached, during the entire iteration process, retain the individual position that minimizes the model loss The λ corresponding to this position is the optimal solution.

[0177] After determining the optimal solution λ, by using the formula w = (X T X + λP) -1 X Ty, solve for the current w using the closed-form solution, and then substitute it into the Ridge regression model y = Xw + ε to optimize and obtain the source domain disease monitoring model.

[0178] In addition, during the process of validating and evaluating the wheat early disease monitoring model, it also includes:

[0179] Input the source domain validation data in the target source domain dataset into the source domain disease monitoring model to obtain the model prediction results;

[0180] Fit the model prediction results with the disease severity labels to obtain the first coefficient of determination;

[0181] Input the source domain validation data in the target source domain dataset into the comparison model to obtain the second coefficient of determination; among them, the comparison model includes a random forest model and a support vector regression model;

[0182] Compare the first coefficient of determination and the second coefficient of determination;

[0183] When the first coefficient of determination is greater than the second coefficient of determination, the source domain disease monitoring model is effective;

[0184] When the first coefficient of determination is less than the second coefficient of determination, the source domain disease monitoring model is ineffective.

[0185] Specifically, take the remaining 1 / 3 of the target source domain dataset except for the source domain training data as the source domain validation data, and then input it into the source domain disease monitoring model after optimizing the regularization parameter to obtain the corresponding model prediction results. Then, fit the model prediction results with the actual disease severity labels to obtain the first coefficient of determination. Then, input the validation data into comparison models such as the random forest (RF) and support vector regression (SVR) models to obtain the second coefficient of determination. Finally, compare the first coefficient of determination and the second coefficient of determination. When the first coefficient of determination is greater than the second coefficient of determination, it indicates that the source domain disease monitoring model is effective; otherwise, it indicates that the source domain disease monitoring model is ineffective. As Figure 12 The example of the source domain monitoring results based on the TCA-ALA-Ridge regression model is given, Figure 13 and Figure 14 The example of the source domain monitoring results based on the SVR and RF regression models is given. It can be seen that the monitoring accuracy R of the model optimized by TCA-ALA-Ridge for the source domain data 2 has been improved to 88%, reaching the target domain monitoring accuracy in the above table of the present invention. The present invention compares the TCA-ALA-Ridge regression algorithm with the SVR and RF conventional algorithms. From Figure 13 、 Figure 14It can be seen that the evaluation effect of model migration of conventional algorithms applied between different data sets is not good. The accuracy of SVR is only 74%, while that of RF only reaches 61%. Therefore, the technical method proposed in the present invention has improved by 14% and 27% respectively compared with the traditional SVR and RF algorithms, indicating that the source domain disease monitoring model is effective. The optimized source domain disease monitoring model can improve the early disease monitoring ability.

[0186] Please refer to Figure 15 , the present invention also provides a wheat early disease monitoring model construction system 11, including: an acquisition unit 111 for acquiring the spectral reflectance data of wheat ear samples and the disease severity of the spectral reflectance data; a division unit 112 for dividing the spectral reflectance data into a source domain data set and a target domain data set according to the disease severity; wherein, the source domain data set includes the first spectral reflectance data corresponding to health and the first disease infection degree, and the target domain data set includes the second spectral reflectance data corresponding to health and the second disease infection degree; a selection unit 113 for respectively screening the sensitive band ranges of the source domain data set and the target domain data set to obtain the first sensitive wavelength reflectance corresponding to the source domain data set and the second sensitive wavelength reflectance corresponding to the target domain data set; and a modeling unit 114 for constructing a source domain disease monitoring model according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance and the disease severity label.

[0187] It should be noted that the wheat early disease monitoring model construction system 11 provided in the above embodiment belongs to the same concept as the wheat early disease monitoring model construction method provided in the above embodiment. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the wheat early disease monitoring system 11 provided in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This will not be limited here either.

[0188] Please refer to Figure 16 , the electronic device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a wheat early disease monitoring model construction program.

[0189] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 12 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 can also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can be used not only to store the application software installed in the electronic device 1 and various types of data, such as the code for constructing the wheat early disease monitoring model, etc., but also to temporarily store the data that has been output or will be output.

[0190] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including the combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting all components of the entire electronic device 1 through various interfaces and lines, and by running or executing the programs or modules stored in the memory 12 (such as the program for constructing the wheat early disease monitoring model, etc.), and calling the data stored in the memory 12, to execute various functions of the electronic device 1 and process data.

[0191] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application program to implement the steps in the above-mentioned method for constructing the wheat early disease monitoring model.

[0192] Exemplarily, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and this instruction segment is used to describe the execution process of the computer program in the electronic device 1. For example, the computer program can be divided into each unit module of the wheat early disease monitoring system 11.

[0193] The integrated unit implemented in the form of software function modules described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above software function modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the wheat early disease monitoring model construction method described in various embodiments of the present application.

[0194] In summary, a method and system for constructing a wheat early disease monitoring model disclosed by the present invention divide spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity, then select the sensitive wavelength reflectance, and then map the source domain and target domain sample features with different disease degrees to a high-dimensional space by using the TCA method to achieve the alignment of the source domain and target domain feature distributions. At the same time, the ALA algorithm can be used to achieve intelligent parameter optimization, and the Ridge regression algorithm is combined to construct a source domain disease monitoring model. This method combines the domain adaptation ability of TCA, the adaptive optimization of ALA, and the regularization advantage of ridge regression, and is applicable to the wheat scab early disease monitoring modeling scenario, that is, it improves the early disease monitoring ability of the source domain disease monitoring model and solves the technical problems such as poor model generalization ability, low transferability, and low accuracy in the traditional method in the wheat scab early monitoring scenario. Therefore, the present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.

[0195] The above embodiments are only illustrative of the principles and effects of the present invention and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for constructing a monitoring model for early diseases of wheat, characterized in that, Including: Obtaining the spectral reflectance data of the wheat ear samples and the disease severity of the spectral reflectance data; Dividing the spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity; wherein, the source domain dataset includes the first spectral reflectance data corresponding to health and the first disease infection degree, and the target domain dataset includes the second spectral reflectance data corresponding to health and the second disease infection degree; Respectively performing sensitive band range screening on the source domain dataset and the target domain dataset to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset; Constructing a source domain disease monitoring model according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance and the disease severity label.

2. The method for constructing a wheat early disease monitoring model according to claim 1, wherein Respectively performing sensitive band range screening on the source domain dataset and the target domain dataset to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset, including: Establishing a first partial least squares regression model according to the source domain dataset within the set band range; Predicting a first disease severity prediction value through the first partial least squares regression model according to the first spectral reflectance data; Obtaining a first maximum determination coefficient according to the first actual disease severity value corresponding to the first spectral reflectance data and the first disease severity prediction value; Establishing a second partial least squares regression model according to the target domain dataset within the set band range; Predicting a second disease severity prediction value through the second partial least squares regression model according to the second spectral reflectance data; Obtaining a second maximum determination coefficient according to the second actual disease severity value corresponding to the second spectral reflectance data and the second disease severity prediction value; Respectively performing sensitive band range screening on the source domain dataset and the target domain dataset according to the first maximum determination coefficient and the second maximum determination coefficient to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

3. The method for constructing a wheat early disease monitoring model according to claim 2, wherein Respectively performing sensitive band range screening on the source domain dataset and the target domain dataset according to the first maximum determination coefficient and the second maximum determination coefficient to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance, including: Performing sensitive band range screening within the set band range according to the first maximum determination coefficient and the second maximum determination coefficient to obtain a target sensitive band range; wherein, the target sensitive band range is the same band range of the source domain dataset and the target domain dataset; Performing sensitive wavelength reflectance selection according to the target sensitive band range to obtain the first sensitive wavelength reflectance and the second sensitive wavelength reflectance.

4. The method for constructing a wheat early disease monitoring model according to claim 1, wherein Constructing a source domain disease monitoring model according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance and the disease severity label, including: Performing standardization processing on the first sensitive wavelength reflectance and the second sensitive wavelength reflectance; Map the standardized first sensitive wavelength reflectance and the second sensitive wavelength reflectance to the same high-dimensional reproducing kernel Hilbert space for feature space alignment, and obtain the target source domain dataset after feature space alignment according to the disease severity label; Optimize the regularization parameter of the regression model through the target source domain dataset to construct the source domain disease monitoring model.

5. The method for constructing a wheat early disease monitoring model according to claim 4, wherein Perform standardization processing on the first sensitive wavelength reflectance and the second sensitive wavelength reflectance, including: According to the first mean value of all the first sensitive wavelength reflectances and the first standard deviation of all the first sensitive wavelength reflectances, perform standardization processing on the first sensitive wavelength reflectance through the z-score method to obtain the standardized first sensitive wavelength reflectance; According to the second mean value of all the second sensitive wavelength reflectances and the second standard deviation of all the second sensitive wavelength reflectances, perform standardization processing on the second sensitive wavelength reflectance through the z-score method to obtain the standardized second sensitive wavelength reflectance.

6. The method for constructing a wheat early disease monitoring model according to claim 4, wherein Map the standardized first sensitive wavelength reflectance and the second sensitive wavelength reflectance to the same high-dimensional reproducing kernel Hilbert space for feature space alignment, and obtain the target source domain dataset after feature space alignment according to the disease severity label, including: Merge the standardized first sensitive wavelength reflectance and the second sensitive wavelength reflectance by rows to obtain a joint matrix; Obtain the dimension information of the joint matrix according to the joint matrix; Obtain a symmetric matrix according to the dimension information and the preset weight coefficient; Obtain a centering matrix according to the dimension information, the identity matrix, and the all-ones matrix; Obtain a kernel matrix according to the standardized first sensitive wavelength reflectance, the second sensitive wavelength reflectance, the preset linear kernel function type, and matrix transpose; Obtain the optimal projection matrix according to the symmetric matrix, the kernel matrix, the centering matrix, the identity matrix, and the regularization parameter; Obtain the target joint matrix that maps the joint matrix to the reproducing kernel Hilbert space according to the optimal projection matrix, the matrix transpose, and the kernel matrix; According to the disease severity label, perform source domain and target domain segmentation and reconstruction on the target joint matrix to obtain the target source domain dataset after feature space alignment.

7. The method for constructing a wheat early disease monitoring model according to claim 4, wherein Optimize the regularization parameter of the regression model through the target source domain dataset to construct the source domain disease monitoring model, including: Obtain a regression model according to the feature matrix of the source domain training data, the disease severity label, and the error in the target source domain dataset; where the regression model includes a regression parameter vector to be solved; Obtain the loss function of the regression model according to the regression model, the L2 regularization term, and the number of samples of the source domain training data; where the loss function includes the regularization parameter; Dynamically adjust the regularization parameter according to the initial search range of the regularization parameter, the random number generation function, and the energy factor to construct the source domain disease monitoring model.

8. The method for constructing a wheat early disease monitoring model according to claim 7, wherein Dynamically adjust the regularization parameter according to the initial search range of the regularization parameter, the random number generation function, and the energy factor to construct the source domain disease monitoring model, including: Generate an initial candidate value of the regularization parameter according to the initial search range of the regularization parameter; Obtain the regression model loss value according to the initial candidate value and the loss function; When the energy factor is greater than the first set value and the generated value of the random number generation function is less than the second set value, perform migratory perturbation global exploration according to the preset number of iterations to obtain the first regularization parameter; When the energy factor is greater than the first set value and the generated value of the random number generation function is greater than or equal to the second set value, perform local fine search according to the preset number of iterations to obtain the second regularization parameter; When the energy factor is less than the first set value and the generated value of the random number generation function is less than the third set value, perform spiral path density search according to the preset number of iterations to obtain the third regularization parameter; When the energy factor is less than the first set value and the generated value of the random number generation function is greater than or equal to the third set value, perform Levy flight hybrid search according to the preset number of iterations to obtain the fourth regularization parameter; Obtain the target regularization parameter corresponding to the minimum regression model loss value under the preset number of iterations according to the first regularization parameter, the second regularization parameter, the third regularization parameter, the fourth regularization parameter, and the loss function of the regression model; Construct the source domain disease monitoring model according to the target regularization parameter.

9. The method for constructing a wheat early disease monitoring model according to claim 4, characterized in that, It also includes: Input the source domain validation data in the target source domain dataset into the source domain disease monitoring model to obtain the model prediction result; Fit the model prediction result with the disease severity label to obtain the first coefficient of determination; Input the source domain validation data in the target source domain dataset into the comparison model to obtain the second coefficient of determination; wherein, the comparison model includes a random forest model and a support vector regression model; Compare the first coefficient of determination and the second coefficient of determination; When the first coefficient of determination is greater than the second coefficient of determination, the source domain disease monitoring model is effective; When the first coefficient of determination is less than the second coefficient of determination, the source domain disease monitoring model is invalid.

10. A system for constructing a monitoring model for early diseases of wheat, characterized in that, It includes: An acquisition unit for acquiring the spectral reflectance data of the wheat ear sample and the disease severity of the spectral reflectance data; A division unit for dividing the spectral reflectance data into a source domain dataset and a target domain dataset according to the disease severity; wherein, the source domain dataset includes the first spectral reflectance data corresponding to health and the first disease infection degree, and the target domain dataset includes the second spectral reflectance data corresponding to health and the second disease infection degree; A selection unit for respectively screening the sensitive band ranges of the source domain dataset and the target domain dataset to obtain the first sensitive wavelength reflectance corresponding to the source domain dataset and the second sensitive wavelength reflectance corresponding to the target domain dataset; and A modeling unit, configured to construct a source domain disease monitoring model according to the first sensitive wavelength reflectance, the second sensitive wavelength reflectance, and the disease severity label.

Citation Information

Patent Citations

  • Winter wheat scab remote sensing identification method based on near-earth hyperspectral technology

    CN111767863A

  • Winter wheat moisture monitoring method and system based on PROSPECT model

    CN111965117A

  • Lake chlorophyll inversion method based on transfer learning and semi-supervised regression model

    CN116824393A

  • Air-ground integrated mining area environment monitoring method and system based on hyperspectral imaging

    CN119147481A

  • Rain, snow, and hail classification monitoring method based on semi-supervised domain adaptation

    WO2021159844A1