Tomato disease diagnosis method based on multi-modal data analysis
Through multimodal data analysis, combined with hyperspectral imaging and microcamera, the growth stage is divided, the network weight is dynamically adjusted, and an interpretable diagnostic report is generated, which solves the problems of high misdiagnosis rate and delayed prevention and treatment in the diagnosis of tomato diseases, and achieves accurate disease diagnosis and timely prevention and treatment.
Patent Information
- Application Number
- CN202510652779.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing tomato disease diagnosis technology relies on a single data source, making it difficult to distinguish disease symptoms from environmental stress, and lacks a multimodal data fusion mechanism adaptive to the growth stage, resulting in a high misdiagnosis rate and delayed prevention and treatment timing.
Synchronous acquisition and preprocessing of multimodal data, combined with hyperspectral imaging and microscopic cameras, the growth stage is divided, adaptive feature extraction and fusion is carried out, hybrid models are constructed, network weights are dynamically adjusted through the gated network, and interpretability diagnostic reports are generated.
Accurate and efficient diagnosis of tomato diseases has been achieved, the rate of misdiagnosis has been reduced, scientific basis is provided to prevent and treat promptly, and the stability of agricultural production has been ensured.
Smart Images

Figure CN120495825A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent agricultural equipment, and in particular to a tomato disease diagnosis method based on multimodal data analysis. Background Art
[0002] Against the backdrop of the rapid development of intelligent agriculture, tomato disease diagnosis technology has gradually evolved from traditional manual experience-based judgment to automation and precision. However, existing technologies still have significant shortcomings in practical applications, as reflected in the following aspects: Existing methods often rely on a single data source (such as visible light images), making it difficult to distinguish between similar manifestations of disease symptoms and environmental stress. For example, nutrient deficiency and fungal diseases both manifest as yellowing and mottling on leaves, making image analysis alone prone to misjudgment. Furthermore, single-modal data cannot capture the multidimensional correlations of disease development (such as the coordinated changes in environmental parameters and physiological indicators), resulting in insufficient model generalization.
[0003] While some studies have attempted to fuse multimodal data (images, environmental parameters, etc.), these fusion methods often rely on fixed weights or simple feature splicing, failing to consider the dynamic impact of plant growth stages on disease characterization. For example, root environmental stability during the seedling stage is more sensitive to disease occurrence, while microscopic pathological features during the peak fruiting stage are more diagnostically valuable. Existing methods lack a weight allocation mechanism that adapts to the growth stage, limiting fusion effectiveness.
[0004] If these issues are not addressed, the misdiagnosis rate will increase, prevention and control will be delayed, and economic losses will result. Therefore, the development of a diagnostic method that is adaptive to growth stages, dynamically integrates multiple modalities, and is compatible with edge computing has become an urgent need in the field of smart agriculture. Summary of the Invention
[0005] Based on the above objectives, the present invention provides a tomato disease diagnosis method based on multimodal data analysis, comprising: Step 1: Synchronous multimodal data acquisition and preprocessing: The hyperspectral imaging device and the microscopic camera are triggered synchronously to collect hyperspectral images of the plant canopy and microscopic images of the stem, respectively. Temperature, electrical conductivity, and dissolved oxygen environmental parameters are also continuously collected in the root zone. The growth stages are divided based on plant height and number of flowering nodes, and the growth stages include seedling stage, flowering and fruiting stage, and peak fruiting stage; Step 2: Adaptive feature extraction and fusion during the growth stage, reflectance correction and leaf segmentation of the hyperspectral image, and extraction of leaf spectral curve features; Perform stomata segmentation on microscopic images and extract dynamic features of stomata density and opening; Calculate dynamic stability indicators of environmental parameters, including temperature coefficient of variation and conductivity change trend; Assign weights to each modal feature according to the current growth stage and map the features into a three-dimensional orthogonal space; Step 3: Hybrid model construction and spatiotemporal analysis: construct a hybrid model that includes spectral, microscopic, and environmental analysis networks, and dynamically adjust the weights of each network through a gating network; Analyze the three-dimensional spatial trajectory of multiple consecutive days and calculate its deviation from the preset health pattern; Step 4: Disease decision generation: When the deviation exceeds the adaptive threshold, a disease warning is triggered, and an interpretable diagnostic report containing multimodal abnormality indicators is generated by combining the hybrid model output.
[0006] Beneficial effects of the present invention: Through the fusion of multimodal data, it is possible to fully capture multi-dimensional information during the plant growth process, thereby improving diagnostic accuracy; secondly, the adaptive division of growth stages and dynamic adjustment mechanism enable the model to better adapt to different growth environments and changes; finally, interpretable diagnostic reports and early warning mechanisms help farmers and relevant personnel take effective prevention and control measures in a timely manner, reduce the impact of diseases on crops, and ensure the stability of agricultural production. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0008] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2 This is a flowchart of the method for pore segmentation and feature extraction in step 2 of the method of the present invention; Figure 3 Flowchart of the steps of the anomaly detection method in the present invention. DETAILED DESCRIPTION
[0009] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It is also noted that, to provide a more detailed description, the following embodiments are best and preferred embodiments, and those skilled in the art may employ alternative methods for implementing certain known technologies. Furthermore, the accompanying drawings are intended only to provide a more detailed description of the embodiments and are not intended to limit the present invention.
[0010] See Figure 1-Figure 3This embodiment of the present invention provides a tomato disease diagnosis method based on multimodal data analysis. Step 1 involves synchronously triggering a hyperspectral imaging device and a microscope camera to capture hyperspectral images of the plant canopy and microscopic images of the stem. Simultaneously, environmental parameter data (such as temperature, electrical conductivity, and dissolved oxygen) is collected in the root zone. This ensures that the collected multimodal data fully reflects the growth and health of the tomatoes. This step advantageously enables the acquisition of comprehensive plant status data at a single point in time, providing accurate input for subsequent analysis.
[0011] In step 2, adaptive feature extraction and fusion based on growth stages are performed. Growth stages are first divided according to plant height and number of flowering nodes. Different growth stages have varying impacts on plant physiological characteristics. Therefore, feature extraction is performed on various data types in this step: hyperspectral image reflectance correction, leaf segmentation and leaf spectral curve extraction, microscopic image stomatal segmentation and stomatal dynamic feature extraction, and dynamic stability indicators of environmental parameters (such as temperature coefficient of variation and conductivity trend). Furthermore, different weights are assigned to each modal feature based on the growth stage, prioritizing features from different growth stages. This step has the beneficial effect of improving diagnostic accuracy through precise feature extraction and weight assignment.
[0012] Step 3 combines hyperspectral, microscopic, and environmental analysis data to construct a hybrid model comprising multiple analytical networks. A gated network dynamically adjusts the weights of each network. Through spatiotemporal analysis, the three-dimensional spatial trajectory of plants over multiple consecutive days is calculated, and the degree of deviation from the pre-set health pattern is then calculated. When a plant's growth trajectory deviates significantly from the healthy pattern, an early warning of the onset of disease can be provided. This step has the beneficial effect of enabling real-time monitoring of plant growth, promptly identifying potential health risks, and improving the predictability of disease diagnosis.
[0013] Finally, in step 4, when the deviation exceeds the set adaptive threshold, a disease warning is triggered, and an interpretable diagnostic report containing multimodal anomaly indicators is generated based on the output of the hybrid model. This report provides a detailed description of the plant's health status, including anomaly indicators in each diagnostic modality, helping to accurately diagnose the type and severity of the disease and provide a basis for subsequent treatment. This step has the beneficial effect of generating scientific and intuitive disease diagnosis reports, assisting agricultural workers in decision-making and reducing the errors and time consumption of manual diagnosis.
[0014] In one possible implementation, plant height is measured using a non-contact laser ranging method, with the measurement point located at the base of the first true leaf below the top of the main stem. This measurement method avoids disturbing the plant and ensures measurement accuracy and repeatability. Each day, three measurements are collected within a fixed time period, and the median is taken as the effective height for that day. This method ensures the stability and representativeness of the height data, effectively avoiding deviations caused by environmental fluctuations or measurement errors.
[0015] Flowering nodes are counted by combining reflectance changes in the characteristic anthocyanin bands in hyperspectral imagery with calyx morphology identification in microscopic imagery. A valid flowering node is counted when floral organ features are detected in the same leaf axil for three consecutive days. This multimodal information fusion technique, combining spectral and microscopic imagery, efficiently and accurately detects flowering nodes, facilitating precise identification of the flowering stage of tomatoes.
[0016] When a plant's height exceeds the threshold for the current stage for five consecutive days and the number of flowering nodes reaches the requirement for the next stage, the system initiates the stage transition verification process. This verification process involves calculating the Leaf Area Index (LAI) and comparing the microscopic characteristics of the apical meristem. The LAI effectively reflects the plant's growth saturation, while the microscopic characteristics of the apical meristem accurately identify the plant's specific growth stage. These verification methods ensure the accuracy of growth stage transitions and avoid misjudgments due to environmental changes or measurement errors.
[0017] These refined operational steps enable highly accurate growth stage delineation. The effective fusion of non-contact laser ranging and hyperspectral imagery not only improves measurement accuracy but also avoids bias caused by manual intervention. Accurate flowering node statistics and stage switching mechanisms ensure that the diagnostic system can precisely adjust to the plant's growth stage and dynamically optimize disease diagnosis strategies based on real-time data. These methods have the beneficial effect of providing an efficient and accurate means of tracking and predicting tomato growth and promptly detecting abnormalities in plant health, thereby providing a scientific basis for disease prevention and control.
[0018] In one possible implementation, dark current correction is performed to eliminate the effects of sensor background noise on hyperspectral data. Before the first acquisition of the day, all lighting is turned off, and a baseline noise image is captured in a completely dark scene. This image is used to correct for the sensor's baseline noise. Subsequently, a pixel-by-pixel subtraction operation is performed on the hyperspectral data collected. This subtraction involves subtracting the corresponding pixel value in the dark current image from the acquired value at each pixel, thereby eliminating the effects of dark current on the image data. This process ensures that the acquired hyperspectral data is not affected by changes in ambient light sources or sensor noise, ensuring data accuracy and reliability.
[0019] To address the impact of varying light levels on reflectance data, an array of ambient light intensity sensors was deployed above the plant canopy to monitor light intensity distribution in different areas in real time. Based on this data, a mapping table was constructed between light intensity and band gain, compensating reflectance in different bands based on light intensity. Specifically for low-light areas (light intensity below 20,000 lux), an exponential gain compensation curve was employed for the 700-1000nm band. This compensation method effectively reduces reflectance deviations caused by uneven lighting, ensuring balanced correction of light effects across all areas, thereby improving data accuracy.
[0020] Leaf region segmentation aims to accurately extract leaf regions from hyperspectral data. During the preliminary segmentation process, the reflectance threshold of the 660nm band is used to perform preliminary segmentation of the leaf region. This step quickly determines the area that may be a leaf by setting the reflectance threshold. However, due to the complex morphology of leaves, the preliminary segmentation often has a certain degree of edge blur or defects. Therefore, an improved U-Net network is further used for edge repair. Adversarial samples are added to the U-Net network during training, thereby enhancing the network's robustness to leaf defect areas, so that even when the leaves are affected by diseases or have irregular shapes, the boundaries of the leaves can be accurately identified.
[0021] In one possible embodiment, during microscopic image processing, the image is first preprocessed using the contrast-limited adaptive histogram equalization (CLAHE) algorithm. This algorithm enhances image contrast, thereby improving the visibility of the stomata region. In specific implementation, the contrast limit threshold is set at 1.5 times the standard deviation of the original image's grayscale levels to ensure moderate contrast after image processing and avoid excessive image enhancement leading to loss of detail. Furthermore, the CLAHE algorithm's block size is set at twice the average stomata diameter to ensure that the image details in each block effectively reflect the structural characteristics of the stomata. This preprocessing method improves image quality, reduces noise interference, and enables more accurate subsequent stomata segmentation and feature extraction.
[0022] Stomata are located using a modified YOLOv5 network. As a real-time object detection algorithm, YOLOv5 can quickly and accurately locate stomata. During training, a transfer learning strategy is employed, leveraging pre-trained weights from a plant microscopic image dataset. This allows the model to better adapt to the characteristics of plant microscopic images, thereby improving the accuracy of stomatal detection. Transfer learning not only accelerates the training process but also improves the performance of stomatal localization by fully leveraging existing pre-trained models, making this method more robust and accurate in practical applications.
[0023] After the stomatal area segmentation is completed, an ellipse fitting is performed on each segmented stomatal area, and the stomatal opening is calculated based on the fitting results. The specific method is to judge the stomatal opening state based on the ratio of the major axis to the minor axis of the fitted ellipse. When the ratio of the major axis to the minor axis exceeds the preset critical value, the stoma is marked as open. This critical value is obtained by statistically analyzing the state distribution of stomata of healthy plants under different light cycles. The preset critical value is adjusted based on experimental data to ensure that it can accurately reflect the degree of stomatal opening, thereby providing a key basis for disease diagnosis.
[0024] In one possible implementation, in multimodal data analysis, the information provided by different modalities (such as hyperspectral, microscopic images, environmental data, etc.) has different effects on disease diagnosis at different growth stages. Therefore, the setting of initial weights is key. The initial weight ratio of each modality is set based on the F1-score performance of each modal feature at different growth stages in the historical data. For example, at different growth stages of tomato plants, the diagnostic information provided by microscopic image features may be more accurate at certain stages, while at other stages, hyperspectral features may have higher diagnostic value. By analyzing historical data and combining the performance of each modal feature at a specific stage, the initial weight is set to ensure that each modal feature is given reasonable attention throughout the diagnostic process.
[0025] To optimize the weight distribution strategy in real time, a reinforcement learning framework was deployed. In this framework, diagnostic accuracy serves as a reward function, and the weights of each modality are continuously adjusted to enhance the model's performance in different diagnostic scenarios. Based on real-time feedback, the reinforcement learning framework optimizes the weight distribution of each modality through trial and error and learning. In practice, the reinforcement learning model adjusts the weight ratio of each modality based on the input data modalities and the results of disease diagnosis. This automatically adapts to the needs of disease diagnosis in new growth stages or under different lighting conditions, improving diagnostic accuracy. This online adjustment mechanism possesses powerful adaptive capabilities, continuously adjusting the weight distribution according to varying environmental conditions and diagnostic needs, ensuring real-time and accurate diagnosis.
[0026] In practical applications, diagnostic results from different modalities may differ, especially when the diagnostic confidence levels of hyperspectral and microscopic features differ significantly. To resolve this conflict, an automated conflict resolution rule was designed. When the difference in diagnostic confidence between hyperspectral and microscopic image features exceeds a preset threshold, the system automatically increases the weight of environmental features for secondary verification, especially when the diagnostic results are ambiguous or face significant uncertainty. This mechanism can improve diagnostic robustness by increasing the weight of environmental features, better determining the type of stress a plant is experiencing, and avoiding misdiagnosis or missed diagnosis due to insufficient information from a single modality.
[0027] In one possible implementation, at the initial stage of constructing the three-dimensional orthogonal space, it is first necessary to perform statistical analysis on the healthy sample data set to calculate the mean and variance of each feature dimension. By calculating the mean and variance of the healthy samples, the distribution of each feature dimension in a healthy state can be obtained, thus laying the foundation for the subsequent standardization process. Next, Z-score standardization is performed to convert the numerical values of all features into standardized data with a mean of 0 and a variance of 1. In this way, all eigenvalues will be adjusted to the same dimensional range, eliminating the differences between different feature scales, making subsequent feature extraction and projection more fair and consistent.
[0028] When optimizing directions in three-dimensional space, a sparse principal component analysis (Sparse PCA) algorithm is used to extract key projection directions. Compared to traditional principal component analysis (PCA), Sparse PCA introduces sparse constraints, ensuring that each principal component contains as few and significant features as possible. This algorithm extracts the projection directions most critical for tomato disease diagnosis. This method further constrains each principal component to contain feature contributions from at most two modalities, ensuring that the correlation between feature dimensions is minimized when constructing a three-dimensional orthogonal space, thereby improving the independence and interpretability of the space. This optimization maximizes the effective information in each dimension of the space, providing a more accurate feature space for subsequent disease identification.
[0029] To ensure the stability and accuracy of the model over long-term use, a dynamic calibration mechanism has been designed. Every month, the system recalculates the projection matrix to adapt to newly collected data samples. This process reoptimizes the spatial construction based on new data at different time points and adjusts the projection direction of the orthogonal space based on the latest dataset. If the KL divergence (Kullback-Leibler Divergence) of the newly collected sample exceeds two standard deviations of the historical distribution, the system triggers an emergency calibration mechanism. KL divergence is used to measure the difference between the distribution of new data and the distribution of historical data. If the distribution of new samples differs significantly from that of historical data, it indicates that the data has changed significantly. In this case, emergency calibration is used to adjust the projection matrix to ensure the accuracy and robustness of the diagnostic model.
[0030] In one possible implementation, the disease diagnosis system first collects multimodal data from a typical disease development cycle. This data includes data collected by multiple sensors on tomato plants at different stages of disease development, such as temperature, humidity, color changes, and leaf status. After this data is standardized, a library of standardized trajectory templates can be generated. Each template represents how the health status of tomatoes changes over time during the typical development of a particular disease. These templates not only capture the development trend of the disease from its early stages to its final stages, but also take into account environmental factors and the differences in the manifestations of different disease types. The template library provides a benchmark for subsequent disease detection and deviation assessment.
[0031] To detect abnormalities in tomato plant diseases, an improved dynamic time warping algorithm is used to calculate similarity between three-dimensional spatial trajectories over multiple consecutive days. The traditional DTW algorithm effectively handles data drift at different time intervals when comparing time series data. However, in practice, different stages of tomato disease may have different impacts on each modal data. Therefore, the improved DTW algorithm introduces a modal weight coefficient. This coefficient is dynamically adjusted based on the contribution of each disease stage to the features of each modal data, ensuring more accurate distance calculations between features at different stages. This improvement allows the system to better identify key features related to disease progression, thereby accurately determining whether a tomato plant is on an abnormal trajectory.
[0032] To provide timely warnings of disease development, a dual-threshold mechanism has been designed. In the short term (3 days), if the trajectory deviation exceeds the first threshold, the system triggers an observation warning, alerting farmers to the possibility of disease. At this stage, the disease may still be in its early stages, requiring further monitoring and data accumulation. In the long term (7 days), if the trajectory deviation exceeds the second threshold, an action warning is triggered, indicating that the disease has progressed beyond normal disease patterns and requires appropriate action, such as spraying pesticides or other agricultural control methods. By establishing both short-term and long-term warning mechanisms, the system can respond promptly based on the degree of deviation to prevent the further spread of the disease.
[0033] In one possible implementation, first, the input multimodal data needs to be effectively feature engineered. The growth stage information is one of the input features and is represented by encoding it into a three-channel hot encoding vector. Three-channel hot encoding can better represent the category information of different growth stages and avoid the sequential deviation that may be caused by traditional numerical encoding. Environmental parameter stability indicators, such as temperature, humidity, light intensity, etc., can balance the impact of extreme values on model training after logarithmic transformation, making the data distribution more stable and conducive to model convergence. At the same time, historical diagnostic results are also converted into a sliding window mean sequence, which represents the trend of diagnostic results over the past period of time. This sequence reflects the evolution trend of tomato diseases and is of great significance for predicting future disease development.
[0034] In terms of network architecture design, a multi-head self-attention mechanism is employed. This mechanism performs weighted averaging across different input feature dimensions, enabling the model to focus on the spatiotemporal correlations between different features. Each attention head corresponds to the feature space of an expert network. This design enables the network to independently learn multiple disease patterns in different input subspaces and automatically select the contribution of each subspace to disease diagnosis. For example, one expert network might focus on changes in environmental parameters, while another might focus on changes in growth stage information. This multi-head structure improves the network's expressive power and enhances its performance in recognizing complex disease patterns.
[0035] To ensure that the weight distribution of each expert network aligns with the domain experience in the prior knowledge base, the Kullback-Leibler (KL) divergence loss function was introduced. KL divergence, a metric that measures the difference between two probability distributions, can force the model's output distribution to be close to the distribution in the prior knowledge base. Specifically, the output weights of the expert networks are adjusted using the KL divergence loss to align the focus of each network with the experience in the domain knowledge base. For example, if the domain knowledge base indicates that certain diseases have a high probability of occurring under certain circumstances, the KL divergence loss function will make the model's predictions more consistent with this experience, thereby improving the credibility and accuracy of the predictions.
[0036] In one possible implementation, each node uses differential privacy technology to add noise to the gradients during local training. This step is intended to protect data privacy and prevent the leakage of personal data during training. Differential privacy technology ensures the privacy of individual data points by adding noise to the calculated gradients. The intensity of the noise is dynamically adjusted based on the sensitivity of the data. For example, for highly sensitive data (such as user personal information), the added noise is stronger, minimizing the impact of individual data points and reducing the risk of privacy leakage. Furthermore, dynamic adjustment of the noise ensures accuracy and stability during training, preventing degradation of model accuracy due to excessive noise.
[0037] After local training is complete, the gradients of each node need to be aggregated on the central server. To ensure the robustness of the aggregation results, the central server adopts a robust aggregation strategy, specifically using the median instead of the mean to handle gradients of outliers. Mean aggregation is susceptible to the influence of outliers (such as abnormal updates caused by faulty or malicious nodes), causing the aggregation results to deviate from the true value. In contrast, median aggregation has strong anti-interference capabilities and can effectively eliminate the influence of outliers, ensuring the quality and stability of model updates. This strategy ensures that even if some nodes experience anomalies, the overall model can still maintain good training results.
[0038] To manage hybrid model versions, a version tree is maintained, which records each updated model version and its corresponding performance metrics. During versioning, if the F1-score improvement on the validation set by a new version of the model is less than 1%, the old version is retained as a rollback candidate. This strategy prevents performance degradation during updates, ensuring that the model's performance improves significantly with each update. If the new version's performance fails to significantly improve, the system automatically rolls back to the old version, avoiding unnecessary performance loss.
[0039] In one possible implementation, during multimodal data analysis, inconsistent conclusions may arise between different modalities. For example, the visual modality and the sensor modality may produce different diagnostic results. To resolve this conflict, the system activates a Bayesian network inference engine, which calculates the maximum a posteriori probability (MAP) by combining environmental parameters (such as temperature, humidity, and light) with operational history (such as agricultural operation records). The Bayesian network utilizes probabilistic reasoning methods, comprehensively considering the reliability of data from different modalities and the influence of the external environment to help determine the most likely disease diagnosis. This reasoning method can effectively resolve conflicts between modalities and provide more accurate and reliable diagnostic conclusions.
[0040] During the disease diagnosis process, if certain characteristic indicators show abnormalities (such as abnormal temperature and humidity or abnormal image features), further analysis is required to determine the source of these abnormal features. At this point, the system performs backpropagation analysis to trace the source of these abnormal features, annotating their spatial location and time-series change nodes in the original data. This traceability mechanism not only identifies the specific data causing abnormal diagnostic results, but also uses time-series change analysis to determine whether these anomalies are caused by short-term environmental changes or long-term trends. Feature traceability helps improve diagnostic transparency and make diagnostic results more interpretable, making it easier for users to understand and take appropriate measures.
[0041] To improve the practicality of diagnostic results, the system constructs a knowledge graph that correlates diseases, environmental conditions, and agricultural practices. Through the knowledge graph, the system can recommend appropriate prevention and control measures based on disease type and environmental parameters. When chemical control is recommended, the system automatically matches the minimum effective concentration and calculates the most appropriate chemical control plan for the current situation. In addition, the system can also recommend the most appropriate agricultural practices based on historical operational and environmental data to ensure the effectiveness and pertinence of the recommended plans. This recommendation mechanism not only helps users take precise prevention and control measures based on diagnostic results, but also reduces the overuse of chemical drugs and reduces negative impacts on the environment by automatically calculating the minimum effective concentration.
[0042] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.
[0043] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A tomato disease diagnosis method based on multimodal data analysis, characterized in that: include: Step 1: Synchronous multimodal data acquisition and preprocessing: The hyperspectral imaging device and the microscopic camera are triggered synchronously to collect hyperspectral images of the plant canopy and microscopic images of the stem, respectively. Temperature, electrical conductivity, and dissolved oxygen environmental parameters are also continuously collected in the root zone. The growth stages are divided based on plant height and number of flowering nodes, and the growth stages include seedling stage, flowering and fruiting stage, and peak fruiting stage; Step 2: Adaptive feature extraction and fusion during the growth stage, reflectance correction and leaf segmentation of the hyperspectral image, and extraction of leaf spectral curve features; Perform stomata segmentation on microscopic images and extract dynamic features of stomata density and opening; Calculate dynamic stability indicators of environmental parameters, including temperature coefficient of variation and conductivity change trend; Assign weights to each modal feature according to the current growth stage and map the features into a three-dimensional orthogonal space; Step 3: Hybrid model construction and spatiotemporal analysis: construct a hybrid model that includes spectral, microscopic, and environmental analysis networks, and dynamically adjust the weights of each network through a gating network; Analyze the three-dimensional spatial trajectory of multiple consecutive days and calculate its deviation from the preset health pattern; Step 4: Disease decision generation: When the deviation exceeds the adaptive threshold, a disease warning is triggered, and an interpretable diagnostic report containing multimodal abnormality indicators is generated by combining the hybrid model output.
2. A tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The division of growth stages in step 1 specifically includes: Plant height was measured using a non-contact laser ranging method. The measuring point was located at the base of the first true leaf below the top of the main stem of the plant. Three measurements were taken at a fixed time each day, and the median was taken as the effective height for that day. The number of flowering nodes is counted by fusing the reflectance mutation detection of anthocyanin characteristic bands in hyperspectral images with the calyx morphology recognition in microscopic images. When the floral organ characteristics are detected in the same leaf axil for three consecutive days, it is counted as a valid flowering node. The trigger mechanism for growth stage switching is: when the plant height exceeds the current stage threshold for five consecutive days and the number of flowering nodes meets the requirements of the next stage, the stage conversion verification process is initiated, which includes leaf area index calculation and comparison of apical meristem microscopic characteristics.
3. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The detailed processing flow of the reflectivity correction described in step 2 includes: Dark current correction: Before the first acquisition of the day, turn off all lighting devices and collect a baseline noise image in a completely dark scene. Perform pixel-by-pixel subtraction on the subsequently collected hyperspectral data to eliminate the sensor background noise. Light compensation: Deploy an array of ambient light intensity sensors above the plant canopy to obtain real-time light intensity distribution in each area, establish a light intensity-band gain mapping table, and use an exponential gain compensation curve for the 700-1000nm band in low-light areas; Leaf region segmentation: After initial segmentation based on the reflectivity threshold of the 660nm band, an improved U-Net network is used for edge repair. Adversarial samples are added to the U-Net network during training to enhance the robustness to leaf defect areas.
4. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The method for pore segmentation and feature extraction in step 2 includes: Microscopic image preprocessing: Adaptive histogram equalization algorithm with limited contrast was used, with the contrast limit threshold set to 1.5 times the standard deviation of the original image grayscale level and the block size set to 2 times the average pore diameter; Stomatal localization: Real-time detection is performed based on an improved YOLOv5 network. A transfer learning strategy is used during training, and the backbone network retains pre-trained weights from a plant microscopic image dataset. Aperture calculation: Ellipse fitting is performed on the segmented stomatal area. When the ratio of the major axis to the minor axis exceeds a preset critical value, it is marked as open. The preset critical value is determined by statistically analyzing the state distribution of stomata of healthy plants under different light cycles.
5. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The optimization strategy for dynamically adjusting the weights of each network described in step 3 includes: Weight initialization: Set the initial weight ratio based on the F1-score performance of each modal feature at different growth stages in historical data; Online adjustment mechanism: Deploy a reinforcement learning framework, use diagnostic accuracy as the reward function, and optimize the weight allocation strategy in real time; Conflict resolution rules: When the diagnostic confidence difference between hyperspectral and microscopic features exceeds a preset threshold, the weight of the environmental features is automatically increased for secondary verification of the stress type.
6. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The construction and optimization method of the three-dimensional orthogonal space includes: Spatial initialization: Calculate the mean and variance of each feature dimension based on the healthy sample dataset and perform Z-score normalization; Direction optimization: A sparse principal component analysis algorithm is used to extract key projection directions, constraining each principal component to contain feature contributions from at most two modes; Dynamic calibration: The projection matrix is recalculated every month, and emergency calibration is triggered when the KL divergence of newly collected samples exceeds 2 standard deviations from the historical distribution.
7. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The anomaly detection method for analyzing the three-dimensional spatial trajectory of multiple consecutive days in step 4 includes: Template construction: Collect multimodal data of typical disease development cycles and generate a standardized trajectory template library; Similarity calculation: An improved dynamic time warping algorithm is used, and a modal weight coefficient is introduced to adjust the feature distance contribution at different stages; Gradual warning: A dual-threshold mechanism is set up. When the short-term trajectory deviation exceeds the first threshold, an observation warning is triggered. When the long-term deviation exceeds the second threshold, a disposal warning is triggered.
8. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The structural design and training method of the gated network include: Input feature engineering: Encode the growth stage into a three-channel hot encoding vector, perform logarithmic transformation on the environmental parameter stability index, and convert the historical diagnosis results into a sliding window mean sequence; Network architecture: A multi-head self-attention mechanism is used to fuse spatiotemporal features, with each attention head corresponding to the feature space of an expert network; Constraints: By adding the Kullback-Leibler divergence loss function, we ensure that the weight distribution of each expert network is consistent with the domain experience in the prior knowledge base.
9. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The specific implementation of the hybrid model updating method includes: Local training: Each node uses differential privacy technology to add noise to the gradient, and the noise intensity is dynamically adjusted according to the data sensitivity level; Aggregation strategy: The central server performs robust aggregation and uses the median instead of the mean to handle abnormal node gradients; Version control: Maintain the hybrid model version tree and retain the old version as a rollback candidate when the F1-score improvement of the new version on the validation set is less than 1%.
10. The tomato disease diagnosis method based on multimodal data analysis according to claim 1, characterized in that: The logic for generating the explainability diagnostic report includes: Conflict resolution: When multimodal conclusions are inconsistent, the Bayesian network inference engine is activated to calculate the maximum a posteriori probability diagnosis result by combining environmental parameters and operation history; Feature tracing: Perform backpropagation analysis on abnormal feature indicators and mark their spatial locations and temporal change nodes in the original data; Prevention and control recommendations: Build a knowledge graph to associate the disease-environment-agricultural operation relationship, and automatically match the minimum effective concentration calculation hybrid model when recommending chemical control.
Citation Information
Cited By
Sakura leaf spot identification method and system based on deep convolutional network
CN120783187A
Lemon health analysis method
CN121053535A
A method of lemon health analysis
CN121053535B
Greenhouse plant product harvesting monitoring method and system based on knowledge graph
CN121121502A
Tobacco virus classification model construction method based on tobacco hyperspectrum
CN121637134A