A high-efficiency leaf area estimation method
By using the self-supervised pre-training framework DINOv2 and the Canopy-Mix token-level hybrid enhancement technique, the problems of destructiveness, time consumption, high annotation cost, and weak generalization ability of leaf area estimation methods are solved, and efficient and accurate leaf area estimation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-04-10
AI Technical Summary
Existing leaf area estimation methods suffer from problems such as destructive measurement, time-consuming and labor-intensive, high labeling costs, susceptibility to ambient light interference, high equipment costs, and weak cross-species generalization ability.
We employ the self-supervised pre-training framework DINOv2 combined with a two-stage transfer learning framework, introduce the Canopy-Mix token-level hybrid augmentation technique, and fuse smooth L1 and Log-Cosh losses. Through self-supervised pre-training, model fine-tuning, and evaluation, we achieve leaf area estimation.
It significantly improves the model's cross-species and environmental generalization ability, reduces annotation costs and equipment investment, shortens measurement time, improves work efficiency, and meets the needs of large-scale breeding analysis.
Smart Images

Figure CN121010779B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of agricultural artificial intelligence, and specifically relates to an efficient method for estimating leaf area. Background Technology
[0002] In the field of agricultural artificial intelligence, the accuracy and efficiency of crop phenotypic analysis directly impact breeding progress and cultivation management decisions. Leaf area estimation methods are the technical foundation of crop phenotypic analysis, and their quantitative results directly serve phenotypic feature analysis, growth model construction, and agricultural decision optimization. Current leaf area estimation methods suffer from significant technical limitations:
[0003] (1) Destructive measurement methods, such as LAI punching and weighing method and AGB drying method, are time-consuming, labor-intensive and cannot be scaled up.
[0004] (2) Traditional computer vision technology relies on manual feature extraction, which is easily affected by ambient lighting and has a relative error of over 35%.
[0005] (3) Existing AI solutions face two major bottlenecks: First, supervised learning models require massive amounts of pixel-level labeled data, which are costly to label and have weak cross-species generalization ability; Second, overlapping and occlusion of leaves lead to systematic underestimation of predicted values, and solutions such as point cloud reconstruction are costly and difficult to promote.
[0006] Therefore, a new and efficient method for estimating leaf area is needed. Summary of the Invention
[0007] This invention proposes an innovative solution to the problems in the prior art. It combines a self-supervised pre-training framework (DINOv2) with a two-stage transfer learning framework. It introduces the Canopy-Mix token-level hybrid enhancement technique and integrates a hybrid strategy of smooth L1 and Log-Cosh loss, which breaks through the limitations of traditional methods in terms of sample size, generalization ability and cost-effectiveness.
[0008] This invention proposes an efficient leaf area estimation method, comprising the following steps:
[0009] S1. Collect data for model training and evaluation. Take photos of the plants in the field from above using a camera. Then, use a transparent glass plate to cover the plant leaves and calibration plate on a black background for taking photos. Calculate the calibration conversion coefficient and segment the leaves to obtain the actual leaf area value.
[0010] S2. Self-supervised pre-training of the model: Based on the DINOv2 architecture, a visual Transformer teacher-student network is built, integrating multi-source plant datasets to generate global and local views. The cross-entropy loss function is used to complete multiple rounds of large-scale self-supervised training.
[0011] S3. Model fine-tuning and leaf area regression: Add a multilayer perceptron regression head to the end of the visual Transformer teacher-student network, unfreeze the backbone network parameters in stages, use Canopy-Mix token fusion enhancement technology and smooth L1+Log-Cosh fusion loss to complete the model parameter fine-tuning;
[0012] S4. Model evaluation and validation: Cross-validation is used to divide the dataset, and statistical error indicators are used to validate the model's generalization ability and batch processing stability.
[0013] S5, Model Deployment and Application: Real-time calculation and output of leaf area value, correlation of fresh weight and dry weight data, supporting cultivation management and variety breeding.
[0014] Preferably, S1 includes:
[0015] S11. Select the main camera, set the main camera parameters, use a vertical bamboo pole with scales as a height aid, control the image acquisition environment, and take a bird's-eye view of the plants in the field under natural light.
[0016] S12. Use a transparent glass plate to cover the plant leaves and the calibration plate on a black background for photography.
[0017] S13. Using a vision and machine learning software library, calculate the calibration conversion coefficient, which is the ratio of the actual size of the calibration board to the number of pixels in the image;
[0018] S14. Segmenting leaf regions based on hue, saturation, and brightness;
[0019] S15. Count the total number of leaf pixels after segmentation, multiply by the calibration conversion coefficient to obtain the actual leaf area value.
[0020] Preferably, S2 includes:
[0021] S21. The DINOv2 self-supervised framework was selected as the technical foundation, and a teacher-student network was built based on the visual Transformer architecture.
[0022] S22. Integrate multi-source heterogeneous plant datasets for pre-training to establish the model's cross-species generalization ability;
[0023] S23. Perform large-scale self-supervised training, configure multiple iteration cycles and batch processing image quantity, and output a pre-trained model with general plant morphology perception capabilities.
[0024] Preferably, S3 includes:
[0025] S31. Add two multilayer perceptron regression heads to the end of the pre-trained visual Transformer teacher-student network, freeze the backbone network parameters, use the AdamW optimizer, configure the hierarchical learning rate, and combine the cosine annealing scheduler to achieve the initial mapping between leaf area and model output.
[0026] S32. Unfreeze the parameters of the last four layers of the visual Transformer teacher-student network, perform end-to-end joint training with the regression head, use the Canopy-Mix token hybrid augmentation technique, and set the weight ratio of smoothed L1 and Log-Cosh hybrid loss to 0.7:0.3.
[0027] The Canopy-Mix token hybrid enhancement technology includes:
[0028] (1) By randomly mixing the Transformer token sequences of global and local images using the Beta distribution, a synthetic sample with multi-image features is generated.
[0029] (2) Perform a weighted average of the leaf area label values using the same mixing ratio as the token.
[0030] Preferably, S4 includes:
[0031] S41. The samples collected in step S1 are randomly divided into multiple groups, and cross-validation is used to perform multiple rounds of iteration.
[0032] S42. Calculate the error index during each iteration, generate an evaluation report, and evaluate the model's performance.
[0033] The error indices include the coefficient of determination, mean absolute error, root mean square error, and relative root mean square error.
[0034] Preferably, S5 includes:
[0035] S51. Use a mobile phone or drone to take aerial photos of the plants in the field.
[0036] S52: Load the prediction model with fine-tuned parameters, receive image input, and output the leaf area value to the terminal or cloud in real time.
[0037] S53. Analyze and apply the predicted output results, correlate leaf fresh weight and dry weight data, and support cultivation management and variety breeding.
[0038] The efficient leaf area estimation method of the present invention has the following beneficial effects:
[0039] (1) By learning the general morphological features of non-rapeseed plants through the self-supervised pre-training framework (DINOv2) and combining it with Canopy-Mix token-level hybrid enhancement, the model’s cross-species / environment generalization ability is significantly improved, solving the domain offset problem of traditional ImageNet pre-trained models.
[0040] (2) The hybrid strategy of integrating smoothed L1 and Log-Cosh loss effectively reduces the impact of outliers and maintains high robustness in complex field backgrounds (such as soil noise and leaf overlap).
[0041] (3) Only a small number of labeled samples are needed to complete model fine-tuning, which greatly reduces the labeling cost. Using smartphones or drone platforms to replace traditional multi-view shooting equipment greatly reduces equipment investment.
[0042] (4) The time for measuring the leaf area of a single plant has been reduced from 5 minutes to 5 seconds using traditional destructive methods, significantly improving work efficiency. The batch processing capacity reaches 1000 plants / hour, meeting the needs of large-scale breeding phenotypic analysis.
[0043] The present invention proposes an efficient leaf area estimation method, which provides an efficient and accurate intelligent tool for plant breeding phenotypic analysis. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating the steps of an efficient leaf area estimation method according to an embodiment of the present invention;
[0045] Figure 2 This is a leaf plan view of an efficient leaf area estimation method according to an embodiment of the present invention. Detailed Implementation
[0046] To further understand the present invention, embodiments of the present invention are described below in conjunction with examples. However, it should be understood that these descriptions are only for further illustrating the features and advantages of the present invention, and not for limiting the scope of the claims of the present invention.
[0047] The following abbreviations or terms are used in embodiments of the present invention:
[0048] LAI: Leaf Area Index, a key physiological indicator reflecting the total area of plant leaves per unit land area.
[0049] AGB: Above Ground Biomass, a core parameter for monitoring crop growth.
[0050] SSL: Self-Supervised Learning, a machine learning paradigm that learns visual representations without requiring labeled data.
[0051] ViT: Vision Transformer, an image processing model architecture based on a self-attention mechanism.
[0052] DINOv2: Self-DIstillation with NO labels v2, an algorithm for self-supervised pre-training.
[0053] RRMSE: Relative Root Mean Square Error, the percentage of prediction error relative to the mean of the true value.
[0054] Canopy-Mix: This invention employs a token-level hybrid enhancement technique that generates enhanced samples with multi-image features by mixing the Transformer token sequences of two images according to a Beta distribution, thereby improving the robustness of the model to complex occlusion scenarios.
[0055] An embodiment of the present invention provides an efficient leaf area estimation method, such as... Figure 1 As shown, it includes the following steps:
[0056] S1. Collect data for model training and evaluation. This includes the following sub-steps:
[0057] S11. Select the main camera, set its parameters, and use a vertical bamboo pole with markings as a height aid to control the image acquisition environment. Perform an overhead shot of the field under natural light conditions.
[0058] In this embodiment of the invention, a commercial mobile phone is selected as the image acquisition device. A main camera is used, configured with a large aperture of f / 1.42, a shutter speed of 1 / 954s, and an ISO sensitivity of 50 to ensure image clarity. During shooting, a bamboo pole with a 60cm graduation is used to mark the shooting height to avoid image distortion caused by drastic changes in height.
[0059] rape( Brassica napus The L. (rapeseed) was sown on October 20, 2024, and the seedlings were transplanted to the field on November 15, 2024, at a planting density of 225,000 plants / hectare. The soil type was paddy soil, and the fertilizer application rate was 750 kg / hectare of rapeseed-specific slow-release fertilizer (N-P2O5-K2O=25:7:8). Field management practices for all plants were consistent throughout the growing season.
[0060] This embodiment of the invention was photographed from above on January 1, 2025, collecting 833 rapeseed plant samples from seedlings to the 5-leaf stage. The main differences between the samples were leaf area, leaf shape, and growth status. A 4cm × 3cm calibration board was placed next to each sample to establish a conversion benchmark between pixels and actual size, ensuring the accuracy of subsequent area calculations. The overhead photography was conducted under natural light in the field, using a bamboo pole with a 60cm graduation to mark the shooting height, and keeping the mobile phone parallel to the ground during shooting.
[0061] S12. Use a transparent glass plate to cover the plant's leaves and calibration plate on a black background for photography.
[0062] After filming in the field, such as Figure 2 As shown, the leaves of the plant were picked and laid flat on a black background along with the calibration board to ensure high contrast. The leaves were then pressed flat with a transparent glass plate to keep them spread out before taking the picture.
[0063] S13. Use vision and machine learning software libraries to perform calibration board recognition and scale calculation.
[0064] The OpenCV library is used to detect the edges of the calibration board. By calculating the ratio of the actual size of the calibration board to the number of pixels in the image, a pixel-to-centimeter calibration conversion coefficient is established. OpenCV is an open-source, cross-platform computer vision and machine learning software library, primarily used for real-time image processing and vision application development.
[0065] S14, Blade region segmentation.
[0066] Leaf pixels were extracted based on the HSV color space model (hue 35-75, saturation 80-255, value 50-255), and soil background was filtered out using a threshold segmentation method. The HSV color space is a three-dimensional color model based on hue, saturation, and value. Morphological operations can further optimize the segmentation results to ensure the integrity of the leaf outline.
[0067] S15, Pixel area conversion.
[0068] The total number of pixels in the segmented leaf is counted and multiplied by the calibration conversion factor to obtain the actual leaf area value. This step achieves non-contact measurement through calibration plate calibration, replacing traditional manual measurement methods such as the perforation method for leaf area measurement.
[0069] Step S1 provides high-precision, structured input data for model training through standardized data collection and automated processing, which is the foundation for building an efficient leaf area estimation system.
[0070] S2, self-supervised pre-training.
[0071] Step S2 completes the function of building the basic feature learning capability in the entire leaf area estimation process. This step, through pre-training on non-target crop images, enables the model to acquire cross-species plant morphology understanding capabilities, providing initialization parameter support for subsequent accurate regression of rapeseed leaf area. The specific implementation includes the following sub-steps:
[0072] S21. Basic framework setup and parameter initialization.
[0073] The DINOv2 self-supervised framework was chosen as the technical foundation. This framework constructs a teacher-student network based on a 12-layer, 768-dimensional visual Transformer architecture (ViT-Base). The teacher network weights are updated using an exponential moving average mechanism, with the momentum coefficient linearly increasing from an initial 0.996 to 1, ensuring the stability of the feature representation during knowledge distillation. This stage requires the completion of model architecture code implementation, GPU resource allocation, and initial weight randomization settings.
[0074] S22. Integration and enhancement of multi-source heterogeneous data.
[0075] To construct a robust and general-purpose feature extractor, this embodiment of the invention integrates images from five publicly available plant datasets. These datasets do not include rapeseed, aiming to allow the model to learn diverse plant morphological features, forming a pre-training data pool containing 25,535 non-rapeseed images. The publicly available plant datasets include:
[0076] (1) CVPPP dataset: 810 top views of rosette leaves of Arabidopsis thaliana and tobacco, showing basic overlap patterns.
[0077] (2) Flavia dataset: 1,907 single-leaf scans of 32 species, showing fine leaf shape and vein details.
[0078] (3) Plant Pathology 2021 dataset: 18,635 images of apple leaves in the field, covering different health conditions.
[0079] (4) Plant Seedlings dataset: 408 top views of seedlings of 12 crops / weeds.
[0080] (5) VegAnn dataset: 3,775 complex field scene images of 26+ vegetable crops.
[0081] Two spatial transformations are performed on each original image: a 224×224 pixel global view is generated to preserve overall structural cognition, while a 96×96 pixel local view is simultaneously generated to enhance detail sensitivity. A cross-entropy loss function is used to constrain the consistency of probability distributions across different views, constructing a spatially invariant feature learning scenario.
[0082] S23. Perform large-scale self-supervised training.
[0083] The training system is configured with 300 iteration cycles, processing 256 images per batch, with an initial learning rate of 1e-4, and employing a cosine annealing strategy. The temperature parameter τ is controlled at 0.07 to balance feature discriminative power and uniformity. During training, the KL divergence loss of the teacher and student networks is monitored in real time. An early stopping mechanism is automatically triggered when the loss decreases by less than 0.1% for five consecutive iterations. The final output is a pre-trained model with general plant morphology perception capabilities. This model addresses the domain shift problem of traditional ImageNet (a large visualization database used for visual object recognition software research) pre-trained models in intermediate layer representations such as leaf edge detection and texture feature extraction, thus improving domain adaptability.
[0084] S3: Model fine-tuning and leaf area regression.
[0085] Step S3 completes the function of transforming general plant morphological features into specialized leaf area prediction capabilities. This step achieves efficient adaptation of the pre-trained model to rapeseed data through a phased parameter unfreezing and joint optimization strategy, ultimately constructing a leaf area estimation model with centimeter-level accuracy. The specific implementation includes the following sub-steps:
[0086] S31. Regression Head Initialization and Feature Transfer.
[0087] Two multilayer perceptron (MLP) layers are added as dedicated regression heads at the end of the ViT-Base backbone network, with the hidden layer dimension set to 512 and ReLU activation function used. After loading the pre-trained visual Transformer weights from step S2, all backbone network parameters are frozen, and only the regression heads are enabled for training.
[0088] The optimizer used is AdamW (an adaptive optimizer for deep learning), configured with a hierarchical learning rate strategy: the backbone network learning rate is set to 5e-5, and the regression head learning rate is increased to 5e-4, with a cosine annealing scheduler to achieve smooth learning rate decay. This stage focuses on regression head training, optimizing the mapping relationship between leaf area and model output through the first 50 iterations to ensure effective utilization of pre-trained features.
[0089] S32. Perform full fine-tuning and joint optimization.
[0090] Unfreeze the parameters of the last four layers of the ViT-Base network and perform end-to-end joint optimization with the trained regression head. Maintain the AdamW optimizer configuration, dynamically monitor the validation set performance, and trigger an early stopping mechanism if the validation loss does not decrease for five consecutive rounds.
[0091] This invention introduces the Canopy-Mix token-level fusion enhancement technique, which mixes the Transformer token sequences of two images (global and local) according to a Beta distribution (α=0.5, β=0.5) to generate enhanced samples with multi-image features, while simultaneously weighting and mixing the leaf area label values. In this embodiment, the Canopy-Mix technique, through token-level fusion, randomly combines token sequences from different images to generate synthetic samples with multi-image features. This fusion method retains some features of the original images while introducing new combination variations, forcing the model to learn more general and abstract feature representations.
[0092] The loss function employs a hybrid strategy of smoothed L1 (handling outliers) and Log-Cosh (improving gradient stability), with a weight distribution of 0.7:0.3, effectively suppressing the interference of outliers on gradient updates. After 150 iterations, the model achieved an accuracy level of R²=0.805 and MAE=22.207cm² in five-fold cross-validation.
[0093] S4. Model Evaluation and Validation.
[0094] Step S4 quantifies model performance and ensures reliability for practical applications. This step verifies the model's accuracy and generalization ability in the rapeseed leaf area estimation task through five-fold cross-validation and multi-dimensional performance statistics, providing a quantitative basis for subsequent deployment. The specific implementation includes the following sub-steps:
[0095] S41. Dataset partitioning and configuration.
[0096] In this embodiment of the invention, 833 collected rapeseed samples were randomly divided into 5 groups, each containing plants at different leaf stages and growth states. For each evaluation, one group was selected as the test set (20%), and the remaining 4 groups were combined into the training set (80%). Five iterations were used to ensure that each sample was tested once, forming a complete cross-validation cycle. This process was implemented using a custom data loader with built-in random seed control to ensure repeatability.
[0097] S42, Perform multi-indicator performance statistics.
[0098] In each round of testing, the model predicts leaf area on the test set images, simultaneously recording the actual and predicted values. Four core metrics are calculated: coefficient of determination (R²) measures the degree of linear correlation, mean absolute error (MAE) reflects the level of absolute error, root mean square error (RMSE) assesses sensitivity to outliers, and relative root mean square error (RRMSE) represents the proportion of standardized error. After five rounds of cross-validation, the average values of each metric are summarized to form the final evaluation report. In typical experiments, this method improves predictive relevance by 26.6% compared to traditional methods.
[0099] S5, Model Deployment and Application.
[0100] This step enables the real-time application of the leaf area estimation model in a field environment through standardized equipment configuration, a precise calibration system, and an efficient inference process. The specific implementation includes the following sub-steps:
[0101] S51. Use mobile phones or drones to take aerial photos of plants in the field; the operation is simple.
[0102] S52, Model Inference Process and Prediction Output.
[0103] The system uses an established model to output the leaf area value in cm² of the captured images. The prediction results can be displayed on the mobile phone screen in real time and transmitted to the cloud via Bluetooth or Wi-Fi. It supports batch processing of 1000 plants / hour.
[0104] S53. Analysis and application of predicted output results.
[0105] The output can be correlated with fresh weight and dry weight, and can also include plant number, predicted value, and timestamp, facilitating subsequent phenotypic analysis. This result directly supports dynamic field cultivation management and breeding screening, accelerating the process of breeding high-yielding and stress-resistant varieties.
[0106] This invention establishes and applies a leaf area prediction model through steps S1-S5. The method in this embodiment is superior to other traditional methods. Data comparisons are shown in Table 1, with the coefficient of determination R between the predicted leaf area and the true leaf area value being... 2 The correlation coefficients (r) between leaf area and leaf fresh weight and dry weight predicted by this method reached 0.805, and were 0.900 and 0.885, respectively.
[0107] Table 1 Comparison of the effects of the method in this embodiment and the traditional method.
[0108]
[0109] The efficient leaf area estimation method of this invention has high accuracy and large processing capacity, and can provide efficient and accurate assistance for plant breeding phenotypic analysis.
[0110] The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A high-efficiency leaf area estimation method, characterized by, The method comprises the following steps: S1, collecting data for model training and evaluation, taking overhead photos of field plants through a camera, then using a transparent glass plate to cover the plant leaves and a calibration plate on a black background plate for shooting, calculating the calibration conversion coefficient, and segmenting the leaves to obtain the actual leaf area value; S2, model self-supervised pre-training, constructing a visual Transformer teacher-student network based on the DINOv2 architecture, integrating a multi-source plant dataset, generating global and local views, using a cross-entropy loss function, and completing multiple rounds of large-scale self-supervised training; S3, model fine-tuning and leaf area regression, adding a multi-layer perception regression head at the end of the visual Transformer teacher-student network, unfreezing the backbone network parameters in stages, using Canopy-Mix token mixing enhancement technology and smooth L1+Log-Cosh mixed loss, and completing model parameter fine-tuning; S4, model evaluation and verification, dividing the dataset using cross-validation, calculating error indicators, and verifying the generalization ability and batch processing stability of the model; S5, model deployment and application, real-time calculation of output leaf area value, correlation of fresh weight and dry weight data, and support for cultivation management and variety selection; The S1 comprises: S11, selecting a main camera, setting the main camera parameters, using a vertical bamboo pole with scales as a height aid, and controlling the image acquisition environment, and taking overhead field photos of plants in a natural light environment; S12, using a transparent glass plate to cover the plant leaves and a calibration plate on a black background plate for shooting; S13, using a visual and machine learning software library to calculate the calibration conversion coefficient; the calibration conversion coefficient is the ratio of the actual size of the calibration plate to the number of pixels in the image; S14, segmenting the leaf area based on hue, saturation, and lightness; S15, calculating the total number of segmented leaf pixels and multiplying the calibration conversion coefficient to obtain the actual leaf area value; The S2 comprises: S21, selecting a DINOv2 self-supervised framework as the technical base, and constructing a teacher-student network based on the visual Transformer architecture; S22, integrating a multi-source heterogeneous plant dataset for pre-training, and establishing the cross-species generalization ability of the model; S23, performing large-scale self-supervised training, configuring multiple iteration cycles and batch processing image quantities, and outputting a pre-trained model with general plant shape perception ability; The S3 comprises: S31, adding a two-layer multi-layer perception regression head at the end of the pre-trained visual Transformer teacher-student network, freezing the backbone network parameters, using the AdamW optimizer, configuring a hierarchical learning rate, combining the cosine annealing scheduler, and realizing the initial mapping of leaf area and model output; S32, unfreezing the last four layers of parameters of the visual Transformer teacher-student network, performing end-to-end joint training with the regression head, using the Canopy-Mix token mixing enhancement technology, and setting the smooth L1 and Log-Cosh mixed loss weight ratio to 0.7:0.3; The Canopy-Mix token mixing enhancement technology comprises: (1) Randomly mix the global and local image Transformer token sequences by Beta distribution to generate synthetic samples with multi-image features; (2) Perform weighted average on leaf area label values using the same mixing ratio as the tokens.
2. The method of claim 1, wherein, The S4 comprises: S41, randomly divide the samples collected in the S1 step into multiple groups, use cross-validation, and perform multiple rounds of iteration; S42, calculate the error index during each round of iteration, form an evaluation report, and evaluate the model effect; The error index includes the coefficient of determination, mean absolute error, root mean square error, and relative root mean square error.
3. The method of claim 1 or 2, wherein, The S5 comprises: S51, use a mobile phone or unmanned aerial vehicle device to take an overhead photograph of the plant in the field; S52, load the fine-tuned parameter prediction model, receive image input, and output the leaf area value to the terminal or cloud in real time; S53, analyze and apply the prediction output result, correlate the leaf fresh weight and dry weight data, and support cultivation management and variety selection.
Citation Information
Patent Citations
Preparation method of lactobacillus acidophilus
CN101096648A
Agricultural unmanned aerial vehicle terrain following method and device based on vision
CN120255536A