Litchi phenotype identification and weight estimation method based on multi-modal learning
By using the LitchiPhenoNet model, which employs both RGB and depth images as bimodal inputs, for litchi fruit detection and segmentation, and combining this with a multimodal regression method for fruit weight estimation, the accuracy issues of litchi phenotypic recognition and weight estimation are resolved, achieving efficient and accurate litchi fruit detection and weight estimation.
Patent Information
- Application Number
- CN202511313444.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies are insufficient for efficient and accurate phenotypic identification and weight estimation of litchi fruits, especially in terms of applicability and accuracy to different varieties. Furthermore, existing models are inadequate in handling the boundaries of fine-grained areas such as the peel edge and the pulp and pit.
A multimodal learning-based method for litchi phenotypic recognition and weight estimation is adopted. By using RGB images and depth images as bimodal inputs, the LitchiPhenoNet model is used for fruit detection and segmentation, and a multimodal regression method is combined for fruit weight estimation. The geometric features and abstract intermediate features extracted by the LitchiPhenoNet model are used for fruit weight prediction.
It significantly improves the accuracy and robustness of litchi fruit detection and segmentation, enhances the accuracy and robustness of fruit weight estimation, and enables automated analysis of high-precision phenotypic traits of different litchi varieties.
Smart Images

Figure CN120877276A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of litchi phenotypic analysis technology, and in particular to a litchi phenotypic recognition and weight estimation method based on multimodal learning. Background Technology
[0002] In recent years, methods for automatically estimating fruit size have been continuously developing, mainly including techniques based on 2D images and 3D point clouds. 2D image methods rely on reference objects or precise distance settings, which limits their practical application. 3D point cloud methods, through sensing technologies such as LiDAR and RGB-D cameras, achieve high-precision measurement of fruit size. In particular, RGB-D cameras, due to their low cost and ability to simultaneously acquire color and depth information, have been widely used in phenotypic analysis of various crops such as apples, mangoes, and grapes. However, litchi peels have complex structures due to their cleavage and significant varietal differences, making it difficult to directly apply existing methods. Furthermore, most fruit weight estimation methods are based on fruit growth curves or geometric parameters and weight relationship models, but due to the diverse morphologies among litchi varieties, these methods have limited predictive effectiveness on litchi, making it difficult to balance applicability and accuracy across different varieties. Therefore, there is a lack of automated analytical methods capable of efficiently and accurately measuring the geometric and weight traits of litchi fruits and seeds of different varieties.
[0003] With the development of computer vision and deep learning technologies, high-performance instance segmentation models such as YOLO and Mask R-CNN have been widely used in the agricultural field, enabling automatic detection and contour segmentation of crops and fruits. Multimodal learning models, by integrating multiple data sources such as RGB and depth, have demonstrated superior performance in tasks such as target detection, segmentation, and feature estimation. Although the YOLO series models can effectively detect and classify fruits and related structures, they still have shortcomings in segmenting fine structures, especially in the boundary processing of fine-grained regions such as the peel edge and the flesh and pit. Insufficient accuracy of segmentation results leads to low stability in the prediction of characterization parameters and the estimation of fruit weight.
[0004] Therefore, we propose a multimodal learning-based method for litchi phenotypic recognition and weight estimation to address the existing problems. Summary of the Invention
[0005] The purpose of this invention is to provide a method for litchi phenotypic recognition and weight estimation based on multimodal learning, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for litchi phenotypic recognition and weight estimation based on multimodal learning, comprising a phenotypic recognition stage and a weight estimation stage. In the phenotypic recognition stage, RGB images and depth images are used as bimodal inputs and input into the LitchiPhenoNet model to realize fruit detection and segmentation. In the weight estimation stage, based on the geometric features and abstract intermediate features directly extracted by LitchiPhenoNet, a multimodal regression method is used to estimate the fruit weight.
[0007] Furthermore, in the weight estimation stage, the intermediate feature map is flattened into a one-dimensional vector as an abstract intermediate feature in the last layer of the LitchiPhenoNet model. The abstract intermediate feature and the geometric feature are then used as inputs to the regression model and processed through two fully connected layers respectively. The abstract intermediate feature and the geometric feature are then concatenated and fused through an additional fully connected layer. Finally, the predicted weight of the fruit is output through the regression network.
[0008] Furthermore, for the phenotypic recognition stage, the bounding box and segmentation mask output by the LitchiPhenoNet model are used to locate the region of interest. The region of interest mask is combined with the calibrated camera intrinsics to extract the geometric phenotypic features of the fruit from the RGB and depth data.
[0009] Furthermore, during the training of the regression network, the regression model optimizes its parameters by minimizing the mean squared error between the predicted value and the actual weight.
[0010] Furthermore, geometric phenotypic features include horizontal diameter, vertical diameter, area, perimeter, minimum circumscribed circle radius, maximum inscribed circle radius, and shape index.
[0011] Furthermore, the LitchiPhenoNet model contains two independent feature extraction branches, which extract features for the RGB modality and the deep modality respectively. The RD fusion module uses a multi-scale channel attention module to perform deep fusion of multi-modal features.
[0012] Furthermore, the horizontal and vertical diameters of the litchi fruit were calculated using bounding box information extracted from the LitchiPhenoNet model.
[0013] Furthermore, the area, perimeter, minimum circumscribed circle radius, maximum inscribed circle radius, and shape index of the litchi fruit are calculated by extracting the contour from the segmentation mask.
[0014] Furthermore, for the calculation of horizontal and vertical diameters, the LitchiPhenoNet model automatically detects litchi fruits in the image and extracts the depth information of the target region; it uses bounding box coordinates to define the region of interest, and the depth data within the region is used to reconstruct the actual physical size; it selects the average depth value within the segmentation mask range to participate in the calculation, and combines the camera's intrinsic parameters to calculate the horizontal and vertical diameters of the litchi.
[0015] Furthermore, the processing flow for calculating area, perimeter, minimum circumcircle radius, maximum incircle radius, and shape index includes mask extraction, contour detection, depth calculation, and feature calculation.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention uses RGB and depth images as bimodal inputs to the LitchiPhenoNet model for fruit detection and segmentation. The RD fusion module significantly enhances multimodal feature representation and segmentation performance. This fusion strategy integrates the texture details of the RGB image with the spatial structure information of the depth image, greatly improving the accuracy and robustness of detection and segmentation in complex environments. Furthermore, the network combines multi-scale feature integration and context aggregation mechanisms to progressively refine feature representation, enhancing adaptability and generalization ability to targets at different scales. Based on the geometric features and abstract intermediate features directly extracted from LitchiPhenoNet, a multimodal regression method is used for fruit weight estimation. This regression model combines information from both visual and physical aspects. By combining these two types of features, the regression model effectively improves the utilization efficiency and prediction accuracy of multimodal data. This multimodal regression framework effectively integrates complementary information from geometric features and model visual features, thereby improving the accuracy and robustness of lychee weight estimation. Attached Figure Description
[0017] Figure 1 This is a flowchart of the algorithm for litchi phenotypic recognition and weight estimation based on multimodal learning in this invention. Figure 2 This is a structural diagram of the LitchiPhenoNet model of the present invention; Figure 3 This is a block diagram showing the module composition of RDFusion, MS-CAM, SPPF, CBS, BottleNeck, and C2f of the present invention; Figure 4 This is a page diagram of the horizontal and vertical diameter estimation software of the present invention; Figure 5 A comparison chart of performance data for different models; Figure 6 This is a comparison of the recognition and segmentation performance of the LitchiPhenoNet model of this invention with other models; Figure 7 A comparison chart showing the estimation of the transverse diameter and lateral diameter of litchi using different models; Figure 8 A comparison chart of weight estimation performance for different input modalities and models. Detailed Implementation
[0018] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. Example 1
[0019] An Intel RealSense D405 binocular depth camera was used to acquire RGB and depth images of litchi fruit. Compared to traditional LiDAR and other devices, this camera is low-cost, compact, and achieves sub-millimeter depth accuracy, making it suitable for high-precision measurement of fruit morphology. The D405 camera integrates two imaging sensors, acquiring high-precision depth information through active stereo structured light technology while simultaneously obtaining high-quality RGB images. A global shutter design effectively reduces motion blur, and pixel-level strict alignment between the RGB and depth images ensures data spatial consistency. In the experiment, the camera was mounted on a stable bracket in an indoor photography studio, with the shooting distance controlled within the range of 200-400 mm to obtain optimal image quality.
[0020] A total of 1197 sets of RGB and depth images were collected, each with a resolution of 1280×720 pixels, covering more than ten lychee varieties, including "Feizixiao," "Ziniangxi," and "Seedless Lychee." Whole fruits, pulp, and seeds were collected separately to reduce estimation bias caused by phenotypic differences between varieties. To verify the accuracy of the proposed model in estimating horizontal diameter, vertical diameter, and fruit weight, the horizontal and vertical diameters of 332 fruits (including seeds) were measured manually, and the weights of 408 samples (including whole fruits, pulp, and seeds) were manually weighed as a reference standard.
[0021] The open-source annotation software LabelMe was used to manually segment and annotate all images, classifying the targets into four categories: whole fruit, half-fruit without seeds, half-fruit with seeds, and seeds. Each annotated object had its contour accurately depicted, and its horizontal and vertical diameter information was recorded. All annotated data was saved in YOLO format. Finally, the entire dataset was randomly divided into training and testing sets at a 4:1 ratio for subsequent model training and performance validation.
[0022] like Figure 1As shown, this invention develops a multimodal learning measurement process that can perform high-precision automated estimation of key geometric phenotypic features (including horizontal diameter, vertical diameter and fruit weight) of litchi samples imaged by binocular depth cameras. It fully utilizes RGB and depth information to extract geometric features and combines a multimodal feature fusion strategy to improve the accuracy of weight prediction. The overall process includes two core stages: bimodal estimation of geometric features and multimodal estimation of fruit weight.
[0023] In the first stage, RGB and depth images are used as bimodal inputs to LitchiPhenoNet, an instance segmentation network specifically designed for litchi phenotypic analysis, to achieve high-precision fruit detection and segmentation. The bounding boxes and segmentation masks output by the model can accurately locate regions of interest (ROIs), providing a reliable foundation for subsequent feature calculations. Combining the ROI mask with calibrated camera intrinsics, the geometric phenotypic features of the fruit can be extracted from the RGB and depth data, including horizontal diameter, vertical diameter, area, perimeter, minimum circumcircle radius, maximum incircle radius, and other shape indices.
[0024] In the second-stage weight estimation regression model, a multimodal regression method is used to estimate fruit weight based on geometric features and abstract intermediate features directly extracted from LitchiPhenoNet. This regression model combines information from both visual and physical perspectives. Geometric features are closely related to the physical morphology of the fruit, while intermediate features are high-dimensional abstract representations extracted through deep learning networks. By combining these two types of features, the regression model can effectively improve the utilization efficiency and prediction accuracy of multimodal data.
[0025] Specifically, the intermediate feature map F extracted from LitchiPhenoNet model ∈R C×H×W Transform it into a one-dimensional vector f through flattening. model ∈R CHW Meanwhile, the geometric eigenvector f geo ∈R D It is obtained directly from geometric features. The two sets of features are each passed through two independent fully connected layers (FC) to generate two high-dimensional feature representations: h model =σ(W model ·f model +b model ), h geo =σ(W geo ·f geo +b geo ), where W model and W geo These are the weight matrices for the two fully connected layers, b model and b geoThese are the bias vectors for the two fully connected layers, and σ represents the activation function, such as ReLU.
[0026] Subsequently, these two sets of features are concatenated into a fused feature vector h. fusion :h fusion =[h model h geo ].
[0027] The fused features are further mapped to the final weight prediction through an additional fully connected layer. : =w fusion ·h fusion +b fusion , where w fusion For the weights of the additional fully connected layer, b fusion This is the bias for the additional fully connected layer.
[0028] The model training objective is to minimize the predicted weight. Mean square error (MSE) between the actual weight y and the mean square error:
[0029] This multimodal regression framework effectively integrates complementary information from geometric and intermediate visual features, thereby improving the accuracy and robustness of litchi weight estimation.
[0030] like Figure 2 and Figure 3 As shown, the LitchiPhenoNet model proposed in this invention innovatively extends the traditional single-input YOLOv8 object detection framework into a deep learning structure that supports dual-modal input, enabling simultaneous processing of RGB and depth images to achieve high-precision detection and segmentation of litchi fruits. To effectively fuse the complementary features of the two modalities, a novel RD fusion module is further proposed, significantly improving multimodal feature representation and segmentation performance. Specifically, the LitchiPhenoNet model includes two independent feature extraction branches, extracting features for the RGB and depth modalities respectively. Then, the RD fusion module uses a multi-scale channel attention module (MS-CAM) to deeply fuse multimodal features. MS-CAM can dynamically weight multi-scale, multi-channel features, highlighting discriminative information and suppressing redundant features. This fusion strategy integrates the texture details of the RGB image and the spatial structure information of the depth image, greatly improving the accuracy and robustness of detection and segmentation in complex environments. Furthermore, the network combines multi-scale feature integration and context aggregation mechanisms to progressively refine feature representation, enhancing adaptability and generalization ability to targets at different scales.
[0031] In summary, the proposed dual-modal input structure and RD fusion module effectively overcome the inherent limitations of the single-modal method, significantly improving the robustness and accuracy of fruit detection and segmentation, and laying a solid foundation for high-precision automatic estimation of phenotypic traits such as transverse diameter, longitudinal diameter, and weight of lychee.
[0032] The RD fusion module proposed in this invention draws on and improves upon the iterative attention feature fusion approach, enabling efficient integration of features from RGB and depth modalities. This module effectively addresses the semantic and scale inconsistencies between the two modalities through two consecutive multi-scale channel attention modules (MS-CAM), thereby obtaining an enhanced fused feature representation.
[0033] The RD fusion module first processes the RGB and depth feature maps X... RGB , X Depth ∈R C×H×W Preliminary integration is performed, and the mathematical expression is as follows: .
[0034] Subsequently, the initial fusion feature Z′ is further refined using another MS-CAM module to generate the final fusion feature Z: .
[0035] The core of the MS-CAM module lies in simultaneously aggregating global and local contextual information, achieved through a channel attention mechanism, mathematically expressed as: M(X RGB-D )=σ(L(X RGB-D )+G(X RGB-D )).
[0036] Global context features G(X) RGB−D )∈R C×1×1 Calculated using Global Average Pooling (GAP): .
[0037] Local context features L(X) RGB−D )∈R C×H×W L(X) is generated through a bottleneck structure consisting of two pointwise convolutions (PWConv), batch normalization (BN), and ReLU activation. RGB-D )=BN(PWConv2(ReLU(BN(PWConv1(X RGB-D ))))).
[0038] Through the aforementioned multi-scale attention mechanism, the RD fusion module can effectively integrate complementary information from different modalities, significantly enhancing the model's discriminative ability and providing a key technical foundation for high-precision estimation of phenotypic traits such as horizontal diameter, vertical diameter, and weight of litchi fruits.
[0039] An instance segmentation model was applied to both RGB and depth images to automatically extract the geometric features of a lychee. The key extracted geometric features included horizontal diameter, vertical diameter, area, perimeter, minimum circumcircle radius, maximum incircle radius, and shape index. To simplify the computation process, a fusion strategy based on pixel-by-pixel information from the segmentation mask and corresponding region depth data was adopted to accurately estimate each geometric feature.
[0040] The horizontal and vertical diameters of the litchi fruits were calculated using bounding box information extracted by the instance segmentation model. First, the model automatically detected litchi fruits in the image and extracted the depth information of the target region. The bounding box coordinates (x1, y1, x2, y2) defined the region of interest (ROI), and the depth data within the region was used to reconstruct the actual physical dimensions. To improve estimation accuracy, the average depth value within the segmentation mask range was selected for calculation. This was combined with the camera's intrinsic parameters (focal length f). x f y The formulas for calculating the horizontal and vertical diameters of a lychee are as follows: horizontal_diameter = (x2 - x1) · depth avg / f x ,vertical_diameter=(y2-y1)·depth avg / f y , where depth avg It is the average depth value within the mask.
[0041] The model is used to perform operations on objects in RGB and depth images and output a result set. For each result in the result set, the following operations are performed: bounding boxes, confidence scores, class labels, and masks are extracted from the results. Non-maximum suppression is applied to the bounding boxes using the confidence scores, and an index is output. For each index i in the index, the following operations are performed: (x1, y1, x2, y2) is extracted from the i-th extracted bounding box; the cropping depth is output from [y1: y2, x1: x2] of the depth image; the cropping mask is output from [y1: y2, x1: x2] of the mask; the average depth value within the mask is obtained by averaging the portions where both the cropping depth and the cropping mask are greater than 0; and the horizontal and vertical diameters are calculated from this average depth value. Figure 4 Page diagram of the horizontal and vertical diameter prediction software Besides the horizontal and vertical diameters, the calculation of other geometric features requires extracting the contours from the segmentation mask. The specific processing flow is as follows: 1. Mask extraction: Convert the segmentation mask into a binary image for subsequent feature analysis; 2. Contour detection: Extract the largest contour in the binary image as the boundary of the lychee fruit; 3. Depth Calculation: Calculate the average depth value within the masked area; 4. Feature Calculation: Based on the contour and average depth, calculate parameters such as area, perimeter, minimum circumcircle radius, maximum incircle radius, and shape index.
[0042] The area of a lychee fruit is estimated by statistically analyzing the total number of pixels within a segmented mask. Combining the camera's focal length and the average depth of the mask region, the actual area is calculated using the formula: area = A pixels ·depth avg 2 / (f x ·f y ) / 20, where A pixels This represents the total number of pixels within the segmentation mask.
[0043] The perimeter is obtained by summing the number of boundary pixels of the maximum contour. The actual perimeter (in centimeters) is converted based on a scaling factor calculated using camera intrinsics and average depth, thus achieving an accurate estimate of the pixel length to the actual physical length: perimeter = P pixels ·(depth avg / f x +depth avg / f y ) / 20, where P pixels This represents the perimeter in pixels.
[0044] By calculating the minimum enclosing circle, the radius of the minimum enclosing circle that completely encloses the outline of the lychee can be obtained. The formula for calculating the actual radius is: min_enclosing_radius = R enclose ·(depth avg / f x +depth avg / f y ) / 20, where R enclose The pixel radius of the smallest circumcircle.
[0045] The radius of the maximum inscribed circle is obtained by performing a distance transform on the segmented mask region, where the maximum local extremum in the distance transform represents the pixel radius of the inscribed circle. Specifically, it is expressed as: max_inscribed_radius = R max_inscribed ·(depth avg / f x +depth avg / f y ) / 20, where R max_inscribed is the pixel radius of the largest inscribed circle obtained through distance transformation.
[0046] The shape index is a measure of fruit elongation, defined as follows: shape_index=(horizontal_diameter+vertical_diameter) / 20.
[0047] The model is used to perform operations on objects in RGB and depth images and output a result set. For each result in the result set, the following operations are performed: extract bounding boxes, confidence scores, class labels, and masks from the results; perform non-maximum suppression on the bounding boxes using the confidence scores and output an index; for each index i in the index, the following operations are performed: extract (x1, y1, x2, y2) from the i-th extracted bounding box; binarize the i-th mask using a threshold and find the contour; if a contour exists, take the largest contour as the contour of the lychee; calculate the area of the lychee contour in pixels and the number of lychee pixels. The number of pixels for the perimeter of the outline, the pixel radius of the minimum circumscribed circle of the lychee outline, and the mask distance transformation are calculated. The maximum value of the distance transformation is used as the pixel radius of the maximum inscribed circle of the lychee outline. The clipping depth is output from the depth image [y1:y2, x1:x2], and the clipping mask is output from the mask [y1:y2, x1:x2]. The depth image with a clipping mask greater than 0 is used as the effective depth set of the mask. If the effective depth set of the mask is not empty, the mean of the effective depth set of the mask is calculated to obtain the average depth value inside the mask. From this, the area, perimeter, minimum circumscribed circle radius, maximum inscribed circle radius, and shape index are calculated.
[0048] To accurately evaluate the model's predictive performance on litchi size (horizontal and vertical diameters) and weight, precision, recall, mean precision (mAP), root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²) were used. 2 Standardized quantitative indicators such as )
[0049] Precision is defined as the proportion of correctly identified positive instances (True positive, TP) out of all instances identified as positive, including false positive instances (False positive, FP): Precision = TP / (TP + FP).
[0050] Recall measures the proportion of correctly identified positive instances (True positive, TP) out of all actual positive instances (including false negative instances, FN). The formula is: Recall = TP / (TP + FN).
[0051] Average precision (mAP) provides a comprehensive evaluation of the precision and recall tradeoffs at different confidence thresholds, and can fully reflect the overall performance of the model in object detection and segmentation tasks: mAP = AP(Precision, Recall).
[0052] In this study, two mAP-based metrics were used: mAP50 and mAP50-95. mAP50 represents the average accuracy calculated with a fixed Intersection over Union (IOU) threshold of 0.5, reflecting a typical level of detection accuracy. In contrast, mAP50-95 is the mean of multiple sets of average accuracies calculated within an IOU threshold range of 0.5 to 0.95 (with a step size of 0.05), providing a more rigorous and comprehensive evaluation of the model's detection performance.
[0053] In addition, to quantitatively evaluate the model's performance in predicting attributes such as the horizontal diameter, vertical diameter, and weight of litchi, indicators such as root mean square error (RMSE) were introduced.
[0054] RMSE quantifies the magnitude of model prediction error by measuring the difference between predicted and actual values. Its calculation formula is as follows: .
[0055] Mean Absolute Error (MAE), as another measure of prediction accuracy, is used to calculate the average magnitude of the absolute values of all prediction errors, regardless of the direction of the error. Its formula is: .
[0056] Coefficient of determination (R) 2 R is used to evaluate the ability of a regression model to explain the variance of observed data. It represents the proportion of variance in the observed data that the model successfully predicts. Its value ranges from 0 to 1, and a higher value indicates a better model fit. 2 The calculation formula is: , where Observation is the mean of the observed values.
[0057] In summary, these evaluation indicators together construct a comprehensive and rigorous evaluation system that can effectively measure the accuracy and practicality of the model in predicting key traits such as litchi size and weight in real-world applications.
[0058] like Figure 5As shown, to systematically evaluate the performance of the proposed LitchiPhenoNet model compared with mainstream instance segmentation models (YOLO5s, YOLO6s, YOLO8s) in the litchi phenotypic recognition task, quantitative and qualitative analyses were performed on the models. All tested models achieved high levels of precision and recall (0.998-1.000) for litchi fruit, seed, and pulp categories, indicating comparable basic detection capabilities. However, at the more discriminative mAP50-95 metric, LitchiPhenoNet demonstrated a significant technical advantage.
[0059] Specifically, the LitchiPhenoNet model achieved an overall Mask mAP50-95 of 0.886, a 2.4% improvement over YOLO5s and a 1.6% improvement over both YOLO6s and YOLO8s. In terms of specific categories, LitchiPhenoNet achieved mAP50-95 values of 0.950 (1.2%-2.0% improvement) for litchi fruit, 0.813 (2.4%-4.4% improvement) for the seed, and 0.895 (0.7%-1.5% improvement) for the pulp, all significantly outperforming other comparative models. These data demonstrate that the model of this invention possesses stronger phenotypic recognition capabilities and better generalization in multi-category and complex contexts.
[0060] like Figure 6 As shown in the visual analysis results, LitchiPhenoNet demonstrates a significant advantage in segmentation accuracy, consistent with the quantitative evaluation results obtained from the mAP50-95 index. Yellow arrows indicate areas with recognition deficiencies. While YOLO models can effectively detect and classify fruits and related structures, they still have shortcomings in segmenting fine structures, especially in handling the boundaries of fine-grained regions such as peel edges and the separation of flesh and pit. In contrast, LitchiPhenoNet can consistently output segmentation results with clear boundaries and rich details, effectively overcoming the limitations of YOLO models in terms of detailed segmentation accuracy. The above qualitative analysis further corroborates the comprehensive performance advantages of LitchiPhenoNet in instance segmentation tasks.
[0061] Therefore, the improved accuracy of LitchiPhenoNet in instance segmentation not only enhances the reliability of the segmentation itself but also lays the foundation for high-precision extraction of key geometric features of litchi (such as horizontal diameter, vertical diameter, perimeter, and area). Considering that the accuracy of the segmentation results directly affects the extraction accuracy of these features, this method effectively improves the accuracy and stability of subsequent fruit weight estimation and phenotypic parameter prediction, and has significant practical application value.
[0062] like Figure 7As shown, the accuracy of the LitchiPhenoNet model in automatically estimating the transverse and lateral diameters (i.e., horizontal and vertical diameters) of litchi fruit and seeds was evaluated. Quantitative analysis results show that LitchiPhenoNet achieved the lowest mean absolute error (MAE) and root mean square error (RMSE) on all relevant traits, comprehensively outperforming various YOLO basic models.
[0063] Specifically, for the transverse diameter of the whole fruit, LitchiPhenoNet's MAE was 1.03 mm and RMSE was 1.36 mm, representing reductions of 2.9% and 5.6% respectively compared to the optimal YOLO variant; the MAE for the lateral diameter was 1.73 mm and RMSE was 2.16 mm, representing reductions of 0.6% and 4.1% respectively. For core measurement, the accuracy improvement was even more significant, with the transverse diameter MAE and RMSE decreasing to 0.62 mm and 0.81 mm respectively, and the lateral diameter MAE and RMSE decreasing to 0.68 mm and 0.88 mm respectively, representing reductions of 7.5%-12.0% compared to the corresponding YOLO series indicators.
[0064] Regarding correlation metrics, the coefficient of determination (R²) of LitchiPhenoNet is... 2 To achieve the optimal value, the transverse diameter R of the whole fruit is [value missing]. 2 The value is 0.97, and the transverse diameter and lateral diameter R of the pit are... 2 Both are 0.98, whole fruit lateral diameter R 2 The accuracy is 0.91, which is at or above the level of the YOLO comparison model. These results demonstrate that our method can achieve automatic estimation of key phenotypic traits of litchi with millimeter-level precision, exhibiting high accuracy and stability.
[0065] To achieve high-precision automatic estimation of litchi fruit weight, a multimodal regression model was employed, fusing the RGB and depth information of litchi fruits with key geometric phenotypic features (including transverse diameter, lateral diameter, area, perimeter, minimum circumscribed circle radius, maximum inscribed circle radius, shape index, and seed-to-fruit ratio). The selection of these geometric features was based on existing research confirming that these phenotypic traits provide effective allometric growth model support for fruit weight prediction.
[0066] like Figure 8 As shown, the coefficient of determination (R²) of the pulp using only the geometric features model (GEO) is... 2 The R² value was 0.9353, while the multimodal regression model (LitchiPhenoNet: RGB-D + GEO) proposed in this invention significantly improved the prediction accuracy, with R² value exceeding 0.9353. 2The accuracy was improved to 0.9554, and the MAE decreased from 1.8345g to 1.5100g, while the RMSE decreased from 2.4101g to 1.9999g. These improvements demonstrate that the multimodal regression framework can effectively utilize the complementary advantages of multi-source information, significantly improving model accuracy even in the prediction of phenotypic traits with weak correlations.
[0067] Compared to YOLO-based models (YOLO5s, YOLO6s, and YOLO8s), the multimodal LitchiPhenoNet proposed in this invention demonstrates higher estimation accuracy in predicting all evaluated phenotypic traits. Specifically, LitchiPhenoNet significantly reduces both MAE and RMSE values, improving whole fruit accuracy by 5.1%–12.2%, core accuracy by 6.1%–40.8%, and pulp accuracy by 7.5%–17.5%; simultaneously, the coefficient of determination (R²) also improves accuracy. 2 The values also reached higher levels, reaching 0.9786 (whole fruit), 0.9371 (fruit pit) and 0.9554 (fruit pulp) respectively.
[0068] These results collectively demonstrate the effectiveness and robustness of the proposed multimodal regression weight estimation framework. The significant improvements observed, particularly in pulp weight estimation, highlight the importance of multimodal feature fusion, providing strong evidence for the accuracy and reliability of litchi phenotypic analysis and demonstrating the significant advantages of litchi phenotypic networks.
[0069] The above specific embodiments are merely several preferred embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A method for litchi phenotypic recognition and weight estimation based on multimodal learning, comprising a phenotypic recognition stage and a weight estimation stage, characterized in that: In the phenotypic recognition stage, RGB images and depth images are used as bimodal inputs and fed into the LitchiPhenoNet model to achieve fruit detection and segmentation. In the weight estimation stage, a multimodal regression method is used to estimate the fruit weight based on the geometric features and abstract intermediate features directly extracted by LitchiPhenoNet.
2. The litchi phenotypic recognition and weight estimation method based on multimodal learning according to claim 1, wherein the weight estimation stage is characterized in that: In the last layer of the LitchiPhenoNet model, the intermediate feature map is flattened into a one-dimensional vector as an abstract intermediate feature. The abstract intermediate feature and the geometric feature are then used as inputs to the regression model and processed through two fully connected layers respectively. The abstract intermediate feature and the geometric feature are then concatenated and fused through an additional fully connected layer. Finally, the predicted weight of the fruit is output through the regression network.
3. The litchi phenotypic recognition and weight estimation method based on multimodal learning according to claim 1, wherein the phenotypic recognition stage is characterized in that: The region of interest (ROI) is located using the bounding box and segmentation mask output by the LitchiPhenoNet model. The ROI mask is then combined with calibrated camera intrinsics to extract the geometric phenotypic features of the fruit from the RGB and depth data.
4. The method for litchi phenotypic recognition and weight estimation based on multimodal learning according to claim 2, characterized in that: During the training of the regression network, the regression model optimizes its parameters by minimizing the mean squared error between the predicted value and the actual weight.
5. The method for litchi phenotypic recognition and weight estimation based on multimodal learning according to claim 3, characterized in that: Geometric phenotypic features include horizontal diameter, vertical diameter, area, perimeter, minimum circumscribed circle radius, maximum inscribed circle radius, and shape index.
6. The method for litchi phenotypic recognition and weight estimation based on multimodal learning according to claim 1, characterized in that: The LitchiPhenoNet model contains two independent feature extraction branches, which extract features for RGB modality and deep modality respectively. The RD fusion module uses a multi-scale channel attention module to perform deep fusion of multi-modal features.
7. The method for litchi phenotypic recognition and weight estimation based on multimodal learning according to claim 5, characterized in that: The horizontal and vertical diameters of the litchi fruit were calculated using bounding box information extracted from the LitchiPhenoNet model.
8. The method for litchi phenotypic recognition and weight estimation based on multimodal learning according to claim 5, characterized in that: The area, perimeter, minimum circumcircle radius, maximum incircle radius, and shape index of lychee fruits are calculated by extracting the contour from the segmentation mask.
9. The method for litchi phenotypic recognition and weight estimation based on multimodal learning according to claim 7, characterized in that: For the calculation of horizontal and vertical diameters, the method is as follows: The LitchiPhenoNet model automatically detects litchi fruits in images and extracts depth information of the target region. It uses bounding box coordinates to define the region of interest, and the depth data within the region is used to reconstruct the actual physical size. The average depth value within the segmentation mask range is selected for calculation, and the horizontal and vertical diameters of the litchi are calculated in combination with the camera's intrinsic parameters.
10. The method for litchi phenotypic recognition and weight estimation based on multimodal learning according to claim 8, characterized in that: For the calculation of area, perimeter, minimum circumscribed circle radius, maximum inscribed circle radius, and shape index, the method is characterized in that: The processing flow includes mask extraction, contour detection, depth calculation, and feature calculation.
Citation Information
Cited By
Image-based legume plant phenotype analysis method and system
CN121438010A
Image-based method and system for phenotyping legume plants
CN121438010B