Method suitable for intelligent recognition and classification of Chinese torreya maturity under unmanned aerial vehicle perspective
By fusing visible light and near-infrared images from the perspective of drones and using a lightweight network, the problems of insufficient feature representation and poor robustness to occlusion in the maturity identification of Torreya grandis were solved, and accurate maturity identification and real-time positioning were achieved in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUZHOU VOCATIONAL TECH COLLEGE
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies for identifying the maturity of Torreya grandis from the perspective of drones suffer from problems such as insufficient feature expression, low accuracy in maturity differentiation, poor robustness to occlusion, insufficient real-time performance, and high equipment costs. They are particularly difficult to achieve accurate identification in complex mountainous orchard environments.
Simultaneous acquisition of visible and near-infrared dual-modal images was adopted. Background interference was shielded by dark channel prior dehazing and canopy layering mask generation. A dual-channel YOLOv8-Lite lightweight network was constructed for feature adaptive alignment and weighted fusion. Combined with a cross-modal attention module and a three-classification decision head, fruit maturity was determined. Occlusion and adhesion were eliminated by non-maximum suppression and clustering correction to achieve geographic coordinate mapping.
It significantly improves the accuracy and real-time performance of Torreya grandis maturity recognition from the perspective of UAVs, and can accurately distinguish between immature, mature and overripe fruit in complex environments, providing stable recognition results and positioning information, and adapting to the real-time reasoning needs of UAV onboard platforms.
Smart Images

Figure CN122289974A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of technology, and specifically to a method for intelligent identification and classification of the maturity of Torreya grandis from the perspective of unmanned aerial vehicles (UAVs). Background Technology
[0002] Torreya grandis, a rare and precious dried fruit and woody oilseed crop unique to my country, has extremely high economic value. Accurate maturity identification is a core prerequisite for intelligent harvesting, yield prediction, and quality grading. Currently, maturity assessment of Torreya grandis mainly relies on manual visual observation, which suffers from strong subjectivity, low efficiency, large errors, and high risks associated with high-altitude operations, making it unsuitable for the needs of a large-scale, standardized modern Torreya grandis cultivation industry. With the rapid development of smart agriculture and drone remote sensing technology, fruit detection methods based on machine vision and deep learning are gradually being applied in the agricultural and forestry fields. However, Torreya grandis fruits are characterized by their small size, dense clustering, color similar to branches and leaves, subtle color changes at maturity, and severe canopy shading. Conventional RGB single-modal vision solutions are prone to missed detections, false detections, and misclassification of maturity, failing to directly meet the accurate identification requirements of drones in high-altitude scenarios. Most current forest fruit identification methods only focus on whether the target is detected, lacking a mechanism for mining weak spectral and textural features for the maturity of Torreya grandis. Their insufficient adaptability to complex mountain orchard environments has become a key technical bottleneck restricting the intelligent operation of Torreya grandis.
[0003] In recent years, multi-source image fusion and lightweight deep learning technologies have provided new pathways for the accurate identification of forest fruits. Visible and near-infrared dual-modal imaging can simultaneously preserve external texture and internal physiological spectral information, showing great potential in fruit quality detection. Existing research mostly uses YOLO series models to achieve fruit target detection, improving recognition accuracy in small targets and occluded scenes through lightweight modifications, attention mechanism embedding, and multi-scale feature fusion. However, for Torreya grandis canopy images from UAV aerial perspectives, there are still problems such as missing preprocessing, weak background interference suppression, and coarse dual-modal feature alignment and fusion strategies. Although near-infrared spectroscopy can reflect differences in the internal composition of fruits, it is easily affected by lighting, posture, and terrain under dynamic UAV acquisition conditions, resulting in insufficient spectral feature stability. At the same time, existing networks have not constructed dedicated feature extraction and decision-making mechanisms for Torreya grandis maturity grading, resulting in blurred classification boundaries and difficulty in distinguishing the subtle differences between immature, mature, and overripe states. It is difficult to balance lightweight models with detection accuracy, and the real-time inference performance on the edge is insufficient to support online recognition and positioning output on UAV-borne platforms.
[0004] Chinese patent CN118262229A discloses a method for detecting alfalfa seed pod maturity based on improved YOLOv8 drone aerial images. This method improves the accuracy of crop maturity detection in the field by optimizing the dataset, improving the model structure, and adjusting the loss function, and achieves the distinction between multiple maturity states, thus improving the detection effect of small targets in complex backgrounds to some extent. However, this technology only uses a single RGB visual modality and does not introduce near-infrared spectral information, which is insufficient for expressing the features of Torreya grandis fruits with weak maturity differences and high integration with the background. It also lacks a canopy layer mask and dual-modal feature fusion mechanism, and its occlusion robustness and maturity fine classification accuracy cannot meet the requirements of Torreya grandis recognition. Chinese patent CN117465707A discloses a tethered drone and harvesting method for fruit picking. This patent optimizes the fruit positioning accuracy by fusing stereo vision and depth information, solves the positioning deviation problem of single RGB image, and improves the target positioning reliability of drone harvesting operations. However, this solution relies on laser depth sensors, which are expensive and lack spectral information. It cannot perform precise classification of Torreya grandis maturity and lacks lightweight network and cross-modal feature alignment strategies. It is easily interfered with in the complex environment of Torreya grandis forests in mountainous areas, and its real-time performance and recognition accuracy are difficult to meet the operational requirements.
[0005] Therefore, there is an urgent need for an intelligent identification and classification method for Torreya grandis maturity that is oriented towards the perspective of drones and integrates visible light and near-infrared dual-modal information, in order to solve the problems of insufficient feature expression, low maturity differentiation accuracy, poor occlusion robustness, insufficient real-time performance and high equipment cost in existing technologies. Summary of the Invention
[0006] Based on the aforementioned technical problems, this application discloses an intelligent identification and classification method for the maturity of Torreya grandis from the perspective of unmanned aerial vehicles, specifically including:
[0007] Simultaneously acquire RGB visible light images and near-infrared irradiance images of the canopy of Torreya grandis from the perspective of UAV, and construct a multi-source heterogeneous input dataset;
[0008] Dark channel prior dehazing and canopy layering mask generation are performed on RGB images to shield non-target branches and leaves and complex background interference.
[0009] After Gaussian difference filtering, the spectral reflectance difference characteristics of Torreya grandis fruit and leaves are extracted from the near-infrared image to generate a fruit saliency heatmap.
[0010] A dual-channel YOLOv8-Lite lightweight network is constructed, with the visible light channel used for small target localization and the near-infrared channel used for maturity feature discrimination. Feature adaptive alignment and weighted fusion are achieved through a cross-modal attention module.
[0011] Based on the spectral threshold and color gradient features corresponding to the maturity of Torreya grandis, a three-classification decision head is constructed to output the category information of immature, mature and overripe fruits in real time.
[0012] Non-maximum suppression and clustering correction are applied to the same cluster of Torreya grandis to eliminate misclassification caused by occlusion and adhesion, and output accurate maturity classification results and location coordinates from the perspective of UAV.
[0013] Preferably, the step of performing dark channel prior dehazing and canopy layering mask generation on the RGB image specifically includes: calculating the dark channel map of the RGB image. The formula is: ,in, Represents pixels In color channels The strength, Therefore A local window centered on the atmospheric scattering model is used to estimate global atmospheric light values. and transmittance Guided filtering is used to refine the transmittance map, and a clear image after dehazing is reconstructed. The formula is: ,in, A constant threshold is used to prevent the denominator from being too small; a vegetation index mask is constructed based on the ratio of the green component to the near-infrared component of the reconstructed image. The tree canopy region is extracted by adaptive threshold segmentation, and the pixel values of the non-tree canopy region are set to zero to generate a tree canopy layering mask.
[0014] Preferably, the step of extracting the spectral reflectance difference features of Torreya grandis fruit and leaves after Gaussian difference filtering of the near-infrared image specifically includes: processing the near-infrared irradiance image... Application of multi-scale Gaussian difference operator To enhance the details of the fruit's edges and textures, the formula is: ,in, The standard deviation is Gaussian kernel function, This represents the convolution operation. and These are the scale parameters for capturing the fruit's outline and internal texture, respectively; combined with the red band reflectance in the RGB image. Near-infrared band reflectivity Calculate the modified normalized difference index To distinguish mature fruit from background leaves, the formula is: ,in, This represents the maximum value of the multi-scale Gaussian difference response. This is the texture suppression coefficient. As a smoothing factor; based on Spatial distribution of fruit saliency heatmap Its pixel value The formula is: ,in, It is the Sigmoid activation function. and These are the spectral weights and the second-order gradient weights, respectively.
[0015] Preferably, the construction of the dual-channel YOLOv8-Lite lightweight network includes a visible light coding branch. With near-infrared coding branch ;
[0016] The visible light encoding branch An improved CSPDarknet structure is adopted, and a coordinate attention mechanism is embedded to enhance the spatial position awareness of small targets;
[0017] The near-infrared coding branch The lightweight MobileNetV3 backbone is used to extract spectral texture features; the two branches output feature maps separately before the feature fusion stage. and .
[0018] Preferably, the visible light coding branch With near-infrared coding branch Both employ dynamically depthwise separable convolution operations, with their convolution kernel weights... Dynamically generated based on channel statistics of the input feature map: ,in, The channel selection coefficients are generated through global average pooling and the fully connected layer. For predefined The group of basic convolutional kernels enables the receptive field size and feature extraction granularity to be adaptively adjusted according to the texture complexity of Torreya grandis fruits at different growth stages.
[0019] Preferably, the step of achieving adaptive feature alignment and weighted fusion through a cross-modal attention module specifically includes: constructing a cross-modal gating attention unit, and... and The layers are concatenated along the channel dimension, and a gated weight matrix is generated through a shared fully connected layer. : Utilizing gating weights Dynamically adjust the contribution of dual-channel features to generate a fused feature map. The formula is: ,in, This represents element-wise multiplication. The relevance coefficient is a learnable coefficient. For the cross-correlation attention operator, the formula is: in, They are respectively from The generated query matrix and by The generated key-value matrix, This is the scaling factor.
[0020] Preferably, the construction of a three-classification decision head based on the spectral threshold and color gradient features corresponding to the maturity of Torreya grandis specifically includes: introducing a multi-task loss function at the end of the detection head. Simultaneously, the positioning regression and maturity classification are optimized, with the following formula: in, For CIoU loss, FocalLoss is the classification loss. The formula for the spectral consistency constraint loss is: ,in, For the network prediction of the first The spectral feature vector of each target This is the standard spectral prototype vector corresponding to the maturity category. Represents the spectral gradient operator. The gradient penalty coefficient is minimized. This forces the features learned by the network to cluster in the spectral space toward the prototype center of their respective maturity categories.
[0021] Preferably, the three-class decision head employs a temperature scaling calibration strategy during the inference phase to adjust the output class probability distribution. Perform calibration: ,in, The first output of the logic layer Unnormalized score The temperature parameters are obtained through optimization using the validation set; combined with the confidence threshold. and Only when or Only when the ripeness is reached will the corresponding picking signal be triggered; otherwise, it will be marked as unripe.
[0022] Preferably, the non-maximum suppression and clustering correction for the same cluster of Torreya grandis specifically includes: constructing a weighted adjacency matrix based on spatial density and confidence. For the set of detection boxes Any two boxes Edge weights between The calculation is as follows: ,in, The coordinates of the detection box center are, For classification confidence, and Hyperparameters for controlling sensitivity; based on The spectral clustering algorithm is used to group spatially adjacent and feature-similar detection boxes into the same fruit cluster. ; in each cluster Within the cluster, the detection box with the smallest distance between the spectral features and the maturity prototype is selected as the representative box, and the remaining redundant boxes in the cluster are removed to complete the occlusion and adhesion correction.
[0023] Preferably, it also includes geographic coordinate mapping, specifically: using the drone's POS data to obtain the latitude and longitude of the shooting time. ,altitude and attitude angle Constructing from image pixel coordinate system To the world geographic coordinate system Projection transformation model: ,in, For the camera intrinsic parameter matrix, For rotation matrix, It is a translation vector. The estimated average height of the canopy; world coordinates Convert to WGS84 geographic coordinates and associate with the output maturity classification results.
[0024] Compared with the prior art, the technical solution of this application has the following technical effects:
[0025] This invention can simultaneously acquire visible light and near-infrared dual-modal images and complete the stable construction of multi-source heterogeneous data. Through image dehazing and canopy layering mask generation, it effectively shields complex background interference, significantly improving the purity of input data and target saliency.
[0026] This invention achieves efficient extraction of dual-modal features through a dual-channel lightweight network and a dynamic convolutional structure. It leverages cross-modal attention units to complete adaptive feature alignment and weighted fusion, thereby enhancing the ability to express the weak maturity features of Torreya grandis and improving the stability of feature perception and discrimination in occlusion and dense growth scenarios.
[0027] Based on multi-task constraints and classification probability calibration mechanisms, this invention can accurately distinguish between three types of Torreya grandis fruits: immature, mature, and overripe. This makes the decision output more consistent with the actual maturity distribution. At the same time, through spectral clustering and optimized post-processing, it effectively eliminates misclassification caused by fruit adhesion and occlusion, ensuring the consistency and reliability of the recognition results.
[0028] This invention can accurately map image pixel coordinates to geospatial coordinates, and output complete Torreya grandis maturity classification results and location information. It is adapted to the real-time reasoning and operation guidance needs of UAV on-board terminals, and provides stable and efficient technical support for intelligent monitoring and precise harvesting of Torreya grandis orchards.
[0029] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings.
[0030] The above and other objects, advantages and features of this application will become more apparent to those skilled in the art from the following detailed description of specific embodiments in conjunction with the accompanying drawings. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0032] Based on the description of the figures and their corresponding technical content in the document, the titles of the figures are as follows:
[0033] Figure 1 Flowchart of the intelligent identification and classification method for maturity of Torreya grandis from the perspective of drones;
[0034] Figure 2 Flowchart of RGB image dark channel dehazing and canopy layering mask generation preprocessing;
[0035] Figure 3 Schematic diagram of Gaussian difference filtering for near-infrared images and generation principle of fruit saliency heatmap;
[0036] Figure 4 A schematic diagram of the dual-channel YOLOv8-Lite lightweight network architecture and cross-modal feature fusion;
[0037] Figure 5 Flowchart of spectral clustering optimization nonmaximum suppression and geographic coordinate mapping post-processing;
[0038] Figure 6 Visual comparison chart of Torreya grandis fruit detection results and tree canopy mask matching;
[0039] Figure 7 Comparison of PR curves for various algorithms at different maturity levels and in foggy weather scenarios;
[0040] Figure 8 A visual comparison chart of the feature responses of multiple algorithms and the cross-modal fusion effect of this application;
[0041] Figure 9 Comparison of geographical spatial distribution and maturity classification results of Torreya grandis fruits;
[0042] Figure 10 Robustness test curves of the algorithm under dual variables of light intensity and fog concentration. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. In the following description, specific details such as specific configurations and components are provided merely to help fully understand the embodiments of this application. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. In addition, for clarity and brevity, descriptions of known functions and structures are omitted in the embodiments.
[0044] It should be understood that the phrase "an embodiment" or "this embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "an embodiment" or "this embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0045] Furthermore, reference numerals and / or letters may be repeated in different examples within this application. Such repetition is for the purpose of simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or settings discussed.
[0046] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" in this article describes another type of relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the related objects before and after it are in an "or" relationship.
[0047] In this article, the term "at least one" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, "at least one of A and B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.
[0048] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion.
[0049] Example 1
[0050] This embodiment mainly describes a method for intelligent identification and classification of the maturity of Torreya grandis from the perspective of a drone, such as Figure 1 As shown, it specifically includes:
[0051] Simultaneously acquire RGB visible light images and near-infrared irradiance images of the canopy of Torreya grandis from the perspective of UAV, and construct a multi-source heterogeneous input dataset;
[0052] Dark channel prior dehazing and canopy layering mask generation are performed on RGB images to shield non-target branches and leaves and complex background interference.
[0053] After Gaussian difference filtering, the spectral reflectance difference characteristics of Torreya grandis fruit and leaves are extracted from the near-infrared image to generate a fruit saliency heatmap.
[0054] A dual-channel YOLOv8-Lite lightweight network is constructed, with the visible light channel used for small target localization and the near-infrared channel used for maturity feature discrimination. Feature adaptive alignment and weighted fusion are achieved through a cross-modal attention module.
[0055] Based on the spectral threshold and color gradient features corresponding to the maturity of Torreya grandis, a three-classification decision head is constructed to output the category and location information of immature, mature and overripe fruits in real time.
[0056] Non-maximum suppression and clustering correction are applied to the same cluster of Torreya grandis to eliminate misclassification caused by occlusion and adhesion, and output accurate maturity classification results and geographic coordinates from the perspective of UAV.
[0057] Furthermore, such as Figure 2 As shown, RGB visible light images and near-infrared irradiance images of the Torreya grandis canopy were simultaneously acquired from a UAV perspective to construct a multi-source heterogeneous input dataset. Using a UAV flight platform equipped with a dual-light pod, and under preset flight altitude and overlap parameters, the exposure times of the RGB full-color camera and the near-infrared narrowband camera were strictly synchronized via hardware trigger signals to ensure millisecond-level synchronization of the two images in the time domain, thus eliminating parallax caused by the high-speed movement of the UAV. The RGB camera acquisition resolution was [resolution missing]. The raw image data was used to record the reflected radiation intensity in the red (600-700nm), green (500-600nm), and blue (400-500nm) bands, capturing the apparent color and texture of the Torreya grandis aril as it changes from bluish-brown to purplish-brown; the near-infrared camera acquired data with a resolution of [missing information]. The raw image data was used to record the near-infrared radiation energy with a center wavelength of 850 nm and a half-width of 40 nm, reflecting the differences in spectral reflectance caused by changes in fruit internal water content and cell wall structure. When constructing the multi-source heterogeneous input dataset, radiometric calibration was performed on the two RAW format data streams, using laboratory-calibrated gain coefficients. and bias Quantize the digital value Converted to physical radiance Eliminate sensor response nonlinearity errors; perform geometric fine correction based on distortion parameters obtained from the calibration plate. The images are distorted, and the SIFT algorithm is used to extract corresponding feature points from the two images to calculate the homography matrix. The RGB and near-infrared images are resampled to a unified pixel coordinate system, forming a spatially strictly registered dual-channel image pair. The images were compared with the latitude and longitude coordinates recorded by the drone POS system. ,altitude Attitude angle Metadata is bound to timestamps to construct a structured tensor dataset containing multidimensional spectral information and spatial location information, which serves as a standardized input for subsequent deep learning models.
[0058] Furthermore, dark channel prior dehazing and canopy layering mask generation are performed on the RGB image to shield non-target branches and leaves and complex background interference. This is achieved by applying a priori dark channel dehazing to the registered RGB image. Perform dark channel prior calculations and define a local window. In pixels Centered on, with radius Calculate the dark channel map for a square region. By utilizing the physical statistical law that the intensity of fog in the Torreya grandis forest area approaches zero in dark passages, the scene transmittance is estimated. Based on atmospheric scattering model ,in The global atmospheric light value is estimated by selecting the average pixel value of the original image corresponding to the top 0.1% of the brightest pixels in the dark channel image. Subsequently, a guided filtering algorithm was used to extract the original image. To obtain a refined transmittance map, the initial rough transmittance map is smoothed to preserve its edges. To avoid halo effects, reconstruct a clear, dehazed image. The formula is ,in To prevent the denominator from being too small and to restore high-frequency details and color saturation attenuated by fog, a vegetation index mask is constructed based on the dehazed image, and normalized differential vegetation index variants are calculated. The adaptive segmentation threshold is automatically calculated by combining Otsu's maximum inter-class variance method. Generate a binary tree canopy mask like Otherwise, it is 0; the mask is multiplied pixel by pixel with the dehazed image, and the pixel values of non-canopy areas such as sky, mountain rocks and ground weeds are forced to zero, retaining only the effective data of the canopy layer, and physically shielding the background noise interference from the input end.
[0059] Furthermore, such as Figure 3 As shown, after Gaussian difference filtering, the spectral reflectance difference characteristics of Torreya grandis fruits and leaves are extracted from the near-infrared image to generate a fruit saliency heatmap. This is then applied to the denoised near-infrared image. Applying the multi-scale difference of Gaussians (DoG) filtering operator, two Gaussian kernel functions with different standard deviations are defined. and ,in Used to capture the small-scale outline edges of Torreya grandis fruits. Used to extract large-scale texture background of leaves and calculate differential response maps. This method utilizes bandpass filtering to enhance high-frequency signals at the fruit edges and suppress low-frequency background noise. It also incorporates visible light red band reflectance. With near-infrared reflectivity Construct a modified normalized difference index ,in This represents the maximum value of the multi-scale difference response. This is the texture suppression coefficient, used to eliminate spurious responses from highly textured leaves. This is the numerical stability factor. Based on... Spatial distribution constructs saliency heatmap Its pixel value The weighted sum of spectral intensity and second-order gradient information is mapped using the Sigmoid activation function, and the formula is: ,in For spectral weights, For edge gradient weights, The generated heatmap The region exhibits a high response peak (close to 1) in the Torreya grandis fruit area and a low response valley (close to 0) in the background area, which serves as an explicit spatial attention prior map input to the subsequent neural network.
[0060] Furthermore, such as Figure 4As shown, a dual-channel YOLOv8-Lite lightweight network is constructed. The visible light channel is used for small target localization, and the near-infrared channel is used for maturity feature discrimination. Feature adaptive alignment and weighted fusion are achieved through a cross-modal attention module. A dual-stream coding architecture is constructed, including a visible light coding branch. With near-infrared coding branch . The branch uses an improved CSPDarknet-Lite backbone network, with the input being... The RGB image is processed into four stages (Stage 1-4). Each stage consists of a Focus module, a CBS convolutional block (Conv+BN+SiLU), and a C2f module. Coordinate attention is embedded in Stage 3 and Stage 4. Orientation-aware feature maps are generated by performing global average pooling in the horizontal and vertical directions, and then mapped through a shared convolutional layer to generate a spatial weight map. This enhances the network's sensitivity to the geometric position of tiny Torreya grandis fruits; Branch input is Near-infrared images and Significance heatmap (Synthesized into 2 channels), using the lightweight MobileNetV3-Small backbone, it extracts spectral texture features unique to the near-infrared band using depthwise separable convolution, reducing the number of parameters. Both branches output feature maps of the same resolution in the neck region. and ( Then, the input is fed into the cross-modal gated attention fusion module (CM-GAM), which first... and splicing along the channel dimension, via Convolution dimensionality reduction to Then, channel description vectors are generated through global average pooling and two fully connected layers (including ReLU activation), and finally, a gated weight matrix is generated through the Sigmoid function. ,in . use Dynamically adjust the feature contribution of the two channels and perform fusion calculation. Among them, the cross-correlation attention operator , , , , To generate learnable correlation coefficients, a fusion feature map is generated that combines precise localization capabilities with deep biochemical discrimination capabilities. Input to the detection head.
[0061] Furthermore, based on the spectral threshold and color gradient features corresponding to the maturity of Torreya grandis, a three-classification decision head is constructed to output the category and location information of immature, mature, and overripe fruits in real time. A decoupled three-classification decision head is designed at the end of the detection network, containing independent regression and classification branches. The regression branch uses three... Convolutional layer predicts the center coordinate offset of the bounding box Width and height scaling factors and confidence score; classification branches are based on three The convolutional layer outputs logits vectors for three categories: immature, mature, and overripe. A multi-task composite loss function is introduced during the training phase. Perform end-to-end backpropagation optimization. Using CIoU loss, the calculation formula is as follows: Minimize the overlap area between the predicted bounding box and the ground truth bounding box, the Euclidean distance between the center points, and the consistency of the aspect ratio; Focal Loss is used. This addresses the imbalance between positive and negative samples by focusing on gradient updates for difficult-to-distinguish samples; the core of this approach lies in spectral consistency constraint loss. It is defined as the predicted spectral feature vector extracted by the intermediate layer of the network. Corresponding to the predefined maturity category standard prototype vector The weighted sum of the squared Euclidean distances and the spectral gradient differences between them, i.e. ,in This represents the gradient operator along the spectral dimension. is the gradient penalty coefficient. This loss term forces the network to cluster the feature distributions of the three fruit classes towards their respective spectral prototype centers in the latent feature space, while preserving the boundaries of spectral gradient differences between classes. During the inference phase, the logits output by the classification branch are Softmax normalized to obtain the probability distribution. By combining the coordinate information output by the regression branch, a structured detection result containing category labels, confidence scores, and bounding box coordinates is output in real time.
[0062] Furthermore, such as Figure 5 As shown, nonmaximum suppression and clustering correction are applied to the same cluster of Torreya grandis to eliminate misclassification caused by occlusion and adhesion. The results output accurate maturity classification and geographic coordinates from the perspective of the UAV. A weighted adjacency matrix based on spatial density and confidence is constructed. For the set of detection boxes Any two boxes in Calculate edge weights ,in The coordinates of the detection box center are, For classification confidence, and Pixels are hyperparameters controlling sensitivity, quantifying the spatial overlap, geometric proximity, and semantic credibility between bounding boxes. Based on Perform spectral clustering algorithm to construct Laplacian matrix ( (for degree matrix), for Perform eigenvalue decomposition and take the first few features. The feature space is formed by the feature vectors corresponding to the smallest feature values. The K-Means algorithm is used to automatically divide the spatially adjacent and feature-similar detection boxes into independent fruit clusters. This simulates the clustered topology of naturally growing Torreya grandis. In each cluster... Internally, instead of simply ranking by confidence level for suppression, the cosine distance between the spectral feature vector of each detection box within a cluster and the corresponding maturity prototype vector is calculated. The box with the smallest distance (i.e., the one with the most typical spectral features) is selected as the unique representative box for that cluster, and the remaining redundant boxes are eliminated. This solves the problem that traditional non-maximum suppression is prone to mistakenly deleting high-value mature fruits under severe occlusion. Finally, the latitude and longitude of the shooting time recorded by the drone's POS data are used. ,altitude and attitude angle Constructing from image pixel coordinate system To the world geographic coordinate system The rigorous projection transformation model: ,in For the camera intrinsic parameter matrix, For rotation matrix, It is a translation vector. The estimated average height of the canopy layer; convert world coordinates to WGS84 geographic coordinates. The revised maturity category labels are then bound to geographic coordinates to output accurate classification results.
[0063] This implementation improves the detection accuracy and maturity discrimination capability of small targets of Torreya grandis in complex mountainous forest environments by simultaneously acquiring dual-modal images, dehazing and masking preprocessing, generating saliency heatmaps, fusing lightweight dual-channel network features, decoupling the three-classification decision head, and performing spectral clustering postprocessing and geographic mapping. It effectively suppresses false positives and false negatives caused by background interference and occlusion adhesion, while achieving lightweight model and real-time inference. It can stably output maturity classification results with geographic coordinates and has strong robustness, high accuracy and practicality for airborne deployment.
[0064] Based on Example 1, this example verifies the applicability of the intelligent identification and classification method for Torreya grandis maturity from an UAV perspective. It constructs a Zhuji Torreya grandis multimodal full-cycle fine dataset (ZJ-SGF-MultiModal-2025), covering three complete maturity seasons. Figure 6 As shown. The data acquisition platform uses a customized hexacopter industrial drone equipped with a rigorously synchronized dual-light pod. The visible light sensor is a Sony IMX477 global shutter camera (resolution 4096×3072, pixel size 1.55μm), and the near-infrared sensor is a modified Global Shutter NIR camera (center wavelength 850nm, half-width 40nm, quantum efficiency >65%). The entire dataset contains 158,420 original image pairs, which were jointly annotated by professional agricultural experts and remote sensing algorithm engineers, generating a total of 482,650 high-confidence bounding box annotations. These annotations cover three states: immature (Green), mature (Purple-Brown), and overripe (Dark-Brown / Decayed). Additional fine-grained attribute labels were added, including occlusion level (no occlusion, mild occlusion <30%, moderate occlusion 30%-60%, severe occlusion >60%) and lighting conditions (front lighting, side lighting, backlighting, diffused light, dense fog). The dataset is divided into training, validation and test sets in a 7:2:1 ratio. A "challenge subset" containing long-tailed distributions such as extreme weather, strong winds and tremors, and severe overlap is specially constructed for the test set to ensure the rigor and authenticity of the evaluation results.
[0065] During the model training and inference phases, this study conducted a comprehensive comparative experiment with eight of the most representative cutting-edge algorithms in the field of agricultural remote sensing. These included: a segmentation algorithm based on color space thresholds (CST-Seg), a texture classifier combining local binary patterns and support vector machines (LBP-SVM), a multispectral exponential regression model using random forests (RF-MSI), a single-stage object detection network YOLOv5x (YOLOv5x-Det), a Faster R-CNN variant incorporating coordinate attention (FR-CAT), a two-stream detection network based on linear weighted fusion (LF-DualNet), a global perception detector based on SwinTransformer (Swin-Det), and a recently proposed simple feature stitching multimodal network (Cat-Fusion). All comparison models were trained or fine-tuned from scratch on the ZJ-SGF-MultiModal-2025 dataset. Hyperparameters were optimized using grid search and run on a unified computing cluster (8×NVIDIA A100 80GB GPU, Intel Xeon Platinum 8380 CPU) to ensure consistency between the hardware environment and the data preprocessing workflow. The evaluation metrics not only include the conventional mean precision (mAP@0.5, mAP@0.5:0.95), recall, and precision, but also introduce dimensions such as small object detection accuracy (AP_small, target area <32×32 pixels), occlusion robustness index (ORI), and single-frame inference time (Latency) to quantitatively evaluate the comprehensive performance of each algorithm in complex real-world scenarios.As shown in Table 1, the comparison results on the overall test set show that CST-Seg and LBP-SVM are limited by the expressive power of handcrafted features, with mAP@0.5 and 0.95 respectively only reaching 38.42% and 45.17%, and the false negative rate is as high as 42.5% under dense fog conditions; RF-MSI, although utilizing multispectral information, lacks spatial context modeling capabilities, and its mAP stagnates at 54.28%; the deep learning single-modal methods YOLOv5x-Det and FR-CAT have improved in localization accuracy, with mAP reaching 67.35% and 70.82% respectively, but their maturity levels are still low. The method performs poorly in classification discrimination, with the confusion matrix showing that it misclassifies a large number of immature fruits as mature fruits. The multimodal fusion methods LF-DualNet and Cat-Fusion, by introducing a near-infrared channel, improve mAP to 73.46% and 75.91%, respectively, but suffer from high repetition rates (RDR) of 18.3% and 15.7% when dealing with severely occluded targets. Swin-Det achieves an mAP of 79.45% thanks to its powerful global modeling capabilities, but its large parameter count (88.5M) results in an inference latency as high as 65ms, which cannot meet the real-time operation requirements of UAVs at the edge. In contrast, the method in this application, with its dual-channel YOLOv8-Lite architecture, cross-modal gated attention mechanism, and spectral consistency constraint loss, achieves 89.73% mAP at 0.5:0.95, a 10.28 percentage point improvement over the second-best Swin-Det, while keeping the parameter count to 18.2M and the inference latency as low as 22ms, achieving the best balance between accuracy and speed. In particular, the proposed method achieved a small target detection (AP_small) metric of 83.45%, far exceeding all other comparative algorithms, demonstrating its unique advantage in identifying small fruits.
[0066] To deeply analyze the fine-grained performance differences of various algorithms under different maturity levels and environmental interference, this study plotted multi-dimensional precision-recall (PR) curves for comparison, such as... Figure 7 As shown, Figure 7 (a) Figure 7 (b) Figure 7 (c) and Figure 7(d) Corresponding to immature fruit, mature fruit, overripe fruit, and dense fog / low light environments, the broken lines in the figure represent CST-Seg, LBP-SVM, RF-MSI, YOLOv5x-Det, FR-CAT, LF-DualNet, Cat-Fusion, Swin-Det, and the method of this application, respectively. In all scenarios, the PR curve of the method of this application encloses the area of all other comparison algorithms, especially maintaining extremely high precision in the high recall region (>0.9). Particularly in the dense fog / low light sub-figure, the curves of other algorithms drop sharply, and the AUC value decreases significantly, while the curve of the method of this application remains firm, almost consistent with the performance under clear weather conditions, intuitively reflecting its excellent environmental adaptability and anti-interference ability. The curves of CST-Seg and LBP-SVM collapsed rapidly after the recall exceeded 0.3, and the precision dropped below 0.2. Although YOLOv5x-Det and FR-CAT showed some improvement, their precision was still less than 0.6 in the high recall range. In contrast, the curve of the method in this application remained high at over 0.88, showing extremely strong background suppression ability. In the mature fruit detection sub-image, the performance of each deep learning model was relatively similar, but the precision of the method in this application remained at 0.82 even under the extreme condition of 0.95 recall, which is about 12 percentage points better than Swin-Det. This is due to the sensitive capture of the biochemical changes inside the fruit by the near-infrared channel. In the overripe fruit detection sub-image, LF-DualNet and Cat-Fusion showed obvious performance fluctuations and sawtooth curves due to the texture blurring and irregular shape caused by fruit decay. In contrast, the method in this application, thanks to the spectral consistency constraint loss, had a smooth and stable curve, and the precision was always above 0.85. Most notably, in the dense fog and low-light environment sub-image, the PR curve area of the other seven comparison algorithms shrank significantly, and the AUC value decreased by an average of more than 35%, indicating that they are extremely sensitive to atmospheric scattering and changes in illumination. In contrast, the method in this application, due to the integration of dark channel prior defogging and canopy layering masking technology, has a curve shape that basically overlaps with that under clear weather conditions, and the AUC value only decreases slightly by 2.1%, verifying its high robustness under extreme weather conditions. This series of curve data intuitively reveals the overwhelming advantage of the method in various sub-scenarios.
[0067] Table 1: Comparison of overall performance of different detection methods on the ZJ-SGF-MultiModal-2025 test set
[0068]
[0069] (Note: RDR stands for Redundant Detection Rate, and ORI stands for Occlusion Robustness Index)
[0070] To address the long-standing technical bottleneck of clustered occlusion of Torreya grandis fruits in the industry, this study designed a specific ablation experiment to evaluate the detection accuracy (Acc) and repetition rate (RDR) of different algorithms under light, medium, and heavy occlusion levels. The results are summarized in Table 2. Data shows that as the occlusion level increases, the performance of traditional algorithms such as CST-Seg and LBP-SVM drops drastically, with accuracy below 20% in heavy occlusion scenarios, and they are almost unable to distinguish individual fruits. While mainstream deep learning models such as YOLOv5x-Det and Swin-Det show some resistance, their accuracy also drops to 52.4% and 59.8% respectively under heavy occlusion, and their RDR soars to 26.5% and 21.3%. This indicates that the traditional non-maximum suppression (NMS) strategy has serious defects in handling closely clustered targets, easily leading to the false deletion or duplicate counting of high-value mature fruits. While LF-DualNet and Cat-Fusion incorporate multi-source information, their performance improvement under heavy occlusion is limited due to a lack of topology awareness. In contrast, our proposed method, by introducing a spatial density-weighted spectral clustering correction algorithm, constructs a logarithmic adjacency matrix and utilizes Laplacian eigenmaps to simulate the natural growth topology of fruits. This method maintains a high accuracy of 78.9% even in heavily occluded scenarios, strictly controlling the RDR to within 5.8%, an improvement of nearly 16 percentage points compared to the second-best, Swin-Det. This result fully demonstrates that our proposed method can effectively resolve complex overlapping relationships, accurately separate adhered individuals, and solve the counting and classification challenges in dense scenes.
[0071] Table 2: Comparison of Detection Accuracy and Repeat Detection Rate under Different Degrees of Occlusion
[0072]
[0073] To intuitively reveal the intrinsic mechanism of the proposed method in feature extraction and fusion, this study utilizes Gradient Weighted Class Activation Mapping (Grad-CAM++) technology to generate multi-stage feature visualization comparisons, such as... Figure 8As shown, the figure contains four sub-figures, from left to right: (a) the original RGB-NIR synthesized input image, (b) the intermediate layer feature response map of the LF-DualNet network, (c) the global attention map of the Swin-Det network, and (d) the cross-modal fusion feature map of the method in this application. All sub-figures adopt the Jet color mapping heatmap format, with blue representing low response regions and red representing high response regions. Observing the original images in column (a), it can be seen that the Torreya grandis fruits are hidden in the complex background of branches and leaves, and some targets are shrouded in fog and obscure each other; in the feature map of LF-DualNet in column (b), the leaf texture and branches produce strong response signals, while the signal in the fruit area is weak and scattered, indicating that simple linear fusion failed to effectively suppress background noise; in column (c), although the attention map of Swin-Det covers the fruit area, the response range is too large, the boundaries are blurred, and it is difficult to distinguish adjacent adhering fruits; in column (d), the feature map of the method in this application shows a completely different effect. The fruit area appears as a bright red compact mass with sharp and clear boundaries. The background noise is completely suppressed to dark blue, and the response intensity of fruits of different maturity levels is distinct. The response of immature fruits is weak, and the response of mature fruits is the strongest. This phenomenon is directly attributed to the adaptive weighting of visible light geometric information and near-infrared spectral information by the cross-modal gating attention module, as well as the guiding role of the prior of saliency heatmap, which enables the network to accurately focus on the target area and achieve efficient feature purification.
[0074] To verify the accuracy of the geographic coordinate back projection and the field usability of the output results, the study selected 500 sample trees with known coordinates in the test area for field point verification, and mapped the algorithm output results to a high-precision spatial distribution map, such as... Figure 9 As shown, the left side displays the scatter plot results of fruit distribution using the Cat-Fusion method, while the right side displays the scatter plot results using the method described in this application. These are overlaid on a high-resolution satellite image base map with a resolution of 0.2 meters. The scatter plot colors strictly correspond to maturity categories (green represents immature, purple represents mature, and brown represents overripe). In the left sub-image, a large number of scatter plots deviate from the actual location of the tree canopy, even appearing on roads, open spaces, or adjacent trees, forming a clear "flying point" phenomenon. Statistical analysis shows that the average positioning error reaches 3.65 meters, and the maturity classification is chaotic, with extreme jumps in maturity within the same canopy. In the right sub-image, the scatter plots are closely clustered within the canopy of the Torreya grandis, and the spatial distribution highly matches the actual forest morphology, exhibiting a clear maturity gradient distribution pattern. The average positioning error is reduced to 0.38 meters. This significant difference is attributed to the rigorous POS data time synchronization, joint calibration of camera intrinsic and extrinsic parameters, and projection transformation model based on digital elevation model (DEM) in the method of this application. This ensures seamless and high-precision conversion from pixel coordinates to WGS84 world geographic coordinates, providing a reliable decision-making basis for subsequent variable harvesting and precision agriculture management.
[0075] Based on the quantitative data and qualitative analysis above, the method in this application comprehensively surpasses the eight existing mainstream technologies in all key indicators, such as... Figure 10 As shown, the mAP variation trend of the proposed method is further illustrated under different light intensities (1000 lx to 10000 lx) and different fog concentrations (visibility 50 m to 500 m). The left Y-axis represents mAP@0.5:0.95 (%), the right Y-axis represents visibility (m), and the X-axis represents light intensity. The nine broken lines in the figure represent nine different algorithms. Among them, the curves of CST-Seg and LBP-SVM fluctuate drastically, with mAP amplitude exceeding 30%; single-modal deep learning methods such as YOLOv5x-Det and FR-CAT also show fluctuations of about 18%; although Swin-Det has certain stability, its performance still declines significantly under low light and high fog conditions. In contrast, the curve of the proposed method is almost horizontal. Under full light range and low visibility conditions, the mAP fluctuation range is controlled within 2.5%, and it remains at a high level of over 88.5%. This result profoundly reveals the insensitivity of the dual-channel architecture, dark channel defogging preprocessing, and spectral consistency constraint mechanism to environmental changes, proving that the method has extremely strong generalization ability and can meet the needs of intelligent monitoring of Torreya grandis industry in all weather and all terrain.
[0076] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any changes, modifications, substitutions, integrations, and parameter changes made to these embodiments within the spirit and principles of the present invention, without departing from the principles and spirit of the present invention, through conventional substitutions or to achieve the same function, fall within the scope of protection of the present invention.
Claims
1. A method for intelligent identification and classification of the maturity of Torreya grandis from the perspective of unmanned aerial vehicles, characterized in that, include: Simultaneously acquire RGB visible light images and near-infrared irradiance images of the canopy of Torreya grandis from the perspective of UAV, and construct a multi-source heterogeneous input dataset; Dark channel prior dehazing and canopy layering mask generation are performed on RGB images to shield non-target branches and leaves and complex background interference. After Gaussian difference filtering, the spectral reflectance difference characteristics of Torreya grandis fruit and leaves are extracted from the near-infrared image to generate a fruit saliency heatmap. A dual-channel YOLOv8-Lite lightweight network is constructed, with the visible light channel used for small target localization and the near-infrared channel used for maturity feature discrimination. Feature adaptive alignment and weighted fusion are achieved through a cross-modal attention module. Based on the spectral threshold and color gradient features corresponding to the maturity of Torreya grandis, a three-classification decision head is constructed to output the category information of immature, mature and overripe fruits in real time. Non-maximum suppression and clustering correction are applied to the same cluster of Torreya grandis to eliminate misclassification caused by occlusion and adhesion, and output accurate maturity classification results and location coordinates from the perspective of UAV.
2. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 1, characterized in that, The process of performing dark channel prior dehazing and canopy layering mask generation on the RGB image specifically includes: calculating the dark channel map of the RGB image. The formula is: ,in, Represents pixels In color channels The strength, Therefore A local window centered on the atmospheric scattering model is used to estimate global atmospheric light values. and transmittance Guided filtering is used to refine the transmittance map, and a clear image after dehazing is reconstructed. The formula is: ,in, A constant threshold is used to prevent the denominator from being too small; a vegetation index mask is constructed based on the ratio of the green component to the near-infrared component of the reconstructed image. The tree canopy region is extracted by adaptive threshold segmentation, and the pixel values of the non-tree canopy region are set to zero to generate a tree canopy layering mask.
3. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 2, characterized in that, The near-infrared image is filtered using a Gaussian difference filter to extract the spectral reflectance difference features of the Torreya grandis fruit and leaves. Specifically, this includes: analyzing the near-infrared irradiance image. Application of multi-scale Gaussian difference operator To enhance the details of the fruit's edges and textures, the formula is: ,in, The standard deviation is Gaussian kernel function, This represents the convolution operation. and These are the scale parameters for capturing the fruit's outline and internal texture, respectively; combined with the red band reflectance in the RGB image. Near-infrared band reflectivity Calculate the modified normalized difference index To distinguish mature fruit from background leaves, the formula is: ,in, This represents the maximum value of the multi-scale Gaussian difference response. This is the texture suppression coefficient. As a smoothing factor; based on Spatial distribution of fruit saliency heatmap Its pixel value The formula is: ,in, It is the Sigmoid activation function. and These are the spectral weights and the second-order gradient weights, respectively.
4. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 1, characterized in that, The construction of the dual-channel YOLOv8-Lite lightweight network includes a visible light coding branch. With near-infrared coding branch ; The visible light encoding branch An improved CSPDarknet structure is adopted, and a coordinate attention mechanism is embedded to enhance the spatial position awareness of small targets; The near-infrared coding branch It uses a lightweight MobileNetV3 backbone to specifically extract spectral texture features; Before the feature fusion stage, the two branches output feature maps respectively. and .
5. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 4, characterized in that, The visible light encoding branch With near-infrared coding branch Both employ dynamically depthwise separable convolution operations, with their convolution kernel weights... Dynamically generated based on channel statistics of the input feature map: ,in, The channel selection coefficients are generated through global average pooling and the fully connected layer. For predefined The group of basic convolutional kernels enables the receptive field size and feature extraction granularity to be adaptively adjusted according to the texture complexity of Torreya grandis fruits at different growth stages.
6. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 5, characterized in that, The feature adaptive alignment and weighted fusion achieved through a cross-modal attention module specifically includes: constructing a cross-modal gated attention unit, and... and The layers are concatenated along the channel dimension, and a gated weight matrix is generated through a shared fully connected layer. : Utilizing gating weights Dynamically adjust the contribution of dual-channel features to generate a fused feature map. The formula is: ,in, This represents element-wise multiplication. The relevance coefficient is a learnable coefficient. For the cross-correlation attention operator, the formula is: in, They are respectively from The generated query matrix and by The generated key-value matrix, This is the scaling factor.
7. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 1, characterized in that, The construction of a three-class classification decision head based on the spectral threshold and color gradient features corresponding to the maturity of Torreya grandis specifically includes: introducing a multi-task loss function at the end of the detection head. Simultaneously, the positioning regression and maturity classification are optimized, with the following formula: in, For CIoU loss, FocalLoss is the classification loss. The formula for the spectral consistency constraint loss is: ,in, For the network prediction of the first The spectral feature vector of each target This is the standard spectral prototype vector corresponding to the maturity category. Represents the spectral gradient operator. The gradient penalty coefficient is minimized. This forces the features learned by the network to cluster in the spectral space toward the prototype center of their respective maturity categories.
8. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 7, characterized in that, The three-class decision head employs a temperature scaling calibration strategy during the inference phase to adjust the output class probability distribution. Perform calibration: ,in, The first output of the logic layer Unnormalized score The temperature parameters are obtained through optimization using the validation set; combined with the confidence threshold. and Only when or Only when the ripeness is reached will the corresponding picking signal be triggered; otherwise, it will be marked as unripe.
9. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 1, characterized in that, The non-maximum suppression and clustering correction for the same cluster of Torreya grandis specifically includes: constructing a weighted adjacency matrix based on spatial density and confidence. For the set of detection boxes Any two boxes Edge weights between The calculation is as follows: ,in, The coordinates of the detection box center are, For classification confidence, and Hyperparameters for controlling sensitivity; based on The spectral clustering algorithm is used to group spatially adjacent and feature-similar detection boxes into the same fruit cluster. ; in each cluster Within the cluster, the detection box with the smallest distance between the spectral features and the maturity prototype is selected as the representative box, and the remaining redundant boxes in the cluster are removed to complete the occlusion and adhesion correction.
10. The intelligent identification and classification method for the maturity of Torreya grandis from the perspective of an unmanned aerial vehicle (UAV) as described in claim 9, characterized in that, This also includes geographic coordinate mapping, specifically: using the drone's POS data to obtain the latitude and longitude of the shooting time. ,altitude and attitude angle ; Constructing from image pixel coordinate system To the world geographic coordinate system Projection transformation model: ,in, For the camera intrinsic parameter matrix, For rotation matrix, It is a translation vector. The estimated average height of the canopy; world coordinates Convert to WGS84 geographic coordinates and associate with the output maturity classification results.
Citation Information
Patent Citations
Mooring unmanned aerial vehicle for fruit picking and harvesting method
CN117465707A
Method for detecting maturity of alfalfa seed pods through unmanned aerial vehicle aerial image based on improved YOLO v8
CN118262229A