Fruit tree flower bud recognition and positioning method
Through deep learning models and geometric constraint technology, the distribution density of buds on the branches of fruit trees is accurately identified and segmented, and the problems of low efficiency and poor accuracy of flower sparse operations in the existing technology are solved, and the support of intelligent flower sparse operations is achieved.
Patent Information
- Application Number
- CN202411960764.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to accurately identify and divide the distribution density of flower buds on the branches of fruit trees, resulting in low efficiency and poor accuracy of flower sparse operations, especially in complex natural scenarios.
Through the deep learning model combined with geometric constraints, a multi-objective recognition segmentation model for fruit tree bud period is established, the images of fruit tree branches during the bud period are obtained, data augmentation and preprocessing are performed, and instance segmentation is used to analyze the morphological characteristics and spatial relationships of the organs, and the distance and density between buds are calculated.
It realizes the accurate identification and segmentation of the distribution density of flower buds on fruit tree branches, improves the efficiency and accuracy of flower sparse operations, and supports intelligent flower sparse operations.
Smart Images

Figure CN120032240A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart agriculture. Specifically, the image data of the bud stage of fruit trees in natural scenes is obtained by a camera, and multiple organs such as buds, branches, and leaves in the obtained images are segmented by instance through a deep learning model. On the basis of instance segmentation, the growth position of the buds on the branches is determined by analyzing the morphological characteristics, spatial relationships and other attributes of each organ, thereby providing support for intelligent operations based on the requirements of flower thinning agriculture based on the "distance method". Background Art
[0002] The bud stage of fruit trees is one of the important links in their growth cycle and is a critical period for the transition from vegetative growth to reproductive growth. During this period, the internal physiological and biochemical changes of fruit trees play a decisive role in the quantity and quality of fruits. Since the physiological functions of leaves are not fully mature at this time, the efficiency of photosynthesis is limited; therefore, various growth organs rely heavily on nutrients stored in the roots in the previous quarter. Since the size, shape and final quality of the fruit depend to a large extent on the adequacy and reasonable distribution of nutrients stored in the tree during the cell division stage of the bud stage, if the flower buds are not thinned, excessive competition among a large number of flower buds will consume limited nutrients, resulting in the failure of individual fruits to develop in high quality. Therefore, it is particularly important to choose the appropriate time and method of flower thinning.
[0003] Appropriate flower thinning measures can make the limited nutrients of the tree more concentratedly supplied to the flower buds with higher growth potential, thereby increasing the number and volume of cells in a single fruit, and further improving the fruit shape index and overall quality. Compared with flower thinning during the bud stage or after flowering, flower thinning during the bud separation stage can avoid potential climate (late spring cold) risks and reduce energy loss, so it is a more economical and efficient flower thinning strategy.
[0004] At present, although there are mechanical and chemical flower thinning methods mainly for the peak flowering period after flowering, manual flower thinning is still the only commonly used operation method during the bud separation period. However, this method is easily affected by factors such as fatigue, low efficiency, and high cost, and misses the optimal flower thinning window. With the increasing aging of society and the shortage of labor, fruit farmers are facing many challenges in management and maintenance at this stage. Therefore, technological innovation is urgently needed to optimize the flower thinning management of fruit trees in order to meet these challenges and promote the development of the industry in an efficient, environmentally friendly, and sustainable direction.
[0005] With the development of agricultural mechanization, the construction of dwarfed and densely planted standardized orchards continues to advance, the integration of agricultural machinery and agronomy continues to deepen, and intelligent operation and management of orchards is becoming increasingly important.
[0006] However, in practice, the branching structure of fruit tree branches is complex, their growth is random, and they are affected by internal and external factors, which increases the difficulty of intelligent and precise flower thinning. Therefore, a fruit tree flower bud recognition and positioning method and system based on the actual physiology, morphology, structure and other aspects of fruit trees, especially the agronomic requirements of flower bud distribution density thinning on the same branch, focuses on how to correctly identify and segment branch segments, flower buds and other targets that are divided into different sizes and numbers due to occlusions such as leaves, determine which branch segments belong to the same branch, estimate the morphology of the obscured branches, determine which flower buds grow on the same branch, and locate the growth points of the flower buds on the branches, finally calculate the distance between the flower buds, evaluate the flower bud distribution density, and determine whether flower thinning is needed based on the measured data, especially a fruit tree flower bud recognition and positioning method based on deep learning and geometric constraints. Summary of the invention
[0007] The technical problem to be solved by the present invention is to evaluate the distance and density between the flower buds on the branches of fruit trees, and to determine whether flower thinning is needed based on the measured data. The present invention is particularly directed to a method for identifying and locating flower buds of fruit trees based on deep learning and geometric constraints.
[0008] To achieve the above objectives, the present invention provides the following technical solutions.
[0009] Model building: Establish a multi-target recognition and segmentation model for fruit tree bud stage: including data set acquisition, data enhancement, and model building.
[0010] Data acquisition: In order to construct a diverse and informative image dataset required for deep learning models, images of fruit tree branches in the bud stage are acquired through imaging devices such as mobile phone cameras, digital cameras and / or depth cameras. These devices are used to shoot under varying natural lighting conditions (sunny, cloudy) and at different time periods (morning, noon, afternoon), and at different distances of 0.3-3.0 m from the branches to obtain and select pictures.
[0011] Due to the differences in imaging characteristics among different imaging devices, such as the width of viewing angle, the length of focal length, and the difference in photosensitivity, the imaging device preferably includes more than two cameras. Through the complementary advantages between the devices, the appearance characteristics and texture details of the target object can be captured in all directions, enriching the diversity and complexity of the data set. This is crucial for the model to deeply understand and accurately identify the various complex characteristics of fruit trees during the bud and fruit thinning periods.
[0012] Data enhancement: Use digital image processing technology to perform targeted preprocessing and enhancement operations on the selected images, including but not limited to color space conversion, geometric transformation (such as flipping, translation, rotation), brightness contrast adjustment, and image fusion. In particular, when performing operations such as rotation, cropping, and moving, strategically remove some leaf areas with a large number of leaves, and keep the visual integrity of other categories such as branches, buds, and flower branches as much as possible while ensuring the overall structure of the image, thereby effectively increasing the pixel occupancy rate of these relatively scarce categories, which is conducive to the model's ability to strengthen the balanced recognition of each category, and ultimately improve the overall generalization performance and robustness of the model, and finally normalize the obtained image into a standardized image. For example, images of 640×480, 640×640, 720×640, 1024×768 or 1920×1024 pixels.
[0013] Furthermore, in order to enrich the diversity of the dataset and ensure that more minority category targets are included, it is preferred to use an image target synthesis data augmentation algorithm to generate new images.
[0014] In the initial image dataset, there is an imbalance in the distribution of the number of objects in each category, such as branches, buds, leaves and flower branches. This difference in pixel density between categories may cause the deep learning model to prefer predicting the quantitatively dominant categories (branches and leaves), while weakening its recognition ability for the scarce categories (flower branches).
[0015] In addition, the model also faces the potential risk of overfitting. In order to enhance the model's adaptability to the diversity of orchard operating environments and its resistance to overfitting, the inherent complexity of static images is taken into account, such as weather conditions (lighting and other factors), irregular shapes of branches, and noise changes caused by aging of hardware equipment. Attention is also paid to problems such as vibration, image capture color gamut variation, shadow occlusion, and object overlap that may occur in mobile devices during actual operation.
[0016] Selectively select several image samples with subtle differences in the number of target objects and a high proportion of small targets (such as flower branches) from the original data set, and divide them into training and validation sets in a ratio of 8:2-8:3. The specific steps are as follows. ① Create a corresponding mask using the polygon area point coordinate information in the json file generated by the annotation of image A. ② Use the mask generated in step ① to extract the corresponding RGB pixels in the image to the black background area. ③ Apply the translation matrix to adapt the mask generated in step ① to the area specified by image B. ④ Extract the RGB pixels based on the mask generated in step ③ to the corresponding position in image B to synthesize a new image.
[0017] Through comprehensive data enhancement and reorganization strategies, we aim to systematically make up for the incomplete coverage of real scenes that may exist in the original data collection stage, and specifically solve the serious imbalance of category pixel distribution. Based on the method described, a dataset of apple tree bud stage containing 3860 images was constructed. In addition, 50 images were randomly selected from unlabeled images as an independent test set.
[0018] Model construction: Establish the YOLO-Bud model, including the input layer, backbone network (Backbone), neck network (Neck), and head network (Head).
[0019] This network structure extracts features through the backbone network, the neck network fuses multi-scale information, and the head network performs the final task output. Specially designed modules such as SPPFX, DSample Lite and C2f-DCN-t enhance the performance and robustness of the model, ensuring that the model can effectively handle multi-scale targets and improve the accuracy of detection and segmentation.
[0020] The input layer: YOLO-Bud model, whose execution process covers the whole process from input image to output target detection result. It receives pictures of any pixel size, and then crops them into standardized pixel images after image preprocessing in the early stage of the model.
[0021] The backbone network (Backbone) described herein inputs the enhanced image, and the backbone network (Backbone) uses 3x3 or 1x1 convolution modules and pooling operations to progressively extract multi-scale features containing rich semantic information and spatial structures layer by layer.
[0022] Specifically, it includes the following execution process: the first convolution layer (p1): repeated once, the feature map scaling factor is 1 / 2; the second convolution layer (p2): repeated once, the feature map scaling factor is 1 / 2; the third deformable convolution layer (C2f-DCN-t): repeated 3 times, the feature map scaling factor is 1; the fourth lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the fifth deformable convolution layer (C2f-DCN-t): repeated 6 times, the feature map scaling factor is 1; the sixth lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the seventh deformable convolution layer (C2f-DCN-t): repeated 3 times, the feature map scaling factor is 1; the eighth lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the last spatial pyramid pooling module (SPPFX): repeated once, the feature map scaling factor is 1.
[0023] The neck network (Neck): In order to further enhance the adaptability and robustness of the model to multi-scale targets, the feature extraction is introduced into the Spatial Pyramid Pooling-Fast-X (SPPFX) pyramid pooling layer, which unifies the multi-scale features through pooling operations of different scales and retains important spatial information.
[0024] Subsequently, the feature information is passed to the Neck network part. Through upsampling and cross-layer feature fusion methods, the spatial position information of shallow features is closely combined with the abstract semantic knowledge of deep features to construct a hierarchical feature representation, thereby improving the accuracy of target positioning and segmentation. The specific execution process includes the following: Upsampling operation (Upsample): The eleventh layer (p4) is repeated once, and the feature map scaling factor is 1 / 2; the thirteenth layer (p3) is repeated once, and the feature map scaling factor is 1 / 2. Feature map concatenation (Concat): The twelfth layer (p4) is repeated once, and the feature map scaling factor is 1; the fourteenth layer (p4) is repeated once, and the feature map scaling factor is 1. Deformable convolution layer (C2f-DCN): The fifteenth layer (p3) is repeated 3 times, and the feature map scaling factor is 1 / 2; the eighteenth layer (p4) is repeated 3 times, and the feature map scaling factor is 1 / 2. Lightweight downsampling module (DSample Lite): The sixteenth layer (p4) is repeated once, and the feature map scaling factor is 1 / 2; the nineteenth layer (p5) is repeated once, and the feature map scaling factor is 1 / 2.
[0025] The head network (Head): The multi-level feature map output by Neck is sent to the detection head for training and reasoning. The detection head predicts the category, position and segmentation mask of each target based on the fused features. It includes three decoupled heads (Decoupled Head), each head is responsible for a specific task output, and each decoupled head includes a deformable convolution layer (C2f-DCN) and feature map splicing (Concat).
[0026] In general, this network structure extracts features through the backbone network, the neck network fuses multi-scale information, and the head network performs the final task output. Specially designed modules such as SPPFX, DSample Lite, and C2f-DCN-t enhance the performance and robustness of the model, ensuring that the model can effectively handle multi-scale targets and improve the accuracy of detection and segmentation.
[0027] Furthermore, the YOLO-Bud model is optimized: wherein, the spatial pyramid pooling feature (SPPF) module in the Backbone architecture of the model is specifically improved: the present invention increases the number of pooling layers to 512 layers (, SPPFX).
[0028] By increasing the number of channels in the middle layer of the model, the model can obtain richer and more detailed multi-scale feature expressions in the deep feature extraction stage. In image segmentation tasks, especially when dealing with problems such as multi-target segmentation in the bud stage of apple trees, which have morphological differences and may be affected by complex backgrounds, multi-level feature fusion can effectively improve the segmentation accuracy of the target.
[0029] Secondly, the C2f module in YOLOv8 was innovatively transformed (C2f-DCN-t), introducing an efficient dynamic sparse operator DCNv4 (Deformable Convolution v4) based on the optimized and upgraded iterative version DCNv3 (Deformable Convolution v3).
[0030] The DCNv series of operators is an advanced convolution operator used in the field of computer vision to deal with the geometric deformation problem of feature maps in convolutional neural networks. It is an improvement on conventional convolution. Its core idea is to allow the convolution kernel to no longer be fixed on a regular grid when performing convolution operations, but to generate dynamic offsets as the input feature map changes. This dynamic offset is calculated by an additional offset field, which dynamically adjusts the position of the convolution kernel according to the local information of the input feature map, so that the model can better capture the complex geometric transformations and object posture changes in the image. DCNv4 has redesigned the core features of dynamic convolution and optimized memory access, effectively reducing redundant operations and thus achieving an increase in computing speed. This improvement not only enables DCNv4 to reduce the consumption of computing resources while maintaining a high degree of flexibility, but also makes it more efficient when processing complex visual tasks. By combining with DCNv4, the C2f module can more effectively utilize the efficient dynamic sparse computing characteristics brought by DCNv4, thereby achieving faster feature calculation and transmission during the forward propagation of the network. This combination not only improves the computing speed of the network, but also enhances the ability to extract irregular image features and the robustness of the network without affecting the expressiveness of the model, enabling it to maintain excellent performance when dealing with various complex visual scenes.
[0031] A lightweight downsampling module DSampleLite (Down Sample Lite) is used as the downsampling method, that is, part of the feature map is spatially downsampled through 3×3 convolution, and the other part uses maximum pooling combined with 1×1 convolution for feature extraction.
[0032] The DSampleLite module design combines average pooling, maximum pooling and channel attention mechanism, which improves the network's feature expression ability and multi-scale feature fusion ability while maintaining a low parameter count.
[0033] The DSampleLite module first performs an average pooling operation on the input feature map to keep the spatial dimension down, then divides the feature map evenly in the channel dimension and performs targeted processing on each part separately.
[0034] The feature map is spatially downsampled via 3×3 convolution, and feature extraction is performed using maximum pooling combined with 1×1 convolution.
[0035] In order to enhance the network's sensitivity to key features and suppress redundant information, the present invention also introduces a lightweight attention mechanism, which applies adaptive weighting to the downsampled feature maps through global average pooling and channel attention operations, thereby achieving feature selection and enhancement.
[0036]
[0037] Shape-IoU is introduced, which emphasizes the accurate modeling of bounding box shape and scale factors.
[0038] Shape-IoU takes into account the impact of the asymmetry and scale diversity of the bounding box's shape on the regression results. By introducing shape-related loss terms, Shape-IoU can more accurately measure the degree of shape matching between the predicted box and the true box, which is particularly important for detection tasks with obvious shape feature changes or small targets. Especially in the multi-target segmentation task of apple tree buds, since the target bounding box has various shapes and small sizes, Shape-IoU can help the model better understand and capture the morphological details of the target object, thereby improving the segmentation accuracy. The formula of Shape-IoU is as follows.
[0039]
[0040]
[0041] Where scale is the scale factor, which is related to the scale of the target in the data set; B and B gt Represent the predicted box and GT box respectively; ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT bounding box. The corresponding bounding box regression loss function is:
[0042] L Shape-IoU =1-IoU+distance shape +0.5×Ω shape (7).
[0043] In addition, Focaler-IoU focuses on the differentiated attention to regression samples of different difficulty levels. The Focaler-IoU formula is as follows.
[0044]
[0045] Where IoU focaler is the reconstructed Focaler-IoU, IoU is Shape-IoU, [d,u]∈[0,1]. By adjusting the values of d(0) and u(0.95), the loss function can be focused on different regression samples. The loss is defined as follows:
[0046] L Focaler-IoU =1-IoU focaler (9).
[0047] In practical applications, this combination can effectively improve the detection performance of the model in complex scenarios.
[0048] Model evaluation: When evaluating the performance of the YOLO-Bud model in the image segmentation task, multiple key quantitative indicators were used for comprehensive evaluation, including precision (P), recall (R), intersection over union (IoU), Dice similarity coefficient (Dice) and mean average precision (mAP). The calculation formulas for each indicator are as follows.
[0049]
[0050] In formula (10)-(14), TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative, respectively. The value of k refers to the number of classes of segmented objects in the model. The number of segmentation classes in this dataset is 4, that is, k = 4. AP is an important indicator for evaluating the performance of a model in multi-class classification tasks.
[0051] The calculation formula of AP is as follows:
[0052] AP = ∫P(R)dR (15).
[0053] Model training: All experiments are based on the PyTorch2.1 framework. The experimental environment is configured with an NVIDIA GeForce RTX 4090 Laptop GPU, 16GB of memory, and an operating system of Ubuntu 20.04. For the training of the YOLO-Bud model, the AdamW optimizer is used, the initial learning rate is set to 0.00125, the momentum parameter remains at 0.9, and the weight decay coefficient is 0.0005. On this basis, the learning rate adjustment strategy adopts the cosine annealing method. The input image is uniformly normalized to a size of 640×640 pixels, and the model is trained for up to 500 cycles. During the training process, if no significant performance improvement is observed for 50 consecutive iterations, the training process is automatically terminated. The accuracy and reliability of model training are ensured by configuring the experimental environment, strictly setting hyperparameters, and adopting effective training strategies. These measures not only help to accelerate the convergence of the model, but also effectively alleviate the overfitting phenomenon and minimize the model from falling into the local optimal solution.
[0054] Training results: The trained YOLO-Bud model will eventually get a weight parameter file for a specific data set. This file is usually a PyTorch model file with the suffix .pt or .pth. It contains all the parameters learned by the model during the training process (such as the weights and bias items of the convolutional layer) and the architecture information of the model. These weight parameters are the key to the model's ability to accurately identify and locate targets in images. During the training process, the model continuously iterates and optimizes these parameters to achieve the best performance (such as accuracy, recall, etc.) on a given data set. After training is complete, this weight parameter file is loaded into the YOLO-Bud model, and then inference (or prediction) is performed to detect targets on new image data. During inference, the model uses these learned parameters to extract image features and output the location and category information of the target.
[0055] Segmentation effect: The results show that the YOLO-Bud model has significant advantages in multiple key performance indicators. This advantage comes from the YOLO-Bud model's optimized design in feature pyramid construction, context information fusion, and balanced processing of local and global information, making it more robust and have better segmentation performance than classic methods in tasks that require fine segmentation, such as the bud stage of apple trees.
[0056] Branch segment affiliation judgment algorithm: When analyzing images of fruit trees at the bud stage to achieve intelligent flower thinning, it is a complex task to accurately locate the branches to which the buds belong in the two-dimensional image. Therefore, the first problem that needs to be solved is how to correctly judge the affiliation of the branch segments segmented by the model and effectively reassemble these independent branch segments into a single branch structure that reflects its biological integrity.
[0057] The principal component analysis (PCA) method was used for the branch segments after instance segmentation, and the first principal component (PCA1) was selected to characterize the local growth direction of the branch segments. In order to determine whether two adjacent branch segments belong to the same physiological branch, two core constraints were proposed: first, the angle between the first principal component vectors of the segmented regions of adjacent branch segments must meet a threshold of less than 20°; second, the angle between the first principal component and the centroid of the branch segment polygon must meet a threshold of less than 15°, achieving a classification accuracy of 96%.
[0058] 2D reconstruction of occluded branches: During the bud growth period, partial occlusion of branches by leaves, buds and other plant organs is common, which poses a challenge to the construction of a complete and accurate branch skeleton, and has a direct impact on the accuracy and reliability of subsequent bud positioning. To address this problem, a branch segment attribution algorithm based on the consistency of the directions of adjacent branch segments, principal component analysis (PCA), B-spline curve fitting and iterative connection strategy is proposed on the basis of completing the branch segment attribution algorithm, aiming to achieve effective splicing of branch segments and finally construct a continuous and accurate branch model.
[0059] According to the consistency of the spatial direction of each branch segment that has been separated and belongs to the same branch, the first principal component analysis (PCA1) direction representing the branch direction is selected, and the straight line equation L1(x,y) is constructed based on this direction and the center of mass of the polygon, as shown in formula (4-12). This straight line divides the coordinate points in the polygonal area of the branch segment (based on the image coordinate system) into two sub-areas L1(x1,y1)>0 and L1(x2,y2)<0 according to the sign of the straight line L1, as shown in the figure. Figure 4 1 and 2. In order to achieve seamless splicing of adjacent branch segments, especially in the restoration of the occluded part, the present invention focuses on selecting key control points in the two sub-regions divided by the straight line L. These control points can represent the boundary features of their respective regions, which is crucial for simulating the edge morphology of the occluded branch segments. Based on these control points, B-spline curves are introduced as a smooth interpolation tool to estimate the continuous curve morphology of the edge of the occluded area. Due to its locality, flexibility and high-order continuity, the B-spline curve can naturally fill the occluded area while keeping the position of the boundary points unchanged, achieve a smooth transition of the edges between the branch segments, and effectively simulate the natural shape of the occluded area.
[0060] L1(x, y)|ax+by+c=0 (16).
[0061] The branch segment regions obtained by model segmentation show the characteristics of being closed and connected in sequence. However, the endpoint curve segments (pseudo-edges) formed at both ends of these closed regions cannot accurately reflect the direction of the growth trend of the edge characteristics of the branches, which interferes with the accuracy of the subsequent branch segment connection to a certain extent. In particular, when the regional points are divided according to the sign of the straight line L1, some endpoint curve points are taken into consideration, and these points have significant differences in spatial distribution from the laws of branch segment edge points, which may lead to large variations in the spatial distribution of edge point coordinates, which is not conducive to achieving a smooth transition of B-spline curve interpolation. In order to ensure the accuracy and smoothness of the branch segment connection, it is necessary to "eliminate" the points contained in the curve endpoints, eliminate the influence of abnormal points on the spatial distribution of edge point coordinates, and ensure that the transition between edge points remains highly smooth during the subsequent B-spline curve interpolation process.
[0062] The present invention constructs a straight line L2(x,y) along the second principal component analysis (PCA2) direction of the polygonal area, taking this direction and the center of mass of the polygon as the reference, and the line segment formed by the two points in the polygonal area represents the pixel length (the center of mass span of the branch segment) of the diameter of the branch segment mapped in the two-dimensional image. In view of the fact that the number of pixels contained in the curve segments at both ends of the branch segment and the center of mass span is similar, the number of control points in the two sub-areas divided by the straight line L1 and the number of pixel points contained in the center of mass span segment are counted. On this basis, by appropriately eliminating points at both ends of the control point set, the endpoint curve pixel points that may be greatly affected by the variation are removed, thereby ensuring that the remaining control points can more accurately represent the edge features of the branch segment.
[0063] Although the two sub-regions divided by L1 may not contain all the pixel points of the endpoint curve, the spatial variability of the edge features at both ends of the branch segment, especially at locations far from the centroid, may also be large. Therefore, in order to reflect the edge features of the branch as realistically as possible and make full use of the advantages of B-spline curve fitting, the pixel coordinate points close to the centroid are retained first, and the edge points with drastic variations are appropriately eliminated.
[0064] In the specific operation, suppose a sub-region sequentially sorted coordinate point set is p, which contains n points, and the centroid span line segment contains m coordinate point sets. At both ends of the coordinate point set p, points with 1.2 times the number of pixels contained in the centroid span line segment are removed to form a new point set p′. This strategy can effectively reduce unnecessary computing resource consumption while ensuring accurate capture of branch segment edge features.
[0065] p={(x 1 ,y 1 ), (x 2,y 2 ),...,(x n ,y n )} (17).
[0066] p′={(x [1.2*m]+1 ,y [1.2*m]+1 ),…,(x n-[1.2*m] ,y n-[1.2*m] )} (18).
[0067] Wherein, [1.2*m] is the maximum integer not exceeding 1.2*m.
[0068] B-spline interpolation is a powerful mathematical tool used to construct a smooth, continuous curve given a set of discrete data points so that the curve passes through all the data points exactly. The key features of the B-spline interpolation algorithm include local support, smooth transitions, and flexibility of arbitrary order. The specific interpolation formula can be expressed as
[0069]
[0070] Where C(t) is the curve value at the interpolation point, P i is the ith control point, is the value of the nth-order B-spline basis function associated with the control point at parameter t. Through this formula, given a set of edge coordinates as control points, the position of the curve is calculated at any parameter value t, and then the complete interpolation curve is obtained.
[0071] When some edges of branches cannot be directly observed due to occlusion, the B-spline interpolation algorithm is used to interpolate the occluded parts with the help of known visible edge coordinate data. It can not only fill in the missing coordinate information, but also ensure that the generated edge feature coordinates maintain a high degree of consistency and smooth transition with the known data points, thereby reproducing the true outline of the occluded area. By applying the B-spline interpolation algorithm to the edge coordinate set, the occluded area can be reasonably inferred and reconstructed while maintaining the continuity of the overall shape. The local support of the algorithm makes the interpolation process rely only on valid data points near the occluded area, avoiding the interference of irrelevant information and ensuring the robustness of the interpolation results.
[0072] The present invention selects a cubic (n=3) B-spline basis function for interpolation calculation, which has C 2 Continuity means that the curve itself and its first and second order derivatives are continuous. This high degree of smoothness ensures that the interpolation result not only has no abrupt changes at the data points, but also transitions naturally between the data points. This is especially important for describing the edges of natural objects such as apple tree branches, where the edges are continuous and gradually changing.
[0073] Determine the subordination of flower buds, branches and branches: When estimating the position distribution of flower buds on the branches to which they belong, the first task is to identify the branch to which each flower bud belongs. Since the morphological information of flower buds and branches in two-dimensional images is complex and easily disturbed by factors such as occlusion and overlap, it is quite challenging to rely solely on these intuitive features to determine the subordination relationship between flower buds and branches. Most flower buds and branches are relatively close to their branches in spatial distribution. The present invention first searches for the nearest neighbor branch area with the target (flower buds and branches) centroid as the base point. To calculate the shortest distance from a point to a polygon, it is necessary to calculate the distance from the point to each edge of the polygon. Take the calculation of the distance from point p to polygon AB as an example. Among them, the vector With vector The angle between them is β, and the following relationship can be obtained from the vector inner product.
[0074]
[0075] Let |AF|=k|AB|
[0076] but This can be obtained.
[0077]
[0078] According to the above formula, we can get the k value, and then get the coordinates of F, and finally get the length of PF. It should be noted that point F does not necessarily fall on line segment AB. There are several cases as follows: (1) k>1, F is on the extension line of AB, and the shortest distance is PB; (2) 0≤k≤1, F is on AB, and the shortest distance is PD; (3) k<0, F is on the reverse extension line of AB, and the shortest distance is PA.
[0079] In the spatial relationship analysis between polygons and points, given the high time complexity of directly calculating the distance from a point to each side of a polygon when there are many sides, it is an efficient processing strategy to prioritize the relative positional relationship between a point and a polygon. By judging the initial position, the subsequent distance calculation process can be effectively simplified. If the analysis determines that the point is inside the polygon, its shortest distance to the polygon may be the distance from the point to the nearest vertex, thus avoiding the calculation of the distance of each edge one by one. On the contrary, if the point is outside the polygon, in order to find the shortest distance, the distance from the point to each side of the polygon must be calculated and compared. In special cases, when a point is located on the edge or vertex of a polygon, the shortest distance from the point to the polygon is zero, and no further distance calculation is required. The strategy of determining the positional relationship between a point and a polygon as a preprocessing step is essentially an effective pruning method that reduces the computational burden when processing a large number of points associated with complex polygons.
[0080] The commonly used method to determine the positional relationship between a point and a polygon is the "ray intersection method". Starting from the point to be detected, draw a ray in any direction (usually horizontal or vertical directions are chosen to simplify calculations). Record the number of intersections between this ray and the edge of the polygon. If the ray enters from one side of the polygon and leaves from the other side, the count increases by one each time it crosses the edge. If the number of intersections is an odd number, the point is inside the polygon; if it is an even number, the point is outside the polygon. In addition, by comparing the slopes or coordinate positions of adjacent points, it is possible to determine whether a point is on the edge.
[0081] Using the above algorithm, after statistical analysis of 50 image samples, the accuracy of flower bud identification was 94.9%. Since the spatial distribution of flower branches is closer to branches, its accuracy of identification is higher than that of flower buds, reaching 96.3%.
[0082] Bud growth point estimation: To determine the growth point of the bud, the present invention proposes a method based on the centroid relationship. The centroids of the buds and branches belonging to the same branch are calculated, and the intersection of these centroid lines and the branches is identified as the growth point of the bud. The proposal of this method depends on clarifying the subordinate relationship between each bud and the corresponding branch. To this end, the nearest neighbor algorithm is used as an initial step to effectively predict the mutual subordinate relationship between most buds and branches. Further statistical analysis reveals a significant feature between the growth direction of the bud and the axial direction of the branch growth, that is, the two form a mostly acute angle relationship (accounting for 98.0%), and when observed along the growth direction of the branch, the centroid of the bud is located in front of the centroid of the branch to which it belongs. It provides additional judgment criteria for the algorithm and strengthens the rationality of the judgment of the subordinate relationship between the bud and the branch. The branch area is small and is often blocked by surrounding leaves, which may affect the accuracy of the centroid coordinates. To address this problem, the intersection of the minimum circumscribed rectangle and the leaf area is used to reconstruct the partially blocked branch area, thereby improving the accuracy of position recognition. The specific algorithm flow is as follows.
[0083]
[0084]
[0085] Through this algorithm, the subordinate relationships between most flower buds and branches are accurately identified.
[0086] This patent is aimed at segmenting flower buds through deep learning models in complex natural scenes, and completing the positioning of flower buds on the branches through geometric constraints. In particular, through the optimized deep learning model, multiple targets in the flower bud period can be identified and segmented, including flower buds, branches, leaves, flower branches, etc. On the basis of multi-target recognition and segmentation, on the basis of analyzing the distribution and individual morphology of flower buds and branches, the patent explores the influence of different data collection field of view points on the spatial positioning between flower buds and the positioning of flower bud thinning areas for scene configurations such as orchard planting patterns and fruit tree branch structures in natural scenes. It can accurately identify flower buds in complex scenes and obtain the growth position of the branches where they are located, and measure the distance between flower buds through the coordinates of the location points where the flower buds grow, thereby constructing an identification and positioning system for the intersection of flower stems and branches in the flower bud period to achieve the purpose of intelligent flower thinning. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 Schematic diagram of the target synthetic data augmentation algorithm; Figure 2 Schematic diagram of the YOLO-Bud model structure; Figure 3 Schematic diagram of bounding box regression feature analysis; Figure 4 This is a schematic diagram of the image segmentation effect in a scene with good lighting; Figure 5 This is a schematic diagram of the image segmentation effect in a dark light scene; Figure 6 The figure shows the curve estimation effect after the third-order B-spline curve interpolation. Figure 7 The flowchart of the two-dimensional reconstruction algorithm for partial occlusion of branches; Figure 8 Schematic diagram of the distance from a point to a polygon; Fig. 9 Schematic diagram of the shortest distance under different k values; Fig.10 It is a schematic diagram of the positional relationship between points and polygons; Fig.11 It is a schematic diagram for judging the subordinate relationship between flower buds and flower branches and estimating the growth point of flower buds; Fig.12 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0088] Preferred embodiments of the present invention are described in detail below.
[0089] Includes model building: Building a multi-target recognition and segmentation model for fruit tree bud stage: including data set acquisition, data enhancement, and model building.
[0090] Data acquisition: In order to construct a diverse and informative image dataset required for deep learning models, images of fruit tree branches in the bud stage are acquired through imaging devices such as mobile phone cameras, digital cameras and / or depth cameras. These devices are used to shoot under varying natural lighting conditions (sunny, cloudy) and at different time periods (morning, noon, afternoon), and at different distances of 0.3-3.0 m from the branches to obtain and select pictures.
[0091] Due to the differences in imaging characteristics among different imaging devices, such as the width of viewing angle, the length of focal length, and the difference in photosensitivity, the imaging device preferably includes more than two cameras. Through the complementary advantages between the devices, the appearance characteristics and texture details of the target object can be captured in all directions, enriching the diversity and complexity of the data set. This is crucial for the model to deeply understand and accurately identify the various complex characteristics of fruit trees during the bud and fruit thinning periods.
[0092] Data enhancement: Use digital image processing technology to perform targeted preprocessing and enhancement operations on the selected images, including but not limited to color space conversion, geometric transformation (such as flipping, translation, rotation), brightness contrast adjustment, and image fusion. In particular, when performing operations such as rotation, cropping, and moving, strategically remove some leaf areas with a large number of leaves, and keep the visual integrity of other categories such as branches, buds, and flower branches as much as possible while ensuring the overall structure of the image, thereby effectively increasing the pixel occupancy rate of these relatively scarce categories, which is conducive to the model's ability to strengthen the balanced recognition of each category, and ultimately improve the overall generalization performance and robustness of the model, and finally normalize the obtained image into a standardized image. For example, images of 640×480, 640×640, 720×640, 1024×768 or 1920×1024 pixels.
[0093] Furthermore, in order to enrich the diversity of the dataset and ensure that more minority category targets are included, it is preferred to use an image target synthesis data augmentation algorithm to generate new images, such as Figure 1 shown.
[0094] In the initial image dataset, there is an imbalance in the distribution of the number of objects in each category, such as branches, buds, leaves and flower branches. This difference in pixel density between categories may cause the deep learning model to prefer predicting the quantitatively dominant categories (branches and leaves), while weakening its recognition ability for the scarce categories (flower branches).
[0095] In addition, the model also faces the potential risk of overfitting. In order to enhance the model's adaptability to the diversity of orchard operating environments and its resistance to overfitting, the inherent complexity of static images is taken into account, such as weather conditions (lighting and other factors), irregular shapes of branches, and noise changes caused by aging of hardware equipment. Attention is also paid to problems such as vibration, image capture color gamut variation, shadow occlusion, and object overlap that may occur in mobile devices during actual operation.
[0096] Selectively select several image samples with subtle differences in the number of target objects and a high proportion of small targets (such as flower branches) from the original data set, and divide them into training and validation sets in a ratio of 8:2-8:3. The specific steps are as follows. ① Use the polygon area point coordinate information in the json file generated by the annotation of image A to create a corresponding mask; ② Use the mask generated in step ① to extract the corresponding RGB pixels in the image to the black background area; ③ Apply the translation matrix to adapt the mask generated in step ① to the area specified by image B; ④ Extract the RGB pixels based on the mask generated in step ③ to the corresponding position in image B to synthesize a new image.
[0097] Through comprehensive data enhancement and reorganization strategies, we aim to systematically make up for the incomplete coverage of real scenes that may exist in the original data collection stage, and specifically solve the serious imbalance of category pixel distribution. Based on the method described, a dataset of apple tree bud stage containing 3860 images was constructed. In addition, 50 images were randomly selected from unlabeled images as an independent test set.
[0098] Model construction: The YOLO-Bud model is built, including the input layer, backbone network (Backbone), neck network (Neck), and head network (Head). This network structure extracts features through the backbone network, the neck network fuses multi-scale information, and the head network performs the final task output. Specially designed modules such as SPPFX, DSample Lite, and C2f-DCN-t enhance the performance and robustness of the model, ensuring that the model can effectively handle multi-scale targets and improve the accuracy of detection and segmentation.
[0099] The input layer: YOLO-Bud model, whose execution process covers the whole process from input image to output target detection result. It receives pictures of any pixel size, and then crops them into standardized pixel images after image preprocessing in the early stage of the model.
[0100] The backbone network (Backbone) described herein inputs the enhanced image, and the backbone network (Backbone) uses 3x3 or 1x1 convolution modules and pooling operations to progressively extract multi-scale features containing rich semantic information and spatial structures layer by layer. The specific execution process includes: the first layer of convolution (p1): repeated once, the feature map scaling factor is 1 / 2; the second layer of convolution (p2): repeated once, the feature map scaling factor is 1 / 2; the third layer of deformable convolution layer (C2f-DCN-t): repeated 3 times, the feature map scaling factor is 1; the fourth layer of lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the fifth layer of deformable convolution layer (C2f-DCN-t): repeated 6 times, the feature map scaling factor is 1; the sixth layer of lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the seventh layer of deformable convolution layer (C2f-DCN-t): repeated 3 times, the feature map scaling factor is 1; the eighth layer of lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the last layer of spatial pyramid pooling module (SPPFX): repeated once, the feature map scaling factor is 1.
[0101] The neck network (Neck): In order to further enhance the adaptability and robustness of the model to multi-scale targets, the feature extraction is introduced into the Spatial Pyramid Pooling-Fast-X (SPPFX) pyramid pooling layer, which unifies the multi-scale features through pooling operations of different scales and retains important spatial information.
[0102] Subsequently, the feature information is passed to the Neck network part. Through upsampling and cross-layer feature fusion methods, the spatial position information of shallow features is closely combined with the abstract semantic knowledge of deep features to construct a hierarchical feature representation, thereby improving the accuracy of target positioning and segmentation.
[0103] The specific execution process includes: Upsample operation (Upsample): The eleventh layer (p4) is repeated once, and the feature map scaling factor is 1 / 2; the thirteenth layer (p3) is repeated once, and the feature map scaling factor is 1 / 2. Feature map concatenation (Concat): The twelfth layer (p4) is repeated once, and the feature map scaling factor is 1. The fourteenth layer (p4) is repeated once, and the feature map scaling factor is 1. Deformable convolution layer (C2f-DCN): The fifteenth layer (p3) is repeated 3 times, and the feature map scaling factor is 1 / 2; the eighteenth layer (p4) is repeated 3 times, and the feature map scaling factor is 1 / 2. Lightweight downsampling module (DSample Lite): The sixteenth layer (p4) is repeated once, and the feature map scaling factor is 1 / 2; the nineteenth layer (p5) is repeated once, and the feature map scaling factor is 1 / 2.
[0104] The head network (Head): The multi-level feature map output by Neck is sent to the detection head for training and reasoning. The detection head predicts the category, position and segmentation mask of each target based on the fused features. It includes three decoupled heads (Decoupled Head), each head is responsible for a specific task output, and each decoupled head includes a deformable convolution layer (C2f-DCN) and feature map splicing (Concat).
[0105] In general, this network structure extracts features through the backbone network, the neck network fuses multi-scale information, and the head network performs the final task output. Specially designed modules such as SPPFX, DSample Lite, and C2f-DCN-t enhance the performance and robustness of the model, ensuring that the model can effectively handle multi-scale targets and improve the accuracy of detection and segmentation.
[0106] Furthermore, the YOLO-Bud model is optimized: Among them, the spatial pyramid pooling feature (SPPF) module in the Backbone architecture of the model is specifically improved: the present invention increases the number of pooling layers to 512 layers (, SPPFX). By increasing the number of channels in the middle layer of the model, the model can be prompted to obtain richer and more detailed multi-scale feature expressions in the deep feature extraction stage. In image segmentation tasks, especially when dealing with problems such as multi-target segmentation in the bud period of apple trees that have morphological differences and may be affected by complex backgrounds, multi-level feature fusion can effectively improve the segmentation accuracy of the target. Secondly, the C2f module in YOLOv8 has been innovatively transformed (C2f-DCN-t), introducing an efficient dynamic sparse operator based on the iterative version DCNv3 (Deformable Convolution v3) optimized and upgraded—DCNv4 (DeformableConvolution v4).
[0107] The DCNv series of operators is an advanced convolution operator used in the field of computer vision to deal with the geometric deformation problem of feature maps in convolutional neural networks. It is an improvement on conventional convolution. Its core idea is to allow the convolution kernel to no longer be fixed on a regular grid when performing convolution operations, but to generate dynamic offsets as the input feature map changes. This dynamic offset is calculated by an additional offset field, which dynamically adjusts the position of the convolution kernel according to the local information of the input feature map, so that the model can better capture the complex geometric transformations and object posture changes in the image. DCNv4 has redesigned the core features of dynamic convolution and optimized memory access, effectively reducing redundant operations and thus achieving an increase in computing speed. This improvement not only enables DCNv4 to reduce the consumption of computing resources while maintaining a high degree of flexibility, but also makes it more efficient when processing complex visual tasks. By combining with DCNv4, the C2f module can more effectively utilize the efficient dynamic sparse computing characteristics brought by DCNv4, thereby achieving faster feature calculation and transmission during the forward propagation of the network. This combination not only improves the computing speed of the network, but also enhances the ability to extract irregular image features and the robustness of the network without affecting the expressiveness of the model, enabling it to maintain excellent performance when dealing with various complex visual scenes.
[0108] A lightweight downsampling module DSampleLite (Down Sample Lite) is used as the downsampling method, that is, part of the feature map is spatially downsampled through 3×3 convolution, and the other part uses maximum pooling combined with 1×1 convolution for feature extraction.
[0109] The DSampleLite module design combines average pooling, maximum pooling and channel attention mechanism, which improves the network's feature expression ability and multi-scale feature fusion ability while maintaining a low parameter count.
[0110] The DSampleLite module first performs an average pooling operation on the input feature map to keep the spatial dimension down, then divides the feature map evenly in the channel dimension and performs targeted processing on each part separately.
[0111] The feature map is spatially downsampled via 3×3 convolution, and feature extraction is performed using maximum pooling combined with 1×1 convolution.
[0112] In order to enhance the network's sensitivity to key features and suppress redundant information, the present invention also introduces a lightweight attention mechanism, which applies adaptive weighting to the downsampled feature maps through global average pooling and channel attention operations, thereby achieving feature selection and enhancement.
[0113] Although the attention mechanism is added, the DSampleLite module does not significantly increase the number of network parameters compared to the full convolution downsampling method because it uses global average pooling and 1×1 convolution with the same number of channels. This strategy ensures that the network can obtain richer feature levels and more powerful representation capabilities under limited computing resources, which is conducive to improving the performance and efficiency of the network in target detection tasks. In addition, in the early stages of the neural network, the Backbone layer is responsible for extracting low-level features and intermediate features, which contain rich and diverse information. Enabling the attention mechanism helps the model distinguish and focus on the most important features in the image, ignore irrelevant noise, and enhance the expressiveness of features. Especially for small target detection or complex background situations, the attention mechanism can improve the recognition accuracy of the model. Turning off the attention mechanism at the Neck layer helps avoid over-complicating the feature fusion process. This is because on the basis of the aggregation of multi-scale information, too many attention operations will bring additional computational overhead and may not significantly improve the detection performance.
[0114] Shape-IoU is introduced, which emphasizes the accurate modeling of bounding box shape and scale factors.
[0115] Compared with traditional IoU and other improved versions, Shape-IoU takes into account the impact of the asymmetry and scale diversity of the bounding box's shape on the regression results. By introducing shape-related loss terms, Shape-IoU can more accurately measure the degree of shape matching between the predicted box and the true box, which is especially important for detection tasks with obvious shape feature changes or tiny targets. Especially in the multi-target segmentation task of apple tree buds, since the target bounding boxes have diverse shapes and small sizes, Shape-IoU can help the model better understand and capture the morphological details of the target object, thereby improving segmentation accuracy.
[0116] Through Figure 3 Parameter Analysis,This method can focus on the shape and scale of the bounding box itself to,calculate the loss, thereby improving the accuracy.
[0117] like Figure 3 , is the bounding box regression feature analysis, green is the anchor generated during the model training process, and yellow is the ground truth. b and b gt are the center points of the anchor and GT box respectively. Therefore, the formula of Shape-IoU can be obtained from Figure 3 It can be derived from: Equation (1)-Equation (6). Where scale is the scale factor, which is related to the scale of the target in the data set; B and B gtDenote the predicted box and the GT box respectively; ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT bounding box. The corresponding bounding box regression loss function is formula (7).
[0118] In addition, Focaler-IoU focuses on the differentiated attention to regression samples of different difficulty levels. Inspired by the idea of focal loss, Focaler-IoU reconstructs the IoU loss through linear interval mapping, allowing the model to pay different attention to easy and difficult samples during training. This is of great significance for solving the problem of category imbalance in target detection tasks, especially when dealing with small, overlapping or occluded targets. Focaler-IoU can ensure that the model focuses more on difficult samples that have a greater impact on the overall detection performance, thereby avoiding the limitation of model performance caused by simple samples dominating the training. The Focaler-IoU formula is as shown in formula (8). In formula (8), IoU focaler is the reconstructed Focaler-IoU, IoU is Shape-IoU, [d,u]∈[0,1]. By adjusting the values of d(0) and u(0.95), the loss function can be focused on different regression samples. The loss is defined as in formula (9).
[0119] Combining the advantages of Shape-IoU and Focaler-IoU, the improved loss function can not only more accurately quantify the similarity between the predicted bounding box and the real bounding box in shape and scale, but also automatically allocate training resources and prioritize the difficult bounding box regression problem that has a greater impact on detection performance. In practical applications, this combination can effectively improve the detection performance of the model in complex scenarios.
[0120] Model evaluation: When evaluating the performance of the YOLO-Bud model in the image segmentation task, multiple key quantitative indicators were used for comprehensive evaluation, including precision (Precision, P), recall (Recall, R), intersection over Union (IoU), Dice Similarity Coefficient (Dice) and mean Average Precision (mAP). Precision (P), as the core indicator for evaluating the prediction accuracy of the model at the pixel level, reflects the proportion of target pixels correctly identified by the model among all pixels predicted as positive, revealing the accuracy of the model in predicting positive instances. Recall (R) measures the model's ability to successfully identify all truly labeled samples (such as objects or features to be segmented in the image), providing substantial insights into the overall effectiveness of evaluating the model in fully detecting positive sample instances. The intersection over union (IoU) is a key criterion for measuring the similarity between the model's predicted segmentation results and the true labeled segmentation results. By calculating the ratio of the intersection area and the union area of the predicted segmentation area and the actual labeled area, it intuitively reflects the degree of fit between the model's predicted segmentation and the actual situation. The higher the IoU value, the better the model's segmentation performance. In addition, the Dice similarity coefficient (Dice) is also an important means of evaluating segmentation performance. It indirectly measures the model's ability to detect subtle target differences by quantifying the overlap between the predicted segmentation and the true segmentation. Finally, the mean average precision (mAP) is a key indicator for evaluating the overall performance of the model in multi-class segmentation tasks. It counts the average accuracy of the model over all categories, thereby comprehensively reflecting the overall performance level of the model in multi-class segmentation scenarios. The calculation formulas for each indicator are as follows (10)-(14). In (10)-(14), TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative, respectively. The value of k refers to the number of classes of segmented objects in the model. The number of segmentation classes in this dataset is 4, i.e., k = 4. AP is an important indicator for evaluating the performance of a model in multi-class classification tasks. The calculation formula of AP is as shown in formula (15).
[0121] Model training: All experiments are based on the PyTorch2.1 framework. The experimental environment is configured with an NVIDIA GeForce RTX 4090 Laptop GPU, 16GB of memory, and an operating system of Ubuntu 20.04. For the training of the YOLO-Bud model, the AdamW optimizer is used, the initial learning rate is set to 0.00125, the momentum parameter remains at 0.9, and the weight decay coefficient is 0.0005. On this basis, the learning rate adjustment strategy adopts the cosine annealing method. The input image is uniformly normalized to a size of 640×640 pixels, and the model is trained for up to 500 cycles. During the training process, if no significant performance improvement is observed for 50 consecutive iterations, the training process is automatically terminated. The accuracy and reliability of model training are ensured by configuring the experimental environment, strictly setting hyperparameters, and adopting effective training strategies. These measures not only help to accelerate the convergence of the model, but also effectively alleviate the overfitting phenomenon and minimize the model from falling into the local optimal solution.
[0122] Training results: The trained YOLO-Bud model will eventually get a weight parameter file for a specific data set. This file is usually a PyTorch model file with the suffix .pt or .pth. It contains all the parameters learned by the model during the training process (such as the weights and bias items of the convolutional layer) and the architecture information of the model. These weight parameters are the key to the model's ability to accurately identify and locate targets in images. During the training process, the model continuously iterates and optimizes these parameters to achieve the best performance (such as accuracy, recall, etc.) on a given data set. After training is complete, this weight parameter file is loaded into the YOLO-Bud model, and then inference (or prediction) is performed to detect targets on new image data. During inference, the model uses these learned parameters to extract image features and output the location and category information of the target.
[0123] Segmentation effect: Table 1 shows the comparison results of the YOLO-Bud model with the classic Mask R-CNN and YOLACT models on the multi-target segmentation task of the apple tree bud stage. The results show that the YOLO-Bud model has significant advantages in multiple key performance indicators. In the comparative experiments shown, the performance differences of several models for the multi-target segmentation task of the apple tree bud stage under good light and low light conditions were investigated. The YOLO-Bud model has advantages in target feature extraction and spatial detail expression, and can maintain good segmentation accuracy in complex environments and a variety of challenging conditions. This advantage comes from the optimized design of the YOLO-Bud model in feature pyramid construction, context information fusion, and balanced processing of local and global information, which makes it more robust and has better segmentation performance than classic methods in tasks that require fine segmentation such as the apple tree bud stage.
[0124] Table 1 shows the performance comparison results of different models.
[0125]
[0126] Algorithm for judging the affiliation of branch segments: When analyzing images of fruit trees during the bud stage to achieve intelligent flower thinning, accurately locating the branches to which the buds belong in the two-dimensional image is a complex task. Since there may be multiple intersecting and overlapping branches in the image, coupled with the occlusion effect between branches and the influence of other organs such as leaves, branches that originally belonged to the same physiological structure may be misjudged as multiple independent branch segments after image segmentation. During the segmentation process, although the YOLO-Bud model can effectively identify and segment individual branch segments, due to its independent segmentation nature and the lack of mutual affiliation information, it is unable to automatically associate and combine these segments to restore a complete and continuous branch structure. Therefore, the primary problem that needs to be solved is how to correctly judge the affiliation of these branch segments segmented by the model, and effectively reassemble these independent branch segments into a single branch structure that reflects their biological integrity.
[0127] For the branch segments after instance segmentation, the study used the principal component analysis (PCA) method and selected the first principal component (PCA1) to characterize the local growth direction of the branch segments. In order to determine whether two adjacent branch segments belong to the same physiological branch, two core constraints were proposed: first, the angle between the first principal component vectors of the segmented regions of adjacent branch segments must meet a threshold of less than 20°; second, the angle between the first principal component and the centroid of the branch segment polygon must meet a threshold of less than 15°, achieving a classification accuracy of 96%.
[0128] 2D reconstruction of occluded branches: During the bud growth period, partial occlusion of branches by leaves, buds and other plant organs is common, which poses a challenge to the construction of a complete and accurate branch skeleton, and has a direct impact on the accuracy and reliability of subsequent bud positioning. To address this problem, a branch segment attribution algorithm based on the consistency of the directions of adjacent branch segments, principal component analysis (PCA), B-spline curve fitting and iterative connection strategy is proposed on the basis of completing the branch segment attribution algorithm, aiming to achieve effective splicing of branch segments and finally construct a continuous and accurate branch model.
[0129] According to the consistency of the spatial direction of each branch segment that has been separated and belongs to the same branch, the first principal component analysis (PCA1) direction representing the branch direction is selected, and the straight line equation L1(x,y) is constructed based on this direction and the center of mass of the polygon, as shown in formula (4-12). This straight line divides the coordinate points in the polygonal area of the branch segment (based on the image coordinate system) into two sub-areas L1(x1,y1)>0 and L1(x2,y2)<0 according to the sign of the straight line L1, as shown in the figure. Figure 4 1 and 2. In order to achieve seamless splicing of adjacent branch segments, especially in the restoration of the occluded part, the present invention focuses on selecting key control points in the two sub-areas divided by the straight line L. These control points can represent the boundary characteristics of their respective areas, which is crucial for simulating the edge morphology of the occluded branch segments. Based on these control points, B-spline curves are introduced as a smooth interpolation tool to estimate the continuous curve morphology of the edge of the occluded area. Due to its locality, flexibility and high-order continuity, the B-spline curve can naturally fill the occluded area while keeping the position of the boundary points unchanged, achieve a smooth transition between the edges of the branch segments, and effectively simulate the natural shape of the occluded area. As shown in formula (16).
[0130] The branch segment regions obtained by model segmentation show the characteristics of being closed and connected in sequence. However, the endpoint curve segments (pseudo-edges) formed at both ends of these closed regions cannot accurately reflect the direction of the growth trend of the edge characteristics of the branches, which interferes with the accuracy of the subsequent branch segment connection to a certain extent. In particular, when the regional points are divided according to the sign of the straight line L1, some endpoint curve points are taken into consideration, and these points have significant differences in spatial distribution from the laws of branch segment edge points, which may lead to large variations in the spatial distribution of edge point coordinates, which is not conducive to achieving a smooth transition of B-spline curve interpolation. In order to ensure the accuracy and smoothness of the branch segment connection, it is necessary to "eliminate" the points contained in the curve endpoints, eliminate the influence of abnormal points on the spatial distribution of edge point coordinates, and ensure that the transition between edge points remains highly smooth during the subsequent B-spline curve interpolation process.
[0131] The present invention constructs a straight line L2(x,y) along the second principal component analysis (PCA2) direction of the polygonal area, taking this direction and the center of mass of the polygon as the reference, and the line segment formed by the two points in the polygonal area represents the pixel length (the center of mass span of the branch segment) of the diameter of the branch segment mapped in the two-dimensional image. In view of the fact that the number of pixels contained in the curve segments at both ends of the branch segment and the center of mass span is similar, the number of control points in the two sub-areas divided by the straight line L1 and the number of pixel points contained in the center of mass span segment are counted. On this basis, by appropriately eliminating points at both ends of the control point set, the endpoint curve pixel points that may be greatly affected by the variation are removed, thereby ensuring that the remaining control points can more accurately represent the edge features of the branch segment.
[0132] Although the two sub-regions divided by L1 may not contain all the pixel points of the endpoint curve, the spatial variability of the edge features at both ends of the branch segment, especially at locations far from the centroid, may also be large. Therefore, in order to reflect the edge features of the branch as realistically as possible and make full use of the advantages of B-spline curve fitting, the pixel coordinate points close to the centroid are retained first, and the edge points with drastic variations are appropriately eliminated.
[0133] In the specific operation, let a sub-region sequentially sorted coordinate point set be p, which contains n points, and the centroid span line segment contains m coordinate point sets. At both ends of the coordinate point set p, points with 1.2 times the number of pixels contained in the centroid span line segment are removed to form a new point set p′. This strategy can effectively reduce unnecessary computing resource consumption while ensuring accurate capture of branch segment edge features. For example, equations (17) and (18).
[0134] Wherein, [1.2*m] is the maximum integer not exceeding 1.2*m.
[0135] B-spline interpolation is a powerful mathematical tool used to construct a smooth and continuous curve given a set of discrete data points so that the curve passes through all the data points exactly. The key features of the B-spline interpolation algorithm include local support, smooth transition, and flexibility of arbitrary order. The specific interpolation formula can be expressed as Equation (19).
[0136] Where C(t) is the curve value at the interpolation point, P i is the ith control point, is the value of the nth-order B-spline basis function associated with the control point at parameter t. Through this formula, given a set of edge coordinates as control points, the position of the curve is calculated at any parameter value t, and then the complete interpolation curve is obtained.
[0137] When some edges of branches cannot be directly observed due to occlusion, the B-spline interpolation algorithm is used to interpolate the occluded parts with the help of known visible edge coordinate data. It can not only fill in the missing coordinate information, but also ensure that the generated edge feature coordinates maintain a high degree of consistency and smooth transition with the known data points, thereby reproducing the true outline of the occluded area. By applying the B-spline interpolation algorithm to the edge coordinate set, the occluded area can be reasonably inferred and reconstructed while maintaining the continuity of the overall shape. The local support of the algorithm makes the interpolation process rely only on valid data points near the occluded area, avoiding the interference of irrelevant information and ensuring the robustness of the interpolation results.
[0138] The present invention selects a cubic (n=3) B-spline basis function for interpolation calculation, which has C 2 Continuity means that the curve itself and its first and second order derivatives are continuous. This high degree of smoothness ensures that the interpolation result not only has no abrupt changes at the data points, but also transitions naturally between the data points. This is especially important for describing the edges of natural objects such as apple tree branches, where the edges are continuous and gradually changing. Figure 6 The curve estimation effect obtained after interpolation by the third-order B-spline curve is shown. As can be seen from the figure, the interpolation curve passes through the provided data points, and the adjacent data points show a continuous and gradual shape without any sharp corners, broken lines or discontinuities, which conforms to C 2 The definition of continuity. It also fits the natural morphological characteristics of the edge of the apple tree branch.
[0139] Based on the above analysis method, an iterative process is constructed to implement the two-dimensional reconstruction process algorithm of partial occlusion of apple tree branches, such as Figure 7 shown.
[0140] Determine the subordination of flower buds, flower branches and branches: When estimating the position distribution of flower buds on the branches to which they belong, the first task is to identify the branch to which each flower bud belongs. Since the morphological information of flower buds and branches in two-dimensional images is complex and easily interfered by factors such as occlusion and overlap, it is quite challenging to rely solely on these intuitive features to determine the subordination relationship between flower buds and branches. Most flower buds and flower branches are relatively close to their branches in spatial distribution. The present invention first searches for the nearest neighbor branch area with the target (flower buds and flower branches) centroid as the base point. To calculate the shortest distance from a point to a polygon, it is necessary to calculate the distance from the point to each edge of the polygon, such as Figure 8 d in 1 -d 5 shown.
[0141] Take the calculation of the distance from point p to polygon AB as an example. With vector The angle between them is β, and the vector inner product can be used to obtain the relationship shown in equations (14) and (15). According to this formula, the k value can be obtained, and then the coordinates of F can be obtained, and finally the length of PF can be obtained. It should be noted that point F does not necessarily fall on line segment AB. Fig. 9 There are several situations: (1) k>1, F is on the extension line of AB, and the shortest distance is PB; (2) 0≤k≤1, F is on AB, and the shortest distance is PD; (3) k<0, F is on the reverse extension line of AB, and the shortest distance is PA.
[0142] In the spatial relationship analysis between polygons and points, given the high time complexity of directly calculating the distance from a point to each side of a polygon when there are many sides, it is an efficient processing strategy to prioritize the relative positional relationship between a point and a polygon. By judging the initial position, the subsequent distance calculation process can be effectively simplified. If the analysis determines that the point is inside the polygon, its shortest distance to the polygon may be the distance from the point to the nearest vertex, thus avoiding the calculation of the distance of each edge one by one. On the contrary, if the point is outside the polygon, in order to find the shortest distance, the distance from the point to each side of the polygon must be calculated and compared. In special cases, when a point is located on the edge or vertex of a polygon, the shortest distance from the point to the polygon is zero, and no further distance calculation is required. The strategy of determining the positional relationship between a point and a polygon as a preprocessing step is essentially an effective pruning method that reduces the computational burden when processing a large number of points associated with complex polygons.
[0143] The commonly used method to determine the positional relationship between a point and a polygon is the "ray intersection method". Starting from the point to be detected, draw a ray in any direction (usually horizontal or vertical directions are chosen to simplify calculations). Record the number of intersections between this ray and the edge of the polygon. If the ray enters from one side of the polygon and leaves from the other side, the count increases by one each time it crosses the edge. If the number of intersections is an odd number, the point is inside the polygon; if it is an even number, the point is outside the polygon. In addition, by comparing the slopes or coordinate positions of adjacent points, it is possible to determine whether a point is on the edge.
[0144] Using the above algorithm, after statistical analysis of 50 image samples, the accuracy of flower bud identification was 94.9%. Since the spatial distribution of flower branches is closer to branches, its accuracy of identification is higher than that of flower buds, reaching 96.3%.
[0145] Bud growth point estimation: To determine the growth point of the bud, the present invention proposes a method based on the centroid relationship. The centroids of the buds and branches belonging to the same branch are calculated, and the intersection of these centroid lines and the branches is identified as the growth point of the bud. The proposal of this method depends on clarifying the subordinate relationship between each bud and the corresponding branch. To this end, the nearest neighbor algorithm is used as an initial step to effectively predict the mutual subordinate relationship between most buds and branches. Further statistical analysis reveals a significant feature between the growth direction of the bud and the axial direction of the branch growth, that is, the two form a mostly acute angle relationship (accounting for 98.0%), and when observed along the growth direction of the branch, the centroid of the bud is located in front of the centroid of the branch to which it belongs. It provides additional judgment criteria for the algorithm and strengthens the rationality of the judgment of the subordinate relationship between the bud and the branch. The branch area is small and is often blocked by surrounding leaves, which may affect the accuracy of the centroid coordinates. To address this problem, the intersection of the minimum circumscribed rectangle and the leaf area is used to reconstruct the partially blocked branch area, thereby improving the accuracy of position recognition.
[0146] After implementing the affiliation judgment algorithm, the affiliation between most flower buds and branches was accurately identified and Fig.11 The white line in the middle marks the relevant objects. When determining the growth point of the flower bud, the midpoint of the intersection of the flower branch and the branch is determined as the growth point of the flower bud. When the flower branch area and the branch area intersect, the growth point of the flower bud is estimated as the extension of the centroid line ( Fig.11 The intersection of the red line in the middle) and the flower branch polygon; otherwise, if there is no intersection, the intersection of the extended line and the branch polygon is taken as the estimated growth point of the flower bud. Fig.11 As shown, the yellow dots represent the estimated bud growth points, while the white dots are the actual growth points.
[0147] from Fig.11From the observation results of (b), it can be seen that although the estimated growth points and the actual growth points show good consistency on most objects, there are slight deviations in certain cases (such as samples b-1 and b-3). This deviation is mainly attributed to the mutual occlusion and interference of natural elements such as branches, twigs and buds in complex scenes, which affect the accurate recognition of contours during image processing. In particular, when branches or buds are partially invisible due to overlap, the contour recognition ability of the model is limited, which leads to the distortion of the branch length in the reconstructed two-dimensional image compared with the actual one, the edge positioning is inaccurate, and finally the predicted position of the intersection point between the centroid line and the ideal edge is offset. In addition, the different shooting angles and the randomness of the plant growth position are also the reasons for the mismatch in the number of flower buds and branches detected and the difficulty in accurately locating the growth points of some flower buds. This difference not only increases the difficulty of identification, but also poses a challenge to the accuracy of growth point estimation. Despite the above challenges, the algorithm implemented in the detected objects still showed an object matching degree of 90.2%, which can effectively distinguish and locate targets such as branches, twigs and buds in complex environments.
[0148] The above description is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several changes and improvements without departing from the creative concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A method for identifying and locating flower buds of fruit trees, characterized in that The following procedures are included: Model building: Building a multi-target recognition and segmentation model for fruit tree bud stage: including data set acquisition, data enhancement, model building, Data acquisition: images of fruit tree branches in the bud stage are acquired by an imaging device, such as a mobile phone camera, a digital camera and / or a depth camera, and the images are acquired and selected by taking pictures at different distances of 0.3-3.0 m from the branches under variable natural lighting conditions, including at least sunny and cloudy days, at different time periods, including at least morning, noon and afternoon; Data enhancement: Use digital image processing technology to perform targeted preprocessing and enhancement operations on the selected images, including but not limited to color space conversion, geometric transformation, such as flipping, translation, rotation, adjusting brightness and contrast, and image fusion methods, strategically remove the leaf areas with a large number of leaves in the image, and maintain the visual integrity of branches, buds and branches while ensuring the overall structure of the image, ultimately improving the overall generalization performance and robustness of the model, and finally normalizing the obtained image into a standardized image; Generate new images using image target synthesis data augmentation algorithm; In addition, we selectively selected several image samples with subtle differences in the number of target objects and a high proportion of small targets from the original data set and divided them into training and validation sets in a ratio of 8:2-8:3; The specific steps are as follows: ① Use the polygon area point coordinate information in the json file generated by the annotation of image A to create the corresponding mask; ② Use the mask generated in step ① to extract the corresponding RGB pixels in the image to the black background area; ③ Apply the translation matrix to adapt the mask generated in step ① to the area specified by image B; ④ Extract RGB pixels based on the mask generated in step ③ to the corresponding positions in image B to synthesize a new image; And specifically solve the problem of serious imbalance in the distribution of category pixels by constructing a fruit tree bud stage dataset containing several images. At the same time, randomly select images with a smaller number of images than the fruit tree bud stage dataset from the unlabeled images as an independent test set. Model construction: Establish a YOLO-Bud model, which includes an input layer, a backbone network (Backbone), a neck network (Neck), and a head network (Head); The input layer receives images of any pixel size and then crops them into images of standardized pixels after image preprocessing in the early stage of the model. The backbone network (Backbone) described above: inputs the enhanced image, and the backbone network (Backbone) uses 3x3 or 1x1 convolution modules and pooling operations to progressively extract multi-scale features containing rich semantic information and spatial structure. The specific execution process includes: The first convolution layer (p1): repeated once, the feature map scaling factor is 1 / 2; the second convolution layer (p2): repeated once, the feature map scaling factor is 1 / 2; the third deformable convolution layer (C2f-DCN-t): repeated 3 times, the feature map scaling factor is 1; the fourth lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the fifth deformable convolution layer (C2f-DCN-t): repeated 6 times, the feature map scaling factor is 1; the sixth lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the seventh deformable convolution layer (C2f-DCN-t): repeated 3 times, the feature map scaling factor is 1; the eighth lightweight downsampling module (DSample Lite): repeated once, the feature map scaling factor is 1 / 2; the last spatial pyramid pooling module (SPPFX): repeated once, the feature map scaling factor is 1; The neck network (Neck): after feature extraction, it is introduced into the Spatial Pyramid Pooling–Fast–X (SPPFX) pyramid pooling layer, which unifies multi-scale features through pooling operations of different scales and retains important spatial information; Subsequently, the feature information is passed to the Neck network part. Through upsampling and cross-layer feature fusion methods, the spatial position information of shallow features is closely combined with the abstract semantic knowledge of deep features to construct a hierarchical feature representation. The specific execution process includes: Upsample: The eleventh layer (p4) is repeated once, and the feature map scaling factor is 1 / 2; the thirteenth layer (p3) is repeated once, and the feature map scaling factor is 1 / 2; Feature map concatenation (Concat): The twelfth layer (p4) is repeated once, and the feature map scaling factor is 1; the fourteenth layer (p4) is repeated once, and the feature map scaling factor is 1; Deformable convolutional layer (C2f-DCN): The fifteenth layer (p3) is repeated 3 times, and the feature map scaling factor is 1 / 2; the eighteenth layer (p4) is repeated 3 times, and the feature map scaling factor is 1 / 2; Lightweight downsampling module (DSample Lite): The sixteenth layer (p4) is repeated once, and the feature map scaling factor is 1 / 2; the nineteenth layer (p5) is repeated once, and the feature map scaling factor is 1 / 2; The head network (Head): The multi-level feature map output by Neck is sent to the detection head for training and reasoning. The detection head predicts the category, position and segmentation mask of each target based on the fusion features. It contains three decoupled heads (Decoupled Head), each head is responsible for a specific task output, and each decoupled head contains a deformable convolution layer (C2f-DCN) and feature map splicing (Concat). Optimization of the YOLO-Bud model: the number of pooling layers of the spatial pyramid pooling feature (SPPF) module in the Backbone architecture of the model is 512 (SPPFX); secondly, for the C2f module in YOLOv8, an efficient dynamic sparse operator DCNv4 (DeformableConvolution v4) is introduced based on the optimized and upgraded iterative version DCNv3 (Deformable Convolution v3); The lightweight downsampling module DSampleLite (Down Sample Lite) is used as the downsampling method; The DSampleLite module first performs an average pooling operation on the input feature map, then evenly divides the feature map in the channel dimension and performs targeted processing on each part: the feature map is spatially downsampled by 3×3 convolution, and feature extraction is performed by using maximum pooling combined with 1×1 convolution; A lightweight attention mechanism is also introduced, which applies adaptive weighting to the downsampled feature maps through global average pooling and channel attention operations; Shape-IoU is introduced to emphasize the accurate modeling of bounding box shape and scale factors; The formula of Shape-IoU is: (1) (2) (3) (4) (5) (6) Where scale is the scale factor, which is related to the scale of the target in the data set; B and B gt Represent the prediction box and GT box; ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are GT The corresponding bounding box regression loss function is: (7) In addition, Focaler-IoU focuses on the differentiated attention to regression samples of different difficulty levels. Focaler-IoU reconstructs the IoU loss through linear interval mapping, allowing the model to pay different attention to easy and difficult samples during training; The Focaler-IoU formula is as follows: (8) in IoU focaler is the reconstructed Focaler-IoU, IoU is Shape-IoU, [ d , u ]∈[0,1]; by adjusting d (0) and u (0.95) to make the loss function focus on different regression samples; The loss is defined as follows: (9); Model training: For the training of the YOLO-Bud model, the AdamW optimizer was used, the initial learning rate was set to 0.00125, the momentum parameter remained at 0.9, and the weight decay coefficient was 0.0005; on this basis, the learning rate adjustment strategy adopted the cosine annealing method; the input image was uniformly normalized to standard pixels, and the model was trained for up to 500 cycles; during the training process, if no obvious performance improvement was observed after 50 consecutive iterations, the training process was automatically terminated; Training results: The final trained YOLO-Bud model is stored in the weight parameter file.
2. The method for identifying and locating flower buds of fruit trees according to claim 1, characterized in that: The principal component analysis (PCA) method was used for the branch segments after instance segmentation, and the first principal component (PCA1) was selected to characterize the local growth direction of the branch segments; Two core constraints are adopted: first, the angle between the first principal component vectors of adjacent branch segment segments must be less than the 20° threshold; second, the angle between the first principal component and the centroid of the branch segment polygon must be less than the 15° threshold; On the basis of completing the branch segment attribution algorithm, the strategy based on the consistency of the directions of adjacent branch segments, principal component analysis (PCA), B-spline curve fitting and iterative connection is adopted; Select the first principal component analysis (PCA1) direction representing the branch direction, and use this direction and the polygon centroid as the reference to construct the straight line equation L1(x,y), as shown in formula (4-12); this straight line divides the coordinate points in the branch segment polygon area (based on the image coordinate system) into two sub-areas L1(x1, y1) > 0 and L1(x2, y2) < 0 according to the sign of the straight line L1; B-spline curves are introduced as smooth interpolation tools to fill the occluded area, achieve smooth transitions between branch segments, and simulate the natural shape of the occluded area. (10); By following the second principal component analysis (PCA2) direction of the polygonal area, taking this direction and the center of mass of the polygon as the reference, a straight line L2(x,y) is constructed, and the line segment formed by the two points in the polygonal area represents the pixel length (the center of mass span of the branch segment) of the diameter of the branch segment mapped in the two-dimensional image. By removing points at both ends of the control point set appropriately, the end point curve pixels that may be greatly affected by the variation are removed. In the specific operation, suppose the coordinate point set of a sub-region is sorted in sequence , which contains n points, the centroid span line segment contains m coordinate point sets, in the coordinate point set Eliminate 1.2 times the number of pixels contained in the centroid span line segment at both ends of the point to form a new point set ; (11) (12) Wherein, [1.2*m] is the largest integer not exceeding 1.2*m; B-spline interpolation is used to construct a smooth and continuous curve given a set of discrete data points, so that the curve passes through all the data points accurately. The B-spline interpolation algorithm formula is expressed as: (13) in is the curve value at the interpolation point, It is i control points, is related to the control point n Step The spline basis function has parameters The value at .
3. The method for identifying and locating flower buds of fruit trees according to claim 2, characterized in that: An iterative process is constructed to realize the two-dimensional reconstruction process algorithm of partial occlusion of fruit tree branches, including the subordinate judgment of flower buds, flower branches and branches, and the estimation of flower bud growth points. The nearest neighbor branch area is found based on the centroid of the target (flower bud and flower branch) to calculate the nearest distance from the point to the polygon. The centroids of the flower bud and flower branch areas belonging to the same branch are calculated, and the intersection of these centroid lines and the branches is identified as the growth point of the flower bud. The specific algorithm flow is as follows: When specifically determining the growth point of the flower bud, the midpoint of the intersection line between the flower branch and the branch is determined as the growth point of the flower bud. When the flower branch area and the branch area intersect, the growth point of the flower bud is estimated to be the intersection point of the extension of the centroid line and the flower branch polygon. On the contrary, if there is no intersection, the intersection of the extended line and the branch polygon is taken as the estimated growth point of the flower bud.
4. The method for identifying and locating flower buds of fruit trees according to claim 3, characterized in that: After completing the attribution of branch segments, perform two-dimensional reconstruction of the blocked branches: According to the consistency characteristics of the branch segments that have been separated and belong to the same branch in the spatial direction, the first principal component analysis (PCA1) direction representing the branch direction is selected. Based on this direction and the polygon centroid, the straight line equation L1(x,y) is constructed. This straight line divides the coordinate points in the polygonal area of the branch segment into two sub-areas L1(x1, y1) > 0 and L1(x2, y2) < 0 according to the sign of the straight line L1. Control points are selected in the two sub-areas divided by the straight line L. Based on these control points, the B-spline curve is introduced as a smooth interpolation tool to estimate the continuous curve shape of the edge of the occluded area, effectively simulating the natural shape of the occluded area. (10), By following the second principal component analysis (PCA2) direction of the polygonal area, a straight line L2(x,y) is constructed based on this direction and the center of mass of the polygon. The line segment formed by the two points in the polygonal area represents the pixel length of the diameter of this branch segment mapped in the two-dimensional image. On this basis, by removing appropriate points at both ends of the control point set, the endpoint curve pixels that may be greatly affected by the variation are removed. By applying the B-spline interpolation algorithm to the edge coordinate set, the occluded area can be reasonably inferred and reconstructed while maintaining the continuity of the overall shape; Select three times ( n =3) The B-spline basis function performs interpolation calculations, which have C² continuity, that is, the curve itself and its first-order and second-order derivatives are continuous.