Blueberry whole growth cycle detection method and system based on improved YOLO model
By improving the YOLO model and combining it with lightweight semantic segmentation and cross-cycle feature memory modules, the problems of small target detection accuracy and occlusion in the detection of the entire blueberry growth cycle were solved, realizing high-precision dynamic monitoring of the blueberry growth cycle and meeting the needs of growth trend prediction in precision agriculture.
Patent Information
- Application Number
- CN202511039909.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies for detecting blueberries throughout their entire growth cycle suffer from insufficient accuracy in detecting small targets, difficulty in extracting features from scenes with target occlusion, and a lack of dynamic correlation analysis throughout the entire growth cycle, resulting in detection results that cannot provide support for predicting growth trends.
An improved YOLO model is adopted, and a lightweight semantic segmentation model and a conditional generative adversarial network are used to accurately extract the ROI region. The target features are enhanced by combining CBAM and a cross-period feature memory module. A cross-modal mapping matrix and LSTM temporal connection are introduced to carry out multi-task joint training and edge optimization, and a multi-factor risk assessment function is constructed.
It achieves high-precision detection of the entire growth cycle of blueberries, accurately predicts growth trends, provides real-time monitoring of growth status, and meets the needs of precision agriculture.
Smart Images

Figure CN120932095A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a method and system for detecting the entire growth cycle of blueberries based on an improved YOLO model. Background Technology
[0002] With the rapid development of smart agriculture, crop growth monitoring technology based on computer vision has become a key means to improve the level of precision planting management. In the field of blueberry planting, the realization of automated detection of the entire growth cycle (budding, flowering, fruiting, and ripening) is of great significance for yield prediction and agricultural operation planning. Traditional agricultural monitoring mainly relies on manual observation or single sensor data, which has significant drawbacks such as strong subjectivity, poor real-time performance, and inability to quantify and analyze data. In recent years, target detection models such as the YOLO series have been widely used in agricultural scenarios due to their high efficiency and accuracy. However, existing technologies still have many shortcomings in meeting the specific needs of the entire growth cycle of blueberries.
[0003] First, during the blueberry budding stage, the bud size is typically less than 10×10 pixels. During the flowering stage, the petals have rich edge details but the target area is low. Traditional YOLO models are prone to missed or false detections when extracting features from small targets. Additionally, the dense foliage of blueberry plants easily causes target occlusion in natural growing environments. Furthermore, uneven lighting in greenhouses (such as switching between cloudy days and artificial lighting) can lead to color shifts in images. Existing models struggle to distinguish between the target and the background when dealing with scenes where foliage occlusion occurs.
[0004] Finally, current detection methods mostly focus on single-cycle identification and lack dynamic correlation analysis of the entire growth cycle. They cannot predict the flowering time window from the current budding state or infer the yield at maturity based on the characteristics of young fruit. This lack of time-series modeling means that the detection results can only provide static classification and cannot provide growth trend prediction support for planting decisions. It is difficult to meet the needs of precision agriculture for dynamic monitoring of crop growth. At present, there is a need for a blueberry full growth cycle detection method and system based on an improved YOLO model. Summary of the Invention
[0005] To address the issues of missing growth cycle continuity and insufficient accuracy in small target detection in traditional blueberry detection methods, this invention provides a blueberry full growth cycle detection method and system based on an improved YOLO model.
[0006] In a first aspect, the present invention provides a method for detecting the entire growth cycle of blueberries based on an improved YOLO model, employing the following technical solution:
[0007] A method for detecting the entire growth cycle of blueberries based on an improved YOLO model, including...
[0008] Collect multi-dimensional image data of blueberries at different growth stages and annotate the acquired multi-dimensional image data;
[0009] Preprocessing is performed on the labeled multi-dimensional images, including semantic segmentation of the image data using a lightweight semantic segmentation model, and multi-scale dynamic input based on different annotated images;
[0010] A YOLO detection model is constructed based on the preprocessed multi-dimensional images, including adding a cross-period feature memory module after the backbone network and calculating the importance score of each channel of the backbone network.
[0011] The completed YOLO detection model is jointly trained on multiple tasks, including constructing a cross-modal mapping matrix of RGB features and NIR features, and aligning similar features in the embedding space by using a contrastive loss function.
[0012] Edge-end inference optimization is performed based on the trained YOLO detection model, including model updates through feature space alignment loss and EWC regularization;
[0013] A multi-factor risk assessment function was constructed based on the output of the YOLO detection model to complete the final detection of blueberry growth status.
[0014] Furthermore, the semantic segmentation of image data using a lightweight semantic segmentation model includes: firstly, extracting the mask matrix of blueberry organs using the lightweight semantic segmentation model; determining regions with mask values greater than a threshold as Regions of Interest (ROIs) and the remaining regions as background regions; applying bicubic interpolation to ROIs and nearest neighbor interpolation downsampling to background regions; unifying image size; and adaptively selecting the input resolution based on the average target size of different labeled images. The determination expressions for different regions are as follows:
[0015]
[0016] Where I represents the original input image, M represents the mask matrix output by the lightweight semantic segmentation model, ⊙ represents element-wise multiplication, Bicubic represents bicubic interpolation, and Nearest represents nearest neighbor interpolation.
[0017] Furthermore, the preprocessing based on the labeled multi-dimensional images also includes constructing a conditional generative adversarial network (GAN). The generator employs a U-Net architecture, taking as input the real image after ROI processing and multi-scale input, along with growth cycle labels. It generates enhanced samples simulating different lighting and occlusion conditions through an encoder-decoder path. The discriminator uses a PatchGAN structure, and the generator and discriminator are alternately optimized using an adversarial loss function. The formula for the adversarial loss function is:
[0018]
[0019] Among them, I real This indicates that the image is indeed of a blueberry. Let represent the simulated image output by the generator, y represent the growth cycle label, D(.,.) represent the discriminator function, and E[.] represent the mathematical expectation operator.
[0020] Furthermore, the YOLO detection model constructed based on the preprocessed multi-dimensional image includes embedding a CBAM attention mechanism in the backbone network based on the CSPDarknet architecture, enhancing target features through channel attention and spatial attention, adding a cross-period feature memory module after the backbone network output, fusing the feature map sequence of consecutive T frames through 3D convolution to obtain temporal features, which are then connected to the neck network, calculating the importance score of each channel in the backbone network, averaging the channel attention maps output by CBAM in the spatial dimension, and dynamically adjusting the channel retention ratio according to the proportion of mature samples in the current batch. The importance score formula is as follows:
[0021]
[0022] Where, ω c This is represented by the importance score of the c-th channel in the backbone network, where H and W represent the height and width of the feature map, respectively. Represented by a double summation symbol, this indicates that the summation is performed over all pixels in the spatial dimension of the feature map. This represents the pixel value of the c-th channel in the channel attention map, with coordinates (h, w).
[0023] Furthermore, the construction of the YOLO detection model based on the preprocessed multi-dimensional image also includes introducing LSTM temporal connections into the bidirectional feature pyramid network of the neck network. The features of the current frame are interacted with the output features of the previous frame through LSTM to construct a spatiotemporal feature fusion pyramid. Additional convolutional branches are used to predict the sampling point offset, dynamically adjusting the sampling kernel position. Finally, three detection heads with different resolutions are established. The LSTM interaction formula is as follows:
[0024]
[0025] in, This is represented as the LSTM output feature of the current frame. This represents the original feature input of frame t in the bidirectional feature pyramid network. This represents the output features of the previous frame's LSTM, where LSTM(.) represents the LSTM function operation, and ω... i Let ω represent the importance weight of the i-th input feature, where ∈ denotes the minimum constant. jIt is represented as the importance weight of the j-th input feature.
[0026] Furthermore, the multi-task joint training of the constructed YOLO detection model includes using quantile regression to predict the remaining time of the blueberry growth cycle, with the loss function being a weighted sum of multiple quantile losses; constructing a cross-modal mapping matrix W for RGB and NIR features; aligning similar features in the embedding space using a contrastive loss function; defining the sequential encoding of the growth cycle; and forcing adjacent frame labels to satisfy the growth order using a temporal loss function. The formula for the contrastive loss function is as follows:
[0027]
[0028] Among them, f rgb f is represented as a feature vector extracted from an RGB image via a backbone network. nir Let f' be the feature vector extracted from the NIR image. nir It represents the negative sample NIR features, W1 represents the cross-modal mapping matrix, and τ represents the temperature parameter.
[0029] Furthermore, the edge-end inference optimization based on the trained YOLO detection model includes calculating the maximum confidence score of each detection head in the predicted probability vector for each growth cycle of blueberries. When the maximum confidence score is greater than a set confidence threshold, the branch calculation for that detection head is retained. The iCaRL algorithm is used, and online model updates are achieved through feature space alignment loss and EWC regularization. The update formula is:
[0030]
[0031] Where, θ t L represents the parameters of the current model. cls Let L represent the classification loss function, η represent the learning rate, and L represent the classification loss function. align L represents the feature space alignment loss. ewc It is represented as the elastic weight consolidation loss, and λ1 and λ2 represent different weight coefficients.
[0032] Furthermore, the edge-end inference optimization based on the trained YOLO detection model also includes using the MOG algorithm to build a background model for each frame, calculating the absolute value of the pixel difference between the current frame and the background model, locating motion regions through inter-frame difference, and performing detection on the motion regions. The formula for the background model is:
[0033] B t (x,y)=MOG(B t-1 ,I t ),
[0034] Among them, B t-1Represented as the background model of the previous frame, I t This represents the current frame image.
[0035] Furthermore, the multi-factor risk assessment function is constructed based on the output results of the YOLO detection model. This includes extracting the growth cycle label, remaining time prediction, and target positioning information output by the YOLO model, constructing a multi-factor risk assessment function, including the growth cycle deviation risk factor and the remaining time fluctuation risk factor, calculating a comprehensive risk index based on the growth cycle deviation risk factor and the remaining time fluctuation risk factor, and finally outputting the growth cycle classification.
[0036] Secondly, a blueberry full growth cycle detection system based on an improved YOLO model includes:
[0037] The data acquisition module is configured to: collect multi-dimensional image data of blueberries at different growth stages and annotate the acquired multi-dimensional image data;
[0038] The preprocessing module is configured to perform preprocessing on the labeled multi-dimensional images, including semantic segmentation of the image data using a lightweight semantic segmentation model and multi-scale dynamic input based on different annotated images;
[0039] The model module is configured to: construct a YOLO detection model based on the preprocessed multi-dimensional image, including adding a cross-period feature memory module after the backbone network and calculating the importance score of each channel of the backbone network;
[0040] The training module is configured to perform multi-task joint training on the constructed YOLO detection model, including constructing a cross-modal mapping matrix of RGB features and NIR features, and aligning similar features in the embedding space through a contrastive loss function.
[0041] The optimization module is configured to perform edge-end inference optimization based on the trained YOLO detection model, including model updates through feature space alignment loss and EWC regularization.
[0042] The output module is configured to construct a multi-factor risk assessment function based on the output results of the YOLO detection model, and complete the final blueberry growth status detection.
[0043] In summary, the present invention has the following beneficial technical effects:
[0044] 1. This invention collects multi-dimensional image data during the blueberry growth process. Through a lightweight semantic segmentation model and a conditional generative adversarial network, it accurately extracts ROI regions and enhances samples to ensure the effective utilization of features at each stage, taking into account the characteristics of small buds during the budding stage, fragile petals during the flowering stage, diverse fruit shapes during the fruiting stage, and color changes during the ripening stage.
[0045] 2. This invention constructs a cGAN architecture that combines a U-Net generator with a PatchGAN discriminator, effectively simulating the appearance changes of blueberries under different lighting and occlusion conditions, significantly expanding the diversity of training samples, introducing growth cycle labels as conditional constraints, ensuring that the generated samples conform to biological laws, avoiding the generation of abnormal samples across cycles, and improving the model's generalization ability to real-world scenarios.
[0046] 3. This invention is based on the CSPDarknet backbone network with embedded CBAM and cross-period feature memory module, combined with the LSTM temporal connection of bidirectional feature pyramid network. It can not only strengthen the key features of blueberry organs (buds, flowers, fruits) at each stage through channel attention, but also use 3D convolution and LSTM to fuse continuous frame information to learn the growth dynamic changes from budding to maturity.
[0047] 4. This invention uses quantile regression to predict the remaining time of each growth stage, aligns RGB and NIR features through cross-modal comparative learning, and uses temporal loss to constrain the growth sequence, ensuring that the model maintains high accuracy in tasks such as predicting growth trends during the budding stage, identifying flower status during the flowering stage, tracking fruit development during the fruiting stage, and determining harvesting time during the ripening stage.
[0048] 5. The dynamic network pruning strategy of this invention closes invalid branches based on the confidence of the detection heads at each stage. In the scenario where large fruits are the main focus during the mature stage, small target detection heads are closed to reduce the amount of computation. The iCaRL algorithm combined with EWC regularization realizes online model updates, adapts to the changes in data distribution at different growth stages, and meets the real-time and stability requirements of full-cycle monitoring. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the overall process of a blueberry full growth cycle detection method based on an improved YOLO model according to an embodiment of the present invention.
[0050] Figure 2 This is a schematic diagram of the DySample dynamic kernel generation process according to an embodiment of the present invention. Detailed Implementation
[0051] The present invention will be further described in detail below with reference to the accompanying drawings.
[0052] Example 1
[0053] Reference Figure 1 This embodiment of a method for detecting the entire growth cycle of blueberries based on an improved YOLO model includes:
[0054] Collect multi-dimensional image data of blueberries at different growth stages and annotate the acquired multi-dimensional image data;
[0055] Preprocessing is performed on the labeled multi-dimensional images, including semantic segmentation of the image data using a lightweight semantic segmentation model, and multi-scale dynamic input based on different annotated images;
[0056] A YOLO detection model is constructed based on the preprocessed multi-dimensional images, including adding a cross-period feature memory module after the backbone network and calculating the importance score of each channel of the backbone network.
[0057] The completed YOLO detection model is jointly trained on multiple tasks, including constructing a cross-modal mapping matrix of RGB features and NIR features, and aligning similar features in the embedding space by using a contrastive loss function.
[0058] Edge-end inference optimization is performed based on the trained YOLO detection model, including model updates through feature space alignment loss and EWC regularization;
[0059] A multi-factor risk assessment function was constructed based on the output of the YOLO detection model to complete the final detection of blueberry growth status.
[0060] Specifically, a method for detecting the entire growth cycle of blueberries based on an improved YOLO model includes the following steps:
[0061] S1. Collect multi-dimensional image data of blueberries at different growth stages and label the acquired multi-dimensional image data;
[0062] like Figure 1As shown, an image acquisition device equipped with a multi-angle adjustable lens was used in a blueberry greenhouse to collect images according to phenological stages (budding, flowering, fruiting, and ripening). Images were collected at fixed times each week, covering the entire growth cycle. The adjustable lens uses an RGB-NIR dual-lens module. The RGB-NIR dual-lens module has a built-in optical beam splitter that separates the incident light into visible light (450-700nm) and near-infrared light (700-850nm), which are received by the RGB lens and NIR sensor respectively, achieving parallel acquisition in the 450-850nm wavelength band. The near-infrared channel is fixed at a wavelength of 800nm, which is sensitive to the sugar content of the fruit. The 800nm absorptive peak is a characteristic absorption peak, and the RGB channels retain the information of the three primary colors of red, green, and blue, which are used to capture color features, such as light green during the budding stage, white petals during the flowering stage, and the sugar content of mature blueberries is negatively correlated with 800nm reflectance. The higher the sugar content, the lower the reflectance. The NIR channel can be used to extract physiological indicators of the fruit to help judge the maturity. During the budding stage, the chlorophyll content of the buds is high, and the green light (G channel) reflection is strong. During the flowering stage, the pigment components of the petals make the red (R) and green (G) channel values higher than the blue (B) channel values. During the fruiting stage, the chlorophyll of the fruit degrades, and the absorption of blue and purple light increases. During the ripening stage, the accumulation of anthocyanins in the fruit makes the B channel value significantly increase, thus completing the acquisition of multi-dimensional data.
[0063] To acquire the multi-dimensional data, the LabelImg tool was used to annotate the 2500 collected images with anchor boxes. The labels were divided into four categories: budding stage, flowering stage, fruiting stage, and maturity stage, and were further divided into training set, validation set, and test set in a 7:2:1 ratio.
[0064] S2. Preprocess the labeled multi-dimensional images, including semantic segmentation of the image data using a lightweight semantic segmentation model and multi-scale dynamic input based on different labeled images;
[0065] Based on the labeled raw image data, a lightweight version of the YOLO-Pose model is used for real-time blueberry organ segmentation. This model is based on an anchor-free architecture and generates a mask matrix through keypoint prediction. Where M(x,y)∈[0,1] represents the probability that pixel (x,y) belongs to a blueberry organ. A mask threshold of 0.5 is set to binarize the mask into a Region of Interest (ROI) and background. For regions where M(x,y)>0.5, a bicubic interpolation algorithm (quadratic polynomial fitting) is used to calculate the new pixel value, as shown in the formula:
[0066]
[0067] Where I'(x,y) represents the pixel value at the target position (x,y) after interpolation, I(x+i,y+j) represents the pixel value at position (x+i,y+j) in the original image, and w(i,j) represents the weight calculated based on the bicubic kernel function, which depends on the distance from the target point (x,y) to the neighboring pixel (x+i,y+j). The weight w(i,j) is inversely proportional to the distance of the 16 adjacent pixels, which can preserve details such as bud tips and petal edges, and avoid the jagged effect when scaling small targets. For regions M(x,y)≤0.5, the nearest pixel value is directly taken to reduce the computation of background texture. At the same time, the ROI is masked by I⊙(1-M) mask to ensure that the background interpolation does not affect the target features. The judgment expression for different regions is:
[0068]
[0069] Where I represents the original input image, M represents the mask matrix output by the lightweight semantic segmentation model, ⊙ represents element-wise multiplication, Bicubic represents bicubic interpolation, and Nearest represents nearest neighbor interpolation.
[0070] The above interpolation strategy is applied to the RGB three-channel and NIR single-channel respectively, and the final output image is 320×320×4 in size. The consistency of spectral dimensions is maintained, which is convenient for subsequent model input. The proportion of bud pixels in the budding stage is <5%. Traditional global bilinear interpolation will cause blurring of details (such as loss of bud tip). ROI-preserving interpolation effectively improves the clarity of bud edge. Low-complexity nearest neighbor interpolation is used in the background area, which reduces the overall computational cost compared to global bicubic interpolation.
[0071] For the image after ROI processing, calculate the average size of all labeled targets:
[0072]
[0073] Among them, (x1) i ,y1 i x2 i ,y2 iLet be the coordinates of the i-th bounding box, and n be the number of targets. The input resolution is adaptively selected based on the average size of the targets in different labeled images. For the budding stage, a 416×416 input is used, which improves the pixel density compared to a 320×320 resolution. This increases the proportion of 5×5 pixel buds in the feature map from 2×2 to 3×3, avoiding feature loss due to downsampling. For the mature stage, a 256×256 input is used. This reduces the feature map size while still preserving sufficient semantic features for large targets at low resolution. The original aspect ratio is maintained during size conversion, and edges are filled with black (RGB=0,0,0; NIR=0) to prevent target deformation. For example, if the original image has an aspect ratio of 4:3, when converting to 416×416, black pixels are used to fill the top and bottom, while the original content is preserved on the left and right sides to ensure that the fruit shape is not distorted.
[0074] Finally, adversarial networks are used for image enhancement. The generator G adopts the U-Net architecture, and the input is the real image I. real Enhanced samples are generated using the periodic label y (one-hot encoded) through an encoding-decoding path. A conditional embedding layer is added to the intermediate layer to ensure that the generated image retains the periodic features corresponding to the label (e.g., when the input label is "flowering period", a white petal region is forcibly generated). The discriminator D adopts a PatchGAN structure and outputs a discrimination matrix to determine the authenticity of each image patch, improving the local realism of the generated image and avoiding global blurring. The loss function L of the adversarial network... D It consists of two parts: the true sample log-likelihood - E[logD(I)] real [,y)] and the log-likelihood of the generated samples The specific loss function formula is as follows:
[0075]
[0076] Among them, I real This indicates that the image is indeed of a blueberry. Let G be the simulated image output by the generator, y be the growth cycle label, D(.,.) be the discriminator function, and E[.] be the mathematical expectation operator. By alternately optimizing G and D, the distribution of the generated samples is made closer to the real data.
[0077] S3. Construct a YOLO detection model based on the preprocessed multi-dimensional image, including adding a cross-period feature memory module after the backbone network and calculating the importance score of each channel of the backbone network.
[0078] The traditional YOLO model is composed of three parts: the backbone network, the neck network, and the head network. This embodiment improves upon the traditional YOLO model to generate a YOLO detection model. The preprocessed blueberry image is used as input to construct the YOLO detection model. The core objectives of this YOLO detection model are to "enhance feature representation capabilities, optimize computational efficiency, and improve multi-scale detection accuracy". It is formed by three major modules: backbone network feature extraction optimization, neck network cross-scale fusion enhancement, and detection head multi-scale adaptation design. Combined with cross-period feature memory and dynamic channel pruning, a complete detection architecture is formed.
[0079] CBAM is embedded in the backbone network based on the CSPDarknet architecture to enhance target features and suppress background noise through channel attention and spatial attention. This includes adding channel attention submodules and spatial attention submodules. The channel attention submodule extracts global information in the channel dimension through global average pooling and global max pooling, processes it through a multilayer perceptron (MLP), sums the results, and then generates a channel attention map through sigmoid activation.
[0080] M c (X)=sigmoid(MLP(AvgPool(X))+MLP(MaxPool(X))),
[0081] Where AvgPool represents the global average pooling operation, MaxPool represents the global max pooling operation, MLP represents the multilayer perceptron, and sigmoid represents the activation function. The spatial attention submodule performs AvgPool and MaxPool operations on the feature map respectively, concatenates them along the channel dimension, processes them through a 7×7 convolutional layer, and generates spatial attention through sigmoid activation.
[0082] M s (X)=sigmoid(f 7×7 ([AvgPool(X);MaxPool(X)])),
[0083] Among them, f 7×7 Using 7×7 convolutions, CBAM makes the feature maps output by the backbone network more prominent in highlighting the key features of blueberry organs.
[0084] CFCM is added after the backbone network output, and the feature map sequence of 6 consecutive frames (T=6) is fused by 3D convolution [F]. t-5 ,F t-4 ,…,F t The formula is:
[0085] F T =3D-Conv([F t-5 ,Ft-4 ,…,F t ]),
[0086] Where 3D-Conv represents a 3D convolution operation, [F t-5 ,F t-4 ,…,F t ] represents the input feature sequence, each F i It is a four-dimensional tensor that inputs the temporal features output by CFCM into the neck network, enabling the model to predict the growth stage.
[0087] Calculate the importance score ω of each channel in the backbone network. c The channel attention map M of the CBAM output c The formula for calculating the mean along the spatial dimension is:
[0088]
[0089] Where, ω c This is represented by the importance score of the c-th channel in the backbone network, where H and W represent the height and width of the feature map, respectively. Represented by a double summation symbol, this indicates that the summation is performed over all pixels in the spatial dimension of the feature map. This represents the pixel value of the c-th channel in the channel attention map, with coordinates (h, w). The higher the score of a channel, the greater its contribution to the target feature.
[0090] Then, based on the proportion of mature samples in the current batch Dynamically adjust the channel retention ratio r:
[0091]
[0092] Here, count(y=4) represents the number of samples labeled 4 (maturity) in the current batch. When the proportion of mature samples is >50%, r>0.7, the weight of the blue channel (corresponding to the color of mature fruit) is increased, and redundant channels related to green (leaves) are pruned. When the proportion of budding samples is high, r<0.7, the green channel (light green buds) is retained, and the calculation of irrelevant channels is reduced.
[0093] Finally, an LSTM unit is introduced into BiFPN to process the current frame features. Output features of the previous frame Through LSTM interaction, and then weighted fusion with features from other scales:
[0094]
[0095] in, This is represented as the LSTM output feature of the current frame. This represents the original feature input of frame t in the bidirectional feature pyramid network. This represents the output features of the previous frame's LSTM, where LSTM(.) represents the LSTM function operation, and ω... i Let ω represent the importance weight of the i-th input feature, where ∈ denotes the minimum constant. j Let represent the importance weight of the j-th input feature. BiFPN fuses features through bidirectional cross-scale connections and normalized weights. Each level of BiFPN includes two paths: a bottom-up path to supplement detailed features and a top-down path to facilitate cross-level fusion, allowing high-level semantic information from higher layers to be conveyed to lower layers, thus improving multi-scale adaptability. The feature fusion expression is as follows:
[0096]
[0097] like Figure 2 As shown, for the feature map output by CBAM, the DySample upsampler predicts the sampling point offset Δp through a convolutional branch, causing the sampling kernel to move in the direction of the maximum feature gradient (such as the high gradient region at the tip of a petal), accurately capturing the detailed features enhanced by CBAM. An additional convolutional branch predicts the sampling point offset Δp, dynamically adjusting the sampling kernel position.
[0098] Δp=Conv(F),k(x,y)=I(p+Δp)·w(x,y),
[0099] Wherein, Conv(F) represents applying a convolution operation to the input feature map F, I(p+Δp) represents extracting the feature value at the offset position (p+Δp), and w(x,y) represents the sampling kernel weight, which is used to weight the feature values of the sampling points. In regions with high curvature, such as the tips of petals during the flowering period, Δp causes the sampling points to shift towards the direction of the maximum gradient, preserving the jagged edge details. Finally, three detection heads with different resolutions are constructed: a high-resolution detection head that receives the low-level features of the backbone network and combines them with the edge details enhanced by CBAM to detect small targets in the budding stage; a medium-resolution detection head that is used for detecting medium-sized targets in the young fruit stage; and a low-resolution detection head that receives the high-level features of the backbone network and uses the semantic information enhanced by CBAM to detect large targets in the mature stage.
[0100] S4. Perform multi-task joint training on the completed YOLO detection model, including constructing a cross-modal mapping matrix of RGB features and NIR features, and aligning similar features in the embedding space by using a contrastive loss function.
[0101] Based on the constructed YOLO detection model, this study further enhances the model's understanding of the entire blueberry growth cycle through cross-modal feature alignment, growth cycle remaining time prediction, and temporal logic constraints. In blueberry growth status detection, RGB images contain visual information such as color and shape, while NIR images reflect physiological indicators such as chlorophyll content and sugar content. However, the feature distributions of the two modalities differ, necessitating feature alignment through cross-modal mapping. A parameterizable mapping matrix W1 is learned through training to map the RGB feature vectors f... rgb Mapped to NIR eigenvector f nir The same embedding space, i.e., f' rgb =f rgb W1 calculates the similarity between the mapped RGB features and NIR features using cosine similarity: A contrastive loss function is calculated based on similarity. This contrastive loss function employs a contrastive learning mechanism, and the formula is as follows:
[0102]
[0103] Among them, f rgb f is represented as a feature vector extracted from an RGB image via a backbone network. nir Let f' be the feature vector extracted from the NIR image. nir Let f represent the negative sample NIR features, W1 represent the cross-modal mapping matrix, and τ represent the temperature parameter, which adjusts the compactness of the feature distribution. The smaller τ is, the more clustered similar features are in the embedding space, and the higher the model's sensitivity to feature differences. In each training iteration, the specific process is as follows: sample RGB-NIR image pairs from the dataset, and extract features f through the backbone network. rgb and f nir Calculate the contrast loss L con And backpropagate to update the W1 parameters, so that similar features are aligned in the embedding space.
[0104] The subsequent prediction algorithm, based on quantile regression, aims to predict the remaining growth cycle time at different quantiles. The loss function is a weighted sum of multi-quantile losses. Where Q is the set of quantiles, α q L represents the weight of the loss for each quantile. q The loss function for the q-th quantile is... Where, ρ q Let y be the quantile loss function. i This represents the actual remaining time. Given the predicted value of the i-th sample at the q-th quantile, the input is the sample's feature vector (including growth cycle labels, image features, etc.). A quantile regression model is used to predict the remaining time at different quantiles, and the multiquantile loss L is calculated.qreg And then backpropagate to update the model parameters.
[0105] The growth cycle is mapped to an ordered integer ord(y) = {1, 2, 3, 4}, where 1, 2, 3, and 4 correspond to the budding stage, flowering stage, fruiting stage, and maturity stage, respectively, ensuring that the encoding is monotonically increasing. By calculating the mean absolute difference between the encodings of adjacent frame tags, reverse order cases are penalized (e.g., the encoding difference between "flowering stage → budding stage" is -1), forcing the model to output results that conform to the growth pattern. The temporal loss function formula is:
[0106] L seq =E[∑ t |ord(y t )-ord(y t-1 )|],
[0107] Where E represents the expected value, which is the average of the temporal losses over a batch of samples to avoid the influence of a single sample anomaly on the overall optimization direction, |ord(y t )-ord(y t-1 )| represents the absolute difference between the sequential encoding of the current frame and the previous frame, ∑ t This is expressed as the sum of the absolute differences of all adjacent frames in the time series. In actual training, the loss functions of the three tasks mentioned above are weighted and summed to obtain the total loss function L:
[0108] L=λ1L con +λ2L qreg +λ3L seq ,
[0109] Where λ1, λ2, and λ3 are different weighting coefficients, and the cross-modal contrastive loss L con With timing loss L seq With lower weights to avoid dominating the training direction, backpropagation is performed based on the total loss L during each training iteration. At the same time, the parameters of the cross-modal mapping matrix, quantile regression model and YOLO detection model are optimized to achieve multi-task collaborative learning.
[0110] S5. Optimize edge-end inference based on the trained YOLO detection model, including updating the model through feature space alignment loss and EWC regularization;
[0111] The improved YOLO detection model includes three detection heads with resolutions of 80×80×128 (small targets), 40×40×256 (medium targets), and 20×20×512 (large targets), corresponding to the output prediction probability vector p for each period. h = [p1, p2, p3, p4], where h is the index of the detection head. For each detection head h, calculate its maximum prediction confidence max(p h), when max(p h If the maximum value of the detection head is greater than 0.3, it is determined that the detection head is effective for the current scene and the calculation is retained; otherwise, it is turned off to reduce computing power consumption. If the maximum value of the mature detection head is greater than 0.3, the calculation is retained. h If the maximum value of the target is greater than 0.7, the current scene is determined to be dominated by large targets, and the small target detection head is turned off. If the maximum value of the target detection head in the nascent stage is greater than 0.7, the small target detection head is turned off. h If the value is greater than 0.5, it is determined that a small target exists. All branches are retained, but the computational priority of the large target detection head is reduced. While the invalid detection head is closed, the freed computing power is used for the super-resolution processing of the valid detection head.
[0112] When a new sample is input, it is processed through the base model θ base The new model θ extracts features f(x; θ) base f(x; θ) and f(x; θ) are kept close by an alignment loss:
[0113]
[0114] This loss allows the new model to retain the feature representation capability of the base model, and each parameter θ is calculated. i The Fisher information matrix is used to quantify the importance of parameters to the old task. The parameter updates are constrained by a regularization term. Key parameters (such as those related to real color recognition) have larger Fisher information matrices and are strongly constrained during updates; non-key parameters (such as background feature parameters) have smaller Fisher information matrices and can be updated freely. Simultaneously, the classification loss, alignment loss, and EWC loss are optimized. The update formula is:
[0115]
[0116] Where, θ t L represents the parameters of the current model. cls Let L represent the classification loss function, η represent the learning rate, and L represent the classification loss function. align L represents the feature space alignment loss. ewc It is represented as the elastic weight consolidation loss, and λ1 and λ2 represent different weight coefficients.
[0117] Finally, for each frame image I t Background model B is constructed using the MOG algorithm. t Each pixel is modeled as a mixture of K=3 Gaussian distributions, and detection is performed only on moving regions. The formula is:
[0118] B t (x,y)=MOG(B t-1 ,I t ),
[0119] Among them, B t-1 Represented as the background model of the previous frame, It Represented as the current frame image, when a new frame is input, the pixel values are matched with the existing Gaussian distribution, and the mean, variance, and weights are updated. If there is a mismatch, a new distribution is created or the old distribution is replaced. For each pixel, its mean and standard deviation are calculated as the threshold benchmark for inter-frame differencing. The current frame I is then calculated. t With background model B t If the absolute value of the pixel difference is greater than 3 times the standard deviation, it is determined to be a moving region. Only the moving region is subjected to the complete detection process (feature extraction + prediction). The background region directly reuses the detection results of the previous frame.
[0120] S6. Construct a multi-factor risk assessment function based on the output of the YOLO detection model to complete the final blueberry growth status detection;
[0121] The growth cycle label is extracted from the output of the YOLO detection model, which represents the current blueberry growth stage detected by the model (budding, flowering, fruiting, and ripening). Remaining time prediction is based on the remaining days of the growth cycle and target location information output by the quantile regression model, along with the coordinates and size of the detection box. This information is used to calculate spatial features such as fruit quantity and distribution density. Then, the actual growth cycle is compared with the standard growth cycle for the same variety and region to assess whether the current stage is lagging or ahead of schedule. For example, if the model predicts "budding," but local phenological data indicates "flowering," then the growth is considered lagging; the greater the deviation, the higher the risk value. A formula can be used to calculate this. The risk factor for deviation from the growth cycle is represented by a quantile regression prediction of the remaining time confidence interval. The stability of the growth progress is evaluated to obtain the remaining time fluctuation risk factor. The neutralization risk index is calculated by weighted summation of the growth cycle deviation risk and the remaining time fluctuation risk. The detection results and risk index are output to realize the detection of the entire growth cycle of blueberry fruit.
[0122] Example 2
[0123] The difference between this embodiment and Embodiment 1 is that this embodiment provides a blueberry full growth cycle detection system based on an improved YOLO model, including:
[0124] The data acquisition module is configured to: collect multi-dimensional image data of blueberries at different growth stages and annotate the acquired multi-dimensional image data;
[0125] The preprocessing module is configured to perform preprocessing on the labeled multi-dimensional images, including semantic segmentation of the image data using a lightweight semantic segmentation model and multi-scale dynamic input based on different annotated images;
[0126] The model module is configured to: construct a YOLO detection model based on the preprocessed multi-dimensional image, including adding a cross-period feature memory module after the backbone network and calculating the importance score of each channel of the backbone network;
[0127] The training module is configured to perform multi-task joint training on the constructed YOLO detection model, including constructing a cross-modal mapping matrix of RGB features and NIR features, and aligning similar features in the embedding space through a contrastive loss function.
[0128] The optimization module is configured to perform edge-end inference optimization based on the trained YOLO detection model, including model updates through feature space alignment loss and EWC regularization.
[0129] The output module is configured to construct a multi-factor risk assessment function based on the output results of the YOLO detection model, and complete the final blueberry growth status detection.
[0130] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for detecting the entire growth cycle of blueberries based on an improved YOLO model, characterized in that, include: Collect multi-dimensional image data of blueberries at different growth stages and annotate the acquired multi-dimensional image data; Preprocessing is performed on the labeled multi-dimensional images, including semantic segmentation of the image data using a lightweight semantic segmentation model, and multi-scale dynamic input based on different annotated images; A YOLO detection model is constructed based on the preprocessed multi-dimensional images, including adding a cross-period feature memory module after the backbone network and calculating the importance score of each channel of the backbone network. The completed YOLO detection model is jointly trained on multiple tasks, including constructing a cross-modal mapping matrix of RGB features and NIR features, and aligning similar features in the embedding space by using a contrastive loss function. Edge-end inference optimization is performed based on the trained YOLO detection model, including model updates through feature space alignment loss and EWC regularization; A multi-factor risk assessment function was constructed based on the output of the YOLO detection model to complete the final detection of blueberry growth status.
2. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 1, characterized in that, The process of semantic segmenting image data using a lightweight semantic segmentation model includes: firstly, extracting a mask matrix of blueberry organs using the lightweight semantic segmentation model; classifying regions with mask values greater than a threshold as Regions of Interest (ROIs) and the remaining regions as background regions; applying bicubic interpolation to ROIs and nearest-neighbor interpolation downsampling to background regions; standardizing image size; and adaptively selecting the input resolution based on the average target size of different labeled images. The expression for determining different regions is as follows: Where I represents the original input image, M represents the mask matrix output by the lightweight semantic segmentation model, ⊙ represents element-wise multiplication, Bicubic represents bicubic interpolation, and Nearest represents nearest neighbor interpolation.
3. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 2, characterized in that, The preprocessing of the labeled multi-dimensional images further includes constructing a conditional generative adversarial network (GAN). The generator uses a U-Net architecture, taking as input real images after ROI processing and multi-scale input, along with growth cycle labels. It generates enhanced samples simulating different lighting and occlusion conditions via an encoder-decoder path. The discriminator uses a PatchGAN structure, and the generator and discriminator are alternately optimized using an adversarial loss function. The formula for the adversarial loss function is: Among them, I real This indicates that the image is indeed of a blueberry. Let represent the simulated image output by the generator, y represent the growth cycle label, D(.,.) represent the discriminator function, and E[.] represent the mathematical expectation operator.
4. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 1, characterized in that, The YOLO detection model constructed based on preprocessed multi-dimensional images includes embedding a CBAM attention mechanism into a backbone network based on the CSPDarknet architecture. This enhances target features through channel attention and spatial attention. A cross-period feature memory module is then added after the backbone network output. A sequence of feature maps from T consecutive frames is fused using 3D convolution to obtain temporal features, which are then integrated into the neck network. The importance score of each channel in the backbone network is then calculated. The average of the channel attention maps output by CBAM is calculated in the spatial dimension. The channel retention ratio is dynamically adjusted based on the proportion of mature samples in the current batch. The importance score formula is as follows: Where, ω c This is represented by the importance score of the c-th channel in the backbone network, where H and W represent the height and width of the feature map, respectively. Represented by a double summation symbol, this indicates that the summation is performed over all pixels in the spatial dimension of the feature map. This represents the pixel value of the c-th channel in the channel attention map, with coordinates (h, w).
5. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 4, characterized in that, The YOLO detection model built based on the preprocessed multi-dimensional image also includes introducing LSTM temporal connections into the bidirectional feature pyramid network of the neck network. The features of the current frame are interacted with the output features of the previous frame through LSTM to construct a spatiotemporal feature fusion pyramid. Additional convolutional branches are used to predict sampling point offsets, dynamically adjusting the sampling kernel position. Finally, three detection heads with different resolutions are established. The LSTM interaction formula is as follows: in, This is represented as the LSTM output feature of the current frame. This represents the original feature input of frame t in the bidirectional feature pyramid network. This represents the output features of the previous frame's LSTM, where LSTM(.) represents the LSTM function operation, and ω... i Let ω represent the importance weight of the i-th input feature, where ∈ denotes the minimum constant. j It is represented as the importance weight of the j-th input feature.
6. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 1, characterized in that, The multi-task joint training of the constructed YOLO detection model includes: predicting the remaining time of the blueberry growth cycle using quantile regression, with the loss function being a weighted sum of multiple quantile losses; constructing a cross-modal mapping matrix W for RGB and NIR features; aligning similar features in the embedding space using a contrastive loss function; defining the sequential encoding of the growth cycle; and forcing adjacent frame labels to satisfy the growth order using a temporal loss function. The formula for the contrastive loss function is as follows: Among them, f rgb f is represented as a feature vector extracted from an RGB image via a backbone network. nir Let f' be the feature vector extracted from the NIR image. nir It represents the negative sample NIR features, W1 represents the cross-modal mapping matrix, and τ represents the temperature parameter.
7. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 1, characterized in that, The edge-end inference optimization based on the trained YOLO detection model includes calculating the maximum confidence score of each detection head in the predicted probability vector for each growth cycle of blueberries. When the maximum confidence score is greater than a set confidence threshold, the branch calculation for that detection head is retained. The iCaRL algorithm is used, and online model updates are achieved through feature space alignment loss and EWC regularization. The update formula is as follows: Where, θ t L represents the parameters of the current model. cls Let L represent the classification loss function, η represent the learning rate, and L represent the classification loss function. align L represents the feature space alignment loss. ewc It is represented as the elastic weight consolidation loss, and λ1 and λ2 represent different weight coefficients.
8. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 7, characterized in that, The edge-end inference optimization based on the trained YOLO detection model also includes building a background model for each frame using the MOG algorithm, calculating the absolute value of the pixel difference between the current frame and the background model, locating motion regions through inter-frame difference, and performing detection on the motion regions. The formula for the background model is: B t (x,y)=MOG(B t-1 ,I t ), Among them, B t-1 Represented as the background model of the previous frame, I t This represents the current frame image.
9. The method for detecting the entire growth cycle of blueberries based on an improved YOLO model according to claim 1, characterized in that, The multi-factor risk assessment function is constructed based on the output results of the YOLO detection model. This includes extracting the growth cycle label, remaining time prediction, and target positioning information output by the YOLO model, constructing a multi-factor risk assessment function, including the growth cycle deviation risk factor and the remaining time fluctuation risk factor, calculating a comprehensive risk index based on the growth cycle deviation risk factor and the remaining time fluctuation risk factor, and finally outputting the growth cycle classification.
10. A blueberry full growth cycle detection system based on an improved YOLO model, comprising the method described in claim 1, characterized in that, include: The data acquisition module is configured to: collect multi-dimensional image data of blueberries at different growth stages and annotate the acquired multi-dimensional image data; The preprocessing module is configured to perform preprocessing on the labeled multi-dimensional images, including semantic segmentation of the image data using a lightweight semantic segmentation model and multi-scale dynamic input based on different annotated images; The model module is configured to: construct a YOLO detection model based on the preprocessed multi-dimensional image, including adding a cross-period feature memory module after the backbone network and calculating the importance score of each channel of the backbone network; The training module is configured to perform multi-task joint training on the constructed YOLO detection model, including constructing a cross-modal mapping matrix of RGB features and NIR features, and aligning similar features in the embedding space through a contrastive loss function. The optimization module is configured to perform edge-end inference optimization based on the trained YOLO detection model, including model updates through feature space alignment loss and EWC regularization. The output module is configured to construct a multi-factor risk assessment function based on the output results of the YOLO detection model, and complete the final blueberry growth status detection.
Citation Information
Cited By
Tongue picture image attitude correction and integrity detection method and system based on key point detection
CN121259061A
Evaluation method, device and equipment for homework correction system
CN121904779A