High-precision identification method for fruit calyx / stem and surface defects with enhanced polarization characteristic difference
By acquiring multi-angle polarization images and combining them with median filtering and histogram equalization, a polarization characterization parameter fusion algorithm was constructed. Combined with visible light images, the YOLO-PVF multimodal detection model was used to solve the problem of high-precision identification of calyx/stalk and surface defects in fruit detection, achieving efficient detection results.
Patent Information
- Application Number
- CN202511001153.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-07-21
Smart Images

Figure CN120766274B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fruit image detection technology, and in particular to a high-precision identification method for fruit calyx / pedicel and surface defects that enhances polarization characteristic differences. Background Technology
[0002] Since 1994, global fruit production has shown a continuous upward trend. China, as one of the world's major fruit producers, produced 327 million tons of fruit in 2023. During harvesting and transportation, surface defects significantly shorten the shelf life of pears, making them more susceptible to bacterial and fungal infections that lead to spoilage and affecting surrounding fruit, causing cross-contamination. Furthermore, external defects causing fruit rot, quality decline, and skin shrinkage significantly reduce their quality and commercial value, resulting in substantial economic losses to the fruit industry. Therefore, improving fruit defect detection technology is of great significance for improving the quality and added value of commercial fruits.
[0003] The fruit stalk / calyx area is usually concave, making it difficult for light to reflect. This results in a low-grayscale area on the fruit image, causing the fruit stalk / calyx to appear highly similar in color to severe defects on the fruit surface, making it difficult to distinguish between defects and the fruit stalk / calyx during inspection. Therefore, accurate identification of the fruit stalk / calyx is crucial for improving the accuracy of fruit defect detection.
[0004] For fruits like apples and pears, where the difference between calyx / pedicel and defects is relatively small, previous studies have attempted to achieve accurate detection using novel imaging technologies such as laser scattering imaging and near-infrared imaging. Zhang et al. (2017) used near-infrared (NIR) coded structured light to analyze the deformation and distribution of NIR lattice projected onto apples, aiming to enhance the geometric and textural features of the apple surface. However, this technique suffers from low efficiency and is difficult to apply in actual production.
[0005] Polarization imaging is an optical imaging method that rapidly analyzes the physical properties of a target object's surface by measuring the polarization state distribution of light waves. Current polarization information processing methods are mainly geared towards target recognition in large outdoor scenes. For example, multi-polarization information fusion algorithms can effectively detect targets concealed in low-light grass and flying objects obscured by clouds. In such scenes, the target-background heterogeneity is significant, resulting in high image contrast after polarization information processing. In contrast, calyx / pedicel and defects, as part of the fruit, have similar tissue structures, leading to lower contrast between calyx / pedicel and defects in the polarization information image of Crown Pear, and the problem of effective polarization information accumulation.
[0006] Therefore, based on polarization imaging technology, existing technologies lack a method for accurately identifying fruit stalks and calyxes that balances high precision and high speed. Summary of the Invention
[0007] To address the problem that existing technologies often misidentify calyxes and pedicels as defects, this invention designs a high-precision identification method for fruit calyxes / pedicels and surface defects that enhances differences in polarization characteristics, balancing high accuracy and high speed. Specifically, it includes a polarization characterization parameter fusion method that enhances the differences in polarization characteristics of fruit calyxes / pedicels and surface defects, and a multimodal calyx, pedicel, and defect detection method based on polarization parameter fusion images and visible light images.
[0008] The method of this invention mainly includes: 1) calculating the polarization degree image and horizontal polarization enhancement image of the calyx, fruit stalk and defect area from the four acquired polarization images; 2) preprocessing the polarization degree image and horizontal polarization enhancement image using median filtering and histogram equalization, then constructing a polarization characterization parameter fusion algorithm and outputting a polarization parameter fusion image; 3) establishing a multimodal model for detecting the calyx, fruit stalk and defect area based on the polarization parameter fusion image and the visible light image.
[0009] More specifically, the technical solution adopted by the present invention to solve its technical problem is as follows:
[0010] 1) Characterization of polarization parameters of fruits:
[0011] An initial dataset DATA containing fruit polarization information was constructed by collecting images of numerous fruits at different polarization angles and processing them. Then, a subset of images was randomly selected from the initial dataset DATA to form a dataset DATA representing the polarization parameters of the fruits. DOP+IE ;
[0012] 2) Fusion of fruit polarization characterization parameters:
[0013] Data set of fruit polarization characterization parameters DOP+IEAnnotation is performed to obtain masks for fruit stalks / calyxes and surface defects. A FusionNet network model for fusing fruit polarization characterization parameters is then constructed, with the fruit polarization characterization parameter dataset DATA as input. DOP+IE The optimal fusion model is obtained by training a loss function using a mask design. The initial dataset DATA is then processed using the optimal fusion model to obtain a polarization parameter fused image (IMG). Fusion ;
[0014] 3) Multimodal detection of fruit calyxes, pedicels, and their defects:
[0015] Using the results from steps 1) and 2), a dataset DATA for detecting defects in the fruit stalk / calyx and surrounding area is constructed. Fusion+90 Based on the DATA dataset for detecting defects in the fruit stalk / calyx and surrounding area Fusion+90 A multimodal detection model, YOLO-PVF, for fruit stalks / calyxes and defects was constructed and trained to obtain the optimal detection model. Finally, the optimal detection model was used to detect the calyxes / stalks and surface defects of fruits.
[0016] Step 1) specifically refers to:
[0017] For each fruit, a split-focus plane polarization camera was used to acquire visible light images at linear polarization angles of 0°, 45°, 90°, and 135°, resulting in four polarized visible light images P0, P1, and P2 at different polarization angles. 45 P 90 P 135 Using the grayscale images I0, I1, and I2 of four polarized visible light images 45 I 90 I 135 The combined processing yields the horizontal polarization enhancement image IE0 and the polarization degree image DOP0 of the current fruit.
[0018] Then, median filtering is applied to obtain the denoised horizontal polarization enhanced image IE1 and polarization degree image DOP1 of the fruit. Next, histogram equalization is applied to the denoised horizontal polarization enhanced image IE1 and polarization degree image DOP1 to obtain the optimized horizontal polarization enhanced image IE2 and polarization degree image DOP2 of the fruit. Finally, a fruit polarization characterization parameter dataset DATA is constructed from the optimized horizontal polarization enhanced images IE2 and polarization degree images DOP2 of all the fruits. DOP+IE .
[0019] Step 1) includes the following steps:
[0020] 1.1) Polarized visible light image acquisition:
[0021] Four polarized visible light images P0, P1, P2, and P3 were acquired using a focal plane polarization camera at polarization angles of 0°, 45°, 90°, and 135°. 45 P90 P 135 ;
[0022] 1.2) Obtaining the grayscale image of the polarization angle:
[0023] Four polarized visible light images P0, P1, P2, P3, P4, P5, P6, P7, P8, P9, P10, P11, P12, P13, P14, P15, P16, P17, P18, P1 ... 45 P 90 P 135 Their respective grayscale images I0, I 45 I 90 I 135 The formula is as follows:
[0024]
[0025] Where R, G, and B represent the R-channel component image, G-channel component image, and B-channel component image of the polarization image, respectively; Gray represents the grayscale image.
[0026] 1.3) Acquisition of polarization degree image and horizontal polarization enhancement image:
[0027] The following formulas are used based on the grayscale images I0, I... of four polarized visible light images. 45 I 90 I 135 The horizontal polarization enhancement image IE0 and polarization degree image DOP0 of the current fruit are calculated as follows:
[0028] S0=I0+I 90
[0029] S1=I0-I 90
[0030] S2=I 45 -I 135
[0031]
[0032] Where S0, S1, and S2 represent the three vector parameters of Stoksk, and I0, I... 45 I 90 I 135 These represent the grayscale images of four polarized visible light images, with DOP0 representing the polarization degree image and IE0 representing the horizontal polarization enhancement image.
[0033] 1.4) Denoising the horizontal polarization enhancement image IE0 and the polarization degree image DOP0: The horizontal polarization enhancement image IE0 and the polarization degree image DOP0 are blurred by median filtering according to the following formula, and the denoised horizontal polarization enhancement image IE1 and polarization degree image DOP1 are respectively used as the horizontal polarization enhancement denoised image IE1 and the polarization degree denoised image DOP1.
[0034]
[0035] Where g(i,j) represents the pixel value after median filtering; f(i,j) represents the pixel value before median filtering; k represents the window radius parameter, usually set to 1; median{} represents the median operation; i and j represent the x and y coordinates of the image; m and n represent the independent values from the set {-k,...,k}. 2 Take the value from;
[0036] 1.5) Histogram equalization of horizontal polarization enhanced denoised image IE1 and polarization degree denoised image DOP1: The following formula is used to transform each image of horizontal polarization enhanced denoised image IE1 and polarization degree denoised image DOP1 to obtain optimized horizontal polarization enhanced image IE2 and polarization degree image DOP2.
[0037] I out (i,j)=I in (i,j) γ
[0038] Among them, I out I represents the pixel value after image transformation. in This represents the pixel value before image transformation processing, and γ represents the power-law transformation parameter. In the specific implementation, γ is set to 0.8.
[0039] 1.6) Constructing the initial dataset and the polarization characterization parameter dataset:
[0040] Each fruit image is paired with an optimized horizontal polarization-enhanced image IE2 and a polarization degree image DOP2. All fruit image pairs constitute the initial dataset DATA. From all processed image pairs in the initial dataset DATA, a subset of image pairs are randomly selected and processed using one of the following methods: rotation, flipping, scaling, etc., to achieve offline data enhancement, thereby constructing the polarization characterization parameter dataset DATA. DOP+IE Each image pair in the polarization characterization parameter dataset consists of a polarization degree image DOP2 and an aligned horizontal polarization enhancement image IE2.
[0041] Step 2) specifically refers to:
[0042] Data set of fruit polarization characterization parameters DOP+IE The horizontal polarization-enhanced image IE2 and the degree of polarization image DOP2 of each image pair were annotated with the fruit stalk / calyx region and the defect region using software to obtain the fruit stalk / calyx mask M. s / c With defect mask M defectSpecifically, the fruit stalk / calyx region and the defect region are marked in the horizontally polarized enhanced image IE2 and the degree of polarization image DOP2, and the two images are aligned.
[0043] Construct a FusionNet network model for fusing fruit polarization characterization parameters and input the fruit polarization characterization parameter dataset DATA obtained in step 1). DOP+IE Horizontal polarization enhancement image IE2 and polarization degree image DOP2, using fruit stalk / calyx mask M s / c With defect mask M defect The loss function Loss is designed to train and validate the fruit polarization representation parameter fusion network model FusionNet to obtain the optimal fusion model. Then, the optimal fusion model is used to process each pair of images in the initial dataset DATA to output the polarization parameter fused image IMG for each fruit. Fusion .
[0044] Step 2) includes the following steps:
[0045] 2.1) Polarization characterization parameter dataset annotation:
[0046] Using annotation software to analyze the fruit polarization characterization parameter dataset DATA DOP+IE In each image pair, the two images are labeled with the fruit stalk / calyx and surface defects, respectively, to obtain the fruit stalk / calyx mask M. s / c and defect mask M defect ;
[0047] The term "fruit stalk / calyx" refers to one of the fruit stalk and the calyx; usually, only one of the fruit stalk and the calyx can be captured in an image.
[0048] 2.2) Constructing FusionNet, a network model for fusing fruit polarization characterization parameters:
[0049] The fruit polarization characterization parameter fusion network model FusionNet includes two polarization feature extraction modules. FE A polarization feature reconstruction module FR Two polarization feature extraction modules FE The two modules have the same topology but different internal parameters, and are used to extract features from the polarization degree image DOP2 and the horizontal polarization enhancement image IE2, respectively. Specifically, they include two parallel modules MODULE with the same network structure but independent internal parameters. FE-DOP and MODULE FE-IE Features were extracted from the polarization degree image DOP2 and the horizontal polarization enhancement image IE2, respectively.
[0050] 2.3) The fruit polarization characterization parameter dataset DATA DOP+IEThe horizontal polarization enhancement image IE2 and the polarization degree image DOP2 of each image pair are respectively input into two polarization feature extraction modules MODULE. FE The respective convolutional features are obtained, namely, horizontal polarization enhanced convolutional features. FE-DOP Convolutional features with polarization degree FE-IE ;
[0051] Then, the horizontal polarization-enhanced convolution feature is applied using the following formula. FE-DOP Convolutional features with polarization degree FE-IE The two convolutional features are concatenated along the channel dimension to obtain the sum feature:
[0052] F eature =Concat(Feature) FE-DOP Feature FE-IE )
[0053] Wherein, Concat represents a concatenation operation at the channel level;
[0054] Finally, the summation feature is input into the polarization feature reconstruction module. FR The polarization parameter fusion image IMG corresponding to the horizontal polarization enhancement image IE2 and the polarization degree image DOP2 is obtained. Fusion .
[0055] During the model training phase, the fruit polarization characterization parameter dataset DATA DOP+IE Image pairs (polarization degree image DOP2 and horizontal polarization enhancement image IE2) are input into the FusionNet model to obtain the optimal fusion model. Then, the optimal fusion model is used to process each image pair in the initial dataset DATA to output a polarization parameter fused image IMG for each fruit. Fusion .
[0056] In the FusionNet network model for fusing fruit polarization characterization parameters in step 2.2):
[0057] Each of the polarization feature extraction modules FE It mainly consists of four modules connected in sequence: a convolution module, a primary residual module (PRM), an intermediate residual module (IRM), and a high-level residual module (ARM); the polarization feature reconstruction module (MODULE) FR It mainly consists of four modules connected in sequence: the first residual module RM1, the second residual module RM2, the third residual module RM3, and the fourth residual module RM4; the polarization feature extraction module MODULE FEThe convolution module in the code mainly consists of a 5×5 convolution operation and a Leaky ReLU activation function connected in sequence.
[0058] Each residual module mainly consists of three convolutional layers and one convolutional skip connection. The input to the residual module is fed into the three convolutional layers and the convolutional skip connection in sequence. The outputs of the three convolutional layers and the output of the convolutional skip connection are added together, passed through an activation function, and then output as the output of the residual module. The first two convolutional layers are mainly composed of a convolution operation, an activation function, and a batch normalization operation connected in sequence, while the last convolutional layer consists of only a single convolution operation. The convolutional skip connection uses only one convolution operation.
[0059] The polarization feature reconstruction module MODULE FR After the fourth residual module RM4, the output is mapped to the 0-1 range using the following formula:
[0060]
[0061] Where x represents the pixel value of the feature map output by the fourth residual module RM4, tanh() represents the Tanh activation function operation, and f(x) represents the pixel value finally output by the polarization feature extraction module.
[0062] In practice, the fruit polarization characterization parameter dataset can be divided into two parts, with the datasets arranged in a 7:2 ratio. DOP+IE The dataset is divided into a training set and a validation set; the polarization characterization parameter dataset DATA is... DOP+IE The training and validation sets are input into the FusionNet model for training to obtain the optimal FusionNet model. Then, the optimal fusion model is used to process each image pair in the initial dataset DATA to output a fused image (IMG) of the polarization parameters of each fruit. Fusion .
[0063] Loss function design for FusionNet, a polarization characterization parameter fusion algorithm:
[0064] In step 2), the loss function Loss will be combined with the fruit stalk / calyx mask M. s / c With defect mask M defect Set it up as follows: For each pair of images, respectively set the fruit stalk / calyx mask M. s / c The true value GT of the fruit stalk / calyx region fusion target is obtained by multiplying the polarization degree image DOP2 pixel by pixel. s / c Defect Mask M defect The true value GT of the defect region fusion target is obtained by multiplying the horizontally polarized enhanced image IE2 pixel by pixel. defect Then, by separately masking the fruit stalk / calyx Ms / c and defect mask M defect IMG images fused with polarization parameters respectively Fusion Pixel-by-pixel multiplication yields the fused fruit stalk / calyx region images. s / c Fusion of defect area images defect ;
[0065] The loss function includes the pedicel / calyx region and the defect region, and the pixel loss L is calculated separately according to the following method. pixel With gradient loss L pixel :
[0066] (a) Pixel loss:
[0067] Based on the true value GT of the fruit stalk / calyx region fusion target s / c The true value GT of the defect region fusion target defect Image Fusion of Fruit Stalk / Cepal Region s / c Fusion of defect area images defect The pixel loss L in the fruit stalk / calyx region and defect region is calculated using the following formula. pixel_s / c L pixel_defect Finally, the weighted sums are used to obtain the final pixel loss L. pixel ;
[0068]
[0069] L pixel =L pixel_s / c +α×L pixel_defect
[0070] Where H and W represent the image length and width, respectively, |||1 represents the L1 norm calculation, and α represents the weighting coefficient, which is 1.5.
[0071] (b) Gradient Loss: The gradient loss L for the fruit stalk / calyx region and the defect region is calculated using the following formulas respectively. grad_s / c L grad_defect Finally, the weighted sums are used to obtain the final gradient loss L. pixel ;
[0072]
[0073] L grad =L grad-s / c +β×L grad-defect (16)
[0074] Where H and W represent the length and width of the image, respectively; The gradient operator is represented by β, which can be used to calculate the gradient of the image using the Sobel operator; β represents the weight coefficient; here, the weight coefficient β is still taken as 1.5.
[0075] The final pixel loss L is obtained by weighted calculation of the losses in each region. pixel With gradient loss L pixel Finally, the total pixel loss L pixel With gradient loss L pixel The sums are used to obtain the total loss L. After obtaining the total loss L, the FusionNet model is trained with the goal of minimizing the total loss L.
[0076] Step 3) specifically involves: processing the polarized visible light image P obtained in step 1). 90 The fused image IMG with the polarization parameters obtained in step 2). Fusion A dataset for detecting defects in fruit stalks / calyxes and surrounding areas was constructed through data augmentation, image annotation, and other processes. Fusion+90 Among them, the polarization parameter fused image IMG Fusion With polarized visible light image P 90 It is aligned; then, the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects is constructed, and the DATA dataset for detecting fruit stalks / calyxes and surrounding defects is used. Fusion+90 The optimal detection model was trained, and finally, the optimal detection model was used to detect defects in the fruit stalk / calyx and surrounding area on the dataset DATA. Fusion+90 The test set is processed to output the detection results of calyx / pedicel and surface defects of the fruit, and the model performance is evaluated.
[0077] Step 3) includes the following steps:
[0078] 3.1) Construct a dataset for detecting defects in the fruit stalk / calyx and surrounding area. Fusion+90 :
[0079] The polarization parameters of each image pair obtained in step 2) are fused into an IMG image. Fusion and the corresponding polarized visible light image P in step 1). 90 All images are scaled to the same size and then copied and subjected to the same processing using one of the following methods: horizontal flip, vertical flip, or angular rotation, for data augmentation. The resulting IMG image is then fused from all original and augmented polarization parameters. Fusion and its corresponding polarized visible light image P 90 Construct a dataset for detecting defects in fruit stalks / calyxes and surrounding areas. Fusion+90 ;
[0080] The data set for detecting defects in the fruit stalk / calyx and surrounding area Fusion+90 In the middle, a polarization parameter fused image IMG Fusion and its corresponding polarized visible light image P 90 This forms an image pair.
[0081] 3.2) Dataset for detecting defects in fruit stalks / calyxes and surrounding areas Fusion+90 Note:
[0082] The LabelImg software was used to annotate the defect detection dataset DATA for the fruit stalk / calyx and surrounding areas. Fusion+90 The fruit stalk / calyx region and defective regions;
[0083] Specifically, the data set for detecting defects in the fruit stalk / calyx and surrounding area... Fusion+90 In the process, the fruit stalk / calyx region and the defect region are labeled for each pair of images.
[0084] 3.3) Construct the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects, and use the labeled data set of fruit stalks / calyxes and surrounding defects. Fusion+90 The data was input into the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects, and a new loss function was constructed for training to obtain the optimal detection model.
[0085] The YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects described in this invention is based on the YOLOv10N model. It is constructed by specific optimization and enhancement of the model for fruit stalks / calyxes and surface defects, and is named YOLO-PVF.
[0086] The YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects specifically includes a backbone network, a neck network, and a head network, including:
[0087] The original backbone network uses two feature extraction networks and three polarization and visible light information interaction modules (PVF). The two feature extraction networks are the polarization feature extraction network BACKBONE. Pol BACKBONE, a network for extracting visible light features Vis The three PVF modules are PVF1, PVF2, and PVF3. The feature F values of multi-scale 90° polarized visible light images are fused using these three modules. Vis Features F of the image fused with polarization parameters Pol ; where Pol represents polarization and Vis represents visible light.
[0088] The visible light feature extraction network BACKBONE VisThe system comprises four sequentially connected feature extraction modules. Each module outputs its own 90° polarization image features as feature maps. The first stage is a visible light primary feature extraction module, mainly composed of two consecutive convolutional modules and a primary feature C2f module connected sequentially. The second stage is a visible light intermediate feature extraction module, mainly composed of one convolutional module and two consecutive intermediate feature C2f modules connected sequentially. The third stage is a visible light advanced feature extraction module, mainly composed of one SCDown module and two consecutive advanced feature C2f modules connected sequentially. The fourth stage is a visible light proprietary feature extraction module, mainly composed of one SCDown module, one SPPA module, and one PSA module connected sequentially.
[0089] The polarization feature extraction network BACKBONE Pol The system comprises three sequentially connected feature extraction modules. Each module outputs its own polarization parameter fusion image features as a feature map. The first stage is a polarization primary feature extraction module, which is mainly composed of two consecutive convolutional modules and a primary feature C2f module connected in sequence. The topology of the polarization primary feature extraction module is the same as that of the visible light primary feature extraction module. The second stage is a polarization intermediate feature extraction module, which is mainly composed of a convolutional module and a mid-level feature C2f module connected in sequence. The third stage is a polarization advanced feature extraction module, which is mainly composed of an SCDown module and a high-level feature C2f module connected in sequence.
[0090] The visible light primary feature map output by the visible light primary feature extraction module and the polarization primary feature map output by the polarization primary feature extraction module are jointly input into the PVF1 module for processing to obtain the first PVF feature map, the first 90° polarized visible light image feature map, and the first polarization parameter fused image feature map. Then, the first 90° polarization image feature map and the first polarization parameter fused image feature map are respectively input into the visible light intermediate feature extraction module and the polarization intermediate feature extraction module.
[0091] The visible light intermediate feature map output by the visible light intermediate feature extraction module and the polarization intermediate feature map output by the polarization intermediate feature extraction module are jointly input into the PVF2 module for processing to obtain the second PVF feature map, the second 90° polarized visible light image feature map, and the second polarization parameter fused image feature map. Then, the second 90° polarization image feature map and the second polarization parameter fused image feature map are respectively input into the visible light advanced feature extraction module and the polarization advanced feature extraction module.
[0092] The visible light advanced feature map output by the visible light advanced feature extraction module and the polarization advanced feature map output by the polarization advanced feature extraction module are jointly input into the PVF3 module for processing to obtain the third PVF feature map. The visible light advanced feature map output by the visible light advanced feature extraction module is input into the visible light proprietary feature extraction module for processing to obtain the visible light proprietary feature map.
[0093] Finally, the neck network structure and head network structure of the YOLOv10N model are reused, combining the first PVF feature map, the second PVF feature map, the third PVF feature map, and the visible light-specific features. Figure 1 The input is fed into the neck network and then through the head network, outputting the fruit stalk / calyx region and the defect region;
[0094] BACKBONE, a polarization feature extraction network on the polarization parameter fusion image side Pol In comparison, the visible light feature extraction network BACKBONE Vis The number of channels was reduced to half of the original number. The repetition count of the intermediate and advanced C2f modules was changed to 1. The SPPF and PSA modules were removed and replaced with the simplified SCDown_Lite module.
[0095] Each of the PVF modules specifically comprises: a visible light feature extraction network BACKBONE. Vis The feature extraction module at a certain stage outputs 90° polarized visible light image features F Vis Feature map and polarization feature extraction network BACKBONE Pol The polarization parameter fusion image features F output by the feature extraction module at the same stage Pol The feature maps are respectively processed by the channel attention module M CA Spatial attention module M SA Each of them obtained its own feature enhancement map F″ Vis and F″ Pol Then the two feature enhancement maps F″ are... Pol and F″ Vis The components are multiplied, then fused using the Sigmoid activation function to generate a fused feature map F. weight Then fuse the feature map F weight Image features F of visible light polarized at 90° respectively Vis Feature map and polarization parameter fusion image features F Pol Multiplying the feature maps respectively yields the 90° polarized visible light image feature map F. Out-Vis Image feature map F fused with polarization parameters Out-Pol Finally, the 90° polarized visible light image feature map F Out-Vis Image feature map F fused with polarization parametersOut-Pol Adding them together yields the PVF feature map F, the output of the PVF module. PVF The feature map F of the 90° polarized visible light image Out-Vis Polarization parameter fusion image feature map F Out-Pol and PVF feature map F PVF This constitutes the output of the PVF module, the PVF feature map F. PVF Input into the neck network;
[0096] The simplified SCDown_Lite module includes: first performing spatial downsampling to reduce the spatial resolution of the feature map, then performing channel downsampling, and finally using a max pooling layer instead of a depthwise separable convolution to achieve spatial downsampling processing.
[0097] In specific implementation, the DATA dataset for detecting defects in the fruit stalk / calyx and surrounding area will be used. Fusion+90 The dataset DATA for detecting defects in fruit stalks / calyxes and surrounding areas is divided into two parts in a 7:2:1 ratio. Fusion+90 The dataset DATA for detecting defects in fruit stalks / calyxes and surrounding areas is divided into training, validation, and test sets. Fusion+90 The training and validation sets are input into the YOLO-PVF model to obtain the optimal YOLO-PVF model. The test set is then input into the trained YOLO-PVF model to obtain the detection results, which are used for model evaluation.
[0098] The new loss function in the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects is constructed as follows: The optimized BCE loss function is constructed according to the following formula, and the optimized BCE loss is calculated:
[0099]
[0100] Where n represents the total number of categories detected, j represents the training rounds, i represents the number of categories detected, and w ij AP represents the weight parameters of the BCE loss for the i-th class in the j-th training epoch. ij-1 y represents the average precision (AP) of the i-th class in the (j-1)-th training epoch (the previous training epoch). i This indicates whether there is a result for the i-th type of target. BCE is the predicted probability of the i-th type of target. total This represents the sum of the BCE losses for each detection category in the j-th training round;
[0101] The optimized BCE loss and the original loss of the YOLOv10N model are weighted and summed to obtain the final loss, which is then used for training.
[0102] The fruit used in this invention can typically be pear, apple, etc.
[0103] The beneficial effects of this invention are:
[0104] This invention proposes a fruit polarization characterization parameter fusion algorithm to obtain a polarization parameter fused image, which improves the contrast of the calyx / stalk region and defect regions. The fused polarization parameter image is then used in conjunction with a 90° polarized image (visible light image) to detect calyx, stalk, and defect regions, overcoming the limitation of high false recognition rates in traditional visible light image detection. Furthermore, this method uses a simple and low-cost device, making it applicable in actual fruit production and grading lines. Attached Figure Description
[0105] Figure 1 This is a flowchart of the processing method of the present invention.
[0106] Figure 2 These are grayscale images of the four polarization angles of this invention.
[0107] Figure 3 This is the calculated initial polarization image obtained according to the present invention.
[0108] Figure 4 This is the calculated initial horizontal polarization enhancement image obtained according to the present invention.
[0109] Figure 5 This is the preprocessed polarization image of the present invention.
[0110] Figure 6 This is the preprocessed horizontal polarization enhanced image of the present invention.
[0111] Figure 7 This is the network diagram of FusionNet, the polarization characterization parameter fusion model of this invention.
[0112] Figure 8 This is the logic diagram for calculating the loss function of the FusionNet loss function in the polarization characterization parameter fusion model of this invention.
[0113] Figure 9 This is a YOLO-PVF network diagram of the multimodal detection model for the calyx and fruit stalk of the crown pear and its defects, as presented in this invention.
[0114] Figure 10 This is a structural diagram of the PVF module in the YOLO-PVF model of this invention.
[0115] Figure 11 This is an example diagram of the defect detection method for crown pears according to the present invention.
[0116] Figure 12 This is the confusion matrix result of an embodiment of the present invention.
[0117] Table 1 shows the training parameters of the polarization information fusion model FusionNet.
[0118] Table 2 shows the training parameters of the YOLO-PVF model, a multimodal detection model for the calyx and fruit stalk of Crown Pear and its defects.
[0119] Table 3 shows the detection results of the calyx, fruit stalk, and defects of the Crown Pear. Detailed Implementation
[0120] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0121] The embodiments of the present invention use Crown pears as the research object, employing 1000 Crown pear samples. The specific process of the embodiments is as follows:
[0122] 1) Characterization of polarization parameters of crown pear.
[0123] 1.1) Acquisition of polarized visible light images: Four polarized visible light images P0, P1, P2, P3, P4, P5, P6, P7, P8, P9, P10, P11, P12, P13, P14, P15, P16, P17, P18, P19, P10, P11, P12, P13, P14, P15, P16, P17, P18 45 P 90 P 135 A total of 3,000 sets of images were obtained, each set of images including four polarized visible light images.
[0124] 1.2) Obtaining grayscale images at different polarization angles: Calculate the four polarized visible light images P0, P1, P2, and P3 using formula (1). 45 P 90 P 135 grayscale images I0, I 45 I 90 I 135 . ( Figure 2 );
[0125]
[0126] Where R, G, and B represent the R-channel component map, G-channel component map, and B-channel component map of the polarized visible light image, respectively; Gray represents the grayscale image.
[0127] 1.3) Acquisition of polarization degree image and horizontal polarization enhancement image: The grayscale images I0, I10, and I20 of the four polarized visible light images were obtained using formulas (2) to (7). 45 I 90 I 135 Calculate the degree of polarization image DOP0 and the horizontal polarization enhancement image IE0:
[0128] S0=I0+I 90 (2)
[0129] S1=I0-I 90 (3)
[0130] S2=I 45 -I 135 (4)
[0131]
[0132] Where S0, S1, and S2 represent the three parameters of the Stokk vector, and I0, I... 45 I 90 I 135 These represent the grayscale images of four polarized visible light images, with DOP0 representing the polarization degree image. Figure 3 ), IE0 represents a horizontally polarized enhanced image ( Figure 4 ).
[0133] 1.4) Denoising the horizontal polarization enhancement image IE0 and the polarization degree image DOP0: The polarization degree image DOP0 and the horizontal polarization enhancement image IE0 are blurred using median filtering to obtain the denoised polarization degree image DOP1 and the horizontal polarization enhancement image IE1 of the calyx and pedicel region of the crown pear.
[0134]
[0135] Where g(i,j) represents the pixel value after median filtering; f(i,j) represents the pixel value before median filtering; and k determines the window radius, which is 1.
[0136] 1.5) Histogram equalization of horizontal polarization enhanced denoised image IE1 and polarization degree denoised image DOP1: Calculate the optimized polarization degree image DOP2 using formula (8). Figure 5 ) and horizontal polarization enhanced image IE2 ( Figure 6 );
[0137] I out (i,j)=I in (i,j) γ (8)
[0138] Where i and j represent the x and y coordinates of the image points, I out I represents the pixel value after image processing. in The pixel value before image processing is represented by γ, which represents the power-law transformation parameter. In this invention, γ is set to 0.8.
[0139] 1.6) Constructing the Crown Pear Polarization Characterization Parameter Dataset: A pair of fruit images is formed by the horizontal polarization enhancement image IE2 and the degree of polarization image DOP2 obtained from the above processing of a single Crown Pear. The initial dataset DATA consists of all Crown Pear image pairs. 300 sets of images are randomly selected from all processed image pairs in the initial dataset DATA, and offline data augmentation is achieved by copying and amplifying them using rotation, flipping, and scaling, forming a Crown Pear Polarization Characterization Parameter Dataset DATA consisting of 3000 sets of images. DOP+IE Each image pair consists of a polarization degree image DOP2 and a horizontal polarization enhancement image IE2, and the polarization degree image DOP2 and the horizontal polarization enhancement image IE2 are aligned images.
[0140] 2) Construction of FusionNet, a fusion model for polarization characterization parameters of crown pear.
[0141] 2.1) Labeling of Crown Pear Polarization Information Dataset: The data set of fruit polarization characterization parameters DATA was labeled using annotation software. DOP+IE In each image pair, the two images are labeled with the fruit stalk / calyx and surface defects, respectively, to obtain the fruit stalk / calyx mask M. s / c and defect mask M defect That is, for each pair of images, both images are labeled with fruit stalks / calyxes and surface defects.
[0142] 2.2) Constructing the polarization characterization parameter fusion algorithm FusionNet ( Figure 7 ):
[0143] FusionNet, a network model for fusing fruit polarization characterization parameters, includes two polarization feature extraction modules. FE A polarization feature reconstruction module FR Two polarization feature extraction modules FE The two modules have the same topology but different internal parameters, and are used to extract features from the polarization degree image DOP2 and the horizontal polarization enhancement image IE2, respectively. Specifically, they include two parallel modules MODULE with the same network structure but independent internal parameters. FE-DOP and MODULE FE-IE Features were extracted from the polarization degree image DOP2 and the horizontal polarization enhancement image IE2, respectively.
[0144] Each polarization feature extraction module FE It mainly consists of four modules connected in sequence: a convolutional module, a primary residual module (PRM), an intermediate residual module (IRM), and a high-level residual module (ARM); and a polarization feature reconstruction module (MODULE). FRIt mainly consists of four modules connected in sequence: the first residual module RM1, the second residual module RM2, the third residual module RM3, and the fourth residual module RM4; the polarization feature extraction module MODULE FE The convolution module in the code mainly consists of a 5×5 convolution operation and a Leaky ReLU activation function connected in sequence.
[0145] Each residual module in the Primary Residual Module (PRM), Intermediate Residual Module (IRM), and Advanced Residual Module (ARM) mainly consists of three convolutional layers (Conv1, Conv2, and Conv3) and one convolutional skip connection. The three convolutional layers are connected in series and then in parallel with a convolutional skip connection. That is, the input of the residual module is fed into the three convolutional layers and the convolutional skip connection respectively. The outputs of the three convolutional layers and the output of the convolutional skip connection are added together, passed through an activation function, and then output as the output of the residual module. The first two convolutional layers, Conv1 and Conv2, are mainly composed of a convolution operation, an activation function, and a batch normalization operation connected in series. The last convolutional layer, Conv3, consists of only a convolution operation. The convolutional skip connection uses only one convolution operation.
[0146] In the specific implementation, convolutional layers Conv1 and Conv3 use 1×1 convolutional kernels, while the convolutional kernel size of convolutional layer Conv2 is 3×3. The activation functions in convolutional layers Conv1 and Conv2 are Leaky ReLU activation functions. The outputs of convolutional layers Conv3 and the convolutions are added together and then processed by the Leaky ReLU activation function.
[0147] In specific implementation, the last activation function of the first residual module RM1 and the second residual module RM2 is ELU, the activation function of the third residual module RM3 is Leaky ReLU, and the Tanh activation function is used in the fourth residual module RM4.
[0148] Polarization Feature Reconstruction Module FR After the fourth residual module RM4, the output is mapped to the 0-1 range using the following formula:
[0149]
[0150] Where x represents the pixel value of the feature map output by the fourth residual module RM4, tanh() represents the Tanh activation function operation, and f(x) represents the pixel value finally output by the polarization feature extraction module.
[0151] Furthermore, the padding in all convolution operations of the network model is set to SAME, and the stride is set to 1.
[0152] 2.3) The fruit polarization characterization parameter dataset DATA DOP+IE The horizontal polarization enhancement image IE2 and the polarization degree image DOP2 of each image pair are respectively input into two polarization feature extraction modules MODULE. FE The respective convolutional features are obtained, namely, horizontal polarization enhanced convolutional features. FE-DOP Convolutional features with polarization degree FE-IE ;
[0153] Then, the horizontal polarization-enhanced convolution feature is applied using the following formula. FE-DOP Convolutional features with polarization degree FE-IE The two convolutional features are concatenated along the channel dimension to obtain the sum feature:
[0154] F eature =Concat(Feature) FE-DOP Feature FE-IE )
[0155] Here, concat represents a concatenation operation along the channel dimension;
[0156] Finally, the summation feature is input into the polarization feature reconstruction module. FR Obtain the fruit polarization characterization parameter dataset DATA DOP+IE The fused polarization parameter image (IMG) for each pair of images Fusion .
[0157] The fruit polarization characterization parameter dataset was divided into two parts, with the datasets in a 7:2 ratio. DOP+IE The dataset is divided into a training set and a validation set; the polarization characterization parameter dataset DATA is... DOP+IE The training and validation sets are input into the FusionNet model for training to obtain the optimal FusionNet model. Then, the optimal fusion model is used to process each image pair in the initial dataset DATA to output a fused image (IMG) of the polarization parameters of each fruit. Fusion .
[0158] 2.4) Loss function design for FusionNet, a polarization characterization parameter fusion algorithm:
[0159] Mask the fruit stalk / calyx M respectively s / c Multiplying the polarization image DOP2 pixel by pixel, the defect mask M defect Multiplying the horizontally polarized enhanced image IE2 pixel by pixel yields the ground truth GT of the fused fruit stalk / calyx region and defect region. s / c and GT defect .
[0160] That is, for each pair of images, the fruit stalk / calyx mask M is respectively... s / c The true value GT of the fruit stalk / calyx region fusion target is obtained by multiplying the polarization degree image DOP2 pixel by pixel. s / c Defect Mask M defect The true value GT of the defect region fusion target is obtained by multiplying the horizontally polarized enhanced image IE2 pixel by pixel. defect .
[0161] Then, by separately masking the fruit stalk / calyx M s / c and defect mask M defect IMG images fused with polarization parameters respectively Fusion Pixel-by-pixel multiplication yields the fused fruit stalk / calyx region and defect region. s / c and Fusion defect The loss function includes the pedicel / calyx region and the defect region.
[0162] The final pixel loss L is obtained by weighted calculation of the losses in each region. pixel With gradient loss L pixel The total pixel loss L pixel With gradient loss L pixel The summation yields the total loss L. The calculation logic diagram of this loss function is as follows: Figure 8 As shown, the pixel loss and gradient loss are calculated as follows:
[0163] (a) Pixel loss: The pixel loss L in the pedicel / calyx region and the defect region is calculated using formulas (11) to (13) respectively. pixel_s / c L pixel_defect and final pixel loss L pixel ;
[0164]
[0165] L pixel =L pixel-s / c +α×L pixel-defect (13)
[0166] Where H and W represent the length and width of the image, respectively; The gradient operator is denoted by α, which can be used to calculate the gradient of an image using the Sobel operator; α represents the weight coefficient; here, the weight coefficient α is still set to 1.5.
[0167] (b) Gradient loss: The gradient loss L in the fruit stalk / calyx region and the defect region is calculated using formulas (14) to (16) respectively. grad_s / c L grad_defect and the final gradient loss L pixel ;
[0168]
[0169] L grad =L grad-s / c +β×L grad-defect (16)
[0170] Where H and W represent the length and width of the image, respectively; The gradient operator is represented here; this invention uses the Sobel operator to calculate the gradient of the image, and the weight coefficient β is still taken as 1.5.
[0171] 2.5) Dataset partitioning for Crown Pear polarization characterization parameters: The dataset DATA is partitioned in a 7:2 ratio. DOP+IE Divided into training set and validation set;
[0172] 2.6) Training and Validation of the Polarization Characterization Parameter Fusion Algorithm FusionNet: The Crown Pear Polarization Characterization Parameter Dataset DATA DOP+IE The FusionNet model is trained using the training and validation sets as inputs, and the following hyperparameters and strategies are used to obtain the optimal FusionNet model.
[0173] Table 1. Training parameters of the FusionNe model for polarization characterization parameters.
[0174]
[0175] 2.7) Input each pair of images (polarization degree image DOP2 and horizontal polarization enhancement image IE2) from the initial dataset DATA into the trained FusionNet model to obtain the polarization parameter fused image IMG. Fusion .
[0176] 3) Construction of YOLO-PVF, a multimodal detection model for fruit calyxes, pedicels and their defects.
[0177] 3.1) Construct a dataset for detecting defects in the fruit stalk / calyx and surrounding area. Fusion+90 :
[0178] The polarization parameters of each image pair obtained in step 2) are fused into an IMG image. Fusion and the corresponding polarized visible light image P in step 1). 90 All images are scaled to the same size and then copied and subjected to the same processing using one of the following methods: horizontal flip, vertical flip, or angular rotation, for data augmentation. The resulting IMG image is then fused from all original and augmented polarization parameters. Fusion and its corresponding polarized visible light image P 90 A dataset of 33,000 images was constructed for detecting defects in fruit stalks / calyxes and surrounding areas. Fusion+90 ;
[0179] Dataset for detecting defects in fruit stalks / calyxes and surrounding areas Fusion+90 In the middle, a polarization parameter fused image IMG Fusion and its corresponding polarized visible light image P 90 This forms an image pair.
[0180] 3.2) Dataset for detecting defects in fruit stalks / calyxes and surrounding areas Fusion+90 Note:
[0181] The LabelImg software was used to annotate the defect detection dataset DATA for the fruit stalk / calyx and surrounding areas. Fusion+90 The fruit stalk / calyx region and defective regions;
[0182] Specifically, the dataset for detecting defects in the fruit stalk / calyx and surrounding area. Fusion+90 In the image, the fruit stalk / calyx region and the defect region are labeled for each image pair.
[0183] 3.3) Dataset for detecting defects in the pedicel / calyx and surrounding area of Crown Pear Fusion+90 Division: The Crown Pear Fruit Stalk / Calyx and Surrounding Defect Detection Dataset DATA was divided into two datasets in a 7:2:1 ratio. Fusion+90 It is divided into training set, validation set and test set;
[0184] 3.4) Constructing a multimodal detection model for fruit stalks / calyxes and defects using YOLO-PVF ( Figure 9 The labeled fruit stalk / calyx and surrounding defect detection dataset DATA Fusion+90 The data was input into the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects, and a new loss function was constructed for training to obtain the optimal detection model.
[0185] The YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects specifically includes a backbone network, a neck network, and a head network. It uses YOLOv10N as the baseline model and adds two feature extraction networks and three polarization and visible light information interaction modules (PVF modules) to the original backbone network. The two feature extraction networks are the polarization feature extraction network BACKBONE. Pol BACKBONE, a network for extracting visible light features Vis The three PVF modules are PVF1, PVF2, and PVF3. The feature F values of multi-scale 90° polarized visible light images are fused using these three modules. Vis Features F of the image fused with polarization parameters Pol ; where Pol represents polarization and Vis represents visible light.
[0186] like Figure 10 As shown, the lower part contains two feature extraction networks: the polarization feature extraction network BACKBONE. Pol The upper part is the visible light feature extraction network BACKBONE. Vis .
[0187] BACKBONE Visible Light Feature Extraction Network Vis The system comprises four sequentially connected feature extraction modules. Each module outputs its own 90° polarization image features as a feature map. The first stage is a visible light primary feature extraction module, mainly composed of two consecutive convolutional modules and a primary feature C2f module connected sequentially. The second stage is a visible light intermediate feature extraction module, mainly composed of one convolutional module and two consecutive intermediate feature C2f modules connected sequentially. The third stage is a visible light advanced feature extraction module, mainly composed of one SCDown module and two consecutive advanced feature C2f modules connected sequentially. The fourth stage is a visible light proprietary feature extraction module, mainly composed of one Spatial-Channel Decoupled Downsampling (SCDown) module, one Spatial Pyramid Pooling-Attention (SPPA) module, and one Partial Self-Attention (PSA) module connected sequentially.
[0188] BACKBONE, a polarization feature extraction network Pol The system comprises three sequentially connected feature extraction modules. Each module outputs its own polarization parameter fusion image features as a feature map. The first stage is a polarization primary feature extraction module, which is mainly composed of two consecutive convolutional modules and a primary feature C2f module connected in sequence. The topology of the polarization primary feature extraction module is the same as that of the visible light primary feature extraction module. The second stage is a polarization intermediate feature extraction module, which is mainly composed of a convolutional module and a mid-level feature C2f module connected in sequence. The third stage is a polarization advanced feature extraction module, which is mainly composed of an SCDown module and a high-level feature C2f module connected in sequence.
[0189] The visible light primary feature map output by the visible light primary feature extraction module and the polarization primary feature map output by the polarization primary feature extraction module are jointly input into the PVF1 module for processing to obtain the first PVF feature map, the first 90° polarized visible light image feature map, and the first polarization parameter fused image feature map. Then, the first 90° polarization image feature map and the first polarization parameter fused image feature map are respectively input into the visible light intermediate feature extraction module and the polarization intermediate feature extraction module.
[0190] The visible light intermediate feature map output by the visible light intermediate feature extraction module and the polarization intermediate feature map output by the polarization intermediate feature extraction module are jointly input into the PVF2 module for processing to obtain the second PVF feature map, the second 90° polarized visible light image feature map, and the second polarization parameter fused image feature map. Then, the second 90° polarization image feature map and the second polarization parameter fused image feature map are respectively input into the visible light advanced feature extraction module and the polarization advanced feature extraction module.
[0191] The visible light advanced feature map output by the visible light advanced feature extraction module and the polarization advanced feature map output by the polarization advanced feature extraction module are jointly input into the PVF3 module for processing to obtain the third PVF feature map. The visible light advanced feature map output by the visible light advanced feature extraction module is input into the visible light proprietary feature extraction module to obtain the visible light proprietary feature map.
[0192] Finally, the neck network structure and head network structure of the YOLOv10N model are reused, combining the first PVF feature map, the second PVF feature map, the third PVF feature map, and the visible light-specific features. Figure 1 The input is fed into the neck network and then through the head network, outputting the fruit stalk / calyx region and the defect region;
[0193] BACKBONE, a polarization feature extraction network on the polarization parameter fusion image side Pol In comparison, the visible light feature extraction network BACKBONE Vis The number of channels was reduced to half of the original number. The repetition count of the intermediate and advanced C2f modules was changed to 1. The SPPF and PSA modules were removed and replaced with the simplified SCDown_Lite module.
[0194] The convolutional module in the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects includes a 3×3 convolutional kernel with a stride of 2, a batch normalization layer, and a SiLU activation function. The polarization feature extraction network BACKBONE... Pol BACKBONE, a network for extracting visible light features Vis The topological structures of the primary feature C2f module, intermediate feature C2f module, and advanced feature C2f module in C2f are the same as those in C2f module in YOLOv10N.
[0195] like Figure 10 As shown, each PVF module specifically consists of the visible light feature extraction network BACKBONE. Vis The feature extraction module at a certain stage outputs 90° polarized visible light image features F Vis Feature map and polarization feature extraction network BACKBONE PolThe polarization parameter fusion image features F output by the feature extraction module at the same stage Pol The feature maps all pass through the channel attention module M CA Spatial attention module M SA Each of them obtained its own feature enhancement map F″ Vis and F″ Pol Then the two feature enhancement maps F″ are... Pol and F″ Vis The components are multiplied, then fused using the Sigmoid activation function to generate a fused feature map F. weight Then fuse the feature map F weight Image features F of visible light polarized at 90° respectively Vis Feature map and polarization parameter fusion image features F Pol Multiplying the feature maps respectively yields the 90° polarized visible light image feature map F. Out-Vis Image feature map F fused with polarization parameters Out-Pol Finally, the 90° polarized visible light image feature map F Out-Vis Image feature map F fused with polarization parameters Out-Pol Adding them together yields the PVF feature map F, the output of the PVF module. PVF The feature map F of the 90° polarized visible light image Out-Vis Polarization parameter fusion image feature map F Out-Pol and PVF feature map F PVF This constitutes the output of the PVF module, the PVF feature map F. PVF The input is fed into the neck network. Among them, the channel attention module M... CA Spatial attention module M SA All are based on the CBAM module.
[0196] In the PVF1 and PVF2 modules, the 90° polarized visible light image feature map F Out-Vis Image feature map F fused with polarization parameters Out-Pol The inputs are fed into the feature extraction network BACKBONE. Pol and BACKBONE Vis In the middle, the 90° polarized visible light image feature map F Out-Vis Image feature map F fused with polarization parameters Out-Pol The summation yields the output PVF feature map F. PVF1 and F PVF2 The input is fed into the neck module of the target detection network.
[0197] In the PVF3 module, only the PVF feature map F, the output of the PVF3 module, is obtained. PVF3 and 90° polarized visible light image feature map F Out-Vis PVF feature map FPVF3 The input is to the neck module of the object detection network, and the 90° polarized visible light image feature map F. Out-Vis Input into the feature extraction network BACKBONE Vis In the middle, the feature map F of the fused image without returning polarization parameters is not returned. Out-Pol .
[0198] like Figure 9 As shown, the simplified SCDown_Lite module includes: first, spatial downsampling to reduce the spatial resolution of the feature map; then, channel downsampling; and finally, using max pooling layers instead of depthwise separable convolutions to achieve spatial downsampling processing.
[0199] 3.5) Construction of the loss function for the YOLO-PVF multimodal detection model for defects in the pedicel / calyx and surrounding area of crown pear: The cross-entropy loss function was optimized, the BCE loss of each category was weighted, and the loss weight of each category was correlated with the accuracy of each category in the model training. The optimized BCE loss function calculation formula is shown in (17) to (18):
[0200]
[0201] Where n represents the total number of categories detected, j represents the training rounds, i represents the number of categories detected, and w ij AP represents the weight parameters of the BCE loss for the i-th class in the j-th training epoch. ij-1 y represents the average precision (AP) of the i-th class in the (j-1)-th training epoch (the previous training epoch). i This indicates whether there is a result for the i-th type of target. BCE is the predicted probability of the i-th type of target. total This represents the sum of the BCE losses for each detection category in the j-th training round;
[0202] 3.6) YOLO-PVF training for the multimodal detection model of crown pear fruit stalk / calyx and surrounding defects: The dataset DATA for detecting crown pear fruit stalk / calyx and surrounding defects... Fusion+90 The YOLO-PVF model is input with the training and validation sets, and trained using the following hyperparameters and strategies to obtain the optimal YOLO-PVF model.
[0203] Table 2. Training parameters of the YOLO-PVF model for rapid detection of calyx, pedicel, and defects in Crown Pear.
[0204]
[0205] 3.7) Input the test set into the trained YOLO-PVF model to obtain the detection results. Figure 11 ) and confusion matrix ( Figure 12 ).
[0206] The detection methods described above were used to test images in the test set. To evaluate the effectiveness of the proposed multimodal detection method for crown pear calyx, fruit stalk, and defects based on enhanced polarization characteristics, it was compared with a crown pear calyx, fruit stalk, and defect detection model that uses visible light images alone. This invention additionally trains a crown pear calyx, fruit stalk, and defect detection model, YOLO10N_RGB, using the YOLO10N model based on visible light images. The dataset, experimental environment, and training parameters are consistent with the YOLO-PVF model for crown pear calyx, fruit stalk, and defect detection based on polarization information fusion.
[0207] Table 3 lists the test results of different models for detecting the calyx, fruit stalk and defects of Crown Pear. The evaluation indicators include precision, recall, mAP@0.5, and mAP@0.5:0.95.
[0208] Table 3. Detection results of calyx, fruit stalk, and defects in different models of *Pyrus pyrifolia*.
[0209]
[0210] The results in Table 3 above show that the detection model for calyx, fruit stalk, and defects in Crown Pear based on polarization information fusion has higher detection accuracy.
[0211] In summary, this invention overcomes the limitations of traditional visible light image detection, which suffers from high false recognition rates for calyxes, fruit stalks, and defects. It constructs a multimodal method for detecting calyxes, fruit stalks, and defects by fusing polarization-based images with visible light images. Furthermore, this method is simple to implement, low in cost, and can be applied in actual fruit production and grading lines.
[0212] The above specific embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.
[0213] The above description is only a preferred embodiment of the present invention. Therefore, all equivalent changes or modifications made to the structure, features and principles described in the claims of this patent application are included in the scope of this patent application.
Claims
1. A high-precision identification method for fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences, characterized in that, Includes the following steps: 1) Characterization of polarization parameters of fruits: An initial dataset containing polarization information of numerous fruits was constructed by acquiring images of them at different polarization angles and then processing them. DATA and from the initial dataset DATA A dataset of fruit polarization characterization parameters was constructed by randomly selecting a portion of images. ; 2) Fusion of fruit polarization characterization parameters: Fruit polarization characterization parameter dataset Annotations are used to obtain masks for fruit stalks / calyxes and surface defects. A FusionNet network model for fusing fruit polarization characterization parameters is then constructed, with the fruit polarization characterization parameter dataset as input. The optimal fusion model is obtained by training a loss function using a mask design, and then applied to the initial dataset using the optimal fusion model. DATA Processing yields polarization parameter fused images ; 3) Multimodal detection of fruit calyxes, pedicels, and their defects: Construct a dataset for detecting defects in fruit stalks / calyxes and surrounding areas using the results from steps 1) and 2). Based on this dataset of defects detected in the pedicel / calyx and surrounding area A multimodal detection model YOLO-PVF for fruit stalks / calyxes and defects was constructed and trained to obtain the optimal detection model. Finally, the optimal detection model was used to detect the calyxes / stalks and surface defects of fruits. Step 2) includes the following steps: 2.1) Polarization characterization parameter dataset annotation: Fruit polarization characterization parameter dataset Each pair of images in the image is labeled with the fruit stalk / calyx and surface defects, resulting in a fruit stalk / calyx mask. and defect masks ; 2.2) Constructing the FusionNet network model for fusing fruit polarization characterization parameters: The fruit polarization characterization parameter fusion network model FusionNet includes two polarization feature extraction modules. A polarization feature reconstruction module Two polarization feature extraction modules They have the same topological structure but different internal parameters, and are used to extract polarization images. DOP 2 and horizontal polarization enhanced image IE 2 Features; 2.3) The dataset of fruit polarization characterization parameters Horizontal polarization enhancement image of each image pair IE 2 and polarization image DOP 2 The inputs are fed into two polarization feature extraction modules respectively. The respective convolutional features are obtained, namely, horizontal polarization enhanced convolutional features. and polarization degree convolution features ; Then, the horizontal polarization-enhanced convolution features are applied using the following formula. and polarization degree convolution features The two convolutional features are concatenated along the channel dimension to obtain the sum feature. Feature : Feature=Concat(Feature FE-DOP ,Feature FE-IE ) Wherein, Concat represents a concatenation operation at the channel level; Finally, the summation characteristic will be used. Feature Input polarization feature reconstruction module Obtain the horizontally polarized enhanced image IE 2 and polarization image DOP 2 Corresponding polarization parameter fused image ; In the FusionNet network model for fusing fruit polarization characterization parameters in step 2.2): Each of the polarization feature extraction modules It consists of four modules connected in sequence: a convolution module, a primary residual module (PRM), an intermediate residual module (IRM), and a high-level residual module (ARM). The polarization feature reconstruction module It consists of four modules connected in sequence: the first residual module RM1, the second residual module RM2, the third residual module RM3, and the fourth residual module RM4. The polarization feature extraction module The convolution module in the code consists of a 5 × 5 convolution operation and a LeakyReLU activation function connected in sequence. Each residual module consists of three convolutional layers and one convolutional skip connection. The input to the residual module is fed into the three convolutional layers and the convolutional skip connection in sequence. The outputs of the three convolutional layers and the output of the convolutional skip connection are added together, passed through an activation function, and then output as the output of the residual module. The first two convolutional layers each consist of a convolution operation, an activation function, and a batch normalization operation connected in sequence, while the last convolutional layer consists of only a single convolution operation. The convolutional skip connection uses only one convolution operation. The polarization feature reconstruction module After the fourth residual module RM4, the output is mapped to the 0-1 range using the following formula: f(x) = (tan h(x) + 1) / 2 Where x represents the pixel value of the feature map output by the fourth residual module RM4, tanh() represents the Tanh activation function operation, and f(x) represents the pixel value finally output by the polarization feature extraction module.
2. The method for high-precision identification of fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences according to claim 1, characterized in that, Step 1) specifically refers to: For each fruit, a split-focus plane polarization camera was used to acquire visible light images at polarization angles of 0°, 45°, 90°, and 135°, resulting in four polarized visible light images. , , , Using grayscale images of four polarized visible light images , , , The combined processing yields a horizontally polarized enhanced image of the current fruit. IE 0 and polarization image DOP 0 ; Then, median filtering is applied to obtain the denoised horizontal polarization enhanced image. IE 1 and polarization image DOP 1 Then, the optimized horizontal polarization enhanced image is obtained through histogram equalization. IE 2 and polarization image DOP 2 ; Optimized horizontal polarization enhanced images of all the numerous fruits IE 2 and polarization image DOP 2 Constructing a dataset of fruit polarization characterization parameters .
3. A high-precision identification method for fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences, as described in claim 1 or 2, characterized in that... Step 1) includes the following steps: 1.1) Polarized visible light image acquisition: Four polarized visible light images with polarization angles of 0°, 45°, 90°, and 135° were acquired using a split-focus plane polarization camera. , , , ; 1.2) Obtaining the grayscale image of the polarization angle: Four polarized visible light images were obtained through processing. , , , Their respective grayscale images , , , ; 1.3) Acquisition of polarization degree image and horizontal polarization enhancement image: The following formulas are used based on the grayscale images of four polarized visible light images. , , , The calculated horizontal polarization-enhanced image of the current fruit IE 0 and polarization image DOP 0 : in, , , These represent the three vector parameters of Stokke. , , , These represent the grayscale images of four polarized visible light images. DOP 0 Represents the polarization degree image. IE 0 Represents a horizontally polarized enhanced image; 1.4) Horizontal polarization enhanced image IE 0 and polarization image DOP 0 Denoising: Apply median filtering to the horizontally polarized enhanced image using the following formula. IE 0 and polarization image DOP 0 Perform blurring operations separately to obtain the respective denoised horizontal polarization enhanced images. IE 1 and polarization image DOP 1 These are respectively used as horizontal polarization enhanced and denoised images. IE 1 Denoising images based on polarization DOP 1 ; in, This represents the pixel value after median filtering; This represents the pixel value before median filtering; k This represents the window radius parameter; `median{}` represents the median operation. i , j Represents the x and y coordinates of the image; m , n These represent independent values from the set {-k, ..., k}. 2 Take the value from; 1.5) Enhance and denoise the horizontally polarized image. IE 1 Denoising images based on polarization DOP 1 Histogram equalization: The following formula is used to transform each image separately to obtain the optimized horizontal polarization enhanced image. IE 2 and polarization image DOP 2 ; I out (i,j)=I in (i,j) γ in, This represents the pixel value after image transformation. This represents the pixel values before image transformation. Indicates the parameters of the power-law transform; 1.6) Constructing the initial dataset and the polarization characterization parameter dataset: Optimized horizontal polarization enhanced image obtained for each fruit IE 2 and polarization image DOP 2 The initial dataset consists of all fruit image pairs. From the initial dataset A subset of image pairs is randomly selected from all image pairs and processed using one of the following methods: rotation, flipping, or scaling, to achieve offline data augmentation and thus construct a dataset of polarization characterization parameters. .
4. The method for high-precision identification of fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences according to claim 1, characterized in that, Step 2) specifically refers to: Fruit polarization characterization parameter dataset Horizontal polarization enhancement images of each image pair in the image IE 2 and polarization image DOP 2 The fruit stalk / calyx region and the defect region were marked to obtain the fruit stalk / calyx mask. With defect mask ; Construct a FusionNet network model for fusing fruit polarization characterization parameters and input the fruit polarization characterization parameter dataset obtained in step 1). Using fruit stalks / calyxes as a mask With defect mask The loss function Loss is designed to train the fruit polarization representation parameter fusion network model FusionNet to obtain the optimal fusion model. Then, the optimal fusion model is used to apply the initial dataset. The image processing outputs a fused image of each fruit by processing each pair of images. .
5. A high-precision identification method for fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences, as described in claim 1 or 4, characterized in that... In step 2), the loss function Loss will be combined with the fruit stalk / calyx mask. With defect mask Set it up as follows: For each pair of images, mask the fruit stalk / calyx separately. With polarization image DOP 2 Pixel-wise multiplication yields the ground truth value of the fused fruit stalk / calyx region. Defect mask The true value of the defect region fusion target is obtained by multiplying the horizontally polarized enhanced image IE2 pixel by pixel. ; Then, by separately masking the fruit stalk / calyx and defect mask Images fused with polarization parameters respectively Pixel-by-pixel multiplication yields the fused fruit stalk / calyx region images. and defect area images ; Pixel loss is calculated as follows: With gradient loss L grad : (a) Pixel loss: Based on the true value of the fruit stalk / calyx region fusion target True value of defect region fusion target Images of the fruit stalk / calyx region and defect area images The pixel loss in the fruit stalk / calyx region and the defect region is calculated using the following formula. , Finally, the weighted sums are used to obtain the final pixel loss. ; in, H and W These represent the length and width of the image, respectively. This indicates the calculation of the L1 norm. Indicates the weighting coefficient; (b) Gradient loss: The gradient loss of the fruit stalk / calyx region and the defect region are calculated using the following formulas respectively. , Finally, the weighted sums are used to obtain the final gradient loss L. grad ; in, H and W These represent the image's length and width, respectively. Represents the gradient operator; Indicates the weighting coefficient; Finally, the total pixel loss is calculated. With gradient loss L grad Add them together to get the total loss L .
6. The method for high-precision identification of fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences according to claim 1, characterized in that, Step 3) specifically involves: processing the polarized visible light image obtained in step 1). The image fused with the polarization parameters obtained in step 2). A dataset for detecting defects in fruit stalks / calyxes and surrounding areas was constructed through data augmentation and image annotation. Among them, polarization parameter fusion image With polarized visible light images It is aligned; then, the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects is constructed, and the fruit stalk / calyx and surrounding defect detection dataset is used. The optimal detection model was trained, and finally used to detect defects in the fruit stalk / calyx and surrounding area. The test set is processed to output the detection results of calyx / pedicel and surface defects of the fruit, and the model performance is evaluated.
7. A high-precision identification method for fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences, as described in claim 1 or 6, characterized in that... Step 3) includes the following steps: 3.1) Construct a dataset for detecting defects in fruit stalks / calyxes and surrounding areas. : The image obtained from step 2) is fused together. and the corresponding polarized visible light image in step 1). All images are uniformly scaled to the same size and subjected to the same processing using one of the following methods: horizontal flip, vertical flip, or angular rotation, for data augmentation. The images are then fused together using all the original and augmented polarization parameters. and its corresponding polarized visible light image Construct a dataset for detecting defects in fruit stalks / calyxes and surrounding areas. ; 3.2) Defect detection dataset for fruit stalks / calyxes and surrounding areas Note: The datasets for detecting defects in the fruit stalk / calyx and surrounding areas are labeled separately. The fruit stalk / calyx region and defective regions; 3.3) Construct the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects, and use the labeled dataset of fruit stalks / calyxes and surrounding defects. The data was input into the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects, and a new loss function was constructed for training to obtain the optimal detection model.
8. The method for high-precision identification of fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences according to claim 7, characterized in that, The YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects specifically includes: The backbone network employs two feature extraction networks and three polarization and visible light information interaction modules (PVF). The two feature extraction networks are polarization feature extraction networks. Visible light feature extraction network The three PVF modules are PVF1, PVF2, and PVF3. Features from multi-scale 90° polarized visible light images are fused using these three modules. Features of images fused with polarization parameters ; The visible light feature extraction network The system includes four sequentially connected feature extraction modules. Each module outputs its own 90° polarization image features as feature maps. The first stage is a visible light primary feature extraction module consisting of two consecutive convolutional modules and a primary feature C2f module connected in sequence. The second stage is a visible light intermediate feature extraction module consisting of one convolutional module and two consecutive intermediate feature C2f modules connected in sequence. The third stage is a visible light advanced feature extraction module consisting of one SCDown_Lite module and two consecutive advanced feature C2f modules connected in sequence. The fourth stage is a visible light proprietary feature extraction module consisting of one SCDown_Lite module, one SPPA module, and one PSA module connected in sequence. The polarization feature extraction network The system includes three sequentially connected feature extraction modules. Each module outputs its own polarization parameter fusion image features as a feature map. The first stage is a polarization primary feature extraction module consisting of two consecutive convolutional modules and a primary feature C2f module connected in sequence. The second stage is a polarization intermediate feature extraction module consisting of a convolutional module and a mid-level feature C2f module connected in sequence. The third stage is a polarization advanced feature extraction module consisting of an SCDown_Lite module and a high-level feature C2f module connected in sequence. The feature map output by the visible light primary feature extraction module and the feature map output by the polarization primary feature extraction module are jointly input into the PVF1 module for processing to obtain the first PVF feature map, the first 90° polarized visible light image feature map, and the first polarization parameter fusion image feature map. Then, the first 90° polarization image feature map and the first polarization parameter fusion image feature map are respectively input into the visible light intermediate feature extraction module and the polarization intermediate feature extraction module. The feature maps output by the visible light intermediate feature extraction module and the polarization intermediate feature extraction module are jointly input into the PVF2 module for processing to obtain a second PVF feature map, a second 90° polarized visible light image feature map, and a second polarization parameter fused image feature map. Then, the second 90° polarized image feature map and the second polarization parameter fused image feature map are respectively input into the visible light advanced feature extraction module and the polarization advanced feature extraction module. The feature map output by the visible light advanced feature extraction module and the feature map output by the polarization advanced feature extraction module are jointly input into the PVF3 module for processing to obtain a third PVF feature map. The feature map output by the visible light advanced feature extraction module is input into the visible light proprietary feature extraction module for processing to obtain a visible light proprietary feature map. Finally, the first PVF feature map, the second PVF feature map, the third PVF feature map and the visible light proprietary feature map are input into the neck network and passed through the head network to output the fruit stalk / calyx region and the defect region. Each PVF module specifically comprises: 90° polarized visible light image features Feature map and polarization parameter fusion image features The feature maps are respectively processed by the channel attention module Spatial attention module Each of them obtained its own feature enhancement map. and Then the two feature enhancement maps were... and The components are multiplied, and then fused using the Sigmoid activation function to generate a fused feature map. Then fuse the feature maps Image features of visible light polarized at 90° respectively Feature map and polarization parameter fusion image features Multiplying the feature maps respectively yields the 90° polarized visible light image feature map. Image feature map fused with polarization parameters Finally, the 90° polarized visible light image feature map was generated. Image feature map fused with polarization parameters Adding them together yields the PVF feature map output by the PVF module. Feature map of visible light image with 90° polarization Polarization parameter fusion image feature map and PVF feature map This constitutes the output of the PVF module; The SCDown_Lite module includes: first performing spatial downsampling, then performing channel downsampling, and finally using a max pooling layer instead of a depthwise separable convolution to achieve spatial downsampling processing.
9. The method for high-precision identification of fruit calyx / pedicel and surface defects based on enhanced polarization characteristic differences according to claim 7, characterized in that, The loss function in the YOLO-PVF multimodal detection model for fruit stalks / calyxes and defects is constructed as follows: The optimized BCE loss function is constructed according to the following formula, and the optimized BCE loss is calculated: in, n This indicates the total number of categories detected. j Indicates training rounds, i Indicates the detection category. Indicates the middle j In the training rounds, the first i The weight parameters of the BCE loss for each category, Indicates the first In the training rounds i Average precision (AP) for each category Indicates whether the first [element] exists. i The result of the class target, It is the first i Predicted probability of class target, BCE total Indicates the first j The sum of BCE losses for each detection category in each training epoch.
Citation Information
Patent Citations
Fruit peel defect detection method and system
CN112686885A