Cotton bale detection method for driverless clamping and holding vehicle

Through the unmanned clamping car, a fixed camera and a mobile camera collects cotton bag image data, a deep backbone network and a multi-scale feature fusion module are used to extract cotton bag features, and a lightweight Anchor-Free detection head is used for target detection, which solves the accuracy and real-time problems of cotton bag detection in complex environments, and realizes high-precision and high-real-time cotton bag detection.

CN120032106APending Publication Date: 2025-05-23TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510103093.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The target detection of cotton bales is difficult to effectively identify in complex environments, especially under conditions such as medium and long distances, strong exposure and low light, traditional visual algorithms are difficult to extract effective features, resulting in insufficient detection accuracy and real-time performance.

Method used

The unmanned clamping car is used to combine a fixed camera and a mobile camera to collect cotton bag image data, and the cotton bag features are extracted through a deep backbone network and a multi-scale feature fusion module, and the target detection is combined with a lightweight Anchor-Free detection head to output the target classification information, coordinates of the target box center point, offset and confidence score.

Benefits of technology

It improves the robustness and real-time nature of the cotton bale detection algorithm, enhances the detection ability of small targets and complex backgrounds, and is suitable for fine-grain detection scenarios of cotton bale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032106A_ABST
    Figure CN120032106A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned clamping vehicle cotton bale detection method, which comprises the following steps: collecting cotton bale image data, and carrying out preprocessing and labeling to generate a cotton bale image data set; constructing a cotton bale recognition neural network model, taking the cotton bale image data set as input, extracting cotton bale image features by using a deep backbone network, performing feature fusion on the image features and a preset reference frame by using a multi-scale feature fusion module, and performing target detection on a fused feature image by using a lightweight Anchor-Free detection head; training the cotton bale identification neural network model, optimizing parameters of the cotton bale identification neural network model, verifying the optimized performance cotton bale identification neural network model by using mAP and Precise indexes, and optimizing the cotton bale identification neural network model for a weak light scene and a complex background; and inputting cotton bale image data acquired in real time into the optimized cotton bale identification neural network model, and outputting the classification information of the target, the coordinates of the center point of the target frame, the offset and the confidence score in combination with the reference frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a cotton bale detection method using an unmanned clamping vehicle. Background Art

[0002] In recent years, object detection technology has made significant progress at home and abroad, mainly through the evolution of region proposal-based methods and single-stage detection methods, and further development combined with emerging technologies such as Transformer. Region proposal methods (such as R-CNN series and Faster R-CNN (RB Girshick. Fast R-CNN. CoRR, abs / 1504.08083, 2015.2, 5, 6, 7.)) have the advantage of high precision through candidate region generation and feature classification, especially suitable for complex backgrounds and precise demand scenes, but the computational complexity is high and the real-time performance is poor.Single-stage detection methods (such as the YOLO (J. Redmon, S. Divvala, R. Girshick, and A. Farhadi. You only look once: Unified, real-time object detection. arXiv preprint arXiv: 1506.02640, 2015.) series) unify the detection task into a regression problem, significantly improving the detection speed. In particular, YOLOv3 (Redmon J. Yolov3: An incremental improvement [J]. arXiv preprint arXiv: 1804.02767, 2018.) and YOLOv4 (Bochkovskiy A, Wang CY, Liao H Y M. Yolov4: Optimal speed and accuracy of object detection [J]. arXiv preprint arXiv: 2004.10934, 2020.) introduce multi-scale feature fusion to enhance the detection capability of small targets; YOLOv6 (Li C, Li L, Jiang H, et al.) al.YOLOv6:A single-stage object detection framework forindustrial applications[J].arXiv preprint arXiv:2209.02976,2022) and YOLOv7(WangC Y,Bochkovskiy A,Liao HY M.YOLOv7:Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C] / / Proceedings of the IEEE / CVFconference on computer vision and pattern recognition.2023:7464-7475) are further optimized in terms of lightweight and real-time performance, adapting to industrial and embedded scenarios, but the detection accuracy in complex scenes is slightly inferior to the region proposal method, and it is difficult to ensure real-time performance under limited computing power.Transformer-based methods (such as the DETR (Y. Zhao et al., "DETRs Beat YOLOs on Real-time Object Detection," 2024 IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 16965-16974, doi: 10.1109 / CVPR52733.2024.01605.) series) achieve fully end-to-end detection by eliminating candidate region designs, and perform well in complex scenes and large-scale datasets, but have high training complexity and lack of real-time performance.

[0003] There are many difficulties in the target detection of cotton bales, especially in complex environments such as medium and long distances, strong exposure and weak light, where the image features of cotton bales are significantly reduced. In this case, the appearance features of cotton bales are not significant enough, making it difficult for traditional visual algorithms to effectively identify them. In addition, the gaps between cotton bales are small, and the surface of cotton bales is usually uneven, which makes the clustering and segmentation process of point cloud data more difficult and the accuracy is difficult to guarantee. Therefore, how to extract effective features from such low-contrast, low-quality images and point cloud data and improve the robustness of the detection algorithm has become the key to solving the problem. Therefore, a high-precision and high-real-time cotton bale detection algorithm is urgently needed. Summary of the invention

[0004] The purpose of the present invention is to provide a cotton bale detection method for an unmanned clamping vehicle in view of the technical defects in the prior art.

[0005] The technical solution adopted to achieve the purpose of the present invention is:

[0006] A method for detecting cotton bales using an unmanned clamping vehicle comprises the following steps:

[0007] Step 1, collecting cotton bale image data in different environments in a cotton ginning mill, preprocessing and annotating the collected cotton bale image data to generate a cotton bale image data set, and dividing the cotton bale image data set into a training set, a validation set, and a test set;

[0008] Step 2, constructing a cotton bale recognition neural network model, inputting a cotton bale image data set into the cotton bale recognition neural network model, extracting cotton bale image features using a deep backbone network combined with batch normalization and nonlinear activation, using a multi-scale feature fusion module to perform feature fusion on the extracted image features and a preset reference frame to obtain a fused feature image, and using a lightweight Anchor-Free detection head to perform target detection on the fused feature image, and outputting the target classification information, target frame center point coordinates, offset and confidence score;

[0009] Step 3, input the training set into the cotton bale recognition neural network model for training, optimize the parameters of the cotton bale recognition neural network model, use the mAP and Precision indicators to verify the performance of the optimized cotton bale recognition neural network model, and optimize the cotton bale recognition neural network model for low-light scenes and complex backgrounds;

[0010] Step 4, input the cotton bale image data collected in real time into the cotton bale recognition neural network model optimized in step 3, and output the classification information of the target, the coordinates of the center point of the target frame, the offset and the confidence score in combination with the reference frame.

[0011] In the above technical scheme, cotton bale image data in different environments in the cotton ginning mill are collected by combining a fixed camera installed near the conveyor line with a mobile camera installed on an unmanned clamping vehicle. Both the fixed camera and the mobile camera are fisheye cameras. The fisheye camera is used to calibrate the camera to determine the internal parameters, external parameters and distortion coefficients of the fixed camera and the mobile camera. The distortion coefficient is used to correct the lens distortion of the fixed camera and the mobile camera to generate corrected cotton bale image data.

[0012] In the above technical solution, the different environments include different lighting conditions such as uniform lighting under natural light, low light environment and local strong light reflection, different background complexities such as clean background or complex background containing interference, and cotton bale stacking methods such as single cotton bale, dense stacking, and multiple stacking angles.

[0013] In the above technical solution, the preprocessing and labeling of the collected cotton bale image data includes the following steps:

[0014] S1.1: perform brightness adjustment, contrast enhancement and background separation processing on the collected cotton bale image data;

[0015] S1.2: Data amplification is performed on the processed cotton bale image data using random rotation, mirror flipping and blur processing techniques;

[0016] S1.3: Use a semi-automatic annotation tool to annotate the cotton bale image data and generate an annotation file in YOLO format, i.e., the cotton bale image dataset. Classify the cotton bale image dataset according to different environments and divide it into training set, validation set, and test set.

[0017] In the above technical solution, the cotton bale image data set includes: target category number, normalized target frame center coordinates and width and height information.

[0018] In the above technical solution, the deep backbone network adopts the residual module of the deep separable convolution and SE attention mechanism; the deep separable convolution includes deep convolution and point-by-point convolution, and the deep convolution performs a separate convolution operation on each input channel to extract spatial features; the point-by-point convolution is performed after the deep convolution, using a 1x1 convolution kernel, and the output of the deep convolution is aggregated on all feature channels to extract the cotton bale image features of the feature channels; the residual module of the SE attention mechanism helps the deep backbone network to adaptively adjust the importance of the feature channels.

[0019] In the above technical solution, the depth convolution expression is as follows:

[0020]

[0021] In the formula, k represents the channel index, M and N represent the convolution kernel size, represents the value of the output feature position (i, j) of the kth channel, Represents the pixel value of the input feature map of the kth channel within the convolution window, Represents the depth convolution kernel weight corresponding to the kth channel;

[0022] The point-by-point convolution expression is as follows:

[0023]

[0024] Where o represents the output channel index, Represents the value of the oth channel at position (i, j) of the output feature map, represents the output of the depthwise convolution (the value of the output feature position (i, j) of the kth channel), Represents the weight of the point-by-point convolution, the linear transformation weight from input channel k to output channel o, c m Represents the number of channels of the input feature map;

[0025] The residual module expression of the SE attention mechanism is as follows:

[0026] y = ReLU6(BN(Depthwise(x)))

[0027] y=x+SE(F(x,W))

[0028] SE(x)=x·σ(W 2 ReLU(W 1 ·GAP(x)))

[0029] In the formula, y represents the final output feature image, ReLU6 represents nonlinear activation, BN represents batch normalization, Depthwise(x) represents the depthwise separable convolution of the input feature image, x represents the feature image after the depthwise separable convolution operation, F(x,W) represents the function of operating the input feature image, including convolution and pooling operations, where W represents the weight to be learned in the operation, GAP(x) represents full average pooling, σ represents the Sigmoid activation function, and W 1 , W 2 Represents the weight matrices of the two fully connected layers in the SE attention mechanism.

[0030] In the above technical solution, the multi-scale feature fusion module is used to fuse the extracted image features to obtain fused features, which include the center coordinates of the basic grid (c x , c y ), offsets Δx and Δy; the multi-scale feature fusion module introduces a feature pyramid network, a path aggregation network and a dilated convolution. The feature pyramid network uses upsampling operations and weighted fusion to effectively combine high-level features with low-level cotton bale image features; the path aggregation network enhances the cotton bale detection performance through bottom-up paths and splicing operations; and the receptive field of small targets is expanded by introducing dilated convolutions.

[0031] In the above technical solution, the upsampling operation expression is as follows:

[0032] P i =Conv(Upsample(P i+1 ))+Conv(C i )

[0033] Where P i represents the fusion feature of the i-th layer in the top-down path, P i+1 Represents the features of the previous layer (higher semantics), Upsample (P i+1 ) represents the nearest neighbor interpolation or bilinear interpolation of P i+1 Upsample to the current layer resolution, C i represents the bottom-up i-th layer feature (higher resolution), Conv(C i ) represents the channel alignment;

[0034] The weighted fusion expression is as follows:

[0035] P i =α·Conv(Upsample((P i+1 ))+β·Conv(C i )

[0036] In the formula, α, β represent learnable parameters, initialized to 0.5, Upsample(P i+1 ) represents the nearest neighbor interpolation or bilinear interpolation of P i+1 Upsample to the current layer resolution, C i represents the bottom-up i-th layer feature (higher resolution), Conv(C i ) represents the channel alignment;

[0037] The bottom-up path expression is as follows:

[0038] P i =Conv(Downsample((P i-1 ))+P i

[0039] Where, Downsample((P i-1 ) represents downsampling the resolution to the current layer through maximum pooling or convolution with a stride of 2;

[0040] The splicing operation expression is as follows:

[0041] P i =Conv(Concat(Downsample(P i-1 ),P i ))

[0042] Where, Downsample((P i-1 ) represents downsampling the resolution to the current layer through maximum pooling or convolution with a stride of 2.

[0043] In the above technical solution, the expression for performing target detection on the fused feature image using a lightweight Anchor-Free detection head is as follows:

[0044] P x =c x +Δx·ε

[0045] P y =c y +Δy·ε

[0046]

[0047]

[0048] p obj =σ(z ob j)

[0049] Where P x , P y Represents the predicted target center point coordinates, c x , c y represents the center coordinates of the base grid, Δx, Δy represent the predicted offsets, ε represents the grid side length, w, h represent the width and height of the predicted target box, respectively, w a , ha represent the width and height of the reference frame respectively, t w ,t h Represent the adjustment values ​​of the width and height of the predicted target box, p c represents the probability that the target cotton bale belongs to category c, z c represents the classification prediction value of the target cotton bale, N represents the total number of target cotton bale categories, and p obj represents the confidence score of the prediction, z obj represents the confidence prediction value, and σ represents the Sigmoid activation function: σ(x) = 1 / (1+e -x ), z j Represents the network's score for category j and is the original output value of the target category prediction.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] 1. The multi-scale feature fusion module of the present invention combines the feature pyramid network (FPN) with the path aggregation network (PANet), and introduces learnable parameters α and β to dynamically adjust the weights of high-level semantic features and underlying resolution features, which can enhance the fusion effect. In addition, the introduction of dilated convolution in the high-resolution feature map can expand the receptive field and improve the detection performance of small targets. It is suitable for fine-grained detection scenarios of cotton bales.

[0052] 2. This invention significantly reduces the computational complexity by removing the anchor frame design and using the Anchor-Free method to directly regress the target center point coordinates and bounding box size. The target center point is dynamically adjusted relative to the grid center by predicting the offset (Δx, Δy), which improves the detection accuracy under complex backgrounds. a ,h a and the predicted adjustment t w ,t h Dynamically adjusting the bounding box can improve the fitting ability of the target box. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Shown is a block diagram of cotton bale detection by the unmanned clamping vehicle of the present invention.

[0054] Figure 2 Shown are schematic diagrams of cotton bale images under different environments according to the present invention.

[0055] Figure 3 The figure shows the cotton bale detection effect under a complex background according to the present invention. DETAILED DESCRIPTION

[0056] The present invention is further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] A method for detecting bales of cotton using an unmanned clamping vehicle, see Figure 1 , including the following steps:

[0058] Step 1: Collect cotton bale image data under different environments in the cotton ginning mill (such as Figure 2 As shown), the collected cotton bale image data is preprocessed and annotated to generate a cotton bale image dataset, and the cotton bale image dataset is divided into a training set, a validation set and a test set.

[0059] This embodiment collects cotton bale image data in different environments in the cotton ginning plant by combining a fixed camera installed near the conveyor line with a mobile camera installed on an unmanned clamping vehicle. The cotton bale image data is high-resolution, multi-perspective cotton bale image data, which can ensure the diversity and coverage of samples (multi-angle, multi-target, multi-level). The different environments include different lighting conditions such as uniform lighting under natural light, low light environment, and local strong light reflection, different background complexities such as clean background or complex background containing interference (equipment, personnel), and different cotton bale stacking methods such as single cotton bale, dense stacking, and multiple stacking angles.

[0060] The fixed camera and mobile camera in this implementation are both fisheye cameras. The fisheye camera is used for camera calibration to determine the internal parameters, external parameters and distortion coefficients of the fixed camera and the mobile camera. The distortion coefficient is used to correct the lens distortion of the fixed camera and the mobile camera to generate corrected cotton bale image data.

[0061] The preprocessing and labeling of the collected cotton bale image data includes the following steps:

[0062] S1.1: Perform brightness adjustment, contrast enhancement and background separation on the collected cotton bale image data.

[0063] S1.2: The processed cotton bale image data is augmented using random rotation, mirror flipping and blur processing techniques.

[0064] S1.3: Use semi-automatic annotation tools (manual annotation and auxiliary annotation tools) to annotate the cotton bale image data and generate an annotation file in YOLO format, i.e., the cotton bale image dataset. Classify the cotton bale image dataset according to different environments (lighting conditions and background complexity) and divide it into training set, validation set, and test set.

[0065] The cotton bale image dataset includes: target category number, normalized target frame center coordinates, and width and height information.

[0066] Step 2, ( Figure 1 The torch model in the example is built with the python pytrorch library. In this embodiment, the packaged ready-made pytrorch library is called to optimize and build a cotton bale recognition neural network model. The cotton bale image data set is input into the cotton bale recognition neural network model. The cotton bale image features (global features (the overall shape of the cotton bale) and local features (the detailed features of the cotton bale)) are extracted using a deep backbone network (multi-layer convolutional neural network, CNN) combined with batch normalization (Batch Normalization) and nonlinear activation (ReLU6). The extracted image features and the preset reference frame are subjected to feature fusion to obtain a fused feature image, which can improve the perception of small targets and complex backgrounds (in the case of different stacking methods and uneven distribution, the scales of the cotton bale image features are different, that is, the size and position of the cotton bale change differently). The fused feature image is subjected to target detection using a lightweight Anchor-Free detection head, and the classification information of the target cotton bale, the coordinates of the center point of the target frame, the offset and the confidence score are output. See Figure 3 The lightweight Anchor-Free detection head avoids the computational complexity of the traditional anchor box method by regressing the target center point coordinates and bounding box size, significantly reducing the computational complexity and improving the adaptability to complex scenes. It not only improves the computational efficiency, but also effectively reduces the computational burden of the model while maintaining a high detection accuracy. Among them, the deep backbone network combined with batch normalization (Batch Normalization) and nonlinear activation (ReLU6) can further improve the extraction efficiency of the deep backbone network and reduce the risk of overfitting.

[0067] The deep backbone network adopts the residual module of the depthwise separable convolution and SE (Squeeze-and-Excitation) attention mechanism. The depthwise separable convolution includes depthwise convolution and pointwise convolution. The depthwise convolution performs a separate convolution operation on each input channel (so that each convolution kernel acts independently on each channel of the input cotton bale image) to extract spatial features, which can reduce the amount of calculation; the pointwise convolution is performed after the depthwise convolution, and a 1x1 convolution kernel is used to aggregate the output of the depthwise convolution on all feature channels (by performing pointwise linear combination on each feature channel) to extract the cotton bale image features of the feature channels, further optimizing the extraction capability of the cotton bale image features; the residual module of the SE (Squeeze-and-Excitation) attention mechanism helps the deep backbone network to adaptively adjust the importance of the feature channels, so as to better focus on the key parts of the cotton bale, especially in a complex background, and effectively highlight the features of the cotton bale itself.

[0068] The depth convolution expression is as follows:

[0069]

[0070] In the formula, k represents the channel index, M and N represent the convolution kernel size, represents the value of the output feature position (i, j) of the kth channel, Represents the pixel value of the input feature map of the kth channel within the convolution window, Represents the depth convolution kernel weight corresponding to the kth channel.

[0071] The point-by-point convolution expression is as follows:

[0072]

[0073] Where o represents the output channel index, Represents the value of the oth channel at position (i, j) of the output feature map, represents the output of the depthwise convolution (the value of the output feature position (i, j) of the kth channel), Represents the weight of the point-by-point convolution, the linear transformation weight from input channel k to output channel o, c m Represents the number of channels of the input feature map.

[0074] The residual module expression of the SE attention mechanism is as follows:

[0075] y = ReLU6(BN(Depthwise(x)))

[0076] y=x+SE(F(x,W))

[0077] SE(x)=x·σ(W 2 ReLU(W 1 ·GAP(x)))

[0078] In the formula, y represents the final output feature image, ReLU6 represents nonlinear activation, BN represents batch normalization, Depthwise(x) represents the depthwise separable convolution of the input feature image, x represents the feature image after the depthwise separable convolution operation, and F(x,W) represents the function that operates on the input feature image, including convolution, pooling and other operations. Among them, W represents the weight that needs to be learned in the operation, GAP(x) represents full average pooling, σ represents the Sigmoid activation function, and W 1 , W 2 Represents the weight matrices of the two fully connected layers in the SE attention mechanism.

[0079] The multi-scale feature fusion module is used to fuse the extracted image features. The multi-scale feature fusion module introduces a feature pyramid network (FPN), a path aggregation network (PANet) and a dilated convolution. The feature pyramid network uses upsampling operations and weighted fusion to effectively combine high-level features with low-level cotton bale image features to improve the detection capability of cotton bales of different sizes. The path aggregation network can enhance the detection performance of cotton bales through bottom-up paths and splicing operations, especially in complex environments, and can more accurately locate the boundaries of cotton bales. The dilated convolution is introduced to expand the receptive field of small targets (small cotton bales), thereby improving the perception capability of small cotton bales and long-distance cotton bales.

[0080] The upsampling operation expression is as follows:

[0081] P i =Conv(Upsample(P i+1 ))+Conv(C i )

[0082] Where P i represents the fusion feature of the i-th layer in the top-down path, and the fusion feature includes the center coordinates of the base grid (c x , c y ) (representing the center coordinates of a candidate region or anchor box (reference box: a preset border)), offsets Δx and Δy (the offset is the offset of the center coordinates of the target box obtained by network regression relative to the center coordinates of the base grid), Pi+1 Represents the features of the previous layer (higher semantics), Upsample (P i+1 ) represents the nearest neighbor interpolation or bilinear interpolation of P i+1 Upsample to the current layer resolution, C i represents the bottom-up i-th layer feature (higher resolution), Conv(C i ) represents the channel alignment.

[0083] The weighted fusion expression is as follows:

[0084] P i =α·Conv(Upsample((P i+1 ))+β·Conv(C i )

[0085] In the formula, α, β represent learnable parameters, initialized to 0.5, Upsample(P i+1 ) represents the nearest neighbor interpolation or bilinear interpolation of P i+1 Upsample to the current layer resolution, C i represents the bottom-up i-th layer feature (higher resolution), Conv(C i ) represents the channel alignment.

[0086] The bottom-up path expression is as follows:

[0087] P i =Conv(Downsample((P i-1 ))+P i

[0088] Where, Downsample((P i-1 ) represents downsampling the resolution to the current layer through maximum pooling or convolution with a stride of 2.

[0089] The splicing operation expression is as follows:

[0090] P i =Conv(Concat(Downsample(P i-1 ), P i ))

[0091] Where, Downsample((P i-1 ) represents downsampling the resolution to the current layer through maximum pooling or convolution with a stride of 2.

[0092] The expression for performing target detection on the fused feature image using the lightweight Anchor-Free detection head is as follows:

[0093] Px =c x +Δx·ε

[0094] P y =c y +Δy·ε

[0095]

[0096] p obj =σ(z obj )

[0097] Where P x , P y Represents the predicted target center point coordinates, c x , c y represents the center coordinates of the base grid, Δx, Δy represent the predicted offsets, ε represents the grid side length, w, h represent the width and height of the predicted target box, respectively, w a 、h a Represent the width and height of the reference frame, t w , t h Represent the adjustment values ​​of the width and height of the predicted target box, p c represents the probability that the target cotton bale belongs to category c, z c represents the classification prediction value of the target cotton bale, N represents the total number of target cotton bale categories, and p obj represents the confidence score of the prediction, z obj represents the confidence prediction value, and σ represents the Sigmoid activation function: σ(x) = 1 / (1+e -x ), z j Represents the network's score for category j and is the original output value of the target category prediction.

[0098] Step 3: Input the training set into the cotton bale recognition neural network model for training, optimize the parameters of the cotton bale recognition neural network model, use the mAP and Precision indicators to verify the performance of the optimized cotton bale recognition neural network model, and optimize the cotton bale recognition neural network model for low-light scenes and complex backgrounds.

[0099] Step 4, input the cotton bale image data collected in real time into the cotton bale recognition neural network model optimized in step 3, and output the classification information of the target, the coordinates of the center point of the target frame, the offset and the confidence score in combination with the reference frame.

[0100] The above is only a preferred embodiment of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for detecting cotton bales using an unmanned clamping vehicle, characterized in that: The following steps are involved: Step 1, collecting cotton bale image data in different environments in a cotton ginning mill, preprocessing and annotating the collected cotton bale image data to generate a cotton bale image data set, and dividing the cotton bale image data set into a training set, a validation set, and a test set; Step 2, constructing a cotton bale recognition neural network model, inputting a cotton bale image data set into the cotton bale recognition neural network model, extracting cotton bale image features using a deep backbone network combined with batch normalization and nonlinear activation, using a multi-scale feature fusion module to perform feature fusion on the extracted image features and a preset reference frame to obtain a fused feature image, and using a lightweight Anchor-Free detection head to perform target detection on the fused feature image, and outputting the target classification information, target frame center point coordinates, offset and confidence score; Step 3, input the training set into the cotton bale recognition neural network model for training, optimize the parameters of the cotton bale recognition neural network model, use the mAP and Precision indicators to verify the performance of the optimized cotton bale recognition neural network model, and optimize the cotton bale recognition neural network model for low-light scenes and complex backgrounds; Step 4, input the cotton bale image data collected in real time into the cotton bale recognition neural network model optimized in step 3, and output the classification information of the target, the coordinates of the center point of the target frame, the offset and the confidence score in combination with the reference frame.

2. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: Cotton bale image data in different environments in the cotton ginning mill are collected by combining a fixed camera installed near the conveyor line with a mobile camera installed on an unmanned clamping vehicle. Both the fixed camera and the mobile camera are fisheye cameras. The fisheye camera is used to calibrate the camera to determine the intrinsic parameters, extrinsic parameters and distortion coefficient of the fixed camera and the mobile camera. The distortion coefficient is used to correct the lens distortion of the fixed camera and the mobile camera to generate corrected cotton bale image data.

3. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: The different environments include different lighting conditions such as uniform lighting under natural light, low light environment and local strong light reflection, different background complexities such as clean background or complex background with interference, and cotton bale stacking methods such as single cotton bale, dense stacking, and multiple stacking angles.

4. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: The preprocessing and labeling of the collected cotton bale image data includes the following steps: S1.1: perform brightness adjustment, contrast enhancement and background separation processing on the collected cotton bale image data; S1.2: Data amplification is performed on the processed cotton bale image data using random rotation, mirror flipping and blur processing techniques; S1.3: Use a semi-automatic annotation tool to annotate the cotton bale image data and generate an annotation file in YOLO format, i.e., the cotton bale image dataset. Classify the cotton bale image dataset according to different environments and divide it into training set, validation set, and test set.

5. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: The cotton bale image dataset includes: target category number, normalized target frame center coordinates, and width and height information.

6. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: The deep backbone network adopts deep separable convolution and residual module of SE attention mechanism; the deep separable convolution includes deep convolution and point-by-point convolution, and the deep convolution performs a separate convolution operation on each input channel to extract spatial features; the point-by-point convolution is performed after the deep convolution, using a 1x1 convolution kernel, and the output of the deep convolution is aggregated on all feature channels to extract the cotton bale image features of the feature channels; the residual module of the SE attention mechanism helps the deep backbone network to adaptively adjust the importance of feature channels.

7. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: The depth convolution expression is as follows: In the formula, k represents the channel index, M and N represent the convolution kernel size, represents the value of the output feature position (i, j) of the kth channel, Represents the pixel value of the input feature map of the kth channel within the convolution window, Represents the depth convolution kernel weight corresponding to the kth channel; The point-by-point convolution expression is as follows: Where o represents the output channel index, Represents the value of the oth channel at position (i, j) of the output feature map, represents the output of the depthwise convolution (the value of the output feature position (i, j) of the kth channel), Represents the weight of the point-by-point convolution, the linear transformation weight from input channel k to output channel o, c m Represents the number of channels of the input feature map; The residual module expression of the SE attention mechanism is as follows: y = ReLU6(BN(Depthwise(x))) y=x+SE(F(x,W)) SE(x)=x·σ(W2·ReLU(W1·GAP(x))) In the formula, y represents the final output feature image, ReLU6 represents nonlinear activation, BN represents batch normalization, Depthwise(x) represents the depthwise separable convolution of the input feature image, x represents the feature image after the depthwise separable convolution operation, F(x,W) represents the function of operating the input feature image, including convolution and pooling operations, where W represents the weight to be learned in the operation, GAP(x) represents full average pooling, σ represents the Sigmoid activation function, and W1 and W2 represent the weight matrices of the two fully connected layers in the SE attention mechanism.

8. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: The multi-scale feature fusion module is used to fuse the extracted image features to obtain fused features. The fused features include the center coordinates of the basic grid (c x , c y ), offsets Δx and Δy; the multi-scale feature fusion module introduces a feature pyramid network, a path aggregation network and a dilated convolution. The feature pyramid network uses upsampling operations and weighted fusion to effectively combine high-level features with low-level cotton bale image features; the path aggregation network enhances the cotton bale detection performance through bottom-up paths and splicing operations; and the receptive field of small targets is expanded by introducing dilated convolutions.

9. The unmanned clamping vehicle cotton bale detection method according to claim 1 is characterized in that: The upsampling operation expression is as follows: P i =Conv(Upsample(P i+1 ))+Conv(C i ) Where P i represents the fusion feature of the i-th layer in the top-down path, P i+1 Represents the features of the previous layer (higher semantics), Upsample (P i+1 ) represents the nearest neighbor interpolation or bilinear interpolation of P i+1 Upsample to the current layer resolution, C i represents the bottom-up i-th layer feature (higher resolution), Conv(C i ) represents the channel alignment; The weighted fusion expression is as follows: P i =α·Conv(Upsample((P i+1 ))+β·Conv(C i ) In the formula, α, β represent learnable parameters, initialized to 0.5, Upsample(P i+1 ) represents the nearest neighbor interpolation or bilinear interpolation of P i+1 Upsample to the current layer resolution, C i represents the bottom-up i-th layer feature (higher resolution), Conv(C i ) represents the channel alignment; The bottom-up path expression is as follows: P i =Conv(Downsample((P i-1 ))+P i Where, Downsample((P i-1 ) represents downsampling the resolution to the current layer through maximum pooling or convolution with a stride of 2; The splicing operation expression is as follows: P i =Conv(Concat(Downsample(P i-1 ),P i )) Where, Downsample((P i-1 ) represents downsampling the resolution to the current layer through maximum pooling or convolution with a stride of 2.

10. The unmanned clamping vehicle cotton bale detection method according to claim 1, characterized in that: The expression for target detection using a lightweight Anchor-Free detection head on the fused feature image is as follows: P x =c x +Δx·ε P y =c y +Δy·e p obj =σ(z obj ) Where P x , P y Represents the predicted target center point coordinates, c x , c y represents the center coordinates of the base grid, Δx, Δy represent the predicted offsets, ε represents the grid side length, w, h represent the width and height of the predicted target box, respectively, w a 、h a Represent the width and height of the reference frame, t w ,t h Represent the adjustment values ​​of the width and height of the predicted target box, p c represents the probability that the target cotton bale belongs to category c, z c represents the classification prediction value of the target cotton bale, N represents the total number of target cotton bale categories, and p obj represents the confidence score of the prediction, z obj represents the confidence prediction value, and σ represents the Sigmoid activation function: σ(x) = 1 / (1+e -x ), z j Represents the network's score for category j and is the original output value of the target category prediction.