Floating object identification method suitable for lightweight environment

Through the improved WL-SGBM algorithm and the improved lightweight YOLOv5 model, combined with the CIOU_sc loss function and EBi-FPNet network, the problem of low garbage detection accuracy in complex lake surface environments is solved, and efficient floating object recognition and detection is achieved, which is suitable for embedded device deployment.

CN119963890APending Publication Date: 2025-05-09JIANGSU UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411969531.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art has low garbage detection accuracy and low efficiency in complex lake environments, and cannot be effectively applied to embedded equipment.

Method used

The improved WL-SGBM algorithm and the improved lightweight YOLOv5 model are used, combined with the CIOU_sc loss function and the EBi-FPNet network, and data is collected and preprocessed through a binocular camera to achieve efficient identification and detection of floating objects.

Benefits of technology

Improves automation and efficiency of lake waste cleaning, reduces model size and increases frame rate, making it suitable for embedded platform deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963890A_ABST
    Figure CN119963890A_ABST
Patent Text Reader

Abstract

The invention discloses a floating object identification method suitable for a lightweight environment, and the method comprises the steps: collecting a plurality of groups of pictures shot by a binocular camera, carrying out the preprocessing of the pictures, and obtaining the internal parameters and external parameters of the binocular camera; after internal parameters and external parameters obtained through calibration of the binocular camera are subjected to stereo correction, a disparity map of a left camera is obtained through an improved WL-SGBM algorithm; a YOLOv5 network model is improved to obtain an improved lightweight model, floating objects are detected and recognized through the improved lightweight model, and coordinate information of a bounding box and types of the floating objects are obtained through a disparity map; and obtaining a data set of water surface floating objects, training the improved lightweight model by using the floating object data set, inputting a to-be-detected floating object data set into the improved lightweight model, and outputting a detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a floating object identification method suitable for a lightweight environment. Background Art

[0002] As an important part of the natural ecosystem, lakes not only provide precious water resources for humans, but also support the survival and reproduction of many biological species. With the rapid development of smart transportation, small-water intelligent unmanned boats have quietly emerged. Among them, small-water cleaning unmanned boats make floating garbage cleaning tasks all-weather and autonomous. At the same time, they can be operated in some dangerous areas that are difficult to reach manually, effectively improving cleaning efficiency and work safety.

[0003] Facing complex water environments, unmanned boats need to accurately identify and locate floating garbage during their travels. Floating objects may have different shapes, colors, and textures, and may be covered by waves or mixed with other objects, which increases the difficulty of identification.

[0004] Previous research results show that great progress has been made in garbage detection, but there are few studies on garbage detection and ranging in lake environments. The existing model has low accuracy and efficiency in garbage detection in complex lake environments, and cannot be more reasonably applied to embedded devices, so the present invention conducts research on lightweight detection of floating objects on the lake surface. Summary of the invention

[0005] The present invention is to solve the problems existing in the above-mentioned prior art and provide a floating object identification method suitable for a lightweight environment. The present invention reduces the model size and increases the frame rate to facilitate embedded platform deployment, thereby improving the automation and efficiency of lake garbage cleaning.

[0006] The technical solutions adopted in the present invention are:

[0007] A floating object identification method suitable for a lightweight environment comprises the following steps:

[0008] S1: Collect several groups of photos taken by the binocular camera, pre-process the photos, and obtain the internal and external parameters of the binocular camera;

[0009] S2: After stereo correction of the intrinsic and extrinsic parameters obtained by calibrating the binocular camera, the disparity map of the left camera is obtained through the improved WL-SGBM algorithm;

[0010] S3: An improved lightweight model is obtained by improving the YOLOv5 network model. The floating objects are detected and identified by the improved lightweight model, and the coordinate information of the bounding box and the category of the floating objects are obtained by the disparity map. The improvements are as follows:

[0011] (1) The original backbone layer CSP-DARnet network is replaced by the improved lightweight ShufflenetV2_cssp network;

[0012] (2) The original PANet network of the neck layer is replaced by the improved EBi-FPNet network;

[0013] (3) Introduce CIOU_sc loss function;

[0014] S4: Obtain a data set of floating objects on the water surface, use the floating object data set to train the improved lightweight model, input the floating object data set to be detected into the improved lightweight model, and output the detection result.

[0015] Furthermore, in S2, an improved WL-SGBM algorithm is obtained by combining the WLS filter and the SGBM algorithm. The optimization process of the improved WL-SGBM algorithm for the disparity map is:

[0016] (21) Parallax calculation:

[0017] Use the SGBM algorithm to calculate the preliminary disparity map, introduce the illumination normalization mechanism into the SGBM algorithm, and add the texture sensitivity factor and edge protection weight into the cost function of the SGBM algorithm;

[0018] (22) Parallax Optimization:

[0019] After obtaining the preliminary disparity map, the WLS filter is used to optimize the disparity map.

[0020] Further, the original ShuffleNetV2 lightweight network of the YOLOv5 network model includes: a CBRM module, a first Shuffle module, a first Shuffle2 module, a second Shuffle module, a second Shuffle2 module, a third Shuffle module, and a third Shuffle2 module;

[0021] The improvements of the improved lightweight ShufflenetV2_cssp are:

[0022] (1) Replace the ordinary convolutions of the first Shuffle module, the second Shuffle module, and the third Shuffle module with 3×3 depth-wise separable convolutions;

[0023] (2) Add C3 module to the first Shuffle2 module, the third Shuffle2 module and after the third Shuffle2 module;

[0024] (3) The improved DSPPF_cs pyramid pooling is introduced at the end of the original ShuffleNetV2 lightweight network.

[0025] Furthermore, the improved DSPPF_cs pyramid pooling is an improvement on the spatial pyramid pooling SPP, and the improvements are as follows:

[0026] (1) Replace the convolution inside SPP with depth-wise separable convolution;

[0027] (2) The 1*1 convolution layer in SPP uses grouped convolution to divide the input channels into several groups, and the channels in each group only operate with the convolution kernels in the same group;

[0028] (3) The brightness of the photograph is determined by introducing the function select_kernel_size(x).

[0029] Furthermore, the improved EBi-FPNet network is obtained by improving BIFPN, and the improvements are as follows:

[0030] (1) For the P4 layer of BIFPN, three-channel features are used and the C3 module is added to process and fuse the features. After fusion, the EMA algorithm is added to calculate the weighted average of the previous output and the current input to smooth the feature map.

[0031] (2) For the P5 layer of BIFPN, the feature map size is adjusted by upsampling and downsampling to ensure the consistency of spatial size. Then, two-channel feature fusion is used to fuse the upsampled feature map from the P4 layer and the downsampled feature map from the P6 layer.

[0032] Furthermore, the CIOU_sc loss function is an improvement on the CIoU loss function, specifically:

[0033] The CIoU_sc loss function is used to process bounding boxes of different scales in combination with size scaling. The center point distance and aspect ratio are added to calculate the distance ρ between the adjusted predicted box and the center point of the real box.

[0034]

[0035] Among them, x' b and y' b is the center point coordinate of the adjusted prediction box, x' bgt and y' bgt is the coordinate of the center point of the adjusted real box, ρ(b′,b′ gt ) is the Euclidean distance between the center point of the predicted box and the true box;

[0036] Calculate the difference in aspect ratio ν

[0037]

[0038] Among them, w' b and h' b is the width and height of the adjusted prediction box, w' gt and h' gt It is the adjusted width and height of the real frame. The inverse tangent function is used to perform nonlinear mapping of the aspect ratio to prevent numerical explosion caused by drastic changes in the aspect ratio.

[0039] The calculation formula of CIOU_sc is as follows:

[0040]

[0041] Among them, IoU is the intersection over union ratio between the predicted box and the real box, α is the weight hyperparameter, and αν is the aspect ratio difference term, which is used to optimize the shape of the target box.

[0042] The present invention has the following beneficial effects:

[0043] ShufflenetV2_cssp design concept: (G1) Remove the Focus layer to avoid multiple slice operations (G2) Avoid multiple use of C3 Layer and C3 Layer with high channels C3 Layer is an improved version of CSP Botleneck proposed by the YOLOv5 author. It is simpler, faster, and lighter, and can achieve better results at similar losses. However, C3 Layer uses multi-way separation convolution. Tests have shown that frequent use of C3 Layer and C3 Layer with a high number of channels takes up more cache space and reduces the running speed.

[0044] The order in which C3 and Shuffle_Block are used may optimize the feature fusion and shuffling process, thereby more effectively utilizing the advantages of these modules. In particular, by using C3 to fuse and enhance features before shuffling, the shuffling operation may be more efficient and meaningful because it is based on richer and enhanced features.

[0045] When performing feature map fusion, EBI-FPN does not simply stack features, but dynamically weights feature maps from different sources using learned weights. This weighting is achieved through trainable parameters, usually calculated through a small "attention" network that dynamically adjusts the weights based on the contextual information of the feature map.

[0046] Before fusion, each feature map is adjusted through an independent 1*1 convolution to match the requirements of the subsequent fusion operation. This not only standardizes the depth of the features, but also provides an additional linear transformation to optimize the combination of features. Normalization layers (such as batch normalization) follow to stabilize the learning process and accelerate convergence. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flow chart of the present invention.

[0048] Figure 2 This is a distance measurement principle diagram of the present invention.

[0049] Figure 3 These are the original disparity map and the improved disparity map of the present invention.

[0050] Figure 4 This is a schematic diagram of the basic network structure of the improved lightweight ShufflenetV2_cssp in the present invention.

[0051] Figure 5 (a) (b) are the architecture diagrams of ShuffleNet; (c) (d) are the architecture diagrams of ShuffleNetV2.

[0052] Figure 6(a) and (b) are schematic diagrams of SPFF and DCPPF_S, respectively.

[0053] Figure 7 This is the architecture diagram of PANet, BiFPN, and EBi-FPN. DETAILED DESCRIPTION

[0054] The present invention will be further described below in conjunction with the accompanying drawings.

[0055] Figure 1 It is a flow chart of the present invention, and its implementation process is as follows:

[0056] S1: Collect 50 groups of photos taken by the binocular camera, use the functions in the camera calibration toolbox of MATLAB to process the images, and obtain the intrinsic and extrinsic parameters of the binocular camera (as shown in Table 1).

[0057] Table 1 Calibration parameters

[0058]

[0059]

[0060] S2: After stereo correction of the intrinsic and extrinsic parameters obtained by calibrating the binocular camera, the disparity map of the left camera is obtained through the improved WL-SGBM algorithm.

[0061] like Figure 2As shown in the figure, P is a point on the object to be measured, OR and OT are two light points of the binocular camera, the imaging points of a point P on the object to be measured on the two camera sensors of the binocular camera are P' and P", f is the focal length of the left and right cameras of the binocular camera, B is the center distance between the light points of the two cameras, Z is the required depth information, and the parallax from point P' to point P" is dis, then

[0062] dis=B-(X r -X t )

[0063] According to the principle of similar triangles:

[0064]

[0065] We can get:

[0066]

[0067] The focal length f and the camera optical center distance B in the formula can be obtained through the calibration process. Therefore, as long as X is obtained r -X t The value of parallax d can be used to obtain the depth information Z of point P on the object to be measured.

[0068] like Figure 3 As shown, the improved WL-SGBM algorithm (i.e., combining the WLS filter and the SGBM cost aggregation algorithm) aims to take advantage of both and provide a more accurate and robust disparity map.

[0069] The optimization process of the improved WL-SGBM algorithm for the disparity map is:

[0070] (21) Parallax calculation:

[0071] The SGBM algorithm is used to calculate the preliminary disparity map. The semi-global matching strategy of the SGBM algorithm is used to estimate the disparity value of each pixel. This method further introduces an illumination normalization mechanism (Illumination-Invariant Compensation Module). Illumination normalization eliminates the unevenness of illumination distribution by performing local contrast enhancement and histogram equalization on the image, thereby ensuring that the disparity calculation results remain stable in shadow or highlight areas. In addition, texture sensitivity factors and edge protection weights are added to the cost function of the SGBM algorithm. By dynamically adjusting the matching weights using local gradients and texture information, the disparity response of object edges and weak texture areas is enhanced, and over-smoothing in sparse texture areas is avoided.

[0072] (22) Parallax Optimization:

[0073] After obtaining the preliminary disparity map, the WLS filter is used to optimize it. The WLS filter adjusts the smoothness according to the edge information of the image, strengthens the details and edge retention in the disparity map, especially in areas with high texture or high contrast, and performs a two-way consistency check (Left-Right Consistency Check, LRCC) on the matching results of the left and right disparity maps. Specifically, for each pixel, the mapping value from the left disparity map to the right disparity map is calculated and compared with the mapping value from the right disparity map to the left disparity map. If the difference exceeds the set threshold, the point is marked as an invalid disparity point. This process greatly improves the accuracy and confidence of the disparity map.

[0074] The disparity map processed by the improved WL-SGBM algorithm can accurately characterize the depth difference between floating objects on the water surface and the background, generating stable ranging results. The core of this method is to combine the semi-global matching strategy of SGBM with the edge-sensitive characteristics of WLS, so that the algorithm can still maintain high accuracy in environments with complex lighting and sparse textures. At the same time, it can adapt to embedded devices through lightweight design to meet real-time application requirements.

[0075] The above disparity calculation and disparity optimization are expressed through functions, specifically:

[0076] SGBM cost aggregation, cost function definition:

[0077]

[0078] Where x represents the pixel coordinates, d represents the distance (depth) information, I1 and I2 represent the pixel values ​​under the left and right viewing angles respectively, W(x) represents the window centered on the pixel x, and Pen(dd") is a penalty term used to constrain the difference between the depth values ​​of adjacent pixels. λ is the weight of the penalty term. D(x,d) is a set of possible depth differences, indicating the possible depth values ​​of a pixel x at the current depth d. L(x,d) represents the minimum cost value at a depth of d at the pixel x. d(x) represents the depth value corresponding to the pixel x.

[0079] After the cost function is improved, the disparity calculation is:

[0080] (1) Adaptive histogram equalization is performed on the left and right images respectively. After the brightness distribution is adjusted to a uniform state in the local area, an illumination normalization mechanism is introduced. The formula for normalizing the image brightness is as follows:

[0081]

[0082] Where I is the pixel intensity, μ and σ are the mean and standard deviation of the image, respectively. The image is preprocessed before calculating the disparity to reduce the impact of illumination changes on the matching accuracy.

[0083] (2) Add texture sensitivity factor and edge protection weight to the cost function C(x,d)

[0084] C"(x,d)=W(p,q)·T(p)·|I left (p)-I right (p)|

[0085] Among them, I left (p) and I right (p) are the brightness values ​​of the pixels in the left and right images respectively. Their difference I left (p)-I right (p) is the core metric of the cost function. The smaller the value, the higher the matching degree. T(p) (texture sensitivity factor): describes the texture intensity of the current pixel. W(p,q) (edge ​​protection weight): describes whether the pixel is in the edge area. By calculating the gradient (image brightness change), the formula is obtained:

[0086]

[0087] Parallax Optimization:

[0088] The WLS filter adjusts the smoothness of the disparity map through a weight matrix:

[0089]

[0090] Where: d(i) is the disparity value, λ controls the global smoothness, W(j) is the weight matrix, and Represent the gradients in the horizontal and vertical directions respectively.

[0091] Generate optimized disparity maps and convert them into depth maps for object detection and 3D scene reconstruction.

[0092] S3: An improved lightweight model is obtained by improving the YOLOv5 network model. The floating objects are detected and identified by the improved lightweight model, and the coordinate information of the bounding box and the category of the floating objects are obtained by the disparity map. The improvements are as follows:

[0093] Improvement (1): The original backbone layer (Bankbone) CSP-DARnet network is replaced by the improved lightweight ShufflenetV2_cssp network;

[0094] Improvement (2): The original PANet network of the neck layer is replaced by the improved EBi-FPNet network;

[0095] Improvement (3): Introduce the CIOU_sc loss function.

[0096] S4: Obtain a dataset of floating objects (bottles and cans), use the floating object dataset to train the improved lightweight network, and obtain a lightweight .pt model. Input the floating object dataset to be detected into the improved lightweight model, and the model outputs the detection results. While maintaining the accuracy, reduce the model size and increase the frame rate to facilitate embedded platform deployment, and improve the automation and efficiency of lake garbage cleaning.

[0097] The improvements involved in S3 are further described below.

[0098] Regarding improvement point (1):

[0099] The original CSP-DARnet network of the backbone layer (Bankbone) is replaced by the improved lightweight ShufflenetV2_cssp network. Figure 4 The improved ShufflenetV2_cssp network structure diagram includes input layer, Backbone layer, Neck layer, and Head layer. The input layer includes a Mosaic data enhancement module and an adaptive anchor frame calculation module connected in series, which are used to process the input data. A picture of the data set with a size of 640×640 is input into the improved ShufflenetV2_cssp network, and after operations such as Mosaic data enhancement at the input end and adaptive anchor frame calculation, it arrives.

[0100] Figure 5 ab and Figure 5 cd are the architecture diagrams of the Shufflnet and ShufflneV2 algorithms in the Backbone layer respectively. In order to balance the accuracy and speed of the network structure while pursuing a more lightweight structure, the improved lightweight model is optimized as follows:

[0101] (1) Replace the ordinary convolutions of the first Shuffle module, the second Shuffle module, and the third Shuffle module with 3×3 depth-wise separable convolutions;

[0102] Depthwise separable convolution operation: First, perform spatial convolution (depthwise convolution) on each channel to extract spatial information. Then integrate the information of different channels through point-by-point convolution (1×1 convolution) to achieve feature fusion. The feature map is downsampled from 640×640 to 320×320.

[0103] (2) C3 modules are added to the first Shuffle2 module, the third Shuffle2 module, and after the third Shuffle2 module, and the feature map size is halved again, from 320×320 to 160×160.

[0104] (3) The improved DSPPF_cs pyramid pooling is introduced at the end of the original ShuffleNetV2 lightweight network. The output feature map size is reduced from 160×160 to 80×80.

[0105] For the improved DSPPF_cs pyramid pooling:

[0106] The improved DSPPF_cs pyramid pooling is an improvement on the spatial pyramid pooling SPP. The feature map enters the improved DSPPF_cs module to complete the multi-scale information fusion. Figure 6a and Figure 6b They are the architecture diagrams of SPPF and DSPPF_cs respectively, and the improvements are as follows:

[0107] (1) Replace the convolution inside SPP with depth-wise separable convolution;

[0108] First, a depthwise convolution is applied to each input channel (one convolution kernel per channel), and then the outputs of these channels are fused using a 1x1 pointwise convolution. The depthwise convolution is first performed with a smaller number of channels, and then the output is processed by pointwise convolution. The separation of depthwise convolution and pointwise convolution significantly reduces the number of multiplication operations and the amount of computation compared to standard convolution.

[0109] (2) The 1*1 convolution layer in SPP uses grouped convolution to divide the input channels into several groups, and the channels in each group only operate with the convolution kernels in the same group;

[0110] Grouped convolution is used for the 1*1 convolution layer to divide the input channels into several groups, here divided into four groups (g=4). The channels in each group only operate with the convolution kernels of the same group, so as to independently learn different feature sets. Grouped convolution reduces the interaction in the convolution operation, thereby reducing the computational burden. Grouped convolution allows different groups to learn features independently, which increases the diversity and robustness of the model and helps the model capture richer information.

[0111] (3) The brightness of the photograph is selected by introducing the function select_kernel_size(x).

[0112] After calculating the average brightness of the input image, select a pooling kernel of the corresponding size (such as 3×3, 5×5, or 7×7) according to the brightness range. Apply the selected kernel to perform multi-scale pooling operations on the feature map to improve the model's adaptability to different scene lighting conditions. Due to the influence of the water surface environment, the pooling operation does not use a fixed kernel size, but dynamically selects the kernel size according to the dimension of the input feature map, and uses the function select_kernel_size(x) to judge the brightness. Adjust the pooling strategy according to the size or feature complexity of the input image to optimize the processing effect. Use a larger pooling kernel to capture more background information under insufficient lighting conditions.

[0113] Regarding improvement point (2):

[0114] The features extracted by the backbone network ShufflenetV2_cssp lightweight network are compressed and fused, and then the processed feature maps are input into the neck layer network.

[0115] For the neck (NECK) layer, EBI-FPN is used to replace the traditional PANet. The information flow of features is optimized through bidirectional features, the feature layer is bidirectionally fused, two-branch and three-branch size feature fusion is added, and multi-scale branch bidirectional feature pyramid fusion is used for the P4 and P5 layer feature maps. The EMA exponential moving average method is introduced to increase data generalization.

[0116] like Figure 7 , the feature maps of two scales with different depths from P4 and P5 respectively come from the backbone network and are transformed by the EBI-FPN fusion framework.

[0117] (1) For the medium target layer (P4), the EBI-FPN3 three-channel fusion strategy is adopted, and the subsequent output of a feature map that fuses cross-scale information is put into the C3 module for processing. After fusion, the EMA algorithm is added to calculate the weighted average of the previous moment output and the current input to smooth the feature map.

[0118] The feature map input from the P4 layer is fused across scales through the following three channels: Channel 1: Directly use the original information of the P4 feature map. Channel 2: Introduce the upsampled feature map from P3 (shallower layer) to provide finer-grained spatial information. Channel 3: Introduce the downsampled feature map from P5 to supplement the semantic information. The feature maps of the three channels are fused in the channel dimension to generate a unified fused feature map.

[0119] The fused feature map is passed into the C3 module: the C3 module further enhances the feature extraction capability and improves the representation of the target.

[0120] Added EMA smoothing algorithm.

[0121] (2) For the large target layer (P5), EBI-FPN2 is used to adjust the feature map size through upsampling and downsampling. After ensuring the consistency of spatial size, two-channel feature fusion is used to fuse the upsampled feature map from the P4 layer and the downsampled feature map from the P6 layer.

[0122] The features of the P5 layer are fused in two channels: Channel 1: Upsampled feature map from P4: The feature map of P4 is upsampled by 2 times to match the feature map of P5. Figure 1 Channel 2: Downsampled feature map from P6: The feature map of P6 is downsampled by 2 times to be the same as the feature map of P5. Figure 1 Consistent space dimensions.

[0123] Fusion operation: The upsampled P4 feature map and the downsampled P6 feature map are weightedly fused pixel by pixel in the channel dimension to generate a new feature map.

[0124] The size of the fused feature map is consistent with that of P5, and the fused feature map of P4 and P5 is input to the downstream Neck layer and object detection head.

[0125] Regarding improvement point (3):

[0126] The CIoU_sc loss function is used to calculate the matching degree between the predicted box and the real box. The improvements are:

[0127] On the basis of CIoU, a scale-aware mechanism is added to process bounding boxes of different scales in combination with size scaling. The center point distance and aspect ratio are added to calculate the distance ρ between the adjusted predicted box and the center point of the real box.

[0128]

[0129] Calculate the difference in aspect ratio ν

[0130]

[0131] The calculation formula of CIOU_sc is as follows:

[0132]

[0133] A sufficient number of floating objects (bottles and cans) are obtained for dataset construction, and the images are input into the optimized network to obtain a lightweight .pt model. The improved lightweight model is applied to the dataset, and the improved lightweight model is compared with the original YOLOv5 algorithm. The evaluation indicators of the present invention use four common evaluation indicators in target detection algorithms, namely, mean average precision (mAP), number of frames per second (FPS), and model size, to evaluate the performance of the algorithm in this paper. The calculation formulas for precision and recall are as follows:

[0134]

[0135] Among them, Precision is the accuracy, Recall is the recall rate, mAP is the average precision of all categories, C is the total number of floating object categories, TP means the number of positive samples correctly identified as positive samples, and FP means the number of negative samples identified as wrong.

[0136] The proposed algorithm is experimentally compared with MobilNetV3, GhostNetV2, and FasterNet, as shown in Table 2.

[0137] Table 2 Comparative experiment

[0138]

[0139] In order to further measure the performance of the improved YOLOv5s lightweight floating object detection model, the same experimental device and parameters were used to conduct experiments under the test data set. The details are shown in Table 3.

[0140] Table 3 Ablation experiment

[0141]

[0142] The model demonstrated an accuracy of 98.6%, the highest among all models, which shows that the model can identify targets very accurately while reducing misjudgments. Although EBi-FPN adds a certain computational burden, the lightweight structure of ShufflenetV2_cssp balances the model's parameter volume and GFLOPs to a low level (1.5M and 3.3GFLOPs). This design enables the model to maintain an efficient data processing speed without placing an excessive burden on computing resources. Overall, the ShufflenetV2_cssp+EBI-FPN model not only improves the accuracy and detection efficiency of the model, but also achieves an effective balance between resource consumption and performance, basically meeting the needs of actual applications.

[0143] The above description is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.

Claims

1. A floating object identification method suitable for lightweight environments, characterized in that: The following steps are involved: S1: Collect several groups of photos taken by the binocular camera, pre-process the photos, and obtain the internal and external parameters of the binocular camera; S2: After stereo correction of the intrinsic and extrinsic parameters obtained by calibrating the binocular camera, the disparity map of the left camera is obtained through the improved WL-SGBM algorithm; S3: An improved lightweight model is obtained by improving the YOLOv5 network model. The floating objects are detected and identified by the improved lightweight model, and the coordinate information of the bounding box and the category of the floating objects are obtained by the disparity map. The improvements are as follows: (1) The original backbone layer CSP-DARnet network is replaced by the improved lightweight ShufflenetV2_cssp network; (2) The original PANet network of the neck layer is replaced by the improved EBi-FPNet network; (3) Introduce CIOU_sc loss function; S4: Obtain a data set of floating objects on the water surface, use the floating object data set to train the improved lightweight model, input the floating object data set to be detected into the improved lightweight model, and output the detection result.

2. The floating object identification method suitable for a lightweight environment as claimed in claim 1, characterized in that: In S2, the improved WL-SGBM algorithm is obtained by combining the WLS filter and the SGBM algorithm. The optimization process of the improved WL-SGBM algorithm for the disparity map is: (21) Parallax calculation: Use the SGBM algorithm to calculate the preliminary disparity map, introduce the illumination normalization mechanism into the SGBM algorithm, and add the texture sensitivity factor and edge protection weight into the cost function of the SGBM algorithm; (22) Parallax Optimization: After obtaining the preliminary disparity map, the WLS filter is used to optimize the disparity map.

3. The floating object identification method suitable for a lightweight environment as claimed in claim 1, characterized in that: The original ShuffleNetV2 lightweight network of the YOLOv5 network model includes: CBRM module, first Shuffle module, first Shuffle2 module, second Shuffle module, second Shuffle2 module, third Shuffle module, third Shuffle2 module; The improvements of the improved lightweight ShufflenetV2_cssp are: (1) Replace the ordinary convolutions of the first Shuffle module, the second Shuffle module, and the third Shuffle module with 3×3 depth-wise separable convolutions; (2) Add C3 module to the first Shuffle2 module, the third Shuffle2 module and after the third Shuffle2 module; (3) The improved DSPPF_cs pyramid pooling is introduced at the end of the original ShuffleNetV2 lightweight network.

4. The floating object identification method suitable for a lightweight environment as claimed in claim 3, characterized in that: The improved DSPPF_cs pyramid pooling is an improvement on the spatial pyramid pooling SPP, and the improvements are as follows: (1) Replace the convolution inside SPP with depth-wise separable convolution; (2) The 1*1 convolution layer in SPP uses grouped convolution to divide the input channels into several groups, and the channels in each group only operate with the convolution kernels in the same group; (3) The brightness of the photograph is determined by introducing the function select_kernel_size(x).

5. The floating object identification method suitable for a lightweight environment as claimed in claim 1, characterized in that: The improved EBi-FPNet network is obtained by improving BIFPN, and the improvements are as follows: (1) For the P4 layer of BIFPN, three-channel features are used and the C3 module is added to process and fuse the features; After fusion, the EMA algorithm is added to calculate the weighted average of the previous moment output and the current input to smooth the feature graph; (2) For the P5 layer of BIFPN, the feature map size is adjusted by upsampling and downsampling to ensure the consistency of spatial size. Then, two-channel feature fusion is used to fuse the upsampled feature map from the P4 layer and the downsampled feature map from the P6 layer.

6. The floating object identification method suitable for a lightweight environment as claimed in claim 1, characterized in that: The CIOU_sc loss function is an improvement on the CIoU loss function, specifically: The CIoU_sc loss function is used to process bounding boxes of different scales in combination with size scaling. The center point distance and aspect ratio are added to calculate the distance ρ between the adjusted predicted box and the center point of the real box. Among them, x' b and y' b is the center point coordinate of the adjusted prediction box, x' bgt and y' bgt is the coordinate of the center point of the adjusted real box, ρ(b′,b′ gt ) is the Euclidean distance between the center point of the predicted box and the true box; Calculate the difference in aspect ratio ν Among them, w' b and h' b is the width and height of the adjusted prediction box, w' gt and h' gt is the adjusted width and height of the real frame; The calculation formula of CIOU_sc is as follows: Among them, IoU is the intersection over union ratio between the predicted box and the real box, α is the weight hyperparameter, and αν is the aspect ratio difference term.