Small target detection method for working image of gantry crane
By building a proprietary data set and optimizing the YOLOv5_OBB model, the attention mechanism and IoU value classification loss function are introduced, and the accuracy problem of small-scale object detection in the working image of the gantry crane is solved, achieving efficient identification and warning in complex environments.
Patent Information
- Application Number
- CN202510311660.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The working image quality of the portal crane is unstable, resulting in low accuracy of small target detection. Especially in severe weather and complex light conditions, it is difficult to effectively identify key objects and provide timely warnings.
Build a proprietary dataset, optimize the YOLOv5_OBB model, introduce Center-BRA and Spatial Group-wise Enhance attention mechanisms, design a classification loss function C_scale Loss based on IoU value, and improve the model's detection ability of small targets.
It improves the recognition accuracy of small and medium-sized targets in the working image of the portal crane, can effectively detect and warn key objects in complex environments, and enhances port operation safety.
Smart Images

Figure CN120259962A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a small target detection method for the working images of portal cranes. Background Art
[0002] In ports or docks, portal cranes are widely used as mechanical equipment for ship cargo handling. As Figure 1 shown, its working function is to move and lift goods to complete the transfer of goods between the cabin and the land. Portal cranes have strong lifting capabilities and the flexibility to handle different types and sizes of goods, which can greatly enhance the efficiency of cargo handling and speed up the turnover of goods. In actual use scenarios, portal crane drivers need to control the gantry crane in the high-altitude cab and use a camera installed on the gantry crane to monitor the working scene below from a top-down perspective to avoid dangerous accidents.
[0003] To ensure the safe operation of portal cranes in ports, it is necessary to detect small targets in the working images of portal cranes. However, due to the special working environment of ports and docks and the camera installation positions, the quality of the working images of portal cranes is unstable, which greatly affects the accuracy of target detection. Specifically, adverse weather, too strong or too weak light, and some uncertain factors generated during the working process of portal cranes reduce the quality of the images transmitted back by the camera. In addition, the extremely uneven object sizes in the working images make it extremely difficult to detect people from the images and give warnings in a timely manner. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention proposes a small target detection method for the working images of portal cranes, which optimizes data collection, improves the upper limit of recognition accuracy, optimizes the model, strengthens important features, and ignores the interference caused by noise or other background factors.
[0005] The present invention adopts the following technical solutions to solve the above problems:
[0006] A small target detection method for the working images of portal cranes, the method comprising the following steps:
[0007] S1: Propose a dataset construction method and construct a proprietary dataset: Propose a dataset construction method for portal crane working images, collect relevant images and construct a proprietary dataset;
[0008] S2: Optimize the YOLOv5_OBB model: Based on the YOLOv5_OBB optimization model with attention mechanism, design and apply the Center-BRA attention module that combines global information and local information;
[0009] S3: Further add an attention mechanism: The YOLOv5_OBB optimized model adds a hybrid attention branch in the Backbone stage and introduces a lightweight Spatial Group-wise Enhance (SGE) attention mechanism into the C3 module in the Neck part;
[0010] S4: Design and introduce a classification loss function with IoU value: A loss function C_scale Loss based on the target size is proposed, introducing an IoU-aware dynamic factor and detection box scale information to make the model focus on detection boxes with smaller sizes or more difficult to be correctly classified;
[0011] S5: Output results: Process the dataset in S1 and screen the final results.
[0012] Furthermore, in S1, the proprietary dataset selects and labels eight categories according to whether the target is common in the working environment of the harbor portal crane and whether it has warning value, including people, slings, trucks, flats, sticks, containers, grabs, and long poles.
[0013] Furthermore, in S2, the processing of the feature map by the Center-BRA attention module can be divided into two stages: local information aggregation and global attention calculation. In the local information aggregation stage, a block-based local attention is used for information interaction in a small area of the feature map, and in the global attention calculation stage, the BRA (bilayer routing attention) in Biformer is adopted.
[0014] Furthermore, the actual calculation process in the local information aggregation stage includes the following steps:
[0015] S211: Divide the input feature map into blocks;
[0016] S212: After the division, perform attention calculation on the adjacent activation values around each small block together with its own body;
[0017] S213: Before finally calculating the attention value, the content brought by the temporary padding needs to be masked to avoid interference of invalid information on the attention weight.
[0018] Furthermore, the global attention calculation process includes the following steps:
[0019] S221: First divide the feature map into multiple regions with larger granularity;
[0020] S222: Calculate the attention for the regions divided in S221.
[0021] Further, in S3, the hybrid attention branch is arranged at the last stage of the Backbone of YOLOv5_OBB, and the hybrid attention branch includes a Center-BRA module and an EMA module.
[0022] Further, in S3, the Spatial Group-wise Enhance (SGE) attention mechanism is arranged in the Bottleneck of the main branch of the C3 module. During the actual operation process, the SGE attention will first divide the input feature map into multiple groups along the channel dimension, and each group of channels corresponds to different semantic sub-features.
[0023] Further, in S4, the loss function C_scale Loss introduces an IoU-aware dynamic factor and a detection box scale-aware dynamic factor in the classification loss, enabling the model to focus on detection boxes with smaller scales and those that are more difficult to be correctly classified.
[0024] Advantages of the present invention:
[0025] 1. Aiming at the special working environment of harbor portal cranes, the present invention proposes a method for constructing a working image dataset of harbor portal cranes and constructs a proprietary dataset. The dataset consists of 30,254 working images of harbor portal cranes with a pixel size of 1960×1080, taking into account the image imaging quality in various situations such as heavy fog, rain and snow, strong light, weak light at night, and normal working environment; according to the actual working needs, the dataset can identify eight types of objects: people, slings, vehicles, flats, sticks, containers, grabs, and long poles; during the actual annotation process of the dataset, rectangular anchor boxes with different rotation angles are used to accurately mark each object, thereby improving the upper limit of recognition accuracy.
[0026] 2. To handle complex background environments, based on the existing model, the present invention proposes an optimized YOLOv5_OBB model based on the attention mechanism, and further designs a Center-BRA attention module that utilizes global information and local information; to make full use of spatial information and information between channels, the model also adds a hybrid attention branch in the Backbone stage; finally, the optimized model introduces a lightweight Spatial Group-wise Enhance attention mechanism in the C3 module of the Neck part, which can further strengthen important features and ignore the interference caused by noise or other background factors.
[0027] 3. Considering that the number of pixels occupied by people in the picture is relatively small, the present invention proposes a loss function C_scale Loss based on the target size. This function mainly introduces the IoU-aware dynamic factor and the detection box scale information, enabling the model to focus on high-quality detection boxes with smaller sizes or higher IoU values. The function assigns a larger weight to the detection boxes with high IoU values and low classification losses, enhancing the model's attention to these accurate detection boxes. In addition, the function weights the detection boxes with high IoU values and high classification losses and those with low IoU values and low classification losses, and the weights are directly related to the magnitudes of the IoU values and the classification losses. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the specific embodiments of the present invention, the following will briefly introduce the drawings required for use in the description of the specific embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0029] Figure 1 It is a physical diagram of a gantry crane in the background art;
[0030] Figure 2 It is a flowchart of this method;
[0031] Figure 3 It is a schematic diagram of local information aggregation calculation;
[0032] Figure 4 It is a schematic diagram of global attention calculation;
[0033] Figure 5 It is a schematic diagram of the structure of the added hybrid attention branch;
[0034] Figure 6 It is a schematic diagram of the structure of adding an attention mechanism in Bottleneck. SPECIFIC EMBODIMENTS
[0035] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, various changes are obvious within the spirit and scope of the present invention defined and determined by the appended claims. All inventions and creations using the concept of the present invention are within the scope of protection.
[0036] It should be noted that the professional terms used in the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the protection scope of the present invention. Unless otherwise specifically stated, various instruments and equipment used in the following embodiments of the present invention can be obtained through market purchase or prepared by existing methods.
[0037] Embodiment 1
[0038] As Figure 2 shown, this solution provides a small target detection method for the working images of portal cranes, which specifically includes the following steps:
[0039] S1: Propose a dataset construction method and construct a proprietary dataset: Propose a dataset construction method for crane working images, collect relevant images and construct a proprietary dataset.
[0040] The proprietary dataset selects and labels eight categories according to whether the target is common in the working environment of seaport portal cranes and whether it has warning value, including people, slings, trucks, flats, sticks, containers, grabs and long poles.
[0041] S2: Optimize the YOLOv5_OBB model: Based on the YOLOv5_OBB optimization model with attention mechanism, design and apply the Center-BRA attention module that uses global information and local information.
[0042] The processing of the feature map by the Center-BRA attention module can be divided into two stages: local information aggregation and global attention calculation. In the local information aggregation stage, a block-based local attention is used to perform information interaction in a small range of regions of the feature map. In the global attention calculation stage, the BRA (bilayer routing attention) in Biformer is adopted.
[0043] The actual calculation process in the local information aggregation stage includes the following steps:
[0044] S211: Divide the input feature map into blocks;
[0045] S212: After the division, perform attention calculation on the adjacent activation values around each small block together with its own body;
[0046] S213: Before finally calculating the attention value, the content brought by the temporary padding needs to be masked to avoid the interference of invalid information on the attention weight.
[0047] The global attention calculation process includes the following steps:
[0048] S221: First divide the feature map into multiple regions with larger granularity;
[0049] S222: Calculate the attention for the regions divided in S221.
[0050] S3: Further add an attention mechanism: The YOLOv5_OBB optimization model adds a hybrid attention branch in the Backbone stage and introduces a lightweight Spatial Group-wise Enhance (SGE) attention mechanism in the C3 module of the Neck part.
[0051] The hybrid attention branch is set at the last stage of the Backbone of YOLOv5_OBB. The hybrid attention branch includes a Center-BRA module and an EMA module.
[0052] The Spatial Group-wise Enhance (SGE) attention mechanism is set in the Bottleneck of the main branch of the C3 module. During the actual operation, the SGE attention will first divide the input feature map along the channel dimension into multiple groups, and each group of channels corresponds to different semantic sub-features.
[0053] S4: Design a classification loss function introducing the IoU value: A loss function C_scale Loss based on the target size is proposed, which introduces an IoU-aware dynamic factor and the detection box scale information, enabling the model to focus on detection boxes with smaller sizes or those that are more difficult to be correctly classified.
[0054] The loss function C_scale Loss introduces an IoU-aware dynamic factor and a detection box scale-aware dynamic factor in the classification loss, allowing the model to focus on detection boxes with smaller scales and those that are more difficult to be correctly classified.
[0055] S5: Output the result: Process the dataset in S1 and screen the final result.
[0056] Embodiment 2
[0057] As a preferred embodiment of the present invention, the method specifically includes the following steps:
[0058] S1: Propose a dataset construction method and construct a proprietary dataset.
[0059] In this embodiment, a dataset dedicated to image recognition of portal cranes in seaports is established. Through actual investigation and analysis, this dataset should have the following four characteristics: a sufficient number of images covering various working conditions, common categories that need to be detected at the work site, a large number of instances under each category, and accurate annotations with correct orientations. Based on this consideration, the dataset of portal crane working images in seaports is different from traditional natural target detection datasets. It selects fewer categories to be more specialized and selects pictures according to the ratio of normal working conditions: strong light: dark environment of 0.8:0.1:0.1, which can reflect the actual working environment of portal cranes in seaports.
[0060] Eight categories are selected and annotated in the dataset of portal crane working images established in this embodiment, including people, slings, trucks, flats, sticks, containers, grabs, and long poles. The selection of categories is based on whether the target of this category is common in the working environment of seaport portal cranes and whether it has warning value.
[0061] The images of the entire dataset are from the working images transmitted back by the cameras of portal cranes in Tianjin Port at different time periods. The size of each image is 1960×1080, and it includes target instances presented in various proportions, orientations, and shapes. The shooting position and shooting time of each image are collected and recorded to ensure that there are no duplicate images in the dataset. Each target instance is framed by a quadrilateral marking box. The specific annotation method is processed using OBB (Oriented BoundingBox). The finally annotated dataset contains 30,254 pictures and is divided into a training set and a test set according to a ratio of 4:1.
[0062] S2: Optimize the YOLOv5_OBB model.
[0063] In this embodiment, a feature extraction module Center - BRA that combines global information and local information is proposed. Its processing of the feature map can be divided into two stages: local information aggregation and global attention calculation.
[0064] In the local information aggregation stage, a block - based local attention is used to perform information interaction in a small - range area of the feature map, so as to obtain a more accurate local pattern, help the model better model local features, and improve the model's performance in details. In the actual calculation process, such as Figure 3As shown in the figure, the Center-BRA module first divides the input feature map into blocks. Assuming the size of the input feature map is W×H and the block size is block_size, then during the actual operation, it will be divided into (W / block_size)×(H / block_size) small blocks in the W and H directions with block_size as the unit, and the size of each small block is block_size×block_size. After the division, the attention calculation is performed on the surrounding adjacent activation values of each small block together with its own body. The range of the selected adjacent feature values here can be adjusted according to actual needs and is set to a. For the small blocks at some special positions in the feature map, such as the small blocks at the corners where there are no adjacent activation values in a certain direction, for this phenomenon, a padding operation needs to be performed to fill the feature map once around, and the filling size is the adjacent range a. Before calculating the attention value finally, the content brought by padding needs to be masked to avoid the interference of invalid information on the attention weight and ensure that the model only focuses on the valid area.
[0065] In the global attention calculation stage of the Center-BRA module designed in this embodiment, BRA (Bi-level Routing Attention) in Biformer is adopted. As an improvement of the Transformer structure, this method reduces the overhead of the original Transformer structure while ensuring performance. BRA is essentially a sparse attention mechanism. As Figure 4 shown, during the actual calculation process, BRA will first divide the feature map into multiple regions with larger granularity and then calculate the attention for these regions. Assuming that a feature map is divided into N blocks in both the W and H directions, if W or H cannot be divided evenly by N, then a padding operation needs to be considered for the initial feature map to fill it to a value that can be divided evenly by N. For the convenience of explanation later, it is defaulted that W and H are numbers that can be divided evenly by N after padding. After the block division operation, a total of N×N blocks will be generated, and the size of each block is (W / N)×(H / N). After obtaining the Q, K, and V matrices that are mapped from the input and are block-divided, a mean operation is performed on the Q and K matrices of each block at the spatial level to obtain the Q and K matrices at the block level, and then the attention score matrix is obtained through QK T calculation. Using this attention score matrix, it is easy to obtain the positions of the k blocks with the highest correlation with a certain block. Then, when calculating the attention, each block can only calculate with the k blocks that are most relevant to it, reducing a large amount of calculation overhead.
[0066] S3: Further add an attention mechanism.
[0067] To improve the quality of the high-level feature maps obtained at the end of the Backbone stage of the model, as Figure 5 shown, in this embodiment, an attention branch is added at the last stage of the Backbone of YOLOv5_OBB. The specific addition position is shown in the following figure. The newly added hybrid attention branch is in the red dashed box; in the specific branch design, this paper uses the Center-BRA module and the EMA (Efficient Multi-Scale Attention) module to form the attention branch.
[0068] As the main module of this branch, Center-BRA has much stronger ability in information utilization than the convolution operation in the initial Backbone; EMA is a module that combines multi-scale feature fusion and lightweight attention mechanism, integrating channel attention mechanism and spatial attention mechanism; in an image, some channels may correspond to some key textures, which are more helpful for determining the contour of the target. Channel attention enhances the feature expression of important channels and suppresses redundant or noisy channels, thereby improving the overall performance of the model; spatial attention generates a spatial weight map by operating on the initial input, and multiplies the spatial weight map with the initial input to highlight the important regions in the image and suppress background interference; moreover, because the spatial weight map is dynamically generated according to the input content and has the ability of dynamic adjustment, it has a certain processing ability for deformed or occluded targets. In addition, the EMA attention also uses 1×1 convolution kernels and 3×3 convolution kernels respectively for feature extraction and feature fusion, further enriching the semantic information in the feature maps.
[0069] To improve the feature detection ability of the traditional C3 module, an attention mechanism is added to the Bottleneck in the main branch of the C3 module. On the one hand, the Bottleneck blocks can be stacked, which can conveniently control the feature extraction ability of the branches. Adding the attention mechanism in the Bottleneck can alleviate the problem of feature information loss caused by the increase in the number of network layers due to the stacking of Bottlenecks. The attention used in this module is the Spatial Group-wise Enhance (SGE) attention mechanism, which is a lightweight spatial attention mechanism. The core is to perform fine-grained enhancement on the feature map by grouping on the channel dimension. During the actual operation process, the SGE attention first divides the input feature map along the channel dimension into multiple groups, and each group of channels corresponds to different semantic sub-features. Each group of feature maps will obtain the global statistical features of the group through global average pooling operation, and this global statistical feature reflects the semantics of the group of feature maps to a certain extent. To enhance important information and suppress redundant information, the value at each spatial position will be calculated for similarity with the global feature to generate the initial similarity value. Then, the similarity value is normalized, and the normalized value is adjusted by a learnable weight and a learnable bias value. Finally, the final attention weight is generated through Sigmoid activation and multiplied by the initial input. As Figure 6 shown, the blue part in the following figure shows the specific optimization position of the Spatial Group-wise Enhance C3 compared with the original C3 module.
[0070] S4: Design a classification loss function that introduces the IoU value.
[0071] IoU (Intersection over Union) is an index in object detection to measure the overlap degree between the predicted bounding box and the ground truth bounding box. Its size directly reflects the accuracy of the predicted box. Using S predict to represent the area of the predicted box S true to represent the area of the ground truth box, and its calculation formula is
[0072]
[0073] In the object detection task, IoU is one of the core metrics for evaluating the model performance and has a direct relationship with some other metrics. For example, mAP (Mean Average Precision), an important metric in the object detection task, is related to the value of IoU. Different IoU thresholds set will directly affect the size of the mAP value. The higher the IoU threshold, the stricter the evaluation criterion. Only when the IoU value of a detection box is greater than the threshold is it considered a correct detection; otherwise, it is considered an incorrect detection. In addition, Precision and Recall are also judged based on the determination results of IoU. IoU is also directly involved in the calculation of the bounding box regression loss, which helps the model optimize the position of the prediction box. In the post-processing stage during the model inference process, IoU is also involved in the calculation process of non-maximum suppression (NMS) to remove redundant detection boxes and only retain the optimal results.
[0074] In the original design of YOLOv5_OBB, the loss consists of four parts, namely classification loss, object confidence loss, rotation angle loss, and bounding box loss. Each part of the loss has a unique function to optimize a certain part of the final result. Some studies have shown that the performance of existing detectors in the field of object detection is restricted by the relatively low connection between the classification score and the localization score. After obtaining the detection box, the classification score and the localization score are calculated independently, which directly leads to a mismatch problem between the classification score and the localization score. For example, there are detection boxes with a relatively high IoU value but a relatively low classification score and detection boxes with a relatively low IoU value but a relatively high classification score. In the NMS stage during the inference process, the sorting is based on the classification score, and finally, the detection box with the highest score greater than the set IoU threshold is retained. In addition, when calculating mAP, since the detection boxes with high classification scores and low IoU suppress the detection boxes with low classification scores and high IoU, the obtained mAP value will decrease when the threshold is set relatively large.
[0075] To address the problems caused by the existing loss calculation method, this embodiment proposes a loss function C_scale Loss based on target size information, which introduces an IoU-aware dynamic factor into the classification loss, enabling the model to focus on high-quality detection boxes with higher IoU values. For the function calculation part, BCEWithLogitsLoss is used as the basis, and different weighting methods are adopted for different IoU values. Specifically, the IoU values in the original range of (0,1) are divided into three segments: (0,0.25) is the low IoU segment, (0.25,0.75) is the normal-sized IoU segment, and (0.75,1) is the high IoU segment. On this basis, the combination of the IoU value and the classification loss is used as an important evaluation criterion, and a larger weight is assigned to the detection boxes with high IoU values and low classification losses to enhance the model's attention to these accurate detection boxes. In addition, two situations that affect subsequent evaluation, namely high IoU value and high classification loss, and low IoU value and low classification loss, need to be processed because these two situations represent difficult samples, and the model needs to perform reinforcement learning on such samples. The specific measure is also to weight the losses of these two types of detection boxes, and the weights are directly related to the sizes of the IoU value and the classification loss.
[0076] Let L cls be the classification loss value of the detection box, and IoU be the IoU value between the detection box and the ground truth box. Then the combined dynamic weight formula is as follows:
[0077]
[0078] In addition, since humans belong to a difficult category among all categories, in order to better identify such targets, weighting is required. In the task of identifying working images of harbor portal cranes, a prominent feature of the human category is its small size, often occupying less than 1% of the pixels in the entire image. Based on this phenomenon, this embodiment sets a weight box_scale based on the size of the detection box to represent the proportion of the size of the annotation box in the feature map. If S_true represents the size of the annotation box and S_feature represents the size of the feature map, then the calculation formula of box_scale is as follows:
[0079]
[0080] The dynamic weight formula for the size of the detection box is as follows:
[0081]
[0082] Assuming the initially calculated loss is L, the loss enhanced by combining the two dynamic weights is as follows:
[0083] C_scale Loss(L) = (mweight +m box )*L
[0084] S5: Output result: Process the dataset in S1 and screen the final result.
[0085] The above has described this embodiment in detail through examples, but the content is only the preferred embodiment of this embodiment and cannot be considered as used to limit the implementation scope of this embodiment; all equivalent changes and improvements made according to the scope of this embodiment application shall still fall within the patent coverage scope of this embodiment.
Claims
1. A small target detection method for the working images of portal cranes, characterized in that: The method includes the following steps: S1: Propose a dataset construction method and construct a proprietary dataset: Propose a method for constructing a dataset of crane working images, collect relevant images and construct a proprietary dataset; S2: Optimize the YOLOv5_OBB model: Based on the YOLOv5_OBB optimization model with attention mechanism, design and apply the Center-BRA attention module that uses global information and local information; S3: Further add an attention mechanism: The YOLOv5_OBB optimization model adds a hybrid attention branch in the Backbone stage and introduces a lightweight Spatial Group-wise Enhance (SGE) attention mechanism in the C3 module of the Neck part; S4: Design and introduce a classification loss function with IoU value: Propose a loss function C_scaleLoss based on the target size, introduce an IoU-aware dynamic factor and detection box scale information, and let the model focus on detection boxes with smaller sizes or more difficult to be correctly classified; S5: Output the result: Process the dataset in S1 and screen the final result.
2. The small target detection method for the working image of a portal crane according to claim 1, wherein: In S1, according to whether the class of target is common in the working environment of the harbor portal crane and whether it has warning value, eight categories are selected and labeled, including people, slings, trucks, flats, sticks, containers, grabs and long poles.
3. A small target detection method for the working image of a portal crane according to claim 1, characterized in that: In S2, the processing of the feature map by the Center-BRA attention module can be divided into two stages: local information aggregation and global attention calculation. In the local information aggregation stage, a block-based local attention is used to perform information interaction in a small area of the feature map, and in the global attention calculation stage, the BRA (bilayer routing attention) in Biformer is adopted.
4. A small target detection method for the working image of a portal crane according to claim 3, characterized in that: The actual calculation process of the local information aggregation stage includes the following steps: S211: Divide the input feature map into blocks; S212: After the division, calculate the attention for each small block together with the adjacent activation values in a circle around it and its own body; S213: Before finally calculating the attention value, the content brought by the temporary padding needs to be masked to avoid interference of invalid information on the attention weight.
5. The small target detection method for the working image of a portal crane according to claim 3, characterized in that: The global attention calculation process includes the following steps: S221: First divide the feature map into multiple regions with larger granularity; S222: Calculate the attention for the regions divided in S221.
6. The small target detection method for the working image of a portal crane according to claim 1, characterized in that: In S3, the hybrid attention branch is set in the last stage of the Backbone of YOLOv5_OBB, and the hybrid attention branch includes a Center-BRA module and an EMA module.
7. A small target detection method for the working image of a portal crane according to claim 1, characterized in that: In S3, the Spatial Group-wise Enhance (SGE) attention mechanism is set in the Bottleneck of the main branch of the C3 module. In the actual operation process, the SGE attention will first divide the input feature map into multiple groups along the channel dimension, and each group of channels corresponds to different semantic sub-features.
8. The small target detection method for the working image of a portal crane according to claim 1, wherein: In S4, the loss function C_scale Loss introduces an IoU-aware dynamic factor and a detection box scale-aware dynamic factor into the classification loss, enabling the model to focus on detection boxes with smaller scales and those that are more difficult to be correctly classified.