Crown block hook identification method and system based on YOLOv8

By improving the YOLOv8 model and combining data collection, lightweight attention mechanism, and self-supervised pre-training technologies, the problems of insufficient accuracy and real-time performance in overhead crane hook recognition were solved, and high-precision, robust, and fast-response overhead crane hook recognition was achieved.

CN120807959AActive Publication Date: 2025-10-17SHANDONG BOANG INFORMATION TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510945083.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in crane hook recognition, weak generalization capabilities, and insufficient real-time performance in complex environments, making it difficult to meet the safety and automation requirements of industrial cranes.

Method used

An improved model based on YOLOv8 is adopted. Through data collection and labeling, lightweight attention mechanism, local detail enhancement branch, dynamic anchor box adaptation mechanism and self-supervised pre-training, combined with model quantization and pruning, the model structure and deployment are optimized to improve recognition accuracy and real-time performance.

Benefits of technology

It improves the accuracy and robustness of overhead crane hook recognition, shortens frame processing time, makes it suitable for real-time deployment on edge devices, and enhances the stability and adaptability of the model in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807959A_ABST
    Figure CN120807959A_ABST
Patent Text Reader

Abstract

The invention discloses a crown block hook identification method and system based on YOLOv8. The method comprises the following steps: S1, data acquisition and labeling; s2, carrying out model structure innovation; s3, a pre-training mechanism is self-supervised; s4, model training and deployment optimization; comprising a data acquisition and labeling module, a model construction and optimization module, a self-supervision pre-training module and a model training and deployment module. According to the invention, the detail extraction capability is enhanced through the LAF and LDEB modules, and the hook identification accuracy is improved; after lightweight optimization of the model, the frame processing time is shortened by more than 30%, and the method is suitable for edge deployment; the target can be stably recognized in complex environments such as strong light, shadow and shielding; and the model structure is high in adaptability, and can be quickly migrated to other lifting target identification tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of industrial hoisting technology, and in particular to a YOLOv8-based overhead trolley hook identification method and system. BACKGROUND

[0002] In the field of industrial hoisting, the overhead trolley hook (also known as the overhead trolley hook or lifting hook) is a key component of the crane (such as the overhead trolley, bridge crane, gantry crane, etc.), used to support, lift and hoist the load. It is usually installed on the lifting arm or hook system of the crane, connected with the ropes, steel wires, chains, etc. of the crane, to complete the hoisting, moving and placing of the goods. The positioning and state recognition of the overhead trolley hook is of great significance to ensure the safety of operation and improve the level of automation.

[0003] At present, the conventional image processing method has limited recognition accuracy in complex environments, and although the deep learning method has made progress, it generally has problems such as low detection accuracy, weak generalization ability, and insufficient real-time performance. Therefore, there is an urgent need for a YOLOv8 improved model-based overhead trolley hook identification method and system to solve the above problems and improve detection accuracy, robustness and inference speed. SUMMARY

[0004] The purpose of the present application is to solve the shortcomings in the prior art and to propose a YOLOv8-based overhead trolley hook identification method and system.

[0005] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions: A YOLOv8-based overhead trolley hook identification method, comprising the following steps: S1: Data acquisition and annotation: use a camera to collect video frames of the overhead trolley hook in different postures, lighting and backgrounds, and perform fine manual annotation to construct a high-quality data set; S2: Model structure innovation: based on the original structure of YOLOv8, the following improvements are proposed: lightweight attention mechanism (Lightweight Attention Fusion, LAF) module, local detail enhancement branch (Local Detail Enhancement Branch, LDEB), dynamic anchor frame adaptive mechanism (Dynamic Anchor Adaptation, DAA), improved loss function; S3: Self-supervised pre-training mechanism (Industrial Pretext Task Pretraining): Construct self-supervised tasks such as "hook direction prediction" and "hook head occlusion prediction"; pre-train on large-scale unlabeled industrial video data, and migrate the pre-training weights to the YOLOv8 backbone network to improve the initial convergence performance and feature expression ability, and improve the model's migration ability and generalization performance in the industrial scene; S4: Model training and deployment optimization: use mixed precision training to improve training efficiency, and model quantization and pruning to adapt to edge device real-time deployment.

[0006] As a further technical solution of the present application, S1 specifically includes: S11: In the data acquisition link, the camera has a frame rate of 30 frames per second. Collect the video frames of the hook of the crown, and ensure that the dynamic changes of the hook of the crown can be fully captured in the time dimension; S12: For data labeling, use fine-grained manual labeling method, specifically as follows: define the labeling accuracy parameter ; Introduce the labeling-consistency algorithm to calculate the consistency index c of the labeling results of different labelers in the same video frame; when building a high-quality data set, use data enhancement algorithms, including random rotation angle , random scaling ratio s and random translation distance d, process the original collected video frames through these algorithms to expand the data set size and enhance the model's recognition ability for different posture hooks of the crown.

[0007] As a further technical solution of the present application, S2 specifically includes: S21: Lightweight attention mechanism (Lightweight Attention Fusion, LAF) module: Introduce a lightweight attention module between Backbone and Neck to effectively enhance the model's perception ability for small targets such as hook edges, hooks, and small details, and suppress background interference; S22: Local detail enhancement branch (Local Detail Enhancement Branch, LDEB): Add an auxiliary branch that focuses on extracting hook edge and connection feature, and improves target recognition ability after fusion with main feature; S23: Dynamic anchor adaptation mechanism (Dynamic Anchor Adaptation, DAA): Optimize anchor parameters through online learning to improve adaptability to hooks of different sizes; S24: Improved loss function: Introduce position-sensitive IoU loss (Pos-IoU Loss) to more accurately evaluate the hook center of gravity prediction error.

[0008] As a further technical solution of the present application, the S21 specifically comprises: S211: defining the channel number of the input feature map as C, the width as W, and the height as H, that is, the input feature map size is ; S212: introducing a convolution operation to calculate the channel attention weight, the convolution kernel size is , the step is , the padding is , and the channel number remains unchanged; S213: calculating the output attention weight feature map ; S214: element-wise multiplying and fusing the attention weight feature map and the input feature map to obtain an enhanced feature map.

[0009] As a further technical solution of the present application, the S22 specifically comprises: adding a parallel local enhancement branch in the original YOLOv8 Neck layer, introducing a Dilated Convolution (Dilated Convolution) in the branch structure to expand the receptive field without losing resolution; combining a Swin Transformer module to introduce a local window attention mechanism to improve local structure expression; and finally fusing into the main feature map through a Concat method to solve the problem of missed detection under hook occlusion, partial blur, or rotated posture.

[0010] As a further technical solution of the present application, the S23 specifically comprises: S231: defining the number of anchor boxes as n _ anchors, the initial anchor box parameters as the width and the height ; S232: optimizing the anchor box parameters through online learning, and the calculation formula is: , wherein t is the iteration number, is the learning rate, and are the real width and height estimation values of the current batch of hooks, respectively.

[0011] As a further technical solution of the present application, S24 specifically comprises: S241: defining the position parameters of the predicted box as the center coordinates , the width , and the height , and the position parameters of the real box as and ; S242: the formula for calculating the position-sensitive IoU loss is: , wherein IoU is the intersection over union of the predicted box and the real box.

[0012] As a further technical solution of the present application, the S3 specifically comprises: S31: constructing a self-supervised task, taking "hook direction prediction" and "hook head occlusion prediction" as examples; S32: pre-training on large-scale unlabeled industrial video data: S321: define the pre-training learning rate as lr _ pre, the number of training rounds as epoch _ pre, the batch size of each training round as batch _ size _ pre; S322: during training, update the model parameters by optimizing the loss function of the above self-supervised task and S33: after pre-training, migrate the pre-trained weights to the YOLOv8 backbone network, and define the transfer learning rate as Ir _ transfer, the number of fine-tuning training rounds as epoch _ fine _ tune, the batch size of each training round as batch _ size _ fine _ tune.

[0013] As a further technical solution of the present application, the S4 specifically comprises: S41: mixed precision training: define the total number of iterations during model training as , the initial learning rate as , and the number of training samples per batch as ; adopt mixed precision training, set the loss scaling factor as to prevent gradient underflow; when updating the gradient, first calculate the scaled gradient , then scale the gradient back to the normal range, and update the model parameters; S42: model quantization: quantize the model to convert 32-bit floating-point (FP32) weights to 8-bit integers (INT8); define the quantization parameters and ; the quantized weight is ; S43: model pruning: adopt a weight-based pruning method, define the pruning rate as prune _ rate (value range 0-1), i.e. prune _ rate proportion of small weights; calculate the absolute value threshold of the weight ​; weights with absolute values less than thr are set to 0; the sparsity of the pruned model is calculated ; S44: Model deployment optimization: define the computing capability parameter of the edge device as compute _ capability, in GFLOPS; according to the device computing capability, the model after quantization and pruning is adapted and optimized to ensure that the inference time of the model on the edge device meets the real-time requirement, that is, the inference time , wherein is the set real-time threshold; at the same time, the detection accuracy of the model on the edge device is ensured to decrease by no more than the set tolerance .

[0014] A YOLOv8-based crown block hook recognition system, comprising a data acquisition and labeling module, a model construction and optimization module, a self-supervised pre-training module, and a model training and deployment module; The data acquisition and labeling module uses a high-definition camera to collect crown block hook video frames from multiple angles, labels key parts, expands the data set by means of data enhancement technology, and improves the adaptability of the model to different scenes; The model construction and optimization module is used to improve on the basis of YOLOv8, integrate a lightweight attention mechanism, a local detail enhancement branch, a dynamic anchor frame adaptive mechanism, and an improved loss function, and strengthen the perception and feature extraction capability of the model for small targets; The self-supervised pre-training module is used to construct a self-supervised task, pre-train on large-scale unlabeled industrial video data, transfer the pre-training weights to the YOLOv8 backbone network, and improve the feature extraction and generalization capability; The model training and deployment module uses mixed precision training to improve efficiency, reduces the computational load and storage requirements through model quantization and pruning, and is finally deployed on an edge device to ensure real-time performance and accuracy.

[0015] The beneficial effects of the present application are: 1. Higher accuracy: the LAF and LDEB modules enhance the detail extraction capability and improve the hook recognition accuracy.

[0016] 2. Strong real-time performance: after lightweight optimization of the model, the frame processing time is shortened by more than 30%, and the model is suitable for edge deployment.

[0017] 3. Strong robustness: can stably identify targets in complex environments such as strong light, shadow, and occlusion.

[0018] 4. Strong scalability: the model structure has strong adaptability and can be quickly migrated to other lifting target recognition tasks. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1A flow chart of a YOLOv8-based head-and-head recognition method is provided for the present application. Figure 2 A module diagram of a YOLOv8-based head-and-head recognition system is provided for the present application. DETAILED DESCRIPTION

[0020] To make the technical means, creative features, purposes and effects of the present application easy to understand, the present application is further described below in conjunction with specific embodiments.

[0021] Please refer to the accompanying Figure 1 A YOLOv8-based head-and-head recognition method, comprising the following steps: S1: Data acquisition and annotation: Use a camera to collect video frames of the head-and-head in different postures, illuminations and backgrounds, and perform fine manual annotation to construct a high-quality data set; S11: In the data acquisition link, the camera collects head-and-head video frames at a frame rate of (unit: frames / second, The value range of is 15-30) to ensure that the dynamic changes of the head-and-head can be fully captured in the time dimension; S12: For data annotation, fine manual annotation is adopted, specifically as follows: Define the annotation accuracy parameter : The number of correctly annotated pixels / the total number of pixels of the head-and-head , which is required to ensure the annotation quality; Introduce an annotation consistency algorithm to calculate the consistency index c of the annotation results of the same video frame by different annotators; c - (the number of annotation differences / the total number of pixels of the head-and-head), when , re-annotation is required until the requirements are met; When constructing a high-quality data set, data enhancement algorithms are used, including random rotation angle ( The value range of is -30° to 30°), random scaling ratio s (the value range of s is 0.8-1.2) and random translation distance d (the value range of d is -20 to 20 pixels), which are used to process the originally collected video frames, expand the data set size, and enhance the recognition ability of the model for head-and-head in different postures.

[0022] S2: Model structure innovation: Based on the original structure of YOLOv8, the following improvements are proposed: lightweight attention mechanism (Lightweight Attention Fusion, LAF) module, local detail enhancement branch (Local Detail Enhancement Branch, LDEB), dynamic anchor adaptation mechanism (Dynamic Anchor Adaptation, DAA), improved loss function; S21: Lightweight Attention Fusion (LAF) module: A lightweight attention module is introduced between Backbone and Neck, which effectively enhances the model's perception ability for small targets (such as hook edges, hooks, and other small details) and suppresses background interference. S211: Define the channel number of the input feature map as C, the width as W, and the height as H, i.e., the input feature map size is ; S212: Introduce convolution operation to calculate channel attention weight, with convolution kernel size , stride , padding , and channel number unchanged; S213: Calculate the output attention weight feature map : , where: is the input feature map, and are the convolution kernel weight and bias parameters, is the activation function; S214: Multiply the attention weight feature map and the input feature map element by element to obtain the enhanced feature map, the calculation formula is: , where represents element-wise multiplication operation.

[0023] S22: Local Detail Enhancement Branch (LDEB): Add an auxiliary branch that focuses on extracting hook edge and connection features, and improves target recognition ability after fusion with the main feature; Add a parallel local enhancement branch to the original YOLOv8 Neck layer, introduce DilatedConvolution (Dilated Convolution) in the branch structure to expand the receptive field without losing resolution; combine with Swin Transformer module to introduce local window attention mechanism to improve local structure expression; finally, through Concat method to fuse into the main feature map, solve the problem of missed detection under hook occlusion, partial blur or rotating posture.

[0024] S221: In the local enhancement branch structure in parallel in the Neck layer, introduce the empty hole convolution, the empty hole rate is defined as d, the empty hole convolution calculation formula is: , wherein: is the empty hole convolution kernel weight, is the empty hole convolution output feature map; S222: Combine the Swin Transformer module, the local window size is defined as , calculate the attention in the local window, the calculation formula is: , wherein: Q and K are query matrix and key matrix, is the dimension of the attention head, is the value matrix, is the local window attention weight matrix; S223: Finally, the local enhancement branch feature map is fused with the backbone feature map by Concat, the calculation formula is: Concat , wherein: is the backbone feature map, is the local enhancement branch output feature map.

[0025] S23: Dynamic anchor adaptive mechanism (Dynamic Anchor Adaptation, DAA): optimize anchor parameters through online learning to improve the adaptability to hooks of different sizes; S231: Define the number of anchors as n _ anchors, the initial anchor parameters are width and height ; S232: Optimize the anchor parameters through online learning, the calculation formula is: , wherein: t is the iteration number, is the learning rate, and are the real width and height estimates of the current batch of hooks, respectively.

[0026] S24: Improve the loss function: introduce the position-sensitive IoU loss (Pos-IoU Loss) to more accurately evaluate the hook center prediction error; S241: Define the position parameters of the predicted box as the center coordinates , width , and height , and the position parameters of the real box are and ; S242: The formula for calculating the position-sensitive IoU loss is: where IoU is the intersection over union of the predicted box and the real box.

[0027] S3: Industrial Pretext Task Pretraining: Construct self-supervised tasks such as "hook direction prediction" and "hook head occlusion prediction"; pretrain on large-scale unlabeled industrial video data, and migrate the pretraining weights to the YOLOv8 backbone network to improve the initial convergence performance and feature expression ability, and improve the model's migration ability and generalization performance in industrial scenarios; S31: Construct self-supervised tasks, taking "hook direction prediction" and "hook head occlusion prediction" as examples; When constructing the self-supervised task "hook direction prediction": S311a: Define the original video frame as l, and perform a rotation operation on l with a rotation angle to generate the rotated image l _ rot; S312a: Construct the rotation label y _ rot, when the rotation angle is _ , y rot corresponds to ; S313a: Design a small Convolutional Neural Network (CNN) model M _ rot to predict the rotation angle, the model structure includes two convolutional layers and one fully connected layer, the convolutional kernel size of each convolutional layer is defined as and , the step size is s1 and s2, the output channel number is c1 and c2, and the number of neurons in the fully connected layer is n; the input of the model M _ rot is I _ rot, and the output is the predicted rotation angle label y _ rot; S314a: Define the loss function as the cross-entropy loss, the calculation formula is: where: is the number of training samples, represents the i-th sample, represents the class of the rotation angle, rot, is the one-hot encoding of the true label, rot, The class probability predicted by the model; When constructing the self-supervised task "hook head occlusion prediction": S311b: Define the original video frame as 1, and use the occlusion block size as The block blocks the hook head area and generates the blocked image I _ occ; S312b: Construct occlusion label y _ occ, when the hook head is blocked, y _ occ is 1, otherwise it is 0; S313b: Design a small CNN model M _ OCC is used to predict whether the hook head is occluded. The model structure includes a convolutional layer and a fully connected layer. The convolution kernel size of the convolutional layer is defined as , the step size is s3, the number of output channels is c3, and the number of neurons in the fully connected layer is n; model M _ The input of occ is I _ occ, the output is the predicted occlusion label y _ occ; S314b: Define loss function is the binary cross entropy loss, and the calculation formula is: ,in: is the number of training samples, Indicates the samples, is the true label, is the occlusion probability predicted by the model; S32: Pre-training on large-scale unlabeled industrial video data: S321: Define the pre-training learning rate lr _ pre, the number of training rounds is epoch _ pre, the batch size of each round of training is batch _ size _ pre; S322: During the training process, by optimizing the loss function of the self-supervised task and , update model parameters; S33: After pre-training is completed, the pre-trained weights are transferred to the YOLOv8 backbone network, and the transfer learning rate is defined as Ir _ transfer, fine-tuning training round number is epoch _ fine _ tune, the batch size of each round of training is batch _ size _ fine _tune; To improve the model's transferability and generalization performance in industrial scenarios, a feature extraction capability improvement indicator is defined , with the formula being: , wherein: represents the feature extraction capability of the YOLOv8 backbone network before transfer, represents the feature extraction capability of the YOLOv8 backbone network after fine-tuning; The feature extraction capability can be evaluated by calculating the similarity between the feature map and the true label on the validation set, with the formula being: where M and N are the height and width of the feature map, F represents the feature map, and L represents the feature representation of the true label.

[0028] Feature extraction evaluation: In addition to calculating the similarity between the feature map and the true label, other evaluation indicators (such as precision, recall, F1-score, etc.) can also be added, which can more comprehensively reflect the performance of the model.

[0029] S4: Model training and deployment optimization: using mixed precision training to improve training efficiency, model quantization and pruning to adapt to real-time deployment on edge devices; S41: Mixed precision training: define the total number of iterations during model training as , the initial learning rate as , and the number of training samples per batch as ; use mixed precision training, set the loss scaling factor to to prevent gradient underflow, with the formula being: where L is the original loss value; when updating the gradient, first calculate the scaled gradient , where W is the model parameter, then scale the gradient back to the normal range: , update the model parameters: , where t is the current iteration number, is the current learning rate, with the formula being: , where power is the power parameter of learning rate decay; S42: Model quantization: quantize the model to convert 32-bit floating-point (FP32) weights to 8-bit integers (INT8); define the quantization parameters and , with the formula being: , wherein: and are the range extremes of the weights, is the quantization bit number (value is 8); the quantized weight is : wherein: original 32-bit floating point number (FP32) weights of the model are shown in Table 1; S43: Model pruning: a weight-based pruning method is adopted, and the pruning rate is defined as prune _ rate (with a value range of 0-1), that is, a proportion of small weights is pruned; the absolute value threshold of the weights is calculated _ percentile , prune _ rate ; the weights with an absolute value less than thr are set to 0; the sparsity of the model after pruning is calculated : ; wherein: is the number of weights with a value of 0 after pruning of the model, that is, the number of pruned weights; is the total number of weights of the model before pruning, that is, the total number of original weight parameters of the model; Pruning strategy: the setting of the pruning rate can be adjusted based on the sparsity requirement of the network. If too much pruning may cause a significant decline in model performance, the pruning process can be adjusted gradually and verified.

[0030] Dynamic pruning: some methods support dynamic pruning, which gradually removes unimportant weights during training.

[0031] S44: Model deployment optimization: define the computing capability parameter of the edge device as compute _ capability, with a unit of GFLOPS; according to the computing capability of the device, the quantized and pruned model is adapted and optimized to ensure that the inference time of the model on the edge device meets the real-time requirement, that is, the inference time , wherein is the set real-time threshold; at the same time, the detection accuracy of the model on the edge device is ensured to decrease by no more than a set tolerance , and the calculation formula is: , wherein: is the detection accuracy of the original model on the validation set, is the detection accuracy of the model deployed on the edge device.

[0032] Hardware adaptation: according to the computing capability (such as memory, computing power, etc.) of the device, appropriate quantization and pruning parameters are selected. Some edge devices may have hardware optimization support for INT8 weights after quantization, which can further accelerate inference.

[0033] ​Accuracy assurance: For the trade-off between accuracy and real-time performance, experiments can be conducted during pruning and quantization to find the optimal balance point between accuracy and speed. Consideration can be given to increasing post-deployment performance monitoring to ensure that real-time performance and accuracy meet the requirements.

[0034] Please refer to Figure 2 A YOLOv8-based crown block hook recognition system includes a data acquisition and labeling module, a model construction and optimization module, a self-supervised pre-training module, and a model training and deployment module. The data acquisition and labeling module uses a high-definition camera to collect crown block hook video frames from multiple angles, labels key parts, and uses data augmentation techniques to expand the data set and improve the model's adaptability to different scenarios. The model construction and optimization module is used to improve YOLOv8, incorporating lightweight attention mechanisms, local detail enhancement branches, dynamic anchor frame adaptive mechanisms, and improved loss functions to enhance the model's ability to perceive and extract features from small targets. The self-supervised pre-training module is used to build a self-supervised task and pre-train on large-scale unlabeled industrial video data, transferring pre-training weights to the YOLOv8 backbone network to improve feature extraction and generalization capabilities. The model training and deployment module uses mixed-precision training to improve efficiency, reduces computational load and storage requirements through model quantization and pruning, and finally deploys on edge devices to ensure real-time performance and accuracy.

[0035] From the above description, it can be seen that the above-mentioned embodiments of the present application achieve the following technical effects: higher accuracy: through the LAF and LDEB modules to enhance the detail extraction capability, improve the hook recognition accuracy.

[0036] Strong real-time performance: After model lightweight optimization, frame processing time is shortened by more than 30%, suitable for edge deployment.

[0037] Strong robustness: can stably identify targets in complex environments such as strong light, shadow, and occlusion.

[0038] Strong scalability: the model structure has strong adaptability and can be quickly migrated to other lifting target recognition tasks.

[0039] Those skilled in the art should understand that the above discussion of any embodiment is only exemplary and is not intended to limit the scope of the present application (including claims) to these examples; under the idea of the present application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the present application as described above. In order to be brief, they are not provided in detail.

[0040] The present application is intended to cover all such alternatives, modifications, and variations as fall within the broad scope of the claims. Accordingly, any and all such alternatives, modifications, equivalents, improvements and the like are intended to be encompassed by the present application.

Claims

1. A method for identifying overhead crane hooks based on YOLOv8, characterized in that: The specific steps include: S1: Data collection and annotation: Use cameras to collect video frames of the overhead crane hook in different postures, lighting and backgrounds, and perform detailed manual annotation to build a high-quality dataset; S2: Model structure innovation: Based on the original YOLOv8 structure, the following improvements are proposed: lightweight attention mechanism module, local detail enhancement branch, dynamic anchor box adaptation mechanism, and improved loss function; S3: Self-supervised pre-training mechanism: Construct self-supervised tasks, perform pre-training on large-scale unlabeled industrial video data, and migrate the pre-trained weights to the YOLOv8 backbone network; S4: Model training and deployment optimization: Use mixed precision training to improve training efficiency, and quantize and prune models to adapt to real-time deployment on edge devices.

2. The method for identifying a crane hook based on YOLOv8 according to claim 1, characterized in that: Said S1 specifically includes: S11: In the data collection phase, the camera uses a frame rate Collect video frames of the overhead crane hook to ensure that the dynamic changes of the overhead crane hook can be fully captured in the time dimension; S12: For data annotation, a refined manual annotation method is used, as follows: Define the annotation accuracy parameters ; Introduce the annotation-consistency algorithm to calculate the consistency index c of the annotation results of different annotators on the same video frame; When building a high-quality dataset, use data enhancement algorithms, including random rotation angle , random scaling ratio s and random translation distance d.

3. The method for identifying a crane hook based on YOLOv8 according to claim 1, characterized in that: The S2 specifically includes: S21: Lightweight Attention Mechanism Module: A lightweight attention module is introduced between Backbone and Neck to effectively enhance the model's ability to perceive small objects; S22: Local detail enhancement branch: A new auxiliary branch is added to focus on extracting hook edge and connection features, which are then integrated with the main features to improve target recognition capabilities. S23: Dynamic anchor frame adaptation mechanism: Optimize anchor frame parameters through online learning to improve adaptability to hooks of different sizes; S24: Improved loss function: Introducing position-sensitive IoU loss to more accurately evaluate the hook center of gravity prediction error.

4. The method for identifying a crane hook based on YOLOv8 according to claim 3, characterized in that: The S21 specifically includes: S211: Define the number of channels of the input feature map as C, width as W, and height as H, that is, the size of the input feature map is ; S212: Introduce convolution operation to calculate channel attention weight, the convolution kernel size is , the step size is , filled with , the number of channels remains unchanged; S213: Calculate the output attention weight feature map ; S214: Multiply the attention weight feature map and the input feature map element by element to obtain an enhanced feature map.

5. The method for identifying a crane hook based on YOLOv8 according to claim 4, characterized in that: The S22 specifically includes: adding parallel local enhancement branches to the original YOLOv8 Neck layer, introducing DilatedConvolution into the branch structure to expand the receptive field without losing resolution; combining with the Swin Transformer module, introducing a local window attention mechanism to improve the expressiveness of the local structure; and finally integrating it into the backbone feature map through the Concat method.

6. The method for identifying a crane hook based on YOLOv8 according to claim 5, characterized in that: The S23 specifically includes: S231: Define the number of anchor boxes as n _ anchors, the initial anchor box parameter is width and height ; S232: Optimize anchor box parameters through online learning. The calculation formula is: , where: t is the number of iterations, is the learning rate, and They are respectively the estimated true width and height of the hooks in the current batch.

7. The method for identifying a crane hook based on YOLOv8 according to claim 6, characterized in that: The S24 specifically includes: S241: Define the position parameters of the prediction box as the center coordinates and width ,high , the position parameters of the real box are and ; S242: The formula for calculating position-sensitive IoU loss is: , where IoU is the intersection over union (IoU) of the predicted box and the true box.

8. The method for identifying an overhead crane hook based on YOLOv8 according to claim 1, characterized in that: The S3 specifically includes: S31: Construct self-supervised tasks, taking "hook direction prediction" and "hook head occlusion prediction" as examples; S32: Pre-training on large-scale unlabeled industrial video data: S321: Define the pre-training learning rate lr _ pre, the number of training rounds is epoch _ pre, the batch size of each round of training is batch _ size _ pre; S322: During the training process, by optimizing the loss function of the self-supervised task and , update model parameters; S33: After pre-training is completed, the pre-trained weights are transferred to the YOLOv8 backbone network, and the transfer learning rate is defined as Ir _ transfer, fine-tuning training round number is epoch _ fine _ tune, the batch size of each round of training is batch _ size _ fine _ tune.

9. The method for identifying an overhead crane hook based on YOLOv8 according to claim 1, characterized in that: The S4 specifically includes: S41: Mixed precision training: Define the total number of iterations during model training as , the initial learning rate is , the number of training samples in each batch is ; Use mixed precision training and set the loss scaling factor to ; When updating the gradient, first calculate the scaled gradient , then shrink the gradient back to the normal range and update the model parameters; S42: Model quantization: quantize the model, convert the 32-bit floating point weights into 8-bit integers, and define the quantization parameters and , the quantized weight is ; S43: Model pruning: Use weight-based pruning method and define pruning rate as prune _ rate, that is, prune _ rate ratio; calculate the absolute value threshold of the weight ; Reset the weights whose absolute value is less than thr to 0; Calculate the sparsity of the pruned model ; S44: Model deployment optimization: Define the computing power parameters of edge devices as compute _ capability; Adapt and optimize the quantized and pruned model based on the computing power of the device to ensure that the inference time of the model on the edge device meets the real-time requirements; at the same time, ensure that the detection accuracy of the model on the edge device does not drop by more than the set tolerance .

10. A system for identifying a crane hook based on YOLOv8, for implementing a method for identifying a crane hook based on YOLOv8 according to any one of claims 1 to 9, characterized in that: It includes data collection and annotation module, model construction and optimization module, self-supervised pre-training module and model training and deployment module; The data acquisition and annotation module uses a high-definition camera to collect video frames of the overhead crane hook from multiple angles and annotate key parts; The model building and optimization module is used to improve YOLOv8 by incorporating a lightweight attention mechanism, a local detail enhancement branch, a dynamic anchor box adaptation mechanism, and an improved loss function; The self-supervised pre-training module is used to construct self-supervised tasks, pre-train on large-scale unlabeled industrial video data, and migrate pre-trained weights to the YOLOv8 backbone network; The model training and deployment module uses mixed precision training to improve efficiency, reduces computational complexity and storage requirements through model quantization and pruning, and is ultimately deployed on edge devices to ensure real-time performance and accuracy.

Citation Information

Patent Citations

  • Improved yolov5-based aerial insulator orientation identification method

    CN115690542A

  • YOLOv7 vehicle identification method based on improved HAT attention mechanism

    CN117253204A

  • Face detection method and system in complex environment based on improved RetinaFace

    CN118196857A

  • Rolov5-based motor vehicle model optimization method

    CN119152343A

  • Improved Yolov10-oriented target detection fusion method

    CN119723272A