A rail-mounted crane visual anti-lifting detection method, system, device and medium

By deploying industrial cameras on both sides of the rail-mounted gantry crane and employing an image processing method that combines YOLOv5 and CNN-Transformer, the problems of blind spots and low recognition accuracy in anti-lifting detection of rail-mounted gantry cranes in ports have been solved, achieving high-precision anti-lifting detection and improving the safety of port operations.

CN122265169APending Publication Date: 2026-06-23QINGDAO PORT INT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO PORT INT CO LTD
Filing Date
2026-03-09
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

In existing technologies, the anti-lifting detection of port rail gantry cranes relies on physical sensors and single visual recognition technology, which has blind spots, low recognition accuracy, and is prone to missed or false judgments, and cannot achieve high-precision all-round safety protection.

Method used

Industrial cameras are symmetrically deployed on the lower crossbeams of the rail-mounted gantry crane on both the sea and land sides. Image data is transmitted to the server via a wired network. The YOLOv5 algorithm and a hybrid CNN-Transformer structure are used for image processing to identify the status of the truck tires off the ground and the alignment of the spreader, triggering an abnormal alarm process.

Benefits of technology

It achieves full-dimensional, no-blind-spot image acquisition, improves the accuracy and reliability of detection, reduces the probability of missed and false detections, and ensures the safety of port operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265169A_ABST
    Figure CN122265169A_ABST
Patent Text Reader

Abstract

The application discloses a rail-mounted crane visual anti-lifting detection method, system, device and medium, mainly relates to the anti-lifting detection technical field, to solve the current port rail-mounted crane anti-lifting detection generally depends on the low precision of physical sensor and single visual identification technology. Including: according to the pixel range of each industrial camera corresponding image data in the lifting height interval and the hanging high altitude interval, cutting the processed image into lifting height slices and hanging high altitude slices according to the pixel range; using YOLOv5 algorithm, detecting the lifting height slice, identifying whether the container truck tire is off the ground and whether the spreader is aligned; using the CNN-Transformer hybrid structure, detecting the hanging high altitude slice, identifying whether the container truck head has a preset abnormal lifting behavior; when it is detected that the container truck tire is off the ground and the spreader is not aligned or the preset abnormal lifting behavior exists, triggering the abnormal alarm process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of anti-lifting detection technology, and in particular to a visual anti-lifting detection method, system, equipment and medium for rail cranes. Background Technology

[0002] In port container handling operations, anti-lifting detection of rail-mounted gantry cranes is a crucial step in ensuring the safety of personnel, equipment, and cargo. Currently, the industry commonly employs a combination of physical sensors and single visual recognition technology, or uses either one separately, for detection. Physical sensors are mostly deployed at key locations such as the rail-mounted gantry crane's spreader and trolley mechanism, detecting parameters such as weight, displacement, and distance to determine if the truck is at risk of being lifted. Single visual recognition technology typically uses a single industrial camera to capture images of the scene, combined with traditional image processing algorithms or simple deep learning algorithms, to initially identify the status of key targets such as the truck's tires and spreader, assisting in completing the anti-lifting monitoring task. Some ports also use video footage obtained from a single camera to assist in manual intervention and achieve basic safety protection.

[0003] However, single-vision recognition technology, which uses only a single industrial camera, has blind spots and cannot fully capture the relative state of the truck and the lifting device. Furthermore, the recognition capabilities of traditional or simple deep learning algorithms are limited, resulting in low accuracy in recognizing subtle states such as truck tires being off the ground or the cab being abnormally raised. Human intervention is susceptible to fatigue and subjective factors, leading to delayed response and a high risk of missed or incorrect judgments. Whether used alone or in combination, it is impossible to achieve high-precision, all-round anti-lifting detection, resulting in significant security vulnerabilities. Summary of the Invention

[0004] This application provides a visual anti-lifting detection method, system, equipment, and medium for rail-mounted gantry cranes, in order to solve the problem that the current anti-lifting detection of rail-mounted gantry cranes in ports generally relies on physical sensors and single visual recognition technology with low accuracy.

[0005] Firstly, this application provides a visual anti-lifting detection method for rail-mounted cranes, the method comprising: Image data is collected by symmetrically deploying industrial cameras on the lower crossbeams of the rail-mounted crane on both the sea and land sides. Image data is transmitted to the server via a wired network; the server then preprocesses the image data to obtain the processed image. Based on the pixel range of the corresponding image data of each industrial camera in the lifting height range and the suspended height range, the processed image is cropped into lifting height slices and suspended height slices according to the pixel range. The YOLOv5 algorithm is used to detect lifting height slices and identify whether the truck tires are off the ground and whether the spreader is aligned. A hybrid CNN-Transformer architecture is used to detect suspended high-altitude slices and identify whether the truck cab exhibits pre-defined abnormal lifting behavior. When it is detected that the truck tires are off the ground and the spreader is not aligned or there is a preset abnormal lifting behavior, the abnormal alarm process is triggered.

[0006] In one implementation of this application, before acquiring image data using industrial cameras symmetrically positioned at the lower crossbeams on both the sea and land sides of the rail-mounted gantry crane, the method further includes: A calibration plate was placed directly below the lifting device, and the Zhang Zhengyou calibration method was used to calibrate the internal and external parameters of each industrial camera, aligning the camera coordinate system with the rail-mounted crane's mechanical coordinate system.

[0007] In one implementation of this application, image data is transmitted to a server via a wired network; the server then preprocesses the image data to obtain a processed image, specifically including: An independent wired network transmission channel is built using a leased Ethernet line. Each industrial camera is connected to an industrial-grade switch via an RJ45 interface, and the switch is connected to the server via fiber optic cable. The server deploys a multi-threaded receiving service, allocates an independent buffer for each industrial camera, and achieves spatiotemporal alignment of each image data through a timestamp synchronization mechanism. After receiving the image data, the server performs distortion correction and coordinate system alignment correction on each frame of the image based on the intrinsic parameter matrix and extrinsic parameter transformation matrix obtained by Zhang Zhengyou calibration method, and obtains the processed image.

[0008] In one implementation of this application, based on the pixel range of the image data corresponding to each industrial camera in the lifting height range and the suspended height range, the processed image is cropped into lifting height slices and suspended height slices according to the pixel range, specifically including: Obtain the pixel range of image data from each industrial camera within the lifting height range and the suspended height range; The processed image is divided into slices according to preset pixel intervals: lifting height intervals generate lifting height slices, and suspended high altitude intervals generate suspended high altitude slices; wherein, the slice boundaries are statically defined by the relative positional relationship between the lifting device and the container truck in the mechanical coordinate system.

[0009] In one implementation of this application, the YOLOv5 algorithm is used to detect lifting height slices and identify whether the truck tires are off the ground and whether the spreader is aligned, specifically including: The image is input into the lightweight YOLOv5s model, and through multi-scale feature fusion and attention mechanism, the output is a set of pixel coordinates of the tire bounding box coordinates and the four preset key points of the hanger latch. After deduplication by nonmaximum suppression, the tire boundary frame coordinates and the four corners of the hanger lock are used to extract the tire center point projection coordinates and the four corners center point coordinates to form a structured inspection result set. By using a preset pixel-to-mechanical coordinate mapping relationship, the projected coordinates of the tire center point are converted into mechanical coordinates; it is determined whether the vertical height in the mechanical coordinates falls within the preset ground clearance range; the horizontal pixel offset between the coordinates of the four corner center points of the latch and the projected coordinates of the tire center point is calculated, converted into actual centimeter offset, and it is determined whether it exceeds the preset alignment distance.

[0010] In one implementation of this application, a CNN-Transformer hybrid structure is used to detect suspended high-altitude slices and identify whether the truck cab exhibits pre-defined abnormal lifting behavior, specifically including: Local texture and edge features are extracted using the CNN backbone in the CNN-Transformer hybrid structure, and multi-layer spatial feature maps are output. The multi-layer spatial feature map is flattened into a sequence and embedded with position encoding. It is then fed into the Transformer encoder, which models the global dependency of the overall shape of the front of the vehicle through a self-attention mechanism and outputs a global semantic vector. The global semantic vector output by the Transformer is fused with the local feature map of the CNN at the channel level to generate a joint representation, which is then compressed into a 128-dimensional feature vector through a fully connected layer. Input the 128-dimensional feature vector into the binary classification head, and output whether the truck head has a preset abnormal lifting behavior.

[0011] Secondly, this application provides a visual anti-lifting detection system for rail-mounted gantry cranes, the system comprising: The acquisition module is used to acquire image data through industrial cameras symmetrically deployed on the lower crossbeams of the rail-mounted gantry crane on both the sea and land sides. The processing module is used to transmit image data to the server via a wired network; the server then preprocesses the image data to obtain the processed image. The cropping module is used to crop the processed image into a lifting height slice and a suspended height slice according to the pixel range of the corresponding image data of each industrial camera in the lifting height range and the suspended height range. The recognition module uses the YOLOv5 algorithm to detect lifting height slices and identify whether the truck tires are off the ground and whether the spreader is aligned; it uses a CNN-Transformer hybrid structure to detect suspended height slices and identify whether the truck cab has any preset abnormal lifting behavior. The triggering module is used to trigger an abnormal alarm process when it is detected that the truck tires are off the ground and the spreader is not aligned or there is a preset abnormal lifting behavior.

[0012] In one implementation of this application, the processing module includes a processing unit. An independent wired network transmission channel is built using a leased Ethernet line. Each industrial camera is connected to an industrial-grade switch via an RJ45 interface, and the switch is connected to the server via fiber optic cable. The server deploys a multi-threaded receiving service, allocates an independent buffer for each industrial camera, and achieves spatiotemporal alignment of each image data through a timestamp synchronization mechanism. After receiving the image data, the server performs distortion correction and coordinate system alignment correction on each frame of the image based on the intrinsic parameter matrix and extrinsic parameter transformation matrix obtained by Zhang Zhengyou calibration method, and obtains the processed image.

[0013] Thirdly, this application provides a visual anti-lifting detection device for rail-mounted gantry cranes, the device comprising: processor; And a memory containing executable code, which, when executed, causes the processor to perform a visual anti-lifting detection method for rail-mounted cranes, as described above.

[0014] Fourthly, this application provides a non-volatile computer storage medium storing computer instructions thereon, which, when executed, implement a visual anti-lifting detection method for rail-mounted cranes as described above.

[0015] As can be seen from the above technical solutions, this application has the following advantages: By symmetrically deploying dual industrial cameras on the seaside and landside lower crossbeams of the rail-mounted crane, the blind spot problem caused by the limited field of view of a single camera is eliminated, enabling full-dimensional, blind-spot-free image acquisition of the relative position of the truck and the spreader. This structural design allows the system to simultaneously acquire complete visual information on the tire status of both sides of the truck and the alignment of the spreader, overcoming the technical shortcomings of traditional monocular vision systems that cannot accurately determine critical risk states such as tire lift-off and spreader misalignment due to their single field of view. Simultaneously, the image data is stably transmitted to the server via a wired network, avoiding the latency and packet loss risks of wireless transmission, ensuring the integrity and real-time nature of the image data, and providing a reliable data foundation for subsequent high-precision analysis.

[0016] By adaptively slicing the processed images into lifting height and suspended height ranges, and employing both YOLOv5 and a CNN-Transformer hybrid structure for layered detection, a refined division of labor and efficient allocation of computational resources for the detection task were achieved. YOLOv5 enables rapid and high-precision identification of the truck tires' ground clearance and the spreader's alignment in the lifting height slice, and its lightweight nature ensures real-time response capabilities in dynamic operating scenarios. Meanwhile, the CNN-Transformer hybrid structure models the spatiotemporal features of abnormal lifting behavior of the truck cab in the suspended height slice, effectively capturing the long-range dependencies between minute vehicle displacements and attitude changes, thus improving the accuracy of identifying concealed risk behaviors. A dual-condition triggering mechanism—that is, triggering an alarm only when "the tires are off the ground and the spreader is not aligned" or "a preset abnormal lifting behavior exists"—filters out false alarm interference from single abnormal signals, reducing the probability of missed and false detections. Attached Figure Description

[0017] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a visual anti-lifting detection method for a rail-mounted crane provided in an embodiment of this application.

[0019] Figure 2 This is a schematic diagram of the internal structure of a visual anti-lifting detection system for rail-mounted cranes provided in an embodiment of this application.

[0020] Figure 3 This is a schematic diagram of the internal structure of a visual anti-lifting detection device for a rail-mounted gantry crane provided in an embodiment of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Those skilled in the art should understand that the embodiments described below are merely preferred embodiments of this disclosure and do not imply that this disclosure can only be implemented through these preferred embodiments. These preferred embodiments are merely used to explain the technical principles of this disclosure and are not intended to limit the scope of protection of this disclosure. Based on the preferred embodiments provided by this disclosure, all other embodiments obtained by those skilled in the art without creative effort should still fall within the scope of protection of this disclosure.

[0023] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0024] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0025] The embodiment provides a visual anti-lifting detection method for rail-mounted cranes, such as... Figure 1 As shown in the embodiments of this application, the method mainly includes the following steps: Step 110: Collect image data by symmetrically deploying industrial cameras on the seaside and landside lower crossbeams of the rail-mounted crane.

[0026] In some embodiments, before acquiring image data using industrial cameras symmetrically positioned at the lower crossbeams on both the sea and land sides of the rail-mounted gantry crane, the method further includes: A calibration plate was placed directly below the lifting device, and the Zhang Zhengyou calibration method was used to calibrate the internal and external parameters of each industrial camera, aligning the camera coordinate system with the rail-mounted crane's mechanical coordinate system.

[0027] It should be noted that the calibration plate uses a high-contrast black and white checkerboard pattern, is made of matte aluminum alloy, and measures 400mm × 400mm with a grid side length of 20mm. It has mounting holes along the edges and is magnetically or bolted to a stable platform directly beneath the lifting device, ensuring no displacement or vibration during calibration. The calibration plate is positioned at the center of the vertical projection of the rail-mounted crane and within the center of the industrial camera's field of view to avoid image edge distortion affecting calibration accuracy.

[0028] The calibration process is automatically triggered before the start of daily operations, when industrial cameras are replaced, or when there are significant changes in ambient lighting. The system controls the uniform illumination of the calibration board area, and the server sequentially acquires multiple frames of calibration board images taken by each camera. The intrinsic and extrinsic parameter transformation matrices for each camera are calculated using the Zhang Zhengyou calibration method. The calibration results are verified for consistency through reprojection error. After confirming that the error is within an acceptable range, the parameters are stored in the server configuration file for subsequent image processing modules to use, achieving precise mapping from pixel coordinates to the rail-mounted crane's mechanical coordinate system.

[0029] Step 120: Transmit the image data to the server via a wired network; the server preprocesses the image data to obtain the processed image.

[0030] This step can be specifically described as follows: An independent wired network transmission channel is built using a leased Ethernet line. Each industrial camera is connected to an industrial-grade switch via an RJ45 interface, and the switch is connected to the server via fiber optic cable. The server deploys a multi-threaded receiving service, allocates an independent buffer for each industrial camera, and achieves spatiotemporal alignment of each image data through a timestamp synchronization mechanism. After receiving the image data, the server performs distortion correction and coordinate system alignment correction on each frame of the image based on the intrinsic parameter matrix and extrinsic parameter transformation matrix obtained by Zhang Zhengyou calibration method, and obtains the processed image.

[0031] It should be further explained that an independent transmission channel is built through a dedicated Ethernet line, and industrial-grade switches and the PTP protocol are used to achieve sub-microsecond time synchronization of image data. This ensures that images acquired by cameras on the sea side and land side are consistent in the temporal dimension, avoiding spatiotemporal misalignment caused by network latency or asynchronous acquisition. The allocation mechanism of the independent buffer prevents multi-channel data contention, ensures the complete reception of image frames, provides a stable and ordered input sequence for subsequent processing, and improves the reliability of the system in high-concurrency operating environments.

[0032] Based on the intrinsic and extrinsic parameter matrices obtained using Zhang Zhengyou's calibration method, distortion correction and coordinate system alignment are performed on the images, uniformly mapping pixel coordinates to the rail-mounted crane's coordinate system. This eliminates geometric distortion introduced by lens distortion and ensures that the images acquired by the two cameras are on the same spatial reference datum. This processing method allows the subsequent YOLOv5 and CNN-Transformer detection modules to directly determine the target state based on physical coordinates, reducing the risk of false detections due to coordinate inconsistencies and improving the determinism and consistency of the detection logic.

[0033] Step 130: Based on the pixel range of the corresponding image data of each industrial camera in the lifting height range and the suspended height range, cut the processed image into lifting height slices and suspended height slices according to the pixel range.

[0034] This step can be specifically described as follows: Obtain the pixel range of image data from each industrial camera within the lifting height range and the suspended height range; The processed image is divided into slices according to preset pixel intervals: lifting height intervals generate lifting height slices, and suspended high altitude intervals generate suspended high altitude slices; wherein, the slice boundaries are statically defined by the relative positional relationship between the lifting device and the container truck in the mechanical coordinate system.

[0035] It should be further explained that spatial region separation for the detection task is achieved by dividing the processed image into slices representing the lifting height and slices representing the suspended height, based on preset pixel intervals. The slice boundaries are statically defined according to the relative positional relationship between the lifting device and the truck in the mechanical coordinate system, ensuring that the image segmentation strictly corresponds to the physical operation stage and avoiding uncertainties introduced by dynamically adjusting the segmentation threshold. This approach allows subsequent algorithms to perform target detection only within a limited area, reducing computational redundancy, improving processing efficiency, and maintaining consistency between the detection logic and the mechanical motion state.

[0036] The lifting height slice focuses on the area near the top of the truck, while the suspended height slice covers the typical spatial range of the truck head during lifting. The coverage areas of both are statically calculated from the calibrated external mechanical structure parameters of the camera, without relying on dynamic adjustments based on real-time target detection results. This design ensures the reproducibility and verifiability of slice division, avoiding slice shifts or missed detections due to changes in image content. It provides a stable and predictable input area for the YOLOv5 and CNN-Transformer modules, enhancing the consistency of the system's behavior judgments across different operating scenarios.

[0037] Step 140: Using the YOLOv5 algorithm, detect the lifting height slice to identify whether the truck tires are off the ground and whether the spreader is aligned; using the CNN-Transformer hybrid structure, detect the suspended height slice to identify whether the truck cab has any preset abnormal lifting behavior.

[0038] In some embodiments, the YOLOv5 algorithm is used to detect lifting height slices and identify whether the truck tires are off the ground and whether the spreader is aligned, specifically including: The image is input into the lightweight YOLOv5s model, and through multi-scale feature fusion and attention mechanism, the output is a set of pixel coordinates of the tire bounding box coordinates and the four preset key points of the hanger latch. After deduplication by nonmaximum suppression, the tire boundary frame coordinates and the four corners of the hanger lock are used to extract the tire center point projection coordinates and the four corners center point coordinates to form a structured inspection result set. By using a preset pixel-to-mechanical coordinate mapping relationship, the projected coordinates of the tire center point are converted into mechanical coordinates; it is determined whether the vertical height in the mechanical coordinates falls within the preset ground clearance range; the horizontal pixel offset between the coordinates of the four corner center points of the latch and the projected coordinates of the tire center point is calculated, converted into actual centimeter offset, and it is determined whether it exceeds the preset alignment distance.

[0039] In some embodiments, a CNN-Transformer hybrid structure is used to detect suspended high-altitude slices and identify whether the truck cab exhibits any pre-defined abnormal lifting behavior, specifically including: Local texture and edge features are extracted using the CNN backbone in the CNN-Transformer hybrid structure, and multi-layer spatial feature maps are output. The multi-layer spatial feature map is flattened into a sequence and embedded with position encoding. It is then fed into the Transformer encoder, which models the global dependency of the overall shape of the front of the vehicle through a self-attention mechanism and outputs a global semantic vector. The global semantic vector output by the Transformer is fused with the local feature map of the CNN at the channel level to generate a joint representation, which is then compressed into a 128-dimensional feature vector through a fully connected layer. Input the 128-dimensional feature vector into the binary classification head, and output whether the truck head has a preset abnormal lifting behavior.

[0040] It should be further explained that a lightweight YOLOv5s model is used to detect lifting height slices. Combined with multi-scale feature fusion and an attention mechanism, this enhances the local response capability to small tire boundaries and key points of the spreader latch. This structure improves pixel-level positioning accuracy without increasing computational burden, ensuring reproducible stability in the output coordinates of the tire center point and the four corners of the latch. Non-maximum suppression effectively eliminates duplicate detection boxes, ensuring that each target outputs only a unique set of coordinates, providing reliable input for subsequent mechanical coordinate transformations and avoiding false positives caused by detection redundancy.

[0041] By fusing local texture features extracted from the CNN backbone with global morphological dependencies modeled by the Transformer encoder at the channel level, a joint representation is generated. This allows the system to retain both local edge variations and overall pose evolution information when recognizing abnormal front-end lifting behavior. This design avoids the limitations of a single feature dimension and reduces the risk of misclassification caused by local occlusion, lighting fluctuations, or background interference. The 128-dimensional feature vector serves as the structured discrimination basis; its dimensions and content are fixed by the network architecture and do not depend on the distribution of external training samples, ensuring that the behavior determination process is traceable and verifiable.

[0042] Based on the pixel-mechanical coordinate mapping relationship, after transforming the tire center point and latch coordinates to a unified mechanical coordinate system, both vertical height judgment and horizontal offset calculation are performed based on static calibration parameters, without introducing dynamic threshold adjustments. This method makes the judgment logic of the two conditions, "tire off the ground" and "spreader aligned," entirely dependent on physical spatial relationships, rather than statistical features of image content, thus improving the certainty and consistency of the judgment results. The binary classification head only performs linear mapping on the joint representation, without introducing a complex nonlinear decision layer, ensuring that the output results of the vehicle front lifting behavior are interpretable and engineering-adjustable.

[0043] Step 150: When it is detected that the truck tires are off the ground and the spreader is not aligned or there is a preset abnormal lifting behavior, the abnormal alarm process is triggered.

[0044] It should be noted that when the tire lift-off determination result and the spreader alignment offset result output by YOLOv5 simultaneously meet the preset conditions, or when the abnormal lifting behavior of the tractor cab output by CNN-Transformer is true, the logic judgment module in the server will trigger an alarm condition. This judgment is based on binary logic operations, and a high-level alarm signal is output only when "tire lift-off and spreader misalignment" is true, or "tractor cab lifting behavior" is true, to avoid false triggering by a single condition. The alarm signal is output to the port operation monitoring system through the server's IO interface, and is simultaneously pushed to the operator terminal and on-site audible and visual alarm devices to ensure timely information transmission.

[0045] After an alarm is triggered, the system enters a locked state and continuously outputs alarm signals until manual confirmation or the operation returns to normal. The system automatically determines whether the alarm conditions are lifted by monitoring subsequent image slices to see if the tires regain contact with the ground, if the spreader latches are realigned, and if the vehicle's front end posture returns to normal. If three consecutive frames fail to meet the alarm triggering conditions, the alarm state is automatically reset to avoid continuous false alarms caused by momentary interference. The entire alarm process is based entirely on the output of existing detection modules, without relying on external manual intervention or additional sensor input, ensuring consistency between the response mechanism and the detection logic loop.

[0046] In addition, this application Figure 2 This application provides a visual anti-lifting detection system for a rail-mounted gantry crane. For example... Figure 2 As shown in the embodiments of this application, the system mainly includes: The acquisition module 210 is used to acquire image data through industrial cameras symmetrically deployed on the lower crossbeams of the rail-mounted gantry crane on both the sea and land sides.

[0047] The processing module 220 is used to transmit image data to the server via a wired network; and to preprocess the image data on the server to obtain a processed image.

[0048] Processing module 220 includes a processing unit. An independent wired network transmission channel is built using a leased Ethernet line. Each industrial camera is connected to an industrial-grade switch via an RJ45 interface, and the switch is connected to the server via fiber optic cable. The server deploys a multi-threaded receiving service, allocates an independent buffer for each industrial camera, and achieves spatiotemporal alignment of each image data through a timestamp synchronization mechanism. After receiving the image data, the server performs distortion correction and coordinate system alignment correction on each frame of the image based on the intrinsic parameter matrix and extrinsic parameter transformation matrix obtained by Zhang Zhengyou calibration method, and obtains the processed image.

[0049] The cutting module 230 is used to cut the processed image into lifting height slices and suspended height slices according to the pixel range of the corresponding image data of each industrial camera in the lifting height range and the suspended height range.

[0050] The recognition module 240 is used to detect lifting height slices using the YOLOv5 algorithm to identify whether the truck tires are off the ground and whether the spreader is aligned; and to detect suspended high-altitude slices using a CNN-Transformer hybrid structure to identify whether the truck cab has any preset abnormal lifting behavior.

[0051] The trigger module 250 is used to trigger an abnormal alarm process when it is detected that the truck tires are off the ground and the spreader is not aligned or there is a preset abnormal lifting behavior.

[0052] The above are method embodiments of this application. Based on the same inventive concept, this application also provides a visual anti-lifting detection device for rail-mounted gantry cranes. Figure 3 As shown, the device includes: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform a visual anti-lifting detection method for a rail-mounted crane as described in the above embodiments.

[0053] Specifically, the server acquires image data using industrial cameras symmetrically positioned on the sea and land sides of the lower crossbeam of the rail-mounted gantry crane. The image data is then transmitted to the server via a wired network. The server preprocesses the image data to obtain processed images. Based on the pixel range of the corresponding image data from each industrial camera within the lifting height and suspended height ranges, the processed images are cropped into lifting height slices and suspended height slices according to the pixel range. The YOLOv5 algorithm is used to detect the lifting height slices, identifying whether the truck tires are off the ground and whether the spreader is aligned. A CNN-Transformer hybrid structure is used to detect the suspended height slices, identifying whether the truck cab exhibits any pre-defined abnormal lifting behavior. When truck tires are detected to be off the ground and the spreader is not aligned, or when pre-defined abnormal lifting behavior is detected, an abnormal alarm process is triggered.

[0054] In addition, this application embodiment also provides a non-volatile computer storage medium storing executable instructions, which, when executed, implement the above-described visual anti-lifting detection method for rail-mounted cranes.

[0055] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual anti-lifting detection method for rail-mounted cranes, characterized in that, The method includes: Image data is collected by symmetrically deploying industrial cameras on the lower crossbeams of the rail-mounted crane on both the sea and land sides. Image data is transmitted to the server via a wired network; the server then preprocesses the image data to obtain the processed image. Based on the pixel range of the corresponding image data of each industrial camera in the lifting height range and the suspended height range, the processed image is cropped into lifting height slices and suspended height slices according to the pixel range. The YOLOv5 algorithm is used to detect lifting height slices and identify whether the truck tires are off the ground and whether the spreader is aligned; a CNN-Transformer hybrid structure is used to detect suspended height slices and identify whether the truck cab has any pre-set abnormal lifting behavior. When it is detected that the truck tires are off the ground and the spreader is not aligned or there is a preset abnormal lifting behavior, the abnormal alarm process is triggered.

2. The visual anti-lifting detection method for rail-mounted cranes according to claim 1, characterized in that, Before acquiring image data using industrial cameras symmetrically positioned at the lower crossbeams on both the sea and land sides of the rail-mounted gantry crane, the method further includes: A calibration plate was placed directly below the lifting device, and the Zhang Zhengyou calibration method was used to calibrate the internal and external parameters of each industrial camera, aligning the camera coordinate system with the rail-mounted crane's mechanical coordinate system.

3. The visual anti-lifting detection method for rail-mounted cranes according to claim 1, characterized in that, Image data is transmitted to the server via a wired network; The image data is preprocessed by the server to obtain the processed image, specifically including: An independent wired network transmission channel is built using a leased Ethernet line. Each industrial camera is connected to an industrial-grade switch via an RJ45 interface, and the switch is connected to the server via fiber optic cable. The server deploys a multi-threaded receiving service, allocates an independent buffer for each industrial camera, and achieves spatiotemporal alignment of each image data through a timestamp synchronization mechanism. After receiving the image data, the server performs distortion correction and coordinate system alignment correction on each frame of the image based on the intrinsic parameter matrix and extrinsic parameter transformation matrix obtained by Zhang Zhengyou calibration method, and obtains the processed image.

4. The visual anti-lifting detection method for rail-mounted cranes according to claim 1, characterized in that, Based on the pixel range of the corresponding image data from each industrial camera within the lifting height and suspended height ranges, the processed images are cropped into lifting height slices and suspended height slices according to the pixel range, specifically including: Obtain the pixel range of image data from each industrial camera within the lifting height range and the suspended height range; The processed image is divided into slices according to preset pixel intervals: lifting height intervals generate lifting height slices, and suspended high altitude intervals generate suspended high altitude slices; wherein, the slice boundaries are statically defined by the relative positional relationship between the lifting device and the container truck in the mechanical coordinate system.

5. The visual anti-lifting detection method for rail-mounted cranes according to claim 1, characterized in that, Using the YOLOv5 algorithm, the system detects lifting height slices to identify whether the truck tires are off the ground and whether the spreader is aligned. Specifically, this includes: The image is input into the lightweight YOLOv5s model, and through multi-scale feature fusion and attention mechanism, the output is a set of pixel coordinates of the tire bounding box coordinates and the four preset key points of the hanger latch. After deduplication by nonmaximum suppression, the tire boundary frame coordinates and the four corners of the hanger lock are used to extract the tire center point projection coordinates and the four corners center point coordinates to form a structured inspection result set. By using a preset pixel-to-mechanical coordinate mapping relationship, the projected coordinates of the tire center point are converted into mechanical coordinates; it is determined whether the vertical height in the mechanical coordinates falls within the preset ground clearance range; the horizontal pixel offset between the coordinates of the four corner center points of the latch and the projected coordinates of the tire center point is calculated, converted into actual centimeter offset, and it is determined whether it exceeds the preset alignment distance.

6. The visual anti-lifting detection method for rail-mounted cranes according to claim 1, characterized in that, A CNN-Transformer hybrid architecture is used to detect suspended high-altitude slices and identify whether the truck cab exhibits any pre-defined abnormal lifting behavior, specifically including: Local texture and edge features are extracted using the CNN backbone in the CNN-Transformer hybrid structure, and multi-layer spatial feature maps are output. The multi-layer spatial feature map is flattened into a sequence and embedded with position encoding. It is then fed into the Transformer encoder, which models the global dependency of the overall shape of the front of the vehicle through a self-attention mechanism and outputs a global semantic vector. The global semantic vector output by the Transformer is fused with the local feature map of the CNN at the channel level to generate a joint representation, which is then compressed into a 128-dimensional feature vector through a fully connected layer. Input the 128-dimensional feature vector into the binary classification head, and output whether the truck head has a preset abnormal lifting behavior.

7. A visual anti-lifting detection system for rail-mounted cranes, characterized in that, The system includes: The acquisition module is used to acquire image data through industrial cameras symmetrically deployed on the lower crossbeams of the rail-mounted gantry crane on both the sea and land sides. The processing module is used to transmit image data to the server via a wired network; the server then preprocesses the image data to obtain the processed image. The cropping module is used to crop the processed image into a lifting height slice and a suspended height slice according to the pixel range of the corresponding image data of each industrial camera in the lifting height range and the suspended height range. The recognition module uses the YOLOv5 algorithm to detect lifting height slices and identify whether the truck tires are off the ground and whether the spreader is aligned; it uses a CNN-Transformer hybrid structure to detect suspended height slices and identify whether the truck cab has any preset abnormal lifting behavior. The triggering module is used to trigger an abnormal alarm process when it is detected that the truck tires are off the ground and the spreader is not aligned or there is a preset abnormal lifting behavior.

8. The visual anti-lifting detection system for rail-mounted cranes according to claim 7, characterized in that, The processing module includes processing units. An independent wired network transmission channel is built using a leased Ethernet line. Each industrial camera is connected to an industrial-grade switch via an RJ45 interface, and the switch is connected to the server via optical fiber. The server deploys a multi-threaded receiving service, allocates an independent buffer for each industrial camera, and uses a timestamp synchronization mechanism to achieve spatiotemporal alignment of image data. After receiving the image data, the server performs distortion correction and coordinate system alignment correction on each frame of the image based on the intrinsic parameter matrix and extrinsic parameter transformation matrix obtained by Zhang Zhengyou calibration method, and obtains the processed image.

9. A visual anti-lifting detection device for rail-mounted cranes, characterized in that, The device includes: processor; And a memory having executable code stored thereon, which, when executed, causes the processor to perform a visual anti-lifting detection method for a rail-mounted crane as described in any one of claims 1-6.

10. A non-volatile computer storage medium, characterized in that, It stores computer instructions, which, when executed, implement a visual anti-lifting detection method for rail-mounted cranes as described in any one of claims 1-6.