Vision-based port machinery equipment moving distance detection method and system

By combining visual positioning technology with GPS calibration, multispectral imaging and machine learning models are used to achieve high-precision positioning of port equipment in harsh environments, solving the problem of inaccurate positioning of traditional GPS in complex environments and improving the operating efficiency and safety of port equipment.

CN120599016APending Publication Date: 2025-09-05FUJIAN ELECTRONIC PORT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510527193.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional GPS technology lacks positioning accuracy and stability in severe weather or complex environments, resulting in inaccurate positioning of port equipment and posing safety risks.

Method used

A vision-based method for detecting the moving distance of port machinery equipment is adopted. Through the camera system and machine learning technology, calibration and error correction are carried out in combination with GPS positioning information. Multispectral imaging modules and machine learning models are used to achieve high-precision positioning under different lighting and weather conditions. SWIR sensors and polarization imaging technology are integrated to improve recognition stability. Multi-source data fusion is carried out in combination with inertial measurement unit data.

Benefits of technology

Achieving millimeter-level equipment positioning accuracy in complex environments improves the operating efficiency and safety of port equipment, and reduces false alarm rates and deployment and operation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599016A_ABST
    Figure CN120599016A_ABST
Patent Text Reader

Abstract

The invention provides a vision-based port machine equipment movement distance detection method and system, and relates to the technical field of computer vision, and the method comprises the following steps: symmetrically deploying industrial camera groups on two sides of a gantry crane, and collecting ground identification images in real time through a multispectral imaging module; radial and tangential distortion coefficients are calculated based on the calibration board, a conversion mapping relation between the bay position number and the image coordinates is established, original point deviation is calibrated by combining GPS positioning data, and a conversion matrix of the bay position number and the physical coordinates is generated; adding a height compensation coefficient to the vertical line landmark based on the trained digital recognition model and the vertical line detection model; a camera synchronously collects images, position number coordinates are extracted through a digital recognition model, a visual positioning result and GPS data are fused according to environment weights, a millimeter-level positioning result is output, and reliable movement distance detection and equipment state monitoring are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, in particular to the field of positioning and monitoring technology of port automation equipment, and specifically to a vision-based port machinery equipment moving distance detection method and system. Background Art

[0002] With the advancement of port automation, gantry cranes and other equipment are increasingly used in yard operations. Accurate equipment positioning is crucial for efficient and safe operations. While traditional GPS technology can provide centimeter-level positioning accuracy, its accuracy and stability can be compromised in harsh weather or complex environments, such as rain, fog, and obstructions, leading to fluctuations in positioning information. Furthermore, positioning errors in GPS systems often pose safety risks to the normal operation of equipment.

[0003] To address the shortcomings of GPS, visual positioning technology is gaining attention as a supplementary or alternative solution. By utilizing cameras and visual processing technology installed on port machinery, the position of equipment can be detected and located in real time, improving operational accuracy and safety. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to propose a vision-based method and system for detecting the moving distance of port machinery equipment. This method, through a camera system and machine learning technology, can achieve high-precision equipment positioning in complex environments such as different lighting and weather conditions, and combine GPS positioning information for calibration and error correction, ultimately providing reliable moving distance detection and equipment status monitoring. The present invention is applied to modern port logistics automation operations, especially for the precise positioning and monitoring of port gantry crane equipment, which can effectively improve the efficiency and safety of yard automation operations.

[0005] To achieve the above object, the present invention provides the following technical solutions: Based on the above objectives, in a first aspect, the present invention provides a method for detecting the moving distance of port machinery equipment based on vision, comprising the following steps: Industrial cameras are deployed symmetrically on both sides of the gantry crane to capture ground marker images in real time using multispectral imaging modules. Calculate the radial and tangential distortion coefficients based on the calibration plate, establish the conversion mapping relationship between the bay number and the image coordinates, calibrate the origin deviation in combination with GPS positioning data, and generate the conversion matrix between the bay number and the physical coordinates; Based on the trained digit recognition model and vertical line detection model, a height compensation coefficient is added to the vertical line landmarks; The camera synchronously captures images, extracts the bay number coordinates through the digital recognition model, locates the center point of the lower edge of the vertical line through the vertical line detection model, and interpolates adjacent landmarks to complete the occluded landmarks; After converting the image coordinates into physical coordinates, the landmark sequence is screened, the median is used to determine the current position, and the Y-axis deviation between the center line of the landmark group and the calibration origin is calculated. If the threshold is exceeded, a track offset alarm is triggered; the visual positioning results are fused with the GPS data according to the environmental weight to output the millimeter-level positioning results.

[0006] As a further solution of the present invention, after deploying the industrial camera group, the pitch angle is adjusted so that the camera field of view covers ≥2 adjacent bay mark marks, and ground mark images under multiple lighting scenes (sunny / cloudy / night) are collected in real time, and digital landmarks containing bay number text and vertical line landmarks containing position offsets are marked.

[0007] As a further solution of the present invention, an industrial camera group integrating visible light and near-infrared spectrum imaging modules is symmetrically deployed on both sides of the gantry crane to automatically switch the imaging mode according to the ambient light intensity.

[0008] As a further solution of the present invention, an industrial camera group is symmetrically deployed on both sides of the gantry crane to integrate a multispectral imaging module, which automatically switches the imaging mode according to the ambient light intensity. The switching logic of the multispectral imaging module is as follows: When the ambient light intensity is greater than 10 4 When the light is lux, near-infrared imaging is enabled to suppress reflections; When the ambient light intensity is less than 50 lux, the visible light and infrared spectra are integrated to enhance the edge features of landmarks through polarization filtering technology; the integrated SWIR sensor penetrates rain and fog interference to improve the recognition rate in low-visibility scenes.

[0009] As a further solution of the present invention, YOLOv5 is used to train a digital recognition model and a vertical line detection model respectively. The digital recognition model is used to identify the shell position number text, and the vertical line detection model is used to locate the center point of the lower edge of the vertical line, and a closed-loop mechanism for misidentification data is established.

[0010] As a further solution of the present invention, the data closed loop mechanism is: When the difference between GPS and visual positioning is greater than 1 dB, it is marked as a misidentified sample; When the GPS is stationary, if the visual fluctuation is greater than 0.2 dB, it is marked as an offset anomaly sample; The abnormal data was added to the training set to iteratively optimize the YOLOv5 model. After three iterations, the model was optimized and the recognition accuracy was improved by ≥12%.

[0011] As a further solution of the present invention, when screening landmark sequences, landmark sequences with a spacing of ≈1 beta and continuous numbers are screened; when outputting millimeter-level positioning results by fusion according to environmental weights, the environmental weight is: GPS weight ≤0.3 in severe weather.

[0012] As a further solution of the present invention, when interpolating adjacent landmarks to complete the occluded landmarks, an adversarial occlusion compensation algorithm is used to process the container occlusion scene. The adversarial occlusion compensation processing steps include: Generate virtual complete images of occluded landmarks through conditional generative adversarial networks (cGAN); The obscured bay number is calculated based on the distance between adjacent landmarks and the continuity of the numbers, and a compensation result with a probability confidence level greater than 95% is generated; at the same time, the ToF depth sensor is triggered to verify the consistency of the 3D spatial coordinates.

[0013] As a further solution of the present invention, when visual positioning results are fused with GPS data according to environmental weights to output millimeter-level positioning results, an improved Kalman filter is used for multi-source data fusion, and inertial measurement unit data is introduced to build a tightly coupled model. The filter state equation is:

[0014] Where, is the system state vector (including position and velocity); It is the IMU measurement input; is the visual / GPS observation value, the noise covariance matrix 、 Dynamically adjust based on environmental sensor data.

[0015] As a further solution of the present invention, the multi-source data fusion further includes: A tightly coupled Kalman filter model is constructed, and IMU angular velocity and acceleration data are introduced into the state equation. The noise covariance matrix is ​​dynamically adjusted based on environmental sensors (rainfall, wind speed). In severe weather, the GPS weight is reduced to 0.3, and the visual weight is increased to 0.7.

[0016] As a further solution of the present invention, if a track deviation alarm is triggered when a threshold is exceeded, the track deviation detection includes the following steps: Calculate the Y-axis deviation between the center line of the landmark group and the calibration origin. If the deviation is greater than 5mm for five consecutive frames of data, an alarm is triggered. Simultaneously, start the fisheye camera array to scan the track seam features, and use the ORB-SLAM3 algorithm to build a sparse semantic map to verify the offset.

[0017] In a second aspect, the present invention further provides a vision-based port machinery equipment movement distance detection system for implementing the above-mentioned port machinery equipment movement distance detection method, the system comprising: Multimodal visual perception array: This array includes a symmetrically arranged industrial camera group, which includes a SWIR camera, a polarization imager, and a ToF depth sensor. Dynamic landmark generation subsystem: uses the ORB-SLAM3 algorithm to extract ground texture features and track joint features to generate a virtual dynamic landmark library; Spatiotemporal Continuous Displacement Modeling Engine: Integrates a bidirectional temporal convolutional network (Bi-TCN) to analyze 200 consecutive frame displacement sequences, predict device motion trends, and compensate for transmission delays. Closed-loop self-calibration module: Set three error thresholds (5mm for instantaneous level, 2cm for trend level, and 1cm for system level) to trigger re-testing, IMU calibration, or full system calibration.

[0018] As a further solution of the present invention, when the dynamic landmark generation subsystem generates a virtual dynamic landmark library, the dynamic landmark generation steps are as follows: A 360-degree fisheye camera array is used to capture ground texture and container edge feature points, and an improved ORB-SLAM3 algorithm is used to cluster and generate virtual dynamic landmarks. A confidence decay mechanism is designed to upgrade stable feature clusters to permanent landmarks, and interfering features are automatically eliminated with an update cycle of ≤30s‌.

[0019] As a further solution of the present invention, the port machinery equipment movement distance detection system also includes a human-computer interaction module, which integrates an AR-HUD display terminal and uses a variable focus optical solution to project 4m-10m depth of field positioning information; supports a wide depth of field projection mode, and displays the bay number, offset alarm and equipment movement trend prediction data in a hierarchical manner.

[0020] As a further solution of the present invention, the closed-loop self-correction module adopts a self-correction mechanism and sets three levels of error thresholds: the instantaneous level (5mm) triggers local re-detection, the trend level (2cm) starts IMU calibration, and the system level (1cm) performs full system calibration; the calibration data is stored through blockchain to ensure traceability credibility.

[0021] In another aspect of the present invention, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, any one of the above-mentioned methods for detecting the moving distance of port machinery equipment based on vision according to the present invention is executed.

[0022] In another aspect of the present invention, a computer-readable storage medium is provided, storing computer program instructions, which, when executed, implement any of the above-mentioned methods for detecting the moving distance of port machinery equipment based on vision according to the present invention.

[0023] Compared with the existing technology, the vision-based port machinery equipment movement distance detection method and system proposed in the present invention has the following beneficial effects: The present invention suppresses reflective interference in strong light through intelligent switching between visible light and near-infrared spectra, and enhances landmark edge features in weak light through multi-spectral fusion, significantly improving recognition stability in harsh environments such as rain, fog, and reflections. It integrates short-wave infrared to penetrate rain and fog interference, and combines polarization imaging to suppress oil reflections on the water surface, thereby improving landmark recognition rates in low-visibility scenes. Through camera collaborative calibration and distortion correction, combined with the lower vertical line center point positioning algorithm, the system error is effectively reduced, and the positioning accuracy is improved to the millimeter level compared to traditional GPS positioning. The Y-axis deviation monitoring of the landmark group center line and the calibration origin is combined with the fisheye camera array and ORB-SLAM3 to build a sparse semantic map, realizing three-dimensional verification of track offset, effectively reducing the false alarm rate, and achieving millimeter-level positioning, high anti-interference and adaptability in complex port environments. At the same time, it reduces deployment and operation and maintenance costs, providing reliable technical support for port automation upgrades.

[0024] These and other aspects of the present application will be more clearly understood in the following description of the embodiments. It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for the exemplary embodiments or related technical descriptions. The drawings are used to provide a further understanding of the present invention and constitute part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the drawings: Figure 1 The present invention is a flowchart of a method for detecting the moving distance of port machinery equipment based on vision.

[0026] Figure 2 This is a schematic diagram of marking landmarks in a vision-based method for detecting the moving distance of port machinery equipment according to an embodiment of the present invention.

[0027] Figure 3 This is a schematic diagram with numbers marked in a method for detecting the moving distance of port machinery equipment based on vision according to an embodiment of the present invention.

[0028] Figure 4 This is a schematic diagram of calibrating camera parameters in a vision-based method for detecting the moving distance of port machinery equipment according to an embodiment of the present invention.

[0029] Figure 5 Schematic diagram of landmark training in a vision-based method for detecting the moving distance of port machinery equipment according to an embodiment of the present invention.

[0030] Figure 6Schematic diagram of digital training in a vision-based method for detecting the moving distance of port machinery equipment according to an embodiment of the present invention.

[0031] Figure 7 This is a schematic diagram of an identification example in a vision-based port machinery equipment movement distance detection method according to an embodiment of the present invention.

[0032] Figure 8 The present invention provides a flowchart of processing antagonistic occlusion compensation in a vision-based port machinery equipment movement distance detection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] Below, the present application is further described in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0034] To make the purpose, technical solutions and advantages of the present invention more clearly understood, the following is a further detailed description of the embodiments of the present invention in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0035] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are intended to distinguish two non-identical entities or non-identical parameters with the same name. Therefore, "first" and "second" are used for convenience of expression only and should not be understood as limitations on the embodiments of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, other steps or units inherent to a process, method, system, product, or device that includes a series of steps or units.

[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0038] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0039] To compensate for the shortcomings of GPS technology, the present invention proposes a vision-based method and system for detecting the movement distance of port machinery equipment. Through a camera system and machine learning technology, it can achieve high-precision equipment positioning in complex environments such as different lighting and weather conditions, and combine GPS positioning information for calibration and error correction, ultimately providing reliable movement distance detection and equipment status monitoring. The present invention is applied to modern port logistics automation operations, especially for the precise positioning and monitoring of port gantry crane equipment, which can effectively improve the efficiency and safety of yard automation operations.

[0040] See also Figure 1 As shown, an embodiment of the present invention provides a method for detecting the moving distance of port machinery equipment based on vision, the method comprising the following steps: Step S10: symmetrically deploy industrial camera groups on both sides of the gantry crane to collect ground marker images in real time through a multispectral imaging module.

[0041] In this step, see Figures 1 to 4 As shown in the figure, after the industrial camera group is deployed, the pitch angle is adjusted so that the camera field of view covers ≥2 adjacent bay markers, and ground marker images under multiple lighting scenes (sunny / cloudy / night) are collected in real time, and digital landmarks containing bay number text and vertical line landmarks containing position offsets are marked.

[0042] In this embodiment, industrial camera groups integrating visible light and near-infrared spectrum imaging modules are symmetrically deployed on both sides of the gantry crane, and the imaging mode is automatically switched according to the ambient light intensity.

[0043] Among them, the industrial camera group integrated multispectral imaging module is symmetrically deployed on both sides of the gantry crane, and the imaging mode is automatically switched according to the ambient light intensity. The switching logic of the multispectral imaging module is: When the ambient light intensity is greater than 10 4 When the light is lux, near-infrared imaging is enabled to suppress reflections; When the ambient light intensity is less than 50 lux, the visible light and infrared spectra are integrated to enhance the edge features of landmarks through polarization filtering technology; the integrated SWIR sensor penetrates rain and fog interference to improve the recognition rate of low-visibility scenes.

[0044] In this embodiment, the position of the hoist camera is first adjusted: cameras are installed on both sides of the gantry crane to ensure that each camera faces the ground and that the two bay markers can be fully visually seen at any time; then a large number of images are taken in various scenes and lighting environments, and the ground markers on the map are marked, distinguishing them into digital landmarks and vertical line landmarks. The training of the two models is divided into more detailed markings of the individual numbers in the digital landmarks; the distortion coefficient of the camera is calculated by marking, with the center being 0 bay, the bay number changing direction being 1, and vice versa. Because the vertical line and the landmark are not actually aligned in a straight line, a vertical line* needs to be added, and the height coefficient of the vertical line* also needs to be saved. The calibrated GPS positioning is input (mainly to obtain the deviation of the middle 0 bay, which can be defaulted to 0 for the center), and then the direct linear method is used to obtain the correspondence between the bay number and the coordinates with the center of the three lower edges as the core.

[0045] Step S20: Calculate radial and tangential distortion coefficients based on the calibration plate, establish a conversion mapping relationship between the Bayer number and the image coordinates, calibrate the origin deviation in combination with GPS positioning data, and generate a conversion matrix between the Bayer number and the physical coordinates.

[0046] In this step, YOLOv5 is used to train the digital recognition model and the vertical line detection model respectively. The digital recognition model is used to identify the bay number text, and the vertical line detection model is used to locate the center point of the lower edge of the vertical line, and a closed-loop mechanism for misidentification data is established.

[0047] In this embodiment, the data closed-loop mechanism is: When the difference between GPS and visual positioning is greater than 1 dB, it is marked as a misidentified sample; When the GPS is stationary, if the visual fluctuation is greater than 0.2 dB, it is marked as an offset anomaly sample; The abnormal data was added to the training set to iteratively optimize the YOLOv5 model. After three iterations, the model was optimized and the recognition accuracy was improved by ≥12%.

[0048] Step S30 : adding a height compensation coefficient to the vertical line landmark based on the trained digit recognition model and vertical line detection model.

[0049] In this step, when screening landmark sequences, select landmark sequences with a spacing of ≈1 beta and continuous numbers; when outputting millimeter-level positioning results by fusion according to the environmental weight, the environmental weight is: GPS weight ≤ 0.3 in severe weather.

[0050] In this embodiment, see Figures 1 to 6 As shown, in the training phase, the large-scale images collected in the previous phase are trained using yolov5. Figure 7As shown, the signs on the ground are identified to obtain the coordinate information of the digital landmarks and vertical line signs in the image. The digital information contained in the digital landmarks is further identified to determine the bay number represented by the sign. According to the correspondence between the bay number and the coordinates calculated in step 3 of the calibration phase, the deviation value corresponding to the center point of the lower edge of each sign is reversely calculated to infer the corresponding position information of the origin, that is, the bay information of the vehicle's current location. In particular, if there is an incomplete recognition, it is also necessary to make certain supplementary calculations using the size of general landmarks, and select the most reasonable set of landmarks with an interval of approximately 1 and a normal order to determine the origin information.

[0051] The Y-axis deviation between the origin and the line containing the landmarks, as determined in the previous step, is calculated to confirm whether the vehicle has deviated from its track. False positives are collected based on the discrepancy between the output and GPS. For example, if the difference between the GPS and the output is greater than 1, it is generally a case of misidentification. If the GPS remains essentially unchanged for a certain period (less than 0.05), a discrepancy of more than 0.2 from the output is also considered an offset.

[0052] Step S40: The camera synchronously captures images, extracts the bay number coordinates through the digital recognition model, locates the center point of the lower edge of the vertical line through the vertical line detection model, and uses adjacent landmark interpolation to complete the blocked landmarks.

[0053] In this step, when the occluded landmarks are interpolated and completed by adjacent landmarks, the adversarial occlusion compensation algorithm is used to handle the container occlusion scene, see Figure 8 As shown in Figure 2, the processing steps for adversarial occlusion compensation include: Step S401: Generate a virtual complete image of the occluded landmark through a conditional generative adversarial network (cGAN); Step S402: Calculate the number of the obscured bay based on the distance between adjacent landmarks and the digital continuity, and generate a compensation result with a probability confidence level greater than 95%. Simultaneously, trigger the ToF depth sensor to verify the consistency of the three-dimensional spatial coordinates.

[0054] Step S50: After converting the image coordinates into physical coordinates, filter the landmark sequence, determine the current position using the median, calculate the Y-axis deviation between the center line of the landmark group and the calibration origin, and trigger a track offset alarm if it exceeds the threshold; fuse the visual positioning results with the GPS data according to the environmental weight to output the millimeter-level positioning results.

[0055] In this step, when visual positioning results and GPS data are fused according to environmental weights to output millimeter-level positioning results, an improved Kalman filter is used for multi-source data fusion, and inertial measurement unit data is introduced to build a tightly coupled model. The filter state equation is:

[0056] Where, is the system state vector (including position and velocity); It is the IMU measurement input; is the visual / GPS observation value, the noise covariance matrix 、 Dynamically adjust based on environmental sensor data.

[0057] In this embodiment, the multi-source data fusion further includes: A tightly coupled Kalman filter model is constructed, and IMU angular velocity and acceleration data are introduced into the state equation. The noise covariance matrix is ​​dynamically adjusted based on environmental sensors (rainfall, wind speed). In severe weather, the GPS weight is reduced to 0.3, and the visual weight is increased to 0.7.

[0058] If the threshold is exceeded, a track deviation alarm is triggered. Track deviation detection includes the following steps: Calculate the Y-axis deviation between the center line of the landmark group and the calibration origin. If the deviation is greater than 5mm for 5 consecutive frames of data, an alarm is triggered. Synchronously start the fisheye camera array to scan the track seam features, and use the ORB-SLAM3 algorithm to build a sparse semantic map to verify the offset.

[0059] Compared with the existing commonly used GPS positioning technology, the present invention shows higher robustness under adverse weather conditions and can significantly improve positioning accuracy, serving as a powerful supplement to GPS technology.

[0060] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0061] It should be understood that, although the above is described in a certain order, these steps are not necessarily performed in sequence according to the above order. Unless clearly stated herein, the execution of these steps does not have strict order restrictions, and these steps can be performed in other orders. Moreover, a part of the steps of the present embodiment may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.

[0062] In a second aspect of the embodiments of the present invention, the present invention further provides a vision-based port machinery equipment movement distance detection system, comprising: Multimodal visual perception array: This array includes a symmetrically arranged industrial camera group, which includes a SWIR camera, a polarization imager, and a ToF depth sensor. Dynamic landmark generation subsystem: uses the ORB-SLAM3 algorithm to extract ground texture features and track joint features to generate a virtual dynamic landmark library; Spatiotemporal Continuous Displacement Modeling Engine: Integrates a bidirectional temporal convolutional network (Bi-TCN) to analyze 200 consecutive frame displacement sequences, predict device motion trends, and compensate for transmission delays. Closed-loop self-calibration module: Set three error thresholds (5mm for instantaneous level, 2cm for trend level, and 1cm for system level) to trigger re-testing, IMU calibration, or full system calibration.

[0063] In this embodiment, when the dynamic landmark generation subsystem generates a virtual dynamic landmark library, the dynamic landmark generation steps are as follows: A 360-degree fisheye camera array is used to capture ground texture and container edge feature points, and an improved ORB-SLAM3 algorithm is used to cluster and generate virtual dynamic landmarks. A confidence decay mechanism is designed to upgrade stable feature clusters to permanent landmarks, and interfering features are automatically eliminated with an update cycle of ≤30s‌.

[0064] In this embodiment, the port machinery equipment movement distance detection system also includes a human-computer interaction module, which integrates an AR-HUD display terminal and uses a variable-focus optical solution to project 4m-10m depth of field positioning information; supports a wide depth of field projection mode, and displays the bay number, offset alarm and equipment movement trend prediction data in a hierarchical manner.

[0065] In this embodiment, the closed-loop self-correction module adopts a self-correction mechanism and sets three levels of error thresholds: the instantaneous level (5mm) triggers local re-detection, the trend level (2cm) starts IMU calibration, and the system level (1cm) performs full system calibration; the calibration data is stored on the blockchain to ensure traceability and credibility.

[0066] Through the above detailed steps, the vision-based port machinery equipment movement distance detection system of the present invention is used to execute the steps of the vision-based port machinery equipment movement distance detection method of the above embodiment, which will not be repeated here. The vision-based port machinery equipment movement distance detection system of the present invention suppresses reflection interference in strong light through intelligent switching between visible light and near-infrared spectra, and enhances landmark edge features in low light by fusing multiple spectra, significantly improving recognition stability in harsh environments such as rain, fog, and reflections. The integrated short-wave infrared penetrates rain and fog interference, and combines with polarization imaging to suppress oil reflections on the water surface, thereby improving landmark recognition rate in low-visibility scenes. Through camera collaborative calibration and distortion correction, combined with the vertical line center point positioning algorithm, the system error is effectively reduced, achieving millimeter-level positioning accuracy compared to traditional GPS positioning. In addition, by monitoring the Y-axis deviation between the center line of the landmark group and the calibration origin, combining the fisheye camera array and ORB-SLAM3 to build a sparse semantic map, three-dimensional verification of track offset is achieved, effectively reducing the false alarm rate. In complex port environments, millimeter-level positioning, high anti-interference and adaptability are achieved, while reducing deployment and operation and maintenance costs, providing reliable technical support for port automation upgrades.

[0067] According to a third aspect of an embodiment of the present invention, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method of any one of the above embodiments is implemented.

[0068] The computer device includes a processor and a memory, and may also include an input system and an output system. The processor, memory, input system, and output system may be connected via a bus or other means. The input system may receive digital or character input and generate signal input related to the migration of the vision-based port machinery equipment movement distance detection. The output system may include a display device such as a display screen.

[0069] As a non-volatile computer-readable storage medium, the memory can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as the program instructions / modules corresponding to the vision-based port machinery equipment movement distance detection method in the embodiment of the present application. The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created by the use of the vision-based port machinery equipment movement distance detection method, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the local module via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0070] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run program code stored in the memory or process data. The processors of the multiple computer devices of the computer device of this embodiment execute various functional applications and data processing of the server by running non-volatile software programs, instructions and modules stored in the memory, that is, implementing the steps of the vision-based port machinery equipment movement distance detection method of the above-mentioned method embodiment.

[0071] It should be understood that, to the extent that they do not conflict with each other, all the embodiments, features and advantages described above for the vision-based port machinery equipment movement distance detection method according to the present invention are also applicable to the vision-based port machinery equipment movement distance detection and storage medium according to the present invention.

[0072] It will also be appreciated by those skilled in the art that the various exemplary logic blocks, modules, circuits and algorithmic steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given of the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or hardware depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the function in various ways for each specific application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the embodiments of the present invention.

[0073] Finally, it should be noted that the computer-readable storage medium (e.g., memory) herein may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. By way of example and not limitation, non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which may act as external cache memory. By way of example and not limitation, RAM is available in various forms, such as synchronous RAM (DRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory devices of the disclosed aspects are intended to include, but are not limited to, these and other suitable types of memory.

[0074] The various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure herein may be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components, designed to perform the functions herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP, and / or any other such configuration.

[0075] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications may be made without departing from the scope of the embodiments disclosed in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present invention may be described or required in individual form, they may also be understood as multiple unless expressly limited to the singular.

[0076] It should be understood that, as used herein, the singular form "a" or "an" is intended to include the plural form as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" refers to any and all possible combinations of one or more of the items listed in association. The serial numbers of the embodiments disclosed in the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0077] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to limit the scope of the disclosure of the present invention (including the claims) to these examples. Within the spirit of the present invention, the technical features of the above embodiments or different embodiments may be combined, and many other variations exist in different aspects of the above embodiments, which are not provided in detail for the sake of clarity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting the moving distance of port machinery equipment based on vision, characterized in that: The method comprises the following steps: Industrial cameras are deployed symmetrically on both sides of the gantry crane to capture ground marker images in real time using multispectral imaging modules. Calculate the radial and tangential distortion coefficients based on the calibration plate, establish the conversion mapping relationship between the bay number and the image coordinates, calibrate the origin deviation in combination with GPS positioning data, and generate the conversion matrix between the bay number and the physical coordinates; Based on the trained digit recognition model and vertical line detection model, a height compensation coefficient is added to the vertical line landmarks; The camera synchronously captures images, extracts the bay number coordinates through the digital recognition model, locates the center point of the lower edge of the vertical line through the vertical line detection model, and interpolates adjacent landmarks to complete the occluded landmarks; After converting the image coordinates into physical coordinates, the landmark sequence is screened, the median is used to determine the current position, and the Y-axis deviation between the center line of the landmark group and the calibration origin is calculated. If the threshold is exceeded, a track offset alarm is triggered; the visual positioning results are fused with the GPS data according to the environmental weight to output the millimeter-level positioning results.

2. The method for detecting the moving distance of port machinery equipment based on vision according to claim 1, characterized in that: After deploying the industrial camera group, adjust the pitch angle so that the camera field of view covers ≥2 adjacent bay markers, collect ground marker images under multiple lighting scenarios in real time, and mark digital landmarks containing bay number text and vertical line landmarks containing position offsets.

3. The method for detecting the moving distance of port machinery equipment based on vision according to claim 2, characterized in that: Industrial camera groups integrated with visible light and near-infrared spectrum imaging modules are symmetrically deployed on both sides of the gantry crane, and the imaging mode is automatically switched according to the ambient light intensity.

4. The method for detecting the moving distance of port machinery equipment based on vision according to claim 3, characterized in that: Industrial camera groups integrated with multispectral imaging modules are symmetrically deployed on both sides of the gantry crane to automatically switch imaging modes based on ambient light intensity. The switching logic of the multispectral imaging module is as follows: When the ambient light intensity is greater than 10 4 When the light is lux, near-infrared imaging is enabled to suppress reflections; When the ambient light intensity is less than 50 lux, the visible light and infrared spectra are fused and the edge features of the landmarks are enhanced through polarization filtering technology; The integrated SWIR sensor penetrates rain and fog interference, improving the recognition rate in low-visibility scenes.

5. The method for detecting the moving distance of port machinery equipment based on vision according to claim 1, characterized in that: YOLOv5 is used to train the digital recognition model and the vertical line detection model respectively. The digital recognition model is used to identify the bay number text, and the vertical line detection model is used to locate the center point of the lower edge of the vertical line, and a closed-loop mechanism for misidentification data is established.

6. The method for detecting the moving distance of port machinery equipment based on vision according to claim 5, characterized in that: The data closed-loop mechanism is: When the difference between GPS and visual positioning is greater than 1 dB, it is marked as a misidentified sample; When the GPS is stationary, if the visual fluctuation is greater than 0.2 dB, it is marked as an offset anomaly sample; The abnormal data was added to the training set to iteratively optimize the YOLOv5 model. After three iterations, the model was optimized and the recognition accuracy was improved by ≥12%.

7. The method for detecting the moving distance of port machinery equipment based on vision according to claim 6, characterized in that: When screening landmark sequences, select landmark sequences with a spacing of ≈1 beta and continuous numbers; when outputting millimeter-level positioning results by fusion according to the environmental weight, the environmental weight is: GPS weight ≤0.3 in severe weather.

8. The method for detecting the moving distance of port machinery equipment based on vision according to claim 1, characterized in that: When interpolating adjacent landmarks to complete the occluded landmarks, an adversarial occlusion compensation algorithm is used to handle container occlusion scenarios. The adversarial occlusion compensation processing steps include: Generate virtual complete images of occluded landmarks through conditional generative adversarial networks; The obscured bay number is calculated based on the distance between adjacent landmarks and the continuity of the numbers, and a compensation result with a probability confidence level of >95% is generated; at the same time, the ToF depth sensor is triggered to verify the consistency of the three-dimensional spatial coordinates.

9. The method for detecting the moving distance of port machinery equipment based on vision according to claim 8, characterized in that: When visual positioning results are fused with GPS data according to environmental weights to output millimeter-level positioning results, an improved Kalman filter is used for multi-source data fusion, and inertial measurement unit data is introduced to build a tightly coupled model. The filter state equation is: Where, is the system state vector; It is the IMU measurement input; is the visual / GPS observation value, the noise covariance matrix 、 Dynamically adjust based on environmental sensor data.

10. A vision-based port machinery equipment moving distance detection system, characterized in that: The system is used to execute the method for detecting the moving distance of port machinery equipment based on vision according to any one of claims 1 to 9, comprising: Multimodal visual perception array: This array includes a symmetrically arranged industrial camera group, which includes a SWIR camera, a polarization imager, and a ToF depth sensor. Dynamic landmark generation subsystem: used to extract ground texture features and track joint features to generate a virtual dynamic landmark library; Spatiotemporal Continuous Displacement Modeling Engine: Integrates a bidirectional temporal convolutional network to analyze 200 consecutive frame displacement sequences, predict device motion trends, and compensate for transmission delays. Closed-loop self-calibration module: Set three levels of error thresholds to trigger re-testing, IMU calibration, or full system calibration.