A near-ground ranging method, medium and device based on virtual large baseline four-eye vision

By using four-eye vision technology to generate a virtual large baseline on a port crane and combining it with a depth estimation neural network, the problems of high cost of lidar and insufficient accuracy of traditional binocular cameras are solved, achieving high-precision and low-cost near-ground ranging and improving the automation level and safety of port operations.

CN121383872BActive Publication Date: 2026-03-20BROAD VISION (XIAMEN) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511960421.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-20
Estimated Expiration
2045-12-24

AI Technical Summary

Technical Problem

In existing port crane ranging solutions, lidar is expensive and has a short lifespan, while traditional binocular cameras have insufficient long-distance ranging accuracy due to their short physical baseline, which cannot meet the requirements of precise operation.

Method used

By using four cameras installed on the high-altitude gantry and hoist, a virtual large baseline is generated through calibration relationships, and image fusion and stereo correction are performed. Combined with a depth estimation neural network model, near-ground distance measurement is achieved.

Benefits of technology

It improves the accuracy of long-distance ranging, reduces costs, and enhances the reliability and safety of ranging, making it suitable for automated port operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121383872B_ABST
    Figure CN121383872B_ABST
Patent Text Reader

Abstract

The application discloses a near-ground ranging method based on virtual large baseline four-eye vision, a medium and equipment, the method synchronously acquires four images when a spreader enters a near-ground operation range; based on a calibration relationship between cameras, high-position images are respectively projected and transformed to a visual angle of a corresponding spreader camera and fused to generate two virtual camera images, so that a virtual large baseline far beyond a physical distance is constructed; virtual image pairs are subjected to stereo correction and cutting to obtain row-aligned stereo image pairs; feature matching and triangulation are performed on the virtual image pairs to generate a sparse reference depth map; the stereo image pairs and the sparse depth map are jointly input into a pre-trained depth estimation neural network model to output a dense depth map, and the vertical distance is obtained in a coordinate system with the spreader as an origin. The application utilizes existing cameras to virtually synthesize an ultra-large baseline, improves long-distance ranging precision, and is low in cost and high in reliability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of port automation, and particularly relates to a near-ground ranging method based on virtual large-baseline four-eye vision, a medium and an equipment. BACKGROUND

[0002] In modern port operations, the spreader of a shore-based container crane needs to accurately perceive the vertical distance between itself and the target below (such as a container or a vehicle), which is directly related to the operation efficiency, safety and automation level.

[0003] At present, the following camera systems are usually installed on port cranes:

[0004] (1) High camera group: installed at a high position (first preset height) of the crane gantry structure, containing two cameras (first camera and second camera). These two cameras have a wide field of view and can overlook the entire operation area.

[0005] (2) Spreader camera group: installed on the spreader that moves vertically with the spreader (second preset height, lower than the first preset height), containing two cameras (third camera and fourth camera).

[0006] However, the existing ranging schemes have significant defects:

[0007] (1) Laser radar scheme: although it has high accuracy, the device cost is high, and in the harsh environment of high salt, high humidity and high vibration in the port, the service life is short and the maintenance cost is extremely high.

[0008] (2) Traditional binocular camera scheme: limited by the physical installation space on the spreader, its baseline length is usually only 0.1-0.3 meters. According to the stereo vision depth error calculation formula, when the working distance reaches dozens of meters, the tiny parallax measurement error will be amplified sharply, resulting in a serious lack of long-distance ranging accuracy, which cannot meet the requirements of accurate operations.

[0009] Therefore, how to make full use of the existing and widely deployed camera hardware of the crane, overcome the limitation of physical space on the baseline, and realize a low-cost, high-precision and high-reliability long-distance vertical ranging scheme has become a technical problem to be solved in the field. SUMMARY

[0010] In view of the above problems, the present application provides a near-ground ranging technical scheme based on virtual large-baseline four-eye vision to solve the technical problems of high cost and short service life of laser radar, and the serious lack of long-distance ranging accuracy caused by the short physical baseline of traditional binocular cameras.

[0011] To achieve the above object, in a first aspect, the application provides a near-ground ranging method based on virtual large-baseline four-eye vision, applied to a container crane, the crane comprising a portal structure and a spreader vertically movable along the portal structure, the portal structure being provided with a first camera and a second camera at a first preset height, and the spreader being provided with a third camera and a fourth camera at a second preset height, the first preset height being higher than the second preset height;

[0012] The method comprises:

[0013] S1: when the spreader enters a preset near-ground operation range, synchronously triggering and acquiring a first image collected by the first camera, a second image collected by the second camera, a third image collected by the third camera and a fourth image collected by the fourth camera;

[0014] S2: based on a calibration relationship between the first camera and the third camera, projecting and transforming the first image to the visual angle of the third camera, and fusing with the third image to generate a first virtual camera image, and based on a calibration relationship between the second camera and the fourth camera, projecting and transforming the second image to the visual angle of the fourth camera, and fusing with the fourth image to generate a second virtual camera image;

[0015] S3: performing stereo correction on the first virtual camera image and the second virtual camera image to obtain a pair of corrected images in line alignment; cropping a common field of view region from the pair of corrected images to obtain a first visual angle image and a second visual angle image in the same size and in line alignment;

[0016] S4: performing feature point matching and triangulation on the first virtual camera image and the second virtual camera image to obtain coordinates and corresponding depth values of a plurality of feature points; according to the coordinate mapping relationship between the plurality of feature points in the first virtual camera image and the first visual angle image, filling the depth values into a two-dimensional matrix in the same size as the first visual angle image to generate a sparse reference depth map in pixel alignment with the first visual angle image, wherein the matrix elements without filling the depth values are empty or zero;

[0017] S5: jointly inputting the first visual angle image, the second visual angle image and the sparse reference depth map into a pre-trained depth estimation neural network model, and outputting a dense depth map in pixel alignment with the first visual angle image from the model;

[0018] S6: based on the depth values of each pixel in the dense depth map, the intrinsic matrix of the first virtual camera, the geometric relationship between the first virtual camera and the third camera and the installation pose of the third camera relative to the spreader, converting the coordinates corresponding to each pixel to a coordinate system with the spreader as the origin, and taking the vertical component of the coordinates as the vertical distance from the pixel to the spreader, to obtain the near-ground ranging result.

[0019] Further, the projecting and transforming the first image to the perspective of the third camera based on the calibration relationship between the first camera and the third camera comprises: calculating a first homography matrix H1 between the first camera and the third camera, and projecting and transforming the first image to the perspective of the third camera based on the first homography matrix H1. The first perspective transformation of the first image is based on the first homography matrix H1. The calculation formula of the first homography matrix H1 is as follows:

[0020] ;

[0021] Wherein, G1 represents the first camera, A1 represents the third camera, h represents the height when the spreader enters the preset near-ground operation range, K A1 represents the intrinsic matrix of the third camera, represents the rotation matrix of the first camera to the third camera at the height h, represents the translation vector from the origin of the coordinate system of the first camera to the origin of the coordinate system of the third camera at the height h, and n represents the unit normal vector of the preset reference plane. represents the vertical distance from the origin of the world coordinate system to the preset reference plane, represents the inverse matrix of the intrinsic matrix of the first camera.

[0022] The projecting and transforming the second image to the perspective of the fourth camera based on the calibration relationship between the second camera and the fourth camera comprises: calculating a second homography matrix H2 between the second camera and the fourth camera, and projecting and transforming the second image to the perspective of the fourth camera based on the second homography matrix H2. The second perspective transformation of the second image is based on the second homography matrix H2. The calculation formula of the second homography matrix H2 is as follows:

[0023] ;

[0024] Wherein, G2 represents the second camera, A2 represents the fourth camera, K A2 represents the intrinsic matrix of the fourth camera, represents the rotation matrix of the second camera to the fourth camera at the height h, represents the translation vector from the origin of the coordinate system of the second camera to the origin of the coordinate system of the fourth camera at the height h, represents the inverse matrix of the intrinsic matrix of the second camera.

[0025] Further, the first virtual camera image is generated in the following manner:

[0026] Aligning the first image after the first perspective transformation with the original image of the third image in the image space to determine the overlapping area between the two images;

[0027] In the overlapping region, corresponding pixels of the first perspective-transformed first image and the original image of the third image are fused pixel by pixel by using a weight-based fusion algorithm to generate the first virtual camera image with an extended field of view;

[0028] The second virtual camera image is generated according to the following manner:

[0029] The second perspective-transformed second image and the original image of the fourth image are aligned in the image space to determine an overlapping region between the two images;

[0030] In the overlapping region, corresponding pixels of the second perspective-transformed second image and the original image of the fourth image are fused pixel by pixel by using a weight-based fusion algorithm to generate the second virtual camera image with an extended field of view.

[0031] Further, the weight-based fusion algorithm includes a multi-band fusion algorithm;

[0032] The weight-based fusion algorithm is used to fuse pixels pixel by pixel to generate the first virtual camera image with an extended field of view, specifically including:

[0033] The first perspective-transformed first image and the original image of the third image are respectively subjected to Laplacian pyramid decomposition;

[0034] In each layer of the Laplacian pyramid, a first fusion weight map is calculated according to the position of the pixel in the overlapping region;

[0035] The pyramid coefficients of the corresponding layers of the first image and the original image of the third image are weighted and fused using the first fusion weight map;

[0036] The weighted and fused Laplacian pyramid is reconstructed to generate the first virtual camera image with an extended field of view;

[0037] The weight-based fusion algorithm is used to fuse pixels pixel by pixel to generate the second virtual camera image with an extended field of view, specifically including:

[0038] The second perspective-transformed second image and the original image of the fourth image are respectively subjected to Laplacian pyramid decomposition;

[0039] In each layer of the Laplacian pyramid, a second fusion weight map is calculated according to the position of the pixel in the overlapping region;

[0040] The pyramid coefficients of the corresponding layers of the second image and the original image of the fourth image are weighted and fused using the second fusion weight map;

[0041] reconstructing the weighted fused Laplacian pyramid to generate the second virtual camera image with an extended field of view.

[0042] Further, performing stereo rectification on the first virtual camera image and the second virtual camera image to obtain a pair of rectified images aligned in rows includes the following steps:

[0043] S31: calculating stereo rectification parameters based on camera parameters associated with the first virtual camera image and the second virtual camera image; the camera parameters include the relative pose between the first virtual camera and the second virtual camera and the respective intrinsic matrix;

[0044] S32: using the calculated stereo rectification parameters to perform geometric transformation on the first virtual camera image and the second virtual camera image respectively to obtain a pair of rectified images aligned in rows, and the calculation formula of the geometric transformation is as follows:

[0045] ;

[0046] ;

[0047] wherein, (x, y) represents the pixel coordinates of the first virtual camera image before rectification, (x', y') represents the pixel coordinates of the first virtual camera image after rectification, (x'', y'') represents the pixel coordinates of the second virtual camera image before rectification, (x''', y''') represents the pixel coordinates of the second virtual camera image after rectification, , , , , represents the inverse matrix of the intrinsic matrix of the first virtual camera, represents the inverse matrix of the intrinsic matrix of the second virtual camera, represents the rectification rotation matrix of the first virtual camera, represents the rectification rotation matrix of the second virtual camera,

[0048] Further, performing feature point matching and triangulation on the first virtual camera image and the second virtual camera image to obtain the coordinates of a plurality of feature points and the corresponding depth values includes:

[0049] extracting feature points and generating corresponding feature descriptors on the first virtual camera image and the second virtual camera image respectively; ​​​​​​​

[0050] The feature descriptors are used to match feature points in the first virtual camera image and the second virtual camera image to obtain an initial set of matching point pairs.

[0051] The initial set of matching points is filtered to obtain a set of matching points with high confidence. Any high-confidence matching point in the first virtual camera image is denoted as p1, and the coordinates of p1 are (u1, v1). The corresponding matching point in the second virtual camera image is denoted as p2, and the coordinates of p2 are (u2, v2).

[0052] For each pair of matching points (p1, p2) in the set of high-confidence matching points, a homogeneous linear equation system is constructed. The singular value decomposition algorithm is used to solve the homogeneous linear equation system to obtain the coordinates P=(X,Y,Z) of the matching points (p1, p2) in three-dimensional space. The homogeneous linear equation system is expressed as follows:

[0053] ;

[0054] in, Represents the projection matrix The Okay, M v1 and M v2 M represents the projection matrices of the first virtual camera and the second virtual camera, respectively. v1 and M v2 The formula for representing is as follows:

[0055] ;

[0056] ;

[0057] in, This represents the intrinsic parameter matrix of the first virtual camera. This represents the intrinsic parameter matrix of the second virtual camera. This represents the extrinsic parameter matrix when the first virtual camera coordinate system is used as the world coordinate system. This represents the relative pose matrix that transforms a point from the first virtual camera coordinate system to the second virtual camera coordinate system;

[0058] Transform the coordinates P=(X,Y,Z) into the camera coordinate system of the first virtual camera to obtain the coordinates P v1 The conversion formula is as follows: ,in, Let these be the coordinates of the optical center of the first virtual camera. This represents the rotation matrix that transforms coordinate P to the camera coordinate system of the first virtual camera;

[0059] Take P v1Z-axis component of the pixel as the depth value Z corresponding to the matching point p1 depth ;

[0060] Iterate all high-confidence matching point pairs to obtain the coordinates (u1, v1) of the plurality of feature points in the first virtual camera image and the corresponding depth value Z of the plurality of feature points. depth The set of Z.

[0061] Further, step S6 specifically comprises:

[0062] S61: Based on the depth value of each pixel in the dense depth map and the intrinsic matrix of the first virtual camera, each pixel is back-projected into the coordinate system of the first virtual camera to obtain a first set of three-dimensional point clouds P d ;

[0063] S62: Based on the geometric relationship between the first virtual camera and the third camera, the first set of three-dimensional point clouds P d is converted into the coordinate system of the third camera to obtain a second set of three-dimensional point clouds P c ;

[0064] S63: Based on the mounting pose of the third camera relative to the spreader, the second set of three-dimensional point clouds P c is converted into the coordinate system with the spreader as the origin to obtain a third set of three-dimensional point clouds P h ;

[0065] S64: Extract the vertical direction coordinate component of each point in the third set of three-dimensional point clouds P h as the vertical distance from the corresponding pixel to the spreader to obtain the near-ground ranging result.

[0066] Further, the depth estimation neural network model is obtained by training in the following manner:

[0067] S81: Construct a training sample set, each sample in the training sample set includes a stereo image pair composed of a first training perspective image and a second training perspective image, and a corresponding sparse training depth map; wherein the sparse training depth map is generated by randomly sampling or noise simulation on a dense ground truth depth map aligned with the pixels of the first training perspective image;

[0068] S82: Input the training sample set into the depth estimation neural network model, and the model performs the following operations:

[0069] Feature extraction and stereo matching are performed on the stereo image pair to generate an initial depth estimate;

[0070] The sparse training depth map is used as guide information to fuse with the initial depth estimate during the stereo matching process or subsequent optimization process.

[0071] output different stages or different resolution depth maps, including a first dense depth map and a second dense depth map; wherein the first dense depth map is an intermediate coarse depth map, and the second dense depth map is a final output fine depth map;

[0072] S83: For the first dense depth map, calculate the difference between the first dense depth map and a first reference dense depth map to obtain a first loss L coarse , the first reference dense depth map is obtained by spatially down-sampling a second reference dense depth map;

[0073] For the second dense depth map, calculate the difference between the second dense depth map and a second reference dense depth map, a second loss L final , the second reference dense depth map is the dense ground truth depth map;

[0074] For the sparse pixel points with valid depth values in the sparse training depth map, calculate the difference between the predicted depth value of the second dense depth map on the corresponding pixel point and the sparse ground truth depth value, a third loss L sparse ;

[0075] Based on the first loss L coarse , the second loss L final and the third loss L sparse , a total loss function is constructed, and the formula of the total loss function is as follows:

[0076] L total =L final +α·L coarse +β·L sparse ;

[0077] Wherein, L total represents the total loss, and α and β are preset weight coefficients;

[0078] S84: Use the total loss L total to update the parameters of the depth estimation neural network model through a back propagation algorithm;

[0079] S85: Repeat steps S82 to S84 until the model converges, and obtain a pre-trained depth estimation neural network model.

[0080] In a second aspect, the present application provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method for near-ground ranging based on virtual large baseline four-vision according to the first aspect of the present application.

[0081] In a third aspect, the present application provides an electronic device having a computer program stored thereon, comprising a processor and a storage medium, wherein the storage medium has a computer program stored thereon, and the computer program, when executed by the processor, implements the virtual large-baseline four-camera vision-based near-ground ranging method according to the first aspect of the present application.

[0082] Unlike the prior art, the above technical solution relates to a virtual large-baseline four-camera vision-based near-ground ranging method, medium and device. The method uses a first camera and a second camera installed on a high gantry, and a third camera and a fourth camera installed on a moving hoist. When the hoist enters the near-ground working range, four images are acquired synchronously. First, based on the calibration relationship between the cameras, the high images are projected and transformed to the visual angle of the corresponding hoist cameras, and are fused with the hoist images to generate two virtual camera images, thereby constructing a virtual large baseline far beyond the physical installation distance. Then, the virtual image pair is subjected to stereo rectification and cropping to obtain a row-aligned stereo image pair. Next, feature matching and triangulation are performed on the virtual image pair to generate a sparse reference depth map aligned with one of the stereo image pair. Finally, the stereo image pair and the sparse depth map are jointly input into a pre-trained depth estimation neural network model to output a dense depth map, and the vertical distance is obtained in the coordinate system with the hoist as the origin. The present application ingeniously uses existing cameras to virtually synthesize an ultra-large baseline, which fundamentally improves the long-distance ranging accuracy, while having the advantages of low cost and high reliability.

[0083] The above summary of the invention is only a summary of the technical solutions of the present application. In order for those skilled in the art to more clearly understand the technical solutions of the present application, and to be able to implement the contents recorded in the specification and drawings, and in order for the above and other purposes, features and advantages of the present application to be more easily understood, the following will be described in conjunction with the specific embodiments of the present application and the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0084] The accompanying drawings are only used to illustrate the principles, implementation manners, applications, characteristics and effects of the specific embodiments of the present application and other related contents, and cannot be considered as limitations of the present application.

[0085] In the drawings of the specification:

[0086] Figure 1 A flowchart of the virtual large-baseline four-camera vision-based near-ground ranging method according to an exemplary embodiment of the present application;

[0087] Figure 2 A flowchart of the first virtual camera image generation method according to an exemplary embodiment of the present application;

[0088] Figure 3A flowchart of a second virtual camera image generation method related to an exemplary embodiment of the present application;

[0089] Figure 4 A flowchart of a method of generating a first virtual camera image based on a weight-based fusion algorithm related to an exemplary embodiment of the present application;

[0090] Figure 5 A flowchart of a method of generating a second virtual camera image based on a weight-based fusion algorithm related to an exemplary embodiment of the present application;

[0091] Figure 6 A flowchart of a stereoscopic correction method of a first virtual camera image and a second virtual camera image related to an exemplary embodiment of the present application;

[0092] Figure 7 A flowchart of a feature point matching and triangulation method performed on a first virtual camera image and a second virtual camera image related to an exemplary embodiment of the present application;

[0093] Figure 8 A flowchart of a near-ground ranging method based on a depth value of each pixel in a dense depth map, an intrinsic matrix of a first virtual camera, and the like related to an exemplary embodiment of the present application;

[0094] Figure 9 A schematic diagram of an electronic device according to an exemplary embodiment of the present application;

[0095] The reference signs involved in the above-mentioned figures are explained as follows:

[0096] 10, electronic device; 101, processor; 102, storage medium. DETAILED DESCRIPTION

[0097] To explain the possible application scenarios, technical principles, specific implementable schemes, and the purposes and effects that can be achieved of the present application in detail, the embodiments described herein are explained in detail below in combination with the specific embodiments listed and the accompanying drawings. The embodiments described herein are only used to more clearly explain the technical schemes of the present application, and therefore cannot be used to limit the protection scope of the present application.

[0098] In a first aspect, the present application provides a near-ground ranging method based on virtual large-baseline four-eye vision, applied to a container crane, the crane comprising a portal structure and a spreader vertically movable along the portal structure, the portal structure being provided with a first camera and a second camera at a first preset height, the spreader being provided with a third camera and a fourth camera at a second preset height, the first preset height being higher than the second preset height;

[0099] As shown in Figure 1 the method comprises:

[0100] S1: when the spreader enters the preset near-ground operation range, synchronously trigger and acquire a first image collected by a first camera, a second image collected by a second camera, a third image collected by a third camera, and a fourth image collected by a fourth camera;

[0101] S2: based on a calibration relationship between the first camera and the third camera, project and transform the first image to the perspective of the third camera, fuse with the third image, generate a first virtual camera image, and based on a calibration relationship between the second camera and the fourth camera, project and transform the second image to the perspective of the fourth camera, fuse with the fourth image, generate a second virtual camera image;

[0102] S3: perform stereo rectification on the first virtual camera image and the second virtual camera image to obtain a pair of rectified images that are row-aligned; crop a common field of view region from the pair of rectified images to obtain a first perspective image and a second perspective image that are the same in size and row-aligned;

[0103] S4: perform feature point matching and triangulation on the first virtual camera image and the second virtual camera image to obtain coordinates of a plurality of feature points and corresponding depth values; according to the coordinate mapping relationship between the plurality of feature points in the first virtual camera image and the first perspective image, fill the depth values into a two-dimensional matrix that is the same in size as the first perspective image to generate a sparse reference depth map that is pixel-aligned with the first perspective image, wherein the matrix elements without filled depth values are empty or zero;

[0104] S5: input the first perspective image, the second perspective image, and the sparse reference depth map into a pre-trained depth estimation neural network model, and output a dense depth map that is pixel-aligned with the first perspective image from the model;

[0105] S6: based on the depth values of each pixel in the dense depth map, the intrinsic matrix of the first virtual camera, the geometric relationship between the first virtual camera and the third camera, and the installation pose of the third camera relative to the spreader, convert the coordinates corresponding to each pixel to a coordinate system with the spreader as the origin, and take the vertical component of the coordinates as the vertical distance from the pixel to the spreader to obtain the near-ground ranging result.

[0106] In step S1, when the height h of the spreader fed back by the crane PLC system is less than a preset threshold value h PLCWhen the distance is less than or equal to 5 meters (i.e., entering the near-ground operation range), a synchronous acquisition mechanism is triggered. The mechanism is linked with the signals of the PLC and each camera to ensure that the first camera (G1), the second camera (G2), the third camera (A1), and the fourth camera (A2) capture a frame of static image at the same time, obtaining the first image, the second image, the third image, and the fourth image, respectively. The core purpose of synchronous acquisition is to avoid scene changes caused by time differences in acquisition and to ensure that the physical scenes corresponding to the four images are completely consistent, providing a unified data source in time and space for subsequent image fusion and stereo matching.

[0107] In step S2, the first camera and the third camera, and the second camera and the fourth camera need to be calibrated in advance. The third camera and the fourth camera are fixed on the spreader to form a rigid binocular system, and the relative pose (rotation matrix and translation vector) is constant. Preferably, the length of the translation vector is about 2.8 meters. The first camera and the second camera are zooming pan-tilt cameras. For each discrete height h (such as 5, 4, 3, 2, 1, and 0.5 meters) of the near-ground operation, the corresponding pan angle and focal length are preset, and the internal and external parameters of each point are calibrated and stored to form a lookup table for real-time calling.

[0108] Then, based on the homography matrix, the high camera image is projected and transformed to the view angle of the spreader camera, and the high camera image after projection and transformation is weighted and fused with the original image of the spreader camera in the overlapping area. The fusion algorithm can use multi-band fusion or Poisson fusion, etc. to ensure that the generated virtual camera image (first virtual camera image and second virtual camera image) has an extended and smooth transition view, combining the advantages of the view of both cameras and the detail clarity.

[0109] In step S3, based on the relative pose and internal parameter matrix of the first virtual camera and the second virtual camera, the Bouguet stereo correction algorithm is used to calculate the correction rotation matrix and the unified internal parameter matrix. The two virtual camera images are re-projected to a common image plane through geometric transformation, so that the projection points of the same physical point in the two images are located on the same horizontal line (epipolar line horizontal), and the two-dimensional search for feature matching is simplified to one-dimensional search. After correction, the maximum overlapping observation area is determined by calculating the intersection of the effective pixels of the two corrected images. The first view image and the second view image are obtained by cropping a rectangular area of a fixed size (such as 640x480 pixels) with the spreader projection center as the reference. The two images have the same size and pixel rows are strictly aligned, meeting the standardized input requirements of stereo matching.

[0110] In step S4, feature points are extracted on the first virtual camera image and the second virtual camera image using a feature point detection algorithm such as ORB, SIFT or SURF, brute force matching is performed through a feature descriptor, and high-confidence matching point pairs are screened through a ratio test (typical threshold 0.75) and cross-validation. Then, based on the projection matrices of the first virtual camera and the second virtual camera, a homogeneous linear equation set is constructed using the triangulation principle of stereo vision, the three-dimensional space coordinates of the matching point pairs are solved through singular value decomposition (SVD), and the depth values of the feature points are calculated. This process relies on the high depth sensitivity of the virtual large baseline (about 2.8 meters) to ensure the geometric accuracy of the depth values. Then, through homography transformation or direct projection calculation, the coordinate mapping relationship between the feature points in the first virtual camera image and the first perspective image is established, the depth values of each feature point are filled into a two-dimensional matrix with the same size as the first perspective image, and the unfilled area is set to zero, forming a sparse reference depth map, which provides reliable geometric prior for the subsequent neural network.

[0111] In step S5, the depth estimation neural network model is an end-to-end architecture optimized for a specific near-ground height h, and the input is the first perspective image (grayscale image), the second perspective image (grayscale image) and the sparse reference depth map. The model fuses image features and depth prior features through a multi-modal feature extraction network, constructs a 4D cost space and is regularized by 3D convolution, then obtains a rough disparity map through a differentiable disparity regression, and finally converts it into a dense depth map through optimization and upsampling, achieving pixel-level accurate depth estimation.

[0112] In step S5, based on the depth values of each pixel in the dense depth map and the first virtual camera intrinsic parameter, the coordinates are back-projected to the first virtual camera coordinate system; combined with the geometric relationship between the first virtual camera and the third camera, the coordinates are converted to the third camera coordinate system; and finally, using the installation pose of the third camera relative to the spreader, the coordinates are converted to the coordinate system with the spreader as the origin. In the spreader coordinate system, the vertical component of the three-dimensional coordinates is the vertical distance from the corresponding pixel to the spreader, and the complete near-ground ranging result is obtained by integrating the vertical distances of all pixels, which directly provides decision basis for the automatic control system of the crane.

[0113] The above scheme directly calls the four-channel camera already installed on the crane, avoiding the high equipment cost and maintenance cost of the laser radar scheme, and also eliminating the need for additional deployment of special ranging equipment. The virtual large baseline (about 2.8 meters) constructed through image fusion improves the depth measurement accuracy compared to the short baseline (0.1-0.3 meters) of the traditional binocular camera, solving the precise ranging demand in the near-ground operation scene. The final output is the vertical distance in the spreader coordinate system, which does not need additional conversion and can be directly connected to the automatic control system of the crane, improving the operation safety and automation level and effectively avoiding collision risks.

[0114] In some embodiments, the projecting and transforming the first image to the perspective of the third camera based on the calibration relationship between the first camera and the third camera comprises: projecting and transforming the first image to the perspective of the third camera based on a first homography matrix between the first camera and the third camera , the first homography matrix is calculated according to the following formula:

[0115] ;

[0116] wherein G1 represents the first camera, A1 represents the third camera, h represents the height of the spreader when entering the preset near-ground operation range, K A1 represents an intrinsic matrix of the third camera, represents a rotation matrix from the first camera to the third camera at the height h, represents a translation vector from the coordinate system origin of the first camera to the coordinate system origin of the third camera at the height h, and n represents a unit normal vector of the preset reference plane, represents the vertical distance from the origin of the world coordinate system to the preset reference plane, represents the inverse matrix of the intrinsic matrix of the first camera.

[0117] The projecting and transforming the second image to the perspective of the fourth camera based on the calibration relationship between the second camera and the fourth camera comprises: projecting and transforming the second image to the perspective of the fourth camera based on a second homography matrix between the second camera and the fourth camera , the second homography matrix is calculated according to the following formula:

[0118] ;

[0119] wherein G2 represents the second camera, A2 represents the fourth camera, K A2 represents an intrinsic matrix of the fourth camera, represents a rotation matrix from the second camera to the fourth camera at the height h, represents a translation vector from the coordinate system origin of the second camera to the coordinate system origin of the fourth camera at the height h, represents the inverse matrix of the intrinsic matrix of the second camera.

[0120] In the present embodiment, the homography matrix is a 3x3 matrix describing the projection transformation relationship between two image planes, and is a core mathematical tool for realizing the projection of the high camera image to the perspective of the spreader camera, including a first homography matrix and a second homography matrix .

[0121] The intrinsic matrix is a matrix representing the internal optical characteristics of the camera, including focal length (f x , f y ) and principal point coordinates (cx, cy). The third camera intrinsic matrix K A1 , the fourth camera intrinsic matrix K A2 are fixed values (one-time calibration). The first camera intrinsic matrix K G1 , the second camera intrinsic matrix K G2 vary dynamically with the preset point and are stored in the lookup table.

[0122] The reference plane can be the horizontal plane on which the target container top surface is located, and its equation is n T •X=d plane , where n=(0, 1, 0) T is the unit normal vector (perpendicular upward), X is the coordinate of any point on the horizontal plane, and d plane is the vertical distance from the origin of the world coordinate system to the horizontal plane, which is usually set as the standard top surface height of the target container (such as 2.5 meters). In near-ground operation, the main operation target of the spreader is the container, and taking its top surface as the reference plane can maximize the accuracy of the projection transformation. At the same time, this plane is approximately an ideal plane, which meets the application premise of the homography matrix. If the operation target is a cargo ship deck, a trailer or other planar target, the value of d plane can be adjusted through system parameter configuration to adapt to different operation scenarios.

[0123] When the height of the spreader is h, the intrinsic matrix K G1 and the extrinsic parameters (rotation matrix, position coordinates) of the first camera (G1) corresponding to the preset point are read from the lookup table; through the PLC and the crane kinematics model, combined with the fixed calibration relationship between the third camera (A1) and the spreader, the real-time extrinsic parameters (rotation matrix R A1 (h), position coordinates P A1 (h)) of A1 in the world coordinate system are calculated.

[0124] The relative pose conversion formula of the first camera to the third camera is as follows: =R A1 (h)•R G1 (h) T , where R G1 (h) is the rotation matrix of the first camera at height h; the calculation formula of the translation vector from the origin of the coordinate system of the first camera to the origin of the coordinate system of the third camera is as follows: =P A1 (h)-R A1 (h)•R G1 (h) T •P G1 (h), where PG1 (h) is the position coordinate of the first camera at height h. Then each pixel coordinate of the first image is substituted into the homography matrix , and the projection transformation of the pixel coordinate is realized by matrix multiplication to obtain the transformed image of the first image under the perspective of the third camera, and the perspective alignment is completed. The calculation process of the second homography matrix is the same.

[0125] The height h of the spreader is included in the homography matrix as a variable, so that the rotation matrix and the translation vector are dynamically adjusted with h. No matter at which specific height (such as 5 meters, 3 meters, 0.5 meters) the spreader is in the near-ground operation range, the corresponding homography matrix can be calculated to realize accurate perspective projection at different heights and ensure the consistency of the ranging accuracy in the entire operation range.

[0126] In the above scheme, the homography matrix strictly follows the projection geometry principle, ensuring that the image of the high camera after perspective transformation is completely matched with the image of the spreader camera in geometric structure, providing a solid foundation for subsequent image fusion. By introducing the height h of the spreader into the homography matrix calculation, the projection transformation can adapt to different operation heights of the spreader in real time, avoiding the height adaptation limitations caused by fixed transformation matrix, and ensuring stable performance in the 0-5 meter operation range. For the dynamic parameters of the first camera and the second camera (zooming camera), through the preset point calibration and lookup table storage method, the dynamic changes of the internal and external parameters are solved, and the accuracy of the homography matrix calculation is ensured, without additional complex processing.

[0127] In some embodiments, as shown in Figure 2 , the first virtual camera image is generated according to the following manner:

[0128] S201: Align the first image after the first perspective transformation with the original image of the third image in the image space, and determine the overlapping area between the two images;

[0129] S202: In the overlapping area, the corresponding pixels of the first image after the first perspective transformation and the original image of the third image are fused pixel by pixel using a weight-based fusion algorithm to generate the first virtual camera image with an expanded field of view;

[0130] As shown in Figure 3 , the second virtual camera image is generated according to the following manner:

[0131] S301: Align the second image after the second perspective transformation with the original image of the fourth image in the image space, and determine the overlapping area between the two images;

[0132] S302: In the overlapping region, corresponding pixels of the second perspective-transformed second image and the original fourth image are fused pixel by pixel by using a weight-based fusion algorithm to generate the second virtual camera image with an extended field of view.

[0133] In this embodiment, the first perspective-transformed first image (G1 transformed image) and the third image (A1 original image) are introduced into the same pixel coordinate system, and the pixel coordinates of corresponding physical points in the two images are accurately aligned based on the geometric mapping relationship of the homography matrix, to ensure scene consistency. Through pixel coordinate traversal and scene feature matching, the pixel region corresponding to the physical scene simultaneously contained in the two images, i.e., the overlapping region, is identified. In the overlapping region, the G1 transformed image provides remote scene information, and the A1 original image provides near-end detail information, and then a fusion weight is assigned to each pixel in the overlapping region. The weight assignment rule can be based on the distance of the pixel to the image edge (e.g., the pixel near the edge of the G1 transformed image has a reduced G1 image weight and an increased A1 image weight, and vice versa for the pixel near the edge of the A1 image), or based on the image sharpness index (the pixel weight of the image with high sharpness is higher). Through weighted average calculation, the fused pixel value is obtained, and the pixel value of the corresponding image is directly retained in the non-overlapping region, to finally generate the first virtual camera image with an extended field of view. The generation method of the second virtual camera image is the same.

[0134] The fusion weight can be dynamically adjusted according to the actual scene, for example, in a strong light environment in a port, if the image of the high camera is overexposed, the weight of the image in the overlapping region can be reduced; if the image of the spreader camera is too close to the distance, resulting in blurred remote details, the weight of the transformed image of the high camera can be increased to ensure the overall quality of the fused image and provide high-quality input for subsequent feature extraction and matching.

[0135] The above scheme avoids the problems of seams, ghosting and brightness mutation caused by simple splicing through weighted fusion, so that the visual effect of the virtual camera image is close to the natural image taken by a single camera, and the reliability of subsequent image processing is improved. The fused image contains both the near-end clear details of the spreader camera and the remote wide field of view of the high camera, solving the problem of limited field of view of a single camera and providing more comprehensive scene information for stereo matching.

[0136] In some embodiments, the weight-based fusion algorithm includes a multi-band fusion algorithm; as Figure 4 As shown, the pixel-by-pixel fusion by using the weight-based fusion algorithm to generate the first virtual camera image with an extended field of view specifically includes:

[0137] S401: Laplacian pyramid decomposition is performed on the first perspective-transformed first image and the original third image, respectively;

[0138] S402: At each layer of the Laplacian pyramid, a first fusion weight map is calculated according to the position of the pixels in the overlapping area;

[0139] S403: The pyramid coefficients of the original image corresponding layers of the first image and the third image are weighted fused using the first fusion weight map;

[0140] S404: The weighted fused Laplacian pyramid is reconstructed to generate the first virtual camera image with an expanded field of view.

[0141] As shown in the first embodiment, the pixel-by-pixel fusion using the weight-based fusion algorithm to generate the second virtual camera image with an expanded field of view specifically includes: Figure 5 S501: The original images of the second image after the second perspective transformation and the fourth image are respectively decomposed into Laplacian pyramids;

[0142] S502: At each layer of the Laplacian pyramid, a second fusion weight map is calculated according to the position of the pixels in the overlapping area;

[0143] S503: The pyramid coefficients of the original image corresponding layers of the second image and the fourth image are weighted fused using the second fusion weight map;

[0144] S504: The weighted fused Laplacian pyramid is reconstructed to generate the second virtual camera image with an expanded field of view.

[0145] In the present embodiment, the multi-band fusion algorithm is a fusion algorithm based on image frequency decomposition, which decomposes the image into subbands of different spatial frequencies (low-frequency layer, high-frequency layer), and reconstructs after independent fusion at each layer, which can ensure the overall structure of the image to be smooth and the details and textures to be clear. In the present application, the Laplacian pyramid fusion algorithm is specifically used.

[0146] Laplacian pyramid decomposition refers to decomposing an image into a series of sub-images of different resolutions (pyramid layers). The bottom layer (low-frequency layer) contains the overall structure and brightness information of the image, and the high layer (high-frequency layer) contains the edge, texture and other detail information of the image. The decomposition process is realized through Gaussian filtering and down-sampling.

[0147] The fusion weight map is a weight matrix generated for each layer of the Laplacian pyramid. Each element in the matrix represents the weight of the corresponding pixel during fusion, which is dynamically calculated based on the position of the pixel in the overlapping area, ensuring smooth transition of each layer fusion.

[0148]

[0149] ​The pyramid coefficients are pixel values of each layer of the Laplacian pyramid, representing image information of a corresponding frequency band. In the fusion process, the corresponding layer coefficients of the two images are weighted and summed by using a weight map to realize the fusion of the information of the frequency band.

[0150] Pyramid reconstruction refers to combining the fused Laplacian pyramid coefficients of each layer in the inverse process, restoring the image to the original resolution through upsampling and Gaussian filtering to obtain the final fusion result.

[0151] The specific process of the Laplacian pyramid fusion of the first virtual camera image is as follows: the original images of the first image and the third image after the first perspective transformation are respectively subjected to Laplacian pyramid decomposition to obtain respective pyramid hierarchical structures. The number of decomposition layers can be adjusted according to the image resolution, and the typical number of decomposition layers is 4-6 layers to ensure that the full-band information from the overall structure to the detailed texture is covered. For each layer of the pyramid, the first fusion weight map is calculated based on the position of the pixel in the overlapping area. For example, in a certain layer, the closer the pixel is to the edge of the third image original image, the higher the weight of the corresponding third image, and the closer the pixel is to the edge of the first image after the perspective transformation, the higher the weight of the corresponding first image, and the weight value changes in a smooth gradient in the overlapping area. For the non-overlapping area, the weight value is set to 1 (corresponding to the image itself) and 0 (the other image). Using the first fusion weight map, the pyramid coefficients of the corresponding layers of the first image and the third image original image are weighted and fused, that is, the fusion coefficient of each pixel = the first image coefficient of the layer x the corresponding weight + the third image coefficient of the layer x (1-the corresponding weight), to realize accurate fusion of the information of each frequency band. Then, the fused Laplacian pyramid coefficients of all levels are reconstructed in the inverse process of the Laplacian pyramid, the resolution of each layer is restored through upsampling, and then Gaussian filtering is performed for smoothing processing, to finally generate the first virtual camera image with an expanded field of view, smooth structure and clear details. The generation mode of the second virtual camera image is similar to that of the first virtual camera image, and only the corresponding parameters need to be replaced in the calculation process.

[0152] Different levels of the Laplacian pyramid correspond to different frequency information, and the fusion strategy can be adjusted according to the characteristics of the frequency band: the low-frequency layer (image structure) adopts a smoother weight transition to ensure the consistency of the overall brightness and structure; the high-frequency layer (edge details) can appropriately increase the weight gradient to retain clear edge texture, so that the structure and details of the fused image are both optimal.

[0153] Compared with traditional pixel domain fusion, Laplacian pyramid fusion can independently optimize the fusion effect in different frequency bands, ensuring smooth transition of the overall structure of the image, effectively preserving the edge, texture and other detail information, and avoiding the problem of blur or artifacts of the fused image. Different fusion rules can be designed for different frequency bands, such as using higher weight resolution for high-frequency detail layers and using wider weight transition band for low-frequency structure layers, to adapt to the quality difference of the high-altitude image and the spreader image in different frequency bands. High-quality virtual camera images contain clear feature points and rich texture information, making subsequent feature extraction more stable, matching more accurate, reducing matching errors, and providing reliable image input for triangulation and depth estimation. Under complex lighting conditions in the port (such as strong light and shadow), the frequency characteristics of different images differ greatly, and multi-band fusion can effectively balance the advantages of each image in the frequency band, reduce the impact of lighting changes on the fusion effect, and improve the environmental robustness of the system.

[0154] In some embodiments, the weight-based fusion algorithm is a Laplacian pyramid fusion algorithm, which specifically includes: taking the original image of the third image as a fusion target region, taking the image gradient of the first image in the overlapping region after the first perspective transformation as a guide field, and solving the Poisson equation to fuse the gradient information of the guide field into the target region while keeping the boundary conditions of the original image of the third image unchanged, to generate the image with an expanded field of view. The Poisson fusion algorithm and the Laplacian pyramid fusion algorithm can be dynamically switched according to the actual working environment and image quality, improving the flexibility and adaptability of the system.

[0155] In some embodiments, as shown in FIG. 4, the method for generating a virtual camera image includes the following steps: Figure 6 As shown in FIG. 4, the method for generating a virtual camera image includes the following steps:

[0156] S31: Calculate the stereo rectification parameters based on the camera parameters associated with the first virtual camera image and the second virtual camera image; the camera parameters include the relative pose between the first virtual camera and the second virtual camera and the respective intrinsic matrices;

[0157] S32: Use the calculated stereo rectification parameters to perform geometric transformation on the first virtual camera image and the second virtual camera image respectively to obtain a pair of row-aligned corrected images, and the calculation formula of the geometric transformation is as follows:

[0158] ;

[0159] ;

[0160] wherein, , ) represents the pixel coordinates of the first virtual camera image before rectification, , ) represents the pixel coordinates of the first virtual camera image after rectification, , ) represents the pixel coordinates of the second virtual camera image before rectification, , ) represents the pixel coordinates of the second virtual camera image after rectification, represents the inverse matrix of the intrinsic matrix of the first virtual camera , represents the inverse matrix of the intrinsic matrix of the second virtual camera , represents the rectification rotation matrix of the first virtual camera, represents the rectification rotation matrix of the second virtual camera, represents the rectified and unified intrinsic matrix.

[0161] In step S31, the relative pose of the first virtual camera and the second virtual camera is directly determined by the relative pose of the third camera A1 and the fourth camera A2 located on the spreader. Since the virtual cameras are generated by the fusion of A1, A2 and the gantry structure, their relative pose remains consistent with the relative pose of A1 and A2. The intrinsic matrix , of the first virtual camera and the second virtual camera respectively inherits the intrinsic matrix , of A1 and A2.

[0162] Then the Bouguet stereo rectification algorithm is used. This algorithm solves the rotation matrix , so that the rectified two virtual camera images satisfy the following constraints: epipolar line is horizontal, i.e. the vertical coordinates of the projection points of the same physical point in the two images are the same = ; the principal point coordinates are consistent and located on the same horizontal line; the focal length is the same to ensure the image scale is unified. The output results are the rectification rotation matrix corresponding to the first virtual camera, corresponding to the second virtual camera, and the rectified and unified intrinsic matrix .

[0163] In step S32, the pixels of the first virtual camera image , are first back-projected to the normalized camera coordinate system of A1 through (i.e. ); then rotated through to make the image epipolar line horizontal; finally, the rectified and unified intrinsic matrix Reprojecting onto the corrected pixel coordinate system, we obtain ( , ).

[0164] For the pixels of the second virtual camera image ( , Similarly, through (Right now Back projection, rotation, and reprojection yield ( , The two corrected images (Rectified1 and Rectified2) have strictly aligned pixel rows, and the projection points of the same physical point are only offset in the horizontal direction (parallax), which greatly facilitates subsequent stereo matching.

[0165] The corrected images possess uniform intrinsic parameters, horizontal epipolar lines, and consistent scale, eliminating the need for subsequent feature matching algorithms (including neural networks) to adapt to pose differences across different cameras, thus improving the algorithm's versatility and stability. Horizontal epipolar line constraints prevent mismatches of non-corresponding points, and combined with the high disparity sensitivity of the virtual large baseline, further enhance the confidence of feature point matching, providing high-quality matching point pairs for subsequent triangulation. Stereo correction parameters are dynamically calculated based on camera calibration parameters, ensuring that standardized corrected images are generated regardless of the lifting device's height within the near-ground working range, guaranteeing stable ranging accuracy across the entire working area.

[0166] In some embodiments, such as Figure 7 As shown, the step of performing feature point matching and triangulation on the first virtual camera image and the second virtual camera image to obtain the coordinates of multiple feature points and their corresponding depth values ​​specifically includes:

[0167] S701: Extract feature points from the first virtual camera image and the second virtual camera image respectively and generate corresponding feature descriptors;

[0168] S702: Match feature points in the first virtual camera image and the second virtual camera image using the feature descriptor to obtain an initial set of matching point pairs;

[0169] S703: Filter the initial set of matching points to obtain a set of high-confidence matching points, wherein any high-confidence matching point in the first virtual camera image is denoted as p1, and the coordinates of p1 are (u1, v1), and the corresponding matching point in the second virtual camera image is denoted as p2, and the coordinates of p2 are (u2, v2).

[0170] S704: For each pair of matching points (p1, p2) in the set of high-confidence matching point pairs, construct a homogeneous linear equation system, solve the homogeneous linear equation system using a singular value decomposition algorithm to obtain the coordinates P = (X, Y, Z) of the matching point (p1, p2) in the three-dimensional space, and the homogeneous linear equation system is represented as follows:

[0171] ;

[0172] wherein, represents the first row of the projection matrix ; M v1 and M v2 represent the projection matrices of the first virtual camera and the second virtual camera, respectively, and the formulas of M v1 and M v2 are as follows:

[0173] ;

[0174] ;

[0175] wherein, represents the intrinsic matrix of the first virtual camera, represents the intrinsic matrix of the second virtual camera, represents the extrinsic matrix when the first virtual camera coordinate system is the world coordinate system, represents the relative pose matrix for transforming a point from the first virtual camera coordinate system to the second virtual camera coordinate system;

[0176] S705: Convert the coordinates P = (X, Y, Z) to the camera coordinate system of the first virtual camera to obtain the coordinates P v1 , and the conversion formula is as follows: wherein, is the optical center coordinates of the first virtual camera, represents the rotation matrix for converting the coordinates P to the camera coordinate system of the first virtual camera;

[0177] S706: Take the Z-axis component of P v1 as the depth value Z depth corresponding to the matching point p1.

[0178] S707: Traverse all high-confidence matching point pairs to obtain a set of coordinates (u1, v1) of a plurality of feature points in the first virtual camera image and the corresponding depth values Z depth .

[0179] In feature point extraction and matching, the ORB algorithm (high efficiency, real-time) is preferred considering the balance between real-time performance and robustness. If high scale invariance is required, SIFT or SURF algorithm can be selected. In the extraction process, by setting the feature point response threshold, high-identifiability feature points are selected to avoid low-quality feature points in weak texture areas. Then, the Hamming distance (binary descriptor) or Euclidean distance (floating-point descriptor) of the feature descriptors of the first virtual camera image and the second virtual camera image is calculated to preliminarily match similar feature points. For each initial matching point pair, the distance ratio of the optimal matching and the suboptimal matching is calculated, and only the matching point pair with a ratio less than 0.75 is retained to eliminate ambiguous matching. If the feature point p1 in the first virtual camera image matches the feature point p2 in the second virtual camera image, and the optimal matching of p2 in the second virtual camera image is p1 in the first virtual camera image, the matching point pair is retained, and further one-way false matching is eliminated. The output result is a high-confidence matching point pair set, each point pair containing (p1, p2).

[0180] Then, based on the projection matrix M v1 and M v2 , a 4x4 overdetermined equation set is constructed for each matching point pair (p1, p2). The core constraint of the equation set is that the three-dimensional point is on the imaging ray of the two cameras, that is, it satisfies the projection relationship. Singular value decomposition (SVD) is used to solve the homogeneous linear equation set. SVD can handle overdetermined equations and obtain the optimal solution, avoiding numerical instability problems caused by direct solving. After solving, the three-dimensional point P=(X, Y, Z) (world coordinate system, i.e., first virtual camera coordinate system) is obtained. Then, the three-dimensional point P is converted to the first virtual camera coordinate system (since M v1 is the world coordinate system, the converted coordinates are invariant), and the Z-axis component is the depth value Z depth of the feature point p1. This value is the physical distance along the camera optical axis.

[0181] Then, all high-confidence matching point pairs are traversed, and the above triangulation process is repeated to obtain the coordinates of each matching point pair in the first virtual camera image and the corresponding depth value, forming a complete depth value set to provide data support for sparse reference depth map generation.

[0182] The above scheme is based on virtual large-baseline triangulation, and the depth measurement error is significantly reduced. Compared with traditional short-baseline binocular cameras, it provides reliable geometric constraints for subsequent neural networks. Through ratio testing and cross-validation, false matching point pairs are effectively eliminated to ensure the accuracy of the depth value and avoid depth estimation deviation caused by false matching. The feature point extraction algorithm has rotation and scale invariance, and can still extract stable feature points even in complex scenes such as port light changes and target occlusion, ensuring the robustness of matching and triangulation.

[0183] In some embodiments, asFigure 8 As shown, step S6 specifically includes:

[0184] S61: Based on the depth value of each pixel in the dense depth map and the intrinsic matrix of the first virtual camera, each pixel is back-projected into the coordinate system of the first virtual camera to obtain a first set of three-dimensional point clouds P d ;

[0185] S62: Based on the geometric relationship between the first virtual camera and the third camera, the first set of three-dimensional point clouds P d is converted into the coordinate system of the third camera to obtain a second set of three-dimensional point clouds P c ;

[0186] S63: Based on the mounting pose of the third camera relative to the spreader, the second set of three-dimensional point clouds P c is converted into the coordinate system with the spreader as the origin to obtain a third set of three-dimensional point clouds P h ;

[0187] S64: The vertical direction coordinate component of each point in the third set of three-dimensional point clouds P h is extracted as the vertical distance from the corresponding pixel to the spreader to obtain the near-ground ranging result.

[0188] In step S61, for a pixel with coordinates (x, y) and depth value depth in the dense depth map, the three-dimensional coordinates P d =(X d ,Y d ,Z d ) of the pixel in the coordinate system of the first virtual camera are calculated as follows:

[0189] X d =(x-c x )×depth / f x ;

[0190] Y d =(y-c y )×depth / f y ;

[0191] Z d =depth;

[0192] where (c x , c y ) are the principal point coordinates, f x , f y are the focal lengths, and all come from the intrinsic matrix K v1 of the first virtual camera.

[0193] In step S62, the first virtual camera is generated by fusing the third camera (A1) and the first camera G1, and the coordinate system thereof is completely consistent with the A1 coordinate system (both the internal parameters and the external parameters are inherited from A1), so the first set of three-dimensional point clouds P d are the same as the second set of three-dimensional point cloud coordinates P c , that is, P c =P d .

[0194] If there is a slight pose deviation (such as geometric error in the fusion process) between the virtual camera and A1, a pre-calibrated correction matrix R cal , T cal can be used for fine adjustment to ensure the coordinate conversion accuracy: P c =R cal ×P d +T cal , and the correction matrix is obtained through a calibration experiment.

[0195] In step S63, based on the installation pose of the third camera (A1) relative to the spreader, the conversion to the spreader coordinate system is performed. After the conversion, the origin of the three-dimensional point cloud P h is coincident with the center of the spreader, and the coordinate axes are consistent with the motion direction of the spreader (for example, the X axis is parallel to the length direction of the spreader, and the Y axis is perpendicular upward), which directly reflects the spatial position of the target relative to the spreader.

[0196] In step S64, in the spreader coordinate system, the vertical direction is the Y axis direction, so the Y coordinate component of each point in the third set of three-dimensional point clouds P h is the vertical distance of the point to the spreader. Extracting the Y coordinate components corresponding to all pixels forms a vertical distance matrix aligned with the pixels of the first view image, which is the near-ground ranging result, which can be directly output to the crane automatic control system for spreader positioning, anti-collision and other operations.

[0197] The above scheme converts the abstract dense depth map into a vertical distance in the spreader coordinate system, directly matches the input requirements of the crane control system, does not require additional data conversion, and improves the operation efficiency. Through a three-level coordinate conversion chain, the camera calibration parameters and the installation pose are strictly followed, so that the measurement result is the true vertical distance relative to the spreader regardless of the height and attitude of the spreader, and there is no system deviation.

[0198] In some embodiments, the depth estimation neural network model is trained in the following manner:

[0199] S81: Construct a training sample set, each sample in the training sample set includes a stereo image pair composed of a first training perspective image and a second training perspective image, and a corresponding sparse training depth map; wherein the sparse training depth map is generated by randomly sampling or noise simulation on a dense ground truth depth map aligned with pixels of the first training perspective image;

[0200] S82: input the training sample set into the depth estimation neural network model, and the model performs the following operations:

[0201] perform feature extraction and stereo matching on the stereo image pair to generate an initial depth estimation;

[0202] fuse the sparse training depth map as guide information with the initial depth estimation in the process of stereo matching or in a subsequent optimization process;

[0203] output depth maps at different stages or different resolutions, including a first dense depth map and a second dense depth map; wherein the first dense depth map is an intermediate coarse depth map, and the second dense depth map is a final output fine depth map;

[0204] S83: for the first dense depth map, calculate the difference between the first dense depth map and a first reference dense depth map to obtain a first loss L coarse , the first reference dense depth map is obtained by spatially downsampling a second reference dense depth map;

[0205] for the second dense depth map, calculate the difference between the second dense depth map and a second reference dense depth map, a second loss L final , the second reference dense depth map is the dense ground truth depth map;

[0206] for the sparse pixel points with valid depth values in the sparse training depth map, calculate the difference between the predicted depth value of the corresponding pixel point in the second dense depth map and the sparse ground truth depth value, a third loss L sparse ;

[0207] construct a total loss function based on the first loss L coarse , the second loss L final and the third loss L sparse , the formula of the total loss function is as follows:

[0208] L total =L final +α·L coarse +β·L sparse ;

[0209] wherein, L totaldenotes the total loss, and a and b are preset weight coefficients;

[0210] S84: using the total loss L total , updating the parameters of the depth estimation neural network model by a back propagation algorithm;

[0211] S85: repeating steps S82 to S84 until the model converges, obtaining a pre-trained depth estimation neural network model.

[0212] The above scheme guides through multi-stage supervision and sparse depth prior, the model not only learns the mapping relationship between image features and depth, but also masters the geometric constraint law, and does not need to be retrained when deployed across scenes, and only needs to load a new scene calibration file. Triple loss supervision ensures that the model not only guarantees the accuracy of the global geometric structure, but also optimizes the local details, and finally the precision of the fine depth map approaches the level of the laser radar, meeting the high-precision requirements of port near-ground operations. Sparse depth prior provides strong geometric constraints for the model, and even in weak texture, occlusion and other traditional stereo matching problem scenes, the model can still generate reliable dense depth maps based on sparse anchors, and the robustness is better than that of pure data-driven models.

[0213] In the embodiment, the depth estimation neural network model includes five core modules, which work together to achieve high-precision depth estimation, specifically including:

[0214] The multi-modal feature extraction network module includes a twin feature network (weight sharing) and a depth prior encoder, which respectively extract image features (low resolution 32 channels, high resolution 16 channels) of a stereo image pair and depth prior features (8 channels) of a sparse depth map, and fuse the image features and the depth prior features to obtain 40-channel fusion features.

[0215] The 4D cost space construction module constructs a 4D cost space with a dimension of [B, 48, 120, 160] (B is the batch size, 48 is the disparity range, and 120x160 is the low resolution feature size) for the low resolution features (downsampling factor 4), and represents the matching cost by calculating the correlation of the left and right features.

[0216] The 3D cost space regularization network module adopts a 3D U-Net architecture, and smoothes and aggregates the context of the 4D cost space through 3D convolution, pooling, deconvolution and skip connection, highlights the low cost value corresponding to the correct disparity, and generates a regularized cost space.

[0217] The differentiable disparity regression module converts the regularized cost space into a probability distribution through a Soft-Argmin operation, calculates the weighted average disparity, and obtains a low resolution rough disparity map (120x160), supporting end-to-end training.

[0218] The depth map optimization and up-sampling module up-samples the coarse disparity map to 480*640 by bilinear interpolation, splices the high-resolution left feature, and outputs a fine disparity map by a lightweight 2D CNN optimization network, and finally converts the fine disparity map into a dense depth map.

[0219] In a second aspect, the present application also provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the virtual large baseline four-vision based ground proximity ranging method according to the first aspect of the present application.

[0220] The computer readable storage medium can be a volatile memory or a non-volatile memory, or both.

[0221] The non-volatile memory can be a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read only memory (CD ROM).

[0222] The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), Direct Rambus Random Access Memory (DRRAM). The computer-readable storage media described in the embodiments of the application are intended to include these and any other suitable types of memory.

[0223] As shown in Figure 9 In a third aspect, the application provides an electronic device 10 comprising a processor 101 and a storage medium 102, wherein the storage medium 102 has stored thereon a computer program which, when executed by the processor 101, implements the method for near-ground ranging based on virtual large-baseline four-eye vision according to the first aspect of the application.

[0224] In some embodiments, the processor can be implemented by software, hardware, firmware or a combination thereof, and can use at least one of circuit, single or multiple Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), controller, microcontroller, microprocessor, so that the processor can execute part or all of the steps or any combination of the steps in the virtual large baseline four-vision based ground proximity ranging method in various embodiments of the present application.

[0225] Finally, it should be noted that although the above embodiments have been described in the specification and drawings of the present application, the patent protection scope of the present application should not be limited thereby. Any technical solutions obtained by replacing or modifying the equivalent structure or equivalent flow based on the essential concept of the present application, using the content described in the specification and drawings of the present application, and directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, etc., are all included in the patent protection scope of the present application.

Claims

1. A near-ground ranging method based on virtual large baseline quad-vision, applied to a container crane, the crane comprising a gantry structure and a spreader capable of vertically moving along the gantry structure, characterized in that, The gantry structure is equipped with a first camera and a second camera at a first preset height, and the lifting device is equipped with a third camera and a fourth camera at a second preset height. The first preset height is higher than the second preset height. The method includes: S1: When the spreader enters the preset near-ground working range, it is simultaneously triggered to acquire the first image captured by the first camera, the second image captured by the second camera, the third image captured by the third camera, and the fourth image captured by the fourth camera. S2: Based on the calibration relationship between the first camera and the third camera, the first image is projected and transformed to the viewpoint of the third camera, and then fused with the third image to generate a first virtual camera image; and based on the calibration relationship between the second camera and the fourth camera, the second image is projected and transformed to the viewpoint of the fourth camera, and then fused with the fourth image to generate a second virtual camera image. S3: Perform stereo correction on the first virtual camera image and the second virtual camera image to obtain a pair of row-aligned corrected images; crop out the common field of view area from the pair of corrected images to obtain a first-view image and a second-view image of the same size and row-aligned. S4: Perform feature point matching and triangulation on the first virtual camera image and the second virtual camera image to obtain the coordinates of multiple feature points and their corresponding depth values; according to the coordinate mapping relationship between the multiple feature points in the first virtual camera image and the first view image, fill the depth values ​​into a two-dimensional matrix with the same size as the first view image to generate a sparse reference depth map aligned with the pixels of the first view image, wherein the matrix elements without filled depth values ​​are empty or zero. S5: Input the first-view image, the second-view image, and the sparse reference depth map into the pre-trained depth estimation neural network model, and the model outputs a dense depth map that is aligned with the pixels of the first-view image. S6: Based on the depth values ​​of each pixel in the dense depth map, the intrinsic parameter matrix of the first virtual camera, the geometric relationship between the first virtual camera and the third camera, and the installation pose of the third camera relative to the hoist, the coordinates corresponding to each pixel are transformed to a coordinate system with the hoist as the origin, and the vertical component of the coordinate is taken as the vertical distance from the pixel to the hoist to obtain the near-ground distance measurement result.

2. The near-ground ranging method based on virtual large baseline tetranoscopic vision as described in claim 1, characterized in that, The process of projecting and transforming the first image onto the viewpoint of the third camera based on the calibration relationship between the first and third cameras includes: based on the first homography matrix between the first and third cameras. A first perspective transformation is performed on the first image, and the first homography matrix... The calculation formula is as follows: ; Where G1 represents the first camera, A1 represents the third camera, h represents the height at which the lifting device enters the preset near-ground working range, and K... A1 This represents the intrinsic parameter matrix of the third camera. This represents the rotation matrix from the first camera to the third camera at height h. This represents the translation vector from the origin of the coordinate system of the first camera to the origin of the coordinate system of the third camera at a height h, where n represents the unit normal vector of the preset reference plane. This represents the perpendicular distance from the origin of the world coordinate system to the preset reference plane. This represents the inverse matrix of the intrinsic parameter matrix of the first camera; The method of projecting and transforming the second image onto the viewpoint of the fourth camera based on the calibration relationship between the second and fourth cameras includes: based on the second homography matrix between the second and fourth cameras. A second perspective transformation is performed on the second image, and the second homography matrix... The calculation formula is as follows: ; Where G2 represents the second camera, A2 represents the fourth camera, and K... A2 This represents the intrinsic parameter matrix of the fourth camera. This represents the rotation matrix from the second camera to the fourth camera at height h. This represents the translation vector from the origin of the second camera's coordinate system to the origin of the fourth camera's coordinate system at height h. This represents the inverse of the intrinsic parameter matrix of the second camera.

3. The near-ground ranging method based on virtual large baseline tetranoscopic vision as described in claim 2, characterized in that, The first virtual camera image is generated in the following manner: Align the first image after the first perspective transformation with the original image of the third image in the image space to determine the overlapping area between the two images; Within the overlapping area, corresponding pixels of the original images of the first image after the first perspective transformation and the third image are fused pixel by pixel using a weight-based fusion algorithm to generate a first virtual camera image with an expanded field of view. The second virtual camera image is generated in the following manner: Align the second image after the second perspective transformation with the original image of the fourth image in the image space to determine the overlapping area between the two images; Within the overlapping area, corresponding pixels of the second image after the second perspective transformation and the original image of the fourth image are fused pixel by pixel using a weight-based fusion algorithm to generate a second virtual camera image with an expanded field of view.

4. The near-ground ranging method based on virtual large baseline tetranoscopic vision as described in claim 3, characterized in that, The weight-based fusion algorithm includes a multi-band fusion algorithm; The step of employing a weight-based fusion algorithm to perform pixel-by-pixel fusion to generate the first virtual camera image with an expanded field of view specifically includes: Laplacian pyramid decomposition is performed on the first image after the first perspective transformation and the original image of the third image, respectively. In each layer of the Laplacian pyramid, a first fusion weight map is calculated based on the position of the pixel within the overlapping region; The first fusion weight map is used to perform weighted fusion of the pyramid coefficients of the corresponding layers of the original images of the first image and the third image; The weighted and fused Laplacian pyramid is reconstructed to generate the first virtual camera image with an expanded field of view; The step of employing a weight-based fusion algorithm to perform pixel-by-pixel fusion to generate the second virtual camera image with an expanded field of view specifically includes: The original images of the second image after the second perspective transformation and the fourth image are respectively subjected to Laplacian pyramid decomposition; In each layer of the Laplacian pyramid, a second fusion weight map is calculated based on the position of the pixel within the overlapping region; The second fusion weight map is used to perform weighted fusion of the pyramid coefficients of the corresponding layers of the original images of the second image and the fourth image; The weighted and fused Laplacian pyramid is reconstructed to generate the second virtual camera image with an expanded field of view.

5. The near-ground ranging method based on virtual large baseline tetranoscopic vision as described in claim 1, characterized in that, To perform stereo correction on the first and second virtual camera images to obtain a pair of row-aligned corrected images, the following steps are included: S31: Calculate stereo correction parameters based on camera parameters associated with the first virtual camera image and the second virtual camera image; the camera parameters include the relative pose between the first virtual camera and the second virtual camera and their respective intrinsic parameter matrices; S32: Using the calculated stereo correction parameters, perform geometric transformations on the first virtual camera image and the second virtual camera image respectively to obtain a pair of row-aligned corrected images. The calculation formula for the geometric transformation is as follows: ; ; in,( , ) represents the pixel coordinates of the first virtual camera image before correction, ( , ) represents the pixel coordinates of the corrected first virtual camera image, ( , ) represents the pixel coordinates of the second virtual camera image before correction, ( , () represents the pixel coordinates of the corrected second virtual camera image. The intrinsic parameter matrix of the first virtual camera The inverse matrix, The intrinsic parameter matrix representing the second virtual camera The inverse matrix, This represents the correction rotation matrix of the first virtual camera. This represents the correction rotation matrix of the second virtual camera. This represents the unified intrinsic parameter matrix after correction.

6. The near-ground ranging method based on virtual large baseline tetranoscopic vision as described in claim 1, characterized in that, The step of performing feature point matching and triangulation on the first virtual camera image and the second virtual camera image to obtain the coordinates of multiple feature points and their corresponding depth values ​​specifically includes: Feature points are extracted from the first virtual camera image and the second virtual camera image, and corresponding feature descriptors are generated. The feature descriptors are used to match feature points in the first virtual camera image and the second virtual camera image to obtain an initial set of matching point pairs. The initial set of matching points is filtered to obtain a set of matching points with high confidence. Any high-confidence matching point in the first virtual camera image is denoted as p1, and the coordinates of p1 are (u1, v1). The corresponding matching point in the second virtual camera image is denoted as p2, and the coordinates of p2 are (u2, v2). For each pair of matching points (p1, p2) in the set of high-confidence matching points, a homogeneous linear equation system is constructed. The singular value decomposition algorithm is used to solve the homogeneous linear equation system to obtain the coordinates P=(X,Y,Z) of the matching points (p1, p2) in three-dimensional space. The homogeneous linear equation system is expressed as follows: ; in, Represents the projection matrix The Okay, M v1 and M v2 M represents the projection matrices of the first virtual camera and the second virtual camera, respectively. v1 and M v2 The formula for representing is as follows: ; ; in, This represents the intrinsic parameter matrix of the first virtual camera. This represents the intrinsic parameter matrix of the second virtual camera. This represents the extrinsic parameter matrix when the first virtual camera coordinate system is used as the world coordinate system. This represents the relative pose matrix that transforms a point from the first virtual camera coordinate system to the second virtual camera coordinate system; Transform the coordinates P=(X,Y,Z) into the camera coordinate system of the first virtual camera to obtain the coordinates P v1 The conversion formula is as follows: ,in, Let P be the optical center coordinates of the first virtual camera; v1 The Z-axis component is used as the depth value Z corresponding to the matching point p1. depth , This represents the rotation matrix that transforms coordinate P to the camera coordinate system of the first virtual camera; By traversing all high-confidence matching point pairs, the coordinates (u1, v1) of multiple feature points in the first virtual camera image and their corresponding depth values ​​Z are obtained. depth A set of.

7. The near-ground ranging method based on virtual large baseline tetranoscopic vision as described in claim 1, characterized in that, Step S6 specifically includes: S61: Based on the depth values ​​of each pixel in the dense depth map and the intrinsic parameter matrix of the first virtual camera, back-project each pixel onto the coordinate system of the first virtual camera to obtain the first set of 3D point clouds P. d ; S62: Based on the geometric relationship between the first virtual camera and the third camera, the first set of 3D point cloud P d Transform to the coordinate system of the third camera to obtain the second set of 3D point cloud P. c ; S63: Based on the installation pose of the third camera relative to the hoist, the second set of three-dimensional point cloud P c Transform to the coordinate system with the lifting device as the origin to obtain the third set of three-dimensional point clouds P. h ; S64: Extract the third set of three-dimensional point cloud P h The vertical coordinate component of each point is used as the vertical distance from the corresponding pixel to the lifting device to obtain the near-ground distance measurement result.

8. The near-ground ranging method based on virtual large baseline tetranoscopic vision as described in claim 1, characterized in that, The depth estimation neural network model is trained in the following way: S81: Construct a training sample set, wherein each sample in the training sample set includes a stereo image pair consisting of a first training viewpoint image and a second training viewpoint image, and a corresponding sparse training depth map; wherein the sparse training depth map is generated by random sampling or noise simulation of a dense ground truth depth map aligned with the pixels of the first training viewpoint image. S82: Input the training sample set into the depth estimation neural network model, and the model performs the following operations: extract features and perform stereo matching on the stereo image pairs to generate an initial depth estimate; The sparse training depth map is used as guiding information and fused with the initial depth estimate during the stereo matching process or subsequent optimization process; Output depth maps at different stages or at different resolutions, including the first dense depth map and the second dense depth map; Wherein, the first dense depth map is an intermediate coarse depth map, and the second dense depth map is the final output fine depth map; S83: For the first dense depth map, calculate the difference between the first dense depth map and the first reference dense depth map to obtain the first loss L. coarse The first reference dense depth map is obtained by spatial downsampling the second reference dense depth map; For the second dense depth map, calculate the difference between the second dense depth map and the second reference dense depth map, and the second loss L. final The second reference dense depth map is the dense true depth map; For sparse pixels with valid depth values ​​in the sparse training depth map, calculate the difference between the predicted depth value and the ground truth depth value in the second dense depth map at the corresponding pixel, and apply a third loss L. sparse ; Based on the first loss L coarse Second loss L final and the third loss L sparse Construct the total loss function, the formula of which is as follows: L total =L final +α·L coarse +β·L sparse ; Among them, L total This represents the total loss, where α and β are preset weighting coefficients; S84: Utilizing total loss L total The parameters of the depth estimation neural network model are updated using the backpropagation algorithm. S85: Repeat steps S82 to S84 until the model converges, and obtain the pre-trained deep estimation neural network model.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the near-ground ranging method based on virtual large baseline quad-vision as described in any one of claims 1 to 8.

10. An electronic device having a computer program stored thereon, characterized in that, It includes a processor and a storage medium, wherein a computer program is stored on the storage medium, and the computer program, when executed by the processor, implements the near-ground ranging method based on virtual large baseline tetracular vision as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image processing method and device

    CN114119701A

  • Four-eye structured light stereoscopic vision imaging method

    CN120953480A