Visual scale correction method and device, computer equipment and storage medium

By introducing TOF sensors and depth calculation values ​​in monocular vision devices, the visual scale of the visual odometer is updated based on the depth ratio, and the scale drift and accuracy degradation of the traditional visual odometer in a uniform motion state is solved, and reliable scale calibration and relative positioning functions are achieved.

CN120028777APending Publication Date: 2025-05-23启元实验室
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411837780.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional visual odometers or visual inertial odometers are prone to scale drift and accuracy degradation in a uniform state of motion, resulting in the inability to accurately correct the visual scale.

Method used

A monocular vision device with a point laser ranging is used to obtain the depth measurement value through the TOF sensor, and the depth calculation value is obtained through the stereo matching between two frames of images adjacent to time, and the visual scale of the visual odometer is updated based on the depth ratio.

Benefits of technology

Continuous and reliable scale calibration is achieved, and the scale drift and accuracy reduction problems of traditional methods in uniform motion states are overcome, ensuring the long-term reliable relative positioning function of outdoor unmanned platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120028777A_ABST
    Figure CN120028777A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing and computer vision, and discloses a vision scale correction method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining two frames of scene images which are collected in a camera motion process and are adjacent in time, and calculating a relative pose corresponding to the collection moment of the two frames of scene images, recording a depth measurement value and a pixel position of the TOF sensor at the first image acquisition moment; selecting an image pixel alignment direction based on the relative pose, performing epipolar correction on the first image and the second image based on the pixel alignment direction, and calculating a sub-pixel position in the corrected first image; and performing stereo matching on the corrected image to obtain a disparity map, calculating a depth calculation value corresponding to the sub-pixel position based on the disparity map, calculating a depth ratio, and updating the visual scale based on the depth ratio. The method solves the problems of scale drift and precision reduction of a visual odometer or a visual inertia odometer in a constant-speed motion state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and computer vision technology, and in particular to a visual scale correction method, device, computer equipment and storage medium. Background Art

[0002] The use of traditional visual odometry / visual inertial odometry on outdoor high-altitude unmanned platforms faces great difficulties due to the difficulty in obtaining effective scale measurement information. First, the monocular camera itself cannot provide absolute scale information for navigation; second, due to the large depth of outdoor scenes, binocular stereo measurement requires a long baseline to work, and the size of the drone limits the length of the baseline; third, the inertial device requires the prior assistance of the accelerometer bias (the offset between the accelerometer output signal and the true value) to restore the scale during initialization, and the real-time bias of the inertial device is difficult to obtain during the system bootstrap or navigation state recovery process, affecting the accuracy of scale estimation; in addition, when the platform is close to constant speed motion, the inertial device will lose its contribution to the scale, resulting in the divergence of the final positioning solution result, which is very likely to cause scale drift and accuracy reduction of the visual odometry or visual inertial odometry in a uniform motion state. The above problems affect the actual use of the visual odometry / visual inertial odometry as an independent navigation unit on outdoor unmanned platforms. Summary of the invention

[0003] In view of this, the present invention provides a visual scale correction method, apparatus, computer equipment and storage medium to solve the problem that the visual scale cannot be accurately corrected due to scale drift and precision degradation of traditional visual odometers or visual inertial odometers in a uniform motion state.

[0004] In a first aspect, the present invention provides a visual scale correction method, which is applied to a monocular vision device with a point laser ranging, wherein the monocular vision device with a point laser ranging comprises a rigidly connected camera and a TOF sensor, and the direction of emitting and receiving lasers of the TOF sensor is within the observation field angle of the camera; the method comprises:

[0005] Acquire two adjacent frames of scene images captured by the camera within a preset time interval during the motion process, and calculate the relative pose corresponding to the two frames of scene image capture time and the relative distance between the two frames of scene images, the two frames of scene images are images containing the laser points emitted by the TOF sensor, including the first image and the second image, and record the depth measurement value and pixel position of the TOF sensor at the time of capturing the first image;

[0006] Selecting a direction for aligning image pixels based on the relative posture, and performing epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculating a pixel position corresponding to a laser projection direction of the TOF sensor in the first image and mapping it to a sub-pixel position in the corrected first image;

[0007] Performing stereo matching on the corrected first image and the corrected second image to obtain a disparity map, and determining whether an average disparity of all pixels in the disparity map is greater than a first specified threshold;

[0008] When the average disparity of all pixels is greater than a first specified threshold, a depth calculation value corresponding to the sub-pixel position is calculated based on the disparity map and the relative distance between the two frames of scene images;

[0009] A depth ratio is determined based on the depth measurement value and the depth calculation value, and a visual scale within a current time window of the visual odometry is updated based on the depth ratio.

[0010] The visual scale correction method provided by the present invention introduces the ranging information obtained by a monocular vision device with a point laser ranging into a visual odometer, obtains the depth measurement value of the corresponding pixel position through a TOF sensor, obtains the depth calculation value of the corresponding pixel position through stereo matching between two temporally adjacent frames of images, determines the depth ratio based on the depth measurement value and the depth calculation value, and updates the visual scale within the current time window of the visual odometer based on the depth ratio, so as to achieve continuous and reliable scale calibration, so that a long-term and reliable relative positioning function of an outdoor unmanned platform can be achieved by combining a monocular camera with a small single-point TOF sensor (Time-of-Flight, a sensor that measures distance using the time-of-flight technology), and solves the problem that the visual scale cannot be accurately corrected due to scale drift and precision reduction of a traditional visual odometer or a visual inertial odometer in a uniform motion state.

[0011] In an optional implementation, before acquiring two adjacent frames of scene images captured by a camera within a preset time interval, the visual scale correction method further includes:

[0012] Calibrate the pixel positions on the image corresponding to the orientation elements in the camera and the laser projection direction of the TOF sensor.

[0013] The visual scale correction method provided by the present invention calibrates the orientation elements in the camera and the pixel positions on the image corresponding to the laser projection direction of the TOF sensor, which is conducive to obtaining the pixel position values ​​on the image corresponding to the laser projection direction of the TOF sensor, improving the precision and accuracy of the scale correction, avoiding the accuracy problems caused by deviations in the shooting angle and position, and providing conditions for accurately obtaining depth measurement values ​​and calculation values.

[0014] In an optional implementation, calculating the relative pose corresponding to the capture moment of two frames of scene images and the relative distance between the two frames of scene images includes:

[0015] A visual odometer is used in a visual odometer reference coordinate system to calculate a first relative pose and a second relative pose corresponding to the first image and the second image acquisition moments, and a relative distance between two frames of scene images is obtained based on the first relative pose and the second relative pose.

[0016] The visual scale correction method provided by the present invention uses a visual odometer to calculate the first relative posture and the second relative posture corresponding to the first image and the second image acquisition time in a visual odometer reference coordinate system, estimates the movement of the camera between the two frames of scene images, and realizes accurate calculation of the first relative posture and the second relative posture corresponding to the first image and the second image acquisition time and the relative distance between the two frames of scene images.

[0017] In an optional implementation, selecting a direction for aligning image pixels based on the relative posture, and performing epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculating the pixel position corresponding to the TOF sensor laser projection direction in the first image and mapping it to the sub-pixel position in the corrected first image includes:

[0018] Based on the relative posture, the vector pointing from the camera center position corresponding to the first image to the camera center position corresponding to the second image is projected onto the XOY plane of the camera coordinate system corresponding to the first image to obtain the projection vector, and the magnitude of the components of the projection vector in the X direction and the Y direction is determined. If the component in the X direction is large, the direction of pixel alignment is selected as the x direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel row; if the component in the Y direction is large, the direction of pixel alignment is selected as the y direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel column;

[0019] Calculate a first correction matrix and a second correction matrix based on the pixel alignment direction, the first relative pose and the second relative pose;

[0020] Perform epipolar correction on the first image based on the first correction matrix to obtain a corrected first image, and perform epipolar correction on the second image based on the second correction matrix to obtain a corrected second image;

[0021] The pixel position corresponding to the laser projection direction of the TOF sensor in the first image is calculated based on the first correction matrix and mapped to the sub-pixel position in the corrected first image.

[0022] The visual scale correction method provided by the present invention can better assist in the matching of two frames of images through epipolar correction. An ideal relative pose relationship pair makes the epipole lie on the same pixel row / pixel column, greatly reducing the complexity of matching adjacent two frames of images and making it easier to find corresponding points. Epipolar correction gives important constraint conditions for corresponding points, compressing the search for corresponding points from the entire image to a single pixel row / pixel column, reducing the search range, and guiding stereo matching.

[0023] In an alternative embodiment, performing stereo matching on the corrected first image and the corrected second image to obtain a disparity map includes:

[0024] Performing stereo matching on the corrected first image and the corrected second image using a binocular stereo matching algorithm to obtain a disparity map.

[0025] The visual scale correction method provided by the present invention performs stereo matching on the corrected first image and the corrected second image using a binocular stereo matching algorithm to obtain a disparity map, making the obtained disparity map represent the difference between the first image and the second image, providing conditions for subsequent calculation of depth calculation values.

[0026] In an alternative embodiment, when the average disparity of all pixels is greater than a first specified threshold, calculating the depth calculation value corresponding to the sub-pixel position based on the disparity map and the relative distance between two frames of scene images includes:

[0027] Taking a preset value pixel image window centered on the sub-pixel position as a rectangular image region;

[0028] Obtaining the disparity value corresponding to each pixel in the rectangular image region based on the disparity map, and taking the relative distance between two frames of scene images as the baseline distance;

[0029] Calculating the depth value corresponding to each pixel in the rectangular image region based on the baseline distance, the camera internal orientation elements, and the disparity value corresponding to each pixel, and determining whether the difference between the maximum depth value and the minimum depth value of all pixels is less than a second specified threshold;

[0030] When it is less than the second specified threshold, obtaining the depth values of adjacent pixels at the sub-pixel position, and using an interpolation algorithm based on the adjacent pixel depth values to obtain the depth calculation value corresponding to the sub-pixel position.

[0031] The visual scale correction method provided by the present invention obtains the disparity value corresponding to each pixel in the rectangular image area based on the disparity map, and takes the relative distance between two frames of scene images as the baseline distance; calculates the depth value corresponding to each pixel in the rectangular image area based on the baseline distance, the camera internal orientation element and the disparity value corresponding to each pixel, and judges whether the difference between the maximum depth value and the minimum depth value of all pixels is less than a second specified threshold; when it is less than the second specified threshold, obtains the depth value of the adjacent pixel at the sub-pixel position, and adopts an interpolation algorithm based on the adjacent pixel depth value to obtain the depth calculation value corresponding to the sub-pixel position. The method has the advantages of high calculation efficiency, accurate and reliable calculation of the depth calculation value and easy implementation, and provides conditions for subsequent determination of the depth ratio.

[0032] In an optional implementation, determining a depth ratio based on the depth measurement value and the depth calculation value, and updating the visual scale within the current time window of the visual odometer based on the depth ratio includes:

[0033] The depth ratio is obtained by calculating the ratio of the depth measurement value to the depth calculation value;

[0034] Get the posture matrix and position coordinates of multiple frames of images calculated by the visual odometer in the current time window;

[0035] Keeping the posture and position of any specified frame image unchanged, calculating the product of the position coordinates of other frame images relative to the specified frame image and the depth ratio to obtain new relative position coordinates;

[0036] The visual scale within the current time window of the visual odometry is updated based on the pose matrix and the new relative position coordinates.

[0037] The visual scale correction method provided by the present invention obtains the posture matrix and position coordinates of multiple frames of images calculated by a visual odometer in the current time window; keeps the posture and position of any specified frame image unchanged, calculates the product of the position coordinates of other frame images except the specified frame image relative to the specified frame image and the depth ratio to obtain new relative position coordinates; achieves the purpose of updating the visual scale in the current time window of the visual odometer based on the posture matrix and the new relative position coordinates, introduces point laser ranging information into the visual odometer through the depth ratio, realizes continuous and reliable scale calibration, and can effectively overcome the scale drift and precision reduction problems of the visual odometer or visual inertial odometer in a uniform motion state.

[0038] In a second aspect, the present invention provides a visual scale correction device, which is applied to a monocular vision device with a point laser ranging, wherein the monocular vision device with a point laser ranging comprises a rigidly connected camera and a TOF sensor, and the TOF sensor emits and receives lasers in a direction within the observation field of view of the camera; the device comprises:

[0039] An image acquisition and posture calculation module is used to acquire two adjacent frames of scene images captured by the camera within a preset time interval during the motion process, and calculate the relative posture corresponding to the acquisition moment of the two frames of scene images and the relative distance between the two frames of scene images. The two frames of scene images are images containing the laser points emitted by the TOF sensor, including the first image and the second image, and the depth measurement value and pixel position of the TOF sensor at the acquisition moment of the first image are recorded at the same time;

[0040] An image correction module is used to select a direction for aligning image pixels based on the relative posture, and to perform epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and to calculate and map the pixel position corresponding to the laser projection direction of the TOF sensor in the first image to the sub-pixel position in the corrected first image;

[0041] An image stereo matching module, used for performing stereo matching on the corrected first image and the corrected second image to obtain a disparity map, and determining whether an average disparity of all pixels in the disparity map is greater than a first specified threshold;

[0042] A depth calculation value calculation module in the laser projection direction, used to calculate the depth calculation value corresponding to the sub-pixel position based on the disparity map and the relative distance between two frames of scene images when the average disparity of all pixels is greater than a first specified threshold;

[0043] The visual scale update module is used to determine the depth ratio based on the depth measurement value and the depth calculation value, and update the visual scale within the current time window of the visual odometer based on the depth ratio.

[0044] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the visual scale correction method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0045] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the visual scale correction method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 is a schematic flow chart of a visual scale correction method according to an embodiment of the present invention;

[0048] Figure 2 is a flow chart of another visual scale correction method according to an embodiment of the present invention;

[0049] Figure 3 is a flow chart of another visual scale correction method according to an embodiment of the present invention;

[0050] Figure 4 is a flow chart of yet another visual scale correction method according to an embodiment of the present invention;

[0051] Figure 5 is a schematic structural diagram of a point laser ranging device according to an embodiment of the present invention;

[0052] FIG6( a ) is a schematic diagram of the pose of a first image before and after epipolar correction according to an embodiment of the present invention;

[0053] FIG6( b ) is a schematic diagram of the pose of the second image before and after epipolar correction according to an embodiment of the present invention;

[0054] Figure 7 is a schematic diagram of an image window when calculating a depth calculation value according to an embodiment of the present invention;

[0055] Figure 8 is a structural block diagram of a visual scale correction device according to an embodiment of the present invention;

[0056] Fig. 9 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0058] According to an embodiment of the present invention, an embodiment of a visual scale correction method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0059] In this embodiment, a visual scale correction method is provided, which can be used for a monocular vision device with a point laser ranging. The monocular vision device with a point laser ranging includes a rigidly connected camera and a TOF sensor, and the direction of the TOF sensor transmitting and receiving lasers is within the observation field angle of the camera, such as Figure 5 As shown, the camera can be a visible light monocular camera, and the TOF sensor can be a single-point TOF sensor. Rigid connection means connecting the camera and the TOF sensor into a fixed and solid whole, keeping the geometric shape of the connecting parts unchanged, and the connecting parts will not produce relative movement when subjected to force. The relative position of the camera and the TOF sensor is known and fixed. Figure 1 is a flow chart of a visual scale correction method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0060] Step S101, obtaining two adjacent frames of scene images captured by the camera within a preset time interval during the motion process, and calculating the relative posture corresponding to the two frames of scene image capture time and the relative distance between the two frames of scene images, the two frames of scene images are images containing the laser points emitted by the TOF sensor, including the first image and the second image, and recording the depth measurement value and pixel position of the TOF sensor at the time of capturing the first image.

[0061] Specifically, the preset time interval is set according to the actual situation and is not specifically limited here. The two frames of scene images taken by the visible light camera at the preset time interval are obtained, and both frames of scene images are images containing the laser points emitted by the TOF sensor. The visual odometer can be used to calculate the relative pose corresponding to the acquisition time of the two frames of scene images, and record the depth measurement value Z of the TOF sensor at the time of shooting the first image (hereinafter referred to as image 1), and the pixel position p in the first image (image 1) corresponding to the measurement direction of the TOF sensor.

[0062] Step S102, select the direction of image pixel alignment based on the relative posture, and perform epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculate the pixel position corresponding to the laser projection direction of the TOF sensor in the first image and map it to the sub-pixel position in the corrected first image.

[0063] Specifically, epipolar correction refers to mathematically aligning the cameras to the same observation plane so that the pixel rows or pixel columns on the cameras are strictly aligned.

[0064] Based on the relative position and posture of the first image and the second image obtained in step S101, a vector pointing from the camera center position corresponding to image 1 to the camera center position corresponding to the second image (hereinafter referred to as image 2) is projected onto the XOY plane of the camera coordinate system corresponding to the first image to obtain a projection vector, and the magnitudes of the components of the projection vector in the X and Y directions are determined. If the component in the X direction is large, the direction of pixel alignment is selected as the x direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel row; if the component in the Y direction is large, the direction of pixel alignment is selected as the y direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel column.

[0065] Based on the selected pixel alignment direction, the relative position and posture of image 1 and image 2, image 1 and image 2 are epipolar corrected to obtain corrected image 1' and image 2'. At the same time, the sub-pixel position p' of pixel position p in image 1 after epipolar correction mapping in image 1' is recorded.

[0066] Step S103 , stereo matching is performed on the corrected first image and the corrected second image to obtain a disparity map, and it is determined whether the average disparity of all pixels in the disparity map is greater than a first specified threshold.

[0067] Specifically, stereo matching is an important task in computer vision, and its goal is to find matching corresponding points from images with different viewpoints. It involves finding corresponding points of the same object in images taken by two or more cameras to reconstruct a three-dimensional scene and obtain a disparity map. Stereo matching is performed on the corrected first image and the corrected second image to find corresponding points and obtain a disparity map. The first specified threshold is T d Determine whether the average disparity of all pixels in the disparity map is greater than the first specified threshold T d Make sure the camera moves a long enough distance in the scene and the binocular epipolar line is long enough so that the depth value calculated subsequently is accurate enough.

[0068] Step S104 : when the average disparity of all pixels is greater than a first specified threshold, a depth calculation value corresponding to the sub-pixel position is calculated based on the disparity map and the relative distance between the two frames of scene images.

[0069] Specifically, when the average disparity of all pixels is greater than the first specified threshold T d When the depth calculation value Z' corresponding to the sub-pixel position p' is obtained by linear interpolation based on the disparity map and the relative distance between image 1 and image 2.

[0070] When it is less than the first specified threshold T d When the camera is in motion, the process returns to step S101 to reacquire two adjacent frames of scene images captured by the camera within a preset time interval during the motion process.

[0071] Step S105 , determining a depth ratio based on the depth measurement value and the depth calculation value, and updating the visual scale within the current time window of the visual odometer based on the depth ratio.

[0072] Specifically, repeat steps S101 to S104 N times to obtain N depth measurement values ​​Z i (i=1,…,N) and the corresponding N depth calculation values ​​Z i '(i=1,…,N). Based on N depth measurement values ​​Z i and N depth calculation values ​​Z i 'Calculate the depth ratio s, and update the visual scale within the current time window based on the depth ratio s.

[0073] The visual scale correction method provided in this embodiment introduces the ranging information obtained by a monocular vision device with a point laser ranging into the visual odometer, obtains a depth measurement value through a TOF sensor, obtains a disparity map by correcting and stereo matching scene images taken at different times, calculates a depth calculation value corresponding to a sub-pixel position corresponding to the laser projection direction of the TOF sensor based on the disparity map and the relative distance between two frames of scene images, finally determines a depth ratio based on the depth measurement value and the depth calculation value, and updates the visual scale within the current time window of the visual odometer based on the depth ratio, thereby achieving continuous and reliable scale calibration, so that a long-term and reliable relative positioning function of an outdoor unmanned platform can be achieved by combining a monocular camera with a small single-point TOF sensor (Time-of-Flight, a sensor that measures distance using time-of-flight technology), thereby solving the problem that the visual scale cannot be accurately corrected due to scale drift and decreased accuracy of traditional visual odometers or visual inertial odometers in a uniform motion state.

[0074] In this embodiment, a visual scale correction method is provided, which can be used for a monocular vision device with a point laser ranging. The monocular vision device with a point laser ranging includes a rigidly connected camera and a TOF sensor, and the direction of the TOF sensor transmitting and receiving lasers is within the observation field angle of the camera, such as Figure 5 As shown, the camera can be a visible light monocular camera, and the TOF sensor can be a single-point TOF sensor. Rigid connection means connecting the camera and the TOF sensor into a fixed and solid whole, keeping the geometric shape of the connecting parts unchanged, and the connecting parts will not produce relative movement when subjected to force. The relative position of the camera and the TOF sensor is known and fixed. Figure 2is a flow chart of a visual scale correction method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0075] Step S201, calibrating the camera internal orientation elements and the pixel positions on the image corresponding to the TOF sensor laser projection direction.

[0076] Specifically, camera calibration refers to the process of determining the camera's internal parameters (such as focal length, principal point coordinates, etc.) and external parameters (such as the camera's position and posture relative to the world coordinate system).

[0077] The camera is calibrated, its intrinsic orientation elements are known (for example, the focal length of the camera is known), and the pixel position p on the image corresponding to the projection direction of the TOF sensor laser is a known quantity.

[0078] Step S202, obtaining two adjacent frames of scene images captured by the camera within a preset time interval during the motion process, and calculating the relative posture corresponding to the two frames of scene images at the time of capture and the relative distance between the two frames of scene images, the two frames of scene images are images containing the laser points emitted by the TOF sensor, including the first image and the second image, and the depth measurement value and pixel position of the TOF sensor at the time of capture of the first image are recorded at the same time.

[0079] Specifically, the above step S202 includes:

[0080] Step a. Calculate the first relative pose and the second relative pose corresponding to the first image and the second image acquisition time using the visual odometer in the visual odometer reference coordinate system, and obtain the relative distance between the two frames of scene images based on the first relative pose and the second relative pose. Pose is the abbreviation of position and attitude.

[0081] The steps of visual odometer calculation refer to the relevant technical content and will not be repeated here. The calculated pose corresponding to image 1 is (R org1 , t org1 ), where R org1 is the rotation matrix, t org1 is the translation vector. The pose corresponding to image 2 is (R org2 , t org2 ). Where R org2 is the rotation matrix, t org2 is the translation vector. The relative distance between image 1 and image 2 can be obtained by subtracting the relative position of image 2 from the relative position of image 1 and taking the modulus.

[0082] Step S203, select the direction of image pixel alignment based on the relative posture, and perform epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculate the pixel position corresponding to the laser projection direction of the TOF sensor in the first image and map it to the sub-pixel position in the corrected first image.

[0083] Specifically, the above step S203 includes:

[0084] Step S2031: Based on the relative posture, a vector pointing from the camera center position corresponding to the first image to the camera center position corresponding to the second image is projected onto the XOY plane of the camera coordinate system corresponding to the first image to obtain a projection vector, and the magnitudes of the components of the projection vector in the X and Y directions are determined. If the component in the X direction is large, the direction of pixel alignment is selected as the x direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel row; if the component in the Y direction is large, the direction of pixel alignment is selected as the y direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel column.

[0085] Specifically, based on the relative posture, the vector O pointing from the camera center position corresponding to image 1 to the camera center position corresponding to image 2 is 1 O 2 =R org1 t org1 -R org2 t org2 Project it onto the XOY plane of the camera coordinate system corresponding to image 1 and get the projection vector O 1 O 2 ',in:

[0086]

[0087] Determine the projection vector O 1 O 2 'The size of the components in the X and Y directions. If the X direction component is larger, the direction of pixel alignment is selected as the x direction of the image coordinate system, that is, the corresponding pixels after epipolar correction are on the same pixel row; if the Y direction component is larger, the direction of pixel alignment is selected as the y direction of the image coordinate system, that is, the corresponding pixels after epipolar correction are on the same pixel column.

[0088] Let the corrected pose corresponding to image 1 be (R new1 , t new1 ), the corrected pose corresponding to image 2 is (R new2 , t new2 ), according to the theory of binocular epipolar correction, the rotation matrices of image 1 and image 2 after correction are R new1 and R new2, needs to be selected within the range that satisfies the following restrictions: (1) If the pixel alignment direction is the x direction of the image coordinate system, then R new1 The first row is parallel to the line connecting the camera center corresponding to image 1 and the camera center corresponding to image 2 (i.e., vector O 1 O 2 =R org1 t org1 -R org2 t org2 If the pixel alignment direction is the y direction of the image coordinate system, then R new1 The second row is parallel to the line connecting the camera center corresponding to image 1 and the camera center corresponding to image 2, that is, R new1 =R new2 In summary, according to the selection of the pixel alignment direction, the corresponding pixels of the corrected first image and the corrected second image are located in the same pixel row or the same pixel column.

[0089] Step S2032: Calculate a first correction matrix and a second correction matrix based on the pixel alignment direction, the first relative posture, and the second relative posture.

[0090] Specifically, the rotation matrix R from the camera coordinate system of the image 1 before correction to the camera coordinate system of the image 1' after correction is upt1 It can be calculated by the following formula:

[0091] R upt1 =R new1 R org1 T (2);

[0092] The rotation matrix R from the camera coordinate system of the image 2 before correction to the camera coordinate system of the image 2' after correction upt2 It can be calculated by the following formula:

[0093] R upt2 =R new2 R org2 T (3);

[0094] The first correction matrix is ​​expressed as follows:

[0095] H 1 =KR upt1 K -1 (4);

[0096] Among them, K is the intrinsic parameter matrix of the camera, R upt1 is the rotation matrix from the camera coordinate system of the image 1 before correction to the camera coordinate system of the image 1' after correction.

[0097] The second correction matrix is ​​expressed as follows:

[0098] H2 =KR upt2 K -1 (5);

[0099] Among them, R upt2 is the rotation matrix from the camera coordinate system of the image 2 before correction to the camera coordinate system of the image 2' after correction.

[0100] Step S2033, performing epipolar correction on the first image based on the first correction matrix to obtain a corrected first image, and performing epipolar correction on the second image based on the second correction matrix to obtain a corrected second image.

[0101] Specifically, as shown in FIG6(a) and FIG6(b), the image 1 is subjected to epipolar correction based on the first correction matrix to obtain a corrected image 1', and the image 2 is subjected to epipolar correction based on the second correction matrix to obtain a corrected image 2'.

[0102] Step S2034, based on the first correction matrix, the pixel position corresponding to the laser projection direction of the TOF sensor in the first image is mapped to the sub-pixel position in the corrected first image.

[0103] Specifically, as shown in FIG6( a ), the calculation formula for mapping from the pixel position p in image 1 to the sub-pixel position p′ in image 1′ is:

[0104] P′=H 1 P (6);

[0105] Where P is the homogeneous coordinate of the pixel position p, P′ is the homogeneous coordinate of the sub-pixel position p′, and H 1 is the first correction matrix.

[0106] Step S204 , stereo matching is performed on the corrected first image and the corrected second image to obtain a disparity map, and it is determined whether the average disparity of all pixels in the disparity map is greater than a first specified threshold.

[0107] Specifically, the above step S204 includes:

[0108] Step b. Using a binocular stereo matching algorithm, stereo matching is performed on the corrected first image and the corrected second image to obtain a disparity map.

[0109] Specifically, stereo matching can be implemented using a binocular matching algorithm, such as a window-based local matching algorithm SAD (Sum of Absolute Differences), NCC (Normalized Cross Correlation), a global matching algorithm BP (Belief Propagation), GC (Grapth Cut) or a semi-global matching algorithm SGBM (Semi-Global BlockMatching, SGBM) and the like. This embodiment uses a semi-global matching algorithm SGBM (Semi-Global BlockMatching, SGBM) to perform binocular stereo matching on image 1' and image 2', and calculates a disparity map. After obtaining the disparity map, calculate whether the average disparity of all pixels is greater than a specified threshold value T d If not, return to step S202 to continue acquiring two adjacent frames of scene images captured by the camera within a preset time interval.

[0110] Step S205: When the average disparity of all pixels is greater than the first specified threshold, the depth calculation value corresponding to the sub-pixel position is calculated based on the disparity map and the relative distance between the two frames of scene images. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0111] Step S206: Determine the depth ratio based on the depth measurement value and the depth calculation value, and update the visual scale in the current time window of the visual odometer based on the depth ratio. Figure 1 Step S105 of the illustrated embodiment will not be described in detail here.

[0112] The visual scale correction method provided in this embodiment calibrates the pixel positions on the image corresponding to the orientation elements in the camera and the laser projection direction of the TOF sensor, which is conducive to obtaining the pixel position value on the image corresponding to the laser projection direction of the TOF sensor, improving the precision and accuracy of the scale correction, avoiding the accuracy problems caused by the deviation of the shooting angle and position, and providing conditions for accurately obtaining the depth measurement value and the calculated value. In the coordinate system where the camera is located, the visual odometer is used to calculate the first relative posture and the second relative posture corresponding to the acquisition time of the first image and the second image, and the movement of the camera between the two frames of scene images is estimated to achieve the accurate calculation of the first relative posture and the second relative posture corresponding to the acquisition time of the first image and the second image and the relative distance between the two frames of images. The polar line correction can better assist the matching of the two frames of images. The ideal relative posture relationship after correction makes the pole points on the same pixel row or the same pixel column, which greatly reduces the computational complexity of matching the two adjacent frames of images, making it easier to find corresponding points. The polar line correction gives important constraints on the corresponding points, and compresses the corresponding point matching from the entire image to a pixel row or pixel column, which reduces the search range and plays a guiding role in stereo matching.

[0113] In this embodiment, a visual scale correction method is provided, which can be used for a monocular vision device with a point laser ranging. The monocular vision device with a point laser ranging includes a rigidly connected camera and a TOF sensor, and the direction of the TOF sensor transmitting and receiving lasers is within the observation field angle of the camera, such as Figure 5 As shown, the camera can be a visible light monocular camera, and the TOF sensor can be a single-point TOF sensor. Rigid connection means connecting the camera and the TOF sensor into a fixed and solid whole, keeping the geometric shape of the connecting parts unchanged, and the connecting parts will not produce relative movement when subjected to force. The relative position of the camera and the TOF sensor is known and fixed. Figure 3 is a flow chart of a visual scale correction method according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0114] Step S301, obtain two adjacent frames of scene images captured by the camera within a preset time interval during the motion process, and calculate the relative posture corresponding to the two frames of scene image capture time and the relative distance between the two frames of scene images. The two frames of scene images are images containing the laser points emitted by the TOF sensor, including the first image and the second image, and record the depth measurement value and pixel position of the TOF sensor at the time of capturing the first image. For details, please refer to Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.

[0115] Step S302, based on the relative posture, select the direction of image pixel alignment, and perform epipolar correction on the first image and the second image based on the pixel alignment direction to obtain the corrected first image and the corrected second image, and calculate the pixel position corresponding to the TOF sensor laser projection direction in the first image and map it to the sub-pixel position in the corrected first image. For details, please refer to Figure 2 Step S203 of the illustrated embodiment will not be described in detail here.

[0116] Step S303, stereo matching is performed on the corrected first image and the corrected second image to obtain a disparity map, and it is determined whether the average disparity of all pixels in the disparity map is greater than a first specified threshold. Figure 2 Step S204 of the illustrated embodiment will not be described in detail here.

[0117] Step S304 : when the average disparity of all pixels is greater than the first specified threshold, a depth calculation value corresponding to the sub-pixel position is calculated based on the disparity map and the relative distance between the two frames of scene images.

[0118] Specifically, the disparity map is used in combination with the relative distance between cameras between the time when image 1 is acquired and the time when image 2 is acquired to calculate the depth calculation values ​​of n pixels in the image area D near the sub-pixel position p'.

[0119] The above step S304 includes:

[0120] Step S3041: a preset pixel image window centered at a sub-pixel position is used as a rectangular image area.

[0121] Specifically, Figure 7 As shown in the figure, a 4×4 pixel image window centered at the sub-pixel position p' is taken as the area D, and D is the area around the pixel p 11 , p 14 , p 41 and p 44 A rectangular image region with vertices.

[0122] Step S3042: obtaining the disparity value corresponding to each pixel in the rectangular image area based on the disparity map, and taking the relative distance between the two frames of scene images as the baseline distance.

[0123] Specifically, the relative distance between image 1 and image 2, ie, the baseline distance, can be obtained by taking the modulus of a vector obtained by subtracting the relative position of image 2 from the relative position of image 1.

[0124] Step S3043, calculating the depth value corresponding to each pixel in the rectangular image area based on the baseline distance, the camera internal orientation element and the disparity value corresponding to each pixel, and determining whether the difference between the maximum depth value and the minimum depth value of all pixels is less than a second specified threshold.

[0125] Specifically, the depth value calculation formula corresponding to each pixel is as follows:

[0126]

[0127] Where b is the baseline distance between two frames of scene images calculated by the visual odometer, f is the focal length of the camera, and d is xy is pixel p xy The corresponding disparity value.

[0128] The second specified threshold is T Z It indicates that judging whether the difference between the maximum depth value and the minimum depth value of all pixels is less than the second specified threshold value is to ensure that the depth value at the laser pointing point is stable, changes smoothly, and there is no depth mutation or depth fault. When it is greater than the second specified threshold value, continue to return to step S301 to obtain two adjacent frames of scene images collected by the camera within a preset time interval.

[0129] Step S3044: when the value is less than the second specified threshold, the depth values ​​of the adjacent pixels of the sub-pixel position are obtained, and an interpolation algorithm is used based on the depth values ​​of the adjacent pixels to obtain a depth calculation value corresponding to the sub-pixel position.

[0130] Specifically, the interpolation algorithm of this embodiment adopts a bilinear interpolation algorithm. The depth calculation value Z' at the sub-pixel position p' is obtained by interpolation. Figure 7 As shown, the four pixels p around the sub-pixel position p' can be used 22 , p 23 , p 32 and p 33 The depth value Z 22 , Z 23 , Z 32 and Z 33 , for the depth value Z 22 , Z 23 , Z 32 and Z 33 The depth calculation value Z' is obtained by bilinear interpolation. The specific content of the bilinear interpolation algorithm can be found in the related art, which will not be described here.

[0131] Step S305 , determining a depth ratio based on the depth measurement value and the depth calculation value, and updating the visual scale within the current time window of the visual odometer based on the depth ratio.

[0132] Specifically, the depth ratio s is used to update the scale of the visual odometer in the current time window. The updating method is to keep a certain frame image F in the current time window. j The position and posture of the other image frames are unchanged relative to image F jThe three components of the position X, Y, and Z are multiplied by the depth ratio s, and all landmark points observed in the current time window are relative to the image F j The three components X, Y, and Z of the position corresponding to the camera coordinate system are multiplied by the depth ratio s, and the updated position is used to update the visual scale within the current time window of the visual odometer.

[0133] The above step S305 includes:

[0134] Step S3051: Calculate the ratio of the depth measurement value to the depth calculation value to obtain a depth ratio.

[0135] Specifically, repeat steps S301 to S304 N times to obtain N depth measurement values ​​Z i (i=1,…,N) and the corresponding N depth calculation values ​​Z i '(i=1,…,N). Based on N depth measurement values ​​Z i and N depth calculation values ​​Z i 'Calculate the depth ratio s, the specific formula is as follows:

[0136]

[0137] Step S3052, obtaining the posture matrix and position coordinates of multiple frames of images calculated by the visual odometer within the current time window.

[0138] Specifically, suppose there are M frames of image F in the current time window 1 , F 2 , …, F M In the visual odometer reference coordinate system, the visual odometer is used to calculate the posture matrix (also called rotation matrix) of each frame image, which is R 1 , R 2 , …, R M , and the position coordinates (there is a fixed conversion relationship between the position in the camera coordinate system and the translation vector, which can be expressed by the translation vector) are O 1 , O 2 , …, O M .

[0139] Step S3053, keeping the posture and position of any designated frame image unchanged, calculating the product of the position coordinates of other frame images except the designated frame image relative to the designated frame image and the depth ratio to obtain new relative position coordinates.

[0140] Specifically, keep the j-th frame image F j The posture and position are fixed, and any other frame image F i Relative to F j The posture matrix is ​​R i *R jT , position R j *(O i -O j ), the updated posture matrix remains unchanged, and the position becomes s*R j *(O i -O j ).

[0141] Step S3054: Update the visual scale within the current time window of the visual odometer based on the posture matrix and the new relative position coordinates.

[0142] Specifically, the visual scale update is converted to the visual odometer reference coordinate system, and the previous image frame F is updated i The posture matrix is ​​R i , position is O i , updated image frame F i The posture matrix is ​​R i , position is O j +s*(O i -O j ).

[0143] Similarly, if there are K landmarks in the current time window, the position of each landmark before the update is L i , then the updated position can be expressed as O j +s*(L i -O j ).

[0144] It should be noted that the visual scale correction method provided in this embodiment can be extended to the combination of a monocular camera + a multi-point TOF sensor. It is only necessary to take a weighted average of the multiple depth ratios s calculated, and update the visual scale within the current time window of the visual odometer based on the depth ratio after the weighted average.

[0145] The visual scale correction method provided in this embodiment obtains the disparity value corresponding to each pixel in the rectangular image area based on the disparity map, and uses the relative distance between the two frames of scene images as the baseline distance; calculates the depth value corresponding to each pixel in the rectangular image area based on the baseline distance, the camera internal orientation element and the disparity value corresponding to each pixel, and determines whether the difference between the maximum depth value and the minimum depth value of all pixels is less than the second specified threshold; when it is less than the second specified threshold, obtains the depth value of the adjacent pixel at the sub-pixel position, and uses the interpolation algorithm based on the depth value of the adjacent pixel to obtain the depth calculation value corresponding to the sub-pixel position, which has the advantages of high calculation efficiency, accurate calculation of depth calculation value and easy implementation. Obtain the posture matrix and position coordinates of multiple frames of images calculated by the visual odometer in the current time window; keep the posture and position of any specified frame image unchanged, calculate the product of the position coordinates of other frame images relative to the specified frame image and the depth ratio to obtain the new relative position coordinates; realizes the purpose of updating the visual scale in the current time window of the visual odometer based on the depth ratio.

[0146] As one or more specific application embodiments of the present invention, Figure 4 The visual scale correction method provided by the present invention is further described in detail, and the specific process is as follows:

[0147] Step (1) obtains two frames of scene images, image 1 and image 2, taken by a visible light camera at a specified time interval during the motion process, and records the depth measurement value Z of the TOF sensor at the time of image 1 shooting, as well as the pixel position p in image 1 corresponding to the measurement direction of the TOF sensor, and calculates the relative position and posture (abbreviated as posture) of the camera corresponding to the two frames of scene image acquisition time through the visual odometer. The calculated posture corresponding to image 1 is (R org1 , t org1 ), where R org1 is the rotation matrix, t org1 is the translation vector. The pose corresponding to image 2 is (R org2 , t org2 ). Where R org2 is the rotation matrix, t org2 is the translation vector.

[0148] Step (2), using the relative position and posture obtained in step (1), select the direction of image pixel alignment, and perform epipolar correction on image 1 and image 2 based on the pixel alignment direction to obtain corrected images 1' and 2'. According to the selection of the pixel alignment direction, the corresponding pixels of image 1' and image 2' are in the same pixel row or the same pixel column. At the same time, the pixel position p corresponding to the laser projection direction of the TOF sensor in image 1 is calculated, and the sub-pixel position p' in image 1' is mapped by epipolar correction.

[0149] Based on the relative pose, the vector O pointing from the camera center position corresponding to image 1 to the camera center position corresponding to image 2 1 O 2 =R org1 t org1 -R org2 t org2 Project it onto the XOY plane of the camera coordinate system corresponding to image 1 and get the projection vector O 1 O 2 ',in:

[0150]

[0151] Determine the projection vector O 1 O 2 'The size of the components in the X and Y directions. If the X direction component is larger, the direction of pixel alignment is selected as the x direction of the image coordinate system, that is, the corresponding pixels after epipolar correction are on the same pixel row; if the Y direction component is larger, the direction of pixel alignment is selected as the y direction of the image coordinate system, that is, the corresponding pixels after epipolar correction are on the same pixel column.

[0152] Let the corrected pose corresponding to image 1 be (R new1 , t new1 ), the corrected pose corresponding to image 2 is (R new2 , t new2 ), according to the theory of binocular epipolar correction, the rotation matrices of image 1 and image 2 after correction are R new1 and R new2 , needs to be selected within the range that satisfies the following restrictions: (1) If the pixel alignment direction is the x direction of the image coordinate system, then R new1 The first row is parallel to the line connecting the camera center corresponding to image 1 and the camera center corresponding to image 2 (i.e., vector O 1 O 2 =R org1 t org1 -R org2 t org2 If the pixel alignment direction is the y direction of the image coordinate system, then R new1 The second row is parallel to the line connecting the camera center corresponding to image 1 and the camera center corresponding to image 2, R new1 =R new2 .

[0153] The rotation matrix R from the camera coordinate system of the image 1 before correction to the camera coordinate system of the image 1' after correction upt1 It can be calculated by the following formula:

[0154] R upt1 =R new1 Rorg1 T (2);

[0155] The rotation matrix R from the camera coordinate system of the image 2 before correction to the camera coordinate system of the image 2' after correction upt2 It can be calculated by the following formula:

[0156] R upt2 =R new2 R org2 T (3);

[0157] The correction matrix from image 1 to image 1' can be expressed as:

[0158] H 1 =KR upt1 K -1 (4);

[0159] Among them, K is the intrinsic parameter matrix of the camera, R upt1 is the rotation matrix from the camera coordinate system of the image 1 before correction to the camera coordinate system of the image 1' after correction.

[0160] The correction matrix from image 2 to image 2' can be expressed as:

[0161] H 2 =KR upt2 K -1 (5);

[0162] Among them, R upt2 is the rotation matrix from the camera coordinate system of the image 2 before correction to the camera coordinate system of the image 2' after correction.

[0163] As shown in Figure 6(a) and Figure 6(b), the calculation formula for mapping the pixel position p in image 1 to the sub-pixel position p' in image 1' is:

[0164] P′=H 1 P (6);

[0165] Where P is the homogeneous coordinate of the pixel position p, P′ is the homogeneous coordinate of the sub-pixel position p′, and H 1 is the first correction matrix.

[0166] Step (3) Perform binocular stereo matching on image 1' and image 2' to obtain a disparity map. Use the Semi-Global Block Matching (SGBM) algorithm to match image 1' and image 2' and calculate the disparity map. After obtaining the disparity map, calculate whether the average disparity of all pixels is greater than the specified threshold T d , if not, return to step (1) to continue acquiring images;

[0167] Step (4) uses the disparity map and the relative distance between the cameras at the time of capturing the two frames of scene images to calculate the depth values ​​of n pixels in the image area D near the sub-pixel position p'.

[0168] Specifically, a 4×4 pixel image window centered at p' is taken as region D, as Figure 7 As shown, D is the pixel p 11 , p 14 , p 41 and p 44 For each pixel p in the area xy , its depth value Z xy The calculation formula is:

[0169]

[0170] Where b is the baseline distance between two frames of scene images calculated by the visual odometer, f is the focal length of the camera, and d is xy is pixel p xy The corresponding disparity value.

[0171] Determine whether the difference between the maximum depth value and the minimum depth value of all pixels in area D is less than the specified threshold T Z , if not, return to step (1) to continue acquiring images;

[0172] Step (5) is to obtain the depth calculation value Z' at the sub-pixel position p' by interpolation. Specifically, the depth calculation value Z' at the sub-pixel position p' can be obtained by using the four pixels p around p'. 22 , p 23 , p 32 and p 33 The depth value Z 22 , Z 23 , Z 32 and Z 33 , the depth calculation value Z' at the sub-pixel position p' is obtained by bilinear interpolation.

[0173] Step (6), repeat steps (1) to (5) N times to obtain N depth measurement values ​​Z i (i=1,…,N) and the corresponding N depth calculation values ​​Z i '(i=1,…,N). The depth ratio s is obtained using the following formula;

[0174]

[0175] Step (7), using the depth ratio s to update the visual scale in the current time window of the visual odometer, the updating method is to keep a certain frame image F in the current time window j The position and posture of the other image frames are unchanged relative to image Fj The three components of the position X, Y, and Z are multiplied by the depth ratio s, and all landmark points observed in the current time window are relative to the image F j The three components X, Y, and Z corresponding to the camera coordinate system are multiplied by the depth ratio s.

[0176] Specifically, suppose there are M frames of image F in the current time window 1 , F 2 , …, F M The pose matrix of the image calculated by the visual odometer relative to the camera coordinate system is: R 1 , R 2 , …, R M , position is O 1 , O 2 , …, O M . Keep the j-th frame image F j The posture and position are fixed, and any other frame F i Relative to F j The posture matrix is ​​R i *R j T , position R j *(O i -O j ), the updated posture matrix remains unchanged, and the position becomes s*R j *(O i -O j ). Transform the scale update to the reference frame and update the previous image frame F i The posture matrix is ​​R i , position is O i , updated image frame F i The posture matrix is ​​R i , position is O j +s*(O i -O j ).

[0177] Similarly, suppose there are K landmarks in the current time window. The position of each landmark before updating is L i , the updated position can be expressed as O j +s*(L i -O j ).

[0178] Step (8), return to step (1) to continue acquiring the scene image.

[0179] The present invention introduces the distance measurement information of point laser into the visual odometer to achieve continuous and reliable scale calibration, so that the combination of a monocular camera and a small single-point TOF sensor can realize the long-term reliable relative positioning function of an outdoor unmanned platform. It can effectively overcome the scale drift and accuracy reduction problems of visual odometers or visual inertial odometers in a uniform motion state.

[0180] In this embodiment, a visual scale correction device is also provided, which is used to implement the above embodiments and preferred implementations, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0181] This embodiment provides a visual scale correction device, which is applied to a monocular vision device with a point laser ranging. The monocular vision device with a point laser ranging includes a rigidly connected camera and a TOF sensor, and the direction of the TOF sensor transmitting and receiving lasers is within the observation field angle of the camera; Figure 8 As shown, including:

[0182] The image acquisition and posture calculation module 801 is used to acquire two adjacent frames of scene images captured by the camera within a preset time interval during movement, and calculate the relative posture corresponding to the two frames of scene image acquisition time and the relative distance between the two frames of scene images. The two frames of scene images are images containing the laser points emitted by the TOF sensor, including the first image and the second image, and the depth measurement value and pixel position of the TOF sensor at the time of capturing the first image are recorded at the same time.

[0183] The image correction module 802 is used to select the direction of image pixel alignment based on the relative posture, and perform epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculate the pixel position corresponding to the TOF sensor laser projection direction in the first image and map it to the sub-pixel position in the corrected first image.

[0184] The image stereo matching module 803 is used to perform stereo matching on the corrected first image and the corrected second image to obtain a disparity map, and determine whether the average disparity of all pixels in the disparity map is greater than a first specified threshold.

[0185] The laser projection direction depth calculation value calculation module 804 is used to calculate the depth calculation value corresponding to the sub-pixel position based on the disparity map and the relative distance between two frames of scene images when the average disparity of all pixels is greater than a first specified threshold;

[0186] The visual scale updating module 805 is used to determine a depth ratio based on the depth measurement value and the depth calculation value, and update the visual scale within the current time window of the visual odometer based on the depth ratio.

[0187] In some optional implementations, the image acquisition and posture calculation module 801 includes:

[0188] The relative posture calculation unit is used to calculate the first relative posture and the second relative posture corresponding to the first image and the second image acquisition time using the visual odometer in the visual odometer reference coordinate system, and obtain the relative distance between the two frames of scene images based on the first relative posture and the second relative posture.

[0189] In some optional implementations, the image correction module 802 includes:

[0190] The alignment unit is used to project the vector pointing from the camera center position corresponding to the first image to the camera center position corresponding to the second image onto the XOY plane of the camera coordinate system corresponding to the first image based on the relative posture, obtain the projection vector, and determine the magnitude of the components of the projection vector in the X direction and the Y direction. If the component in the X direction is large, the direction of pixel alignment is selected as the x direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel row; if the component in the Y direction is large, the direction of pixel alignment is selected as the y direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel column.

[0191] The correction matrix calculation unit is used to calculate the first correction matrix and the second correction matrix based on the pixel alignment direction, the first relative posture and the second relative posture.

[0192] The epipolar correction unit is used to perform epipolar correction on the first image based on the first correction matrix to obtain a corrected first image, and to perform epipolar correction on the second image based on the second correction matrix to obtain a corrected second image.

[0193] The pixel position mapping unit is used to calculate the pixel position corresponding to the laser projection direction of the TOF sensor in the first image based on the first correction matrix and map it to the sub-pixel position in the corrected first image.

[0194] In some optional implementations, the image stereo matching module 803 includes:

[0195] The binocular stereo matching unit is used to perform stereo matching on the corrected first image and the corrected second image using a binocular stereo matching algorithm to obtain a disparity map.

[0196] In some optional implementations, the depth calculation value calculation module 804 includes:

[0197] The image area determination unit is used to determine a preset pixel image window centered at a sub-pixel position as a rectangular image area.

[0198] The disparity value and baseline distance calculation unit is used to obtain the disparity value corresponding to each pixel in the rectangular image area based on the disparity map, and use the relative distance between two frames of scene images as the baseline distance.

[0199] The depth value calculation unit is used to calculate the depth value corresponding to each pixel in the rectangular image area based on the baseline distance, the camera internal orientation element and the disparity value corresponding to each pixel, and determine whether the difference between the maximum depth value and the minimum depth value of all pixels is less than a second specified threshold.

[0200] The interpolation unit is used to obtain the depth value of the adjacent pixel at the sub-pixel position when it is less than the second specified threshold, and obtain the depth calculation value corresponding to the sub-pixel position by using an interpolation algorithm based on the adjacent pixel depth value.

[0201] In some optional implementations, the visual scale update module 805 includes:

[0202] The depth ratio calculation unit is used to calculate the ratio of the depth measurement value and the depth calculation value to obtain the depth ratio.

[0203] The pose calculation unit is used to obtain the pose matrix and position coordinates of multiple frames of images calculated by the visual odometer within the current time window.

[0204] The position updating unit is used to keep the posture and position of any specified frame image unchanged, and calculate the product of the position coordinates of other frame images relative to the specified frame image and the depth ratio to obtain new relative position coordinates.

[0205] The visual scale update unit is used to update the visual scale within the current time window of the visual odometer based on the posture matrix and the new relative position coordinates.

[0206] In some optional implementations, the visual scale correction device further includes:

[0207] The calibration module is used to calibrate the orientation elements in the camera and the pixel positions on the image corresponding to the laser projection direction of the TOF sensor.

[0208] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0209] The visual scale correction device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0210] The embodiment of the present invention also provides a computer device having the above Figure 8 The visual scale correction device shown.

[0211] See also Fig. 9 , Fig. 9 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Fig. 9 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig. 9 A processor 10 is taken as an example.

[0212] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0213] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0214] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0215] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0216] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.

[0217] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0218] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading over a network from an original storage in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0219] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A visual scale correction method, characterized in that: The method is applied to a monocular vision device with a point laser ranging, wherein the monocular vision device with a point laser ranging comprises a rigidly connected camera and a TOF sensor, and the directions of emitting and receiving lasers of the TOF sensor are within the observation field angle of the camera; the method comprises: Acquire two adjacent frames of scene images captured by the camera within a preset time interval during the motion process, and calculate the relative posture corresponding to the capture moments of the two frames of scene images and the relative distance between the two frames of scene images, wherein the two frames of scene images are images containing laser points emitted by the TOF sensor, including a first image and a second image, and simultaneously record the depth measurement value and pixel position of the TOF sensor at the capture moment of the first image; Selecting a direction for aligning image pixels based on the relative posture, and performing epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculating a pixel position corresponding to a laser projection direction of the TOF sensor in the first image and mapping it to a sub-pixel position in the corrected first image; Performing stereo matching on the corrected first image and the corrected second image to obtain a disparity map, and determining whether an average disparity of all pixels in the disparity map is greater than a first specified threshold; When the average disparity of all pixels is greater than a first specified threshold, calculating a depth calculation value corresponding to the sub-pixel position based on the disparity map and a relative distance between two frames of scene images; A depth ratio is determined based on the depth measurement value and the depth calculation value, and a visual scale within a current time window of the visual odometry is updated based on the depth ratio.

2. The method according to claim 1, characterized in that Before acquiring two adjacent frames of scene images captured by the camera within a preset time interval, the method further includes: The camera internal orientation elements and the pixel positions on the image corresponding to the TOF sensor laser projection direction are calibrated.

3. The method according to claim 1, characterized in that Calculating the relative position and distance between the two frames of scene images at the time of acquisition includes: A visual odometer is used in a visual odometer reference coordinate system to calculate a first relative pose and a second relative pose corresponding to the first image and the second image acquisition moments, and a relative distance between two frames of scene images is obtained based on the first relative pose and the second relative pose.

4. The method according to claim 3, characterized in that Selecting a direction for aligning image pixels based on the relative posture, and performing epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculating a pixel position corresponding to a laser projection direction of a TOF sensor in the first image and mapping it to a sub-pixel position in the corrected first image includes: Based on the relative posture, a vector pointing from the camera center position corresponding to the first image to the camera center position corresponding to the second image is projected onto the XOY plane of the camera coordinate system corresponding to the first image to obtain a projection vector, and the magnitudes of the components of the projection vector in the X direction and the Y direction are determined. If the component in the X direction is large, the direction of pixel alignment is selected as the x direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel row; if the component in the Y direction is large, the direction of pixel alignment is selected as the y direction of the image coordinate system, that is, the corresponding pixels after the epipolar correction are on the same pixel column; Calculate a first correction matrix and a second correction matrix based on the pixel alignment direction, the first relative posture and the second relative posture; Perform epipolar correction on the first image based on the first correction matrix to obtain a corrected first image, and perform epipolar correction on the second image based on the second correction matrix to obtain a corrected second image; The pixel position corresponding to the laser projection direction of the TOF sensor in the first image is calculated based on the first correction matrix and mapped to the sub-pixel position in the corrected first image.

5. The method according to claim 1, characterized in that Performing stereo matching on the corrected first image and the corrected second image to obtain a disparity map includes: A binocular stereo matching algorithm is used to perform stereo matching on the corrected first image and the corrected second image to obtain a disparity map.

6. The method according to claim 2, characterized in that When the average disparity of all pixels is greater than a first specified threshold, calculating the depth calculation value corresponding to the sub-pixel position based on the disparity map and the relative distance between the two frames of scene images includes: A preset pixel image window centered at the sub-pixel position is used as a rectangular image area; Based on the disparity map, the disparity value corresponding to each pixel in the rectangular image area is obtained, and the relative distance between the two frames of scene images is used as the baseline distance; Calculate the depth value corresponding to each pixel in the rectangular image area based on the baseline distance, the camera internal orientation element and the disparity value corresponding to each pixel, and determine whether the difference between the maximum depth value and the minimum depth value of all pixels is less than a second specified threshold; When it is less than the second specified threshold, the depth values ​​of the adjacent pixels at the sub-pixel position are obtained, and an interpolation algorithm is used based on the depth values ​​of the adjacent pixels to obtain a depth calculation value corresponding to the sub-pixel position.

7. The method according to claim 1, characterized in that Determining a depth ratio based on the depth measurement value and the depth calculation value, and updating a visual scale within a current time window of the visual odometer based on the depth ratio includes: The depth ratio is obtained by calculating the ratio of the depth measurement value to the depth calculation value; Get the posture matrix and position coordinates of multiple frames of images calculated by the visual odometer in the current time window; Keeping the posture and position of any specified frame image unchanged, calculating the product of the position coordinates of other frame images relative to the specified frame image and the depth ratio to obtain new relative position coordinates; The visual scale within the current time window of the visual odometry is updated based on the pose matrix and the new relative position coordinates.

8. A visual scale correction device, characterized in that: Applicable to a monocular vision device with point laser ranging, the monocular vision device with point laser ranging includes a rigidly connected camera and a TOF sensor, and the TOF sensor emits and receives lasers in a direction within the observation field of view of the camera; the device includes: An image acquisition and posture calculation module is used to acquire two adjacent frames of scene images captured by the camera within a preset time interval during the motion process, and calculate the relative posture corresponding to the acquisition moment of the two frames of scene images and the relative distance between the two frames of scene images. The two frames of scene images are images containing the laser points emitted by the TOF sensor, including the first image and the second image, and at the same time record the depth measurement value and pixel position of the TOF sensor at the acquisition moment of the first image; An image correction module, configured to select a direction of image pixel alignment based on the relative posture, and perform epipolar correction on the first image and the second image based on the pixel alignment direction to obtain a corrected first image and a corrected second image, and calculate a pixel position corresponding to a laser projection direction of the TOF sensor in the first image and map it to a sub-pixel position in the corrected first image; An image stereo matching module, used for performing stereo matching on the corrected first image and the corrected second image to obtain a disparity map, and determining whether an average disparity of all pixels in the disparity map is greater than a first specified threshold; A depth calculation value calculation module in the laser projection direction, used to calculate the depth calculation value corresponding to the sub-pixel position based on the disparity map and the relative distance between two frames of scene images when the average disparity of all pixels is greater than a first specified threshold; A visual scale update module is used to determine a depth ratio based on the depth measurement value and the depth calculation value, and to update a visual scale within a current time window of the visual odometer based on the depth ratio.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the visual scale correction method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the visual scale correction method according to any one of claims 1 to 7.