A device and method for three-dimensional reconstruction of visual guidance under binocular camera and laser radar fusion
By using vision-guided binocular camera and LiDAR fusion technology, and leveraging sparse laser constraints and adaptive parameter optimization, high-precision 3D reconstruction in complex environments was achieved, reducing data redundancy and computational load, and improving the stability and efficiency of 3D reconstruction.
Patent Information
- Application Number
- CN202610536431.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-24
AI Technical Summary
In existing 3D reconstruction technologies, the fusion scheme of binocular vision and LiDAR fails to effectively utilize the sparse depth information of LiDAR, resulting in unstable reconstruction accuracy in complex environments. Furthermore, the large scanning redundancy of LiDAR increases the computational and transmission burden.
By employing a vision-guided approach, the laser ranging unit is directionally scanned through a beam deflection scanning unit. Combined with sparse laser constraints and parameter adaptive fusion optimization, a dense disparity map and depth model are generated, reducing invalid sampling in non-interested areas and improving the stability and accuracy of stereo matching.
It improves matching stability and reconstruction accuracy in complex areas such as weak texture, strong reflection and occlusion, reduces data redundancy and backend computing load, and improves the overall efficiency of 3D reconstruction.
Smart Images

Figure CN122454038A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual recognition technology, and in particular to a three-dimensional reconstruction device and method that integrates a binocular camera and a lidar under visual guidance. Background Technology
[0002] 3D reconstruction technology is widely used in fields such as autonomous driving, reverse engineering, and robot navigation. Existing 3D reconstruction solutions are mainly divided into two categories: one is a pure vision solution (such as binocular stereo vision), which has high resolution and low cost, but is prone to matching failure in areas with weak texture (such as white walls and glass), and model parameters (such as the smoothing term penalty coefficient) need to be manually adjusted, making it difficult to adapt to complex environments; the other is a LiDAR solution, which has high ranging accuracy and strong anti-interference ability, but high-resolution LiDAR is expensive, and existing LiDARs mostly use uniform scanning across the entire field of view. In scenes where targets are sparse or the background is dominant, a large number of laser pulses fall on non-interested areas such as the sky and the ground, resulting in low effective sampling rate and large amount of redundant data, which increases the pressure on backend computing and transmission.
[0003] While existing technologies offer fusion solutions for binocular vision and LiDAR, they are limited to spatial calibration and coordinate alignment at the data level, or simple post-fusion after each device independently estimates depth. In these solutions, the LiDAR employs a full-field uniform scanning or fixed scanning strategy, which cannot avoid invalid sampling of non-interested areas, resulting in significant data redundancy and computational burden. Furthermore, they do not fully utilize the high-confidence sparse depth information of LiDAR, and stereo matching still relies on manual parameter tuning based on experience. This can easily lead to unstable matching and fluctuations in reconstruction accuracy in complex environments such as weak textures, reflections, and occlusions.
[0004] Therefore, there is an urgent need for a fusion-based 3D reconstruction scheme to solve the problems of large redundancy in lidar scanning and unstable 3D reconstruction accuracy caused by the failure to utilize sparse laser depth information during the fusion process in existing technologies. Summary of the Invention
[0005] To address the aforementioned technical problem in the prior art, namely the unstable accuracy of 3D reconstruction based on the fusion of binocular cameras and lidar, this application provides a vision-guided 3D reconstruction apparatus and method based on the fusion of binocular cameras and lidar.
[0006] In a first aspect of this application, a 3D reconstruction apparatus that fuses a binocular camera and a lidar under vision guidance is provided, comprising:
[0007] A binocular vision imaging unit is used to capture two-dimensional images of the left and right views of the field of view.
[0008] A beam deflection scanning unit is disposed in the emission optical path of the laser ranging unit. It is used to receive the scanning control signal output by the main control processing unit and change the emission direction of the ranging laser beam according to the scanning control signal, so that the ranging laser beam is pointed to the region of interest and scans the region of interest.
[0009] A laser ranging unit is used to emit the ranging laser beam and receive the echo signal, and output the distance information of each scanning point in the region of interest in the direction of the echo signal.
[0010] The main control processing unit is used for:
[0011] Receive the two-dimensional image output by the binocular vision imaging unit;
[0012] Detect whether the region of interest exists in the two-dimensional image; if the region of interest does not exist, return to the step of receiving the two-dimensional image output by the binocular vision imaging unit.
[0013] If the region of interest exists, the pixel positions of the region of interest in the two-dimensional image are extracted, and based on the preset calibration parameters and the pixel positions, the azimuth and elevation angles of each scanning point in the region of interest relative to the reference axis are obtained. The reference axis is the emission center axis of the ranging laser beam when the beam deflection scanning unit is in the zero position. The reference axis is jointly calibrated and determined by the binocular vision imaging unit, the beam deflection scanning unit and the laser ranging unit.
[0014] A scanning control signal is generated based on the azimuth and elevation angles of each scanning point, and the scanning control signal is output to the beam deflection scanning unit.
[0015] The system receives the distance information output by the laser ranging unit and obtains a three-dimensional reconstruction model of the region of interest based on the distance information and angle information, wherein the angle information includes the azimuth angle and the pitch angle.
[0016] Optionally, when the main control processing unit obtains the 3D reconstruction model of the region of interest based on the distance information and angle information, it is used to:
[0017] Based on the distance and angle information, sparse 3D point cloud data corresponding to the region of interest is calculated.
[0018] The sparse 3D point cloud data is projected onto the image plane to obtain sparse disparity constraints. An objective function is constructed to characterize the deviation between the stereo matching algorithm results and the sparse disparity constraints. The optimal parameters of the stereo matching algorithm are obtained by minimizing the objective function.
[0019] Based on the stereo matching algorithm using the optimal parameters, stereo matching is performed on the left view and the right view to obtain a dense disparity map;
[0020] Based on the dense disparity map, a dense depth is obtained, and a three-dimensional reconstruction model of the region of interest is generated according to the dense depth.
[0021] In a second aspect of this application, a method for 3D reconstruction using a vision-guided binocular camera and LiDAR fusion is provided. This method is based on the aforementioned vision-guided binocular camera and LiDAR fusion 3D reconstruction device. The method is applied to a main control processing unit, and includes:
[0022] Obtain two-dimensional images of the left and right views of the field of view;
[0023] Detect whether a region of interest exists in the two-dimensional image;
[0024] If the region of interest does not exist, return to the step of obtaining two-dimensional images of the left and right views of the field of view;
[0025] If the region of interest exists, the pixel position of the region of interest in the two-dimensional image is extracted, and the azimuth and pitch angles of the region of interest relative to the reference axis are obtained based on the preset calibration parameters and the pixel position. The reference axis is the emission center axis of the ranging laser beam when the beam deflection scanning unit is in the zero position. The reference axis is jointly calibrated and determined by the binocular vision imaging unit, the beam deflection scanning unit and the laser ranging unit.
[0026] A scanning control signal is generated based on the azimuth and elevation angles, and the scanning control signal is output to the beam deflection scanning unit.
[0027] The distance information of the region of interest is obtained, and a three-dimensional reconstruction model of the region of interest is obtained based on the distance information and angle information, wherein the angle information includes the azimuth angle and the pitch angle.
[0028] Optionally, obtaining the 3D reconstruction model of the region of interest based on the distance and angle information includes:
[0029] Based on the distance and angle information, sparse 3D point cloud data corresponding to the region of interest is calculated.
[0030] The sparse 3D point cloud data is projected onto the image plane to obtain sparse disparity constraints. An objective function is constructed to characterize the deviation between the result error of the stereo matching algorithm and the sparse disparity constraints. The optimal parameters of the stereo matching algorithm are obtained by minimizing the objective function.
[0031] Based on the stereo matching algorithm using the optimal parameters, stereo matching is performed on the left view and the right view to obtain a dense disparity map;
[0032] Based on the dense disparity map, a dense depth is obtained, and a three-dimensional reconstruction model of the region of interest is generated according to the dense depth.
[0033] The visual-guided binocular camera and LiDAR fusion 3D reconstruction device and method provided in this application differs from the traditional binocular stereo matching method that relies on manual experience to set parameters. It improves the robustness of weak texture reconstruction by using parameter adaptive fusion optimization driven by sparse laser constraints. It uses sparse high-confidence distance points provided by LiDAR to form depth consistency constraints, and infers stereo matching parameters adapted to the current scene through parameter adaptive optimization. This realizes a fusion mechanism that guides dense visual reconstruction with sparse and accurate depth constraints, improving the matching stability and reconstruction accuracy of complex areas such as weak texture, strong reflection, and occlusion, and reducing parallax holes and mismatches. Through visual-guided directional scanning, redundancy is reduced and the backend load is lessened: based on visual recognition and location of the region of interest, a scanning control signal is generated to guide the laser ranging unit to perform directional scanning within the region of interest, reducing invalid sampling of non-interested areas such as the sky and ground, reducing data redundancy and transmission pressure, and reducing the computational load of backend depth fusion and point cloud processing, thereby improving the overall scanning and processing efficiency. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 A structural block diagram of a vision-guided binocular camera and lidar fusion 3D reconstruction device provided in an embodiment of this application;
[0036] Figure 2 This is a flowchart illustrating the three-dimensional reconstruction method using a vision-guided binocular camera and LiDAR fusion, as provided in an embodiment of this application. Detailed Implementation
[0037] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0038] The following will describe in detail, with reference to the accompanying drawings, a vision-guided binocular camera and lidar fusion 3D reconstruction device according to an embodiment of this application. Figure 1 This application provides a vision-guided binocular camera and LiDAR fusion 3D reconstruction device as an embodiment of the present application. For example... Figure 1 As shown, the vision-guided binocular camera and lidar fusion 3D reconstruction device of the first embodiment of this application includes:
[0039] The binocular vision imaging unit 101 is used to capture two-dimensional images of the left and right views of the field of view area;
[0040] The main control processing unit 102 is communicatively connected to the binocular vision imaging unit 101, the laser ranging unit 104, and the beam deflection scanning unit 103.
[0041] The beam deflection scanning unit 103 is disposed in the emission optical path of the laser ranging unit. It is used to receive the scanning control signal output by the main control processing unit and change the emission direction of the ranging laser beam according to the scanning control signal, so that the ranging laser beam is pointed to the region of interest and scans the region of interest.
[0042] The laser ranging unit 104 is used to emit the ranging laser beam and receive the echo signal, and output the distance information of the region of interest in the direction of the echo signal.
[0043] The main control processing unit 102 is used for:
[0044] Receive the two-dimensional image output by the binocular vision imaging unit;
[0045] Detect whether the region of interest exists in the two-dimensional image; if the region of interest does not exist, return to the step of receiving the two-dimensional image output by the binocular vision imaging unit.
[0046] If the region of interest exists, the pixel positions of the region of interest in the two-dimensional image are extracted, and based on the preset calibration parameters and the pixel positions, the azimuth and elevation angles of each scanning point in the region of interest relative to the reference axis are obtained. The reference axis is the emission center axis of the ranging laser beam when the beam deflection scanning unit is in the zero position. The reference axis is jointly calibrated and determined by the binocular vision imaging unit, the beam deflection scanning unit and the laser ranging unit.
[0047] A scanning control signal is generated based on the azimuth and elevation angles of each scanning point, and the scanning control signal is output to the beam deflection scanning unit.
[0048] The system receives the distance information output by the laser ranging unit and obtains a three-dimensional reconstruction model of the region of interest based on the distance information and angle information, wherein the angle information includes the azimuth angle and the pitch angle.
[0049] The binocular vision imaging unit 101 can be composed of two synchronously triggered cameras or an integrated binocular camera module, which captures the left and right views of the field of view. Figure 2 3D images can be obtained using conventional techniques, so they will not be elaborated upon here.
[0050] The laser ranging unit 104 can be in the form of pulse time-of-flight (TOF) ranging, or it can be in the form of phase ranging, frequency modulated continuous wave (FMCW) ranging, etc.
[0051] The beam deflection scanning unit 103 can be implemented using deflection devices that alter the reflection, refraction, or diffraction characteristics of the optical path, such as fast-reflecting mirrors / galvanometers, MEMS micromirrors, two-beam wedge prisms, rotating polygon mirrors, and optical phased arrays (OPAs). Specifically, the beam deflection scanning unit can change the reflection characteristics of the optical path by controlling the orientation of reflective optical elements (such as fast-reflecting mirrors and MEMS mirrors); change the refraction characteristics by controlling the deflection amount of refractive elements (such as two-beam wedge prisms); and change the diffraction characteristics by controlling the phase distribution or frequency response of diffraction elements (such as OPAs and liquid crystal metasurfaces). The above description is merely a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
[0052] The main control processing unit 102 can be implemented using an embedded processor, an industrial control computer, a GPU / FPGA acceleration unit, or a combination thereof. Specifically, when detecting whether a region of interest exists in the two-dimensional image, the main control processing unit 102 is used to:
[0053] The input two-dimensional image is preprocessed, including distortion correction, grayscale normalization, and size normalization, and may also include noise suppression;
[0054] The preprocessed 2D image is input into a preset target detection model (such as YOLO series, SSD, RetinaNet or Faster R-CNN) to extract image features and output candidate target boxes, candidate target categories and corresponding confidence scores. A segmentation mask can also be output at the same time.
[0055] Candidate bounding boxes with confidence scores greater than a preset confidence threshold are filtered, and non-maximum suppression is performed on the filtered candidate bounding boxes to remove overlapping and redundant candidate bounding boxes, and low-confidence candidate regions can also be removed.
[0056] The candidate bounding boxes are subjected to non-maximum suppression (NMS). Then, it is determined whether any candidate bounding boxes belonging to a preset target category exist within the NMS-suppressed candidate bounding boxes. If they do, a region of interest (ROI) is determined to exist in the 2D image; otherwise, no ROI exists. When a ROI exists, its location parameters can also be output.
[0057] Specifically, when the main control processing unit 102 extracts the pixel position of the region of interest in the two-dimensional image and obtains the azimuth and pitch angles of the region of interest relative to the optical axis of the device based on preset calibration parameters and the pixel position, it is used to:
[0058] If multiple candidate regions of interest (ROIs) exist in the 2D image, and each ROI corresponds to a candidate bounding box, then the candidate bounding boxes corresponding to each ROI are sorted according to a preset priority rule. The ROI corresponding to the candidate bounding box with the highest priority is selected as the ROI. The preset priority rule is as follows: ROIs are sorted from highest to lowest confidence level, with the highest confidence level corresponding to the highest priority; when the corresponding confidence levels are the same, ROIs are sorted from largest to smallest area, with the largest area corresponding to the highest priority; when the areas are the same, ROIs are sorted from closest to farthest distance between the center of the ROI and the center of the 2D image, with the closest distance corresponding to the highest priority; if only one ROI exists in the 2D image, then that area is the ROI.
[0059] Based on the candidate bounding box or segmentation mask corresponding to the region of interest, extract the pixel set of the region of interest. The pixel range enclosed by the candidate bounding box corresponding to the region of interest is the pixel range of the region of interest, and the set of pixels contained in the pixel range is the pixel set of the region of interest. The candidate target box can be represented as, ,in, This represents the top-left pixel coordinates of the candidate target bounding box. This indicates the width of the candidate target bounding box. This represents the height of the candidate bounding box; similarly, the set of pixels contained in the corresponding region of the segmentation mask is the set of pixels in the region of interest.
[0060] The center pixel coordinates of the region of interest are calculated. (Coordinates in the image coordinate system) , , Given the coordinates of pixels within the region of interest in the image coordinate system, when using a rectangular candidate bounding box, the coordinates of the center pixel of the region of interest can be expressed as: , ;
[0061] Constructing the camera intrinsic parameter matrix , and These represent the edges in the intrinsic parameter matrix of the stereo camera, respectively. Axial direction and along Focal length along the axial direction, and Let x and y represent the x and y coordinates of the camera principal point in the image coordinate system, respectively. The center pixel coordinates are converted from the image coordinate system to a normalized direction vector in the camera coordinate system, calculated using the following formula: ,in, For normalized direction vectors, , ;
[0062] The azimuth angle of the region of interest relative to the reference axis is calculated based on the normalized direction vector. and pitch angle , , Or based on the center point of the 3D point cloud of the region of interest. Calculate the azimuth angle and the pitch angle , , The center point coordinates of the 3D point cloud of the region of interest can be calculated using conventional point cloud center point coordinate calculation methods, and therefore will not be elaborated upon. The azimuth angle represents the deflection angle of the region of interest relative to the reference axis in the horizontal plane, and the pitch angle represents the deflection angle of the region of interest relative to the reference axis in the vertical plane.
[0063] Specifically, when the main control processing unit 102 generates a scanning control signal based on the azimuth and elevation angles, it is used to:
[0064] Get the current azimuth angle and pitch angle And according to the target azimuth angle and target pitch angle Calculate the angle deviation: , ;
[0065] A horizontal control quantity is generated based on the angular deviation. and pitch control The horizontal and pitch control variables are generated using proportional control, proportional-integral control, or proportional-integral-derivative control methods. , , For parameters of a P controller, PI controller, or PID controller, when using a P controller, The value is 0 when using a PI controller. If the value is 0, then the P control, PI control, or PID control here is a closed-loop control.
[0066] According to the driving method of the beam deflection scanning unit, the horizontal control quantity and the pitch control quantity are converted into driving control signals. When the beam deflection scanning unit adopts voltage driving, the driving control signal is the horizontal scanning driving voltage. and pitch scan drive voltage The conversion formula is: , , The driving control signal can also be the driving current, which is a preset driving mapping parameter; when the beam deflection scanning unit uses step driving, the driving control signal is the number of step pulses in the horizontal direction. Pitch step pulse count The conversion formula is: , ,in, and The preset single-step angular resolution, This represents the rounding function; the drive control signal can also be a pulse frequency.
[0067] The drive control signal is combined with preset scan trigger flag, scan speed parameter, and dwell time parameter to generate a scan control signal frame. The scan control signal frames constitute the scan control signal, wherein, This indicates a horizontal control command. This indicates the pitch control command. Indicates the scanning speed. Indicates the length of stay. Indicates the laser emission trigger flag;
[0068] The scanning control signal is sent to the beam deflection scanning unit to control the beam deflection scanning unit to rotate the ranging laser beam to the direction corresponding to the target azimuth and target elevation angles. After reaching the preset angle deviation range, a laser scanning trigger signal is output to scan the region of interest. At the same time, the laser scanning position feedback signal can be received during the scanning process, and the scanning control signal can be closed-loop corrected according to the feedback result.
[0069] Specifically, when the main control processing unit 102 obtains the three-dimensional reconstruction model of the region of interest based on the distance information and angle information, it is used to:
[0070] Based on the distance and angle information, sparse 3D point cloud data corresponding to the region of interest is calculated.
[0071] The sparse 3D point cloud data is projected onto the image plane to obtain sparse disparity constraints. An objective function is constructed to characterize the deviation between the result error of the stereo matching algorithm and the sparse disparity constraints. The optimal parameters of the stereo matching algorithm are obtained by minimizing the objective function.
[0072] Based on the stereo matching algorithm using the optimal parameters, stereo matching is performed on the left view and the right view to obtain a dense disparity map;
[0073] Based on the dense disparity map, a dense depth is obtained, and a three-dimensional reconstruction model of the region of interest is generated according to the dense depth.
[0074] Specifically, when the main control processing unit 102 calculates sparse 3D point cloud data corresponding to the region of interest based on the distance and angle information, it is used for:
[0075] Based on the distance and angle information, the distance values of each scanning point in the region of interest, as well as the corresponding azimuth and elevation angles, are obtained.
[0076] Based on the distance, azimuth, and elevation angles of each scan point, the three-dimensional coordinates of each scan point in the scan coordinate system are calculated. Specifically, the distance value can be used as the radial distance, and the azimuth and elevation angles can be used as spherical coordinate angle parameters. Each scan point can be transformed into a polar coordinate system to obtain its three-dimensional coordinates in the scan coordinate system. For example, for the ... Each scan point, its three-dimensional coordinates It can be calculated using the following formula: , , ,in, Indicates the first Distance value of each scan point, Indicates the first The azimuth angle of each scan point Indicates the first To make the calculation more accurate and efficient, the elevation angle of each scanning point can be synchronized with time and validated first. Abnormal sampling points that exceed the preset ranging range, angle range or signal strength threshold can be removed before calculation.
[0077] Based on the extrinsic parameter matrix obtained from joint calibration, the 3D coordinates of each scanning point are transformed to the camera coordinate system, and aggregated to form sparse 3D point cloud data corresponding to the region of interest in the camera coordinate system. The extrinsic parameter matrix... ,in, The extrinsic parameter matrix is the one from the scan coordinate system to the camera coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector. , ,in, These are rotation matrix elements used to describe the pose relationship between the scan coordinate system and the camera coordinate system; , , These are translation vector elements used to describe the position offset of the origin of the scan coordinate system in the camera coordinate system. The three-dimensional coordinates of each scanning point in the camera coordinate system , , , and The first The x-axis, y-axis, and z-axis coordinates of each scan point in the camera coordinate system. For the first The three-dimensional coordinates of each scan point in the scan coordinate system , , and The first The x-axis, y-axis, and z-axis coordinates of each scanning point in the scanning coordinate system. The joint calibration is a joint calibration of the binocular vision imaging unit, the laser ranging unit, and the beam deflection scanning unit. The extrinsic parameter matrix obtained based on the joint calibration can be obtained as follows: the calibration object is placed in the common measurement area of the binocular vision imaging unit and the laser ranging unit, and multiple sets of calibration data under different poses are collected; the pose information of the calibration object in the camera coordinate system is obtained by the binocular vision imaging unit, and the pose information of the calibration object in the scanning coordinate system is obtained by the laser ranging unit and the beam deflection scanning unit; then, based on multiple sets of corresponding pose relationships, the rigid transformation relationship from the scanning coordinate system to the camera coordinate system is solved, thereby obtaining the extrinsic parameter matrix.
[0078] Wherein, the objective function is ,in, Let be the parameter vector to be optimized. , , A normalization parameter representing the cost of grayscale or color consistency. A normalized parameter representing the cost of local structural consistency. The weighting coefficients represent the cost of local structural consistency. To smooth the constraint weight parameters, This refers to the reference point set, which is the set of effective pixels formed by projecting a sparse three-dimensional point cloud obtained by the laser ranging unit onto a two-dimensional image. Indicates adoption After performing stereo matching, at the pixel point The disparity value obtained at that location, This indicates the value obtained by the laser ranging unit at the pixel point. Reference parallax value at that location Represents pixels The pre-set confidence weights, This represents the preset smoothing constraint coefficient. Indicates the parallax field at the pixel. gradient magnitude at that point A preset positive number is used to prevent the denominator or square root term from degenerating. In the objective function, the first term is used to constrain the deviation between the matching disparity and the laser ranging reference disparity, and the second term is used to suppress local disparity oscillations and improve the stability of the obtained parameters in weak texture regions and boundary regions.
[0079] When the main control processing unit 102 optimizes the stereo matching algorithm to obtain the optimal parameters with the objective function minimization as the optimization goal, it is specifically used for:
[0080] A surrogate model for the objective function is established using Gaussian process regression; this surrogate model is used to characterize the objective function. With the parameter vector to be optimized The mapping relationship between them, the proxy model is represented as , where the symbol To show obedience, Represents a Gaussian process, with the mean function taken as... The covariance function is taken as a rational quadratic kernel function. ,in, and These represent the first and second parameters in the parameter search space, respectively. The group of parameter vectors to be optimized and the first The set of parameter vectors to be optimized. Different value samples, and The preset mean function parameters, For signal variance, The length scale is used to characterize the rate at which the covariance function decays with the distance between the parameter vectors to be optimized. The greater the distance between two parameter vectors to be optimized in the parameter space, the faster the covariance function decays. The kernel function shape parameter, when already obtained Group samples When, for any parameter vector to be optimized The predicted mean and predicted variance are respectively and , Indicates the first Group of parameter vectors to be optimized The corresponding objective function value, The sample covariance matrix, Let covariance be the vector between the parameter vector to be optimized and the obtained samples. To observe the noise variance, , This represents the mean function value vector at the obtained sample points. Based on the obtained sample set, posterior inference is performed on the surrogate model to obtain the predicted mean of the objective function output by the surrogate model at the parameter vector to be optimized. and prediction variance ;
[0081] The optimal parameters are calculated using the acquisition function based on the aforementioned proxy model, including:
[0082] Setting the search space for parameters ;
[0083] In the parameter search space The internal data is obtained using Latin hypercube sampling or random sampling methods. Initial sample of parameters to be optimized ;
[0084] Sample each set of initial parameter vectors to be optimized Input the stereo matching algorithm model and calculate the corresponding disparity results. Based on the aforementioned parallax results Calculate the objective function value The initial parameter sample With the corresponding objective function value Forming initial sample pairs , The initial sample set is composed of initial sample pairs. Based on the initial sample set Train the surrogate model to obtain any parameter vector to be optimized in the current iteration round. Predicted mean function and prediction variance function ;
[0085] Construct the acquisition function ,in, For the exploration coefficient;
[0086] In the parameter search space Inner solution To obtain the new parameters for the current iteration round. ;
[0087] New parameters The results are fed back to the stereo matching algorithm model to recalculate the disparity and the corresponding objective function value. ;
[0088] New sample pairs Add to the initial sample set The proxy model is then updated based on the updated initial sample set, which is the new sample set formed by adding new samples to the initial sample set. The proxy model is then updated based on this new sample set, i.e., the objective function is updated. With the parameter vector to be optimized The mapping relationship between them;
[0089] Determine if the preset iteration stopping condition is met. If it is, output the current historical best parameters as the optimal parameters. If not, return to the constructed acquisition function. The process continues to the next iteration, where the iteration stops when the iteration is continuous. The next iteration simultaneously satisfies , and ,in, The relative change threshold of the preset objective function, For parameter change threshold, For the uncertainty threshold of the proxy model, To prevent extremely small positive numbers with a denominator of zero, or the iteration stopping condition being that the number of iterations reaches the maximum number of iterations. The current historical optimal parameter is the parameter that minimizes the objective function value among all the parameter samples that have been actually evaluated up to the current iteration. Therefore, after obtaining a new parameter, it is necessary to first input the new parameter into the stereo matching algorithm to calculate the corresponding objective function value, and then compare it with the existing historical optimal objective function value. If it is smaller, the new parameter is updated to the current historical optimal parameter; otherwise, the original historical optimal parameter remains unchanged.
[0090] Specifically, when the main control processing unit 102 performs stereo matching on the left view and the right view based on the stereo matching algorithm using the optimal parameters to obtain a dense disparity map, it is used to:
[0091] For each pixel in the left view and the corresponding candidate disparity Calculate the overall matching cost ,in, , , Candidate parallax To provide pixels within the preset parallax search range A discrete disparity hypothesis value is set. The preset pixel-level adaptive weights satisfy... , This represents the grayscale value or color vector at the corresponding pixel in the left view. This represents the grayscale value or color vector at the corresponding pixel in the right view. This represents the sum of the absolute values of the differences in grayscale values or color vectors between corresponding pixels. This represents the local structure descriptor of the neighborhood of the corresponding pixel in the left view. A local structure descriptor representing the neighborhood of the corresponding pixel in the right view. Let represent the Hamming distance between two local structure descriptors. Represents pixels The confidence weighting coefficient for lidar guidance at the location. This indicates that the sparse 3D point cloud, after being projected onto the left view, has pixels... The reference parallax value formed at that location, This represents a default positive number to prevent the square root term from degenerating;
[0092] Constructing a global energy function ,in, , , This represents a dense parallax field, composed of the parallax values corresponding to each pixel. Represents a set of pixels. Represents the set of adjacent pixel pairs. Represents pixels Corresponding candidate parallax, pixels With pixels Adjacent;
[0093] The message quantity between adjacent pixels is updated iteratively using belief propagation, and the update formula is as follows: ,in, Indicates the first In the next iteration, the pixel To adjacent pixels Sent and corresponding to pixels Candidate disparity values The amount of messages at that time Indicates the pixel All candidate disparity values within the preset disparity search range Take the minimum value. Represents pixels Belongs to pixels The set of neighboring pixels and After completing After round belief propagation iterations, calculate each pixel In the corresponding candidate parallax Confidence value below , Represents pixels The set of neighboring pixels, the Round iteration refers to the number of rounds of message updates when solving the global energy function using confidence propagation, i.e., the number of confidence propagation iterations. The global energy function can be solved using any of the following methods: confidence propagation, semi-global matching, graph cut optimization, or dynamic programming.
[0094] For each pixel Select the candidate disparity that minimizes the confidence value as the initial disparity value. The initial disparity values of all pixels constitute the initial disparity map;
[0095] The initial disparity map is subjected to left-right consistency detection to remove occluded points and mismatched points. Specifically, the disparity values of the initial disparity map reflect the horizontal displacement relationship between pixels in the left view and their corresponding pixels in the right view, which can be used for left-right consistency detection. In one possible implementation, the left-right consistency detection can be performed using a left-right disparity cross-validation method. Specifically, based on the same stereo matching parameters, an initial left disparity map corresponding to the left view and an initial right disparity map corresponding to the right view are generated respectively. For any pixel in the left view... Based on its initial disparity value Determine the corresponding pixel position in the right view. ,get Corresponding initial disparity value ,Compare and The difference, when Pixel determination Through consistency checks, when If the corresponding position exceeds the boundary, determine the pixel. An invalid point is determined as follows: if the corresponding position is outside the boundary, or if there is a foreground parallax occlusion relationship at the corresponding position in the right view, it is determined as an occluded point; otherwise, it is determined as a mismatched point.
[0096] The resulting voids are filled using neighborhood-guided interpolation to restore missing disparity. Weighted median filtering, edge-preserving filtering, or bilateral filtering are then used for refinement to remove isolated noise and preserve depth boundaries, resulting in a dense disparity map. It should be noted that before stereo matching of the left and right views, epipolar correction, image registration, and grayscale or color normalization can be performed on the left and right views to ensure that corresponding points lie on their respective scan lines.
[0097] Specifically, when the main control processing unit 102 obtains the dense depth based on the dense disparity map and generates a 3D reconstruction model of the region of interest based on the dense depth, it is used for:
[0098] For each pixel in the dense parallax map Based on the baseline length of the binocular camera and camera internal parameters Calculate pixel points Corresponding density depth value ,in and Representing pixels The horizontal and vertical coordinates in the image coordinate system of the dense parallax map, and These represent the edges in the intrinsic parameter matrix of the stereo camera, respectively. Axial direction and along Focal length along the axial direction, and These represent the x and y coordinates of the camera principal point in the image coordinate system, respectively. Indicates the baseline length of the stereo camera. Represents the pixels in the dense parallax map The corresponding disparity value, This is a preset minimum effective disparity threshold, used to prevent depth divergence caused by excessively small disparity values.
[0099] For the density depth value Outlier removal and hole repair are performed to obtain a continuous dense depth map; outlier removal can be performed using a neighborhood depth difference threshold method, and hole repair can be performed using a neighborhood interpolation method.
[0100] Extract the pixel set corresponding to the region of interest. Extracting pixel sets from the continuous dense depth map The corresponding dense depth value is used to generate a depth submap of the region of interest. ;
[0101] The pixels within the region of interest are back-projected from the image coordinate system to the camera coordinate system to obtain the 3D point cloud of the region of interest. The back-projection formula is as follows: , , Thus, a three-dimensional point set is obtained. , , and Representing pixels x-axis coordinates, y-axis coordinates, and z-axis coordinates in the camera coordinate system;
[0102] Outlier removal and smoothing filtering are performed on the 3D point set to obtain an optimized region of interest (ROI) point cloud, which constitutes the 3D reconstruction model of the ROI. Specifically, outlier removal is first performed on the 3D point set, followed by smoothing filtering on the remaining points. Outlier removal can employ either statistical outlier removal or radius-based outlier removal. Specifically, the average distance between each 3D point and its neighboring points can be calculated. If the average distance is greater than a preset distance threshold, or the number of neighboring points within a preset radius is less than a preset number threshold, the 3D point is identified as an outlier and removed. Smoothing filtering can employ any of moving least squares smoothing, bilateral filtering, or mean filtering to suppress local noise and maintain the continuity of the point cloud surface.
[0103] The visually guided binocular camera and LiDAR fusion 3D reconstruction device provided in this application improves the robustness of weak texture reconstruction through parameter adaptive fusion optimization of sparse laser constraints; it utilizes the sparse high-confidence distance points of the LiDAR to form depth consistency constraints, and infers stereo matching parameters adapted to the current scene through parameter adaptive optimization, realizing a fusion mechanism of sparse and accurate depth constraints guiding dense visual reconstruction, improving the matching stability and reconstruction accuracy of complex areas (weak texture, strong reflection, occlusion), and reducing parallax holes and mismatches; it reduces redundancy and alleviates backend load through visually guided directional scanning: based on visual recognition to locate the region of interest, a scanning control signal for laser beam deflection is generated to guide the laser ranging unit to scan directionally in that area, reducing invalid sampling of non-interested areas such as the sky and ground, reducing data redundancy and transmission pressure, reducing backend computing load, and improving overall scanning and processing efficiency.
[0104] A second aspect of this application provides a three-dimensional reconstruction method that fuses a binocular camera and a lidar sensor. The method is based on the aforementioned three-dimensional reconstruction device that fuses a binocular camera and a lidar sensor, and is applied to the main control processing unit of that device. Figure 2 This document illustrates a flowchart of a 3D reconstruction method that fuses binocular cameras and lidar, as provided in an embodiment of this application. Figure 2As shown, the 3D reconstruction method based on the fusion of binocular camera and lidar in the second embodiment of this application includes:
[0105] Step S201: Obtain two-dimensional images of the left and right views of the field of view;
[0106] Step S202: Detect whether there is a region of interest in the two-dimensional image. If the region of interest does not exist, return to the step of obtaining two-dimensional images of the left and right views of the field of view.
[0107] Step S203: If the region of interest exists, extract the pixel position of the region of interest in the two-dimensional image, and obtain the azimuth and elevation angles of the region of interest relative to the reference axis based on the preset calibration parameters and the pixel position. The reference axis is the emission center axis of the ranging laser beam when the beam deflection scanning unit is in the zero position. The reference axis is jointly calibrated and determined by the binocular vision imaging unit, the beam deflection scanning unit and the laser ranging unit.
[0108] Step S204: Generate a scanning control signal based on the azimuth angle and the elevation angle, and output the scanning control signal to the beam deflection scanning unit;
[0109] Step S205: Obtain distance information of the region of interest, and obtain a three-dimensional reconstruction model of the region of interest based on the distance information and angle information, wherein the angle information includes the azimuth angle and the pitch angle.
[0110] Specifically, in step S201, the main control processing unit receives the two-dimensional image output by the binocular vision imaging unit, thereby obtaining two-dimensional images of the left and right views of the field of view.
[0111] Specifically, in step S202, detecting whether there is a region of interest in the two-dimensional image includes:
[0112] The input two-dimensional image is preprocessed, including distortion correction, grayscale normalization, and size normalization.
[0113] The preprocessed 2D image is input into a preset target detection model (such as YOLO series, SSD, RetinaNet or Faster R-CNN) to extract image features and output candidate bounding boxes, candidate target categories and corresponding confidence scores.
[0114] Candidate bounding boxes with confidence scores greater than a preset confidence threshold are filtered, and non-maximum suppression is performed on the filtered candidate bounding boxes to remove overlapping and redundant boxes.
[0115] Non-maximum suppression is applied to the candidate target boxes. Then, it is determined whether there are candidate target boxes belonging to a preset target category in the candidate target boxes after non-maximum suppression. If there are, it is determined that there is a region of interest in the two-dimensional image. If not, it is determined that there is no region of interest in the two-dimensional image.
[0116] Specifically, in step S203, the pixel positions of the region of interest in the two-dimensional image are extracted, and based on preset calibration parameters and the pixel positions, the azimuth and pitch angles of the region of interest relative to the optical axis of the device are obtained, including:
[0117] If multiple candidate regions of interest (ROIs) exist in the 2D image, and each ROI corresponds to a candidate bounding box, then the candidate bounding boxes corresponding to each ROI are sorted according to a preset priority rule. The ROI corresponding to the candidate bounding box with the highest priority is selected as the ROI. The preset priority rule is as follows: ROIs are sorted from highest to lowest confidence level, with the highest confidence level corresponding to the highest priority; when the corresponding confidence levels are the same, ROIs are sorted from largest to smallest area, with the largest area corresponding to the highest priority; when the areas are the same, ROIs are sorted from closest to farthest distance between the center of the ROI and the center of the 2D image, with the closest distance corresponding to the highest priority; if only one ROI exists in the 2D image, then that area is the ROI.
[0118] Based on the candidate bounding box or segmentation mask corresponding to the region of interest, extract the pixel set of the region of interest. Candidate target boxes can be represented as ,in, This represents the top-left pixel coordinates of the candidate target bounding box. This indicates the width of the candidate target bounding box. This represents the height of the candidate bounding box; similarly, the set of pixels within the range enclosed by the candidate bounding box is the set of pixels in the region of interest, and the set of pixels contained in the corresponding region of the segmentation mask is the set of pixels in the region of interest.
[0119] The center pixel coordinates of the region of interest are calculated. , , , Given the coordinates of pixels within the region of interest in the image coordinate system, when using a rectangular candidate bounding box, the coordinates of the center pixel of the region of interest can be expressed as: , ;
[0120] Constructing the camera intrinsic parameter matrix , and These represent the edges in the intrinsic parameter matrix of the stereo camera, respectively. Axial direction and along Focal length along the axial direction, and Let x and y represent the x and y coordinates of the camera principal point in the image coordinate system, respectively. The center pixel coordinates are converted from the image coordinate system to a normalized direction vector in the camera coordinate system, calculated using the following formula: ,in, For normalized direction vectors, , ;
[0121] The azimuth angle of the region of interest relative to the reference axis is calculated based on the normalized direction vector. and pitch angle , , Or based on the center point of the 3D point cloud of the region of interest. Calculate the azimuth angle and the pitch angle , , Wherein, the azimuth angle represents the deflection angle of the region of interest relative to the reference axis in the horizontal plane, and the pitch angle represents the deflection angle of the region of interest relative to the reference axis in the vertical plane.
[0122] Specifically, step S204 generates a scanning control signal based on the azimuth and elevation angles, including:
[0123] Get the current azimuth angle and pitch angle And according to the target azimuth angle and target pitch angle Calculate the angle deviation: , ;
[0124] A horizontal control quantity is generated based on the angular deviation. and pitch control The horizontal and pitch control variables are generated using proportional control, proportional-integral control, or proportional-integral-derivative control methods. , , For parameters of a P controller, PI controller, or PID controller, when using a P controller, The value is 0 when using a PI controller. If the value is 0, then the PI control, PI control, or PID control here is a closed-loop control.
[0125] According to the driving method of the beam deflection scanning unit, the horizontal control quantity and the pitch control quantity are converted into driving control signals. When the beam deflection scanning unit adopts voltage driving, the driving control signal is the horizontal scanning driving voltage. and pitch scan drive voltage The conversion formula is: , , The driving control signal can also be the driving current, which is a preset driving mapping parameter; when the beam deflection scanning unit uses step driving, the driving control signal is the number of step pulses in the horizontal direction. Pitch step pulse count The conversion formula is: , ,in, and The preset single-step angular resolution, This represents the rounding function; the drive control signal can also be a pulse frequency.
[0126] The drive control signal is combined with preset scan trigger flag, scan speed parameter, and dwell time parameter to generate a scan control signal frame. The scan control signal frames constitute the scan control signal, wherein, This indicates a horizontal control command. This indicates the pitch control command. Indicates the scanning speed. Indicates the length of stay. Indicates the laser emission trigger flag;
[0127] The scanning control signal is sent to the beam deflection scanning unit, which controls the beam deflection scanning unit to rotate the ranging laser beam to the direction corresponding to the target azimuth and target elevation angles. After reaching the preset angle deviation range, a laser scanning trigger signal is output to scan the region of interest. At the same time, during the scanning process, the scanning speed parameters and dwell time parameters can be dynamically adjusted according to the position feedback signal and distance sampling status to improve the sampling accuracy and scanning stability of the region of interest.
[0128] Specifically, in step S205, the main control processing unit receives the distance and angle information output by the laser ranging unit, thereby acquiring the distance and angle information of the region of interest. The angle information includes the azimuth and elevation angles. Specifically, step S205 may include:
[0129] Step S2051: Based on the distance and angle information, sparse 3D point cloud data corresponding to the region of interest is calculated.
[0130] Step S2052: Project the sparse 3D point cloud data onto the image plane to obtain sparse disparity constraints. Construct an objective function to characterize the consistency between the error of the stereo matching algorithm result and the sparse disparity constraints. Optimize the stereo matching algorithm by minimizing the objective function.
[0131] Step S2053: Perform stereo matching on the left view and the right view based on the stereo matching algorithm using the optimal parameters to obtain a dense disparity map;
[0132] Step S2054: Obtain dense depth based on the dense disparity map, and generate a three-dimensional reconstruction model of the region of interest based on the dense depth.
[0133] Specifically, step S2051 calculates sparse 3D point cloud data corresponding to the region of interest based on the distance and angle information, including:
[0134] Based on the distance and angle information, the distance values of each scanning point in the region of interest, as well as the corresponding azimuth and elevation angles, are obtained.
[0135] Based on the distance, azimuth, and elevation angles of each scan point, the three-dimensional coordinates of each scan point in the scan coordinate system are calculated. Specifically, the distance value can be used as the radial distance, and the azimuth and elevation angles can be used as spherical coordinate angle parameters. Each scan point can be transformed into a polar coordinate system to obtain its three-dimensional coordinates in the scan coordinate system. For example, for the ... Each scan point, its three-dimensional coordinates It can be calculated using the following formula: , , ,in, Indicates the first Distance value of each scan point, Indicates the first The azimuth angle of each scan point Indicates the first The elevation angle of each scan point;
[0136] Based on the extrinsic parameter matrix obtained from joint calibration, the 3D coordinates of each scanning point are transformed to the camera coordinate system, and aggregated to form sparse 3D point cloud data corresponding to the region of interest in the camera coordinate system. The extrinsic parameter matrix... ,in, The extrinsic parameter matrix is the one from the scan coordinate system to the camera coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector. , ,in, These are rotation matrix elements used to describe the pose relationship between the scan coordinate system and the camera coordinate system; , , These are translation vector elements used to describe the position offset of the origin of the scan coordinate system in the camera coordinate system. The three-dimensional coordinates of each scanning point in the camera coordinate system , , , and The first The x-axis, y-axis, and z-axis coordinates of each scan point in the camera coordinate system. For the first The three-dimensional coordinates of each scan point in the scan coordinate system , , and The first The x-axis, y-axis, and z-axis coordinates of each scanning point in the scanning coordinate system. The joint calibration is the joint calibration of the binocular vision imaging unit, the laser ranging unit, and the beam deflection scanning unit, or the extrinsic parameter matrix can be obtained by referring to conventional joint calibration.
[0137] Specifically, the objective function can be ,in, Let be the parameter vector to be optimized. , , A normalization parameter representing the cost of grayscale or color consistency. A normalized parameter representing the cost of local structural consistency. The weighting coefficients represent the cost of local structural consistency. To smooth the constraint weight parameters, This refers to the reference point set, which is the set of effective pixels formed by projecting a sparse three-dimensional point cloud obtained by the laser ranging unit onto a two-dimensional image. Indicates adoption After performing stereo matching, at the pixel point The disparity value obtained at that location, This indicates the value obtained by the laser ranging unit at the pixel point. Reference parallax value at that location Represents pixels The pre-set confidence weights, This represents the preset smoothing constraint coefficient. Indicates the parallax field at the pixel. gradient magnitude at that point A default positive number to prevent the denominator or square root term from degenerating.
[0138] Specifically, in step S2051, with minimizing the objective function as the optimization objective, the optimal parameters of the stereo matching algorithm are obtained, which may include:
[0139] Step A: Use Gaussian process regression to build a surrogate model for the objective function. The surrogate model is used to predict the value and uncertainty of the objective function, and it is used to characterize the objective function. With the parameter vector to be optimized The mapping relationship between them, the proxy model is represented as , where the symbol To show obedience, Represents a Gaussian process, with the mean function taken as... The covariance function is taken as a rational quadratic kernel function. ,in, and These represent the first and second parameters in the parameter search space, respectively. The group of parameter vectors to be optimized and the first The set of parameter vectors to be optimized. Different value samples, and The preset mean function parameters, For signal variance, The length scale is used to characterize the rate at which the covariance function decays with the distance between the parameter vectors to be optimized. The greater the distance between two parameter vectors to be optimized in the parameter space, the faster the covariance function decays. The kernel function shape parameter, when already obtained Group samples When, for any parameter vector to be optimized The predicted mean and predicted variance are respectively and , Indicates the first Group of parameter vectors to be optimized The corresponding objective function value, The sample covariance matrix, Let covariance be the vector between the parameter vector to be optimized and the obtained samples. To observe the noise variance, , This represents the mean function value vector at the obtained sample points. Based on the obtained sample set, posterior inference is performed on the surrogate model to obtain the predicted mean of the objective function output by the surrogate model at the parameter vector to be optimized. and prediction variance ;
[0140] Step B involves calculating optimized new parameters based on the agent model using a data acquisition function, including:
[0141] Setting the search space for parameters ;
[0142] In the parameter search space The internal data is obtained using Latin hypercube sampling or random sampling methods. Initial sample of parameters to be optimized ;
[0143] Sample each set of initial parameter vectors to be optimized Input the stereo matching algorithm model and calculate the corresponding disparity results. Based on the aforementioned parallax results Calculate the objective function value The initial parameter sample With the corresponding objective function value Forming initial sample pairs , The initial sample set is composed of initial sample pairs. Based on the initial sample set Train the surrogate model to obtain any parameter vector to be optimized in the current iteration round. Predicted mean function and prediction variance function ;
[0144] Construct the acquisition function ,in, For the exploration coefficient;
[0145] In the parameter search space Inner solution To obtain the new parameters for the current iteration round. ;
[0146] New parameters The results are fed back to the stereo matching algorithm model to recalculate the disparity and the corresponding objective function value. ;
[0147] New sample pairs Add to the initial sample set A new sample set is formed, and then the proxy model is updated based on the new sample set, that is, the objective function is updated. With the parameter vector to be optimized The mapping relationship between them;
[0148] Determine if the preset iteration stopping condition is met. If it is, output the current historical best parameters as the optimal parameters. If not, return to the constructed acquisition function. The process continues to the next iteration, where the iteration stops when the iteration is continuous. The next iteration simultaneously satisfies , and ,in, The relative change threshold of the preset objective function, For parameter change threshold, For the uncertainty threshold of the proxy model, To prevent extremely small positive numbers with a denominator of zero, or the iteration stopping condition being that the number of iterations reaches the maximum number of iterations. .
[0149] Specifically, step S2053 involves performing stereo matching on the left view and the right view based on a stereo matching algorithm using the optimal parameters to obtain a dense disparity map, including:
[0150] For each pixel in the left view and the corresponding candidate disparity Calculate the overall matching cost ,in, , , Candidate parallax To provide pixels within the preset parallax search range A discrete disparity hypothesis value is set. The preset pixel-level adaptive weights satisfy... , This represents the grayscale value or color vector at the corresponding pixel in the left view. This represents the grayscale value or color vector at the corresponding pixel in the right view. This represents the sum of the absolute values of the differences in grayscale values or color vectors between corresponding pixels. This represents the local structure descriptor of the neighborhood of the corresponding pixel in the left view. A local structure descriptor representing the neighborhood of the corresponding pixel in the right view. Let represent the Hamming distance between two local structure descriptors. Represents pixels The confidence weighting coefficient for lidar guidance at the location. This indicates that the sparse 3D point cloud, after being projected onto the left view, has pixels... The reference parallax value formed at that location, This represents a default positive number to prevent the square root term from degenerating;
[0151] Constructing a global energy function ,in, , , This represents a dense parallax field, composed of the parallax values corresponding to each pixel. Represents a set of pixels. Represents the set of adjacent pixel pairs. Represents pixels Corresponding candidate parallax, pixels With pixels Adjacent;
[0152] The message quantity between adjacent pixels is updated iteratively using belief propagation, and the update formula is as follows: ,in, Indicates the first In the next iteration, the pixel To adjacent pixels Sent and corresponding to pixels Candidate disparity values The amount of messages at that time Indicates the pixel All candidate disparity values within the preset disparity search range Take the minimum value. Represents pixels Belongs to pixels The set of neighboring pixels and After completing After round belief propagation iterations, calculate each pixel In the corresponding candidate parallax Confidence value below , Represents pixels The set of neighboring pixels, the Round iteration refers to the number of confidence propagation iterations when solving the global energy function using confidence propagation;
[0153] For each pixel Select the candidate disparity that minimizes the confidence value as the initial disparity value. The initial disparity values of all pixels constitute the initial disparity map;
[0154] Perform left-right consistency detection on the initial disparity map to remove occluded points and mismatched points;
[0155] The resulting voids are filled using neighborhood-guided interpolation, and then refined using weighted median filtering, edge-preserving filtering, or bilateral filtering to obtain a dense disparity map.
[0156] Specifically, step S2054 obtains the dense depth based on the dense disparity map, and generates a 3D reconstruction model of the region of interest based on the dense depth, including:
[0157] For each pixel in the dense parallax map Based on the baseline length of the binocular camera and camera internal parameters Calculate pixel points Corresponding density depth value ,in and Representing pixels The horizontal and vertical coordinates in the image coordinate system of the dense parallax map, and These represent the edges in the intrinsic parameter matrix of the stereo camera, respectively. Axial direction and along Focal length along the axial direction, and These represent the x and y coordinates of the camera principal point in the image coordinate system, respectively. Indicates the baseline length of the stereo camera. Represents the pixels in the dense parallax map The corresponding disparity value, This is a preset minimum effective disparity threshold, used to prevent depth divergence caused by excessively small disparity values.
[0158] For the density depth value Outlier removal and hole repair are performed to obtain a continuous dense depth map; outlier removal can be performed using a neighborhood depth difference threshold method, and hole repair can be performed using a neighborhood interpolation method.
[0159] Extract the pixel set corresponding to the region of interest. Extracting pixel sets from the continuous dense depth map The corresponding dense depth value is used to generate a depth submap of the region of interest. ;
[0160] The pixels within the region of interest are back-projected from the image coordinate system to the camera coordinate system to obtain the 3D point cloud of the region of interest. The back-projection formula is as follows: , , Thus, a three-dimensional point set is obtained. , , and Representing pixels x-axis coordinates, y-axis coordinates, and z-axis coordinates in the camera coordinate system;
[0161] Outlier removal and smoothing filtering are performed on the 3D point set to obtain an optimized region of interest (ROI) point cloud, which constitutes the 3D reconstruction model of the ROI. Specifically, outlier removal is performed on the 3D point set first, followed by smoothing filtering on the remaining points. Outlier removal can employ either statistical outlier removal or radius-based outlier removal; smoothing filtering can employ either moving least squares smoothing, bilateral filtering, or mean filtering. In one possible implementation, a mesh connection is established when the Euclidean distance between adjacent 3D points is less than a preset distance threshold and the depth difference is less than a preset depth threshold; when the Euclidean distance between adjacent 3D points is greater than the preset distance threshold or the depth difference is greater than the preset depth threshold, the connection is canceled to avoid erroneous modeling across abrupt depth changes.
[0162] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related explanations of the method described above can be referred to the corresponding process in the foregoing system embodiments, and therefore will not be repeated here.
[0163] The visual-guided binocular camera and LiDAR fusion 3D reconstruction method provided in this application improves the robustness of weak texture reconstruction through parameter adaptive fusion optimization driven by sparse laser constraints: it constructs depth consistency constraints using sparse high-confidence distance points provided by LiDAR, and infers stereo matching parameters adapted to the current scene through parameter adaptive optimization, realizing a fusion mechanism that guides dense visual reconstruction with sparse and accurate depth constraints. This improves the matching stability and reconstruction accuracy of complex areas such as weak texture, strong reflection, and occlusion, and reduces parallax holes and mismatches. It also reduces data redundancy and backend load through visual-guided directional scanning: based on visual recognition to locate the region of interest, a scanning control signal for laser beam deflection is generated to guide the laser ranging unit to perform directional scanning in the region, reducing invalid sampling of non-interested areas such as the sky and ground, reducing data redundancy and transmission pressure, and reducing the computational load of backend depth fusion and point cloud processing, thereby improving the overall scanning and processing efficiency.
[0164] It should be noted that the vision-guided binocular camera and lidar fusion 3D reconstruction system provided in the first embodiment of this application is only an example of the above-described functional unit division. In practical applications, the above functions can be assigned to different functional units as needed, that is, the units or steps in the embodiments of this application can be further decomposed or combined. For example, the units in the above embodiments can be merged into one unit, or further divided into multiple sub-units to complete all or part of the functions described above. The names of the units and steps involved in the embodiments of this application are only for distinguishing the various units or steps and are not considered as an improper limitation of this application.
[0165] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. The program corresponding to the software units and method steps can be placed in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a unit, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0167] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0168] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0169] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0170] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A three-dimensional reconstruction device that fuses a vision-guided binocular camera with lidar, characterized in that, include: A binocular vision imaging unit is used to capture two-dimensional images of the left and right views of the field of view. A beam deflection scanning unit is disposed in the emission optical path of the laser ranging unit. It is used to receive the scanning control signal output by the main control processing unit and change the emission direction of the ranging laser beam according to the scanning control signal, so that the ranging laser beam is pointed to the region of interest and scans the region of interest. A laser ranging unit is used to emit the ranging laser beam and receive the echo signal, and output the distance information of each scanning point in the region of interest in the direction of the echo signal. The main control processing unit is used for: Receive the two-dimensional image output by the binocular vision imaging unit; Detect whether the region of interest exists in the two-dimensional image; if the region of interest does not exist, return to the step of receiving the two-dimensional image output by the binocular vision imaging unit. If the region of interest exists, the pixel positions of the region of interest in the two-dimensional image are extracted, and based on the preset calibration parameters and the pixel positions, the azimuth and elevation angles of each scanning point in the region of interest relative to the reference axis are obtained. The reference axis is the emission center axis of the ranging laser beam when the beam deflection scanning unit is in the zero position. The reference axis is jointly calibrated and determined by the binocular vision imaging unit, the beam deflection scanning unit and the laser ranging unit. A scanning control signal is generated based on the azimuth and elevation angles of each scanning point, and the scanning control signal is output to the beam deflection scanning unit. The system receives the distance information output by the laser ranging unit and obtains a three-dimensional reconstruction model of the region of interest based on the distance information and angle information, wherein the angle information includes the azimuth angle and the pitch angle.
2. The three-dimensional reconstruction device based on the fusion of a vision-guided binocular camera and lidar as described in claim 1, characterized in that, When the main control processing unit obtains the 3D reconstruction model of the region of interest based on the distance and angle information, it is used for: Based on the distance and angle information, sparse 3D point cloud data corresponding to the region of interest is calculated. The sparse 3D point cloud data is projected onto the image plane to obtain sparse disparity constraints. An objective function is constructed to characterize the deviation between the stereo matching algorithm results and the sparse disparity constraints. The optimal parameters of the stereo matching algorithm are obtained by minimizing the objective function. Based on the stereo matching algorithm using the optimal parameters, stereo matching is performed on the left view and the right view to obtain a dense disparity map; Based on the dense disparity map, a dense depth is obtained, and a three-dimensional reconstruction model of the region of interest is generated according to the dense depth.
3. A three-dimensional reconstruction method that fuses a vision-guided binocular camera with lidar, characterized in that, The 3D reconstruction method is based on the vision-guided binocular camera and lidar fusion 3D reconstruction device described in claims 1-2. The vision-guided binocular camera and lidar fusion 3D reconstruction device includes a binocular vision imaging unit, a laser ranging unit, a beam deflection scanning unit, and a main control processing unit. The vision-guided binocular camera and lidar fusion 3D reconstruction method is applied to the main control processing unit, including: Obtain two-dimensional images of the left and right views of the field of view; Detect whether a region of interest exists in the two-dimensional image; If the region of interest does not exist, return to the step of obtaining two-dimensional images of the left and right views of the field of view; If the region of interest exists, the pixel position of the region of interest in the two-dimensional image is extracted, and the azimuth and pitch angles of the region of interest relative to the reference axis are obtained based on the preset calibration parameters and the pixel position. The reference axis is the emission center axis of the ranging laser beam when the beam deflection scanning unit is in the zero position. The reference axis is jointly calibrated and determined by the binocular vision imaging unit, the beam deflection scanning unit and the laser ranging unit. A scanning control signal is generated based on the azimuth angle and the elevation angle, and the scanning control signal is output to the beam deflection scanning unit; The distance information of the region of interest is obtained, and a three-dimensional reconstruction model of the region of interest is obtained based on the distance information and angle information, wherein the angle information includes the azimuth angle and the pitch angle.
4. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to claim 3, characterized in that, When obtaining the 3D reconstruction model of the region of interest based on the distance and angle information, the process includes: Based on the distance and angle information, sparse 3D point cloud data corresponding to the region of interest is calculated. The sparse 3D point cloud data is projected onto the image plane to obtain sparse disparity constraints. An objective function is constructed to characterize the deviation between the result error of the stereo matching algorithm and the sparse disparity constraints. The optimal parameters of the stereo matching algorithm are obtained by minimizing the objective function. Based on the stereo matching algorithm using the optimal parameters, stereo matching is performed on the left view and the right view to obtain a dense disparity map; Based on the dense disparity map, a dense depth is obtained, and a three-dimensional reconstruction model of the region of interest is generated according to the dense depth.
5. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to claim 4, characterized in that, The step of calculating sparse 3D point cloud data corresponding to the region of interest based on the distance and angle information includes: Based on the distance and angle information, the distance values of each scanning point in the region of interest, as well as the corresponding azimuth and elevation angles, are obtained. Based on the distance, azimuth, and elevation angles of each scanning point, the three-dimensional coordinates of each scanning point based on the scanning coordinate system are calculated. Based on the extrinsic parameter matrix obtained from the joint calibration, the 3D coordinates of each scanning point are transformed to the camera coordinate system. And aggregated to form sparse 3D point cloud data corresponding to the region of interest in the camera coordinate system, wherein the extrinsic parameter matrix ,in, The extrinsic parameter matrix is the one from the scan coordinate system to the camera coordinate system. It is a 3×3 rotation matrix. It is a 3×1 translation vector. , ,in, These are rotation matrix elements used to describe the pose relationship between the scan coordinate system and the camera coordinate system; , , These are translation vector elements used to describe the position offset of the origin of the scan coordinate system in the camera coordinate system. The three-dimensional coordinates of each scanning point in the camera coordinate system , , , and The first The x-axis, y-axis, and z-axis coordinates of each scan point in the camera coordinate system. For the first The three-dimensional coordinates of each scan point in the scan coordinate system , , and The first The x-axis, y-axis, and z-axis coordinates of each scan point in the scan coordinate system.
6. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to claim 5, characterized in that, The objective function is: ,in, Let be the parameter vector to be optimized. , , A normalization parameter representing the cost of grayscale or color consistency. A normalized parameter representing the cost of local structural consistency. The weighting coefficients represent the cost of local structural consistency. To smooth the constraint weight parameters, This refers to the reference point set, which is the set of effective pixels formed by projecting a sparse three-dimensional point cloud obtained by the laser ranging unit onto a two-dimensional image. Indicates adoption After performing stereo matching, at the pixel point The disparity value obtained at that location, This indicates the value obtained by the laser ranging unit at the pixel point. Reference parallax value at that location Represents pixels The pre-set confidence weights, This represents the preset smoothing constraint coefficient. Indicates the parallax field at the pixel. gradient magnitude at that point Preset positive numbers to prevent the denominator or radical term from degenerating; The optimization of the stereo matching algorithm, with minimizing the objective function as the optimization objective, yields the optimal parameters, including: A surrogate model for the objective function is established using Gaussian process regression; this surrogate model is used to characterize the objective function. With the parameter vector to be optimized The mapping relationship between them, the proxy model is represented as , where the symbol To show obedience, Represents a Gaussian process, with the mean function taken as... The covariance function is taken as a rational quadratic kernel function. ,in, and These represent the first and second parameters in the parameter search space, respectively. The set of parameter vectors to be optimized and the first The set of parameter vectors to be optimized. Different value samples, and The preset mean function parameters, For signal variance, The length scale is used to characterize the decay rate of the covariance function as the distance between the parameter vectors to be optimized changes. The kernel function shape parameter, when already obtained Group samples When, for any parameter vector to be optimized The predicted mean and predicted variance are respectively and , Indicates the first Group of parameter vectors to be optimized The corresponding objective function value, The sample covariance matrix, Let covariance be the vector between the parameter vector to be optimized and the obtained samples. To observe the noise variance, , This represents the mean function value vector at the obtained sample points. Based on the obtained sample set, posterior inference is performed on the surrogate model to obtain the predicted mean of the objective function output by the surrogate model at the parameter vector to be optimized. and prediction variance ; The optimal parameters are calculated using the acquisition function based on the aforementioned proxy model, including: Setting the search space for parameters ; In the parameter search space The internal data is obtained using Latin hypercube sampling or random sampling methods. Initial sample of parameters to be optimized ; Sample each set of initial parameter vectors to be optimized Input the stereo matching algorithm model and calculate the corresponding disparity results. Based on the aforementioned parallax results Calculate the objective function value The initial parameter sample With the corresponding objective function value Forming initial sample pairs , The initial sample set is composed of initial sample pairs. ; Based on the initial sample set Train the surrogate model to obtain any parameter vector to be optimized in the current iteration round. Predicted mean function and prediction variance function ; Construct acquisition function ,in, To explore coefficients; In the parameter search space Inner solution To obtain the new parameters for the current iteration round. ; New parameters The results are fed back to the stereo matching algorithm model to recalculate the disparity and the corresponding objective function value. ; New sample pairs Add to the initial sample set The proxy model is then updated based on the updated initial sample set. Determine if the preset iteration stopping condition is met. If it is, output the current historical best parameters as the optimal parameters. If not, return to the constructed acquisition function. The process continues to the next iteration, where the iteration stops when the iteration is continuous. The next iteration simultaneously satisfies , and ,in, The relative change threshold of the preset objective function, For parameter change threshold, For the uncertainty threshold of the proxy model, To prevent extremely small positive numbers with a denominator of zero, or the iteration stopping condition being that the number of iterations reaches the maximum number of iterations. .
7. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to claim 6, characterized in that, The stereo matching algorithm based on the optimal parameters performs stereo matching on the left view and the right view to obtain a dense disparity map, including: For each pixel in the left view and the corresponding candidate disparity Calculate the overall matching cost ,in, , , Candidate parallax To provide pixels within the preset parallax search range A discrete disparity hypothesis value is set. The preset pixel-level adaptive weights satisfy... , This represents the grayscale value or color vector at the corresponding pixel in the left view. This represents the grayscale value or color vector at the corresponding pixel in the right view. This represents the sum of the absolute values of the differences in grayscale values or color vectors between corresponding pixels. This represents the local structure descriptor of the neighborhood of the corresponding pixel in the left view. A local structure descriptor representing the neighborhood of the corresponding pixel in the right view. Let represent the Hamming distance between two local structure descriptors. Represents pixels The confidence weighting coefficient for lidar guidance at the location. This indicates that the sparse 3D point cloud, after being projected onto the left view, has pixels... The reference parallax value formed at that location, This represents a default positive number to prevent the square root term from degenerating; Constructing a global energy function ,in, , , Represents a dense parallax field. Represents a set of pixels. Represents the set of adjacent pixel pairs. Represents pixels Corresponding candidate parallax, pixels With pixels Adjacent; The message quantity between adjacent pixels is updated iteratively using belief propagation, and the update formula is as follows: ,in, Indicates the first In the next iteration, the pixel To adjacent pixels Sent and corresponding to pixels Candidate disparity values The amount of messages at that time Indicates the pixel All candidate disparity values within the preset disparity search range Take the minimum value. Represents pixels Belongs to pixels The set of neighboring pixels and After completing After the next belief propagation iteration, calculate each pixel In the corresponding candidate parallax Confidence value below , Represents pixels The set of neighboring pixels; For each pixel Select the candidate disparity that minimizes the confidence value as the initial disparity value. The initial disparity values of all pixels constitute the initial disparity map; Perform left-right consistency detection on the initial disparity map to remove occluded points and mismatched points; The resulting voids are filled using neighborhood-guided interpolation, and then refined using weighted median filtering or edge-preserving filtering to obtain a dense disparity map.
8. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to claim 7, characterized in that, The process of obtaining dense depth based on the dense disparity map and generating a 3D reconstruction model of the region of interest based on the dense depth includes: For each pixel in the dense parallax map Based on the baseline length of the binocular camera and camera internal parameters Calculate pixel points Corresponding density depth value ,in and Representing pixels The horizontal and vertical coordinates in the image coordinate system of the dense parallax map, and These represent the edges in the intrinsic parameter matrix of the stereo camera, respectively. Axial direction and along Focal length along the axial direction, and These represent the x and y coordinates of the camera principal point in the image coordinate system, respectively. Indicates the baseline length of the stereo camera. Represents the pixels in the dense parallax map The corresponding disparity value, This is a preset minimum effective disparity threshold, used to prevent depth divergence caused by excessively small disparity values. For the density depth value Outlier removal and hole repair are performed to obtain a continuous dense depth map; outlier removal uses a neighborhood depth difference threshold discrimination method, and hole repair uses a neighborhood interpolation method. Extract the pixel set corresponding to the region of interest. Extracting pixel sets from the continuous dense depth map The corresponding dense depth value is used to generate a depth submap of the region of interest. ; The pixels within the region of interest are back-projected from the image coordinate system to the camera coordinate system to obtain the 3D point cloud of the region of interest. The back-projection formula is as follows: , , Thus, a three-dimensional point set is obtained. , , and Representing pixels x-axis coordinates, y-axis coordinates, and z-axis coordinates in the camera coordinate system; Outlier removal and smoothing filtering are performed on the three-dimensional point set to obtain an optimized region of interest point cloud, which constitutes the three-dimensional reconstruction model of the region of interest.
9. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to any one of claims 1-8, characterized in that, The detection of whether a region of interest exists in the two-dimensional image includes: The input two-dimensional image is preprocessed, including distortion correction, grayscale normalization, and size normalization. The preprocessed 2D image is input into a preset target detection model to extract image features and output candidate target boxes, candidate target categories and corresponding confidence scores. Candidate bounding boxes with confidence scores greater than a preset confidence threshold are filtered, and non-maximum suppression is performed on the filtered candidate bounding boxes to remove overlapping and redundant boxes. Determine whether there is a candidate target box belonging to a preset target category in the candidate target box after non-maximum suppression processing. If it exists, determine that there is a region of interest in the two-dimensional image. If it does not exist, determine that there is no region of interest in the two-dimensional image.
10. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to claim 9, characterized in that, The step of extracting the pixel position of the region of interest in the two-dimensional image, and obtaining the azimuth and pitch angles of the region of interest relative to the optical axis of the device based on preset calibration parameters and the pixel position, includes: If multiple candidate regions of interest exist in the two-dimensional image, the candidate bounding boxes corresponding to each candidate region of interest are sorted according to a preset priority rule. The candidate region of interest corresponding to the candidate bounding box with the highest priority is selected as the region of interest. The preset priority rule is as follows: the candidate bounding boxes are sorted from high to low confidence, with the highest confidence level corresponding to the highest priority; when the corresponding confidence levels are the same, the candidate bounding boxes are sorted from large to small area, with the largest area corresponding to the highest priority; when the areas are the same, the candidate bounding boxes are sorted from near to far distance between the center of the candidate bounding box and the center of the two-dimensional image, with the closest distance corresponding to the highest priority. The candidate region of interest is the region corresponding to the obtained candidate bounding box. Based on the candidate bounding box or segmentation mask corresponding to the region of interest, extract the pixel set of the region of interest. ; The center pixel coordinates of the region of interest are calculated. ; Constructing the camera intrinsic parameter matrix , and These represent the edges in the intrinsic parameter matrix of the stereo camera, respectively. Axial direction and along Focal length along the axial direction, and These represent the x and y coordinates of the camera principal point in the image coordinate system, respectively. The center pixel coordinates are converted from the image coordinate system to a normalized direction vector in the camera coordinate system, calculated using the following formula: ,in, For normalized direction vectors, , ; The azimuth angle of the region of interest relative to the reference axis is calculated based on the normalized direction vector. and pitch angle , , Or based on the center point of the 3D point cloud of the region of interest. Calculate the azimuth angle and the pitch angle , , Wherein, the azimuth angle represents the deflection angle of the region of interest relative to the reference axis in the horizontal plane, and the pitch angle represents the deflection angle of the region of interest relative to the reference axis in the vertical plane.
11. The three-dimensional reconstruction method based on the fusion of a vision-guided binocular camera and lidar according to claim 10, characterized in that, The step of generating a scanning control signal based on the azimuth and elevation angles includes: Get the current azimuth angle and pitch angle And according to the target azimuth angle and target pitch angle Calculate the angle deviation: , ; A horizontal control quantity is generated based on the angular deviation. and pitch control The horizontal control and pitch control quantities are generated using proportional control, proportional-integral control, or proportional-integral-derivative control methods. , , For controller parameters; According to the driving method of the beam deflection scanning unit, the horizontal control quantity and the pitch control quantity are converted into driving control signals. When the beam deflection scanning unit adopts voltage driving, the driving control signal is the horizontal scanning driving voltage. and pitch scan drive voltage The conversion formula is: , , As preset driving mapping parameters, when the beam deflection scanning unit adopts step driving, the driving control signal is the number of step pulses in the horizontal direction. Pitch step pulse count The conversion formula is: , ,in, and The preset single-step angular resolution, This represents the rounding function; The drive control signal is combined with preset scan trigger flag, scan speed parameter, and dwell time parameter to generate a scan control signal frame. The scan control signal frames constitute the scan control signal, wherein, This indicates a horizontal control command. This indicates the pitch control command. Indicates the scanning speed. Indicates the length of stay. Indicates the laser emission trigger flag; The scanning control signal is sent to the beam deflection scanning unit to control the beam deflection scanning unit to rotate the ranging laser beam to the direction corresponding to the target azimuth and target elevation angles, and outputs a laser scanning trigger signal after reaching a preset angle deviation range to scan the region of interest.