Data set construction method and device of monocular speckle depth model, equipment and storage medium

By using a binocular infrared speckle camera and a depth baseline model to generate pseudo-labels, and combining the calibration results to convert them into monocular disparity maps, the problem of insufficient training data for monocular speckle depth models is solved, and the efficiency of dataset construction and the scale of real-world scene data are improved.

CN121544985APending Publication Date: 2026-02-17GOERTEK OPTICAL TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511745404.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

The lack of a large-scale training dataset for monocular infrared speckle depth technology results in a severe shortage of training data for monocular speckle depth models.

Method used

By acquiring binocular speckle image pairs from a binocular infrared speckle camera, binocular disparity map pseudo-labels are generated using a binocular depth base model. Combined with binocular stereo vision calibration results and monocular baseline calibration results, the binocular disparity map pseudo-labels are converted into monocular disparity map pseudo-labels, thus constructing a dataset for training a monocular speckle depth model.

Benefits of technology

By effectively utilizing data from binocular infrared speckle cameras, the efficiency of constructing training datasets for monocular speckle depth models has been improved, and the scale of real-world scene data available for model training has been increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121544985A_ABST
    Figure CN121544985A_ABST
Patent Text Reader

Abstract

The invention discloses a data set construction method and device of a monocular speckle depth model, equipment and a storage medium, and relates to the technical field of computer vision, and the disclosed data set construction method of the monocular speckle depth model comprises the steps: obtaining each binocular speckle image pair collected by a binocular infrared speckle camera; processing each binocular speckle image pair through a binocular depth basic model to obtain a binocular disparity map pseudo label of each binocular speckle image pair; based on the binocular stereoscopic vision calibration result, the monocular baseline calibration result and each reference speckle image pair acquired by the binocular infrared speckle camera, converting each binocular parallax image pseudo label into each monocular parallax image pseudo label; and forming a data set for training a monocular speckle depth model according to each binocular speckle image pair, each reference speckle image pair and each monocular disparity map pseudo label. According to the invention, the construction efficiency of the training data set of the monocular speckle depth model based on the monocular infrared speckle depth technology can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method, apparatus, device and storage medium for constructing datasets for monocular speckle depth models. Background Technology

[0002] With the development of deep learning technology in the field of computer vision, data-driven depth perception methods have become the core of applications such as 3D reconstruction and autonomous driving.

[0003] Currently, depth estimation models are typically trained by collecting a large amount of image and ground-value depth data. After fully training the model using a dataset containing real depth information, inputting a new image will output the corresponding depth map. However, in the specific depth sensing technique of monocular infrared speckle, due to the difficulty in obtaining ground-value depth data, the dataset size is often small, and a significant portion of the data is synthesized through virtual modeling. This results in a lack of large-scale training datasets for monocular speckle depth models based on monocular infrared speckle depth technology, leading to a severe shortage of real-world scene data available for model training.

[0004] In summary, how to improve the efficiency of constructing training datasets for monocular speckle depth models based on monocular infrared speckle depth technology has become a pressing technical problem in this field. Summary of the Invention

[0005] The main objective of this application is to provide a method, apparatus, device, and storage medium for constructing a dataset for a monocular speckle depth model, aiming to improve the efficiency of constructing a training dataset for a monocular speckle depth model based on monocular infrared speckle depth technology.

[0006] To achieve the above objectives, this application proposes a method for constructing a dataset for a monocular speckle depth model. The method includes: Acquire each pair of binocular speckle images captured by a binocular infrared speckle camera; Each pair of binocular speckle images is processed by a binocular depth model to obtain a pseudo-label of the binocular disparity map for each pair of binocular speckle images. Based on the binocular stereo vision calibration results, the monocular baseline calibration results, and each reference speckle image pair acquired by the binocular infrared speckle camera, each binocular disparity map pseudo-label is converted into a monocular disparity map pseudo-label, wherein the monocular disparity map pseudo-label corresponds to the virtual monocular infrared speckle device obtained by decomposition from the binocular infrared speckle camera. The dataset for training the monocular speckle depth model is constructed based on each of the binocular speckle image pairs, each of the reference speckle image pairs, and each of the monocular disparity map pseudo-labels.

[0007] In one embodiment, before the step of converting each binocular disparity map pseudo-label into each monocular disparity map pseudo-label based on the binocular stereo vision calibration results, the monocular baseline calibration results, and each reference speckle image pair acquired by the binocular infrared speckle camera, the method further includes: Acquire pairs of calibration plate images from multiple perspectives captured by the binocular infrared speckle camera, and acquire pairs of reference speckle images from different distances captured by the binocular infrared speckle camera. The binocular stereo vision calibration results are calculated based on the images of each calibration board. The monocular baseline calibration result is calculated based on each of the reference speckle image pairs and the binocular stereo vision calibration result.

[0008] In one embodiment, the step of calculating the binocular stereo vision calibration result based on each of the calibration board images includes: The first calibration parameters of the binocular infrared speckle camera are calculated based on the image pairs of each calibration plate. The first calibration parameters include the initial intrinsic parameter matrix and distortion coefficients of the binocular infrared speckle camera. The first calibration parameters are imported into the stereo calibration function for calculation to obtain the second calibration parameters. The second calibration parameters include the rotation matrix, translation matrix, basic matrix, essential matrix, and stereo baseline length of the binocular infrared speckle camera. The first calibration parameters and the second calibration parameters are imported into the stereo correction function for calculation to obtain the third calibration parameters, which include the correction rotation matrix, correction intrinsic parameter matrix and reprojection matrix of the binocular infrared speckle camera; The first calibration parameter, the second calibration parameter, and the third calibration parameter are imported into the distortion correction function for calculation to obtain the distortion correction mapping table of the binocular infrared speckle camera.

[0009] In one embodiment, the step of calculating the monocular baseline calibration result based on each of the reference speckle image pairs and the binocular stereo vision calibration result includes: Based on the distortion correction mapping table, stereo correction is performed on each of the reference speckle image pairs to obtain each corrected reference image pair; Multiple sampling points are selected on each of the calibration reference image pairs, and the disparity of each calibration reference image pair is calculated based on each sampling point; Calculate the three-dimensional coordinates of each sampling point based on the disparity and the reprojection matrix; Based on the three-dimensional coordinates of the same sampling point at multiple distances, multiple three-dimensional straight lines are fitted to obtain the line. The coordinates of the speckle projector are calculated based on the linear equations of each of the three-dimensional lines. The monocular baseline length of the virtual monocular infrared speckle device is calculated based on the speckle projector coordinates and the translation matrix.

[0010] In one embodiment, the step of calculating the speckle projector coordinates based on the linear equations of each of the three-dimensional lines includes: Substitute the coordinates of the speckle projector to be solved into the linear equations of each of the three-dimensional lines to construct an overdetermined system of equations. Solve the least squares solution of the overdetermined system of equations to obtain the coordinates of the speckle projector.

[0011] In one embodiment, the step of converting the binocular disparity map pseudo-labels into monocular disparity map pseudo-labels based on binocular stereo vision calibration results and monocular baseline calibration results includes: Based on the calibration intrinsic parameter matrix and binocular baseline length in the binocular stereo vision calibration results, each binocular disparity map pseudo-label is converted into a binocular depth map. Based on the calibration intrinsic parameter matrix, the monocular baseline length in the monocular baseline calibration result, and the distance between each reference speckle image pair and the reference plane during acquisition, each binocular depth map is converted into a pseudo-label for each monocular disparity map.

[0012] In one embodiment, the step of constructing a dataset for training a monocular speckle depth model based on each of the binocular speckle image pairs, each of the reference speckle image pairs, and each of the monocular disparity map pseudo-labels includes: A dataset for training a monocular speckle depth model is constructed based on each of the binocular speckle image pairs, each of the reference speckle image pairs, and each of the monocular disparity map pseudo-labels. The dataset includes each set of data sample pairs, and each set of data sample pairs includes a speckle image, a reference speckle image corresponding to the speckle image, and a monocular disparity map pseudo-label.

[0013] Furthermore, to achieve the above objectives, this application also proposes a dataset construction device for a monocular speckle depth model, which includes: The image pair acquisition module is used to acquire each pair of binocular speckle images captured by the binocular infrared speckle camera; A binocular pseudo-label generation module is used to process each of the binocular speckle image pairs through a binocular depth basic model to obtain the binocular disparity map pseudo-labels for each of the binocular speckle image pairs. The monocular pseudo-label generation module is used to convert each of the binocular disparity map pseudo-labels into monocular disparity map pseudo-labels based on the binocular stereo vision calibration results, the monocular baseline calibration results, and each reference speckle image pair acquired by the binocular infrared speckle camera. The monocular disparity map pseudo-labels correspond to the virtual monocular infrared speckle device obtained by decomposing the binocular infrared speckle camera. The dataset construction module is used to construct a dataset for training a monocular speckle depth model based on each of the binocular speckle image pairs, each of the reference speckle image pairs, and each of the monocular disparity map pseudo-labels.

[0014] In addition, to achieve the above objectives, this application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the dataset construction method for the monocular speckle depth model as described above.

[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the dataset construction method for the monocular speckle depth model as described above.

[0016] This application proposes a dataset construction method for a monocular speckle depth model. The method involves acquiring pairs of binocular speckle images captured by a binocular infrared speckle camera; processing these pairs using a binocular depth base model to obtain pseudo-labels for each pair of binocular disparity maps; converting these pseudo-labels into pseudo-labels for monocular disparity maps based on binocular stereo vision calibration results, monocular baseline calibration results, and reference speckle image pairs captured by the binocular infrared speckle camera; and constructing a dataset for training the monocular speckle depth model based on these pairs of binocular speckle images, the reference pairs of speckle images, and the pseudo-labels for monocular disparity maps.

[0017] In summary, this application acquires binocular speckle image pairs from a binocular infrared speckle camera, generates binocular disparity map pseudo-labels using a binocular depth baseline model, and converts these pseudo-labels into monocular disparity map pseudo-labels by combining relevant calibration results and benchmark speckle image pairs. Finally, it constructs a dataset for training a monocular speckle depth model. This approach effectively utilizes data from a binocular infrared speckle camera, generating data suitable for training a monocular speckle depth model through reasonable transformation and processing. This improves the efficiency of dataset construction and increases the scale of real-world scene data available for model training. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating the data set construction method for the monocular speckle depth model of this application (Example 1). Figure 2 A schematic diagram of the binocular infrared speckle camera structure provided in Embodiment 1 of the method for constructing the dataset for the monocular speckle depth model of this application; Figure 3 A schematic diagram of the sampling point location distribution provided in Embodiment 2 of the method for constructing the dataset for the monocular speckle depth model of this application; Figure 4 A schematic diagram of the position of a three-dimensional straight line provided in Embodiment 2 of the method for constructing a dataset for the monocular speckle depth model of this application; Figure 5 A simplified flowchart illustrating the second embodiment of the method for constructing the dataset for the monocular speckle depth model of this application; Figure 6 This is a schematic diagram of the module structure of the dataset construction device for the monocular speckle depth model according to an embodiment of this application; Figure 7 This is a schematic diagram of the hardware operating environment involved in the dataset construction method of the monocular speckle depth model in this application embodiment.

[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0023] With the development of deep learning technology in the field of computer vision, data-driven depth perception methods have become the core of applications such as 3D reconstruction and autonomous driving.

[0024] Currently, depth estimation models are typically trained by collecting a large amount of image and ground-value depth data. After fully training the model using a dataset containing real depth information, inputting a new image will output the corresponding depth map. However, in the specific depth sensing technique of monocular infrared speckle, due to the difficulty in obtaining ground-value depth data, the dataset size is often small, and a significant portion of the data is synthesized through virtual modeling. This results in a lack of large-scale training datasets for monocular speckle depth models based on monocular infrared speckle depth technology, leading to a severe shortage of real-world scene data available for model training.

[0025] In summary, how to improve the efficiency of constructing training datasets for monocular speckle depth models based on monocular infrared speckle depth technology has become a pressing technical problem in this field.

[0026] This application provides a solution for acquiring binocular speckle image pairs captured by a binocular infrared speckle camera; processing each binocular speckle image pair using a binocular depth baseline model to obtain a pseudo-label for each binocular disparity map; based on binocular stereo vision calibration results, monocular baseline calibration results, and each reference speckle image pair captured by the binocular infrared speckle camera, converting each pseudo-label for the binocular disparity map into a pseudo-label for the monocular disparity map, wherein the pseudo-label for the monocular disparity map corresponds to a virtual monocular infrared speckle device obtained by decomposing the binocular infrared speckle camera; and constructing a dataset for training a monocular speckle depth model based on each binocular speckle image pair, each reference speckle image pair, and each pseudo-label for the monocular disparity map.

[0027] In summary, this embodiment acquires binocular speckle image pairs from a binocular infrared speckle camera, generates binocular disparity map pseudo-labels using a binocular depth baseline model, and converts these pseudo-labels into monocular disparity map pseudo-labels by combining relevant calibration results and benchmark speckle image pairs. Finally, it constructs a dataset for training a monocular speckle depth model. This approach effectively utilizes data from a binocular infrared speckle camera, generating data suitable for training a monocular speckle depth model through reasonable transformation and processing. This improves the efficiency of dataset construction and increases the scale of real-world scene data available for model training.

[0028] It should be noted that the executing entity in this embodiment can be an electronic device with data processing, network communication, and program execution functions, such as a computer, host computer, controller, etc., or an electronic device capable of performing the above functions. The following description uses an electronic device as an example to illustrate this embodiment and the subsequent embodiments.

[0029] Based on this, embodiments of this application provide a method for constructing a dataset for a monocular speckle depth model, referring to... Figure 1 , Figure 1This is a flowchart illustrating the first embodiment of the dataset construction method for the monocular speckle depth model of this application.

[0030] In this embodiment, the method for constructing the dataset for the monocular speckle depth model includes steps S10 to S40: Step S10: Acquire each pair of binocular speckle images captured by the binocular infrared speckle camera; It should be noted that a binocular infrared speckle camera is a type of binocular depth camera that uses infrared speckle as its active light source. Its main hardware structure typically consists of two infrared cameras and a speckle projector, such as... Figure 2 As shown, the optical axis of its speckle projector is parallel to the optical axes of the left and right infrared cameras. Theoretically, the speckle projector and the left and right infrared cameras can each form a monocular infrared speckle device. A binocular speckle image pair refers to two infrared speckle images acquired at the same time by the left and right infrared cameras of a binocular infrared speckle camera. The images contain depth information of the scene because the speckle pattern will show different displacements as the distance to the object changes.

[0031] By placing a binocular infrared speckle camera in a preset working environment, and controlling the camera to capture a series of image pairs from multiple angles and distances, it is possible to cover different scene conditions, such as different lighting, object shapes and textures, in order to obtain rich training data.

[0032] Step S20: Process each pair of binocular speckle images using the binocular depth base model to obtain the pseudo-labels of the binocular disparity maps for each pair of binocular speckle images. It should be noted that the binocular depth foundation model is a pre-trained deep learning model, such as the FoundationStereo foundation model (a name for a binocular depth foundation model), which can compute disparity maps from a set of binocular speckle image pairs. The binocular disparity map pseudo-label refers to the disparity map generated by the binocular depth foundation model, which represents the pixel displacement between corresponding points in a set of binocular speckle image pairs. However, since it is based on model estimation rather than actual measurement, it is called a pseudo-label in this embodiment.

[0033] The acquired binocular speckle image pairs are input into a pre-trained binocular depth model. The model will output a disparity map for each image pair through steps such as feature extraction, cost calculation, and disparity optimization. These disparity maps serve as pseudo-labels, providing depth information of the scene.

[0034] Step S30: Based on the binocular stereo vision calibration results, monocular baseline calibration results, and each reference speckle image pair acquired by the binocular infrared speckle camera, convert each binocular disparity map pseudo-label into each monocular disparity map pseudo-label. The monocular disparity map pseudo-label corresponds to the virtual monocular infrared speckle device obtained by decomposition from the binocular infrared speckle camera. It should be noted that the binocular stereo vision calibration result refers to the camera parameters obtained through binocular stereo vision calibration processing, including intrinsic parameter matrices, distortion mapping tables, reprojection matrices, binocular translation matrices, and binocular baseline lengths, used to describe the camera's geometric characteristics. The monocular baseline calibration result refers to the calculated baseline length of a virtual monocular infrared device (consisting of a speckle projector and left / right infrared cameras). The reference speckle image pair refers to speckle image pairs acquired at different known distances, used for calibration and verification. The monocular disparity map pseudo-label refers to the disparity map applicable to the monocular speckle depth model, obtained by converting the binocular disparity map pseudo-label.

[0035] Using the parameters in the binocular stereo vision calibration results, the binocular disparity map pseudo-labels are converted into binocular depth maps. The depth map represents the actual distance of each pixel. Then, combining the monocular baseline length in the monocular baseline calibration results and the distance of the reference speckle image to the reference plane during acquisition, the binocular depth map is converted into monocular disparity map pseudo-labels through a geometric transformation formula.

[0036] Step S40: Construct a dataset for training a monocular speckle depth model based on each pair of binocular speckle images, each pair of reference speckle images, and each monocular disparity map pseudo-label.

[0037] It should be noted that a dataset refers to a collection of data used to train a machine learning model, which typically includes input images and their corresponding labels.

[0038] All binocular speckle image pairs, reference speckle image pairs, and monocular disparity map pseudo-labels are organized and paired to form sets of data samples. Each sample includes one speckle image as input, a corresponding reference speckle image as a reference, and a monocular disparity map pseudo-label as a supervision signal. The dataset can be divided according to the ratio of training, validation, and test sets to ensure the generalization ability of the model during training. In addition, data augmentation techniques such as rotation, scaling, and color transformation can be applied to increase the diversity and robustness of the data.

[0039] In one feasible embodiment, step S40 may include step S401: Step S401: Based on each pair of binocular speckle images, each pair of reference speckle images, and each pseudo-label of monocular disparity map, a dataset for training the monocular speckle depth model is constructed. The dataset includes each set of data sample pairs, and each set of data sample pairs includes a speckle image, a reference speckle image corresponding to the speckle image, and a pseudo-label of monocular disparity map.

[0040] It's important to note that data sample pairs are the basic units of the dataset, comprising a speckle image, a reference speckle image, and a monocular disparity map pseudo-label. The speckle image is one image from a pair of binocular speckle images; for example, the left image might be chosen as input. The reference speckle image is a reference image corresponding to the same scene, providing additional contextual information. The monocular disparity map pseudo-label serves as a supervision signal for the disparity map. In practice, all these elements need to be paired and organized to ensure that each sample pair is consistent in time, viewpoint, and distance.

[0041] Thus, in this embodiment, by utilizing real-world scene data collected by a binocular infrared speckle camera and combining it with a binocular depth model to generate binocular disparity map pseudo-labels, and then converting the binocular disparity map pseudo-labels into monocular disparity map pseudo-labels through calibration and conversion techniques, the binocular disparity map pseudo-labels are adapted to the training of the monocular speckle depth model. This effectively solves the problem of insufficient training data for the monocular infrared speckle depth model, avoids the high cost and difficulty of directly collecting the true monocular depth values, and improves the efficiency of dataset construction.

[0042] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Furthermore, steps A10 to A30 may be included before step S30: Step A10: Acquire the calibration plate image pairs acquired by the binocular infrared speckle camera from multiple perspectives, and acquire the reference speckle image pairs acquired by the binocular infrared speckle camera from different distances. It should be noted that calibration plate image pairs refer to image pairs acquired by a binocular infrared speckle camera at different viewing angles using a standard calibration plate (such as a checkerboard or dot pattern), which are used to calculate camera parameters.

[0043] Place the calibration plate within the field of view of the binocular infrared speckle camera and capture image pairs from multiple angles and positions to ensure coverage of the camera's entire field of view. Additionally, project speckle patterns onto planes at different distances (e.g., from 0.5 meters to 5 meters) and acquire reference speckle image pairs using the binocular infrared speckle camera to obtain n reference speckle images A from the left infrared camera. i ,i∈(1,n), n reference speckle images B from the right infrared camera i For i∈(1,n), the reference speckle images corresponding to the same sequence number are taken simultaneously.

[0044] Step A20: Calculate the binocular stereo vision calibration results based on the images of each calibration plate; It should be noted that the calibration algorithm is used to extract corner points or feature points in the calibration board image. Then, the intrinsic parameter matrix and distortion coefficient of the left and right cameras in the binocular infrared speckle camera are calculated by optimization. Subsequently, the rotation matrix and translation matrix between the left and right cameras are calculated as stereo vision parameters. These parameters obtained from the calibration are used as the binocular stereo vision calibration results for subsequent image correction and depth calculation.

[0045] In one feasible embodiment, step A20 may include steps A201 to A204: Step A201: Calculate the first calibration parameters of the binocular infrared speckle camera based on the image pairs of each calibration board. The first calibration parameters include the initial intrinsic parameter matrix and distortion coefficients of the binocular infrared speckle camera. It should be noted that the first calibration parameters include the initial intrinsic parameter matrix and distortion coefficients of the binocular infrared speckle camera. The intrinsic parameter matrix describes the internal geometric characteristics of the camera, such as focal length and principal point, while the distortion coefficients represent the radial and tangential distortion of the lens.

[0046] Using calibration software or libraries, such as OpenCV (OpenSource ComputerVisionLibrary, an open-source cross-platform computer vision library), process calibration board image pairs, detect corner points and optimize using the least squares method, and calculate the intrinsic parameter matrix and distortion coefficients for each camera.

[0047] For example, key point information is extracted from the calibration board image captured by the left infrared camera, and imported together with the calibration board parameters into the single-camera calibration function to obtain the intrinsic parameters I of the left camera. l and distortion coefficient D l The same steps are followed for the right infrared camera to obtain the right camera's intrinsic parameter I. r and distortion coefficient D r ; Step A202: The first calibration parameters are imported into the stereo calibration function for calculation to obtain the second calibration parameters. The second calibration parameters include the rotation matrix, translation matrix, basic matrix, essential matrix and stereo baseline length of the binocular infrared speckle camera. It should be noted that the stereo calibration function is an algorithmic function used to calculate the relative position and orientation between the two cameras. The second calibration parameters include the rotation matrix, translation matrix, fundamental matrix, essential matrix, and binocular baseline length. The rotation and translation matrices describe the rotation and translation of the right camera relative to the left camera. The fundamental and essential matrices are used for stereo matching, and the binocular baseline length is the physical distance between the centers of the two cameras.

[0048] The first calibration parameter is input into the stereo calibration function, which calculates the extrinsic parameters through the corresponding points to obtain the rotation matrix R between the left and right cameras. lr Translation matrix T lrThe fundamental matrix F and the essential matrix E. The translation matrix T... lr = (X lr ,Y lr Z lr ), representing the coordinates of the optical center of the right camera in the optical coordinate system of the left camera, i.e., the binocular baseline length T.

[0049] Step A203: The first calibration parameters and the second calibration parameters are imported into the stereo correction function for calculation to obtain the third calibration parameters. The third calibration parameters include the correction rotation matrix, correction intrinsic parameter matrix and reprojection matrix of the binocular infrared speckle camera. It should be noted that the stereo calibration function is used to align the images from both cameras to the same plane, simplifying stereo matching. The third calibration parameters include the calibration rotation matrix, the calibration intrinsic parameter matrix, and the reprojection matrix. The calibration rotation matrix rotates the image to align with the epipolar lines, the calibration intrinsic parameter matrix is ​​the calibrated camera intrinsic parameters, and the reprojection matrix is ​​used to map 2D points to 3D space. In practice, stereo calibration functions such as cv2.stereoRectify (a function in OpenCV used for stereo calibration) can be used to calculate these parameters to ensure row alignment of the left and right images.

[0050] For example, the first calibration parameters and the second calibration parameters are imported into the stereo correction function to obtain the correction rotation matrix R of the left and right cameras. l and R r Correction intrinsic parameter matrix K l and K r Because the intrinsic parameters of the left and right cameras are the same after calibration, i.e., K l =K r This will be denoted as K from now on. ; Among them, f x f is the focal length (in pixels) along the X-axis. y The focal length (in pixels) along the Y-axis, (c x ,c y () are the coordinates of the main point.

[0051] The reprojection matrix Q can also be obtained; ; in, , All of these are values ​​in the calibration intrinsic parameter matrix K. Let K be the principal point X-axis coordinate of the other camera in the binocular system. l =K r ,so T is the binocular baseline length; .

[0052] Step A204: The first calibration parameter, the second calibration parameter, and the third calibration parameter are imported into the distortion correction function for calculation to obtain the distortion correction mapping table of the binocular infrared speckle camera.

[0053] It should be noted that the distortion correction mapping table M l and M r It is a lookup table used for quickly correcting image distortion. In practice, functions such as cv2.initUndistortRectifyMap (a function in OpenCV used for image distortion correction and image calibration) can be used to generate a mapping table, which maps the original image coordinates to the corrected coordinates and is applied to all acquired images to eliminate distortion.

[0054] Thus, through a hierarchical calibration process, the camera parameters are gradually refined and corrected, ensuring the accuracy and consistency of the binocular vision system, thereby improving the accuracy of the dataset and the model training effect.

[0055] Step A30: Calculate the monocular baseline calibration result based on each reference speckle image pair and the binocular stereo vision calibration result.

[0056] The stereo calibration results of binocular vision are used to perform stereo correction on the reference speckle image pair to eliminate distortion and align the images. Then, feature points are selected on the corrected images, disparity maps are calculated, and disparity is converted into three-dimensional coordinates by combining the reprojection matrix. By fitting three-dimensional points at different distances, the position of the speckle projector is obtained. Finally, the monocular baseline length is calculated, and this monocular baseline length is used as the monocular baseline calibration result for subsequent dataset construction to ensure the geometric consistency of the virtual monocular device.

[0057] In one feasible embodiment, step A30 may include steps A301 to A306: Step A301: Perform stereo correction on each reference speckle image pair based on the distortion correction mapping table to obtain each corrected reference image pair; Stereo correction transforms the left and right images to the same plane, aligning corresponding points in the same row. In practice, this is achieved using a distortion correction mapping table M. l and M r For each reference speckle image pair A i and B i Perform remapping, eliminate distortion and align the image, and generate a corrected reference image pair C. i and D i ,i∈(1,n), which facilitates subsequent disparity calculation.

[0058] Step A302: Select multiple sampling points on each calibration reference image pair, and calculate the disparity of each calibration reference image pair based on each sampling point; Taking a set of calibration reference image pairs as an example, m sampling points L are selected from the calibration reference image C1 of the left camera. 1j For example, when m=9, the sampling point locations are distributed as follows: j∈(1,m). Figure 3 As shown, sampling points should not be too close to the image edges to avoid not finding corresponding points during template matching. For any point L 1j coordinates (u) 1j ,v 1j Using template matching methods, including but not limited to common methods such as SSD (Sum of Squared Differences), NCC (Normalized Cross-Correlation), and ZNCC (Zero-Normalized Cross-Correlation), the v value of the right camera reference speckle map D1 is obtained. 1j Search for the corresponding point R on the row 1j =(r 1j ,w 1j Then calculate the disparity d of the corrected reference image pair. 1j =u 1j -r 1j Similarly, disparity is calculated for each pair of calibration reference images.

[0059] Step A303: Calculate the three-dimensional coordinates of each sampling point based on the parallax and reprojection matrix; Taking a set of calibration reference image pairs as an example, for each sampling point, the three-dimensional coordinates of the sampling point are calculated using a formula based on the disparity and reprojection matrix. Specifically, for the three-dimensional coordinates P of a sampling point on the left camera calibration reference image C1, 1j = (X 1j ,Y 1j Z 1j The calculation formula is as follows:

[0060]

[0061] Where X, Y, and Z represent the spatial coordinate components of a point in three-dimensional space, corresponding to the x-axis, y-axis, and z-axis coordinates in the three-dimensional coordinate system, respectively; W is the scaling factor for homogeneous coordinates. And so on, the three-dimensional coordinates of all m sampling points on C1 are calculated.

[0062] Similarly, for m sampling points L on C1 1j According to the template matching method in C i Search for the corresponding point and find all corresponding points L. ij , i∈(2,n), j∈(1,m); for C iAll m corresponding points L ij According to the template matching method in D i Search for the corresponding point R ij Given i∈(2,n) and j∈(1,m), the three-dimensional point coordinates P of all the calibration reference image pairs are then calculated. ij .

[0063] Step A304: Based on the three-dimensional coordinates of the same sampling point at multiple distances, fit multiple three-dimensional straight lines to obtain them; The coordinates of the three-dimensional point P ij The sampling points are divided into m groups based on their index j, and a three-dimensional straight line is fitted to each group. Specifically, the coordinates P of the first group of three-dimensional points are used as the basis for the fitting. i1 Taking i∈(1,n) as an example, according to the three-dimensional linear symmetric equation, each three-dimensional point P i1 All satisfied ,make Based on the characteristics of the speckle projector, the center of the projector must be located at the intersection of m three-dimensional straight lines. The positions of the fitted three-dimensional straight lines are illustrated as follows: Figure 4 As shown.

[0064] Step A305: Calculate the coordinates of the speckle projector based on the linear equations of each three-dimensional line; All three-dimensional straight lines intersect at the speckle projector point. Substitute the coordinates of the speckle projector to be solved into the equations of each straight line to obtain the coordinates of the speckle projector.

[0065] In one feasible embodiment, step A305 may include steps A3051 to A3052: Step A3051: Substitute the coordinates of the speckle projector to be solved into the equations of the three-dimensional lines to construct an overdetermined system of equations. An overdetermined system of equations is a system of equations in which the number of equations exceeds the number of unknowns. For each three-dimensional straight line, its equation can be expressed as a point-directed equation or a parametric equation. By substituting the coordinates of the speckle projector as a common point into the equations of all straight lines, a set of linear equations is formed. Due to the large number of straight lines, an overdetermined system of equations is constructed.

[0066] Step A3052: Solve the least squares solution of the overdetermined system of equations to obtain the coordinates of the speckle projector.

[0067] It should be noted that least squares solution is an optimization method used to find the solution that minimizes the sum of squared residuals. In practice, numerical methods such as SVD (singular value decomposition) are used to solve the overdetermined system of equations to obtain the optimal estimate of the speckle projector coordinates, ensuring the accuracy and robustness of the coordinates.

[0068] Specifically, the overdetermined system of equations can be written in matrix form as follows:

[0069] The overdetermined system of equations in matrix form is decomposed using SVD to find its non-zero solutions, yielding a1, b1, and c1. Similarly, all parameters (X, Y, F, Z) of the equations for the m three-dimensional lines corresponding to the coordinates of all m sets of three-dimensional points are calculated. cj ,Y cj Z cj ,a j ,b j ,c j ),j∈(1,m).

[0070] Let the coordinates of the speckle projector be P. t =(X t ,Y t Z t Substituting this point into the equations of m three-dimensional lines, and rearranging, we obtain the overdetermined system of equations in matrix form as follows:

[0071] Solve the above equation using least squares to obtain the coordinates P of the speckle projector. t Because of solving P ij The process is based on the left camera coordinate system, therefore the fitted P t It is also located in the left camera coordinate system. In the left camera coordinate system, the optical center coordinates of the left camera are (0,0,0), and the optical center coordinates of the right camera are the coordinates of the binocular translation matrix T. lr = (X lr ,Y lr Z lr Projector coordinates P t =(X t ,Y t Z t ).

[0072] Step A306: Calculate the monocular baseline length of the virtual monocular infrared speckle device based on the speckle projector coordinates and translation matrix.

[0073] Using information from the translation matrix and the speckle projector coordinates, the monocular baseline length is obtained through geometric calculations. The monocular baseline length includes the baseline length L of the monocular speckle module consisting of the left camera and the projector. l The baseline length L of the monocular speckle module composed of the right camera and projector r The expression is as follows: ; .

[0074] Thus, through precise 3D coordinate fitting and geometric calculation, accurate speckle projector positions and monocular baseline lengths were obtained, ensuring the authenticity of the virtual monocular device parameters. This reduced errors during pseudo-label conversion, improved the reliability of monocular disparity map pseudo-labels, and ultimately enhanced the quality of the dataset and the performance of model training.

[0075] Based on this, step S30 may include steps S301 to S302: Step S301: Based on the calibration intrinsic parameter matrix and binocular baseline length in the binocular stereo vision calibration results, convert each binocular disparity map pseudo-label into each binocular depth map; It should be noted that the value of each pixel in the binocular depth map represents the actual distance of each pixel from the camera. In practice, the disparity-depth conversion formula is used to calculate the pixel depth value, i.e., v. Where v represents the depth value, To correct the intrinsic parameter matrix, T is the binocular baseline length, and s is the binocular disparity map pseudo-label S. lk and S rk The disparity value of any pixel. The depth map V is obtained by calculating the depth values ​​of all pixels. lk and V rk .

[0076] It's worth noting that the depth map is considered a true representation of the surrounding 3D environment in the camera coordinate system; therefore, V lk and V rk It is also equivalent to the depth map obtained by the monocular depth algorithm.

[0077] Step S302: Based on the calibration intrinsic parameter matrix, the monocular baseline length in the monocular baseline calibration results, and the distance between each reference speckle image and the reference plane during acquisition, convert each binocular depth map into a pseudo-label for each monocular disparity map.

[0078] Monocular disparity map pseudo-labels are disparity maps suitable for monocular speckle depth models. Specifically, they first use the reference plane distance (i.e., the known object distances when acquiring the reference speckle image pair) as a reference for calibrating depth values. Then, using the monocular baseline length and the calibration intrinsic parameter matrix, the binocular depth map is converted into a monocular disparity map pseudo-label through geometric relationships.

[0079] For example, using the left camera depth map V l and monocular reference image C i For example, let C i The distance between the reference plane and V during shooting is e. l For any pixel with a depth value of v, the corresponding monocular parallax is... Calculate the disparity values ​​of all pixels to obtain the monocular disparity pseudo-label W for the left camera corresponding to this reference speckle map. ln baseline speckle maps can generate n monocular parallax pseudo-labels; similarly, the right camera can also generate n monocular parallax pseudo-labels.

[0080] Thus, by adapting binocular depth information to a monocular speckle depth model, high-quality monocular disparity map pseudo-labels are generated, avoiding the difficulty of directly collecting monocular depth data and improving the efficiency of dataset construction.

[0081] For example, to help understand the implementation process of the dataset construction method for the monocular speckle depth model obtained by combining this embodiment with the above embodiment one, please refer to... Figure 5 , Figure 5 A simplified flowchart illustrating a method for constructing a dataset for a monocular speckle depth model is provided, specifically: First, each pair of stereo speckle images acquired by a stereo infrared speckle camera is obtained. Each pair includes an infrared speckle image from the left camera and an infrared speckle image from the right camera. The stereo depth model is then used to process each pair of stereo speckle images to obtain stereo disparity map pseudo-labels. Each set of stereo disparity map pseudo-labels includes a stereo disparity image from the left camera and a stereo disparity image from the right camera. Then, the stereo disparity map pseudo-labels are converted into stereo depth maps. Each set of stereo depth maps includes a depth image from the left camera and a depth image from the right camera. Then, pseudo-labels for monocular disparity maps are generated based on the binocular depth and each reference speckle image pair. The reference speckle image pair includes the left camera monocular reference image and the right camera monocular reference image. The pseudo-labels for monocular disparity maps include the left camera monocular disparity map and the right camera monocular disparity map. Finally, a dataset for training the monocular speckle depth model is constructed. The left camera infrared speckle image, any left camera monocular reference image, and its corresponding left camera monocular disparity map can constitute a set of training data. Similarly, the right camera infrared speckle image, any right camera monocular reference image, and its corresponding right camera monocular disparity map can also constitute a set of training data. The infrared speckle image and the monocular reference image serve as model inputs, and the monocular disparity map serves as the ground truth pseudo-label for the model output. Furthermore, data with different projectors and different baselines (adjusting the projector position) can be added to enrich the dataset and improve the robustness of subsequent model training results.

[0082] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the dataset construction method of the monocular speckle depth model of this application. Any simple transformations based on this technical concept are within the protection scope of this application.

[0083] This application also provides a dataset construction device for a monocular speckle depth model. Please refer to... Figure 6 The dataset construction device for the monocular speckle depth model includes: Image pair acquisition module 10 is used to acquire each binocular speckle image pair acquired by the binocular infrared speckle camera; The binocular pseudo-label generation module 20 is used to process each binocular speckle image pair through the binocular depth basic model to obtain the binocular disparity map pseudo-labels for each binocular speckle image pair. The monocular pseudo-label generation module 30 is used to convert each binocular disparity map pseudo-label into a monocular disparity map pseudo-label based on the binocular stereo vision calibration results, the monocular baseline calibration results, and each reference speckle image pair acquired by the binocular infrared speckle camera. The monocular disparity map pseudo-label corresponds to the virtual monocular infrared speckle device obtained by decomposition from the binocular infrared speckle camera. The dataset construction module 40 is used to construct a dataset for training a monocular speckle depth model based on each pair of binocular speckle images, each pair of reference speckle images, and each monocular disparity map pseudo-label.

[0084] Optionally, the dataset construction device for the monocular speckle depth model also includes a calibration result acquisition module (not shown), which is used for: Acquire pairs of calibration plate images from multiple perspectives captured by a binocular infrared speckle camera, and acquire pairs of reference speckle images from different distances captured by a binocular infrared speckle camera. The binocular stereo vision calibration results are calculated based on the images of each calibration board. The monocular baseline calibration result is calculated based on each reference speckle image pair and the binocular stereo vision calibration result.

[0085] Optionally, the calibration result acquisition module is also used for: The first calibration parameters of the binocular infrared speckle camera are calculated based on the image pairs of each calibration plate. The first calibration parameters include the initial intrinsic parameter matrix and distortion coefficients of the binocular infrared speckle camera. The first calibration parameters are imported into the stereo calibration function for calculation to obtain the second calibration parameters. The second calibration parameters include the rotation matrix, translation matrix, basic matrix, essential matrix and stereo baseline length of the binocular infrared speckle camera. The first and second calibration parameters are imported into the stereo correction function for calculation to obtain the third calibration parameter, which includes the correction rotation matrix, correction intrinsic parameter matrix and reprojection matrix of the binocular infrared speckle camera. The first, second, and third calibration parameters are imported into the distortion correction function for calculation, resulting in the distortion correction mapping table for the binocular infrared speckle camera.

[0086] Optionally, the calibration result acquisition module is also used for: Stereo correction is performed on each pair of reference speckle images based on the distortion correction mapping table to obtain each corrected reference image pair. Multiple sampling points are selected on each calibration reference image pair, and the disparity of each calibration reference image pair is calculated based on each sampling point; Calculate the three-dimensional coordinates of each sampling point based on the parallax and reprojection matrix; Based on the three-dimensional coordinates of the same sampling point at multiple distances, multiple three-dimensional straight lines are fitted to obtain the line. The coordinates of the speckle projector are calculated based on the linear equations of each three-dimensional line; The monocular baseline length of the virtual monocular infrared speckle device is calculated based on the speckle projector coordinates and translation matrix.

[0087] Optionally, the calibration result acquisition module is also used for: Substitute the coordinates of the speckle projector to be solved into the equations of the three-dimensional lines to construct an overdetermined system of equations. Solve the least squares solution of the overdetermined system of equations to obtain the coordinates of the speckle projector.

[0088] Optionally, the monocular pseudo-tag generation module 30 is also used for: Based on the calibration intrinsic parameter matrix and binocular baseline length in the binocular stereo vision calibration results, the pseudo-labels of each binocular disparity map are converted into each binocular depth map; Based on the calibration intrinsic parameter matrix, the monocular baseline length in the monocular baseline calibration results, and the distance between each reference speckle image and the reference plane at the time of acquisition, each binocular depth map is converted into a pseudo-label for each monocular disparity map.

[0089] Optionally, the dataset building module 40 is also used for: The dataset for training the monocular speckle depth model is constructed based on each pair of binocular speckle images, each pair of reference speckle images, and each monocular disparity map pseudo-label. The dataset includes each set of data sample pairs, and each set of data sample pairs includes a speckle image, a reference speckle image corresponding to the speckle image, and a monocular disparity map pseudo-label.

[0090] The monocular speckle depth model dataset construction apparatus provided in this application adopts the monocular speckle depth model dataset construction method in the above embodiments, which can improve the construction efficiency of training datasets for monocular speckle depth models based on monocular infrared speckle depth technology. Compared with the prior art, the beneficial effects of the monocular speckle depth model dataset construction apparatus provided in this application are the same as those of the monocular speckle depth model dataset construction method provided in the above embodiments, and other technical features in the monocular speckle depth model dataset construction apparatus are the same as those disclosed in the monocular speckle depth model dataset construction method in the above embodiments, and will not be repeated here.

[0091] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the dataset construction method for the monocular speckle depth model in the first embodiment described above.

[0092] The following is for reference. Figure 7 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as laptops, PDAs (Personal Digital Assistants), PADs (Portable Application Description), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as desktop computers. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0093] like Figure 7 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0094] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0095] The electronic device provided in this application adopts the dataset construction method of the monocular speckle depth model in the above embodiments, which can improve the construction efficiency of the training dataset of the monocular speckle depth model based on monocular infrared speckle depth technology. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as the beneficial effects of the dataset construction method of the monocular speckle depth model provided in the above embodiments, and other technical features in the electronic device are the same as the features disclosed in the dataset construction method of the monocular speckle depth model in the previous embodiment, and will not be repeated here.

[0096] It should be understood that the various parts disclosed in the embodiments of this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0097] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0098] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the dataset construction method for the monocular speckle depth model in the above embodiments.

[0099] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0100] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0101] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire each pair of binocular speckle images captured by a binocular infrared speckle camera; process each pair of binocular speckle images using a binocular depth baseline model to obtain pseudo-labels for each pair of binocular disparity maps; based on binocular stereo vision calibration results, monocular baseline calibration results, and each pair of reference speckle images captured by the binocular infrared speckle camera, convert each pseudo-label for a binocular disparity map into a pseudo-label for a monocular disparity map, wherein the pseudo-label for a monocular disparity map corresponds to a virtual monocular infrared speckle device obtained by decomposing the binocular infrared speckle camera; and construct a dataset for training a monocular speckle depth model based on each pair of binocular speckle images, each pair of reference speckle images, and each pseudo-label for a monocular disparity map.

[0102] Computer program code for performing the operations of the embodiments of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0104] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0105] The readable storage medium provided in this application embodiment is a computer-readable storage medium. This medium stores computer-readable program instructions (i.e., a computer program) for executing the dataset construction method of the monocular speckle depth model described above, which can improve the efficiency of constructing the training dataset of the monocular speckle depth model based on monocular infrared speckle depth technology. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the dataset construction method of the monocular speckle depth model provided in the above embodiments, and will not be repeated here.

[0106] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for constructing a dataset of a monocular speckle depth model, characterized in that, The dataset construction method of the monocular speckle depth model comprises: acquiring each binocular speckle image pair collected by a binocular infrared speckle camera; processing each binocular speckle image pair by a binocular depth base model to obtain a respective binocular disparity map pseudo-label of each binocular speckle image pair; based on binocular stereo vision calibration results, monocular baseline calibration results and each reference speckle image pair collected by the binocular infrared speckle camera, converting each binocular disparity map pseudo-label into a respective monocular disparity map pseudo-label, wherein the monocular disparity map pseudo-label corresponds to a virtual monocular infrared speckle device decomposed from the binocular infrared speckle camera; according to each binocular speckle image pair, each reference speckle image pair and each monocular disparity map pseudo-label, constructing a dataset for training a monocular speckle depth model.

2. The monoscopic speckle depth model dataset construction method of claim 1, wherein, Before the step of converting each binocular disparity map pseudo-label into a respective monocular disparity map pseudo-label based on binocular stereo vision calibration results, monocular baseline calibration results and each reference speckle image pair collected by the binocular infrared speckle camera, the method further comprises: acquiring each calibration board image pair collected by the binocular infrared speckle camera from multiple viewing angles, and acquiring each reference speckle image pair collected by the binocular infrared speckle camera from different distances; calculating binocular stereo vision calibration results according to each calibration board image pair; calculating monocular baseline calibration results according to each reference speckle image pair and the binocular stereo vision calibration results.

3. The monoscopic speckle depth model dataset construction method of claim 2, wherein, The step of calculating binocular stereo vision calibration results according to each calibration board image pair comprises: calculating first calibration parameters of the binocular infrared speckle camera based on each calibration board image pair, wherein the first calibration parameters comprise an initial intrinsic matrix and distortion coefficients of the binocular infrared speckle camera; introducing the first calibration parameters into a stereo calibration function for calculation to obtain second calibration parameters, wherein the second calibration parameters comprise a rotation matrix, a translation matrix, an essential matrix and a binocular baseline length of the binocular infrared speckle camera; introducing the first calibration parameters and the second calibration parameters into a stereo rectification function for calculation to obtain third calibration parameters, wherein the third calibration parameters comprise a rectified rotation matrix, a rectified intrinsic matrix and a re-projection matrix of the binocular infrared speckle camera; introducing the first calibration parameters, the second calibration parameters and the third calibration parameters into a de-distortion rectification function for calculation to obtain a de-distortion rectification mapping table of the binocular infrared speckle camera.

4. The method of claim 3, wherein the monoscopic speckle depth model is constructed by, The step of calculating monocular baseline calibration results according to each reference speckle image pair and the binocular stereo vision calibration results comprises: performing stereo rectification on each reference speckle image pair based on the de-distortion rectification mapping table to obtain each rectified reference image pair; selecting multiple sampling points on each rectified reference image pair and calculating a disparity of each rectified reference image pair according to each sampling point; calculating three-dimensional coordinates of each sampling point according to the disparity and the re-projection matrix; fitting to obtain multiple three-dimensional straight lines based on the three-dimensional coordinates of the same sampling point at multiple distances; calculating speckle projector coordinates according to straight line equations of each three-dimensional straight line; A monocular baseline length of a virtual monocular infrared speckle device is calculated according to the speckle projector coordinates and the translation matrix.

5. The method of claim 4, wherein, The step of calculating the speckle projector coordinates according to the straight line equations of the three-dimensional straight lines comprises: substituting the to-be-solved speckle projector coordinates into the straight line equations of the three-dimensional straight lines to construct an over-determined equation set; solving the least square solution of the over-determined equation set to obtain the speckle projector coordinates.

6. The data set construction method of a monocular speckle depth model according to any one of claims 1 to 5, wherein, The step of converting the binocular disparity map pseudo-labels into monocular disparity map pseudo-labels based on the binocular stereo vision calibration result and the monocular baseline calibration result comprises: converting each binocular disparity map pseudo-label into a binocular depth map based on a corrected intrinsic parameter matrix and a binocular baseline length in the binocular stereo vision calibration result; converting each binocular depth map into each monocular disparity map pseudo-label based on the corrected intrinsic parameter matrix, a monocular baseline length in the monocular baseline calibration result, and a reference plane distance when each reference speckle image pair is collected.

7. The monoscopic speckle depth model dataset construction method of claim 1, wherein, The step of constructing a data set for training a monocular speckle depth model according to each binocular speckle image pair, each reference speckle image pair, and each monocular disparity map pseudo-label comprises: constructing a data set for training a monocular speckle depth model according to each binocular speckle image pair, each reference speckle image pair, and each monocular disparity map pseudo-label, wherein the data set comprises a plurality of groups of data sample pairs, and each group of data sample pairs comprises one speckle image, a reference speckle image corresponding to the speckle image, and a monocular disparity map pseudo-label.

8. A data set construction apparatus of a monocular speckle depth model, characterized by, The data set construction apparatus of the monocular speckle depth model comprises: an image pair acquisition module configured to acquire each binocular speckle image pair collected by a binocular infrared speckle camera; a binocular pseudo-label generation module configured to process each binocular speckle image pair by a binocular depth base model to obtain a binocular disparity map pseudo-label corresponding to each binocular speckle image pair; a monocular pseudo-label generation module configured to convert each binocular disparity map pseudo-label into a monocular disparity map pseudo-label based on a binocular stereo vision calibration result, a monocular baseline calibration result, and each reference speckle image pair collected by the binocular infrared speckle camera, wherein the monocular disparity map pseudo-label corresponds to a virtual monocular infrared speckle device decomposed from the binocular infrared speckle camera; a data set construction module configured to construct a data set for training a monocular speckle depth model according to each binocular speckle image pair, each reference speckle image pair, and each monocular disparity map pseudo-label.

9. An electronic device, comprising: The electronic device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the monocular speckle depth model data set construction method according to any one of claims 1 to 7.

10. A storage medium, characterized by The storage medium is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the monocular speckle depth model data set construction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Monocular depth estimation method, device and equipment and storage medium

    CN108961327A

  • Unsupervised monocular depth estimation method based on generative adversarial network

    CN110443843A

  • Depth image generation method and system, electronic equipment and readable storage medium

    CN117152223A

  • Novel knowledge distillation method from binocular parallax to monocular depth

    CN119478003A

  • Depth image generation method and device

    US20210150747A1