Data set construction method and device, electronic equipment and storage medium

By acquiring three-dimensional models and image pairs using a three-dimensional scanner, determining the camera position and generating a parallax map, the problems of difficult and poor accuracy in the construction of data sets in the prior art are solved, and a higher accuracy data set construction is achieved.

CN120219151APending Publication Date: 2025-06-27CHENGDU XIANLIN 3D TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510397486.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing method of constructing structures matching data sets has the problem of high data construction difficulty and poor data construction accuracy.

Method used

By obtaining the three-dimensional model and image pair of the target object, the camera position is determined, and the points on the three-dimensional model are projected on the image plane based on the camera position are obtained, and the first disparity map is filtered to obtain, and the second disparity map is obtained for constructing the body matching data set.

Benefits of technology

It reduces the difficulty of building a data set, improves the accuracy of the built data set, and provides more accurate and stable deep data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219151A_ABST
    Figure CN120219151A_ABST
Patent Text Reader

Abstract

The invention provides a data set construction method and device, electronic equipment and a storage medium, and is suitable for the field of data processing, and the method comprises the steps: obtaining a three-dimensional model and an image pair obtained through the three-dimensional scanning of a target object; determining a camera pose by using the three-dimensional model and the image pair; projecting points on the three-dimensional model to an image plane corresponding to the image pair based on the camera pose to obtain a first disparity map; and filtering pixels in the first disparity map to obtain a second disparity map used for constructing the stereo matching data set. According to the scheme, the camera pose is determined by using the three-dimensional model and the image pair obtained in the three-dimensional scanning process, and the second disparity map used for constructing the stereo matching data set is generated based on the camera pose and the projection, so that the construction difficulty of the data set is reduced, and the precision of the constructed data set is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method, device, electronic device and storage medium for constructing a data set. Background Art

[0002] Stereo matching is one of the key problems in computer vision and robotics, and is widely used in fields such as depth perception and 3D reconstruction.

[0003] Existing stereo matching data sets are relatively scarce and difficult to produce. To produce a high-quality stereo matching data set, precise sensors are required to obtain accurate 3D information. However, on the one hand, the acquisition process and annotation process of the stereo matching data set are relatively complex, and on the other hand, the data in the real environment often contains noise, which results in the matching accuracy being affected.

[0004] In summary, the existing methods for constructing stereo matching data sets have problems such as large data construction difficulty and poor data construction accuracy. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method, device, electronic device and storage medium for constructing a data set to solve problems such as large data construction difficulty and poor data construction accuracy existing in the existing methods for constructing stereo matching data sets.

[0006] To achieve the above object, embodiments of the present invention provide the following technical solutions:

[0007] A first aspect of an embodiment of the present invention discloses a method for constructing a data set, the method including:

[0008] Obtaining a 3D model and an image pair obtained by 3D scanning a target object;

[0009] Determining a camera pose by using the 3D model and the image pair;

[0010] Projecting points on the 3D model onto the image plane corresponding to the image pair based on the camera pose to obtain a first disparity map;

[0011] Filtering pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching data set.

[0012] Preferably, the image pair includes: a left image taken by a left camera and a right image taken by a right camera;

[0013] Determining a camera pose by using the 3D model and the image pair includes:

[0014] Sampling 3D object points from the 3D model;

[0015] Using the coordinates on the left image corresponding to the points of the three-dimensional object and the camera parameters corresponding to the left image, minimizing a preset reprojection error function to obtain the camera pose of the left camera;

[0016] Using the coordinates of the three-dimensional object points on the right image and the camera parameters corresponding to the right image, minimizing the reprojection error function to obtain the camera pose of the right camera, where the camera parameters include internal parameters and external parameters.

[0017] Preferably, based on the camera pose, projecting the points on the three-dimensional model onto the image plane corresponding to the image pair to obtain a first disparity map, including:

[0018] Using the camera poses of the left camera and the right camera, respectively converting the true coordinates of the three-dimensional object points into the left camera coordinate system and the right camera coordinate system;

[0019] Based on the internal parameters of the left camera and the true coordinates of the three-dimensional object points converted into the left camera coordinate system, projecting the three-dimensional object points onto the image plane of the left camera to obtain the coordinates of the pixels in the left image;

[0020] Based on the internal parameters of the right camera and the true coordinates of the three-dimensional object points converted into the right camera coordinate system, projecting the three-dimensional object points onto the image plane of the right camera to obtain the coordinates of the pixels in the right image;

[0021] Calculating the difference between the coordinates of the corresponding pixels projected from the same three-dimensional object point in the left image and the right image to obtain a first disparity map.

[0022] Preferably, filtering the pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching dataset, including:

[0023] Calculating the depth error and disparity error of each pixel in the first disparity map;

[0024] Removing the pixels in the first disparity map whose depth error exceeds a first threshold and / or whose disparity error exceeds a second threshold to obtain a second disparity map for constructing a stereo matching dataset.

[0025] Preferably, the process of calculating the depth error of each pixel in the first disparity map includes:

[0026] Converting the disparity values of each pixel in the first disparity map into depth values;

[0027] Using the three-dimensional model to obtain a depth map and rendering the depth map to obtain true depth values;

[0028] Calculate the difference between the depth value of each pixel in the first disparity map and the true depth value to obtain the depth error of each pixel in the first disparity map.

[0029] Preferably, the process of calculating the disparity error of each pixel in the first disparity map includes:

[0030] Perform stereo matching on the image pair to obtain an estimated disparity;

[0031] Calculate the difference between the disparity value of each pixel in the first disparity map and the estimated disparity to obtain the disparity error of each pixel in the first disparity map.

[0032] Preferably, after sampling three-dimensional object points from the three-dimensional model, it further includes:

[0033] Perform epipolar rectification on the left image and the right image.

[0034] A second aspect of the embodiments of the present invention discloses a dataset construction device, the device includes:

[0035] An acquisition unit, configured to acquire a three-dimensional model and an image pair obtained by three-dimensional scanning of a target object;

[0036] A determination unit, configured to determine the camera pose by using the three-dimensional model and the image pair;

[0037] A projection unit, configured to project the points on the three-dimensional model onto the image plane corresponding to the image pair based on the camera pose to obtain a first disparity map;

[0038] A filtering unit, configured to filter the pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching dataset.

[0039] A third aspect of the embodiments of the present invention discloses an electronic device, including: a processor and a memory, the processor and the memory are connected by a communication bus; wherein, the processor is configured to call and execute a program stored in the memory; the memory is configured to store a program, and the program is used to implement the dataset construction method disclosed in the first aspect of the embodiments of the present invention.

[0040] A fourth aspect of the embodiments of the present invention discloses a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the dataset construction method disclosed in the first aspect of the embodiments of the present invention.

[0041] Based on the dataset construction method, apparatus, electronic device, and storage medium provided in the embodiments of the present invention described above, the method is as follows: Obtain a 3D model and an image pair obtained by performing 3D scanning on a target object; Use the 3D model and the image pair to determine the camera pose; Based on the camera pose, project the points on the 3D model onto the image plane corresponding to the image pair to obtain a first disparity map; Filter the pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching dataset. This solution uses the 3D model and the image pair obtained during the 3D scanning process to determine the camera pose, and then generates a second disparity map for constructing a stereo matching dataset based on the camera pose and projection, reducing the difficulty of dataset construction and improving the accuracy of the constructed dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.

[0043] Figure 1 It is a flowchart of a dataset construction method provided in an embodiment of the present invention;

[0044] Figure 2 It is a flowchart of determining the camera pose provided in an embodiment of the present invention;

[0045] Figure 3 It is a flowchart of obtaining the first disparity map provided in an embodiment of the present invention;

[0046] Figure 4 It is a flowchart of obtaining the second disparity map provided in an embodiment of the present invention;

[0047] Figure 5 It is another flowchart of a dataset construction method provided in an embodiment of the present invention;

[0048] Figure 6 It is an example diagram of the workflow of the dataset construction method provided in an embodiment of the present invention;

[0049] Figure 7 It is a structural block diagram of a dataset construction apparatus provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] In this application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0052] Stereo Matching is one of the key problems in computer vision and robotics, and is widely used in fields such as depth perception, 3D reconstruction, and visual SLAM. Although significant progress has been made in this field through deep learning algorithms in recent years, existing stereo matching datasets still have limitations, resulting in limitations in their effectiveness and universality in practical applications.

[0053] Existing stereo matching datasets are relatively scarce and difficult to produce. Most publicly available stereo matching datasets mainly focus on outdoor or specific scenarios (such as urban landscapes and road environments with medium complexity). The data in these scenarios has certain reference value for algorithm evaluation, but in practical applications, they cannot cover all complex real-world scenarios (especially the lack of industrial scene datasets). These environments have unique challenges (such as smooth surfaces, complex textures, and low-light conditions, etc.), which all make traditional stereo matching algorithms face certain difficulties.

[0054] Producing a high-quality stereo matching dataset requires precise sensors (such as lidar) to obtain accurate 3D information. However, the acquisition process and annotation process of stereo matching datasets are relatively complex and costly. In addition, the data in the real environment often contains noise, which also affects the matching accuracy and increases the difficulty of producing stereo matching datasets.

[0055] It has been found through research that the disadvantages of various traditional methods for producing stereo matching datasets are as follows:

[0056] (1) Data acquisition method based on a stereo camera: First, the accuracy of a stereo camera is limited by the calibration of the camera and the performance of the sensor. Especially in scenes with less texture or poor lighting conditions, parallax calculation may become difficult. Second, sufficient overlapping areas are required between image pairs to ensure rich parallax information, which is difficult to achieve in some specific scenarios (such as narrow spaces or environments with severe object occlusion). In addition, the baseline and viewing angle settings of the stereo camera also affect the accuracy of depth calculation. Especially in long-distance or large-scale scenes, the parallax change may be small, resulting in inaccurate depth information.

[0057] (2) Data acquisition method based on the fusion of LiDAR and vision sensors: The calibration process of LiDAR and cameras is complex and requires precise calibration to ensure the correct registration of point clouds and images. LiDAR devices are usually expensive and are greatly affected by environmental factors. The resolution difference between LiDAR point clouds and camera images may lead to inaccurate alignment in some scenes, thereby affecting the accuracy of depth ground truth.

[0058] (3) Data acquisition method based on synthetic data: There may be a gap between the depth maps and parallax maps generated by the synthetic data generation method and real-world scenes. Objects in virtual scenes are usually regular and have smooth surfaces, making it difficult to simulate complex textures, reflections, lighting changes, etc. that exist in the real environment. This may result in the generated data not being able to fully represent the challenges in actual applications. Synthetic data lacks the complexity of the real world, resulting in weak generalization ability of the trained algorithms in actual applications.

[0059] To solve the problems of poor accuracy, weak environmental adaptability, and scarcity of industrial datasets in the production of traditional stereo matching datasets, this solution proposes a dataset construction method, device, electronic device, and storage medium. The camera pose is determined using the 3D model and image pairs obtained during the 3D scanning process, and then a second parallax map for constructing a stereo matching dataset is generated based on the camera pose and projection, reducing the difficulty of dataset construction and improving the accuracy of the constructed dataset.

[0060] Compared with the traditional stereo matching dataset production method, the 3D scanner adopted in this solution can provide higher-precision depth information. By using the 3D scanner to generate a 3D model of the target object, the errors generated between the stereo camera and the LiDAR sensor are avoided. The camera pose calculation and projection generation of parallax ground truth during the scanning process can effectively reduce parallax errors and provide more accurate and stable depth data. The following will explain this solution in detail through various embodiments.

[0061] See Figure 1 , which shows a flowchart of a dataset construction method provided by an embodiment of the present invention. The dataset construction method includes the following steps:

[0062] Step S101: Obtain the 3D model and image pair obtained by performing 3D scanning on the target object.

[0063] In the process of specifically implementing step S101, when using a high-precision 3D scanner (only an example) to scan the target object, save the image pair (also called the left and right image pair) during the scanning process and the final 3D model. Obtain the 3D model and image pair obtained by performing 3D scanning on the target object.

[0064] It should be noted that the device used for 3D scanning of the target object (such as a high-precision 3D scanner) includes at least a left camera and a right camera; the image pair includes: the left image obtained by the left camera (that is, the photo taken by the left camera) and the right image obtained by the right camera (that is, the photo taken by the right camera).

[0065] Step S102: Determine the camera poses using the 3D model and the image pair.

[0066] It should be noted that the production of the true disparity value (that is, the subsequent second disparity map) first requires the estimation of the camera pose, the purpose of which is to determine the position and orientation of the camera (such as the left camera) in the 3D space.

[0067] In the process of specifically implementing step S102, using the 3D model and the image pair obtained by performing 3D scanning on the target object, combined with the global optimization algorithm, determine the camera poses of the left camera and the right camera.

[0068] It can be understood that the process of determining the camera pose depends on the camera parameters of the camera (including internal parameters and external parameters). The internal parameters of the camera are obtained through calibration; through the camera pose estimation based on the global optimization algorithm, use the multi-view geometry method to estimate the relative pose of different views with respect to the first frame, and calculate the accurate spatial relationship under different views.

[0069] Step S103: Based on the camera poses, project the points on the 3D model onto the corresponding image planes of the image pair to obtain the first disparity map.

[0070] In the process of specifically implementing step S103, after determining the camera poses of the left camera and the right camera, based on the camera poses of the left camera and the right camera, project the points on the 3D model onto the corresponding image planes of the image pair to obtain the first disparity map, and the first disparity map contains the disparity values of the pixels at each pixel position.

[0071] Among them, the disparity value of the pixel at a certain pixel position in the first disparity map represents: the difference in the pixel position of the left image and the right image at this pixel position.

[0072] Step S104: Filter the pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching dataset.

[0073] In the specific implementation process of step S104, based on the disparity values of the pixels in the first disparity map, the pixels in the first disparity map are filtered, so as to obtain a second disparity map for constructing a stereo matching dataset.

[0074] In the embodiment of the present invention, the camera pose is determined by using the three-dimensional model and the image pair obtained during the three-dimensional scanning process, and then a second disparity map for constructing a stereo matching dataset is generated based on the camera pose and projection, reducing the construction difficulty of the dataset and improving the accuracy of the constructed dataset.

[0075] It should be noted that through the above embodiment of the present invention Figure 1 As can be seen from the content in step S102, the camera pose is determined by the global optimization algorithm. Before explaining the process of determining the camera pose, the global optimization algorithm is described here first:

[0076] The global optimization algorithm first regards the camera pose and the positions of the three-dimensional object points (also simply referred to as three-dimensional points) as unknown parameters, and defines the optimization objective through the reprojection error. Each three-dimensional object point needs to be matched with the corresponding two-dimensional pixel point (or two-dimensional image point).

[0077] It should be noted that the three-dimensional object points in space can be converted into the corresponding two-dimensional pixel points by using the camera projection model.

[0078] The global optimization algorithm optimizes the camera pose and the positions of the three-dimensional object points by minimizing the reprojection error of all two-dimensional pixel points. The specific process of the global optimization algorithm is as follows: Assume that there are n three-dimensional object points in space, and m images are taken around these three-dimensional object points. Then the coordinates (spatial coordinates) of the i-th three-dimensional object point seen in the j-th image are denoted as x ij .

[0079] The global optimization algorithm aims to optimize the parameter estimation of multiple cameras and structures, so as to find reasonable parameters to accurately calculate the coordinates of n three-dimensional object points in m images. The camera parameters (including internal parameters and external parameters) of the j-th image are represented by a j and each three-dimensional object point i is represented by a vector b i .

[0080] The core problem of the global optimization algorithm is to minimize the reprojection error function shown in formula (1).

[0081] (1);

[0082] In formula (1), the function Q(aj , b i ) represents the three-dimensional object point b i The projection coordinates under the camera parameter a j , Q(a j ,b i ) is equivalent to the predicted image coordinates, x ij is equivalent to the observed image coordinates, and the function d(x, y) represents the Euclidean distance between the observed image coordinates and the predicted image coordinates.

[0083] The camera pose can be obtained by minimizing Equation (1). After knowing the camera pose, combined with the internal parameters, the two-dimensional to three-dimensional and three-dimensional to two-dimensional transformations can be achieved.

[0084] The above is the related description of the global optimization algorithm.

[0085] Based on the above global optimization algorithm, refer to Figure 2 , which shows the flowchart of determining the camera pose provided by the embodiment of the present invention, Figure 2 including the following steps:

[0086] Step S201: Sample three-dimensional object points from the three-dimensional model.

[0087] In the process of specifically implementing step S201, sampling from the three-dimensional model of the target object can obtain three-dimensional object points.

[0088] In some embodiments, after sampling the three-dimensional object points, rectify the epipolar lines of the left image and the right image to remove the distortion between the images under the perspectives of the left camera and the right camera, so that the corresponding feature points in the two images (the corresponding left image and right image) are on the same horizontal line.

[0089] Step S202: Use the coordinates on the left image corresponding to the three-dimensional object points and the camera parameters corresponding to the left image to minimize a preset reprojection error function to obtain the camera pose of the left camera.

[0090] It should be noted that the camera parameters include internal parameters and external parameters.

[0091] In the process of specifically implementing step S202, use the coordinates on the left image corresponding to the three-dimensional object points and the camera parameters corresponding to the left image, and use the global optimization algorithm to minimize the reprojection error function, thereby calculating the camera pose of the left camera for each frame.

[0092] Step S203: Use the coordinates of the three-dimensional object points on the right image and the camera parameters corresponding to the right image to minimize the reprojection error function to obtain the camera pose of the right camera.

[0093] In the process of specifically implementing step S203, the global optimization algorithm is used to minimize the reprojection error function by using the coordinates of the three-dimensional object points on the right image and the corresponding camera parameters of the right image, so as to calculate the camera pose of the right camera for each frame.

[0094] The above embodiments of the present invention Figure 2 are related descriptions about determining the camera pose.

[0095] For the above embodiments of the present invention Figure 1 Regarding the camera pose involved in step S103, project the points on the three-dimensional model onto the image planes corresponding to the image pair to obtain the first disparity map. Refer to Figure 3 which shows the flowchart of obtaining the first disparity map provided by the embodiments of the present invention. Figure 3 It includes the following steps:

[0096] Step S301: Use the camera poses of the left camera and the right camera to respectively convert the true coordinates of the three-dimensional object points into the left camera coordinate system and the right camera coordinate system.

[0097] In the process of specifically implementing step S301, use the camera pose of the left camera calculated by the global optimization algorithm to convert the true coordinates of the three-dimensional object points into the left camera coordinate system, where the true coordinates of the three-dimensional object points are denoted as P world =(X, Y, Z).

[0098] Use the camera pose of the right camera calculated by the global optimization algorithm to convert the true coordinates of the three-dimensional object points into the right camera coordinate system.

[0099] It should be noted that the left camera coordinate system and the right camera coordinate system differ by the baseline distance B in the X component of the translation vector.

[0100] Specifically, convert the true coordinates of the three-dimensional object points into the left camera coordinate system through formula (2).

[0101] P left =R left P world +T left (2);

[0102] In formula (2), R represents the rotation matrix and T represents the translation matrix; convert the points in the world coordinate system to the left camera coordinate system, and P left is the coordinate of the three-dimensional object point in the left camera coordinate system (i.e., the true coordinates of the three-dimensional object point converted to the left camera coordinate system), R left is the rotation matrix from the world to the left camera coordinate system, T left is the translation matrix from the world to the left camera coordinate system, and P worldare the true coordinates of the three-dimensional object points (coordinates in the world coordinate system).

[0103] It should be noted that for the specific process of converting the true coordinates of the three-dimensional object points to the right camera coordinate system, reference can be made to the above description regarding "converting the true coordinates of the three-dimensional object points to the left camera coordinate system", and it will not be elaborated here again.

[0104] Step S302: Based on the internal parameters of the left camera and the true coordinates of the three-dimensional object points converted to the left camera coordinate system, project the three-dimensional object points onto the image plane of the left camera to obtain the coordinates of the pixels in the left image.

[0105] In the specific implementation process of step S302, based on the "internal parameters of the left camera" and the "true coordinates of the three-dimensional object points converted to the left camera coordinate system", project the three-dimensional object points onto the image plane of the left camera to obtain the coordinates of the pixels in the left image.

[0106] Among them, the coordinates of the pixels in the left image are denoted as u left , u left =K left *P left , where u left represents the coordinate of the pixel u obtained by projecting the three-dimensional object points onto the image plane of the left camera, K left is the internal parameter of the left camera, and P left is the three-dimensional object point in the left camera coordinate system.

[0107] Step S303: Based on the internal parameters of the right camera and the true coordinates of the three-dimensional object points converted to the right camera coordinate system, project the three-dimensional object points onto the image plane of the right camera to obtain the coordinates of the pixels in the right image.

[0108] In the specific implementation process of step S303, based on the "internal parameters of the right camera" and the "true coordinates of the three-dimensional object points converted to the right camera coordinate system", project the three-dimensional object points onto the image plane of the right camera to obtain the coordinates of the pixels in the right image.

[0109] Among them, the coordinates of the pixels in the right image are denoted as u right , u right =K right *P right , where u right represents the coordinate of the pixel u obtained by projecting the three-dimensional object points onto the image plane of the right camera, K right is the internal parameter of the right camera, and P right is the three-dimensional object point in the right camera coordinate system.

[0110] Step S304: Calculate the difference between the coordinates of the corresponding pixels obtained by projecting the same three-dimensional object point in the left image and the right image to obtain the first disparity map.

[0111] In the process of specifically implementing step S304, after obtaining the coordinates of the pixels in the left image and the right image, calculate the difference (i.e., the disparity value) between the coordinates of the corresponding pixels projected from the same three-dimensional object point in the left image and the right image, to obtain a first disparity map. The first disparity map contains the disparity values (denoted as d) of the pixels at each pixel position; the disparity value of the pixel at a certain pixel position in the first disparity map represents the difference between the pixel positions in the left image and the right image.

[0112] Specifically, the disparity value d of the pixel is calculated by "d = u left - u right ", where u left and u right represent the coordinates of the pixel u projected from the same three-dimensional object point onto the image planes of the left camera and the right camera. Subtracting u left from u right can obtain the disparity value d, and a disparity value d is calculated for each three-dimensional object point.

[0113] As can be seen from the content of the above steps S301 to S304, projecting the three-dimensional object point onto the cameras (left camera and right camera) can obtain a depth map, and the first disparity map can be obtained by conversion of this depth map.

[0114] The above embodiments of the present invention Figure 3 are related descriptions on how to obtain the first disparity map.

[0115] Regarding the above embodiments of the present invention Figure 1 Step S104 involves obtaining a second disparity map. Refer to Figure 4 , which shows a flowchart of obtaining the second disparity map provided by the embodiments of the present invention. Figure 4 It includes the following steps:

[0116] Step S401: Calculate the depth error and disparity error of each pixel in the first disparity map.

[0117] It should be noted that the process of disparity filtering aims to improve the quality of the disparity map, and improves the accuracy by removing pixels with a large difference between the depth value and the actual depth. Therefore, it is necessary to first calculate the depth error and disparity error of each pixel in the first disparity map.

[0118] In the process of specifically implementing step S401, the specific method for calculating the depth error of each pixel in the first disparity map is: convert the disparity value d of each pixel in the first disparity map into a depth value. Specifically, the calculated disparity value d is converted into a depth value Z by "Z = f * B / d", where f represents the camera focal length, B represents the baseline distance, and both f and B are fixed parameters obtained through camera calibration.

[0119] A depth map is obtained using a three-dimensional model, and the true depth value is obtained by rendering the depth map. Specifically, the depth map can be projected using the three-dimensional model and camera parameters, and then the true depth value (denoted as Z) can be obtained by rendering the depth map. true )

[0120] The depth value Z of each pixel in the first disparity map is obtained, and the true depth value Z is obtained. true After that, the difference between the depth value of each pixel in the first disparity map and the true depth value is calculated to obtain the depth error of each pixel in the first disparity map. The depth error of each pixel is "|Z - Z true |".

[0121] It should be noted that the "depth error" of each pixel in the first disparity map is the depth error at the position of each pixel in the first disparity map, and each pixel position corresponds to a three-dimensional object point.

[0122] In some other embodiments, the specific method for calculating the disparity error of each pixel in the first disparity map is: performing stereo matching on the image pair (left image and right image) to obtain an estimated disparity (denoted as d SGM ), calculating the difference between the disparity value of each pixel in the first disparity map and the estimated disparity to obtain the disparity error of each pixel in the first disparity map. The disparity error of each pixel is "|d - d SGM |".

[0123] Specifically, the semi-global matching algorithm (SGM) is used to match the left image and the right image, and the estimated disparity (d SGM ) is calculated. The difference between the disparity value of each pixel in the first disparity map and the estimated disparity is calculated, thereby obtaining the disparity error of each pixel in the first disparity map.

[0124] It should be noted that the "disparity error" of each pixel in the first disparity map is the disparity error at the position of each pixel in the first disparity map.

[0125] Step S402: Remove the pixels in the first disparity map whose depth error exceeds the first threshold and / or whose disparity error exceeds the second threshold to obtain a second disparity map for constructing a stereo matching data set.

[0126] In the process of specifically implementing step S402, the depth error (|Z - Z true |) and the disparity error (|d - d SGMAfter that, pixels with depth errors exceeding the first threshold and / or parallax errors exceeding the second threshold are removed from the first parallax map to obtain a second parallax map for constructing a stereo matching dataset, and this second parallax map is the true parallax value. This second parallax map can be used for deep learning to facilitate providing accurate references for the training and evaluation of deep learning algorithms.

[0127] Specifically, for the first parallax map, pixels with depth errors (|Z - Z true |) exceeding a preset first threshold and / or pixels with parallax errors (|d - d SGM |) exceeding a preset second threshold are removed, thereby obtaining the second parallax map.

[0128] Through the parallax filtering method given in the above steps S401 and S402, abnormal parallax and abnormal external parameter estimation caused by projection errors and incorrect pose calculations are effectively removed, thereby improving the accuracy of the final second parallax map (true parallax value).

[0129] The above Figure 4 is the related description of obtaining the second parallax map.

[0130] To further understand the above Figures 1 to 4 shown content, through Figure 5 to illustrate the present solution from an overall perspective by way of example, see Figure 5 , which shows another flowchart of a dataset construction method provided by an embodiment of the present invention, Figure 5 including the following steps:

[0131] Step S501: Camera calibration to obtain the camera parameters of the left camera and the right camera.

[0132] Step S502: When using a high-precision 3D scanner to scan the target object, save the image pairs during the scanning process and the final 3D model.

[0133] Step S503: Epipolar rectification, and use a global optimization algorithm to obtain the camera pose.

[0134] Step S504: Based on the camera pose, project the points on the 3D model onto the image planes corresponding to the image pairs to obtain the first parallax map.

[0135] Step S505: Convert the first parallax map into a depth map and compare it with the rendered depth map, and filter pixels with depth errors and / or parallax errors exceeding the threshold to obtain the second parallax map.

[0136] It should be noted that for the execution principles of steps S501 to S506, reference can be made to the above explanations of the embodiments of the present invention Figures 1 to 4 and will not be elaborated here.

[0137] Generally speaking, the workflow involved in this solution is as Figure 6 shown, Figure 6 The workflow example diagram of the dataset construction method shown includes the following parts: camera calibration, saving image pairs, calculating the camera pose using a global optimization algorithm, and back-projecting the model onto the camera to obtain a second disparity map; Figure 6 For the execution principles of each part in the shown workflow, reference can be made to the Figures 1 to 4 explanation in the above embodiments of the present invention, which will not be elaborated here.

[0138] In practical applications, the dataset construction method proposed in this solution can cover more complex scenarios. By using a 3D scanner to obtain the 3D model of the target object, data can be collected in an industrial environment (speckle scene) to generate disparity ground truths in complex texture and low-light scenarios, thereby providing more representative and diverse industrial stereo matching data.

[0139] In addition, this solution uses 3D scanning and camera pose estimation, which can accurately obtain the 3D model and generate disparity ground truths, reduce camera calibration errors, improve the accuracy of the data, and provide reliable training data and evaluation data for deep learning algorithms.

[0140] Corresponding to the dataset construction method provided in the above embodiments of the present invention, refer to Figure 7 , the embodiments of the present invention also provide a structural block diagram of a dataset construction device, which includes: an acquisition unit 100, a determination unit 200, a projection unit 300, and a filtering unit 400;

[0141] The acquisition unit 100 is used to acquire the 3D model and image pairs obtained by 3D scanning the target object.

[0142] The determination unit 200 is used to determine the camera pose using the 3D model and image pairs.

[0143] The projection unit 300 is used to project the points on the 3D model onto the image plane corresponding to the image pairs based on the camera pose to obtain a first disparity map.

[0144] The filtering unit 400 is used to filter the pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching dataset.

[0145] In the embodiments of the present invention, the 3D model and image pairs obtained during the 3D scanning process are used to determine the camera pose, and then a second disparity map for constructing a stereo matching dataset is generated based on the camera pose and projection, reducing the difficulty of dataset construction and improving the accuracy of the constructed dataset.

[0146] Preferably, in combination with Figure 7The content shown is that the image pair includes a left image captured by a left camera and a right image captured by a right camera; the determination unit 200 includes a sampling module, a first processing module, and a second processing module, and the execution principles of each module are as follows:

[0147] The sampling module is used to sample three-dimensional object points from the three-dimensional model.

[0148] The first processing module is used to minimize a preset reprojection error function by using the coordinates of the three-dimensional object points corresponding to the left image on the left image and the camera parameters corresponding to the left image to obtain the camera pose of the left camera.

[0149] The second processing module is used to minimize the reprojection error function by using the coordinates of the three-dimensional object points on the right image and the camera parameters corresponding to the right image to obtain the camera pose of the right camera, and the camera parameters include internal parameters and external parameters.

[0150] Preferably, the determination unit 200 further includes:

[0151] The rectification module is used to perform epipolar rectification on the left image and the right image.

[0152] Preferably, in combination with Figure 7 The content shown, the projection unit 300 includes a conversion module, a first projection module, a second projection module, and a calculation module, and the execution principles of each module are as follows:

[0153] The conversion module is used to respectively convert the true coordinates of the three-dimensional object points to the left camera coordinate system and the right camera coordinate system by using the camera poses of the left camera and the right camera.

[0154] The first projection module is used to project the three-dimensional object points onto the image plane of the left camera based on the internal parameters of the left camera and the true coordinates of the three-dimensional object points converted to the left camera coordinate system to obtain the coordinates of the pixels in the left image.

[0155] The second projection module is used to project the three-dimensional object points onto the image plane of the right camera based on the internal parameters of the right camera and the true coordinates of the three-dimensional object points converted to the right camera coordinate system to obtain the coordinates of the pixels in the right image.

[0156] The calculation module is used to calculate the difference between the coordinates of the corresponding pixels projected by the same three-dimensional object point in the left image and the right image to obtain the first disparity map.

[0157] Preferably, in combination with Figure 7 The content shown, the filtering unit 400 includes a calculation module and a culling module, and the execution principles of each module are as follows:

[0158] The calculation module is used to calculate the depth error and disparity error of each pixel in the first disparity map.

[0159] In a specific implementation, the calculation module is specifically configured to: convert the disparity values of the pixels in the first disparity map into depth values; obtain a depth map using a three-dimensional model, and render the depth map to obtain true depth values; calculate the difference between the depth values of the pixels in the first disparity map and the true depth values to obtain the depth error of each pixel in the first disparity map.

[0160] The calculation module is specifically configured to: perform stereo matching on the image pair to obtain an estimated disparity; calculate the difference between the disparity values of the pixels in the first disparity map and the estimated disparity to obtain the disparity error of each pixel in the first disparity map.

[0161] The elimination module is used to eliminate the pixels in the first disparity map whose depth error exceeds the first threshold and / or whose disparity error exceeds the second threshold, so as to obtain a second disparity map for constructing a stereo matching data set.

[0162] Preferably, an embodiment of the present invention further provides an electronic device, including: a processor and a memory, and the processor and the memory are connected through a communication bus; wherein, the processor is configured to call and execute a program stored in the memory; the memory is configured to store a program, and the program is used to implement the data set construction method provided by the above method embodiment.

[0163] Preferably, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the data set construction method provided by the above method embodiment.

[0164] In summary, an embodiment of the present invention provides a data set construction method, device, electronic device and storage medium, which use the three-dimensional model and image pair obtained during the three-dimensional scanning process to determine the camera pose, and then generate a second disparity map for constructing a stereo matching data set based on the camera pose and projection, reducing the construction difficulty of the data set and improving the accuracy of the constructed data set.

[0165] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to a method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiment. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0166] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0167] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing a data set, characterized in that: The method comprises: Acquire a three-dimensional model and an image pair obtained by three-dimensionally scanning the target object; Determine a camera pose using the three-dimensional model and the image pair; Based on the camera pose, projecting points on the three-dimensional model onto an image plane corresponding to the image pair to obtain a first disparity map; Pixels in the first disparity map are filtered to obtain a second disparity map for constructing a stereo matching data set.

2. The method according to claim 1, characterized in that The image pair includes: a left image taken by a left camera and a right image taken by a right camera; Determining a camera pose using the three-dimensional model and the image pair includes: Sampling three-dimensional object points from the three-dimensional model; Minimize a preset reprojection error function using the coordinates corresponding to the three-dimensional object points on the left image and the camera parameters corresponding to the left image to obtain a camera pose of the left camera; The reprojection error function is minimized by using the coordinates of the three-dimensional object points on the right image and the camera parameters corresponding to the right image to obtain the camera pose of the right camera, wherein the camera parameters include intrinsic parameters and extrinsic parameters.

3. The method according to claim 2, characterized in that Based on the camera pose, projecting the points on the three-dimensional model onto an image plane corresponding to the image pair to obtain a first disparity map, comprising: Using the camera poses of the left camera and the right camera, respectively convert the real coordinates of the three-dimensional object point into a left camera coordinate system and a right camera coordinate system; Based on the intrinsic parameters of the left camera and the real coordinates of the three-dimensional object points converted to the left camera coordinate system, projecting the three-dimensional object points to the image plane of the left camera to obtain the coordinates of the pixels in the left image; Based on the intrinsic parameters of the right camera and the real coordinates of the three-dimensional object points converted to the right camera coordinate system, projecting the three-dimensional object points to the image plane of the right camera to obtain the coordinates of the pixels in the right image; The difference between the coordinates of corresponding pixels in the left image and the right image obtained by projecting the same three-dimensional object point is calculated to obtain a first disparity map.

4. The method according to any one of claims 1 to 3, characterized in that: Filtering pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching data set includes: Calculating a depth error and a disparity error for each pixel in the first disparity map; Pixels whose depth errors exceed a first threshold and / or whose disparity errors exceed a second threshold are removed from the first disparity map to obtain a second disparity map for constructing a stereo matching data set.

5. The method according to claim 4, characterized in that The process of calculating the depth error of each pixel in the first disparity map includes: Converting the disparity value of each pixel in the first disparity map into a depth value; Obtaining a depth map using the three-dimensional model, and rendering the depth map to obtain a true depth value; A difference between a depth value of each pixel in the first disparity map and the true depth value is calculated to obtain a depth error of each pixel in the first disparity map.

6. The method according to claim 4, characterized in that The process of calculating the disparity error of each pixel in the first disparity map includes: performing stereo matching on the image pair to obtain an estimated disparity; The difference between the disparity value of each pixel in the first disparity map and the estimated disparity is calculated to obtain the disparity error of each pixel in the first disparity map.

7. The method according to claim 2, characterized in that After sampling the three-dimensional object points from the three-dimensional model, the method further includes: Epipolar correction is performed on the left image and the right image.

8. A data set construction device, characterized in that: The device comprises: An acquisition unit, used to acquire a three-dimensional model and an image pair obtained by three-dimensionally scanning a target object; A determination unit, configured to determine a camera pose using the three-dimensional model and the image pair; A projection unit, configured to project points on the three-dimensional model onto an image plane corresponding to the image pair based on the camera posture, to obtain a first disparity map; A filtering unit is used to filter the pixels in the first disparity map to obtain a second disparity map for constructing a stereo matching data set.

9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; The memory is used to store a program, and the program is used to implement the data set construction method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for constructing a data set according to any one of claims 1 to 7 is implemented.