A method for creating a 3D database based on three-dimensional reconstruction
Through a three-dimensional reconstruction method, video images are acquired and processed, models are established and adjusted to match the point cloud of the target object, and a high-precision 3D database is generated, which solves the problems of narrow application scope and low accuracy of databases in the prior art, and achieves a widely applicable high-precision database establishment.
Patent Information
- Application Number
- CN202111257311.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-10-27
AI Technical Summary
It is difficult to efficiently produce widely applicable high-precision 3D databases in the prior art, especially in industrial production. The existing databases cannot fully provide any target object information in real scenarios, and it is difficult to meet the needs of industrial production.
Using a three-dimensional reconstruction method, the camera acquires videos of different angles of the target object, extracts images and performs three-dimensional reconstruction, and obtains the point cloud, camera external parameters and camera internal parameters of the image. Then, import the bounding box model and the object equal-scale model, adjust its parameters and poses, make it coincide with the point cloud of the target object, read its three-dimensional coordinates and generate a mask diagram, and finally enter the relevant data into the database.
It realizes the establishment of a high-precision database for any object, with a wide range of application and high accuracy, and can meet the needs of any target object information in industrial production.
Smart Images

Figure CN114022542B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a method for manufacturing a 3D database based on three-dimensional reconstruction. Background Art
[0002] With the development of industry, vision systems play a crucial role in industrial production environments. In such environments, robots or robotic arms are often required to automatically identify and accurately position target parts and perform operations such as grasping, part welding, and painting. In recent years, vision-based pose estimation has also been on the rise, especially pose estimation methods based on deep learning, which perform better than traditional methods in terms of both accuracy and speed. Such technologies often require a large amount of data, mainly including images and the pixel coordinates of the target object corresponding to each image. Different from traditional 2D annotation, the target data annotation for pose estimation refers to the coordinates of the vertices of the cube frame enclosing the object in the image, that is, the coordinates of at least eight spatial points.
[0003] In the actual industrial inspection field, producing a large amount of high-quality training data is time-consuming and costly. Especially when calculating the true pose of an object by manual measurement, this method often uses sensors in cooperation, requires a large amount of manpower and material resources, and there are certain differences between the obtained pose and the true pose. Using 3D software to render the target object to obtain its corresponding true values such as pose, the pose obtained by this method is relatively accurate, but there are differences between the objects in the virtual scene and the objects in the real scene, and it is difficult to imitate the real application scenario.
[0004] In addition, existing public databases such as the Rigid Pose database cover very limited objects and cannot fully provide information on any target object in the real scene, making it difficult to meet the needs of industrial production. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method for manufacturing a 3D database based on three-dimensional reconstruction.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A method for manufacturing a 3D database based on three-dimensional reconstruction includes the following steps:
[0008] S1. Use a camera to obtain videos of the target object from different angles, extract images frame by frame, and establish a bounding box model and an object proportional model according to the actual size of the target object;
[0009] S2. Perform three-dimensional reconstruction on the images to obtain the point cloud, external camera parameters, and internal camera parameters of the images;
[0010] S3. Import the bounding box model and the object proportional model obtained in S1 into the point cloud of the image. Change the parameters of the bounding box model so that it exactly encloses the point cloud of the target object, and read the three-dimensional coordinates of the eight vertices of the bounding box model at this time as the initial three-dimensional coordinates.
[0011] S4. Change the size and pose of the object proportional model so that it coincides with the point cloud of the target object, and read the surface coordinates of the object proportional model at this time.
[0012] S5. According to the camera model function, combine the initial three-dimensional coordinates, the model surface coordinates, the camera extrinsic parameters, and the camera intrinsic parameters to obtain the pixel coordinates and generate a mask image.
[0013] S6. Enter the initial three-dimensional coordinates, the mask image, the object pixel coordinates, and the type of the target object into the database.
[0014] Furthermore, the steps for obtaining the point cloud of the image include: obtaining a sparse point cloud, and obtaining a dense point cloud based on the sparse point cloud as the point cloud of the image.
[0015] Furthermore, the steps for obtaining the sparse point cloud are as follows:
[0016] A1. Extract the feature point coordinates of all images, match the feature points between all images, and use the random sample consensus method to remove the incorrect matching pairs.
[0017] A2. Use the constraint matrix to obtain the matching feature points with the largest camera baseline as the maximum image pair.
[0018] A3. According to the coordinates of the maximum image pair, use the random sample consensus eight-point method to calculate the essential matrix.
[0019] A4. Decompose the essential matrix to obtain the camera pose.
[0020] A5. According to the feature point coordinates and the camera pose, calculate the three-dimensional point coordinates. Substitute the three-dimensional point coordinates and the camera pose into the error projection equation, and use the bundle adjustment method to optimize the three-dimensional point coordinates and the camera pose. Combine the optimized three-dimensional point coordinates to obtain the sparse point cloud, and set the optimized camera pose as the camera extrinsic parameters.
[0021] Furthermore, the constraint matrix used in step A2 is expressed as follows:
[0022]
[0023] where F represents the constraint matrix, and x', y', z' and x, y, z represent the feature point coordinates of two frames of images.
[0024] Furthermore, the steps for obtaining the dense point cloud are as follows:
[0025] Using a multi-view stereo vision production system, the external camera parameters and the overall image obtained from the sparse point cloud are input to obtain a dense point cloud.
[0026] Furthermore, the categories of feature points include FAST corner points, SIFT corner points, ORB corner points, and Harris corner points.
[0027] Furthermore, the error projection equation is expressed as follows:
[0028]
[0029] In the formula, g(C, X) represents minimizing the reprojection equation, and its parameters represent all three-dimensional point coordinates X to be optimized and all camera poses C. n represents the number of frames of images selected, m represents the number of three-dimensional points in the sparse point cloud, q ij represents the pixel coordinates of the feature point corresponding to the jth three-dimensional point in the ith frame of the image, and P(C i , X j ) represents the projection coordinates of the jth three-dimensional point X j combined with the ith camera pose C i in the ith frame of the image, where when X j has a projection in the ith frame of the image, ω ij = 1, otherwise ω ij = 0.
[0030] Furthermore, the ICP matching algorithm is used to change the pose of the object's isometric model.
[0031] Furthermore, the camera is a monocular camera.
[0032] Furthermore, the camera model function is expressed as follows:
[0033]
[0034] In the formula, K represents the camera internal parameters, R and T represent the camera external parameters, and Z c represents the z-direction coordinate of K(RP W + T), and P W represents the surface coordinate or the boundary box vertex coordinate;
[0035] When Pw represents the boundary box vertex coordinate, the calculated u and v represent the object pixel coordinates;
[0036] When Pw represents the surface coordinate, the calculated u and v represent all the coordinates of the object mask image.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] 1. The present invention extracts images of objects in real scenes, builds models, and performs three-dimensional reconstruction, and obtains database-related information such as pixel coordinates through calculations of various functions. This method can establish databases for any object, with a wide range of applications and high accuracy.
[0039] 2. When establishing the point cloud, the present invention first establishes a sparse point cloud based on image features and then establishes a dense point cloud, further ensuring the accuracy of image information extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic flow diagram of the present invention.
[0041] Figure 2 It is a schematic diagram of the target image obtained by the present invention.
[0042] Figure 3 It is a schematic diagram of the border model created by the present invention.
[0043] Figure 4 It is a schematic diagram of the equal-proportion model of the object created by the present invention.
[0044] Figure 5a It is a schematic diagram of extracting feature points from a certain frame of image of the present invention
[0045] Figure 5b It is a schematic diagram of extracting feature points from another frame of image of the present invention.
[0046] Figure 6 It is a schematic diagram of feature point matching of the present invention.
[0047] Figure 7 It is a schematic diagram of the matching after using the random sample consensus method of the present invention.
[0048] Figure 8 It is a schematic diagram of the sparse point cloud of the present invention.
[0049] Figure 9 It is a schematic diagram of the dense point cloud of the present invention.
[0050] Figure 10 It is a schematic diagram of the border model including the point cloud of the target object.
[0051] Figure 11 It is a schematic diagram of the coincidence of the equal-proportion model of the object and the point cloud of the target object.
[0052] Figure 12 It is a schematic diagram of pixel coordinate annotation of the present invention.
[0053] Figure 13 It is a schematic diagram of the mask map generated by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0054] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0055] This embodiment provides a method for creating a 3D database based on three-dimensional reconstruction, and the process is as Figure 1 shown, which specifically includes the following steps:
[0056] Step S1: Use a camera as a visual sensor to shoot a video with a target object. The camera can be a monocular camera, a binocular camera or a depth camera. Since in the three-dimensional reconstruction steps used in the present invention, only unordered pictures need to be input and the external and internal parameters of the camera do not need to be provided in advance, a monocular camera is preferably selected when selecting pictures, which is simple to operate and has a relatively fast processing speed. After shooting the video, extract the images of the video, which can be obtained by single-frame shooting or by intercepting the video stream frame by frame. Try to make the overlap area between every two frames of images greater than 30%, the rotation angle of the images is between 30 degrees and 45 degrees, and each point of the object can be observed by at least three frames of images. The collected images are as Figure 1 shown.
[0057] In addition, it is also necessary to prepare a model for the target object. According to the actual parameters of the target object, use 3D software to make a suitable hollow cuboid model as the bounding box, and make a proportional model for generating the mask map. The bounding box model and the object proportional model are respectively as Figure 3 and Figure 4 shown.
[0058] Step S2: Perform three-dimensional reconstruction on the images. In computer vision, three-dimensional reconstruction refers to obtaining a three-dimensional model of the environment or object through a specific process based on a series of photos of the environment or object from different perspectives. The specific steps of three-dimensional reconstruction can be as follows:
[0059] First, extract the feature points of the images. The purpose of feature extraction is to describe an image with a small amount of information, and then be able to estimate the movement of the camera as accurately and stably as possible. The categories of the extracted feature points can be FAST, SIFT, ORB, Harris corners, etc. The feature extraction results of the first frame and the second frame are as Figure 5a and Figure 5b shown.
[0060] Then, match the feature points extracted from all the images. The matching relationship diagram is as Figure 6 shown. Feature matching is to find the corresponding relationship between the feature points of two pictures according to the similarity of the feature points. After matching, it is necessary to remove the wrong matching pairs by the random sample consensus method. The matching relationship diagram after removal is asFigure 7 as shown
[0061] After obtaining the initial matching relationship, a geometric constraint matrix needs to be added, and this geometric constraint completely depends on the objective facts in the scene. The pixel coordinates (x, y), (x', y') between two matched frames of images can be associated through the fundamental matrix F, and the pixel coordinates of the matching pairs that meet the conditions need to satisfy the following formula:
[0062]
[0063] After the constraint, obtain the matching feature points with the largest camera baseline as the largest image pair, and calculate the intrinsic matrix using the random sample consensus eight-point method according to the pixel coordinates of the largest image pair.
[0064] After obtaining the intrinsic matrix, decompose it to obtain the camera pose R and T, where R represents the rotation information and T represents the displacement.
[0065] Perform distortion correction on the image. According to the corrected feature point coordinates and camera pose, calculate the three-dimensional point coordinates. Substitute the three-dimensional point coordinates and camera pose into the error projection equation, and the expression of the objective optimization equation for error projection is as follows:
[0066]
[0067] In the formula, g(C, X) is the minimized reprojection equation, and its parameters are all three-dimensional points X to be optimized and all camera poses C. n represents the number of frames of images selected, m represents the number of three-dimensional points in the sparse point cloud, q ij represents the pixel coordinates of the feature point corresponding to the jth three-dimensional point in the ith frame of the image, and P(C i , X j ) represents the projection coordinates of the jth three-dimensional point X j combined with the ith camera pose C i in the ith frame of the image, where when X j has a projection in the ith frame, ω ij = 1, otherwise ω ij = 0. Use the bundle adjustment method to optimize the three-dimensional point coordinates and camera pose. Combine the optimized three-dimensional point coordinates to obtain a sparse point cloud. The optimized camera pose R and T are the external parameters of the camera. As Figure 8 shown, where the triangular shape at the top is a schematic diagram of the external parameters of the camera.
[0068] After obtaining the sparse point cloud, use the multi-view stereo vision production system. Input the external parameters of the camera and the overall image obtained from the sparse point cloud to obtain a dense point cloud as the point cloud of the image, as Figure 9 shown.
[0069] Step S3: Convert the bounding box model and the object proportional model obtained in Step S1 into the STL format, and import them into the point cloud using point cloud auxiliary tools such as cloudcompare. Adjust the size and pose of the bounding box model so that it just contains the target object, as Figure 10 shown, and read the coordinates of the eight fixed points of the bounding box at this time as the initial three-dimensional coordinates of the target object.
[0070] Step S4: Change the size of the object proportional model so that it is the same as the size of the point cloud of the target object, and change the pose of the object proportional model through the ICP matching algorithm so that it coincides with the point cloud of the target object. The adjusted model is as Figure 11 shown. The surface represented by the model in STL format is composed of a number of closed and connected triangles. Read the vertex coordinates of all triangles as the surface coordinates, and the set of points inside all triangles represents the surface of the target object.
[0071] Step S5: According to the camera model function, the expression is as follows:
[0072]
[0073] In the formula, K represents the camera internal parameter, R and T represent the camera external parameters, P W can represent the initial three-dimensional coordinates or the surface coordinates, Z c represents the coordinate calculation in the z direction of K(RP W +T); when P W represents the initial three-dimensional coordinates, the obtained u and v represent the object pixel coordinates, and the annotation result of the object pixel coordinates is as Figure 12 shown.
[0074] When P W represents the surface coordinates, the obtained u and v represent the coordinates of the object mask map. Adjust the pixels of the object mask map coordinates and the pixels of the remaining coordinates in the image respectively, so as to generate the mask map, as Figure 13 shown.
[0075] Step S6: Enter the initial three-dimensional coordinates, mask map, object pixel coordinates and the type of the target object into the database. Thus, the production of the database is completed.
[0076] In this embodiment, by performing image extraction, model establishment and three-dimensional reconstruction on the objects in the real scene, and calculating through various-level functions to obtain database-related information such as pixel coordinates, this method can establish a database for any object, with a wide application range and high accuracy.
[0077] This embodiment further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for making a 3D database based on three-dimensional reconstruction as mentioned in the embodiments of the present invention. Any combination of one or more computer-readable media can be adopted. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device.
[0078] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A method for making a 3D database based on three-dimensional reconstruction, characterized in that, It includes the following steps: S1. Use a camera to obtain videos of the target object from different angles, extract images frame by frame, and establish a bounding box model and an object proportional model according to the actual size of the target object; S2. Perform 3D reconstruction on the images to obtain the point cloud, external camera parameters, and internal camera parameters of the images; The method for obtaining the said point cloud includes: obtaining a sparse point cloud, and obtaining a dense point cloud based on the sparse point cloud as the point cloud of the image, wherein the steps for obtaining the said sparse point cloud include: A1. Extract the coordinate of feature points of all images, match the feature points between all images, and use the random sample consensus method to remove incorrect matching pairs; A2. Use a constraint matrix to obtain the matched feature points with the maximum camera baseline as the maximum image pair; A3. Calculate the essential matrix using the random sample consensus eight-point method according to the coordinates of the maximum image pair; A4. Decompose the said essential matrix to obtain the camera pose; A5. Calculate the 3D point coordinates based on the feature point coordinates and the camera pose, substitute the 3D point coordinates and the camera pose into the error projection equation, optimize the 3D point coordinates and the camera pose using the bundle adjustment method, combine the optimized 3D point coordinates to obtain a sparse point cloud, and set the optimized camera pose as the external camera parameters; S3. Import the bounding box model and the object proportional model obtained in S1 into the point cloud of the image, change the parameters of the bounding box model so that it exactly contains the point cloud of the target object, and read the 3D coordinates of the eight vertices of the bounding box model at this time as the initial 3D coordinates; S4. Change the size and pose of the object proportional model to make it coincide with the point cloud of the target object, and read the surface coordinates of the object proportional model at this time; S5. According to the camera model function, combine the initial 3D coordinates, the model surface coordinates, the external camera parameters, and the internal camera parameters to obtain the pixel coordinates and generate a mask image; S6. Enter the initial 3D coordinates, the mask image, the object pixel coordinates, and the type of the target object into the database.
2. The method for manufacturing a 3D database based on three-dimensional reconstruction according to claim 1, wherein, The expression of the constraint matrix in step A2 is as follows: where F represents the constraint matrix, and x', y', z' and x, y, z represent the coordinate of feature points of two frames of images.
3. A method for creating a 3D database based on three-dimensional reconstruction according to claim 1, characterized in that, The steps for obtaining the said dense point cloud are as follows: Use a multi-view stereo vision production system to input the external camera parameters and the overall images obtained from the sparse point cloud to obtain a dense point cloud.
4. A method for making a 3D database based on three-dimensional reconstruction according to claim 1, characterized in that, The categories of feature points include FAST corner points, SIFT corner points, ORB corner points, and Harris corner points.
5. A method for making a 3D database based on three-dimensional reconstruction according to claim 1, characterized in that The expression of the error projection equation is as follows: In the formula, g(C, X) represents the minimized reprojection equation, where its parameters represent all the three-dimensional point coordinates X to be optimized and all the camera poses C, n represents the number of selected image frames, m represents the number of three-dimensional points in the sparse point cloud, and q ij represents the pixel coordinates of the feature point corresponding to the j-th three-dimensional point in the i-th image frame, P(C i , X j ) represents the projection coordinates of the j-th three-dimensional point X j combined with the i-th camera pose C i in the i-th image frame, where when X j has a projection in the i-th image frame, ω ij = 1, otherwise ω ij = 0.
6. A method for creating a 3D database based on three-dimensional reconstruction according to claim 1, characterized in that Use the ICP matching algorithm to change the pose of the object proportional model.
7. A method for manufacturing a 3D database based on three-dimensional reconstruction according to claim 1, characterized in that, The said camera is a monocular camera.
8. A method for creating a 3D database based on three-dimensional reconstruction according to claim 1, characterized in that The expression of the said camera model function is as follows: where K represents the camera internal parameters, R and T represent the camera external parameters, and Z c represents the z - coordinate of K(RP W +T), and P W represents the surface coordinates or the coordinates of the bounding box vertices; When Pw represents the coordinate of the bounding box vertex, the calculated u and v represent the object pixel coordinates; When Pw represents the surface coordinates, the calculated u and v represent all the coordinates of the object mask image.