Three-dimensional modeling method and device based on visual fusion, terminal and storage medium
Through the multi-binar camera system and spatial transformation fusion algorithm, the low accuracy problems caused by occlusion and lighting changes in three-dimensional modeling are solved, and a more efficient and accurate three-dimensional modeling effect is achieved.
Patent Information
- Application Number
- CN202411993065.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing three-dimensional modeling technology leads to low modeling accuracy when dealing with occlusion and illumination changes between objects.
A multi-binar camera system based on visual fusion is used to shoot target objects from multiple angles after internal and external parameter calibration, and the parallax map is fused through a spatial transformation fusion algorithm to calculate point cloud data and build a three-dimensional model.
It improves the integrity and accuracy of three-dimensional modeling in complex scenes and object structures, and reduces the algorithm complexity and modeling cost.
Smart Images

Figure CN119919582A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer vision and three-dimensional modeling, and in particular to a three-dimensional modeling method, device, terminal and storage medium based on vision fusion. Background Art
[0002] With the continuous development of computer technology, 3D reconstruction technology has received more and more attention. Among them, 3D reconstruction technology based on visual information is a relatively common method. As an important field of computer vision, the realization of non-contact measurement, recognition and rapid modeling technology is an important development direction. This technology provides strong support for many fields such as medicine, industry, virtual reality, augmented reality, cultural relics protection, and security monitoring by restoring the 3D shape and structure information of objects from 2D images or video sequences. In the current field of 3D modeling technology, traditional 3D modeling methods have many defects. When a monocular camera performs 3D modeling, due to the lack of direct means of obtaining depth information, complex algorithms and a large amount of computing resources are often required to estimate the depth, and the accuracy is limited. Although some high-precision 3D scanning equipment, such as laser scanners, can obtain accurate depth information, the equipment is expensive and large in size, which is not suitable for large-scale promotion and some cost-sensitive application scenarios. In addition, in the actual 3D modeling process, the problem of visual occlusion has always been one of the key factors that plague the development of technology. Whether in complex object modeling or large scene modeling, the mutual occlusion between objects and the invisible areas caused by the object's own structure will make the acquired image information incomplete, which will seriously affect the accuracy and integrity of the 3D model. At the same time, changes in lighting conditions will also greatly interfere with modeling. Different lighting intensities, directions, and uneven lighting will cause changes in the reflective properties of the object surface, which in turn affects the quality of the image and the accuracy of feature extraction, ultimately reducing the accuracy and reliability of 3D modeling. Summary of the invention
[0003] The present application provides a three-dimensional modeling method, device, terminal and storage medium based on visual fusion to solve the problem of low three-dimensional modeling accuracy of objects due to occlusion or different lighting between objects in the prior art.
[0004] In a first aspect, the present application provides a three-dimensional modeling method based on visual fusion, comprising:
[0005] Using a multi-binocular camera system calibrated with internal and external parameters to shoot a target object at multiple angles to obtain a disparity map of the target object shot by each binocular camera, the multi-binocular camera system includes a plurality of binocular cameras, each binocular camera includes a left camera and a right camera;
[0006] Using a spatial transformation fusion algorithm, each disparity map is fused to obtain a complete disparity map of the target object;
[0007] The complete disparity map of the target object is used to calculate point cloud data, and a visualized three-dimensional model of the target object is constructed based on the point cloud data.
[0008] In a second aspect, the present application provides a three-dimensional modeling device based on visual fusion, comprising:
[0009] A disparity map acquisition module is used to shoot a target object at multiple angles using a multi-binocular camera system calibrated with internal and external parameters to obtain a disparity map of the target object shot by each binocular camera, wherein the multi-binocular camera system includes a plurality of binocular cameras, and each binocular camera includes a left camera and a right camera;
[0010] A disparity map fusion module is used to fuse each disparity map using a spatial transformation fusion algorithm to obtain a complete disparity map of the target object;
[0011] The three-dimensional model construction module is used to calculate point cloud data using the complete disparity map of the target object, and to construct a visual three-dimensional model of the target object based on the point cloud data.
[0012] In a third aspect, the present application provides a terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method described in the first aspect or any possible implementation of the first aspect are implemented.
[0013] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect or any possible implementation of the first aspect.
[0014] The present application provides a 3D modeling method, device, terminal and storage medium based on visual fusion. The method comprises: using a multi-binocular camera system after internal and external parameter calibration to shoot a target object from multiple angles to obtain a disparity map of the target object shot by each binocular camera. The multi-binocular camera system comprises a plurality of binocular cameras, and each binocular camera comprises a left camera and a right camera. Using a spatial transformation fusion algorithm, each disparity map is fused to obtain a complete disparity map of the target object. Using the complete disparity map of the target object, point cloud data is calculated, and based on the point cloud data, a visual 3D model of the target object is constructed. This application uses multiple binocular cameras to shoot the target object from different angles, and can obtain more information about the target object. Even if there is occlusion in some areas, it can be supplemented by image information from other cameras to obtain a complete 3D modeling result, which greatly improves the integrity and accuracy of the visualized 3D model in complex scenes and object structures; at the same time, while ensuring the modeling accuracy, it reduces the complexity of the algorithm and improves the modeling efficiency; in addition, this application uses multiple binocular cameras as core equipment. Binocular cameras have the characteristics of low cost, small size and easy integration in the market. Through reasonable layout and efficient algorithm design, it can greatly reduce the cost of the entire 3D modeling system while ensuring the modeling accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0016] Figure 1 It is a flowchart of the implementation of the three-dimensional modeling method based on visual fusion provided in the embodiment of the present application;
[0017] Figure 2 is a structural schematic diagram of a three-dimensional modeling device based on visual fusion provided in an embodiment of the present application;
[0018] Figure 3 It is a schematic diagram of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0020] In order to make the purpose, technical solutions and advantages of the present application clearer, specific embodiments will be described below in conjunction with the accompanying drawings.
[0021] Figure 1 The implementation flow chart of the 3D modeling method based on visual fusion provided in the embodiment of the present application is described in detail as follows:
[0022] In step 101, a multi-binocular camera system after internal and external parameter calibration is used to shoot a target object at multiple angles to obtain a disparity map of the target object shot by each binocular camera. The multi-binocular camera system includes multiple binocular cameras, and each binocular camera includes a left camera and a right camera.
[0023] In an embodiment of the present application, a calibrated multi-binocular camera system is used to shoot a target object at multiple angles to obtain a disparity map of the target object shot by each binocular camera.
[0024] Among them, the multi-binocular camera is a system composed of multiple binocular cameras.
[0025] In a possible implementation, before using the multi-binocular camera system after internal and external parameter calibration to shoot the target object at multiple angles and obtain the disparity map of the target object shot by each binocular camera, the method may further include:
[0026] Arrange two adjacent binocular cameras at a fixed distance and a preset angle;
[0027] Use the checkerboard calibration plate to calibrate the intrinsic parameters of each binocular camera;
[0028] After the intrinsic parameter calibration of each binocular camera is completed, the extrinsic parameters of each binocular camera are calibrated separately.
[0029] Optionally, before shooting, the application also needs to calibrate the internal and external parameters of the multi-binocular camera system, that is, using a multi-binocular camera system, multiple binocular cameras are arranged at different positions and angles to perform stereo calibration of the internal and external parameters. The specific calibration process is as follows:
[0030] First, two adjacent binocular cameras are arranged at a fixed distance and a preset angle to construct a multi-binocular camera system. Then, the intrinsic parameters of each binocular camera in the multi-binocular camera system are calibrated using a checkerboard calibration plate. Finally, the extrinsic parameters of each binocular camera with the intrinsic parameter calibrated are calibrated.
[0031] For internal parameter calibration, the embodiment of the present application uses a checkerboard calibration plate to calibrate the internal parameters of each binocular camera.
[0032] First, the distortion correction formula is used to calculate the distortion coefficient of each binocular camera, where the distortion coefficient includes a first distortion coefficient (i.e., the radial distortion coefficient of the left and right cameras) and a second distortion coefficient.
[0033] The distortion correction formula is:
[0034]
[0035] Among them, x is the horizontal coordinate after correction, y is the vertical coordinate after correction, and r is the distance between the point and the coordinate. is the horizontal coordinate between the distortion correction, is the vertical coordinate between the distortion correction, k1, k2, k3 are the first distortion coefficients, and p1, p2 are the second distortion coefficients.
[0036] Then, a checkerboard calibration board is used to collect image information at different positions and angles, and constraints are constructed through the corner points of the calibration board. Finally, the camera's intrinsic parameter matrix K is estimated by the least squares method, that is;
[0037]
[0038] Among them, K is the internal parameter matrix, f x and f y is the focal length parameter, c x and c y is the principal point coordinate, and s is the coordinate axis tilt parameter.
[0039] For external parameter calibration, the external parameter calibration of each binocular camera refers to the relative position and posture of the left and right cameras, that is, the rotation matrix R and the translation vector T. The specific steps are:
[0040] First, the corner points of the images taken by the left and right cameras are matched to ensure that the positions of each pair of matching points in the two images taken by the left and right cameras are corresponding.
[0041] The rotation matrix R and the translation vector T require the use of the basic matrix F and the essential matrix E. The basic matrix F describes the geometric relationship between the two camera coordinate systems. Based on the matching point pairs in the image, the basic matrix F can be estimated by the eight-point method. The relationship between the essential matrix E and the basic matrix F is:
[0042]
[0043] Among them, F is the basic matrix, E is the essential matrix, K r is the intrinsic parameter matrix of the right camera, K l is the intrinsic parameter matrix of the left camera.
[0044] Because we know the basic matrix F, we can use the formula Get the essential matrix E.
[0045] After obtaining the essential matrix E, the embodiment of the present application obtains four solutions (combinations of two rotation matrices and translation vectors) through SVD decomposition, and finally selects the correct solution through the spatial position of the matching point, thereby obtaining the external parameter calibration result, namely the rotation matrix R and the translation vector T.
[0046] The external parameter calibration process for multiple binocular cameras is similar to that for a single binocular camera. The goal is to determine the relative position and posture between multiple binocular cameras, that is, the system's rotation matrix R and translation vector T. The simplified process is to use multiple groups of binocular cameras to shoot multiple groups of calibration images of the same scene, extract common corner points or feature points, and use the essential matrix and basic matrix to solve the rotation matrix and translation vector between the systems.
[0047] In a possible implementation, a multi-stereo camera system after internal and external parameter calibration is used to shoot a target object at multiple angles to obtain a disparity map of the target object shot by each binocular camera, which may include:
[0048] For each binocular camera, perform the following steps:
[0049] Obtain a left image taken by a left camera of the binocular camera, obtain a right image taken by a right camera of the binocular camera, and perform Gaussian filtering on the left image and the right image;
[0050] The scale-invariant feature transformation algorithm and the accelerated robust feature algorithm are used to extract feature points on the left image and the right image respectively, and a feature descriptor is generated for each feature point;
[0051] Calculate the Euclidean distance between the feature descriptor on the left image and the feature descriptor on the right image, and determine the matching feature descriptors on the left image and the right image based on the minimum distance;
[0052] Using the matching feature descriptors on the left image and the right image, calculate the disparity of the feature points corresponding to all the matching feature descriptors;
[0053] Based on each disparity, a disparity map of the target object photographed by the binocular camera is obtained.
[0054] Among them, the scale-invariant feature transform (SIFT) algorithm is a machine vision algorithm used to detect and describe local features in images. It searches for extreme points in the spatial scale and extracts their position, scale, and rotation invariants.
[0055] Speeded Up Robust Features (SURF) is a robust image recognition and description algorithm that can be used for computer vision tasks such as object recognition and 3D reconstruction.
[0056] The embodiment of the present application uses the principle of binocular stereo vision to calculate the disparity map of each binocular camera to the target object. The process includes: first selecting feature points, then performing stereo matching, and generating the core disparity map of three-dimensional modeling through the stereo matching algorithm. In the generated disparity map, each pixel can be calculated by triangulation to calculate its distance in three-dimensional space (i.e., depth information). For each binocular camera, the specific disparity map calculation process is as follows:
[0057] The left image captured by the left camera of the binocular camera is obtained, and the right image captured by the right camera of the binocular camera is obtained. Then, the left image and the right image are preprocessed by Gaussian filtering to reduce the influence of noise on feature extraction.
[0058] Then, feature selection is performed. In terms of feature extraction, SIFT and SURF detection algorithms are combined to find feature points. The fast detection advantage of SURF is used to preliminarily locate possible feature point areas on each image. Then, the SIFT algorithm is used in these areas for more refined extreme point detection and screening. This ensures that high-quality feature points can be found while improving the efficiency of the entire feature point extraction process. After using SURF for preliminary positioning and SIFT for fine detection to screen out feature points, SIFT calculates the gradient direction histogram in the neighborhood of the feature point to form a feature descriptor. This histogram can capture the directional distribution information of pixels around the feature point, making the feature point scale and rotation invariant, which makes it more stable under different lighting and partial occlusion.
[0059] After obtaining the feature descriptor, feature matching is performed. The embodiment of the present application adopts a distance measurement method, namely, the Euclidean distance, to measure the similarity between two feature descriptors. The smaller the distance, the more the feature points match, and the minimum distance is selected. Because there are many matching points, the nearest neighbor ratio method is used to eliminate false matches. This matching process can help us find the corresponding relationship between the two images.
[0060] After the matching is completed, the disparity of the feature points corresponding to all matching feature descriptors is calculated to finally obtain the disparity map of the target object captured by each binocular camera.
[0061] The embodiment of the present application adopts a hybrid feature extraction algorithm of SIFT and SURF, which can more stably extract common feature points of the images generated by each binocular camera under different lighting and partial occlusion conditions, so as to match the left and right images of each binocular camera to produce a more accurate disparity map.
[0062] In addition, in the matching process, the embodiments of the present application fully consider the spatial position relationship and local geometric structure of the feature points, and use a matching optimization algorithm based on graph theory to reduce the occurrence of false matches and improve matching accuracy.
[0063] In step 102, each disparity map is fused using a spatial transformation fusion algorithm to obtain a complete disparity map of the target object.
[0064] In the embodiment of the present application, the spatial transformation fusion algorithm may be a weighted average method or an optimal matching algorithm.
[0065] In the embodiment of the present application, each disparity map calculated in step 101 is fused using a spatial transformation fusion algorithm to obtain a complete disparity map of the captured target object.
[0066] In a possible implementation, each disparity map is fused using a spatial transformation fusion algorithm to obtain a complete disparity map of the target object, which may include:
[0067] Using the least squares method, the geometric transformation matrix between each two disparity images is calculated respectively;
[0068] Using each geometric transformation matrix, transform the pixel coordinates of the two disparity images of the corresponding geometric transformation matrix into the world coordinate system;
[0069] Using the preset fusion weights of each binocular camera, each disparity map transformed into the world coordinate system is fused to obtain a complete disparity map of the target object.
[0070] Optionally, the purpose of the embodiment of the present application is to fuse each disparity map into a unified three-dimensional coordinate system. First, image registration is performed, and the corresponding geometric transformation matrix (including translation, rotation, and scaling) is estimated using all matching feature points of each binocular camera obtained in the feature extraction stage.
[0071] Taking two binocular cameras as an example, the embodiment of the present application uses the least squares method to solve the geometric transformation matrix. Assuming the coordinates of the feature point in the first disparity map are (x1, y1), and the coordinates of the corresponding matching feature point in the second disparity map are (x2, y2), the geometric transformation matrix A can be expressed as:
[0072]
[0073] Among them, A is the geometric transformation matrix, that is, a 11 、a 12 、a 13 、a 21 、a 22 、a 23are the parameters of the geometric transformation matrix, x1 is the horizontal coordinate of the feature point in the first disparity image, y1 is the vertical coordinate of the feature point in the first disparity image, x2 is the horizontal coordinate of the feature point in the second disparity image, and y2 is the vertical coordinate of the feature point in the second disparity image.
[0074] By using enough matching points, the parameters of the geometric transformation matrix A can be solved to determine the transformation relationship between the two disparity images.
[0075] Once the geometric transformation matrix is obtained, all pixel coordinates (including feature points and non-feature points) in one of the disparity maps can be transformed according to the geometric transformation matrix so that the two disparity maps are aligned in the same world coordinate system.
[0076] Then the depth information is fused, namely parallax fusion.
[0077] The preset fusion weight of the corresponding stereo camera is determined based on factors such as the position, angle, and image quality of each stereo camera. For example, if the shooting angle of a stereo camera is more perpendicular to the surface of the target object and the image clarity is high, then the disparity map of the stereo camera can be given a higher weight during fusion. The importance of the weight is not limited to this. In other scenarios, factors such as the geometric relationship between the stereo camera and the target object, the resolution and contrast of the image, etc. need to be comprehensively considered.
[0078] Spatial transformation fusion plays a key role in disparity fusion. The embodiment of the present application uses a weighted average method or an optimal matching method to fuse two sets of disparity maps to obtain a complete disparity map.
[0079] It should be noted that if the disparity maps of the two perspectives overlap a lot, the depth map of one of the binocular cameras will be given priority during depth fusion. Otherwise, the best data will be selected based on the error metric (such as the accuracy of the disparity).
[0080] In step 103, the complete disparity map of the target object is used to calculate point cloud data, and a visual three-dimensional model of the target object is constructed based on the point cloud data.
[0081] In the embodiment of the present application, the complete disparity map of the target object obtained by fusion in step 102 is used to calculate point cloud data, and the point cloud data and the positions of the corresponding point cloud data are used to construct a visual three-dimensional model of the target object.
[0082] In a possible implementation, using the complete disparity map of the target object to calculate the point cloud data may include:
[0083] Using the triangulation principle, the three-dimensional coordinates of each pixel on the complete disparity map are calculated respectively;
[0084] Using each set of 3D coordinates, point cloud data is constructed.
[0085] Optionally, the embodiment of the present application utilizes the triangulation principle to convert the disparity value of each pixel on the complete disparity map into a three-dimensional coordinate.
[0086] First, calculate the three-dimensional coordinates. According to the complete disparity map, use the triangulation principle to calculate the three-dimensional coordinates of each pixel on the complete disparity map. The details are as follows:
[0087] Knowing the baseline length B of the binocular camera, the camera angle f, and the disparity d in the complete disparity map, the formula Calculate the depth value Z of the feature point. Then, through the stereo matching structure and the depth value Z of each pixel in the disparity map, the coordinates of the corresponding pixel in three-dimensional space can be calculated. For each pixel point (x, y), the corresponding three-dimensional coordinates (X, Y, z) are calculated by the following formula, namely:
[0088]
[0089] Among them, c x and c y is the principal coordinate of the image (i.e. the origin of the camera coordinate), f x and f y is the camera focal length.
[0090] The complete coordinates of each pixel in three-dimensional space are calculated to form point cloud data.
[0091] In one possible implementation, during the fusion process, for point cloud data with occluded areas, the embodiment of the present application analyzes the point cloud features and geometric relationships of adjacent non-occluded areas, uses the least squares method to fit the point clouds of adjacent non-occluded areas into quadratic surfaces, and uses methods based on surface fitting and interpolation to fill and repair, ultimately completing the visual three-dimensional modeling of the target object.
[0092] The embodiment of the present application processes overlapping areas by adopting a spatial transformation fusion algorithm during parallax fusion, assigning different weights to each point cloud according to factors such as the camera's position, posture, and imaging quality, and fusion of multiple point clouds into a complete, high-precision point cloud model. In the process of filling in non-parallax areas such as occlusions, for data in occluded areas, the point cloud features and geometric relationships of adjacent non-occluded areas are analyzed, and surface fitting and interpolation methods are used to fill and repair, so as to improve the integrity and accuracy of the constructed visual 3D model in the occluded areas.
[0093] The present application provides a 3D modeling method based on visual fusion, which uses a multi-binocular camera system calibrated with internal and external parameters to shoot a target object from multiple angles to obtain a disparity map of the target object shot by each binocular camera, wherein the multi-binocular camera system includes multiple binocular cameras, and each binocular camera includes a left camera and a right camera; uses a spatial transformation fusion algorithm to fuse each disparity map to obtain a complete disparity map of the target object; uses the complete disparity map of the target object to calculate point cloud data, and constructs a visual 3D model of the target object based on the point cloud data. This application uses multiple binocular cameras to shoot the target object from different angles, and can obtain more information about the target object. Even if there is occlusion in some areas, it can be supplemented by image information from other cameras to obtain a complete 3D modeling result, which greatly improves the integrity and accuracy of the visualized 3D model in complex scenes and object structures; at the same time, while ensuring the modeling accuracy, it reduces the complexity of the algorithm and improves the modeling efficiency; in addition, this application uses multiple binocular cameras as core equipment. Binocular cameras have the characteristics of low cost, small size and easy integration in the market. Through reasonable layout and efficient algorithm design, it can greatly reduce the cost of the entire 3D modeling system while ensuring the modeling accuracy.
[0094] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0095] The following is an embodiment of the device of the present application. For details not described in detail, please refer to the corresponding method embodiment described above.
[0096] Figure 2 The following is a schematic diagram of the structure of a 3D modeling device based on visual fusion provided in an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown, which are described in detail as follows:
[0097] like Figure 2 As shown, the three-dimensional modeling device 2 based on visual fusion includes:
[0098] The disparity map acquisition module 21 is used to shoot the target object at multiple angles using the multi-binocular camera system after internal and external parameter calibration to obtain the disparity map of the target object shot by each binocular camera, wherein the multi-binocular camera system includes multiple binocular cameras, and each binocular camera includes a left camera and a right camera;
[0099] The disparity map fusion module 22 is used to fuse each disparity map using a spatial transformation fusion algorithm to obtain a complete disparity map of the target object;
[0100] The three-dimensional model construction module 23 is used to calculate point cloud data using the complete disparity map of the target object, and to construct a visual three-dimensional model of the target object based on the point cloud data.
[0101] The present application provides a 3D modeling device based on visual fusion. The device uses a multi-binocular camera system calibrated with internal and external parameters to shoot a target object from multiple angles to obtain a disparity map of the target object shot by each binocular camera. The multi-binocular camera system includes multiple binocular cameras, and each binocular camera includes a left camera and a right camera. A spatial transformation fusion algorithm is used to fuse each disparity map to obtain a complete disparity map of the target object. The complete disparity map of the target object is used to calculate point cloud data, and a visual 3D model of the target object is constructed based on the point cloud data. This application uses multiple binocular cameras to shoot the target object from different angles, and can obtain more information about the target object. Even if there is occlusion in some areas, it can be supplemented by image information from other cameras to obtain a complete 3D modeling result, which greatly improves the integrity and accuracy of the visualized 3D model in complex scenes and object structures; at the same time, while ensuring the modeling accuracy, it reduces the complexity of the algorithm and improves the modeling efficiency; in addition, this application uses multiple binocular cameras as core equipment. Binocular cameras have the characteristics of low cost, small size and easy integration in the market. Through reasonable layout and efficient algorithm design, it can greatly reduce the cost of the entire 3D modeling system while ensuring the modeling accuracy.
[0102] In a possible implementation, the device may further include a calibration module, which may be used to:
[0103] Arrange two adjacent binocular cameras at a fixed distance and a preset angle;
[0104] Use the checkerboard calibration plate to calibrate the intrinsic parameters of each binocular camera;
[0105] After the intrinsic parameter calibration of each binocular camera is completed, the extrinsic parameters of each binocular camera are calibrated separately.
[0106] In a possible implementation, the internal parameters include distortion coefficients and an internal parameter matrix, and the calibration module can be specifically used for:
[0107] The distortion coefficient of each binocular camera is calculated using the distortion correction formula;
[0108] Each binocular camera is calibrated using a checkerboard calibration plate, and the intrinsic parameter matrix is calculated using the least squares method.
[0109] The distortion correction formula is:
[0110]
[0111] Among them, x is the horizontal coordinate after correction, y is the vertical coordinate after correction, and r is the distance between the point and the coordinate. is the horizontal coordinate between the distortion correction, is the vertical coordinate between the distortion correction, k1, k2, k3 are the first distortion coefficients, and p1, p2 are the second distortion coefficients.
[0112] In a possible implementation, the disparity map acquisition module may be used to:
[0113] For each binocular camera, perform the following steps:
[0114] Obtain a left image taken by a left camera of the binocular camera, obtain a right image taken by a right camera of the binocular camera, and perform Gaussian filtering on the left image and the right image;
[0115] The scale-invariant feature transformation algorithm and the accelerated robust feature algorithm are used to extract feature points on the left image and the right image respectively, and a feature descriptor is generated for each feature point;
[0116] Calculate the Euclidean distance between the feature descriptor on the left image and the feature descriptor on the right image, and determine the matching feature descriptors on the left image and the right image based on the minimum distance;
[0117] Using the matching feature descriptors on the left image and the right image, calculate the disparity of the feature points corresponding to all the matching feature descriptors;
[0118] Based on each disparity, a disparity map of the target object photographed by the binocular camera is obtained.
[0119] In a possible implementation, the disparity map fusion module may be used to:
[0120] Using the least squares method, the geometric transformation matrix between each two disparity images is calculated respectively;
[0121] Using each geometric transformation matrix, transform the pixel coordinates of the two disparity images of the corresponding geometric transformation matrix into the world coordinate system;
[0122] Using the preset fusion weights of each binocular camera, each disparity map transformed into the world coordinate system is fused to obtain a complete disparity map of the target object.
[0123] In a possible implementation, the disparity map fusion module may also be used for:
[0124] The geometric transformation matrix is calculated by the second formula, which is:
[0125]
[0126] Among them, A is the geometric transformation matrix, x1 is the horizontal coordinate of the feature point in the first disparity image, y1 is the vertical coordinate of the feature point in the first disparity image, x2 is the horizontal coordinate of the feature point in the second disparity image, and y2 is the vertical coordinate of the feature point in the second disparity image.
[0127] In a possible implementation, the 3D model building module can be used to:
[0128] Using the triangulation principle, the three-dimensional coordinates of each pixel on the complete disparity map are calculated respectively;
[0129] Using each set of 3D coordinates, point cloud data is constructed.
[0130] Figure 3 is a schematic diagram of a terminal provided in an embodiment of the present application. Figure 3 As shown, the terminal 3 of this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, the steps in the above-mentioned three-dimensional modeling method based on visual fusion are implemented, for example Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are realized, for example, Figure 2 The functions of each module are shown.
[0131] Exemplarily, the computer program 32 may be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 30 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 32 in the terminal 3. For example, the computer program 32 may be divided into Figure 2 The modules shown.
[0132] The terminal 3 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal 3 may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 3 It is only an example of terminal 3 and does not constitute a limitation on terminal 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal may also include input and output devices, network access devices, buses, etc.
[0133] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0134] The memory 31 may be an internal storage unit of the terminal 3, such as a hard disk or memory of the terminal 3. The memory 31 may also be an external storage device of the terminal 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal 3. Further, the memory 31 may also include both an internal storage unit of the terminal 3 and an external storage device. The memory 31 is used to store the computer program and other programs and data required by the terminal. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0135] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0136] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0137] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0138] In the embodiments provided in the present application, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0139] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0141] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can be executed by the processor to implement the steps of the above-mentioned three-dimensional modeling method embodiment based on visual fusion. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable media does not include electrical carrier signals and telecommunication signals.
[0142] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A three-dimensional modeling method based on visual fusion, characterized in that: include: Using a multi-binocular camera system calibrated with internal and external parameters to shoot a target object at multiple angles to obtain a disparity map of the target object shot by each binocular camera, the multi-binocular camera system includes a plurality of binocular cameras, each binocular camera includes a left camera and a right camera; Using a spatial transformation fusion algorithm, each disparity map is fused to obtain a complete disparity map of the target object; The complete disparity map of the target object is used to calculate point cloud data, and a visualized three-dimensional model of the target object is constructed based on the point cloud data.
2. The three-dimensional modeling method based on visual fusion according to claim 1, characterized in that: Before the multi-binocular camera system calibrated with internal and external parameters performs multi-angle shooting on the target object to obtain a disparity map of the target object shot by each binocular camera, the method further includes: Arrange two adjacent binocular cameras at a fixed distance and a preset angle; Use the checkerboard calibration plate to calibrate the intrinsic parameters of each binocular camera; After the intrinsic parameter calibration of each binocular camera is completed, the extrinsic parameters of each binocular camera are calibrated separately.
3. The three-dimensional modeling method based on visual fusion according to claim 2 is characterized in that: The internal parameters include distortion coefficients and internal parameter matrices. The internal parameters of each binocular camera are calibrated using a checkerboard calibration plate, including: The distortion coefficient of each binocular camera is calculated using the distortion correction formula; Each binocular camera is calibrated using a checkerboard calibration plate, and the intrinsic parameter matrix is calculated using the least squares method. Wherein, the distortion correction formula is: Among them, x is the horizontal coordinate after correction, y is the vertical coordinate after correction, and r is the distance between the point and the coordinate. is the horizontal coordinate between the distortion correction, is the vertical coordinate between the distortion correction, k1, k2, k3 are the first distortion coefficients, and p1, p2 are the second distortion coefficients.
4. The three-dimensional modeling method based on visual fusion according to claim 1, characterized in that: The multi-binocular camera system after the internal and external parameters are calibrated is used to shoot the target object at multiple angles to obtain a disparity map of the target object shot by each binocular camera, including: For each stereo camera, perform the following steps: Acquire a left image taken by a left camera of the binocular camera, and acquire a right image taken by a right camera of the binocular camera, and perform Gaussian filtering on the left image and the right image; Using a scale-invariant feature transformation algorithm and an accelerated robust feature algorithm, feature points on the left image and the right image are extracted respectively, and a feature descriptor is generated for each feature point; Calculating the Euclidean distance between the feature descriptor on the left image and the feature descriptor on the right image, and determining the matching feature descriptors on the left image and the right image based on the minimum distance; Using the matched feature descriptors on the left image and the right image, calculating the disparity of feature points corresponding to all matched feature descriptors; Based on each disparity, a disparity map of the target object photographed by the binocular camera is obtained.
5. The three-dimensional modeling method based on visual fusion according to claim 1, characterized in that: The method of fusing each disparity map by using a spatial transformation fusion algorithm to obtain a complete disparity map of the target object includes: Using the least squares method, the geometric transformation matrix between each two disparity images is calculated respectively; Using each geometric transformation matrix, transform the pixel coordinates of the two disparity images of the corresponding geometric transformation matrix into the world coordinate system; The preset fusion weight of each binocular camera is used to fuse the disparity maps transformed into the world coordinate system to obtain a complete disparity map of the target object.
6. The three-dimensional modeling method based on visual fusion according to claim 5, characterized in that: The least square method is used to calculate the geometric transformation matrix between each two disparity images, including: The geometric transformation matrix is calculated by a second formula, wherein the second formula is: Among them, A is the geometric transformation matrix, x1 is the horizontal coordinate of the feature point in the first disparity image, y1 is the vertical coordinate of the feature point in the first disparity image, x2 is the horizontal coordinate of the feature point in the second disparity image, and y2 is the vertical coordinate of the feature point in the second disparity image.
7. The three-dimensional modeling method based on visual fusion according to claim 1, characterized in that: The step of calculating point cloud data by using the complete disparity map of the target object includes: Using the triangulation principle, the three-dimensional coordinates of each pixel point on the complete disparity map are calculated respectively; The point cloud data is constructed using each set of three-dimensional coordinates.
8. A three-dimensional modeling device based on visual fusion, characterized in that: include: A disparity map acquisition module is used to shoot a target object at multiple angles using a multi-binocular camera system calibrated with internal and external parameters to obtain a disparity map of the target object shot by each binocular camera, wherein the multi-binocular camera system includes a plurality of binocular cameras, and each binocular camera includes a left camera and a right camera; A disparity map fusion module is used to fuse each disparity map using a spatial transformation fusion algorithm to obtain a complete disparity map of the target object; The three-dimensional model construction module is used to calculate point cloud data using the complete disparity map of the target object, and to construct a visual three-dimensional model of the target object based on the point cloud data.
9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the three-dimensional modeling method based on visual fusion as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the three-dimensional modeling method based on visual fusion as described in any one of claims 1 to 7 above are implemented.
Citation Information
Cited By
Printed circuit board (PCB) positioning and multifunctional testing integrated device based on visual identification
CN120510108A
Tracking type scanning method and system based on monocular and binocular fusion
CN121353346A
Contour correction calibration method, system and equipment based on binocular vision and medium
CN122281785A
Contour correction calibration method, system, device and medium based on binocular vision
CN122281785B