A vision-based point cloud image fusion method

By extracting the geometric features of point cloud data and the visual features of image data for matching and screening, a three-dimensional model is constructed and optimized, the problem of low matching accuracy in point cloud image fusion method is solved, high-quality three-dimensional model generation is achieved, and accurate data support is provided.

CN119445005BActive Publication Date: 2025-07-18NANJING SIWEI VECTOR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510045771.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-07-18
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The existing point cloud image fusion method is prone to errors in the calculation of feature points during matching, resulting in low matching accuracy and inability to provide accurate and reliable data support for the recognition and judgment of target objects.

Method used

By extracting the geometric features of point cloud data and visual features of image data, feature matching and screening are performed, the initial three-dimensional model is constructed, and optimization is performed to generate high-quality three-dimensional models.

Benefits of technology

It improves the fusion accuracy of point cloud data and image data, reduces the probability of mismatch, and can comprehensively and accurately reflect the detailed information of the target object such as the shape, surface, color and texture, providing more accurate and reliable data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445005B_ABST
    Figure CN119445005B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image data processing. The present invention discloses a vision-based point cloud image fusion method, which includes extracting geometric features and visual features from point cloud data and image data, performing feature matching on point cloud feature points and image feature points, importing the geometric features and visual features into the fusion points in the model architecture to construct an initial three-dimensional model, and generating a high-quality three-dimensional model after optimizing the initial three-dimensional model. Compared with the prior art, the present invention screens the matching feature points after the matching operation of the point cloud feature points and the image feature points, which not only improves the fusion accuracy of the point cloud data and the image data, but also reduces the probability of mis-matching between the two. By fusing the geometric features and the visual features, the details such as the shape, surface, color and texture of the target objects in various different types of scenes can be comprehensively reflected, so as to construct a high-quality three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and more specifically, to a vision-based point cloud image fusion method. Background Art

[0002] With the rapid development of computer vision and three-dimensional reconstruction technologies, vision-based point cloud image fusion methods have shown great application potential in multiple fields such as autonomous driving, robot navigation, and industrial design. Since point cloud data lacks color and texture, and image data is difficult to directly represent the three-dimensional geometric structure of an object, there are limitations in using only point cloud data or image data alone. Therefore, it is necessary to perform fusion processing on point cloud data and image data through a point cloud image fusion method to generate a high-quality three-dimensional model with both precise geometric shape and color texture information.

[0003] The patent application with the publication number CN116579967A discloses a three-dimensional point cloud image fusion method and system based on computer vision. Through a fusion analysis module, it can perform decision analysis on the point cloud image fusion method, calculate a decision coefficient through comprehensive analysis and calculation of various processing parameters of the fusion task, and execute corresponding fusion tasks at an appropriate fusion timing to improve the efficiency of image fusion processing. Through a segmentation optimization module, it can perform optimization analysis on the process of point cloud image fusion processing, thereby obtaining central values, grid values, and random values through the optimization coefficients of fusion features processed by different segmentation modes, and adjusting the selection weights of image segmentation modes according to the central values, grid values, and random values to improve the efficiency of image segmentation processing;

[0004] Existing point cloud image fusion methods evaluate and analyze the accuracy of data matching fusion by calculating the reliability of the fusion match between point cloud data and image data. For example, in the above patent application, it calculates a decision coefficient through comprehensive analysis of various processing parameters involved in the fusion task in point cloud data and image data, and determines the fusion timing of the fusion task through the decision coefficient to achieve the matching fusion operation between point cloud data and image data. However, since a large amount of matching data of different types of feature points needs to be processed and calculated when point cloud data and image data are matched, when there are a large number of different types and similar feature points that need to be matched and calculated, it is easy to have the phenomenon of incorrect matching calculation of feature points, which in turn leads to low matching accuracy between point cloud data and image data and cannot provide more accurate and reliable data support for the subsequent recognition and judgment of the target object.

[0005] In view of this, the present invention proposes a vision-based point cloud image fusion method to solve the above problems. Summary of the Invention

[0006] In order to overcome the above defects of the prior art and to achieve the above objectives, the present invention provides the following technical solutions: a point cloud image fusion method based on vision, applied to a fusion server, comprising:

[0007] S1: acquiring point cloud data of the target object, and extracting geometric features from the preprocessed point cloud data after preprocessing the point cloud data, wherein the geometric features include shape features, surface features, and edge features;

[0008] S2: acquiring image data of the target object, and extracting visual features from the preprocessed image data after preprocessing the image data, wherein the visual features include color features, texture features, and boundary features;

[0009] S3: Identify the point cloud feature points of the point cloud data and the image feature points of the image data, perform feature matching on the point cloud feature points and the image feature points based on the matching criteria, and select the registration feature points;

[0010] S4: Build a model architecture based on point cloud data, and import geometric features and visual features into the fusion points in the model architecture to build an initial 3D model;

[0011] S5: After optimizing the initial three-dimensional model, a high-quality three-dimensional model is generated.

[0012] Furthermore, the preprocessing method of point cloud data is:

[0013] Import all point cloud data into the PCL library and load it, set the neighborhood size and outlier ratio, and determine the neighborhood of each point cloud data corresponding to the point location;

[0014] In the same neighborhood, measure the distance between any two points one by one, and add up all the distance values and average them to get the point distance value. The points with a point distance value greater than the outlier ratio are recorded as outliers, and the point cloud data corresponding to the outliers are removed.

[0015] The remaining point cloud data is converted into a grid structure through grid filtering technology, and the standard radius of the grid structure is set. The center point of the grid structure is identified through computer vision, and a circle is drawn with the center point as the base point and the standard radius as the radius to generate a grid circle.

[0016] The points outside the grid circle are recorded as out-of-circle points, and the point cloud data corresponding to the out-of-circle points are eliminated to obtain the preprocessed point cloud data.

[0017] Furthermore, the method for extracting shape features, surface features and edge features is as follows:

[0018] A scale space is constructed, and Gaussian blurred images of different scales are generated through Gaussian blur processing and downsampling processing;

[0019] Among Gaussian blurred images of different scales, calculate the differences between adjacent-scale Gaussian blurred images one by one to obtain difference-of-Gaussian (DoG) images, mark the gray values of the points in the DoG images, and record the points corresponding to the maximum and minimum gray values as extreme points.

[0020] Calculate the variance of the gray values of all points in the region where the extreme points are located, which is denoted as the contrast, and remove the extreme points with a contrast less than a preset contrast threshold. Record the remaining extreme points as key points.

[0021] Calculate the gradients of the remaining points in the region where the key points are located in the X-axis and Y-axis directions through the Sobel edge detection algorithm to obtain the gradient direction and magnitude. Use the horizontal axis to represent the gradient direction and the vertical axis to represent the gradient magnitude, and count the histogram.

[0022] Select the gradient direction with the largest peak in the histogram as the main direction of the key point, and combine it with the magnitude of the key point to generate a feature descriptor.

[0023] Identify the key words of the feature descriptors one by one, and record the feature descriptors with key words of shape, surface, and edge as shape features, surface features, and edge features respectively.

[0024] Furthermore, the preprocessing method for image data is as follows:

[0025] Import all the image data into the graphics processing unit, and identify the central dividing line of all the image data through computer vision technology.

[0026] Continuously adjust the tilt angle of the image data through the tilt correction algorithm until the central dividing line of the image data coincides with the standard dividing line to generate corrected image data.

[0027] Remove the noise of the corrected image data through the median filtering algorithm, and enhance the edge information of the denoised image data through the Laplacian algorithm to generate preprocessed image data.

[0028] Furthermore, the extraction methods for color features, texture features, and boundary features are as follows:

[0029] Convert the preprocessed image data from the RGB color space to the HSV color space, use the color automatic segmentation technology to divide the image data into adjacent A sub-regions, use one of the color components in the HSV color space as the index of the sub-region, and convert the image data into a binary color index set to generate color features.

[0030] Randomly select a pixel point from A sub-regions as the target point, and compare the pixel value of the target point with the pixel values of the pixel points in the sub-regions that are annularly adjacent to the target point in the clockwise direction;

[0031] When the pixel value of the target point is greater than or equal to the pixel values of the pixel points in the annularly adjacent sub-regions, assign the pixel value of 1 to the pixel points in the annularly adjacent sub-regions;

[0032] When the pixel value of the target point is less than the pixel values of the pixel points in the annularly adjacent sub-regions, assign the pixel value of 0 to the pixel points in the annularly adjacent sub-regions to generate a binary number, calculate the frequencies of 0 and 1 appearing in the A sub-regions in turn, generate A histograms, and connect the A histograms to generate texture features;

[0033] Perform edge detection on the preprocessed image data through the Canny algorithm to obtain an edge image, and identify the boundary orientation and boundary curvature of the edge image through computer vision technology. After combining the boundary orientation and boundary curvature, obtain boundary features.

[0034] Furthermore, the matching criterion is: the smaller the Euclidean distance between the geometric feature point and the image feature point, the higher the matching degree between the geometric feature point and the image feature point.

[0035] Furthermore, the feature matching method for the point cloud feature point and the image feature point is as follows:

[0036] Pre-calibrate the rotation matrix and translation vector between the lidar and the high-resolution camera, and use the rotation matrix and translation vector as the conversion standard to convert the point cloud feature points from the radar coordinate system to the camera coordinate system to generate the feature point coordinates in the camera coordinate system;

[0037] The expression of the feature point in the camera coordinate system is:

[0038] ;

[0039] In the formula, is the feature point coordinate in the camera coordinate system, is the rotation matrix, is the feature point coordinate in the radar coordinate system, is the translation vector;

[0040] Pre-calibrate the camera internal parameter matrix, and convert the feature point coordinates in the camera coordinate system to image pixel coordinates through the camera internal parameter matrix;

[0041] The expression of the image pixel coordinates is:

[0042] ;

[0043] In the formula, is the image pixel coordinate, is the camera internal parameter matrix;

[0044] By using the Euclidean distance formula, calculate the Euclidean distance between the point cloud feature points and the image feature points in the multi-dimensional space one by one. Denote the point cloud feature point and the image feature point corresponding to the minimum Euclidean distance as the matching feature points, and perform matching on the matching feature points.

[0045] Further, the method for screening the registered feature points is as follows:

[0046] Preset an error threshold and the number of iterations of the RANSAC algorithm, and initialize an empty set, denoted as the inlier set;

[0047] Successively take all the matching feature points as the initial samples, estimate the transformation matrix through the initial samples, and calculate the distance from the points in the transformation matrix to the transformation plane, denoted as the projection error;

[0048] Denote the matching feature points corresponding to the initial samples with a projection error less than the error threshold as inliers, and import the inliers into the inlier set. Count the number of inliers in the inlier set to obtain the initial value;

[0049] Successively denote the remaining matching feature points as verification samples, calculate the projection errors of the verification samples, identify the inliers from the remaining matching feature points, and import the identified inliers into the inlier set;

[0050] Real-time count the number of inliers in the inlier set to obtain the real-time value. When the real-time value is greater than the initial value, update the transformation matrix and the inlier set;

[0051] Repeat the above steps until the preset number of iterations is reached or the real-time value no longer increases. Select the transformation matrix corresponding to the maximum real-time value as the optimal transformation matrix, and denote the inliers in the inlier set of the optimal transformation matrix as the registered feature points.

[0052] Further, the method for constructing the initial three-dimensional model is as follows:

[0053] Based on the point cloud data, construct a model architecture corresponding to the point cloud data through three-dimensional reconstruction technology. Mark the original positions corresponding to the point cloud data one by one in the model architecture, and sequentially number the original positions in ascending order according to the marking sequence to generate fused points;

[0054] In the order from small to large by number, successively fuse the shape features, surface features, edge features, color features, texture features, and boundary features corresponding to the registered feature points through the data fusion algorithm to obtain the fused features;

[0055] Assign the fused features to the corresponding fusion points one by one. After coloring the fusion points after assignment, an initial 3D model is constructed.

[0056] Furthermore, the method for generating a high-quality 3D model is as follows:

[0057] Identify the boundary of the initial 3D model through computer vision technology. After drawing a line along the position of the boundary, a model boundary line is generated.

[0058] Mark the fusion points located on the model boundary line as key points, and mark the geometric features and visual features on the key points as key features.

[0059] Enhance the key features through a feature enhancement algorithm to generate detailed features, and mark the grid where the detailed features are located as the detailed area.

[0060] Perform regional simplification on the initial 3D model through a model simplification algorithm, and restore the detailed area and detailed features through a detail restoration technique to generate a high-quality 3D model.

[0061] The technical effects and advantages of a vision-based point cloud image fusion method of the present invention:

[0062] By extracting the geometric features in the point cloud data and the visual features in the image data, the present invention can provide a direct and clear matching object for the matching of the point cloud data and the image data, and screen the matching feature points after the matching operation of the point cloud feature points and the image feature points. This not only improves the fusion accuracy of the point cloud data and the image data, but also greatly reduces the probability of mis-matching between the two. By fusing the geometric features and the visual features, the present invention can comprehensively and accurately reflect the detailed information such as the shape, surface, color, and texture of the target object in a variety of different types of scenes, thereby constructing a high-quality 3D model that can truly and accurately reflect the point cloud information and image information of the target object, ensuring that the fused 3D model has stronger details, higher clarity, and better realism, and further providing more accurate and reliable data support for the subsequent recognition and judgment of the target object. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a flowchart of a vision-based point cloud image fusion method provided by Embodiment 1 of the present invention;

[0064] Figure 2 It is a module diagram of a vision-based point cloud image fusion system provided by Embodiment 2 of the present invention;

[0065] Figure 3 It is a structural diagram of an electronic device provided by Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0067] Embodiment 1: Please refer to Figure 1 As shown, a vision-based point cloud image fusion method described in this embodiment is applied to a fusion server and includes:

[0068] S1: Obtain the point cloud data of the target object. After preprocessing the point cloud data, extract geometric features from the preprocessed point cloud data.

[0069] The target object refers to a three-dimensional object that needs to perform point cloud image fusion operations and serves as an object that provides basic and original data for point cloud image fusion. The point cloud data is a data set obtained by scanning the target object with a scanning device and can represent the coordinates of each point of the target object in three-dimensional space, thus serving as one of the data for subsequent point cloud image fusion. Specifically, the scanning device includes, but is not limited to, three-dimensional scanners, stereo cameras, lidar, etc. In this embodiment, the point cloud data is obtained by scanning the target object with a lidar.

[0070] Since the point cloud data is obtained by directly scanning the target image with a scanning device, the quantity of the point cloud data is relatively large at this time, and there are also problems with the noise of the point cloud data. In order to improve the subsequent use effect of the point cloud data, it is necessary to perform preprocessing operations on the point cloud data to ensure the high quality and consistency of the point cloud data.

[0071] The preprocessing method of the point cloud data is as follows:

[0072] Import all the point cloud data into the PCL library and load it. Set the neighborhood size and the outlier ratio, and determine the neighborhood of the area where each point of the point cloud data is located. The neighborhood size is used to calculate the number of points within the neighborhood of each point. The outlier ratio is the threshold for determining whether a point is an outlier. Usually, the outlier ratio is calculated based on the distance mean and standard deviation of the points within the neighborhood, so as to remove those points that are too far away within the neighborhood and reduce the quantity of the point cloud data.

[0073] Within the same neighborhood, measure the distance value between any two points one by one, and after accumulating all the distance values and taking the average, obtain the point distance value. Mark the points with a point distance value greater than the outlier ratio as outliers, and remove the point cloud data corresponding to the outliers.

[0074] The remaining point cloud data is converted into a grid structure through grid filtering technology, and the standard radius of the grid structure is set. The center point of the grid structure is identified through computer vision, and a circle is drawn with the center point as the base point and the standard radius as the radius to generate a grid circle. The standard radius is a numerical representation of the maximum scanning radius of the grid structure and provides a basis for the subsequent identification of points outside the circle.

[0075] The points located outside the grid circle are marked as points outside the circle, and the point cloud data corresponding to the points outside the circle is removed to obtain the preprocessed point cloud data.

[0076] The preprocessed point cloud data can maintain high quality and consistency. However, the information and features contained in the point cloud data at this time are numerous and of various types. In order to accurately and quickly obtain the direct information for subsequent point cloud image fusion from the point cloud data, it is necessary to extract geometric features from the preprocessed point cloud data so that the geometric features can clearly represent the features such as the shape, size, and boundary of the target object in the point cloud data.

[0077] Geometric features include shape features, surface features, and edge features. Shape features refer to the features in the point cloud data that can reflect the overall or local geometric shape of the target object. Surface features refer to the geometric shape features on the surface of the target object in the point cloud data, including but not limited to details such as protrusions, depressions, and textures. Edge features refer to the features at the surface or structure boundary of the target object in the point cloud data.

[0078] The extraction methods for shape features, surface features, and edge features are as follows:

[0079] A scale space is constructed, and Gaussian blurred images of different scales are generated through Gaussian blur processing and downsampling processing. Gaussian blur processing is to reduce noise and details and highlight the main features. In a high-dimensional space, this can be achieved by applying Gaussian blur to each dimension of the point cloud data separately. Downsampling can be achieved through methods such as voxel grid downsampling to reduce the number of data points while maintaining the main shape and structure of the point cloud data.

[0080] In Gaussian blurred images of different scales, the differences between adjacent-scale Gaussian blurred images are calculated one by one to obtain Gaussian difference images, and the gray values of the points in the Gaussian difference images are marked. The points corresponding to the maximum and minimum gray values are marked as extreme points.

[0081] The variance of the gray values of all points in the region where the extreme points are located is calculated and denoted as the contrast. Extreme points with a contrast less than a preset contrast threshold are removed, and the remaining extreme points are denoted as key points. The preset contrast threshold is the minimum value of the variance of the gray values of the points marked as extreme points, which can provide a basis for the subsequent identification and determination of extreme points.

[0082] Calculate the gradients of the remaining points in the region where the keyword is located in the X-axis direction and the Y-axis direction through the Sobel edge detection algorithm, obtain the gradient direction and amplitude, use the horizontal axis to represent the gradient direction, use the vertical axis to represent the gradient amplitude, and count the histogram;

[0083] Select the gradient direction with the largest peak in the histogram as the main direction of the key point, and combine it with the amplitude of the key point to generate a feature descriptor;

[0084] Identify the keywords of the feature descriptors one by one, and record the feature descriptors with keywords of shape, surface, and edge as shape features, surface features, and edge features respectively. Keywords are used for a concise representation of the true meaning of the descriptors, so that shape features, surface features, and edge features can be quickly and accurately identified.

[0085] S2: Obtain the image data of the target object. After preprocessing the image data, extract the visual features from the preprocessed image data;

[0086] Image data refers to the data set obtained by photographing the target object through a camera device, which can represent the color, texture, and detail information of the target object, and provides a high-quality image data source for subsequent image processing and 3D reconstruction; specifically, the image data in this embodiment is obtained by photographing multiple angles of the target object with a high-resolution camera.

[0087] Since the image data is obtained by directly photographing the target image through a camera device, the image quality and clarity of the image data may have certain defects at this time. In order to improve the subsequent use effect of the image data, it is necessary to perform preprocessing operations on the image data;

[0088] The preprocessing method of the image data is as follows:

[0089] Import all the image data into the graphics processing unit, and identify the central dividing line of all the image data through computer vision technology; the central dividing line refers to the line passing through the center point of the image data and capable of bisecting the image data in the vertical direction, and is used as the basis for judging whether the image data is tilted;

[0090] Continuously adjust the tilt angle of the image data through the tilt correction algorithm until the central dividing line of the image data coincides with the standard dividing line to generate the corrected image data; the standard dividing line refers to the central dividing line of the image data in the state of keeping vertical and not tilted, and is used as the basis for subsequent adjustment of the real-time tilt state of the image data;

[0091] Remove the noise of the corrected image data through the median filtering algorithm, and enhance the edge information of the denoised image data through the Laplace algorithm to generate the preprocessed image data.

[0092] The preprocessed image data can maintain a high level in terms of clarity and contrast. However, the types of information contained in the preprocessed image data are miscellaneous and large in quantity. In order to represent the preprocessed image data accurately and concisely, it is necessary to extract visual features from the preprocessed image data so that the visual features can represent the feature information regarding the visual direction in the preprocessed image data;

[0093] Visual features include color features, texture features, and boundary features; color features refer to the features in the image data that can reflect the overall or local color information of the target object, texture features refer to the features in the image data that can reflect the information on the boundary trend of the target object area, and boundary features refer to the features in the image data that can reflect the distribution information of the boundary lines of the target object;

[0094] The extraction methods for color features, texture features, and boundary features are as follows:

[0095] Convert the preprocessed image data from the RGB color space to the HSV color space, use the color automatic segmentation technology to divide the image data into adjacent A sub-regions, use one of the color components in the HSV color space as the index of the sub-region, convert the image data into a binary color index set, and generate color features;

[0096] Randomly select a pixel point from the A sub-regions as the target point, and compare the pixel value of the target point with the pixel values of the pixel points in the sub-regions that are annularly adjacent to the target point in the clockwise direction; the clockwise direction can ensure the orderliness of the comparison of the pixel points in the annularly adjacent sub-regions and avoid the phenomenon of chaotic and disorderly comparison process;

[0097] When the pixel value of the target point is greater than or equal to the pixel value of the pixel point in the annularly adjacent sub-region, assign the pixel value of the pixel point in the annularly adjacent sub-region as 1;

[0098] When the pixel value of the target point is less than the pixel value of the pixel point in the annularly adjacent sub-region, assign the pixel value of the pixel point in the annularly adjacent sub-region as 0, generate a binary number, calculate the frequencies of 0 and 1 appearing in the A sub-regions in turn, generate A histograms, and connect the A histograms to generate texture features;

[0099] Perform edge detection on the preprocessed image data through the Canny algorithm to obtain an edge image, and identify the boundary trend and boundary curvature of the edge image through computer vision technology. After combining the boundary trend and boundary curvature, obtain boundary features. The boundary trend is used to represent the distribution direction of the boundary position of the image data, and the boundary curvature is used to represent the bending curvature of the boundary position of the image data, so as to accurately represent the boundary features of the image data.

[0100] S3: Identify the point cloud feature points of the point cloud data and the image feature points of the image data. Based on the matching criterion, perform feature matching on the point cloud feature points and the image feature points, and filter out the registration feature points;

[0101] The point cloud feature points refer to the positions where the geometric features are located in the point cloud data, such that each point cloud feature point has a unique and independent feature descriptor. And since the point cloud feature points are the positions in the point cloud data where the features are relatively prominent and obvious, therefore, the key points in the point cloud data are the point cloud feature points.

[0102] The image feature points refer to the positions where the visual features are located in the image data, such that each image feature point has a unique and independent feature descriptor. And since the image feature points are the positions in the image data where the features are relatively prominent and obvious, therefore, the target points in the image data are the image feature points.

[0103] After obtaining the point cloud feature points and the image feature points, since the point cloud feature points are the position representations of the target object in terms of the point cloud, and the image feature points are the position representations of the target object in terms of the image, in order to perform high-similarity screening and matching on the point cloud feature points and the image feature points, it is necessary to perform a matching operation on the point cloud feature points and the image feature points, and under the limitation of the matching criterion, achieve the effective matching of the point cloud feature points and the image feature points, and then find the positions with extremely high similarity in the point cloud data and the image data;

[0104] The matching criterion is: the smaller the Euclidean distance between the geometric feature points and the image feature points, the higher the matching degree between the geometric feature points and the image feature points; thus, accurate matching can be performed on the geometric feature points in the point cloud data and the image feature points in the image data;

[0105] The feature matching method for the point cloud feature points and the image feature points is:

[0106] Pre-calibrate the rotation matrix and translation vector between the lidar and the high-resolution camera, and use the rotation matrix and translation vector as the conversion standard to convert the point cloud feature points from the radar coordinate system to the camera coordinate system, generating the feature point coordinates in the camera coordinate system; the rotation matrix and translation vector, as the external parameters between the lidar and the high-resolution camera, provide the information for the rotation and translation of the point cloud feature points and are used to achieve the coordinate conversion between the radar coordinate system and the camera coordinate system;

[0107] The expression of the feature points in the camera coordinate system is:

[0108] ;

[0109] In the formula, is the feature point coordinate in the camera coordinate system, is a rotation matrix, is the coordinate of the feature point in the radar coordinate system, is a translation vector;

[0110] Pre-calibrate the camera internal parameter matrix, and convert the coordinate of the feature point in the camera coordinate system to the image pixel coordinate through the camera internal parameter matrix;

[0111] The expression of the image pixel coordinate is:

[0112] ;

[0113] In the formula, is the image pixel coordinate, is the camera internal parameter matrix;

[0114] Through the Euclidean distance formula, calculate the Euclidean distance between the point cloud feature point and the image feature point in the multi-dimensional space one by one. Record the point cloud feature point and the image feature point corresponding to the minimum value of the Euclidean distance as the matching feature points, and perform matching on the matching feature points.

[0115] After the feature matching between the point cloud feature point and the image feature point, there may be a phenomenon that the Euclidean distance of the point cloud feature point and the image feature point after feature matching is not the minimum value when performing feature matching. Therefore, it will cause errors in the matching results. In order to avoid the phenomenon of matching errors, it is necessary to verify and screen the point cloud feature points and the image feature points after feature matching, so as to obtain the registration feature points;

[0116] The screening method of the registration feature points is:

[0117] Pre-set the error threshold and the number of iterations of the RANSAC algorithm, and initialize an empty set, denoted as the inlier set; through the set error threshold and the number of iterations of the RANSAC algorithm, it can provide an operation basis for the matching verification and screening of the registration feature points, ensuring that all registration feature points can be screened safely and quickly;

[0118] Take all the matching feature points as the initial samples in turn, estimate the transformation matrix through the initial samples, and calculate the distance from the point in the transformation matrix to the transformation plane, denoted as the projection error;

[0119] Record the matching feature points corresponding to the initial samples with the projection error less than the error threshold as inliers, and import the inliers into the inlier set, count the number of inliers in the inlier set, and obtain the initial value;

[0120] Take the remaining matching feature points as the verification samples in turn, calculate the projection error of the verification samples, identify the inliers from the remaining matching feature points, and import the identified inliers into the inlier set;

[0121] Statistically obtain the number of inliers in the inlier set in real time to obtain a real-time value. When the real-time value is greater than the initial value, update the transformation matrix and the inlier set;

[0122] Repeat the above steps until the preset number of iterations is reached or the real-time value no longer increases. Select the transformation matrix corresponding to the maximum real-time value as the optimal transformation matrix, and mark the inliers in the inlier set of the optimal transformation matrix as registration feature points.

[0123] Through continuous iterative calculation of the above RANSAC algorithm, it is possible to verify the matching results of the point cloud feature points and the image feature points after feature matching, and eliminate the feature points with low matching accuracy to ensure the high accuracy of the feature matching between the point cloud feature points and the image feature points.

[0124] S4: Based on the point cloud data, construct a model architecture, and import the geometric features and visual features into the fusion points in the model architecture to construct an initial three-dimensional model;

[0125] The model architecture is a three-dimensional simulation model architecture of the target object initially constructed according to the point cloud data, which only has geometric features and no visual features, and serves as the basis for the construction of the subsequent initial three-dimensional model, and can provide a reasonable model for the import of geometric features and visual features;

[0126] The fusion point refers to the data fusion position where the point cloud data and the image data can achieve position correspondence and information correspondence when they are fused in the model architecture, and serves as the node of the subsequent initial three-dimensional model to ensure the accuracy and rationality of the point cloud data and the image data during fusion and improve the accuracy of the subsequent initial three-dimensional model;

[0127] The initial three-dimensional model refers to a three-dimensional simulation model of the target object constructed by initially fusing geometric features and visual features, and serves as the basis for subsequent optimization and recognition of the target object;

[0128] The construction method of the initial three-dimensional model is as follows:

[0129] Based on the point cloud data, construct a model architecture corresponding to the point cloud data through three-dimensional reconstruction technology. Mark the original positions corresponding to the point cloud data one by one in the model architecture, and sequentially number the original positions in ascending order according to the marking sequence to generate fusion points; the model architecture is a basic model without any substantial meaning and serves as the basis for the construction of the subsequent initial three-dimensional model, and the original position is the position in the model architecture used to provide a position for the fusion of the point cloud data and the image data and serves as the precursor of the fusion point;

[0130] In the order from small to large by number, sequentially fuse the shape features, surface features, edge features, color features, texture features, and boundary features corresponding to the registration feature points through a data fusion algorithm to obtain fusion features;

[0131] Assign the fusion features to the corresponding fusion points one by one. After point cloud coloring the assigned fusion points, an initial three-dimensional model is constructed. After point cloud coloring the fusion points, the positions corresponding to each point cloud data can be assigned corresponding colors, so that the initial three-dimensional model can have a relatively consistent appearance color with the target object.

[0132] It should be noted that the constructed initial three-dimensional model can perform initial and reasonable three-dimensional simulation on the point cloud data and image data of the target object, ensuring that the target object can be initially presented by three-dimensional simulation.

[0133] S5: After optimizing the initial three-dimensional model, a high-quality three-dimensional model is generated;

[0134] Since the initial three-dimensional model is only an initialized three-dimensional simulation model of the target object, at this time, the initial three-dimensional model may have problems such as rough surface, unclear texture, and prominent details, resulting in the inability of the initial three-dimensional model to accurately and realistically perform three-dimensional simulation on the target object. Therefore, it is necessary to optimize the initial three-dimensional model to improve the overall performance of the initial three-dimensional model, so as to obtain a high-quality three-dimensional model;

[0135] The method for generating a high-quality three-dimensional model is as follows:

[0136] Identify the boundary of the initial three-dimensional model through computer vision technology, draw a line along the position of the boundary, and generate a model boundary line;

[0137] Mark the fusion points located on the model boundary line as key points, and mark the geometric features and visual features on the key points as key features;

[0138] Enhance the key features through a feature enhancement algorithm to generate detail features, and mark the grid where the detail features are located as the detail area; By generating detail features and detail areas, the detail information with important significance and obvious features in the three-dimensional model can be strengthened and highlighted, so that the accuracy between the three-dimensional model and the target object can be higher, and the detail information of the target object can be represented with higher accuracy;

[0139] Perform regional simplification on the initial three-dimensional model through a model simplification algorithm, and restore the detail area and detail features through a detail restoration technology to generate a high-quality three-dimensional model.

[0140] It should be noted that, in addition to enhancing the details of the initial 3D model, surface smoothing and texture mapping operations can also be performed on the initial 3D model according to actual optimization requirements. In the surface smoothing operation, to eliminate the noise and unevenness generated during the data fusion process, the Kalman filtering method is used for dynamic smoothing processing, effectively improving the smoothness of the surface of the initial 3D model and enhancing the realism of the initial 3D model. To handle the sampling problem in texture mapping, bilinear interpolation technology is used to improve the accuracy and effect of texture mapping, so as to enhance the realism and visual effect of the initial 3D model. The surface smoothing operation and the texture mapping operation are both existing technologies in this field, and the relevant algorithms are also widely used.

[0141] In this embodiment, by extracting the geometric features in the point cloud data and the visual features in the image data, a direct and clear matching object can be provided for the matching of the point cloud data and the image data. After the matching operation of the point cloud feature points and the image feature points, the screening of the matching feature points is carried out, which not only improves the fusion accuracy of the point cloud data and the image data, but also greatly reduces the probability of false matching between the two. By fusing the geometric features and the visual features, the detailed information such as the shape, surface, color, and texture of the target object in various different types of scenes can be comprehensively and accurately reflected, so as to construct a high-quality 3D model that can truly and accurately reflect the point cloud information and image information of the target object, ensuring that the details of the fused 3D model are stronger, the clarity is higher, and the realism is better, and further providing more accurate and reliable data support for the subsequent recognition and judgment of the target object.

[0142] Embodiment Two: Please refer to Figure 2 As shown in the figure, for the parts not described in detail in this embodiment, please refer to the description content of Embodiment One. A vision-based point cloud image fusion system is provided, which is applied to a fusion server and is used to implement a vision-based point cloud image fusion method, including a first feature extraction module, a second feature extraction module, a feature matching module, a data fusion module, and a model optimization module. Among them, each module is connected by a wired or wireless network method;

[0143] The first feature extraction module is used to obtain the point cloud data of the target object. After preprocessing the point cloud data, geometric features are extracted from the preprocessed point cloud data. The geometric features include shape features, surface features, and edge features;

[0144] The second feature extraction module is used to obtain the image data of the target object. After preprocessing the image data, visual features are extracted from the preprocessed image data. The visual features include color features, texture features, and boundary features;

[0145] A feature matching module, configured to identify the point cloud feature points of the point cloud data and the image feature points of the image data, perform feature matching on the point cloud feature points and the image feature points based on a matching criterion, and filter out the registered feature points;

[0146] A data fusion module, configured to construct a model architecture based on the point cloud data, and import geometric features and visual features into the fusion points in the model architecture to construct an initial three-dimensional model;

[0147] A model optimization module, configured to generate a high-quality three-dimensional model after optimizing the initial three-dimensional model.

[0148] Embodiment 3: Please refer to Figure 3 As shown, this embodiment publicly provides an electronic device, including a processor and a memory;

[0149] Wherein, a computer program that can be called by the processor is stored in the memory;

[0150] The processor executes the implemented method for visual-based point cloud image fusion by calling the computer program stored in the memory.

[0151] Since the electronic device introduced in this embodiment is the electronic device adopted for implementing the method for visual-based point cloud image fusion in Embodiment 1 of the present application, based on the method for visual-based point cloud image fusion introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device adopted for the method for visual-based point cloud image fusion in the embodiments of the present application, it falls within the scope of protection of the present application.

[0152] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention.

Claims

1. A vision-based point cloud image fusion method, applied to a fusion server, characterized in that Including: S1: Obtain the point cloud data of the target object. After preprocessing the point cloud data, extract geometric features from the preprocessed point cloud data. The geometric features include shape features, surface features, and edge features; S2: Obtain the image data of the target object. After preprocessing the image data, extract visual features from the preprocessed image data. The visual features include color features, texture features, and boundary features; S3: Identify the point cloud feature points of the point cloud data and the image feature points of the image data. Based on the matching criterion, perform feature matching on the point cloud feature points and the image feature points, and filter out the registration feature points; S4: Build a model architecture based on the point cloud data, and import the geometric features and visual features into the fusion points in the model architecture to build an initial three-dimensional model; S5: After optimizing the initial three-dimensional model, generate a high-quality three-dimensional model; The extraction methods of color features, texture features, and boundary features are as follows: Convert the preprocessed image data from the RGB color space to the HSV color space, use the color automatic segmentation technology to divide the image data into adjacent A sub-regions, use one of the color components in the HSV color space as the index of the sub-region, convert the image data into a binary color index set to generate color features; Randomly select a pixel point from the A sub-regions as the target point, and compare the pixel values of the target point and the pixel points in the sub-regions that are annularly adjacent to the target point in the clockwise direction; When the pixel value of the target point is greater than or equal to the pixel value of the pixel points in the annularly adjacent sub-regions, assign the pixel points in the annularly adjacent sub-regions a value of 1; When the pixel value of the target point is less than the pixel value of the pixel points in the annularly adjacent sub-regions, assign the pixel points in the annularly adjacent sub-regions a value of 0 to generate a binary number, calculate the frequencies of 0 and 1 appearing in the A sub-regions in turn to generate A histograms, and connect the A histograms to generate texture features; Perform edge detection on the preprocessed image data through the Canny algorithm to obtain an edge image, and identify the boundary trend and boundary curvature of the edge image through computer vision technology. After combining the boundary trend and boundary curvature, obtain boundary features; The matching criterion is: The smaller the Euclidean distance between the point cloud feature point and the image feature point, the higher the matching degree between the point cloud feature point and the image feature point; The feature matching method for the point cloud feature points and the image feature points is: Calibrate the rotation matrix and translation vector between the lidar and the high-resolution camera in advance, and use the rotation matrix and translation vector as the conversion standard to convert the point cloud feature points from the radar coordinate system to the camera coordinate system to generate the feature point coordinates in the camera coordinate system; The expression of the feature point in the camera coordinate system is: ; wherein, is the coordinate of the feature point in the camera coordinate system, is the rotation matrix, is the coordinate of the feature point in the radar coordinate system, is the translation vector; the camera internal parameter matrix is pre-calibrated, and the coordinate of the feature point in the camera coordinate system is converted into the image pixel coordinate through the camera internal parameter matrix; the expression of the image pixel coordinate is: ; In the formula, is the image pixel coordinate, is the camera intrinsic matrix; by using the Euclidean distance formula, the Euclidean distances between the point cloud feature points and the image feature points in the multi-dimensional space are calculated one by one. The point cloud feature point and the image feature point corresponding to the minimum Euclidean distance are recorded as the matching feature points, and the matching feature points are matched; The method for screening registration feature points is as follows: preset the error threshold and the number of iterations of the RANSAC algorithm, and initialize an empty set, denoted as the inlier set; sequentially take all the matching feature points as the initial samples, estimate the transformation matrix through the initial samples, and calculate the distance from the points in the transformation matrix to the transformation plane, denoted as the projection error; mark the matching feature points corresponding to the initial samples with projection errors less than the error threshold as inliers, and import the inliers into the inlier set, count the number of inliers in the inlier set to obtain the initial value; Sequentially mark the remaining matching feature points as verification samples, calculate the projection errors of the verification samples, identify the inliers from the remaining matching feature points, and import the identified inliers into the inlier set; Real-time count the number of inliers in the inlier set to obtain the real-time value. When the real-time value is greater than the initial value, update the transformation matrix and the inlier set; repeat the above steps until the preset number of iterations is reached or the real-time value no longer increases and then stop. Select the transformation matrix corresponding to the maximum real-time value as the optimal transformation matrix, and mark the inliers in the inlier set of the optimal transformation matrix as the registration feature points; The method for constructing the initial 3D model is as follows: based on the point cloud data, construct a model architecture corresponding to the point cloud data through 3D reconstruction technology, mark the original positions corresponding to the point cloud data one by one in the model architecture, and sequentially number the original positions in ascending order according to the marking order to generate fused points; in the order from small to large by number, sequentially perform feature fusion on the shape features, surface features, edge features, color features, texture features, and boundary features corresponding to the registration feature points through the data fusion algorithm to obtain fused features; assign the fused features to the corresponding fused points one by one, and after point cloud coloring of the assigned fused points, construct the initial 3D model; The method for generating a high-quality 3D model is as follows: identify the boundary of the initial 3D model through computer vision technology, draw a line along the position of the boundary, and generate the model boundary line; mark the fused points located on the model boundary line as key points, and mark the geometric features and visual features on the key points as key features; Enhance the key features through the feature enhancement algorithm to generate detailed features, and mark the grid where the detailed features are located as the detailed area; perform regional simplification on the initial 3D model through the model simplification algorithm, and restore the detailed area and detailed features through the detail restoration technology to generate a high-quality 3D model.

2. The method for fusing point cloud images based on vision according to claim 1, wherein, The method for preprocessing the point cloud data is as follows: import all the point cloud data into the PCL library and load it, set the neighborhood size and the outlier ratio, and determine the neighborhood of the area where the position corresponding to each point cloud data is located; In the same neighborhood, the distance between any two points is measured one by one, and all the distance values are accumulated and averaged to obtain the point distance value. The points with point distance values greater than the outlier ratio are recorded as outliers, and the point cloud data corresponding to the outliers are eliminated; the remaining point cloud data are converted into a grid structure through grid filtering technology, and the standard radius of the grid structure is set. The center point of the grid structure is identified through computer vision, and a circle is drawn with the center point as the base point and the standard radius as the radius to generate a grid circle; the points located outside the grid circle are recorded as points outside the circle, and the point cloud data corresponding to the points outside the circle are eliminated to obtain the preprocessed point cloud data.

3. A vision-based point cloud image fusion method according to claim 2, characterized in that, The method for extracting shape features, surface features and edge features is as follows: construct a scale space, and generate Gaussian blurred images of different scales through Gaussian blur processing and downsampling processing; in Gaussian blurred images of different scales, calculate the differences between Gaussian blurred images of adjacent scales one by one to obtain Gaussian difference images, and mark the grayscale values of the points in the Gaussian difference images, and record the points corresponding to the maximum and minimum grayscale values as extreme points; calculate the variance of the grayscale values of all points in the area where the extreme points are located, record it as contrast, and eliminate the extreme points whose contrast is less than the preset contrast threshold , and record the remaining extreme points as key points; calculate the gradients of the remaining points in the region where the keyword is located in the X-axis direction and the Y-axis direction through the Sobel edge detection algorithm, obtain the gradient direction and amplitude, and use the horizontal axis to represent the gradient direction and the vertical axis to represent the gradient amplitude, and calculate the histogram; select the gradient direction with the largest peak in the histogram as the main direction of the key point, and generate a feature descriptor after combining it with the amplitude of the key point; identify the keywords of the feature descriptor one by one, and record the feature descriptors with the keywords of shape, surface and edge as shape features, surface features and edge features respectively.

4. A vision-based point cloud image fusion method according to claim 3, characterized in that The preprocessing method of image data is as follows: all image data are imported into the graphics processor, and the central dividing line of all image data is identified through computer vision technology; the tilt angle of the image data is continuously adjusted through the tilt correction algorithm until the central dividing line of the image data coincides with the standard dividing line, thereby generating corrected image data; the noise of the corrected image data is removed through the median filtering algorithm, and the edge information of the denoised image data is enhanced through the Laplace algorithm, thereby generating preprocessed image data.

Citation Information

Patent Citations

  • Three-dimensional point cloud image fusion method and system based on computer vision

    CN116579967A

  • Method and system for constructing three-dimensional model through fusion of laser point cloud and single-lens image

    CN117853656A