A high-precision three-dimensional topography measurement method and system based on a speckle matching network

By using a speckle matching network-based method, the problems of long computation time and data loss in traditional speckle matching methods are solved, and fast and high-precision three-dimensional topography measurement is achieved. In particular, edge information is effectively protected in areas with large gradient changes.

CN115482268BActive Publication Date: 2026-01-13SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211142765.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-20
Publication Date
2026-01-13
Estimated Expiration
2042-09-20

AI Technical Summary

Technical Problem

Traditional speckle matching methods suffer from problems such as long computation time and data loss, making it difficult to meet the needs of high-precision 3D topography measurement. In particular, when processing larger resolution images, errors are prone to occur in areas with large edge and gradient changes.

Method used

A speckle matching network-based approach is adopted, including a feature extraction module, a cost volume construction module, a 3D aggregation module, and an edge refinement module. 3D topography measurement is performed using a high-performance GPU, the speckle dataset is used for preprocessing, and an end-to-end network is constructed for 3D topography measurement.

Benefits of technology

It achieves fast and high-precision 3D topography measurement with edge information protection, single-frame measurement error of less than 0.03mm, standard deviation of complex surface measurement of about 0.088mm, and significantly shortens the calculation time, thus significantly improving measurement efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115482268B_ABST
    Figure CN115482268B_ABST
Patent Text Reader

Abstract

The application discloses a high-precision three-dimensional topography measurement method and system based on a speckle matching network, wherein the measurement method comprises the following steps: acquiring a left surface image and a right surface image of an object to be measured; constructing a relationship between a target sub-region in the right surface image and a reference sub-region of the left surface image based on a mapping function; constructing a speckle dataset based on the left surface image and the right surface image, and pre-processing the speckle dataset; constructing a speckle matching network, and performing three-dimensional topography measurement based on the speckle dataset and the speckle matching network. The measurement system comprises a high-precision dataset acquisition device capable of acquiring a high-precision dataset and a high-performance GPU used for network training. The application uses a special speckle projection device to acquire a high-precision speckle binocular dataset containing thousands of data, and realizes fast, high-precision and edge information protection three-dimensional topography measurement. Compared with a traditional measurement method, the error is obviously reduced, and the training parameters and inference time are greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of three-dimensional topography measurement, and particularly relates to a high-precision three-dimensional topography measurement method and system based on a speckle matching network. BACKGROUND

[0002] Fast, high-precision and non-contact three-dimensional topography measurement plays an important role in scientific research and industrial application fields. Generally, three-dimensional optical measurement can be mainly divided into two categories: passive and active, and the main classification basis is whether to project light field. The passive method has lower matching accuracy in weak texture area, on the contrary, the active light projection measurement technology improves the measurement accuracy, and is more widely applied in fields including engineering field, material science, and topography measurement field.

[0003] Structured light is a typical active light measurement method, which is widely used due to its high measurement accuracy, high point cloud density and full-field measurement characteristics. According to different projection patterns, structured light technology includes fringe projection method and speckle projection method. The fringe projection method adopts multiple continuous phase shift fringe projection to perform phase unwrapping. In comparison, the speckle projection method only needs to project a speckle pattern, and this technology is widely used in consumer products and handheld measurement systems. The key part of the speckle measurement technology, binocular stereo matching technology mainly includes local matching method, global matching method and semi-global matching method. However, the matching accuracy of these methods is low, and they are not suitable for high-precision measurement fields. In the speckle measurement algorithm, there is a method based on digital speckle correlation, which can achieve the same measurement accuracy as the fringe projection method, and meets the accuracy requirements of topography measurement. The DSC method is first applied to the measurement of in-plane displacement and stress of materials. If it is applied to the field of topography measurement, there are two main problems to be solved. First, the calculation time and efficiency are inevitable problems. The surface of the object in topography measurement is usually very complex, and the disparity search range is large, so the iteration process is very time-consuming, which makes it difficult for this method to be practically applied. In order to speed up the measurement, the inverse Gauss-Newton (IC-GN) algorithm is proposed to increase the speed of sub-pixel matching. The sub-pixel interpolation lookup table is used for efficient interpolation operation. Shao proposed a pixel screening strategy, which can achieve a measurement speed of 30000 points / s, but this order of magnitude of measurement speed still cannot meet the demand of fast measurement. The second major problem is that the three-dimensional measurement based on the principle of DSC may have errors in the areas with large gradient changes and object edges, resulting in edge blurring and data missing, which reduces the integrity and reliability of three-dimensional measurement. Chen proposed a new edge detection method, which combines edge sub-regions to realize complete measurement. There is also a new multi-step stereo DSC algorithm that can preserve clear edge information. Wu proposed a phase-shift-assisted DSC method, which can realize high-precision dynamic three-dimensional measurement. However, due to the limitations of the DSC method itself, the effect of this edge protection still cannot meet the requirements of topography measurement. In summary, these improved methods based on the traditional DSC principle cannot fundamentally solve the above two problems.

[0004] Recently, deep learning techniques have many applications in the field of vision, and convolutional neural networks have been widely used in optical measurement, such as passive stereo vision technology. Researchers have proposed many related network architectures, which have greatly improved the matching accuracy, matching success rate and speed compared with traditional methods. Some studies only replace some steps in binocular matching, and some studies construct an end-to-end network structure to directly output the results. From feature extraction, cost volume construction, similarity function variants to the addition of new modules, researchers continue to improve the accuracy and efficiency of binocular matching from various aspects of the network. However, the matching accuracy of these network designs is not high, and the data sets they are mainly used for training cannot be directly applied to the field of topography measurement. The Scene Flow data set simulates different objects and scenes by computer, and there is still a gap between the real industrial scene. The KITTI data has the characteristics of large scale and sparse texture, and its accuracy is in the order of microns, which cannot meet the accuracy requirements of topography measurement. In the field of using deep learning combined with speckle images for precision measurement, Yin proposed an end-to-end network that can achieve high-precision disparity prediction from a single frame of speckle image, but there is still a lot of optimization and simplification space in the network structure. And the data set of this work is constructed by the method of stripe projection, which must require a digital optical processor (DLP), increasing the number of pictures and the complexity of the equipment required for database establishment. Hieu Nguyen also proposed a method of using convolutional network and speckle image for three-dimensional reconstruction, although this method can directly output the corresponding three-dimensional topography map, but it cannot be evaluated for precision, and it cannot get the real three-dimensional point cloud. And the image resolution of these two studies is 640x480, with the advancement of technology, we need to study larger resolution pictures. SUMMARY

[0005] The purpose of the present application is to provide a high-precision three-dimensional topography measurement method and system based on a speckle matching network to solve the problems of long calculation time and data loss in traditional DSC methods. At the same time, it can process larger resolution pictures, so that the measurement accuracy is high and the calculation consumption is less.

[0006] In order to achieve the above-mentioned purpose, the present application provides a high-precision three-dimensional topography measurement method based on a speckle matching network, comprising:

[0007] obtaining a left surface image and a right surface image of an object to be measured;

[0008] constructing the relationship between the target sub-area in the right surface image and the reference sub-area of the left surface image based on the mapping function;

[0009] constructing a speckle data set based on the left surface image and the right surface image, and preprocessing the speckle data set;

[0010] A speckle matching network is constructed, and three-dimensional topography measurement is performed based on the preprocessed speckle dataset and the speckle matching network.

[0011] Optionally, the speckle matching network comprises a feature extraction module, a cost volume construction module, a three-dimensional aggregation module, and an edge refinement module.

[0012] Two hollow convolution blocks are arranged after the network of the feature extraction module.

[0013] The three-dimensional aggregation module comprises a three-dimensional hourglass subnetwork and a pre-aggregation module.

[0014] Optionally, the process of performing three-dimensional topography measurement based on the preprocessed speckle dataset and the speckle matching network comprises:

[0015] Feature extraction is performed on the preprocessed speckle dataset to obtain unary features.

[0016] Based on the unary features, a four-dimensional cost volume is constructed using a group correlation method and a feature concatenation method.

[0017] The four-dimensional cost volume is aggregated to obtain a four-dimensional tensor.

[0018] The four-dimensional tensor is converted into a possibility distribution along the parallax dimension to obtain a disparity map.

[0019] Edge refinement is performed on the disparity map.

[0020] The disparity map is reconstructed into a three-dimensional point cloud, and three-dimensional topography measurement is performed based on the three-dimensional point cloud.

[0021] Optionally, in the process of performing feature extraction on the preprocessed speckle dataset, the feature tensor of the speckle dataset is downsampled by 1 / 4.

[0022] Optionally, in the process of aggregating the four-dimensional cost volume, a coarse disparity map is output based on the pre-aggregation module as a reference, aggregation is performed based on the three-dimensional hourglass subnetwork, and low-dimensional information is introduced based on the pre-aggregation module to provide parameters for the aggregation process.

[0023] Optionally, in the process of performing edge refinement on the disparity map, edge information of a left surface map is introduced based on the edge refinement module to guide the edge refinement of the disparity map, and missing information in the disparity map is predicted based on the edge information of the left surface map.

[0024] In another aspect to achieve the above object, the application provides a high-precision three-dimensional topography measurement system based on a speckle matching network, comprising a dataset acquisition device and a high-performance GPU.

[0025] The dataset acquisition device is used to acquire a speckle dataset of the object to be measured, and a speckle database is constructed based on the speckle dataset.

[0026] The high-performance GPU is used to perform three-dimensional topography measurement based on the speckle dataset.

[0027] Optionally, the dataset acquisition device comprises a binocular camera and a speckle projection module.

[0028] The binocular camera is used to acquire left and right surface images of the object to be measured.

[0029] The speckle projection module comprises a vertical cavity surface emitting laser, a projection lens and a photolithographic speckle mask, and is used to construct a speckle dataset based on the left and right surface images.

[0030] The technical effects of the present application are as follows:

[0031] The present application proposes a speckle matching network based on the DSC method, which can realize fast, high-precision and edge information protection three-dimensional topography measurement. In order to make the network measurement more reliable, a high-precision speckle binocular dataset containing thousands of data is collected by using a dedicated speckle projection device. The proposed end-to-end network can realize plane measurement error less than 0.03mm using only a single frame of speckle image, and the measurement standard deviation of complex surface is about 0.088mm. The inference speed of a single frame is about 0.13s. Compared with the DSC method, the experimental results show that the method proposed in the present application can solve the two main problems of the traditional method under the premise of ensuring the measurement accuracy. Compared with other learning methods, the proposed method reduces the EPE error from 0.2569 to 0.0481 pixels, and the training parameters and inference time are also greatly reduced. BRIEF DESCRIPTION OF DRAWINGS

[0032] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The accompanying drawings do not constitute an inappropriate limitation on the present application. In the drawings:

[0033] Figure 1 A high-precision three-dimensional topography measurement method flowchart in the embodiment of the present application;

[0034] Figure 2 A speckle dataset schematic diagram in the embodiment of the present application;

[0035] Figure 3 A speckle matching neural network structure diagram in the embodiment of the present application;

[0036] Figure 4 An edge thinning process schematic diagram in the embodiment of the present application;

[0037] Figure 5 A three-dimensional reconstruction comparison chart of the method and the traditional DSC method in the embodiment of the application;

[0038] Figure 6 A three-dimensional reconstruction result comparison chart of three learning-based model outputs in the embodiment of the application;

[0039] Figure 7 A three-dimensional deviation distribution comparison chart of four methods in the embodiment of the application; wherein (a) is a three-dimensional deviation distribution chart of the traditional DSC method, (b) is a three-dimensional deviation distribution chart of the Adjusted-GCNet method, (c) is a three-dimensional deviation distribution chart of the 3D-MobileStereoNet method, and (d) is a three-dimensional deviation distribution chart of the method of the application;

[0040] Figure 8 A disparity map result comparison chart of the application and non-application of an edge refinement module in the embodiment of the application;

[0041] Figure 9 A data set collection device schematic diagram in the embodiment of the application. DETAILED DESCRIPTION

[0042] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0043] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0044] Embodiment one

[0045] As shown in Figures 1-9 , the present embodiment provides a high-precision three-dimensional topography measurement method and system based on a single-frame speckle matching network.

[0046] The method of the present embodiment can generate a fast, high-precision and edge-protected disparity map through high-performance GPU training. A dedicated speckle projection module projects a speckle pattern onto the surface of the object to be measured, and the left and right images can be simultaneously collected by the binocular camera. The obtained speckle image pair is first limited to a one-dimensional region by limiting the search range of the corresponding points through rectification. Then, the rectified image is input into the end-to-end network proposed in the present embodiment for disparity map prediction. According to the principle of triangulation, the disparity map can be reconstructed into the corresponding three-dimensional point cloud. Figure 1The flowchart of the method is shown. In deep learning tasks, it is very important to build a high-precision dataset. The following text mainly describes the basic principle of the DSC method, uses the sub-pixel precision disparity map generated by it as the training ground value, and describes the composition of the dataset.

[0047] DSC constructs the relationship between the target sub-region in the right image and the reference sub-region in the left image using a mapping function, which is usually called a shape function.

[0048]

[0049] Where (x,y) and (x′,y′) are the coordinates of the midpoints of the reference subregion and the target subregion, respectively. p=(u,u x ,u y ,v,v x ,v y ) T This is a mapping parameter vector. The mapped coordinates (x′, y′) are usually sub-pixel level, requiring cubic B-spline interpolation for grayscale interpolation to improve accuracy.

[0050] The similarity between the reference and target sub-regions can be calculated using zero mean variance (ZNSSD), a criterion function that is insensitive to deformation and brightness variations. The Δp function constructed using ZNSSD can be expressed as:

[0051]

[0052] in f(x,y) and g(x′,y′) are the gray values ​​of corresponding points in the reference sub-region and the target sub-region, respectively. The size of the sub-region is (2M+1)×(2M+1). and It represents the average grayscale value of the two sub-regions. Δp is the incremental parameter, which can be iteratively optimized using the inverse Gauss-Newton algorithm.

[0053]

[0054]

[0055] Where H is the Hessian matrix. These are the grayscale gradient values ​​of the reference sub-region along the x and y axes. It is the Jacobian matrix of the shape function. The mapping parameter increment Δp is calculated, and the process is repeated until the iterative convergence condition is met, achieving sub-pixel matching between the reference sub-region and the target sub-region. The calculation result can be regarded as the ground truth of the speckle stereo dataset.

[0056] Constructing a high-precision speckle binocular dataset

[0057] The data set acquisition device contains two calibrated cameras and a projection module, and the working scene schematic diagram is shown in Figure 9 Unlike the structured light technique which usually needs a structured and expensive DLP, in this study, a dedicated speckle projection module is used to project high-quality speckle patterns. The speckle projection module contains a vertical cavity surface emitting laser (VCSEL), a projection lens and a photolithographic speckle mask, and the device is smaller in size and easier to control. Compared with the stripe projection technique which needs to project more than a dozen stripe patterns, the speckle data set acquisition only needs two speckle left and right images for each pair of data. Through the acquisition device, a high-precision speckle binocular database is established, which includes 1050 pairs of data (each group of data contains left image, right image and labeled disparity map), and the acquisition objects include some plaster figures and simple geometric bodies. The data set is shown in Figure 2 As shown in the figure. 800 pairs of data in the data set are divided for training, 200 pairs are used for verification, and the remaining 50 pairs of pictures are used for testing. Due to the limited field of view of the binocular camera, there are some invalid areas in the original pictures collected. Therefore, according to the principle of disparity calculation, the whole data set is transversely cropped as a pretreatment. The picture size input into the network is cropped from the original 1280x960 to 1024x960.

[0058] Speckle matching network for three-dimensional topographic measurement

[0059] The end-to-end speckle matching network is the focus of the embodiment. Although the neural network is not as completely step clear as the traditional mathematical algorithm, the proposed network is not a so-called black box structure, and each step of the network is designed with reference to the binocular matching principle and the structured light measurement principle. The network proposed in the embodiment contains four main modules: feature extraction module, cost volume construction module, three-dimensional aggregation module and edge refinement module. Figure 3 The overall network structure diagram is shown.

[0060] Network structure design for fast and high-precision measurement

[0061] Firstly, a pair of corrected speckle left and right images is input into the weight-shared twin network for preliminary feature extraction. Other stereo matching networks trained with public data sets usually have a large number of output channels (32 channels or 64 channels) in the convolution layer in the feature extraction stage, in order to extract different types of features from three-channel RGB images. However, too many unnecessary channels will lead to greater GPU consumption and longer calculation time, and usually the extracted channel information is repetitive. In this study, it is assumed that the object to be measured in the scene lacks texture information and has no relevant color attributes. The projected speckle pattern is also a grayscale image and does not contain color features. Therefore, Figure ThreeThe two-dimensional convolution module in the feature extraction stage reduces the number of channels to reduce memory occupation, but ensures the depth of feature extraction. The image pair after preliminary convolution processing continues to be convolved by seven residual blocks to expand the receptive field of each pixel output tensor. On this basis, high-dimensional semantic information is collected, which is more comprehensive and efficient than traditional artificial feature description. At the end of the feature extraction subnetwork, two dilated convolution blocks are set to expand the receptive field without increasing redundant parameters. Considering the huge amount of calculation and resource consumption in the subsequent steps, 1 / 4 downsampling operation must be performed on the feature tensor.

[0062] In the next stage, the unary features extracted by a series of convolution layers are used to construct a four-dimensional cost volume. The initial four-dimensional cost volume CxHxWxD (channel x height x width x disparity) is a collection of semantic features. Due to the limitations in actual measurement, there is a preset search range for disparity. The height and width are the size information of the feature map. The aggregation of features from different convolution layers ensures that the final feature map information is complete and representative, stacking them into a large 160-channel set and aggregating them later. Unlike direct three-dimensional convolution, this network uses a group correlation method. All feature channels are divided into different groups, and the correlation between each group is calculated. Setting the number of groups to 20 means that all feature channels are evenly divided into 20 groups along the channel dimension, with 8 channels in each feature group. In each group, group correlation is calculated according to formula 5

[0063]

[0064] where <.,.> is the dot product, f l and f r are the left and right features, respectively. g is the number of feature groups, and d is the disparity search range. The group correlation method is used to construct the cost volume because if only the correlation of the feature map is calculated, the semantic information of the left and right images will be lost. The correlation calculation of the features is based on the correlation calculation of each single channel in the disparity, which lacks the semantic information above and below. The group correlation method not only uses correlation but also optimizes the feature concatenation method, which is commonly used in other matching networks. However, if only the concatenation method is used like other networks, the cost volume will lose the information of feature similarity, which will require more parameters and computational complexity to learn the similarity of the previous features in the subsequent aggregation network. The construction method of Gwc combines the advantages of the above two methods to ensure low parameter quantity and high matching accuracy, which is suitable for the network proposed in this paper. The number of concatenated channels is set to 12, and the final cost volume is [N g ,D max / 4, H / 4, W / 4], where the denominator is 4 because of the downsampling operation in the previous step.

[0065] In the cost aggregation module, the 4D cost volume constructed previously is aggregated by using a 3D convolutional layer. The aggregation of features across three dimensions (height, width, and disparity) can maximize the use of geometric and semantic information. Although the initial resolution has been reduced by 1 / 4 in the feature extraction module, this 3D convolution operation still occupies a large amount of GPU memory. Other 3D aggregation modules in matching networks usually stack a large number of hourglass-type 3D modules to better learn semantic features, but in this paper, considering the inference speed and GPU occupancy, a large number of 3D convolution modules cannot be stacked in the case of a large resolution input image (1024x960). Based on this, an additional pre-aggregation module is proposed to output a coarse disparity map as a reference. As can be seen in the loss function formula (6), the tensor is then input into the only lightweight 3D hourglass aggregation module. The pre-aggregation module will be connected to the subsequent modules to introduce low-dimensional information. Low-dimensional information contains more global information, which can provide parameters for subsequent 3D aggregation. The cascaded output method can improve the accuracy. The final output is a 4D tensor in one channel, which is converted into a possibility distribution along the disparity dimension by using the SoftMax function. Such a network structure can maximize the simplification of 3D convolution parameters while ensuring the effectiveness of the aggregation step. In summary, through this simple and efficient structure, high-precision matching left and right images can be obtained.

[0066] Network structure design for edge information protection

[0067] There is a common shortcoming in both traditional DSC methods and classic binocular matching networks, which is that the predicted disparity map has information missing in areas with large edge and gradient changes. This problem can be appropriately solved by adding an edge refinement module. Deep learning can implicitly infer the disparity value of the data loss area according to global features, and is not affected by the lack of reference information. However, the unguided and unrestricted disparity result is chaotic at the edge, with a blurred edge lower than expected. To solve this problem, a module is designed to introduce the complete edge information of the left image to guide the edge refinement of the disparity map. The position of the disparity map is set according to the reference left image, so the outline of the disparity map should be consistent with the input left image. Therefore, the missing information in the disparity map can be inferred from the corresponding left image. The process of edge refinement is shown in Figure One Figure Four

[0068] The disparity regression loss function contains three parts, as shown in formula (6). In the joint loss function, information is iterated and optimized, and the smooth L1 excitation function is used for training.

[0069] Loss (joint) =α1Loss coarse +α2Loss​​disparity + a3Loss edge (6)

[0070]

[0071]

[0072] where a1, a2, a3 are the coefficients of each loss function. D represents the disparity. The original left image edge information is fused by a simple residual convolution block to correct the predicted disparity map. This structure is very simple and occupies little space, but effectively solves the problems of edge missing and information missing in large gradient transformation areas. In summary, this module performs edge perception and protection operations by integrating the left image and guiding information.

[0073] Experiments and Results

[0074] A binocular speckle projection measurement system was built to verify the effectiveness of the network proposed in this embodiment. Two binocular cameras (model: Basler acA1300-30gm, resolution: 1280x960) and two lenses (Computer 8mm 1:1.42 / 3) were used, and the baseline distance of the binocular camera was about 144mm. The dedicated single-frame speckle projection module has been described in detail above, and the distance between the measurement device and the measured object was about 300mm. Considering the actual application and the limitation of computing resources, the disparity range of our system was specified as 0-255 pixels, and the actual measurement range was within 275mm-370mm. All experiments were deployed on the Windows 10 platform, using Python language and Pytorch (Facebook) framework, and applied to the NVIDIA A100-PCIE-80GB. The Adam optimizer was used in the training process, and a total of 200 rounds of training were performed with an initial learning rate of (0.0001). The number of samples for one training was set to 1, and the GPU memory occupied during network training was 20GB.

[0075] Comparison Experiment with DSC Method

[0076] The compared DSC algorithm was also deployed on a high-performance GPU, and the details of the algorithm have been described in detail in Section 2.1. Although the DSC method can reconstruct the three-dimensional surface with high precision, it still has essential problems in application. Due to the limitations of binocular principle and the requirement of initial seed point estimation of DSC method, the calculated disparity map has the phenomenon of information missing in the edge and large gradient transformation area, for example Figure 5 as shown.

[0077] The deep learning-based method proposed in this embodiment can overcome the deficiencies in data integrity. Based on the information surrounding the missing data and the left image of the original input, the lost disparity information can be reasonably inferred. The comparison of the disparity maps of the two object classes and the corresponding reconstructed 3D models are shown below. Figure Six As shown in the figure, the disparity map (after removing invalid background information) of the plaster cast figure calculated by the DSC method has missing information in the hair, nose, and other edge areas, as indicated by the red boxes in the figure. In contrast, our method can effectively predict the disparity values ​​in these areas.

[0078] Although data gaps may occur, the DSC method still maintains high accuracy and can be considered ground truth in the recovered data. A complete portion of the data from the DSC method is selected, as shown in the blue area of ​​Disparity Map:2. Comparing the disparity values ​​in the blue area between the two methods, the average deviation is 0.0224 pixels, demonstrating that the two methods are of the same order of magnitude in terms of accuracy. This experiment confirms that the proposed network can achieve sub-pixel accuracy prediction of dense disparity maps.

[0079] Comparative experiments with other deep learning methods

[0080] Two representative networks were selected for comparative experiments with the network proposed in this paper. GC-Net is a classic binocular stereo matching network that constructs a cost volume by concatenating each unary feature and uses 3D convolutional operations for feature aggregation and filtering. This model has a very high computational cost and is difficult to apply directly to a 1024×960 resolution. Therefore, some minor adjustments need to be made to the number of channels and the network structure. We name this method Adjusted-GCNet. MobileStereoNet is a lightweight network that uses 2D or 3D MobileNet blocks to reduce computational cost while maintaining accuracy. However, this model mainly uses depthwise separable convolutions, which actually consumes more GPU memory during training. It is worth noting that the disparity map calculated by the DSC method is considered the ground truth due to its high absolute accuracy, and the points compared are all valid points. Under the same experimental conditions, the 3D shape reconstruction results of two plaster portrait models were predicted using Adjusted-GCNet, 3D-Mobile StereoNet, and the network proposed in this paper, respectively. The results are shown in [the table / image]. Figure 6 In the 3D model diagram, it is clear that the surfaces recovered by the other two learning-based methods are rougher and have more errors, while our results show a smoother surface and perform better in terms of data integrity and accuracy.

[0081] In addition, we conducted numerical evaluations including point-to-point error (EPE), root mean square error (RMSE), and the probability of an error greater than 1 pixel (>1px). These can all evaluate the similarity to the ground truth, and Table 1 shows the comparison of these data. Our method has an EPE of 0.0481, which is significantly lower than the other two methods. Furthermore, the error greater than 1px is only 2.80% of the ground truth, and the RMSE is also an order of magnitude lower than the other two networks. It can be inferred that the speckle matching network proposed in this paper performs better than the other two matching networks in terms of accuracy and data integrity.

[0082] Table 1

[0083]

[0084] Comparison experiment of inference time and computational cost

[0085] This section primarily discusses computational costs. The computational cost of MACs (Multiply-Accumulated MACs), the number of training parameters, and the GPU memory usage are all important considerations. In particular, inference time determines whether this method can achieve fast or even real-time measurement. Table 2 presents a comparison of the inference time and computational costs of the four methods.

[0086] Table 2

[0087]

[0088] Because the DSC method requires continuous iterations to reach the set threshold range, the entire computation is very time-consuming. Table 2 clearly shows that the proposed method takes less than 2% of the time of the traditional DSC method, achieving rapid measurement. Compared to Adjusted-GCNet and 3D-Mobile StereoNet, this method reduces computational parameters by 10.9 times and 8.17 times, respectively. While the structure of 3D-Mobile StereoNet effectively reduces the number of parameters and the GPU usage during testing, our network can also be considered a lightweight network in terms of various parameters. Furthermore, inference time is more important in practical applications; inference time is the average prediction time for a single frame of a 1024×960 resolution disparity map. This method has the greatest time advantage, requiring only 0.12-0.13 seconds per frame, exceeding the 5-6 seconds of the traditional DSC method and the 3-4 seconds of Adjusted-GCNet. In summary, the proposed method can effectively improve the speed of disparity map prediction.

[0089] Measurement accuracy estimation

[0090] In this section, we measure a precision-manufactured workpiece to verify the absolute accuracy of the method. The complex workpiece surface includes curved surfaces and three planes at different heights.

[0091] Four methods were used to predict the disparity map of the workpiece, which was then transformed into a 3D point cloud in the corresponding world coordinate system. The results of the digital-model matching were all placed in the same coordinate system, and the 3D deviation distribution map was displayed. Figure 7 The error plot shows that the deviations produced by both the DSC method and the method proposed in this invention are very small, verifying the absolute accuracy of the measurement.

[0092] The global bias data of the model are listed in Table 3. The point cloud reconstructed by DSC has the highest matching degree with the digital model, and the global bias value is also the smallest. At the same time, the method proposed in this invention also performs well on multiple datasets, with a standard deviation only 0.007 mm smaller than that of the DSC method, which can also meet the measurement needs of most industries.

[0093] Table 3

[0094]

[0095]

[0096] Because the model has a complex structure, including protrusions and depressions, the standard deviation of the global measurement is relatively large. However, if the flatness of the three planes is measured individually, the method proposed in this invention can achieve a standard deviation of 0.015 mm. The measurement results of the standard deviations of the three planes are shown in Table 4.

[0097] Table 4

[0098]

[0099] ablation experiment

[0100] Based on previous experiments, it can be seen that both the group correlation method and the edge refinement sub-network improve the efficiency and accuracy of the network. Therefore, ablation experiments verify the rationality of the network structure proposed in this invention in these two aspects. To explore the effectiveness of the Gwc method in constructing the cost body, a control experiment was conducted by directly applying concatenated features without the group correlation method to construct the cost body. Table 5 demonstrates that the network with group correlation significantly reduces network parameter values ​​and GPU usage. The inference time per frame also decreased from 2-3s to 0.12-0.13s.

[0101] Table 5

[0102]

[0103]

[0104] To verify whether the edge refinement module can improve accuracy and enhance edge sharpness of the predicted disparity map, a comparison of two disparity maps with and without the edge refinement module is presented. Figure 8 In this set of comparison images, the background of the parallax map is preserved, and only the region of interest is extracted, which makes the contrast more obvious.

[0105] Depend on Figure 8 It can be seen that the entire disparity map without edge refinement is more blurry, and the edges are less clear compared to the result. Introducing information from the corresponding left image can reasonably adjust the edge disparity, and overall, the matching error rate is smaller compared to the ground truth. From these two ablation experiments, it can be inferred that the network structure proposed in this paper is suitable for binocular speckle 3D topography measurement tasks.

[0106] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A high-precision three-dimensional topography measurement method based on a single-frame speckle matching network, characterized in that, The method comprises the following steps: obtaining a left surface image and a right surface image of an object to be measured; constructing a relationship between a target sub-region in the right surface image and a reference sub-region of the left surface image based on a mapping function; constructing a speckle dataset based on the left surface image and the right surface image, and preprocessing the speckle dataset; constructing a speckle matching network, and performing three-dimensional topography measurement based on the preprocessed speckle dataset and the speckle matching network; wherein the speckle matching network comprises a feature extraction module, a cost volume construction module, a three-dimensional aggregation module, and an edge refinement module; two hollow convolution blocks are arranged after the network of the feature extraction module; the three-dimensional aggregation module comprises a three-dimensional hourglass sub-network and a pre-aggregation module; the process of performing three-dimensional topography measurement based on the preprocessed speckle dataset and the speckle matching network comprises: extracting features from the preprocessed speckle dataset to obtain unary features; constructing a four-dimensional cost volume based on the unary features using a group correlation method and a feature concatenation method; aggregating the four-dimensional cost volume to obtain a four-dimensional tensor; transforming the four-dimensional tensor into a possibility distribution along the parallax dimension to obtain a disparity map; refining the edges of the disparity map; reconstructing the disparity map into a three-dimensional point cloud, and performing three-dimensional topography measurement based on the three-dimensional point cloud.

2. The single-frame speckle matching network based high precision three-dimensional profile measurement method according to claim 1, characterized in that, In the process of extracting features from the preprocessed speckle dataset, the feature tensor of the speckle dataset is down-sampled by 1 / 4.

3. The single-frame speckle matching network based high precision three-dimensional profile measurement method according to claim 1, characterized in that, In the process of aggregating the four-dimensional cost volume, a rough disparity map is output based on the pre-aggregation module as a reference, aggregation is performed based on the three-dimensional hourglass sub-network, and low-dimensional information is introduced based on the pre-aggregation module to provide parameters for the aggregation process.

4. The single-frame speckle matching network based high precision three-dimensional profile measurement method according to claim 1, characterized in that, In the process of refining the edges of the disparity map, edge information of the left surface image is introduced based on the edge refinement module to guide the edge refinement of the disparity map, and missing information in the disparity map is predicted based on the edge information of the left surface image.

5. A high-precision three-dimensional topography measurement system based on a single-frame speckle matching network, characterized in that, The system for implementing the method of any one of claims 1-4 comprises a dataset acquisition device and a high-performance GPU; the dataset acquisition device is used to obtain a speckle dataset of an object to be measured, and construct a speckle database based on the speckle dataset; the high-performance GPU is used to perform three-dimensional topography measurement based on the speckle dataset.

6. The single-frame speckle matching network based high precision 3D profile measurement system according to claim 5, characterized in that, The dataset acquisition device comprises a binocular camera and a speckle projection module; the binocular camera is used to obtain a left surface image and a right surface image of an object to be measured; the speckle projection module comprises a vertical cavity surface emitting laser, a projection lens, and a photolithographic speckle mask, and is used to construct a speckle dataset based on the left surface image and the right surface image.