A structured light three-dimensional measurement method based on an absolute phase end-to-end prediction network

By constructing the F2PWholeNet network and imitating the phase solution process of the physical model, the complexity and efficiency problems of grating-to-absolute phase prediction in deep learning methods are solved, and high-precision and robust absolute phase prediction is achieved.

CN118328900BActive Publication Date: 2025-10-10TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410425994.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-10
Estimated Expiration
2044-04-10

AI Technical Summary

Technical Problem

Existing deep learning methods have problems with complex training process and limited improvement in measurement efficiency in the prediction from single-frame grating to absolute phase, and are unable to effectively learn the long linear mapping details of multiple links in traditional phase solution methods.

Method used

An absolute phase end-to-end prediction network F2PWholeNet is constructed. Through the numerator graph branch network, denominator graph branch network, wrapped phase graph branch network, order graph branch network, unwrapping module and trunk branch, the phase solution process in the physical model is imitated, and the multi-frequency N-step phase shift method and multi-frequency unwrapped phase method are used for supervised training.

Benefits of technology

The prediction accuracy and robustness from grating to absolute phase are improved, the complexity of the training and measurement process is reduced, and efficient pixel-by-pixel phase prediction is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118328900B_ABST
    Figure CN118328900B_ABST
Patent Text Reader

Abstract

A structured light three-dimensional measurement method based on an absolute phase end-to-end prediction network, comprising: S1, constructing an absolute phase end-to-end prediction network F2PWholeNet, which comprises a molecule graph branch network, a denominator graph branch network, a wrapped phase graph branch network, a level graph branch network, a first to third fusion module, an unwrapping module and a backbone branch to imitate the phase solution process in the physical model; wherein: S2, training the F2PWholeNet network so that it can learn the mapping relationship from the grating graph to the absolute phase at one time; S3, using the trained F2PWholeNet network to predict the absolute phase of the new grating graph to realize high-precision three-dimensional shape reconstruction. F2PWholeNet can greatly improve the prediction accuracy of the absolute phase P(F2P) from the grating F due to the process supervision of the absolute phase prediction of multiple links of the whole traditional method physical process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to structured light three-dimensional measurement, and in particular to a structured light three-dimensional measurement method based on an absolute phase end-to-end prediction network. Background Art

[0002] Structured light 3D measurement technology has been widely used in industry and scientific research due to its advantages such as simple hardware configuration, high accuracy, high point density, fast speed, and low price. In recent years, as the global manufacturing industry transitions to intelligent manufacturing, and faced with the current state of traditional structured light technology's inability to effectively balance multiple performance indicators such as speed and accuracy while simultaneously improving performance, a large number of researchers have developed structured light 3D measurement technology based on deep learning, achieving performance far exceeding that of traditional methods in multiple indicators such as speed and accuracy. Fringe projection profilometry is the most widely used method in structured light 3D measurement technology. The measurement process can be divided into three major steps: absolute phase solution, system calibration parameter acquisition, and 3D reconstruction using the combination of absolute phase and system calibration parameters. The prediction of absolute phase from gratings is a current research hotspot in the field of structured light deep learning, as it conforms to the processing paradigm of computer vision, which has a broad research foundation.

[0003] In the traditional method, the standard N-step phase shift method solves the numerator and denominator in turn, calculates the wrapped phase by the arctangent function, and combines the multi-frequency phase shift method (or multi-frequency heterodyne method) to obtain the absolute phase by the order relationship between different frequencies. It has the advantages of high precision and high stability, and is the mainstream method of absolute phase calculation. The methods of combining deep learning for absolute phase prediction mainly include the following four categories: the first category is to predict the numerator and denominator of the wrapped phase by a deep learning network for a single frame of grating, and then perform division operation based on the traditional method or another deep learning network to calculate the wrapped phase or absolute phase. The second category is to predict other step number gratings of different frequencies and the same frequency by a deep learning network for a single frame of grating, and then calculate the absolute phase based on the traditional method after obtaining the predicted complete multi-frequency N-step grating. The third category is to predict the wrapped order by a deep learning network for grating, and then calculate and predict the absolute phase based on the predicted wrapped order and the traditional method or other deep learning network. The fourth category is to directly predict the absolute phase by a deep learning network for a single frame of grating. From the current research status of deep learning technology in structured light three-dimensional measurement, because of the long-distance ill-posed problem from grating to absolute phase prediction, most people focus on the first three categories in the task of single frame input and absolute phase acquisition, that is, to solve the absolute phase by using multiple networks, multiple training processes, and deep learning methods combined with traditional algorithms. The disadvantages of these three methods are complex training process and limited efficiency improvement from grating to absolute phase calculation. For the fourth category of research, calculating the absolute phase from a single frame of grating is a long-distance ill-posed problem, and deep learning networks such as Unet and HRnet based on encoder-decoder structure can only learn relatively general prediction features in a single frame of grating image, and cannot learn the long-thread nonlinear mapping details of multiple links in the traditional phase calculation method.

[0004] How to make the network learn the multi-link physical model mechanism and semantic information from a single frame of grating to absolute phase, and far better than the existing end-to-end image prediction regression task model relying on computer vision basic paradigm is a challenging problem that has not been solved.

[0005] It should be noted that the information disclosed in the above background section is only for understanding the background of the present application, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] The main purpose of the present application is to overcome the defects of the above background technology, and provide a structured light three-dimensional measurement method based on an absolute phase end-to-end prediction network.

[0007] To achieve the above purpose, the present application adopts the following technical solutions:

[0008] A structured light 3D measurement method based on an absolute phase end-to-end prediction network, comprising:

[0009] S1. Construct an end-to-end absolute phase prediction network F2PWholeNet, which includes a numerator graph branch network, a denominator graph branch network, a wrapped phase graph branch network, a level graph branch network, the first to third fusion modules, an unwrapping module, and a trunk branch to simulate the phase solution process in the physical model; wherein:

[0010] The molecular graph branch network is supervised by the true molecular graph and predicts the molecular graph from the input raster image;

[0011] The denominator graph branch network is supervised by the true value denominator graph and predicts the denominator graph from the input raster image;

[0012] The first fusion module fuses a selected input combination, the input combination being selected from: a predicted numerator graph, a denominator graph, and an original input raster graph;

[0013] The wrapped phase image branch network is supervised by the true wrapped phase image and predicts the wrapped phase image using a selected input combination, the input combination being selected from: a grating image and an output of the first fusion module;

[0014] The second fusion module fuses a selected input combination, the input combination being selected from: the predicted wrapped phase map and the output of the first fusion module;

[0015] The level graph branch network is supervised by the ground truth parcel level graph and predicts the parcel level graph using a selected input combination, the input combination being selected from: a raster image and an output of the second fusion module;

[0016] The unwrapping module uses the predicted wrapped phase map and the wrapped order map to perform phase solution to obtain the predicted absolute phase;

[0017] The third fusion module fuses a selected input combination, the input combination being selected from: the predicted parcel level map, the predicted absolute phase, and the output of the second fusion module;

[0018] The main branch, i.e., the absolute phase branch, is supervised by the true absolute phase image and predicts the absolute phase image using a selected input combination, the input combination being selected from: the output of the third fusion module;

[0019] S2. Train the F2PWholeNet network to enable it to learn the mapping relationship from grating image to absolute phase in one go;

[0020] S3. Use the trained F2PWholeNet network to perform absolute phase prediction on the new grating image to achieve high-precision 3D shape reconstruction.

[0021] Furthermore, one or more of the following operations are met:

[0022] The first fusion module fuses the predicted numerator graph and denominator graph with the original input raster graph;

[0023] The wrapped phase image branch network is supervised by the true wrapped phase image and predicts the wrapped phase image using the grating image and the output of the first fusion module;

[0024] The second fusion module fuses the predicted wrapped phase map and the output of the first fusion module;

[0025] The level graph branch network is supervised by the true parcel level graph and predicts the parcel level graph using the raster image and the output of the second fusion module;

[0026] The unwrapping module uses the predicted wrapped phase map and the wrapped order map to perform phase solution to obtain the predicted absolute phase;

[0027] The third fusion module fuses the predicted package level map, the predicted absolute phase and the output of the second fusion module;

[0028] The main branch, namely the absolute phase branch, is supervised by the true absolute phase image, and the absolute phase image is predicted based on the output of the third fusion module and the grating image.

[0029] Furthermore, the unwrapping module performs phase calculation according to the following formula:

[0030] Unwrap=O3+2π*O4

[0031] Among them, Unwrap is the predicted absolute phase, O3 is the predicted output of the wrapped phase map, and O4 is the predicted output of the wrapped level map.

[0032] Furthermore, a multi-frequency N-step phase shift method is adopted when constructing a data set for training the F2PWholeNet network, and a grating image of any phase in the highest-frequency grating image is used as input data, and the high-frequency numerator image, high-frequency denominator image, high-frequency wrapped phase image, wrapped order image and absolute phase image corresponding to the input data are used as supervision signals of the numerator image branch network, the denominator image branch network, the wrapped phase image branch network, the order image branch network and the main branch respectively.

[0033] Furthermore, the data input-truth value pairs are solved using the twelve-step phase shift method and the multi-frequency unwrapping phase method.

[0034] Furthermore, the F2PWholeNet network adopts the following loss function:

[0035]

[0036]

[0037] Loss n =L1loss+αL2loss

[0038]

[0039] Among them, the mean absolute error L1loss and the mean square error L2loss are used as the basic loss functions of the branch, y gt and y i Represents the true value and prediction results of each network, Loss n is the loss function of the branch network n, the single-path basic loss function and the total loss function are Loss n and Loss_total, n represents the number of branches, α is the hyperparameter of the basic loss function, a n are the hyperparameters of the five loss functions, and n ranges from 1 to 5.

[0040] A computer-readable storage medium stores a computer program, which implements the method when executed by a processor.

[0041] A computer program product is provided, wherein the computer program product is executed by a processor to implement the method described above.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] This paper provides a structured light 3D measurement method based on an end-to-end absolute phase prediction network (F2PWholeNet) with full-process supervision of physical models. The network is designed to learn a multi-step physical model from grating images to absolute phase, significantly outperforming existing end-to-end image prediction and regression models. In the prediction network, four branches are sequentially integrated with the grating according to the physical process as input to the next branch, effectively improving the problems of insufficient effective information, high prediction difficulty, and long distance in end-to-end prediction from a single grating to absolute phase.

[0044] The traditional N-step phase shift plus multi-frequency phase shift method requires at least 6 frames of grating images (two frequencies, three phase shift images for each frequency) to solve the wrapped phase. Different phase shift images of the same frequency contain phase information of independent temporal continuity between pixels. Different frequency images contain spatial information of the object's wrapped phase expansion. Therefore, a single frame of grating cannot directly obtain an accurate absolute phase image in terms of phase temporal continuity and spatial continuity. Deep learning prediction directly from a single frame of grating to absolute phase is an extremely challenging task. The current end-to-end absolute phase prediction has limited effect and large errors. The present invention uses four branches to imitate the physical model as the intermediate process supervision from grating to absolute phase prediction. The numerator branch and denominator branch supervision help the network learn as much as possible of the phase time information and spatial information contained in the grating of other steps in the dataset construction process, compensate for the end-to-end prediction error of the wrapped phase map, and provide more robust semantic information and phase information to the prediction of the wrapped phase map. The accurate prediction of the wrapped phase can guide the network parameters to suppress noise, enhance the network's robustness to ambient light and object surface reflectivity, and also increase the network's ability to predict pixel-by-pixel phase. The unfolding level branch supervision can help the network unfold the phase at the high-frequency information of the object, guiding the network's ability to unfold pixel by pixel.

[0045] In summary, the multi-channel supervised network structure of the present invention increases the robustness, accuracy, and reliability of deep learning networks for absolute phase prediction, and its end-to-end prediction features make the pixel-by-pixel prediction process fast and efficient. The F2PWholeNet of the present invention requires only a single training phase, eliminating the need for multiple training phases and the need for stage-by-stage prediction during the measurement process, greatly improving training and measurement efficiency. Its main advantages are:

[0046] 1. Compared with the existing single-branch network from grating to absolute phase, F2PWholeNet can significantly improve the prediction accuracy from grating F to absolute phase P (F2P) by integrating multiple links of the entire physical process of traditional methods to supervise the absolute phase prediction process.

[0047] 2. Compared with the current method of sequentially calculating and predicting a single branch and then concatenating the results to resolve the absolute phase, F2PWholeNet greatly improves the training and prediction efficiency of F2P due to its end-to-end strategy.

[0048] 3. The F2PWholeNet strategy has no restrictions on the network itself. The network of each branch can be any network. It has strong applicability and inclusiveness and can integrate and use all effective networks in the existing computer vision field.

[0049] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 Figure 1 is a network structure diagram of F2PwholeNet according to an embodiment of the present invention, including: (a) network structure diagram; (b) introduction to the input and output of fusion module C1; (c) introduction to the input and output of fusion module C2; ​​(d) introduction to the input and output of fusion module C3; (e) implementation of the unwrap module; (f) introduction to the specific meaning of true value supervision; and (d) introduction to the specific meaning of branch network output.

[0051] Figure 2 This is an example display of a dataset image pair for an embodiment of the present invention, wherein (a) the object being measured; (b) the network input map; (c) the network truth map; (d) the sub-network 1 supervision map - the numerator map; (e) the sub-network 2 supervision map - the denominator map; (f) the sub-network 3 supervision map - the wrapped phase map.

[0052] Figure 3 This is the ResUnet network and core feature extraction module of an embodiment of the present invention.

[0053] Figure 4 2 is a comparative experimental scheme diagram of an embodiment of the present invention.

[0054] Figure 5 Comparison of the MAE prediction results of the F2PWholeNet strategy and the single-branch ResUnet strategy of the embodiment of the present invention, where (a) F2PWholeNet predicted MAE; (b) single-branch ResUnet predicted MAE result.

[0055] Figure 6 This is a flowchart of a structured light three-dimensional measurement method based on an absolute phase end-to-end prediction network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.

[0057] See Figure 6 The embodiment of the present invention provides a structured light three-dimensional measurement method based on an absolute phase end-to-end prediction network, comprising:

[0058] S1. Construct an end-to-end absolute phase prediction network F2PWholeNet, which includes a numerator graph branch network, a denominator graph branch network, a wrapped phase graph branch network, a level graph branch network, the first to third fusion modules, an unwrapping module, and a trunk branch to simulate the phase solution process in the physical model; wherein:

[0059] The molecular graph branch network is supervised by the true molecular graph and predicts the molecular graph from the input raster image;

[0060] The denominator graph branch network is supervised by the true value denominator graph and predicts the denominator graph from the input raster image;

[0061] The first fusion module fuses a selected input combination, the input combination being selected from: a predicted numerator graph, a denominator graph, and an original input raster graph;

[0062] The wrapped phase image branch network is supervised by the true wrapped phase image and predicts the wrapped phase image using a selected input combination, the input combination being selected from: a grating image and an output of the first fusion module;

[0063] The second fusion module fuses a selected input combination, the input combination being selected from: the predicted wrapped phase map and the output of the first fusion module;

[0064] The level graph branch network is supervised by the ground truth parcel level graph and predicts the parcel level graph using a selected input combination, the input combination being selected from: a raster image and an output of the second fusion module;

[0065] The unwrapping module uses the predicted wrapped phase map and the wrapped order map to perform phase solution to obtain the predicted absolute phase;

[0066] The third fusion module fuses a selected input combination, the input combination being selected from: the predicted parcel level map, the predicted absolute phase, and the output of the second fusion module;

[0067] The main branch, i.e., the absolute phase branch, is supervised by the true absolute phase image and predicts the absolute phase image using a selected input combination, the input combination being selected from: the output of the third fusion module;

[0068] S2. Train the F2PWholeNet network to enable it to learn the mapping relationship from grating image to absolute phase in one go;

[0069] S3. Use the trained F2PWholeNet network to perform absolute phase prediction on the new grating image to achieve high-precision 3D shape reconstruction.

[0070] Direct prediction from grating to absolute phase is a current research hotspot in the field of structured light deep learning. Existing regression networks such as Unet and HRnet can only learn the relatively general prediction features of a single frame image itself, and cannot learn the long-distance nonlinear mapping details of traditional methods. Based on the physical model of the absolute phase solution of the traditional multi-frequency phase shift algorithm and the multi-frequency heterodyne method, the present invention constructs a new network and construction strategy for absolute phase prediction that includes supervision of the entire physical process - F2PWholeNet. The core of this network structure strategy is to add four branches (the numerator graph is the supervision branch, the denominator graph is the supervision branch, the wrapped phase is the supervision branch, and the order graph is the supervision branch) on the basis of the existing end-to-end process from grating F to absolute phase P prediction (F2P). The four branches are fused with the grating in turn according to the physical process as the input of the next branch, thereby effectively improving the problems of insufficient effective information, high prediction difficulty and long distance in the end-to-end prediction of single grating to absolute phase. Furthermore, for this network structure, the loss function for fusing each branch is optimized and designed. The F2PWholeNet of the present invention only requires one training phase, and does not need to be divided into multiple training phases for training in sequence. The measurement process also does not require stage-by-stage prediction, which greatly improves the training and measurement efficiency.

[0071] Specific embodiments of the present invention are further described below.

[0072] The end-to-end absolute phase prediction of structured light based on deep learning mainly includes four contents: the construction of the F2P dataset, the establishment of the network, the training of the network constrained by the loss function, and the measurement and verification of the test set based on the trained network parameters.

[0073] N-step phase shifting and multi-frequency unwrapping methods guide dataset construction

[0074] Accurate and effective true value supervision is the foundation of deep learning network prediction. Currently, in the field of structured light phase resolution, the absolute phase truth value is calculated using a twelve-step phase shift plus multi-frequency unwrapping phase method. The standard N-step phase shift method is one of the most widely used, most stable, and most accurate methods in structured light 3D measurement. The formula is as follows:

[0075]

[0076] Among them I n The nth (n∈[1,N]) raster image captured by the (x,y) camera, A(x,y) represents the background light intensity, B(x,y) represents the modulated light intensity determined by the reflectivity of the object being measured, represents the phase determined by the surface height of the object being measured, and N represents the total number of phase shift steps.

[0077] The molecular graph is represented as:

[0078]

[0079] The denominator graph is represented as:

[0080]

[0081] Wrap the phase and perform the inverse tangent operation on the numerator and denominator graphs to obtain them.

[0082]

[0083] After the inverse tangent operation, it is wrapped between [-π, π] and represented by φ(x, y).

[0084] Time phase unwrapping is used to recover the precise absolute phase. Multi-frequency time domain phase unwrapping has better reliability. Next, it is necessary to perform wrapped phase unwrapping on the wrapped phase φ(x,y). h (x,y),Φ l (x,y) are low-frequency absolute phase and high-frequency absolute phase respectively, φ h (x, y) and φ l (x, y) represent the high-frequency wrapping phase and the low-frequency wrapping phase, respectively, k h and k l They represent the high-frequency unwrapping level and the low-frequency unwrapping level respectively.

[0085]

[0086] Usually, k l is 0, φ l (x,y) is a single-cycle wrapped phase, so we only need to solve k h , k h The formula is as follows:

[0087]

[0088] Among them, λ l ,λ h Represent the pixel periods of the high-frequency grating and the low-frequency grating respectively, and Round represents the rounding operation. The absolute phase of the object under test is obtained through the above process.

[0089] In summary, the multi-frequency N-step phase shift method is used to solve the data input-ground truth pair. A single-frame grating image of any phase of the N-step phase shift image in the highest-frequency grating image is used as the input of the dataset. The high-frequency numerator image, high-frequency denominator image, high-frequency wrapped phase image, wrapped order image, and absolute phase image are used as supervision for branch 1, branch 2, branch 3, branch 4, and main branch 5, respectively.

[0090] F2PwholeNet network structure

[0091] Four branches mimicking physical models serve as intermediate supervision for the process from grating to absolute phase prediction. The numerator and denominator branches help the network learn as much as possible of the temporal and spatial phase information contained in the gratings at other steps in the dataset construction process, compensating for the end-to-end prediction error of the wrapped phase image and providing more robust semantic and phase information for the wrapped phase image prediction. Accurate prediction of the wrapped phase guides the network parameters' ability to suppress noise, enhances the network's robustness to ambient light and surface reflectivity, and improves the network's ability to predict pixel-by-pixel phase. The unwrapping-level branch supervision aids the network in phase unwrapping at high-frequency locations, guiding the network's pixel-by-pixel unwrapping capabilities. This multi-branch supervised network architecture increases the robustness, accuracy, and reliability of the deep learning network for absolute phase prediction, while the end-to-end prediction approach makes the pixel-by-pixel prediction process fast and efficient.

[0092] The process of the entire network structure strategy is as follows Figure 1 As shown, the following is a detailed introduction:

[0093] Branch 1: The input is a single-frame high-frequency grating, which completes the prediction from the grating to the molecular graph. The output O1 is the predicted molecular graph, which is supervised by the true molecular graph S1. Loss1 is the loss function of branch 1. Represents the direction and content of the branch network back propagation. Where a1 represents the weight parameter of the loss function importance.

[0094] Branch 2: The input is a single-frame high-frequency grating, and the output O2 is the predicted denominator map, which is supervised by the true value S2 denominator map. The branch loss function is Loss2, Represents the direction and content of the branch network back propagation. Where a2 represents the weight parameter of the loss function importance.

[0095] The function of branch 1 and branch 2 is to assist the parcel phase prediction to compensate and correct the difficulty and error of the long distance from single image to direct prediction of parcel phase.

[0096] C1 fusion module: see Figure 1 (b) The input of the fusion module is the numerator branch, the output of the denominator branch O1, O2 and the input grating. The output of the fusion module is fed into the wrapped phase branch, the order branch and the absolute phase branch.

[0097] Branch 3: The input is the output of sub-network 1 (predicted numerator map O1), the output of sub-network 2 (predicted denominator map O2) and the fusion of a single-frame high-frequency grating (after fusion module C1). Output 3 is the predicted wrapped phase map O3, which is supervised by the true wrapped phase S3. The branch loss function is Loss3. Represents the direction and content of the branch network back propagation. Where a3 represents the weight parameter of the loss function importance.

[0098] The precise wrapped phase supervision branch enables the network to obtain the phase information of all phase-shifted grating images at that frequency during training when only a single frame of grating image is used as input.

[0099] Branch 4: The input is the output of branch 3 (predicted wrapped phase map O3), a single frame high frequency grating, the output of the numerator branch, the output of the denominator branch O1, O2, after fusion module C2 (see Figure 1 (c) Perform multi-input feature fusion. Output 4 is the predicted parcel level map O4, which is supervised by the true parcel level S4. The branch loss function is Loss4, Represents the direction and content of the branch network backpropagation.

[0100] Branch 3 predicts the wrapping phase O3, and branch 4 predicts the wrapping level O4. O3 and O4 are unwrap (see Figure 1 (e)) module performs phase calculation to obtain the predicted absolute phase O5. The specific implementation of the Unwrap module is as follows:

[0101] Unwrap=O3+2π*O4

[0102] The F2PWholeNet strategy provides parcel-level branch supervision, which enables network training to obtain effective information on the fringe order, fundamentally improving the phase ambiguity problem, and allows each spatial pixel to be expanded independently, effectively improving the spatial discontinuity problem in the modulation phase expansion process.

[0103] Branch 5: Input is the numerator branch, the output of the denominator branch O1, O2, the output of sub-network 3 (predicted wrapped phase map O3), the output of sub-network 4 (predicted wrapped level map O4), the output of the unwrap module and the single-frame high-frequency grating after fusion module C3 (see Figure 1 (d)) is fused and the output is the predicted absolute phase map O5, which is supervised by the true absolute phase S5. The branch loss function is Loss5, Represents the direction and content of the branch network back propagation. Where a5 represents the weight parameter of the loss function importance.

[0104] The F2PWholeNet strategy is based on the full-process supervision of the physical model. Multiple branches predict different phase solution links, and the branch networks are trained collaboratively. Due to the consistency of the principles of the physical model, they guide each other and promote the accuracy of training. In addition, this strategy divides the high-difficulty long-distance prediction into multiple low-difficulty short-distance small branch predictions, which improves the accuracy, speed and robustness of the phase solution. The branch network can be any regression network. Compared with the existing end-to-end network from high-frequency grating to absolute phase prediction, the input of the backbone network 1 adds the necessary time phase information and spatial information required in the phase recovery process due to the addition of the predicted package phase map and the predicted package level map information and content, which reduces the difficulty of network training and improves the robustness and accuracy of the network output. It should be noted that although the above-mentioned network structure is divided into multiple branches, all branches constitute a large network, and the large network can be trained once. The present application also protects the model of training multiple sub-networks separately and then fusing them.

[0105] Loss function:

[0106] The loss function measures the difference between the model's prediction and the true value, helping the optimization algorithm adjust the model's parameters to more accurately predict the target variable. It is a crucial component of the training process. In a preferred embodiment, a proprietary loss function for a multi-branch physical supervision network structure is proposed.

[0107] The mean absolute error L1loss and mean square error L2loss are used as the basic loss functions of the branch (the loss function is not limited to the above two), y gt and y i Represent the true value and prediction result of each network respectively. The loss function of sub-network n is named Loss n , so the single-path basic loss function and total loss function of the total system are Loss n and Loss_total:

[0108]

[0109]

[0110] Loss n =L1loss+αL2loss

[0111]

[0112] Among them, n represents the number of branches, α is the hyperparameter of the basic loss function, and a n(n = 1…5) are the hyperparameters of the five loss functions. Based on Loss_total, the loss function is calculated at each training iteration, and the model parameters are updated using algorithms such as gradient descent to minimize the loss. Ultimately, the training and test Loss_total converge, completing the training. This results in the trained single-input, single-output network parameters—the F2PwholeNet weights.

[0113] During the verification process, simply inputting a single-frame high-frequency phase-shift grating into the trained network parameters of F2PwholeNet can obtain a high-precision, high-stability, and high-robust absolute phase map for subsequent phase-to-3D mapping, ultimately completing fast and high-precision 3D measurement.

[0114] Alternative embodiment:

[0115] 1. F2PWholeNet is not only a multi-branch network but also a network construction strategy. The branch network and the backbone network can be the same or different. It is not limited to using the same regression network, such as Unet, and can be applied to all regression networks.

[0116] 2. The fusion of grating and branch outputs (such as the fusion of grating and numerator and denominator outputs) is not limited to contact connections and can be other fusion modules.

[0117] 3. This physical model is not limited to the standard N-step phase shift, nor is it limited to the multi-frequency method for unwrapping the phase, and the number of branches is not limited to 4. Under other mathematical models and physical models for absolute phase solution, each link can be a branch to monitor the absolute phase process.

[0118] 4. The training strategy of F2PWholeNet is not limited to the use of all four branches. The four branches can be randomly combined for process supervision. For example, in addition to the numerator branch, denominator branch, and wrapped phase branch, the order branch can be combined with the absolute phase branch, and two or more branches can be combined with each other.

[0119] 5. One of the extension solutions of the present invention is to train multiple sub-branches separately. After the weight of each sub-branch is trained, the weights of multiple sub-branches trained separately are loaded to assist the backbone network training to obtain the absolute phase map.

[0120] 6. The input of each branch is not limited to Figure 1 The content of the network diagram can be any combination of the outputs of all previous links. For example, the input of the wrapped phase diagram can be [grating, numerator, denominator], [grating, numerator], [grating, denominator], or grating, etc. The input of each layer can be any combination of the output of each branch of the physical model and the output of other branches.

[0121] 7. The network input is not limited to a single-frame grating pattern. It can also be multiple grating images, or a combination of a grating image and a background intensity image, a modulation image, a level image, and so on. The type and combination of input images do not affect the protection of the network structure.

[0122] 8. The unwrap module is not limited to the formulas in this article. Various variations of the unwrap module can be implemented without adding the unwrap module, without affecting the protection of the network structure.

[0123] Examples

[0124] use Figure 1 The F2PWholeNet consists of four branches: the numerator branch, the denominator branch, the wrapped phase branch, and the absolute phase branch shown in FIG.

[0125] (1) Dataset Description

[0126] A real-world dataset was constructed using a Tengju Technology III projector with an 8-bit resolution of 1280×720 and a Balser aca2040-120um grayscale camera as a structured light 3D measurement device. Since absolute phase accuracy directly determines 3D measurement accuracy, this dataset only illustrates the improvement in absolute phase accuracy. A four-frequency, twelve-step phase shifting method was used to obtain the absolute phase truth. The dataset was constructed by measuring a large number of real-world objects. The training dataset contains ~1000 image pairs and the test dataset contains ~100 image pairs. The images are 640×512 pixels in size. The dataset includes over 200 plaster statues, wooden objects, and polystyrene artifacts, including basic geometric shapes (spheres, triangular pyramids, and prisms) and a rich assortment of human and animal figurines. The volume of the measured objects is within 10×10×10 cm. A single-frame high-frequency grating with a fringe frequency of 64 and an initial phase of 4π / 3 was selected as input for the real-world dataset. The high-frequency numerator graph, denominator graph, and wrapped phase graph serve as auxiliary supervision, and the absolute phase is used as the true value of the dataset. (The network demonstrated in this example is Figure 1 The F2PWholeNet strategy consists of the numerator branch 1, denominator branch 2, wrapped phase branch 3 and absolute phase branch 5). Figure 2 A set of data in the dataset is shown, including: (a) the object being tested; (b) the network input map; (c) the network truth map; (d) the supervision map of sub-network 1 - the numerator map; (e) the supervision map of sub-network 2 - the denominator map; (f) the supervision map of sub-network 3 - the wrapped phase map.

[0127] (2) Network Description

[0128] The sub-network uses ResUnet as Figure 3As shown in the figure, ResUnet retains the U-shaped structure of UNet, including encoder, decoder and skip connection operation (Concatenate), which allows features to be extracted from the input image and restored to the original size. The numbers in the figure represent the changes in the number of feature map channels. Max Pooling performs the maximum pooling operation to halve the size of the feature map; Conv1×1 convolution is used to change the number of feature map channels; Trans-Conv3×3 performs transposed convolution upsampling to double the size of the feature map; Resblock is the core feature extraction module of ResUnet. Feature map represents the feature map. In addition, the algorithm also combines skip connections to retain the texture information of the input image, thereby ensuring that each feature map before upsampling in the decoding stage contains more underlying features.

[0129] (3) Training parameter description:

[0130] Based on the above dataset and network, the Pytorch framework (version 2.0.1+cu117+

[0131] The network was trained with Python 3.9.13), two NVIDIA GeForce RTX 3090 GPUs, Adam optimizer, CosineAnnealingLR learning rate descent method, initial learning rate 0.01, 200 epochs, 8 batch size, and basic branch loss function.

[0132] Loss n =L1loss+0.5L2loss

[0133] Total loss function:

[0134]

[0135] Among them, Loss n (n = 1…4) represents the loss of different branches, and a1, a2, a3, and a4 represent the loss function weights of the numerator branch, denominator branch, wrapped phase branch, and absolute phase backbone network, respectively. Based on the numerical ratio between experience and the true branch values, a1, a2, a3, and a4 are set to 1, 1, 1, and 500, respectively.

[0136] (4) Test results:

[0137] Based on more than 80 test sets, the F2PWholeNet end-to-end network with three branches, numerator, denominator, and wrapped phase, was compared with the end-to-end network with a single ResUnet branch (see the structural diagram of the comparison strategy for details). Figure 4 ) to compare the experimental results. Figure 5, shows the comparison of MAE prediction results of F2PWholeNet strategy and single-branch ResUnet, where: (a) F2PWholeNet predicted MAE. (b) Single-branch ResUnet predicted MAE results. Figure 5 As can be seen, the end-to-end prediction results of the F2PWholeNet strategy are better than the MAE prediction results of the single-branch ResUnet. In addition, in more than 100 test sets, the average MAE of the F2PWholeNet strategy is 40% higher than the MAE of the single-branch ResUnet.

[0138] The significant advantages of this invention over traditional technologies are:

[0139] The present invention integrates the physical model of the absolute phase solution of the traditional multi-frequency phase shift algorithm and the multi-frequency heterodyne method, and constructs a proprietary strategy for constructing an end-to-end prediction network that includes supervision of the entire physical process of absolute phase solution - F2PWholeNet. The core of this network structure strategy is to add four branches (branch 1 for supervision by the numerator graph, branch 2 for supervision by the denominator graph, branch 3 for supervision by the wrapped phase, and branch 4 for supervision by the order graph) on the basis of the existing end-to-end process from grating F to absolute phase P prediction (F2P). The four branches are sequentially fused with the grating as the input of the next branch according to the physical model solution process, thereby effectively improving the problems of insufficient effective information, high prediction difficulty, and long distance in the end-to-end prediction from a single grating to absolute phase. Specifically:

[0140] Advantage 1:

[0141] During phase modulation, the standard N-step phase shifting method offers greater robustness, high resolution, and high accuracy. Compared to Fourier transform profilometry's single-frame modulation phase calculation, its performance at edges and discontinuous transitions is more accurate. In this context, the F2PWholeNet strategy provides a parcel phase supervision branch. Given only a single-frame grating image as input, the network receives phase information from all phase-shifted grating images at that frequency during training. Because the parcel phase supervision truth is obtained through twelve-step phase shifting, this truth guides the network during training to increase its suppression of harmonic noise, enhancing its tolerance and robustness to ambient light and surface reflectivity. This guides the network to account for phase variations in each pixel, enabling independent pixel-by-pixel calculation and prediction, improving overall phase accuracy and its ability to resolve high-frequency phase details.

[0142] Advantage 2:

[0143] Furthermore, the single wrapped phase prediction branch suffers from discontinuous jumps due to the division operation during the solution of the wrapped phase truth, making end-to-end prediction difficult for a single network. The auxiliary supervision of the numerator and denominator branches provides more robust phase information of the N-step phase-shifted grating. The combined phase and semantic information of the numerator and denominator reduces the difficulty of long-range semantic prediction for a single image and improves the accuracy of wrapped phase prediction.

[0144] Advantage 3:

[0145] The wrapped phase value after phase modulation is limited to the range [-π, π]. Therefore, to eliminate periodic phase discontinuities, the wrapped phase needs to be unwrapped to obtain the absolute phase. During the phase unwrapping process, the single-frame spatial phase unwrapping method is limited by the premise of phase continuity and cannot handle large surface discontinuities and isolated objects. The F2PWholeNet strategy provides wrapping-level branch supervision, allowing network training to obtain additional information about the fringe order, fundamentally improving the phase ambiguity problem and allowing each spatial pixel to be unwrapped independently, effectively improving the spatial discontinuity problem during the modulation phase unwrapping process.

[0146] Advantage 4:

[0147] The F2PWholeNet strategy, thanks to its full-process supervision of the physical model, combines multiple branch networks into a single large network. Predicting different absolute phase solution steps requires only a single training phase, eliminating the need for separate training phases. This improves training and measurement speed and efficiency. Because the principles of the physical model are coherent and mutually guided, the synergy between the sub-networks promotes training accuracy. Furthermore, this strategy breaks down challenging long-distance predictions into multiple, less challenging, short-distance, small branch predictions, improving the accuracy, speed, and robustness of phase solution.

[0148] Advantage 5:

[0149] This supervision strategy is applicable to all existing regression networks. With the support of multi-way supervision, existing regression networks can build a more robust large network based on this network structure, thereby predicting higher-precision and more stable absolute phase.

[0150] The main innovative contributions of this invention include:

[0151] 1. In the field of structured light 3D measurement, for the design of an end-to-end absolute phase prediction network, we propose a multi-branch network architecture strategy for achieving end-to-end absolute phase prediction from a single grating frame, supervised by a physical model based on the principles of multi-frequency N-step phase shifting and multi-frequency heterodyning. Specifically, the network structure includes a numerator branch, a denominator branch, a wrapping branch, and an absolute phase branch as the supervision of the corresponding physical model. Specific modules, such as the unwrapping module and the fusion module, assist the network in feature fusion.

[0152] 2. The network composition strategy includes a backbone network from grating to absolute phase and four branch networks, branch 1 is the numerator branch, branch 2 is the denominator prediction branch, branch 3 is the wrapped phase prediction branch, and branch 4 is the secondary prediction branch (it should be noted that if the physical supervision method for unwrapping the phase uses the multi-frequency heterodyne method, the four branches are changed to only the first three branches 1, 2, and 3). The four network sub-branches predict the intermediate links required to calculate the absolute phase in sequence according to the traditional physical model and mathematical model. The main network's direct prediction from grating to absolute phase has been greatly improved in accuracy and precision based on the auxiliary supervision of the four branches, and the interpretability of the entire network has been enhanced. Due to the connectivity of the object model and the guidance of the true value, the four sub-networks and one backbone network play a positive role in promoting each other, and ultimately obtain an end-to-end fast, high-precision, and high-stability absolute phase acquisition network.

[0153] 3. It should be pointed out that the present invention not only proposes a network structure design strategy, but also specific networks. The branch network can be any regression network. As long as it is a new network based on this network structure design strategy, it is within the protection scope.

[0154] 4. Based on this fully supervised end-to-end absolute phase prediction network structure based on this physical model, we propose a strategy for designing the network's total loss function by summing the loss functions of multiple branches according to their weighted proportions. The proper design of this total loss function is a key step in the network's learning process towards the true value. The loss functions of each branch and the main branch are combined with weights to form this loss function, which guides the network's prediction direction and content.

[0155] In terms of interpretability and network prediction performance, the F2PWholeNet strategy of this invention mimics the physical model and uses the results of each intermediate link as process supervision for the sub-network, allowing the network to learn more phase, time, and spatial information, significantly improving the network robustness and prediction accuracy of end-to-end training. Because the principles of the physical model are mutually guided by coherence, the synergy between branches promotes training accuracy. Furthermore, this strategy divides difficult long-distance predictions into multiple, less difficult, short-distance, small branch predictions, improving the accuracy, speed, and robustness of phase solution.

[0156] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.

[0157] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.

[0158] An embodiment of the present invention further provides a processor, which executes a computer program and at least performs the method described above.

[0159] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0160] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0161] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0163] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0164] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0165] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0166] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0167] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0168] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.

Claims

1. A structured light 3D measurement method based on an absolute phase end-to-end prediction network, characterized in that: include: S1. Construct an end-to-end absolute phase prediction network F2PWholeNet, which includes a numerator graph branch network, a denominator graph branch network, a wrapped phase graph branch network, a level graph branch network, a first fusion module, a second fusion module, a third fusion module, an unwrapping module, and a trunk branch to simulate the phase solution process in the physical model; wherein: The molecular graph branch network is supervised by the true molecular graph and predicts the molecular graph from the input raster image; The denominator graph branch network is supervised by the true value denominator graph and predicts the denominator graph from the input raster image; The first fusion module fuses a selected input combination, the input combination being selected from: a predicted numerator graph, a denominator graph, and an original input raster graph; The wrapped phase image branch network is supervised by the true wrapped phase image and predicts the wrapped phase image using a selected input combination, the input combination being selected from: a grating image and an output of the first fusion module; The second fusion module fuses a selected input combination, the input combination being selected from: the predicted wrapped phase map and the output of the first fusion module; The level graph branch network is supervised by the ground truth parcel level graph and predicts the parcel level graph using a selected input combination, the input combination being selected from: a raster image and an output of the second fusion module; The unwrapping module uses the predicted wrapped phase map and the wrapped order map to perform phase solution to obtain the predicted absolute phase; The third fusion module fuses a selected input combination, the input combination being selected from: the predicted parcel level map, the predicted absolute phase, and the output of the second fusion module; The main branch, i.e., the absolute phase branch, is supervised by the true absolute phase image and predicts the absolute phase image using a selected input combination, the input combination being selected from: the output of the third fusion module; S2. Train the F2PWholeNet network to enable it to learn the mapping relationship from grating image to absolute phase in one go; S3. Use the trained F2PWholeNet network to perform absolute phase prediction on the new grating image to achieve high-precision 3D shape reconstruction.

2. The structured light 3D measurement method based on an absolute phase end-to-end prediction network according to claim 1, characterized in that: One or more of the following conditions must be met: The first fusion module fuses the predicted numerator graph and denominator graph with the original input raster graph; The wrapped phase image branch network is supervised by the true wrapped phase image and predicts the wrapped phase image using the grating image and the output of the first fusion module; The second fusion module fuses the predicted wrapped phase map and the output of the first fusion module; The level graph branch network is supervised by the true parcel level graph and predicts the parcel level graph using the raster image and the output of the second fusion module; The third fusion module fuses the predicted package level map, the predicted absolute phase and the output of the second fusion module; The main branch, namely the absolute phase branch, is supervised by the true absolute phase map, and the absolute phase map is predicted according to the output of the third fusion module.

3. The structured light 3D measurement method based on an absolute phase end-to-end prediction network according to claim 2, wherein: The unwrapping module performs phase calculation according to the following formula: Unwrap=O3+2π*O4 Among them, Unwrap is the predicted absolute phase, O3 is the predicted output of the wrapped phase map, and O4 is the predicted output of the wrapped level map.

4. The structured light 3D measurement method based on an absolute phase end-to-end prediction network according to any one of claims 1 to 3, characterized in that: When constructing the data set for training the F2PWholeNet network, the multi-frequency N-step phase shift method is adopted, and the grating image of any phase in the highest frequency grating image is used as the input data, and the high-frequency numerator image, high-frequency denominator image, high-frequency wrapped phase image, wrapped order image and absolute phase image corresponding to the input data are used as the supervision signals of the numerator image branch network, the denominator image branch network, the wrapped phase image branch network, the order image branch network and the trunk branch respectively.

5. The structured light 3D measurement method based on an absolute phase end-to-end prediction network according to claim 4, characterized in that: The data input-truth value pairs are solved using the twelve-step phase shift method and the multi-frequency unwrapping phase method.

6. The structured light 3D measurement method based on an absolute phase end-to-end prediction network according to any one of claims 1 to 3, characterized in that: The F2PWholeNet network adopts the following loss function: ; Among them, the mean absolute error L1 loss and mean square error L2 loss As the basic loss function of the branch, and Represent the true value and prediction results of each network respectively, is the loss function of the branch network n, the single-path basic loss function and the total loss function are and , n represents the number of branches, α is the hyperparameter of the basic loss function, are the hyperparameters of the five loss functions, and i ranges from 1 to 5.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the structured light three-dimensional measurement method based on the absolute phase end-to-end prediction network is implemented.

8. A computer program product, characterized in that When the computer program product is run by a processor, the structured light three-dimensional measurement method based on an absolute phase end-to-end prediction network according to any one of claims 1 to 6 is implemented.