Stripe structured light three-dimensional reconstruction method and device based on self-supervised learning
Through self-supervised learning methods, an end-to-end solution-phase network is built, which solves the problems of large data requirements and insufficient accuracy of deep learning structured optical systems, and realizes high-precision three-dimensional reconstruction.
Patent Information
- Application Number
- CN202510436080.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
Existing structured light systems based on deep learning require a large amount of data training, the model is difficult to migrate, and the reconstruction accuracy is insufficient.
Using a self-supervised learning method, by obtaining the deformation stripe diagram and phase information of different objects, an encoder is built and connected to a symmetric decoder, an end-to-end dephasing network is built, and a deep learning model is trained. The stripe-phase features are extracted only by a small amount of data.
High-precision 3D reconstruction with less data is realized, reducing the difficulty of model migration, improving reconstruction accuracy, and suitable for three-dimensional perception and precise measurement.
Smart Images

Figure CN120339512A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to three-dimensional structured light measurement technology, and particularly to a method and device for three-dimensional reconstruction of fringe structured light based on self-supervised learning. Background Art
[0002] The precise measurement technology of three-dimensional topography is an important cornerstone of modern scientific research and industrial applications. Its core goal is to achieve efficient acquisition and accurate processing of three-dimensional information. This reflects humanity's continuous improvement in the ability to explore and understand the objective world, providing rich and accurate data features for various downstream tasks.
[0003] As a non-contact, high-speed, and high-precision measurement method, three-dimensional structured light shows broad application prospects in many fields with its unique advantages, such as face recognition, three-dimensional design, automotive manufacturing, virtual reality, and animation modeling. Among the structured light measurement methods, the fringe profile measurement method has developed into a widely used and technically mature method due to its high speed, high precision, and strong anti-interference ability.
[0004] The fringe profile structured light technology usually uses a DLP (Digital Light Processing) projector to project a prefabricated fringe pattern carrying phase information onto the surface of the object to be measured. The camera captures the fringe image modulated by the surface topography of the object, and obtains the phase information through subsequent fringe demodulation steps, thereby realizing the reconstruction of the three-dimensional coordinates of the object. However, traditional methods often need to project multiple repeated phase-shifted patterns to improve the measurement accuracy, resulting in low measurement efficiency.
[0005] In recent years, fringe phase analysis methods based on deep learning have emerged widely. Deep learning models, with their powerful feature extraction capabilities, can achieve high-precision three-dimensional reconstruction using fewer fringe images. However, deep learning methods usually require a large amount of data to extract effective features, and constructing a large-scale real dataset is time-consuming and laborious. Although simulated data can be used to assist training, there are differences between simulated data and real scenarios, and the generalization ability of the extracted features is limited.
[0006] At the same time, the fringe features extracted by deep learning usually depend on the fringe morphology and position under specific spatial parameters, and it is difficult to effectively reuse the depth features between different scenarios. The effect of traditional feature transfer methods is limited. This problem severely restricts the wide application of deep learning-based structured light technology in actual scenarios. In addition, the features extracted by current deep learning methods have not fully focused on the phase features themselves, further resulting in less than ideal reconstruction accuracy.
[0007] Therefore, existing deep learning-based structured light systems require a large amount of data for training, the models are difficult to transfer, and the accuracy is insufficient.
[0008] It should be noted that the information disclosed in the above background art section is only used to understand the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0009] The present invention provides a three-dimensional reconstruction method of stripe structured light based on self-supervised learning, which solves the problems that the existing structured light system based on deep learning requires a large amount of data for training, the model is difficult to migrate, and the accuracy is insufficient.
[0010] The present invention adopts the following technical solutions:
[0011] In a first aspect, there is provided a three-dimensional reconstruction method of stripe structured light based on self-supervised learning, including the following steps:
[0012] S1. Obtain the deformed fringe patterns on the surfaces of several different objects, and solve for the phase information of the deformed fringe patterns;
[0013] S2. Construct an encoder, and train the encoder with the deformed fringe patterns and the phase information in step S1 through self-supervised learning to obtain a self-supervised learning encoder;
[0014] S3. Connect the self-supervised learning encoder to a symmetric decoder to construct an end-to-end phase unwrapping network, input the deformed fringe pattern in step S1 into the end-to-end phase unwrapping network, and constrain it with the phase ground truth to train the end-to-end phase unwrapping network to obtain a deep learning model;
[0015] S4. Input the deformed fringe pattern on the surface of the object to be measured obtained into the deep learning model, output the phase information of the deformed fringe pattern on the surface of the object to be measured, and analyze the phase information to obtain the three-dimensional point cloud map of the object to be measured.
[0016] Further, step S1 includes:
[0017] S11. Obtain the deformed fringe patterns on the surfaces of several different objects: use a DLP projector to project an N-frequency M-step phase-shifted fringe pattern onto the surfaces of several different objects, and use a camera to collect the deformed fringe patterns modulated by the surfaces of different objects; where M≥3;
[0018] S12. Solve for the phase information of the deformed fringe patterns: solve the sine component map and cosine component map corresponding to each frequency of the deformed fringe pattern through the M-step phase-shift method. Where, for a certain frequency of fringe, the deformed fringe pattern, sine component map, and cosine component map are respectively expressed as formula (1), formula (2), and formula (3):
[0019]
[0020] Among them, the subscript c represents the deformed fringe pattern collected by the camera, I represents the gray value of the deformed fringe pattern, k represents the number of phase shift steps, A represents the ambient light intensity, and B represents the projection light intensity.
[0021] Further, in step S2, the first deformed fringe pattern, sine component diagram, and cosine component diagram of each frequency corresponding to several different objects are used as the input of the encoder.
[0022] Further, in step S2, in a self-supervised learning manner, the encoder is trained with the deformed fringe pattern and its corresponding phase information. Through the contrastive learning loss constraint, the spatial distance between the deformed fringe features and the phase features of the corresponding object is minimized in the feature space, and the spatial distance between the deformed fringe features and the phase features of different objects is maximized, and the self-supervised learning encoder capable of accurately extracting fringe-phase features is trained.
[0023] Further, the encoder is a four-layer residual convolutional neural network.
[0024] Further, in the contrastive learning loss constraint of step S2, the features of the same object are used as positive sample pairs, the features of different objects are used as negative sample pairs, and the contrastive learning loss of multiple features is used as the loss of the encoder. Contrastive learning can pull positive sample pairs closer to each other and push negative sample pairs farther apart in the feature space. The loss function of any two features is:
[0025]
[0026] where t is the number of negative sample pairs, h θ is the spatial distance function used to measure two features, v represents the convolutional feature vector extracted by the neural network, the subscripts 1 and 2 respectively represent the features corresponding to the same object, and finally the contrastive learning loss of any two features is obtained. A total of three contrastive learning losses are used in step S2, and the total loss is:
[0027]
[0028] Through several rounds of training, the encoder can accurately learn the features of the deformed fringe pattern, sine component diagram, and cosine component diagram, thereby obtaining the self-supervised learning encoder.
[0029] Further, in step S3, the L1 loss function and the L2 loss function are used to constrain the output of each pixel point, and the loss function is:
[0030]
[0031] L = L1 + L2 (10)
[0032] Among them, L1 and L2 respectively represent the MAE loss function and the MSE loss function, f represents the frequency, and H and W respectively represent the height and width of the feature map.
[0033] Furthermore, step S4 includes:
[0034] S41. Use the DLP projector in the structured light three-dimensional measurement system to project the N-frequency 1-step phase-shifted fringe pattern onto the surface of the object to be measured, and use the camera to collect the deformed fringe pattern modulated by the surface of the object to be measured;
[0035] S42. Input the deformed fringe pattern in step S41 into the deep learning model obtained in step S3, and output the corresponding sine component map S and cosine component map C;
[0036] S43. Through the arctangent function, calculate the wrapped phase from the sine component map and cosine component map obtained in step S42, where the wrapped phase is expressed as:
[0037]
[0038] S44. Unwrap the wrapped phase to obtain the absolute phase, and obtain the three-dimensional point cloud map of the object to be measured through the absolute phase. Among them, the absolute phase is solved through the wrapped phases of different frequencies, and the formula is as follows:
[0039]
[0040] Among them, is the wrapped phase of the high-frequency fringe, is the wrapped phase of the low-frequency fringe, and λ is the frequency.
[0041] Furthermore, in step S44, through the cubic polynomial model calibrated by the camera and the structured light three-dimensional measurement system, the absolute phase of the highest frequency is mapped to the three-dimensional space coordinates, so as to obtain the three-dimensional point cloud map of the object to be measured.
[0042] In a second aspect, the present invention also provides a structured light three-dimensional reconstruction device based on self-supervised learning, including: a memory for storing executable instructions; a processor for executing the executable instructions stored in the memory to implement a structured light three-dimensional reconstruction method based on self-supervised learning as described in the first aspect.
[0043] The beneficial effects of the present invention are as follows: The present invention trains the encoder through self-supervised learning, accurately extracts the features between the fringes and the phase. This feature extraction method requires less data volume and has higher accuracy, and has the following advantages:
[0044] 1. Introduce self-supervised learning into the extraction of structured light fringe features. Only a small amount of data is required to learn the features for accurately solving the phase, which facilitates model migration and expands the application of deep learning in the field of structured light.
[0045] 2. The features based on self-supervised learning are more suitable for the fringe solving process and can achieve higher solving accuracy.
[0046] The present invention can be applied to precise measurements in fields such as three-dimensional perception, automotive manufacturing, and aerospace. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a flowchart of the three-dimensional reconstruction method of fringe structured light based on self-supervised learning according to an embodiment of the present invention.
[0048] Figure 2 The phase-shift fringe patterns of 4 frequencies in step S1 of the embodiment of the present invention.
[0049] Figures 3a - 3f They are respectively the deformed fringe pattern, sine component diagram, cosine component diagram, wrapped phase, absolute phase of the highest frequency, and three-dimensional point cloud diagram of the object to be measured in step S4 of the embodiment of the present invention.
[0050] Figure 4 It is a diagram of the structured light three-dimensional measurement system according to an embodiment of the present invention.
[0051] Figure 5 It is a schematic diagram of training the encoder in step S2 of the embodiment of the present invention.
[0052] Figure 6 It is a schematic diagram of training the end-to-end phase unwrapping network in step S3 of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0054] This embodiment provides a three-dimensional reconstruction method of fringe structured light based on self-supervised learning, introducing contrastive learning into the field of structured light, constraining the fringe features and phase features at the feature level, so that the trained encoder can extract features that better conform to the fringe-phase mapping relationship with fewer samples. The method of the present invention can effectively extract the features of fringe images and phase information relying only on less data, thereby realizing high-precision phase solving and three-dimensional reconstruction.
[0055] The structured light three-dimensional measurement system used in this embodiment is as Figure 4As shown in the figure, it mainly includes a DLP projector, a camera, and the object to be measured. Among them, the DLP projector, the camera, and the object to be measured form a triangular relationship. The DLP projector projects a clear projection fringe pattern (phase-shifted fringe pattern) onto the surface of the object to be measured, and the camera captures the deformed fringe pattern modulated by the object surface.
[0056] Based on the above-mentioned hardware system (structured light three-dimensional measurement system) after spatial calibration, first, a small sample data set is collected through the structured light three-dimensional measurement system; secondly, the features of fringe-phase are learned using self-supervised learning; then, the encoder learned by self-supervised learning is incorporated into the classical end-to-end phase unwrapping network and continues to be trained; finally, the trained end-to-end phase unwrapping network is applied to the structured light measurement system for three-dimensional reconstruction of the object to be measured. The flowchart of the method in this embodiment is as Figure 1 shown, including the following steps:
[0057] S1: Obtain the deformed fringe patterns on the surfaces of several different objects, and solve the phase information of the deformed fringe patterns; specifically, it includes:
[0058] S11: Use the above-mentioned structured light measurement system, and use the DLP projector to project 4-frequency 12-step phase-shifted fringe patterns (as Figure 2 shown in the figure, where the frequencies of figures (a), (b), (c), and (d) are 1, 4, 16, and 64 respectively) onto several different objects (about 100 groups), and use the camera to capture the deformed fringe patterns modulated by the object surface;
[0059] S12: Use the 12-step phase-shift method to obtain the phase information (sine component map and cosine component map) corresponding to each frequency of the deformed fringe pattern. Among them, for a certain frequency of fringes, the deformed fringe pattern The sine component map S and the cosine component map C can be respectively expressed as the following formulas (1), (2), and (3):
[0060]
[0061] where the subscript c represents the deformed fringe pattern captured by the camera, I represents the gray value of the deformed fringe pattern, k represents the number of steps of phase shift, A represents the ambient light intensity, B represents the projection light intensity, and S and C are respectively the sine component map and the cosine component map obtained by calculating the deformed fringe pattern of M-step phase shift.
[0062] Preferably, the first-step fringe (i.e., the first deformed fringe pattern), the sine component map, and the cosine component map corresponding to each frequency of these objects are used as the training set for training the encoder in step S2.
[0063] S2: Construct an encoder, and train the encoder with the deformed fringe pattern and the phase information in step S1 through self-supervised learning to obtain a self-supervised learning encoder; specifically:
[0064] Build an encoder, as Figure 5 shown, use the training set obtained in step S1 as the input of the encoder, and train the encoder through self-supervised learning to extract the fringe feature F I of the object, the sine feature F s and the cosine feature F C . During the training process, through the contrast learning loss constraint, minimize the spatial distance between the deformed fringes and the phase features (sine feature and cosine feature) of the corresponding object in the feature space, and maximize the spatial distance between the fringes and the phase features (sine feature and cosine feature) of different objects. That is, contrast learning pulls closer the fringe feature, sine feature, and cosine feature of the same object in the feature space and pushes away these features of different objects. After several rounds of training, the encoder can well learn the fringe feature of the deformed fringe pattern, the sine feature of the sine component map, and the cosine feature of the cosine component map, and provide features for the subsequent high-precision end-to-end fringe-phase solving network.
[0065] Preferably, in this example, a four-layer residual convolutional neural network (Residual Convolutional Neural Network) is used as the encoder. The N deformed fringe maps, sine component maps, and cosine component maps in step S1 are used as the input of the encoder, and the corresponding fringe feature F I , sine feature F s , and cosine feature F C are output respectively. The features of the same object are used as positive sample pairs, and the features of different objects are used as negative sample pairs. The contrast learning loss of multiple features is used as the loss of the encoder. Contrast learning can pull closer the positive sample pairs and push away the negative sample pairs in the feature space. The loss function of any two features is:
[0066]
[0067] where t is the number of negative sample pairs, h θ is the function for measuring the spatial distance between two features, V represents the convolutional feature vector extracted by the neural network, and the subscripts 1 and 2 respectively represent the features corresponding to the same object (V1 and V2 can represent the fringe feature, sine feature, or cosine feature). The present invention uses the Euclidean distance. Finally, the contrast learning loss of any two features is obtained. In this embodiment, a total of three contrast learning losses are used, and the total loss is:
[0068]
[0069] Through several rounds of training, the encoder can accurately learn the fringe features, sine features, and cosine features corresponding to the deformed fringe pattern, sine component pattern, and cosine component pattern. These learned features provide a basis for subsequent high-precision calculation of the fringe phase.
[0070] S3: Connect the self-supervised learning encoder to a symmetric decoder to construct an end-to-end phase unwrapping network. Input the deformed fringe pattern in step S1 into the end-to-end phase unwrapping network and constrain it with the phase ground truth to train the end-to-end phase unwrapping network to obtain a deep learning model. Specifically:
[0071] In the above step S2, an accurate feature representation for fringe-phase has been obtained. That is, in step S2, with only a small amount of data, the trained encoder has accurately obtained the fringe features, sine features, and cosine features. Further, in step S3, an end-to-end phase unwrapping network containing the above-trained encoder is constructed, as Figure 6 shown, connect the self-supervised learning encoder to a symmetric decoder to construct an end-to-end phase unwrapping network (such as U-Net). The role of this end-to-end phase unwrapping network is to analyze the deformed fringe pattern for fast and high-precision measurement. Therefore, the input of this end-to-end phase unwrapping network is the deformed fringe pattern of each frequency, and the output is the sine component pattern and cosine component pattern corresponding to these deformed fringe patterns.
[0072] Different from some deep learning-based phase unwrapping methods (such as U-Net), the encoder of the present invention can already extract the features of fringes and phases well, so there is no need for a large amount of data sets to train the network. This end-to-end phase unwrapping network is a regression task, and the core is to accurately solve the sine component pattern and cosine component pattern. Therefore, preferably, in step S3, the L1 loss function (MAE loss function) and L2 loss function (MSE loss function) are used to constrain the output of each pixel point, and trained for several rounds, so that the end-to-end phase unwrapping network can calculate the phase information (sine component pattern and cosine component pattern) of the input deformed fringe pattern, thereby obtaining a deep learning model. The loss functions of MAE (mean absolute error) and MSE (mean square error) are respectively:
[0073]
[0074] L = L1 + L2 (10)
[0075] where L1 and L2 respectively represent the MAE loss function and MSE loss function, f represents the frequency, and H and W respectively represent the height and width of the feature map.
[0076] Through step S3, a deep learning model for accurately solving the fringe phase is obtained. Since the encoder of this deep learning model effectively learns the features of fringes and phases, it can analyze fringes more accurately compared to classical deep learning phase unwrapping networks.
[0077] Step S4: Input the obtained deformed fringe pattern on the surface of the object to be measured into the deep learning model trained in step S3, output the phase information of the deformed fringe pattern on the surface of the object to be measured, and analyze this phase information to obtain the three-dimensional point cloud map of the object to be measured. Specifically:
[0078] Based on the above step S3, the obtained deep learning model can accurately analyze the phase. In actual measurement, for the object to be measured, step S4 includes:
[0079] S41. Use a DLP projector to project the N-frequency one-step phase-shifted fringe pattern onto the surface of the object to be measured, and use a camera to collect the deformed fringe pattern modulated by the surface of the object to be measured;
[0080] S42. Take the deformed fringe pattern collected by the camera and modulated by the surface of the object to be measured as the input of the deep learning model obtained in step S3, and obtain the sine component map and cosine component map corresponding to this deformed fringe pattern;
[0081] S43. Through the arctangent function, calculate the N-frequency wrapped phase from the N-frequency sine component map and cosine component Wrapped phase Can be expressed as:
[0082]
[0083] S44. Unwrap the N wrapped phases to obtain the absolute phase, and combine the highest-frequency absolute phase Φ with the parameters of the structured light three-dimensional measurement system to obtain the three-dimensional point cloud map of the object to be measured (specifically, through the cubic polynomial model calibrated by the camera and the structured light three-dimensional measurement system, map the highest-frequency absolute phase into spatial three-dimensional coordinates, thereby obtaining the three-dimensional point cloud map of the object to be measured). Among them, the absolute phase can be solved from the wrapped phases of different frequencies, and the formula is as follows:
[0084]
[0085] Among them, Is the wrapped phase of the high-frequency fringe, Is the wrapped phase of the low-frequency fringe, and λ is the frequency.
[0086] The deformed fringe pattern, sine component map, cosine component map, wrapped phase, highest-frequency absolute phase, and three-dimensional point cloud map of the object to be measured in step S4 are respectively as Figures 3a - 3f Shown.
[0087] The three-dimensional reconstruction method of fringe structured light based on self-supervised learning of the present invention solves the problems of large data requirements and insufficient reconstruction accuracy of existing deep learning methods. The present invention introduces self-supervised learning into three-dimensional measurement of structured light, and can effectively realize feature modeling of fringe-phase with only a small amount of data, greatly reducing the deployment difficulty of deep learning models. In addition, this method can clearly focus on and extract the phase features of fringes during the feature learning stage, and can more accurately realize phase calculation, thereby improving the measurement accuracy of three-dimensional reconstruction. Compared with the classical phase solving method based on deep learning, the present invention has the following two advantages:
[0088] (1) The method of the present invention can extract fringe features based on only a few picture samples, without the need to learn according to a large number of samples. The present invention compares the method of the present invention with the method based on the UNet network trained from scratch (hereinafter referred to as the classical method). As shown in Table 1, it is the phase unwrapping accuracy of the method of the present invention and the classical method. At the same accuracy, the method of the present invention only requires about 1 / 5 of the data volume, which greatly reduces the feasibility of the application of deep learning methods in structured light technology.
[0089] (2) The method of the present invention has higher phase calculation accuracy because the method of the present invention has accurately extracted fringe and phase features during the feature extraction stage, while the network trained from scratch is only constrained at the output layer of the network. As shown in Table 1, with sufficient data (about 700 pictures), the accuracy of the method of the present invention is improved by about 10% compared with the classical method.
[0090] Table 1:
[0091]
[0092] A three-dimensional reconstruction device of fringe structured light based on self-supervised learning includes: a memory for storing executable instructions; a processor for executing the executable instructions stored in the memory to implement the three-dimensional reconstruction method of fringe structured light based on self-supervised learning as described above.
[0093] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0094] In addition, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0095] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention.
[0096] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A three-dimensional reconstruction method of stripe structured light based on self-supervised learning, characterized in that It includes the following steps: S1. Obtain the deformed fringe patterns of the surfaces of several different objects, and solve to obtain the phase information of the deformed fringe patterns; S2. Construct an encoder, and train the encoder with the deformed fringe patterns and the phase information in step S1 through self-supervised learning to obtain a self-supervised learning encoder; S3. Connect the self-supervised learning encoder to a symmetric decoder to construct an end-to-end phase unwrapping network. Input the deformed fringe pattern in step S1 into the end-to-end phase unwrapping network, and constrain it with the phase ground truth to train the end-to-end phase unwrapping network to obtain a deep learning model; S4. Input the deformed fringe pattern of the surface of the object to be measured obtained into the deep learning model, output the phase information of the deformed fringe pattern of the surface of the object to be measured, and analyze the phase information to obtain the three-dimensional point cloud map of the object to be measured.
2. The method for three-dimensional reconstruction of stripe structured light based on self-supervised learning according to claim 1, wherein Step S1 includes: S11. Obtain the deformed fringe patterns of the surfaces of several different objects: Use a DLP projector to project an N-frequency M-step phase-shifted fringe pattern onto the surfaces of several different objects, and use a camera to collect the deformed fringe patterns modulated by the surfaces of different objects; where M≥3; S12. Solve to obtain the phase information of the deformed fringe pattern: Solve the sine component map and cosine component map corresponding to each frequency of the deformed fringe pattern through the M-step phase-shift method. For a certain frequency of fringes, the deformed fringe pattern, sine component map, and cosine component map are respectively expressed as formulas (1), (2), and (3): where the subscript c represents the deformed fringe pattern collected by the camera, I represents the gray value of the deformed fringe pattern, k represents the number of phase-shift steps, A represents the ambient light intensity, and B represents the projection light intensity.
3. The method for three-dimensional reconstruction of stripe structured light based on self-supervised learning according to claim 1 or 2, wherein In step S2, the first deformed fringe pattern, sine component map, and cosine component map of each frequency corresponding to several different objects are used as the input of the encoder.
4. The method for three-dimensional reconstruction of stripe structured light based on self-supervised learning according to claim 1 or 2, characterized in that, In step S2, in a self-supervised learning manner, the encoder is trained with the deformed fringe pattern and its corresponding phase information. Through contrast learning loss constraint, the spatial distance between the deformed fringe feature and the phase feature of the corresponding object is minimized in the feature space, and the spatial distance between the deformed fringe feature and the phase features of different objects is maximized. The self-supervised learning encoder that can accurately extract the fringe-phase features is trained.
5. The method for three-dimensional reconstruction of stripe structured light based on self-supervised learning according to claim 4, wherein The encoder is a four-layer residual convolutional neural network.
6. The method for three-dimensional reconstruction of stripe structured light based on self-supervised learning according to claim 5, characterized in that, In the contrast learning loss constraint of step S2, the features of the same object are used as positive sample pairs, and the features of different objects are used as negative sample pairs. The contrast learning loss of multiple features is used as the loss of the encoder. Contrast learning can pull the positive sample pairs closer to each other and push the negative sample pairs farther apart in the feature space. The loss function of any two features is: where t is the number of negative sample pairs, and h θ is a spatial distance function for measuring two features, v represents the convolutional feature vector extracted by the neural network, and the subscripts 1 and 2 respectively represent the features corresponding to the same object. Finally, the contrastive learning loss of any two features is obtained. A total of three contrastive learning losses are used in step S2, and the total loss is: Through several rounds of training, the encoder can accurately learn the features of the deformed fringe pattern, sine component map, and cosine component map, so as to obtain the self-supervised learning encoder.
7. The method for three-dimensional reconstruction of stripe structured light based on self-supervised learning according to claim 1, characterized in that, In step S3, the L1 loss function and L2 loss function are used to constrain the output of each pixel point, and the loss function is: L = L1 + L2 (10) Among them, L1 and L2 respectively represent the MAE loss function and the MSE loss function, f represents the frequency, and H and W respectively represent the height and width of the feature map.
8. The method for three-dimensional reconstruction of stripe structured light based on self-supervised learning according to claim 1, wherein Step S4 includes: S41. Use the DLP projector in the structured light three-dimensional measurement system to project the N-frequency 1-step phase-shifted fringe pattern onto the surface of the object to be measured, and use the camera to collect the deformed fringe pattern modulated by the surface of the object to be measured; S42. Input the deformed fringe pattern in step S41 into the deep learning model obtained in step S3, and output the corresponding sine component map S and cosine component map C; S43. By using the arctangent function, the wrapped phase is obtained from the sine component diagram and the cosine component diagram obtained in step S42, where the wrapped phase is expressed as: S44. Unwrap the wrapped phase to obtain the absolute phase, and obtain the three-dimensional point cloud map of the object to be measured through the absolute phase, where the absolute phase is solved by the wrapped phases of different frequencies, and the formula is as follows: Among them, is the wrapped phase of the high-frequency fringes, is the wrapped phase of the low-frequency fringes, and λ is the frequency.
9. The method for three-dimensional reconstruction of fringe structured light based on self-supervised learning according to claim 8, wherein In step S44, through the cubic polynomial model calibrated by the camera and the structured light three-dimensional measurement system, the absolute phase of the highest frequency is mapped into the three-dimensional space coordinates, so as to obtain the three-dimensional point cloud map of the object to be measured.
10. A three-dimensional reconstruction device for stripe structured light based on self-supervised learning, characterized in that, It includes: A memory for storing executable instructions; A processor for executing the executable instructions stored in the memory to implement a three-dimensional reconstruction method of fringe structured light based on self-supervised learning according to any one of claims 1-9.