Hybrid stripe coding method based on end-to-end deep learning

By adopting the hybrid stripe encoding method in end-to-end deep learning, the problem of depth estimation accuracy gap caused by periodicity of traditional stripe encoding methods is solved, and a higher depth estimation accuracy is achieved.

CN120198774APending Publication Date: 2025-06-24SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510263954.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

There is a gap in accuracy in the existing end-to-end stripe depth estimation methods. The traditional stripe encoding method periodically makes it difficult for the network to distinguish each stripe, resulting in interference in depth estimation.

Method used

Using a hybrid stripe encoding method based on end-to-end deep learning, projected stripes are spatially divided into different subdomains, 32-frequency stripes and 64-frequency stripes are inserted into different subdomains, and mixed stripe encoding is generated through a specific encoding method.

Benefits of technology

The accuracy of depth estimation is significantly improved. Whether it is to train real data on simulated data or verify real data on real data, hybrid stripe encoding can improve the accuracy of depth estimation compared with traditional periodic stripe encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198774A_ABST
    Figure CN120198774A_ABST
Patent Text Reader

Abstract

The invention provides a mixed stripe coding method based on end-to-end deep learning, and relates to the technical field of mixed stripe coding, and the method comprises the following steps: S1, projection stripes are obtained: a plurality of groups of stripes are projected to a measured object through a projector, and then the projection stripes are obtained; s2, obtaining a mixed stripe code: dividing the projection stripe into different sub-domains in space, inserting a 32-frequency stripe and a 64-frequency stripe into the different sub-domains, and obtaining the mixed stripe code after coding; and S3, verifying the mixed stripe code: constructing a virtual data set and a real data set, and verifying the validity of the mixed stripe code. According to the hybrid stripe coding method based on end-to-end deep learning, whether real data verification is trained on analog data or real data verification is trained on real data, the hybrid stripe coding as network input can significantly improve the precision of depth estimation compared with traditional periodic stripe coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hybrid fringe coding, and specifically provides a hybrid fringe coding method based on end-to-end deep learning. Background Art

[0002] Due to its high speed and high precision, fringe projection profilometry has been widely used in 3D imaging. Although multiple sets of fringes can ensure the reconstruction accuracy, it will inevitably reduce the measurement speed. In recent years, with the continuous development of deep learning, more and more work has combined deep learning with traditional fringe projection profilometry, using the convolutional network of deep learning to replace the complex operation process in traditional fringe projection profilometry, so as to reduce the number of input fringes and improve the efficiency. Some of these works use deep learning networks to directly estimate depth from fringes, replacing all processes of traditional methods, which is called end-to-end fringe depth estimation. These works can directly estimate depth through only one or more fringes, resulting in a significant improvement in efficiency.

[0003] However, compared with traditional methods, there is a large accuracy gap, which is also the main shortcoming of current end-to-end fringe depth estimation. Due to the periodicity of traditional fringe coding methods, this periodicity makes it impossible for the network to effectively distinguish the difference between each fringe, thus introducing certain interference in depth estimation. It is often necessary to collect a large amount of data to enable the network to distinguish the difference between each fringe. Summary of the Invention

[0004] The present invention aims to provide a hybrid fringe coding method based on end-to-end deep learning. Whether trained on simulated data and verified on real data or trained on real data and verified on real data, the hybrid fringe coding can significantly improve the depth estimation accuracy compared with traditional periodic fringe coding as the network input.

[0005] To achieve the above effects, the present invention provides the following technical solution: A hybrid fringe coding method based on end-to-end deep learning, including the following steps:

[0006] S1. Obtain projection fringes: Project multiple sets of fringes onto the object to be measured by a projector, and then obtain the projection fringes;

[0007] It further includes the following steps:

[0008] S2. Obtain hybrid fringe coding: Divide the projection fringes into different sub-domains in space, insert 32-frequency fringes and 64-frequency fringes in different sub-domains, and obtain hybrid fringe coding after encoding, where the encoding method is:

[0009]

[0010] S = (2T, 3T] ∪ (7T, 10T] ∪ (12T, 13T] ∪ (14T, 16T] ∪

[0011] (18T, 19T] ∪ (23T, 26T] ∪ (28T, 29T] ∪ (30T, 32T].

[0012] Wherein: represents the projected fringe, (x, y) represents the pixel coordinates of the projected fringe, A P represents the background intensity of the projected fringe, B P represents the modulation degree of the projected fringe, N represents the number of steps of phase shift, n represents the serial number of phase shift, T represents the length of one period of 32-frequency fringes, and S represents the interval where 64-frequency fringes are located.

[0013] Further, the steps for obtaining the projected fringe in step S1 are as follows:

[0014] S101: Use matlab to collect and generate a three-step phase-shifted fringe coding image;

[0015] S102: Write the three-step phase-shifted fringe coding image into the DLP optical engine;

[0016] S103: Use the DLP optical engine to project the three-step phase-shifted fringe coding image onto the object to be measured;

[0017] S104: Use a camera to capture the projected fringe on the object to be measured.

[0018] Further, in step S101, collecting the three-step phase-shifted fringe coding image includes five collection periods, and the five collection periods are 1 - 4 - 16 - 32 - 64 respectively.

[0019] Further, it also includes:

[0020] S3. Verify the hybrid fringe coding: Construct a virtual data set and a real data set to verify the effectiveness of the hybrid fringe coding;

[0021] S4. Verify the hybrid fringe coding scenario: Verify the training of simulated data with real data and the training of real data with real data respectively on two scenarios. After multiple rounds of training, obtain the comparison results of the hybrid fringe coding, and then determine the accuracy of the hybrid fringe coding.

[0022] Further, in step S3, the data in the virtual data set is collected from a simulated FPP system, and the data in the real data set is collected from a structured light projection measurement system.

[0023] Further, in step S3, the parameters of the simulated FPP system are the same as those of the real FPP system, and the structured light projection measurement system includes an infrared camera and a DLP optical engine.

[0024] Further, in step S3, the depth ground truth of the real dataset is obtained by 1-4-16-64 phase unwrapping, and the depth ground truth of the virtual dataset is directly generated by Blender software.

[0025] Further, in step S4, the training process is as follows:

[0026] S401: Input the projection stripes into the network;

[0027] S402: The network outputs the predicted depth value of the projection stripes through an encoder and a decoder;

[0028] S404: Supervise the predicted depth value of the projection stripes using the depth ground truth of the real data;

[0029] S404: Train the network by updating the network parameters through backpropagation.

[0030] Further, in step S4, the traditional 64-frequency stripes, the traditional 32-frequency stripes, and the hybrid-encoded stripes are respectively input, and the evaluation metrics are RMSE and MAE.

[0031] Further, the RMSE and MAE are as follows:

[0032]

[0033] where D ij represents the predicted depth value, represents the ground truth depth, n represents the number of pictures, and ij represents the pixel coordinates.

[0034] The present invention provides a hybrid stripe coding method based on end-to-end deep learning, which has the following beneficial effects: Whether training on simulated data and validating on real data or training on real data and validating on real data, the hybrid stripe coding as the network input can significantly improve the accuracy of depth estimation compared with the traditional periodic stripe coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flowchart of a hybrid stripe coding method based on end-to-end deep learning according to the present invention;

[0036] Figure 2 is a schematic diagram of a PSP three-dimensional imaging system of a hybrid stripe coding method based on end-to-end deep learning according to the present invention;

[0037] Figure 3 Schematic diagrams of traditional periodic fringe encoding and hybrid periodic fringe encoding of a hybrid fringe encoding method based on end-to-end deep learning according to the present invention;

[0038] Figure 4 Schematic diagram of a real structured light projection measurement system of a hybrid fringe encoding method based on end-to-end deep learning according to the present invention;

[0039] Figure 5 Schematic diagrams of real datasets and simulated datasets of a hybrid fringe encoding method based on end-to-end deep learning according to the present invention;

[0040] Figure 6 Schematic diagram of the reconstruction result of an isolated object of a hybrid fringe encoding method based on end-to-end deep learning according to the present invention. Detailed implementation manners

[0041] Please refer to Figure 1-6 , the present invention provides a technical solution: a hybrid fringe encoding method based on end-to-end deep learning, including the following steps:

[0042] S1. Obtain projection fringes: Project multiple groups of fringes onto an object to be measured through a projector, and then obtain the projection fringes;

[0043] It further includes the following steps:

[0044] S2. Obtain hybrid fringe encoding: Divide the projection fringes into different sub-domains in space, insert 32-frequency fringes and 64-frequency fringes in different sub-domains, and obtain hybrid fringe encoding after encoding, where the encoding method is:

[0045]

[0046] S = (2T, 3T] ∪ (7T, 10T] ∪ (12T, 13T] ∪ (14T, 16T] ∪

[0047] (18T, 19T] ∪ (23T, 26T] ∪ (28T, 29T] ∪ (30T, 32T].

[0048] In the formula: Inp represents the projection fringes, (x, y) represents the pixel coordinates of the projection fringes, A P represents the background intensity of the projection fringes, B P represents the modulation degree of the projection fringes, N represents the number of phase shift steps, n represents the serial number of the phase shift, T represents the length of one period of the 32-frequency fringes, and S represents the interval where the 64-frequency fringes are located.

[0049] Specifically, the steps for step S1 to obtain the projection fringes are:

[0050] S101: Collect and generate three-step phase-shifted fringe encoded images using Matlab;

[0051] S102: Write the three-step phase-shifted fringe encoded images into the DLP optical engine;

[0052] S103: Use the DLP optical engine to project the three-step phase-shifted fringe encoded images onto the object to be measured;

[0053] S104: Use a camera to capture the projected fringes on the object to be measured.

[0054] Specifically, in step S101, collecting the three-step phase-shifted fringe encoded images includes five acquisition periods, and the five acquisition periods are 1 - 4 - 16 - 32 - 64 respectively.

[0055] Specifically, it further includes:

[0056] S3. Verify the hybrid fringe encoding: Construct a virtual dataset and a real dataset to verify the effectiveness of the hybrid fringe encoding;

[0057] S4. Hybrid fringe encoding scenario verification: Verify the simulated data training real data and the real data training real data respectively on two scenarios. After multiple rounds of training, obtain the comparison results of the hybrid fringe encoding, and then determine the accuracy of the hybrid fringe encoding.

[0058] Specifically, in step S3, the data in the virtual dataset is collected from a simulated FPP system, and the data in the real dataset is collected from a structured light projection measurement system.

[0059] Specifically, in step S3, the parameters of the simulated FPP system are the same as those of the real FPP system, and the structured light projection measurement system includes an infrared camera and a DLP optical engine.

[0060] Specifically, in step S3, the depth ground truth of the real dataset is obtained through 1 - 4 - 16 - 64 phase unwrapping, and the depth ground truth of the virtual dataset is directly generated by Blender software.

[0061] Specifically, in step S4, the training process is as follows:

[0062] S401: Input the projected fringes into the network;

[0063] S402: The network outputs the predicted depth value of the projected fringes through an encoder and a decoder;

[0064] S404: Use the depth ground truth of the real data to supervise the predicted depth value of the projected fringes;

[0065] S404: Update the network parameters through backpropagation to train the network.

[0066] Specifically, in step S4, the traditional 64-frequency fringes, the traditional 32-frequency fringes, and the hybrid-coded fringes are respectively input, and the evaluation metrics are RMSE and MAE.

[0067] Specifically, the RMSE and MAE are as follows:

[0068]

[0069] where D ij represents the predicted depth value, represents the true depth, n represents the number of pictures, and ij represents the pixel coordinates.

[0070] The method of the embodiment is used for detection and analysis, and compared with the prior art. According to the comparison, it can be obtained that when the embodiment is used, whether it is training on simulated data and validating on real data or training on real data and validating on real data, the hybrid fringe coding as the network input can significantly improve the accuracy of depth estimation compared with the traditional periodic fringe coding.

[0071] The present invention provides a hybrid fringe coding method based on end-to-end deep learning:

[0072] The fringe is divided into different sub-domains in space, and fringes with different frequencies are inserted in different sub-domains. Considering the continuity of the fringes, in this study, the fringes are co-coded using 32-frequency and 64-frequency fringes. The specific coding method is as follows:

[0073]

[0074] S = (2T, 3T] ∪ (7T, 10T] ∪ (12T, 13T] ∪ (14T, 16T] ∪

[0075] (18T, 19T] ∪ (23T, 26T] ∪ (28T, 29T] ∪ (30T, 32T].

[0076] where Inp is the projected fringe, (x, y) represents the pixel coordinates of the projected fringe, A P represents the background intensity of the projected fringe, B P represents the modulation degree of the projected fringe, N represents the number of phase-shift steps, which is equal to 3 in this study, n represents the phase-shift serial number, n = 0, 1, 2 in this study, T represents the length of one period of the 32-frequency fringe, and S represents the interval where the 64-frequency fringe is located.

[0077] The encoded fringe is compared with the traditional periodic fringe as Figure 3 shown, where (a) is the traditional fringe period coding, and (b) is the hybrid period coding proposed by the present invention.

[0078] To verify the effectiveness of the hybrid periodic coding proposed in this study, two datasets were constructed in this study, namely the virtual dataset and the real dataset. The real data was collected from a self-built structured light projection measurement system, which consists of an infrared camera and a DLP (Digital Light Processing) optical engine. The infrared camera is a Hikvision camera (model MV-CA016-10UM, black and white), with a resolution of 1440*1080 pixels, the camera sensor model is IMX273, the type is CMOS, the DLP optical engine model is Texas Instruments DLP LightCrafter 4500, the projection pattern size is 912×1140 pixels, and the system is as Figure 4 shown.

[0079] The system working process is as follows: First, different three-step phase-shift fringe coding images are generated, and different fringe coding images are written into the optical engine using the optical engine program. Then, the fringe pattern is projected onto the object using the optical engine, and finally, the corresponding object fringe image is captured using the camera.

[0080] Using this system, data was collected for 42 different plaster models respectively, and a total of 9000 groups of scene data were collected. Each scene has 18 fringe pictures, which are three-step phase-shift fringe maps of five frequencies of 1-4-16-32-64 and the three-step phase-shift fringe map of the hybrid coding proposed in this study. The real data was divided into a training set, a validation set, and a test set according to a ratio of 7:1:1. The simulated data was collected from a simulated FPP system using the same parameters as the real FPP system, and was collected on 21 model files, with a total of 14880 groups of data collected. The finally constructed dataset is as Figure 5 shown, where (a) is the constructed real dataset and (b) is the constructed simulated dataset.

[0081] The training process is as follows: First, the scene fringe image is input into the network. The network outputs the predicted depth value through the encoder and decoder, and then the depth ground truth is used to supervise the depth value predicted by the network. The network parameters are updated through backpropagation to train the network. The depth ground truth of the real data is obtained through 1-4-16-64 phase unwrapping, and the depth ground truth of the simulated data is directly generated by the simulation software.

[0082] This study was verified in two scenarios respectively. The first scenario was to train on 14880 simulated data and verify on 1000 real data. The second scenario was to train on 7000 real data and verify on 1000 real data. The traditional 64-frequency fringe image, the traditional 32-frequency fringe image, and the hybrid coding fringe image were input respectively, and the evaluation metrics were RMSE and MAE, as follows:

[0083]

[0084] Among them, D ij represents the predicted depth value, represents the true depth value, n represents the number of pictures, ij represents the pixel coordinates. Finally, after 100 rounds of training, the comparison results in the first and second scenarios are shown in Table 1 and Table 2 respectively:

[0085] Table 1 Comparison Results of Simulated Data Training and Real Data Validation

[0086] RMSE (mm) MAE (mm) 32-frequency periodic fringe coding 10.998 9.961 64-frequency periodic fringe coding 12.341 11.065 Hybrid fringe coding 4.286 2.297

[0087] Table 2 Comparison Results of Real Data Training and Real Data Validation

[0088] RMSE (mm) MAE (mm) 32-frequency periodic fringe coding 7.981 7.171 64-frequency periodic fringe coding 7.521 6.741 Hybrid fringe coding 1.489 0.664

[0089] As can be seen from the above table, whether it is training on simulated data and validating on real data or training on real data and validating on real data, using hybrid stripe coding as the network input can significantly improve the accuracy of depth estimation compared to traditional periodic stripe coding.

[0090] The quantitative accuracy results of using hybrid stripe coding as the network input for validation on 1000 test sets are good. To verify the qualitative indicators of hybrid coding in a single scenario, in this study, an isolated object scenario was selected from 1000 test scenarios. The network trained with hybrid coding was used to predict the depth map of the scenario, and the scene depth map was reconstructed into a three-dimensional scene using the camera internal parameters. The results are as Figure 6 shown.

[0091] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A hybrid stripe encoding method based on end-to-end deep learning, comprising the following steps: S1. Obtaining projection fringes: projecting multiple groups of fringes onto the object to be measured via a projector, and then obtaining the projection fringes; It is characterized in that The following steps are also included: S2. Obtaining mixed fringe coding: Divide the projected fringe into different sub-domains in space, insert 32-frequency fringe and 64-frequency fringe into different sub-domains, and obtain mixed fringe coding after coding, wherein the coding method is: S=(2T,3T]∪(7T,10T]∪(12T,13T]∪(14T,16T]∪ (18T,19T]∪(23T,26T]∪(28T,29T]∪(30T,32T]. Where: represents the projected stripes, (x, y) represents the pixel coordinates of the projected stripes, and A P represents the background intensity of the projected fringes, B P represents the modulation degree of the projected stripes, N represents the number of phase shift steps, n represents the serial number of the phase shift, T represents the length of one cycle of the 32-frequency stripes, and S represents the interval where the 64-frequency stripes are located.

2. A hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: The steps of obtaining the projection fringes in step S1 are: S101: Use MATLAB to collect and generate three-step phase-shift fringe coded images; S102: writing the three-step phase-shift fringe coded image into a DLP optical engine; S103: Projecting the three-step phase-shift fringe coding image onto the object to be measured using a DLP optical machine; S104: Using a camera to photograph the projection fringes on the object to be measured.

3. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: In step S101, collecting the three-step phase-shift fringe coded image includes five collection cycles, and the five collection cycles are 1-4-16-32-64 respectively.

4. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: Also includes: S3, verifying the hybrid stripe coding: constructing a virtual data set and a real data set to verify the validity of the hybrid stripe coding; S4. Hybrid stripe coding scenario verification: The simulation data training real data and the real data training real data are verified in two scenarios respectively. After multiple rounds of training, the comparison results of the hybrid stripe coding are obtained, and then the accuracy of the hybrid stripe coding is determined.

5. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: In step S3, the data in the virtual data set are collected from the simulated FPP system, and the data in the real data set are collected from the structured light projection measurement system.

6. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: In step S3, the parameters of the simulated FPP system are the same as those of the real FPP system, and the structured light projection measurement system includes an infrared camera and a DLP optical machine.

7. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: In step S3, the true depth value of the real data set is obtained by 1-4-16-64 phase unwrapping, and the true depth value of the virtual data set is directly generated by blender software.

8. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: In step S4, the training process is: S401: Inputting the projection fringes into a network; S402: The network outputs the predicted depth value of the projected stripes through an encoder and a decoder; S404: Using the true depth value of the real data to supervise the predicted depth value of the projected fringes; S404: Train the network by updating the network parameters through back propagation.

9. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: In step S4, the traditional 64-frequency stripes, the traditional 32-frequency stripes and the hybrid coding stripes are input respectively, and the evaluation indicators are RMSE and MAE.

10. The hybrid stripe encoding method based on end-to-end deep learning according to claim 1, characterized in that: The RMSE and MAE are shown below: Where D ij represents the predicted depth value, Represents the true depth, n represents the number of images, and ij represents the pixel coordinates.