Absolute phase unwrapping model training method and three-dimensional reconstruction method
By training the absolute phase expansion model, using the wrapping phase generation model and the stripe rank segmentation model, the error propagation and calculation problems of the phase expansion method in the existing three-dimensional visual imaging technology are solved, and fast and accurate three-dimensional image reconstruction is achieved.
Patent Information
- Application Number
- CN202510247429.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-27
AI Technical Summary
In the existing three-dimensional visual imaging technology, the phase expansion method has problems such as large error propagation and calculation and slow speed. Especially when noise interference, the phase expansion error is large, making it difficult to achieve fast and accurate three-dimensional image reconstruction.
A training method for absolute phase expansion model is proposed. By obtaining training data, including sample stripe diagrams, real labels and masks, iterative training is used to use the wrapping phase generation model and the stripe rank segmentation model to generate a model that can output accurate absolute phases.
It achieves the improvement of efficiency and accuracy of three-dimensional reconstruction of objects, and can maintain high accuracy in the case of noise interference, and is suitable for dynamic and rapidly changing three-dimensional reconstruction scenarios.
Smart Images

Figure CN120219913A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the technical field of image processing, and particularly relate to a training method for an absolute phase unwrapping model and a three-dimensional reconstruction method. Background Art
[0002] The most intuitive manifestation of the feature information of an object is the three-dimensional surface information of the object. Obtaining three-dimensional information is usually achieved through three-dimensional vision imaging technology. Three-dimensional vision imaging methods have a wide range of applications in fields such as the metaverse, machining, defect detection, medical image imaging, virtual / augmented reality, and digital twin. Different development directions have emerged for three-dimensional vision imaging technology according to the application fields, and their aims are all to improve the accuracy, speed, practicality, and convenience of three-dimensional imaging in different scenarios.
[0003] Currently, there are many technical route branches for three-dimensional vision imaging technology. According to whether it contacts the imaging object, three-dimensional vision imaging technology is roughly divided into contact imaging and non-contact imaging. According to the different media used in the imaging method, non-contact imaging is further divided into optical three-dimensional imaging, electromagnetic three-dimensional imaging, and acoustic three-dimensional imaging. Among them, optical three-dimensional imaging has gradually become a key research direction in the field of three-dimensional imaging due to its advantages such as fast speed, high accuracy, and convenient measurement. According to whether it actively emits light signals during imaging, optical three-dimensional imaging is further divided into active optical imaging and passive optical imaging. Among them, Fringe Projection Profilometry (FPP) has become one of the most popular active optical three-dimensional imaging technologies due to its simple hardware configuration, flexible implementation, and high measurement accuracy.
[0004] The imaging method of fringe projection profilometry is to recover the depth information of the imaged object from the fringe projection image, and this process involves the phase unwrapping processing of the fringes. In phase unwrapping, there are also: path-dependent phase unwrapping methods and path-independent phase unwrapping methods. Although the path-dependent phase unwrapping method is fast, the error will spread in the local area. At the same time, although the path-independent method has certain anti-noise performance and the unwrapped phase is smoother, its computational complexity is large, the phase unwrapping speed is slow, and secondly, when affected by noise interference, there is an overall error in the phase. There is an urgent need for a fast and accurate phase unwrapping method to perform three-dimensional image reconstruction quickly and accurately. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a training method for an absolute phase unwrapping model and a three-dimensional reconstruction method, so as to provide an absolute phase unwrapping model, and using this model can improve the efficiency and accuracy of three-dimensional reconstruction of an object.
[0006] In the first aspect of the embodiments of the present application, a training method for an absolute phase unwrapping model is provided. The absolute phase unwrapping model includes: a wrapped phase generation model and a fringe order segmentation model connected to each other;
[0007] Obtain training data; the training data includes sample fringe patterns, true labels corresponding to the sample fringe patterns, and masks corresponding to each training object; multiple training objects are included in the sample fringe patterns;
[0008] Input the sample fringe pattern into the wrapped phase generation model to obtain a predicted wrapped phase;
[0009] Input the predicted wrapped phase and the predicted edge map corresponding to the predicted wrapped phase into the fringe order segmentation model to obtain the predicted fringe orders corresponding to each training object;
[0010] Iteratively train the wrapped phase generation model and the fringe order segmentation model according to the predicted wrapped phase, the predicted fringe orders, the masks, and the true labels until the training of the wrapped phase generation model and the fringe order segmentation model is completed;
[0011] Wherein, the predicted edge map is generated from the predicted wrapped phase, and the sample fringe pattern includes a left camera sample image and a right camera sample image collected by a pre-set binocular structured light projection system.
[0012] In the second aspect of the embodiments of the present application, a three-dimensional reconstruction method is provided, including:
[0013] Obtain an absolute phase unwrapping model, a to-be-processed fringe pattern collected by a pre-set binocular structured light projection system, and projection parameters of the projection system;
[0014] Input the to-be-processed fringe pattern into the absolute phase unwrapping model;
[0015] Receive the target absolute phase output by the absolute phase unwrapping model for the to-be-processed fringe pattern;
[0016] Perform three-dimensional reconstruction according to the target absolute phase and the projection parameters to generate three-dimensional reconstruction data matching the to-be-processed fringe pattern;
[0017] Wherein the absolute phase unwrapping model is trained according to the method described in the first aspect of the embodiments of the present application.
[0018] In some embodiments of the second aspect, the target absolute phase includes an individual left target phase and an individual right target phase of each target object; the three-dimensional reconstruction based on the target absolute phase and the projection parameters to generate three-dimensional reconstruction data matching the to-be-processed fringe pattern includes:
[0019] For each of the target objects, calculate a target disparity based on the pixel points in the individual left target phase and the pixel points in the individual right target phase;
[0020] Determine the depth information of each of the target objects based on the target disparity and the projection parameters;
[0021] Convert the depth information based on the projection parameters to generate three-dimensional reconstruction data of the to-be-processed fringe pattern.
[0022] A third aspect of the embodiments of the present application provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the absolute phase unwrapping model training method as described in the first aspect above or the three-dimensional reconstruction method as described in the second aspect above.
[0023] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the absolute phase unwrapping model training method as described in the first aspect above or the three-dimensional reconstruction method as described in the second aspect above.
[0024] A fifth aspect of the embodiments of the present application provides a computer program product including a computer program, and when the computer program is run, it implements the absolute phase unwrapping model training method as described in the first aspect above or the three-dimensional reconstruction method as described in the second aspect above.
[0025] The embodiments of the present application have the following beneficial effects:
[0026] In an embodiment of the present application, training data is obtained; the training data includes sample fringe patterns, true labels corresponding to the sample fringe patterns, and masks corresponding to each training object; multiple training objects are included in the sample fringe patterns; the sample fringe patterns are input into the wrapped phase generation model to obtain predicted wrapped phases; the predicted wrapped phases and the predicted edge maps corresponding to the predicted wrapped phases are input into the fringe order segmentation model to obtain predicted fringe orders corresponding to each training object; the wrapped phase generation model and the fringe order segmentation model are iteratively trained according to the predicted wrapped phases, the predicted fringe orders, the masks, and the true labels until the training of the wrapped phase generation model and the fringe order segmentation model is completed; the predicted edge maps are generated from the predicted wrapped phases, and the sample fringe patterns include left camera sample images and right camera sample images collected by a pre-set binocular structured light projection system, so as to realize the generation of an absolute phase unwrapping model including a mutually connected wrapped phase generation model and a fringe order segmentation model. The trained absolute phase unwrapping model can output accurate absolute phases for the fringe patterns it receives. The wrapped phase generation model and the fringe order segmentation model are deep learning models. By using the trained absolute phase model for absolute phase unwrapping, a deep learning-based absolute phase unwrapping method is provided. The deep learning-based absolute phase unwrapping method has obvious advantages compared with traditional non-deep learning methods due to its learning of a large amount of data sets, and at the same time, its robustness is higher due to the training of data sets in a large number of scenarios. The deep learning-based absolute phase unwrapping method is applicable to three-dimensional reconstruction in specific scenarios, such as the three-dimensional reconstruction of workpieces in the industrial robot arm grasping workpieces. Moreover, due to the advantages of deep learning, reconstruction with fewer image numbers can be achieved, and it has good practice and performance in dynamic sorting scenarios. Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a schematic diagram of the training of an absolute phase unwrapping model provided by an embodiment of the present application;
[0029] Figure 2 It is a schematic diagram of a sample fringe pattern provided by an embodiment of the present application;
[0030] Figure 3It is a schematic diagram of the setting of a binocular structured light projection system provided by an embodiment of the present application;
[0031] Figure 4 It is a schematic diagram of wrapped phase provided by an embodiment of the present application,
[0032] Figure 5 It is a schematic diagram of absolute phase provided by an embodiment of the present application;
[0033] Figure 6 It is a schematic diagram of fringe order provided by an embodiment of the present application;
[0034] Figure 7 It is a schematic diagram of a marking interface provided by an embodiment of the present application;
[0035] Figure 8 It is a schematic diagram of a marking result provided by an embodiment of the present application;
[0036] Figure 9A It is a schematic diagram of a mask provided by an embodiment of the present application;
[0037] Figure 9B It is another schematic diagram of a mask provided by an embodiment of the present application;
[0038] Figure 9C It is yet another schematic diagram of a mask provided by an embodiment of the present application;
[0039] Figure 10 It is a schematic diagram of the molecular term of wrapped phase provided by an embodiment of the present application;
[0040] Figure 11 It is a schematic diagram of the molecular term of wrapped phase provided by an embodiment of the present application;
[0041] Figure 12 It is a method for generating a data set provided by an embodiment of the application;
[0042] Figure 13 It is a schematic diagram of the data processing process flow of a fringe order segmentation model provided by an embodiment of the present application;
[0043] Figure 14 It is a schematic diagram of the wrapped phase of the left camera provided by an embodiment of the present application;
[0044] Figure 15 It is a schematic diagram of the wrapped phase of the right camera provided by an embodiment of the present application;
[0045] Figure 16 It is an edge map of the wrapped phase of the left camera provided by an embodiment of the present application;
[0046] Figure 17It is an edge map of the wrapped phase of the right camera provided by an embodiment of the present application;
[0047] Figure 18 It is a fringe order map obtained without using a mask provided by an embodiment of the present application;
[0048] Figure 19 It is a fringe order map obtained without using a mask provided by an embodiment of the present application;
[0049] Figure 20 It is a schematic diagram of the processing flow of a wrapped phase generation model provided by an embodiment of the present application;
[0050] Figure 21 It is a schematic diagram of a dense residual connection module provided by an embodiment of the present application;
[0051] Figure 22 It is a schematic diagram of the three-dimensional reconstruction effect provided by an embodiment of the present application;
[0052] Figure 23 It is another schematic diagram of the three-dimensional reconstruction effect provided by an embodiment of the present application;
[0053] Figure 24 It is a schematic diagram of a three-dimensional reconstruction method provided by an embodiment of the present application;
[0054] Figure 25 It is a schematic diagram of a training device for an absolute phase unwrapping model provided by an embodiment of the present application;
[0055] Figure 26 It is a schematic diagram of a three-dimensional reconstruction device provided by an embodiment of the present application;
[0056] Figure 27 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0057] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0058] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0059] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0060] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0061] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0062] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0063] One of the objectives of the embodiments of the present invention is to provide a model training method such that the generated model can perform an end-to-end phase unwrapping method. By using the absolute phase unwrapping model after completion of training, three-dimensional reconstruction of an object can be performed based on a single-frame fringe pattern. The embodiments of this application are applicable to the field of three-dimensional reconstruction of various rapidly changing scenarios, such as conveyor belt sorting, dynamic object reconstruction, and other fields.
[0064] The technical solutions of this application will be described below through specific embodiments.
[0065] Referring to Figure 1 , a schematic diagram of the training of an absolute phase unwrapping model provided by the embodiments of this application is shown. The absolute phase unwrapping model includes: a wrapped phase generation model and a fringe order segmentation model connected to each other;
[0066] The absolute phase unwrapping model may include a wrapped phase generation model and a fringe order segmentation model. The wrapped phase generation model can output a matching wrapped phase for the fringe pattern received by the absolute phase unwrapping model and transmit the wrapped phase to the fringe order segmentation model. The fringe order segmentation model can process the received data including the wrapped phase and output a fringe order matching the data it receives. The absolute phase unwrapping model can obtain the absolute phase based on the wrapped phase output by the wrapped phase generation model and the fringe order output by the fringe separation model, thereby realizing the absolute phase unwrapping of the fringe pattern.
[0067] The embodiments of the present application may specifically include the following steps:
[0068] Step 101, obtaining training data;
[0069] The training data includes sample fringe patterns, true labels corresponding to the sample fringe patterns, and masks corresponding to each training object; the sample fringe patterns contain multiple training objects;
[0070] Referring to Figure 2 , a schematic diagram of a sample fringe pattern provided by an embodiment of the present application is shown; the sample fringe pattern includes a left camera sample image and a right camera sample image collected by a pre-set binocular structured light projection system.
[0071] Referring to Figure 3 , a schematic diagram of the setting of a binocular structured light projection system provided by an embodiment of the present application is shown. In order to obtain training data for training the absolute phase unwrapping model, a binocular structured light projection system can be set up in a real scene. The binocular structured light system includes a left camera (Camera 1), a right camera (Camera 2), and a structured light projector (Projector). The structured light projector can irradiate the training object. The left camera and the right camera can capture the training object to collect the fringe pattern irradiated on the training object. Correspondingly, the fringe pattern collected by the left camera is the left camera sample image (as Figure 3 shown), and the fringe pattern collected by the right camera is the right camera sample image.
[0072] In order to evaluate the accuracy of the absolute phase unwrapping model during the training process and after the training is completed, it is necessary to first determine the true label corresponding to the sample fringe pattern. The true label can be calculated for the fringe pattern.
[0073] In the embodiments of the present application, the absolute phase unwrapping model can output the wrapped phase and the fringe order for multiple objects included in a single fringe pattern, so as to perform absolute phase unwrapping for multiple objects in a single fringe pattern. To avoid the error of fringe order segmentation caused by the interval between objects, at the same time, the absolute phase unwrapping model can output the absolute phase corresponding to each object, and a mask set for the training object is introduced into the training data, and different masks respectively correspond to different training objects in the fringe pattern.
[0074] As another example, the fringe pattern can also be obtained by simulating and photographing using a virtual binocular structured light projection system, that is, the left camera sample image and the right camera sample image can be simulation images.
[0075] Step 102: Input the sample fringe pattern into the wrapped phase generation model to obtain the predicted wrapped phase;
[0076] Input the sample fringe pattern into the wrapped phase generation model, and the wrapped phase generation model can process the sample fringe pattern to obtain the predicted wrapped phase corresponding to the sample fringe pattern. The predicted wrapped phase includes the real part and the imaginary part of the sample fringe pattern.
[0077] Refer to Figure 4 , which shows a schematic diagram of the wrapped phase provided by the embodiments of the present application. Refer to Figure 5 , which shows a schematic diagram of the absolute phase provided by the embodiments of the present application. In the embodiments of the present application, the wrapped phase and the absolute phase can be displayed in the form of an image, that is, the predicted wrapped phase can be an image.
[0078] Step 103: Input the predicted wrapped phase and the predicted edge map corresponding to the predicted wrapped phase into the fringe order segmentation model to obtain the predicted fringe orders corresponding to each training object;
[0079] The predicted edge map is generated from the predicted wrapped phase. To improve the feature recognition and segmentation accuracy of the fringe order segmentation model for the predicted wrapped phase and reduce mis-segmentation, the predicted edge map corresponding to the predicted wrapped phase can be generated first, and the predicted wrapped phase obtained in step 102 and the predicted edge map generated based on the predicted wrapped phase are input into the fringe order segmentation model, and the fringe order segmentation model outputs the predicted fringe orders corresponding to the training objects according to the predicted wrapped phase and the predicted edge map.
[0080] As an example, the Sobel operator can be used to calculate the gradient Ix in the x direction and the gradient Iy in the y direction of the predicted wrapped phase, and then the edge map of the predicted wrapped phase map is obtained through .
[0081] Refer to Figure 6, which shows a schematic diagram of stripe levels provided by an embodiment of the present application. In the embodiment of the present application, the predicted stripe levels can be displayed in the form of an image, that is, the predicted stripe levels can be an image.
[0082] Step 104: Iteratively train the wrapped phase generation model and the stripe level segmentation model according to the predicted wrapped phase, the predicted stripe levels, the mask, and the true label until the training of the wrapped phase generation model and the stripe level segmentation model is completed;
[0083] After each training object is labeled, a corresponding mask can be generated for each training object respectively.
[0084] The wrapped phase generation model and the stripe level segmentation model are trained using the predicted wrapped phase, the predicted stripe levels, the mask, and the true label. The wrapped phase generation model can be iteratively updated according to the predicted wrapped phase and the true label until the training is completed, and the stripe level segmentation model can be iteratively updated according to the predicted wrapped phase, the true label, and the mask until the training is completed.
[0085] After the training of the wrapped phase generation model and the stripe level segmentation model is completed, the absolute phase unwrapping model can accurately output the absolute phase matching the received fringe pattern for the fringe pattern.
[0086] In the embodiments of the present application, training data is obtained; the training data includes sample fringe patterns, true labels corresponding to the sample fringe patterns, and masks corresponding to each training object; the sample fringe patterns contain multiple training objects; the sample fringe patterns are input into the wrapped phase generation model to obtain predicted wrapped phases; the predicted wrapped phases and the predicted edge maps corresponding to the predicted wrapped phases are input into the fringe order segmentation model to obtain predicted fringe orders corresponding to each training object; the wrapped phase generation model and the fringe order segmentation model are iteratively trained according to the predicted wrapped phases, the predicted fringe orders, the masks, and the true labels until the training of the wrapped phase generation model and the fringe order segmentation model is completed; the predicted edge maps are generated from the predicted wrapped phases, and the sample fringe patterns include left camera sample images and right camera sample images collected by a pre-set binocular structured light projection system, so as to implement an absolute phase unwrapping model including a mutually connected wrapped phase generation model and a fringe order segmentation model. The trained absolute phase unwrapping model can output accurate absolute phases for the fringe patterns it receives. The wrapped phase generation model and the fringe order segmentation model are deep learning models. By using the trained absolute phase model for absolute phase unwrapping, an absolute phase unwrapping method based on deep learning is provided. The absolute phase unwrapping method based on deep learning has obvious advantages compared with traditional non-deep learning methods because of its learning of a large amount of data sets, and at the same time, its robustness is higher due to the training of data sets in a large number of scenarios. The absolute phase unwrapping method based on deep learning is applicable to 3D reconstruction in specific scenarios, such as 3D reconstruction of workpieces in industrial robotic arm workpiece grasping. Moreover, due to the advantages of deep learning, reconstruction with fewer image numbers can be achieved, and it has good practice and performance in dynamic sorting scenarios.
[0087] As an example, referring to Figure 7 , a schematic diagram of a labeling interface provided by the embodiments of the present application is shown; referring to Figure 8 , a schematic diagram of a labeling result provided by the embodiments of the present application is shown. Referring to Figure 9A , Figure 9B , Figure 9C , respectively, schematic diagrams of masks provided by the embodiments of the present application are exemplified. Any fringe pattern containing the same training object can be labeled, and the boundaries of the training objects are labeled in the fringe pattern to obtain a labeling result as shown in Figure 8 . According to the labeling result, masks corresponding to each training object are generated (taking Figure 7 as an example, including labeling three objects respectively, so as to obtain masks corresponding to each individual object, as shown in Figure 9A , 9B , 9C respectively).
[0088] In some implementation manners of the embodiments of the present application, the real label is obtained by calculating the feature fringe pattern; the real label includes: real wrapped phase, real absolute phase, real fringe order, and real parallax;
[0089] The feature fringe pattern and the sample fringe pattern correspond to different frequencies and / or different phases, and the frequencies between the feature fringe patterns with different frequencies and the sample fringe pattern are in an arithmetic progression.
[0090] As an example, 100 objects (as training objects in the fringe pattern) and 5000 data sets of scenes can be created in advance for generating training data. For the sample fringe patterns in the training data, 6 sine fringe frequencies (in hertz) are encoded as 47, 51, 56, 62, 69, and 77 respectively. In order to expand the data set to meet the requirements of various fringe frequencies in industry, the fringes of each frequency are phase-shifted 4 steps, that is, each step moves π / 2. Among them, every three frequencies can be combined to obtain a data set. For example, the data set with frequencies 47, 51, and 56 is a group, and the data set with frequencies 51, 56, and 62 is a group, etc.
[0091] Each training object is illuminated by 12 fringe patterns and then captured by a camera. In the embodiments of the present application, by setting an absolute phase unwrapping model with a specific structure, the absolute phase unwrapping model only needs a single-frame fringe pattern to output the corresponding wrapped phase and fringe order, and obtain the absolute phase. Therefore, at least one of the above 12 fringe images is required as the input of the absolute phase unwrapping model. Taking the fringe frequencies of 56, 62, and 69 as an example, assuming that one of the four images with a frequency of 56 is a sample fringe pattern, then the images of the remaining phases of the frequency of 56 except the sample fringe pattern are feature fringe patterns, and the four fringe images of each phase of the remaining frequencies 62 and 69 are also feature fringe patterns. The input is the first one (sample fringe pattern) with different phases in the fringe pattern with a frequency f = 56, and the other 11 images (characteristic fringe patterns) are used to generate the real label.
[0092] Taking the fringe frequencies of 56, 62, and 69 as an example, each training data sample includes an input matrix: the first one in the fringe pattern with f = 56, and four output matrices: the numerator of the wrapped phase (refer to Figure 10 , which shows a schematic diagram of the numerator term of a wrapped phase provided by the embodiments of the present application, Figure 10 is Figure 3 as the numerator term of the wrapped phase corresponding to the sample fringe pattern) and the denominator (refer to Figure 11 , which shows a schematic diagram of the numerator term of a wrapped phase provided by the embodiments of the present application, Figure 11 is Figure 3is the numerator term of the wrapped phase corresponding to the sample fringe pattern), the fringe order, and the absolute phase. The width of each data set is 1280 pixels and the height is 1024 pixels.
[0093] As an example, referring to Figure 12 , a method for generating a data set provided by an application embodiment is shown. Specifically, it may include: setting up a binocular structured light projection system as shown in Figure 2 to capture a fringe pattern.
[0094] Using the three-frequency four-step phase-shifting method to calculate the true wrapped phase, true fringes, and true absolute phase. The input of the above-mentioned wrapped phase generation model is a single-frame deformed fringe pattern, and the output is the real part and the imaginary part of the fringe pattern. The true value of the output can be obtained by means of the phase-shifting method. Assuming that the number of phase-shifting steps of the phase-shifting method is N, then there is
[0095]
[0096] Im represents the real part of the characteristic fringe pattern, Re represents the imaginary part of the characteristic fringe pattern, and the true wrapped phase is calculated through the real part and the imaginary part of the characteristic fringe pattern. Thus, the true wrapped phase can be calculated as:
[0097]
[0098] Then, the fringe order k and the absolute phase Φ are calculated by the multi-frequency heterodyne method, and the calculation formulas are as follows:
[0099]
[0100] where φ1 and φ2 are the wrapped phases of two adjacent frequencies (taking 56 and 62 as examples). To calculate the true fringe order, it is also necessary to calculate:
[0101] The equivalent period synthesized by the multi-frequency heterodyne method:
[0102] where the period T of the fringe pattern = 1280 / f, 1280 is the resolution (width) of the projection fringe pattern in the horizontal direction, and the calculation formula for the true fringe order k is:
[0103]
[0104] The round(·) function is used for rounding.
[0105]
[0106] Φ = φ1 + 2k1π (Formula (7))
[0107] In the embodiment of the present application, the three-frequency four-step phase-shifting method is used, so it is necessary to calculate the wrapped phase: φ 1,2,3 , φ1,2 , and φ 2,3 Similarly, for T and order k, three values also need to be calculated.
[0108] The dataset also includes a mask (Mask) determined as shown in Figures 3 - 5 . In practical applications, the mask of each independent object can be generated by using annotation software to annotate the feature fringe pattern.
[0109] The true disparity can be generated from the true absolute phase.
[0110] In some embodiments, the training data can also be other data calculated from one or more of the true wrapped phase, the true absolute phase, the true fringe order, and the true disparity.
[0111] In some implementation manners of the embodiments of the present application, the iterative training of the wrapped phase generation model and the fringe order segmentation model according to the predicted wrapped phase, the predicted fringe order, the mask, and the true label includes:
[0112] Determine a first loss function corresponding to the wrapped phase generation model and a second loss function corresponding to the fringe order segmentation model;
[0113] Iteratively train the wrapped phase generation model according to the sample fringe pattern, the true label, and the first loss function;
[0114] Iteratively train the fringe order segmentation model according to the predicted wrapped phase, the true label, the mask, and the second loss function;
[0115] The first loss function includes an output consistency loss term, a phase consistency constraint term, and a feature consistency loss term; the second loss function includes a cross-entropy loss term, a Dice coefficient loss term, and a phase matching loss term.
[0116] In the embodiments of the present application, a first loss function can be set for the wrapped phase generation model. The input of the wrapped phase generation model of the sample fringe pattern is used, and the first loss function is used to evaluate the difference between the output of the wrapped phase generation model and the true label, and iterative training is performed.
[0117] A second loss function can be set for the fringe order segmentation model. The predicted wrapped phase and the predicted fringe pattern obtained according to the predicted wrapped phase are used as the input of the fringe order segmentation model. The second loss function is used to evaluate the difference between the output of the fringe order segmentation model and the true label, and iterative training is performed.
[0118] Specifically, the first loss function is obtained by weighting the output consistency loss term, the phase consistency constraint term, and the feature consistency loss term, and the second loss function is obtained by weighted summation of the cross-entropy loss term, the Dice coefficient loss term, and the phase matching loss term.
[0119] In the embodiment of the present application, the wrapped phase generation task is a regression problem, and the wrapped phase generation model is a regression model. Each pixel in the input and output of the wrapped phase generation model has a quantitative result. Therefore, the output consistency loss term, the phase consistency constraint term, and the feature consistency loss term in the first loss function can be determined in the following manner.
[0120] Output consistency loss term: The loss function of the regression model usually needs to minimize the gap between the network prediction value and the output true value. In the embodiment of the present application, the mean absolute error function (MAE) is used to measure the gap between the output of the convolutional neural network and the label. The calculation formula is as follows:
[0121]
[0122] where y and respectively represent the true label and the output quantity predicted by the wrapped phase generation model. The subscripts Im and Re are the real part term and the imaginary part term respectively, N is the total number of effective pixel points in a single-frame fringe pattern, and ||·||1 represents calculating the absolute error between the true label and the predicted output.
[0123] Phase consistency constraint term: The output of the wrapped phase generation model is used for subsequent arctangent operation to solve the wrapped phase. Due to the sign degeneracy caused by the cosine function in the fringe intensity distribution, both φ and 2π - φ can be regarded as the correct wrapped phase. Therefore, minimizing the error is adopted to ensure the accuracy of phase demodulation. The formula is as follows:
[0124]
[0125] where y φ represents the phase true value (true wrapped phase) calculated by the phase shift method, is the network output (predicted wrapped phase), y φ and are both obtained by calculating the corresponding phase through the following formula:
[0126]
[0127] where min(·) represents taking the minimum value in the parameter list, y φ is obtained by performing arctangent calculation on y Im and y Re and is obtained from and is obtained by performing the arctangent calculation.
[0128] Feature consistency loss. Since the loss function that minimizes the pixel prediction error, such as MAE, will prompt the wrapped phase generation model to output the error of each pixel towards the average value, resulting in the loss of high-frequency details in the prediction result. To avoid the over-smoothing effect of the output caused by using only the per-pixel prediction error function as the loss function, the embodiment of the present application sets a perceptual loss module, inputs the output of the wrapped phase generation model and the ground truth label into a feature network respectively to extract features and minimize the gap between the features. As an example, a pre-trained VGG19 network is used as the feature extraction network, and the output of the middle layer of this network is used as the feature loss. The calculation formula is:
[0129]
[0130] where f i represents the feature input of the first five pooling operations of the convolutional module in the VGG19 network, and the weight parameters λ i are set to 1, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively.
[0131] In a specific implementation, the weighting coefficients of the output consistency loss term, the phase consistency constraint term, and the feature consistency loss term can be set according to actual requirements to determine the first loss function.
[0132] As an example, the first loss function is:
[0133] Loss = Loss L1 + 0.5Loss φ + 0.5Loss f
[0134] Referring to Figure 13 , a schematic flowchart of the data processing process of a stripe order segmentation model provided by an embodiment of the present application is shown.
[0135] The stripe order segmentation model of the embodiment of the present application adopts a neural network structure as shown in Figure 13 . The stripe order segmentation model is a single-input and dual-output structure, and the structure of each channel is the same as that of the wrapped phase generation network. The input of the stripe order segmentation model is the normalized wrapped phase of the left and right cameras (referring to Figure 14 , a schematic diagram of the wrapped phase of a left camera provided by an embodiment of the present application is shown; referring to Figure 15 , a schematic diagram of the wrapped phase of a right camera provided by an embodiment of the present application is shown) and the edge map (referring to Figure 16 , an edge map of the wrapped phase of a left camera provided by an embodiment of the present application is shown; referring to Figure 17, showing an edge map of the wrapped phase of the right camera provided by the embodiment of the present application; the stacking in the channel dimension, that is, the input of the fringe order segmentation model is a four-channel data. After being encoded by the encoder and decoded by the double decoder, the outputs of the two decoders are respectively the fringe orders corresponding to the left camera fringe map and the fringe orders corresponding to the right camera fringe map, and the dimension is the segmentation category. Taking f = 56 as an example, the segmentation category is 57 (including the background). The input image and the output channel have the same resolution size. In order to calculate the matching loss of each independent continuous object respectively, the output is multiplied by the mask obtained in the previous dataset acquisition stage at the same time, and each object (such as the above-mentioned training object) corresponds to a mask.
[0136] For the second loss function in the embodiment of the present application, it includes a cross-entropy loss term, a Dice coefficient loss term, and a phase matching loss term.
[0137] Cross-entropy loss term: In the image segmentation task, the cross-entropy loss is a key loss function and is widely used to evaluate the accuracy of the model's classification of each pixel. Similar to the classification task, the cross-entropy loss is used to quantify the difference between the pixel category distribution predicted by the model and the true label distribution. Its mathematical form can be expressed as:
[0138]
[0139] where i represents the index of the category, n represents the number of categories, and P(i) and Q(i) respectively represent the probabilities of the i-th category in the true distribution and the predicted distribution.
[0140] Dice coefficient loss term: Based on the Dice coefficient, it aims to improve the ability of the segmentation model to separate the foreground (i.e., the object in the fringe map) from the background. The Dice coefficient loss term directly optimizes the segmentation quality by maximizing the overlap between the prediction and the true label, thereby improving the performance of the model on the few-class samples. The formula is as follows:
[0141]
[0142] where X represents the pixel label of the true segmentation image, and Y represents the pixel category of the model-predicted segmentation image. |X∩Y| is the intersection between X and Y, which is approximately the dot product between the pixels of the predicted image and the pixels of the true label image, and the dot product results are added together. |X| and |Y| respectively represent the number of elements of X and Y.
[0143] Phase matching loss term: As Figure 13As shown, the fringe order output by the fringe order segmentation model needs to be used to calculate the predicted absolute phase with the predicted wrapped phase output by the wrapped phase generation model. In 3D reconstruction, it is necessary to find the absolute phase matching points and combine the geometric relationships to obtain the disparity, so as to obtain the depth information. At the same time, the calculated absolute phase needs to be multiplied by the mask obtained in the dataset acquisition stage to obtain the individual output of the absolute phase of each independent individual. At the same time, phase matching is performed on the same object to obtain the matching loss. The purpose of the matching loss is to avoid the order segmentation error caused by the interval between objects.
[0144] Referring to Figure 18 , a fringe order map obtained without using a mask provided by an embodiment of the present application is shown; referring to Figure 19 , a fringe order map obtained without using a mask provided by an embodiment of the present application is shown. For example: for an entire image, the original fringe order is from 1 to 57. However, in order to facilitate the fringe order segmentation model to learn simpler features and avoid the order segmentation error caused by the interval between objects, the fringe order of each continuous object is set to be from 1 to N (the size of N is based on the number of fringes occupied by the object). In this case, if the phase matching is still based on the entire image, it is very likely that the phases at multiple positions are the same (as shown in Figure 18 ), resulting in matching confusion. Therefore, for each individual continuous object, the embodiment of the present application first uses a mask for segmentation, and then performs separate phase matching for each mask, so as to obtain the fringe order distribution corresponding to each individual object (as shown in Figure 19 ). After obtaining the disparity, the disparities corresponding to each mask are added together, and the final correct disparity and depth results can also be obtained.
[0145] The binocular matching loss compares the corresponding pixels in the left and right phase views to ensure that the model can accurately segment the fringe order, so that matching points can be found in the calculated absolute phase. The matching loss forces the model to learn the consistency between views during the training process, helping to reduce the matching errors caused by fringe order classification errors. It improves the robustness of the model to noise and occlusion, etc., and helps the model perform more reliably in real-world applications.
[0146] In specific implementation, the weighting coefficients of the cross-entropy loss term, Dice coefficient loss term, and phase matching loss term can be set according to actual needs to determine the second loss function.
[0147] As an example, the second loss function is:
[0148] Loss = Loss cross + Loss diceloss + Loss disloss
[0149] Among them, LOSS disloss is the phase matching loss term.
[0150] The determination process of the matching loss term is further described below. In some implementation manners of the embodiments of the present application, the true label includes a true disparity; the phase matching loss term is determined through the following steps:
[0151] Generate a predicted absolute phase according to the predicted wrapped phase and the predicted fringe order.
[0152] Generate an individual predicted phase for each of the training objects according to the predicted absolute phase and the mask; the individual predicted phase includes an individual left predicted phase and an individual right predicted phase.
[0153] Determine the number of valid pixels and the predicted disparity according to the pixel points in the body left predicted phase and the pixel points in the individual right predicted phase.
[0154] Determine the phase matching loss term according to the predicted disparity, the number of valid pixels, and the true disparity.
[0155] The specific absolute phase matching process is to calculate the disparity by means of binary search. First, relevant parameters are initialized, including the nearest and farthest depth ranges (znear and zfar), and the maximum disparity dmax and minimum disparity dmin calculated according to the projection matrices P1 and P2 (obtained from camera calibration) and the reconstruction matrix Q (obtained from camera calibration). Then, the algorithm traverses each row of the left image (individual left predicted phase) and the right image (individual right predicted phase), determines the valid pixels and the predicted disparity through the corresponding pixels, and the pixels of the same corresponding point in the individual left predicted phase and the individual right predicted phase, and further determines the phase matching loss term.
[0156] In some implementation manners of the embodiments of the present application, the determining the number of valid pixels and the predicted disparity according to the pixel points in the body left predicted phase and the pixel points in the individual right predicted phase includes:
[0157] Determine the first non-zero pixel in the body left predicted phase.
[0158] Traverse the first non-zero pixel to determine a second non-zero pixel in the individual right predicted phase that matches the first non-zero pixel.
[0159] Calculate the predicted disparity according to the mutually matching first non-zero pixel and second non-zero pixel.
[0160] Determine that the number of mutually matching first non-zero pixels and second non-zero pixels for which the predicted disparity satisfies the disparity threshold is the number of valid pixels.
[0161] During the process of traversing the left and right images, record the positions of non-zero pixels in each row and store them in the valueL and valueR lists respectively. Then, for each non-zero pixel in the left image, the algorithm uses the binary search method in the right image to find the matching pixel, and judges its disparity by comparing the pixel values of the two. If the calculated disparity meets the set threshold (for example: threshold = 0.03), the matching result and its disparity value are stored in the disparity list. Finally, the algorithm returns the result in the form of a NumPy array, ensuring that each element contains the corresponding coordinate points and disparity values.
[0162] Then it is necessary to minimize the gap between the disparity calculated from the network prediction value and the true disparity. In the embodiments of the present application, the mean absolute error function (Mean Absolute Error, MAE) is used to measure the gap between the output of the convolutional neural network and the label:
[0163]
[0164] where y dis is the true disparity, is the disparity calculated from the predicted value, and N is the total number of effective pixel points in a single-frame fringe pattern.
[0165] In some implementation manners of the embodiments of the present application, the wrapped phase generation model has a multi-layer U-shaped network structure; the U-shaped network structure is provided with a residual module or a dense residual connection module;
[0166] The residual module includes a plurality of consecutive first convolutional layers, and the output end of the first convolutional layer is connected to a non-linear activation function layer;
[0167] The dense residual connection module includes a residual component, a second convolutional layer, and a skip connection module connected in sequence. The residual component includes a plurality of residual modules connected in sequence; the input ends of the respective residual modules before the last residual module are connected to the output ends of the respective residual modules after it; the input end of the first residual module is connected to the skip connection module.
[0168] Refer to Figure 20 , which shows a schematic diagram of the processing flow of a wrapped phase generation model provided by the embodiments of the present application. The wrapped phase generation model in the embodiments of the present application adopts the convolutional neural network structure as shown in Figure 20 . The input of the wrapped phase generation model is a single-frame normalized fringe pattern, and finally a stacked output of two channels of the numerator and denominator of the predicted wrapped phase can be obtained. The input image and the output channels have the same resolution size. After normalization processing, the input fringes are in the same gray level interval, which is beneficial to improving the accuracy of the neural network.
[0169] The main architecture of the product neural network that wraps the phase generation model is a five - layer U - Net structure. The feature processing of each layer is realized by a residual module or a residual dense connection module (Residual Dense Block, RDB) to achieve feature reuse. The feature information learned by the network from a single - frame fringe pattern is mainly the gradient relationship between adjacent pixels within the receptive field. Since the dataset contains fringe patterns of different frequencies, and the number of pixels contained in a single fringe period in sparse fringe modulation patterns and dense fringe modulation patterns is different, the neural network should have a sufficient receptive field and multi - scale ability to process fringe patterns of different frequencies simultaneously. U - Net is a typical encoder - decoder structure that extracts features of different scales by scaling the feature maps multiple times. The residual module only contains two consecutive convolutional layers with a kernel size of 3×3, and each convolutional layer is followed by a non - linear activation function ReLU layer.
[0170] Refer to Figure 21 , which shows a schematic diagram of a residual dense connection module provided by an embodiment of the present application. The specific composition of the above - mentioned RDB is as Figure 21 shown. The dense stacking operation enables the network to extract and reuse more features, and dilated convolutions are used to expand the receptive field. Due to the holes in the convolutional kernel, the image features covered by the dilated convolutional layer are not continuous, and in consecutive stacked dilated convolutional layers, the dilation rates cannot have a common divisor greater than 1, otherwise the checkerboard effect will occur. Therefore, the convolution dilation rates of the RDB module in this article are set to 1, 2, 3, 2, 1 from left to right in sequence.
[0171] After training the absolute phase unwrapping model, the trained phase unwrapping model can be used for three - dimensional reconstruction. Referring to the process of binocular stereo matching described in the phase matching loss term above, assuming that the disparity obtained through binocular matching is disparity, first normalize the disparity (for easy display, the disparity image can be scaled to the range of 0 to 255), and the depth of each pixel can be obtained through the following formula:
[0172]
[0173] where f is the focal length of the camera, B is the camera baseline distance (the distance between the centers of the two cameras), and d is the disparity, that is, disparity.
[0174] According to the calculated depth information, by finding the corresponding points of the binocular absolute phase, combining the calibrated system parameters and the principle of triangulation ranging, the spatial coordinates of the object points in the camera coordinate system are solved, and the position of each pixel is converted into three - dimensional coordinates to generate a point cloud. The coordinates of the points can be obtained through the following formula:
[0175]
[0176]
[0177] Among them, (u, v) are pixel coordinates, and (c x , c y ) is the optical center of the camera.
[0178] Referring to Figure 22 , a schematic diagram of a three-dimensional reconstruction effect provided by an embodiment of the present application is shown. In the embodiment of the present application, a trained phase unwrapping model can be used to perform three-dimensional reconstruction on an independent object in a single-frame fringe pattern, and the obtained three-dimensional reconstruction effect is as Figure 22 shown.
[0179] Referring to Figure 23 , another schematic diagram of a three-dimensional reconstruction effect provided by an embodiment of the present application is shown. In the embodiment of the present application, a trained phase unwrapping model can be used to perform three-dimensional reconstruction on each independent object in a single-frame fringe pattern, and the obtained three-dimensional reconstruction effect is as Figure 23 shown.
[0180] Referring to Figure 24 , a schematic diagram of a three-dimensional reconstruction method provided by an embodiment of the present application is shown, which specifically may include the following steps:
[0181] Step 2401: Obtain an absolute phase unwrapping model, a to-be-processed fringe pattern collected by a pre-set binocular structured light projection system, and projection parameters of the projection system;
[0182] Step 2402: Input the to-be-processed fringe pattern into the absolute phase unwrapping model;
[0183] Step 2403: Receive the target absolute phase output by the absolute phase unwrapping model for the to-be-processed fringe pattern;
[0184] Step 2404: Perform three-dimensional reconstruction based on the target absolute phase and the projection parameters to generate three-dimensional reconstruction data matching the to-be-processed fringe pattern;
[0185] The absolute phase unwrapping model is trained according to the training method of the absolute phase unwrapping model as described above.
[0186] In an implementation manner of the embodiment of the present application, the target absolute phase includes an individual left target phase and an individual right target phase of each target object; the performing three-dimensional reconstruction based on the target absolute phase and the projection parameters to generate three-dimensional reconstruction data matching the to-be-processed fringe pattern includes:
[0187] For each of the target objects, calculate the target disparity based on the pixel points in the individual left target phase and the pixel points in the individual right target phase;
[0188] Based on the target disparity and the projection parameters, determine the depth information of each of the target objects;
[0189] Convert the depth information according to the projection parameters to generate three-dimensional reconstruction data corresponding to the to-be-processed fringe pattern.
[0190] For the three-dimensional reconstruction embodiment, the processing process of the absolute phase unwrapping model for the to-be-processed fringe pattern is similar to the processing process of the absolute phase unwrapping model for the sample fringe pattern during training. For the relevant parts, refer to the embodiment of the training method of the absolute phase unwrapping model, and details will not be elaborated here.
[0191] For the three-dimensional reconstruction process, refer to the embodiment of the training method of the absolute phase unwrapping model. The three-dimensional reconstruction process using the trained absolute phase unwrapping model includes a binocular matching process and a disparity calculation process, which will not be elaborated here.
[0192] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0193] Refer to Figure 25 , which shows a schematic diagram of a training device for an absolute phase unwrapping model provided by an embodiment of the present application. The absolute phase unwrapping model includes: a wrapped phase generation model and a fringe order segmentation model connected to each other; the device may specifically include:
[0194] A training data acquisition module 2501, configured to acquire training data; the training data includes a sample fringe pattern, a true label corresponding to the sample fringe pattern, and a mask corresponding to each training object; the sample fringe pattern includes a plurality of training objects;
[0195] A wrapped phase generation model input module 2502, configured to input the sample fringe pattern into the wrapped phase generation model to obtain a predicted wrapped phase;
[0196] A fringe order segmentation model input module 2503, configured to input the predicted wrapped phase and a predicted edge map corresponding to the predicted wrapped phase into the fringe order segmentation model to obtain a predicted fringe order corresponding to each training object;
[0197] A training module 2504, configured to iteratively train the wrapped phase generation model and the fringe order segmentation model according to the predicted wrapped phase, the predicted fringe order, the mask, and the ground truth label until the training of the wrapped phase generation model and the fringe order segmentation model is completed; wherein, the predicted edge map is generated from the predicted wrapped phase, and the sample fringe pattern includes a left camera sample image and a right camera sample image collected by a pre-set binocular structured light projection system.
[0198] In some implementation manners of the embodiments of the present application, the training module 2504 includes:
[0199] A loss function determination sub-module, configured to determine a first loss function corresponding to the wrapped phase generation model and a second loss function corresponding to the fringe order segmentation model; a wrapped phase generation model training sub-module, configured to iteratively train the wrapped phase generation model according to the sample fringe pattern, the ground truth label, and the first loss function; a fringe order segmentation model training sub-module, configured to iteratively train the fringe order segmentation model according to the predicted wrapped phase, the ground truth label, the mask, and the second loss function; the first loss function includes an output consistency loss term, a phase consistency constraint term, and a feature consistency loss term; the second loss function includes a cross-entropy loss term, a Dice coefficient loss term, and a phase matching loss term.
[0200] In some implementation manners of the embodiments of the present application, the ground truth label includes a ground truth disparity; the phase matching loss term is determined by the following modules: a predicted absolute phase generation module, configured to generate a predicted absolute phase according to the predicted wrapped phase and the predicted fringe order; an individual predicted phase generation module, configured to generate an individual predicted phase for each of the training objects according to the predicted absolute phase and the mask; the individual predicted phase includes an individual left predicted phase and an individual right predicted phase; a pixel processing module, configured to determine the number of valid pixels and the predicted disparity according to the pixel points in the individual left predicted phase and the pixel points in the individual right predicted phase; a phase matching loss term determination module, configured to determine the phase matching loss term according to the predicted disparity, the number of valid pixels, and the ground truth disparity.
[0201] In some implementation manners of the embodiments of the present application, the pixel processing module includes: a first non-zero pixel determination sub-module, configured to determine a first non-zero pixel in the left volume prediction phase; a first non-zero pixel traversal sub-module, configured to traverse the first non-zero pixel to determine a second non-zero pixel matching the first non-zero pixel in the right individual prediction phase; a predicted disparity calculation sub-module, configured to calculate a predicted disparity based on the mutually matching first non-zero pixel and second non-zero pixel; and an effective pixel number determination sub-module, configured to determine that the number of mutually matching first non-zero pixels and second non-zero pixels for which the predicted disparity satisfies a disparity threshold is the effective pixel number.
[0202] In some implementation manners of the embodiments of the present application, the true label is obtained by calculating a feature fringe pattern; the true label includes: a true wrapped phase, a true absolute phase, a true fringe order, and a true disparity; the feature fringe pattern corresponds to different frequencies and / or different phases from the sample fringe pattern, and the frequencies between the feature fringe patterns of different frequencies and the sample fringe pattern are in an arithmetic progression.
[0203] In some implementation manners of the embodiments of the present application, the wrapped phase generation model has a multi-layer U-shaped network structure; the U-shaped network structure is provided with a residual module or a dense residual connection module; the residual module includes a plurality of consecutive first convolutional layers, and the output end of the first convolutional layer is connected to a non-linear activation function layer; the dense residual connection module includes a residual component, a second convolutional layer, and a skip connection module connected in sequence, the residual component includes a plurality of residual modules connected in sequence; the input end of each residual module before the last residual module is connected to the output end of each residual module after it; the input end of the first residual module is connected to the skip connection module.
[0204] A training device for an absolute phase unwrapping model provided by the embodiments of the present application can implement each step in the foregoing embodiments of the training method of the absolute phase unwrapping model when the device is applied.
[0205] Refer to Figure 26 , which shows a schematic diagram of a three-dimensional reconstruction device provided by the embodiments of the present application, and specifically may include:
[0206] An absolute phase unwrapping model acquisition module 2601, configured to acquire an absolute phase unwrapping model, a to-be-processed fringe pattern collected by a pre-set binocular structured light projection system, and projection parameters of the projection system;
[0207] A to-be-processed fringe pattern input module 2602, configured to input the to-be-processed fringe pattern into the absolute phase unwrapping model;
[0208] The target absolute phase receiving module 2603 is configured to receive the target absolute phase output by the absolute phase unwrapping model for the to-be-processed fringe pattern.
[0209] The 3D reconstruction module 2604 is configured to perform 3D reconstruction based on the target absolute phase and the projection parameters to generate 3D reconstruction data matching the to-be-processed fringe pattern; wherein the absolute phase unwrapping model is trained according to the training method of the absolute phase unwrapping model as described above.
[0210] In an implementation manner of the embodiment of the present application, the target absolute phase includes an individual left target phase and an individual right target phase of each target object; the 3D reconstruction module 2604 includes: a target disparity calculation sub-module, configured to calculate a target disparity for each of the target objects based on the pixel points in the individual left target phase and the pixel points in the individual right target phase; a depth information determination sub-module, configured to determine the depth information of each of the target objects based on the target disparity and the projection parameters; and a 3D reconstruction data generation sub-module, configured to convert the depth information according to the projection parameters to generate 3D reconstruction data of the to-be-processed fringe pattern.
[0211] A 3D reconstruction device provided by an embodiment of the present application can implement each step in the foregoing 3D reconstruction method embodiment when this device is applied.
[0212] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the description in the method embodiment section.
[0213] Refer to Figure 27 , which shows a schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 27 shown, the electronic device 2700 in the embodiment of the present application includes: a processor 2710, a memory 2720, and a computer program 2721 stored in the memory 2720 and executable on the processor 2710. When the processor 2710 executes the computer program 2721, the steps in each embodiment of the foregoing training method of the absolute phase unwrapping model are implemented, and / or, the steps in each embodiment of the 3D reconstruction method are implemented. Alternatively, when the processor 2710 executes the computer program 2721, the functions of each module / unit sub-module in each embodiment of the foregoing training of the absolute phase unwrapping model are implemented, and / or, the functions of each module / sub-module in the 3D reconstruction device embodiment are implemented.
[0214] The electronic device 2700 may be a computing device such as a desktop computer or a cloud server. The electronic device 2700 may include, but is not limited to, a processor 2710 and a memory 2720. Those skilled in the art can understand that Figure 27This is merely an example of the electronic device 2700 and does not constitute a limitation on the electronic device 2700. It may include more or fewer components than those shown, or combine certain components, or have different components. For example, the electronic device 2700 may also include input / output devices, network access devices, buses, etc.
[0215] The processor 2710 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0216] The memory 2720 may be an internal storage unit of the electronic device 2700, such as the hard disk or memory of the electronic device 2700. The memory 2720 may also be an external storage device of the electronic device 2700, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the electronic device 2700. Further, the memory 2720 may also include both the internal storage unit and the external storage device of the electronic device 2700. The memory 2720 is used to store the computer program 2721 and other programs and data required by the electronic device 2700. The memory 2720 may also be used to temporarily store data that has been output or is to be output.
[0217] The embodiments of this application also disclose a computer-readable storage medium storing a computer program, which when executed by a processor implements the absolute phase unwrapping model training method as described in the foregoing embodiments of the first aspect or the 3D reconstruction method as described in the foregoing embodiments.
[0218] The fifth aspect of the embodiments of this application provides a computer program product including a computer program, which when run enables the absolute phase unwrapping model training method as described in the foregoing embodiments or the 3D reconstruction method as described in the foregoing embodiments.
[0219] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A training method for an absolute phase unwrapping model, characterized in that: The absolute phase unwrapping model includes: an interconnected wrapping phase generation model and a fringe level segmentation model; Acquire training data; the training data includes a sample fringe image, a true label corresponding to the sample fringe image, and a mask corresponding to each of the training objects; the sample fringe image includes a plurality of training objects; Inputting the sample fringe pattern into the wrapping phase generation model to obtain a predicted wrapping phase; Inputting the predicted wrapping phase and the predicted edge map corresponding to the predicted wrapping phase into the fringe level segmentation model to obtain the predicted fringe level corresponding to each of the training objects; Iteratively training the wrapped phase generation model and the fringe order segmentation model according to the predicted wrapped phase, the predicted fringe order, the mask, and the true label until the training of the wrapped phase generation model and the fringe order segmentation model is completed; The predicted edge map is generated by the predicted wrapped phase, and the sample fringe map includes a left camera sample map and a right camera sample map collected by a preset binocular structured light projection system.
2. The method according to claim 1, characterized in that: The iterative training of the wrapping phase generation model and the fringe level segmentation model according to the predicted wrapping phase, the predicted fringe level, the mask and the true label comprises: Determine a first loss function corresponding to the wrapped phase generation model and a second loss function corresponding to the fringe level segmentation model; Iteratively training the wrapped phase generation model according to the sample fringe image, the true label and the first loss function; Iteratively training the stripe level segmentation model according to the predicted wrapping phase, the true label, the mask and the second loss function; The first loss function includes an output consistency loss term, a phase consistency constraint term, and a feature consistency loss term; the second loss function includes a cross entropy loss term, a Dice coefficient loss term, and a phase matching loss term.
3. The method according to claim 2, characterized in that The true label includes the true disparity; the phase matching loss term is determined by the following steps: generating a predicted absolute phase according to the predicted wrapping phase and the predicted fringe order; Generating an individual predicted phase for each of the training objects according to the predicted absolute phase and the mask; the individual predicted phase includes an individual left predicted phase and an individual right predicted phase; Determine the number of valid pixels and the predicted disparity according to the pixel points in the left predicted phase of the body and the pixel points in the right predicted phase of the body; A phase matching loss term is determined according to the predicted disparity, the number of valid pixels and the actual disparity.
4. The method according to claim 3, characterized in that The determining of the number of valid pixels and the predicted disparity according to the pixel points in the body left predicted phase and the pixel points in the body right predicted phase comprises: determining a first non-zero pixel in the body left prediction phase; traversing the first non-zero pixels to determine a second non-zero pixel matching the first non-zero pixel in the individual right prediction phase; Calculating a predicted disparity according to a first non-zero pixel and a second non-zero pixel that match each other; The number of first non-zero pixels and second non-zero pixels that match each other and whose predicted disparity satisfies a disparity threshold is determined as the number of valid pixels.
5. The method according to any one of claims 1 to 4, characterized in that: The true label is obtained by calculating the characteristic fringe image; The real labels include: real package phase, real absolute phase, real fringe order, and real parallax; The characteristic fringe pattern and the sample fringe pattern correspond to different frequencies and / or different phases, and the frequencies between the characteristic fringe pattern and the sample fringe pattern of different frequencies are distributed in an equidistant manner.
6. The method according to any one of claims 1 to 4, characterized in that: The wrapped phase generation model has a multi-layer U-shaped network structure; the U-shaped network structure is provided with a residual module or a dense residual connection module; The residual module comprises a plurality of consecutive first convolutional layers, and the output end of the first convolutional layer is connected to the nonlinear activation function layer; The dense residual connection module includes a residual component, a second convolutional layer and a jump connection module connected in sequence, and the residual component includes multiple residual modules connected in sequence; the input end of each residual module located before the last residual module is connected to the output end of each residual module located after it; the input end of the first residual module is connected to the jump connection module.
7. A three-dimensional reconstruction method, characterized in that: include: Acquire an absolute phase unwrapping model, a fringe pattern to be processed collected by a preset binocular structured light projection system, and projection parameters of the projection system; Inputting the fringe pattern to be processed into the absolute phase unwrapping model; Receiving a target absolute phase output by the absolute phase unwrapping model for the fringe pattern to be processed; Performing three-dimensional reconstruction according to the target absolute phase and the projection parameters to generate three-dimensional reconstruction data matching the fringe pattern to be processed; The absolute phase unwrapping model is trained according to the method described in claims 1-6.
8. The method according to claim 7, characterized in that The target absolute phase includes an individual left target phase and an individual right target phase of each target object; The three-dimensional reconstruction is performed according to the target absolute phase and the projection parameters to generate three-dimensional reconstruction data matching the fringe pattern to be processed, including: For each of the target objects, calculating a target disparity according to pixel points in the individual left target phase and pixel points in the individual right target phase; Determining depth information of each of the target objects according to the target disparity and the projection parameters; The depth information is converted according to the projection parameters to generate three-dimensional reconstruction data corresponding to the fringe pattern to be processed.
9. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method as described in any one of claims 1-6 or 7-8.
10. A computer program product, characterized in that The invention comprises a computer program, which, when being executed, enables the method as claimed in any one of claims 1 to 6 or 7 to 8 to be performed.