3D profile measurement method and system based on semi-supervised reversible phase unwrapping network
Patent Information
- Application Number
- CN202610930840.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-06-26
AI Technical Summary
[0007]基于上述背景技术所提出的问题,本发明的目的在于提供基于半监督可逆相位偏折网络的3D轮廓测量方法及系统,解决了现有深度学习方法多采用全监督学习策略,依赖大量精确配对的条纹和3D轮廓数据集进行训练,不适用于实际产线的问题
[0046]本发明仅需约主流全监督方法26%的标注数据即可达到更优的重建精度,在高反光显示面板在线质量检测领域具有广阔应用前景。
Smart Images

Figure CN122453860B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of screen contour image processing technology, specifically to a 3D contour measurement method and system based on a semi-supervised reversible phase deflection network. Background Technology
[0002] In the manufacturing process of display panels for smartphones, wearable devices, computers, and televisions, the microscopic three-dimensional morphology of the screen surface directly determines the product's appearance quality and user experience. Defects such as scratches, dents, creases, foreign objects, and dirt not only affect visual aesthetics but can also lead to touch malfunctions or display abnormalities. With the popularization of new form factors such as full-screen, foldable, and curved screens, the increased screen curvature and enhanced reflectivity place higher demands on three-dimensional defect detection. If defects are missed and flow into subsequent assembly stages, it will lead to the scrapping of the entire device, resulting in significant material waste and cost losses. Therefore, developing a high-precision, high-efficiency 3D contour detection system for screens is of great significance for improving yield rates and reducing production costs.
[0003] Currently, mainstream screen inspection methods in the industry can be divided into two categories: contact and non-contact. Contact methods, such as probe profilometers and coordinate measuring machines (CMMs), offer high accuracy but are slow, easily scratch the screen surface, and are unsuitable for full-line inspection requirements. Among non-contact methods, 2D machine vision inspection is the most common. It uses a high-resolution camera to acquire images of the screen surface and then uses traditional image processing or deep learning classification networks to identify defects. However, 2D methods struggle to quantify depth information of defects, such as pit depth and crease undulations, and are easily affected by the specular effect of highly reflective surfaces. The specular reflection of the screen glass maps light sources and environmental textures into the image, confusing them with real defects and leading to missed or false detections. While laser triangulation can acquire depth information, it has poor adaptability to specular surfaces, and laser stripes are prone to overexposure or signal loss in high-light areas. Structured light projection suffers from low signal-to-noise ratio because the projected sinusoidal stripes are difficult for the camera to effectively receive due to the specular reflection characteristics of the screen surface. All of the above methods have limitations when dealing with highly reflective and specular surfaces.
[0004] Phase Measuring Deflectometry (PMD) utilizes the modulation effect of screen surface normal changes on reflected light to infer the gradient of the measured surface from deformed fringes, and then integrates this gradient to obtain the 3D morphology of the surface. This method is naturally suitable for detecting highly reflective and mirror-like surfaces, and theoretically can solve the problem of screen detection. However, traditional PMD methods rely on complex computational processes such as phase unwrapping and gradient integration. Phase unwrapping is susceptible to noise interference, resulting in phase jump errors, and the gradient integration process is sensitive to boundary conditions and integration paths, easily affected by error accumulation, thus reducing reconstruction accuracy and impacting subsequent depth-based quality detection. In addition, traditional methods require accurate system calibration and multi-directional phase-shifting fringe projection, and each measurement is time-consuming, making it difficult to balance accuracy and efficiency in practical measurement applications.
[0005] In recent years, deep learning methods have been introduced into the field of phase measurement deflection, significantly improving measurement efficiency and accuracy by learning intermediate processes through neural networks or directly achieving end-to-end mapping from deformed fringe patterns to 3D contours. Based on the different tasks undertaken by the networks, existing research can be broadly categorized into three types: phase extraction and enhancement, gradient integral solving, and end-to-end direct reconstruction. Regarding phase extraction and enhancement, Suresh et al. introduced a deep learning-based unfolded phase map enhancement method. Their PMENet neural network takes the unfolded phase predicted by Fourier transform as input and labels the high-precision unfolded phase obtained through eighteen phase shifts, effectively reducing phase noise in the single-frame Fourier transform method. Feng et al. from Nanjing University of Science and Technology introduced Bayesian convolutional neural networks into the field of fringe analysis to evaluate the uncertainty of the network model and data, providing a confidence measure for the phase solution results. In terms of gradient integral solving, Dou et al. from China Jiliang University designed a neural network-based surface reconstruction scheme for the integral process of recovering the surface shape from the gradient in phase measurement deflection. The input of this network is the gradient in the X and Y directions, and the output is the height of the object to be measured. The authors used both the proposed method and the traditional Southwell integral algorithm to reconstruct the 3D surface shape of freeform surfaces, and compared the results with real surface shape data obtained by the ZYGO interferometer. The results show that the proposed neural network scheme can achieve 3D reconstruction results close to those of the Southwell integral algorithm, with a reconstruction time of only about 0.81 seconds, significantly improving computational efficiency. Furthermore, some studies have attempted to construct end-to-end fringe-to-depth mapping models, directly learning the complex nonlinear mapping relationship between deformed fringe patterns and 3D contours, bypassing intermediate steps such as phase unwrapping and gradient integration, further simplifying the measurement process.
[0006] However, the aforementioned deep learning methods mostly employ fully supervised learning strategies, relying on large amounts of precisely paired fringe and 3D contour datasets for training. In actual production lines, creating high-precision fringe depth pairing datasets faces significant challenges. On one hand, expensive laser profilometers or white-light interferometers are needed to obtain the true depth values of the screen surface; these devices are costly and slow, making large-scale deployment difficult. On the other hand, the registration process is extremely cumbersome; the deformed fringe map and the true depth map must be precisely aligned at the pixel level, and any slight offset will cause the network to learn incorrect mappings. Since the scale of screen surface defects such as scratches and pits is typically in the micrometer to millimeter range, the registration error must be controlled at the sub-pixel level, which is almost impossible to achieve in batches in practice. Furthermore, publicly available phase measurement deflection datasets are limited in size and cover only a single screen type, making it difficult to cover complex measurement environments with multiple models, curvatures, and materials. Summary of the Invention
[0007] Based on the problems mentioned above, the purpose of this invention is to provide a 3D contour measurement method and system based on a semi-supervised reversible phase deflection network. This solves the problem that existing deep learning methods mostly adopt fully supervised learning strategies, rely on a large number of precisely paired stripe and 3D contour datasets for training, and are not suitable for actual production lines.
[0008] This invention is achieved through the following technical solution:
[0009] The first aspect of this invention provides a 3D contour measurement method based on a semi-supervised reversible phase deflection network, comprising the following steps:
[0010] The original deformed fringe pattern collected and modulated by the specular reflection of the screen under test is divided into labeled data and unlabeled data;
[0011] The unlabeled data is input into the second reconstructor of the teacher branch of the reversible phase deflection network to generate a predicted 3D profile. Pseudo-labels are then selected from the predicted 3D profiles and used as supervision signals for the student branch of the reversible phase deflection network.
[0012] The labeled and unlabeled data are input into the first reconstructor of the student branch of the reversible phase deflection network to generate the reconstructed 3D profile.
[0013] The reconstructed 3D contour is re-rendered as a re-rendered stripe pattern using a renderer with a student branch of a reversible phase deflection network.
[0014] In the above technical solution, unlabeled data is first input into the teacher branch, and the second reconstructor in the teacher branch performs forward processing on the unlabeled data to transform the stripe pattern into a 3D contour, generating a predicted 3D contour. Then, a designed pseudo-label threshold screening mechanism selects high-quality samples from the predicted 3D contour as pseudo-labels, and these pseudo-labels are used as supervision signals for the student branch, thereby achieving semi-supervised reversible phase deflection networks. Compared with fully supervised reversible phase deflection networks, this achieves high-precision screen 3D contour reconstruction driven by a small amount of labeled data.
[0015] After pseudo-labels are generated, a small amount of labeled data and a large amount of unlabeled data are randomly input into the student branch. The student branch includes a first reconstructor and a renderer. Under the supervision of the pseudo-labels, the student branch uses the first reconstructor to perform forward processing on the unlabeled data to transform the stripe pattern into a 3D contour, generating a reconstructed 3D contour, thus achieving forward reconstruction of stripes and 3D contours. Then, the reconstructed 3D contour is input into the renderer in the student branch for reverse processing, generating a re-rendered stripe pattern, thus achieving reverse rendering of 3D contours and stripes. The bidirectional mapping characteristics of reversible networks are used to achieve both forward reconstruction of stripes and 3D contours and reverse rendering of 3D contours and stripes.
[0016] In one alternative embodiment, the first reconstructor, the second reconstructor, and the renderer are all built based on a U-shaped reversible network;
[0017] In this process, the forward process of the U-shaped reversible network is used as the first reconstructor of the student branch and the second reconstructor of the teacher branch to achieve forward modulation from the deformed stripe pattern to the 3D contour.
[0018] Using the reverse process of the U-shaped reversible network as a renderer, the reverse modulation of 3D contours into deformable stripe patterns is achieved.
[0019] In one optional embodiment, the U-shaped reversible network includes: a plurality of sequentially connected bijective function modules, each bijective function module including reversible convolution and affine coupling operation units.
[0020] In one optional embodiment, positive modulation from deformed fringe pattern to 3D contour is achieved through a plurality of sequentially connected bijective function modules, including the following steps:
[0021] The deformed stripe pattern is split into a first positive dual-path feature and a second positive dual-path feature by channel dimension through reversible convolution;
[0022] The first and second positive dual-path features are input into the lightweight affine coupling module for feature extraction, resulting in the transformed first and second positive dual-path features.
[0023] The first positive dual-path feature and the second positive dual-path feature of the transformation are spliced together to obtain the 3D contour.
[0024] In one optional embodiment, the first positive dual-path features and the second positive dual-path features are input into a lightweight affine coupling module for feature extraction, including the following steps:
[0025] The second positive dual-path feature is input into the first lightweight affine coupling module. The output data of the first lightweight affine coupling module is added element by element to the first positive dual-path feature to obtain the transformed first positive dual-path feature.
[0026] The first positive dual-path feature of the transformation is input into the second lightweight affine coupling module and the third lightweight affine coupling module respectively. In the second lightweight affine coupling module, the first positive dual-path feature of the transformation is subjected to element-wise multiplication and element-wise addition operations in sequence through a nonlinear control function.
[0027] The data obtained after processing by the nonlinear control function is multiplied element-wise with the second positive dual-path feature, and the data obtained by element-wise multiplication is added element-wise with the output data of the third lightweight affine coupling module to obtain the transformed second positive dual-path feature.
[0028] In one optional embodiment, the lightweight affine coupling module includes: a downsampling convolutional layer, four cascaded depthwise separable pooling convolutional blocks, a pyramidal double self-similarity module, four cascaded depthwise separable deconvolutional blocks, and an upsampling convolutional layer connected in sequence.
[0029] The four cascaded depthwise separable pooling convolutional blocks correspond one-to-one with the four cascaded depthwise separable deconvolutional blocks, and the output of each depthwise separable pooling convolutional block is fed to the input of the corresponding depthwise separable deconvolutional block through same-scale jump connections.
[0030] In one alternative embodiment, a downsampling convolutional layer is used to perform preliminary feature extraction on the input features;
[0031] The extracted preliminary features are downsampled using four cascaded depthwise separable pooling convolutional blocks to obtain downsampled features;
[0032] The pyramid dual self-similarity module is used to enhance the downsampled features and generate enhanced features.
[0033] The enhanced features are upsampled through four cascaded depthwise separable deconvolution blocks to generate upsampled features;
[0034] The upsampled features are integrated across channels using an upsampled convolutional layer.
[0035] In one alternative embodiment, the loss function of the reversible phase deflection network includes: a supervised loss for labeled data and an unsupervised loss for unlabeled data, with the contributions of the supervised loss and the unsupervised loss balanced by a weighting coefficient;
[0036] Specifically, a bidirectional design is used to construct the supervised loss and the unsupervised loss.
[0037] In one optional embodiment, the supervised loss includes: a forward reconstruction loss and a backward rendering loss; wherein the supervised loss is constructed using a bidirectional design, including:
[0038] Obtain the true 3D contour value of the labeled data, calculate the labeled data contour loss between the true 3D contour value of the labeled data and the corresponding reconstructed 3D contour using root mean square error, and construct the forward reconstruction loss using the labeled data contour loss.
[0039] Obtain the true value of the stripe pattern of the labeled data, and use L1 loss and structural similarity to calculate the L1 loss and similarity loss of the labeled data stripe pattern between the true value of the labeled data stripe pattern and the corresponding re-rendered stripe pattern. Combine the L1 loss and similarity loss of the labeled data stripe pattern to construct the inverse rendering loss.
[0040] A second aspect of the present invention provides a 3D contour measurement system based on a semi-supervised reversible phase deflection network, comprising:
[0041] The segmentation module is used to divide the original deformed fringe pattern, which is modulated by the specular reflection of the screen under test, into labeled data and unlabeled data.
[0042] A semi-supervised module is used to input the unlabeled data into the second reconstructor of the teacher branch of the reversible phase deflection network, generate a predicted 3D profile, filter out pseudo-labels from the predicted 3D profile, and use the pseudo-labels as supervision signals for the student branch of the reversible phase deflection network.
[0043] The stripe-3D contour module is used to input labeled and unlabeled data into the first reconstructor of the student branch of the reversible phase deflection network to generate the reconstructed 3D contour.
[0044] The 3D Contour-Stripes module is used to re-render the reconstructed 3D contours into a re-rendered stripe pattern using a renderer with a student branch of a reversible phase deflection network.
[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0046] This invention requires only about 26% of the annotation data of mainstream fully supervised methods to achieve better reconstruction accuracy, and has broad application prospects in the field of online quality inspection of high reflectivity display panels. Attached Figure Description
[0047] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:
[0048] Figure 1 A schematic diagram of the measuring device provided in an embodiment of the present invention;
[0049] Figure 2 This is a diagram of the reversible phase deflection network structure provided in an embodiment of the present invention;
[0050] Figure 3 This is a schematic diagram of the reconstruction and rendering process in a U-shaped reversible network provided in an embodiment of the present invention;
[0051] Figure 4 This is a schematic diagram of a bijective function module provided in an embodiment of the present invention;
[0052] Figure 5 This is an overall structural diagram of the lightweight affine coupling module structure provided in an embodiment of the present invention;
[0053] Figure 6 This is a diagram of a depth-separable pooling convolutional block structure provided in an embodiment of the present invention;
[0054] Figure 7 A diagram of a depth-separable deconvolution block structure provided in an embodiment of the present invention;
[0055] Figure 8 This is a structural diagram of a pyramid-shaped self-similar module provided in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0057] Embodiment 1 of this invention provides a 3D contour measurement method based on a semi-supervised reversible phase deflection network, such as... Figure 2 As shown, the 3D contour measurement method based on a semi-supervised reversible phase deflection network includes the following steps:
[0058] The original deformed fringe pattern collected and modulated by the specular reflection of the screen under test is divided into labeled data and unlabeled data;
[0059] The unlabeled data is input into the second reconstructor of the teacher branch of the reversible phase deflection network to generate a predicted 3D profile. Pseudo-labels are then selected from the predicted 3D profiles and used as supervision signals for the student branch of the reversible phase deflection network.
[0060] The labeled and unlabeled data are input into the first reconstructor of the student branch of the reversible phase deflection network to generate the reconstructed 3D profile.
[0061] The reconstructed 3D contour is re-rendered as a re-rendered stripe pattern using a renderer with a student branch of a reversible phase deflection network.
[0062] like Figure 1 As shown, the original deformed fringe pattern is acquired through the coordinated efforts of an image acquisition unit, an image projection unit, a timing control unit, and auxiliary components (such as a worktable). Specifically, the image projection unit uses a programmable screen light source as its core device to project a preset sinusoidal fringe pattern onto the surface of the screen being measured (including but not limited to mobile phones, wearable devices, computer displays, and television display panels). During the acquisition process, the programmable screen light source flexibly adjusts the fringe frequency, phase, and brightness contrast according to the measurement requirements to ensure that the projected fringes have sufficient optical quality.
[0063] The image acquisition unit includes a camera and a matching light source. The matching light source provides auxiliary illumination for the camera to ensure the clarity and consistency of image acquisition. The camera can be a 500W pixel CCD high-speed camera, equipped with a 35mm focal length lens to acquire the deformed fringe pattern modulated by the specular reflection of the screen under test as the original deformed fringe pattern.
[0064] During image acquisition, there is a timing difference between the image acquisition unit and the image projection unit. Therefore, this embodiment also includes a timing control unit, which is responsible for the synchronous triggering control of the image acquisition unit and the image projection unit. This ensures that the light source projection of the image projection unit is strictly synchronized with the camera exposure of the image acquisition unit, enabling the simultaneous acquisition of a distorted image frame while projecting one frame of stripes, thus eliminating phase errors caused by timing deviations. In this embodiment, the timing control unit is implemented based on an integrated FPGA control board.
[0065] The 3D contour measurement method provided in this embodiment is implemented based on a data processing unit, which is an industrial computer equipped with a high-performance graphics processor. A reversible phase deflection network is deployed on the industrial computer, and the 3D contour measurement method is implemented based on the reversible phase deflection network.
[0066] Specifically, the data processing unit receives the deformed fringe pattern data transmitted by the image acquisition unit, uses this data as the original deformed fringe pattern, and divides it into labeled and unlabeled data. The reversible phase deflection network includes a student branch and a teacher branch. First, the unlabeled data is input into the teacher branch, and the second reconstructor in the teacher branch performs forward processing from deformed fringe pattern to 3D contour on the unlabeled data to generate the predicted 3D contour. Then, a designed pseudo-label threshold screening mechanism selects high-quality samples from the predicted 3D contour as pseudo-labels, and uses these pseudo-labels as the supervision signal for the student branch, thus achieving semi-supervised reversible phase deflection network. Compared with fully supervised reversible phase deflection network, this achieves high-precision screen 3D contour reconstruction driven by a small amount of labeled data.
[0067] After pseudo-labels are generated, a small amount of labeled data and a large amount of unlabeled data are randomly input into the student branch. The student branch includes a first reconstructor and a renderer. Under the supervision of the pseudo-labels, the student branch uses the first reconstructor to perform forward processing on the unlabeled data to transform the stripe pattern into a 3D contour, generating a reconstructed 3D contour, thus achieving forward reconstruction of stripes and 3D contours. Then, the reconstructed 3D contour is input into the renderer in the student branch for reverse processing, generating a re-rendered stripe pattern, thus achieving reverse rendering of 3D contours and stripes. The bidirectional mapping characteristics of reversible networks are used to achieve both forward reconstruction of stripes and 3D contours and reverse rendering of 3D contours and stripes.
[0068] Furthermore, before using the reversible phase deflection network for 3D contouring, it is necessary to pre-train and formally train the reversible phase deflection network. First, the student branch (first reconstructor and renderer) is pre-trained using labeled data, that is, the weights of the first reconstructor and renderer are initialized, and then these weights are shared to the second reconstructor of the teacher branch to complete the initialization of the reversible phase deflection network.
[0069] During the formal training phase, unlabeled data is first input into the second reconstructor of the teacher branch. The second reconstructor processes the unlabeled data with the aforementioned shared weights to generate predicted 3D contours. Then, a pseudo-label thresholding mechanism is used to select high-quality samples from the predicted 3D contours as supervision signals for the student branches. The process is as follows:
[0070] ;
[0071] in, This indicates the weight of the teacher branch. This indicates unlabeled data. Indicates the second reconstructor. This represents a threshold filtering operation. For the filtered pseudo-labels, This indicates a compound operation performed from back to front.
[0072] In this embodiment, the pseudo-label threshold screening mechanism includes: unlabeled data is screened using a designed threshold screening mechanism to select high-confidence prediction results as pseudo-labels for training, thereby providing effective constraints for the network in the absence of real labels. To address the potential noise and unreliability issues in pseudo-labels, this embodiment uses local variance as a confidence metric, where low variance indicates a stable and reliable region, while high variance typically indicates occlusion or noise in the region. In the actual screening process, for each pixel in the pseudo 3D contour label, the average value of its neighboring pixels is first calculated to reflect the spatial coherence of the 3D contour. Then, the confidence score, defined as the reciprocal of the variance, is compared with a preset threshold; only results higher than the threshold are retained as valid pseudo-labels.
[0073] After pseudo-labels are generated, the student branch processes the labeled and unlabeled data under the supervision of the monitoring signal. Specifically, a small amount of labeled data and a large amount of unlabeled data are randomly input into the student branch to prevent the reversible phase deflection network from overfitting to the labeled data and to make full use of the distribution information of the unlabeled data.
[0074] Then, the first reconstructor processes a small amount of labeled data and a large amount of unlabeled data to reconstruct the deformed stripe pattern into a reconstructed 3D contour. The process is as follows:
[0075] ;
[0076] in, This indicates the weight of the student branch. This represents labeled data. This indicates unlabeled data. Indicates the second reconstructor. and These represent the 3D contours corresponding to the labeled and unlabeled data reconstructed by the first reconstructor, respectively.
[0077] The reconstructed 3D contour is then input into the renderer, which uses the inverse process of a reversible phase deflection network to re-render it as a deformable stripe pattern. The process is as follows:
[0078] ;
[0079] in, Indicates the renderer. and These are re-rendered stripe patterns for the labeled data and the unlabeled data, respectively.
[0080] Furthermore, during the training process, the student's branch weights The teacher branch weights are updated iteratively in each period. The update strategy is as follows:
[0081] ;
[0082] in, It is the first The weight of the student branch in each cycle, It is the weight of the updated teacher branch.
[0083] This periodic synchronization ensures that the teacher branch gradually absorbs the latest knowledge from the student branch, forming a more stable and robust model representation.
[0084] Embodiment 2 of the present invention proposes a method for constructing a first reconstructor, a second reconstructor, and a renderer based on Embodiment 1. In this embodiment, the first reconstructor, the second reconstructor, and the renderer are all constructed based on a U-shaped reversible network.
[0085] In this method, the forward process of the U-shaped reversible network is used as the first reconstructor of the student branch and the second reconstructor of the teacher branch to achieve forward modulation from the deformed stripe pattern to the 3D contour.
[0086] By using the reverse process of the U-shaped reversible network as a renderer, the reverse modulation of 3D contours into deformable stripe patterns is achieved.
[0087] It should be noted that the core of this invention for screen 3D contour measurement lies in the forward modulation of the fringe pattern-3D contour and the inverse modulation of the 3D contour-fringe pattern. To achieve forward and inverse modulation, existing neural networks typically require two independent subnetworks to approximate the fringe pattern. and Mapping, which can accumulate errors from one mapping into another, leads to inaccurate bijective transformations, where... For the original stripe data space, For 3D contour data space.
[0088] In this invention, the reversible property of the reversible phase deflection network is utilized, and bidirectional mapping can be achieved using only a single network.
[0089] Based on the above concept, in this embodiment, the first reconstructor, the second reconstructor, and the renderer in the reversible phase deflection network are all constructed based on a U-shaped reversible network (U-INN). The forward process of the U-shaped reversible network is used as the first and second reconstructors to reconstruct the 3D contour of the screen from the deformed stripes; the reverse process of the U-shaped reversible network is used as the renderer to complete the reverse modulation of the 3D contour to the stripe pattern.
[0090] Furthermore, the U-shaped reversible network includes: several sequentially connected bijective function modules, each of which includes reversible convolution and affine coupling operation units.
[0091] It should be noted that the core idea of the U-shaped reversible network design is to use a reversible neural network to establish a bijective transformation between the stripe space and the 3D contour space. Therefore, the core of the bidirectional mapping of the U-shaped reversible network lies in finding a reversible bijective function to achieve an accurate mapping between the original stripe data space and the 3D contour data space.
[0092] like Figure 3 As shown, the U-shaped reversible network in this embodiment includes several bijective functions connected in sequence. In this embodiment, the bijective functions are implemented through affine coupling.
[0093] Traditional affine coupling operations typically employ a fixed input splitting method (such as by channel or spatial dimension), making it difficult for the model to fully blend all channel information and limiting its flexibility. To overcome these shortcomings, this embodiment introduces invertible convolution before the affine coupling operation to achieve global channel blending. This operation breaks the limitation of a fixed splitting order by permuting the input channels, thereby enhancing the model's representational capabilities.
[0094] Furthermore, positive modulation from deformed fringe patterns to 3D contours is achieved through several sequentially connected bijective function modules, including the following steps:
[0095] The deformed stripe pattern is split into a first positive dual-path feature and a second positive dual-path feature by channel dimension through reversible convolution;
[0096] The first and second positive dual-path features are input into the lightweight affine coupling module for feature extraction, resulting in the transformed first and second positive dual-path features.
[0097] The first positive dual-path feature and the second positive dual-path feature of the transformation are spliced together to obtain the 3D contour.
[0098] Specifically, the process of inverse modulation from 3D contour to deformed stripe pattern is achieved through several sequentially connected bijective functions, as follows: Figure 4 As shown, the deformed stripe pattern is used as input data. Input, through reversible Convolution (i.e.) Figure 4 InvertibleConv1*1), according to channel dimension, input data Perform channel segmentation and split it into the first positive dual-path feature. Second positive dual-path feature The process is described in two parts as follows:
[0099] ;
[0100] in, It is reversible The weight matrix of convolution, This indicates a channel splitting operation. This indicates a point-by-point multiplication operation.
[0101] Then, the split dual-path features are input into the lightweight affine coupling module. Specifically, the second positive dual-path features are input into the module. Input to the first lightweight affine coupling module In the middle, and its output data is compared with the first positive dual-path feature. By performing element-wise addition, the first positive dual-path feature of the transformation is obtained. Then, the first positive dual-path feature of the conversion is... Input to the second lightweight affine coupling module and the third lightweight affine coupling module In the second lightweight affine coupling module The first positive dual-path characteristic of the conversion is obtained through a nonlinear control function. The data is processed and then compared with the second positive dual-path feature. Perform element-wise multiplication, and finally combine the resulting data with the third lightweight affine coupling module. The output data is added element by element to obtain the second positive dual-path feature of the transformation. The process is represented as follows:
[0102] ;
[0103] in, , , For lightweight coupled modules with shared structures, It is a nonlinear control function. It is a sigmoid activation operation. It is a scale-limiting operation, which internally sets hyperparameters to limit the scale range. Limit the data size to between.
[0104] In this embodiment, the three operations—exponential transformation of the nonlinear control function, scale constraint, and Sigmoid activation—are combined. Specifically, Sigmoid activation first transforms the input features (first positive dual-path features) into... The input features are mapped to an interval and then adjusted to a range through linear transformation to achieve smooth normalization control. The scale constraint operation uses adjustable hyperparameters to constrain the output range, which avoids numerical instability while retaining sufficient expressive power. The exponential activation transforms the linear transformation into nonlinear scaling, which further enhances the nonlinear fitting ability of the model while ensuring positive definiteness of the output.
[0105] The organic combination of these three elements effectively breaks through the expression bottleneck of traditional affine coupling while meeting the mathematical rigor requirements of reversible networks, enabling U-shaped reversible networks to learn more complex feature mapping relationships under numerically stable conditions.
[0106] Finally, the first positive dual-path feature of the conversion will be... and the second positive dual-path feature of the conversion Perform feature stitching to generate a 3D contour. As the output of the bijective function, the process is represented as follows:
[0107] ;
[0108] in, It is a channel splicing operation.
[0109] It should be noted that the above process enables the second reconstructor to process unlabeled data to generate predicted 3D contours, and the first reconstructor to process labeled and unlabeled data to generate reconstructed 3D contours.
[0110] Furthermore, the inverse modulation of the 3D contour into a deformed fringe pattern is achieved through several sequentially connected bijective function modules, including the following steps:
[0111] The 3D contour is split into a first inverse dual-path feature and a second inverse dual-path feature;
[0112] The first and second inverse dual-path features are input into the lightweight affine coupling module for feature extraction, resulting in the transformed first and second inverse dual-path features.
[0113] The transformed first inverse dual-path features and the transformed second inverse dual-path features are concatenated by reversible convolution to obtain a deformed stripe pattern.
[0114] It should be noted that the reverse process of a bijective function is the inverse operation of a forward mapping, and its process is represented as follows:
[0115] ;
[0116] in, Reversible The inverse of the weight matrix of the convolution.
[0117] It should be noted that the above process enables the renderer to re-render the reconstructed 3D outline into a re-rendered stripe pattern.
[0118] Based on embodiment 2, embodiment 3 of the present invention provides a design for a lightweight affine coupling module, wherein the lightweight affine coupling module includes: a downsampling convolutional layer, four cascaded depthwise separable pooling convolutional blocks, a pyramidal double self-similarity module, four cascaded depthwise separable deconvolutional blocks, and an upsampling convolutional layer connected in sequence.
[0119] The four cascaded depthwise separable pooling convolutional blocks correspond one-to-one with the four cascaded depthwise separable deconvolutional blocks, and the output of each depthwise separable pooling convolutional block is fed to the input of the corresponding depthwise separable deconvolutional block through same-scale skip connections.
[0120] Lightweight affine coupling modules such as Figure 5 As shown, the lightweight affine coupling module adopts a U-shaped structure, consisting of four sampling blocks that integrate depthwise separable convolutions and a pyramid depthwise separable convolutional block (PDSB).
[0121] Furthermore, feature extraction is performed using a lightweight affine coupling module, including the following steps:
[0122] The input features are initially extracted using a downsampling convolutional layer;
[0123] The extracted preliminary features are downsampled using four cascaded depthwise separable pooling convolutional blocks to obtain downsampled features;
[0124] The downsampled features are enhanced using a pyramid dual self-similarity module to generate enhanced features;
[0125] The enhanced features are upsampled through four cascaded depthwise separable deconvolution blocks to generate upsampled features;
[0126] Upsampling convolutional layers are used to integrate cross-channel information from upsampled features.
[0127] Specifically, firstly, the input features (i.e., the second positive dual-path features mentioned above) are... The first positive dual-path feature of the conversion ), after a 3*3 downsampling convolutional layer (i.e. Figure 5 The Conv3*3 in the input is used for initial extraction. By expanding the number of channels to 8 times that of the input, the model's ability to capture features is enhanced.
[0128] Subsequently, the initially extracted features were processed through... Figure 6The four cascaded depthwise separable pooling convolutional blocks (DSPBs) shown double the number of channels and halve the feature map scale with each DSPB, achieving feature downsampling.
[0129] The downsampling process for each depthwise separable pooling convolutional block includes:
[0130] First, a depthwise convolution operation is performed on the initially extracted features (or the output features of the previous level depthwise separable pooling convolution block). Then, the initially extracted features (or the output features of the previous level depthwise separable pooling convolution block) are combined with the intermediate features obtained after the depthwise convolution operation. Perform instance normalization;
[0131] Secondly, regarding intermediate features Pointwise convolution is performed, followed by max pooling. Considering that stripe patterns have significant individual style features and batch-to-batch statistical differences, this embodiment uses instance normalization for sample-by-sample standardization, so that the model focuses on learning the essential features of the samples rather than batch correlation, thereby improving the network's performance in stripe to 3D contour style conversion.
[0132] Finally, the features obtained after max pooling are activated using the Leaky ReLU activation function, and used as the output features of this level of depthwise separable pooling convolutional block. To further enhance the network's nonlinear representation capability, this embodiment places the activation operation after the pooling operation. By applying a nonlinear transformation to the downsampled abstract features, computational complexity is reduced while ensuring that the activation operation focuses on the representative features selected by pooling, achieving more efficient and targeted feature transformation.
[0133] The downsampling process of depthwise separable pooling convolutional blocks can be represented as follows:
[0134] ;
[0135] in, for Features initially extracted by the convolutional layer For the first The output of depthwise separable pooling convolutional blocks, It is a depthwise convolution operation. It is a pointwise convolution operation. This indicates the instance normalization operation. This is a max pooling operation with a scale of 2. It is the Leaky ReLU activation function.
[0136] According to the above formula and Figure 6 The downsampling process of depthwise separable pooling convolutional blocks can be understood as: through depthwise convolutional layers ( Figure 6 In DWconv (1*5+5*1), the values are... The features initially extracted by the convolutional layer or the output of the previous level's depthwise separable pooling convolutional block are subjected to depthwise convolution, and then element-wise added to the data obtained from the depthwise convolution. The element-wise added data is then subjected to instance normalization (i.e., through...). Figure 6 (Instance normalization is performed in the IN layer to obtain intermediate features of instance normalization) ; through pointwise convolutional layers ( Figure 6 PWconv in the middle), max pooling layer ( Figure 6 MaxPooling and activation layers (in the context of MaxPooling) Figure 6 LRelu (in the context of instance normalization) is an intermediate feature. Perform pointwise convolution, max pooling, and activation operations sequentially to obtain the final result. Output of depthwise separable pooling convolution blocks .
[0137] A Pyramid Dual Self-Similarity Block (PDSB) was designed between the downsampling and upsampling paths, such as... Figure 8 As shown, this module integrates depthwise separable convolutions with different hole rates through a multi-branch structure, which expands the receptive field while maintaining computational efficiency and enhances the model's ability to perceive multi-scale features.
[0138] Specifically, the processing procedure of the pyramid dual self-similarity module can be represented as follows:
[0139] ;
[0140] in, This is the output of the fourth-level depthwise separable pooling convolution block. This is the output of the pyramid dual self-similarity module. Separable convolutions at different depths with varying porosity.
[0141] According to the above formula and Figure 8 The processing of the pyramid dual self-similarity module can be understood as: inputting the output of the 4th level depthwise separable pooling convolutional block into the pointwise convolutional layer ( Figure 8 In PWconv), then depthwise separable convolutions with different dilation rates are used respectively. Figure 8In this embodiment, the dilatancy ratio D is 1, 2, or 3. Depthwise separable convolution is performed on the data after each convolutional layer. The data after pointwise convolution and the data after the depthwise separable convolution are then concatenated. Finally, instance normalization is performed on the concatenated data (i.e., through...). Figure 8 Instance normalization is performed in the IN layer), and pointwise convolution operations are performed (through...). Figure 8 PWconv in the middle) and activation operations (via Figure 8 LRelu in (the text is incomplete and cannot be translated).
[0142] The upsampling process is symmetrically passed through a four-level depthwise separable deconvolution block (DSDB), such as... Figure 7 As shown, spatial resolution is gradually restored.
[0143] Specifically, the processing of a four-level depthwise separable deconvolution block can be represented as follows:
[0144] ;
[0145] in, It is a deconvolution operation with a scale of 2. In the first Intermediate output of instance normalization processing in depthwise separable deconvolution blocks. No. The output of the deconvolution block can be separated at different depths.
[0146] According to the above formula and Figure 7 The processing of a level 4 depthwise separable deconvolution block can be understood as: performing a deconvolution operation on the data output by the pyramid dual self-similarity module (through... Figure 7 Deconv (in the middle) combines the data after deconvolution with the intermediate features normalized to the corresponding instances. Feature concatenation is performed, and the concatenated data is then subjected to pointwise convolution operations (via...). Figure 7 PWconv in the middle) and activation operations (via Figure 7 LRelu (in which the intermediate output is obtained) Perform depthwise separable convolution on the intermediate output (through...) Figure 7 (DWconv); The data after depthwise separable convolution is element-wise summed with the intermediate output, and then the element-wise summed data is sequentially subjected to pointwise convolution (via... Figure 7 PWconv in the middle) and activation operations (via Figure 7 LRelu in the middle), finally obtained the first Output of depth-separable deconvolution blocks .
[0147] A 1*1 upsampling convolutional layer is used to integrate cross-channel information of the upsampling features to complete feature reconstruction.
[0148] Based on Example 1, Embodiment 4 of this invention proposes a method for constructing a loss function for a reversible phase deflection network. Specifically, the loss function of the reversible phase deflection network includes: supervised loss for labeled data and unsupervised loss for unlabeled data, and the contributions of supervised loss and unsupervised loss are balanced by weighting coefficients.
[0149] Among them, a two-way design is used to construct supervised loss and unsupervised loss.
[0150] It should be noted that the core of this invention lies in the forward reconstruction of stripes-3D contours and the reverse rendering of 3D contours-stripes. Based on this, this embodiment designs a bidirectional hybrid loss function that integrates forward reconstruction and reverse rendering. A stripe consistency cyclic constraint is established through an invertible network, and a weighted coefficient is used to balance the contributions of supervised and unsupervised terms, which effectively alleviates the problem of insufficient labeled data and improves the accuracy and reliability of 3D reconstruction.
[0151] Specifically, the total loss consists of the supervision loss based on labeled data. And unsupervised loss based on unlabeled data It consists of two parts, weighted by coefficients. Balancing their contributions, total loss It is expressed as follows:
[0152] .
[0153] Furthermore, the supervised loss includes: forward reconstruction loss and backward rendering loss; wherein, the supervised loss is constructed using a bidirectional design, including:
[0154] Obtain the true 3D contour value of the labeled data, calculate the labeled data contour loss between the true 3D contour value of the labeled data and the corresponding reconstructed 3D contour using root mean square error, and construct the forward reconstruction loss using the labeled data contour loss.
[0155] Obtain the original deformed stripe map of the labeled data. Use L1 loss and structural similarity to calculate the L1 loss of the labeled data stripe map and the similarity loss of the labeled data stripe map between the original deformed stripe map of the labeled data and the corresponding re-rendered stripe map. Combine the L1 loss of the labeled data stripe map and the similarity loss of the labeled data stripe map to construct the inverse rendering loss.
[0156] L1 loss, also known as absolute value loss or mean absolute error, measures the average of the absolute values of the element-by-element differences between the predicted and the true values. It is simple to calculate and has better robustness to outliers.
[0157] It should be noted that the supervision loss adopts a two-way design, including a forward reconstruction loss that optimizes the accuracy of the 3D contour. and reverse re-rendering loss Forward reconstruction losses and reverse re-rendering loss Through hyperparameters Balance, monitor losses It is expressed as follows:
[0158] ;
[0159] While optimizing the accuracy of 3D contours, we also optimize the fidelity of stripe re-rendering.
[0160] The goal of the forward reconstruction process is to enable the network to reconstruct the 3D contour. As close as possible to the true value .
[0161] Since reconstruction error directly affects the final measurement accuracy, this embodiment uses the root mean square error (RMSE) as the loss function. RMSE is more sensitive to larger differences and effectively penalizes predictions that deviate significantly from the true value, thus guiding the network to converge quickly to a reasonable reconstruction direction. Therefore, RMSE is used to calculate the annotation data contour loss between the true value of the 3D contour in the labeled data and the corresponding reconstructed 3D contour. The forward reconstruction loss is then constructed using this annotation data contour loss, as expressed below:
[0162] ;
[0163] The reverse rendering process requires the network to re-render the reconstructed 3D contours back to the original fringe pattern to enhance physical consistency. To maintain both fringe intensity accuracy and structural fidelity, the reverse rendering loss in this embodiment is constructed from both L1 loss and structural similarity (SSIM) loss. The reverse rendering loss is expressed as follows:
[0164] ;
[0165] in, This represents the L1 loss of the labeled data. This represents the loss due to similarity in labeled data structures.
[0166] Preserve the light and dark details of the stripes by using pixel-level constraints, thus reducing the blurring problem in the re-rendered results; By maximizing the structural similarity index between stripes, the local texture and edge structure of the re-rendered stripes are ensured to match the original stripes, avoiding situations where only brightness is matched but deformation is distorted. Specifically, L1 loss and structural similarity are used to calculate the L1 loss of the labeled data stripe image and the similarity loss of the labeled data stripe image between the original deformed stripe image and the corresponding re-rendered stripe image. The inverse rendering loss is constructed by combining the L1 loss of the labeled data stripe image and the similarity loss of the labeled data stripe image.
[0167] ;
[0168] ;
[0169] in, This represents the original deformed stripe pattern of the labeled data; This represents the re-rendered stripe pattern corresponding to the original deformed stripe pattern.
[0170] Furthermore, unsupervised loss, similar to supervised loss, also consists of 3D contour terms and fringe terms. It is expressed as follows:
[0171] ;
[0172] in, Forward reconstruction loss is the loss due to unsupervised loss. This is the inverse rendering loss of the unsupervised loss.
[0173] During forward propagation, only stripe data capable of generating high-confidence pseudo-labels are fed into the student branch for 3D reconstruction and re-rendering. The resulting 3D contours are constrained by the corresponding pseudo-labels, specifically through the unsupervised forward reconstruction loss. Achieve unsupervised forward reconstruction loss It is expressed as follows:
[0174] ;
[0175] in, This represents the reconstructed 3D contour of unlabeled data. This represents the true value of the 3D contour without labeled data.
[0176] In the corresponding reverse process, L1 loss and structural similarity loss are also applied to ensure the consistency between the rendered stripes and the original input, thereby forming a reliable self-supervised signal.
[0177] ;
[0178] in, This indicates L1 loss for unlabeled data; This represents the loss due to unlabeled data structure similarity. This represents the re-rendered stripe pattern corresponding to the original deformed stripe pattern with unlabeled data. This represents the original deformed stripe pattern with unlabeled data.
[0179] Embodiment 5 of the present invention provides a 3D contour measurement system based on a semi-supervised reversible phase deflection network, comprising:
[0180] The segmentation module is used to divide the original deformed fringe pattern, which is modulated by the specular reflection of the screen under test, into labeled data and unlabeled data.
[0181] The semi-supervised module is used to input unlabeled data into the second reconstructor of the teacher branch of the reversible phase deflection network, generate the predicted 3D profile, filter out pseudo-labels from the predicted 3D profile, and use the pseudo-labels as the supervision signal of the student branch of the reversible phase deflection network.
[0182] The stripe-3D contour module is used to input labeled and unlabeled data into the first reconstructor of the student branch of the reversible phase deflection network to generate the reconstructed 3D contour.
[0183] The 3D Contour-Stripes module is used to re-render the reconstructed 3D contours into a re-rendered stripe pattern using a renderer with a student branch of a reversible phase deflection network.
[0184] Embodiment 6 of the present invention provides experimental results of a 3D contour measurement method based on a semi-supervised reversible phase deflection network.
[0185] This embodiment uses 1523 sets of deformable stripe-3D contour paired samples for training. The sample size is 352×640 pixels, and the stripe frequency is 53 periods / image width. The dataset is divided into training, validation, and test sets in an 8:1:1 ratio. The training data contains 318 labeled pairs, and the rest are unlabeled data.
[0186] A reversible phase deflection network was constructed based on the PyTorch deep learning framework. The Adam optimizer was used to optimize the network parameters, with an initial learning rate of 0.0001. The reversible phase deflection network employed a two-stage training strategy: first, 100 epochs of supervised pre-training was performed using labeled data to initialize the network parameters and generate initial pseudo-labels; then, 200 epochs of semi-supervised training were performed using both labeled and unlabeled data. The hyperparameters were set to... , , , All experiments were conducted with a batch size of 2 and under identical hardware configurations to ensure comparability and reproducibility of results.
[0187] To verify the end-to-end reconstruction performance of the proposed reversible phase deflection network, this experiment selected a screen with multiple defects (such as pits, scratches, bubbles, and dirt) for 3D reconstruction. Analysis of the deformed fringe patterns captured by the camera and the corresponding reconstruction results shows that the proposed method can faithfully recover the overall shape, spatial structure, and key geometric features of the screen. The reconstructed surface is continuous and smooth, without obvious geometric distortion or structural breakage. This indicates that the proposed network possesses the ability to globally model complex screen surfaces and effectively capture the overall contour information of the screen. Further magnification of the results reveals that the proposed method can achieve high-quality reconstruction of small, slight, and barely perceptible 3D contour anomalies. Specifically, for bubble areas, the reconstruction results accurately reflect local bulges; for pit areas, the depth and edge transitions of the pits are clearly preserved; and for dirty areas, even with local changes in surface reflectivity, the network can still stably recover its 3D morphology. The above results fully demonstrate that this method not only achieves accurate restoration of the overall outline of the screen at the macro level, but also exhibits keen perception and high-fidelity reconstruction capability for various types and degrees of surface defects at the micro level, verifying its effectiveness and practicality in end-to-end 3D reconstruction tasks.
[0188] Example 7 provides a comparative experiment of a 3D contour measurement method based on a semi-supervised reversible phase deflection network to comprehensively evaluate the overall performance of LIRNet in the screen 3D contour reconstruction task.
[0189] The comparative experiments selected four representative end-to-end reconstruction methods as comparison models, including AEN, U-Net, hNet, and DLALNet, and conducted comprehensive qualitative and quantitative comparisons on real datasets. To further analyze the performance of each method, two representative samples with concave and convex surfaces were selected for comparison.
[0190] Traditional encoding / decoding networks, such as AEN, suffer from information loss during feature extraction due to their underlying architecture, leading to a decline in overall reconstruction quality, noticeable fluctuations in the output surface, and severe edge jaggedness. U-Net mitigates edge jaggedness to some extent with its skip connection structure, but its reconstructed surface is still affected by periodic fluctuations in the stripes. hNet and DLALNet further improve the continuity of the reconstructed surface by introducing multi-scale feature fusion strategies; however, limited by the inherent local receptive field of the CNN architecture, they struggle to effectively capture global structural information, resulting in insufficient reconstruction accuracy at the edges of pits. In contrast, the LIRNet proposed in this invention, based on a reversible neural network structure, achieves lossless feature transfer through a bidirectional information flow design, effectively alleviating the information loss problem during reconstruction. The information coupling and interaction mechanism in the bijective process significantly improves the network's feature utilization, enabling it to more accurately recover the overall contour of the screen. Error distribution maps of each method further validate the superiority of this method. The error maps of the previous methods show numerous red and yellow high-error regions, while the error map of LIRNet is dominated by large areas of dark blue low-value regions, fully demonstrating its optimal reconstruction fidelity.
[0191] In the task of reconstructing convex surfaces, AEN, due to insufficient feature extraction capabilities, produces overly smooth reconstruction results, making the convex structure almost invisible. While U-Net shows some improvement, the overall predicted values for convex regions are low, and the edge transitions are too smooth. hNet and DLALNet perform relatively well, but edge jaggedness remains, with DLALNet's reconstruction results also slightly affected by the periodic fluctuations of the stripes. In contrast, the LIRNet proposed in this invention achieves the best reconstruction results among all methods. AEN's reconstruction curve has almost no fluctuations, indicating its failure to effectively reconstruct the convex structure of the screen; U-Net and DLALNet's contour curves fluctuate significantly, while the curve distribution of this method best matches the true values. The error distribution map further verifies this conclusion. The above comparison results show that LIRNet outperforms other methods in both the overall contour and geometric structure reconstruction of objects.
[0192] Because human visual perception is subjective, Table 1 provides a quantitative analysis of network performance based on 153 test samples, focusing on reconstruction accuracy and network complexity. Reconstruction accuracy was evaluated using two key metrics: Mean Relative Error (MRE) and Structural Similarity Index Measure (SSIM). All values were rigorously calculated using statistical methods. Network complexity was assessed using two metrics: the number of model parameters and the number of floating-point operations (FLOPs), reflecting the model's computational efficiency and deployment feasibility in practical applications. Experimental data show that DLALNet slightly outperforms the SSIM metric in reconstruction accuracy, while our proposed method achieves the optimal MRE while maintaining competitiveness in other accuracy metrics. Notably, our proposed method exhibits a significant advantage in network complexity, reducing the number of parameters by 40.4% compared to the second-best method and having the lowest number of floating-point operations among all compared methods. Furthermore, our proposed method requires only 318 sets of labeled data, approximately 26.1% of the data requirements of other methods. Comprehensive analysis shows that although DLALNet slightly outperforms in the SSIM metric, this difference is negligible in terms of visual performance. Furthermore, DLALNet's high parameter and computational cost limit its deployment feasibility in practical applications. In contrast, LIRNet achieves optimal reconstruction performance while maintaining low data dependency and lightweight requirements, striking the best balance between accuracy, efficiency, and practicality. These experimental results fully validate the significant advantages of the method presented in this embodiment for practical engineering applications.
[0193] Table 1. Reconstruction accuracy and network complexity metrics for different methods
[0194]
[0195] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A 3D contour measurement method based on a semi-supervised reversible phase deflection network, characterized in that, Includes the following steps: The original deformed fringe pattern collected and modulated by the specular reflection of the screen under test is divided into labeled data and unlabeled data; The unlabeled data is input into the second reconstructor of the teacher branch of the reversible phase deflection network to generate a predicted 3D profile. Pseudo-labels are then selected from the predicted 3D profiles and used as supervision signals for the student branch of the reversible phase deflection network. The labeled and unlabeled data are input into the first reconstructor of the student branch of the reversible phase deflection network to generate the reconstructed 3D profile. The reconstructed 3D contour is re-rendered as a re-rendered stripe pattern using the renderer of the student branch of the reversible phase deflection network. The first reconstructor, the second reconstructor, and the renderer are all built based on a U-shaped reversible network; In this process, the forward process of the U-shaped reversible network is used as the first reconstructor of the student branch and the second reconstructor of the teacher branch to achieve forward modulation from the deformed stripe pattern to the 3D contour. Using the reverse process of the U-shaped reversible network as a renderer, the reverse modulation of 3D contours into deformable stripe patterns is achieved.
2. The 3D contour measurement method based on a semi-supervised reversible phase deflection network according to claim 1, characterized in that, The U-shaped reversible network includes: a plurality of bijective function modules connected in sequence, each of the bijective function modules including a reversible convolution and an affine coupling operation unit.
3. The 3D contour measurement method based on a semi-supervised reversible phase deflection network according to claim 2, characterized in that, The forward modulation from the deformed stripe pattern to the 3D contour is achieved by a series of sequentially connected bijective function modules, including the following steps: The deformed stripe pattern is split into a first positive dual-path feature and a second positive dual-path feature by channel dimension through reversible convolution; The first and second positive dual-path features are input into the lightweight affine coupling module for feature extraction to obtain the transformed first positive dual-path features and the transformed second positive dual-path features. The first positive dual-path feature and the second positive dual-path feature of the transformation are spliced together to obtain the 3D contour.
4. The 3D contour measurement method based on a semi-supervised reversible phase deflection network according to claim 3, characterized in that, The first and second forward dual-path features are input into the lightweight affine coupling module for feature extraction, including the following steps: The second positive dual-path feature is input into the first lightweight affine coupling module, and the output data of the first lightweight affine coupling module is added element by element to the first positive dual-path feature to obtain the transformed first positive dual-path feature. The first positive dual-path feature of the transformation is input into the second lightweight affine coupling module and the third lightweight affine coupling module respectively. In the second lightweight affine coupling module, the first positive dual-path feature of the transformation is subjected to element-wise multiplication and element-wise addition operations in sequence through a nonlinear control function. The data obtained after processing by the nonlinear control function is multiplied element-wise with the second positive dual-path feature, and the data obtained by element-wise multiplication is added element-wise with the output data of the third lightweight affine coupling module to obtain the transformed second positive dual-path feature.
5. The 3D contour measurement method based on a semi-supervised reversible phase deflection network according to claim 3, characterized in that, The lightweight affine coupling module includes: a downsampling convolutional layer, four cascaded depthwise separable pooling convolutional blocks, a pyramidal double self-similarity module, four cascaded depthwise separable deconvolutional blocks, and an upsampling convolutional layer connected in sequence. The four cascaded depthwise separable pooling convolutional blocks correspond one-to-one with the four cascaded depthwise separable deconvolutional blocks, and the output of each depthwise separable pooling convolutional block is fed to the input of the corresponding depthwise separable deconvolutional block through same-scale jump connections.
6. The 3D contour measurement method based on a semi-supervised reversible phase deflection network according to claim 5, characterized in that, The downsampling convolutional layer is used to perform preliminary feature extraction on the input features; The extracted preliminary features are downsampled using the four cascaded depthwise separable pooling convolutional blocks to obtain downsampled features; The downsampled features are enhanced using the pyramid dual self-similarity module to generate enhanced features; The enhanced features are upsampled through the four cascaded depthwise separable deconvolution blocks to generate upsampled features; The upsampling convolutional layer is used to integrate cross-channel information of the upsampling features.
7. The 3D contour measurement method based on a semi-supervised reversible phase deflection network according to claim 1, characterized in that, The loss function of the reversible phase deflection network includes: the supervised loss of the labeled data and the unsupervised loss of the unlabeled data, and the contributions of the supervised loss and the unsupervised loss are balanced by a weighting coefficient; Specifically, a bidirectional design is used to construct the supervised loss and the unsupervised loss.
8. The 3D contour measurement method based on a semi-supervised reversible phase deflection network according to claim 7, characterized in that, The supervised loss includes: forward reconstruction loss and backward rendering loss; wherein, the supervised loss is constructed using a bidirectional design, including: Obtain the true 3D contour value of the labeled data, calculate the labeled data contour loss between the true 3D contour value of the labeled data and the corresponding reconstructed 3D contour using root mean square error, and construct the forward reconstruction loss using the labeled data contour loss. Obtain the true value of the stripe pattern of the labeled data, and use L1 loss and structural similarity to calculate the L1 loss and similarity loss of the labeled data stripe pattern between the true value of the labeled data stripe pattern and the corresponding re-rendered stripe pattern. Combine the L1 loss and similarity loss of the labeled data stripe pattern to construct the inverse rendering loss.
9. A 3D contour measurement system based on a semi-supervised reversible phase deflection network, used to implement the 3D contour measurement method based on a semi-supervised reversible phase deflection network as described in any one of claims 1 to 8, characterized in that, include: The segmentation module is used to divide the original deformed fringe pattern, which is modulated by the specular reflection of the screen under test, into labeled data and unlabeled data. A semi-supervised module is used to input the unlabeled data into the second reconstructor of the teacher branch of the reversible phase deflection network, generate a predicted 3D profile, filter out pseudo-labels from the predicted 3D profile, and use the pseudo-labels as supervision signals for the student branch of the reversible phase deflection network. The stripe-3D contour module is used to input labeled and unlabeled data into the first reconstructor of the student branch of the reversible phase deflection network to generate the reconstructed 3D contour. The 3D Contour-Stripes module is used to re-render the reconstructed 3D contours into a re-rendered stripe pattern using a renderer with student branches of a reversible phase deflection network. The first reconstructor, the second reconstructor, and the renderer are all built based on a U-shaped reversible network; In this process, the forward process of the U-shaped reversible network is used as the first reconstructor of the student branch and the second reconstructor of the teacher branch to achieve forward modulation from the deformed stripe pattern to the 3D contour. Using the reverse process of the U-shaped reversible network as a renderer, the reverse modulation of 3D contours into deformable stripe patterns is achieved.
Citation Information
Patent Citations
Unmanned aerial vehicle reconnaissance image target detection and positioning method based on deep learning
CN122244734A