A method for in-situ three-dimensional topography sensing and process mapping of laser processing pits
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWESTERN POLYTECHNICAL UNIV
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-04
AI Technical Summary
[0005]针对现有技术的上述缺陷,本发明的目的在于提供一种激光加工凹坑原位三维形貌感知与工艺映射方法,旨在解决现有多聚焦深度重建方法对焦堆栈输入帧数约束强、弱纹理微结构重建精度低、工业部署难度大,以及激光工艺-形貌映射模型泛化能力不足、缺乏物理一致性约束、无法与形貌感知环节闭环建模融合的技术问题
本发明采用帧维解耦的网络设计,将焦堆栈帧数维度与特征提取过程完全解耦,解除了现有深度学习模型对焦堆栈输入帧数的固定约束,可适配工业现场因成像条件变化导致的帧数调整需求,大幅提升了部署灵活性。
Smart Images

Figure CN122510441A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of laser precision machining technology, machine vision 3D reconstruction technology and intelligent manufacturing data modeling technology, and in particular to a method for in-situ 3D shape perception and process mapping of laser processing pits. Background Technology
[0002] In aero-engine maintenance, blade film etching holes are often blocked by carbon deposits or slag. High-precision layered laser etching is required to effectively remove these blockages while avoiding damage to the hole wall substrate material. The material removal behavior of laser processing is highly sensitive to process parameters such as laser power and defocusing amount. Actual processing requires multiple rounds of trial and error to adjust parameters and achieve layer-by-layer removal. Establishing a mapping relationship between laser process parameters and the three-dimensional morphology of the pits enables parameter planning before processing and dynamic interlayer adjustments during processing, ensuring processing accuracy and aerodynamic performance of the ducts. This process relies on high-precision in-situ morphology perception and highly robust process-morphology mapping capabilities.
[0003] Currently, in-situ 3D topography perception of microscale pits mainly employs multi-focus 3D reconstruction technology, which can be divided into traditional methods and deep learning-based methods. Traditional multi-focus 3D reconstruction methods rely on manually designed focus evaluation functions, which are insufficiently sensitive to microstructures such as laser pits with weak texture, low contrast, but high-frequency abrupt boundary changes. The reconstruction results are prone to depth jumps or edge blurring, failing to meet the inspection requirements of precision machining. Deep learning-based multi-focus depth estimation (DFF) methods can automatically learn focus responses through convolutional neural networks, improving reconstruction accuracy. However, existing DFF models still have many shortcomings: First, the number of input focus stack frames must be strictly consistent with the training phase, making it unable to adapt to the frame number adjustment requirements caused by changes in industrial imaging conditions, resulting in extremely poor flexibility; Second, feature extraction is only performed in the spatial domain, without explicitly modeling high-frequency components of the image, resulting in insufficient focus discrimination ability for weakly textured pits; Third, the backbone network has high computational complexity and a large number of parameters, making it difficult to deploy in a lightweight manner on resource-constrained industrial platforms; Fourth, the model output is only a pixel-level depth map, lacking structured geometric semantic expression, and cannot be directly used for process mapping modeling.
[0004] In laser process-topography mapping technology, current mainstream solutions mostly rely on fitting empirical formulas to offline experimental data or using shallow data-driven models to establish a black-box mapping between process parameters and topography parameters. These methods have fundamental limitations: First, the mapping process is completely decoupled from 3D reconstruction, making it impossible to use topography prediction errors to guide reconstruction optimization, easily leading to the accumulation of overall system errors; second, the model lacks physical constraints, and its generalization ability drops sharply when training data coverage is insufficient or operating conditions change, even resulting in predictions that violate physical laws; third, it fails to construct physically interpretable parametric representations for laser pits, making it impossible to achieve closed-loop modeling fusion of topography perception and process decision-making. In summary, existing technologies lack a fully integrated solution that organically combines multi-focus imaging, 3D reconstruction, parametric representation, standardized 3D model output, and process mapping. They cannot directly generate closed 3D mesh files that can seamlessly integrate with CAD / CAM machining systems, and the reconstruction model has poor adaptability to weakly textured microstructures, while the process mapping model is prone to predictions that violate physical laws, thus hindering the intelligentization and efficiency improvement of laser precision machining. Summary of the Invention
[0005] To address the aforementioned shortcomings of existing technologies, the present invention aims to provide a method for in-situ 3D topographic perception and process mapping of laser-processed pits. This method addresses the technical problems of existing multi-focus depth reconstruction methods, such as strong constraints on the number of input frames for the focus stack, low accuracy in reconstructing weak texture microstructures, and difficulty in industrial deployment. It also addresses the limitations of existing laser process-topographic mapping models, such as insufficient generalization ability, lack of physical consistency constraints, and inability to integrate with closed-loop modeling in the topographic perception stage.
[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: a method for in-situ three-dimensional topography perception and process mapping of laser-processed pits, comprising the following steps: S1: During the training phase, full-focus images of the pits on the surface of the material after laser processing and corresponding real depth maps are acquired. Based on the optical imaging rules, a multi-focus stack image sequence is synthesized and preprocessed to output a standardized multi-focus stack. During the inference phase, an industrial camera is moved along the optical axis to acquire the actual multi-focus stack image sequence of the pits on the surface of the material after laser processing. After grayscale conversion, noise reduction and registration preprocessing, a standardized multi-focus stack is output. S2: The preprocessed multi-focus stack is input into a lightweight multi-focus depth estimation network. The input focus stack is processed through a frame-dimensional decoupling mechanism. Combined with frequency-domain perceptual feature enhancement deeply coupled with the frame-dimensional decoupling mechanism, and global-local dual-branch 3D inter-layer interaction, a normalized depth map with the same spatial size as the input image is output. The frequency-domain perceptual feature enhancement is achieved through a wavelet frequency-domain enhanced feature extraction module, which can specifically enhance the feature response of multi-directional high-frequency abrupt change information at the boundary of the laser pit. The global-local dual-branch 3D inter-layer interaction is achieved through an adaptive gated fusion unit to realize dynamic weighted feature fusion, which can adaptively adapt to the depth reconstruction requirements of different regions of the pit. S3: Apply edge-preserving guided filtering to the normalized depth map, and combine it with the imaging system calibration parameters to map the relative depth to an absolute depth with real physical scale, generating a 3D point cloud of the laser pit. S4: Based on the anisotropic 2D Gaussian height field model, nonlinear least squares fitting with boundary constraints is performed on the 3D point cloud to extract the structured geometric parameters of the pits. S5: Construct a neural network mapping model that integrates material embedding encoding and physical regularization constraints. Using laser processing parameters and material type as inputs, and pit morphology parameters from the structured geometric parameters extracted in step S4 as outputs, the model is trained using a composite loss function. The pit morphology parameters include the depth amplitude parameter A and the lateral scale parameter σ. x and σ y The composite loss function includes a data fitting term and a physical regularization term customized for the laser ablation scenario, which can constrain the model prediction results to conform to the basic physical laws of laser-material interaction and eliminate anti-physical prediction results.
[0007] Furthermore, the training phase of step S1 specifically involves: acquiring a full-focus color image and a micrometer-level true depth map using a confocal microscope; normalizing the true depth map and dividing it into N discrete depth layers; applying Gaussian blur with different standard deviations to the full-focus image to generate a set of blurred images; and then, based on the correspondence between the depth layers and the focal plane, constructing a clear synthetic image of the corresponding depth layer region for each target focal plane, ultimately generating a multi-focus stack containing N images.
[0008] Furthermore, the frame-dimensional decoupling mechanism in step S2 is specifically as follows: the input multi-focus stack tensor [B,N,H,W], where B is the batch size, N is the number of focus stack frames, and H and W are the image space dimensions, is reshaped into [B×N,1,H,W], so that the number of focus frames N is fully integrated into the batch dimension to complete single-frame feature extraction; after single-frame feature extraction is completed, the feature tensor is reshaped into [B,N,C,H,W], where C is the number of feature channels, restoring the independent representation of the frame dimension.
[0009] Furthermore, the frequency domain sensing feature enhancement in step S2 is achieved through the wavelet frequency domain enhancement feature extraction module. Specifically, the single-frame image features are divided into identity branches, spatial branches, and frequency domain branches according to the channel ratio. The spatial branch uses anisotropic depth separable convolution to extract features, and the frequency domain branch extracts high-frequency components of the image through differentiable Haar wavelet decomposition. Finally, the three feature branches are spliced and fused to output a single-frame feature map.
[0010] Further, the global-local dual-branch 3D inter-layer interaction in step S2 is specifically as follows: the continuous variation features of adjacent focal layers are extracted through the lightweight 3D convolution of the local branch, the long-range dependency features across focal layers are extracted through the lightweight self-attention mechanism of the global branch, and the dual-branch features are dynamically weighted and fused through the adaptive gated fusion unit to output the focus response volume; the focus response volume is subjected to Softmax and weighted summation along the focal plane dimension to generate a normalized depth map normalized to the range of [0,1].
[0011] Further, step S3 specifically involves: performing guided filtering on the normalized depth map using the depth map itself as the guiding image; and linearly transforming the filtered relative depth into absolute depth based on the physical distance between adjacent focal points and the horizontal physical size of a single pixel calibrated by the imaging system, thereby generating a three-dimensional point cloud.
[0012] Furthermore, in step S4, the mathematical expression for the anisotropic 2D Gaussian height field model is: In the formula, (x0, y0) represents the center position of the pit, z0 represents the local base surface reference height, and Q is the covariance matrix; x0, y0, z0, A, and σ are obtained through nonlinear least squares optimization fitting. x σ y A total of 7-dimensional structured geometric parameters, ρ.
[0013] Further, in step S5, the material embedding encoding specifically involves: mapping discrete material types into low-dimensional dense vectors through a learnable embedding module; concatenating the material embedding vectors with the standardized laser processing parameter vectors; inputting the concave morphology parameters into the backbone of a shared neural network; and outputting the concave morphology parameters. The laser processing parameters include laser power P, pulse frequency f, number of pulses Np, and defocusing amount z.
[0014] Furthermore, the composite loss function in step S5 includes a data fitting term L. mse With physical regularization term L phys The expression is: L=L mse +λ·L phys Wherein, λ is the regularization weight coefficient; the data fitting term is the mean square error between the predicted pit morphology parameters and the experimentally measured pit morphology parameters; the physical regularization term is a ReLU-type penalty term based on the laser-material interaction law.
[0015] Furthermore, after step S4 and before step S5, a closed 3D mesh generation step is included to conclude the shape perception stage: Based on the 3D point cloud, an open triangular mesh surface is generated using Poisson surface reconstruction; the boundary loops of the pits are identified and the bottom vertices are constructed; a closed triangular mesh entity that meets the requirements of manifold and watertightness is generated and exported as an STL, OBJ, or STEP format file. This step is the concluding step of the shape perception stage; after completion, the laser process-shape mapping modeling stage in step S5 begins.
[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention employs a frame-dimensional decoupled network design, completely decoupling the focus stack frame number dimension from the feature extraction process. This removes the fixed constraint of existing deep learning models on the number of input frames for the focus stack, making it adaptable to the frame number adjustment needs caused by changes in imaging conditions in industrial settings and significantly improving deployment flexibility.
[0017] This invention enhances the ability to focus on and discriminate weakly textured microstructures by explicitly modeling the features of high-frequency abrupt change regions such as the edges of laser pits through a wavelet frequency domain enhancement feature extraction module deeply coupled with a frame-dimensional decoupling mechanism. Combined with anisotropic depth separable convolution, it significantly improves the ability to focus on and discriminate weakly textured microstructures. At the same time, it adopts a global-local dual-branch 3D inter-layer interaction mechanism and achieves dynamic weighted feature fusion through an adaptive gated fusion unit. This takes into account both the local continuous changes of adjacent focal layers and the global dependencies across focal layers. While greatly reducing the number of network parameters and computational complexity, it ensures the accuracy and continuity of deep reconstruction and can be deployed in a lightweight manner on embedded industrial platforms.
[0018] This invention, based on an anisotropic 2D Gaussian height field model, compresses high-dimensional 3D point clouds into 7-dimensional structured geometric parameters with clear physical semantics, realizing the transformation from raw perceptual data to structured knowledge; where x0, y0, and z0 are used to describe the spatial position and local reference height of the pit in the detection coordinate system, and A and σ are used to describe the spatial position and local reference height of the pit in the detection coordinate system. x σ y ρ is used to describe the depth, lateral width, and anisotropic morphological features of the pit, and serves as the main output target for process mapping modeling, thereby providing an interpretable geometric interface for process decision-making and realizing the organic integration of morphological perception and process decision-making.
[0019] The process mapping model constructed in this invention achieves unified modeling of multi-material ablation responses through learnable material embedding encoding, avoiding the high cost of modeling each material separately. At the same time, a physical regularization term based on the laser-material interaction law is introduced into the loss function to constrain the model's prediction results to conform to basic physical laws, effectively suppressing the anti-physics prediction problem that is prone to occur in pure data-driven models, and significantly improving the model's generalization ability and robustness under small sample and cross-material conditions.
[0020] This invention constructs an integrated technical solution for the entire process, from multi-focus image acquisition, three-dimensional morphology reconstruction, structured parameter extraction, standardized three-dimensional model output to process mapping modeling. It can realize in-situ high-precision morphology perception and process modeling of laser-processed pits, and can support closed-loop control and process optimization of laser precision machining. It is especially suitable for application scenarios with extremely high requirements for processing accuracy and reliability, such as laser declogging of air film holes in aero-engine blades.
[0021] This invention, through a closed 3D mesh generation step, can directly output standardized 3D model files that meet the requirements of manifold and watertightness. It can seamlessly connect with mainstream CAD / CAM systems such as UG and SolidWorks without additional model post-processing, greatly improving the applicability and ease of use of this invention in industrial settings. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the overall technical process of the laser processing pit in-situ three-dimensional shape perception and process mapping method described in this invention. Figure 2 This is a schematic diagram showing the blockage state of the air film pores on an aero-engine blade. Figure 2 (a) is a schematic diagram of an aero-engine blade. Figure 2 (b) is a schematic diagram of the blockage state of the air film pores; Figure 3 This is a schematic diagram showing the relationship between focal length and the height distribution on the object's surface. Figure 4 This is a diagram illustrating the overall architecture of the lightweight multi-focus depth estimation network described in this invention. Figure 5 This is a schematic diagram of the wavelet frequency domain enhanced feature extraction module described in this invention; Figure 6 This is a schematic diagram of the structure of the global-local dual-branch 3D inter-layer interaction module described in this invention; Figure 7 This is a comparative diagram of laser-induced pit depth maps, in which... Figure 7 (a) is the original depth map. Figure 7 (b) is the depth map after guided filtering; Figure 8 A schematic diagram of the three-dimensional point cloud of the laser-induced indentation; Figure 9 A schematic diagram of the parametric modeling results for laser-induced pits; Figure 10 This is a schematic diagram of a 3D mesh of a laser-induced indentation, where... Figure 10 (a) is a schematic diagram of the original 3D point cloud of the laser-induced indentation. Figure 10 (b) is a schematic diagram of the open triangular mesh surface obtained based on point cloud reconstruction. Figure 10 (c) is a schematic diagram of a closed triangular mesh entity that meets the requirements of manifold and watertightness after boundary sealing treatment; Figure 11 This is a framework diagram of the process mapping model that integrates material embedding and physical regularization as described in this invention. Detailed Implementation
[0023] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment uses the in-situ morphology perception and process mapping of pits during the femtosecond laser unclogging process of film cooling holes in DD6 nickel-based superalloy blades of aero-engines as an application scenario to provide a complete and clear description of the implementation methods of the present invention. The described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] The method for in-situ three-dimensional topography perception and process mapping of laser-processed pits described in this embodiment is illustrated in the attached diagram of the specification. Figure 1 As shown, the specific implementation steps are as follows: Step S1: Acquisition and Preprocessing of Multifocus Image Sequences The core of this step is to generate a multi-focus stack that conforms to the laws of optical imaging, providing standardized input for subsequent depth reconstruction; this embodiment addresses the blockage state of aero-engine blades and film cooling holes, such as... Figure 2 As shown, the specific implementation process is as follows: Raw data acquisition: The pits on the surface of the DD6 nickel-based superalloy sample after femtosecond laser processing were observed using a Keyence VK-X series laser confocal microscope. Two types of raw data were acquired simultaneously: one full-focus color image. All depth regions are clearly imaged; a corresponding true depth map is also included. The depth value is in micrometers and fully reflects the three-dimensional morphology of the pit. In this embodiment, the image resolution is H=W=512 pixels.
[0025] Depth Layering and Blurred Image Generation: The Relationship Between Focal Length and Object Surface Height Distribution in Multifocal Imaging Figure 3 As shown; the true depth map D true(x,y) is linearly normalized to the grayscale range of [0,255], and divided into N discrete depth layers with equal depth intervals. In this embodiment, N is 20, but it can also be adjusted within the range of 20~36 according to imaging requirements; for the fully focused image I AiF Apply Gaussian blurs of different scales to (x, y) to generate a set of blurred images corresponding one-to-one with the depth layer. The standard deviation of the blur kernel increases with the depth layer index. The specific formula is as follows: In the formula, The standard deviation is σ k The Gaussian kernel is used, and in practice, discrete Gaussian kernels with odd-numbered sizes such as 3×3, 5×5, etc., are used for approximation. * represents a two-dimensional convolution operation. During the synthesis process, Gaussian blur with different standard deviations is used to approximate the defocusing degradation effect of the optical system. This assumption can effectively approximate the real physical degradation process within a shallow depth of field in confocal or depth-of-field synthesis microscopy systems. For more complex imaging systems, a measured PSF can be used instead of the Gaussian kernel for synthesis to further improve the model's generalization accuracy on specific devices.
[0026] Multi-focal stack synthesis: Based on the one-to-one correspondence between depth layers and focal planes, a synthesized image is constructed for the i-th target focal plane. This ensures that only the regions corresponding to the depth layer remain sharp in the synthesized image, while the remaining regions gradually blur with increasing distance from the focal plane, conforming to the defocusing degradation law of actual optical imaging. For each pixel position (x, y), let its corresponding depth layer be l(x, y), then the formula for calculating the pixel value of the synthesized image is: .
[0027] Output standardized data: The final output generates a multi-focus stack containing N images, converting all images to single-channel grayscale format for use as input to a subsequent lightweight multi-focus depth estimation network; simultaneously, it retains the true depth map D. true The original fully focused image I serves as a supervisory label for network training. AiF This can be used as optional auxiliary information for visualization or edge-guided filtering. The above synthesis method is used in the network training stage. In the inference stage of in-situ detection in industrial sites, it is not necessary to obtain a real depth map. Instead, an industrial camera is used to move along the optical axis at equal intervals to capture multi-focal distance images of the pits on the surface of the material after laser processing, acquiring a sequence of actual multi-focal stack images. This sequence is then preprocessed sequentially with grayscale conversion, Gaussian denoising, and inter-frame registration based on feature matching to output a standardized multi-focal stack, which is directly input into the trained lightweight multi-focal depth estimation network for depth map reconstruction.
[0028] Step S2: Multi-focus Depth Map Reconstruction This step achieves end-to-end depth reconstruction using a lightweight multi-focus depth estimation network. The overall network architecture is shown in the attached diagram in the manual. Figure 4 As shown, the core includes a frame-dimensional decoupling mechanism, a wavelet frequency domain enhanced feature extraction module, a global-local dual-branch 3D inter-layer interaction module, and a depth map generation module. The specific implementation process is as follows: Input format definition: The network input is the multi-focus stack output in step S1. The input tensor format is (B,N,H,W), where B is the batch size, N is the number of focus stack frames, and H and W are the spatial dimensions of a single frame image. In this embodiment, B=8 and H=W=512.
[0029] Frame-dimensional decoupling: The fixed constraint on the number of frames in the network focus stack is removed through the frame-dimensional decoupling mechanism. The specific operation is as follows: First, frame-batch fusion is performed to reshape the input tensor (B,N,H,W) into (B×N,1,H,W), so that the number of focal frames N is fully integrated into the batch dimension. All subsequent single-frame feature extraction operations are performed based on a single grayscale image. Core operations such as convolution, pooling, and normalization are only related to the spatial size and number of channels of a single frame, and have no relation to the number of focal frames N, thus completely removing the constraint of the number of frames on feature extraction. After single-frame feature extraction is completed, a frame dimension restoration operation is performed to reshape the extracted single-frame feature tensor (B×N,8,H,W) into (B,N,8,H,W), restoring the independent representation of frame dimension N. This reshaping operation is an adaptive dimension transformation, and N can be flexibly adjusted within a set focal length sampling range according to actual imaging requirements. During the inference stage, only the actual physical position of the corresponding focal plane or the distance between adjacent focal points needs to be input to complete the depth scale mapping without retraining the network.
[0030] Wavelet frequency domain enhanced feature extraction: The wavelet frequency domain enhanced feature extraction module strengthens the feature response to high-frequency structures with weak textures, such as pit edges. The module structure is shown in the attached figure in the manual. Figure 5 As shown, the specific implementation is as follows: The reshaped single-frame feature input module is subjected to 1×1 convolution to achieve adaptive channel adjustment, and then divided into identity branch, spatial branch and frequency domain branch according to the channel ratio, with a total output channel count C. out Split into C res +4×g c =C out C res g is the number of identical branch channels. c This refers to the number of channels in a single branch. The identity branch directly preserves the input features to avoid gradient vanishing; the spatial branch is split into three parallel sub-branches, which respectively use anisotropic depthwise separable convolutions of 3×3 square, 1×11 horizontal and 11×1 vertical. First, depthwise convolution is performed to extract spatial features, and then pointwise convolution is performed to fuse channel information, capturing global texture and edge features in the horizontal and vertical directions respectively. The frequency domain branch performs a differentiable Haar wavelet one-dimensional decomposition on the input features to obtain low-frequency component LL, horizontal high-frequency component LH, vertical high-frequency component HL, and diagonal high-frequency component HH. The three types of high-frequency components are retained and then fused by 1×1 convolution. The result is upsampled to the original input resolution to achieve explicit enhancement of high-frequency information. Finally, the output features of the identity branch, the spatial branch, and the frequency domain branch are concatenated. GroupNorm normalization and GELU activation function are used throughout to ensure training stability. Multi-scale features are extracted through five rounds of convolution-pooling downsampling. Features from corresponding layers of the encoder and decoder are fused via residual skip connections, ultimately outputting a single-frame feature map with enhanced high-frequency information and unified dimensions. This module is deeply coupled with the frame-dimensional decoupling mechanism and is specifically designed for processing the features of the decoupled single-frame image. Unlike the conventional approach of performing frequency domain transformation on the entire feature map, this invention's frequency domain branch performs anisotropic deep separable convolution on the decoupled spatial branch, combined with the horizontal high-frequency (LH), vertical high-frequency (HL), and diagonal high-frequency (HH) components explicitly extracted by differentiable Haar wavelet decomposition. This achieves targeted enhancement of multi-directional high-frequency abrupt changes at the laser pit boundary, while the identity branch ensures stable model response to smooth background regions. The specific ratio and concatenation method of the three feature branches together constitute a multi-view feature extraction mechanism specifically designed to enhance the focusing and discrimination capabilities of weakly textured microstructures.
[0031] Global-local dual-branch 3D inter-layer interaction: The inter-layer interaction module models the focus response relationship between focus stack frames. The module structure is shown in the attached diagram in the manual. Figure 6 As shown, the specific implementation is as follows: Local branching: Two layers of 3×3×3 kernel-sized 3D convolutions are used with padding set to 1 and no pooling operation. The number of feature channels remains unchanged at 8 dimensions. The continuous variation pattern between adjacent focal layers is extracted along the focal plane dimension to characterize the smooth evolution process of the focal layer response in the local neighborhood and enhance the local continuity of depth prediction. Global Branch: A lightweight self-attention mechanism is constructed along the focal plane dimension at each spatial location. First, the feature tensor is reshaped into (B×H×W,N,8). Query features Q, key features K, and value features V are generated along the focal plane dimension (N) through 1×1×1 3D convolutions. Then, the global association weights of the focal layer are calculated using the self-attention formula. In the formula, d is the feature dimension, and in this embodiment, d=8. Finally, the attention weight is multiplied by the value feature V to reshape it back to the original tensor dimension (B,8,N,H,W), capturing the long-range dependency of the focal layer across the long sequence range, and making up for the missing global optimal focal layer discrimination ability when relying only on local convolution. Adaptive Gated Fusion: The adaptive gated fusion unit dynamically weights and fuses the output features of the local and global branches. First, a 1×1×1 3D convolution maps the output features of both branches to weighted features of (B,1,N,H,W). After Sigmoid activation, the fusion weights α∈[0,1] are obtained. Then, feature fusion is completed using the following formula: S=α·S local +(1-α)·S global In the formula, S local For local branch output features, S global The global branch output features are used. The key to this adaptive gating fusion mechanism is that the fusion weight α is not a pre-set fixed hyperparameter, but rather dynamically predicted and generated by the network based on the input bi-branch feature map through a 1×1×1 3D convolutional layer and a sigmoid activation function. This allows the network to automatically increase the local branch weights to ensure depth continuity when facing the center region of a pit with smooth depth; and to automatically increase the global branch weights when facing the edge region of a pit with abrupt depth changes, referencing a wider range of focusing information to make accurate judgments. This significantly differs from the fixed-ratio feature addition or concatenation operations in existing technologies. After fusion, the channel dimension is compressed to 1 dimension through a 1×1×1 3D convolution, outputting a focusing response volume with a shape of (B, N, H, W).
[0032] Depth map generation: Perform a Softmax operation on the focus response volume along the focal plane dimension to generate the probability distribution of the optimal focus position for each pixel. The indices of each focus position are weighted and summed with the probability as the weight to obtain the relative depth value of each pixel. Finally, the values are normalized to the range of [0,1] and the output is a normalized depth map with the same spatial size as the input image.
[0033] Step S3: Depth Map Post-processing and 3D Point Cloud Generation This step completes the artifact removal and physical scale mapping of the depth map, generating a 3D point cloud with realistic geometric dimensions. The specific implementation process is as follows: Guided filtering for artifact removal: Using the depth map itself as the guide image, guided filtering is performed on the normalized depth map output in step S2. The filter window radius is r=8, and the regularization parameter is ε=0.01. While suppressing noise and stair-step artifacts, it effectively preserves important geometric structures such as pit boundaries. A comparison of the depth maps before and after filtering is shown in the attached figure in the specification. Figure 7 As shown.
[0034] Physical Scale Mapping and Point Cloud Generation: Based on the calibration parameters of the imaging system, the normalized relative depth values are mapped to absolute physical depth. In this embodiment, the physical distance between adjacent focal points Δz = 10 μm calibrated by the imaging system, and the lateral physical size of a single pixel Δx = Δy = 0.5 μm / pixel; each pixel position (i,j) is converted into a three-dimensional spatial point, calculated using the formula: Pij =(i·Δx,j·Δy,z ij ) In the formula, z ij The actual depth value is obtained by inverse linear transformation of the filtered depth value, in micrometers; the final generated 3D point cloud is shown in the attached figure in the instruction manual. Figure 8 As shown, it is geometrically continuous, artifact-free, and possesses accurate physical scale, providing reliable input for subsequent parametric modeling.
[0035] Step S4: Structured and parametric modeling of pit morphology This step compresses the high-dimensional 3D point cloud into low-dimensional structured parameters with clear physical semantics. The specific implementation process is as follows: Point cloud preprocessing: The 3D point cloud generated in step S3 is subjected to region segmentation, statistical filtering for noise reduction and local plane fitting to remove outlier noise points and extract sub-point clouds containing only the pit region to ensure that the pit is located on an approximately horizontal local substrate surface.
[0036] Model Fitting and Parameter Extraction: Based on an anisotropic 2D Gaussian height field model, nonlinear least squares fitting with boundary constraints is performed on the pit point cloud. The mathematical expression of the model is as follows: In the formula, (x0, y0) is the center position of the pit, z0 is the local base surface reference height, A is the depth amplitude parameter and A < 0, and Q is the depth amplitude parameter σ. x σ y The covariance matrix formed by the correlation coefficient ρ; this embodiment uses the Levenberg-Marquardt (LM) algorithm for nonlinear optimization, with boundary constraints set: A∈[-200,0]μm, σ x σ y Given ∈[10,200]μm, ρ∈[-0.95,0.95], x0, y0, z0, A, and σ are automatically fitted. x σ y The 7-dimensional structured geometric parameters (ρ, ρ, etc.) and fitting results are shown in the attached figure in the instruction manual. Figure 9 As shown.
[0037] Parameter semantic definition: The extracted 7-dimensional parameters have clear engineering semantics, where (x0, y0) is used to locate the machining center deviation, z0 is used to characterize the local substrate surface reference height, |A| is the maximum ablation depth of the pit, directly reflecting the amount of material removed, and σ x σ yThe lateral widening scale of the pit is characterized and can be converted into an equivalent radius or processed area. ρ is used to describe the orientation correlation and anisotropy characteristics of the pit's elliptical morphology. This 7-dimensional parameter vector constitutes a compact, interpretable low-dimensional representation of the pit morphology, where A, σ... x σ y ρ has a more direct relationship with laser process parameters and material response, therefore A and σ are selected. x σ y ρ serves as the output target of the subsequent process-morphology mapping model.
[0038] Closed 3D Mesh Generation: This step is performed after the core parameter extraction in step S4 and before the start of step S5. It is the final step in the shape perception process. Based on the original pit point cloud generated in step S3, a high-fidelity open triangular mesh surface is generated using the Poisson surface reconstruction algorithm, with the reconstruction depth parameter set to 8. The α-shape algorithm is used to extract the projection boundary ring of the pit in the XY plane, with the α value set to 3 times the pixel size. The bottom vertices are constructed along the normal of the base plane, and the side triangular patches are generated by connecting the boundary ring and the bottom vertices to form a closed triangular mesh entity that meets the requirements of manifold and watertightness, as shown in the attached figure in the specification. Figure 10 As shown; finally, the closed mesh is exported as a standard format such as STL, OBJ, STEP, etc., which can be directly imported into CAD / CAM systems such as UG and SolidWorks for Boolean operations and machining path planning.
[0039] Step S5 Laser Process - Topography Mapping Modeling This step constructs a neural network mapping model that integrates material embedding encoding and physical regularization constraints. The model framework is shown in the attached diagram in the manual. Figure 11 As shown, the specific implementation process is as follows: Input data preprocessing: The model input contains two types of information, one of which is laser processing parameters, including laser power P, pulse frequency f, and number of pulses N. p The defocusing amount z, in this embodiment, the process parameter range is: P∈[10,50]W, f∈[100kHz,1MHz], N p ∈[1,100], z∈[-50,+50]μm, linearly normalize all process parameters to the range of [0,1]; another category is material type, including DD6 nickel-based high-temperature alloy, Ti-6Al-4V titanium alloy, 304 stainless steel and other commonly used metal materials for laser processing.
[0040] Material embedding encoding: Mapping discrete material types into 8-dimensional low-dimensional dense vectors through learnable embedding layers. This transforms discrete material information into a continuous, differentiable representation, enabling the model to learn the differences and similarities between ablation responses of different materials through the embedding space, thereby improving cross-material adaptability. The material embedding vector is concatenated with the standardized process parameter vector and input together into a shared neural network backbone.
[0041] Neural Network Backbone Design: This embodiment uses a 3-layer residual MLP as the network backbone. Each layer contains a linear layer, LayerNorm normalization, and ReLU activation function. The hidden layer dimensions are 64, 32, and 16, respectively. The final output is 4-dimensional pit morphology parameters, which are compared with A and σ extracted in step S4. x σ y ρ corresponds to .
[0042] Composite Loss Function and Model Training: A composite loss function, including a data fitting term and a physical regularization term, is used for model optimization. The expression for the loss function is: L = L mse +λ·L phys In the formula, λ is the regularization weight coefficient, and in this embodiment, λ=0.1; Data fitting term L mse To predict the mean square error between the pit morphology parameters and the experimental measurements, the specific expression is as follows: In the formula, , , , Let A and σ be the model's predicted values. x σ y ρ represents the true value measured in the experiment; Physical regularization term L phys This is a ReLU-type penalty term based on the interaction between laser and metal materials. Local partial derivatives are calculated at each training sample using automatic differentiation to apply a soft penalty to predictions that violate physical laws. The specific expression is as follows: In the formula, A is the pit depth amplitude, P is the laser power, and N is the laser power. p Let |z| be the number of pulses, |z| be the absolute value of the defocus amount, and R be the value of the pulse count. eq For σ x The equivalent radius of the pit derived from σᵧ; Let be the partial derivative of depth amplitude with respect to laser power. The partial derivative of depth amplitude with respect to the number of pulses. The partial derivative of depth amplitude with respect to the absolute value of defocus. It is the partial derivative of the equivalent radius with respect to the absolute value of the defocusing amount.
[0043] This physical regularization term, calculated using a ReLU-type penalty term and the squared L2 norm, strongly constrains the core monotonicity of the laser ablation process: the first term constrains the fundamental principle that higher laser power and greater single-pulse energy lead to deeper ablation depth; the second term constrains the fundamental principle that more pulses and greater cumulative energy lead to deeper ablation depth; the third term constrains the fundamental principle that increased absolute value of defocusing and the focal point moving further away from the material surface lead to decreased energy density and shallower ablation depth; and the fourth term constrains the fundamental principle that increased absolute value of defocusing and laser spot divergence lead to increased equivalent radius of the ablation pit. After normalizing the input process parameters and output morphology parameters, the physical regularization term adopts the squared L2 norm form, possessing a similar numerical scale and stable gradient characteristics to the data fitting term. By minimizing the aforementioned penalty term for local partial derivatives during training, this model can fundamentally prevent anti-physics predictions such as "increased power leading to decreased ablation depth," which violate the physical principles of laser-material interactions, significantly different from general physical regularization schemes in existing technologies that lack clear scenario constraints.
[0044] This constraint ensures that the model's local response near any input point conforms to common sense in engineering, avoiding anti-physical predictions.
[0045] Model Training and Inference: This embodiment constructs 81 datasets through orthogonal experimental design, covering different combinations of materials and process parameters, and divides them into training and test sets in an 8:2 ratio; the Adam optimizer is used, with an initial learning rate set to 1×10⁻⁶. -4 The batch size is 32, the training rounds are 2000, and an early stopping strategy with a pause threshold of 50 is set to avoid overfitting. After training, during the inference phase, only the target material type and laser process parameters need to be input, and the model can output the corresponding pit structured geometric parameters to achieve the mapping between process and morphology.
[0046] Technical effect verification of this embodiment Compared with existing technologies, the method described in this embodiment has significant technical advantages: In the 3D topography reconstruction stage, compared with the mainstream DDFFNet model, the number of network parameters in this invention is only 1 / 585, the inference speed of a single image sequence is increased by 22.62 times, the root mean square error (RMSE) of depth reconstruction is reduced by 76.63%, the reconstruction accuracy of weak texture pit edges is significantly improved, and it supports focal stack input of any number of frames without retraining; In the process mapping stage, compared with the traditional Gaussian support vector regression (G-SVR) model, the model of this invention has a higher prediction R-value within the training data coverage range. 2 It can reach above 0.98, and in cross-material adaptation scenarios, the prediction error is reduced by 60%, with no inverse physics prediction results, and the generalization ability and robustness are significantly improved.
[0047] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and these variations still fall within the protection scope of the present invention.
Claims
1. A method for in-situ three-dimensional topographic perception and process mapping of laser-processed pits, characterized in that, Includes the following steps: S1: During the training phase, acquire full-focus images of the pits on the surface of the material after laser processing and the corresponding real depth maps. Based on the optical imaging rules, synthesize a multi-focus stack image sequence and complete preprocessing to output a standardized multi-focus stack. During the inference phase, an industrial camera is moved along the optical axis to acquire the actual multi-focus stack image sequence of the pits on the surface of the material after laser processing. After grayscale conversion, noise reduction and registration preprocessing, a standardized multi-focus stack is output. S2: The preprocessed multi-focus stack is input into a lightweight multi-focus depth estimation network. The input focus stack is processed through a frame-dimensional decoupling mechanism. Combined with frequency-domain perceptual feature enhancement deeply coupled with the frame-dimensional decoupling mechanism, and global-local dual-branch 3D inter-layer interaction, a normalized depth map with the same spatial size as the input image is output. The frequency-domain perceptual feature enhancement is achieved through a wavelet frequency-domain enhanced feature extraction module, which can specifically enhance the feature response of multi-directional high-frequency abrupt information at the boundary of the laser pit. The global-local dual-branch 3D inter-layer interaction is achieved through an adaptive gated fusion unit to realize dynamic weighted feature fusion, which adaptively adapts to the depth reconstruction requirements of different regions of the pit. S3: Apply edge-preserving guided filtering to the normalized depth map, and combine it with the imaging system calibration parameters to map the relative depth to an absolute depth with real physical scale, generating a 3D point cloud of the laser pit. S4: Based on the anisotropic 2D Gaussian height field model, nonlinear least squares fitting with boundary constraints is performed on the 3D point cloud to extract the structured geometric parameters of the pits. S5: Construct a neural network mapping model that integrates material embedding encoding and physical regularization constraints. Using laser processing parameters and material type as inputs, and pit morphology parameters from the structured geometric parameters extracted in step S4 as outputs, the model is trained using a composite loss function. The pit morphology parameters include the depth amplitude parameter A and the lateral scale parameter σ. x and σ y And the correlation coefficient ρ.
2. The method according to claim 1, characterized in that, The training phase of step S1 is specifically as follows: a confocal microscope is used to acquire a full-focus color image and a micrometer-level real depth map. After normalizing the real depth map, it is divided into N discrete depth layers. Gaussian blur with different standard deviations is applied to the full-focus image to generate a set of blurred images. Based on the correspondence between the depth layer and the focal plane, a clear synthetic image of the corresponding depth layer region is constructed for each target focal plane. Finally, a multi-focus stack containing N images is generated.
3. The method according to claim 1, characterized in that, The frame-dimensional decoupling mechanism in step S2 is as follows: the input multi-focus stack tensor [B,N,H,W], where B is the batch size, N is the number of focus stack frames, and H and W are the image space dimensions, is reshaped into [B×N,1,H,W], so that the number of focus frames N is fully integrated into the batch dimension to complete the single-frame feature extraction. After single-frame feature extraction is completed, the feature tensor is reshaped into [B,N,C,H,W], where C is the number of feature channels, restoring the frame dimension-independent representation.
4. The method according to claim 1, characterized in that, The frequency domain sensing feature enhancement in step S2 is achieved through the wavelet frequency domain enhancement feature extraction module. Specifically, the single-frame image features are divided into identity branches, spatial branches, and frequency domain branches according to the channel ratio. The spatial branch uses anisotropic depth separable convolution to extract features, and the frequency domain branch extracts high-frequency components of the image through differentiable Haar wavelet decomposition. Finally, the three branches are spliced and fused to output a single-frame feature map.
5. The method according to claim 1, characterized in that, The global-local dual-branch 3D interlayer interaction in step S2 is as follows: the continuous variation features of adjacent focal layers are extracted by the lightweight 3D convolution of the local branch, the long-range dependency features across focal layers are extracted by the lightweight self-attention mechanism of the global branch, and the dual-branch features are dynamically weighted and fused by the adaptive gated fusion unit to output the focus response volume; the focus response volume is subjected to Softmax and weighted summation along the focal plane dimension to generate a normalized depth map standardized to the range of [0,1].
6. The method according to claim 1, characterized in that, Step S3 specifically involves: performing guided filtering on the normalized depth map using the depth map itself as the guiding image; and linearly transforming the filtered relative depth into absolute depth based on the physical distance between adjacent focal points and the horizontal physical size of a single pixel calibrated by the imaging system, thereby generating a three-dimensional point cloud.
7. The method according to claim 1, characterized in that, In step S4, the mathematical expression for the anisotropic 2D Gaussian height field model is: In the formula, (x0, y0) represents the center position of the pit, z0 represents the local base surface reference height, and Q is the covariance matrix; x0, y0, z0, A, and σ are obtained through nonlinear least squares optimization fitting. x σ y A total of 7-dimensional structured geometric parameters, ρ.
8. The method according to claim 1, characterized in that, In step S5, the material embedding encoding specifically involves mapping discrete material types into low-dimensional dense vectors through a learnable embedding module, concatenating the material embedding vectors with the standardized laser processing parameter vectors, inputting them into the backbone of a shared neural network, and outputting pit morphology parameters. The laser processing parameters include laser power P, pulse frequency f, number of pulses Np, and defocusing amount z.
9. The method according to claim 1, characterized in that, The composite loss function in step S5 includes a data fitting term L. mse With physical regularization term L phys The expression is: L=L mse +λ·L phys Wherein, λ is the regularization weight coefficient; the data fitting term is the mean square error between the predicted pit morphology parameters and the experimentally measured pit morphology parameters; the physical regularization term is a ReLU-type penalty term based on the laser-material interaction law.
10. The method according to claim 1, characterized in that, After step S4 and before step S5, the method also includes a closed 3D mesh generation step to conclude the shape perception step: based on the 3D point cloud, an open triangular mesh surface is generated by reconstructing the Poisson surface, the concave boundary ring is identified and the bottom vertex is constructed, a closed triangular mesh entity that meets the requirements of manifold and watertightness is generated, and exported as an STL, OBJ or STEP format file.