Two-photon photoetching part quality inspection method based on three-dimensional shift window multi-head self-attention

By adopting a quality inspection method based on three-dimensional shift window multi-head self-attention in the two-photon lithography process, the problem of relying on manual experience in part quality inspection is solved, real-time automated detection of part quality is realized, and large-scale industrial applications are supported.

CN120219394AInactive Publication Date: 2025-06-27GUIZHOU UNIV

Patent Information

Application Number
CN202510698832.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Real-time detection of part quality in two-photon lithography processes relies on manual experience, lacks efficient automation methods, and traditional offline detection is inefficient and difficult to meet the needs of large-scale production.

Method used

The two-photon lithography part quality inspection method based on three-dimensional shift window multi-head self-attention is adopted. Video data is collected in real time through the camera, divided into non-overlapping 3D blocks, features are extracted using 3D convolution, and a Transformer submodule including 3D window self-attention and 3D shift window self-attention is designed to build a Video-SWTrans hierarchical Transformer architecture to achieve end-to-end optimization and real-time detection.

Benefits of technology

Real-time automated detection of part quality is realized, the quality inspection cost and process debugging cycle are significantly reduced, and the large-scale industrial application of two-photon lithography technology is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219394A_ABST
    Figure CN120219394A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of additive manufacturing quality detection, and provides a two-photon photoetching part quality detection method based on three-dimensional shift window multi-head self-attention, which is used for automatically detecting the quality of parts in a two-photon photoetching process. The method comprises the following steps: acquiring a TPL processing process video through a camera, and dividing a video sequence by adopting a three-dimensional sliding window; the method comprises the following steps: constructing a Transform architecture containing 3D shift window multi-head self-attention and hierarchical feature fusion, and extracting spatial and temporal features; and a normal and defective part video training model is utilized to realize classification of cured, uncured and damaged states. Through cooperative calculation of global attention and a local window, the problems that a traditional CNN model sensing field is limited and LSTM calculation efficiency is low are solved, the detection accuracy reaches 96.38%, and real-time monitoring of four industrial scenes is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of additive manufacturing quality inspection, and particularly relates to a quality inspection method for two-photon lithography parts based on three-dimensional shifted window multi-head self-attention. Background Art

[0002] Two-Photon Lithography (TPL) is an advanced additive manufacturing technology based on nonlinear optical effects, capable of fabricating complex three-dimensional structures at the micron or even nanoscale, and is widely used in fields such as biomedical devices, micro-optical components, and flexible electronics. However, the industrial application of the TPL process faces a core bottleneck: the real-time detection of part quality (such as cured, uncured, damaged states) highly relies on manual experience, and there is a lack of efficient automation means for the dynamic adjustment of laser dose parameters. Since the curing effect of photosensitive resin during the TPL processing is affected by the coupling of multiple parameters such as laser intensity and scanning speed, traditional offline detection methods are inefficient and difficult to meet the demands of large-scale production. Therefore, the development of a high-precision and low-latency online quality inspection technology has become the key to promoting the industrialization of TPL.

[0003] Currently, deep learning algorithms based on computer vision are widely used in industrial quality inspection, but there are still significant limitations in the TPL scenario: 1) Traditional 3D-CNN models are difficult to model long-range spatio-temporal dependencies due to the limitation of fixed-size convolutional kernels, resulting in missed detection of tiny defects; 2) Hybrid architectures such as CNN-LSTM cannot process high-frame-rate videos in parallel due to their serial computing characteristics, and the inference latency is significant; 3) The global self-attention mechanism of Transformer architectures such as ViViT has a quadratic growth in computational complexity and is difficult to deploy to resource-constrained industrial devices; 4) The classification accuracy of existing generative models (such as GAN, VAE) drops significantly for uncured and micro-damaged states when there is a lack of fault samples. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a quality inspection method for two-photon lithography parts based on three-dimensional shifted window multi-head self-attention to solve the problems in the prior art. The technical solution adopted by the present invention is as follows: A quality inspection method for two-photon lithography parts based on three-dimensional shifted window multi-head self-attention includes the following steps: Step 1, real-time collect video data of the two-photon lithography process through a camera, sample it along the time axis to generate 10 equally long frame sequences, and construct a training set, a validation set, and a test set including cured, uncured, and damaged part states; Step 2, divide the video into non-overlapping 3D blocks, extract initial features through 3D convolution and linearly map them to a high-dimensional embedding space, add learnable class tokens and 3D position encoding, and retain spatio-temporal structure information; Step 3: Design a Transformer sub-module including 3D window self-attention and 3D shifted window self-attention; Step 4: Construct a Video-SWTrans hierarchical Transformer architecture, including a hierarchical Transformer structure with four stages; Step 5: Map the global information of multi-level spatio-temporal feature extraction to specific quality classification labels, and achieve end-to-end optimization through the cross-entropy loss function; Step 6: Deploy the model to an edge computing platform, perform inference optimization through the TensorRT engine, achieve parallel inference through multi-threaded pipeline processing, and synchronously record the spatio-temporal coordinates of defects.

[0005] Furthermore, in Step 1, the video resolution is 110×110 pixels, single-channel grayscale; the processing area of 25 independent parts is segmented on the X-Y plane, each part corresponds to a sub-video, and an equi-length 10-frame sequence is sampled from each sub-video along the time axis, with 5 frames overlapping between adjacent windows; the label of the sequence is determined according to the physical irreversibility rule: if there is at least one frame marked as "damaged" in the sequence, the whole is marked as "damaged"; if there is a "cured" frame and no damage, it is marked as "cured"; otherwise, it is marked as "uncured".

[0006] Furthermore, in Step 2, the preprocessed video sequence , is divided into non-overlapping 3D blocks, each block has a time span of 2 frames and a spatial size of 4×4 pixels; where, is time, is height, is width, is the number of channels; Extract block features for each block through 3D convolution operation, then increase the number of channels from 16 to 96 through linear transformation, and obtain the initial token after processing, with the dimension of ; Among them, the 3D convolution kernel size is , , is the embedding channel number, and the formula is expressed as: ; Before inputting the initial token into the Transformer network, splice the class token with it and add position encoding to form the final embedding representation, so as to convert the video data into the input token suitable for processing by the Transformer sub-module, which is expressed as: ; Among them, is the initial value of the learnable 3D position encoding tensor and follows a uniform distribution .

[0007] Furthermore, step three includes: Step 3.1, design 3D window multi-head self-attention to capture local spatio-temporal features, and divide the input tokens into multiple non-overlapping 3D windows with a size of , where is the spatial region size, is the number of consecutive frames; Use the multi-head self-attention mechanism to calculate the attention weights within each 3D window: Use the weight matrix , , to perform a linear transformation on the input tokens to generate the query , key and value vectors . The calculation method of a single attention head is: ; Among them, calculate the dot product of and to obtain the attention score, and use the Softmax function to normalize the attention score to obtain the attention weight, is the dimension of the key vector; Apply step 3.1 to multiple different , , vectors, and then splice and linearly transform the results to obtain the finally fused token B. The calculation formula is: ; Among them, is the number of attention heads, is the final output weight matrix; Add this bias to the attention weight when calculating self-attention. The formula is as follows: ; Among them, is the 3D relative position bias, which is adjusted according to the relative positions of the tokens in space and time; Step 3.2, based on the 3D window multi-head self-attention, perform a shift operation on the window. After the shift, the boundaries between the windows are broken, and a masking mechanism is introduced. For non-adjacent windows, set the attention weights to zero using the mask; Step 3.3, Integrate 3D window multi-head self-attention and 3D shifted window multi-head self-attention into the Transformer sub-module. The computational representation of data in a single Transformer sub-module is as follows: ; ; ; ; where, W-MSA3D is the 3D window multi-head self-attention, W-MSA3D is the 3D shifted window multi-head self-attention; represents the input and output tokens, and respectively represent the number and dimension of the current tokens; is the layer normalization operation; represents the intermediate token formed by adding the transformed token and the original token through residual connection; MLP is the multi-layer perceptron.

[0008] Furthermore, Step 4 includes: Step 4.1, Construct 2 Transformer sub-modules in Stage 1. After performing the Transformer sub-module operations on the token sequence obtained in Step 2 twice, maintain the resolution at ; Step 4.2, Construct 2 Transformer sub-modules in Stage 2. Increase the number of attention heads to , introduce the block merging module to splice the input tokens in spatial blocks and expand them along the channel dimension to , compress the channels to through layer normalization and linear projection, and reduce the output resolution to , and perform the Transformer sub-module operations twice in sequence; Step 4.3, Construct 18 Transformer sub-modules in Stage 3. Increase the number of attention heads to , apply the block merging module again, expand the channels to , reduce the output resolution to , and perform the Transformer sub-module operations 18 times in sequence; Step 4.4, Construct 2 Transformer sub-modules in Stage 4. Increase the number of attention heads to , apply the block merging module for the last time, the channels reach the maximum value , reduce the output resolution to , and perform the Transformer sub-module operations twice in sequence.

[0009] Further, in step five, in the final classification stage of the Video-SWTrans model, the model extracts global spatio-temporal features through the class tokens in the fourth stage of the Transformer architecture. These tokens are aggregated into high-dimensional vectors along the temporal and spatial dimensions through global average pooling, and then mapped to a 3D classification space through a multi-layer perceptron containing two linear layers and a GELU activation function. Finally, the output values are converted into a normalized probability distribution through the Softmax function, and the cross-entropy loss function is adopted: ; where is the one-hot encoding of the true label, is the predicted probability. Combining with the Adam optimizer, the model parameters are optimized end-to-end to achieve fully automatic inference from the original video input to quality classification.

[0010] Further, in step six, the trained Video-SWTrans model is deployed to the computing platform for inference optimization, receiving four video stream inputs in real time, and parallel inference is achieved through multi-threaded pipeline processing. When the "damaged" category is detected, the laser intensity is immediately reduced by 10%-20% and an audible and visual alarm is triggered, and at the same time, the defect spatio-temporal coordinates are recorded. When "uncured" is detected, the scanning speed is dynamically adjusted in fixed steps through a PID controller.

[0011] Further, the detection performance of the Video-SWTrans model is quantitatively evaluated using four metrics: Accuracy, Precision, Recall, and F1-score: ; ; ; ; where TN, TP, FN, and FP represent the numbers of true negative, true positive, false negative, and false positive samples, respectively.

[0012] The present invention has the following beneficial effects: The Video-SWTrans constructed in the present invention is a video Transformer framework based on three-dimensional shifted window multi-head self-attention (SW-MSA(3D)) and hierarchical feature fusion. Through local window calculation and periodic displacement mechanism, it can efficiently learn spatio-temporal distribution features under the condition of only normal part video data training. Aiming at different photosensitive resins, lithography modes and part geometries in two-photon lithography process, it significantly reduces the quality inspection cost and process debugging cycle. The application of this method realizes the real-time automatic detection of part quality, providing technical support for the large-scale industrialization of two-photon lithography technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is the flowchart of the method of the present invention; Figure 2 is the principle block diagram of the two-photon lithography part quality inspection method based on three-dimensional shifted window multi-head self-attention of the present invention; Figure 3 is the structure diagram of the core module of the model in the embodiment of the two-photon lithography part quality inspection method based on three-dimensional shifted window multi-head self-attention of the present invention; Figure 4 is the comparison diagram of the performance of the Video-SWTrans model and other different methods in the experiment in the embodiment of the two-photon lithography part quality inspection method based on three-dimensional shifted window multi-head self-attention of the present invention; Figure 5 is the comparison diagram of the performance indexes of the Video-SWTrans model and other different methods in scenario 3 (cone + ring writing mode) in the embodiment of the two-photon lithography part quality inspection method based on three-dimensional shifted window multi-head self-attention of the present invention; Figure 6 is the comparison diagram of the performance of the Video-SWTrans model under different settings in the embodiment of the two-photon lithography part quality inspection method based on three-dimensional shifted window multi-head self-attention of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The following will combine the Figures 1 - 6 in the embodiments of the present invention, and clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. If not specifically specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0015] The two-photon lithography part quality inspection method based on three-dimensional shifted window multi-head self-attention, as Figure 1Shown as follows: This method captures the spatio-temporal dependence of video data through a three-dimensional shifted window multi-head self-attention mechanism (SW-MSA(3D)), combines the advantages of hierarchical representation and parallel computing, and realizes the automatic classification of part quality (uncured, cured, damaged) in the two-photon lithography process, including the following steps: Step 1: Real-time collect video data of the two-photon lithography process through a camera, sample it along the time axis to generate 10 equally long frame sequences, and construct training sets, validation sets, and test sets including the states of cured, uncured, and damaged parts through physical irreversibility rules (such as "if there is a damaged frame, it is marked as damaged"); Step 2: Divide the video into non-overlapping 3D blocks, extract initial features through 3D convolution and linearly map them to a high-dimensional embedding space, add learnable class tokens (Class Token) and 3D position encoding, and retain spatio-temporal structure information to lay the foundation for subsequent Transformer network processing; Step 3: Design a Transformer sub-module containing 3D window self-attention (W-MSA) and 3D shifted window self-attention (SW-MSA); Step 4: Construct a Video-SWTrans hierarchical Transformer architecture, the core of which is a hierarchical Transformer structure including four stages; Step 5: Map the global information extracted from multi-level spatio-temporal features to specific quality classification labels (uncured / cured / damaged), and achieve end-to-end optimization through the cross-entropy loss function; Step 6: Deploy the model to an edge computing platform, implement FP16 quantization and layer fusion optimization through the TensorRT engine, dynamically adjust laser parameters and trigger an alarm for exceptions, and synchronously record the spatio-temporal coordinates of defects.

[0016] Specifically, in Step 1, the original video resolution is 110×110 pixels, single-channel grayscale. Divide the processing area of 25 independent parts on the X-Y plane, each part corresponds to a sub-video, sample each sub-video along the time axis to generate 10 equally long frame sequences, and the adjacent windows overlap by 5 frames to ensure time continuity. The label of the sequence is determined according to the physical irreversibility rule: if there is at least one frame marked as "damaged" in the sequence, the whole is marked as "damaged"; if there is a "cured" frame and no damage, it is marked as "cured"; otherwise, it is marked as "uncured"; Specifically, in Step 2, the preprocessed video sequence ( (10 frames × 110×110×1)) is divided into non-overlapping 3D blocks, the time span of each block is 2 frames, and the spatial size is 4×4 pixels. Through 3D convolution operation ( ) (kernel size Extract block features with a size of (2×4×4), a stride of 4, and 16 output channels, and then perform a linear transformation ( ) to increase the number of channels from 16 to 96, and the feature size becomes (where , is the embedding channel number) to generate initial tokens . The formula is expressed as: ; Before inputting the initial tokens into the Transformer network, concatenate the class token with them and add positional encoding to form the final embedding representation to convert the video data into input tokens suitable for processing by the Transformer sub-module , which is expressed as: ; Among them, is the initial value of the learnable 3D positional encoding tensor and follows a uniform distribution ; Specifically, step three includes: Step 3.1, design 3D window self-attention to capture local spatio-temporal features and reduce computational complexity. Divide the input tokens into multiple non-overlapping 3D windows with a size of ( is the spatial region size, is the number of consecutive frames). Within each 3D window, use the multi-head self-attention mechanism to calculate the attention weights. The specific process is as follows: Use weight matrices , , with learnable parameters to perform a linear transformation on the input tokens to generate queries ( ), keys ( ), and value vectors ( ). The calculation method of a single attention head is: ; Among them, calculate the dot product of and to obtain the attention scores, and use the Softmax function to normalize the attention scores to obtain the attention weights, is the dimension of the key vector, which is used to scale the dot product result to prevent gradient disappearance.

[0017] Apply the above process of step 3.1 to multiple different , , Vectors, and then splice and linearly transform the results. Denote the finally fused feature as B, and its calculation formula is: ; Among them, is the number of attention heads, is the final output weight matrix.

[0018] To better capture spatial and temporal position relationships, 3D relative position biases are introduced. This bias is added to the attention weights when calculating self-attention, and the specific formula is as follows: ; Among them, is the 3D relative position bias, which is adjusted according to the relative spatial and temporal positions between features. This bias can be optimized during the learning process to adapt to different position relationships; Step 3.2, Design 3D shifted window self-attention. Based on 3D window self-attention, perform a shifting operation on the window. After shifting, the boundaries between windows are broken, allowing information interaction between adjacent windows. To avoid information confusion between different windows after shifting, a masking mechanism is introduced. For non-adjacent windows, use a mask to set their attention weights to zero. This can ensure that information only flows between adjacent windows and avoid information chaos. By shifting the window, the model can capture broader spatio-temporal dependencies and make up for the locality limitation of W-MSA(3D); Step 3.3, Integrate 3D window self-attention and 3D shifted window self-attention into the Transformer sub-module. The propagation of data in a single Transformer sub-module can be expressed as: ; ; ; ; Among them, represents the input and output tokens, and represent the number and dimension of the current tokens respectively; is the layer normalization operation. Through the above process, the Transformer sub-module can effectively capture spatio-temporal features in video data and improve the training stability and performance of the model through operations such as residual connection and layer normalization.

[0019] Specifically, step four includes: Step 4.1: There are 2 Transformer sub-modules in Stage 1. After performing the Transformer sub-module operations on the token sequence obtained in Step 2 twice, the resolution is maintained at , and the generation of basic spatio-temporal tokens is completed, laying the foundation for subsequent stages; Step 4.2: There are 2 Transformer sub-modules in Stage 2, and the number of attention heads is increased to . In Stage 2, a Patch Merging module is introduced to splice the input tokens in spatial patches and expand them along the channel dimension to , and then compress the channels to through layer normalization and linear projection. The output resolution is reduced to , and the Transformer sub-module operations are performed twice in sequence to enhance the token representation ability; Step 4.3: There are 18 Transformer sub-modules in Stage 3, and the number of attention heads is increased to . In Stage 3, the Patch Merging module is applied again, and the channels are expanded to , and the output resolution is reduced to . The Transformer sub-module operations are performed 18 times in sequence to deeply explore the spatio-temporal dependencies of the tokens; Step 4.4: There are 2 Transformer sub-modules in Stage 4, and the number of attention heads is increased to . In Stage 4, the Patch Merging module is applied for the last time, and the channels reach the maximum value , and the output resolution is reduced to . The Transformer sub-module operations are performed twice in sequence to ensure that the output tokens can fully represent the spatio-temporal features of the video data.

[0020] Specifically, in Step 5, in the final classification stage of the Video-SWTrans model, the model extracts global spatio-temporal features through the class tokens of the fourth Transformer stage. These tokens are aggregated into a high-dimensional vector along the time and space dimensions through global average pooling, and then mapped to a 3-dimensional classification space (uncured / cured / damaged) through a multi-layer perceptron (MLP) containing two linear layers and a GELU activation function; finally, the output values are converted into a normalized probability distribution through the Softmax function, and the cross-entropy loss function is adopted: ; where is the one-hot encoding of the true label, For the predicted probability, the Adam optimizer (learning rate 0.0001, weight decay 0.01) is combined to optimize the model parameters end-to-end, realizing fully automated inference from the original video input to quality classification; Specifically, in step six, the trained Video-SWTrans model is deployed to the computing platform for inference optimization, reducing the model's video memory requirement from 3.2GB to 1.8GB and increasing the inference speed to 32 FPS. Four video streams are received in real time (resolution 110×110, frame rate 30 FPS), and parallel inference is achieved through multi-threaded pipeline processing. The CPU utilization rate ≤ 70%, and the end-to-end latency ≤ 200ms (including video decoding, preprocessing, model inference, and control signal generation). When the "damaged" category is detected, the laser intensity is immediately reduced by 10%-20% and an audible and visual alarm is triggered. At the same time, the spatio-temporal coordinates of the defect (X-Y-Z position and timestamp) are recorded; when "uncured" is detected, the scanning speed is dynamically adjusted in fixed steps through a PID controller to ensure a smooth change in the light intensity gradient. The computing platform is an existing technology, namely a computer, computer hardware, and edge device; for example, the NVIDIA Jetson AGX Xavier edge computing platform.

[0021] The detection performance of the Video-SWTrans model is quantitatively evaluated using four metrics: accuracy, precision, recall, and F1-score. A high accuracy indicates that the model performs well in correctly classifying samples. However, when dealing with an imbalanced dataset, it may produce biases. A higher precision means that the classifier has fewer false positive (FP) errors, but this may lead to some false negative (FN) samples. A higher recall means that the classifier has fewer false negative (FN) errors, but this may lead to some false positive (FP) samples. The F1 score is the harmonic mean of precision and recall, which provides a balanced evaluation of the classifier's ability. The F1 score is suitable for evaluating the performance of a classifier on an imbalanced dataset because it balances the trade-off between precision and recall. The specific calculation method is as follows: ; ; ; ; Where: TN, TP, FN, and FP represent the numbers of true negative, true positive, false negative, and false positive samples respectively.

[0022] The method for detecting the quality of two-photon lithography parts based on three-dimensional shifted window multi-head self-attention proposed by the present invention, as Figure 2As shown below, the specific implementation process is described in detail through the following embodiments. In this embodiment, the application scenario is the two-photon lithography production line of a certain precision manufacturing enterprise. The enterprise needs to quickly determine the laser dose threshold under various process parameters and monitor the part quality in real time to avoid under-curing or over-curing defects.

[0023] Embodiment: In the data preparation stage, the enterprise collected video data under four typical manufacturing scenarios, covering different part shapes (cuboid and cone), photoresist types (IP-DIP and custom materials), and writing modes (ring and stacking modes). The original videos were recorded by high-resolution industrial cameras, documenting the entire process of the laser beam scanning the photosensitive resin. To extract effective features, the videos were segmented into sub-videos according to the part spatial positions, with each sub-video corresponding to the manufacturing process of a single part. Subsequently, 10 frames were sampled at equal intervals from the sub-videos to construct a time series, with adjacent sequences overlapping by 5 frames to enhance the temporal continuity. The video frame annotation follows strict rules: if any frame in the sequence shows damage (such as bubbles or structural fractures), the entire sequence is marked as "damaged"; if there are cured frames and no damaged frames, it is marked as "cured"; otherwise, it is marked as "uncured". The final constructed dataset contains 37,444 samples, which are divided into a training set and a test set according to the scenarios. Among them, Scenario 3 (cone + ring writing mode) becomes the key test scenario for verifying the robustness of the model due to its small data volume and high structural complexity.

[0024] In the model construction and training stage, a Video Transformer (Video-SWTrans) architecture based on three-dimensional shifted window multi-head self-attention is adopted. The input video sequence (size 10×110×110×1) is first processed by a PatchEmbedding module, which divides the video into spatio-temporal patches using a 3D convolution operation with a kernel size of 2×4×4 and a stride of 4, and maps them to high-dimensional tokens. At the same time, learnable class tokens and position encodings are introduced to enhance the model's perception ability of global semantics and spatio-temporal positions. The processed tokens are input into a four-level hierarchical Transformer structure, where each level of the Transformer contains multiple Transformer sub-modules, and each sub-module is alternately composed of window multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA), as Figure 3As shown in the figure. W-MSA divides the input into non-overlapping three-dimensional windows (such as 8×7×7), calculates self-attention within the windows to capture local spatio-temporal correlations; SW-MSA re-divides the windows through window position offsets (4 frames offset on the time axis and 3 pixels offset on the spatial axis), and combines a masking mechanism to achieve cross-window information interaction, thereby expanding the global perception range. After each self-attention layer, there is a multi-layer perceptron (MLP) containing two fully-connected layers and the GELU activation function, and DropPath (ratio 0.2) is introduced to prevent overfitting. Finally, the features of the class token pass through a fully-connected layer to output the classification result, realizing the accurate recognition of three quality states: "uncured", "cured", and "damaged".

[0025] In the experimental verification stage, the model is trained on a computing platform, using the Adam optimizer with an initial learning rate of 0.0001, and the batch size is set to 23 to adapt to the video memory capacity. The training cycle is 200 epochs, and an early stopping strategy (patience value of 10 epochs) is used to avoid overfitting. At the same time, Dropout (ratio 0.2) and label smoothing (coefficient 0.1) are introduced to enhance the generalization ability. The experimental results are as Figure 4 shown. Video-SWTrans achieves an accuracy of 96.38% and an F1 value of 96.46% on the total dataset, which is significantly improved compared to baseline models such as ViViT (accuracy of 94.84%) and 3D-ResNet (accuracy of 93.23%). In scenario 3 with the least amount of data (cone + ring writing mode), as Figure 5 shown, the model still maintains an accuracy of 93.83%, which has a significant advantage over 3D-DenseNet (90.96%), verifying its strong adaptability to small-sample data. Ablation experiments show that, as Figure 6 shown, removing the three-dimensional relative position bias or the self-attention scaling mechanism respectively leads to a 0.32% and 0.26% decrease in accuracy, confirming the necessity of each module. The application of this method realizes the real-time and automated detection of part quality, providing technical support for the large-scale industrialization of two-photon lithography technology.

[0026] The Video-SWTrans constructed in the present invention is a video Transformer framework based on three-dimensional shifted window multi-head self-attention (SW-MSA(3D)) and hierarchical feature fusion. Through local window calculation and periodic displacement mechanism, it efficiently learns spatio-temporal distribution features under the tokenized sequence-level supervised learning framework. During the training phase, the model uses the labeled data of four types of industrial scenarios (including uncured, cured, and damaged state video sequences) to optimize the parameters of the classification head through cross-entropy loss. During the testing phase, the model directly outputs the category probabilities of uncured, cured, and damaged, and combines with an adaptive confidence threshold (adaptive confidence threshold = baseline mean + α × standard deviation, where α is dynamically adjusted according to the scenario) to achieve fine-grained classification. For different photosensitive resins, lithography modes, and part geometries in two-photon lithography processes, the average detection accuracy of this solution in four industrial scenarios reaches 96.38%, which is 1.54% higher than that of ViViT. The inference speed reaches 32 FPS. After being deployed to edge devices, the system latency is less than 200 ms, and the defect missed detection rate is reduced to 2.1%, significantly reducing the quality inspection cost and process debugging cycle.

[0027] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations, variations, modifications, and substitutions made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for quality inspection of two - photon lithography parts based on three - dimensional shifted window multi - head self - attention, characterized in that, It includes the following steps: Step 1: Real-time collect video data of the two-photon lithography process through a camera, sample along the time axis to generate 10 equal-length frame sequences, and construct training sets, validation sets, and test sets including the states of cured, uncured, and damaged parts; Step 2: Divide the video into non-overlapping 3D blocks, extract initial features through 3D convolution and linearly map them to a high-dimensional embedding space, add learnable class tokens and 3D position encoding, and retain spatio-temporal structure information; Step 3: Design a Transformer sub-module including 3D window self-attention and 3D shifted window self-attention; Step 4: Construct a Video-SWTrans hierarchical Transformer architecture, including a hierarchical Transformer structure with four stages; Step 5: Map the global information extracted by multi-level spatio-temporal features to specific quality classification labels, and achieve end-to-end optimization through the cross-entropy loss function; Step 6: Deploy the model to an edge computing platform, perform inference optimization through the TensorRT engine, achieve parallel inference through multi-threaded pipeline processing, and synchronously record the spatio-temporal coordinates of defects.

2. The method for inspecting quality of two-photon lithography parts based on three-dimensional shifted window multi-head self-attention according to claim 1, wherein In Step 1, the video resolution is 110×110 pixels, single-channel grayscale; the processing areas of 25 independent parts are segmented on the X-Y plane, each part corresponds to a sub-video, and each sub-video is sampled along the time axis to generate 10 equal-length frame sequences, with 5 overlapping frames between adjacent windows; the label of the sequence is determined according to the rule of physical irreversibility: if there is at least one frame marked as "damaged" in the sequence, the whole is marked as "damaged"; if there is a "cured" frame and no damage, it is marked as "cured"; otherwise, it is marked as "uncured".

3. The method for quality inspection of two-photon lithography parts based on three-dimensional shifted window multi-head self-attention according to claim 1, characterized in that In step two, the preprocessed video sequence is divided into non-overlapping 3D blocks, each block having a time span of 2 frames and a spatial size of 4×4 pixels; where is time, is height, is width, is the number of channels; Extract block features for each block through 3D convolution operations, and then increase the number of channels from 16 to 96 through a linear transformation. After processing, obtain the initial tokens The dimension of ; Among them, the 3D convolution kernel size is , , is the number of embedding channels, and the formula is expressed as: ; Before inputting the initial token into the Transformer network, concatenate the class token with it and add positional encoding to form the final embedding representation, so as to convert the video data into input tokens suitable for processing by the Transformer sub-module , which is expressed as: ; Among them, is the initial value of the learnable three-dimensional position encoding tensor and follows a uniform distribution .

4. The method for inspecting two-photon lithography parts based on three-dimensional shifted window multi-head self-attention according to claim 1, wherein Step 3 includes: Step 3.1, design 3D window multi-head self-attention to capture local spatio-temporal features, and divide the input tokens into multiple non-overlapping 3D windows of size , where is the spatial region size, and is the number of consecutive frames; Calculate the attention weights using the multi-head self-attention mechanism within each 3D window: Use a weight matrix with learnable parameters , , Perform a linear transformation on the input tokens to generate queries , keys and value vectors . The calculation method for a single attention head is as follows: ; Among them, calculate and to obtain the dot product, get the attention scores, and use the Softmax function to normalize the attention scores to obtain the attention weights, is the dimension of the key vector; Apply step 3.1 to multiple different , , vectors, and then splice and linearly transform the results to obtain the finally fused token B. The calculation formula is: ; Among them, is the number of attention heads, is the final output weight matrix; When calculating self-attention, add this bias to the attention weight, and the formula is as follows: ; Among them, is the 3D relative position offset, which is adjusted according to the relative positions of the markers in space and time; Step 3.2, based on the 3D window multi-head self-attention, perform a shift operation on the window. After the shift, the boundaries between windows are broken, and a masking mechanism is introduced. For non-adjacent windows, use the mask to set the attention weight to zero; Step 3.3, integrate the 3D window multi-head self-attention and the 3D shifted window multi-head self-attention into the Transformer sub-module. The calculation representation of data in a single Transformer sub-module is: ; ; ; ; Among them, W-MSA3D is the 3D window multi-head self-attention, W-MSA3D is the 3D shifted window multi-head self-attention; represents the input and output tokens, and represent the number and dimension of the current tokens respectively; is the layer normalization operation; represents the intermediate token formed by adding the transformed token and the original token through residual connection; MLP is the multi-layer perceptron.

5. The method for inspecting two-photon lithography parts based on three-dimensional shifted window multi-head self-attention according to claim 2, wherein Step 4 includes: Step 4.1, construct 2 Transformer sub-modules in the first stage. After performing the Transformer sub-module operations on the token sequence obtained in Step 2 twice, keep the resolution as ; Step 4.2, in stage two, construct 2 Transformer sub-modules, increase the number of attention heads to , introduce a block merging module to splice the input tokens according to spatial blocks and expand them along the channel dimension to , compress the channels to through layer normalization and linear projection, and reduce the output resolution to , and perform the Transformer sub-module operations twice in sequence; Step 4.3, in stage three, construct 18 Transformer sub-modules, and increase the number of attention heads to , and apply the block merging module again, expand the channels to , reduce the output resolution to , and perform the operations of the 18 Transformer sub-modules in sequence; Step 4.4, construct 2 Transformer sub-modules in Stage 4, and increase the number of attention heads to , and apply the block merging module for the last time, with the number of channels reaching the maximum value , and the output resolution is reduced to , and perform the operations of the Transformer sub-module twice in sequence.

6. The method for inspecting the quality of two-photon lithography parts based on three-dimensional shifted window multi-head self-attention according to claim 5, wherein In Step 5, through the final classification stage of the Video-SWTrans model, the model extracts global spatio-temporal features through the class tokens in the fourth stage of the Transformer architecture. This token is aggregated into a high-dimensional vector along the time and space dimensions through global average pooling, and then mapped to a 3D classification space through a multi-layer perceptron containing two linear layers and a GELU activation function; finally, the output value is converted into a normalized probability distribution through the Softmax function, and the cross-entropy loss function is adopted: ; where is the one-hot encoding of the true label, is the predicted probability. The model parameters are optimized end-to-end by combining with the Adam optimizer to achieve fully automatic inference from the original video input to quality classification.

7. The dual-photon lithography part quality inspection method based on three-dimensional shifted window multi-head self-attention according to claim 6, wherein: In step six, the trained Video-SWTrans model is deployed to the computing platform for inference optimization, receiving four video stream inputs in real time, and realizing parallel inference through multi-threaded pipeline processing; when the "damaged" category is detected, immediately reduce the laser intensity by 10%-20% and trigger an acousto-optic alarm, and record the defect spatio-temporal coordinates at the same time; when "uncured" is detected, the scanning speed is dynamically adjusted in fixed steps through a PID controller.

8. The method for inspecting two-photon lithography parts based on three-dimensional shifted window multi-head self-attention according to claim 7, wherein The detection performance of the Video-SWTrans model is quantitatively evaluated using four indicators: accuracy, precision, recall, and F1-score. ; ; ; ; Among them, TN, TP, FN, and FP represent the numbers of true negative, true positive, false negative, and false positive samples respectively.

Citation Information

Patent Citations

  • Abnormality prediction method, training method, device and equipment of abnormality prediction model

    CN115082490A

  • Thyroid ultrasound contrast nodule benign and malignant classification method based on 3D ConvFormer

    CN117237739A

  • Double-flow gating violence detection method and system based on expansion 3D convolutional network and Transform

    CN119942407A

  • Automatic unlabeled pancreas image segmentation system based on adversarial learning

    WO2023098289A1

Cited By

  • Training method, prediction method, device and equipment of two-photon image prediction model

    CN120580564A

  • A training method and a prediction method of a two-photon image prediction model, a device and equipment

    CN120580564B

  • Transform-based double-branch photovoltaic panel fault identification method

    CN121640174A