Optical path installation and adjustment error detection method and device based on sliding window self-attention network

By using sliding window self-attention network in optical path installation error detection, the problem of insufficient detection accuracy and generalization performance in the prior art is solved, and the capabilities of high-precision, dynamic measurement and multi-objective prediction are achieved, which significantly improves detection efficiency and reduces hardware costs.

CN120071094APending Publication Date: 2025-05-30BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228303.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture complex global features and multi-dimensional interactions in optical path mounting error detection, resulting in insufficient detection accuracy and generalization performance.

Method used

Using a sliding window self-attention network method, by introducing a sliding window mechanism, the network can extract features in local windows and communicate information across windows, taking into account both local and global modeling capabilities.

Benefits of technology

It realizes high-precision optical path mounting error detection, can dynamically measure and decouple the impact of multiple mounting errors, has multi-objective prediction capabilities, significantly improves detection efficiency, and reduces hardware requirements and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071094A_ABST
    Figure CN120071094A_ABST
Patent Text Reader

Abstract

The invention discloses an optical path installation and adjustment error detection method and device based on a sliding window self-attention network, which have the advantages of high precision, dynamic measurement and multi-target prediction capability, can realize end-to-end optical path installation and adjustment error detection, remarkably improves the optical path installation and adjustment efficiency, reduces the hardware requirement and cost required by optical path installation and adjustment, and improves the optical path installation and adjustment accuracy. The system can be operated on a portable device. The method comprises the following steps: (1) building a sliding window self-attention network main body architecture; (2) different installation and adjustment errors are set in the light path, corresponding image surface wavefront maps are collected, and a data set used for training and testing the sliding window self-attention network is made; (3) training the sliding window self-attention network, and performing iterative training on the sliding window self-attention network by using the data set to enable the sliding window self-attention network to reach an available state; and (4) the sliding window self-attention network processes the wavefront image of the to-be-measured image surface, the wavefront image of the to-be-measured image surface is input into the neural network, and the neural network outputs a corresponding optical path adjustment error value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optoelectronic detection, and in particular to a method for detecting optical path alignment errors based on a sliding window self-attention network, and also to a device for detecting optical path alignment errors based on a sliding window self-attention network, which is used to establish a sliding window self-attention neural network detection model that inputs a single-frame image plane wavefront map and outputs multiple optical path alignment error values. Background Art

[0002] With the increasing complexity of optical systems, such as in applications like interferometers and lidars, precise optical path alignment is particularly crucial. The alignment level of an optical system is often the most critical link restricting the system performance, and the optical path alignment accuracy directly affects the imaging quality and measurement accuracy. In high-resolution imaging or precision measurement, tiny optical path deviations can lead to problems such as aberrations and spot errors, thereby affecting the overall performance of the system. In addition, the development of automated alignment technology has further improved production efficiency and system reliability. Therefore, adopting certain computer-aided alignment technologies to ensure high-precision optical path alignment is the core element for realizing high-performance optical systems and improving the process level.

[0003] Traditional computer-aided alignment technologies include wavefront detection and misalignment calculation. Wavefront detection obtains the image plane wavefront map of the device under test through a wavefront sensing system; misalignment calculation inversely calculates the misalignment amount based on the multi-field image plane wavefront map to guide the optical path alignment. Common calculation algorithms include the sensitivity matrix method, the inverse optimization method, and the vector aberration model, etc. The sensitivity matrix method generates a matrix using the relationship between Zernike coefficients and misalignment amounts. It has a simple principle and wide application, but is easily affected by the coupling of degrees of freedom, resulting in non-linear ill-conditioned problems; the inverse optimization method iteratively optimizes the measured Zernike coefficients through an evaluation function, solving the ill-conditioned error problem, but it is time-consuming and difficult to apply to complex systems; the vector aberration model represents wave aberration with vectors and calculates the misalignment amount through the aberration field offset vector. Although it has high accuracy, the modeling is complex.

[0004] In recent years, deep learning has achieved remarkable results in multiple fields with its powerful data processing capabilities, and people have also begun to try to apply it to the detection of optical path alignment errors. The introduction of deep learning transforms the traditional optical path error detection method that relies on physical models into a data-driven method. By training a neural network with a large amount of data, it automatically learns the complex non-linear relationship between optical path alignment errors and the image plane wavefront map, thus getting rid of the dependence on precise physical models. This method is expected to more efficiently solve complex problems in the detection of optical path alignment errors.

[0005] Currently, the deep learning optical path alignment error detection method uses a convolutional neural network to learn the features related to the alignment error in the image plane wavefront map. The influence of the optical path alignment error on the image plane wavefront is global. The local receptive field of the convolutional neural network limits its ability to capture complex global features. When dealing with the optical path alignment error, it is difficult for it to learn the long-range dependencies in the wavefront image. Especially when there are multiple alignment errors coupled, the multi-dimensional interaction of the errors will limit the detection ability and generalization performance of the network. Summary of the Invention

[0006] To overcome the defects of the prior art, the technical problem to be solved by the present invention is to provide an optical path alignment error detection method based on a sliding window self-attention network, which has the advantages of high precision, dynamic measurement, and multi-target prediction ability. It can realize end-to-end optical path alignment error detection, significantly improve the efficiency of optical path alignment, and reduce the hardware requirements and costs for optical path alignment, and can run on portable devices.

[0007] The technical solution of the present invention is as follows: This optical path alignment error detection method based on a sliding window self-attention network includes the following steps:

[0008] (1) Build the main architecture of the sliding window self-attention network, including: input layer, image segmentation layer, linear mapping layer, feature extraction block, feature downsampling layer, and output layer;

[0009] (2) Make a data set. Set different alignment errors in the optical path and collect the corresponding image plane wavefront maps to make a data set for training and testing the sliding window self-attention network;

[0010] (3) Train the sliding window self-attention network. Use the data set to iteratively train the sliding window self-attention network to make it reach a usable state;

[0011] (4) The sliding window self-attention network processes the image plane wavefront map to be measured. Input the image plane wavefront map to be measured into the neural network, and the neural network outputs the corresponding optical path alignment error value.

[0012] The present invention constructs a sliding window self-attention network capable of learning the mapping relationship between the image wavefront and the optical path alignment error. The network introduces a sliding window mechanism, which can not only extract features in the local window, but also perform cross-window information exchange, taking into account both local and global modeling capabilities. Therefore, the network can achieve high-precision prediction of the optical path alignment error. The network can realize the calculation of the optical path alignment error only relying on a single image wavefront map, and can achieve dynamic measurement. And the network can decouple the influence of multiple alignment errors on the image wavefront, and realize the simultaneous output of multiple alignment errors by a single network, with the ability of multi-target prediction. In addition, the network realizes the end-to-end calculation of the optical path alignment error, significantly improving the detection efficiency. At the same time, this method reduces the hardware requirements and costs, enabling it to operate effectively on portable devices.

[0013] There is also provided an optical path alignment error detection device based on a sliding window self-attention network, which includes:

[0014] A network main structure building module configured to build the main structure of the sliding window self-attention network, including: an input layer, an image segmentation layer, a linear mapping layer, a feature extraction block, a feature downsampling layer, and an output layer;

[0015] A data set making module configured to set different alignment errors in the optical path and collect corresponding image wavefront maps, and make a data set for training and testing the sliding window self-attention network;

[0016] A sliding window self-attention network training module configured to iteratively train the sliding window self-attention network using the data set to make it reach an available state;

[0017] An alignment error output module configured to input the image wavefront map into the sliding window self-attention network, and the neural network outputs the corresponding optical path alignment error value. Description of the Drawings

[0018] Figure 1 is a flowchart of the method for detecting the optical path alignment error based on the sliding window self-attention network of the present invention.

[0019] Figure 2 is a structure diagram of the sliding window self-attention network.

[0020] Figure 3 is a structure diagram of the feature extraction block.

[0021] Figure 4 is a diagram of the aspherical surface measurement optical path.

[0022] Figure 5 is an image wavefront map.

[0023] Figure 6It is a scatter plot of the residual error distribution of the network prediction result. Figure 6 a is the scatter plot of the residual error distribution of the X tilt, Figure 6 b is the scatter plot of the residual error distribution of the Y eccentricity, Figure 6 c is the scatter plot of the residual error distribution of the Z translation. Specific implementation mode

[0024] As Figure 1 shown, this optical path alignment error detection method based on the sliding window self-attention network includes the following steps:

[0025] (1) Build the main architecture of the sliding window self-attention network, including: input layer, image segmentation layer, linear mapping layer, feature extraction block, feature downsampling layer and output layer;

[0026] (2) Make a data set, set different alignment errors in the optical path and collect the corresponding image plane wavefront maps, and make a data set for training and testing the sliding window self-attention network;

[0028] (3) Train the sliding window self-attention network, and use the data set to iteratively train the sliding window self-attention network to make it reach an available state;

[0029] (4) The sliding window self-attention network processes the image plane wavefront map to be measured, inputs the image plane wavefront map to be measured into the neural network, and the neural network outputs the corresponding optical path alignment error value.

[0030] The present invention constructs a sliding window self-attention network that can learn the mapping relationship between the image plane wavefront and the optical path alignment error. The network introduces a sliding window mechanism, which can not only extract features in the local window, but also perform cross-window information exchange, taking into account both local and global modeling capabilities. Therefore, the network can achieve high-precision prediction of the optical path alignment error. The network can realize the calculation of the optical path alignment error only relying on a single image plane wavefront map, and can realize dynamic measurement. And the network can decouple the influence of multiple alignment errors on the image plane wavefront, and realize that a single network outputs multiple alignment errors at the same time, with the ability of multi-target prediction. In addition, the network realizes the end-to-end calculation of the optical path alignment error, significantly improving the detection efficiency. At the same time, this method reduces the hardware requirements and costs, making it able to effectively run on portable devices.

[0031] Preferably, in the step (1), the input layer accepts the original image as input;

[0032] The image segmentation layer evenly and non-overlappingly cuts the input image into 4×4 image blocks, and then flattens each image block into a one-dimensional vector;

[0033] The linear mapping layer takes the vector after flattening the image patch as input, and maps each input vector to a vector of length C through a fully connected layer, and outputs features with a length and width that are 1 / 4 of the length and width of the input image and a number of channels of C. Figure 1 ;

[0034] Features Figure 1 After passing through 2 feature extraction blocks, they are mapped to features Figure 2 Features Figure 2 The size is consistent with that of features Figure 1 ;

[0035] The feature downsampling layer 1 takes features Figure 2 as input, doubles the number of channels while halving the length and width of features Figure 2 by merging adjacent vectors, and then maps them to features Figure 3 after passing through 2 feature extraction blocks;

[0036] The feature downsampling layer 2 takes features Figure 3 as input, doubles the number of channels while halving the length and width of features Figure 3 by merging adjacent vectors, and then maps them to features Figure 4 after passing through 18 feature extraction blocks;

[0037] The feature downsampling layer 3 takes features Figure 4 as input, doubles the number of channels while halving the length and width of features Figure 4 by merging adjacent vectors, and then maps them to features Figure 5 after passing through 2 feature extraction blocks;

[0038] The output layer takes features Figure 5 as input, uses a global average pooling layer to compress features Figure 5 into a vector of length M, then passes it into a fully connected layer to map it to a vector of length n, and finally outputs n predicted optical path alignment error values through a Sigmoid activation function.

[0039] Preferably, in the step (1),

[0040] The feature extraction block includes the following structure:

[0041] 1) Layer Normalization LN (Layer Normalization)

[0042] To maintain the stability of features during transmission, perform normalization processing on features. The formula is:

[0043]

[0044] where μ is the mean, σ is the standard deviation, and ε is an extremely small number to prevent the denominator from being 0;

[0045] 2) Window-based Multi-head Self-

[0046] Attention

[0047] The input feature map is divided into several non-overlapping windows, each window with a size of M*M. The multi-head self-attention calculation MSA (Multi-head Self-

[0048] Attention) is performed on the features within each window. The calculation formula is:

[0049]

[0050] where softmax(·) is the normalized exponential function, Q is the query matrix, K is the key matrix, V

[0051] is the value matrix, and d k is the scaling factor;

[0052] 3) Multi-Layer Perceptron (MLP)

[0053] MLP includes two linear layers and a non-linear activation function, which is used to perform further non-linear transformation on the features;

[0054] 4) Shifted Window Multi-head Self-

[0055] Attention

[0056] SW-MSA adds a shifted window mechanism on the basis of the standard W-MSA. After W-

[0057] MSA, by shifting the window by half of the window size, enables features originally belonging to different windows to interact with each other, thereby capturing cross-window information;

[0058] 5) Residual connection

[0059] The output after feature mapping is directly added to the input features to prevent gradient disappearance and promote the flow of features.

[0060] Preferably, as Figure 3As shown, in the step (1), the process of feature mapping in the feature extraction block is as follows: the input features first go through the first feature mapping, including a LN and a W-MSA operation; then through the second feature mapping including a LN and an MLP operation; then through the third feature mapping, including a LN and a SW-MSA operation; and finally through the fourth feature mapping, including a LN and an MLP operation; a residual connection is added for each feature mapping.

[0061] Preferably, in the step (3), the Adam optimizer is used when training the network, and the loss function uses RMSE, and the formula is:

[0062]

[0063] where N is the number of output alignment error values, y i represents the true value of the i-th alignment error, and y i ′ is the network output value of the i-th alignment error.

[0064] Those of ordinary skill in the art can understand that all or part of the steps in implementing the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the steps of the method of the above embodiments, and the storage medium can be: ROM / RAM, magnetic disk, optical disk, memory card, etc. Therefore, corresponding to the method of the present invention, the present invention also simultaneously includes an optical path alignment error detection device based on a sliding window self-attention network. This device is usually represented in the form of functional modules corresponding to the steps of the method. The device includes:

[0065] A network main structure building module, which is configured to build the main structure of the sliding window self-attention network, including: an input layer, an image segmentation layer, a linear mapping layer, a feature extraction block, a feature downsampling layer, and an output layer;

[0066] A data set production module, which is configured to set different alignment errors in the optical path and collect corresponding image plane wavefront maps to produce a data set for training and testing the sliding window self-attention network;

[0067] A sliding window self-attention network training module, which is configured to iteratively train the sliding window self-attention network using the data set to make it reach an available state;

[0068] An alignment error output module, which is configured to input the image plane wavefront map into the sliding window self-attention network, and the neural network outputs the corresponding optical path alignment error value.

[0069] Preferably, in the network main structure building module, the input layer receives the original image as input;

[0070] The image segmentation layer evenly and non - overlappingly cuts the input image into 4×4 image patches, and then flattens each image patch into a one - dimensional vector;

[0071] The linear mapping layer accepts the vector after the image patch is flattened as input, and maps each input vector to a vector of length C through a fully - connected layer, outputting features with a length and width that are 1 / 4 of the length and width of the input image and a channel number of C Figure 1 ;

[0072] Features Figure 1 After passing through 2 feature extraction blocks, it is mapped to features Figure 2 , Features Figure 2 The size is consistent with features Figure 1 ;

[0073] The feature down - sampling layer 1 accepts features Figure 2 as input, and doubles the channel number while halving the length and width of features Figure 2 by merging adjacent vectors, and then after passing through 2 feature extraction blocks, it is mapped to features Figure 3 ;

[0074] The feature down - sampling layer 2 accepts features Figure 3 as input, and doubles the channel number while halving the length and width of features Figure 3 by merging adjacent vectors, and then after passing through 18 feature extraction blocks, it is mapped to features Figure 4 ;

[0075] The feature down - sampling layer 3 accepts features Figure 4 as input, and doubles the channel number while halving the length and width of features Figure 4 by merging adjacent vectors, and then after passing through 2 feature extraction blocks, it is mapped to features Figure 5 ;

[0076] The output layer accepts features Figure 5 as input, uses the global average pooling layer to compress features Figure 5 into a vector of length M, and then passes it into a fully - connected layer to map it to a vector of length n, and finally outputs n predicted optical path alignment error values through the Sigmoid activation function.

[0077] Preferably, in the network main structure building module,

[0078] The feature extraction block includes the following structure:

[0079] 1) Layer normalization LN

[0080] To maintain the stability of features during transmission, perform normalization processing on features. The formula is:

[0081]

[0082] Among them, μ is the mean, σ is the standard deviation, and ε is an extremely small number to prevent the denominator from being zero;

[0083] 2) Window Multi-Head Self-Attention (W-MSA)

[0084] The input feature map is divided into several non-overlapping windows, and the size of each window is M*M.

[0085] Perform multi-head self-attention calculation (MSA) on the features within each window. The calculation formula is:

[0086]

[0087] where softmax(·) is the normalization exponential function, Q is the query matrix, K is the key matrix, V

[0088] is the value matrix, and d k is the scaling factor;

[0089] 3) Multi-Layer Perceptron (MLP)

[0090] MLP includes two linear layers and a non-linear activation function, which is used to perform further non-linear transformation on the features;

[0091] 4) Sliding Window Multi-Head Self-Attention (SW-MSA)

[0092] SW-MSA adds a sliding window mechanism on the basis of the standard W-MSA. After W-

[0093] MSA, by shifting the window by half of the window size, enables features originally belonging to different windows to interact with each other, thereby capturing cross-window information;

[0094] 5) Residual connection

[0095] Directly add the output after feature mapping to the input features to prevent gradient disappearance and promote the flow of features.

[0096] Preferably, in the network body structure building module, the process of feature mapping in the feature extraction block is as follows: the input features first go through the first feature mapping, including an LN and a W-MSA operation; then go through the second feature mapping, including an LN and an MLP operation; then go through the third feature mapping, including an LN and an SW-MSA operation; finally go through the fourth feature mapping, including an LN and an MLP operation; a residual connection is added to each feature mapping.

[0097] Preferably, in the sliding window self-attention network training module, the Adam optimizer is used when training the network, and the loss function uses RMSE. The formula is:

[0098]

[0099] Among them, N is the number of output alignment error values, and y i represents the true value of the i-th alignment error, and y i ′ is the network output value of the i-th alignment error.

[0100] The following details a specific implementation example of the present invention. A method for detecting the alignment error of an aspherical measurement optical path based on a sliding window self-attention network is implemented as follows:

[0101] The process of establishing a sliding window self-attention network for detecting the alignment error of an aspherical measurement optical path is as Figure 1 shown, and the specific implementation steps are as follows:

[0102] Step 1: Build a sliding window self-attention network

[0103] In the example, build a sliding window self-attention network as Figure 2 shown. The size of the input original image is set to 256*256, and the number of output predicted optical path alignment error values is 3.

[0104] Step 2: Make a dataset

[0105] In the example, use Zemax software to build an aspherical surface measurement optical path as Figure 4 shown. Randomly set the alignment errors of the measured mirror, including: X tilt misalignment, Y eccentricity misalignment, and Z translation misalignment. Among them, the setting range of the X tilt misalignment is (-40″, 40″), the setting range of the Y eccentricity misalignment is (-40μm, 40μm), and the setting range of the Z translation misalignment is (-40μm, 40μm). Run the optical path tracing and record the image plane wavefront diagram as Figure 5 shown, with a resolution of 256*256. Repeat the experiment to obtain 2500 groups of data. Among them, the image plane wavefront diagram is used as the dataset sample, and the corresponding alignment error is used as the dataset label to jointly form the dataset. Use 2000 groups of data to train the network and 500 groups of data to test the detection accuracy of the network.

[0106] Step 3: Train the sliding window self-attention network

[0107] In the example, use 2000 groups of data to train the sliding window self-attention network. The training batch size is 32, and a total of 300 rounds of iterative training are performed. The learning rate is set to 1×10 -4 .

[0108] Step 4: The sliding window self-attention network processes the image plane wavefront diagram to be measured

[0109] In the example, 500 groups of test data are processed using the trained sliding window self-attention network, as Figure 6 shown. The average prediction accuracy of the network for the X tilt misalignment is 0.24″, for the Y eccentricity misalignment is 1.87 μm, and for the Z translation misalignment is 1.25 μm.

[0110] The beneficial effects of the present invention are as follows:

[0111] 1. The optical path alignment error detection method based on the sliding window self-attention network disclosed by the present invention introduces a sliding window mechanism into the network, enabling the method to balance local and global modeling capabilities and achieve high-precision detection of optical path alignment errors.

[0112] 2. The optical path alignment error detection method based on the sliding window self-attention network disclosed by the present invention can achieve the calculation of optical path alignment errors only relying on a single image plane wavefront map and can realize dynamic measurement.

[0113] 3. The optical path alignment error detection method based on the sliding window self-attention network disclosed by the present invention can decouple the influence of multiple alignment errors on the image plane wavefront, realize the simultaneous output of multiple alignment errors by a single network end-to-end, and has the ability of multi-object prediction, significantly improving the detection efficiency.

[0114] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for detecting optical path adjustment errors based on a sliding window self-attention network, characterized in that: It includes the following steps: (1) Build the main architecture of the sliding window self-attention network, including: input layer, image segmentation layer, linear mapping layer, feature extraction block, feature downsampling layer and output layer; (2) Produce a data set, set different alignment errors in the optical path and collect corresponding image wavefront images to produce a data set for training and testing the sliding window self-attention network; (3) Train the sliding window self-attention network and iteratively train the sliding window self-attention network using the dataset to make it usable. (4) The sliding window self-attention network processes the wavefront image of the image plane to be measured, inputs the wavefront image of the image plane to be measured into the neural network, and the neural network outputs the corresponding optical path adjustment error value.

2. The optical path adjustment error detection method based on sliding window self-attention network according to claim 1 is characterized in that: In the step (1), the input layer receives the original image as input; The image segmentation layer cuts the input image into 4×4 image blocks evenly and without overlap, and then flattens each image block into a one-dimensional vector; The linear mapping layer accepts the flattened vector of the image block as input, maps each input vector to a vector of length C through a fully connected layer, and outputs a feature map 1 with a length and width of 1 / 4 of the length and width of the input image and a number of channels of C; After passing through two feature extraction blocks, feature map 1 is mapped to feature map 2, and the size of feature map 2 is consistent with that of feature map 1; Feature downsampling layer 1 accepts feature map 2 as input, reduces the length and width of feature map 2 by half and doubles the number of channels by merging adjacent vectors, and then maps it to feature map 3 after passing through two feature extraction blocks; Feature downsampling layer 2 accepts feature map 3 as input, reduces the length and width of feature map 3 by half and doubles the number of channels by merging adjacent vectors, and then maps it to feature map 4 after passing through 18 feature extraction blocks; Feature downsampling layer 3 accepts feature map 4 as input, reduces the length and width of feature map 4 by half and doubles the number of channels by merging adjacent vectors, and then maps it to feature map 5 after passing through two feature extraction blocks; The output layer accepts feature map 5 as input, uses the global average pooling layer to compress feature map 5 into a vector of length M, and then passes it into a fully connected layer to map it into a vector of length n. Finally, it outputs n predicted optical path adjustment error values ​​through the Sigmoid activation function.

3. The optical path adjustment error detection method based on sliding window self-attention network according to claim 2 is characterized in that: In the step (1), The feature extraction block contains the following structure: 1) Layer Normalization LN In order to maintain the stability of features during the transmission process, the features are normalized. The formula is: Among them, μ is the mean, σ is the standard deviation, and ε is a very small number to prevent the denominator from being 0; 2) Window Multi-Head Self-Attention W-MSA The input feature map is divided into several non-overlapping windows, each of which is of size M*M. Multi-head self-attention MSA is calculated for the features in each window, and the calculation formula is: where softmax(·) is the normalized exponential function, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the scaling factor; 3) Multilayer Perceptron MLP MLP consists of two linear layers and a nonlinear activation function to perform further nonlinear transformation on features; 4) Sliding Window Multi-Head Self-Attention SW-MSA SW-MSA adds a sliding window mechanism to the standard W-MSA. After W-MSA, by offsetting the window by half the window size, features originally belonging to different windows can interact with each other, thereby capturing cross-window information; 5) Residual Connection The output after feature mapping is directly added to the input features to prevent gradient vanishing and promote the flow of features.

4. The optical path adjustment error detection method based on sliding window self-attention network according to claim 3 is characterized in that: In step (1), the process of feature mapping in the feature extraction block is as follows: the input feature first undergoes a first feature mapping, including an LN and a W-MSA operation; then undergoes a second feature mapping, including an LN and a MLP operation; then after the third feature mapping, including an LN and a SW- MSA operation; finally, the fourth feature mapping, including an LN and an MLP Operation; a residual connection is added to each feature map.

5. The optical path adjustment error detection method based on sliding window self-attention network according to claim 4 is characterized in that: In step (3), Adam is used to train the network. The optimizer uses RMSE as the loss function, and the formula is: Where N is the number of output adjustment errors, y i represents the true value of the i-th adjustment error, y i ′The network output value of the ith adjustment error.

6. An optical path adjustment error detection device based on a sliding window self-attention network, characterized in that: It includes: The main network structure building module is configured to build the main structure of the sliding window self-attention network, including: input layer, image segmentation layer, linear mapping layer, feature extraction block, feature downsampling layer and output layer; A data set production module, which is configured to set different alignment errors in the optical path and collect corresponding image wavefront images to produce data sets for training and testing the sliding window self-attention network; A sliding window self-attention network training module, which is configured to iteratively train the sliding window self-attention network using a dataset to achieve a usable state; The adjustment error output module is configured to input the image wavefront image into the sliding window self-attention network, and the neural network outputs the corresponding optical path adjustment error value.

7. The optical path adjustment error detection device based on sliding window self-attention network according to claim 6 is characterized in that: In the network main structure building module, the input layer accepts the original image as input; The image segmentation layer cuts the input image into 4×4 image blocks evenly and without overlap, and then flattens each image block into a one-dimensional vector; The linear mapping layer accepts the flattened vector of the image block as input, maps each input vector to a vector of length C through a fully connected layer, and outputs a feature map 1 with a length and width of 1 / 4 of the length and width of the input image and a number of channels of C; After passing through two feature extraction blocks, feature map 1 is mapped to feature map 2, and the size of feature map 2 is consistent with that of feature map 1; Feature downsampling layer 1 accepts feature map 2 as input, reduces the length and width of feature map 2 by half and doubles the number of channels by merging adjacent vectors, and then maps it to feature map 3 after passing through two feature extraction blocks; Feature downsampling layer 2 accepts feature map 3 as input, reduces the length and width of feature map 3 by half and doubles the number of channels by merging adjacent vectors, and then maps it to feature map 4 after passing through 18 feature extraction blocks; Feature downsampling layer 3 accepts feature map 4 as input, reduces the length and width of feature map 4 by half and doubles the number of channels by merging adjacent vectors, and then maps it to feature map 5 after passing through two feature extraction blocks; The output layer accepts feature map 5 as input, uses the global average pooling layer to compress feature map 5 into a vector of length M, and then passes it into a fully connected layer to map it into a vector of length n. Finally, it outputs n predicted optical path adjustment error values ​​through the Sigmoid activation function.

8. The optical path adjustment error detection device based on sliding window self-attention network according to claim 7 is characterized in that: In the network main structure building module, The feature extraction block contains the following structure: 1) Layer Normalization LN In order to maintain the stability of features during the transmission process, the features are normalized. The formula is: Among them, μ is the mean, σ is the standard deviation, and ε is a very small number to prevent the denominator from being 0; 2) Window Multi-Head Self-Attention W-MSA The input feature map is divided into several non-overlapping windows, each of which is of size M*M. Multi-head self-attention MSA is calculated for the features in each window, and the calculation formula is: where softmax(·) is the normalized exponential function, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the scaling factor; 3) Multilayer Perceptron MLP MLP consists of two linear layers and a nonlinear activation function to perform further nonlinear transformation on features; 4) Sliding Window Multi-Head Self-Attention SW-MSA SW-MSA adds a sliding window mechanism to the standard W-MSA. After W-MSA, by offsetting the window by half the window size, features originally belonging to different windows can interact with each other, thereby capturing cross-window information; 5) Residual Connection The output after feature mapping is directly added to the input features to prevent gradient vanishing and promote the flow of features.

9. The optical path adjustment error detection device based on sliding window self-attention network according to claim 8 is characterized in that: In the network main structure building module, the process of feature mapping in the feature extraction block is as follows: the input feature first undergoes the first feature mapping, including an LN and a W-MSA operation; then undergoes the second feature mapping, including an LN and an MLP operation; then undergoes the third feature mapping, including an LN and a SW-MSA operation; and finally undergoes the fourth feature mapping, including an LN and an MLP operation; A residual connection is added to each feature map.

10. The optical path adjustment error detection method based on sliding window self-attention network according to claim 9 is characterized in that: In the sliding window self-attention network training module, the Adam optimizer is used to train the network, and the loss function uses RMSE, and the formula is: Where N is the number of output adjustment errors, y i represents the true value of the i-th adjustment error, y i ′The network output value of the ith adjustment error.