Complex optical system based on deep feature decoupling and autonomous alignment and diagnosis method
By employing deep feature decoupling and principal component analysis, combined with neural networks and spot data processing, the debugging challenges of complex optical systems in cold start conditions were solved, enabling efficient and accurate optical system alignment and diagnosis, and improving system stability and diagnostic capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
- Filing Date
- 2026-05-06
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional methods struggle to achieve effective convergence in the cold-start state of complex optical systems. Gradient vanishing and strong coupling of multidimensional parameters lead to long debugging times and difficulty in precise alignment. Pure artificial intelligence models have limited accuracy in the minimal deviation domain.
By employing a complex optical system based on deep feature decoupling, and combining a deep convolutional neural network with a ResNet-18 architecture and principal component analysis, high-dimensional parameter reduction adjustment and orthogonal scanning are achieved through spot data acquisition and preprocessing. Combined with neural network feature decoupling and real-time state evaluation, the optical system can achieve autonomous alignment and diagnosis.
It enables efficient transition from misalignment to the global optimal parameter pool, shortens debugging time, improves alignment accuracy and system stability, reduces hardware cost and complexity, and provides accurate diagnostic and early warning capabilities.
Smart Images

Figure CN122131479A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of optical instrument control and artificial intelligence, and in particular to a complex optical system based on deep feature decoupling and an autonomous alignment and diagnostic method. Background Technology
[0002] Platforms such as large-scale neutral atom quantum computing and gravitational wave detection have extremely high requirements for the spatiotemporal stability of complex optical architectures (such as multi-degree-of-freedom resonant cavity systems). They are extremely sensitive to thermal drift and mechanical vibration in the environment. Even a tiny optical path misalignment can lead to a precipitous drop in system performance or complete system collapse.
[0003] Traditional gradient-based control methods (such as hill climbing, genetic algorithms, and jitter locking) often fail completely in the "cold start" state—that is, when the system is not yet close to the resonance condition or has deviated significantly from the optimal operating point. This failure phenomenon stems from two fundamental physical mechanisms: (1) Gradient vanishing (information blind spot): When the system is far from the resonance region, the error signal (such as the PDH signal or the transmission spot) has clear gradient information only within a very narrow parameter range, while in most other parameter spaces the signal is almost flat or dominated by noise; (2) Strong coupling and degeneracy of multidimensional parameters: There is a highly nonlinear strong coupling between spatial alignment (such as beam tilt, position offset) and mode matching (such as lens axial position). Specific parameter combinations may produce extremely similar physical responses in the error signal, making it difficult to distinguish in measurement which hardware dimension is misaligned.
[0004] In this context, relying entirely on human experience for "blind tuning" is not only extremely time-consuming but also makes it very difficult to recover the optimal state. Gradient-based iterative algorithms are prone to getting stuck in local optima or diverging directly, making it difficult to achieve effective convergence. On the other hand, while existing pure artificial intelligence solutions have a global perspective, they suffer from inherent model validation loss and hardware lag in the domain of minimal deviation (submicron level), making it difficult to achieve 100% physical alignment accuracy using deep learning models alone. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention aims to provide a complex optical system and an autonomous alignment and diagnostic method based on deep feature decoupling, in order to solve the problems of "curse of dimensionality" caused by strong coupling of high-dimensional optical parameters, limited accuracy of pure artificial intelligence models in the domain of extremely small deviations, and the heavy reliance on experience and lack of specificity in traditional debugging.
[0006] To achieve the above objectives, this invention proposes a complex optical system based on deep feature decoupling, including a resonant cavity matching system, a spot data acquisition system, and a preprocessing system;
[0007] The resonant cavity supporting system includes a laser source, fiber optic coupler, half-wave plate, polarization beam splitter module, quarter-wave plate, plano-convex lens, first reflector, second reflector, and FP resonant cavity arranged sequentially. The spot data acquisition and preprocessing system includes a main industrial camera and a secondary industrial camera. The light emitted from the laser source is sequentially incident on the FP resonant cavity through the fiber optic coupler, half-wave plate, polarization beam splitter module, quarter-wave plate, plano-convex lens, first reflector, and second reflector. The main industrial camera is used to capture the physical evolution of the transmitted light spot in the FP resonant cavity. The main industrial camera is connected to the preprocessing system, which performs temporal locking and morphological normalization processing to synthesize an average intensity projection feature map (the transmitted light spot changes periodically; the light spot of one period in the video is extracted and processed by the average intensity projection algorithm). The reflected light from the FP resonant cavity is sequentially incident on the secondary industrial camera through the second reflector, first reflector, plano-convex lens, and then through the quarter-wave plate and polarization beam splitter module.
[0008] The plano-convex lens is mounted on an electric linear displacement stage. The electric linear displacement stage drives the plano-convex lens to move along the optical axis, which can adjust the beam waist size and position of the incident light. The first and second reflectors are respectively mounted on an electric frame. The electric frame drives the corresponding reflectors to rotate, which can adjust the angle and position of the light incident on the FP resonant cavity. The auxiliary industrial camera is used to monitor the position of the reflected light spot. The position of the reflected light spot is used as feedback to realize the closed-loop movement of the electric frame (because the electric frame itself is an open loop, the actual distance is different when the clockwise and counterclockwise rotations are synchronized. The position of the reflected light spot is used as feedback to realize the closed-loop movement of the electric frame).
[0009] The above scheme also includes a neural network feature decoupling and inference model, and a real-time state assessment and hybrid scanning system. The neural network feature decoupling and inference model includes a deep convolutional neural network based on the ResNet-18 architecture. It can take the average intensity projection feature map as input, learn through the network to understand the complex interference phenomena caused by the loss of synchronization of the penetrating light spot, and invert the active adjustment angles of the two motorized frames and the axial coordinates of a motorized linear displacement stage in the latent space, ultimately achieving dimensionality reduction and truncation of the high-dimensional representation. The real-time state assessment and hybrid scanning system can use visual purity as a benchmark, and improve the fundamental mode purity to the target threshold through mode switching and orthogonal one-dimensional scanning to achieve control closed loop.
[0010] In the above scheme: when the purity of the fundamental model reaches the first preset threshold, the real-time state evaluation and hybrid scanning system can automatically cut off the large-span neural network inference loop and seamlessly switch to the dimensionality reduction local fine-tuning mode: strictly perform orthogonal sequential one-dimensional linear scanning on the five control dimensions; each dimension is optimized independently, and locked immediately when the maximum purity point is captured, until the global fundamental model purity exceeds the second preset threshold, realizing the control loop. That is, the task of the neural network has been completed, and the subsequent switch to two frames and four dimensions for separate linear scanning can reach 90% of the second preset threshold, and the whole process ends.
[0011] In the above scheme: when the purity of the fundamental model reaches the first preset threshold, the purity of the fundamental model is >60%; when the purity of the fundamental model exceeds the second preset threshold, the purity of the fundamental model is >90%.
[0012] This invention also proposes an autonomous alignment and diagnosis method for complex optical systems based on depth feature decoupling, including the aforementioned autonomous alignment and diagnosis method for complex optical systems based on depth feature decoupling, and further including the following steps:
[0013] S1. Collect the transmitted and reflected light spot data of the FP resonant cavity and construct a training dataset (a video frame) containing light spot images and position parameters.
[0014] S2. Preprocess the training dataset to generate an average intensity projection feature map, input the average intensity projection feature map into the neural network feature decoupling and inference model, infer the multi-degree-of-freedom adjustment parameters, achieve coarse alignment, and make the purity of the fundamental model reach the first preset threshold.
[0015] S3. Switch to the orthogonal serialization one-dimensional linear scan fine-tuning mode, optimize and lock the extreme point in each dimension, so that the purity of the fundamental model reaches the second preset threshold, and complete the closed-loop alignment with a 100% success rate.
[0016] S4. Based on the adjusted data and principal component analysis, decouple the system's sensitive dimensions and output the optical system's sensitivity diagnostic results. Conversely, through principal component analysis, it is possible to analyze which dimensions contribute the most to the system's impact and their respective proportions.
[0017] In the above scheme: step S2, the preprocessing includes the following steps:
[0018] 1) Obtain the video stream to be processed and calculate the global average brightness of each frame in the video stream; sort the frames according to their average brightness and extract the frames with the lowest brightness. The frame is used as a noise sample frame; the noise sample frames are stacked and averaged at the pixel level to generate a single-frame reference noise image that represents the basic background noise of the system.
[0019] 2) Traverse all valid frames of the video stream, convert each frame into a high-precision floating-point grayscale image, and accumulate it pixel by pixel to generate the total energy distribution map of the video stream;
[0020] 3) Perform multi-stage denoising processing on the image sequentially: First, perform non-local mean denoising to smooth Poisson and Gaussian noise and preserve target edges; second, perform residual background suppression to force pixels below the preset background baseline threshold to zero; finally, perform median filtering to eliminate residual high-frequency impulse noise and output the final average intensity projection image.
[0021] In the above scheme: the purity of the fundamental model is calculated through the following steps:
[0022] 1) Extract a single complete resonant sweep cycle using a time-domain algorithm;
[0023] 2) Introduce morphological feature operators such as roundness index, Gaussian fit of energy distribution and spatial centroid to perform strict mode discrimination on continuous frames within the period, thereby accurately separating the fundamental mode and higher-order transverse mode components.
[0024] 3) The ratio of the integral energy of the frame corresponding to the fundamental mode within the period to the total integral energy of all resonant modes is defined as the fundamental mode purity of the system.
[0025] In the above scheme: In step S2, the sensitivity diagnosis extracts the dominant direction of the system's optical response variance through principal component analysis, and defines the first principal component as the most sensitive physical dimension of the system for fault location and structural reinforcement guidance, prioritizing fault location.
[0026] The beneficial effects of this invention are:
[0027] This invention represents a leap from experience-based "blind tuning" to a "precise targeted diagnosis" optimal troubleshooting mechanism. It is not merely an alignment tool, but a data-driven "microscopic diagnostic microscope" for optical systems. Utilizing the massive data mapping patterns captured by neural networks in latent space, combined with the orthogonal decoupling technique of Principal Component Analysis (PCA), it mathematically extracts the dominant variance direction (i.e., the first principal component) that most significantly affects the system's optical response from the entangled chaos of 5 degrees of freedom (5-DOF). This endows the system with unparalleled early warning and diagnostic capabilities: during routine maintenance or crash recovery, the system can clearly identify "where is the most sensitive and has the greatest impact on the system's spatiotemporal stability" within the multidimensional complex hardware. Engineers no longer need to blindly experiment based on intuition and human experience; instead, they can prioritize checking and reinforcing the most sensitive parts of the system's output guidance (such as the specific lens axis or the pitch angle of a key reflector), greatly reducing troubleshooting time and completely ending the "black box" era of optical debugging.
[0028] An unprecedented "single-shot hit" breakthrough overcomes efficiency bottlenecks: Facing the "blind zone" of severely deviated resonance, this invention completely abandons the hundreds of steps of local random walks required by traditional algorithms. Relying on the forward and inverse mapping established by the end-to-end model, this invention achieves a "transition" from an out-of-order state to the globally optimal parameter pool. Although in rare cases it may not be possible to truly "pass on the first try" due to hardware lag and extremely low signal-to-noise ratio, in physical verification, an average of only 2-3 iterations of neural network macroscopic inference are needed to instantly boost the coupling efficiency, which was on the verge of collapse, to a safe zone of 60% or even higher, reducing alignment time by several orders of magnitude.
[0029] AI-driven adjustment has many error points, resulting in poor performance during fine-tuning. However, it can help quickly adjust to near the optimal position. A subsequent linear scan can achieve a spot fundamental mode purity of over 90% with 100% accuracy. Initially, it relies on the "global perspective" provided by data-driven approaches to escape local pitfalls (breaking the ice from 0% to 60%). Later, within the most error-sensitive local smooth manifolds, it switches to rigorous traditional physical orthogonal scanning (achieving a sprint from 60% to 90%+). This dual architecture effectively avoids the convergence blind spots of a single technical path, establishing a 100% convergence success rate for handling complex optical systems from misalignment to perfect coupling. Finally, in the adjustment phase, the system uses only an industrial camera as a sensor, significantly reducing hardware costs and system complexity. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of the optical path of the present invention.
[0031] Figure 2 It is a control flowchart.
[0032] Figure 3 This is a graph showing the convergence relationship between pattern purity and iteration steps.
[0033] Figure 4 This is a plot of variance explained by principal component analysis.
[0034] Figure 5 This is a dimensionality reduction trajectory diagram comparing this method with traditional methods. Detailed Implementation
[0035] A complex optical system based on deep feature decoupling mainly consists of a resonant cavity supporting system, a spot data acquisition system, a neural network feature decoupling and inference model, a real-time state assessment and hybrid scanning system, and a preprocessing system.
[0036] The resonant cavity supporting system includes, in sequence, a laser source, an optical fiber coupler 1, a half-wave plate 2, a polarization beam splitter 4, a quarter-wave plate 5, a plano-convex lens 6, a first reflector 8, a second reflector 9, and an FP resonant cavity 10. The spot data acquisition and preprocessing system includes a main industrial camera 11 and a secondary industrial camera 3.
[0037] The light emitted from the laser source passes sequentially through fiber coupler 1, half-wave plate 2, polarization beam splitter 4, quarter-wave plate 5, plano-convex lens 6, first reflector 8, and second reflector 9 before entering the FP resonant cavity 10. The main industrial camera 11 captures the physical evolution of the transmitted light spot in the FP resonant cavity 10. The main industrial camera 11 is connected to a preprocessing system, which performs temporal locking and morphological normalization to synthesize an average intensity projection feature map (the transmitted light spot changes periodically; one period of the light spot is extracted from the video and processed using an average intensity projection algorithm).
[0038] The reflected light from the FP resonant cavity 10 passes sequentially through the second reflector 9, the first reflector 8, and the plano-convex lens 6, and then through the 1 / 4 wave plate 5 and the polarization beam splitter 4 before being incident on the auxiliary industrial camera 3.
[0039] A plano-convex lens 6 is mounted on an electrically driven linear displacement stage 7. The stage 7 moves the lens 6 along the optical axis, adjusting the beam waist and position of the incident light. A first reflecting mirror 8 and a second reflecting mirror 9 are mounted on an electrically driven frame. The frame rotates the corresponding mirrors, adjusting the angle and position of the light incident on the FP resonant cavity 10. An auxiliary industrial camera 3 monitors the position of the reflected light spot, using this position as feedback to achieve closed-loop movement of the electrically driven frame (because the frame is open-loop, clockwise and counterclockwise rotations have different actual distances when the rotation counts are the same; using the reflected light spot position as feedback enables closed-loop movement of the frame).
[0040] The neural network feature decoupling and inference model includes a deep convolutional neural network based on the ResNet-18 architecture (an existing model). It can take the average intensity projection feature map as input, learn through the network the complex interference appearance caused by the loss of synchronization of the penetrating light spot, and invert the active adjustment angle of the two motorized frames and the axial coordinate of a motorized linear displacement stage in the latent space, ultimately achieving dimensionality reduction and truncation of the high-dimensional representation. The real-time state assessment and hybrid scanning system can use visual purity as a benchmark, and improve the fundamental mode purity to the target threshold through mode switching and orthogonal one-dimensional scanning to achieve control closed loop.
[0041] When the fundamental model purity reaches the first preset threshold, the real-time state evaluation and hybrid scanning system automatically cuts off the large-span neural network inference loop and seamlessly switches to a dimensionality reduction local fine-tuning mode: strictly performing orthogonal sequential one-dimensional linear scanning on the five control dimensions; each dimension independently optimizes, locking the point where the purity maximum is captured, until the global fundamental model purity exceeds the second preset threshold, thus achieving control loop closure. That is, the neural network's task is complete, and subsequent switching to separate linear scanning of the two frames and four dimensions achieves 90% of the second preset threshold, ending the entire process.
[0042] Specifically, when the purity of the fundamental model reaches the first preset threshold, the purity of the fundamental model is >60%; when the purity of the fundamental model exceeds the second preset threshold, the purity of the fundamental model is >90%.
[0043] An autonomous alignment and diagnostic method for complex optical systems based on depth feature decoupling, such as:
[0044] First, data for training the neural network model is collected. The motorized linear displacement stage 7 and the first and second reflecting mirrors 8 and 9 mounted on the motorized mirror frame are manually adjusted to near-optimal positions, where the fundamental mode purity is around 90%. Then, the motorized linear displacement stage 7 and the first and second reflecting mirrors 8 and 9 are randomly adjusted in five dimensions. After each dimension is adjusted, the position of the reflected light spot is recorded using the auxiliary industrial camera 3 to help the five dimensions return to their initial positions. After the five dimensions are adjusted, the main industrial camera 11 captures video information of the transmitted light spot for approximately 2 seconds. After this, the five dimensions automatically return to their initial positions. This process is repeated 3000 times to obtain 3000 sets of transmitted light spot video information and their corresponding five-dimensional position information, which are used as training data.
[0045] The collected training data is converted into images using a high-fidelity extraction method for spot images based on temporal noise floor evaluation and integral projection, including the following steps.
[0046] 1. Acquire the video stream to be processed, and calculate the global average brightness of each frame in the video stream; sort the frames according to their average brightness and extract the frames with the lowest brightness. The frame is used as a noise sample frame; the noise sample frames are averaged at the pixel level to generate a single-frame reference noise image representing the basic background noise of the system. The extracted image is used to train the neural network as a training set.
[0047] 2. Traverse all valid frames of the video stream, convert each frame into a high-precision floating-point grayscale image, and accumulate the images pixel by pixel to generate the total energy distribution map of the video stream.
[0048] 3. Perform multi-stage denoising processing on the image sequentially: First, perform non-local mean denoising to smooth Poisson and Gaussian noise and preserve target edges; second, perform residual background suppression to force pixels below the preset background baseline threshold to zero; finally, perform median filtering to eliminate residual high-frequency impulse noise and output the final average intensity projection (AIP) image.
[0049] In the quantitative calculation of fundamental mode purity, since the intensity of the transmitted light spot exhibits periodic dynamic evolution with the frequency sweep of the resonant cavity, we first accurately extract a single complete resonant frequency sweep cycle using a time-domain algorithm. Subsequently, we introduce morphological feature operators such as the roundness index, the Gaussian fit of the energy distribution, and the spatial centroid to rigorously discriminate the continuous frames within this cycle, thereby accurately separating the fundamental mode and higher-order transverse mode components. Finally, the ratio of the integral energy of the frame corresponding to the fundamental mode within this cycle to the total integral energy of all resonant modes is rigorously defined as the absolute fundamental mode purity of the system. Here, we use a trained neural network to predict the light spot image. The image is extracted from the video in real time, and the extraction steps are consistent with the high-fidelity extraction method for light spot images based on temporal noise floor evaluation and integral projection described above. Simultaneously, the video undergoes another real-time processing, namely the processing method described in this paper. The images processed by the two methods have different uses: the former is used to infer dimensional parameters, while the latter is used to measure the fundamental mode purity of the light spot (i.e., to determine the accuracy of the neural network's prediction, whether it has reached the set first threshold; if it fails once, it is repeated until the target is achieved).
[0050] like Figure 2 During the laser alignment or adjustment process, the transmitted light spot video acquired in real time is first processed into an average intensity projection (AIP) image, which is then input into a ResNet-18 neural network model for training. The output model can predict the positional parameters corresponding to five dimensions from a single AIP image. After several iterations, when the fundamental mode purity is >60%, the control system switches to linear scanning of each of the five dimensions individually, taking the position of the maximum value of the fundamental mode purity, and the entire adjustment process ends. Figure 3 The diagram shows the convergence relationship between pattern purity and iteration steps.
[0051] Based on the predicted data, we extract the principal component eigenvalues using principal component analysis (PCA). In this embodiment, the eigenvalue This precisely represents the severity of the system's optical response. The first principal component (PC1) accurately corresponds to the direction of maximum global variance in the data projection, meaning the system's optical response exhibits the most dramatic gradient change in this direction. Therefore, PC1 defines the system's most critical sensitivity point, providing an absolute physical priority reference for subsequent optical path isolation design and servo control bandwidth allocation.
[0052] Figure 4 The variance explanation plot shows that the cumulative variance explained by the first three principal components (PC1-PC3) reaches 91.8%, indicating that the seemingly large and complex 5-dimensional mechanical control space is actually deeply degenerate in its effective optical response within a 3-dimensional feature subspace. The first principal component (PC1) alone accounts for 58.3% of the total system variance, indicating that the parameters of PC1 have the greatest impact on the system and represent the most sensitive physical direction. Figure 5 The diagram shows the dimensionality reduction trajectory of the method of this invention and the traditional method, illustrating the efficiency of the method of this invention in convergence adjustment.
Claims
1. A complex optical system based on depth feature decoupling, characterized in that: This includes a resonant cavity supporting system, a spot data acquisition system, and a preprocessing system; The resonant cavity system includes a laser source, a fiber optic coupler (1), a half-wave plate (2), a polarization beam splitter (4), a quarter-wave plate (5), a plano-convex lens (6), a first reflector (8), a second reflector (9), and an FP resonant cavity (10) arranged sequentially; the spot data acquisition and preprocessing system includes a main industrial camera (11) and a secondary industrial camera (3); the light emitted from the laser source passes sequentially through the fiber optic coupler (1), the half-wave plate (2), the polarization beam splitter (4), and the quarter-wave plate (5). Wave plate (5), plano-convex lens (6), first reflector (8), and second reflector (9) are incident on FP resonant cavity (10). The main industrial camera (11) is used to capture the physical evolution of the transmitted light spot of FP resonant cavity (10). The main industrial camera (11) is connected to the preprocessing system. The preprocessing system performs time-domain locking and morphological normalization processing to synthesize the average intensity projection feature map. The reflected light from FP resonant cavity (10) passes through the second reflector (9), the first reflector (8), and the plano-convex lens (6) in sequence, and then passes through the 1 / 4 wave plate (5) and polarization beam splitting module (4) to be incident on the auxiliary industrial camera (3). The plano-convex lens (6) is mounted on an electric linear displacement stage (7). The plano-convex lens (6) is moved along the optical axis by the electric linear displacement stage (7), which can adjust the beam waist size and position of the incident light. The first reflector (8) and the second reflector (9) are mounted on an electric frame. The corresponding reflector is rotated by the electric frame, which can adjust the angle and position of the light incident on the FP resonant cavity (10). The auxiliary industrial camera (3) is used to monitor the position of the reflected light spot and use the position of the reflected light spot as feedback to realize the closed-loop movement of the electric frame.
2. The complex optical system based on depth feature decoupling according to claim 1, characterized in that: The system includes a neural network feature decoupling and inference model, and a real-time state assessment and hybrid scanning system. The neural network feature decoupling and inference model includes a deep convolutional neural network based on the ResNet-18 architecture. It can take the average intensity projection feature map as input, learn the complex interference appearance caused by the misalignment of the transmitted light spot through the network, and invert the active adjustment angles of the two motorized frames and the axial coordinates of a motorized linear displacement stage in the latent space, ultimately achieving dimensionality reduction and truncation of the high-dimensional representation. The real-time state assessment and hybrid scanning system can improve the fundamental mode purity to the target threshold by using visual purity as a benchmark and through mode switching and orthogonal one-dimensional scanning, thereby achieving a control closed loop.
3. The complex optical system based on depth feature decoupling according to claim 2, characterized in that: When the purity of the fundamental model reaches the first preset threshold, the real-time state evaluation and hybrid scanning system can automatically cut off the large-span neural network inference closed loop and seamlessly switch to the dimensionality reduction local fine-tuning mode: strictly perform orthogonal sequential one-dimensional linear scanning on the five control dimensions; each dimension is optimized independently, and locked immediately when the maximum purity extreme point is captured, until the global fundamental model purity exceeds the second preset threshold, thus realizing the control closed loop.
4. The complex optical system based on depth feature decoupling according to claim 3, characterized in that: When the purity of the fundamental model reaches the first preset threshold, the purity of the fundamental model is >60%; when the purity of the fundamental model exceeds the second preset threshold, the purity of the fundamental model is >90%.
5. An autonomous alignment and diagnostic method for complex optical systems based on depth feature decoupling, characterized in that: The autonomous alignment and diagnostic method for complex optical systems based on depth feature decoupling, as described in any one of claims 1-4, further includes the following steps: S1. Collect the transmitted and reflected light spot data of the FP resonant cavity (10) and construct a training dataset containing light spot images and position parameters. S2. Preprocess the training dataset to generate an average intensity projection feature map, input the average intensity projection feature map into the neural network feature decoupling and inference model, infer the multi-degree-of-freedom adjustment parameters, achieve coarse alignment, and make the purity of the fundamental model reach the first preset threshold. S3. Switch to the orthogonal serialization one-dimensional linear scan fine-tuning mode, optimize and lock the extreme point in each dimension, so that the purity of the fundamental model reaches the second preset threshold, and complete the closed-loop alignment with a 100% success rate. S4. Based on the adjustment data and principal component analysis, decouple the system's sensitivity dimension and output the optical system sensitivity diagnosis results.
6. The autonomous alignment and diagnosis method for complex optical systems based on depth feature decoupling according to claim 5, characterized in that, In step S2, the preprocessing includes the following steps: 1) Obtain the video stream to be processed and calculate the global average brightness of each frame in the video stream; sort the frames according to their average brightness and extract the frames with the lowest brightness. The frame is used as a noise sample frame; the noise sample frames are stacked and averaged at the pixel level to generate a single-frame reference noise image that represents the basic background noise of the system. 2) Traverse all valid frames of the video stream, convert each frame into a high-precision floating-point grayscale image, and accumulate it pixel by pixel to generate the total energy distribution map of the video stream; 3) Perform multi-stage denoising processing on the image sequentially: First, perform non-local mean denoising to smooth Poisson and Gaussian noise and preserve target edges; second, perform residual background suppression to force pixels below the preset background baseline threshold to zero; finally, perform median filtering to eliminate residual high-frequency impulse noise and output the final average intensity projection image.
7. The autonomous alignment and diagnosis method for complex optical systems based on depth feature decoupling according to claim 5 or 6, characterized in that: The purity of the fundamental model is calculated through the following steps: 1) Extract a single complete resonant sweep cycle using a time-domain algorithm; 2) Introduce morphological feature operators such as roundness index, Gaussian fit of energy distribution and spatial centroid to perform strict mode discrimination on continuous frames within the period, thereby accurately separating the fundamental mode and higher-order transverse mode components. 3) The ratio of the integral energy of the frame corresponding to the fundamental mode within the period to the total integral energy of all resonant modes is defined as the fundamental mode purity of the system.
8. The autonomous alignment and diagnosis method for complex optical systems based on depth feature decoupling according to claim 5, characterized in that: In step S2, the sensitivity diagnosis extracts the dominant direction of the system's optical response variance using principal component analysis, and defines the first principal component as the most sensitive physical dimension of the system for fault location and structural reinforcement guidance.