Coherent free-space optical communication system using the HS-VIT model
By introducing the HS-VIT model into the CFSOC system, combined with the hierarchical spatial enhancement mechanism and sparse cue channel self-attention, the problem of low wavefront aberration correction accuracy under complex turbulent conditions is solved, and higher communication performance and stability are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-13
AI Technical Summary
Existing deep learning control algorithms suffer from low correction accuracy and poor stability when correcting wavefront aberrations under complex turbulent conditions, which limits the communication performance of the CFSOC system.
A coherent free-space optical communication system employing the HS-VIT model, combined with a high-speed camera, wavefront controller, and wavefront corrector, introduces a hierarchical spatial enhancement mechanism and sparse cue channel self-attention in the Transformer encoder layer through the HS-VIT model, thereby improving the accuracy and stability of wavefront aberration correction.
It significantly improves the correction accuracy and real-time performance of the Sensorless AO system under complex turbulent conditions, enhances the communication performance of the CFSOC system, reduces the bit error rate, and improves mixing efficiency.
Smart Images

Figure CN121396327B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of optical communication technology, and in particular to a coherent free-space optical communication system employing the HS-VIT model. Background Technology
[0002] With the continuous growth of demand for ultra-high bandwidth and ultra-low latency applications in sixth-generation communication systems, CFSOC (Coherent Free Space Optical Communication) has been widely used in communication scenarios such as ground-to-satellite communication, inter-satellite communication, security communication, and emergency disaster relief communication due to its advantages of independent spectrum resources, high transmission rate, strong anti-interference ability, and relatively low power consumption.
[0003] However, when the laser signal is transmitted through the atmospheric channel, it is affected by atmospheric turbulence, which causes problems such as intensity flicker and wavefront aberration. Ultimately, this leads to a decrease in the mixing efficiency (ME) of CFSOC, an increase in the bit error rate (BER), and a significant decline in communication performance.
[0004] Using adaptive optics (AO) systems at the receiver of a CFSOC system allows for real-time correction of the effects of atmospheric turbulence. However, the large size and complex structure of traditional AO systems significantly limit their potential applications. Therefore, employing sensorless AO systems to compensate for aberrations is more beneficial for CFSOC systems in various scenarios, such as… Figure 1 The figure shows a schematic diagram of the CFSOC system structure with a Sensorless AO system.
[0005] Existing deep learning control algorithms in sensorless AO systems, such as CNNs (Convolutional Neural Networks), use point spread function (PSF) images of the wavefront and distorted wavefront Zernike coefficients as input for training. The trained model is then applied to the sensorless AO system to predict the Zernike coefficients after phase compensation, thereby performing phase compensation on the wavefront. Alternatively, ResNets (Residual Networks) classify wavefront aberrations using PSF images, determine the corresponding Zernike coefficients based on the classification, and perform targeted phase compensation.
[0006] However, due to the high complexity of wavefront distortion and the structural constraints of the network model itself, the two control algorithms mentioned above still suffer from low correction accuracy and poor stability when correcting wavefront aberrations under complex turbulent conditions, which limits the communication performance of the CFSOC system and restricts its application. Summary of the Invention
[0007] In order to achieve higher wavefront aberration correction accuracy of the Sensorless AO system under complex turbulent conditions and improve the communication performance of the CFSOC system, this invention provides a coherent free space optical communication system using the HS-VIT model.
[0008] A coherent free-space optical communication system employing the HS-VIT model includes a wavefront-free adaptive optics system for correcting wavefront distortion. The wavefront-free adaptive optics system comprises:
[0009] A high-speed camera is used to acquire wavefront PFS images and transmit the preprocessed PFS images to the wavefront controller.
[0010] A wavefront controller is used to identify the wavefront features carried by the PFS image and generate a control voltage signal for the wavefront corrector based on the wavefront features.
[0011] The wavefront corrector receives the control voltage signal and performs wavefront distortion correction, then outputs the corrected optical signal to the beam splitter.
[0012] The wavefront controller identifies wavefront features carried by PFS images by internally constructing an HS-VIT model. The HS-VIT model establishes a hierarchical spatial enhancement mechanism by combining convolutional simulation self-attention on the basic architecture of the VIT model, and applies it to the Transformer encoder layer in the VIT model. Furthermore, it uses a self-modulated feature aggregation method to process the enhanced features at the end of the Transformer encoder layer through sparse cue channel self-attention, thereby intelligently distinguishing key information and background noise in turbulent wavefront features.
[0013] Technical effects:
[0014] This invention first establishes a hierarchical spatial enhancement mechanism by combining the Vision Transformer model with convolutional simulation of self-attention. This mechanism is then applied to the Transformer encoder layer of the Vision Transformer model, repeatedly applying convolutional simulation of self-attention at different depths within the encoder layer, forming a progressive spatial feature enhancement process. Simultaneously, a self-modulated feature aggregation method is used to replace the traditional feature aggregation method of the Vision Transformer model. This method refines the enhanced features at the end of the Transformer encoder layer through sparse cue channel self-attention, intelligently distinguishing key information from background noise in turbulent wavefront features. This enables the HS-VIT model to more accurately aggregate spatial features in complex atmospheric turbulence environments, providing more reliable feature support for wavefront aberration correction, significantly improving the correction accuracy of the Sensorless AO system under complex turbulent conditions, and further enhancing the communication performance of the CFSOC system.
[0015] To further evaluate the performance of the HS-VIT model in correcting wavefront aberrations under complex turbulent conditions, CNN, ResNet, and existing VIT models were selected for comparison with this embodiment. These four network models were used to independently perform 100 simulations under three different turbulent conditions, and the mean changes of RMS, ME, and BER after correction were compared.
[0016] like Figure 11 As shown, the HS-VIT model significantly outperforms the other three deep learning network models under different turbulent conditions, maintaining a clear advantage. Under weak turbulence conditions, these network models can effectively correct wavefront aberrations. As turbulence intensity increases, wavefront distortion becomes more complex, and traditional network models exhibit structural limitations, failing to identify and correct fine-grained local wavefront features. Particularly under strong turbulence conditions, only the HS-VIT model can improve the system's mean error rate (ME) to above 0.8 after correction, ultimately reaching an average of 0.9055. Simultaneously, the HS-VIT model can reduce the bit error rate to an average of 10. -11 Of the other three network models, the original VIT model performed best, which fully demonstrates the advantages of its basic structure.
[0017] Under strong turbulent conditions, the HS-VIT model achieved performance improvements of 20.9%, 23.8%, and 50.9% compared to the VIT, CNN, and ResNet models, respectively. Furthermore, tests showed that the correction times for HS-VIT, VIT, CNN, and ResNet were 0.32ms, 0.27ms, 5ms, and 0.7ms, respectively. Compared to the CNN and ResNet models, HS-VIT achieved better correction results with less correction time, demonstrating stronger real-time performance. It should be noted that although the HS-VIT model, an improvement on the existing VIT model, increased the correction time, the 20.9% performance improvement achieved at the cost of only 0.05ms is entirely within acceptable limits. In conclusion, compared to the other three network models, the HS-VIT model exhibits stronger correction accuracy and real-time performance in correcting wavefront aberrations under complex turbulent conditions. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the structure of a CFSOC system with a sensorless AO system, which is a prior art technology.
[0019] Figure 2 This is a flowchart of the operation of the HS-VIT model in the wavefront controller.
[0020] Figure 3 This is a block diagram of the Transformer block and Transformer_ESAC block in the Transformer encoder layer.
[0021] Figure 4 This is a block diagram of the ESAC structure.
[0022] Figure 5 This is a block diagram of the SPCSA structure.
[0023] Figure 6 This is a structural design diagram for a 32-unit adaptive DM.
[0024] Figure 7 This is a schematic diagram of the Zernike coefficients generated under three turbulent conditions.
[0025] Figure 8 Comparison of wavefront phase planes before and after correction using the HS-VIT model under three turbulent conditions.
[0026] Figure 9 Comparison of PSF images before and after correction using the HS-VIT model under three turbulent conditions.
[0027] Figure 10Box plots show the changes in system parameters before and after 100 simulations using the HS-VIT model under three turbulent conditions.
[0028] Figure 11 The graph shows a comparison of the calibration results of four network models under three turbulent conditions.
[0029] Figure 12 The diagram shows the calibration effect of HS-VIT on the experimental platform.
[0030] Figure 13 This is a comparison graph of the ME of the system before and after correction under strong turbulence conditions in 20 sets of experiments.
[0031] Figure 14 This is a comparison graph of the ME of the system before and after correction under moderate turbulence conditions in 20 sets of experiments.
[0032] Figure 15 This is a comparison graph of the ME of the system before and after correction under weak turbulence conditions in 20 sets of experiments. Detailed Implementation
[0033] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings and preferred embodiments.
[0034] In the existing CFSOC system, such as Figure 1 As shown, the optical signal emitted by the laser source at the transmitting end is amplified by a modulator and an optical amplifier before being transmitted through the atmospheric channel. During transmission, the optical signal is affected by atmospheric turbulence, which causes changes in the phase and amplitude of the optical signal, resulting in wavefront distortion. To compensate for the effects of atmospheric turbulence, a sensorless AO system is introduced at the receiving end to correct the wavefront distortion. The sensorless AO system mainly consists of a high-speed camera, a wavefront controller, and a wavefront corrector (usually a deformable mirror, DM). The sensorless AO system corrects the distorted laser signal through the principle of interference. The core function of the high-speed camera is to acquire the wavefront pre-image (PFS) of the wavefront, which serves as the input to the wavefront controller. The wavefront controller runs the HS-VIT model to identify the wavefront features carried in the PFS image and generates a control voltage signal for the wavefront corrector based on these features. This allows the wavefront corrector to correct the wavefront distortion, resulting in the corrected optical signal. The corrected optical signal is combined with the local oscillator beam, and then the mixed laser beam is transmitted to the demodulator for demodulation after passing through a mirror, lens, photodetector, and optical amplifier. Then, the optical signal is converted into a digital signal in the digital signal processor, and finally the digital signal is converted into data output.
[0035] The effectiveness of the CFSOC system is evaluated using three metrics: mixing efficiency (ME), root mean square value (RMS), and bit error rate (BER). These metrics are of significant guiding importance for the measurement and analysis of network models and the determination of performance fitness in sensorless AO systems.
[0036] 1) Mixing efficiency (ME):
[0037] Because the optical signal is affected by factors such as atmospheric turbulence before being mixed with the local oscillator (LO), the wavelength difference (ME) becomes one of the important evaluation indicators for CFSOC systems. It is numerically quantified as the Strehl ratio (SR) of the far-field focal plane image. The ME change of a CFSOC system with homodyne detection before and after phase distortion compensation can be analyzed by calculating the SR of the wavefront PSF image.
[0038] In a CFSOC system using zero-difference detection, assuming the LO light is a plane wave, the optical signal intensity is constant, and the ME can be expressed as:
[0039] (1)
[0040] In formula (1), the variable and These represent the amplitudes of the optical signal and the LO light, respectively. ,
[0041] and These represent the wavefront phases of the optical signal and the LO light, respectively; U represents the incident field of the photodetector.
[0042] 2) Root Mean Square (RMS):
[0043] The root mean square (RMS) value is a mathematical measure of the square root of the average of the mean of the input wavefront phase aberrations and the square root of the deviations. The RMS value can be expressed by the following mathematical formula:
[0044] (2)
[0045] in This represents the variance of the phase difference.
[0046] 3) Bit Error Rate (BER)
[0047] Bit error rate (BER) is an important indicator for measuring the communication performance of a CFSOC system. Under atmospheric turbulence interference conditions, the mathematical expression for the BER of a binary phase shift keying (BPSK) receiver system is:
[0048] (3)
[0049] in This indicates the quantum efficiency of the detector. ME represents the number of photons received per bit, and ME represents the mixing efficiency of the system. This represents the complementary error function.
[0050] In this embodiment, the specific operation flow of the HS-VIT model in the wavefront controller is as follows: Figure 2 As shown, the PSF images acquired by the high-speed camera are first preprocessed: they are downsampled to 224×224 to fit the structure of VIT, and then labeled. During the labeling process, each image is divided into fixed 16×16 square image blocks, resulting in a total of 196 image blocks.
[0051] Each image patch is flattened and transformed into a one-dimensional feature vector through a linear layer, then projected onto a 768-dimensional embedding space, enabling the network to extract meaningful input data. Next, a learnable classification token is added to the branch of each token sequence, representing global information for all image patches.
[0052] Furthermore, since the images obtained from image patches lack positional information, positional encodings are incorporated into the block embedding process to ensure that the network can interpret the relative spatial relationships between image patches. Position-encoded tokens, along with learnable special tokens for classification, are fed into the Transformer encoder layer, which consists of 12 layers of Transformer blocks.
[0053] like Figure 3As shown, each Transformer block primarily consists of a multi-head attention mechanism, which enables the model to capture global dependencies among tokens; and a multi-layer perceptron (MLP) feedforward network, which performs token-by-token nonlinear transformations on the attention-aggregated features, thereby enhancing the model's expressive power and feature abstraction capabilities. The remaining residual connections and layer normalization ensure efficient and stable training of the entire network. While the standard Transformer block performs well in global interactions due to its reliance on self-attention, this mechanism is relatively weak in responding to local features, making it particularly vulnerable to wavefront aberrations arising under strong turbulent conditions.
[0054] To address this, a hierarchical spatial enhancement mechanism is proposed. Preferably, Emulating Self-Attention with Convolution (ESAC) is incorporated into the Transformer block structure at layers 3, 6, and 9 to effectively enhance the network's robustness in recognizing wavefront aberration features under strong turbulence conditions. This embodiment aims to inject spatial enhancement at key stages of feature evolution. Therefore, at layer 3 (shallow end): before local detail information is over-smoothed or lost, ESAC is used for the first time to enhance its spatial awareness, providing more solid underlying features for subsequent semantic abstraction. At layer 6 (middle core): when features are aggregated at the middle level, ESAC is used again to strengthen the spatial structural relationships of features, ensuring that key relative positional information is not lost during abstraction. At layer 9 (deep entry): before features enter the final high-level semantic integration, a final spatial refinement is performed to ensure that the features subsequently passed to SPCSA and the regression head are spatially clear and semantically explicit.
[0055] This design enables the model to repeatedly apply spatial feature enhancement at different depth stages of the network, effectively compensating for the inherent defects of the original Transformer block architecture in local feature extraction.
[0056] Simultaneously, a Sparse Prompt Channel self-attention (SPCSA) technique is added to the end of the Transformer encoder layer to replace the feature aggregation method of the existing Vision Transformer model. This refines and aggregates the high-level features enhanced by the Transformer encoder layer, retaining only the most critical attention connections, enabling the HS-VIT model to achieve efficient global interaction. Finally, a regression head is used to enhance the nonlinear representation capability of the HS-VIT model, enabling it to learn the complex mapping relationship from the visual features of the PSF image to the control voltage of the wavefront corrector.
[0057] Under strong turbulent conditions, the PSF image of the wavefront exhibits highly localized speckle characteristics, requiring the network to have the ability to sensitively capture local details in order to achieve accurate estimation of wavefront aberrations. To this end, ESAC effectively alleviates the excessive bias of the VIT architecture in the process of modeling wavefront aberrations in strong turbulent conditions towards global dependencies.
[0058] like Figure 4 As shown, 196 image tokens after the first residual connection are extracted in the Transformer block and reshaped into a 14×14 spatial grid to form the input of ESAC. The ESAC divides the input into straight-through branches along the channel. and the input of the attention branch Preferably, to reduce memory access and operator overhead, subsequent operations can be applied only to the attention branches of the first 16 channels.
[0059] First of all Global average pooling is performed, followed by 1×1 dimensionality reduction convolution, and then Gaussian error linear unit activation function and 1×1 dimensionality increase convolution to obtain a dynamic convolution kernel of fixed size. :
[0060] (4)
[0061] in, DK Represents a dynamic convolution kernel; GAP represents global average pooling. This represents a 1×1 dimensionality-reduced convolution; This represents the calculation of the activation function of the Gaussian error linear unit; This represents a 1×1 up-dimensional convolution. The dynamic convolution kernel generates weight parameters in real time based on the input features, achieving content-adaptive local feature extraction.
[0062] Unlike previous cross-layer sharing proposals, this paper adopts a layer-independent large kernel design based on the characteristics of different network layers. The model employs a 13×13 convolutional kernel size, establishing spatial dependencies between the 16 input channels and 16 output channels. Each ESAC block independently learns the parameters of the large convolutional kernel, allowing ESACs in different layers to learn convolutional kernels suitable for the feature representation of that layer. This design enables the model to optimize its spatial augmentation based on the unique properties of features at each layer. Compared to the original cross-layer sharing mechanism, this method allows it to learn the optimal convolutional kernel adapted to the spatial features of its own layer.
[0063] Will LK and DK respectively with Parallel stacking of convolutions yields convolutional features. :
[0064] (5)
[0065] in, DK Represents a dynamic convolution kernel; LK Indicates a large convolution kernel; This represents the input to the attention branch; This represents the convolution operator; This represents the convolutional features after stacking.
[0066] Finally, Features are concatenated with the attention branch, then fused using a 1×1 convolution to obtain multi-scale features. As the output of ESAC:
[0067] (6)
[0068] in, This represents the convolutional features after stacking; This indicates the input to the direct branch; Concat indicates the feature concatenation operation. Represents feature fusion convolution; This represents the final multi-scale features obtained.
[0069] In summary, ESAC incorporates the local sensitivity of convolution operations into the Transformer encoder structure. By introducing a 3×3 dynamic weighted convolution kernel and a 13×13 large convolution kernel, it significantly enhances the ability to extract local features while retaining the advantages of VIT global modeling. This enables the model to capture local details in wavefront PSF images more accurately, ultimately achieving more refined wavefront aberration correction under strong turbulence conditions.
[0070] Typically, VIT relies on a single classification token or simple global pooling for feature aggregation at the end of the encoder layer. This approach has significant shortcomings in regression tasks: on the one hand, the classification token may not be able to fully capture the key information of all spatial locations; on the other hand, global average pooling treats all features equally, lacking targeted attention to important spatial regions. Therefore, a self-modulated feature aggregation method is used after the Transformer encoder layer to replace the original feature aggregation method. This method dynamically identifies and enhances spatial features through SPCSA, achieving more accurate high-level feature integration.
[0071] The structure of the SPCSA is as follows Figure 5 As shown, the image token portion of the feature tensor after processing by the Transformer encoder layer is first extracted and reshaped into a spatial format to obtain high-level features. As input to SPCSA, this high-level feature aggregates the rich spatial semantic information learned by all Transformer encoder layers. 1×1 convolutions (Conv) and 3×3 deep convolutions (DWConv) are applied to aggregate the high-level features across channels, generating query (Q), key (K), and value (V) matrices, as shown in equations (7) to (9):
[0072] (7)
[0073] (8)
[0074] (9)
[0075] In the above formula, AF Represents high-level features of the input; Represents a 1×1 convolution; This represents a 3×3 depthwise convolution; Q, K, and V represent the query, key, and value matrices, respectively.
[0076] Reshape Q, K, V into a multi-head attention format The number of attention heads (h) is 8, the number of channels (d) per attention head is 96, and the total number of spatial locations (HW) is 196. A dense attention matrix is generated by performing a dot product operation on Q and K. To mitigate information redundancy and noise interference in dense attention matrices, SPCSA employs a self-modulated Top-K sparsity strategy. Specifically, this strategy uses the EPGO algorithm to dynamically determine the number of attention connections required for each sample, retaining only the K most critical spatial dependencies and filtering out irrelevant interactions that may impair feature quality.
[0077] Table 1 shows the pseudocode for the EPGO algorithm:
[0078] Table 1
[0079]
[0080] Finally, the output is obtained through a weighted attention matrix, as shown in formula (10). Specifically, K values are retained, and the remaining values are all set to negative infinity. When the softmax function is executed, these unimportant values are set to 0. Then... Remodeling As the final output.
[0081] (10)
[0082] Where V represents the value matrix; M represents the dense attention matrix; Top-K represents the Top-K selection operation modulated by EPGO; Softmax is a normalization function that maps each element of a vector to an exponential proportional probability, enabling the network output to represent the class probability distribution; out represents the final feature output.
[0083] In summary, SPCSA significantly improves the feature aggregation quality of the VIT architecture. By employing a self-modulated Top-K sparse strategy, it ensures accurate focusing on key spatial regions, providing a cleaner and more discriminative feature representation for the subsequent regression head. This effectively enhances the regression accuracy and generalization ability of the model.
[0084] During the training of the HS-VIT model, the learning rate is set to 10. -4 The batch size is set to 128. The learning rate determines the step size of parameter updates and training stability, while the batch size affects the accuracy of gradient estimation and the model's generalization ability. Both work together in the optimization process, having a crucial impact on the model's convergence speed, stability, and final performance.
[0085] To verify the application effect of the HS-VIT model in the Sensorless AO system, this embodiment uses the MATLAB software platform for simulation analysis.
[0086] A continuous curved surface mirror (DM) is used as the wavefront corrector, and wavefront correction is achieved by controlling the voltage to change the shape of the DM. The HS-VIT model is used as the control algorithm to change the voltage of each actuator in the DM. Assuming the initial wavefront aberration of the laser after transmission through the atmospheric channel is denoted as , the HS-VIT model yields a vector Y, which represents the control voltages of each actuator in the DM mirror. The algorithm continuously updates the solution and generates a compensated phase. The remaining phase aberration can be obtained by the difference between the initial phase and the compensated phase. In the simulation, it is assumed that the laser wavelength is , the ratio of the optical lens diameter to the focal length is 1, the number of photons per qubit is 12, the Airy mode radius is , the detector quantum efficiency is 1, and a 32-element adaptive DM is used as the wavefront corrector. Its structural design diagram is shown below. Figure 6 As shown, in this 32-unit adaptive DM, the crosslink value of the actuator is 0.2, the normalized distance coefficient between the actuators is 0.392, and the Gaussian function... The value is 2, where the initial voltage of each driver is set to 0.
[0087] Create a dataset:
[0088] Parameters are typically used. To quantify the intensity of atmospheric turbulence, where D represents the aperture size of the receiving system. Fried parameters characterize atmospheric coherence. To further explore the influence of atmospheric turbulence on wavefront aberrations, the parameters are adjusted... The values are used to simulate Zernike polynomial aberrations of different orders. Zernike polynomials decompose the phase of a distorted wavefront into a sum of weighted orthogonal polynomials, each representing a specific aberration. Normalized atmospheric turbulence intensity is typically classified into three main levels: under weak turbulence conditions, Approximately 2; approximately 10 under moderate turbulence conditions; and under strong turbulence conditions, Greater than 15.
[0089] To evaluate the correction performance of the HS-VIT model under different turbulent conditions, this embodiment will... The values were set to 5, 10, and 20 respectively to simulate performance under weak, moderate, and strong turbulence conditions. Wavefront phase It can be represented as:
[0090] (11)
[0091] in, For piston term coefficients, corresponding to the constant term of the Zernike polynomial, it represents the overall translation of the wavefront in the axial direction, without changing the relative shape of the wavefront, but only causing the overall phase of the wavefront to shift by a constant value. For the first The coefficients of a Zernike polynomial mode reflect the weight of the corresponding mode in the wavefront phase distribution. The larger the coefficient, the more significant the wavefront distortion contribution of the corresponding mode. For the first A Zernike polynomial is a system of orthogonal polynomials defined on the unit circle (after normalization), with different... Corresponding to different wavefront distortion modes, by combining different and coefficient It can fit complex wavefront phase distributions.
[0092] Zernike coefficients represent the magnitude of wavefront aberrations; larger coefficients indicate greater aberrations. In the Zernike polynomial, the first three terms represent piston aberrations and tilt along the X and Y directions, respectively. These aberrations can be directly corrected by a beam steering unit (BSU). In actual communication, the atmospheric transmission environment is highly complex; using Zernike polynomials of orders 4-36 can more realistically simulate wavefront aberrations. In the simulation analysis, based on Roddier's method, from... Initial Zernike coefficients were randomly generated, and higher-order aberrations were constructed using Zernike polynomials from the 4th to the 36th term. Figure 7 This represents three sets of Zernike polynomials randomly generated under conditions of strong turbulence, moderate turbulence, and weak turbulence.
[0093] Simulation verification:
[0094] according to Figure 7 The original wavefront phase plane and point spread function (PSF) plots of the 4th to 36th modes of the Zernike polynomial under strong, moderate, and weak turbulence conditions were obtained, as shown in the figure. Figure 8 and Figure 9 As shown in (a), (b), and (c), the algorithm corrected the wavefront under three different turbulent conditions. The phase plane and PSF plot of the residual wavefront aberration after correction are shown in Figure 1. Figure 8 and Figure 9 As shown in (d), (e) and (f).
[0095] like Figure 8 As shown, the phase distribution of the residual wavefront decreases significantly after correction using the HS-VIT model. This means that a considerable proportion of wavefront aberrations can be successfully corrected.
[0096] PSFs can visualize abstract wavefront aberrations, and their shape, energy distribution, and scattered spot size directly reflect the magnitude of the wavefront aberrations. The SR of the PSF image captured by the high-speed camera is calculated, and then the ME is further obtained according to Eq.(1). The wavefront aberration correction effect is characterized by using PSFs and MEs.
[0097] like Figure 9 As shown in (a), (b) and (c), the initial ME values of the system under the three turbulent conditions are 0.0002, 0.0416 and 0.5009, respectively. Figure 9 Figures (d), (e), and (f) show that after HS-VIT correction, the wavefront shape becomes more refined and the energy more concentrated. Simultaneously, the ME values increase to 0.9464, 0.9810, and 0.9976, respectively. The changes in the PSF plots before and after correction demonstrate that HS-VIT has the ability to correct wavefront aberrations under different turbulent conditions, and most wavefront aberrations are compensated after HS-VIT correction. The correction accuracy of the HS-VIT model gradually improves as the turbulence intensity decreases. In conclusion, the Sensorless AO system based on the HS-VIT model can effectively suppress wavefront distortion caused by complex atmospheric turbulence.
[0098] Furthermore, considering the randomness of the method, respectively in Figure 7 The simulations were independently repeated 100 times under the three turbulent conditions shown. Figure 10 This indicates the changes in ME, RMS, and BER of the system before and after correction.
[0099] Box plots can be used to describe the dispersion of data, and the correction effect of the HS-VIT model on wavefront aberrations can also be visually reflected through box plots. Among them, the interquartile range (IQR) is used to measure the dispersion of the middle 50% of the data in the box plot, and outliers are data that are outside the range of 1.5 times the IQR.
[0100] like Figure 10 As shown in (a), the HS-VIT model effectively reduces wavefront aberrations under all three turbulent conditions. Under strong turbulence, the RMS of the CFSOC system before correction was 3.6158, which was reduced to an average of 0.2978 after correction using the HS-VIT model; under moderate turbulence, the RMS before correction was 1.8125, which was reduced to an average of 0.1769 after correction; and under weak turbulence, the RMS before correction was 0.8236, which was reduced to an average of 0.0470 after correction.
[0101] like Figure 10 As shown in (b), the average ME value of the system can be increased to about 0.9 or higher under all three turbulent conditions. Under strong turbulence, the average ME value increases to 0.9093; under moderate and weak turbulence, the average ME values increase to 0.9666 and 0.9977, respectively.
[0102] In typical communication scenarios, the system BER is usually required to be below 10. -6 In special scenarios such as satellite communication links, the BER requirement needs to be as low as 10. -9 .like Figure 10As shown in (c), under strong turbulent conditions, the system BER decreases from an average of 0.4948 to approximately 10. -11 Under moderate and weak turbulent conditions, the BER decreased from an average of 0.0905 and 10, respectively. -7 All averaged down to about 10 -12 .
[0103] from Figure 10 As can be seen, the RMS, ME, and BER of the system were improved under all three turbulent conditions, indicating that the HS-VIT model can exhibit good correction accuracy and adaptability under complex turbulent conditions. With the increase of turbulence intensity, the distribution of wavefront aberrations becomes irregular and complex, thus the range of IQR gradually increases. This further tests the correction capability of the HS-VIT model. The occurrence of outliers indicates that the HS-VIT model also experiences extremely rare anomalies in the aberration correction process, but its overall performance is stable. Overall, the HS-VIT model shows stable and significant correction effects in multiple simulations, demonstrating high stability and robustness. Applying the HS-VIT model to a sensorless AO system can effectively correct wavefront aberrations of optical signals under complex turbulent conditions. The sensorless AO system based on the HS-VIT model can significantly reduce the bit error rate of the CFSOC system and improve communication performance.
[0104] To verify the ability of the HS-VIT model-based sensorless AO system to correct actual aberrations, wavefront aberration data was collected on the experimental platform for subsequent testing. During the experiment, the Greenwood frequency was changed by adjusting the rotation speed of the phase screen. The laser beam was collimated by the phase screen and lens before being corrected by the sensorless AO system. A wavelength division multiplexing (WDM) receiver split the beam into two parts: one part was used for imaging, and the other part was used to measure wavefront aberration data.
[0105] In respectively Experiments were conducted under turbulent conditions of 5, 10, and 20. During the experiments, the computer first received PSF images acquired from a high-speed camera, and then used the HS-VIT model to generate control voltages for the DM to perform wavefront aberration correction. Figure 12 The results show that the HS-VIT model can effectively correct wavefront distortion.
[0106] In this embodiment, 20 sets of experiments were conducted under three different turbulent conditions using the HS-VIT model. The changes in ME values before and after correction are shown below. Figures 13-15As shown in the figure. Experimental results show that under strong turbulence conditions, the ME value can be improved to above 0.8 in each experimental group. Under strong turbulence conditions, the average ME value before correction is 0.00061, and the average ME value after correction is 0.9073. Under moderate and weak turbulence conditions, the ME value can be improved to above 0.9 in each experimental group. Under moderate turbulence conditions, the average ME value before correction is 0.0373, and the average ME value after correction is 0.9714; under weak turbulence conditions, the average ME value before correction is 0.5062, and the average ME value after correction is 0.9971. Experimental data indicate that the HS-VIT model provides an effective and more robust method for wavefront aberration correction under complex turbulence conditions. Applying the HS-VIT model to a sensorless AO system can significantly improve the system's communication performance.
[0107] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A coherent free-space optical communication system employing the HS-VIT model, comprising a wavefront-free adaptive optics system for correcting wavefront distortion, wherein the wavefront-free adaptive optics system includes: A high-speed camera is used to acquire wavefront PFS images and transmit the preprocessed PFS images to the wavefront controller. The wavefront controller is used to identify the wavefront features carried by the PFS image and generate the control voltage signal for the wavefront corrector based on the wavefront features. The wavefront corrector receives the control voltage signal and performs wavefront distortion correction, then outputs the corrected optical signal to the beam splitter. The wavefront controller is characterized by identifying wavefront features carried by PFS images by internally constructing an HS-VIT model. The HS-VIT model establishes a hierarchical spatial enhancement mechanism by combining convolutional simulation self-attention on the basic architecture of the VIT model, and applies it to the Transformer encoder layer in the VIT model. Furthermore, it uses a self-modulation feature aggregation method to process the enhanced features at the end of the Transformer encoder layer through sparse cue channel self-attention, thereby intelligently distinguishing key information and background noise in turbulent wavefront features. The operation process of the HS-VIT model in the wavefront controller includes: the preprocessed PFS image is labeled to generate multiple image patches. Each image patch is flattened and transformed into a one-dimensional feature vector through a linear layer, then projected onto a 768-dimensional embedding space. Learnable special tokens for classification are added to the branches of each token sequence to represent global information of all image patches, which are then input into a Transformer encoder layer consisting of 12 layers of Transformer blocks. Sparse cue channels with self-attention are added at the end of the Transformer encoder layer. Finally, a regression head is used to enhance the nonlinear representation capability of the HS-VIT model, enabling it to learn the complex mapping relationship from the visual features of the PFS image to the control voltage of the wavefront corrector. Each of the 12 Transformer blocks in the Transformer encoder layer contains a multi-head attention mechanism and a multilayer perceptron feedforward network, as well as residual connections and layer normalization; convolutional simulations of self-attention are added to the Transformer block structure in layers 3, 6, and 9.
2. The coherent free-space optical communication system using the HS-VIT model according to claim 1, characterized in that, The convolution simulates self-attention, which segments the input into direct branches along the channels. and the input of the attention branch First of all Global average pooling is performed, followed by 1×1 dimensionality reduction convolution, and then Gaussian error linear unit activation function and 1×1 dimensionality increase convolution to obtain a dynamic convolution kernel of fixed size. : , in, DK Represents a dynamic convolution kernel; GAP represents global average pooling. This represents a 1×1 dimensionality-reduced convolution; This represents the calculation of the activation function of the Gaussian error linear unit; This represents a 1×1 up-dimensional convolution; the dynamic convolution kernel generates weight parameters in real time based on the input features, achieving content-adaptive local feature extraction; a large 13×13 convolution kernel is used. ,Will LK and DK respectively with Parallel stacking of convolutions yields convolutional features. : , in, DK Represents a dynamic convolution kernel; LK Represents a large convolution kernel; This represents the input to the attention branch; This represents the convolution operator; This represents the convolutional features after stacking; Finally, Features are concatenated with the attention branch, then fused using a 1×1 convolution to obtain multi-scale features. As the output of convolution simulating self-attention: , in, This represents the convolutional features after stacking; This indicates the input to the direct branch; Concat indicates the feature concatenation operation. Represents feature fusion convolution; This represents the final multi-scale features obtained.
3. The coherent free-space optical communication system employing the HS-VIT model according to claim 1, characterized in that, The sparse cue channel self-attention structure applies 1×1 convolutions and 3×3 depthwise convolutions to aggregate high-level features across channels, generating a query, key, and value matrix: , , , in, AF This represents extracting the image token portion of the feature tensor after processing by the Transformer encoder layer and reshaping it into a spatial format as input high-level features; Represents a 1×1 convolution; This represents a 3×3 depthwise convolution; Q , K and V Represent the query, key, and value matrices respectively; Q , K , V Remodeling into a multi-head attention format The number of attention heads, h, is 8; the number of channels per attention head, d, is 96; and the total number of spatial locations, HW, is 196. Q and K Perform dot product operations to generate a dense attention matrix. Meanwhile, the sparse cue channel self-attention employs a self-modulated Top-K sparsity strategy, dynamically determining the number of attention connections required for each sample using the EPGO algorithm, and finally obtaining the output through a weighted attention matrix. , in, V The value matrix is represented by M; the dense attention matrix is represented by Top-K; Top-K represents the Top-K selection operation modulated by EPGO; Softmax is a normalization function that maps each element of a vector to an exponentially proportional probability, enabling the network output to represent the class probability distribution. out This represents the final feature output.
4. The coherent free-space optical communication system employing the HS-VIT model according to claim 3, characterized in that, The EPGO algorithm first performs layer normalization on the input dense attention matrix to obtain features; It will be projected into a hidden space through the first fully connected layer; The ReLU activation function is used to enhance the ability to represent nonlinear features. The second fully connected layer maps the features back to the target dimension, generating preliminary cue features; The Sigmoid activation function is used to process features, compressing all values into the range [0,1]. Flatten the features into a one-dimensional vector; Calculate the average value p of all elements in the one-dimensional vector, where p represents the proportion of information that the network believes the current image needs to retain; Multiply p by the total number of spatial locations of the original high-level features, 196, to obtain the dynamic K value; Returns a dynamic K value, retaining only the K most critical spatial dependencies and filtering out irrelevant interactions that may impair feature quality.
Citation Information
Patent Citations
High-capacity optical communication adaptive damage compensation and demodulation system
CN120150821A
Multi-mode optical fiber imaging method based on dual-encoder network and attention mechanism
CN121074176A