Frequency spectrum processing and deep learning combined underwater orbital angular momentum pattern recognition method

Through a combination of spectrum processing and deep learning, semi-spectral search, energy spectrum optimization, LSTM and improved CapsNet extract the spatiotemporal characteristics of underwater OAM signals, solving the problem of insufficient recognition accuracy and robustness in complex underwater environments in traditional methods, and achieving high-precision OAM pattern recognition.

CN120508806APending Publication Date: 2025-08-19GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510614050.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional OAM pattern recognition methods are difficult to achieve high-precision recognition in complex underwater environments. They are affected by multipath effect, noise interference and environmental dynamic changes, resulting in a decrease in pattern recognition accuracy and robustness.

Method used

Combining spectrum processing and deep learning, the purification signal is performed through semi-spectral search and energy spectrum optimization, spatial and temporal features are extracted using LSTM and improved CapsNet, combined with spatial and temporal cross attention mechanism and feature feedback and residual enhancement units, and finally pattern recognition is performed through SVM classifiers.

Benefits of technology

It significantly improves the accuracy and robustness of OAM pattern recognition, adapts to changes in different underwater environments, and improves the accuracy and stability of pattern recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508806A_ABST
    Figure CN120508806A_ABST
Patent Text Reader

Abstract

The invention discloses a spectrum processing and deep learning combined underwater orbital angular momentum pattern recognition method, which comprises the following steps of: firstly, processing a signal by utilizing half-spectrum search and energy spectrum optimization, and constructing two-dimensional OAM (Orbital Angular Momentum) multi-channel data; then, sending the OAM multichannel data into a deep learning model to extract OAM features; and finally, sending the OAM features into a pre-trained classifier, carrying out matching and classification on the features, and outputting corresponding OAM mode categories. Based on a deep learning and spectrum processing collaborative optimization scheme, the precision and robustness of OAM mode recognition can be remarkably improved, technical support is provided for underwater high-capacity communication and environmental perception, and the method is suitable for multiple fields of underwater high-capacity communication, underwater target detection and recognition, marine environment monitoring, underwater imaging and the like. The method has important theoretical significance and wide practical application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of underwater communication technology, and in particular to an underwater orbital angular momentum pattern recognition method combining spectrum processing and deep learning. Background Art

[0002] With the growing demand for marine resource development and environmental protection, underwater communications can support marine resource development, assist in seabed exploration and facility maintenance, and monitor environmental data in real time, detect pollution and issue early warnings in marine environmental protection. However, it faces challenges such as limited spectrum resources, severe noise interference and significant multipath effects. Orbital Angular Momentum (OAM) technology provides a new direction for underwater communications due to its unique orthogonality and infinite number of modes. OAM modes carry information through different spiral phase structures, and each mode corresponds to a unique spiral phase. Vortex beams carrying OAM modes can distinguish different modes through their topological charge. In theory, infinite orthogonal modes can be achieved, thereby significantly improving the capacity of communication systems. However, in practical applications, when OAM modes are transmitted underwater, they are extremely susceptible to phase distortion caused by multipath effects, signal spectrum susceptibility to noise interference, and signal distortion, resulting in a decrease in the accuracy and robustness of pattern recognition. Furthermore, the spatiotemporal dynamics of the underwater environment, such as phase and amplitude fluctuations caused by turbulence, and the temporal fluctuations of the beam phase due to temperature gradients and salinity changes, further complicate OAM pattern recognition. Traditional OAM pattern recognition models fail to account for the spatiotemporal characteristics of the signal and suffer from shortcomings in extracting spatial or temporal features, making it difficult to achieve high-precision pattern recognition in complex underwater environments. This has become a key bottleneck restricting the development of underwater communications. Summary of the Invention

[0003] The present invention aims to solve the problem that traditional OAM pattern recognition methods are difficult to achieve high-precision recognition in complex underwater environments, and provides an underwater orbital angular momentum pattern recognition method that combines spectrum processing and deep learning.

[0004] Compared with the prior art, the present invention has the following characteristics:

[0005] The underwater orbital angular momentum pattern recognition method combining spectrum processing and deep learning includes the following steps:

[0006] Step 1: Optical-to-electrical conversion of the underwater transmitted optical signal carrying OAM into a time-domain electrical signal, and Fourier transform of the time-domain electrical signal to obtain a frequency-domain signal;

[0007] Step 2: Perform a half-spectrum search on the frequency domain signal to obtain a purified spectrum signal, and perform energy spectrum optimization on the purified spectrum signal based on an energy spectrum optimization function of a minimum mean square error criterion to obtain an optimized spectrum signal;

[0008] Step 3: Convert the optimized spectrum signal back to the time domain through inverse Fourier transform to obtain the optimized time domain signal;

[0009] Step 4: Use the short-time Fourier transform method to obtain the time series data of the optimized time domain signal in the time dimension, and use the Canny edge detection method to obtain the spatial edge data of the optimized time domain signal in the spatial dimension. Then, use the time series data and spatial edge data as independent channels to construct two-dimensional OAM multi-channel data;

[0010] Step 5: Feed the OAM multi-channel data into the deep learning model to extract OAM features;

[0011] The deep learning model includes a spatiotemporal feature extraction unit, a spatiotemporal cross-attention feature fusion unit, and a feature feedback and residual enhancement unit; the spatiotemporal feature extraction unit serves as the input of the deep learning model, the output of the spatiotemporal feature extraction unit is connected to the input of the spatiotemporal cross-attention feature fusion unit, the output of the spatiotemporal cross-attention feature fusion unit is connected to the input of the feature feedback and residual enhancement unit, and the output of the feature feedback and residual enhancement unit serves as the output of the deep learning model;

[0012] First, the spatiotemporal feature extraction unit is used to extract temporal and spatial features from the OAM multi-channel data. The spatiotemporal feature extraction unit includes a parallel LSTM and an improved CapsNet. The LSTM extracts temporal features from the time series data in the OAM multi-channel data, while the improved CapsNet extracts spatial features from the spatial edge data in the OAM multi-channel data. Then, the spatiotemporal cross-attention feature fusion unit is used to adaptively weight the temporal and spatial features to obtain spatiotemporal features. Finally, the feature feedback and residual enhancement unit are used to enhance and optimize the spatiotemporal features to obtain the OAM features.

[0013] Step 6: The OAM features are fed into a pre-trained SVM classifier to match and classify the features and output the corresponding OAM mode category.

[0014] The above-mentioned improved CapsNet consists of a convolutional layer, a main capsule layer, a digital capsule layer, a regularization layer and two CA attention mechanism layers; the input of the convolutional layer is used as the input of the improved CapsNet, the output of the convolutional layer is connected to the input of the main capsule layer through the first CA attention mechanism layer, the output of the main capsule layer is connected to the input of the digital capsule layer through the second CA attention mechanism layer, the output of the digital capsule layer is connected to the input of the regularization layer, and the output of the regularization layer is used as the output of the improved CapsNet.

[0015] The above-mentioned spatiotemporal cross-attention feature fusion unit consists of two multi-head attention mechanism layers, an overlay layer and a normalization layer; the inputs of the two multi-head attention mechanism layers are jointly used as the input of the spatiotemporal cross-attention feature fusion unit, the outputs of the two multi-head attention mechanism layers are simultaneously connected to the input of the overlay layer, the output of the overlay layer is connected to the input of the normalization layer, and the output of the normalization layer is used as the output of the spatiotemporal cross-attention feature fusion unit.

[0016] The above-mentioned feature feedback and residual enhancement unit consists of a fully connected layer and a gating unit; the inputs of the fully connected layer and the gating unit are jointly used as the input of the feature feedback and residual enhancement unit, the output of the fully connected layer is connected to the input of the gating unit, and the output of the gating unit is used as the output of the feature feedback and residual enhancement unit.

[0017] Compared with the prior art, the present invention has the following characteristics:

[0018] 1. Preprocessing by combining half-spectrum search and energy spectrum optimization: Half-spectrum search technology accurately locates key spectral regions carrying useful information in OAM signals, effectively reducing invalid data processing while filtering out noise and interference, making the signal more reliable and pure. Energy spectrum optimization further adjusts the signal's energy distribution and enhances the energy concentration of useful features in the signal, making the signal features more prominent and clear, providing purer and more stable data for subsequent pattern recognition.

[0019] 2. Collaborative Modeling of LSTM and CapsNet: Efficient extraction of the spatiotemporal features of underwater OAM signals is achieved through the Long Short-Term Memory network (LSTM) and Capsule Network (CapsNet). The gating mechanism and memory units of the LSTM network ensure accurate tracking and analysis of signal time series changes. The Coordinate Attention Mechanism (CA) module introduced in CapsNet enhances CapsNet's sensitivity to spatial edge data, enabling it to accurately capture and identify the unique spatial structure of OAM patterns. The collaborative work of LSTM and CapsNet improves the recognition accuracy of spatial features while strengthening the processing capability of temporal features.

[0020] 3. Feature fusion using the spatiotemporal cross-attention mechanism: This mechanism introduces a spatiotemporal cross-attention mechanism consisting of two multi-head self-attention mechanisms in parallel to process the features of LSTM and CapsNet respectively. Through the dynamic interaction of temporal and spatial edge data, it achieves deep fusion of temporal and spatial edge data, thereby more comprehensively understanding and capturing the dynamic spatiotemporal dependencies during feature fusion, and improving recognition accuracy and stability.

[0021] 4. Dynamic feature feedback and enhancement: Dynamic feature adjustment and enhancement are achieved through fully connected layers, gating units, and residual connections. The fully connected layers integrate spatiotemporal features, and residual connections are used to enhance feature expression and alleviate the gradient vanishing problem. At the same time, the gating units dynamically adjust the information flow to adapt to different inputs, enhancing the model's adaptability and feature expression capabilities.

[0022] 5. Improved dynamic adaptability: The integration of dynamic feedback mechanisms and multi-dimensional features enables the deep learning model to perceive environmental changes and adjust accordingly, maintaining high recognition accuracy and stability in different underwater environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Schematic diagram of the underwater orbital angular momentum pattern recognition method combining spectrum processing and deep learning. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific examples and accompanying drawings. The present invention specifically includes the following steps:

[0025] See also Figure 1 The present invention proposes an underwater orbital angular momentum pattern recognition method that combines spectrum processing with deep learning. By integrating spectrum processing and deep learning models for collaborative optimization, a new solution is provided for OAM pattern recognition in complex underwater environments. The method specifically includes the following steps:

[0026] Step 1: The optical signal transmitted underwater and carrying OAM is converted into a time-domain electrical signal x(t) by photoelectric conversion, and the time-domain electrical signal x(t) is Fourier transformed to obtain a frequency-domain signal X(f), where the Fourier transform is:

[0027]

[0028] Step 2: Perform a half-spectrum search on the frequency domain signal X(f) to obtain the purified spectrum signal X′(f). Perform energy spectrum optimization on the purified spectrum signal X′(f) using an energy spectrum optimization function based on the minimum mean square error criterion to obtain the optimized spectrum signal X″(f).

[0029] 1) Half-spectrum search

[0030] In the frequency domain, multipath effects and mirror reflections can cause pseudo peaks to appear in the signal spectrum. Therefore, a frequency search range [f min ,f max ] and perform half-spectrum search with a certain frequency step Δf.

[0031] For each search frequency f i =f min +iΔf, calculate the power spectral density P(f) of the signal at this frequency by Welch method i ):

[0032]

[0033] Among them, X k (f i ) is the Fourier transform result after segmented windowing of the frequency domain signal X(f), and K is the number of segments.

[0034] By analyzing the power spectral density, the frequency bands where the pseudo peaks are located are identified, and then these frequency bands are suppressed using band-stop filters to purify the signal spectrum and reduce the impact of noise and interference on subsequent processing.

[0035] 2) Energy spectrum optimization

[0036] The energy spectrum optimization function based on the minimum mean square error (MMSE) criterion is:

[0037]

[0038] Where F represents the effective frequency range determined by the half-spectrum search, X′(f) is the purified spectrum signal, is the estimated ideal frequency domain signal.

[0039] By iteratively optimizing the optimization function, the weight of each frequency component is iteratively adjusted using the gradient descent method to optimize the frequency domain signal X″(f);

[0040]

[0041] Where α is the learning rate, is the gradient of the optimization function with respect to the frequency f.

[0042] Frequency domain processing effectively suppresses the spectrum distortion caused by the underwater environment, providing pure and stable data for signal recognition.

[0043] Step 3: Convert the optimized spectrum signal X″(f) back to the time domain through inverse Fourier transform to obtain the optimized time domain signal x′(t):

[0044]

[0045] Step 4: Use the short-time Fourier transform method to obtain the time series data of the optimized time domain signal in the time dimension, and use the Canny edge detection method to obtain the spatial edge data of the optimized time domain signal in the spatial dimension. Then, use the time series data and spatial edge data as independent channels to construct two-dimensional OAM multi-channel data.

[0046] On the one hand, the optimized time domain signal is analyzed using the Short-Time Fourier Transform (STFT) method to obtain the time series data of the OAM mode in the time dimension. STFT can be expressed as:

[0047]

[0048] Where x(n) is the discrete signal, w(n) is the window function, and m is the time index. Set the window function parameters and the STFT time step to accurately capture changes in the signal at different time points.

[0049] The Canny edge detection algorithm is then used to obtain the spatial profile of the optimized time-domain signal and the spatial edge data of the OAM mode in the spatial dimension. The Canny edge detection algorithm incorporates Gaussian smoothing, gradient calculation, non-maximum suppression, and double thresholding to generate a grayscale spatial distribution feature map.

[0050] Afterwards, the time series data and spatial edge data are used as independent channels to construct two-dimensional OAM multi-channel data.

[0051] Step 5: Feed the OAM multi-channel data into the deep learning model to extract OAM features.

[0052] The deep learning model includes a spatiotemporal feature extraction unit, a spatiotemporal cross-attention feature fusion unit, and a feature feedback and residual enhancement unit; the spatiotemporal feature extraction unit serves as the input of the deep learning model, the output of the spatiotemporal feature extraction unit is connected to the input of the spatiotemporal cross-attention feature fusion unit, the output of the spatiotemporal cross-attention feature fusion unit is connected to the input of the feature feedback and residual enhancement unit, and the output of the feature feedback and residual enhancement unit serves as the output of the deep learning model.

[0053] Step 5.1: Use the spatiotemporal feature extraction unit to extract the temporal and spatial features of the OAM multi-channel data.

[0054] The spatiotemporal feature extraction unit includes parallel LSTM and improved CapsNet. LSTM is used to extract time series data from OAM multi-channel data to obtain temporal features, and improved CapsNet is used to extract spatial edge data from OAM multi-channel data to obtain spatial features.

[0055] 1) Temporal feature extraction

[0056] LSTM is a time-recurrent neural network designed to address the long-term dependency issues of traditional recurrent neural networks (RNNs). Its core feature is that it dynamically controls the flow of information through gating mechanisms (forget gate, input gate, and output gate), effectively capturing long-term dependencies in sequential data. This paper uses traditional LSTM to extract time series data from multi-channel data.

[0057] The hidden layer dimension of LSTM is set to 256 to capture the time-varying characteristics of the signal. To match the spatial edge data dimension, the output dimension is adjusted from [T, 256] to [256, T] through 1×1 convolution, where T is the time step, to capture the time-varying characteristics of the OAM signal. LSTM has a memory unit inside, and its core gating mechanism is to use the input gate i t 、Forget Gate t , output gate o t and memory unit C t To achieve long-term and short-term processing of the signal, the calculation formula is:

[0058] f t =σ(W f ·h t-1 +U f x t )

[0059] i t =σ(W i ·h t-1 +U i x t )

[0060]

[0061] o t =σ(W o ·h t-1 +U o x t )

[0062] Where σ is the sigmoid activation function, ⊙ represents element-wise multiplication, W is the weight matrix, b is the bias vector, and x t is the input at the current moment, h t-1 It is the hidden state at the previous moment.

[0063] 2) Spatial feature extraction

[0064] CapsNet is a novel neural network architecture designed to address the challenges faced by traditional convolutional neural networks (CNNs) in processing spatial hierarchical relationships and pose variations in images. Its core concept is the introduction of "capsules," small groups of neurons. Each capsule is responsible for identifying a specific type of object in an image and its attributes, such as position, pose, and scale, to better capture the spatial relationships between objects in the image. This paper uses an improved CapsNet to extract spatial edge data from multi-channel data.

[0065] The improved CapsNet consists of a convolutional layer (Conv1), a primary capsule layer (PrimaryCaps), a digit capsule layer (DigitCaps), a regularization layer (L2), and two CA attention layers (CA). The input of the convolutional layer serves as the input of the improved CapsNet. The output of the convolutional layer is connected to the input of the primary capsule layer via the first CA attention layer. The output of the primary capsule layer is connected to the input of the digit capsule layer via the second CA attention layer. The output of the digit capsule layer is connected to the input of the regularization layer, and the output of the regularization layer serves as the output of the improved CapsNet. The improved CapsNet adds a CA attention layer to the CapsNet network. The CA attention mechanism is a lightweight attention mechanism for mobile networks that aims to enhance feature expression capabilities without increasing computational cost. Unlike the traditional channel attention mechanism, CA embeds position information in the channel attention, allowing for better spatial attention.

[0066] The bottom convolutional layer of the improved CapsNet uses a 3×3 convolution kernel and a stride of 2. A CA attention layer is added after the convolutional layer. The CA attention layer first embeds the coordinate information of the input features, performs horizontal-vertical global pooling on the features output by the convolutional layer, and generates direction-sensitive attention weights.

[0067] CA attention layer horizontal and vertical pooling and attention weight generation:

[0068]

[0069] Attention(h,w)=σ(Conv([z h ,z w ]))

[0070] Among them, z h (c) and z w (c) represent horizontal and vertical pooling features respectively, and σ is the Sigmoid activation function.

[0071] Features passed through the CA attention layer are fed into the main capsule layer to extract primary capsule features. A CA attention layer is then inserted between the main capsule layer and the digital capsule layer to perform coordinate attention recalibration on the primary capsule features output by the main capsule layer. This coordinate attention recalibration is performed on the capsule features before dynamic routing, and the optimized capsules are transferred through the dynamic routing algorithm to ensure the accuracy of high-order feature combinations. The digital capsule layer enables efficient extraction of spatial hierarchical features.

[0072] CapsNet dynamic routing mechanism is:

[0073]

[0074] b ij ←b ij +CA(u i )·v j

[0075] Among them, u i is the output vector of the capsule, s i is the weighted input vector and Squash is the activation function.

[0076] CA attention layer weighted primary capsule output u i , improve routing stability. In the dynamic routing protocol, the coupling coefficient update formula from capsule i to capsule j is:

[0077]

[0078] Among them, c ij represents the coupling coefficient from capsule i to capsule j, b ij is the initial routing log-likelihood.

[0079] The capsule dimension is set to 64, and temporal interpolation is used to adjust the spatial edge data dimension from [M, 64] to [T, 64], aligning it with the LSTM time series length, where M is the number of spatial positions. The improved CapsNet extracts the spatial phase features of the OAM mode, achieving robust modeling against spatial transformations such as mode rotation and scaling, improving the ability to identify spatial edge data in the signal.

[0080] Step 5.2: Use the spatiotemporal cross-attention feature fusion unit to adaptively weight the temporal features and spatial features to obtain spatiotemporal features.

[0081] After extracting the temporal and spatial features of the OAM pattern, a spatiotemporal cross-attention feature fusion unit is introduced to fuse the features. This unit consists of two multi-head attention layers, an overlay layer, and a normalization layer. The inputs of the two multi-head attention layers serve as the input of the spatiotemporal cross-attention feature fusion unit. The outputs of the two multi-head attention layers are connected to the input of the overlay layer, which in turn is connected to the input of the normalization layer. The output of the normalization layer serves as the output of the spatiotemporal cross-attention feature fusion unit.

[0082] Two multi-head attention mechanism layers are run in parallel, processing the temporal features of the LSTM output and the spatial features of the improved CapsNet output, respectively, to generate the spatiotemporal features and the QKV vectors of the spatiotemporal features. The multi-head attention mechanism layer maps the input features to the query, key, and value vectors, respectively, where the calculation is:

[0083]

[0084] MultiHead(Q,K,V)=Concat(head1,…,head h )W O

[0085] head i =Attention(QW i Q ,KW i K ,VW i V )

[0086] They are obtained by linear transformation of input features, i.e. Q = XW Q , K=XW K , V=XW V , where X is the input feature matrix, W Q 、W K 、W V is the weight matrix, d k is the dimension of the key vector.

[0087] The LSTM-generated Q is used to explore key information about the temporal evolution of OAM. K and V are derived from linear transformations of temporal features, representing a temporal information index and actual feature data at each time step, respectively, to facilitate the exploration of temporal dependencies and dynamic changes. CapsNet generates Q, focusing on key spatial features of orbital angular momentum patterns. K and V are derived from linear transformations of spatial features, representing spatial information index and actual features at each spatial location, respectively, to capture correlations between spatial edge features. Through the dynamic interaction of temporal and spatial features, the temporal and spatial features are mutually attentive, achieving deep fusion of temporal and spatial features. The two sets of QKV vectors, calculated using their respective multi-head attention mechanisms, are cross-fused to enhance the interaction and compactness of spatiotemporal features, improving the accuracy and stability of underwater OAM pattern recognition. Finally, layer normalization is applied to normalize the feature data to prevent vanishing and exploding gradients during propagation, ensuring stable model training, deep integration of spatiotemporal features, and capturing dynamic spatiotemporal dependencies.

[0088] Step 5.3: Use feature feedback and residual enhancement units to enhance and optimize the spatiotemporal features to obtain OAM features.

[0089] The feature feedback and residual enhancement unit consists of a fully connected layer and a gating unit. The inputs of the fully connected layer and the gating unit serve as the input of the feature feedback and residual enhancement unit. The output of the fully connected layer is connected to the input of the gating unit, and the output of the gating unit serves as the output of the feature feedback and residual enhancement unit. The fused spatiotemporal features are linearly transformed by the fully connected layer to achieve feature integration and dimensionality conversion, and are mapped to a feature space adapted to subsequent tasks. With the help of residual connections and gating units, the fusion weights of the feedback signal and the original input are dynamically adjusted. Residual connections allow the original input features to be directly superimposed with the features transformed by the fully connected layer, thereby alleviating the gradient vanishing problem and enabling the model to learn richer feature representations.

[0090] The gate unit is based on the hidden state H at the previous moment t-1 and the current input X t , by updating gate z t and reset gate r t and candidate hidden states H t Calculation, dynamically adjust the current hidden state This enhances the model's adaptability to different input features and tasks.

[0091] z t =σ(W z ·[H t-1 ,X t ]+b z )

[0092] r t=σ(W r ·[H t-1 ,X t ]+b r )

[0093]

[0094] Among them, W z 、W r 、 is the weight matrix, [H t-1 ,X t ] is the concatenation operation, and b is the bias term.

[0095] The gating unit dynamically adjusts the fusion weight of the feedback signal (candidate state) and the original input feature (current hidden state), so that the model can better adapt to different input features and task requirements, and further improve the feature expression capability.

[0096] Step 6: The OAM features are fed into a pre-trained SVM classifier to match and classify the features and output the corresponding OAM mode category.

[0097] The classifier uses the Support Vector Machine (SVM) classification algorithm. The optimization problem of SVM can be expressed as:

[0098]

[0099] Among them, w is the normal vector of the hyperplane, b is the bias term, and x i is the input feature vector, y i is the corresponding category label. By finding the complexity of the minimized hyperplane, Represents, ensuring that all training samples are correctly classified.

[0100] The performance of the present invention is verified by experiments below.

[0101] Experimental Preparation and Environmental Simulation: An experimental platform was constructed, including an orbital angular momentum (OAM) pattern generator (equipped with a spatial light modulator (SLM), an underwater transmission channel simulator, and a signal spectrum processing and identification device. The OAM pattern generator utilizes a spatial light modulator (SLM) to load a Laguerre-Gaussian (LG) beam phase hologram, generating OAM mode optical signals with topological charges of l = +1, l = +5, l = +10, l = ±1, l = ±5, and l = ±10. To simulate the effects of a real ocean environment on optical signal transmission, the underwater transmission channel simulator utilizes an ocean turbulence phase screen to construct a turbulent transmission channel. The generated LG beam is transmitted through this simulated channel, resulting in a distorted OAM mode optical signal that carries turbulence-induced phase distortion and amplitude fluctuations. The system was set within a range of 0.1-10 NTU to simulate environments with varying degrees of turbidity; the multipath delay was adjusted between 1-20 ms to simulate varying degrees of multipath effects; and the signal-to-noise ratio was controlled between 5-20 dB to test the system's performance under varying noise levels. The signal spectrum processing and recognition device was based on the Pycharm integrated development environment, an Intel Core i9-10920X CPU, 32GB of DDR4 memory, and an NVIDIA RTX6000 GPU (24GB of video memory). The software environment utilized the Windows 10 operating system, Python 3.6, and the deep learning framework Pytorch. It was also equipped with libraries such as numpy 1.19.5, scipy 1.5.4, and torchvision 0.8.2. The device was responsible for spectrum analysis, feature extraction, and pattern recognition of signals transmitted through a simulated underwater environment to verify the performance of the proposed method.

[0102] Comparative experiments were conducted. Signals collected but not processed and purified were fed into a single LSTM model and a single CapsNet model, respectively, as a control group. The data signals, processed after half-spectrum search and energy spectrum optimization, were fed into the improved CapsNet-LSTM fusion model of the present invention as an experimental group. Table 1 compares the OAM pattern recognition accuracy of each model under three different experimental conditions.

[0103] Table 1

[0104]

[0105] Experiments show that the recognition accuracy of the method of the present invention for OAM modes with different topological charges exceeds 95%, and is significantly better than the two single-model methods.

[0106] In summary, the collaborative optimization scheme based on deep learning and spectrum processing in this invention can significantly improve the accuracy and robustness of OAM pattern recognition, provide technical support for underwater high-capacity communication and environmental perception, and is suitable for multiple fields such as underwater high-capacity communication, underwater target detection and identification, marine environment monitoring, and underwater imaging. It has important theoretical significance and broad practical application prospects.

[0107] It should be noted that although the embodiments of the present invention described above are illustrative, they are not intended to limit the present invention. Therefore, the present invention is not limited to the above-mentioned specific embodiments. Without departing from the principles of the present invention, any other embodiments obtained by those skilled in the art under the guidance of the present invention are deemed to be within the protection of the present invention.

Claims

1. An underwater orbital angular momentum pattern recognition method combining spectrum processing and deep learning, characterized by: The steps are as follows: Step 1: Optical-to-electrical conversion of the underwater transmitted optical signal carrying OAM into a time-domain electrical signal, and Fourier transform of the time-domain electrical signal to obtain a frequency-domain signal; Step 2: Perform a half-spectrum search on the frequency domain signal to obtain a purified spectrum signal, and perform energy spectrum optimization on the purified spectrum signal based on an energy spectrum optimization function of a minimum mean square error criterion to obtain an optimized spectrum signal; Step 3: Convert the optimized spectrum signal back to the time domain through inverse Fourier transform to obtain the optimized time domain signal; Step 4: Use the short-time Fourier transform method to obtain the time series data of the optimized time domain signal in the time dimension, and use the Canny edge detection method to obtain the spatial edge data of the optimized time domain signal in the spatial dimension. Then, use the time series data and spatial edge data as independent channels to construct two-dimensional OAM multi-channel data; Step 5: Feed the OAM multi-channel data into the deep learning model to extract OAM features; The deep learning model includes a spatiotemporal feature extraction unit, a spatiotemporal cross-attention feature fusion unit, and a feature feedback and residual enhancement unit; the spatiotemporal feature extraction unit serves as the input of the deep learning model, the output of the spatiotemporal feature extraction unit is connected to the input of the spatiotemporal cross-attention feature fusion unit, the output of the spatiotemporal cross-attention feature fusion unit is connected to the input of the feature feedback and residual enhancement unit, and the output of the feature feedback and residual enhancement unit serves as the output of the deep learning model; First, the spatiotemporal feature extraction unit is used to extract temporal and spatial features from the OAM multi-channel data. The spatiotemporal feature extraction unit includes a parallel LSTM and an improved CapsNet. The LSTM extracts temporal features from the time series data in the OAM multi-channel data, while the improved CapsNet extracts spatial features from the spatial edge data in the OAM multi-channel data. Then, the spatiotemporal features are adaptively weighted and fused with the temporal and spatial features using the spatiotemporal cross attention feature fusion unit. Finally, the spatiotemporal features are enhanced and optimized using the feature feedback and residual enhancement unit to obtain the OAM features. Step 6: The OAM features are fed into a pre-trained SVM classifier to match and classify the features and output the corresponding OAM mode category.

2. The underwater orbital angular momentum pattern recognition method combining spectrum processing and deep learning according to claim 1 is characterized in that: The improved CapsNet consists of a convolutional layer, a main capsule layer, a digital capsule layer, a regularization layer and two CA attention mechanism layers; the input of the convolutional layer is used as the input of the improved CapsNet, the output of the convolutional layer is connected to the input of the main capsule layer through the first CA attention mechanism layer, the output of the main capsule layer is connected to the input of the digital capsule layer through the second CA attention mechanism layer, the output of the digital capsule layer is connected to the input of the regularization layer, and the output of the regularization layer is used as the output of the improved CapsNet.

3. The underwater orbital angular momentum pattern recognition method combining spectrum processing and deep learning according to claim 1 is characterized in that: The spatiotemporal cross-attention feature fusion unit consists of two multi-head attention mechanism layers, a superposition layer, and a normalization layer; The inputs of the two multi-head attention mechanism layers are jointly used as the input of the spatiotemporal cross attention feature fusion unit. The outputs of the two multi-head attention mechanism layers are simultaneously connected to the input of the superposition layer. The output of the superposition layer is connected to the input of the normalization layer. The output of the normalization layer is used as the output of the spatiotemporal cross attention feature fusion unit.

4. The underwater orbital angular momentum pattern recognition method combining spectrum processing and deep learning according to claim 1 is characterized in that feature feedback The residual enhancement unit consists of a fully connected layer and a gating unit; the inputs of the fully connected layer and the gating unit are jointly used as the inputs of the feature feedback and residual enhancement unit, the output of the fully connected layer is connected to the input of the gating unit, and the output of the gating unit is used as the output of the feature feedback and residual enhancement unit.

Citation Information

Cited By

  • Industrial bearing vibration time sequence signal fault prediction method and system fusing attention mechanism and LSTM

    CN121144702A