Enhanced OFDM-DCSK chaotic communication method fusing deep reinforcement learning
By integrating the enhanced OFDM-DCSK chaotic communication method with deep reinforcement learning, using TCN, self-attention and Transformer network for feature extraction and long-term dependency modeling, and introducing the DRQN optimized demodulation strategy, the problems of high computational complexity and difficulty in modeling long time series dependencies in the existing technology are solved, and efficient and robust chaotic encryption communication is achieved.
Patent Information
- Application Number
- CN202510798049.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-09
AI Technical Summary
Existing deep learning-based chaotic encryption methods without reference signals have high computational complexity, large training data requirements, and severe energy efficiency challenges in multi-user MIMO and time-varying fading channel scenarios. In addition, traditional TDNN networks have difficulty modeling long-term dependencies, RNNs have the problem of vanishing gradients, and the efficiency of signal spatiotemporal feature fusion is low, resulting in high bit error rates.
An enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning is adopted. Time series data features are extracted through the TCN module, and the self-attention mechanism and Transformer network are combined to capture long-term dependencies. The DRQN reinforcement learning mechanism is used to optimize the demodulation strategy to achieve multi-scale feature capture and dynamic adaptation.
It significantly improves the system performance under complex channel conditions, reduces computational complexity, improves demodulation accuracy and environmental adaptability, and maintains high-efficiency communication performance.
Smart Images

Figure CN120614591A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technology, and in particular to an enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning. Background Art
[0002] With the increasing demand for wireless communication security, chaotic communication, with its unique non-periodicity and initial value sensitivity, is demonstrating significant value in the field of secure information transmission. Traditional chaotic encryption methods rely primarily on complex modulation and demodulation techniques and signal processing algorithms, using mathematical transformations such as precoding and singular value decomposition to improve system performance. However, these methods generally suffer from inherent limitations such as high implementation complexity and strong dependence on channel state information. Particularly in rapidly time-varying channel environments, their bit error rate performance often fails to meet practical requirements.
[0003] In recent years, deep learning technology has made significant progress in the field of chaotic communications. In particular, neural network-based assisted demodulation methods offer new solutions for chaotic encryption systems that do not require a reference signal. These methods typically employ an end-to-end processing architecture, directly mapping the received encrypted signal to the decrypted data. In the feature extraction stage, specialized deep learning neural network structures are designed to extract key features from the time domain signal. In the time series modeling stage, time series models such as recurrent neural networks are used to capture the temporal correlation of the signal. Finally, in the classification decision stage, a fully connected network is used to achieve a nonlinear mapping of features to decrypted data.
[0004] Current research focuses on optimizing deep learning models for feature extraction and time series modeling. Existing deep learning-based chaotic encryption methods for reference-free signals utilize a TDNN (Time Delay Neural Network) combined with a bidirectional LSTM (Long Short-Term Memory) architecture. The TDNN extracts local time-domain features of the signal, while the bidirectional LSTM models sequential dependencies.
[0005] The disadvantages of the deep learning-based reference signal-free chaotic encryption method in the above-mentioned existing technology include: in multi-user MIMO and time-varying fading channel scenarios, the training process of the deep neural network involves high-dimensional parameter optimization, which leads to a sharp increase in computational complexity; secondly, model training requires a large amount of labeled data, which is often difficult to meet in actual communication systems; thirdly, the deep network structure faces severe energy efficiency challenges when deployed on edge devices.
[0006] Existing methods have inherent flaws in time-domain feature extraction: traditional TDNN networks struggle to model long-term temporal dependencies, and RNNs suffer from the vanishing gradient problem. This, coupled with the lack of effective feature interaction mechanisms, results in inefficient signal spatiotemporal feature fusion and high system bit error rates. However, these architectures still have shortcomings in long-term dependency modeling and multi-scale feature fusion, limiting their performance under complex channel conditions. Summary of the Invention
[0007] An embodiment of the present invention provides an enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning to achieve safer and more robust chaotic encryption communication.
[0008] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions.
[0009] An enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning, comprising:
[0010] At the transmitting end, the chaotic sequence and user data are chaotically encrypted to obtain a chaotic reference signal, which is then modulated onto an OFDM multi-subcarrier for transmission.
[0011] At the receiving end, the chaotic reference signal is demodulated by utilizing the autocorrelation and incoherent demodulation principles of the chaotic sequence, and the time series data features of the demodulated signal are extracted through the basic time series convolutional network (TCN) module;
[0012] With the goal of maximizing long-term rewards, the deep recurrent Q network DRQN reinforcement learning mechanism is used to continuously explore and learn the optimal decision-making strategy based on the characteristics of the time series data and environmental information.
[0013] Preferably, at the transmitting end, chaotically encrypting the chaotic sequence and user data to obtain a chaotic reference signal, modulating the chaotic reference signal onto an OFDM multi-subcarrier and transmitting it, comprises:
[0014] The chaotic sequence is generated using the second-order Chebyshev polynomial function, as shown in formula (1):
[0015]
[0016] Among them, x k represents the kth element of the chaotic sequence, x k-1 is the k-1th element of the chaotic sequence, and the initial value x0 is randomly selected in the interval [-1,1];
[0017] Generate binary phase shift keying BPSK modulated user data, the user data is expressed as: i ∈{-1,+1}, where i is the symbol index;
[0018] Multiply the user data with the chaotic sequence to obtain the modulation symbol s(t):
[0019] s(t)=b(t)·x(t) (2)
[0020] OFDM modulation is performed on the modulation symbol s(t), the modulation symbol s(t) is distributed to N subcarriers, a cyclic prefix CP is added to obtain an OFDM symbol, and the OFDM symbol is transmitted.
[0021] Preferably, at the receiving end, the chaotic reference signal is demodulated using the autocorrelation of the chaotic sequence and the incoherent demodulation principle, including:
[0022] At the receiving end, the CP of the received signal is removed and a fast Fourier transform is performed to obtain the frequency domain representation of the received signal. Through channel estimation and equalization, the equalized symbol is obtained. The expression is:
[0023]
[0024] in, is the channel estimation value, ∈ is a constant;
[0025] Construct a time series training dataset, in which the sample data values are equalized symbols The real part of :
[0026]
[0027] The label of the sample data is the value mapped from the user data to [0,1]:
[0028]
[0029] Preferably, the extracting time series data features of the demodulated signal by the TCN module includes:
[0030] A basic temporal convolutional network (TCN) module based on the convolutional neural network (CNN) is constructed. For the sample data in the temporal training dataset, a local TCN module is used to capture local features using a smaller convolution kernel and a smaller dilation factor. A global TCN module is used to capture global features using a larger convolution kernel and a larger dilation factor. Local features and global features are fused through the self-attention mechanism, and query Q, key K, and value V vectors are generated through linear transformation. The calculation formula is:
[0031] Q=Linear query (Concat(Local,Global)) (10)
[0032] K=Linearkey (Concat(Local,Global)) (11)
[0033] V=Linear value (Concat(Local,Global)) (12)
[0034] Among them, Linear query 、Linear key and Linear value It is a linear transformation function used to map the concatenated features to different spaces to generate query, key, and value vectors. Local is a local feature, and Global is a global feature.
[0035] The attention score is calculated by the dot product of Q and K and normalized by the Softmax() function. The formula is:
[0036]
[0037] Among them, d is the feature dimension, and the attention score represents the query vector Q i With the key vector K j The similarity of these scores is normalized into a probability distribution through the Softmax() function, so that each value vector V j The weight is between 0 and 1;
[0038] The V vector is weighted and summed using the attention weight to obtain the fused feature representation. The formula is:
[0039] Attention(Q,K,V)=Softmax(α)·V (14)
[0040] Among them, Softmax(a) is the normalized attention weight, V is the value vector matrix;
[0041] The long-term dependencies of feature sequences are captured through the Transformer layer. The calculation formula of the Transformer layer is:
[0042] MultiHead(Q,K,V)=Concat(head1,…,head h )W O (15)
[0043] in,
[0044]
[0045] Attention is a self-attention mechanism method. Each head iRepresents an independent self-attention head, projects the query, key, and value vectors into different subspaces through linear transformation, and concatenates the results of each self-attention head and uses the output weight matrix W O Perform linear transformation to obtain the output result of the Transformer layer;
[0046] The Transformer layer processes the position information in the sequence data by introducing position encoding. The calculation formula of position encoding is:
[0047]
[0048] Among them, pos is the position index, indicating the position of the element in the sequence; i is the dimension index, indicating the dimension of the feature; d model is the model dimension.
[0049] Preferably, the method further comprises:
[0050] The TCN module introduces a bidirectional long short-term memory (LSTM) network (BiLSTM). BiLSTM includes forward and reverse LSTMs. BiLSTM processes both forward and reverse information of the time series. LSTM units control the flow of information through a gating mechanism to handle long-range dependencies in time series. The hidden state of the STM unit is calculated as follows:
[0051]
[0052] in, is the hidden state of the forward LSTM, representing the contextual information from the beginning of the sequence to the current position; is the hidden state of the reverse LSTM, which represents the contextual information from the end of the sequence to the current position; [;] represents the concatenation operation, which combines the forward and reverse hidden states into a complete bidirectional hidden state;
[0053] The output of BiLSTM is mapped to the output space of the task through the fully connected layer. The calculation formula is:
[0054] y=σ(Wx+b) (20)
[0055] Among them, σ is the ReLU activation function, which is used to introduce nonlinear characteristics so that the model can learn complex mapping relationships; W is the weight matrix, which is used for linear transformation; b is the bias vector, which is used to adjust the baseline value of linear transformation; x is the input feature vector, that is, the output of the bidirectional LSTM; y is the output vector, which represents the model's prediction result for the input. The prediction result y is used as the time series data feature of the user data. The prediction result y and the demodulated prediction result based on y together constitute the input data of the input feature of the DRQN reinforcement learning part.
[0056] Preferably, the goal of maximizing long-term rewards is to continuously explore and learn the optimal decision-making strategy based on the time series data features and environmental information through the DRQN reinforcement learning mechanism, including:
[0057] Setting the input features of DRQN reinforcement learning includes two parts: one is the final output vector y of the deep learning model in the previous stage, and the other is the result of demodulation prediction based on y, that is, the predicted category obtained by the argmax() operation. The current reward value is determined by comparing the predicted value with the true value.
[0058] In DRQN model training, s t represents the time series features of the input, a t Represents the predicted model output, which is used to calculate the reward and update the Q value, reward r t The predicted category is compared with the true label and the calculation formula is:
[0059]
[0060] in, is the predicted category, y t is the true label;
[0061] Q(s t ,a t )The update formula is as follows:
[0062]
[0063] The Q value is updated by calculating the Q value of the current state, the maximum Q value of the next state, the target Q value, and the Q loss, and updating the model parameters through back propagation and the optimizer. During training, each epoch model iterates multiple times on the training set, calculates the output and hidden state, calculates the classification loss and Q loss, and sums the weighted sum to obtain the total loss. The DRQN model parameters are updated through back propagation to obtain the trained DRQN model;
[0064] Assume state s t 、Action a t and reward r t Represent the current features, demodulation decisions, and rewards for demodulation results respectively. The goal of the trained DRQN model is to maximize the long-term reward G t , the calculation formula is:
[0065]
[0066] Where γ is a discount factor that weighs the importance of current and future rewards;
[0067] The DRQN model updates the Q value through Q learning. DRQN learns the optimal demodulation strategy π* by interacting with the environment, making the long-term reward G t maximize.
[0068] As can be seen from the technical solutions provided by the above embodiments of the present invention, an enhanced OFDM-DCSK chaotic communication architecture that integrates deep reinforcement learning has been designed. This solution innovatively employs a hierarchical feature extraction module to capture multi-scale time-frequency features, enhances feature interaction through an attention mechanism, and utilizes a deep reinforcement learning framework for dynamic model optimization, significantly improving system performance.
[0069] Additional aspects and advantages of the present invention will be set forth in part in the following description, will become apparent from the following description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0071] Figure 1 A schematic diagram illustrating an implementation principle of an enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning provided by an embodiment of the present invention;
[0072] Figure 2 A processing flow chart of an enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning provided by an embodiment of the present invention;
[0073] Figure 3 A structural diagram of a time series feature extraction model provided by an embodiment of the present invention;
[0074] Figure 4 This is a structural diagram of a demodulation strategy optimization model based on reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0075] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limiting the present invention.
[0076] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or couplings. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.
[0077] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless defined as such herein, will not be interpreted in an idealized or overly formal sense.
[0078] To facilitate understanding of the embodiments of the present invention, several specific embodiments will be further explained below with reference to the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.
[0079] The implementation principle diagram of an enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning provided by an embodiment of the present invention is as follows: Figure 1As shown, the system is mainly divided into three stages: signal modulation and transmission, time series feature extraction and analysis, and demodulation strategy optimization based on reinforcement learning. In the signal modulation and transmission stage, the DTSAT-DRQN (Dual TCN, Self-attention, Transformer, Deep Recurrent Q-Network) model proposed in this invention adopts a modulation mechanism that combines OFDM (Orthogonal Frequency Division Multiplexing) and DCSK (Differential Chaos Shift Keying). At the transmitting end, the chaotic sequence is combined with the user data to achieve chaotic encryption modulation of the signal. The chaotic reference signal and the required transmission signal are transmitted on the OFDM multi-subcarrier through a specific modulation method. The receiving end demodulates the signal by utilizing the autocorrelation of the chaotic sequence and the incoherent demodulation principle. This modulation method combines the frequency diversity of OFDM and the anti-interference capability of DCSK, and can maintain high spectrum efficiency and reliability in complex channel environments such as multipath fading.
[0080] During the time series feature extraction and analysis phase, the DTSAT-DRQN model introduces a bidirectional temporal convolutional network (DualTCN) to enhance its ability to extract time series data features. Convolution kernels with varying dilation rates expand the receptive field and enable multi-scale feature extraction. Furthermore, the DTSAT-DRQN model integrates a self-attention mechanism with a Transformer network to adaptively capture long-range dependencies in time series. The multi-head attention mechanism captures dependencies from different representation subspaces, avoiding information loss and improving model performance. The integrated DTSAT-DRQN model can more accurately understand and represent signal features, providing more comprehensive and richer information support for subsequent demodulation decisions.
[0081] During the demodulation strategy optimization phase, the DTSAT-DRQN model incorporates a deep recurrent Q-network (DRQN) reinforcement learning mechanism to enhance its dynamic adaptability. DRQN enables the model to interact with the dynamic communication environment and, through a reward mechanism, guides the model to learn the optimal demodulation strategy. Aiming to maximize long-term rewards, DRQN continuously explores and utilizes environmental information to learn the optimal decision-making strategy. This mechanism enables the model to adjust the demodulation strategy in real time, achieving long-term improvements in demodulation performance in various communication scenarios.
[0082] The processing flow chart of an enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning provided by an embodiment of the present invention is as follows: Figure 2 As shown, the processing steps include the following:
[0083] Step S1: In the signal modulation and transmission stage, the chaotic sequence is first generated using the second-order Chebyshev polynomial function. The principle is shown in formula (1):
[0084]
[0085] Among them, x k represents the kth element of the chaotic sequence, x k-1 is the k-1th element of the chaotic sequence, and the initial value x0 is randomly selected in the interval [-1,1].
[0086] Step S2: Generate BPSK (Binary Phase Shift Keying) modulated user data, the user data is represented as: i ∈{-1,+1}, where i is the symbol index.
[0087] Step S3: Multiply the user data by the chaotic sequence to complete the modulation process and obtain the modulation symbol s(t):
[0088] s(t)=b(t)·x(t) (2)
[0089] Step S4: Perform OFDM modulation on the modulation symbol s(t), distribute it across N subcarriers, and add a cyclic prefix (CP) to obtain an OFDM symbol. To reduce inter-symbol interference (ISI), the CP length is N / 4, and the total OFDM symbol length is N + CP.
[0090] Step S5: Add noise according to the selected channel model. For the Rayleigh channel, the expression of the received signal is:
[0091]
[0092] Where h is the Rayleigh channel coefficient, which obeys the complex Gaussian distribution For the high-speed railway channel model (Rician fading), the expression of the received signal is:
[0093]
[0094] Among them, h LOS is the sight distance component, h NLOS is the non-line-of-sight component, and K is the Rician factor. For the Nakagami fading channel, the expression of the received signal is:
[0095]
[0096] Where θ is the Nakagami fading channel coefficient, which obeys the Nakagami distribution Nakagami(m,Ω).
[0097] Step S6: At the receiving end, remove the CP and perform FFT (fast Fourier transform) to obtain the frequency domain representation of the received signal. Through channel estimation and equalization, the equalized symbol is obtained. The expression is:
[0098]
[0099] in, is the channel estimate, and ∈ is a small constant used to prevent division by zero.
[0100] Step S7: Construct a time series training data set. The value of the sample data in the training data set is the real part of the equalized symbol:
[0101]
[0102] The label of the sample data is the value mapped from the user data to [0,1]:
[0103]
[0104] Through these steps, a time series training dataset can be generated for training a deep learning model to demodulate and analyze modulated signals.
[0105] Step S8: Figure 3 This is a structural diagram of a time series feature extraction model provided by an embodiment of the present invention. In the time series feature extraction and analysis stage, a TCN (Temporal Convolutional Network) module based on CNN (Convolutional Neural Networks) is first constructed. Dilated convolution is used to expand the receptive field while keeping the length of the output sequence unchanged. The operating formula of dilated convolution is:
[0106] y i =σ(W*x i +b) (9)
[0107] Where * represents the convolution operation, W is the weight of the convolution kernel, which determines the shape and weight distribution of the convolution kernel; b is the bias, which is used to adjust the baseline value of the convolution result; σ is the ReLU nonlinear activation function, which is used to introduce nonlinear characteristics to enable the model to learn complex patterns.
[0108] The local temporal convolutional network (Local TCN) is based on TCN and captures short-term dependencies of time series by using smaller convolution kernels (kernel_size=3) and smaller dilation factors (dilation_base=2). Smaller convolution kernels and dilation factors enable the model to focus on a smaller time window, thereby better capturing local features. The global temporal convolutional network (Global TCN) is also based on TCN, but uses larger convolution kernels (kernel_size=5) and larger dilation factors (dilation_base=3). Its main function is to capture the global temporal patterns of time series (such as long-term dependencies). Larger convolution kernels and dilation factors enable the model to cover a wider time range, thereby better capturing global features. Next, the attention mechanism module fuses the features of local and global TCN through the self-attention mechanism. Specifically, the local features are first concatenated with the global features, and then the query (Query), key (Key) and value (Value) vectors are generated through linear transformation. The calculation formula is:
[0109] Q=Linear query (Concat(Local,Global)) (10)
[0110] K=Linear key (Concat(Local,Global)) (11)
[0111] V=Linear value (Concat(Local,Global)) (12)
[0112] Among them, Linear query 、Linear key and Linear value It is a linear transformation function used to map the concatenated features into different spaces to generate query, key, and value vectors.
[0113] The attention score is calculated by the dot product of Q and K and normalized by the Softmax() function. The formula is:
[0114]
[0115] Where d is the feature dimension (the dimension of Q and K). The attention score represents the query vector Q i With the key vector K j The similarity of these scores is normalized into a probability distribution through the Softmax() function, so that each value vector V j The weight is between 0 and 1.
[0116] Finally, the attention weight is used to perform weighted summation on the V vector to obtain the fused feature representation, which is:
[0117] Attention(Q,K,V)=Softmax(α)·V (14)
[0118] Here, Softmax(a) is the normalized attention weight, and V is the value vector matrix. The weighted sum operation linearly combines the value vectors according to the attention weights, resulting in a feature representation that combines local and global features. Subsequently, a Transformer layer is constructed, which includes a multi-head self-attention mechanism and a feedforward network to capture long-term dependencies in feature sequences and further enhance global features. The calculation formula for multi-head self-attention is:
[0119] MultiHead(Q,K,V)=Concat(head1,…,head h )W O (15)
[0120] in,
[0121]
[0122] Attention is a self-attention mechanism method, as mentioned above. Each head i Represents an independent self-attention head, projects the query, key, and value vectors into different subspaces through linear transformation, and then calculates self-attention to capture the relationship between features from different angles. Finally, by concatenating the results of these self-attention heads and using the output weight matrix W O Perform linear transformation to obtain the result of multi-head self-attention. In addition, by introducing position encoding, the Transformer layer can process the position information in the sequence data. The calculation formula of position encoding is:
[0123]
[0124] Among them, pos is the position index, indicating the position of the element in the sequence; i is the dimension index, indicating the dimension of the feature; d modelis the model dimension, that is, the total dimension of the feature vector. Position encoding encodes the position information into a feature vector through sine and cosine functions, and then adds it to the input features, allowing the Transformer layer to distinguish elements at different positions in the sequence. To further extract contextual information in the time series, a bidirectional LSTM network (BiLSTM) is introduced into the model. Its structure contains LSTMs in two directions (forward and reverse), which can simultaneously process the forward and reverse information of the time series, thereby capturing long-term dependencies in the sequence. The LSTM unit controls the flow of information through a gating mechanism, effectively handling long-range dependencies in the time series. The calculation formula for its hidden state is:
[0125]
[0126] in, is the hidden state of the forward LSTM, representing the contextual information from the beginning of the sequence to the current position; is the hidden state of the reverse LSTM, which represents the contextual information from the end of the sequence to the current position; [;] represents the concatenation operation, which combines the forward and reverse hidden states into a complete bidirectional hidden state, thereby fusing the forward and reverse information of the time series.
[0127] Step S9: Map the output of BiLSTM to the output space of the task (such as classification or regression) through the fully connected layer. The calculation formula is:
[0128] y=σ(Wx+b) (20)
[0129] Here, σ is the ReLU activation function, which introduces nonlinearity and enables the model to learn complex mapping relationships. W is the weight matrix used for linear transformations. b is the bias vector used to adjust the baseline value of linear transformations. x is the input feature vector, i.e., the output of the bidirectional LSTM. y is the output vector, representing the model's prediction of the input. Before the fully connected layer, the model also introduces a batch normalization layer to accelerate training and improve the model's generalization capabilities.
[0130] The feature extraction stage can effectively capture local and global temporal patterns in time series through the collaborative work of the above modules, and perform in-depth feature extraction and fusion of time series through components such as attention mechanism, Transformer layer and bidirectional LSTM network, ultimately achieving accurate modeling and analysis of time series data.
[0131] Step S10: Figure 4This is a structural diagram of a demodulation strategy optimization model based on reinforcement learning in an embodiment of the present invention. In the demodulation strategy optimization stage based on reinforcement learning, the present invention uses a deep recurrent Q network (DRQN) to optimize the demodulation strategy. DRQN learns the optimal demodulation strategy by interacting with the environment. Assume that state s t 、Action a t and reward r t Represent the current features, demodulation decision and reward of demodulation result respectively.
[0132] Input state s for reinforcement learning t It is the result y output by the previous deep learning steps, i.e., the fusion feature representation of bidirectional TCN output, attention fusion features, Transformer temporal coding, and BiLSTM output state. These features jointly encode the local and global information, temporal dependencies, and historical context required for signal demodulation, enabling DRQN to learn robust decision-making strategies in dynamic channels. Action state a t is the demodulation result of the model output y, that is, the argmax() value of y, r t Is the reward value, if the predicted value is equal to the true value, it is given 1, otherwise it is 0.
[0133] Step S11: The goal of DRQN is to maximize the long-term reward G t , the calculation formula is:
[0134]
[0135] Here, γ is a discount factor that is used to weigh the importance of current and future rewards.
[0136] DRQN updates the Q value through Q learning, Q(s t ,a t )The update formula is as follows:
[0137]
[0138] Finally, DRQN learns the optimal demodulation strategy π* to maximize the long-term reward.
[0139] The solution process of the above optimal demodulation strategy π includes the following: First, the final vector y output by the deep learning model and the demodulation prediction category obtained by y through the argmax() operation are directly spliced as the input feature of reinforcement learning to provide a data basis for subsequent decision-making. Secondly, by comparing the predicted category with the true label, when the two are consistent, a reward value of 1 is given, and when they are inconsistent, other set values are given. The next decision direction is directly adjusted based on the reward result. Furthermore, strictly follow the Q value update formula and use the current reward r t, discount factor γ and the maximum Q value of the next state, calculate the target Q value, and then combine the learning rate α to calculate the Q(s) of the current state. t ,a t ) is updated. In the model training phase, in each training cycle, the model output and hidden state are iteratively calculated based on the training set data for multiple times, and the classification loss and Q loss are calculated respectively. The two are added together according to the pre-set weights to obtain the total loss, and then the DRQN model parameters are updated through the back propagation algorithm. At the same time, in order to maximize the long-term reward G t As the goal, the discount factor γ is used to balance the current reward and future reward, and the impact of each decision on the long-term benefit is calculated. Finally, combining the environmental information and time series data characteristics, the action that maximizes the Q value is selected as the demodulation decision at each decision. Then, according to the reward value of the environmental feedback, the subsequent decision actions are continuously adjusted, and the demodulation strategy is gradually optimized to achieve efficient demodulation in complex time series data processing tasks.
[0140] Step S12: In model training, the present invention uses the above feature extraction module as the basic model and combines it with DRQN: t Represented by the input time series features, a t The model outputs predictions, which are used to calculate rewards and update Q values; the reward r t The predicted category is compared with the true label and the calculation formula is:
[0141]
[0142] in, is the predicted category, y t The Q-value is updated by calculating the Q-value of the current state, the maximum Q-value of the next state, the target Q-value, and the Q-loss. The model parameters are then updated through backpropagation and the optimizer. During training, the model iterates multiple times over the training set, calculating the output and hidden states, the classification loss, and the Q-loss. The weighted sum of these sums yields the total loss. The parameters are then updated through backpropagation, and the target model parameters are periodically updated to stabilize training.
[0143] Through the above training process, the DRQN model can learn the optimal demodulation strategy π*, so that the long-term reward G t Maximization, the model combines the advantages of deep learning and reinforcement learning, can effectively process time series data, and performs well in demodulation tasks.
[0144] In summary, the embodiments of the present invention aim to combine the advantages of the Dual TCN network in multi-scale feature extraction, the self-attention mechanism and the Transformer's ability in long-distance dependency modeling, and the characteristics of deep reinforcement learning in dynamic environment adaptation to solve key problems in existing chaotic communication systems, such as insufficient feature extraction, difficulty in modeling long-term dependencies, and poor adaptability of static models. Specifically, the present invention first implements parallel extraction of multi-scale features of the signal through the Dual TCN architecture, where local branches capture transient features and global branches model long-term trends. Then, the combination of self-attention and Transformer effectively solves the shortcomings of traditional methods in modeling long-term dependencies. Finally, the introduced DRQN reinforcement learning framework gives the system dynamic adaptability, enabling it to automatically adjust the demodulation strategy according to real-time channel conditions. This innovative design that integrates multiple technologies not only greatly improves the performance of the system under complex channel conditions, but also maintains high computational efficiency.
[0145] Experimental results show that compared with existing solutions, the architecture of the present invention shows significant advantages in key performance indicators such as demodulation accuracy, environmental adaptability and system robustness while maintaining computational efficiency, providing a reliable technical path for the engineering implementation of high-security chaotic communication systems.
[0146] Those skilled in the art will appreciate that the accompanying drawings are merely schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.
[0147] From the above description of the embodiments, it can be seen that those skilled in the art can clearly understand that the present invention can be implemented by means of software plus the necessary general-purpose hardware platform. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.
[0148] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative efforts.
[0149] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. An enhanced OFDM-DCSK chaotic communication method integrating deep reinforcement learning, characterized in that: include: At the transmitting end, the chaotic sequence and user data are chaotically encrypted to obtain a chaotic reference signal, which is then modulated onto an OFDM multi-subcarrier for transmission. At the receiving end, the chaotic reference signal is demodulated by utilizing the autocorrelation and incoherent demodulation principles of the chaotic sequence, and the time series data features of the demodulated signal are extracted through the basic time series convolutional network (TCN) module; With the goal of maximizing long-term rewards, the deep recurrent Q network DRQN reinforcement learning mechanism is used to continuously explore and learn the optimal decision-making strategy based on the characteristics of the time series data and environmental information.
2. The method according to claim 1, characterized in that The chaotic sequence and user data are chaotically encrypted at the transmitting end to obtain a chaotic reference signal, and the chaotic reference signal is modulated onto an OFDM multi-subcarrier for transmission, including: The chaotic sequence is generated using the second-order Chebyshev polynomial function, as shown in formula (1): Among them, x k represents the kth element of the chaotic sequence, x k-1 is the k-1th element of the chaotic sequence, and the initial value x0 is randomly selected in the interval [-1,1]; Generate binary phase shift keying BPSK modulated user data, the user data is expressed as: i ∈{-1,+1}, where i is the symbol index; Multiply the user data with the chaotic sequence to obtain the modulation symbol s(t): s(t)=b(t)·x(t) (2) OFDM modulation is performed on the modulation symbol s(t), the modulation symbol s(t) is distributed to N subcarriers, a cyclic prefix CP is added to obtain an OFDM symbol, and the OFDM symbol is transmitted.
3. The method according to claim 2, characterized in that The method of demodulating the chaotic reference signal at the receiving end by utilizing the autocorrelation of the chaotic sequence and the incoherent demodulation principle includes: At the receiving end, the CP of the received signal is removed and a fast Fourier transform is performed to obtain the frequency domain representation of the received signal. Through channel estimation and equalization, the equalized symbol is obtained. The expression is: in, is the channel estimation value, ∈ is a constant; Construct a time series training dataset, in which the sample data values are equalized symbols The real part of : The label of the sample data is the value mapped from the user data to [0,1]:
4. The method according to claim 3, characterized in that The extraction of time series data features from the demodulated signal through the TCN module includes: A basic temporal convolutional network (TCN) module based on the convolutional neural network (CNN) is constructed. For the sample data in the temporal training dataset, a local TCN module is used to capture local features using a smaller convolution kernel and a smaller dilation factor. A global TCN module is used to capture global features using a larger convolution kernel and a larger dilation factor. Local features and global features are fused through the self-attention mechanism, and query Q, key K, and value V vectors are generated through linear transformation. The calculation formula is: Q=Linear query (Concat(Local,Global)) (10)K=Linear key (Concat(Local,Global))(11)V=Linear value (Concat(Local,Global)) (12) Among them, Linear query 、Linear key and Linear value It is a linear transformation function used to map the concatenated features to different spaces to generate query, key, and value vectors. Local is a local feature, and Global is a global feature. The attention score is calculated by the dot product of Q and K and normalized by the Softmax() function. The formula is: Among them, d is the feature dimension, and the attention score represents the query vector Q i With the key vector K j The similarity of these scores is normalized into a probability distribution through the Softmax() function, so that each value vector V j The weight of is between 0 and 1; The V vector is weighted and summed using the attention weight to obtain the fused feature representation. The formula is: Attention(Q,K,V)=Softmax(α)·V (14) Among them, Softmax(a) is the normalized attention weight, V is the value vector matrix; The long-term dependencies of feature sequences are captured through the Transformer layer. The calculation formula of the Transformer layer is: MultiHead(Q,K,V)=Concat(head1,…,head h )W O (15) in, Attention is a self-attention mechanism method. Each head i Represents an independent self-attention head, projects the query, key, and value vectors into different subspaces through linear transformation, and concatenates the results of each self-attention head and uses the output weight matrix W O Perform linear transformation to obtain the output result of the Transformer layer; The Transformer layer processes the position information in the sequence data by introducing position encoding. The calculation formula of position encoding is: Among them, pos is the position index, indicating the position of the element in the sequence; i is the dimension index, indicating the dimension of the feature; d model is the model dimension.
5. The method according to claim 4, characterized in that The method further comprises: The TCN module introduces a bidirectional long short-term memory (LSTM) network (BiLSTM). BiLSTM includes forward and reverse LSTMs. BiLSTM processes both forward and reverse information of the time series. LSTM units control the flow of information through a gating mechanism to handle long-range dependencies in time series. The hidden state of the STM unit is calculated as follows: in, is the hidden state of the forward LSTM, representing the contextual information from the beginning of the sequence to the current position; is the hidden state of the reverse LSTM, which represents the contextual information from the end of the sequence to the current position; [;] represents the concatenation operation, which combines the forward and reverse hidden states into a complete bidirectional hidden state; The output of BiLSTM is mapped to the output space of the task through the fully connected layer. The calculation formula is: y=σ(Wx+b) (20) Among them, σ is the ReLU activation function, which is used to introduce nonlinear characteristics so that the model can learn complex mapping relationships; W is the weight matrix, which is used for linear transformation; b is the bias vector, which is used to adjust the baseline value of linear transformation; x is the input feature vector, that is, the output of the bidirectional LSTM; y is the output vector, which represents the model's prediction result for the input. The prediction result y is used as the time series data feature of the user data. The prediction result y and the demodulated prediction result based on y together constitute the input data of the input feature of the DRQN reinforcement learning part.
6. The method according to claim 5, characterized in that The goal of maximizing long-term rewards is to continuously explore and learn the optimal decision-making strategy based on the characteristics of the time series data and the environment information through the DRQN reinforcement learning mechanism, including: Setting the input features of DRQN reinforcement learning includes two parts: one is the final output vector y of the deep learning model in the previous stage, and the other is the result of demodulation prediction based on y, that is, the predicted category obtained by the argma() operation. The current reward value is determined by comparing the predicted value with the actual value. In DRQN model training, s t represents the time series features of the input, a t Represents the predicted model output, which is used to calculate the reward and update the Q value, reward r t The predicted category is compared with the true label and the calculation formula is: in, is the predicted category, y t is the true label; Q(s t ,a t )The update formula is as follows: The Q value is updated by calculating the Q value of the current state, the maximum Q value of the next state, the target Q value, and the Q loss, and updating the model parameters through back propagation and the optimizer. During training, each epoch model iterates multiple times on the training set, calculates the output and hidden state, calculates the classification loss and Q loss, and sums the weighted sum to obtain the total loss. The DRQN model parameters are updated through back propagation to obtain the trained DRQN model; Assume state s t 、Action a t and reward r t Represent the current features, demodulation decisions, and rewards for demodulation results respectively. The goal of the trained DRQN model is to maximize the long-term reward G t , the calculation formula is: Where γ is a discount factor that weighs the importance of current and future rewards; The DRQN model updates the Q value through Q learning. DRQN learns the optimal demodulation strategy π* by interacting with the environment, making the long-term reward G t maximize.