Waveform construction method of low-interception smart data link based on PPO-MMN-OCDM

CN122533906APending Publication Date: 2026-08-07MIANYANG NETOP TELECOM EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MIANYANG NETOP TELECOM EQUIP
Filing Date
2026-06-25
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]面对上述严峻挑战,当前的技术弊端集中体现在模式固化、形式单一以及特征明显等方面

Benefits of technology

其一,本发明提出了一种基于OCDM与人工噪声深度融合的物理层安全传输机制。通过对数据信息注入随机置换的人工噪声,挖掘OCDM的时频双重弥散特性,实现信号能量在二维域内高斯化均匀分布,并赋予OCDM子载波随机初始相位,以此有效抑制对信号特征谱分解,循环谱相关分析及高阶累积量的截获能力,从波形层面实现了信号隐于噪声的通信设计,完成了抗截获波形设计;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122533906A_ABST
    Figure CN122533906A_ABST
Patent Text Reader

Abstract

The application discloses a low-interception intelligent data link waveform construction method based on PPO-MMN-OCDM, and comprises the following steps: S1, constructing a data link wireless resource dynamic allocation framework based on a proximal policy optimization algorithm PPO, and modeling optimization of a modulation and coding strategy as a constrained Markov decision process; S2, initializing system parameters after receiving input information; S3, based on the initialized system parameters, performing Turbo coding, interleaving and QAM modulation on information sequences in the input information, injecting artificial noise MMN after random permutation to obtain a corresponding new symbol vector; S4, based on initial phases of OCDM subcarriers in the input information, performing OCDM modulation on the new symbol vector to obtain an initial transmission waveform; and S5, based on the data link wireless resource dynamic allocation framework, performing iteration on the transmission waveform until the iteration is completed to output a real-time low-interception data link waveform. The application can maximize the throughput capacity of the system and complete intelligent resource scheduling design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data link communication. More specifically, this invention relates to a low-interception smart data link waveform construction method based on PPO-MMN-OCDM. Background Technology

[0002] With the accelerating evolution of warfare, the battlefield environment exhibits unprecedentedly complex electromagnetic characteristics, profoundly impacting data link communications. In particular, the widespread use of high-power jammers and strong electromagnetic pulse (ESP) devices severely undermines the effectiveness and reliability of information transmission. However, their significant power characteristics also make them easily detectable and located, leading to their destruction by fire. In contrast, smart jamming devices lack significant power characteristics and possess strong concealment. They typically detect and capture slices of received signals, analyze and process them to precisely interfere with specific parts of the signal, such as synchronization headers, pilots, cyclic prefixes, and data headers, thereby paralyzing the communication process. This smart jamming method has become the primary mode of warfare in current and future conflicts.

[0003] Faced with these severe challenges, the current technological shortcomings are mainly reflected in rigid patterns, monotonous forms, and obvious characteristics. Specifically, existing data links rely on predefined communication systems and static adversarial rules, lacking the ability to dynamically evolve the system and adaptively adjust based on communication effects. This forces the system to rely solely on performance margins to resist external factors, failing to effectively mobilize system resources and further impacting communication efficiency. Simultaneously, the design of data link waveforms exhibits significant dimensional characteristics that can be reverse-analyzed, such as high-order cumulants, frequency domain energy peaks, and time domain periodicity, leading to a predicament of passive exposure and diminished effectiveness in information exchange. Summary of the Invention

[0004] One object of the present invention is to solve at least the above-mentioned problems and / or defects, and to provide at least the advantages described below.

[0005] To achieve these objectives and other advantages of the present invention, a low-interception smart data link waveform construction method based on PPO-MMN-OCDM is provided, comprising: S1. Construct a dynamic allocation framework for data link radio resources based on the near-end policy optimization algorithm PPO, and model the optimization of modulation and coding strategy as a constrained Markov decision process. S2. After receiving the input information, initialize the system parameters; S3. Based on the initialized system parameters, the information sequence in the input information is Turbo encoded, interleaved, and QAM modulated, and then injected with randomly permuted artificial noise MMN to obtain the corresponding new symbol vector. S ; S4. Based on the initial phase of the OCDM subcarrier in the input information, adjust the new symbol vector. S OCDM modulation is performed to obtain the initial transmit waveform; S5. Based on the dynamic allocation framework of wireless resources of data link, the transmitted waveform is iterated. After obtaining the final policy function and cost function, the optimal action is generated and the modulation order and coding rate of the data link are updated synchronously. The initial phase of the OCDM subcarrier in the next iteration is generated according to the artificial noise key, and then returned to S3. The real-time low intercept data link waveform is output after the iteration ends. In S1, the data link wireless resource dynamic allocation framework refers to: based on the dynamically adjustable resource status of the data link, and within the PPO algorithm framework, constructing a framework including state variables. S A Action space A Reward function R reward The basic elements, and the reward function R reward Characterized by the following formula: In the above formula, ω t , ω e , ω e , ω b For each weight, R tg For throughput rewards, R dl As a time delay reward, R EE Energy efficiency penalties R ber Error rate penalty.

[0006] Preferably, in S1, the state variable S A It is characterized by the following formula: In the above formula, SINR The signal-to-interference-plus-noise ratio (SIR) of the received signal. P tx For received signals and transmitted power; Action space A It is characterized by the following formula: In the above formula, M The modulation order, R TFor encoding bitrate; Throughput Rewards R tg It is characterized by the following formula: In the above formula, R t Indicates the system's real-time throughput. R max This indicates the system's maximum throughput; Delay Reward R dl It is characterized by the following formula: In the above formula, t op Indicates the current transmission delay. T op This indicates the system's maximum acceptable latency; Energy efficiency penalty R EE It is characterized by the following formula: In the above formula, α The power coefficient representing the effective signal. β Indicates the noise power factor; Bit error rate penalty R ber It is characterized by the following formula: In the above formula, BER This indicates the bit error rate.

[0007] Preferably, in S2, the input information includes: information sequence, modulation order, coding rate, artificial noise, random permutation matrix, power allocation factor, OCDM subcarrier initial phase, learning rate, truncation factor, maximum number of training samples, data pool capacity, and training batch size. The initialization includes: modulation order, coding rate, artificial noise key, initial wave phase of OCDM subcarrier, value network parameters, policy network parameters, current policy network parameters, data pool, and initial network state.

[0008] Preferably, in S3, the system parameters are initial information or optimal action. Before Turbo encoding, the method further includes: determining the current coding rate and modulation order based on the system parameters, and performing CRC check on the information sequence. The Turbo encoding refers to the process by which the Turbo encoding module encodes the information sequence based on the encoding code rate to obtain the corresponding encoded bitstream; The Turbo encoding module includes: an RSC component encoder, an interleaver, a matrix eliminator, a multiplexer, and a decision unit; QAM modulated transmitted signal X ( t It is characterized by the following formula: In the above formula, p ( t ), q ( t These are two independent baseband signals formed after the coded bitstream is mapped during QAM modulation. and These are two mutually orthogonal carriers generated during QAM modulation.

[0009] Preferably, in S3, the new symbol vector S It is characterized by the following formula: In the above formula, X This indicates that the transmitted signal is modulated by QAM. Z This represents the structured noise in the design, and Z Each element is independently and identically distributed in , P Denotes a random permutation matrix, and P Each row and each column of the array contains exactly one element that is 1, and all other elements are 0. α The power coefficient representing the effective signal. β This represents the noise power figure, and .

[0010] Preferably, in S4, the OCDM modulation process is as follows: New symbol vector S When loaded onto the subcarriers of the Chirp, the synthesized continuous-domain OCDM time-domain signal is obtained. s ( t It is represented by the following formula: In the above formula, N represents the number of subcarriers. S ( k ) indicates the first k Complex data symbols carrying artificial noise on each subcarrier. Indicates the first k The continuous time of a linear frequency modulated signal, where T represents the duration of a symbol; For time domain signals s ( t Discrete sampling is performed, and the resulting OCDM signal is... s [n This can be represented by the following formula: In the above formula, It is the phase rotation factor. It is a quadratic phase term. It is the rearranged quadratic phase term; Define the discrete Fresnel transformation matrix as ,but elements in Represented as: In the above formula, It is the phase rotation factor. It is an inverse quadratic phase term; Then, the discretized OCDM modulated signal Characterized by the following formula: In the above formula, This represents the inverse discrete Fresnel transform.

[0011] The present invention has at least the following beneficial effects: Firstly, this invention proposes a physical layer secure transmission mechanism based on the deep fusion of OCDM and artificial noise. By injecting randomly permuted artificial noise into the data information, the time-frequency dual dispersion characteristics of OCDM are exploited to achieve a Gaussian uniform distribution of signal energy in the two-dimensional domain. Furthermore, the OCDM subcarriers are given random initial phases, thereby effectively suppressing the interception of signal feature spectrum decomposition, cyclic spectrum correlation analysis, and higher-order cumulants. This achieves a communication design where the signal is hidden in noise at the waveform level, completing the anti-interception waveform design. Secondly, this invention constructs a dynamic allocation framework for data link wireless resources based on a near-end policy optimization algorithm. The optimization of modulation and coding strategy is modeled as a constrained Markov decision process. This mechanism utilizes the state perception capability of deep reinforcement learning to maximize the system throughput under the multi-dimensional constraints of communication latency, energy efficiency and bit error rate, and completes the intelligent resource scheduling design. Thirdly, this invention designs a reasonable engineering synchronization scheme, clarifies the design of the physical frame structure, solves the problem of confirming the dynamic optimization information of the system when facing clever interception interference under the time-varying optimization framework, enhances the feasibility and robustness of the data link in complex electromagnetic environments, and completes the engineering feasibility design.

[0012] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description

[0013] Figure 1 This is a structural design block diagram of the Turbo encoder of the present invention; Figure 2 This is a block diagram of the QAM modulation and demodulation structure of the present invention; Figure 3 This is a QAM constellation diagram corresponding to different modulation orders in this invention; Figure 4 This is a complete flowchart of the anti-interception waveform generation process in this invention. Figure 5 This is a physical frame structure design diagram of the present invention; Figure 6 This is a diagram illustrating the PPO algorithm structure and data flow of the present invention. Figure 7 This is a flowchart of the low intercept smart data link waveform construction method based on PPO-MMN-OCDM in this invention. Detailed Implementation

[0014] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.

[0015] The low-interception intelligent waveform construction method of the present invention mainly includes three aspects in its core ideas: First, the artificial noise MMN is randomly permuted and injected into the modulated QAM symbols. By adjusting the weights, the noise power is made much stronger than the signal power, thus achieving signal characteristic spectrum flattening. Secondly, the symbols are mapped onto the subcarriers of the orthogonal chirped modulated (OCDM) signal to decompose the time-frequency energy of the information to be transmitted, so that the signal presents uniformity in the two-dimensional plane and the initial phase of each subcarrier has random dynamics, ensuring that the transmitted signal does not exhibit typical cyclic stationary characteristics, thereby achieving the design of anti-interception signal.

[0016] Third, a dynamic allocation framework for data link wireless resources based on the near-end policy optimization algorithm (PPO) is constructed. The optimization of modulation and coding strategies is modeled as a constrained Markov decision process, which maximizes system throughput under multi-dimensional constraints of communication delay, energy efficiency, and bit error rate.

[0017] Furthermore, the low intercept smart data link waveform construction method for PPO-MMN-OCDM provided by this invention, such as... Figure 7 As shown, it mainly includes the following: Step 1: Input the system information sequence, modulation order, coding rate, artificial noise, random permutation matrix, power allocation factor, initial phase of OCDM subcarriers, learning rate, truncation factor, maximum number of training samples, data pool capacity, and training batch size; initialize the modulation order, coding rate, artificial noise key, initial phase of OCDM subcarriers, value network parameters, policy network parameters, current policy network parameters, data pool, and initial network state. Step 2: Based on the initial information or optimal action, determine the current coding rate and modulation order, perform CRC check on the information sequence, then perform Turbo coding on the information sequence according to the determined coding rate, interleave the encoded sequence to obtain a new information sequence, perform QAM modulation on the new information sequence according to the determined modulation order to obtain modulated symbols, inject randomly permuted artificial noise into the symbols to obtain a new symbol vector, and perform OCDM modulation on the new symbol vector according to the initial phase of the input subcarrier to obtain the initial transmit waveform; This step mainly includes the following: 1. Turbo Encoding and Decoding Design Using Turbo coding as the channel error correction coding method, it exhibits decoding performance close to the Shannon limit, effectively enabling information error correction and detection in harsh channel environments. Figure 1 As shown, the Turbo encoder is implemented by consisting of an RSC component encoder, an interleaver, a matrix eliminator, a multiplexer, and a decision unit.

[0018] Taking the encoding process as an example, let the sequence of information to be sent be... The RSC component encoder / decoder is implemented using a convolutional encoder with a feedback structure. The encoding process is mainly used to append check bits after the information bits. , C To verify the information set, the decoding process acts inversely. The decoder performs maximum a posteriori calculation in the lattice graph and eliminates uncertainty through a loop structure, thereby extracting the valid information sequence. The number of decoders corresponds to the number of encoders, completing the inverse processes of encoding and decoding respectively. The interleaving / deweaving unit uses a group interleaving method, i.e., information is written row by row and read column by column. Its purpose is to prevent continuous error effects in the information flow and alleviate the pressure of error detection and correction through dispersion. Deweaving is the reverse of the above process, which can restore the original order of the information sequence. The matrix eliminator is used to adjust the coding rate. It is a matrix composed of elements 0 and 1. When an element is 1, the corresponding encoded bit is allowed to be sent as an output bit. When an element is 0, it is not allowed to be sent. The designed coding rate of this invention is... As one of the dynamic control variables of the system, it is used to dynamically control the system's error detection and correction capabilities according to environmental conditions; the multiplexer concatenates the original bits and encoded bits in sequence to form the final encoded bit stream. .

[0019] 2. QAM modulation and demodulation design After channel coding, orthogonal amplitude modulation (OAM) is used to modulate the information stream. OAM is a combined amplitude and phase modulation technique that enables two orthogonal parallel data transmissions within the same bandwidth. It is mainly used in high-speed data transmission systems and has good spectral efficiency and power efficiency. Figure 2 As shown, QAM modulation enables two orthogonal parallel data transmissions within the same bandwidth, where... and These are two independent baseband signals formed after the encoded bitstream has been mapped. Indicates the signal phase, and the two signals are passed through shaping filters and then... and Multiplying two mutually orthogonal carrier waves yields the QAM-modulated transmit signal: .

[0020] It should be noted that after the signal passes through the channel, it is assumed that after channel estimation and frequency offset correction at the receiver, a carrier with the same frequency and phase is generated. and The output signals of the two mutually orthogonal demodulators are: After passing through the low-pass filter, the resulting upper and lower branch signals are as follows: This completes the transmission of the baseband signal.

[0021] For QAM modulation, if the input baseband signal is multi-level, a high-order QAM system can be constructed. This invention uses the signal modulation order M as one of the dynamic control variables of the system, namely… These correspond to 4QAM (QPSK), 16QAM, and 64QAM respectively, and their corresponding constellations are as follows: Figure 3 As shown.

[0022] 3. Injection and Separation Design of Artificial Noise MMN To ensure excellent signal anti-interception capability and prevent deliberate intrusion by clever interference, this invention employs frequency domain noise modulation to inject artificial noise into the modulated symbols. This injected artificial noise, after random permutation, submerges the effective signal within a flat Gaussian white noise group. A new symbol vector at the transmitting end is constructed based on affine transformation. S for: in, This represents the symbol vector after QAM modulation. This represents the structured noise of the design, where each element is independently and identically distributed. , This represents a random permutation matrix, where each row and each column has exactly one element that is 1, and all other elements are 0. It is used to shuffle the index order of noise, preventing the enemy from identifying the signal through local statistical characteristics. Z and P Through artificial noise key K generate, α , β These are the power coefficients of the effective signal and the noise power coefficients, respectively. ,Right now , and These represent the signal power and noise power, respectively, used to ensure that the artificial noise power is much greater than the signal power, so that the received signal covariance matrix... The characteristic spectrum is flattened to improve the signal's concealment ability.

[0023] It should be noted that at the receiving end, for noise separation, after the receiver receives the signal, it is aware of the structured noise. Z With random permutation matrix P The generation principle allows for the deterministic elimination of artificial noise from the equalized signal, because... It is a non-random deterministic sequence, directly generated by the key K, and does not introduce new random errors; only channel estimation errors exist, i.e.: in, This is the actual channel estimation matrix. The Fresnel equivalent channel matrix, E If the estimation error is represented, then the corresponding expression for the solution signal is: in, T y Indicates the received signal Y The result after Fresnel transformation. This represents the inverse of the actual channel estimation matrix. n It is a Gaussian white noise vector. Represents the Fresnel transform. For noise impact, This represents the error impact term. In traditional estimation, the introduction of strong artificial noise introduces severe numerical ill-conditioning, leading to significant estimation errors and noise in the system. However, based on the prior signal structure, deep learning algorithms such as CNNs to extract local features or Transformers to capture global correlations can approximate the lower bound of the optimal estimation solution, thereby accurately restoring the channel state while suppressing noise amplification. At this point, the transmitted signal can be recovered, perfectly eliminating various noises and errors. The residuals superimposed during this process can be gradually eliminated in subsequent coherent accumulation and Turbo decoding, thus achieving the acquisition of net information.

[0024] 4. OCDM signal modulation and demodulation design After injecting artificial noise, the effective signal exhibits the characteristics of Gaussian white noise at the original frequency. To further improve the signal's concealment capability, this invention introduces OCDM modulation to decompose the signal energy into a two-dimensional time-frequency spectrum, thereby deceiving the enemy's spectrum monitoring.

[0025] OCDM signal modulation uses a set of mutually orthogonal chirp signals as subcarriers to carry out data transmission. The difference between OCDM and OFDM signal modulation is that the base signal of OFDM signal is a subcarrier with a fixed frequency, while the base signal of OCDM signal is a linearly changing frequency.

[0026] For a containing N In an OCDM system with 1 subcarrier, the symbol duration is T , No. k The continuous-time expression of a linear frequency modulated signal is defined as follows: in, k It is a subcarrier index. It is the phase rotation factor, used to ensure mathematical symmetry. It is a quadratic phase term that determines the changes in the signal and the changes in the signal frequency.

[0027] OCDM signals can be multiplexed, mainly for different subcarriers. m and n Its inner product is 0, that is: In the data domain, we first set the new symbol vector S Loading onto the subcarriers of the Chirp signal means associating the modulated constellation points with the subcarriers of the Chirp signal. Assume the... k The complex data symbols carrying artificial noise on each subcarrier are: S ( k Then the synthesized continuous domain OCDM time-domain signal S (t It is the weighted sum of all subcarriers, that is: To facilitate processing by digital systems, the above signals need to be sampled. Assuming the system sampling rate is B, and B = N / T, meaning one symbol has N sampling points, then the discrete OCDM signal... s [ n ]for: The discretized expression above is a linear transformation, and the discrete Fresnel transformation matrix is ​​defined as follows: Its elements for: The OCDM modulation of the new symbol vector based on the initial phase of the input subcarrier can be characterized by the following formula: in, Indicates the symbol after OCDM modulation. This represents the symbol vector after artificial noise has been injected. Represents the inverse discrete Fresnel transform; To conceal the cyclic stationary characteristics of the signal, the initial phase of each subcarrier in the OCDM signal is randomly and dynamically adjusted, specified by XORing the system time with an artificial noise key. In this case, the subcarriers have different initial phases. , This represents the initial phase vector of the subcarrier.

[0028] It should be noted that at the receiving end, it is assumed that the received signal... Y The expression is: in, H express N × N The multipath channel matrix, due to the cyclic prefix structure in OCDM, is generally represented as a cyclic matrix. Applying a Fresnel transform to the received signal, its expression is: in, Indicates the channel HA unitary transform is performed, and the result is called the Fresnel equivalent channel matrix, which is in diagonal form. This transformation minimizes channel interference between OCDM subcarriers. Simultaneously, the diagonal matrix transforms the channel equalization process into point-by-point division, directly eliminating the high complexity of channel matrix inversion. After noise cancellation, matched filtering is used, leveraging the orthogonality of the chirp subcarriers in the OCDM, to perform coherent accumulation at the symbol level, thereby achieving OCDM demodulation and enhancing the received signal-to-noise ratio.

[0029] like Figure 4 As shown, in the construction of the anti-interception waveform, CRC check is a cyclic redundancy check, which is used to further improve the error detection capability of the sequence. Adding and de-interleaving are independent of the encoding module, which is used to further enhance the error correction capability of the sequence.

[0030] To achieve a balance between low interception and practical engineering feasibility, this invention employs a fixed modulation and coding strategy and high-speed frequency hopping based on a return-to-zero mechanism to drive the synchronization header. Specifically, the frequency hopping sequence is designed to terminate at a predetermined fixed frequency point, ensuring basic consistency of the initial phase of the subcarriers. This provides effective prior information for dynamic changes in data segments, ensuring the robustness of the communication link. After synchronization is established, the system enters the data payload transmission phase. To improve transmission efficiency and enhance the link's environmental adaptability, the modulation and coding scheme is dynamically and intelligently controlled at this stage, further supporting optimal transmission performance for the anti-interception link. The specific physical frame structure design is as follows: Figure 5 As shown, the frequency hopping process can complete the transmission of information through multiple hop groups. Each hop group contains L hops. The physical frame structure of each hop consists of power adjustment, synchronization header, cyclic prefix, data header, and data packet. Among them, power adjustment is used to wait for the hardware system to complete power adjustment, synchronization header is used for frame synchronization, cyclic prefix is ​​used to ensure signal integrity, and data header is used to transmit marking information and information on the modulation order and coding rate of each iteration.

[0031] In summary, this step creates a pseudo-rank deficiency in the signal subspace by additive spreading of artificial noise in the frequency domain, thereby countering enemy interception analysis. The quadratic phase term in the OCDM transform distributes the concentrated energy in the frequency domain evenly across the entire time domain symbol, further enhancing the signal's concealment capability.

[0032] Step 3: Transmit waveforms in the working channel and save the interaction data of this iteration to the data pool; for example... Figure 6As shown, the near-end policy optimization architecture consists of a policy optimizer, a policy network, a current policy network, a value optimizer, and a value optimization network. The value network and policy network are used for intelligent learning, evaluating the quality of actions and generating actions. The current policy network assists the policy network in achieving diverse online learning. The policy optimizer and value optimizer provide gradient directions for the policy network and value network, enabling effective learning. After obtaining the interaction data, the advantage function needs to be calculated and the result stored in the trajectory set. The value network and policy network are then trained based on the trajectory set and data pool data, the policy network parameters are updated, and the final policy function and value function are generated. The optimal action is generated, and the data link modulation order and coding rate are updated synchronously. The initial phase of the OCDM subcarrier in the next iteration is generated based on the artificial noise key, and the process returns to step two for another iteration to guide the generation of anti-interception waveforms until the generation is complete.

[0033] In this step, near-end policy optimization addresses the trust region constraint by modifying the surrogate objective function, thus controlling the KL divergence between the old and new policies within a certain range before and after iteration. This solves the problem of large-scale policy gradient changes during the learning process. Specifically, the policy gradient estimate... The calculation expression is: in, This indicates the calculation of the expected value. The policy function is constructed from a deep neural network. θ Its parameters, for t Moment State The estimated value of the dominance function under the following conditions. for t The specific actions of the intelligent agent at any given moment. This is the evaluation function. Now, consider another... A strategy According to the importance sampling theorem, the corresponding policy gradient is: in, This represents the evaluation function corresponding to another strategy. , Indicates action a According to strategy , If this occurs, then the policy gradient loss function corresponding to this policy is: in, Indicates importance weight, If there is no difference between the old and new strategies, then the strategy... , The expected value of the KL divergence is: Here, υ represents the upper bound of the policy gap stability, i.e., the trust region. The policy optimization process then involves finding the maximum value within the trust region. .

[0034] To effectively address the aforementioned optimization problem, a cut surrogate function is defined to... Placed Within this range, thus enabling the strategy With strategy To be as close as possible, its expression is defined as: in, Represents the shearing function. ε The cutoff factor is expressed as follows: The dominance function is calculated using generalized dominance estimation, and its expression is as follows: Where t is the time marker in the trajectory. for TD (0) Difference, where λ is the adjustment weight for different reward levels. For infinitely long time-domain learning tasks, generally only the time domain is considered. T Within this range, the corresponding advantage function estimation expression is: In the cycle T In the final step, the corresponding value function is obtained through the value network. By back-estimating the value estimate corresponding to each previous time marker, the corresponding advantage function can be obtained.

[0035] Within the PPO algorithm framework, and considering the dynamically adjustable resource status of the data chain, the required basic elements are set as follows: (1) State variable design State variables S A The design includes the signal-to-interference-plus-noise ratio of the received signal. SINR and transmission power P tx This is used to monitor the quality of the channel, i.e.: (2) Action space design Discrete Actions A Spatial design includes modulation order M Selection and encoding bitrate R TThe selection involves two parts to adjust the dynamic adaptability of the data link according to environmental conditions, namely: (3) Reward function design reward function R reward It is the core of driving the system to achieve intelligent operation. The design of the reward function includes several parts such as throughput reward, latency reward, energy efficiency reward, and bit error rate penalty. Its specific definition is: Throughput Rewards R tg Defined as: in, R t Indicates the system's real-time throughput. R max This indicates the system's maximum throughput.

[0036] Delay Reward R EE Defined as: in, t op Indicates the current transmission delay. T op This indicates the maximum acceptable latency of the system.

[0037] Energy efficiency penalty R EE Defined as: in, α The power corresponding to the effective signal, β Corresponding noise power.

[0038] Bit error rate penalty R ber Defined as: in, BER This indicates the bit error rate.

[0039] The reward function is then designed as follows: in, ω t , ω e , ω e , ω b The adjustment weights for each item can be adjusted according to actual application needs.

[0040] The above solution is merely an illustration of a preferred example and is not limited thereto. When implementing this invention, appropriate substitutions and / or modifications can be made according to the user's needs.

[0041] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be applied to various fields suitable for the present invention. Other modifications can be readily made by those skilled in the art. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and examples shown and described herein.

Claims

1. A low-interception smart data link waveform construction method based on PPO-MMN-OCDM, characterized in that, include: S1. Construct a dynamic allocation framework for data link radio resources based on the near-end policy optimization algorithm PPO, and model the optimization of modulation and coding strategy as a constrained Markov decision process. S2. After receiving the input information, initialize the system parameters; S3. Based on the initialized system parameters, the information sequence in the input information is Turbo encoded, interleaved, and QAM modulated, and then injected with randomly permuted artificial noise MMN to obtain the corresponding new symbol vector. S ; S4. Based on the initial phase of the OCDM subcarrier in the input information, adjust the new symbol vector. S OCDM modulation is performed to obtain the initial transmit waveform; S5. Based on the dynamic allocation framework of wireless resources of data link, the transmitted waveform is iterated. After obtaining the final policy function and cost function, the optimal action is generated and the modulation order and coding rate of the data link are updated synchronously. The initial phase of the OCDM subcarrier in the next iteration is generated according to the artificial noise key, and then returned to S3. The real-time low intercept data link waveform is output after the iteration ends. In S1, the data link wireless resource dynamic allocation framework refers to: based on the dynamically adjustable resource status of the data link, and within the PPO algorithm framework, constructing a framework including state variables. S A Action space A Reward function R reward The basic elements, and the reward function R reward Characterized by the following formula: In the above formula, ω t , ω e , ω e , ω b For each weight, R tg For throughput rewards, R dl As a time delay reward, R EE Energy efficiency penalties R ber Error rate penalty.

2. The low-interception smart data link waveform construction method based on PPO-MMN-OCDM as described in claim 1, characterized in that, In S1, the state variable S A It is characterized by the following formula: In the above formula, SINR The signal-to-interference-plus-noise ratio (SIR) of the received signal. P tx For received signals and transmitted power; Action space A It is characterized by the following formula: In the above formula, M The modulation order, R T For encoding bitrate; Throughput Rewards R tg It is characterized by the following formula: In the above formula, R t Indicates the system's real-time throughput. R max This indicates the system's maximum throughput; Delay Reward R dl It is characterized by the following formula: In the above formula, t op Indicates the current transmission delay. T op This indicates the system's maximum acceptable latency; Energy efficiency penalty R EE It is characterized by the following formula: In the above formula, α The power coefficient representing the effective signal. β Indicates the noise power factor; Bit error rate penalty R ber It is characterized by the following formula: In the above formula, BER This indicates the bit error rate.

3. The low-interception smart data link waveform construction method based on PPO-MMN-OCDM as described in claim 1, characterized in that, In S2, the input information includes: information sequence, modulation order, coding rate, artificial noise, random permutation matrix, power allocation factor, OCDM subcarrier initial phase, learning rate, truncation factor, maximum number of training samples, data pool capacity, and training batch size; The initialization includes: modulation order, coding rate, artificial noise key, initial wave phase of OCDM subcarrier, value network parameters, policy network parameters, current policy network parameters, data pool, and initial network state.

4. The low-interception smart data link waveform construction method based on PPO-MMN-OCDM as described in claim 1, characterized in that, In S3, the system parameters are initial information or optimal action. Before Turbo encoding, it also includes: determining the current coding rate and modulation order based on the system parameters, and performing CRC check on the information sequence. The Turbo encoding refers to the process by which the Turbo encoding module encodes the information sequence based on the encoding code rate to obtain the corresponding encoded bitstream; The Turbo encoding module includes: an RSC component encoder, an interleaver, a matrix eliminator, a multiplexer, and a decision unit; QAM modulated transmitted signal X ( t It is characterized by the following formula: In the above formula, p ( t ), q ( t These are two independent baseband signals formed after the coded bitstream is mapped during QAM modulation. and These are two mutually orthogonal carriers generated during QAM modulation.

5. The low-interception smart data link waveform construction method based on PPO-MMN-OCDM as described in claim 1, characterized in that, In S3, the new symbol vector S It is characterized by the following formula: In the above formula, X This indicates that the transmitted signal is modulated by QAM. Z This represents the structured noise in the design, and Z Each element is independently and identically distributed in , P Denotes a random permutation matrix, and P Each row and each column of the array contains exactly one element that is 1, and all other elements are 0. α The power coefficient representing the effective signal. β This represents the noise power figure, and .

6. The low-interception smart data link waveform construction method based on PPO-MMN-OCDM as described in claim 5, characterized in that, In S4, the OCDM modulation process is as follows: New symbol vector S When loaded onto the subcarriers of the Chirp, the synthesized continuous-domain OCDM time-domain signal is obtained. s ( t It is represented by the following formula: In the above formula, N represents the number of subcarriers. S ( k ) indicates the first k Complex data symbols carrying artificial noise on each subcarrier. Indicates the first k The continuous time of a linear frequency modulated signal, where T represents the duration of a symbol; For time domain signals s ( t Discrete sampling is performed, and the resulting OCDM signal is... s [ n This can be represented by the following formula: In the above formula, It is the phase rotation factor. It is a quadratic phase term. It is the rearranged quadratic phase term; Define the discrete Fresnel transformation matrix as ,but elements in Represented as: In the above formula, It is the phase rotation factor. It is an inverse quadratic phase term; Then, the discretized OCDM modulated signal Characterized by the following formula: In the above formula, This represents the inverse discrete Fresnel transform.