An ai-based end-to-end wireless communication system and method

By employing an AI-based end-to-end wireless communication method, and utilizing hybrid modeling and adversarial training to generate adversarial distortion modes, the decoding failure and bit error rate issues of traditional systems in complex environments are resolved, achieving stable, high-quality communication and adaptive optimization in complex environments.

CN121665278BActive Publication Date: 2026-05-08TIANYUAN RUIXIN COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANYUAN RUIXIN COMM TECH CO LTD
Filing Date
2026-02-06
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional modular communication systems struggle to effectively cope with joint distortion in complex, time-varying wireless communication environments, leading to decoding failures and soaring bit error rates in real-world scenarios.

Method used

An AI-based end-to-end wireless communication method is adopted. By integrating differentiable physical simulation and data-driven residual learning into a hybrid model, a joint distortion representation is constructed. An adversarial training mechanism is used to generate adversarial distortion patterns, and the communication process parameters are adjusted in real time for multiple rounds of training until stability is achieved.

Benefits of technology

Maintain stable and reliable high-quality communication in complex environments, reduce the risk of communication interruption and performance degradation, and achieve adaptive optimization that dynamically adapts to various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121665278B_ABST
    Figure CN121665278B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of wireless communication, and particularly discloses an end-to-end wireless communication system and method based on AI, which collects signals and channels and hardware data in a real scene to construct an original data set; based on the data set, a hybrid modeling method of fusing a differentiable physical simulation and data-driven residual learning is used to construct a parameter-adjustable joint distortion representation; the representation is combined with internal states of a communication process, and a reinforcement learning idea is used to dynamically adjust parameters to generate an adversarial distortion mode; the mode is used for multi-round adversarial training of the communication process, communication parameters and distortion generation strategies are alternately optimized until a stable and robust state is reached; finally, the trained communication process is deployed on an actual link, and parameter adaptive matching and online fine-tuning are realized through environment sensing; the application can effectively simulate the coupling distortion of channels and hardware, and improves the communication reliability and generalization performance in a complex dynamic wireless environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and more specifically to an AI-based end-to-end wireless communication system and method. Background Technology

[0002] As wireless communication technology advances to higher frequency bands and more complex scenarios, traditional modular communication system design faces severe challenges. Traditional methods typically involve the separate design and optimization of transmitters, channel coding and decoding, modulation and demodulation based on explicit mathematical models (such as Shannon's theory and specific modulation and coding schemes). This approach performs well in ideal or steady-state channels, but it struggles to effectively address the complex, time-varying, and difficult-to-model "joint distortion" caused by the coupling effect of multipath, Doppler effects, and inherent RF hardware defects (such as power amplifier nonlinearity and phase noise) in environments such as dense urban areas and high-speed mobile environments.

[0003] In recent years, artificial intelligence (AI) technology has provided a new paradigm for the design of communication physical layers, namely end-to-end communication systems based on deep neural networks. These systems treat the transmitter and receiver as a whole for joint optimization, aiming to learn patterns in data to bypass explicit mathematical models and automatically find efficient and reliable encoding and decoding strategies. Traditional methods cannot accurately model and reproduce the complex joint distortions generated in real-world wireless environments due to the real-time coupling between rapidly changing channel characteristics caused by high-speed movement and complex reflections, and inherent hardware impairments such as power amplifier nonlinearity and phase noise. This means that systems trained in idealized simulation environments may experience decoding failures and soaring bit error rates when deployed in real-world scenarios (such as dense urban vehicle networks) and faced with unknown combinations of distortions. Summary of the Invention

[0004] The purpose of this invention is to provide an AI-based end-to-end wireless communication system and method to solve the problems mentioned above.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] An AI-based end-to-end wireless communication method includes the following steps:

[0007] S1: Collect the signal sequences transmitted between the transmitting end and the receiving end under different actual scenarios, and record the corresponding channel state information and hardware working status data to form the original dataset;

[0008] S2: Based on the original dataset, an integrated and continuously adjustable joint distortion representation is constructed by integrating differentiable physical simulation and data-driven residual learning into a hybrid modeling method.

[0009] S3: Taking the joint distortion representation and the internal state information of the current end-to-end communication process as input, adjust multiple parameters in the joint distortion representation in real time to generate an adversarial distortion mode that reduces the performance of the current communication process;

[0010] S4: Using an adversarial distortion mode, the end-to-end communication process is trained in multiple rounds. In each round, the parameters of the communication process are first optimized with the current distortion mode, and then the generation strategy of the distortion mode is updated based on the performance feedback of the optimized communication process until the communication process reaches a stable state under various distortion conditions.

[0011] S5: Apply the end-to-end communication process obtained after training to the actual wireless communication link to complete the intelligent transmission and recovery of the entire process from sending signals to receiving signals.

[0012] As a further aspect of the present invention: S2 specifically includes:

[0013] A differentiable physical simulation structure based on the electromagnetic propagation equation is constructed. The channel parameters and hardware parameters extracted from the original dataset in the previous step are used as inputs, and the basic distortion response is output.

[0014] The residual learning network is trained using the original dataset. The residual learning network takes the basic distortion response of the previous stage and the corresponding actual measurement signal as input and learns to output high-dimensional residual features.

[0015] By weighted fusion of the basic distortion response and high-dimensional residual features, a joint distortion characterization with continuously adjustable parameters is formed.

[0016] As a further aspect of the present invention: the step of training the residual learning network using the original dataset specifically includes:

[0017] The original dataset is aligned by matching the basic distortion response with the corresponding actual measurement signal according to its timestamp and signal frame structure to form paired training samples.

[0018] A residual learning network with a bidirectional interactive structure is constructed. The residual learning network processes the basic distortion response and the actual measurement signal separately through parallel feature extraction channels, and fuses the two features through a cross-attention mechanism in the deep layer to calculate high-dimensional residual features.

[0019] A progressive training strategy based on curriculum learning is adopted. The residual learning network is initially trained using training samples with weaker distortion, and then training samples with stronger distortion are gradually introduced until the network can stably output high-dimensional residual features that cover the entire distortion range.

[0020] As a further aspect of the present invention: S3 specifically includes:

[0021] Extract internal state information representing instantaneous decoding confidence and signal feature distribution from the current end-to-end communication process;

[0022] Based on internal state information, calculate the performance degradation gradient of the communication process under the current joint distortion representation;

[0023] Based on the performance degradation gradient, the channel multipath delay spread, hardware nonlinearity intensity, and noise basis parameters in the joint distortion characterization are adjusted synchronously and in a correlated manner.

[0024] Substitute the adjusted parameter set into the joint distortion characterization, and synthesize and output the adversarial distortion mode that maximizes the local bit error rate of communication in real time.

[0025] As a further aspect of the present invention: the performance degradation gradient of the computational communication process under the current joint distortion representation specifically includes:

[0026] Extract the confidence distribution and characteristic statistics of the received signal from the internal state information;

[0027] Based on confidence distribution and characteristic statistics, a scalar evaluation function is constructed to reflect the instantaneous robustness level of the communication process.

[0028] Deterministic perturbations are applied to key parameters in the joint distortion characterization, and the corresponding changes in the scalar evaluation function are obtained through forward computation.

[0029] The performance degradation gradient is calculated based on the ratio of the change to the deterministic perturbation.

[0030] As a further aspect of the present invention: S4 specifically includes:

[0031] Based on the current adversarial distortion mode, a competitive adaptive modulation training is performed on the end-to-end communication process;

[0032] The performance feedback of the trained communication process in adversarial distortion mode is collected, and the challenge intensity of the adversarial distortion mode to the communication process is evaluated. Based on this two-way evaluation result, the policy parameters used to generate the adversarial distortion mode are updated.

[0033] Repeat the training and update steps, and monitor the dynamic balance between performance feedback and challenge intensity in real time. When both remain within the preset stable range in multiple rounds of training, the communication process is considered to have reached a stable state.

[0034] As a further aspect of the present invention: the updating of the policy parameters used to generate the adversarial distortion mode specifically includes:

[0035] The performance feedback and challenge intensity are normalized to obtain quantitative performance degradation indicators and distortion challenge levels, respectively.

[0036] The performance degradation index and the distortion challenge level are input into the dynamic policy update function, which calculates the policy adjustment vector.

[0037] The policy adjustment vector is used to iteratively update the policy parameters on which the distortion pattern is based.

[0038] As a further aspect of the present invention: S5 specifically includes:

[0039] Monitor the real-time environmental status of the actual wireless communication link and extract the physical channel characteristics and hardware operating point of the current link;

[0040] Based on the real-time environmental state, the system adaptively selects or fuses parameter combinations that match the current conditions from various parameter configurations accumulated during the communication process in the training process.

[0041] In actual wireless communication links, signal transmission is performed using an end-to-end communication process initialized with the selected parameter combination, and its transmission effect data is collected in real time.

[0042] Based on transmission performance data, the parameter combinations are adjusted online until the performance indicators of intelligent transmission and recovery throughout the entire process reach the preset optimal working threshold.

[0043] As a further aspect of the present invention: the online adjustment of the parameter combination specifically includes:

[0044] Real-time analysis of transmission performance data is performed to extract the current signal transmission bit error rate trend and the distortion characteristics of the received signal constellation diagram;

[0045] The bit error rate trend and distortion characteristics are matched with a lightweight experience base built from historical successful adjustment records to obtain a set of alternative parameter fine-tuning vectors.

[0046] From the candidate parameter fine-tuning vectors, select the one with the minimum evaluation cost, perform a superposition operation on the parameter combination, and generate new online working parameters.

[0047] Continue transmitting using the new online working parameters until the calculated performance index stops improving after multiple iterations, at which point the optimal working threshold has been reached.

[0048] An AI-based end-to-end wireless communication system includes:

[0049] The data processing module collects signal sequences transmitted between the transmitter and receiver in different real-world scenarios, and records the corresponding channel status information and hardware operating status data to form the raw dataset.

[0050] The joint distortion representation construction module, based on the original dataset, constructs an integrated joint distortion representation with continuously adjustable parameters by integrating a hybrid modeling method that combines differentiable physical simulation and data-driven residual learning.

[0051] The adversarial distortion mode generation module takes the joint distortion representation and the internal state information of the current end-to-end communication process as input, adjusts multiple parameters in the joint distortion representation in real time, and generates an adversarial distortion mode that reduces the performance of the current communication process.

[0052] The adversarial iterative training module uses an adversarial distortion mode to train the end-to-end communication process in multiple rounds. In each round, the parameters of the communication process are first optimized with the current distortion mode, and then the generation strategy of the distortion mode is updated based on the performance feedback of the optimized communication process until the communication process reaches a stable state under various distortion conditions.

[0053] The intelligent communication deployment module applies the end-to-end communication process obtained after training to the actual wireless communication link, completing the entire process of intelligent transmission and recovery from sending signals to receiving signals.

[0054] The beneficial effects of this invention are:

[0055] (1) This invention constructs a joint distortion representation that integrates physical laws and introduces an adversarial training mechanism to actively generate distortion patterns that degrade the current system performance for "stress testing". This process forces the communication system to continuously learn and overcome various extreme and rare distortion combinations during training, thereby gaining strong adaptability to cope with unforeseen complex scenarios. This enables the finally deployed system to maintain stable and reliable high-quality communication in dynamic environments such as dense urban areas and high-speed movement, effectively reducing the risk of communication interruption or performance drop due to sudden environmental changes or hardware non-idealities.

[0056] (2) This invention constructs a rich parameter configuration library with environment labels during the training phase and adopts an adaptive parameter fusion and online fine-tuning mechanism based on real-time environment awareness in actual deployment. The system can quickly match the current environment state, initialize the optimal parameters, and make small and precise adjustments based on real-time transmission feedback during operation until the local optimal performance is achieved. This process not only saves the huge overhead of retraining for each new scenario, but also enables a single general model to dynamically adapt to various specific deployment conditions, achieving the effect of "train once, adaptive optimization everywhere", and improving the engineering practical value and feasibility of large-scale deployment of the method. Attached Figure Description

[0057] The invention will now be further described with reference to the accompanying drawings.

[0058] Figure 1 This is a flowchart of the method of the present invention;

[0059] Figure 2 This is a system block diagram of the present invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Please see Figure 1 As shown, this invention is an AI-based end-to-end wireless communication method, comprising the following steps:

[0062] S1: Collect the signal sequences transmitted between the transmitting end and the receiving end under different actual scenarios, and record the corresponding channel state information and hardware working status data to form the original dataset;

[0063] S2: Based on the original dataset, an integrated and continuously adjustable joint distortion representation is constructed by integrating differentiable physical simulation and data-driven residual learning into a hybrid modeling method.

[0064] S3: Taking the joint distortion representation and the internal state information of the current end-to-end communication process as input, adjust multiple parameters in the joint distortion representation in real time to generate an adversarial distortion mode that reduces the performance of the current communication process;

[0065] S4: Using an adversarial distortion mode, the end-to-end communication process is trained in multiple rounds. In each round, the parameters of the communication process are first optimized with the current distortion mode, and then the generation strategy of the distortion mode is updated based on the performance feedback of the optimized communication process until the communication process reaches a stable state under various distortion conditions.

[0066] S5: Apply the end-to-end communication process obtained after training to the actual wireless communication link to complete the intelligent transmission and recovery of the entire process from sending signals to receiving signals.

[0067] In S1, fixed transmitting and receiving nodes are deployed in various representative real-world physical environments, including but not limited to dense urban streets, highway sections, and industrial areas with known strong radio frequency interference. The transmitting nodes use their radio frequency front-ends to periodically transmit a series of pre-set known training signal sequences with specific modulation formats and frame structures. Simultaneously, the receiving nodes use their radio frequency front-ends to synchronously receive signals transmitted through actual spatial channels. In this process, the conversion between the baseband digital form and the radio frequency analog form is performed by the digital-to-analog converter, analog-to-digital converter, and radio frequency modulation / demodulation unit integrated in each transmitting and receiving node.

[0068] While receiving signals, external dedicated channel sounding equipment or the channel estimation function built into the receiving node itself are used to measure and extract the channel state information experienced during transmission. Specifically, this extraction process involves processing the received known training signal sequence, calculating the correlation between this sequence and the locally stored original sequence in the time and frequency domains to obtain an estimate of the channel's impulse response, and then analyzing this impulse response to further calculate several key channel parameters, including multipath delay spread, Doppler shift spread, average path loss, and Rice factor. These parameters are recorded as numerical vectors in real time.

[0069] The system synchronously monitors and records the hardware operating status data of the transmitting and receiving nodes at the moment of signal transmission and reception. This data is obtained by reading the values ​​of various sensors and registers embedded in the device's RF front-end chip, specifically including: the operating bias voltage and current of the power amplifier, used to calculate its current nonlinear operating point; the phase error signal power of the phase-locked loop, used to characterize the phase noise level of the local oscillator; and the output codebook statistical distribution of the analog-to-digital converter when there is no signal input, used to evaluate its quantization error and floor noise characteristics.

[0070] To ensure data consistency, all acquisition operations are coordinated by a central control unit. This control unit sends synchronization trigger signals to transmitting nodes, receiving nodes, and channel sounding devices, and generates a unique timestamp for each signal transmission event. All signal samples, channel parameter vectors, and hardware status data acquired from different sources are correlated using this timestamp as an index.

[0071] The associated data is formatted to remove invalid data segments caused by device malfunctions or synchronization failures, and the remaining valid data is stored as a structured raw dataset. Each data sample in this dataset contains: a known transmitted signal sequence index, a corresponding received signal sequence, a channel state parameter vector, and a set of hardware operating status data for both the transmitting and receiving ends.

[0072] In S2, the step of constructing a differentiable physical simulation structure based on the electromagnetic propagation equation is performed. Pre-recorded channel parameters (including multipath delay, path loss, and Doppler shift) and hardware parameters (including power amplifier nonlinearity coefficients and phase noise variance) are extracted from the original dataset. This parameter set is used as input and fed into a computational flow constructed according to physical laws. The core of this flow is solving the simplified electromagnetic wave propagation equation, and the calculation process is as follows: Based on the input delay parameters, a channel impulse response consisting of multiple paths is constructed; the digital baseband representation of the transmitted signal is convolved with this impulse response to simulate multipath effects; subsequently, the amplitude of the convolution result is scaled according to the path loss parameters. Next, based on the input nonlinearity coefficients, a polynomial function is applied to the amplitude component of the signal to simulate power amplifier nonlinearity, and based on the phase noise variance, a phase perturbation generated by a specific random process is superimposed on the phase component of the signal. Finally, the flow outputs a signal sequence called the "basic distortion response," which contains only the distortion component determined by the input parameters and calculated based on physical principles.

[0073] The step involves training a residual learning network using the original dataset. This process comprises three ordered steps. The first step is alignment: the base distortion response sequence calculated in the previous steps is paired with the actual measured signal sequence in the original dataset, ensuring strict temporal alignment and thus forming paired samples for training. The second step is network construction and feature fusion: a network with bidirectional interactive processing capabilities is designed. This network contains two parallel processing paths: one receives the base distortion response sequence, and the other receives the corresponding actual measured signal sequence. Each path consists of multiple cascaded computation layers used to extract abstract features from the input sequence. Deep within the two paths, a cross-attention computation unit is introduced. The specific computation process of this cross-attention unit is as follows: first, the feature data extracted from the first processing path is designated as the query data group, and the feature data extracted from the second processing path are designated as the key data group and the value data group, respectively. Next, the attention weight distribution is calculated. Specifically, the transposes of the query data group and the key data group are multiplied to obtain an initial attention score matrix. Each element in this score matrix is ​​divided by the square root of a scaling factor calculated based on the key data group dimension to adjust the numerical range. Then, a normalized exponential function is applied along the last dimension of the adjusted score matrix, ensuring the sum of the weights at each position is one, thus obtaining the final attention weight matrix. Finally, feature fusion is performed by multiplying the obtained attention weight matrix with the value data group. The output is the high-dimensional residual feature after deep fusion. In the above process, the specific data dimensions of the query data group, key data group, and value data group can be set according to the actual network design. A common example dimension structure is: batch size, sequence length, and feature dimension.

[0074] The fused feature vector is the "high-dimensional residual feature" output by the network, representing the more complex distortion residues that the physical simulation failed to accurately model. The third step is progressive training: using a course-based learning strategy, a subset of samples with relatively small differences between the basic distortion response and the actual measured signal (i.e., "weak distortion") is first selected from the aligned training samples. This subset is used for initial network training. After the network converges on this subset, the training set is gradually expanded, introducing samples with greater differences ("stronger distortion"), until all training samples are used to complete the network training, enabling it to stably output high-dimensional residual features for various distortion conditions.

[0075] The process involves a weighted fusion of the basic distortion response and the high-dimensional residual features. The trained residual learning network, for any given input parameters, can output the basic distortion response and the corresponding high-dimensional residual features in parallel. During fusion, an adjustable weight coefficient is assigned to the high-dimensional residual features. The initial value of this weight coefficient is set based on the statistical measure of the average difference between the physical simulation results and the measured signals in the original dataset. The specific calculation process of the fusion operation is as follows: a scalar multiplication operation is performed between the high-dimensional residual feature tensor and the weight coefficient to obtain a weighted residual feature; then, this weighted residual feature is added to the basic distortion response along the feature dimension to generate the final, continuously adjustable "joint distortion representation". By changing the input channel and hardware parameters, and adjusting the weight coefficient of the residual features, the distortion characteristics of this representation can be continuously controlled.

[0076] In S3, the first step is to extract the internal state information of the end-to-end communication process. After the current communication process completes one signal reception and decoding attempt, the output result is obtained from its internal decoding function unit. Specifically, two data points are extracted: First, the set of probability values ​​corresponding to each possible information symbol in the decoding result, i.e., the "confidence distribution". The entropy of this distribution is calculated as an indicator to quantify the "instantaneous decoding confidence"; the higher the entropy value, the greater the decoding uncertainty. Second, the root mean square of the amplitude value and the standard deviation of the phase value are calculated from the signal processed by the receiving function unit, as "signal feature statistics". These two types of data together constitute the internal state information describing the health and stability of the current communication process.

[0077] The second step involves calculating the performance degradation gradient based on internal state information. First, a scalar evaluation function is constructed. This function is calculated by adding a weighted coefficient multiplied by the root mean square of the signal amplitude to the previously obtained decoding confidence entropy value, and then subtracting the product of another weighted coefficient and the standard deviation of the signal phase. The output value of this function is defined as a quantitative indicator of the "instantaneous robustness level," with a decrease in its value indicating communication performance degradation. Next, the performance degradation gradient is calculated. Parameters to be adjusted in the joint distortion characterization are selected, including channel multipath delay spread, hardware nonlinearity strength, and noise floor. For each parameter, a small, fixed positive number (i.e., a "deterministic perturbation") is added to its current value. Then, using this perturbed parameter set, the entire process from signal transmission to reception and decoding is rerun, and a new instantaneous robustness level value is calculated. The new function value is subtracted from the original function value before the perturbation to obtain the "change." Finally, this change is divided by the previously applied small positive number (the magnitude of the deterministic perturbation), and the result is the "performance degradation gradient" component in the direction of that parameter. Perform this operation sequentially on all parameters to be adjusted to obtain a set of gradient values.

[0078] The third step involves adjusting the joint distortion representation parameters based on the performance degradation gradient. Each parameter is updated according to the gradient vector calculated in the previous step. The update rule is: the new value of each parameter equals its old value, plus the product of a fixed step size coefficient and the corresponding gradient component. The key correlation here is that the update operation is performed "synchronously," meaning all parameters are adjusted simultaneously based on the gradient calculated at the same time; furthermore, the step size coefficient is dynamically scaled according to the overall magnitude of all gradient components, ensuring coordinated adjustment amplitudes and preventing drastic changes in a single parameter from disrupting the inherent correlation structure of the distortion mode. Through this step, the channel multipath delay spread, hardware nonlinearity intensity, and noise basis parameters are correlated and adjusted to a new numerical combination.

[0079] The fourth step involves synthesizing and outputting the adversarial distortion pattern. The adjusted parameter set obtained in the previous step is input into the computational process of the joint distortion representation constructed in step S2. This process calculates and generates a corresponding distortion signal superposition pattern or channel response sequence in real time based on the new parameters. This newly generated pattern, whose parameters are optimized along the direction that causes the scalar evaluation function (instantaneous robustness level) to decrease the fastest (i.e., the gradient direction), is designed to most effectively interfere with the current communication process. Its effect is to achieve a local maximization of the communication bit error rate measured under this pattern relative to the previous pattern. This is the final generated "adversarial distortion pattern," which is output in real time for subsequent training.

[0080] In S4, the first step is to perform competitive adaptive modulation training based on the current adversarial distortion mode. The adversarial distortion mode generated in step S3 is applied to the transmission and reception stages of the end-to-end communication process. The specific training operation is as follows: First, a batch of raw information data to be transmitted is input into the communication process, allowing it to complete the entire process from encoding, transmission, channel transmission to reception and decoding under the superimposed condition of current adversarial distortion. Next, the difference between the decoding result of this transmission and the raw information data is calculated, and this difference is quantified using the cross-entropy loss function. Then, the gradient descent optimization method is used to update all adjustable parameters in the communication process (including the weights of each step of encoding and decoding) according to the loss function to reduce the loss value. The "competitive adaptation" mechanism introduced in this process is reflected in: not only minimizing the basic loss, but also adding the entropy of the decoded output probability distribution as a regularization term to the loss calculation, encouraging the communication process to maintain decoding diversity under severe distortion, thereby enhancing its prediction and adaptation potential for unknown distortions. One parameter update constitutes one round of training.

[0081] The second step involves collecting performance feedback, evaluating the challenge intensity, and updating the distortion generation strategy. After training and updating the communication process according to the current distortion mode as described above, performance evaluation is performed immediately. First, "performance feedback" is collected: using a set of independent test data, the updated communication process is run again under the same adversarial distortion mode, and its bit error rate is calculated. Simultaneously, the constellation diagram of the received signal is analyzed, and the dispersion of its symbol point clusters (inter-symbol interference) is calculated. The bit error rate and inter-symbol interference are weighted and summed, and the result is recorded as the original performance feedback value F. Second, the "challenge intensity" is evaluated: the average relative change of key parameters of the current adversarial distortion mode (such as delay spread and nonlinear coefficients) with their corresponding parameter values ​​in the previous training cycle is calculated, and this magnitude is recorded as the original challenge intensity value C. Then, F and C are normalized: each is divided by a maximum reference value determined statistically based on historical data before the start of training, resulting in a "performance degradation index" f and a "distortion challenge level" c, both between 0 and 1. Finally, the policy parameters are updated: a dynamic policy update function is defined, whose inputs are f and c. Its internal calculation rule is as follows: first, an intermediate vector is calculated. The direction of this vector is dominated by the index f (i.e., the larger the value of f, the more clearly the vector points in the direction that makes communication performance more prone to deterioration), and the magnitude of the vector is modulated by the level c (i.e., the larger the value of c, the larger the magnitude). The final output of this function is called the "policy adjustment vector". The policy parameters previously used to control the gradient calculation and parameter adjustment rules in step S3 are then weighted and added to this policy adjustment vector (the weight coefficient is usually set to a positive number less than 1, such as 0.01), thus completing one iteration update of the policy parameters. This will directly affect the generation method of the next round of adversarial distortion modes.

[0082] The third step involves repeated iterations and a stable state determination. The first step (training the communication process) and the second step (evaluating and updating the strategy) are repeated, forming an alternating iteration. After each completion of the second step, the latest performance degradation metric *f* and distortion challenge level *c* are monitored in real time. A "stable interval" is preset, for example, the lower limit of *f* is 0.2 and the upper limit is 0.4, and the lower limit of *c* is 0.3 and the upper limit is 0.5. The determination condition is: in N consecutive iterations (N is a preset positive integer, such as 5), the calculated *f* value and *c* value in each iteration simultaneously fall within their respective stable intervals. When this condition is met, it is determined that a dynamic equilibrium has been reached between the communication process and the distortion pattern generation strategy, and the performance of the communication process under various distortion conditions no longer fluctuates drastically. At this point, it can be considered to have reached a "stable state," and the adversarial iterative training process is terminated.

[0083] In S5, the first stage of step S5 is monitoring and feature extraction. In a real-world deployment environment, a real-time environmental monitoring process is initiated immediately after the communication node starts. This process performs two tasks in parallel: First, physical channel features are extracted. The receiving node processes the known pilot signal sequence it receives from the transmitting node. By calculating the cross-correlation function in the time domain between the received sequence and the locally stored original pilot sequence, multiple peaks and their corresponding time delay positions are identified, thereby determining the current multipath channel delay spread. Simultaneously, by performing phase difference analysis on multiple consecutive pilot symbols, the Doppler frequency shift spread of the channel was estimated. Furthermore, the path loss is calculated based on the ratio of the average power of the received signal to the transmitted power. Second, the hardware operating point is calibrated. The coefficients fed back from the digital predistortion unit at the power amplifier driver of the transmitting node are read. These coefficients, through inverse operations, can be used to deduce the nonlinear characteristic parameters of the current power amplifier, and are recorded as a set of coefficients to characterize the nonlinearity intensity. Simultaneously, the output variance of the phase detector in the phase-locked loop of the receiving node is monitored. This serves as an indicator of phase noise level. The extracted values ​​mentioned above... , , Nonlinear coefficient set and Together, they form a feature vector describing the "real-time environmental state".

[0084] The second stage of step S5 is adaptive parameter configuration selection and fusion. During the training phase, the end-to-end communication process generates a large number of internal parameter configurations (i.e., network weight sets) for various simulated distortion conditions, and stores them together with the simulated channel and hardware parameters (i.e., "condition labels") used to generate these configurations, forming a parameter configuration library. In this stage, the currently monitored instantaneous environmental state feature vector is compared with the condition label vectors of all historical configurations in this library. The Mahalanobis distance between the current state vector and each historical label vector is calculated to account for the impact of variance in different feature dimensions. The K historical configurations with the smallest Mahalanobis distance are selected as candidates. Instead of directly selecting a single configuration, these candidate configurations are weighted and fused to obtain a parameter combination with a higher degree of matching to the current complex real-world environment. The fusion weights of each candidate configuration are... The condition is dynamically determined based on how closely its condition label is related to the current state. The calculation formula is as follows: ;

[0085] in, Representing the The Mahalanobis distance between the historical condition labels of each candidate configuration and the feature vector of the current instantaneous environment state. It is a temperature coefficient used to control the smoothness of the weight distribution; its value is typically set to all... The median. Finally, the new combination of initialization parameters. From this Weight parameters of each candidate configuration The weighted sum is obtained as follows: This process enables an adaptive mapping from a discrete experience base to a continuous environmental state.

[0086] The third stage of step S5 involves collecting actual transmission and effect data based on the initialization configuration. This includes combining the parameters obtained from the above fusion process. The data is loaded into the transmit and receive functional units of the end-to-end communication process. Then, the actual user data is transmitted. During transmission, the receiving node continuously collects two types of "transmission effect data": the first is the bit error rate (BER) at the link layer. This data is obtained by comparing the transmitted known reference data block with the received decoded data block, statistically analyzing the proportion of error bits, and calculating the average BER within a window of 1000 transmitted data blocks. The data is processed by demodulating the received data symbols, plotting their constellation diagram, and calculating the root mean square value of the Euclidean distance between all symbol points in the constellation diagram and their ideal decision positions. The second category is physical layer signal quality data. This is known as the "constellation chart distortion feature." Simultaneously, the symmetry index of the distribution of constellation points in each quadrant is calculated. .

[0087] The fourth stage of step S5 is online parameter adjustment and convergence determination. This stage constitutes a closed-loop optimization cycle. First, the collected transmission effect data is analyzed in real time. The nearest... Bit error rate sequence of time windows linear regression slope This serves as a "bit error rate trend." Simultaneously, constellation diagram distortion characteristics are considered. and symmetry Combined into a single comprehensive distortion index: ,in and The preset positive weight coefficients are used. Next, the "Lightweight Experience Library" is queried. This library is dynamically built during training and previous online phases, and each record contains three parts: the state before the adjustment is triggered, the parameter fine-tuning vector used, and so on. And the resulting increase in "performance indicators" The current Match the records with the old states in the experience base, calculate the cosine similarity, and select the record with the highest similarity. 1 record, its corresponding This constitutes a candidate fine-tuning vector set. Third, the optimal fine-tuning vector is selected from the candidate set. For each candidate vector... Calculate an "assessment cost" The calculation formula is as follows: ;

[0088] in, It is the cosine similarity between the current state and the previous state of this record. It is the Euclidean norm of the fine-tuning vector (representing the adjustment range). This is the efficiency improvement brought about by recording history. It is a positive tradeoff coefficient. The choice makes... The fine-tuning vector with the smallest value Fourth, update execution parameters: update the current working parameters. Updated to ,in It is a decay factor less than 1, used to ensure the stability of the adjustment. (Used) Continue transmission. Fifth, calculate and monitor "performance indicators". This indicator is defined as a comprehensive evaluation of transmission performance, and its calculation method is as follows: ;

[0089] in, This is the bit error rate of the current window. It adjusts the weight. This indicator. A higher value indicates better overall performance. The online adjustment cycle will continue, recalculating after each complete "parse-match-select-update" process and the collection of new, stable transmission performance data. .when Values ​​in continuous For example In the adjustment loop, the absolute value of each change is less than a preset small positive number. (For example When the system determines that "the performance index has reached the preset optimal working threshold," parameter adjustment stops, and the communication process continues using the current parameter combination. It operates stably, completing the entire intelligent transmission and recovery process from sending to receiving signals.

[0090] Please see Figure 2 As shown, an AI-based end-to-end wireless communication system includes:

[0091] The data processing module collects signal sequences transmitted between the transmitter and receiver in different real-world scenarios, and records the corresponding channel status information and hardware operating status data to form the raw dataset.

[0092] The joint distortion representation construction module, based on the original dataset, constructs an integrated joint distortion representation with continuously adjustable parameters by integrating a hybrid modeling method that combines differentiable physical simulation and data-driven residual learning.

[0093] The adversarial distortion mode generation module takes the joint distortion representation and the internal state information of the current end-to-end communication process as input, adjusts multiple parameters in the joint distortion representation in real time, and generates an adversarial distortion mode that reduces the performance of the current communication process.

[0094] The adversarial iterative training module uses an adversarial distortion mode to train the end-to-end communication process in multiple rounds. In each round, the parameters of the communication process are first optimized with the current distortion mode, and then the generation strategy of the distortion mode is updated based on the performance feedback of the optimized communication process until the communication process reaches a stable state under various distortion conditions.

[0095] The intelligent communication deployment module applies the end-to-end communication process obtained after training to the actual wireless communication link, completing the entire process of intelligent transmission and recovery from sending signals to receiving signals.

[0096] The working principle of this invention is as follows: First, transmit and receive signal sequences are collected in different real-world scenarios, and the channel state and hardware operating state are recorded simultaneously to construct an original dataset. Based on this dataset, an integrated joint distortion representation with continuously adjustable parameters is constructed by using a hybrid modeling method that integrates differentiable physical simulation and data-driven residual learning. Then, this joint distortion representation is combined with the internal state information of the current end-to-end communication process, and the distortion representation parameters are adjusted in real time to generate an adversarial distortion mode that can reduce the current communication performance. Subsequently, the generated adversarial distortion mode is used to perform multiple rounds of adversarial iterative training on the end-to-end communication process. In each round, the communication process parameters are first optimized to adapt to the distortion, and then the generation strategy of the distortion mode is updated based on the optimized performance feedback until the communication process reaches a stable state under various distortion conditions. Finally, the trained and stable communication process is deployed on an actual wireless link. By monitoring the environmental state in real time, the parameter configuration is adaptively selected or fused, and closed-loop fine-tuning is performed based on the online collected performance data during transmission, realizing intelligent and robust signal transmission and recovery from transmission to reception.

[0097] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. An AI-based end-to-end wireless communication method, characterized in that, Includes the following steps: S1: Collect the signal sequences transmitted between the transmitting end and the receiving end under different actual scenarios, and record the corresponding channel state information and hardware working status data to form the original dataset; S2: Based on the original dataset, a hybrid modeling method integrating differentiable physical simulation and data-driven residual learning is used to construct an integrated and continuously adjustable joint distortion representation, specifically including: A differentiable physical simulation structure based on the electromagnetic propagation equation is constructed. The channel parameters and hardware parameters extracted from the original dataset in the previous step are used as inputs, and the basic distortion response is output. The residual learning network is trained using the original dataset. The residual learning network takes the basic distortion response of the previous stage and the corresponding actual measurement signal as input and learns to output high-dimensional residual features. The basic distortion response and high-dimensional residual features are weighted and fused to form a joint distortion characterization with continuously adjustable parameters; S3: Taking the joint distortion representation and the internal state information of the current end-to-end communication process as input, adjust multiple parameters in the joint distortion representation in real time to generate an adversarial distortion mode that reduces the performance of the current communication process; S4: Using an adversarial distortion mode, the end-to-end communication process is trained in multiple rounds. In each round, the parameters of the communication process are first optimized with the current distortion mode, and then the generation strategy of the distortion mode is updated based on the performance feedback of the optimized communication process until the communication process reaches a stable state under various distortion conditions. S5: Applying the end-to-end communication process obtained after training to the actual wireless communication link, completing the entire intelligent transmission and recovery process from sending to receiving signals, specifically including: Monitor the real-time environmental status of the actual wireless communication link and extract the physical channel characteristics and hardware operating point of the current link; Based on the real-time environmental state, the system adaptively selects or fuses parameter combinations that match the current conditions from various parameter configurations accumulated during the communication process in the training process. In actual wireless communication links, signal transmission is performed using an end-to-end communication process initialized with the selected parameter combination, and its transmission effect data is collected in real time. Based on transmission performance data, the parameter combinations are adjusted online until the performance indicators of intelligent transmission and recovery throughout the entire process reach the preset optimal working threshold.

2. The AI-based end-to-end wireless communication method according to claim 1, characterized in that, The training of the residual learning network using the original dataset specifically includes: The original dataset is aligned by matching the basic distortion response with the corresponding actual measurement signal according to its timestamp and signal frame structure to form paired training samples. A residual learning network with a bidirectional interactive structure is constructed. The residual learning network processes the basic distortion response and the actual measurement signal separately through parallel feature extraction channels, and fuses the two features through a cross-attention mechanism in the deep layer to calculate high-dimensional residual features. A progressive training strategy based on curriculum learning is adopted. The residual learning network is initially trained using training samples with weaker distortion, and then training samples with stronger distortion are gradually introduced until the network can stably output high-dimensional residual features that cover the entire distortion range.

3. The AI-based end-to-end wireless communication method according to claim 1, characterized in that, S3 specifically includes: Extract internal state information representing instantaneous decoding confidence and signal feature distribution from the current end-to-end communication process; Based on internal state information, calculate the performance degradation gradient of the communication process under the current joint distortion representation; Based on the performance degradation gradient, the channel multipath delay spread, hardware nonlinearity intensity, and noise basis parameters in the joint distortion characterization are adjusted synchronously and in a correlated manner. Substitute the adjusted parameter set into the joint distortion characterization, and synthesize and output the adversarial distortion mode that maximizes the local bit error rate of communication in real time.

4. The AI-based end-to-end wireless communication method according to claim 3, characterized in that, The performance degradation gradient of the computational communication process under the current joint distortion representation specifically includes: Extract the confidence distribution and characteristic statistics of the received signal from the internal state information; Based on confidence distribution and characteristic statistics, a scalar evaluation function is constructed to reflect the instantaneous robustness level of the communication process. Deterministic perturbations are applied to key parameters in the joint distortion characterization, and the corresponding changes in the scalar evaluation function are obtained through forward computation. The performance degradation gradient is calculated based on the ratio of the change to the deterministic perturbation.

5. The AI-based end-to-end wireless communication method according to claim 1, characterized in that, S4 specifically includes: Based on the current adversarial distortion mode, a competitive adaptive modulation training is performed on the end-to-end communication process; The performance feedback of the trained communication process in adversarial distortion mode is collected, and the challenge intensity of the adversarial distortion mode to the communication process is evaluated. Based on this two-way evaluation result, the policy parameters used to generate the adversarial distortion mode are updated. Repeat the training and update steps, and monitor the dynamic balance between performance feedback and challenge intensity in real time. When both remain within the preset stable range in multiple rounds of training, the communication process is considered to have reached a stable state.

6. The AI-based end-to-end wireless communication method according to claim 5, characterized in that, The updated policy parameters used to generate the adversarial distortion mode specifically include: The performance feedback and challenge intensity are normalized to obtain quantitative performance degradation indicators and distortion challenge levels, respectively. The performance degradation index and the distortion challenge level are input into the dynamic policy update function, which calculates the policy adjustment vector. The policy adjustment vector is used to iteratively update the policy parameters on which the distortion pattern is based.

7. The AI-based end-to-end wireless communication method according to claim 1, characterized in that, The online adjustment of parameter combinations specifically includes: Real-time analysis of transmission performance data is performed to extract the current signal transmission bit error rate trend and the distortion characteristics of the received signal constellation diagram; The bit error rate trend and distortion characteristics are matched with a lightweight experience base built from historical successful adjustment records to obtain a set of alternative parameter fine-tuning vectors. From the candidate parameter fine-tuning vectors, select the one with the minimum evaluation cost, perform a superposition operation on the parameter combination, and generate new online working parameters. Continue transmitting using the new online working parameters until the calculated performance index stops improving after multiple iterations, at which point the optimal working threshold has been reached.

8. An AI-based end-to-end wireless communication system, characterized in that, An AI-based end-to-end wireless communication method according to any one of claims 1-7, comprising: The data processing module collects signal sequences transmitted between the transmitter and receiver in different real-world scenarios, and records the corresponding channel status information and hardware operating status data to form the raw dataset. The joint distortion representation construction module, based on the original dataset, constructs an integrated and continuously adjustable joint distortion representation by fusing differentiable physical simulation and data-driven residual learning. Specifically, it includes: A differentiable physical simulation structure based on the electromagnetic propagation equation is constructed. The channel parameters and hardware parameters extracted from the original dataset in the previous step are used as inputs, and the basic distortion response is output. The residual learning network is trained using the original dataset. The residual learning network takes the basic distortion response of the previous stage and the corresponding actual measurement signal as input and learns to output high-dimensional residual features. The basic distortion response and high-dimensional residual features are weighted and fused to form a joint distortion characterization with continuously adjustable parameters; The adversarial distortion mode generation module takes the joint distortion representation and the internal state information of the current end-to-end communication process as input, adjusts multiple parameters in the joint distortion representation in real time, and generates an adversarial distortion mode that reduces the performance of the current communication process. The adversarial iterative training module uses an adversarial distortion mode to train the end-to-end communication process in multiple rounds. In each round, the parameters of the communication process are first optimized with the current distortion mode, and then the generation strategy of the distortion mode is updated based on the performance feedback of the optimized communication process until the communication process reaches a stable state under various distortion conditions. The intelligent communication deployment module applies the end-to-end communication process learned through training to the actual wireless communication link, completing the entire process of intelligent transmission and recovery from sending to receiving signals. Specifically, this includes: Monitor the real-time environmental status of the actual wireless communication link and extract the physical channel characteristics and hardware operating point of the current link; Based on the real-time environmental state, the system adaptively selects or fuses parameter combinations that match the current conditions from various parameter configurations accumulated during the communication process in the training process. In actual wireless communication links, signal transmission is performed using an end-to-end communication process initialized with the selected parameter combination, and its transmission effect data is collected in real time. Based on transmission performance data, the parameter combinations are adjusted online until the performance indicators of intelligent transmission and recovery throughout the entire process reach the preset optimal working threshold.

Citation Information

Patent Citations

  • Electric power communication multi-route intelligent planning method and device based on scene classification

    CN120811964A

  • End-to-end optimization method, signal constellation geometry and probability joint shaping optimization method in end-to-end intelligent communication system, and communication device based on neural network

    CN121283515A