A signal shielding method based on deep reinforcement learning

By employing a signal shielding method based on deep reinforcement learning, combining dynamic feature perception and multi-domain feature fusion layers with interference strategy generation, the accuracy, real-time performance, and power consumption issues of signal shielding technology in multi-frequency, high-density electromagnetic environments are resolved, achieving efficient and intelligent signal shielding effects.

CN120811542BActive Publication Date: 2025-11-21CHINA YOUKE COMM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511254408.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-21
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing signal shielding technologies suffer from limitations when facing multi-band, high-density, and dynamic electromagnetic environments. These limitations include static frequency band configuration, high power consumption, and protocol circumvention risks, making it difficult to achieve accuracy, real-time performance, and scalability. Furthermore, traditional improvement solutions such as spectrum sensing and threshold-triggered power control are ineffective in complex scenarios.

Method used

A signal shielding method based on deep reinforcement learning is adopted. A signal interference strategy generation model is constructed through a dynamic feature perception layer, a multi-domain feature fusion layer and an interference strategy generation layer. Combined with a generative adversarial network to expand the dataset, multi-dimensional feature extraction and interference strategy generation of spectrum, time domain, spatial domain and protocol semantics are realized. The interference frequency and power allocation are dynamically adjusted. Combined with environmental noise compensation and temperature control, power consumption is optimized.

Benefits of technology

It improves the accuracy, real-time performance, and scalability of signal shielding, reduces power consumption, adapts to rapid frequency shifts and complex electromagnetic environments, supports future 6G and quantum-safe communication scenarios, and achieves precise, compliant, and low-power intelligent signal shielding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811542B_ABST
    Figure CN120811542B_ABST
Patent Text Reader

Abstract

The application provides a signal shielding method based on deep reinforcement learning in the technical field of artificial intelligence and communication, and comprises the following steps: S1, a large amount of historical wireless communication signal construction data sets are acquired to train a signal interference strategy generation model; S2, the signal interference strategy generation model is deployed to a signal shielding device; S3, a signal shielding instruction is verified and analyzed to obtain shielding parameters, real-time wireless communication signals are collected, the real-time wireless communication signals and the shielding parameters are input into the deployed signal interference strategy generation model, and real-time signal interference strategies are obtained; and S4, the signal shielding device modulates real-time interference signals based on the real-time signal interference strategies, the real-time interference signals are amplified by a power amplifier, and then are emitted to the outside through a radio frequency antenna, so that signal shielding is performed. The application has the advantages that the accuracy, real-time performance and expansibility of signal shielding are greatly improved, and the power consumption of signal shielding is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and communication technology, and in particular to a signal shielding method based on deep reinforcement learning. Background Technology

[0002] With the rapid development of wireless communication technology, the global mobile communication system has formed a multi-standard coexistence pattern including 2G (GSM / CDMA), 3G (WCDMA / TD-SCDMA), 4G (LTE), and 5G (NR). According to data from the International Telecommunication Union (ITU), current commercial wireless communication frequency bands cover the range of 400MHz to 6GHz, specifically including core frequency bands such as GSM 900MHz / 1800MHz, CDMA 800MHz, 3G 2100MHz, 4G 2600MHz, and 5G 3.5GHz / 4.9GHz. At the same time, the popularization of short-range communication technologies such as Wi-Fi (2.4GHz / 5GHz), Bluetooth (2.4GHz), and the Internet of Things (NB-IoT / eMTC) has made the electromagnetic environment multi-band, high-density, and dynamic, posing a severe challenge to signal shielding technology.

[0003] Currently, signal shielding techniques mainly include signal-to-noise ratio (SNR) degradation modes and downlink signal interference modes. The principle of SNR degradation modes is to inject wideband noise to lower the receiver's Eb / N0 (bit energy to noise density ratio) below the demodulation threshold (e.g., QPSK requires Eb / N0 ≥ 6dB), thereby achieving signal shielding. However, SNR degradation modes suffer from drawbacks such as low interference efficiency (wideband noise injection may lead to significant wasted power consumption, resulting in relatively low interference efficiency), poor real-time performance (difficult to respond quickly to signal changes in high-dynamic environments), and weak anti-interference capability (for signals employing advanced modulation and demodulation techniques (such as spread spectrum), wideband noise injection may not effectively cover their time-frequency resource units). The downlink signal jamming mode is the mainstream approach due to its drawbacks. The principle of the downlink signal jamming mode is to generate a co-frequency interference signal with the same carrier frequency as the target signal and a power boost of 6-10dB, and then cover the target frequency band by sweeping the frequency (typical sweep rate is 50-200MHz / s). This mechanism utilizes the receiver's automatic gain control (AGC) feature to inject interference pulses during the time slot synchronization period of the base station's downlink signal, making it impossible for the mobile terminal to correctly demodulate the DCI format of PBCH (Physical Broadcast Channel) and PDCCH (Physical Downlink Control Channel), thus causing the RRC connection establishment to fail.

[0004] However, signal jammers using downlink signal interference mode face the following technical bottlenecks in actual deployment: (1) Static frequency band configuration: Signal jammers mostly use a pre-set frequency point library, which cannot adapt to the rapid frequency migration brought about by dynamic spectrum allocation (DSA) and carrier aggregation (CA) technology. (2) High power consumption of signal jamming (low interference efficiency): The full-band continuous frequency sweep mode results in a high proportion of invalid power consumption. Actual measurement data shows that the power consumption of a typical 5-band jamming device can reach 120W under continuous operation. The thermal design power (TDP) exceeds the standard, causing the junction temperature of the device to rise to more than 85℃, which seriously affects the MTBF (mean time between failures). (3) Protocol risk avoidance: New communication systems use cognitive radio (CR) technology to realize spectrum sensing and frequency hopping communication. For example, LoRa's CSS (Chirp Spread Spectrum) modulation can achieve a receiving sensitivity of -148dBm in a 125kHz channel. Traditional frequency sweeping interference is difficult to effectively cover the time and frequency resource units corresponding to its spreading factor (SF7-SF12).

[0005] To address the above issues, the industry has attempted several improvement solutions: (a) Frequency band identification technology based on spectrum sensing, such as using USRP (Universal Software Radio Peripheral) to build a frequency sweep receiver and performing spectrum energy detection through FFT (Fast Fourier Transform). However, due to the complexity of FFT calculation and the working mode of the frequency sweep receiver, this method suffers from insufficient real-time performance, with a single full-band scan taking more than 2 seconds, which cannot match the time slot structure of 5G NR (0.5ms time slot period). (b) Threshold-triggered power control, which dynamically adjusts the interference power based on RSSI (Received Signal Strength Indication), but does not consider the beamforming characteristics in multi-user MIMO (Multiple-Input Multiple-Output) scenarios, resulting in a mismatch between the interference beam and the target user's spatial dimension. (c) Protocol feature matching interference, such as selectively interfering with the CFI (Control Format Indication) field of PDCCH, but this requires deep parsing of the DCI format and faces compatibility risks brought about by the frequent evolution of 3GPP standards.

[0006] Therefore, how to provide a signal masking method based on deep reinforcement learning to improve the accuracy, real-time performance, and scalability of signal masking while reducing its power consumption has become an urgent technical problem to be solved. Summary of the Invention

[0007] The technical problem to be solved by this invention is to provide a signal masking method based on deep reinforcement learning, which improves the accuracy, real-time performance and scalability of signal masking, and reduces the power consumption of signal masking.

[0008] This invention provides a signal masking method based on deep reinforcement learning, comprising the following steps:

[0009] Step S1: Construct a signal interference strategy generation model based on the dynamic feature perception layer, the multi-domain feature fusion layer, and the interference strategy generation layer, and set the loss function of the signal interference strategy generation model;

[0010] The dynamic feature perception layer is constructed based on a spectrum dynamic coding module, a time-domain feature coding module, a spatial-domain feature coding module, a protocol semantic coding module, and a signal strength coding module. The spectrum dynamic coding module, the time-domain feature coding module, the spatial-domain feature coding module, the protocol semantic coding module, and the signal strength coding module are respectively used to extract spectrum features, time features, spatial features, communication protocol features, and signal strength features from the input wireless communication signal.

[0011] The multi-domain feature fusion layer is constructed based on a feature interaction enhancement module and a dynamic weight allocation module. The feature interaction enhancement module is used to construct a three-dimensional correlation matrix based on spectral features, temporal features, and spatial features. The dynamic weight allocation module is used to evaluate the contribution of the three-dimensional correlation matrix, communication protocol features, and signal strength features to perform fusion and obtain fused features.

[0012] The interference strategy generation layer is constructed based on a multi-objective decision-making module, a beam optimization module, and an output module. The multi-objective decision-making module is used to generate an interference frequency selection matrix and a power level vector based on the fusion characteristics. The beam optimization module is used to generate beamforming weights based on the fusion characteristics. The output module is used to generate a signal interference strategy based on the interference frequency selection matrix, the power level vector, and the beamforming weights.

[0013] Step S2: Acquire a large number of historical wireless communication signals, preprocess and label each of the historical wireless communication signals to construct a signal dataset, and expand the sample size of the signal dataset by using a generative adversarial network.

[0014] Step S3: Divide the signal dataset into a training set, a validation set, and a test set, and train, validate, and test the signal interference strategy generation model using the training set, validation set, and test set.

[0015] Step S4: Compress the signal interference strategy generation model that has passed the test, and deploy the compressed signal interference strategy generation model to the signal jammer.

[0016] Step S5: After the signal jammer is started, the input signal jamming command is obtained, the signal jamming command is verified and parsed to obtain the jamming parameters, the real-time wireless communication signal is collected, the real-time wireless communication signal is preprocessed, and the real-time wireless communication signal and the jamming parameters are input into the deployed signal interference strategy generation model to obtain the real-time signal interference strategy.

[0017] Step S6: The signal jammer modulates the corresponding real-time interference signal based on the real-time signal interference strategy. After the power amplifier performs power amplification on the real-time interference signal through digital predistortion technology, it transmits the real-time interference signal to the outside through the radio frequency antenna to perform signal directional shielding operation.

[0018] Step S7: The signal jammer collects environmental noise data and equipment temperature in real time, performs power compensation on the real-time interference signal based on the environmental noise data, and dynamically adjusts the working status of the power amplifier and fan based on the equipment temperature.

[0019] Step S8: The signal jammer records the signal jamming log in real time, and performs federated optimization of the signal interference strategy generation model based on the signal jamming log.

[0020] The advantages of this invention are:

[0021] 1. A signal jamming strategy generation model is constructed through a dynamic feature perception layer, a multi-domain feature fusion layer, and a jamming strategy generation layer. A large number of historical wireless communication signals are acquired to construct a signal dataset, and the sample size of the signal dataset is expanded by a generative adversarial network. The signal dataset is divided into training, validation, and test sets to train, validate, and test the signal jamming strategy generation model. The tested signal jamming strategy generation model is compressed and deployed to a signal jammer. After the signal jammer is activated, the input signal jamming command is verified and parsed to obtain the jamming parameters. Real-time wireless communication signals are collected, preprocessed, and then transmitted to the signal jammer. The signal jamming strategy generation model is deployed based on the shielding parameters to obtain a real-time signal jamming strategy. Then, the signal jammer modulates the corresponding real-time jamming signal based on the real-time jamming strategy. The power amplifier performs power amplification on the real-time jamming signal using digital predistortion technology, and then transmits the real-time jamming signal externally through the radio frequency antenna to perform directional signal jamming. The signal jammer collects environmental noise data and equipment temperature in real time, performs power compensation on the real-time jamming signal based on the environmental noise data, dynamically adjusts the operating status of the power amplifier and fan based on the equipment temperature, and records the signal jamming log in real time. The signal jamming strategy generation model is then federated and optimized based on the signal jamming log. In other words, the signal jamming strategy generation model generates a signal jamming strategy through a pre-trained model. The system generates signal interference strategies and corresponding interference signals for signal shielding. The dynamic feature perception layer of the signal interference strategy generation model extracts frequency migration features and carrier aggregation patterns, and combines this with the interference frequency selection matrix generated by the multi-objective decision module of the interference strategy generation layer. This allows for dynamic identification and adaptation to rapid frequency migration, rather than relying on a pre-set frequency library. The three-dimensional correlation matrix in the multi-domain feature fusion layer models the cross-domain dependencies of spectrum, time, and space, adjusting the interference frequency in real time to adapt to dynamic spectrum allocation scenarios. The interference strategy generation layer achieves precise power allocation through power level vectors and beamforming weights, avoiding full-band frequency sweeping. Power compensation for real-time interference signals is performed using environmental noise data, based on equipment... Temperature-based dynamic adjustment of the power amplifier's operating state further optimizes power consumption and reduces ineffective radiation. Resource consumption loss in the loss function directly constrains the resource usage of power and beam weights, ensuring low power consumption is the primary goal during model training. Communication protocol features are extracted through the protocol semantic encoding module of the dynamic feature perception layer, and combined with signal type labeling, these features can be effectively identified. Signal interference strategies generated by the interference strategy generation layer target specific protocol features, and federated tuning continuously adapts to protocol evolution. Compression of the signal interference strategy generation model before deployment effectively reduces model size and computational load, ultimately significantly improving the accuracy, real-time performance, and scalability of signal shielding while greatly reducing its power consumption.

[0022] 2. By integrating five types of coding modules—spectrum, time domain, spatial domain, protocol semantics, and signal strength—through the dynamic feature perception layer, holographic feature extraction of wireless signals can be achieved, avoiding the limitations of single feature analysis. For example, protocol semantics coding can identify vulnerabilities in specific communication protocols (such as Wi-Fi 6 and 5G NR), improving the targeting of interference.

[0023] 3. By modeling the multi-domain correlation of signals through a three-dimensional correlation matrix (spectrum × time × space) and combining it with a dynamic weight allocation module to evaluate the feature contribution, feature fusion for adaptive scenarios is achieved. This design is significantly better than traditional static weight fusion, especially in complex electromagnetic environments (such as dense urban multipath scenarios).

[0024] 4. By employing a multi-objective decision-making module and a beam optimization module in the interference strategy generation layer, the interference frequency selection matrix, power level vector, and beamforming weights are simultaneously optimized to balance interference efficiency and energy consumption control. For example, in low signal-to-noise ratio scenarios, high-frequency interference is prioritized while reducing ineffective power allocation.

[0025] 5. By expanding the signal dataset through Generative Adversarial Networks (GANs), the problem of insufficient rare signal samples in real-world scenarios (such as sudden emergency communication signals) is effectively solved, improving the model's generalization ability; by aggregating the signal blocking logs of multiple signal jammers through a federated tuning mechanism to update the signal interference strategy generation model, knowledge sharing and privacy protection can coexist; for example, signal jammers in different regions can collaboratively optimize the protocol semantic recognition module without uploading the original signal data; by combining deep reinforcement learning strategies and digital predistortion (DPD) technology in the interference signal generation stage, the nonlinear distortion of the power amplifier is suppressed, ensuring the focus of the interference signal in the target frequency band and reducing the risk of out-of-band radiation.

[0026] 6. By linking real-time environmental noise compensation with temperature control, dynamic adjustment of interference power and thermal management of equipment can be achieved; for example, in high-temperature environments, the power amplifier output is automatically reduced and heat dissipation is accelerated to prevent the signal jammer from shutting down.

[0027] 7. By forming a millisecond-level closed loop from real-time wireless communication signal acquisition to real-time interference signal transmission, and by combining the preprocessing of real-time wireless communication signals with a lightweight model deployment, the real-time tracking requirements for frequency hopping signals can be met.

[0028] 8. Through the dynamic collaboration of deep reinforcement learning and multi-domain feature fusion (spectrum, time domain, spatial domain, protocol semantics, and signal strength), high-precision signal interference strategy generation is achieved. Its core advantages are: multi-dimensional perception (holographic feature extraction and three-dimensional correlation matrix modeling), dynamic optimization (feature weight allocation and multi-objective decision-making based on attention mechanism), real-time closed-loop control (millisecond-level interference strategy generation and digital predistortion power compensation), and edge intelligent deployment (combination of model compression and federated tuning technology). It also has high energy efficiency (dynamic power adjustment and temperature adaptive management), strong robustness (anti-adversarial attack design and hardware fault tolerance mechanism), and environmental friendliness (directional beam reduces electromagnetic pollution). It can be expanded to support future 6G and quantum secure communication scenarios, and achieve accurate, compliant, and low-power intelligent signal shielding in military, civilian, and emergency scenarios.

[0029] 9. The signal is represented in a multi-dimensional way through five coding modules: spectrum, time domain, spatial domain, protocol semantics, and signal strength. It covers the frequency domain dynamics (multi-scale convolution + gated loop), time domain burstiness (time attention), spatial domain topology (graph attention), protocol standardization, and strength dynamics. The feature interaction enhancement module uses a hypergraph neural network to model cross-domain correlation. The constructed three-dimensional correlation matrix breaks through the limitations of traditional two-dimensional feature interaction and effectively captures the correlation of time-frequency-space joint domain features.

[0030] 10. The spectrum dynamic coding module innovatively integrates local features (multi-scale convolution) and global dependencies (multi-head self-attention), captures migration patterns through frequency point clustering, and achieves refined analysis of frequency point dynamic features by combining a multi-resolution matching mechanism; the signal strength coding adopts an adaptive normalization layer to realize the dynamic mapping between RSSI and MCS, overcoming the problem of insufficient adaptability of traditional fixed normalization to rapid changes in signal strength.

[0031] 11. The dynamic weight allocation module realizes the dynamic evaluation of feature contribution through the policy gradient method and uses a value network for online weight adjustment, which is more adaptable to the time-varying characteristics of complex interference scenarios compared with the fixed weight fusion strategy; the multi-objective decision module realizes the joint optimization of interference frequency selection (discrete decision) and power level adjustment (continuous decision) through a dual-branch network structure, avoiding the error accumulation problem of traditional cascaded decision.

[0032] 12. By combining a differentiable convex optimization layer with the SDP relaxation method, gradient differentiability is achieved while ensuring the physical constraints of beamforming, breaking through the bottleneck of traditional convex optimization methods that cannot be trained end-to-end; by introducing the main lobe direction cosine similarity metric into the beam optimization loss function, it is more consistent with the spatial distribution characteristics of beams than Euclidean distance.

[0033] 13. By designing the loss function to integrate four objectives—interference effect (cross-entropy + power deviation), resource consumption (L1 + F norm), beam direction (cosine similarity), and feature fusion error—and optimizing the multi-objective trade-off coefficients through grid search, a system-level performance balance is achieved. By introducing a dynamic regularization mechanism (λ1, λ2, λ3 coefficients), overfitting of the model is effectively prevented. The Lres loss term simultaneously constrains power consumption and beam energy dispersion, reflecting the concept of green communication.

[0034] 14. By innovatively integrating the BERT pre-trained model and the 3GPP knowledge graph through the protocol semantic coding module, the semantic understanding capabilities of large-scale language models are utilized, while prior knowledge of communication standards is injected, significantly improving the accuracy of protocol feature parsing; by introducing carrier aggregation pattern matching in frequency point migration modeling, physical layer features and higher-layer protocol states are organically combined, enhancing adaptability to modern communication systems (such as 5G NR).

[0035] 15. The aggregation parsing unit decouples frequency migration features and carrier aggregation modes through the feature reconstruction subunit, realizing the separation of physical layer dynamic features and protocol layer configuration features, and enhancing the interpretability of the model for complex signal interference scenarios; the spatial domain feature coding module independently extracts signal topology (graph attention network) and beam spatial modes (spatial spectrum estimation), avoiding parameter redundancy problems caused by spatial feature coupling in traditional methods.

[0036] 16. By integrating multi-dimensional feature fusion (spectrum dynamic coding, temporal feature coding, spatial feature coding, protocol semantic coding, and signal strength coding) with a reinforcement learning dynamic decision-making mechanism, refined analysis of signal features and intelligent interference decision-making are achieved. Its core advantages are reflected in: innovatively combining multi-scale convolution, gated recurrent units, and multi-head attention mechanisms to construct a cross-domain dynamic feature coding network, breaking through the limitations of traditional single-domain feature modeling; using differentiable convex optimization layers and hypergraph neural networks to achieve joint optimization of beamforming and feature association, improving computational efficiency while ensuring physical constraints; and effectively balancing interference performance, resource efficiency, and system robustness through dynamic weight allocation and multi-objective loss function design, while also possessing 3GPP multi-protocol compatibility, hardware deployment friendliness, and cross-scenario migration capabilities, providing a high-precision, low-latency, green and energy-saving end-to-end solution for intelligent spectrum countermeasures in complex electromagnetic environments.

[0037] 17. By annotating the signal dataset, physical layer features such as frequency shifting, carrier aggregation, signal burst characteristics, and time-frequency interleaving are covered, while protocol layer information such as communication protocols and signal types are also included, enabling full-dimensional signal modeling. Through feature annotation such as beam space mode vectors and signal interference strategies, the signal characteristics of complex communication scenarios such as multi-antenna systems (MIMO) and dynamic spectrum sharing can be effectively analyzed.

[0038] 18. By amplifying, filtering, and down-converting the signal, environmental noise and hardware distortion are eliminated, and the signal-to-noise ratio (SNR) is improved, providing high-fidelity data for subsequent annotation and modeling.

[0039] 19. By using multi-dimensional signal feature annotation (covering parameters such as frequency shifting, carrier aggregation, and communication protocols from the physical layer to the protocol layer) and standardized preprocessing (signal amplification, filtering, and down-conversion), the accuracy of data representation and signal-to-noise ratio are significantly improved. Combined with pre-trained generative adversarial networks (GANs) to generate adversarial samples, the bottleneck of scarce samples in real communication signal acquisition is overcome, and the robustness of AI models to multiple scenarios (such as channel fading and multi-device interference) is enhanced.

[0040] 20. By using k-fold cross-validation combined with an 8:1:1 ratio to divide the dataset, both data utilization efficiency and generalization ability verification are taken into account. k-fold cross-validation effectively reduces the randomness bias of a single data split and improves model stability. The golden ratio of 8:1:1 ensures that the training data is sufficient (80%) while the validation set (10%) and the test set (10%) are clearly defined in terms of independence. The dual validation mechanism (validation set + test set) forms a closed-loop feedback, which significantly reduces the risk of overfitting.

[0041] 21. By establishing a probabilistic model through a Bayesian optimization framework to guide the parameter search direction, the search efficiency is effectively improved compared with grid search / random search. At the same time, it optimizes heterogeneous parameters such as learning rate, regularization parameters, and network structure parameters, breaking through the limitations of traditional single-parameter tuning. It introduces random dropout rate as an optimization variable, creatively solving the parameter adaptation problem of Dropout mechanism in deep learning.

[0042] 22. Ensure basic model performance meets standards by setting a preset loss threshold; prevent ineffective training resource consumption by setting early stopping conditions, which is particularly suitable for complex models in high-dimensional parameter spaces; automatically expand the training set when validation fails, achieving synergistic optimization of data augmentation and model training.

[0043] 23. By innovatively integrating k-fold cross-validation with an 8:1:1 data stratification mechanism and combining it with a Bayesian optimization framework for intelligent collaborative tuning of multi-dimensional hyperparameters, a closed-loop training system was constructed, including dynamic early stopping thresholds and adaptive expansion of the training set. At the same time, multi-dimensional evaluation indicators such as frequency hit rate and false interference rate were introduced in the testing phase. This effectively solved the technical pain points of low training efficiency and insufficient generalization ability of traditional interference strategy models, significantly improved the model convergence speed, policy generation accuracy, and anti-interference performance in practical applications, especially in electromagnetic countermeasures scenarios, achieving a balance between the accuracy of interference strategies and the efficiency of resource utilization.

[0044] 24. By adopting an absolute value sorting + threshold zeroing pruning strategy, the number of model parameters is effectively reduced while retaining core features. The threshold is dynamically adjusted through a loop verification mechanism to solve the problem of erroneous deletion of key weights caused by traditional pruning. By introducing an offset compensation quantization algorithm (superior to conventional linear quantization) and optimizing the step size parameter through a closed loop of secondary performance verification, the compression ratio is effectively improved while reducing the loss of model accuracy.

[0045] 25. By using two-stage model compression (adaptive pruning based on dynamic threshold adjustment and quantization parameter optimization), the complexity of the signal interference strategy generation model is effectively reduced. Combined with containerized deployment and dual drift compensation training mechanism, the computational efficiency and resource utilization are significantly improved while ensuring model accuracy. Its multi-level verification system and modular design enhance the robustness of the system in complex environments such as electromagnetic interference, providing a high-efficiency deployment solution for signal jammers in resource-constrained scenarios.

[0046] 26. By using timestamp-segmented concatenation of the master key to generate a temporary key (the first 6 bits are used for parameter ciphertext, and the last 6 bits are used for verification ciphertext), a dynamic key generation mechanism is implemented, ensuring that each encryption key is different, effectively preventing system-wide security risks caused by key leakage; the parameter ciphertext uses MD5+AES-128 double encryption, and the verification ciphertext uses a combination of three encryption algorithms: HMAC+SM4+ChaCha20, forming a heterogeneous encryption protection layer, significantly increasing the difficulty of brute-force attacks; obfuscation methods such as character shifting (right shift by 7 bits) and alphanumeric substitution (swapping 6 / 9, A / C, 8 / B) are used to increase the complexity of ciphertext feature analysis.

[0047] 27. Verify the consistency of original data such as account, password and timestamp through hash value comparison (SHA256) to prevent tampering during transmission; ensure the timeliness of instructions through timestamp verification to avoid replay attacks of historical instructions; establish operation permission control through dual authentication of account and password to prevent unauthorized devices from accessing the system.

[0048] 28. By appending 4 random letters to the end of the UTF-8 encoded data, the fixed structure of the data block is disrupted, increasing the difficulty of ciphertext pattern analysis; by setting the verification ciphertext to use three-level encryption processing (SM4 encryption → hexadecimal conversion → character shift → character replacement → ChaCha20 encryption), a multi-layered encryption nested structure is formed, which greatly increases the cost of reverse analysis.

[0049] 29. By integrating a dynamic key generation mechanism (timestamp segmentation of the master key combined with MD5 / HMAC algorithms) with heterogeneous encryption algorithms (AES-128 / SM4 / ChaCha20 collaborative encryption), multiple anti-cracking barriers are constructed. At the same time, a multi-dimensional security verification system of hash verification, timestamp verification, and account password authentication is adopted to ensure the integrity, timeliness, and legality of instruction transmission. Its innovative anti-reverse design (random data expansion, character shifting and permutation, and three-level encryption nesting) significantly improves the difficulty of ciphertext parsing. Combined with a smart signal interference model to dynamically generate beamforming strategies, it achieves accurate signal shielding while taking into account system security and energy efficiency optimization.

[0050] 30. By using a preset table of noise intensity and power amplifier gain, the ambient noise intensity can be matched in real time, and the gain of the power amplifier can be dynamically adjusted to achieve accurate interference signal compensation. This avoids the problems of overcompensation (energy waste) or undercompensation (shielding failure) in the traditional fixed gain mode. Based on the hierarchical compensation strategy of noise intensity, the power amplifier power is increased only when necessary, reducing the overall energy consumption of the signal shield while maintaining a stable signal shielding effect.

[0051] 31. By setting up a graded heat dissipation mechanism, in the active heat dissipation stage (temperature < threshold): the heat dissipation intensity is dynamically adjusted according to the functional relationship between equipment temperature and fan speed, avoiding continuous full-speed fan operation and reducing mechanical wear and noise; in the emergency protection stage (temperature ≥ threshold): the fan is forced to run at full speed, while the power amplifier bias voltage is gradually reduced according to the voltage gradient, and the temperature is rapidly reduced through the dual path of "reducing the power of the heat source + enhancing heat dissipation" to prevent hardware overheating damage; the dynamic temperature control strategy reduces the impact of drastic temperature fluctuations on electronic components, inhibits material aging caused by long-term high temperature, and improves equipment reliability.

[0052] 32. By directly matching noise intensity and gain value through a pre-stored lookup table, the computational overhead of real-time calculation of power amplifier parameters is avoided, reducing processor load and improving response speed; by encapsulating the relationship between fan speed and temperature into a configurable function, the control logic is simplified, and parameters (such as function slope) can be flexibly adjusted to adapt to different heat dissipation design requirements.

[0053] 33. By gradually reducing the power amplifier bias voltage through a preset voltage gradient, signal distortion or circuit impact caused by sudden voltage drops can be avoided, ensuring a stable transition of the system to a safe state; a redundant protection layer is provided through a temperature threshold triggering mechanism: fan speed adjustment is used for cooling first, and power amplifier voltage adjustment is only intervened in extreme cases to maintain the shielding function to the maximum extent.

[0054] 34. Adaptive signal compensation is achieved through dynamic matching of noise intensity and power amplifier gain, optimizing energy consumption while ensuring shielding effectiveness; combined with the equipment temperature graded control mechanism, functional fan speed adjustment and gradient power amplifier voltage control work together to dissipate heat, avoiding the risk of hardware overheating and extending equipment life; its lookup table method and preset function design significantly reduce computational complexity and improve real-time response capability, while parameter configurability and modular architecture give the solution high scalability, ultimately achieving a multi-dimensional balance of performance, efficiency and reliability in complex electromagnetic environments.

[0055] 35. By recording multi-dimensional data in real time (shielding parameters, real-time wireless communication signals, real-time signal interference strategies, real-time interference signals, environmental noise data, equipment temperature, shielding time, account information, and shielding effect feedback), a complete profile of the equipment's operating status is formed, providing high-quality feature input for the signal interference strategy generation model and enhancing the accuracy of the interference strategy.

[0056] 36. By adopting a data volume threshold triggering mechanism, resource waste caused by frequent training is avoided; system reliability is ensured through model backup and version rollback design to prevent service interruption due to training failure; and the utilization rate of computing resources is optimized through an intelligent scheduling strategy for training during idle periods.

[0057] 37. By replacing the original data transmission with parameter aggregation, the data privacy protection requirements are met; by participating in global model optimization through distributed signal jammers, the generalization ability of the model to different environments (such as electromagnetic environments in different regions) is improved; and by realizing continuous model evolution through edge-cloud collaborative architecture, the limitation of insufficient data samples from a single device is overcome.

[0058] 38. By collecting multi-dimensional signal jamming logs in real time to build an incremental dataset, and combining a threshold-based triggering mechanism and idle time scheduling strategy, low-power and high-reliability incremental training of the local model is achieved. Innovatively, a federated learning architecture is introduced. Under the premise of protecting data privacy, the parameter aggregation of distributed devices is used to achieve global model co-evolution, forming a closed-loop optimization system of "environmental perception-policy generation-effect feedback-federated tuning". It has dynamic environment adaptability (such as noise adaptation and heat dissipation linkage), system stability guarantee (version rollback and verification mechanism) and multi-device collaborative generalization advantage, which significantly improves the accuracy and scenario universality of signal interference strategy.

[0059] 39. Accurate signal feature extraction and cross-domain correlation are achieved through a dynamic feature perception layer (integrating five-dimensional features of spectrum, time domain, spatial domain, protocol semantics, and signal strength) and a multi-domain feature fusion layer (based on hypergraph neural network and reinforcement learning dynamic weight allocation). Efficient interference strategies are generated by combining a multi-objective decision module (two-branch network to generate interference frequency points and power strategies) and a beam optimization module (differentiable convex optimization to improve directivity). Generative adversarial networks are used to enhance data diversity, and lightweight adaptation is achieved through model pruning, quantization, and containerized deployment. At the same time, multiple encryption (AES-128, SM4, ChaCha20) and dynamic verification mechanisms are used to ensure command security. The model is optimized in real time based on the federated learning framework, and it also has the capabilities of environmental noise adaptive compensation, intelligent temperature control, and log traceability. Finally, it forms significant advantages in interference accuracy (frequency hit rate), resource utilization (power dynamic optimization), real-time response (millisecond-level inference), anti-countermeasure capability (narrow beam directionality), and compliance (frequency constraints), achieving intelligent, low-power, and highly concealed directional signal shielding. Attached Figure Description

[0060] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0061] Figure 1 This is a flowchart of a signal masking method based on deep reinforcement learning according to the present invention. Detailed Implementation

[0062] The overall approach of the technical solution in this application is as follows: A pre-trained signal interference strategy generation model is used to generate a signal interference strategy. Based on this strategy, corresponding interference signals are generated to perform signal shielding. Because the dynamic feature perception layer of the signal interference strategy generation model extracts frequency point migration features and carrier aggregation modes, combined with the interference frequency point selection matrix generated by the multi-objective decision module of the interference strategy generation layer, it can dynamically identify and adapt to rapid frequency point migration, rather than relying on a pre-set frequency point library. The three-dimensional correlation matrix in the multi-domain feature fusion layer models the cross-domain dependencies of spectrum, time, and space, adjusting the interference frequency points in real time to adapt to dynamic spectrum allocation scenarios. The interference strategy generation layer achieves precise power allocation through power level vectors and beamforming weights, avoiding full-band frequency sweeping. Through a loop... Environmental noise data is used to compensate for the power of real-time interference signals. The power amplifier's operating state is dynamically adjusted based on the device temperature to further optimize power consumption and reduce ineffective radiation. The resource consumption loss in the loss function directly constrains the resource usage of power and beam weights, ensuring that the model training is aimed at low power consumption. The protocol semantic encoding module of the dynamic feature perception layer extracts communication protocol features and, combined with signal type labeling, can effectively identify communication protocol features. The signal interference strategy generated by the interference strategy generation layer targets specific protocol features and continuously adapts to protocol evolution through federated tuning. By compressing the signal interference strategy generation model before deployment, the model size and computational load are effectively reduced, thereby improving the accuracy, real-time performance, and scalability of signal shielding and reducing the power consumption of signal shielding.

[0063] Please refer to Figure 1 As shown, a preferred embodiment of the signal masking method based on deep reinforcement learning of the present invention includes the following steps:

[0064] Step S1: Construct a signal interference strategy generation model based on the dynamic feature perception layer, the multi-domain feature fusion layer, and the interference strategy generation layer, and set the loss function of the signal interference strategy generation model;

[0065] The dynamic feature perception layer is constructed based on a spectrum dynamic coding module, a time-domain feature coding module, a spatial-domain feature coding module, a protocol semantic coding module, and a signal strength coding module. The spectrum dynamic coding module, the time-domain feature coding module, the spatial-domain feature coding module, the protocol semantic coding module, and the signal strength coding module are respectively used to extract spectrum features, time features, spatial features, communication protocol features, and signal strength features from the input wireless communication signal.

[0066] By integrating five types of coding modules—spectrum, time domain, spatial domain, protocol semantics, and signal strength—into a dynamic feature perception layer, holographic feature extraction of wireless signals can be achieved, avoiding the limitations of single feature analysis. For example, protocol semantic coding can identify vulnerabilities in specific communication protocols (such as Wi-Fi 6 and 5G NR), improving the targeting of interference.

[0067] The multi-domain feature fusion layer is constructed based on a feature interaction enhancement module and a dynamic weight allocation module. The feature interaction enhancement module is used to construct a three-dimensional correlation matrix based on spectral features, temporal features, and spatial features. The dynamic weight allocation module is used to evaluate the contribution of the three-dimensional correlation matrix, communication protocol features, and signal strength features to perform fusion and obtain fused features.

[0068] By modeling the multi-domain correlation of signals using a three-dimensional correlation matrix (spectrum × time × space) and combining it with a dynamic weight allocation module to evaluate feature contribution, adaptive feature fusion for specific scenarios is achieved. This design is significantly superior to traditional static weight fusion, especially in complex electromagnetic environments (such as dense urban multipath scenarios).

[0069] The interference strategy generation layer is constructed based on a multi-objective decision-making module, a beam optimization module, and an output module. The multi-objective decision-making module is used to generate an interference frequency selection matrix and a power level vector based on the fusion characteristics. The beam optimization module is used to generate beamforming weights based on the fusion characteristics. The output module is used to generate a signal interference strategy based on the interference frequency selection matrix, the power level vector, and the beamforming weights.

[0070] By employing a multi-objective decision-making module and a beam optimization module in the interference strategy generation layer, the interference frequency selection matrix, power level vector, and beamforming weights are simultaneously optimized to balance interference efficiency and energy consumption control. For example, in low signal-to-noise ratio scenarios, high-frequency interference is prioritized while reducing ineffective power allocation.

[0071] Step S2: Acquire a large number of historical wireless communication signals, preprocess and label each of the historical wireless communication signals to construct a signal dataset, and expand the sample size of the signal dataset by using a generative adversarial network.

[0072] Step S3: Divide the signal dataset into a training set, a validation set, and a test set, and train, validate, and test the signal interference strategy generation model using the training set, validation set, and test set.

[0073] Step S4: Compress the signal interference strategy generation model that has passed the test, and deploy the compressed signal interference strategy generation model to the signal jammer.

[0074] Step S5: After the signal jammer is started, the input signal jamming command is obtained, the signal jamming command is verified and parsed to obtain the jamming parameters, the real-time wireless communication signal is collected, the real-time wireless communication signal is preprocessed, and the real-time wireless communication signal and the jamming parameters are input into the deployed signal interference strategy generation model to obtain the real-time signal interference strategy.

[0075] Step S6: The signal jammer modulates the corresponding real-time interference signal based on the real-time signal interference strategy. After the power amplifier performs power amplification on the real-time interference signal through digital predistortion technology, it transmits the real-time interference signal to the outside through the radio frequency antenna to perform signal directional shielding operation.

[0076] By forming a millisecond-level closed loop from real-time wireless communication signal acquisition to real-time interference signal transmission, and by preprocessing the real-time wireless communication signal and deploying a lightweight model in coordination, the real-time tracking requirements for frequency hopping signals can be met.

[0077] Step S7: The signal jammer collects environmental noise data and equipment temperature in real time, performs power compensation on the real-time interference signal based on the environmental noise data, and dynamically adjusts the working status of the power amplifier and fan based on the equipment temperature.

[0078] By linking real-time environmental noise compensation with temperature control, dynamic adjustment of interference power and thermal management of equipment can be achieved; for example, in high-temperature environments, the power amplifier output is automatically reduced and heat dissipation is accelerated to prevent the signal jammer from shutting down.

[0079] Step S8: The signal jammer records the signal jamming log in real time, and performs federated optimization of the signal interference strategy generation model based on the signal jamming log.

[0080] By expanding the signal dataset through Generative Adversarial Networks (GANs), the problem of insufficient rare signal samples in real-world scenarios (such as sudden emergency communication signals) is effectively addressed, thereby improving the model's generalization ability. A federated tuning mechanism aggregates signal jamming logs from multiple signal jammers to update the signal interference strategy generation model, achieving both knowledge sharing and privacy protection. For example, signal jammers in different regions can collaboratively optimize the protocol semantic recognition module without uploading the original signal data. By combining deep reinforcement learning strategies and digital predistortion (DPD) technology in the interference signal generation stage, nonlinear distortion of the power amplifier is suppressed, ensuring the focusing of the interference signal in the target frequency band and reducing the risk of out-of-band radiation.

[0081] By leveraging deep reinforcement learning and the dynamic fusion of multi-domain features (spectrum, time domain, spatial domain, protocol semantics, and signal strength), high-precision signal interference strategy generation is achieved. Its core advantages lie in: multi-dimensional perception (holographic feature extraction and 3D correlation matrix modeling), dynamic optimization (feature weight allocation and multi-objective decision-making based on attention mechanism), real-time closed-loop control (millisecond-level interference strategy generation and digital predistortion power compensation), and edge intelligent deployment (combination of model compression and federated tuning technology). It also features high energy efficiency (dynamic power adjustment and temperature adaptive management), strong robustness (anti-adversarial attack design and hardware fault tolerance mechanism), and environmental friendliness (directional beams reduce electromagnetic pollution). It can be expanded to support future 6G and quantum-secure communication scenarios, enabling precise, compliant, and low-power intelligent signal shielding in military, civilian, and emergency scenarios.

[0082] In step S1, the spectrum dynamic coding module is constructed based on a dynamic coding unit, a frequency domain correlation unit, and an aggregation parsing unit;

[0083] The spectrum dynamic coding module innovatively integrates local features (multi-scale convolution) and global dependencies (multi-head self-attention), captures migration patterns through frequency point clustering, and achieves refined analysis of frequency point dynamic features by combining a multi-resolution matching mechanism. The signal strength coding adopts an adaptive normalization layer to realize the dynamic mapping between RSSI and MCS, overcoming the problem of insufficient adaptability of traditional fixed normalization to rapid changes in signal strength.

[0084] The dynamic coding unit is used to extract local frequency domain features and temporal dynamic features from the input wireless communication signal through multi-scale convolutional subunits and gated recurrent subunits, and output time-frequency joint features including the local frequency domain features and temporal dynamic features. The multi-scale convolutional subunit uses convolutional kernels of different sizes (such as 1×3, 1×5, 1×7) to extract local frequency domain features in parallel, and fuses the multi-scale local frequency domain features through residual connections. The gated recurrent subunit is used to encode the signal time slices of the wireless communication signal to capture the short-term dynamic changes of frequency point migration, i.e., temporal dynamic features.

[0085] The frequency domain association unit is used to divide adjacent frequency points into similar clusters through a frequency point clustering subunit to capture the temporal dependency of frequency point migration from the time-frequency joint features. A multi-head self-attention subunit calculates cross-frequency point attention weights on the time-frequency joint features to model the global dependency between frequency points, thereby capturing the frequency domain correlation of frequency point migration. Based on the temporal dependency and frequency domain correlation, a dynamic frequency point encoding vector is generated. The frequency point clustering subunit is constructed based on a dynamic threshold-gated clustering algorithm. The dynamic frequency point encoding vector is used to characterize the frequency point migration path and the frequency domain association strength.

[0086] The aggregation parsing unit is used to perform multi-resolution matching of carrier aggregation mode on the frequency point dynamic coding vector through the pattern matching subunit, and to decouple the frequency point migration feature from the frequency point dynamic coding vector through the feature reconstruction subunit, and output the spectral features including the frequency point migration feature and the carrier aggregation mode.

[0087] The pattern matching subunit uses a set of dynamic convolution kernels (which can be understood as "templates" of different sizes and shapes) to slide across the frequency point dynamic coding vector, searching for combination patterns of discrete carrier segments at different resolutions, i.e., carrier aggregation patterns. For example, if certain frequency point combinations are found to occur frequently, these combination patterns represent different carrier aggregation patterns (carrier aggregation features). Since frequency point migration features and carrier aggregation patterns are usually mixed together in the frequency point dynamic coding vector, the feature reconstruction subunit uses an orthogonally constrained projection matrix to decouple the frequency point migration features and carrier aggregation features in the frequency point dynamic coding vector, generating dimension-independent joint spectral features (frequency point migration features and carrier aggregation patterns), which can then be analyzed and processed separately.

[0088] The aggregation parsing unit decouples frequency migration features and carrier aggregation modes through the feature reconstruction subunit, thereby separating physical layer dynamic features from protocol layer configuration features and enhancing the interpretability of the model for complex signal interference scenarios. The spatial domain feature encoding module independently extracts signal topology (graph attention network) and beam spatial modes (spatial spectrum estimation), avoiding parameter redundancy problems caused by spatial feature coupling in traditional methods.

[0089] The temporal feature encoding module is used to extract temporal features, including signal burst characteristics and time-frequency interleaving patterns, from the input wireless communication signal through a temporal convolutional network and a multi-head temporal attention mechanism.

[0090] The temporal convolutional network first divides the wireless communication signal into many small segments using a sliding window approach, extracting the local features of each segment. Then, the multi-head temporal attention mechanism identifies and integrates the more important parts from these local features. For example, if the temporal convolutional network finds that several segments of the wireless communication signal have special waveforms, the multi-head temporal attention mechanism will determine whether these special waveforms are related to the overall pattern of the wireless communication signal. Then, it combines this important information to form a more comprehensive temporal feature. In this way, it can extract temporal features from the wireless communication signal that include both signal burst characteristics (local burst characteristics) and time-frequency interleaving patterns (overall patterns).

[0091] The spatial feature encoding module is used to extract spatial features, including signal topology features and beam spatial mode vectors, from the input wireless communication signal through graph attention network and spatial spectrum estimation.

[0092] Wireless communication signals are transmitted through multiple-input multiple-output (MIMO) channels. The MIMO channel matrix contains signal transmission information between multiple transmitting and receiving antennas. Graph attention networks model the transmission relationships of wireless communication signals between these antennas as a graph, where each antenna is a node, and the signal transmission strength and phase information between antennas are represented as edges between nodes. Graph attention networks analyze the relationships between these nodes through an attention mechanism, automatically learning which signal relationships between antennas are more important, and thus assigning higher weights to these important relationships. For example, if the signal transmission between certain antennas is strong and stable, the relationship between these antennas will be given more attention. In this way, graph attention networks can extract the topological features of signal transmission, that is, how wireless communication signals influence and transmit each other between different antennas.

[0093] Spatial spectrum estimation uses the wireless communication signals received by the antenna array to estimate the direction of arrival (DOA). It determines the direction of arrival by analyzing the phase difference and amplitude difference of the wireless communication signal on different antennas. For example, if a wireless communication signal arrives at the antenna array from a specific direction, the wireless communication signals received by different antennas will have a certain phase difference. By calculating these phase differences, spatial spectrum estimation can generate a beam spatial mode vector. This beam spatial mode vector represents the intensity distribution of the wireless communication signal in different directions, which can intuitively reflect the propagation direction and intensity of the wireless communication signal, helping to understand where the wireless communication signal comes from and its distribution in space.

[0094] The signal topology features extracted by the graph attention network and the beam spatial mode vector generated by spatial spectrum estimation are complementary. The graph attention network focuses on the relationships between signals and their propagation paths, while spatial spectrum estimation focuses on the direction and intensity distribution of signals. When the two are combined, the graph attention network can use its extracted topology features to optimize the accuracy of spatial spectrum estimation. For example, if the graph attention network finds that the signal relationships between certain antennas are very important, then spatial spectrum estimation can more accurately estimate the signals in the directions of these antennas. At the same time, the beam spatial mode vector generated by spatial spectrum estimation can also provide the graph attention network with additional contextual information, helping it to better understand the propagation of signals in space.

[0095] The protocol semantic encoding module is used to extract communication protocol features from the input wireless communication signal using BERT units and the 3GPP standard knowledge graph.

[0096] BERT (Bidirectional Encoder Representations from Transformers) is a natural language processing model based on the Transformer architecture. It can process sequential data such as text. In wireless communication signal processing, protocol control fields (such as the CFI field in DCI) can be regarded as a special "text" sequence. These fields contain key information of the communication protocol, such as channel format indication. 3GPP (3rd Generation Partnership Project) is an organization that develops mobile communication standards. Its standard knowledge graph is a structured representation of various protocol rules, parameters and other knowledge in the 3GPP standard. It is like a knowledge base that contains various semantic information of the communication protocol, such as the meaning of different fields and the relationship between fields. The communication protocol feature extraction process is as follows: (1) BERT unit processes protocol control fields: First, the protocol control fields in the wireless communication signal (such as the CFI field in DCI) are input into the BERT unit. The BERT unit encodes these fields, understands the sequence information in the fields, and extracts the semantic features of the fields. For example, the CFI field may contain information about the size of the control area in the subframe. The BERT unit can convert this information into a semantic representation in vector form. This representation can reflect the meaning of the field. (2) Combine the semantic features extracted by the BERT unit with the 3GPP standard knowledge graph for protocol intent reasoning: Combine the semantic features extracted by the BERT unit with the 3GPP standard knowledge graph. The knowledge in the 3GPP standard knowledge graph can further interpret and reason about the semantic features extracted by the BERT unit. For example, the 3GPP standard knowledge graph stores the channel format meanings corresponding to different values ​​of the CFI field. When the BERT unit extracts the semantic features of the CFI field as a specific value, the 3GPP standard knowledge graph can use its stored rules to infer the communication protocol intent corresponding to this value. This intent may be about how to allocate control area resources or how to control data transmission, etc. (3) Extract communication protocol features: Through this collaborative work, communication protocol features are finally extracted from the wireless communication signal. These features include the semantic information and protocol intent of the protocol field, which can help subsequent protocol parsing, signal processing and other operations, thereby better understanding and processing wireless communication signals.

[0097] By innovatively integrating the BERT pre-trained model and the 3GPP knowledge graph through the protocol semantic coding module, the semantic understanding capabilities of large-scale language models are utilized, while prior knowledge of communication standards is injected, significantly improving the accuracy of protocol feature parsing. By introducing carrier aggregation pattern matching in frequency point migration modeling, physical layer features and higher-layer protocol states are organically combined, enhancing adaptability to modern communication systems (such as 5G NR).

[0098] The signal strength encoding module is used to dynamically map the RSSI value and MCS carried in the input wireless communication signal through an adaptive normalization layer, thereby extracting signal strength features.

[0099] The signal strength of wireless communication signals varies due to various factors (such as distance and obstacles). The Received Signal Strength Indicator (RSSI) value is used to measure signal strength. Meanwhile, the Modulation-Coding Scheme (MCS) determines the signal transmission method; different MCSs correspond to different signal quality and transmission efficiency. Since RSSI values ​​can range widely, and the relationship between RSSI and MCS varies under different environments and devices, the task of the adaptive normalization layer is to dynamically adjust this relationship. Specifically, the adaptive normalization layer continuously observes the currently received signals and automatically adjusts a mapping relationship based on the RSSI values ​​and corresponding MCSs. This mapping relationship acts like a "translator," converting RSSI values ​​into a standardized signal strength characteristic, thus facilitating the comparison and analysis of different signal strengths. For example, assuming RSSI values ​​vary between -50 and -100 in one environment and between -30 and -80 in another, the adaptive normalization layer dynamically adjusts the mapping relationship according to these different ranges, ensuring a reasonable and comparable signal strength characteristic in each environment.

[0100] The signal is represented in a multi-dimensional way through five coding modules: spectrum, time domain, spatial domain, protocol semantics, and signal strength. It covers the signal's frequency domain dynamics (multi-scale convolution + gated loop), time domain burstiness (time attention), spatial domain topology (graph attention), protocol standardization, and strength dynamics. The feature interaction enhancement module uses a hypergraph neural network to model cross-domain correlations. The constructed three-dimensional correlation matrix breaks through the limitations of traditional two-dimensional feature interaction and effectively captures the correlation of features in the time-frequency-space joint domain.

[0101] The feature interaction enhancement module is used to model cross-domain feature associations of spectral features, temporal features, and spatial features through a hypergraph neural network, thereby constructing a three-dimensional association matrix;

[0102] A typical graph neural network (Graph Neural Network) is like a network of points and lines, where points represent data and lines represent relationships between points. Hypergraph Neural Networks (HNNs), however, can handle not only points and lines but also more complex structures. For example, a point can have more complex relationships with multiple points simultaneously. This is similar to a social network, which can represent not only relationships between two people but also complex interactions between a group of people. Spectral features, temporal features, and spatial features are originally independent, like three different worlds. A HNN acts like a super "bridge architect," connecting these three worlds by constructing a complex network of "bridges." In this network, each feature is considered a node, and the HNN establishes various connections based on the complex relationships between these features. These connections include not only simple one-to-one relationships but also more complex one-to-many or many-to-many relationships. In this way, the HNN can capture the complex interactions between these features. A three-dimensional correlation matrix can be imagined as a three-dimensional table with three dimensions, corresponding to the spectrum, time, and space. In this table, the value at each position represents the correlation strength between the corresponding spectrum, time, and space features. For example, if a value at a position in the matrix is ​​very high, it means that the correlation between the features is very strong under that specific combination of spectrum, time, and space.

[0103] The dynamic weight allocation module is used to evaluate the contribution of the three-dimensional correlation matrix, communication protocol features, and signal strength features to the current interference scenario through a value network based on the policy gradient method of reinforcement learning, so as to fuse them to obtain fused features.

[0104] Policy gradient methods are a type of reinforcement learning. Reinforcement learning is like training a puppy: rewarding it for good behavior and punishing it for bad behavior. Policy gradient methods learn in this way, making decisions based on current features and the environment, and then adjusting their decision-making strategy based on reward signals, making good decisions more likely to be selected and gradually eliminating bad decisions. The value network acts like an evaluator, assessing the contribution of different features (such as a 3D correlation matrix, communication protocol features, and signal strength features) to the current interference scenario. These features are like different factors; some have a greater impact on the interference scenario than others. The value network assigns a score, informing the dynamic weight allocation module of the importance of each feature, and then weights and sums them based on importance to obtain the fused feature.

[0105] The dynamic weight allocation module achieves dynamic evaluation of feature contribution through the policy gradient method and uses a value network for online weight adjustment, which is more adaptable to the time-varying characteristics of complex interference scenarios compared with the fixed weight fusion strategy. The multi-objective decision module achieves joint optimization of interference frequency selection (discrete decision) and power level adjustment (continuous decision) through a dual-branch network structure, avoiding the error accumulation problem of traditional cascaded decision.

[0106] The multi-objective decision module is used to reason about the fused features through a dual-branch network structure to generate an interference frequency selection matrix and a power level vector.

[0107] The dual-branch network structure is a common neural network architecture that is widely used in multi-task learning, multi-objective decision-making and other scenarios. This structure improves the efficiency and specificity of the model by dividing the network into two independent branches, each responsible for handling different tasks or sub-tasks.

[0108] The beam optimization module uses a differentiable convex optimization layer and SDP relaxation method to infer the fused features in order to generate beamforming weights.

[0109] Beamforming is a wireless communication technology that focuses signals in a specific direction by adjusting the direction of an antenna beam, thereby improving signal strength and reducing interference. Beamforming weights are key parameters controlling the beam direction. Mathematically, differentiability means that a function has a derivative at every point, allowing optimization algorithms such as gradient descent to adjust parameters. In beam optimization, differentiability allows beamforming weights to be adjusted through backpropagation, minimizing errors. Convex optimization is a special type of optimization problem where both the objective function and constraints are convex. A key characteristic of convex optimization is that a local optimum is the global optimum, making the optimization process more efficient and reliable. The role of a differentiable convex optimization layer is to transform the complex beamforming weight calculation problem into a convex optimization problem, thus efficiently finding the optimal solution. Semidefinite programming (SDP) is a special type of convex optimization problem where the variable is a positive semidefinite matrix. In beam optimization, SDP can be used to handle complex constraints, such as power constraints and interference constraints. Relaxation methods are techniques for transforming complex non-convex problems into convex problems. In beam optimization, the original beamforming weight calculation problem may be non-convex and difficult to solve directly. By using the SDP relaxation method, these non-convex constraints can be relaxed into convex constraints, thereby transforming the problem into a solvable SDP problem.

[0110] The beamforming module works as follows: First, the fused features are input into the differentiable convex optimization layer. The convex optimization algorithm is used to process these fused features to obtain a preliminary beamforming weight. Then, the SDP relaxation method is used to further optimize the preliminary beamforming weight to find the optimal beamforming weight.

[0111] By combining a differentiable convex optimization layer with the SDP relaxation method, gradient differentiability is achieved while ensuring the physical constraints of beamforming, thus overcoming the bottleneck of traditional convex optimization methods being unable to train end-to-end. Furthermore, by introducing the main lobe direction cosine similarity metric into the beam optimization loss function, it better matches the spatial distribution characteristics of beams compared to Euclidean distance.

[0112] The formula for the loss function is:

[0113] L=α*L int +β*L res +γ*L beam +δ*L fusion ;

[0114] Where L represents the loss value of the loss function; L int Indicates the loss of interference effect, used to measure the accuracy of interference frequency selection and power allocation; L res L represents resource consumption loss, used to constrain the total transmit power and beamforming weight complexity; beam L represents beam optimization loss, used to measure the deviation of the main lobe direction from the ideal direction; fusion α represents the feature fusion loss, used to constrain the rationality of multi-domain feature fusion; α, β, γ, and δ all represent multi-objective trade-off coefficients, optimized through grid search.

[0115] ;

[0116] ;

[0117] ;

[0118] ;

[0119] Where f represents the frequency point; This indicates the probability that frequency point f is selected; Indicates characteristic functions; Represents the set of target interference frequencies; , as well as Both represent regularization intensity coefficients; This represents the vector prediction value of the power level; Indicates the theoretically optimal power level; Represents the squared Euclidean distance between vectors; Represents the beamforming weight matrix; Represents the L1 norm; Represents the square of the Frobenius norm; Indicates the main lobe prediction direction of the beam; Indicates the ideal direction of the main lobe of the beam; Represents a three-dimensional incidence matrix; Indicates the characteristics of the communication protocol; Indicates signal strength characteristics; Indicates fusion characteristics; This indicates dynamic weighted multiplicative fusion; This indicates dynamic weighted additive fusion.

[0120] By designing the loss function to integrate four objectives—interference effect (cross-entropy + power deviation), resource consumption (L1 + F norm), beam direction (cosine similarity), and feature fusion error—and optimizing the multi-objective trade-off coefficients through grid search, a system-level performance balance is achieved. The introduction of a dynamic regularization mechanism (λ1, λ2, λ3 coefficients) effectively prevents model overfitting. The Lres loss term simultaneously constrains power consumption and beam energy dispersion, reflecting the concept of green communication.

[0121] By integrating multi-dimensional feature fusion (spectrum dynamic coding, temporal feature coding, spatial feature coding, protocol semantic coding, and signal strength coding) with a reinforcement learning dynamic decision-making mechanism, this system achieves refined analysis of signal features and intelligent interference decision-making. Its core advantages are: innovatively combining multi-scale convolution, gated recurrent units, and multi-head attention mechanisms to construct a cross-domain dynamic feature coding network, breaking through the limitations of traditional single-domain feature modeling; employing differentiable convex optimization layers and hypergraph neural networks to achieve joint optimization of beamforming and feature association, improving computational efficiency while ensuring physical constraints; and effectively balancing interference performance, resource efficiency, and system robustness through dynamic weight allocation and multi-objective loss function design, while also possessing 3GPP multi-protocol compatibility, hardware deployment friendliness, and cross-scenario migration capabilities, providing a high-precision, low-latency, and energy-efficient end-to-end solution for intelligent spectrum countermeasures in complex electromagnetic environments.

[0122] Step S2 specifically involves:

[0123] A large number of historical wireless communication signals are acquired, and each of the historical wireless communication signals is preprocessed, including at least signal amplification, filtering, downconversion, and sampling. The preprocessed historical wireless communication signals are then labeled with at least the following characteristics: frequency shift features, carrier aggregation mode, signal burst characteristics, time-frequency interleaving mode, signal topology features, beam space mode vector, communication protocol features, signal strength features, signal type, and signal interference strategy. A dataset is constructed based on the labeled historical wireless communication signals, and the sample size of the signal dataset is expanded by a generator in a pre-trained generative adversarial network.

[0124] The annotations on the signal dataset cover physical layer features such as frequency shifting, carrier aggregation, signal burst characteristics, and time-frequency interleaving, while also incorporating protocol layer information such as communication protocols and signal types, enabling full-dimensional signal modeling. Through feature annotations such as beam space mode vectors and signal interference strategies, the signal characteristics of complex communication scenarios such as multiple-antenna-machine systems (MIMO) and dynamic spectrum sharing can be effectively analyzed.

[0125] By amplifying, filtering, and down-converting the signal, environmental noise and hardware distortion are eliminated, improving the signal-to-noise ratio (SNR) and providing high-fidelity data for subsequent annotation and modeling.

[0126] By using multi-dimensional signal feature annotation (covering parameters such as frequency shifting, carrier aggregation, and communication protocols from the physical layer to the protocol layer) and standardized preprocessing (signal amplification, filtering, and down-conversion), the accuracy of data representation and signal-to-noise ratio are significantly improved. Combined with pre-trained generative adversarial networks (GANs) to generate adversarial samples, the bottleneck of scarce samples in real communication signal acquisition is overcome, and the robustness of AI models to multiple scenarios (such as channel fading and multi-device interference) is enhanced.

[0127] Step S3 specifically involves:

[0128] Based on k-fold cross-validation, the signal dataset is divided into a training set, a validation set, and a test set in an 8:1:1 ratio. The signal interference strategy generation model is trained using the training set. During training, hyperparameters including the learning rate, regularization parameter, network structure parameter, batch size, training epochs, activation function, optimizer, and random dropout rate are optimized using a Bayesian optimization framework until the loss value of the loss function is less than a preset loss threshold or a preset early stopping condition is met.

[0129] The trained signal interference strategy generation model is validated using the validation set to determine whether the accuracy of the interference strategy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded for continued training. If yes, the validation passes, and the following parameters are calculated using the test set: frequency hit rate, false interference rate, power utilization rate, power dynamic adjustment response time, beam pointing accuracy, strategy conflict rate, and single inference time. These parameters are used to test the validated signal interference strategy generation model. If the test fails, the training set is expanded for continued training. If the test passes, training ends.

[0130] By using k-fold cross-validation combined with an 8:1:1 dataset split, the dual requirements of data utilization efficiency and generalization ability verification are met. k-fold cross-validation effectively reduces the randomness bias of a single data split and improves model stability. The golden ratio of 8:1:1 ensures that the training data is sufficient (80%) while the validation set (10%) and test set (10%) are clearly defined in terms of independence verification. The dual validation mechanism (validation set + test set) forms a closed-loop feedback, which significantly reduces the risk of overfitting.

[0131] By establishing a probabilistic model through a Bayesian optimization framework to guide the parameter search direction, the search efficiency is effectively improved compared to grid search / random search. At the same time, it optimizes heterogeneous parameters such as learning rate, regularization parameters, and network structure parameters, breaking through the limitations of traditional single-parameter tuning. It also introduces the random dropout rate as an optimization variable, creatively solving the parameter adaptation problem of the Dropout mechanism in deep learning.

[0132] The model's basic performance is guaranteed by setting a preset loss threshold; ineffective training resource consumption is prevented by setting early stopping conditions, which is particularly suitable for complex models in high-dimensional parameter spaces; the training set is automatically expanded when validation fails, realizing the synergistic optimization of data augmentation and model training.

[0133] By innovatively integrating k-fold cross-validation with an 8:1:1 data stratification mechanism and combining it with a Bayesian optimization framework for intelligent collaborative tuning of multi-dimensional hyperparameters, a closed-loop training system was constructed, including dynamic early stopping thresholds and adaptive expansion of the training set. At the same time, multi-dimensional evaluation indicators such as frequency hit rate and false interference rate were introduced in the testing phase. This effectively solved the technical pain points of low training efficiency and insufficient generalization ability of traditional interference strategy models, significantly improved the model convergence speed, policy generation accuracy, and anti-interference performance in practical applications, especially in electromagnetic countermeasures scenarios, achieving a balance between interference strategy accuracy and resource utilization efficiency.

[0134] Step S4 specifically involves:

[0135] For all weight parameters of the signal interference strategy generation model that has passed the test, sort them according to the absolute value of the weight values. Set the weight values ​​below a preset weight threshold to zero to perform a pruning operation on the signal interference strategy generation model. Perform a performance verification on the pruned signal interference strategy generation model. If the performance verification fails, adjust the weight threshold and repeat the pruning operation; if the performance verification passes, then:

[0136] The signal interference strategy generation model is quantized based on a preset quantization step size and offset. The quantized signal interference strategy generation model is then subjected to a second performance verification. If the performance verification fails, the quantization step size and offset are adjusted and the quantization operation is re-executed. If the performance verification passes, the compression of the signal interference strategy generation model is completed.

[0137] The compressed signal interference strategy generation model is deployed to the signal jammer using containerization technology, and data drift compensation training and model drift compensation training are performed on the signal interference strategy generation model deployed on the signal jammer.

[0138] By employing an absolute value sorting and threshold zeroing pruning strategy, the number of model parameters is effectively reduced while preserving core features. The threshold is dynamically adjusted through a loop verification mechanism to solve the problem of erroneous deletion of key weights caused by traditional pruning. By introducing an offset compensation quantization algorithm (superior to conventional linear quantization) and optimizing the step size parameter through a secondary performance verification closed loop, the compression ratio is effectively improved while reducing the loss of model accuracy.

[0139] The two-stage model compression (adaptive pruning based on dynamic threshold adjustment and quantization parameter optimization) effectively reduces the complexity of the signal interference strategy generation model. Combined with containerized deployment and dual drift compensation training mechanism, it significantly improves computational efficiency and resource utilization while ensuring model accuracy. Its multi-level verification system and modular design enhance the robustness of the system in complex environments such as electromagnetic interference, providing a high-efficiency deployment solution for signal jammers in resource-constrained scenarios.

[0140] Step S5 specifically involves:

[0141] After the signal jammer is activated, it acquires the input signal jamming command carrying encrypted parameters and encrypted verification data. The encrypted parameters are obtained by encrypting the jamming parameters using a first encryption rule. The encrypted verification data is obtained by encrypting the account, password, timestamp, and hash value using a second encryption rule. The hash value is obtained by calculating hash256 from the account, password, and timestamp. The jamming parameters include at least the jamming duration, jamming range, and type of jammed signal.

[0142] The signal jammer parses the signal jamming command to obtain parameter ciphertext and verification ciphertext. It decrypts the verification ciphertext using the second encryption rule to obtain the account, password, timestamp, and hash value. It then performs integrity verification on the account, password, and timestamp using the hash value, performs timeliness verification using the timestamp, and performs legality verification on the account and password. If all verifications pass, the jamming parameters are obtained by decrypting the parameter ciphertext using the first encryption rule.

[0143] The signal jammer collects real-time wireless communication signals, performs filtering preprocessing on the real-time wireless communication signals, and inputs the real-time wireless communication signals and jamming parameters into the deployed signal interference strategy generation model to obtain a real-time signal interference strategy including interference frequency selection matrix, power level vector and beamforming weights.

[0144] The first encryption rule is as follows: obtain the first 6 digits of the current timestamp and concatenate them with a preset master key to obtain the first concatenated data; calculate the MD5 value of the first concatenated data using the MD5 algorithm as the first temporary key; encode the masking parameter using UTF-8 to obtain encoded data, and append 4 random letters to the end of the encoded data to generate an extended data block; encrypt the extended data block using the first temporary key using the AES-128 algorithm to obtain the parameter ciphertext.

[0145] The second encryption rule is as follows:

[0146] The last 6 digits of the current timestamp are concatenated with a preset master key to obtain the second concatenated data. The MAC value of the second concatenated data is calculated using the HMAC algorithm as the second temporary key. The account, password, timestamp, and hash value are encrypted using the second temporary key via the SM4 algorithm to obtain first-level encrypted data. The first-level encrypted data is converted to hexadecimal data. Each character of the hexadecimal data is shifted 7 bits to the right to obtain second-level encrypted data. The numbers 6 and 9, A and C, and 8 and B in the second-level encrypted data are swapped to obtain third-level encrypted data. The third-level encrypted data is then encrypted using the second temporary key via the ChaCha20 algorithm to obtain the verification ciphertext.

[0147] By using timestamp-segmented concatenation of the master key to generate a temporary key (the first 6 bits are used for parameter ciphertext, and the last 6 bits are used for verification ciphertext), a dynamic key generation mechanism is implemented, ensuring that each encryption key is different, effectively preventing system-wide security risks caused by key leakage; the parameter ciphertext uses MD5+AES-128 double encryption, and the verification ciphertext uses a combination of three encryption algorithms: HMAC+SM4+ChaCha20, forming a heterogeneous encryption protection layer, significantly increasing the difficulty of brute-force attacks; obfuscation techniques such as character shifting (right shift by 7 bits) and alphanumeric substitution (swapping 6 / 9, A / C, 8 / B) are used to increase the complexity of ciphertext feature analysis.

[0148] Verify the consistency of original data such as account, password, and timestamp through hash value comparison (SHA256) to prevent tampering during transmission; ensure the timeliness of instructions through timestamp verification to avoid replay attacks of historical instructions; and establish operation access control through dual authentication of account and password to prevent unauthorized devices from accessing the system.

[0149] By appending 4 random letters to the end of the UTF-8 encoded data, the fixed structure of the data block is disrupted, increasing the difficulty of ciphertext pattern analysis. By setting the verification ciphertext to use three-level encryption processing (SM4 encryption → hexadecimal conversion → character shift → character replacement → ChaCha20 encryption), a multi-layered encryption nested structure is formed, which greatly increases the cost of reverse analysis.

[0150] By integrating a dynamic key generation mechanism (timestamp-segmented master key combined with MD5 / HMAC algorithms) with heterogeneous encryption algorithms (AES-128 / SM4 / ChaCha20 collaborative encryption), multiple anti-cracking barriers are constructed. At the same time, a multi-dimensional security verification system of hash verification, timestamp verification, and account password authentication is adopted to ensure the integrity, timeliness, and legality of command transmission. Its innovative anti-reverse design (random data expansion, character shifting and permutation, and three-level encryption nesting) significantly increases the difficulty of ciphertext parsing. Combined with an intelligent signal interference model to dynamically generate beamforming strategies, it achieves accurate signal shielding while taking into account system security and energy efficiency optimization.

[0151] Step S7 specifically involves:

[0152] The signal jammer is pre-set with a lookup table for storing the correspondence between noise intensity range and power amplifier gain, and a fan control function with the independent variable being the equipment temperature and the dependent variable being the fan speed.

[0153] The signal jammer collects environmental noise data in real time, matches the power amplifier gain from a lookup table based on the noise intensity of the environmental noise data, and controls the power amplifier to perform power compensation on the real-time interference signal based on the power amplifier gain.

[0154] The signal jammer collects the device temperature in real time. When the device temperature is less than a preset temperature threshold, the device temperature is input into the fan control function to obtain the corresponding fan speed. The fan's working state is dynamically adjusted based on the fan speed. When the device temperature is greater than or equal to the temperature threshold, the fan speed is adjusted to the maximum, and the bias voltage of the power amplifier is continuously reduced based on a preset voltage gradient until the device temperature is less than the temperature threshold.

[0155] By using a preset table of noise intensity and power amplifier gain, the ambient noise intensity can be matched in real time, and the gain of the power amplifier can be dynamically adjusted to achieve precise interference signal compensation. This avoids the problems of overcompensation (energy waste) or undercompensation (shielding failure) in the traditional fixed gain mode. Based on the hierarchical compensation strategy of noise intensity, the power amplifier power is increased only when necessary, reducing the overall energy consumption of the signal shield while maintaining a stable signal shielding effect.

[0156] By setting up a tiered heat dissipation mechanism, in the active heat dissipation stage (temperature < threshold): the heat dissipation intensity is dynamically adjusted based on the functional relationship between device temperature and fan speed, avoiding continuous full-speed fan operation and reducing mechanical wear and noise; in the emergency protection stage (temperature ≥ threshold): the fan is forced to run at full speed, while the power amplifier bias voltage is gradually reduced according to the voltage gradient, rapidly cooling down through a dual path of "reducing heat source power + enhancing heat dissipation" to prevent hardware overheating damage; and through dynamic temperature control strategies, the impact of drastic temperature fluctuations on electronic components is reduced, material aging caused by long-term high temperatures is suppressed, and device reliability is improved.

[0157] By directly matching noise intensity and gain values ​​through a pre-stored lookup table, the computational overhead of real-time calculation of power amplifier parameters is avoided, reducing processor load and improving response speed. By encapsulating the relationship between fan speed and temperature into a configurable function, the control logic is simplified, and parameters (such as function slope) can be flexibly adjusted to adapt to different heat dissipation design requirements.

[0158] By gradually reducing the power amplifier bias voltage through a preset voltage gradient, signal distortion or circuit impact caused by sudden voltage drops are avoided, ensuring a stable transition of the system to a safe state. A temperature threshold triggering mechanism provides redundant protection layers: fan speed adjustment is used for cooling first, and power amplifier voltage adjustment is only intervened in extreme cases to maintain shielding function to the maximum extent.

[0159] Adaptive signal compensation is achieved through dynamic matching of noise intensity and power amplifier gain, optimizing energy consumption while ensuring shielding effectiveness. Combined with a graded temperature control mechanism, functional fan speed adjustment and gradient power amplifier voltage control work together to dissipate heat, avoiding the risk of hardware overheating and extending equipment life. Its lookup table method and preset function design significantly reduce computational complexity and improve real-time response capability, while parameter configurability and modular architecture give the solution high scalability, ultimately achieving a multi-dimensional balance of performance, efficiency and reliability in complex electromagnetic environments.

[0160] Step S8 specifically involves:

[0161] The signal jammer records at least the following in real time: jamming parameters, real-time wireless communication signals, real-time signal interference strategies, real-time interference signals, environmental noise data, equipment temperature, jamming time, account information, and signal jamming effect feedback.

[0162] An incremental dataset is constructed based on the signal jamming logs. When the amount of data in the incremental dataset exceeds a preset threshold, the signal jamming strategy generation model is backed up. During idle periods, the signal jamming strategy generation model is trained using the incremental dataset. The performance of the trained signal jamming strategy generation model is then verified. If the verification fails, a version rollback operation is performed. If the verification passes, the update parameters of the signal jamming strategy generation model are extracted.

[0163] The signal jammer uploads the updated parameters to the server through the federation gateway. The server aggregates the updated parameters uploaded by each signal jammer to obtain global parameters, and then sends the global parameters to the signal jammer through the federation gateway. The signal jammer updates its local signal interference strategy generation model based on the received global parameters in order to perform federated optimization of the signal interference strategy generation model.

[0164] By recording multi-dimensional data in real time (shielding parameters, real-time wireless communication signals, real-time signal interference strategies, real-time interference signals, environmental noise data, equipment temperature, shielding time, account information, and shielding effect feedback), a complete profile of the equipment's operating status is formed, providing high-quality feature input for the signal interference strategy generation model and enhancing the accuracy of the interference strategy.

[0165] By adopting a data volume threshold triggering mechanism, resource waste caused by frequent training is avoided; system reliability is ensured through model backup and version rollback design to prevent service interruption due to training failure; and intelligent scheduling strategy for training during idle periods optimizes the utilization of computing resources.

[0166] By replacing the original data transmission with parameter aggregation, the data privacy protection requirements are met; by participating in global model optimization through distributed signal jammers, the generalization ability of the model to different environments (such as electromagnetic environments in different regions) is improved; and by achieving continuous model evolution through an edge-cloud collaborative architecture, the limitation of insufficient data samples from a single device is overcome.

[0167] By collecting multi-dimensional signal jamming logs in real time to construct an incremental dataset, and combining a threshold-based triggering mechanism with an idle time scheduling strategy, low-power and high-reliability incremental training of the local model is achieved. Furthermore, an innovative federated learning architecture is introduced. Under the premise of protecting data privacy, the parameter aggregation of distributed devices is used to achieve global model co-evolution, forming a closed-loop optimization system of "environmental perception - policy generation - effect feedback - federated tuning". It has dynamic environmental adaptability (such as noise adaptation and heat dissipation linkage), system stability assurance (version rollback and verification mechanism), and multi-device collaborative generalization advantages, which significantly improves the accuracy and scenario universality of signal interference strategies.

[0168] Accurate signal feature extraction and cross-domain correlation are achieved through a dynamic feature perception layer (integrating five-dimensional features of spectrum, time domain, spatial domain, protocol semantics, and signal strength) and a multi-domain feature fusion layer (based on hypergraph neural network and reinforcement learning dynamic weight allocation). Efficient interference strategies are generated by combining a multi-objective decision module (a dual-branch network generates interference frequency points and power strategies) and a beam optimization module (differentiable convex optimization improves directivity). Generative adversarial networks are used to enhance data diversity, and lightweight adaptation is achieved through model pruning, quantization, and containerized deployment. At the same time, multiple encryption (AES-128, SM4, ChaCha20) and dynamic verification mechanisms are used to ensure command security. The model is optimized in real time based on a federated learning framework, and it also has the capabilities of adaptive environmental noise compensation, intelligent temperature control, and log traceability. Finally, it has significant advantages in interference accuracy (frequency hit rate), resource utilization (power dynamic optimization), real-time response (millisecond-level inference), anti-countermeasure capability (narrow beam directionality), and compliance (frequency constraints), achieving intelligent, low-power, and highly concealed directional signal shielding.

[0169] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A signal masking method based on deep reinforcement learning, characterized in that: Includes the following steps: Step S1: Construct a signal interference strategy generation model based on the dynamic feature perception layer, the multi-domain feature fusion layer, and the interference strategy generation layer, and set the loss function of the signal interference strategy generation model; The dynamic feature perception layer is constructed based on a spectrum dynamic coding module, a time-domain feature coding module, a spatial-domain feature coding module, a protocol semantic coding module, and a signal strength coding module. The spectrum dynamic coding module, the time-domain feature coding module, the spatial-domain feature coding module, the protocol semantic coding module, and the signal strength coding module are respectively used to extract spectrum features, time features, spatial features, communication protocol features, and signal strength features from the input wireless communication signal. The multi-domain feature fusion layer is constructed based on a feature interaction enhancement module and a dynamic weight allocation module. The feature interaction enhancement module is used to construct a three-dimensional correlation matrix based on spectral features, temporal features, and spatial features. The dynamic weight allocation module is used to evaluate the contribution of the three-dimensional correlation matrix, communication protocol features, and signal strength features to perform fusion and obtain fused features. The interference strategy generation layer is constructed based on a multi-objective decision-making module, a beam optimization module, and an output module. The multi-objective decision-making module is used to generate an interference frequency selection matrix and a power level vector based on the fusion characteristics. The beam optimization module is used to generate beamforming weights based on the fusion characteristics. The output module is used to generate a signal interference strategy based on the interference frequency selection matrix, the power level vector, and the beamforming weights. Step S2: Acquire a large number of historical wireless communication signals, preprocess and label each of the historical wireless communication signals to construct a signal dataset, and expand the sample size of the signal dataset by using a generative adversarial network. Step S3: Divide the signal dataset into a training set, a validation set, and a test set, and train, validate, and test the signal interference strategy generation model using the training set, validation set, and test set. Step S4: Compress the signal interference strategy generation model that has passed the test, and deploy the compressed signal interference strategy generation model to the signal jammer. Step S5: After the signal jammer is started, the input signal jamming command is obtained, the signal jamming command is verified and parsed to obtain the jamming parameters, the real-time wireless communication signal is collected, the real-time wireless communication signal is preprocessed, and the real-time wireless communication signal and the jamming parameters are input into the deployed signal interference strategy generation model to obtain the real-time signal interference strategy. Step S6: The signal jammer modulates the corresponding real-time interference signal based on the real-time signal interference strategy. After the power amplifier performs power amplification on the real-time interference signal through digital predistortion technology, it transmits the real-time interference signal to the outside through the radio frequency antenna to perform signal directional shielding operation. Step S7: The signal jammer collects environmental noise data and equipment temperature in real time, performs power compensation on the real-time interference signal based on the environmental noise data, and dynamically adjusts the working status of the power amplifier and fan based on the equipment temperature. Step S8: The signal jammer records the signal jamming log in real time, and performs federated optimization of the signal interference strategy generation model based on the signal jamming log.

2. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: In step S1, the spectrum dynamic coding module is constructed based on a dynamic coding unit, a frequency domain correlation unit, and an aggregation parsing unit; The dynamic coding unit is used to extract local frequency domain features and temporal dynamic features from the input wireless communication signal through multi-scale convolution subunits and gated recurrent subunits, and output time-frequency joint features including the local frequency domain features and temporal dynamic features. The frequency domain association unit is used to divide adjacent frequency points into the same type of cluster through the frequency point clustering subunit, so as to capture the temporal dependency relationship of frequency point migration from the time-frequency joint features. The multi-head self-attention subunit calculates cross-frequency point attention weights on the time-frequency joint features to model the global dependency relationship between frequency points, thereby capturing the frequency domain correlation of frequency point migration. Based on the temporal dependency relationship and the frequency domain correlation, a frequency point dynamic coding vector is generated. The aggregation parsing unit is used to perform multi-resolution matching of carrier aggregation mode on the frequency point dynamic coding vector through the pattern matching subunit, and to decouple the frequency point migration feature from the frequency point dynamic coding vector through the feature reconstruction subunit, and output the spectral features including the frequency point migration feature and the carrier aggregation mode. The temporal feature encoding module is used to extract temporal features, including signal burst characteristics and time-frequency interleaving patterns, from the input wireless communication signal through a temporal convolutional network and a multi-head temporal attention mechanism. The spatial feature encoding module is used to extract spatial features, including signal topology features and beam spatial mode vectors, from the input wireless communication signal through graph attention network and spatial spectrum estimation. The protocol semantic encoding module is used to extract communication protocol features from the input wireless communication signal using BERT units and the 3GPP standard knowledge graph. The signal strength encoding module is used to dynamically map the RSSI value and MCS carried in the input wireless communication signal through an adaptive normalization layer, thereby extracting signal strength features. The feature interaction enhancement module is used to model cross-domain feature associations of spectral features, temporal features, and spatial features through a hypergraph neural network, thereby constructing a three-dimensional association matrix; The dynamic weight allocation module is used to evaluate the contribution of the three-dimensional correlation matrix, communication protocol features, and signal strength features to the current interference scenario through a value network based on the policy gradient method of reinforcement learning, so as to fuse them to obtain fused features. The multi-objective decision module is used to reason about the fused features through a dual-branch network structure to generate an interference frequency selection matrix and a power level vector. The beam optimization module uses a differentiable convex optimization layer and an SDP relaxation method to infer the fused features in order to generate beamforming weights.

3. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: The formula for the loss function is: L=α*L int +β*L res +γ*L beam +δ*L fusion ; Where L represents the loss value of the loss function; L int Indicates the loss of interference effect; L res Indicates resource consumption and loss; L beam L represents beam optimization loss; fusion α represents the feature fusion loss; α, β, γ, and δ all represent multi-objective trade-off coefficients, optimized through grid search. ; ; ; ; Where f represents the frequency point; This indicates the probability that frequency point f is selected; Indicates characteristic functions; Represents the set of target interference frequencies; , as well as Both represent regularization intensity coefficients; This represents the vector prediction value of the power level; Indicates the theoretically optimal power level; Represents the squared Euclidean distance between vectors; Represents the beamforming weight matrix; Represents the L1 norm; Represents the square of the Frobenius norm; Indicates the main lobe prediction direction of the beam; Indicates the ideal direction of the main lobe of the beam; Represents a three-dimensional incidence matrix; Indicates the characteristics of the communication protocol; Indicates signal strength characteristics; Indicates fusion characteristics; This indicates dynamic weighted multiplicative fusion; This indicates dynamic weighted additive fusion.

4. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: Step S2 specifically involves: A large number of historical wireless communication signals are acquired, and each of the historical wireless communication signals is preprocessed, including at least signal amplification, filtering, down-conversion, and sampling. The preprocessed historical wireless communication signals are then labeled with at least the following characteristics: frequency migration features, carrier aggregation mode, signal burst characteristics, time-frequency interleaving mode, signal topology features, beam space mode vector, communication protocol features, signal strength features, signal type, and signal interference strategy. A dataset is constructed based on the labeled historical wireless communication signals, and the sample size of the signal dataset is expanded by a generator in a pre-trained generative adversarial network.

5. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: Step S3 specifically involves: Based on k-fold cross-validation, the signal dataset is divided into a training set, a validation set, and a test set in an 8:1:1 ratio. The signal interference strategy generation model is trained using the training set. During training, hyperparameters including the learning rate, regularization parameter, network structure parameter, batch size, training epochs, activation function, optimizer, and random dropout rate are optimized using a Bayesian optimization framework until the loss value of the loss function is less than a preset loss threshold or a preset early stopping condition is met. The trained signal interference strategy generation model is validated using the validation set to determine whether the accuracy of the interference strategy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded for continued training. If yes, the validation passes, and the following parameters are calculated using the test set: frequency hit rate, false interference rate, power utilization rate, power dynamic adjustment response time, beam pointing accuracy, strategy conflict rate, and single inference time. These parameters are used to test the validated signal interference strategy generation model. If the test fails, the training set is expanded for continued training. If the test passes, training ends.

6. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: Step S4 specifically involves: For all weight parameters of the signal interference strategy generation model that has passed the test, sort them according to the absolute value of the weight values. Set the weight values ​​below a preset weight threshold to zero to perform a pruning operation on the signal interference strategy generation model. Perform a performance verification on the pruned signal interference strategy generation model. If the performance verification fails, adjust the weight threshold and repeat the pruning operation; if the performance verification passes, then: The signal interference strategy generation model is quantized based on a preset quantization step size and offset. The quantized signal interference strategy generation model is then subjected to a second performance verification. If the performance verification fails, the quantization step size and offset are adjusted and the quantization operation is re-executed. If the performance verification is successful, the compression of the signal interference strategy generation model is completed. The compressed signal interference strategy generation model is deployed to the signal jammer using containerization technology. Data drift compensation training and model drift compensation training are then performed on the signal interference strategy generation model deployed on the signal jammer.

7. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: Step S5 specifically involves: After the signal jammer is activated, it acquires the input signal jamming command carrying ciphertext of parameters and ciphertext of verification; the ciphertext of parameters is obtained by encrypting the jamming parameters using a first encryption rule; The verification ciphertext is obtained by encrypting the account, password, timestamp, and hash value using the second encryption rule; the hash value is obtained by calculating hash256 from the account, password, and timestamp; the blocking parameters include at least the blocking duration, blocking range, and blocking signal type; The signal jammer parses the signal jamming command to obtain parameter ciphertext and verification ciphertext. It decrypts the verification ciphertext using the second encryption rule to obtain the account, password, timestamp, and hash value. It then performs integrity verification on the account, password, and timestamp using the hash value, performs timeliness verification using the timestamp, and performs legality verification on the account and password. If all verifications pass, the jamming parameters are obtained by decrypting the parameter ciphertext using the first encryption rule. The signal jammer acquires real-time wireless communication signals, performs filtering preprocessing on the real-time wireless communication signals, and then inputs the real-time wireless communication signals and jamming parameters into the deployed signal interference strategy generation model to obtain a real-time signal interference strategy including an interference frequency selection matrix, a power level vector, and beamforming weights.

8. The signal masking method based on deep reinforcement learning as described in claim 7, characterized in that: The first encryption rule is as follows: obtain the first 6 digits of the current timestamp and concatenate them with a preset master key to obtain the first concatenated data; calculate the MD5 value of the first concatenated data using the MD5 algorithm as the first temporary key; encode the masking parameter using UTF-8 to obtain encoded data, and append 4 random letters to the end of the encoded data to generate an extended data block; encrypt the extended data block using the first temporary key using the AES-128 algorithm to obtain the parameter ciphertext. The second encryption rule is as follows: The last 6 digits of the current timestamp are concatenated with a preset master key to obtain the second concatenated data. The MAC value of the second concatenated data is calculated using the HMAC algorithm as the second temporary key. The account, password, timestamp, and hash value are encrypted using the second temporary key via the SM4 algorithm to obtain first-level encrypted data. The first-level encrypted data is converted to hexadecimal data. Each character of the hexadecimal data is shifted 7 bits to the right to obtain second-level encrypted data. The numbers 6 and 9, A and C, and 8 and B in the second-level encrypted data are swapped to obtain third-level encrypted data. The third-level encrypted data is then encrypted using the second temporary key via the ChaCha20 algorithm to obtain the verification ciphertext.

9. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: Step S7 specifically involves: The signal jammer is pre-set with a lookup table for storing the correspondence between noise intensity range and power amplifier gain, and a fan control function with the independent variable being the equipment temperature and the dependent variable being the fan speed. The signal jammer collects environmental noise data in real time, matches the power amplifier gain from a lookup table based on the noise intensity of the environmental noise data, and controls the power amplifier to perform power compensation on the real-time interference signal based on the power amplifier gain. The signal jammer collects the device temperature in real time. When the device temperature is lower than the preset temperature threshold, the device temperature is input into the fan control function to obtain the corresponding fan speed. The working state of the fan is dynamically adjusted based on the fan speed. When the device temperature is greater than or equal to the temperature threshold, the fan speed is adjusted to the maximum, and the bias voltage of the power amplifier is continuously reduced based on a preset voltage gradient until the device temperature is less than the temperature threshold.

10. The signal masking method based on deep reinforcement learning as described in claim 1, characterized in that: Step S8 specifically involves: The signal jammer records at least the following in real time: jamming parameters, real-time wireless communication signals, real-time signal interference strategies, real-time interference signals, environmental noise data, equipment temperature, jamming time, account information, and signal jamming effect feedback. An incremental dataset is constructed based on the signal jamming logs. When the amount of data in the incremental dataset exceeds a preset threshold, the signal jamming strategy generation model is backed up. During idle periods, the signal jamming strategy generation model is trained using the incremental dataset. The performance of the trained signal jamming strategy generation model is then verified. If the verification fails, a version rollback operation is performed. If the verification passes, the update parameters of the signal jamming strategy generation model are extracted. The signal jammer uploads the updated parameters to the server through the federation gateway. The server aggregates the updated parameters uploaded by each signal jammer to obtain global parameters, and then sends the global parameters to the signal jammer through the federation gateway. The signal jammer updates its local signal interference strategy generation model based on the received global parameters in order to perform federated optimization of the signal interference strategy generation model.

Citation Information

Patent Citations

  • Intelligent space frequency spectrum anti-interference method based on deep learning network

    CN117240384A

  • Method and system for blocking, shielding and adaptively identifying mobile phone signals

    CN119341679A