Signal shielding method based on deep reinforcement learning

Through the signal shielding method of deep reinforcement learning, frequency migration is dynamically identified and precise interference strategies are generated, which solves the accuracy, real-time and power consumption problems of signal shielding technology in multi-band and high-density electromagnetic environments, and achieves low-power and efficient signal shielding effects.

CN120811542AActive Publication Date: 2025-10-17CHINA YOUKE COMM TECH

Patent Information

Application Number
CN202511254408.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-17
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing signal shielding technology faces the risks of static frequency band configuration, high power consumption and protocol avoidance when facing multi-band, high-density and dynamic electromagnetic environments. It is difficult to achieve accuracy, real-time and scalability. Traditional improvement solutions such as spectrum sensing and threshold-triggered power control are not effective in complex scenarios.

Method used

A signal shielding method based on deep reinforcement learning is adopted. A signal interference strategy generation model is constructed through the dynamic feature perception layer, multi-domain feature fusion layer and interference strategy generation layer. The data set is expanded by combining the generative adversarial network to realize multi-dimensional feature extraction and dynamic interference strategy generation in the spectrum, time domain, spatial domain and protocol semantics. Multi-objective decision-making and beam optimization modules are used for precise power allocation and beamforming. Combined with federal tuning and environmental noise compensation, the interference strategy is dynamically adjusted.

Benefits of technology

It improves the accuracy, real-time performance and scalability of signal shielding, reduces power consumption, can adapt to rapid frequency migration and complex electromagnetic environments, and meet the needs of future communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811542A_ABST
    Figure CN120811542A_ABST
Patent Text Reader

Abstract

The invention provides a signal shielding method based on deep reinforcement learning, and belongs to the technical field of artificial intelligence and communication crossing. The method comprises the steps that S1, a large number of historical wireless communication signals are acquired to construct a data set to train a signal interference strategy generation model; s2, deploying the signal interference strategy generation model to a signal shielding device; s3, verifying and analyzing the signal shielding instruction to obtain a shielding parameter, collecting a real-time wireless communication signal, and inputting the real-time wireless communication signal and the shielding parameter into a deployed signal interference strategy generation model to obtain a real-time signal interference strategy; and S4, the signal shielding device modulates the real-time interference signal based on the real-time signal interference strategy, and the power amplifier performs power amplification on the real-time interference signal and then transmits the real-time interference signal to the outside through the radio frequency antenna so as to perform signal shielding. The method has the advantages that the accuracy, the real-time performance and the expansibility of signal shielding are greatly improved, and the power consumption of signal shielding is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and communication, and particularly relates to a signal shielding method based on deep reinforcement learning. BACKGROUND

[0002] With the rapid development of wireless communication technology, the global mobile communication system has formed a multi-standard coexistence pattern including 2G (GSM / CDMA), 3G (WCDMA / TD-SCDMA), 4G (LTE) and 5G (NR). According to the data of the International Telecommunication Union (ITU), the current commercial wireless communication frequency band has covered the range of 400MHz to 6GHz, including the core frequency bands of GSM 900MHz / 1800MHz, CDMA800MHz, 3G 2100MHz, 4G 2600MHz and 5G 3.5GHz / 4.9GHz, and at the same time, the popularity of short-range communication technologies such as Wi-Fi (2.4GHz / 5GHz), Bluetooth (2.4GHz) and Internet of Things (NB-IoT / eMTC) makes the electromagnetic environment present the characteristics of multi-band, high density and dynamic, which poses a severe challenge to signal shielding technology.

[0003] At present, the signal shielding technology mainly includes signal-to-noise ratio degradation mode and downlink signal interference mode. The principle of the signal-to-noise ratio degradation mode is to make the Eb / N0 (bit energy to noise density ratio) of the receiving end lower than the demodulation threshold (such as QPSK requires Eb / N0≥6dB) through wideband noise injection, so as to achieve the purpose of shielding signals; however, due to the disadvantages of low interference efficiency (wideband noise injection may cause a large amount of invalid power consumption, and the interference efficiency is relatively low), poor real-time performance (in a high dynamic environment, it is difficult to quickly respond to signal changes), and weak anti-interference ability (for signals using advanced modulation and demodulation technology (such as spread spectrum technology), wideband noise injection may not effectively cover its time-frequency resource unit), therefore, the downlink signal interference mode is mainly used; the principle of the downlink signal interference mode is to generate a same-frequency interference signal with the same carrier frequency as the target signal and with a power increase of 6-10dB, to cover the target frequency band in a frequency sweeping manner (the typical frequency sweeping rate is 50-200MHz / s), and the mechanism uses the automatic gain control (AGC) characteristics of the receiver to inject interference pulses in the time slot synchronization period of the base station downlink signal, so that the mobile terminal cannot correctly demodulate the DCI format of PBCH (physical broadcast channel) and PDCCH (physical downlink control channel), thereby causing the RRC connection establishment to fail.

[0004] However, the signal shielder applying the downlink signal interference mode faces the following technical bottlenecks in actual deployment: (1) Static frequency band configuration: The signal shielder usually adopts a preset frequency point library mode, which cannot adapt to the frequency point migration caused by dynamic spectrum allocation (DSA) and carrier aggregation (CA) technologies. (2) High signal shielding power consumption (low interference efficiency): The full-band continuous frequency sweeping mode causes a high proportion of invalid power consumption. The measured data shows that the power consumption of a typical 5-frequency-band shielding device can reach 120W in a continuous working state, and the thermal design power (TDP) exceeds the standard, causing the device junction temperature to rise above 85℃, which seriously affects the MTBF (mean time between failures). (3) Protocol avoidance risk: New communication systems use cognitive radio (CR) technology to realize spectrum sensing and frequency hopping communication. For example, the CSS (Chirp Spread Spectrum) modulation of LoRa can achieve a receiving sensitivity of -148dBm in a 125kHz channel, and traditional frequency sweeping interference cannot effectively cover the time-frequency resource units corresponding to the spreading factor (SF7-SF12).

[0005] In view of the above problems, the industry has tried various improvement schemes: (a) Frequency band identification technology based on spectrum sensing, such as using USRP (universal software radio peripheral) to build a frequency sweeping receiver and performing frequency spectrum energy detection through FFT (fast Fourier transform). However, due to the complexity of FFT calculation and the working mode of the frequency sweeping receiver, this method has the defect of insufficient real-time performance, with a single full-band scanning time of more than 2s, which cannot match the slot structure (0.5ms slot period) of 5G NR. (b) Threshold triggered power control, which dynamically adjusts the interference power according to RSSI (received signal strength indication), but does not consider the beamforming characteristics in the multi-user MIMO (multiple-input multiple-output) scenario, resulting in mismatch between the interference beam and the spatial dimension of the target user. (c) Protocol feature matching interference, such as selectively interfering with the CFI (control format indicator) field of PDCCH, but it needs to deeply analyze the DCI format and faces the compatibility risk brought by frequent evolution of 3GPP standards.

[0006] Therefore, how to provide a signal shielding method based on deep reinforcement learning to improve the accuracy, real-time performance and scalability of signal shielding and reduce the power consumption of signal shielding has become a technical problem to be solved. SUMMARY

[0007] The technical problem to be solved by the present application is to provide a signal shielding method based on deep reinforcement learning to improve the accuracy, real-time performance and scalability of signal shielding and reduce the power consumption of signal shielding.

[0008] The present application provides a signal shielding method based on deep reinforcement learning, comprising the following steps: Step S1, constructing a signal interference strategy generation model based on a dynamic feature perception layer, a multi-domain feature fusion layer, and an interference strategy generation layer, and setting a loss function of the signal interference strategy generation model; The dynamic feature perception layer is constructed based on a spectrum dynamic encoding module, a time domain feature encoding module, a space domain feature encoding module, a protocol semantic encoding module, and a signal strength encoding module; the spectrum dynamic encoding module, the time domain feature encoding module, the space domain feature encoding module, the protocol semantic encoding module, and the signal strength encoding module are respectively used to extract spectrum features, time features, space features, communication protocol features, and signal strength features from input wireless communication signals; The multi-domain feature fusion layer is constructed based on a feature interaction reinforcement module and a dynamic weight distribution module; the feature interaction reinforcement module is used to construct a three-dimensional correlation matrix according to the spectrum features, the time features, and the space features; and the dynamic weight distribution module is used to evaluate the contribution degrees of the three-dimensional correlation matrix, the communication protocol features, and the signal strength features to obtain fused features through fusion; The interference strategy generation layer is constructed based on a multi-objective decision module, a beam optimization module, and an output module; the multi-objective decision module is used to generate an interference frequency point selection matrix and a power level vector according to the fused features; the beam optimization module is used to generate a beamforming weight value according to the fused features; and the output module is used to generate a signal interference strategy according to the interference frequency point selection matrix, the power level vector, and the beamforming weight value; Step S2, obtaining a large number of historical wireless communication signals, preprocessing and labeling each of the historical wireless communication signals to construct a signal dataset, and expanding the sample size of the signal dataset through a generative adversarial network; Step S3, dividing the signal dataset into a training set, a validation set, and a test set, and training, validating, and testing the signal interference strategy generation model through the training set, the validation set, and the test set; Step S4, compressing the signal interference strategy generation model that passes the test, and deploying the compressed signal interference strategy generation model to a signal shield; Step S5, after the signal shield is started, obtaining an input signal shielding instruction, verifying and analyzing the signal shielding instruction to obtain shielding parameters, collecting real-time wireless communication signals, preprocessing the real-time wireless communication signals, inputting the real-time wireless communication signals and the shielding parameters into the deployed signal interference strategy generation model, and obtaining a real-time signal interference strategy; Step S6, the signal shield modulates a corresponding real-time interference signal based on the real-time signal interference strategy, a power amplifier performs a power amplification operation on the real-time interference signal through a digital pre-distortion technology, and the real-time interference signal is transmitted externally through a radio frequency antenna to perform a signal directional shielding operation. Step S7, the signal shielder collects environmental noise data and equipment temperature in real time, compensates the power of the real-time interference signal based on the environmental noise data, and dynamically adjusts the working state of the power amplifier and the fan based on the equipment temperature; Step S8, the signal shielder records a signal shielding log in real time, and performs federal optimization on the signal interference strategy generation model based on the signal shielding log.

[0009] The advantages of the present application are: 1. Constructing a signal interference strategy generation model through a dynamic feature perception layer, a multi-domain feature fusion layer, and an interference strategy generation layer; obtaining a large number of historical wireless communication signals to construct a signal dataset, and expanding the sample size of the signal dataset through a generative adversarial network; dividing the signal dataset into a training set, a validation set, and a test set to train, validate, and test the signal interference strategy generation model; compressing the signal interference strategy generation model that passes the test, and deploying the compressed signal interference strategy generation model to a signal shield; after the signal shield is started, verifying and analyzing the input signal shielding instruction to obtain shielding parameters, collecting real-time wireless communication signals, pre-processing the real-time wireless communication signals, inputting the real-time wireless communication signals and the shielding parameters into the deployed signal interference strategy generation model, and obtaining a real-time signal interference strategy; then the signal shield modulates the corresponding real-time interference signal based on the real-time signal interference strategy, a power amplifier performs a power amplification operation on the real-time interference signal through a digital pre-distortion technology, and the signal shield transmits the real-time interference signal to the outside through a radio frequency antenna to perform a signal directional shielding operation; the signal shield collects environmental noise data and device temperature in real time, compensates the real-time interference signal based on the environmental noise data, dynamically adjusts the working state of the power amplifier and the fan based on the device temperature, and records a signal shielding log in real time, and optimizes the signal interference strategy generation model based on the signal shielding log; that is, the signal interference strategy is generated through the pre-trained signal interference strategy generation model, the corresponding interference signal is executed based on the signal interference strategy to perform signal shielding, and the dynamic feature perception layer of the signal interference strategy generation model can dynamically identify and adapt to the fast migration of frequency points through the extraction of frequency point migration features and carrier aggregation modes, and the interference frequency point selection matrix generated by the multi-objective decision module of the interference strategy generation layer, instead of relying on a preset frequency point library; the three-dimensional correlation matrix in the multi-domain feature fusion layer models the cross-domain dependence of spectrum, time, and space, adjusts the interference frequency points in real time, and adapts to the dynamic spectrum allocation scene; the interference strategy generation layer realizes precise power allocation through a power level vector and a beamforming weight value, avoids full-band sweeping, compensates the real-time interference signal based on environmental noise data, dynamically adjusts the working state of the power amplifier based on the device temperature, further optimizes power consumption, and reduces invalid radiation; the resource consumption loss in the loss function directly constrains the resource occupation of the power and beam weight value, ensuring that the model is trained with low power consumption as the goal; the protocol semantic coding module of the dynamic feature perception layer extracts communication protocol features, and the signal interference strategy generated by the interference strategy generation layer targets specific protocol features, and continuously adapts to protocol evolution through federated optimization; the model volume is effectively reduced and the calculation amount is reduced through compression before deployment of the signal interference strategy generation model, which greatly improves the accuracy, real-time performance, and scalability of signal shielding, and greatly reduces the power consumption of signal shielding.

[0010] 2、Through the dynamic feature perception layer, five types of coding modules such as spectrum, time domain, space domain, protocol semantics and signal strength are integrated to realize holographic feature extraction of wireless signals, avoiding the limitations of single feature analysis; for example, protocol semantic coding can identify vulnerabilities of specific communication protocols (such as Wi-Fi 6, 5G NR), and improve the pertinence of interference.

[0011] 3、Through the three-dimensional correlation matrix (spectrum x time x space) modeling of the multi-domain correlation of the signal, combined with the dynamic weight distribution module to evaluate the feature contribution, the feature fusion of the adaptive scene is realized, which is significantly better than the traditional static weight fusion, especially in complex electromagnetic environments (such as urban dense multipath scenes).

[0012] 4、Through the interference strategy generation layer, the multi-objective decision module and the beam optimization module are used to simultaneously optimize the interference frequency point selection matrix, power level vector and beamforming weight, taking into account the interference efficiency and energy consumption control; for example, in low signal-to-noise ratio scenarios, high frequency band interference is preferred, while invalid power allocation is reduced.

[0013] 5、Through the generation of adversarial network (GAN), the problem of insufficient rare signal samples in actual scenarios (such as emergency communication signals) is effectively solved, and the model generalization ability is improved; through the federal optimization mechanism, the signal shielding logs of multiple signal shielders are aggregated to update the signal interference strategy generation model, realizing the coexistence of knowledge sharing and privacy protection; for example, signal shielders in different regions can cooperate to optimize the protocol semantic recognition module without uploading original signal data; through the interference signal generation stage, combined with deep reinforcement learning strategy and digital pre-distortion (DPD) technology, the non-linear distortion of power amplifier is suppressed, ensuring the focusing of interference signals in the target frequency band, and reducing the risk of out-of-band radiation.

[0014] 6、Through the real-time compensation of environmental noise and temperature control linkage, the dynamic adjustment of interference power and equipment thermal management are realized; for example, in high temperature environment, the power amplifier output is automatically reduced and the heat dissipation is accelerated to avoid the shutdown of the signal shield.

[0015] 7、Through the formation of a millisecond-level closed loop from real-time wireless communication signal collection to real-time interference signal emission, through the pre-processing of real-time wireless communication signals and the deployment of lightweight models, the real-time tracking demand for frequency hopping signals is met.

[0016] 8. Through deep reinforcement learning and dynamic coordination of multi-domain feature fusion (spectrum, time domain, space domain, protocol semantics and signal strength), high-precision signal interference strategy generation is realized, and the core advantages are: multi-dimensional perception (holographic feature extraction and three-dimensional correlation matrix modeling), dynamic optimization (feature weight distribution based on attention mechanism and multi-objective decision), real-time closed-loop control (millisecond-level interference strategy generation and digital pre-distortion power compensation), and edge intelligent deployment (model compression combined with federated optimization technology), while having high energy efficiency (dynamic power adjustment and temperature adaptive management), strong robustness (anti-adversarial attack design and hardware fault tolerance mechanism) and environmental friendliness (directional beam reduces electromagnetic pollution), which can support future 6G and quantum secure communication scenarios, and realize precise, compliant and low-power intelligent signal shielding in military, civilian and emergency scenarios.

[0017] 9. Signal multi-dimensional representation is realized through five-way encoding modules of spectrum, time domain, space domain, protocol semantics and signal strength, covering signal frequency domain dynamics (multi-scale convolution + gated recurrent), time domain burstiness (time attention), space domain topology (graph attention), protocol specification and intensity dynamics; through the feature interaction reinforcement module, a hypergraph neural network is used to model cross-domain correlation, and a three-dimensional correlation matrix is constructed, which breaks through the limitation of traditional two-dimensional feature interaction and effectively captures the time-frequency-space joint domain feature correlation.

[0018] 10. Through the spectrum dynamic coding module, local features (multi-scale convolution) and global dependencies (multi-head self-attention) are innovatively fused, frequency point clustering is used to capture migration patterns, and a multi-resolution matching mechanism is used to realize fine analysis of frequency point dynamic features; signal strength coding uses an adaptive normalization layer to realize dynamic mapping of RSSI and MCS, overcoming the lack of adaptability of traditional fixed normalization to rapid changes in signal strength.

[0019] 11. The dynamic weight distribution module realizes dynamic evaluation of feature contribution degree through the policy gradient method, and uses a value network for online weight adjustment, which is more suitable for the time-varying characteristics of complex interference scenarios than the fixed weight fusion strategy; the multi-objective decision module realizes the joint optimization of interference frequency point selection (discrete decision) and power level adjustment (continuous decision) through a double-branch network structure, avoiding the error accumulation problem of traditional cascading decision.

[0020] 12. By combining the differentiable convex optimization layer with the SDP relaxation method, the gradient is derivable while ensuring the physical constraints of beamforming, breaking the bottleneck of traditional convex optimization methods that cannot be trained end-to-end; by introducing the main lobe direction cosine similarity measure into the beam optimization loss function, it is more consistent with the beam space distribution characteristics than the Euclidean distance.

[0021] 13. By integrating four major objectives in the loss function design, including interference effect (cross entropy + power deviation), resource consumption (L1+F norm), beam direction (cosine similarity), and feature fusion error, the multi-objective trade-off coefficients are optimized through grid search to achieve system-level performance balance; by introducing a dynamic regularization mechanism (λ1, λ2, λ3 coefficients), the model is effectively prevented from overfitting. The Lres loss term simultaneously constrains power consumption and beam energy dispersion, reflecting the concept of green communication.

[0022] 14. The protocol semantic encoding module innovatively integrates the BERT pre-trained model and the 3GPP knowledge graph, leveraging the semantic understanding capabilities of large-scale language models while incorporating prior knowledge of communication standards, significantly improving the accuracy of protocol feature analysis. By introducing carrier aggregation mode matching in frequency migration modeling, the physical layer features are organically combined with the high-level protocol status, enhancing adaptability to modern communication systems (such as 5G NR).

[0023] 15. The aggregation analysis unit decouples the frequency migration features and carrier aggregation mode through the feature reconstruction sub-unit, realizes the separation of physical layer dynamic features and protocol layer configuration features, and enhances the model's interpretability for complex signal interference scenarios; the spatial feature encoding module independently extracts the signal topology (graph attention network) and beam spatial modality (spatial spectrum estimation), avoiding the parameter redundancy problem caused by spatial feature coupling in traditional methods.

[0024] 16. Through multi-dimensional feature fusion (spectral dynamic coding, time domain feature coding, spatial domain feature coding, protocol semantic coding and signal strength coding) and reinforcement learning dynamic decision-making mechanism, refined analysis of signal features and intelligent interference decision-making are achieved. Its core advantages are reflected in: innovative combination of multi-scale convolution, gated recurrent unit and multi-head attention mechanism to construct a cross-domain dynamic feature coding network, breaking through the limitations of traditional single-domain feature modeling; using differentiable convex optimization layer and hypergraph neural network to realize joint optimization of beamforming and feature association, while ensuring physical constraints and improving computational efficiency; through dynamic weight allocation and multi-objective loss function design, it effectively balances interference performance, resource efficiency and system robustness, and has 3GPP multi-protocol compatibility, hardware deployment friendliness and cross-scenario migration capabilities, providing a high-precision, low-latency, green and energy-saving end-to-end solution for intelligent spectrum confrontation in complex electromagnetic environments.

[0025] 17. By annotating the signal data set, the system covers physical layer features such as frequency migration, carrier aggregation, signal burst characteristics, time-frequency interleaving, and incorporates protocol layer information such as communication protocol and signal type to achieve full-dimensional signal modeling. By annotating features such as beam space modal vectors and signal interference strategies, the system can effectively analyze the signal characteristics of complex communication scenarios such as multi-antenna systems (MIMO) and dynamic spectrum sharing.

[0026] 18、Through signal amplification, filtering and down-conversion operations, environmental noise and hardware distortion are eliminated, and the signal SNR is improved to provide high-fidelity data for subsequent labeling and modeling.

[0027] 19、Through multi-dimensional signal feature labeling (covering frequency point migration, carrier aggregation, communication protocol parameters from physical layer to protocol layer) and standardized preprocessing (signal amplification, filtering, down-conversion), the data representation accuracy and SNR are significantly improved; combined with pre-trained generative adversarial network (GAN) to generate adversarial samples, breaking the sample scarcity bottleneck of real communication signal collection, enhancing the robustness of AI model to multiple scenarios (such as channel fading, multi-device interference).

[0028] 20、Through k-fold cross-validation method combined with 8:1:1 data set division ratio, both data utilization efficiency and generalization ability verification are considered; k-fold cross-validation effectively reduces the randomness bias of single data division, improves model stability; the golden ratio of 8:1:1 ensures sufficient training data (80%) while the independence of validation set (10%) and test set (10%) is verified; the double verification mechanism (validation set + test set) forms a closed loop feedback, which significantly reduces the risk of overfitting.

[0029] 21、Through the Bayesian optimization framework to establish a probability model to guide the parameter search direction, compared with grid search / random search to effectively improve the search efficiency; at the same time, the learning rate, regularization parameter, network structure parameter and other heterogeneous parameters are optimized, breaking through the limitations of traditional single parameter optimization; introduce random dropout rate as an optimization variable, creatively solve the parameter self-adaptation problem of Dropout mechanism in deep learning.

[0030] 22、Through pre-set loss threshold to ensure that the basic performance of the model meets the standard; through setting early stopping condition to prevent invalid training resource consumption, especially suitable for complex models with high-dimensional parameter space; if the verification fails, automatically expand the training set to realize the cooperative optimization of data enhancement and model training.

[0031] 23、Through the innovative integration of k-fold cross-validation and 8:1:1 data layering mechanism, combined with the Bayesian optimization framework for intelligent collaborative optimization of multi-dimensional hyperparameters, a closed-loop training system is established, which includes dynamic early stopping threshold and training set adaptive expansion. At the same time, in the test stage, multi-dimensional evaluation indicators such as frequency point hit rate and false interference rate are introduced, effectively solving the technical pain points of low training efficiency and insufficient generalization ability of traditional interference strategy models, significantly improving the model convergence speed, strategy generation accuracy and anti-interference performance in practical applications, especially in the electromagnetic countermeasure scene, achieving the balance optimization of interference strategy precision and resource utilization efficiency.

[0032] 24、By adopting the pruning strategy of absolute value sorting + threshold zeroing, the model parameter quantity is effectively reduced while the core features are retained, the threshold is dynamically adjusted through the loop verification mechanism, and the problem of key weight deletion caused by traditional pruning is solved; by introducing the offset compensation quantization algorithm (better than the conventional linear quantization), the step size parameter is optimized through secondary performance verification closed loop, which effectively improves the compression ratio on the premise of reducing the model accuracy loss.

[0033] 25、Through two-stage model compression (adaptive pruning based on dynamic threshold adjustment and quantization parameter optimization), the complexity of the signal interference strategy generation model is effectively reduced, combined with containerized deployment and double drift compensation training mechanism, the calculation efficiency and resource utilization are significantly improved while ensuring the model accuracy; its multi-level verification system and modular design enhance the robustness of the system in complex environments such as electromagnetic interference, providing a high-energy efficient deployment solution for signal shielding devices in resource-constrained scenarios.

[0034] 26、By using timestamp segmentation to splice the main key to generate a temporary key (the first 6 bits are used for parameter ciphertext, and the last 6 bits are used for verification ciphertext), a dynamic key generation mechanism is realized, and the key is different each time, effectively preventing the security risk of the whole system caused by key leakage; the parameter ciphertext uses MD5 + AES-128 double encryption, and the verification ciphertext uses HMAC + SM4 + ChaCha20 three encryption algorithms, forming a heterogeneous encryption protection layer, significantly improving the difficulty of brute force cracking; through character shifting (right shift 7 bits), number and letter permutation (6 / 9, A / C, 8 / B transposition), etc. Confusion means, increase the complexity of ciphertext feature analysis.

[0035] 27、By comparing the hash value (SHA256) to verify the consistency of the account, password, and timestamp original data, prevent tampering during transmission; through timestamp verification to ensure the timeliness of the instruction, avoid historical instruction replay attacks; through double authentication of account and password to establish operation permission control, prevent illegal devices from accessing.

[0036] 28、By appending 4 random letters at the end of UTF-8 encoded data to destroy the fixed structure of data blocks and increase the difficulty of ciphertext pattern analysis; by setting the verification ciphertext to use three-level encryption processing (SM4 encryption → hexadecimal conversion → character shift → character replacement → ChaCha20 encryption), a multi-layer encryption nested structure is formed, which greatly increases the reverse analysis cost.

[0037] 29. Through the integration of dynamic key generation mechanism (timestamp segmented splicing master key combined with MD5 / HMAC algorithm) and heterogeneous encryption algorithm (AES-128 / SM4 / ChaCha20 collaborative encryption), a multiple anti-cracking barrier is constructed. At the same time, a multi-dimensional security verification system of hash verification, time stamp verification and account password authentication is adopted to ensure the integrity, timeliness and legality of command transmission; its innovative anti-reverse design (random data expansion, character displacement and three-level encryption nesting) significantly increases the difficulty of ciphertext parsing, and combines with intelligent signal interference model to dynamically generate beamforming strategy, while achieving precise signal shielding, taking into account system security and energy efficiency optimization.

[0038] 30. By presetting the comparison table of noise intensity and power amplifier gain, the environmental noise intensity can be matched in real time, and the gain of the power amplifier can be dynamically adjusted to achieve accurate interference signal compensation, avoiding the over-compensation (energy waste) or under-compensation (shielding failure) problems in the traditional fixed gain mode; the hierarchical compensation strategy based on noise intensity only increases the power amplifier power when necessary, reducing the overall energy consumption of the signal blocker while maintaining a stable signal shielding effect.

[0039] 31. By setting up a hierarchical heat dissipation mechanism, in the active heat dissipation stage (temperature < threshold): the heat dissipation intensity is dynamically adjusted through the functional relationship between the device temperature and the fan speed to avoid the fan from running at full speed continuously, reducing mechanical wear and noise; in the emergency protection stage (temperature ≥ threshold): the fan is forced to run at full speed, and the power amplifier bias voltage is gradually reduced according to the voltage gradient. The temperature is quickly reduced through the dual paths of "reducing the power of the heat source + enhancing heat dissipation" to prevent hardware overheating and damage; a dynamic temperature control strategy is used to reduce the impact of drastic temperature fluctuations on electronic components, inhibit material aging caused by long-term high temperature, and improve equipment reliability.

[0040] 32. Directly match noise intensity and gain value through pre-stored lookup table, avoiding the computational overhead of real-time calculation of power amplifier parameters, reducing processor load and improving response speed; by encapsulating the relationship between fan speed and temperature into a configurable function, simplifying control logic, supporting flexible adjustment of parameters (such as function slope) to adapt to different cooling design requirements.

[0041] 33. The amplifier bias voltage is gradually reduced by a preset voltage gradient to avoid signal distortion or circuit impact caused by voltage drops, ensuring the system's stable transition to a safe state; a redundant protection level is provided through a temperature threshold trigger mechanism: heat dissipation is prioritized through fan speed regulation, and amplifier voltage adjustment is only intervened in extreme cases to maximize the shielding function.

[0042] 34. Adaptive signal compensation is achieved by dynamic matching of noise intensity and power amplifier gain, which optimizes energy consumption while ensuring shielding effectiveness. The device temperature is regulated in stages to functionally adjust fan speed and cooperatively control power amplifier voltage to dissipate heat, thereby avoiding hardware overheating risks and prolonging device life. The lookup table method and preset function design significantly reduce computational complexity and improve real-time response capability. The parameter configurability and modular architecture provide high scalability, ultimately achieving a multi-dimensional balance of performance, efficiency, and reliability in complex electromagnetic environments.

[0043] 35. By recording multi-dimensional data (shielding parameters, real-time wireless communication signals, real-time signal interference strategies, real-time interference signals, environmental noise data, device temperature, shielding time, account, and shielding effect feedback), a complete device operating state portrait is formed, providing high-quality feature input for signal interference strategy generation models and enhancing the precision of interference strategies.

[0044] 36. By using a data volume threshold triggering mechanism, resource waste caused by frequent training is avoided. Model backup and version rollback design ensures system reliability, preventing service interruption caused by training failure. The intelligent scheduling strategy for idle period training optimizes the utilization of computing resources.

[0045] 37. By using parameter aggregation instead of raw data transmission, data privacy protection requirements are met. Distributed signal shielding devices participate in global model optimization, improving the model's generalization ability in different environments (such as different regional electromagnetic environments). The edge-cloud collaborative architecture enables continuous model evolution, overcoming the limitations of insufficient data samples from a single device.

[0046] 38. By collecting multi-dimensional signal shielding logs in real time to build incremental data sets, and combining threshold-based triggering mechanisms and idle period scheduling strategies for low-power, high-reliability incremental training of local models, and innovatively introducing a federated learning architecture, the global model is collaboratively evolved using parameter aggregation from distributed devices while protecting data privacy. This forms a closed-loop optimization system of "environment perception-strategy generation-effect feedback-federated optimization," with dynamic environmental adaptability (such as noise adaptation and heat linkage), system stability guarantees (version rollback and verification mechanisms), and multi-device collaborative generalization advantages, significantly improving the precision and scene adaptability of signal interference strategies.

[0047] 39、Through the dynamic feature perception layer (integrating five-dimensional features of spectrum, time domain, space domain, protocol semantics and signal strength) and the multi-domain feature fusion layer (based on hypergraph neural network and dynamic weight distribution of reinforcement learning), precise signal feature extraction and cross-domain association are realized, combined with the multi-objective decision module (double branch network generates interference frequency points and power strategy) and the beam optimization module (differentiable convex optimization improves directivity) to generate efficient interference strategy; the data diversity is enhanced by using the generative adversarial network, and the lightweight adaptation is realized through model pruning, quantization and containerized deployment, while the instruction safety is protected by using multiple encryption (AES-128, SM4, ChaCha20) and dynamic verification mechanism, relying on the federal learning framework to optimize the model in real time, with the ability of environmental noise adaptive compensation, temperature intelligent regulation and log traceability, finally forming significant advantages in interference accuracy (frequency point hit rate), resource utilization (power dynamic optimization), real-time response (millisecond-level inference), anti-countermeasure ability (narrow beam direction) and compliance (frequency point constraint), realizing intelligent, low-power and highly concealed directional signal shielding. BRIEF DESCRIPTION OF DRAWINGS

[0048] The application will be further described below with reference to the accompanying drawings in conjunction with the embodiments.

[0049] Figure 1 is a flowchart of a signal shielding method based on deep reinforcement learning according to the application. DETAILED DESCRIPTION

[0050] The technical solutions in the embodiments of the present application have the following general idea: a signal interference strategy is generated by a pre-trained signal interference strategy generation model, a corresponding interference signal is generated based on the signal interference strategy, and signal shielding is performed, since the dynamic feature perception layer of the signal interference strategy generation model extracts frequency point migration features and carrier aggregation modes, and combines the interference frequency point selection matrix generated by the multi-target decision module of the interference strategy generation layer, the dynamic frequency point migration can be dynamically identified and adapted, rather than relying on a preset frequency point library; the three-dimensional correlation matrix in the multi-domain feature fusion layer models the cross-domain dependence of spectrum, time and space, adjusts the interference frequency points in real time, and adapts to the dynamic spectrum allocation scene; the interference strategy generation layer realizes accurate power allocation by using a power level vector and a beamforming weight value, avoids full-band sweeping, and the like; the real-time interference signal is power compensated by using environmental noise data, the working state of a power amplifier is dynamically adjusted based on the device temperature, power consumption is further optimized, and invalid radiation is reduced; the resource consumption loss in the loss function directly restricts the resource occupation of the power and beam weight, so that the model training is targeted at low power consumption; the protocol semantic coding module of the dynamic feature perception layer extracts communication protocol features, and the signal interference strategy generated by the interference strategy generation layer is targeted at specific protocol features, and continuously adapts to protocol evolution through federal optimization; the model volume is effectively reduced and the calculation amount is reduced by compressing the signal interference strategy generation model before deployment, so as to improve the accuracy, real-time performance and expansibility of signal shielding, and reduce the power consumption of signal shielding.

[0051] Please refer to Figure 1 The preferred embodiment of the signal shielding method based on deep reinforcement learning of the present application comprises the following steps: In step S1, a signal interference strategy generation model is constructed based on a dynamic feature perception layer, a multi-domain feature fusion layer and an interference strategy generation layer, and a loss function of the signal interference strategy generation model is set; The dynamic feature perception layer is constructed based on a spectrum dynamic coding module, a time domain feature coding module, a space domain feature coding module, a protocol semantic coding module and a signal strength coding module; the spectrum dynamic coding module, the time domain feature coding module, the space domain feature coding module, the protocol semantic coding module and the signal strength coding module are respectively used to extract spectrum features, time features, space features, communication protocol features and signal strength features from the input wireless communication signal; The five types of coding modules of the dynamic feature perception layer, i.e., spectrum, time domain, space domain, protocol semantic and signal strength, are integrated to realize holographic feature extraction of the wireless signal, and avoid the limitations of single feature analysis; for example, the protocol semantic coding can identify the vulnerabilities of specific communication protocols (such as Wi-Fi 6 and 5G NR), and improve the interference targeting.

[0052] The multi-domain feature fusion layer is constructed based on a feature interaction reinforcement module and a dynamic weight distribution module; the feature interaction reinforcement module is used to construct a three-dimensional correlation matrix according to spectral features, time features and spatial features; and the dynamic weight distribution module is used to evaluate the contribution degrees of the three-dimensional correlation matrix, communication protocol features and signal strength features to obtain fused features through fusion; The multi-domain correlation of signals is modeled through a three-dimensional correlation matrix (spectrum × time × space), and the feature contribution degrees are evaluated through a dynamic weight distribution module to realize adaptive scene feature fusion. This design is significantly better than traditional static weight fusion, and performs better in complex electromagnetic environments (such as urban dense multipath scenes).

[0053] The interference strategy generation layer is constructed based on a multi-objective decision module, a beam optimization module and an output module; the multi-objective decision module is used to generate an interference frequency point selection matrix and a power level vector according to fused features; the beam optimization module is used to generate beamforming weights according to fused features; and the output module is used to generate a signal interference strategy according to the interference frequency point selection matrix, the power level vector and the beamforming weights; The multi-objective decision module and the beam optimization module are used in the interference strategy generation layer to simultaneously optimize the interference frequency point selection matrix, the power level vector and the beamforming weights, taking into account interference efficiency and energy consumption control. For example, in a low signal-to-noise ratio scene, high-frequency band interference is preferentially selected, and invalid power distribution is reduced.

[0054] In step S2, a large number of historical wireless communication signals are acquired, and after preprocessing and labeling, a signal data set is constructed, and the sample size of the signal data set is expanded through a generative adversarial network; In step S3, the signal data set is divided into a training set, a validation set and a test set, and a signal interference strategy generation model is trained, validated and tested through the training set, the validation set and the test set; In step S4, the signal interference strategy generation model that passes the test is compressed, and the compressed signal interference strategy generation model is deployed to a signal shield. In step S5, after the signal shield is started, an input signal shielding instruction is acquired, the signal shielding instruction is verified and analyzed to obtain shielding parameters, real-time wireless communication signals are collected, the real-time wireless communication signals are preprocessed, the real-time wireless communication signals and the shielding parameters are input into the deployed signal interference strategy generation model, and a real-time signal interference strategy is obtained. In step S6, the signal shield modulates a corresponding real-time interference signal based on the real-time signal interference strategy, a power amplifier performs power amplification on the real-time interference signal through digital pre-distortion technology, and the real-time interference signal is transmitted externally through a radio frequency antenna to perform a signal directional shielding operation. A millisecond-level closed loop is formed by collecting real-time interference signal emission from real-time wireless communication signals, and a real-time wireless communication signal is preprocessed in combination with a lightweight model deployment collaboration to meet the real-time tracking requirements of frequency hopping signals.

[0055] In step S7, the signal shield real-time collects environmental noise data and equipment temperature, compensates the power of the real-time interference signal based on the environmental noise data, and dynamically adjusts the working state of the power amplifier and the fan based on the equipment temperature; Through real-time compensation of environmental noise and temperature control linkage, dynamic adjustment of interference power and equipment thermal management are realized; for example, in a high-temperature environment, the power amplifier output is automatically reduced and the heat dissipation is accelerated to avoid signal shield downtime.

[0056] In step S8, the signal shield real-time records signal shielding logs, and federated optimizes the signal interference strategy generation model based on the signal shielding logs.

[0057] Through the generation of a generative adversarial network (GAN) to expand the signal data set, the problem of insufficient rare signal samples in actual scenarios (such as emergency communication signals) is effectively solved, and the model generalization ability is improved; through the federated optimization mechanism, the signal shielding logs of multiple signal shields are aggregated to update the signal interference strategy generation model, realizing the coexistence of knowledge sharing and privacy protection; for example, signal shields in different regions can cooperate to optimize the protocol semantic recognition module without uploading original signal data; through the combination of deep reinforcement learning strategy and digital pre-distortion (DPD) technology in the interference signal generation stage, the non-linear distortion of the power amplifier is suppressed, the focusing of the interference signal in the target frequency band is ensured, and the risk of out-of-band radiation is reduced.

[0058] Through the dynamic collaboration of deep reinforcement learning and multi-domain feature fusion (spectrum, time domain, space domain, protocol semantics and signal strength), high-precision signal interference strategy generation is realized, and its core advantages are: multi-dimensional perception (holographic feature extraction and three-dimensional correlation matrix modeling), dynamic optimization (feature weight distribution based on attention mechanism and multi-objective decision), real-time closed-loop control (millisecond-level interference strategy generation and digital pre-distortion power compensation), and edge intelligence deployment (model compression combined with federated optimization technology), while having high energy efficiency (dynamic power adjustment and temperature adaptive management), strong robustness (anti-adversarial attack design and hardware fault tolerance mechanism) and environmental friendliness (directional beam reduces electromagnetic pollution), which can be extended to support future 6G and quantum secure communication scenarios, and realize precise, compliant and low-power intelligent signal shielding in military, civilian and emergency scenarios.

[0059] In step S1, the frequency spectrum dynamic coding module is constructed based on a dynamic coding unit, a frequency domain correlation unit and an aggregation analysis unit; The spectrum dynamic coding module innovatively fuses local features (multi-scale convolution) and global dependence (multi-head self-attention), captures migration patterns through frequency point clustering, and realizes fine analysis of frequency point dynamic features by combining a multi-resolution matching mechanism; the signal strength coding adopts an adaptive normalization layer to realize dynamic mapping of RSSI and MCS, overcoming the lack of adaptability of traditional fixed normalization to rapid changes in signal strength.

[0060] The dynamic coding unit is configured to extract local frequency domain features and time sequence dynamic features from the input wireless communication signal through a multi-scale convolution subunit and a gated recurrent subunit, and output time-frequency joint features including the local frequency domain features and the time sequence dynamic features; the multi-scale convolution subunit extracts local frequency domain features in parallel using convolution kernels of different sizes (such as 1x3, 1x5, and 1x7), and fuses multi-scale local frequency domain features through residual connection; the gated recurrent subunit is configured to encode signal time sequence slices of the wireless communication signal, capture short-time dynamic changes of frequency point migration, and output time sequence dynamic features. The frequency domain correlation unit is configured to divide adjacent frequency points into the same cluster through a frequency point clustering subunit to capture time sequence dependence of frequency point migration from the time-frequency joint features, calculate cross-frequency point attention weights of the time-frequency joint features through a multi-head self-attention subunit to model global dependence between frequency points, and capture frequency domain correlation of frequency point migration, and generate a frequency point dynamic coding vector based on the time sequence dependence and the frequency domain correlation; the frequency point clustering subunit is constructed based on a dynamic threshold gated clustering algorithm; the frequency point dynamic coding vector is configured to represent a frequency point migration path and a frequency domain correlation strength. The aggregation analysis unit is configured to perform multi-resolution matching of carrier aggregation patterns on the frequency point dynamic coding vector through a pattern matching subunit, and decouple frequency point migration features from the frequency point dynamic coding vector through a feature reconstruction subunit, and output spectrum features including the frequency point migration features and the carrier aggregation patterns. The pattern matching subunit uses a set of dynamic convolution kernels (which can be understood as "templates" of different sizes and shapes) to slide on the frequency point dynamic coding vector to find combination rules of discrete carrier bands at different resolutions, i.e., carrier aggregation patterns; for example, some frequency point combinations are found to frequently occur, and these combination rules represent different carrier aggregation patterns (carrier aggregation features). Since the frequency point migration features and the carrier aggregation patterns are usually mixed together in the frequency point dynamic coding vector, the feature reconstruction subunit uses an orthogonal constraint projection matrix to decouple the frequency point migration features and the carrier aggregation features in the frequency point dynamic coding vector, and generates dimension-independent joint spectrum features (frequency point migration features and carrier aggregation patterns) which can be analyzed and processed separately.

[0061] The polygenic analysis unit decouples the frequency point migration features and the carrier aggregation mode through a feature reconstruction subunit, realizes the separation of the physical layer dynamic features and the protocol layer configuration features, and enhances the explainability of the model to the complex signal interference scene; the spatial feature coding module independently extracts the signal topology structure (graph attention network) and the beam spatial mode (spatial spectrum estimation), avoiding the parameter redundancy problem caused by the coupling of the spatial features in the traditional method.

[0062] The time domain feature coding module is configured to extract time features including signal burst characteristics and time-frequency interleaving patterns from the input wireless communication signal through a time domain convolution network and a multi-head time attention mechanism. The time domain convolution network first divides the wireless communication signal into many small segments through a sliding window method, extracts the local features of each small segment, and then the multi-head time attention mechanism finds out the more important parts from these local features and integrates them; for example, the time domain convolution network finds that the waveforms of several small segments in the wireless communication signal are special, and the multi-head time attention mechanism judges whether these special waveforms are related to the mode of the whole wireless communication signal, and then combines these important information to form a more comprehensive time feature, so that the time feature containing both signal burst characteristics (local burst characteristics) and considering time-frequency interleaving mode (overall mode) can be extracted from the wireless communication signal.

[0063] The spatial feature coding module is configured to extract spatial features including signal topology structure features and beam spatial mode vectors from the input wireless communication signal through a graph attention network and a spatial spectrum estimation. The wireless communication signal is transmitted through a multiple-input multiple-output (MIMO) channel, and the MIMO channel matrix contains the signal transmission information between multiple transmitting antennas and receiving antennas; the graph attention network models the transmission relationship of the wireless communication signal between these antennas into a graph, each antenna is a node, and the signal transmission strength and phase information between the antennas are edges between the nodes; the graph attention network analyzes the relationship between the nodes through the attention mechanism, automatically learns which antenna signal relationship is more important, and gives higher weight to these important relationships; for example, the signal transmission between some antennas is strong and stable, and these antenna relationships are focused on, in this way, the graph attention network can extract the topology structure features of the signal transmission, that is, how the wireless communication signal influences and transmits between different antennas.

[0064] Spatial spectrum estimation utilizes the wireless communication signals received by an antenna array to estimate the direction of arrival (DOA) of the signals, which determines the direction from which the signals come by analyzing the phase and amplitude differences of the wireless communication signals at different antennas; for example, a wireless communication signal arrives at an antenna array from a certain direction, then the wireless communication signals received by different antennas will have certain phase differences, by calculating these phase differences, spatial spectrum estimation can generate a beam spatial mode vector, which represents the intensity distribution of the wireless communication signals in different directions, can intuitively reflect the propagation direction and intensity of the wireless communication signals, and help understand where the wireless communication signals come from and the distribution in space.

[0065] The signal topology structure features extracted by the graph attention network and the beam spatial mode vector generated by the spatial spectrum estimation are complementary; the graph attention network focuses on the relationship and transmission path between signals, while the spatial spectrum estimation focuses on the direction and intensity distribution of the signals; when the two are combined, the graph attention network can use the topology structure features it extracts to optimize the accuracy of the spatial spectrum estimation; for example, if the graph attention network finds that the relationship between certain antennas is very important, then the spatial spectrum estimation can more accurately estimate the signals in the direction of these antennas; at the same time, the beam spatial mode vector generated by the spatial spectrum estimation can also provide additional context information for the graph attention network, helping it better understand the propagation of signals in space.

[0066] The protocol semantic encoding module is configured to extract communication protocol features from the input wireless communication signals through the BERT unit and the 3GPP standard knowledge graph; BERT (Bidirectional Encoder Representations from Transformers) is a natural language processing model based on the Transformer architecture, capable of processing sequential data such as text. In wireless communication signal processing, protocol control fields (such as the CFI field in DCI) can be considered as a special "text" sequence, which contains key information of the communication protocol, such as channel format indication. 3GPP (3rd Generation Partnership Project) is an organization that develops mobile communication standards. Its standard knowledge graph is a structured representation of various protocol rules, parameters, and other knowledge in 3GPP standards, like a knowledge base, containing various semantic information of communication protocols, such as the meaning of different fields and the relationship between fields. The communication protocol feature extraction process is as follows: (1) BERT unit processing protocol control fields: First, the protocol control fields in wireless communication signals (such as the CFI field in DCI) are input into the BERT unit, which encodes and processes these fields, understands the sequence information in the fields, and extracts the semantic features of the fields. For example, the CFI field may contain information about the size of the control region in the subframe. The BERT unit can convert this information into a vector form of semantic representation, which can reflect the meaning of the field. (2) Combine 3GPP standard knowledge graph for protocol intent reasoning: Combine the semantic features extracted by the BERT unit with the 3GPP standard knowledge graph. The knowledge in the 3GPP standard knowledge graph can further interpret and reason the semantic features extracted by the BERT unit. For example, the 3GPP standard knowledge graph stores the meaning of the channel format corresponding to different values of the CFI field. When the BERT unit extracts the semantic features of the CFI field with a certain value, the 3GPP standard knowledge graph can use its stored rules to infer the communication protocol intent corresponding to this value. This intent may be about how to allocate control region resources or how to control data transmission, etc. (3) Extract communication protocol features: Through this collaborative work, communication protocol features are finally extracted from wireless communication signals. These features include semantic information of protocol fields and protocol intent, which can help subsequent protocol parsing, signal processing, and other operations, thereby better understanding and processing wireless communication signals.

[0067] By innovatively integrating the BERT pre-training model and the 3GPP knowledge graph through the protocol semantic encoding module, the semantic understanding ability of large-scale language models is utilized, and communication standard prior knowledge is injected, significantly improving the accuracy of protocol feature parsing. By introducing carrier aggregation mode matching in the frequency migration modeling, the physical layer features are organically combined with high-level protocol states, enhancing the adaptability to modern communication systems (such as 5G NR).

[0068] The signal strength encoding module is configured to dynamically map the RSSI value and the MCS carried in the input wireless communication signal through an adaptive normalization layer, and then extract the signal strength feature; The signal strength of a wireless communication signal can vary due to various factors such as distance, obstacles, etc. The RSSI (Received Signal Strength Indicator) value is used to measure the strength of the signal. At the same time, the MCS (Modulation and Coding Scheme) determines the transmission method of the signal, and different MCS corresponds to different signal quality and transmission efficiency. Since the range of RSSI value can be very large, and the relationship between RSSI value and MCS will also be different in different environments and devices, the task of the adaptive normalization layer is to dynamically adjust this relationship. Specifically, the adaptive normalization layer will continuously observe the current received signals, and then automatically adjust a mapping relationship according to the RSSI values and corresponding MCS of these signals. This mapping relationship is like a "translator" that translates the RSSI value into a standardized signal strength feature, so that it can be more convenient to compare and analyze the strength of different signals. Assuming that in one environment, the RSSI value varies between -50 and -100, and in another environment, the RSSI value varies between -30 and -80, the adaptive normalization layer will dynamically adjust the mapping relationship according to these different ranges, so that in each environment, a reasonable and comparable signal strength feature can be obtained.

[0069] Through the five-way encoding modules of spectrum, time domain, space domain, protocol semantics, and signal strength, the signal is represented in multiple dimensions, covering the frequency domain dynamics (multi-scale convolution + gated recurrent), time domain burstiness (time attention), spatial topology (graph attention), protocol specification, and intensity dynamics. Through the feature interaction reinforcement module, the hypergraph neural network is used to model cross-domain association, and the three-dimensional association matrix breaks through the limitation of traditional two-dimensional feature interaction, effectively capturing the time-frequency-space joint domain feature correlation.

[0070] The feature interaction reinforcement module is configured to model cross-domain feature association of the spectrum feature, the time feature, and the space feature through a hypergraph neural network, and then construct a three-dimensional association matrix. A typical graph neural network is like a network of points and lines, where points represent data and lines represent relationships between them. However, a hypergraph neural network can handle not only points and lines, but also more complex structures. For example, a point can have complex relationships with multiple points simultaneously. This is like a social network that can represent not only the relationship between two people but also the complex interactions between a group of people. Spectral features, temporal features, and spatial features are originally independent, like three different worlds. The hypergraph neural network is like a super "bridge builder," connecting these three worlds by constructing a complex "bridge" network. In this network, each feature is treated as a node, and the hypergraph neural network establishes various connections based on the complex relationships between these features. These connections include not only simple "one-to-one" relationships, but also more complex "one-to-many" and "many-to-many" relationships. In this way, the hypergraph neural network can capture the complex interactions between these features. The three-dimensional correlation matrix can be imagined as a three-dimensional table with three dimensions, corresponding to spectrum, time, and space. In this table, the value of each position represents the strength of the correlation between the corresponding spectrum, time, and space features. For example, if a position in the matrix has a high value, it means that the correlation between the features is very strong under that specific spectrum, time, and space combination.

[0071] The dynamic weight allocation module is used to evaluate the contribution of the three-dimensional association matrix, communication protocol features, and signal strength features to the current interference scenario through a value network based on a policy gradient method of reinforcement learning, so as to fuse them to obtain a fusion feature; The policy gradient method is a reinforcement learning method. Reinforcement learning is like training a puppy: rewarding it for correct behavior and punishing it for errors. The policy gradient method learns in this way. It makes decisions based on the current features and environment, and then adjusts its decision-making strategy based on reward signals, making good decisions more likely to be selected and gradually eliminating bad ones. The value network acts like an evaluator, evaluating the contribution of different features (such as the three-dimensional correlation matrix, communication protocol features, and signal strength features) to the current interference scenario. These features are like different factors, some of which have a greater impact on the interference scenario, while others have a smaller impact. The value network assigns a score to the dynamic weight allocation module to indicate the importance of each feature. The weighted sum of these scores is then used to generate the fused features.

[0072] The dynamic weight allocation module realizes dynamic evaluation of feature contribution degree through a policy gradient method, and adjusts weights online through a value network, which is more suitable for time-varying characteristics of complex interference scenes than a fixed weight fusion strategy; the multi-objective decision module realizes joint optimization of interference frequency point selection (discrete decision) and power level adjustment (continuous decision) through a double-branch network structure, thereby avoiding error accumulation problems of traditional cascading decisions.

[0073] The multi-objective decision module is configured to infer the fused features through the double-branch network structure to generate an interference frequency point selection matrix and a power level vector. The double-branch network structure is a common neural network architecture, which is widely used in multi-task learning, multi-objective decision and other scenarios. This structure divides the network into two independent branches, each of which is responsible for processing different tasks or sub-tasks, thereby improving the efficiency and pertinence of the model.

[0074] The beam optimization module infers the fused features through the differentiable convex optimization layer and the SDP relaxation method to generate beamforming weights. Beamforming is a wireless communication technology that adjusts the direction of the antenna beam to concentrate the signal in a specific direction, thereby improving the signal strength and reducing interference. Beamforming weights are key parameters for controlling the direction of the beam. In mathematics, differentiability means that the function has a derivative at every point, and gradient descent and other optimization algorithms can be used to adjust the parameters. In beam optimization, differentiability allows the beamforming weights to be adjusted through backpropagation to minimize the error. Convex optimization is a special optimization problem whose objective function and constraint conditions are convex. An important feature of convex optimization problems is that the local optimal solution is the global optimal solution, which makes the optimization process more efficient and reliable. The role of the differentiable convex optimization layer is to convert the complex beamforming weight calculation problem into a convex optimization problem, thereby efficiently finding the optimal solution. SDP (Semi-Definite Programming) is a special convex optimization problem whose variable is a semi-definite matrix. In beam optimization, SDP can be used to handle complex constraints such as power constraints and interference constraints. Relaxation method is a technique for converting complex non-convex problems into convex problems. In beam optimization, the original beamforming weight calculation problem may be non-convex and difficult to solve directly. Through the SDP relaxation method, these non-convex constraints can be relaxed to convex constraints, thereby converting the problem into a solvable SDP problem.

[0075] That is, the workflow of the beam optimization module is as follows: first, input the fused features into the differentiable convex optimization layer, process these fused features through the convex optimization algorithm to obtain a preliminary beamforming weight; then use the SDP relaxation method to further optimize the preliminary beamforming weight to find the optimal beamforming weight.

[0076] By combining the differentiable convex optimization layer with the SDP relaxation method, the gradient is derived under the premise of guaranteeing the physical constraints of beamforming, breaking the bottleneck that traditional convex optimization methods cannot be trained end-to-end; by introducing the main lobe direction cosine similarity metric in the beam optimization loss function, it is more consistent with the beam space distribution characteristics compared with the Euclidean distance.

[0077] The formula of the loss function is: L = α * L int + β * L res + γ * L beam + δ * L fusion ; Wherein, L represents the loss value of the loss function; L int represents the interference effect loss, which is used to measure the accuracy of the interference frequency point selection and power allocation; L res represents the resource consumption loss, which is used to constrain the total transmit power and beamforming weight complexity; L beam represents the beam optimization loss, which is used to measure the deviation of the beam main lobe direction from the ideal direction; L fusion represents the feature fusion loss, which is used to constrain the rationality of multi-domain feature fusion; α, β, γ and δ all represent multi-objective weighting coefficients, which are optimized by grid search; ; ; ; ; Wherein, f represents the frequency point; represents the probability that the frequency point f is selected; represents the indicator function; represents the target interference frequency point set; 、 and all represent the regularization strength coefficient; represents the power level vector prediction value; represents the theoretical optimal power level; represents the Euclidean distance square of the vector; represents the beamforming weight matrix; represents the L1 norm; represents the Frobenius norm square; represents the beam main lobe prediction direction; represents the beam main lobe ideal direction; represents the three-dimensional correlation matrix; represents the communication protocol feature; represents the signal strength feature; represents the fusion feature; Dynamic weight multiplication fusion is indicated. Dynamic weight addition fusion is indicated.

[0078] By designing fusion interference effects (cross-entropy + power deviation), resource consumption (L1 + F norm), beam direction (cosine similarity), and feature fusion error into the loss function, the multi-objective weighting coefficients are optimized through grid search to balance the system-level performance; By introducing a dynamic regularization mechanism (λ1, λ2, λ3 coefficients), the model overfitting is effectively prevented, and the Lres loss term simultaneously constrains power consumption and beam energy dispersion, embodying the green communication concept.

[0079] Through multi-dimensional feature fusion (spectrum dynamic coding, time domain feature coding, spatial domain feature coding, protocol semantic coding, and signal strength coding) and a dynamic decision mechanism of reinforcement learning, fine analysis of signal features and intelligent interference decision are realized, and the core advantages are as follows: a cross-domain dynamic feature coding network is constructed by innovatively combining multi-scale convolution, a gated recurrent unit, and a multi-head attention mechanism, breaking through the limitations of traditional single-domain feature modeling; a differentiable convex optimization layer and a hypergraph neural network are used to realize joint optimization of beamforming and feature association, while ensuring physical constraints and improving computational efficiency; through dynamic weight distribution and multi-objective loss function design, the interference performance, resource efficiency, and system robustness are effectively balanced, and 3GPP multi-protocol compatibility, hardware deployment friendliness, and cross-scene migration ability are combined, providing an end-to-end solution with high precision, low latency, and green energy saving for intelligent spectrum confrontation in complex electromagnetic environments.

[0080] The step S2 specifically includes: A large number of historical wireless communication signals are acquired, each of the historical wireless communication signals is preprocessed including at least signal amplification, filtering, signal down-conversion, and signal sampling, each of the preprocessed historical wireless communication signals is labeled including at least frequency point migration feature, carrier aggregation mode, signal burst characteristic, time-frequency interleaving mode, signal topology structure feature, beam space modal vector, communication protocol feature, signal strength feature, signal type, and signal interference strategy, a data set is constructed based on each of the labeled historical wireless communication signals, and a generator in a pre-trained generative adversarial network is used to expand the sample amount of the signal data set; Through the labeled content of the signal data set, frequency point migration, carrier aggregation, signal burst characteristics, time-frequency interleaving, and other physical layer features are covered, and communication protocols, signal types, and other protocol layer information are included, realizing full-dimensional modeling of signals; through the feature labeling of the beam space modal vector and the signal interference strategy, the signal characteristics of complex communication scenarios such as multi-antenna systems (MIMO) and dynamic spectrum sharing can be effectively analyzed.

[0081] Through signal amplification, filtering and down-conversion operations, environmental noise and hardware distortion are eliminated, the signal signal-to-noise ratio (SNR) is improved, and high-fidelity data is provided for subsequent labeling and modeling; Through signal amplification, filtering and down-conversion operations, environmental noise and hardware distortion are eliminated, the signal signal-to-noise ratio (SNR) is improved, and high-fidelity data is provided for subsequent labeling and modeling.

[0082] Through multi-dimensional signal feature labeling (covering frequency point migration, carrier aggregation, communication protocol and other parameters from the physical layer to the protocol layer) and standardized preprocessing (signal amplification, filtering, down-conversion), the data representation accuracy and signal-to-noise ratio are significantly improved; Combined with pre-trained generative adversarial network (GAN), the adversarial samples are generated, the sample scarcity bottleneck of real communication signal collection is broken through, and the robustness of the AI model to multiple scenarios (such as channel fading and multi-device interference) is enhanced.

[0083] The step S3 is specifically: Based on k-fold cross-validation method, the signal data set is divided into training set, validation set and test set in the ratio of 8:1:1, the signal interference strategy generation model is trained through the training set, and in the training process, the hyperparameters including at least learning rate, regularization parameter, network structure parameter, batch size, training round, activation function, optimizer and random drop rate are optimized through the Bayesian optimization framework until the loss value of the loss function is less than the preset loss threshold or the preset early stopping condition is reached; The trained signal interference strategy generation model is verified through the validation set to determine whether the interference strategy accuracy is greater than the preset accuracy threshold, if not, the verification fails, and the training set is expanded for further training; if yes, the verification passes, and: the frequency point hit rate, the false interference rate, the power utilization rate, the power dynamic adjustment response time, the beam pointing accuracy, the strategy conflict rate and the single inference time consumption are calculated through the test set to test the signal interference strategy generation model that passes the verification, if the test fails, the training set is expanded for further training; if the test passes, the training ends.

[0084] Through k-fold cross-validation method combined with 8:1:1 ratio to divide the data set, the dual needs of data utilization efficiency and generalization ability verification are considered; k-fold cross-validation effectively reduces the randomness deviation of single data division, and improves the stability of the model; The golden ratio of 8:1:1 ensures that the training data is sufficient (80%) while the independence of the validation set (10%) and the test set (10%) is verified; The double verification mechanism (validation set + test set) forms a closed loop feedback, which significantly reduces the risk of overfitting.

[0085] A probability model is established through a Bayesian optimization framework to guide the parameter search direction, which effectively improves the search efficiency compared with grid search / random search; meanwhile, heterogeneous parameters such as learning rate, regularization parameter, network structure parameter, etc. are optimized, breaking through the limitations of traditional single parameter optimization; the random dropout rate is introduced as an optimization variable, creatively solving the parameter self-adaption problem of the Dropout mechanism in deep learning.

[0086] The preset loss threshold is used to ensure that the basic performance of the model meets the standard; the early stopping condition is set to prevent invalid training resource consumption, which is particularly suitable for complex models with high-dimensional parameter space; when the verification fails, the training set is automatically expanded to realize the cooperative optimization of data enhancement and model training.

[0087] By innovatively combining k-fold cross-validation and 8:1:1 data layering mechanism, and intelligently cooperating the multi-dimensional hyperparameters through the Bayesian optimization framework, a closed-loop training system is constructed, which includes dynamic early stopping threshold and training set adaptive expansion. In the test stage, multiple-dimensional evaluation indexes such as frequency hit rate and false interference rate are introduced, effectively solving the technical pain points of low training efficiency and insufficient generalization ability of traditional interference strategy models, significantly improving the model convergence speed, strategy generation accuracy and anti-interference performance in practical applications, especially achieving the balance optimization of interference strategy precision and resource utilization efficiency in the electromagnetic countermeasure scene.

[0088] The step S4 is specifically: The weight parameters of the signal interference strategy generation model that passes the test are sorted according to the absolute value of the weight value, and the weight values below the preset weight threshold are set to zero to perform pruning operation on the signal interference strategy generation model. After the pruning operation, the signal interference strategy generation model is verified for performance, and if the performance verification fails, the weight threshold is adjusted to perform pruning operation again; if the performance verification passes, then: The signal interference strategy generation model is quantized based on the preset quantization step and offset, and the quantized signal interference strategy generation model is verified for performance again. If the performance verification fails, the quantization step and offset are adjusted to perform quantization operation again; if the performance verification passes, the compression of the signal interference strategy generation model is completed. The compressed signal interference strategy generation model is deployed to the signal shield device through containerization technology, and the signal interference strategy generation model deployed to the signal shield device is trained for data drift compensation and model drift compensation. By adopting the pruning strategy of absolute value sorting + threshold zeroing, the model parameter quantity is effectively reduced while the core features are retained, the threshold is dynamically adjusted through the loop verification mechanism, and the problem of key weight deletion caused by traditional pruning is solved; by introducing the offset compensation quantization algorithm (better than the conventional linear quantization), the step size parameter is optimized through secondary performance verification closed loop, which effectively improves the compression ratio on the premise of reducing the model accuracy loss.

[0089] The complexity of the signal interference strategy generation model is effectively reduced through two-stage model compression (adaptive pruning based on dynamic threshold adjustment and quantization parameter optimization), and the computing efficiency and resource utilization are significantly improved while ensuring the model accuracy through containerized deployment and double-drift compensation training mechanism; the multi-level verification system and modular design enhance the robustness of the system in complex environments such as electromagnetic interference, providing a high-energy efficient deployment solution for signal shielders in resource-constrained scenarios.

[0090] The step S5 is specifically: After the signal shielder is started, the input signal shielding instruction carrying the parameter ciphertext and the verification ciphertext is obtained; the parameter ciphertext is obtained by encrypting the shielding parameter through a first encryption rule; the verification ciphertext is obtained by encrypting the account, password, timestamp, and hash value through a second encryption rule; the hash value is obtained by hashing the account, password, and timestamp; the shielding parameter at least includes shielding duration, shielding range, and shielding signal type; The signal shielder parses the signal shielding instruction to obtain the parameter ciphertext and the verification ciphertext, decrypts the verification ciphertext to obtain the account, password, timestamp, and hash value through the second encryption rule, and then performs integrity verification on the account, password, and timestamp through the hash value, time-limited verification through the timestamp, and legality verification through the account and password. If all verifications are passed, the parameter ciphertext is decrypted through the first encryption rule to obtain the shielding parameter; The signal shielder collects real-time wireless communication signals, pre-processes the real-time wireless communication signals through filtering, and then inputs the real-time wireless communication signals and the shielding parameter into the deployed signal interference strategy generation model to obtain a real-time signal interference strategy including an interference frequency point selection matrix, a power level vector, and a beamforming weight value; The first encryption rule is specifically: the first 6 digits of the current timestamp are concatenated with a preset master key to obtain first concatenated data, and the MD5 value of the first concatenated data is calculated as a first temporary key through the MD5 algorithm; the shielding parameter is UTF-8 encoded to obtain encoded data, and 4 random letters are appended at the end of the encoded data to generate an extended data block; the first temporary key is called through the AES-128 algorithm to encrypt the extended data block to obtain the parameter ciphertext; The second encryption rule is specifically: The last 6 digits of the current timestamp are concatenated with the preset master key to obtain second concatenated data, and the MAC value of the second concatenated data is calculated as a second temporary key through an HMAC algorithm; the account, password, timestamp, and hash value are encrypted through an SM4 algorithm to obtain first-level encrypted data by calling the second temporary key, the first-level encrypted data is converted into hexadecimal data, each character of the hexadecimal data is cyclically shifted 7 bits to the right to obtain second-level encrypted data, the numbers 6 and 9 in the second-level encrypted data are transposed, the letters A and C are transposed, and the number 8 and the letter B are transposed to obtain third-level encrypted data, and the third-level encrypted data is encrypted through a ChaCha20 algorithm by calling the second temporary key to obtain verification ciphertext.

[0091] By using the timestamp segmentation and concatenation of the master key to generate a temporary key (the first 6 digits are used for parameter ciphertext, and the last 6 digits are used for verification ciphertext), a dynamic key generation mechanism is realized, and the key is different each time, which effectively prevents the security risk of the whole system caused by key leakage; the parameter ciphertext is double-encrypted by MD5+AES-128, and the verification ciphertext is combined using three encryption algorithms of HMAC+SM4+ChaCha20, forming a heterogeneous encryption protection layer, which significantly improves the difficulty of brute force cracking; through character shifting (7-bit right shift), number-letter permutation (6 / 9, A / C, 8 / B transposition), and other confusion means, the complexity of ciphertext feature analysis is increased.

[0092] The consistency of the account, password, and timestamp original data is verified through hash value comparison (SHA256), to prevent tampering during transmission; the time stamp verification ensures the timeliness of the instruction, avoiding historical instruction replay attacks; the account password double authentication establishes operation permission control, preventing illegal device access.

[0093] By appending 4 random letters at the end of the UTF-8 encoded data, the fixed structure of the data block is destroyed, and the difficulty of ciphertext pattern analysis is increased; by setting the verification ciphertext to be processed by three-level encryption (SM4 encryption→hexadecimal conversion→character shift→character replacement→ChaCha20 encryption), a multi-layer encryption nested structure is formed, which greatly improves the reverse analysis cost.

[0094] Through the dynamic key generation mechanism (time stamp segmentation splicing master key combined with MD5 / HMAC algorithm) and the heterogeneous encryption algorithm fusion (AES-128 / SM4 / ChaCha20 cooperative encryption), a multi-prevention cracking barrier is constructed, and a multi-dimensional security verification system of hash check, timestamp verification and account password authentication is adopted to ensure the integrity, timeliness and legality of the instruction transmission; the innovative anti-reverse design (random data expansion, character bit displacement and three-level encryption nesting) significantly improves the difficulty of ciphertext analysis, and the intelligent signal interference model is used to dynamically generate a beamforming strategy, which realizes accurate signal shielding while considering system security and energy efficiency optimization.

[0095] The step S7 is specifically: The signal shielding device is preconfigured with a control table for storing the corresponding relationship between noise intensity intervals and power amplifier gains, and a fan control function with a device temperature as an independent variable and a fan speed as a dependent variable; The signal shielding device collects environmental noise data in real time, matches the power amplifier gain based on the noise intensity of the environmental noise data from the control table, and controls the power amplifier to perform power compensation on the real-time interference signal based on the power amplifier gain; The signal shielding device collects the device temperature in real time, inputs the device temperature into the fan control function to obtain the corresponding fan speed when the device temperature is less than a preset temperature threshold, and dynamically adjusts the working state of the fan based on the fan speed; when the device temperature is greater than or equal to the temperature threshold, the fan speed of the fan is adjusted to the maximum, and the bias voltage of the power amplifier is continuously reduced based on a preset voltage gradient until the device temperature is less than the temperature threshold; Through the control table of the preset noise intensity and the power amplifier gain, the environmental noise intensity can be matched in real time, the gain of the power amplifier can be dynamically adjusted, accurate interference signal compensation can be realized, and the problems of overcompensation (energy waste) or undercompensation (shielding failure) under the traditional fixed gain mode can be avoided; based on the hierarchical compensation strategy of the noise intensity, the power amplifier power is only increased when necessary, the overall energy consumption of the signal shielding device is reduced, and the stable signal shielding effect is maintained.

[0096] By setting a hierarchical heat dissipation mechanism, in the active heat dissipation stage (temperature < threshold value): the heat dissipation intensity is dynamically adjusted through the functional relationship between the device temperature and the fan speed, the fan is prevented from running at full speed continuously, and mechanical wear and noise are reduced; in the emergency protection stage (temperature ≥ threshold value): the fan is forced to run at full speed, and the power amplifier bias voltage is gradually reduced according to the voltage gradient, the power of the heat source is reduced and the heat dissipation is enhanced through the double paths of “reducing the power of the heat source + enhancing the heat dissipation” to quickly cool down and prevent hardware from being damaged due to overheating; through the dynamic temperature control strategy, the impact of temperature fluctuations on electronic components is reduced, material aging caused by long-term high temperature is inhibited, and the reliability of the device is improved.

[0097] The noise intensity and gain value are directly matched through a pre-stored control table to avoid real-time calculation of the power amplifier parameters, reduce the processor load, and improve the response speed. The relationship between the fan speed and the temperature is packaged as a configurable function to simplify the control logic and support flexible adjustment of parameters (such as the function slope) to adapt to different heat dissipation design requirements.

[0098] The power amplifier bias voltage is gradually reduced through a pre-set voltage gradient to avoid signal distortion or circuit impact caused by voltage drop, and to ensure stable transition of the system to a safe state. A temperature threshold trigger mechanism provides a redundant protection level: the fan speed is adjusted for heat dissipation, and only in extreme cases is the power amplifier voltage adjusted to maximize the shielding function.

[0099] Adaptive signal compensation is achieved through dynamic matching of noise intensity and power amplifier gain to optimize energy consumption while ensuring shielding effectiveness. A temperature grading control mechanism is combined to regulate the fan speed adjustment and gradient power amplifier voltage control to synergistically dissipate heat, avoiding hardware overheating risks and extending device life. The lookup table method and pre-set function design significantly reduce computational complexity and improve real-time response capability. The parameter configurability and modular architecture provide high scalability, ultimately achieving a multi-dimensional balance of performance, efficiency, and reliability in complex electromagnetic environments.

[0100] The step S8 specifically includes: The signal shielding device records a signal shielding log in real time, which includes at least shielding parameters, real-time wireless communication signals, real-time signal interference strategies, real-time interference signals, environmental noise data, device temperature, shielding time, account, and shielding effect feedback. Based on the signal shielding log, an incremental data set is constructed. When the data volume of the incremental data set exceeds a pre-set quantity threshold, the signal interference strategy generation model is backed up, and in the idle period, the signal interference strategy generation model is trained through the incremental data set. The trained signal interference strategy generation model is performance verified. If the verification fails, a version rollback operation is performed. If the verification passes, the update parameters of the signal interference strategy generation model are extracted. The signal shielding device uploads the update parameters to the server through the federation gateway. The server aggregates the update parameters uploaded by each signal shielding device to obtain global parameters. The global parameters are distributed to the signal shielding device through the federation gateway. The signal shielding device updates the local signal interference strategy generation model based on the received global parameters to perform federated optimization of the signal interference strategy generation model.

[0101] By recording multi-dimensional data (shielding parameters, real-time wireless communication signals, real-time signal interference strategies, real-time interference signals, environmental noise data, device temperature, shielding time, account, and shielding effect feedback) in real time, a complete device operation state image is formed, providing high-quality feature input for signal interference strategy generation model, and enhancing the accuracy of interference strategy.

[0102] By adopting a data volume threshold triggering mechanism, resource waste caused by frequent training is avoided; by model backup + version rollback design, system reliability is ensured, and service interruption caused by training failure is prevented; by intelligent scheduling strategy optimization of idle period training, the utilization rate of computing resources is optimized.

[0103] By replacing raw data transmission with parameter aggregation, data privacy protection requirements are met; by participating in global model optimization with distributed signal shielding devices, the generalization ability of the model to different environments (such as different regional electromagnetic environments) is improved; by edge-cloud collaborative architecture, continuous evolution of the model is realized, breaking through the limitations of insufficient data samples of a single device.

[0104] By real-time collection of multi-dimensional signal shielding logs to build incremental data sets, combined with a threshold-based triggering mechanism and an idle period scheduling strategy, low-power and high-reliability incremental training of local models is realized, and a federated learning architecture is innovatively introduced to realize collaborative evolution of the global model based on parameter aggregation of distributed devices under the premise of protecting data privacy. A closed-loop optimization system is formed, with dynamic environmental adaptability (such as noise self-adaptation and heat dissipation linkage), system stability guarantee (version rollback and verification mechanism), and multi-device collaborative generalization advantage, significantly improving the precision and scene universality of signal interference strategies.

[0105] By dynamic feature perception layer (integrating five-dimensional features of spectrum, time domain, space domain, protocol semantics, and signal strength) and multi-domain feature fusion layer (based on hypergraph neural network and dynamic weight distribution of reinforcement learning), precise signal feature extraction and cross-domain association are realized, combined with multi-objective decision module (double-branch network generates interference frequency points and power strategy) and beam optimization module (differentiable convex optimization improves directivity) to generate efficient interference strategies; by using generative adversarial networks to enhance data diversity, and through model pruning, quantization, and containerized deployment to realize lightweight adaptation, while using multiple encryption (AES-128, SM4, ChaCha20) and dynamic verification mechanism to ensure instruction security, relying on the federated learning framework to optimize the model in real time, with environmental noise adaptive compensation, temperature intelligent control, and log traceability, the system finally forms significant advantages in interference accuracy (frequency point hit rate), resource utilization (power dynamic optimization), real-time response (millisecond-level inference), anti-countermeasure capability (narrow beam direction), and compliance (frequency point constraints), realizing intelligent, low-power, and highly concealed directional signal shielding.

[0106] While the foregoing describes specific embodiments of the application, one of ordinary skill in the art will further appreciate that the specific exemplary embodiments described are meant to be illustrative only and are not intended to limit the scope of the application. Changes, modifications, and equivalents which would occur to one skilled in the art upon a reading of the foregoing description are meant to be encompassed within the scope of the application.

Claims

1. A signal shielding method based on deep reinforcement learning, characterized by: The steps include: Step S1: constructing a signal interference strategy generation model based on the dynamic feature perception layer, the multi-domain feature fusion layer, and the interference strategy generation layer, and setting a loss function of the signal interference strategy generation model; The dynamic feature perception layer is constructed based on a spectrum dynamic coding module, a time domain feature coding module, a spatial domain feature coding module, a protocol semantic coding module, and a signal strength coding module; the spectrum dynamic coding module, the time domain feature coding module, the spatial domain feature coding module, the protocol semantic coding module, and the signal strength coding module are respectively used to extract spectrum features, time features, spatial features, communication protocol features, and signal strength features from the input wireless communication signal; The multi-domain feature fusion layer is constructed based on a feature interaction enhancement module and a dynamic weight allocation module; the feature interaction enhancement module is used to construct a three-dimensional correlation matrix based on spectral features, temporal features, and spatial features; the dynamic weight allocation module is used to evaluate the contribution of the three-dimensional correlation matrix, communication protocol features, and signal strength features to fuse them to obtain fused features; The interference strategy generation layer is constructed based on a multi-objective decision module, a beam optimization module, and an output module; the multi-objective decision module is used to generate an interference frequency selection matrix and a power level vector based on fusion features; the beam optimization module is used to generate beamforming weights based on fusion features; the output module is used to generate a signal interference strategy based on the interference frequency selection matrix, the power level vector, and the beamforming weights; Step S2: obtaining a large number of historical wireless communication signals, preprocessing and annotating each of the historical wireless communication signals to construct a signal dataset, and expanding the sample size of the signal dataset by generating an adversarial network; Step S3: dividing the signal data set into a training set, a validation set, and a test set, and training, validating, and testing the signal interference strategy generation model using the training set, the validation set, and the test set; Step S4: compressing the signal interference strategy generation model that has passed the test, and deploying the compressed signal interference strategy generation model to a signal jammer; Step S5: After the signal blocker is started, it obtains an input signal blocking instruction, verifies and analyzes the signal blocking instruction to obtain a blocking parameter, collects a real-time wireless communication signal, pre-processes the real-time wireless communication signal, and then inputs the real-time wireless communication signal and the blocking parameter into a deployed signal interference strategy generation model to obtain a real-time signal interference strategy; Step S6: The signal blocker modulates the corresponding real-time interference signal based on the real-time signal interference strategy, and after the power amplifier performs a power amplification operation on the real-time interference signal through the digital predistortion technology, transmits the real-time interference signal to the outside through the radio frequency antenna to perform a signal directional shielding operation; Step S7: The signal blocker collects environmental noise data and device temperature in real time, performs power compensation on the real-time interference signal based on the environmental noise data, and dynamically adjusts the working states of the power amplifier and the fan based on the device temperature; Step S8: The signal blocker records the signal blocking log in real time, and performs federal tuning on the signal interference strategy generation model based on the signal blocking log.

2. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: In the step S1, the spectrum dynamic coding module is constructed based on a dynamic coding unit, a frequency domain correlation unit and an aggregation analysis unit; The dynamic encoding unit is used to extract local frequency domain features and time series dynamic features from the input wireless communication signal through a multi-scale convolution subunit and a gated loop subunit, and output a time-frequency joint feature including the local frequency domain features and the time series dynamic features; The frequency domain association unit is used to divide adjacent frequency points into similar clusters through the frequency point clustering subunit to capture the temporal dependency of frequency point migration from the time-frequency joint features, calculate the cross-frequency point attention weights on the time-frequency joint features through the multi-head self-attention subunit to model the global dependency between frequency points, and further capture the frequency domain correlation of frequency point migration, and generate the frequency point dynamic coding vector based on the temporal dependency and frequency domain correlation; The aggregation analysis unit is configured to perform multi-resolution matching of the carrier aggregation mode on the frequency dynamic coding vector through the pattern matching subunit, decouple the frequency migration feature from the frequency dynamic coding vector through the feature reconstruction subunit, and output the frequency migration feature and the spectrum feature of the carrier aggregation mode; The time domain feature encoding module is used to extract time features including signal burst characteristics and time-frequency interleaving patterns from the input wireless communication signal through a time domain convolutional network and a multi-head time attention mechanism; The spatial feature encoding module is used to extract spatial features including signal topology structure features and beam space modal vectors from the input wireless communication signal through a graph attention network and spatial spectrum estimation; The protocol semantic encoding module is used to extract communication protocol features from the input wireless communication signal through the BERT unit and the 3GPP standard knowledge graph; The signal strength encoding module is used to dynamically map the RSSI value carried in the input wireless communication signal to the MCS through an adaptive normalization layer, thereby extracting the signal strength feature; The feature interaction enhancement module is used to model cross-domain feature associations among spectral features, temporal features, and spatial features through a hypergraph neural network, thereby constructing a three-dimensional association matrix. The dynamic weight allocation module is used to evaluate the contribution of the three-dimensional association matrix, communication protocol features, and signal strength features to the current interference scenario through a value network based on a policy gradient method of reinforcement learning, so as to fuse them to obtain a fusion feature; The multi-objective decision module is used to reason about the fusion features through a dual-branch network structure to generate an interference frequency selection matrix and a power level vector; The beam optimization module infers the fused features through a differentiable convex optimization layer and an SDP relaxation method to generate beamforming weights.

3. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: The formula of the loss function is: L=α*L int +β*L res +γ*L beam +δ*L fusion ; Among them, L represents the loss value of the loss function; L int Indicates the loss of interference effect; L res Indicates resource consumption loss; L beam represents the beam optimization loss; L fusion represents the feature fusion loss; α, β, γ, and δ represent multi-objective trade-off coefficients, which are optimized by grid search; ; ; ; ; Wherein, f represents the frequency; represents the probability of frequency f being selected; represents the characteristic function; Indicates the target interference frequency set; 、 as well as Both represent the regularization strength coefficient; represents the power level vector prediction value; Indicates the theoretical optimal power level; Represents the square of the Euclidean distance of the vector; represents the beamforming weight matrix; represents the L1 norm; represents the Frobenius norm squared; Indicates the predicted direction of the beam main lobe; represents the ideal direction of the beam main lobe; represents a three-dimensional incidence matrix; Indicates communication protocol characteristics; Indicates signal strength characteristics; represents fusion features; Represents dynamic weight multiplicative fusion; Represents dynamic weighted additive fusion.

4. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: The step S2 is specifically as follows: A large number of historical wireless communication signals are obtained, and each of the historical wireless communication signals is preprocessed by at least signal amplification, filtering, signal down-conversion, and signal sampling. Each of the preprocessed historical wireless communication signals is labeled with at least frequency migration characteristics, carrier aggregation mode, signal burst characteristics, time-frequency interleaving mode, signal topology structure characteristics, beam space modal vector, communication protocol characteristics, signal strength characteristics, signal type, and signal interference strategy. A data set is constructed based on each of the labeled historical wireless communication signals, and the sample size of the signal data set is expanded by a generator in a pre-trained generative adversarial network.

5. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: The step S3 is specifically as follows: Based on the k-fold cross-validation method, the signal dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:

1. The signal interference strategy generation model is trained using the training set. During the training process, the hyperparameters including at least the learning rate, regularization parameter, network structure parameter, batch size, training rounds, activation function, optimizer, and random dropout rate are optimized using a Bayesian optimization framework until the loss value of the loss function is less than a preset loss threshold or a preset early stopping condition is met; The trained signal interference strategy generation model is verified through the verification set to determine whether the accuracy of the interference strategy is greater than the preset accuracy threshold. If not, the verification fails, and the training set is expanded to continue training; if so, the verification passes, and: the frequency hit rate, false interference rate, power utilization rate, power dynamic adjustment response time, beam pointing accuracy, strategy conflict rate and single inference time are calculated through the test set to test the verified signal interference strategy generation model. If the test fails, the training set is expanded to continue training; if the test passes, the training ends.

6. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: The step S4 is specifically as follows: All weight parameters of the signal interference strategy generation model that have passed the test are sorted according to the absolute value of the weight value, and the weight values ​​below the preset weight threshold are set to zero to perform a pruning operation on the signal interference strategy generation model. A performance verification is performed on the signal interference strategy generation model after the pruning operation. If the performance verification fails, the weight threshold is adjusted and the pruning operation is performed again; if the performance verification passes, then: Performing a quantization operation on the signal interference strategy generation model based on a preset quantization step size and an offset, performing a secondary performance verification on the quantized signal interference strategy generation model, and if the performance verification fails, adjusting the quantization step size and the offset and re-performing the quantization operation; If the performance verification is passed, the compression of the signal interference strategy generation model is completed; The compressed signal interference strategy generation model is deployed to a signal blocker through containerization technology, and data drift compensation training and model drift compensation training are performed on the signal interference strategy generation model deployed on the signal blocker.

7. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: The step S5 is specifically as follows: After the signal blocker is started, the signal blocker obtains an input signal block instruction carrying a parameter ciphertext and a verification ciphertext; the parameter ciphertext is obtained by encrypting the block parameter using a first encryption rule; The verification ciphertext is obtained by encrypting the account number, password, timestamp and hash value using the second encryption rule; the hash value is obtained by performing hash256 calculation on the account number, password and timestamp; the shielding parameters include at least shielding duration, shielding range and shielding signal type; The signal blocker parses the signal blocking instruction to obtain a parameter ciphertext and a verification ciphertext, decrypts the verification ciphertext using the second encryption rule to obtain an account number, a password, a timestamp, and a hash value, performs an integrity check on the account number, password, and timestamp using the hash value, performs an age check using the timestamp, and then performs a legitimacy check using the account number and password. If all checks pass, the parameter ciphertext is decrypted using the first encryption rule to obtain a blocking parameter; The signal jammer collects real-time wireless communication signals, pre-processes the real-time wireless communication signals by filtering, and then inputs the real-time wireless communication signals and the jamming parameters into the deployed signal interference strategy generation model to obtain a real-time signal interference strategy including an interference frequency selection matrix, a power level vector, and beamforming weights.

8. The signal shielding method based on deep reinforcement learning according to claim 7, characterized in that: The first encryption rule is specifically as follows: obtaining the first 6 digits of the current timestamp and concatenating them with a preset master key to obtain first concatenated data, calculating the MD5 value of the first concatenated data using the MD5 algorithm as the first temporary key; performing UTF-8 encoding on the masking parameter to obtain encoded data, and appending 4 random letters to the end of the encoded data to generate an extended data block; and encrypting the extended data block using the first temporary key using the AES-128 algorithm to obtain a parameter ciphertext. The second encryption rule is specifically: Obtain the last 6 digits of the current timestamp and concatenate them with the preset master key to obtain second concatenated data, calculate the MAC value of the second concatenated data through the HMAC algorithm as the second temporary key; call the second temporary key through the SM4 algorithm to encrypt the account number, password, timestamp and hash value to obtain first-level encrypted data, convert the first-level encrypted data into hexadecimal data, shift each character of the hexadecimal data right by 7 bits to obtain second-level encrypted data, swap the numbers 6 and 9 in the second-level encrypted data, swap the letters A and C, and swap the number 8 and B to obtain third-level encrypted data, and call the second temporary key through the ChaCha20 algorithm to encrypt the third-level encrypted data to obtain a verification ciphertext.

9. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: The step S7 is specifically as follows: The signal blocker presets a comparison table for storing the corresponding relationship between noise intensity intervals and power amplifier gains, and presets a fan control function with the independent variable being the device temperature and the dependent variable being the fan speed; The signal blocker collects environmental noise data in real time, matches the power amplifier gain from a comparison table based on the noise intensity of the environmental noise data, and controls the power amplifier based on the power amplifier gain to perform power compensation on the real-time interference signal; The signal blocker collects the device temperature in real time. When the device temperature is lower than a preset temperature threshold, the device temperature is input into a fan control function to obtain a corresponding fan speed, and the fan operating state is dynamically adjusted based on the fan speed. When the device temperature is greater than or equal to a temperature threshold, the fan speed is adjusted to a maximum, and the bias voltage of the power amplifier is continuously reduced based on a preset voltage gradient until the device temperature is less than the temperature threshold.

10. The signal shielding method based on deep reinforcement learning according to claim 1, characterized in that: The step S8 is specifically as follows: The real-time record of the signal jammer includes at least the signal jamming log of the jamming parameters, real-time wireless communication signals, real-time signal jamming strategies, real-time jamming signals, environmental noise data, device temperature, jamming time, account number, and jamming effect feedback; An incremental data set is constructed based on the signal shielding log. When the amount of data in the incremental data set exceeds a preset quantity threshold, the signal interference strategy generation model is backed up. During an idle period, the signal interference strategy generation model is trained using the incremental data set. The performance of the trained signal interference strategy generation model is verified. If the verification fails, a version rollback operation is performed. If the verification passes, updated parameters of the signal interference strategy generation model are extracted. The signal blocker uploads the updated parameters to the server through the federal gateway. The server aggregates the updated parameters uploaded by each signal blocker to obtain global parameters, and sends the global parameters to the signal blocker through the federal gateway. The signal blocker updates the local signal interference strategy generation model based on the received global parameters to perform federal tuning on the signal interference strategy generation model.

Citation Information

Patent Citations

  • Intelligent space frequency spectrum anti-interference method based on deep learning network

    CN117240384A

  • Method and system for blocking, shielding and adaptively identifying mobile phone signals

    CN119341679A

  • Consistency testing method and device for single-to-wire communication

    CN120128515A

  • Systems and methods of sensor data fusion

    US20250124112A1

Cited By

  • Dynamic assembly impedance control method and system based on deep reinforcement learning

    CN121411383A

  • Signal reconnaissance, positioning and interference method based on deep learning

    CN121664354A

  • Multi-domain cooperative anti-interference method and device, medium and equipment

    CN121984638A

  • Multi-target adaptive interference system based on array antenna

    CN122093012A

  • Wireless signal shielding method for network control

    CN122226201A