Adaptive generation method and system for UAV countermeasure signals in air-ground collaboration
Through the adaptive generation method of air-ground collaborative drone counter signal, the secondary encoding structure and genetic algorithm are used to optimize the interference strategy, and resource allocation is combined with game theory and reinforcement learning. The problems of adaptability and insufficient resource utilization in the existing drone counter technology are solved, and the efficient and low-power drone interference effect is achieved.
Patent Information
- Application Number
- CN202510686556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing drone countermeasures lack adaptability and are difficult to effectively deal with drones with different communication protocols. The air-to-ground coordination mechanism is insufficient, resulting in low interference efficiency, low resource utilization, and countermeasures are easily detected and have high power consumption.
Adaptive generation method of drone counter signal with air-ground collaboration is adopted. By obtaining the target drone communication signal, the interference strategy is optimized using the secondary coding structure and the genetic algorithm of the composite fitness function, and resource optimization allocation is combined with game theory model and reinforcement learning algorithm to generate adaptive interference signals.
It realizes efficient interference effects for different communication protocols, reduces power consumption, enhances the concealment and overall efficiency of the system, and can cope with complex and changeable drone invasion scenarios.
Smart Images

Figure CN120223235B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to UAV countermeasure technology, and in particular to a method and system for adaptively generating UAV countermeasure signals in air-ground collaboration. Background Art
[0002] To address the threat posed by illegal drones, drone countermeasures have become a research hotspot. Existing drone countermeasures suffer from numerous shortcomings. Traditional electromagnetic interference methods typically employ jamming signals with fixed frequencies and waveforms, lacking adaptability to different communication protocols. This results in low jamming efficiency and limited effectiveness against high-end drones with frequency hopping or encrypted communication capabilities. Existing countermeasures systems generally lack air-ground coordination mechanisms, relying either solely on ground-based equipment, resulting in limited coverage, or relying on a single platform for jamming, resulting in inefficient resource utilization and difficulty forming a comprehensive protection network. Existing countermeasures systems often neglect the combined considerations of power efficiency and anti-detection when generating jamming signals, resulting in high energy consumption and susceptibility to detection, particularly in scenarios requiring long-term continuous operation. Summary of the Invention
[0003] The embodiments of the present invention provide a method and system for adaptively generating countermeasure signals of unmanned aerial vehicles (UAVs) in air-ground collaboration, which can solve the problems in the prior art.
[0004] A first aspect of an embodiment of the present invention provides a method for adaptively generating countermeasure signals for air-ground coordinated UAVs, comprising:
[0005] Obtain the communication signal of the target UAV, extract the signal characteristic parameters and analyze them to obtain the communication protocol type;
[0006] A two-level coding structure is used to encode signal characteristic parameters, where the first level represents the interference strategy type and the second level represents specific parameters. The encoded parameters are optimized using a genetic algorithm with a composite fitness function, which is used to evaluate interference effect, power consumption efficiency, and anti-detection performance. A corresponding basic template is selected from a preset interference waveform template library based on the communication protocol type, and the optimized parameters are used for parameter matching and adjustment to generate an interference signal template.
[0007] Sending the interference signal template to the air countermeasure equipment and the ground countermeasure equipment;
[0008] Optimizing resource allocation for air-ground collaboration includes: constructing an air-ground collaboration strategy based on a game theory model; dynamically allocating resources based on the air-ground collaboration strategy using a reinforcement learning algorithm that optimizes power and computing resources in real time based on device status and target characteristics;
[0009] Based on the results of resource optimization allocation, control the air and ground countermeasure equipment to generate interference signals and transmit them to the target UAV.
[0010] In an optional embodiment,
[0011] A two-level coding structure is used to encode the signal characteristic parameters, where the first level represents the interference strategy type and the second level represents the specific parameters including:
[0012] The interference strategy type is first-level encoded using a 4-bit binary code to obtain a strategy code, wherein the interference strategy types include noise interference strategy, deception interference strategy, blocking interference strategy and combined interference strategy;
[0013] Specific parameters are subjected to secondary encoding using a variable-length coding structure to obtain parameter codes, wherein the specific parameters include frequency domain parameters, time domain parameters, and power parameters. The frequency domain parameters include interference bandwidth ratio, frequency offset, and power spectrum shape factor. The time domain parameters include interference duration, delay, and repetition period. The power parameters include interference power, power allocation coefficient, and shape factor. The coding bit width of the parameter code is optimized based on the importance of the parameters to obtain an optimized code, wherein the coding bit width is determined according to the parameter value range and quantization accuracy requirements.
[0014] The strategy code and the optimization code are adaptively mapped according to target features to obtain a final feature parameter code.
[0015] In an optional embodiment,
[0016] The encoded parameters are optimized using a genetic algorithm with a composite fitness function. The composite fitness function is used to evaluate the interference effect, power consumption efficiency and anti-detection performance, including:
[0017] Constructing a composite fitness function, the composite fitness function including: an interference effect evaluation item for evaluating the signal-to-interference ratio, bit error rate, and interference coverage; a power consumption efficiency evaluation item for evaluating power efficiency, time efficiency, and energy utilization, where the energy utilization is determined by the time-domain integral ratio of effective interference power to total interference power; and an anti-detection evaluation item for evaluating detection probability, exposure time, and low interception characteristics, where the low interception characteristics are determined by interference bandwidth ratio, power fluctuation, and transmission pattern characteristics;
[0018] The composite fitness function uses a dynamic weight coefficient to perform a weighted combination of the evaluation items, and the dynamic weight coefficient calculates a weight adjustment amount according to the deviation between the target performance and the actual performance, the target performance including the expected interference effect, power consumption efficiency and anti-detection index, and the actual performance including the interference effect, power consumption efficiency and anti-detection index measured in real time;
[0019] The strategy coding and optimization coding are iteratively optimized using a genetic algorithm to obtain an optimal solution, and the optimized interference strategy and parameter configuration are output; the genetic algorithm adopts an adaptive mutation rate, and the adaptive mutation rate decreases as the fitness value increases.
[0020] In an optional embodiment,
[0021] The determination of dynamic weight coefficient includes:
[0022] Establish a performance indicator system, including signal-to-interference ratio, bit error rate, and interference coverage in the interference effect indicator set; power efficiency, time efficiency, and energy utilization in the power consumption efficiency indicator set; detection probability, exposure time, and low interception characteristics in the anti-detection indicator set, and determine the ideal target value and acceptable range for each indicator;
[0023] Use normalized deviation to calculate function δ ij =(T ij -A ij ) / N ij Calculate the performance deviation of each indicator, where T ij is the target performance value, A ij is the actual measured value, N ij is the normalization factor; the weighted sum of the deviations of similar indicators is used to obtain the comprehensive deviation Δ i ;
[0024] The weight adjustment is calculated using a nonlinear bias response function:
[0025]
[0026] Where ΔW i is the adjustment amount of the weight corresponding to the i-th indicator, α is the adjustment amplitude coefficient, β is the nonlinear index used to control the curvature of the response function, γ is the sensitivity coefficient used to control the response degree to small deviations, τ min is the minimum response threshold for filtering out noise, τ max is the maximum response threshold used to limit abnormal situations; sign(Δ i ) is the sign function of the bias, which determines the direction of weight adjustment;
[0027] The weights are updated smoothly and normalized to ensure that the sum of the weights is 1 and the weight values are subject to upper and lower limits.
[0028] In an optional embodiment,
[0029] Building an air-ground collaboration strategy based on the game theory model includes:
[0030] Obtaining state vectors of air countermeasure equipment and ground countermeasure equipment, and constructing an air-ground countermeasure game theory model; wherein the state vectors include power, position, coverage, and energy state;
[0031] Taking the air countermeasure equipment and the ground countermeasure equipment as game participants, constructing a game strategy space including power selection, interference direction and resource allocation according to the state vector, wherein the game strategy space is used to constrain the strategy selection range of the game participants;
[0032] Constructing a multi-objective utility function, the multi-objective utility function including an interference effect evaluation function obtained by weighted combination of signal-to-interference ratio evaluation, coverage effect evaluation, and directionality evaluation, and a resource consumption evaluation function obtained by weighted combination of power consumption, communication overhead, and maneuvering cost; combining the multi-objective utility function with a resource penalty factor to construct a game payoff matrix; and solving the optimal response strategy under Nash equilibrium based on the game payoff matrix;
[0033] Calculating the air-ground coordination advantage based on the optimal response strategy and constructing a strategy selection probability distribution, wherein the air-ground coordination advantage is determined by the ratio of the combined countermeasure benefit to the sum of the individual countermeasure benefits;
[0034] A state value function is established according to the strategy selection probability distribution. The state value function calculates the long-term expected return of the game payment matrix based on a discount factor. The air-ground collaborative strategy is obtained according to the gradient of the state value function to the strategy selection probability distribution.
[0035] In an optional embodiment,
[0036] The reinforcement learning algorithm optimizes the allocation of power and computing resources in real time based on device status and target characteristics, including:
[0037] Obtaining device status and target features, and constructing a resource status tensor including platform type and resource type, wherein the platform type includes air platform and ground platform, and the resource type includes power resource and computing resource;
[0038] Constructing a variational autoencoder to encode the resource state tensor to generate latent variables, and decoding the latent variables into communication messages between air-ground platforms;
[0039] Based on the resource state tensor and the communication messages between the air-ground platform, a dual-path deep Q network with a value stream and an advantage stream is designed. The value stream evaluates the state value of resource allocation, and the advantage stream evaluates the action advantage of resource allocation. The state value and action advantage are combined to obtain the Q value of the resource allocation strategy. A proximal policy optimization algorithm is used to train the policy network. The policy network includes an actuator network and a critic network. The actuator network outputs the mean and standard deviation of the resource allocation action distribution that meets the air-ground collaboration requirements according to the constraints of the air-ground collaboration strategy. The critic network estimates the state value and calculates the action advantage based on the current state.
[0040] A hierarchical priority experience replay pool is constructed, sample priorities are calculated based on temporal difference error, and experience samples are sampled according to the sample priorities. A meta-learning framework is used to update the policy, and an air-ground consistency evaluation function is constructed to impose additional constraints on the training of the policy network. Based on the trained policy network, the power and computing resources of the air and ground platforms are optimized in real time.
[0041] In an optional embodiment,
[0042] The air-ground consistency evaluation functions include:
[0043] The air-ground consistency evaluation function includes a resource balance evaluation item, a task completion evaluation item and a collaborative effectiveness evaluation item. The resource balance evaluation item calculates the difference in resource allocation between the air platform and the ground platform. The task completion evaluation item evaluates the task execution effect based on the platform resource utilization. The collaborative effectiveness evaluation item is evaluated based on the communication quality and collaborative response time between the air-ground platforms. The conformity of the resource allocation action with the air-ground collaborative strategy is calculated based on the weighted combination of the resource balance evaluation item, the task completion evaluation item and the collaborative effectiveness evaluation item.
[0044] A second aspect of an embodiment of the present invention provides an adaptive generation system for UAV countermeasure signals for air-ground collaboration, comprising:
[0045] The first unit is used to obtain the communication signal of the target UAV, extract the signal characteristic parameters and analyze them to obtain the communication protocol type;
[0046] The second unit is configured to encode signal characteristic parameters using a two-level encoding structure, wherein the first level represents the interference strategy type and the second level represents specific parameters; optimize the encoded parameters using a genetic algorithm with a composite fitness function, wherein the composite fitness function is used to evaluate interference effect, power consumption efficiency, and anti-detection performance; select a corresponding basic template from a preset interference waveform template library according to the communication protocol type, and use the optimized parameters to perform parameter matching and adjustment to generate an interference signal template;
[0047] The third unit is used to send the interference signal template to the air countermeasure device and the ground countermeasure device;
[0048] The fourth unit is used to optimize resource allocation for air-ground collaboration, including: constructing an air-ground collaboration strategy based on a game theory model; dynamically allocating resources according to the air-ground collaboration strategy using a reinforcement learning algorithm, wherein the reinforcement learning algorithm optimizes the allocation of power and computing resources in real time based on device status and target characteristics;
[0049] The fifth unit is used to control the air and ground countermeasure equipment to generate interference signals and transmit them to the target UAV based on the resource optimization allocation results.
[0050] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:
[0051] processor;
[0052] a memory for storing processor-executable instructions;
[0053] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0054] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0055] The present invention realizes accurate identification and adaptive interference of the communication signal characteristics of the target UAV. Through the optimization of the two-level coding structure and the composite fitness function genetic algorithm, it can achieve high-efficiency interference effects for different communication protocols while maintaining low power consumption and strong concealment.
[0056] This invention adopts an air-ground collaborative strategy and a resource optimization allocation mechanism based on game theory, and realizes the dynamic allocation of power and computing resources between air and ground countermeasure equipment through a reinforcement learning algorithm, which significantly improves the overall efficiency and adaptability of the countermeasure system and can cope with complex and changeable drone intrusion scenarios.
[0057] This invention combines signal processing, artificial intelligence and resource optimization technologies to form a complete set of drone countermeasure technology solutions, which not only improves the accuracy and success rate of countermeasures, but also reduces energy consumption, enhances the concealment and reliability of the system, and is of great value to improving the air defense security of important areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Schematic diagram of the flow of the method for adaptively generating countermeasure signals of UAVs in air-ground collaboration according to an embodiment of the present invention;
[0059] Figure 2 A comparison chart of information transmission accuracy under different communication loads for different communication schemes;
[0060] Figure 3 This is a comparison chart of optimized performance. DETAILED DESCRIPTION
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0062] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0063] Figure 1 FIG. 1 is a flow chart of a method for adaptively generating countermeasure signals of a UAV in air-ground collaboration according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0064] Obtain the communication signal of the target UAV, extract the signal characteristic parameters and analyze them to obtain the communication protocol type;
[0065] A two-level coding structure is used to encode signal characteristic parameters, wherein the first level represents the interference strategy type and the second level represents specific parameters; the encoded parameters are optimized using a genetic algorithm with a composite fitness function, which is used to evaluate interference effect, power consumption efficiency, and anti-detection performance; a corresponding basic template is selected from an interference waveform template library based on the communication protocol type, and the optimized parameters are used for parameter matching and adjustment to generate an interference signal template;
[0066] Sending the interference signal template to the air countermeasure equipment and the ground countermeasure equipment;
[0067] Optimizing resource allocation for air-ground collaboration includes: constructing an air-ground collaboration strategy based on a game theory model; dynamically allocating resources based on the air-ground collaboration strategy using a reinforcement learning algorithm that optimizes power and computing resources in real time based on device status and target characteristics;
[0068] Based on the results of resource optimization allocation, control the air and ground countermeasure equipment to generate interference signals and transmit them to the target UAV.
[0069] In an optional embodiment, a two-level coding structure is used to encode the signal characteristic parameters, wherein the first level represents the interference strategy type, and the second level represents specific parameters including:
[0070] The interference strategy type is first-level encoded using a 4-bit binary code to obtain a strategy code, wherein the interference strategy types include noise interference strategy, deception interference strategy, blocking interference strategy and combined interference strategy;
[0071] Specific parameters are subjected to secondary encoding using a variable-length coding structure to obtain parameter codes, wherein the specific parameters include frequency domain parameters, time domain parameters, and power parameters. The frequency domain parameters include interference bandwidth ratio, frequency offset, and power spectrum shape factor. The time domain parameters include interference duration, delay, and repetition period. The power parameters include interference power, power allocation coefficient, and shape factor. The coding bit width of the parameter code is optimized based on the importance of the parameters to obtain an optimized code, wherein the coding bit width is determined according to the parameter value range and quantization accuracy requirements.
[0072] For example, the interference strategy type is encoded using a 4-bit binary code. The interference strategy types specifically include noise interference strategy, deception interference strategy, blocking interference strategy, and combined interference strategy. The encoding of each type is as follows:
[0073] Noise interference strategy: 0001;
[0074] Deception jamming strategy: 0010;
[0075] Blocking interference strategy: 0100;
[0076] Combined interference strategy: 1000;
[0077] When a combined jamming strategy is required, such as a combined jamming strategy using both noise jamming and deception jamming, the code is: 1001.
[0078] The specific parameters are coded in two levels using a variable length coding structure to form parameter coding. The specific parameters include frequency domain parameters, time domain parameters, and power parameters.
[0079] For each parameter, its encoding bit width determines the numerical range that needs to be represented according to the parameter's value range, then determines the minimum resolution according to the required quantization accuracy requirements, and finally calculates the minimum number of binary bits required to meet these two conditions.
[0080] Frequency domain parameters include interference bandwidth ratio, frequency offset, and power spectrum shape factor. The interference bandwidth ratio represents the ratio of the interference signal bandwidth to the target signal bandwidth. The value range is 0.1 to 10, the quantization accuracy is 0.1, and it requires 7 binary bits to represent (can represent the range of 0 to 12.7). The frequency offset represents the offset between the center frequency of the interference signal and the center frequency of the target signal. The value range is -50MHz to +50MHz, the quantization accuracy is 1MHz, and it requires 7 binary bits to represent (can represent the range of -64MHz to +63MHz). The power spectrum shape factor is used to describe the spectral distribution characteristics of the interference signal. The value range is 0 to 1, the quantization accuracy is 0.01, and it requires 7 binary bits to represent (can represent the range of 0 to 1.27).
[0081] Time domain parameters include interference duration, delay, and repetition period. Interference duration indicates the duration of a single interference signal, with a value range of 1ms to 1000ms, a quantization accuracy of 1ms, and a 10-bit binary representation (which can represent a range of 0 to 1023ms). Delay indicates the delay time of the interference signal relative to the target signal, with a value range of 0 to 100ms, a quantization accuracy of 0.2ms, and a 9-bit binary representation (which can represent a range of 0 to 102.4ms). Repetition period indicates the time interval between repeated occurrences of the interference signal, with a value range of 10ms to 10000ms, a quantization accuracy of 10ms, and a 10-bit binary representation (which can represent a range of 0 to 10230ms).
[0082] Power parameters include interference power, power allocation coefficient, and waveform factor. Interference power represents the transmit power of the interference signal, with a value range of 0 to 50 dBm, a quantization accuracy of 0.2 dBm, and requires 8 binary bits to represent (can represent a range of 0 to 51 dBm). The power allocation coefficient represents the power allocation ratio between channels in the event of multi-channel interference, with a value range of 0 to 1, a quantization accuracy of 0.01, and requires 7 binary bits to represent (can represent a range of 0 to 1.27). The waveform factor is used to describe the waveform characteristics of the interference signal, with a value range of 0.5 to 5, a quantization accuracy of 0.1, and requires 6 binary bits to represent (can represent a range of 0 to 6.3).
[0083] The encoding bit width of the parameter encoding is optimized based on the parameter importance. The parameters are ranked from high to low in the following order: interference power, interference bandwidth ratio, frequency offset, interference duration, delay, power spectrum shape factor, repetition period, power allocation coefficient, and waveform factor. The original bit width is maintained for the parameters with high importance, while the bit width of the parameters with low importance is appropriately reduced. The adjusted bit width is as follows:
[0084] Interference power: maintained at 8 bits; Interference bandwidth ratio: maintained at 7 bits; Frequency offset: maintained at 7 bits; Interference duration: adjusted to 8 bits (accuracy adjusted to approximately 4ms); Delay: adjusted to 8 bits (accuracy adjusted to approximately 0.4ms); Power spectrum shape factor: adjusted to 6 bits (accuracy adjusted to approximately 0.02); Repetition period: adjusted to 8 bits (accuracy adjusted to approximately 40ms); Power allocation coefficient: adjusted to 5 bits (accuracy adjusted to approximately 0.03); Shape factor: adjusted to 5 bits (accuracy adjusted to approximately 0.15);
[0085] Through the above adjustments, the total bit width of parameter encoding is reduced from 71 bits to 62 bits, improving the encoding efficiency.
[0086] Adaptively mapping the strategy code and optimization code based on the target characteristics yields the final feature parameter code. The adaptive mapping process selects an appropriate parameter subset based on the target characteristics and adaptively adjusts the interference requirements in different scenarios. For example, for low-altitude, slow-moving targets, the parameter subsets selected for adaptive mapping include: interference power, interference bandwidth ratio, frequency offset, and interference duration. The corresponding parameter coding bit widths are 8 bits, 7 bits, 7 bits, and 8 bits, respectively, for a total of 30 bits. The strategy code remains unchanged at 4 bits, resulting in a final feature parameter code of 34 bits. For high-speed, maneuvering targets, the parameter subsets selected for adaptive mapping include: interference power, interference bandwidth ratio, frequency offset, interference duration, delay, and repetition period. The corresponding parameter coding bit widths are 8 bits, 7 bits, 7 bits, 8 bits, 8 bits, and 8 bits, respectively, for a total of 46 bits. The strategy code remains unchanged at 4 bits, resulting in a final feature parameter code of 50 bits. For targets in complex electromagnetic environments, adaptive mapping selects a subset of parameters, including interference power, interference bandwidth ratio, frequency offset, interference duration, delay, power spectrum shape factor, repetition period, power allocation coefficient, and waveform factor—all 62 bits in total. The strategy code remains unchanged at 4 bits, resulting in a 66-bit encoding of the final characteristic parameters.
[0087] For example, for an X-band radar target, a noise jamming strategy is required with an interference power of 30 dBm, an interference bandwidth ratio of 2.5, a frequency offset of -15 MHz, an interference duration of 200 ms, a delay of 20 ms, a power spectrum shape factor of 0.8, a repetition period of 500 ms, a power allocation coefficient of 0.6, and a waveform factor of 2.0. According to the above parameter values, the strategy is coded as "0001", the parameter is coded as "0011110000110011110001001100100011001001001100101001001101101", and after adaptive mapping, the final feature parameter is coded as "0001001111000011001 111000100110010001100100100110010101101".
[0088] The two-level coding structure and adaptive mapping method proposed in the present invention achieve efficient digital representation of interference strategies and parameters. By combining the first-level coding of strategy types with the second-level coding of variable-length parameters, the problems of limited expressive power, coding redundancy and poor adaptability of traditional coding methods are solved. The coding bit width is innovatively optimized based on parameter importance, which reduces the overall coding length while ensuring the accuracy of key parameters, effectively improving coding efficiency. The adaptive mapping mechanism can dynamically select appropriate parameter subsets according to different target characteristics, making the coding structure flexible and adaptable to diverse interference scenarios.
[0089] In an optional embodiment, the encoded parameters are optimized using a genetic algorithm with a composite fitness function, where the composite fitness function is used to evaluate the interference effect, power consumption efficiency, and anti-detection performance, including:
[0090] Constructing a composite fitness function, the composite fitness function including: an interference effect evaluation item for evaluating the signal-to-interference ratio, bit error rate, and interference coverage; a power consumption efficiency evaluation item for evaluating power efficiency, time efficiency, and energy utilization, where the energy utilization is determined by the time-domain integral ratio of effective interference power to total interference power; and an anti-detection evaluation item for evaluating detection probability, exposure time, and low interception characteristics, where the low interception characteristics are determined by interference bandwidth ratio, power fluctuation, and transmission pattern characteristics;
[0091] The composite fitness function uses a dynamic weight coefficient to perform a weighted combination of the evaluation items, and the dynamic weight coefficient calculates a weight adjustment amount according to the deviation between the target performance and the actual performance, the target performance including the expected interference effect, power consumption efficiency and anti-detection index, and the actual performance including the interference effect, power consumption efficiency and anti-detection index measured in real time;
[0092] The strategy coding and optimization coding are iteratively optimized using a genetic algorithm to obtain an optimal solution, and the optimized interference strategy and parameter configuration are output; the genetic algorithm adopts an adaptive mutation rate, and the adaptive mutation rate decreases as the fitness value increases.
[0093] For example, in the interference effect evaluation item, the signal-to-interference ratio evaluation is calculated by the power ratio of the interference signal to the useful signal at the target receiver. For example, the interference power measured at the target receiver (such as -45dBm) is divided by the useful signal power (such as -75dBm), and the signal-to-interference ratio is 30dB. The bit error rate evaluation is obtained by monitoring the bit error performance of the target communication system. For example, under the influence of interference, the bit error rate of the target system decreases from 10 -5 Increased to 10 -2The interference coverage is verified through actual measurements, recording the area where effective interference can be achieved at different distances (e.g., 100m, 200m, 500m).
[0094] In the power efficiency evaluation, power efficiency is calculated as the ratio of effective interference power to input power. For example, if the input power is 100W and the measured effective interference power is 85W, the power efficiency is 85%. Time efficiency is determined by the ratio of the interference duration to the total mission time. For example, if the effective interference time is 1.8 hours in a 2-hour mission, the time efficiency is 90%. Energy utilization is calculated by integrating the effective interference power in the time domain as the ratio of the total interference power. In practice, the system samples the interference power every 10 milliseconds and records the interference power data for 24 hours. Assuming a total transmit power of 50W and an accumulated effective interference power of 42W / hour, the energy utilization is 84%.
[0095] In the anti-detection evaluation item, the detection probability is obtained through Monte Carlo simulation. In 1,000 simulations, the enemy detection equipment successfully detected the jammer 150 times, and the detection probability is 15%. The exposure time refers to the average time the jammer is detected by the enemy equipment. The test results show that the average exposure time is 3.5 seconds. The low interception characteristic is determined by three sub-indicators: interference bandwidth ratio (the ratio of the interference bandwidth to the target system bandwidth. For example, if the interference bandwidth is 5MHz and the target system bandwidth is 20MHz, the interference bandwidth ratio is 0.25); power fluctuation (the ratio of the standard deviation of the interference power to the average power. For example, if the standard deviation is 2W and the average power is 40W, the power fluctuation is 5%); the emission pattern characteristics include: main lobe direction (the direction of the maximum radiation power of the antenna); main lobe width (the angle when the radiation power is reduced to half of the maximum value (-3dB)). The following parameters are used for the evaluation of low-acquisition (LOA) characteristics: range (e.g., main lobe width of 30° (azimuth) × 20° (elevation)); sidelobe level (the ratio of the power in the secondary radiation direction to the maximum power in the main lobe, usually expressed in dB, such as a maximum sidelobe level of -25dB); front-to-back ratio (the ratio of the power in the main lobe direction to the power in the opposite direction, such as a front-to-back ratio of 18dB); null depth (the ratio of the minimum power point in the radiation pattern to the maximum power, such as a null depth of -40dB); and directivity (a dimensionless coefficient that characterizes the antenna's ability to concentrate radiation, such as a directivity of 12dB). In evaluating LOA characteristics, an ideal emission pattern should have a narrow mainlobe, low sidelobes, and a deep null. For example, a narrow-beam directional antenna (mainlobe width of 10°, sidelobe level of -30dB) has better LOA characteristics than a wide-beam antenna (mainlobe width of 60°, sidelobe level of -15dB).
[0096] The composite fitness function uses dynamic weight coefficients to perform a weighted combination of the evaluation items. The initial values of the weight coefficients are set to 0.5 for interference effect, 0.3 for power efficiency, and 0.2 for anti-detection. The system calculates the deviation between the target performance and the actual performance every 1 minute and adjusts the weight coefficients. For example, when the actual interference effect (Signal-to-Interference Ratio 25dB) is lower than the target value (Signal-to-Interference Ratio 30dB), the system increases the interference effect weight from 0.5 to 0.6 and reduces the other weights accordingly (power efficiency weight is reduced to 0.25, and anti-detection weight is reduced to 0.15). The amount of weight adjustment is proportional to the performance deviation, and the maximum adjustment range is 20% of the initial weight.
[0097] The genetic algorithm optimization process is as follows:
[0098] The first step is to initialize the population. Fifty individuals are randomly generated, each containing an encoded jamming strategy and parameters. The jamming strategy encoding includes the jamming method (e.g., 0 for noise jamming, 1 for drag jamming, 2 for repetitive jamming), the time pattern (e.g., 0 for continuous, 1 for periodic, 2 for random), and the frequency pattern (e.g., 0 for fixed frequency, 1 for frequency hopping, 2 for frequency sweeping). The encoding parameters include the jamming power (range: 20-100W, accuracy: 0.1W), the jamming frequency (range: 2-6GHz, accuracy: 1MHz), and the frequency hopping rate (range: 100-1000 hops / second, accuracy: 10 hops / second).
[0099] The second step is to evaluate the fitness. The composite fitness value is calculated for each individual. For example, if individual A has an interference effect score of 0.85, a power efficiency score of 0.75, and an anti-detection score of 0.60, and the current weight coefficients (0.6, 0.25, 0.15) are used, the composite fitness value is 0.78.
[0100] The third step is the selection operation. Using the roulette wheel method to select individuals, individuals with higher fitness are more likely to be selected. For example, the individual with the highest fitness (0.88) has an 8.8% probability of being selected, while the individual with the lowest fitness (0.35) has a only 3.5% probability of being selected.
[0101] Step 4: Crossover. Two parent individuals are randomly selected and crossover is performed with an 80% probability. The crossover point is randomly chosen. For example, if parent A's encoding is "101|10110" and parent B's encoding is "010|01001", if the crossover point is after the third position, the resulting offspring are "10101001" and "01010110".
[0102] The fifth step is the mutation operation. Each gene position of the individual is mutated with a probability of mutation rate. An adaptive mutation rate is used, with an initial mutation rate of 5%. When the population reaches a maximum fitness of 0.8, the mutation rate is reduced to 3%. When the fitness reaches 0.9, the mutation rate is further reduced to 1% to ensure convergence stability.
[0103] Step 6: Iterative optimization. Repeat steps 2 to 5 until the termination condition is met: the preset number of iterations reaches 200, or the optimal fitness improves by less than 0.1% for 20 consecutive generations.
[0104] Finally, the optimal solution is output. The individual with the highest fitness is selected from the final population and decoded to obtain the optimized interference strategy and parameter configuration. Specifically, the interference strategy is "repetitive interference (strategy code: 2) + periodic time pattern (code: 1) + frequency hopping pattern (code: 1)." Key parameters include an interference power of 75.3W, a center frequency of 3.5GHz, a frequency coverage range of 3.2-3.8GHz, a frequency hopping rate of 800 hops / second, a duty cycle of 65%, and a radiation pattern with a main lobe width of 15° and a sidelobe level of -28dB.
[0105] The interference waveform template library contains basic interference templates designed for various mainstream communication protocols, such as WiFi, Bluetooth, 4G / 5G, and satellite communication. The signal recognition module first identifies the target communication protocol type, such as identifying a WiFi signal targeting the IEEE 802.11ac protocol. The basic WiFi interference template is then retrieved from the template library. This template pre-sets the basic interference parameters for WiFi signals, including spectral characteristics (center frequency in the 5.2GHz band, 80MHz bandwidth), time domain characteristics (periodic frame structure, 1ms frame length), and modulation characteristics (quadrature amplitude modulation format, symbol rate). The parameters optimized by the genetic algorithm are then matched and adjusted with the basic template. For example, the optimized interference power of 75.3W is applied to the template; the frequency hopping pattern parameters are combined with the base template to enable the interference signal to hop at a rate of 800 hops / second within the 5.15-5.35 GHz band; the optimized 65% duty cycle parameters are applied to the time domain structure to set the duration of the interference signal and the ratio of quiet time; and the pattern configuration (main lobe width 15°, side lobe level -28dB) is applied to the antenna parameters to achieve directional interference. The system further optimizes the template based on the evaluation results of the composite fitness function. When it detects that the WiFi signal uses a dynamic frequency selection mechanism, it automatically adjusts the interference bandwidth ratio and power allocation of the interference template to enhance the interference strength on critical channels. When the target system uses forward error correction coding, the system adjusts the temporal characteristics of the interference signal accordingly, causing the interference duration to exceed the code interleaving period, thereby reducing error correction capabilities. The resulting interference signal template contains not only waveform parameters but also transmission control parameters, forming a complete interference strategy. The system saves this template and transmits it to the signal generation module, which is used to generate optimized interference signals for the target communication system in real time to achieve efficient and accurate interference effects.
[0106] The present invention realizes the intelligent configuration of interference system parameters by constructing a composite fitness function that integrates interference effect, power consumption efficiency and anti-detection performance, and combining it with a dynamic weight adjustment mechanism and genetic algorithm optimization; it overcomes the limitations of single indicator optimization and manual parameter adjustment in traditional interference system parameter configuration, and can find the optimal configuration scheme under multi-objective constraints; the dynamic weight coefficient mechanism can adaptively adjust the evaluation focus according to the deviation between actual performance and target performance, and the adaptive mutation rate design effectively balances the global search capability and local convergence speed of the algorithm, significantly improving the overall effectiveness and environmental adaptability of the interference system.
[0107] In an optional implementation, determining the dynamic weight coefficient includes:
[0108] Establish a performance indicator system, including signal-to-interference ratio, bit error rate, and interference coverage in the interference effect indicator set; power efficiency, time efficiency, and energy utilization in the power consumption efficiency indicator set; detection probability, exposure time, and low interception characteristics in the anti-detection indicator set, and determine the ideal target value and acceptable range for each indicator;
[0109] Use normalized deviation to calculate function δ ij =(T ij -A ij ) / N ij Calculate the performance deviation of each indicator, where T ij is the target performance value, A ij is the actual measured value, N ij is the normalization factor; the weighted sum of the deviations of similar indicators is used to obtain the comprehensive deviation Δ i ;
[0110] The weight adjustment is calculated using a nonlinear bias response function:
[0111]
[0112] Where ΔW i is the adjustment amount of the weight corresponding to the i-th indicator, α is the adjustment amplitude coefficient, β is the nonlinear index used to control the curvature of the response function, γ is the sensitivity coefficient used to control the response degree to small deviations, τ min is the minimum response threshold for filtering out noise, τ max is the maximum response threshold used to limit abnormal situations; sign(Δ i ) is the sign function of the bias, which determines the direction of weight adjustment;
[0113] The weights are updated smoothly and normalized to ensure that the sum of the weights is 1 and the weight values are subject to upper and lower limits.
[0114] For example, a complete performance indicator system is established, including three categories of indicators: interference effect indicator set, power consumption efficiency indicator set, and anti-detection indicator set. The interference effect indicator set includes signal-to-interference ratio, bit error rate, and interference coverage; the power consumption efficiency indicator set includes power efficiency, time efficiency, and energy utilization; and the anti-detection indicator set includes detection probability, exposure time, and low interception characteristics.
[0115] For each specific indicator, set an ideal target value and acceptable range. For example, for the signal-to-interference ratio (SIR), the ideal target value is set at 6dB, with an acceptable range of 3dB to 10dB; for the bit error rate (BER), the ideal target value is set at 0.25, with an acceptable range of 0.15 to 0.35; for the interference coverage, the ideal value is set at 95%, with an acceptable range of 85% to 100%. For power efficiency, the ideal value is set at 0.8, with an acceptable range of 0.6 to 0.9; for time efficiency, the ideal value is set at 0.75, with an acceptable range of 0.6 to 0.85; for energy utilization, the ideal value is set at 0.7, with an acceptable range of 0.5 to 0.8; for detection probability, the ideal value is set at 0.15, with an acceptable range of 0.1 to 0.25; for exposure time, the ideal value is set at 30 seconds, with an acceptable range of 20 to 60 seconds; and for low acquisition characteristics, the ideal value is set at 0.85, with an acceptable range of 0.7 to 0.95.
[0116] The performance deviation of each indicator is calculated using a normalized deviation calculation function. This function normalizes the difference between the target performance value and the actual measured value using a normalization factor. The normalized deviations of similar indicators are weighted and summed to obtain the combined deviation of each indicator. Assuming the internal weights of the interference effect indicator set are 0.4 for signal-to-interference ratio, 0.35 for bit error rate, and 0.25 for interference coverage, if the normalized deviations of these three indicators are 0.57, -0.5, and 0.2, respectively, the combined deviation of the interference effect indicator is 0.4 × 0.57 + 0.35 × (-0.5) + 0.25 × 0.2 = 0.108. Similarly, the combined deviations of the power efficiency and anti-detection indicators are calculated.
[0117] The weight adjustment is calculated using a nonlinear deviation response function, which contains multiple parameters: the adjustment amplitude coefficient α is usually set to 0.15 to control the maximum amplitude of each adjustment; the nonlinear exponent β is set to 2.0 to control the curvature of the response function; the sensitivity coefficient γ is set to 0.5 to control the degree of response to small deviations; the minimum response threshold τ is set to 0.5. min Set to 0.05 to filter out noise; the maximum response threshold τ max Set to 0.85 to limit the impact of abnormal conditions. sign(Δ i ) is used to extract the sign of a real number. For any non-zero real number, the sign function returns the sign of the number; for zero, it returns zero. In the weight adjustment formula, sign(Δ i ) is used to determine the direction of weight adjustment: when the comprehensive deviation Δ i When it is a positive value (indicating that the actual performance is lower than the target performance), sign(Δ i ) returns +1, the weight will increase; when the comprehensive deviation Δ i When it is a negative value (indicating that the actual performance exceeds the target performance), sign(Δ i ) returns -1, the weight will decrease; when the comprehensive deviation Δi When it is equal to zero (indicating that the actual performance is exactly equal to the target performance), sign(Δ i ) returns 0 and the weight remains unchanged.
[0118] Finally, the weights are updated and normalized smoothly. The updated weights are checked to see if they exceed the preset upper and lower limits. If so, they are adjusted to the boundary values. For example, if the upper limit of the interference effect indicator weight is 0.45 and the lower limit is 0.25, then 0.395 is within a reasonable range and does not require adjustment.
[0119] During actual operation, the above weight update process is performed once every predetermined time period (such as 5 minutes), so that the system can dynamically adjust the weight of each indicator according to the performance deviation, thereby optimizing the overall performance of the system under different working environments.
[0120] The present invention realizes the precision and adaptability of jamming system performance evaluation by establishing a comprehensive performance indicator system and an intelligent dynamic weight adjustment mechanism; adopts normalized deviation calculation and nonlinear response function to enable weight adjustment to accurately reflect the difference between actual performance and target performance; this method overcomes the limitations of traditional fixed weight evaluation methods, and can automatically optimize the weight ratio of each indicator according to changes in the working environment, while ensuring the jamming effect and taking into account power consumption efficiency and anti-detection performance, significantly improving the adaptability, work efficiency and overall combat effectiveness of the jamming system in complex electromagnetic environments.
[0121] In an optional implementation, constructing an air-ground collaboration strategy based on a game theory model includes:
[0122] Obtaining state vectors of air countermeasure equipment and ground countermeasure equipment, and constructing an air-ground countermeasure game theory model; wherein the state vectors include power, position, coverage, and energy state;
[0123] Taking the air countermeasure equipment and the ground countermeasure equipment as game participants, constructing a game strategy space including power selection, interference direction and resource allocation according to the state vector, wherein the game strategy space is used to constrain the strategy selection range of the game participants;
[0124] Constructing a multi-objective utility function, the multi-objective utility function including an interference effect evaluation function obtained by weighted combination of signal-to-interference ratio evaluation, coverage effect evaluation, and directionality evaluation, and a resource consumption evaluation function obtained by weighted combination of power consumption, communication overhead, and maneuvering cost; combining the multi-objective utility function with a resource penalty factor to construct a game payoff matrix; and solving the optimal response strategy under Nash equilibrium based on the game payoff matrix;
[0125] Calculating the air-ground coordination advantage based on the optimal response strategy and constructing a strategy selection probability distribution, wherein the air-ground coordination advantage is determined by the ratio of the combined countermeasure benefit to the sum of the individual countermeasure benefits;
[0126] A state value function is established according to the strategy selection probability distribution. The state value function calculates the long-term expected return of the game payment matrix based on a discount factor. The air-ground collaborative strategy is obtained according to the gradient of the state value function to the strategy selection probability distribution.
[0127] Exemplarily, the state vector of the air countermeasure device includes: transmission power (e.g., adjustable from 20W to 200W), geographic location coordinates (latitude, longitude and altitude, such as N30°11′24″, E120°13′49″, altitude 500 meters), coverage range (circular area with a radius of 1000 meters) and energy status (80% remaining power). The state vector of the ground countermeasure device includes: transmission power (e.g., adjustable from 50W to 300W), geographic location coordinates (e.g., N30°11′26″, E120°13′52″, altitude 5 meters), coverage range (fan-shaped area with a radius of 500 meters) and energy status (65% remaining power). Status acquisition is updated in real time through a wireless communication link, with an update period of 100 milliseconds.
[0128] An air-to-ground countermeasure game theory model is constructed, with the air countermeasure device as player A and the ground countermeasure device as player B. Based on the state vector, a game strategy space is constructed: a power selection strategy set (air devices: {40W, 80W, 120W, 160W}; ground devices: {75W, 150W, 225W, 300W}), a jamming direction strategy set (air devices: {omnidirectional, 30° sector, 60° sector, 90° sector}; ground devices: {45° sector, 90° sector, 180° sector, 360°}), and a resource allocation strategy set (air devices: {20%, 40%, 60%, 80%}; ground devices: {25%, 50%, 75%, 100%}). These strategy combinations result in 4×4×4=64 air countermeasure strategies and 4×4×4=64 ground countermeasure strategies, for a total of 64×64=4096 joint strategy combinations.
[0129] When constructing a multi-objective utility function, the interference effect evaluation function is first calculated. For the signal-to-interference ratio (SIR) evaluation, the score is calculated using three parameters: signal strength, spectrum coverage width, and interference incidence angle. For example, when the airborne device selects 120W power, a 60° sector, and a 60% resource allocation, the average SIR within the target area is calculated to be -12dB, corresponding to a score of 85. The coverage effect evaluation considers the coverage ratio of the target area; under the same conditions, the coverage score is 78. The directionality evaluation considers energy concentration and is scored 92. A weighted combination of the three (weights of 0.5, 0.3, and 0.2, respectively) yields a total interference effect score of 84.9. The resource consumption evaluation function uses a weighted combination of power consumption, communication overhead, and maneuvering cost. Under these conditions, the power consumption score is 65 (high consumption), the communication overhead score is 88 (low overhead), and the maneuvering cost score is 72 (medium consumption). With weights of 0.6, 0.2, and 0.2, respectively, the total resource consumption score is 71.4.
[0130] The game payoff matrix is constructed by combining the interference effect and resource consumption evaluations with a resource penalty factor (set to 0.7). The calculation formula is: Payoff = Interference Effect Score - Resource Penalty Factor × Resource Consumption Score. Under these conditions, the payoff for a single aerial device is 84.9 - 0.7 × 71.4 = 34.92. The payoffs for all strategy combinations are iteratively calculated to complete the 4096 × 2 payoff matrix.
[0131] Based on the constructed 4096×2 game payoff matrix, the optimal response iteration method is used to solve the Nash equilibrium: The strategy probability distributions for the air and ground devices are randomly initialized, generating initial probability vectors for each of the 64 strategies. Each iteration performs two steps: First, the ground device's current strategy is fixed, and the optimal response strategy for the air device is calculated. Specifically, the expected payoff of each possible air device strategy under the current ground device strategy is calculated, and the strategy with the largest payoff is selected as the air device's updated strategy. Finally, the updated air device strategy is fixed, and the optimal response strategy for the ground device is calculated. Convergence conditions are set as follows: the change in strategy between two consecutive iterations is less than a threshold of 0.01 or the maximum number of iterations is reached. In this example, the algorithm converged after 23 iterations, resulting in the Nash equilibrium strategy combinations: the air device chooses (80W power, 60° sector, 40% resource allocation) and the ground device chooses (150W power, 90° sector, 50% resource allocation), with corresponding payoffs of 42.5 and 39.8, respectively. This strategy combination represents the optimal response strategy of both devices in a comprehensive game considering interference effects and resource consumption.
[0132] The air-ground coordination advantage was calculated based on the optimal response strategy. Specifically, the Nash equilibrium was calculated by comparing the combined counterattack payoff (e.g., 78.6 points) to the individual counterattack payoff (42.5 points for airborne units and 39.8 points for ground units). The ratio of the combined payoff to the individual payoff yielded an air-ground coordination advantage of 78.6 / (42.5 + 39.8) = 0.96, indicating that coordinated action has a significant advantage over individual action. Subsequently, a strategy selection probability distribution was constructed based on the air-ground coordination advantage, and the softmax method was used to convert the expected payoff of each strategy into a selection probability. The calculation results show that the optimal strategy (80W power, 60° sector, 40% resource allocation) for airborne units has a probability of 0.35, while the suboptimal strategy (120W power, 60° sector, 60% resource allocation) has a probability of 0.28. The optimal strategy (150W power, 90° sector, 50% resource allocation) for ground units has a probability of 0.33.
[0133] The state-value function is directly established based on the strategy selection probability distribution. This is done by multiplying the probability of selecting each possible strategy by the corresponding payoff value, and then summing the results to obtain the expected value of the current state. For example, the state value of an aerial device is calculated as: 0.35 × 42.5 (optimal strategy payoff) + 0.28 × 40.2 (suboptimal strategy payoff) + ... = 41.3. When establishing the state-value function, a discount factor of 0.85 is set to calculate the long-term expected payoff of the game payoff matrix, which is the current payoff plus the discounted value of future payoffs. Based on the gradient of the state-value function with respect to the strategy selection probability distribution, the system uses a policy gradient method for optimization, specifically by increasing the probability of selecting high-payoff strategies and decreasing the probability of selecting low-payoff strategies. After 200 iterative optimizations, the final air-ground collaboration strategy was obtained: the aerial equipment selected (80W power, 60° sector, 40% resource allocation) with a probability of 40%, (120W power, 60° sector, 60% resource allocation) with a probability of 35%, and the alternative strategy with a probability of 25%; the ground equipment selected (150W power, 90° sector, 50% resource allocation) with a probability of 45%, (225W power, 90° sector, 75% resource allocation) with a probability of 30%, and the alternative strategy with a probability of 25%.
[0134] Existing technologies primarily use fixed preset strategies or simple adaptive methods for countermeasure decisions, each making decisions based on local information and lacking a systematic coordination mechanism. The present invention models air and ground countermeasure equipment as game participants, constructing a multidimensional strategy space encompassing power selection, interference direction, and resource allocation; designs a multi-objective utility function that combines interference effects and resource consumption; quantifies the synergistic benefits by calculating the air-ground collaborative advantage; and optimizes collaborative decisions using a state-value function and a policy gradient method. This game-theory-based approach establishes air-ground countermeasure decisions within a rigorous mathematical framework, enabling theory to guide practice. The present invention significantly enhances the synergy of air-ground countermeasures, enabling both devices to make optimal responses based on the Nash equilibrium principle, avoiding the decision-making conflicts found in traditional methods; significantly enhances the countermeasure effect, achieving higher interference efficiency under the same resource constraints through the optimization of multi-objective utility functions; and improves resource utilization efficiency. Through game equilibrium calculation and air-ground collaborative advantage evaluation, unnecessary resource consumption is reduced while ensuring the countermeasure effect.
[0135] In an optional embodiment, the reinforcement learning algorithm optimizes the allocation of power resources and computing resources in real time based on the device state and target characteristics, including:
[0136] Obtaining device status and target features, and constructing a resource status tensor including platform type and resource type, wherein the platform type includes air platform and ground platform, and the resource type includes power resource and computing resource;
[0137] Constructing a variational autoencoder to encode the resource state tensor to generate latent variables, and decoding the latent variables into communication messages between air-ground platforms;
[0138] Based on the resource state tensor and the communication messages between the air-ground platform, a dual-path deep Q network with a value stream and an advantage stream is designed. The value stream evaluates the state value of resource allocation, and the advantage stream evaluates the action advantage of resource allocation. The state value and action advantage are combined to obtain the Q value of the resource allocation strategy. A proximal policy optimization algorithm is used to train the policy network. The policy network includes an actuator network and a critic network. The actuator network outputs the mean and standard deviation of the resource allocation action distribution that meets the air-ground collaboration requirements according to the constraints of the air-ground collaboration strategy. The critic network estimates the state value and calculates the action advantage based on the current state.
[0139] A hierarchical priority experience replay pool is constructed, sample priorities are calculated based on temporal difference error, and experience samples are sampled according to the sample priorities. A meta-learning framework is used to update the policy, and an air-ground consistency evaluation function is constructed to impose additional constraints on the training of the policy network. Based on the trained policy network, the power and computing resources of the air and ground platforms are optimized in real time.
[0140] For example, device status and target feature information are obtained as input data for the resource status tensor. Device status includes current battery level, processor load, memory usage, signal strength, etc.; target features include task priority, computational complexity, and time constraints. For example, a drone formation system consists of five drones of different models on the air platform and three computing nodes on the ground platform. Status information is collected for each of these nodes. For example, drone 1 has a battery level of 85%, a CPU load of 40%, and is performing a target tracking task with a priority of 3.
[0141] The resource status tensor contains two dimensions: platform type and resource type. Platform types are categorized as airborne and ground platforms; resource types include power resources and computing resources. For each airborne platform, its available power percentage, computing unit utilization, remaining storage space, and other information are recorded; for ground platforms, its power supply status, computing load, and network bandwidth utilization are recorded. Each metric is normalized to ensure that the value ranges from 0 to 1, forming a resource status tensor of dimension [8, 10], where 8 represents the number of platforms and 10 represents the resource status feature dimension of each platform.
[0142] The variational autoencoder consists of an encoder and a decoder. The encoder contains a three-layer fully connected network with 80 nodes in the input layer (i.e., the dimension of the flattened resource state tensor), 64 nodes in the hidden layer, and 32 nodes in the hidden layer. The output layer generates 16-dimensional latent variables and their mean and variance. The decoder also uses a three-layer fully connected network to map the latent variables back to the original dimensional space and generate communication messages between the air-ground platform. The communication messages contain information such as resource demand predictions, task priority adjustment suggestions, and coordination strategy parameters. For example, when UAV 2's computing resources are insufficient, it will send a task offload request to ground computing node 1, containing information such as task type, data size, and expected completion time.
[0143] Based on the resource state tensor and communication messages, a two-way deep Q-network is designed. This network consists of a shared feature extraction layer and separate value and advantage streams. The feature extraction layer consists of two convolutional layers with 16 and 32 filters, a 3×3 kernel size, and one fully connected layer with 128 nodes. The value stream, consisting of two fully connected layers with 64 and 1 node, is used to estimate the overall value of the current resource allocation state. The advantage stream, also consisting of two fully connected layers with 64 nodes and an action space dimension (20 in this example, representing the possible resource allocation combinations), is used to assess the advantage of each action relative to the average. The outputs of the value and advantage streams are combined to produce a final Q-value, representing the expected reward of taking each resource allocation action in the current state. The two-way deep Q-network generates Q-value estimates for all possible resource allocation actions in the current state through a forward propagation every 100 milliseconds. These Q-values are stored in an action-value vector of size 20 and subsequently used to guide the action selection and evaluation process of the policy network.
[0144] The resource allocation policy network is trained using the proximal policy optimization (PPO) algorithm. The policy network directly receives the Q-value output of the dual-pass deep Q-network as additional input to optimize action selection. The policy network consists of an actor network (ActorNetwork) and a critic network (Critic Network), which share the first two feature extraction layers but have different output layers. The actor network contains three fully connected layers, with 128, 64, and 40 nodes, respectively. Its inputs are the resource state tensor and the action Q-value vectors provided by the dual-pass deep Q-network. Its outputs are the mean vector (20 dimensions) and the standard deviation vector (20 dimensions) of the action distribution, which together define the Gaussian policy distribution for resource allocation. Specifically, a fusion layer is designed after the actor network's second hidden layer (64 nodes). This layer concatenates the 20-dimensional Q-value vector from the dual-pass deep Q-network with the current hidden layer representation. Feature fusion is then performed through an additional fully connected layer (64 nodes) to ensure that the Q-value information directly influences policy generation. The 40-dimensional final output corresponds to two types of decisions: power and computing resource allocation across eight platforms (five airborne and three ground-based), for a total of 16 decision dimensions. In addition to the two parameters, mean and variance, this results in a total of 40 output nodes. The actuator network receives the constraints of the aforementioned air-ground collaboration strategy (e.g., implemented through additional terms in the loss function) and combines them with action evaluation information provided by Q-values to ensure that the output resource allocation actions meet the requirements of air-ground collaboration and have high long-term benefits. For example, it ensures that the computing resource allocation ratio between air and ground platforms is within a reasonable range (for example, using a SoftMax function to normalize the allocation ratio) and that at least 30% of the computing resources on the ground platform are reserved for critical task processing (achieved by setting a minimum threshold).
[0145] The critic network also consists of three fully connected layers, with 128, 64, and 1 nodes, respectively. Its inputs are similarly the resource state tensor and the Q-value vector from the dual-pass deep Q-network. Its output is an estimated state value V, which is used to estimate the long-term benefits of the current resource allocation state. In its implementation, the critic network incorporates a comparison mechanism that compares its own estimated state value V with the maximum Q-value estimated by the dual-pass deep Q-network. A consistency loss term is then calculated to encourage consistency between the two value estimates, thereby improving evaluation accuracy. By comparing the critic network's state value estimate with the actual reward, a temporal difference error is calculated, which is then used to calculate the action advantage value. This calculation process references the advantage flow design of the dual-pass deep Q-network, but uses the critic's own value estimate. The action advantage value is calculated by subtracting the state value function V from the action-state value function Q. It represents the degree of advantage of a particular action relative to the average state value. The PPO algorithm uses this advantage value (along with information provided by the dual-pass Q-network) to guide the actuator network's policy direction, updating policy parameters by maximizing a specific objective function. The objective function is the smaller of the ratio of the probability of the new policy to the old policy multiplied by the advantage value, or the ratio of the probability of the new policy to the old policy multiplied by the advantage value after clipping. The new policy probability ratio is the ratio of the probability of an action under the new policy to the probability of an action under the old policy. The clipping parameter is set to 0.2 in this example to limit the step size of the policy update and ensure training stability.
[0146] When the system is running, the dual-path deep Q network and the policy network form a complementary collaborative relationship: after each calculation and update of the Q network, its output Q value is immediately used by the policy network to guide the next decision; and the actions generated by the policy network and their actual effects are used to update the parameters of the Q network, forming a closed-loop optimization process.
[0147] A hierarchical priority experience replay pool is constructed to store and sample training data. The replay pool capacity is set to 10,000 records, each of which contains five components: current state, action performed, reward obtained, next state, and termination flag. Sample priority is calculated based on temporal difference error (TDE). Samples with larger TDEs contain more valuable information and receive higher priority. In actual operation, samples that significantly improve system performance after a certain resource allocation have a TDE of 0.85, far higher than the average level of 0.32, and therefore are given a higher sampling probability.
[0148] A meta-learning framework is used for policy updates, adjusting network parameters via gradient descent. During each training cycle, 128 records are sampled from the experience replay pool based on priority, and the loss functions of the actuator and critic networks are calculated and updated. Furthermore, an air-ground consistency evaluation function is constructed to constrain the policy network, ensuring that resource allocation meets the requirements for inter-air and ground platform collaboration, such as ensuring balanced communication bandwidth allocation between the ground command and control center and the UAVs, and a reasonable distribution of computational tasks. After 1000 rounds of training, the policy network is determined to have converged, and resource allocation is optimized based on the trained policy network. For example, if UAV 3 detects insufficient computational resources for target recognition, the policy network automatically offloads 50% of its computational tasks to ground node 2. The UAV's power configuration is also adjusted, allocating the remaining 25% of power to the communication module, improving data transmission efficiency.
[0149] Figure 2 This is a comparison chart of the information transmission accuracy of different communication schemes under different communication loads. The figure shows the comparison of the information transmission accuracy of three different communication schemes under different communication loads. The horizontal axis represents the air-ground platform communication load (Mbps), ranging from 0 to 60Mbps; the vertical axis represents the information transmission accuracy (%), ranging from 50% to 99%. The performance of the three communication schemes is as follows: The present invention (VAE communication) is represented by a circular mark: it maintains the highest transmission accuracy under all communication load conditions, and can still maintain an accuracy of 92.1% even under high load (60Mbps) conditions. Under low load conditions, it can reach an accuracy of 98.5%, showing excellent communication performance and stability. Traditional direct communication is represented by a diamond mark: the performance of this scheme decreases rapidly as the load increases. The accuracy is 91.8% at low load, but drops sharply to 56.3% at high load (60Mbps), indicating its obvious limitations under high communication pressure. The performance based on compressed sensing is represented by square marks: it is between the first two, reaching 95.3% under low load and 72.3% under high load, showing a medium level of anti-interference ability and coding efficiency.
[0150] Figure 3 To optimize the performance comparison chart, this chart compares the training effects and convergence performance of three different reinforcement learning optimization algorithms. The horizontal axis represents the training rounds (0-1200 rounds), and the vertical axis represents the resource allocation efficiency score (0-100 points). The present invention (dual-path Q network + policy network) is represented by a circular mark: it shows the fastest convergence speed and the highest final performance. After 400 rounds of training, the efficiency score reached 68.3, and finally converged to a high efficiency score of 97.8 after 1200 rounds. The single DQN network is represented by a square mark: the convergence speed is the slowest. The single PPO algorithm is represented by a diamond mark: the performance is between the first two, showing a medium level of air-ground collaborative advantage.
[0151] The existing technology mainly adopts fixed rule strategies or simple adaptive algorithms for resource allocation, and lacks a unified coordination mechanism. The hybrid architecture of the integrated variational autoencoder, dual-path deep Q network and policy network of the present invention realizes the unified optimization allocation of air-ground platform resources. The present invention designs a dual-path Q network and policy network collaborative mechanism so that the Q value directly participates in policy decision-making; introduces an air-ground consistency evaluation function to incorporate resource balance, task completion and collaborative effectiveness into a unified evaluation framework. At the same time, through the hierarchical priority experience replay and meta-learning framework, the algorithm's adaptability and generalization to dynamic environments are enhanced. The present invention realizes the unified optimization of power resources and computing resources, avoiding resource allocation conflicts between platforms; the system response speed is significantly accelerated, and resource strategies can be adjusted in real time under complex and changing environments; the quality of task completion is significantly improved, especially under resource-constrained conditions, efficient execution can be maintained; the overall energy efficiency of the system is improved, while ensuring task performance while reducing energy consumption.
[0152] In an optional embodiment, the space-land consistency evaluation function includes:
[0153] The air-ground consistency evaluation function includes a resource balance evaluation item, a task completion evaluation item and a collaborative effectiveness evaluation item. The resource balance evaluation item calculates the difference in resource allocation between the air platform and the ground platform. The task completion evaluation item evaluates the task execution effect based on the platform resource utilization. The collaborative effectiveness evaluation item is evaluated based on the communication quality and collaborative response time between the air-ground platforms. The conformity of the resource allocation action with the air-ground collaborative strategy is calculated based on the weighted combination of the resource balance evaluation item, the task completion evaluation item and the collaborative effectiveness evaluation item.
[0154] For example, the resource balance assessment item is used to calculate the differences in resource allocation between air and ground platforms. A set of air and ground platforms can be set, with each platform containing four resource types: computing resources, storage resources, energy resources, and communication resources. For each resource type, the variance of the resource allocation between the air and ground platforms is calculated. For example, if the computing resource allocation ratio on the air platforms is [0.6, 0.7, 0.5] and the ratio on the ground platforms is [0.8, 0.9, 0.75], then the resource allocation variance is 0.037. Similarly, the variance values for the other resource types are calculated, and finally the weighted sum of the variance values for the four resource types is calculated to obtain the resource balance assessment value. In actual applications, the smaller the resource allocation variance, the higher the balance assessment value, indicating a more reasonable resource allocation. Assume there are three air platforms and four ground platforms, and the weights of the four resource types are 0.3, 0.2, 0.3, and 0.2, respectively. The computing resource allocation ratios for the air platform are [0.65, 0.72, 0.58], and for the ground platform are [0.82, 0.85, 0.78, 0.80]. The storage resource allocation ratios for the air platform are [0.70, 0.75, 0.68], and for the ground platform are [0.60, 0.65, 0.62, 0.67]. The energy resource allocation ratios for the air platform are [0.55, 0.60, 0.58], and for the ground platform are [0.88, 0.92, 0.85, 0.90]. The communication resource allocation ratios for the air platform are [0.80, 0.85, 0.82], and for the ground platform are [0.75, 0.78, 0.72, 0.80]. The calculated variances for each resource type are 0.026, 0.005, 0.093, and 0.002, respectively. The weighted resource balance assessment value is 0.035, indicating relatively balanced resource allocation.
[0155] The task completion evaluation assesses task execution performance based on platform resource utilization. Resource utilization data for all platforms is collected, including CPU utilization, memory utilization, energy consumption, and bandwidth utilization. For each task T, its resource utilization metrics are calculated on each platform. A target completion threshold is set to determine whether the task has met the completion criteria. For example, a task is executed on five platforms, and the resource utilization of each platform is [0.75, 0.82, 0.68, 0.79, 0.73]. If the resource utilization threshold is set to 0.65, the task meets the completion criteria on all platforms. Suppose there are 10 tasks in the system, and the average resource utilization of each task on different platforms is [0.78, 0.72, 0.80, 0.65, 0.81, 0.75, 0.69, 0.82, 0.77, 0.70]. Setting the completion threshold to 0.65 results in a calculated task completion of 100%. Raising the threshold to 0.70 reduces the task completion to 90%. In addition, task priorities can be weighted, such as a critical task weight of 0.6 and an ordinary task weight of 0.4, and the final weighted calculation can be used to obtain the task completion assessment value.
[0156] The collaborative effectiveness evaluation is based on the communication quality and collaborative response time between air-ground platforms. Communication quality can be quantified using metrics such as data transmission success rate, signal strength, and bit error rate. Collaborative response time is determined by measuring the time interval between command issuance and execution. The quality of the communication link between each pair of air-ground platforms is monitored. For example, if the data transmission success rate between air platform A and ground platform B is 95%, the signal strength is -65dBm, and the bit error rate is 0.002, the communication quality score can be calculated as 0.85. The response time of each collaborative task is also recorded. For example, if the time from command issuance to completion of a collaborative reconnaissance task is 1.5 seconds, and the preset threshold is 2 seconds, the response time score is 0.75. Assume that the air-ground system consists of three air platforms and four ground platforms, forming 12 communication links. The communication quality scores for each link are [0.88, 0.92, 0.85, 0.90, 0.83, 0.87, 0.91, 0.86, 0.89, 0.84, 0.93, 0.88], with an average communication quality score of 0.88. The collaborative task response time scores are [0.82, 0.78, 0.85, 0.80, 0.83], with an average response time score of 0.82. Assuming a communication quality weight of 0.6 and a response time weight of 0.4, the calculated collaborative effectiveness evaluation value is 0.856.
[0157] By weighting the three evaluation metrics above, we calculated the degree to which resource allocation actions matched the air-ground coordination strategy. We set a weight of 0.3 for resource balance, 0.4 for task completion, and 0.3 for coordination effectiveness. The overall evaluation function is then calculated as the weighted sum of these factors.
[0158] In this case, assuming a resource balance evaluation of 0.85, a task completion evaluation of 0.92, and a collaborative effectiveness evaluation of 0.88, the overall air-ground consistency evaluation function is 0.3 × 0.85 + 0.4 × 0.92 + 0.3 × 0.88 = 0.887. This value indicates that the current resource allocation action is highly consistent with the air-ground collaborative strategy, and that resource allocation is reasonable and effective.
[0159] During the application of the evaluation function, the weights of various evaluation items can be dynamically adjusted based on the specific scenario. For example, in resource-constrained environments, the weight of resource balance can be increased; in scenarios with high timeliness requirements, the weight of task completion can be increased; and in complex collaborative scenarios, the weight of collaborative effectiveness can be increased. By properly setting weight parameters, the evaluation function can be better adapted to the needs of different application scenarios, improving the overall performance of the air-ground collaborative system.
[0160] By designing an air-ground consistency evaluation function, this paper integrates the three key dimensions of resource balance, task completion, and collaborative effectiveness, achieving a precise assessment of the degree of conformity between resource allocation actions and air-ground collaborative strategies. This evaluation function can effectively coordinate resource allocation differences between aerial and ground platforms, optimize task execution, and improve inter-platform communication quality and collaborative response efficiency. By dynamically adjusting the weights of each evaluation item, the system can adapt to the needs of different application scenarios, improve resource utilization and task completion quality, and enhance the overall performance and adaptability of the air-ground collaborative system.
[0161] A second aspect of an embodiment of the present invention provides an adaptive generation system for UAV countermeasure signals for air-ground collaboration, comprising:
[0162] The first unit is used to obtain the communication signal of the target UAV, extract the signal characteristic parameters and analyze them to obtain the communication protocol type;
[0163] The second unit is configured to encode signal characteristic parameters using a two-level encoding structure, wherein the first level represents the interference strategy type and the second level represents specific parameters; optimize the encoded parameters using a genetic algorithm with a composite fitness function, wherein the composite fitness function is used to evaluate interference effect, power consumption efficiency, and anti-detection performance; select a corresponding basic template from a preset interference waveform template library according to the communication protocol type, and use the optimized parameters to perform parameter matching and adjustment to generate an interference signal template;
[0164] The third unit is used to send the interference signal template to the air countermeasure device and the ground countermeasure device;
[0165] The fourth unit is used to optimize resource allocation for air-ground collaboration, including: constructing an air-ground collaboration strategy based on a game theory model; dynamically allocating resources according to the air-ground collaboration strategy using a reinforcement learning algorithm, wherein the reinforcement learning algorithm optimizes the allocation of power and computing resources in real time based on device status and target characteristics;
[0166] The fifth unit is used to control the air and ground countermeasure equipment to generate interference signals and transmit them to the target UAV based on the resource optimization allocation results.
[0167] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:
[0168] processor;
[0169] a memory for storing processor-executable instructions;
[0170] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0171] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0172] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The adaptive generation method of UAV countermeasure signals for air-ground collaboration is characterized by: include: Obtain the communication signal of the target UAV, extract the signal characteristic parameters and analyze them to obtain the communication protocol type; A two-level coding structure is used to encode signal characteristic parameters, wherein the first level represents the interference strategy type and the second level represents the specific parameters, including: the interference strategy type is encoded using a 4-bit binary code to obtain a strategy code, and the interference strategy type includes a noise interference strategy, a deception interference strategy, a blocking interference strategy and a combined interference strategy; the specific parameters are encoded using a variable-length coding structure to obtain a parameter code, and the specific parameters include frequency domain parameters, time domain parameters and power parameters, and the frequency domain parameters include interference bandwidth ratio, frequency offset and power spectrum shape factor, the time domain parameters include interference duration, delay and repetition period, and the power parameters include interference power , power allocation coefficient and waveform factor; based on the importance of the parameters, the coding bit width of the parameter coding is optimized to obtain the optimized coding, and the coding bit width is determined according to the parameter value range and the quantization accuracy requirement; the strategy coding and the optimization coding are adaptively mapped according to the target characteristics to obtain the final feature parameter coding; the encoded parameters are optimized using a genetic algorithm with a composite fitness function, and the composite fitness function is used to evaluate the interference effect, power consumption efficiency and anti-detection performance; according to the communication protocol type, a corresponding basic template is selected from a preset interference waveform template library, and the optimized parameters are used to match and adjust the parameters to generate an interference signal template; Sending the interference signal template to the air countermeasure equipment and the ground countermeasure equipment; Executing resource optimization allocation for air-ground collaboration, including: constructing an air-ground collaboration strategy based on a game theory model, including: obtaining state vectors of air countermeasure equipment and ground countermeasure equipment, constructing an air-ground countermeasure game theory model, wherein the state vector includes power, position, coverage range and energy state; taking the air countermeasure equipment and the ground countermeasure equipment as game participants, constructing a game strategy space including power selection, interference direction and resource allocation according to the state vector, wherein the game strategy space is used to constrain the strategy selection range of the game participants; constructing a multi-objective utility function, wherein the multi-objective utility function includes an interference effect evaluation function obtained by weighted combination of signal-to-interference ratio evaluation, coverage effect evaluation and directionality evaluation, and a resource consumption evaluation function obtained by weighted combination of power consumption, communication overhead and maneuvering cost , the multi-objective utility function and the resource penalty factor are combined to construct a game payment matrix, and the optimal response strategy under the Nash equilibrium is solved based on the game payment matrix; the air-ground coordination advantage is calculated based on the optimal response strategy, and a strategy selection probability distribution is constructed, and the air-ground coordination advantage is determined by the ratio of the sum of the joint countermeasure benefit and the individual countermeasure benefit; a state value function is established according to the strategy selection probability distribution, and the state value function calculates the long-term expected benefit of the game payment matrix based on the discount factor, and the air-ground coordination strategy is obtained according to the gradient of the state value function to the strategy selection probability distribution; a reinforcement learning algorithm is used to dynamically allocate resources according to the air-ground coordination strategy, and the reinforcement learning algorithm optimizes the allocation of power resources and computing resources in real time based on the device status and target characteristics; Based on the results of resource optimization allocation, control the air and ground countermeasure equipment to generate interference signals and transmit them to the target UAV.
2. The method according to claim 1, characterized in that The encoded parameters are optimized using a genetic algorithm with a composite fitness function. The composite fitness function is used to evaluate the interference effect, power consumption efficiency and anti-detection performance, including: Constructing a composite fitness function, the composite fitness function including: an interference effect evaluation item for evaluating the signal-to-interference ratio, bit error rate, and interference coverage; a power consumption efficiency evaluation item for evaluating power efficiency, time efficiency, and energy utilization, where the energy utilization is determined by the time-domain integral ratio of effective interference power to total interference power; and an anti-detection evaluation item for evaluating detection probability, exposure time, and low interception characteristics, where the low interception characteristics are determined by interference bandwidth ratio, power fluctuation, and transmission pattern characteristics; The composite fitness function uses a dynamic weight coefficient to perform a weighted combination of the evaluation items, and the dynamic weight coefficient calculates a weight adjustment amount according to the deviation between the target performance and the actual performance, the target performance including the expected interference effect, power consumption efficiency and anti-detection index, and the actual performance including the interference effect, power consumption efficiency and anti-detection index measured in real time; The strategy coding and optimization coding are iteratively optimized using a genetic algorithm to obtain an optimal solution, and the optimized interference strategy and parameter configuration are output; the genetic algorithm adopts an adaptive mutation rate, and the adaptive mutation rate decreases as the fitness value increases.
3. The method according to claim 2, characterized in that The determination of dynamic weight coefficient includes: Establish a performance indicator system, including signal-to-interference ratio, bit error rate, and interference coverage in the interference effect indicator set; power efficiency, time efficiency, and energy utilization in the power consumption efficiency indicator set; detection probability, exposure time, and low interception characteristics in the anti-detection indicator set, and determine the ideal target value and acceptable range for each indicator; Use normalized deviation to calculate function δ ij =(T ij -A ij ) / N ij Calculate the performance deviation of each indicator, where T ij is the target performance value, A ij is the actual measured value, N ij is the normalization factor; the weighted sum of the deviations of similar indicators is used to obtain the comprehensive deviation Δ i ; The weight adjustment is calculated using a nonlinear bias response function: Where ΔW i is the adjustment amount of the weight corresponding to the i-th indicator, α is the adjustment amplitude coefficient, β is the nonlinear index used to control the curvature of the response function, γ is the sensitivity coefficient used to control the response degree to small deviations, τ min is the minimum response threshold for filtering out noise, τ max is the maximum response threshold used to limit abnormal situations; sign(Δ i ) is the sign function of the bias, which determines the direction of weight adjustment; The weights are updated smoothly and normalized to ensure that the sum of the weights is 1 and the weight values are subject to upper and lower limits.
4. The method according to claim 1, wherein The reinforcement learning algorithm optimizes the allocation of power and computing resources in real time based on device status and target characteristics, including: Obtaining device status and target features, and constructing a resource status tensor including platform type and resource type, wherein the platform type includes air platform and ground platform, and the resource type includes power resource and computing resource; Constructing a variational autoencoder to encode the resource state tensor to generate latent variables, and decoding the latent variables into communication messages between air-ground platforms; Based on the resource state tensor and the communication messages between the air-ground platform, a dual-path deep Q network with a value stream and an advantage stream is designed. The value stream evaluates the state value of resource allocation, and the advantage stream evaluates the action advantage of resource allocation. The state value and action advantage are combined to obtain the Q value of the resource allocation strategy. A proximal policy optimization algorithm is used to train the policy network. The policy network includes an actuator network and a critic network. The actuator network outputs the mean and standard deviation of the resource allocation action distribution that meets the air-ground collaboration requirements according to the constraints of the air-ground collaboration strategy. The critic network estimates the state value and calculates the action advantage based on the current state. A hierarchical priority experience replay pool is constructed, sample priorities are calculated based on temporal difference error, and experience samples are sampled according to the sample priorities. A meta-learning framework is used to update the policy, and an air-ground consistency evaluation function is constructed to impose additional constraints on the training of the policy network. Based on the trained policy network, the power and computing resources of the air and ground platforms are optimized in real time.
5. The method according to claim 4, characterized in that The air-ground consistency evaluation functions include: The air-ground consistency evaluation function includes a resource balance evaluation item, a task completion evaluation item and a collaborative effectiveness evaluation item. The resource balance evaluation item calculates the difference in resource allocation between the air platform and the ground platform. The task completion evaluation item evaluates the task execution effect based on the platform resource utilization. The collaborative effectiveness evaluation item is evaluated based on the communication quality and collaborative response time between the air-ground platforms. The conformity of the resource allocation action with the air-ground collaborative strategy is calculated based on the weighted combination of the resource balance evaluation item, the task completion evaluation item and the collaborative effectiveness evaluation item.
6. An adaptive generation system for UAV countermeasure signals in air-ground coordination, used to implement the method according to any one of claims 1 to 5, characterized in that: include: The first unit is used to obtain the communication signal of the target UAV, extract the signal characteristic parameters and analyze them to obtain the communication protocol type; The second unit is configured to encode signal characteristic parameters using a two-level encoding structure, wherein the first level represents the interference strategy type and the second level represents specific parameters; optimize the encoded parameters using a genetic algorithm with a composite fitness function, wherein the composite fitness function is used to evaluate interference effect, power consumption efficiency, and anti-detection performance; select a corresponding basic template from a preset interference waveform template library according to the communication protocol type, and use the optimized parameters to perform parameter matching and adjustment to generate an interference signal template; The third unit is used to send the interference signal template to the air countermeasure device and the ground countermeasure device; The fourth unit is used to optimize resource allocation for air-ground collaboration, including: constructing an air-ground collaboration strategy based on a game theory model; dynamically allocating resources according to the air-ground collaboration strategy using a reinforcement learning algorithm, wherein the reinforcement learning algorithm optimizes the allocation of power and computing resources in real time based on device status and target characteristics; The fifth unit is used to control the air and ground countermeasure equipment to generate interference signals and transmit them to the target UAV based on the resource optimization allocation results.
7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Unmanned aerial vehicle trajectory optimization method and system based on bionic algorithm and BP neural network
CN116782269A
Unmanned aerial vehicle communication anti-interference decision model and method based on deep reinforcement learning
CN118138173A