Air-ground cooperative unmanned aerial vehicle countering signal self-adaptive generation method and system
Through the adaptive generation method of air-ground collaborative drone counter signal, the secondary encoding structure and genetic algorithm are used to optimize parameters, combined with game theory and reinforcement learning algorithm resource allocation, the problems of low counter efficiency and high power consumption in the existing technology are solved, and the efficient and low power consumption drone counter effect is achieved.
Patent Information
- Application Number
- CN202510686556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing drone countermeasures technology has problems such as fixed frequency, lack of adaptability, low resource utilization efficiency, difficulty in forming a three-dimensional protection network, and high power consumption and easy detection.
Adaptive generation method of air-ground collaboration drone counter signal is adopted. By obtaining the communication signals of the target drone, signal characteristic parameters are extracted and communication protocol types are analyzed, parameters are optimized using secondary encoding structures and genetic algorithms to generate interference signal templates, and through the optimization allocation of air-ground collaboration resources, resources are dynamically allocated based on game theory and reinforcement learning algorithms, and the counter equipment is controlled to generate interference signals.
It realizes efficient interference effects for different communication protocols, reduces power consumption, enhances concealment, improves the overall efficiency and adaptability of the counter system, and can cope with complex and changeable drone invasion scenarios.
Smart Images

Figure CN120223235A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to unmanned aerial vehicle (UAV) countermeasure technology, and particularly to an adaptive generation method and system for UAV countermeasure signals with air-ground cooperation. Background Art
[0002] To cope with the threats posed by illegal UAVs, UAV countermeasure technology has become a current research hotspot. Existing UAV countermeasure technologies have various deficiencies. Traditional electromagnetic interference methods usually use interference signals with fixed frequencies and fixed waveforms, lacking the adaptive ability for different communication protocols, resulting in low interference efficiency and limited effectiveness against high-end UAVs with frequency hopping or encrypted communication capabilities. Existing countermeasure systems generally lack an air-ground cooperation mechanism. Either they rely only on ground equipment with limited coverage, or they rely on a single platform to implement interference, with low resource utilization efficiency and difficulty in forming a three-dimensional protection network. Existing countermeasure systems often neglect the comprehensive consideration of power consumption efficiency and anti-detection performance during the generation of interference signals, leading to high energy consumption of countermeasure devices and easy detection, especially performing poorly in scenarios that require long-term continuous operation. Summary of the Invention
[0003] Embodiments of the present invention provide an adaptive generation method and system for UAV countermeasure signals with air-ground cooperation, which can solve the problems in the prior art.
[0004] In a first aspect of embodiments of the present invention, an adaptive generation method for UAV countermeasure signals with air-ground cooperation is provided, including: Obtaining the communication signal of a target UAV, extracting signal characteristic parameters and analyzing to obtain the type of communication protocol; Encoding the signal characteristic parameters using a two-level coding structure, where the first level represents the type of interference strategy and the second level represents specific parameters; optimizing the encoded parameters using a genetic algorithm with a composite fitness function, where the composite fitness function is used to evaluate interference effect, power consumption efficiency, and anti-detection performance; selecting a corresponding basic template from a preset interference waveform template library according to the type of communication protocol, and performing parameter matching and adjustment using the optimized parameters to generate an interference signal template; Sending the interference signal template to air countermeasure devices and ground countermeasure devices; Performing resource optimization allocation with air-ground cooperation, including: constructing an air-ground cooperation strategy based on a game theory model; using a reinforcement learning algorithm to perform dynamic resource allocation according to the air-ground cooperation strategy, where the reinforcement learning algorithm performs real-time optimization allocation of power resources and computing resources based on device status and target characteristics; Controlling air and ground countermeasure devices to generate interference signals and transmit them to the target UAV according to the resource optimization allocation result.
[0005] In an alternative embodiment, The signal feature parameters are encoded using a two - level encoding structure, where the first level represents the type of interference strategy, and the second level represents specific parameters, including: The type of interference strategy is encoded at the first level using a 4 - bit binary code to obtain a strategy code. The types of interference strategies include a noise interference strategy, a deception interference strategy, a blocking interference strategy, and a combined interference strategy; The specific parameters are encoded at the second level using a variable - length encoding structure to obtain a parameter code. The specific parameters include frequency - domain parameters, time - domain parameters, and power parameters. The frequency - domain parameters include the interference bandwidth ratio, frequency offset, and power spectrum shape factor. The time - domain parameters include the interference duration, time delay, and repetition period. The power parameters include the interference power, power distribution coefficient, and waveform factor; The encoding bit - width of the parameter code is optimized based on the parameter importance to obtain an optimized code, and the encoding bit - width is determined according to the parameter value range and quantization accuracy requirements; According to the target characteristics, the strategy code and the optimized code are adaptively mapped to obtain the final feature parameter code.
[0006] In an alternative embodiment, A genetic algorithm with a composite fitness function is used to optimize the encoded parameters. The composite fitness function is used to evaluate the interference effect, power consumption efficiency, and anti - detection performance, including: A composite fitness function is constructed. The composite fitness function includes: an interference effect evaluation term for evaluating the signal - to - interference ratio, bit error rate, and interference coverage range; a power consumption efficiency evaluation term for evaluating the power efficiency, time efficiency, and energy utilization rate, where the energy utilization rate is determined by the time - domain integral ratio of the effective interference power to the total interference power; an anti - detection performance evaluation term for evaluating the detection probability, exposure time, and low - intercept characteristic, where the low - intercept characteristic is determined by the interference bandwidth ratio, power fluctuation degree, and emission pattern characteristics; The composite fitness function uses dynamic weight coefficients to perform weighted combination on the evaluation terms. The dynamic weight coefficients calculate the weight adjustment amount according to the deviation between the target performance and the actual performance. The target performance includes the desired interference effect, power consumption efficiency, and anti - detection performance indicators, and the actual performance includes the real - time measured interference effect, power consumption efficiency, and anti - detection performance indicators; The genetic algorithm is used to iteratively optimize the strategy code and the optimized code to obtain the optimal solution, and output the optimized interference strategy and parameter configuration; The genetic algorithm uses an adaptive mutation rate, and the adaptive mutation rate decreases as the fitness value increases.
[0007] In an alternative embodiment, The determination of the dynamic weight coefficient includes: Establish a performance index system, including the signal-to-interference ratio, bit error rate, and interference coverage range in the interference effect index set, the power efficiency, time efficiency, and energy utilization rate in the power consumption efficiency index set, and the detection probability, exposure time, and low intercept characteristics in the anti-detection index set, and determine the ideal target value and acceptable range for each index; Use the normalized deviation calculation function δ ij =(T ij -A ij ) / N ij Calculate the performance deviation of each index, where T ij is the target performance value, A ij is the actual measured value, and N ij is the normalization factor; perform a weighted sum of the deviations of similar indexes to obtain the comprehensive deviation Δ i ; Use the non-linear deviation response function to calculate the weight adjustment amount: ; Among them, ΔW i is the adjustment amount of the weight corresponding to the i-th type of index, α is the adjustment amplitude coefficient, β is the non-linear exponent used to control the curvature of the response function, γ is the sensitivity coefficient used to control the response degree to small deviations, τ min is the minimum response threshold used to filter out noise, τ max is the maximum response threshold used to limit abnormal situations; sign(Δ i ) is the sign function of the deviation, which determines the direction of weight adjustment; Perform smooth update and normalization of the weights to ensure that the sum of the weights is 1, and the weight values are constrained by the upper and lower limits.
[0008] In an alternative implementation, Construct an air-ground cooperation strategy based on the game theory model, including: Obtain the state vectors of the airborne countermeasure equipment and the ground countermeasure equipment, and construct an air-ground countermeasure game theory model. The state vectors include power, position, coverage range, and energy state; Take the airborne countermeasure equipment and the ground countermeasure equipment as game participants, and construct a game strategy space including power selection, interference direction, and resource allocation according to the state vectors. The game strategy space is used to constrain the strategy selection range of the game participants; Construct a multi-objective utility function. The multi-objective utility function includes an interference effect evaluation function obtained by weighted combination of signal-to-interference ratio evaluation, coverage effect evaluation, and directivity evaluation, and a resource consumption evaluation function obtained by weighted combination of power consumption, communication overhead, and maneuvering cost. Combine the multi-objective utility function with a resource penalty factor to construct a game payoff matrix, and solve the optimal response strategy under the Nash equilibrium based on the game payoff matrix; Calculate the air-ground cooperation superiority degree based on the optimal response strategy, and construct a probability distribution of strategy selection. The air-ground cooperation superiority degree is determined by the ratio of the sum of the joint countermeasure benefits to the separate countermeasure benefits; Establish a state value function according to the probability distribution of strategy selection. The state value function calculates the long-term expected benefit of the game payment matrix based on the discount factor, and obtains the air-ground cooperation strategy according to the gradient of the probability distribution of strategy selection with respect to the state value function.
[0009] In an alternative implementation manner, The reinforcement learning algorithm performs real-time optimal allocation of power resources and computing resources based on the device state and target characteristics, including: Obtain the device state and target characteristics, and construct a resource state tensor including platform types and resource types. The platform types include aerial platforms and ground platforms, and the resource types include power resources and computing resources; Construct a variational autoencoder to encode the resource state tensor to generate a latent variable, and decode the latent variable into communication messages between the aerial and ground platforms; Based on the resource state tensor and the communication messages between the aerial and ground platforms, design a two-way deep Q-network with a value stream and an advantage stream. The value stream evaluates the state value of resource allocation, and the advantage stream evaluates the action advantage of resource allocation. Combine the state value and the action advantage to obtain the Q value of the resource allocation strategy; Train the policy network using the proximal policy optimization algorithm. The policy network includes an actor network and a critic network. The actor network outputs the mean and standard deviation of the resource allocation action distribution that meets the requirements of air-ground cooperation according to the constraints of the air-ground cooperation strategy, and the critic network estimates the state value based on the current state and calculates the action advantage; Construct a hierarchical priority experience replay pool, calculate the sample priority based on the temporal difference error, and sample the experience samples according to the sample priority; Use the meta-learning framework for policy update, and construct an air-ground consistency evaluation function to impose additional constraints on the training of the policy network; According to the trained policy network, perform real-time optimal allocation of the power resources and computing resources of the aerial platform and the ground platform.
[0010] In an alternative implementation manner, The air-ground consistency evaluation function includes: The vacant land consistency evaluation function includes a resource balance evaluation item, a task completion evaluation item, and a collaborative efficiency evaluation item. The resource balance evaluation item calculates the difference in resource allocation between the aerial platform and the ground platform. The task completion evaluation item evaluates the task execution effect based on the platform resource utilization rate. The collaborative efficiency evaluation item is evaluated based on the communication quality and collaborative response time between the aerial and ground platforms. The compliance degree between the resource allocation action and the aerial-ground collaboration strategy is calculated according to the weighted combination of the resource balance evaluation item, the task completion evaluation item, and the collaborative efficiency evaluation item.
[0011] In the second aspect of the embodiments of the present invention, there is provided an adaptive generation system for unmanned aerial vehicle countermeasure signals for aerial-ground collaboration, including: A first unit for acquiring the communication signal of the target unmanned aerial vehicle, extracting signal feature parameters, and analyzing to obtain the communication protocol type; A second unit for encoding the signal feature parameters using a two-level coding structure, where the first level represents the interference strategy type and the second level represents specific parameters; using a genetic algorithm with a composite fitness function to optimize the encoded parameters, and the composite fitness function is used to evaluate the interference effect, power consumption efficiency, and anti-detection performance; selecting a corresponding basic template from a preset interference waveform template library according to the communication protocol type, and performing parameter matching and adjustment using the optimized parameters to generate an interference signal template; A third unit for sending the interference signal template to the aerial countermeasure device and the ground countermeasure device; A fourth unit for performing resource optimization allocation for aerial-ground collaboration, including: constructing an aerial-ground collaboration strategy based on a game theory model; using a reinforcement learning algorithm to perform dynamic resource allocation according to the aerial-ground collaboration strategy, and the reinforcement learning algorithm performs real-time optimization allocation of power resources and computing resources based on the device state and target characteristics; A fifth unit for controlling the aerial and ground countermeasure devices to generate interference signals and transmit them to the target unmanned aerial vehicle according to the resource optimization allocation result.
[0012] In the third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0013] In the fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0014] The present invention realizes the precise recognition and adaptive interference of the communication signal characteristics of the target unmanned aerial vehicle (UAV). Through the optimization of the secondary coding structure and the genetic algorithm with a composite fitness function, it can achieve high-efficiency interference effects for different communication protocols, while maintaining low power consumption and strong concealment.
[0015] The present invention adopts an air-ground cooperation strategy and a resource optimization allocation mechanism based on game theory, and realizes the dynamic allocation of power and computing resources between air and ground countermeasure devices through a reinforcement learning algorithm, significantly improving the overall efficiency and adaptability of the countermeasure system, and being able to cope with complex and changeable UAV intrusion scenarios.
[0016] The present invention combines signal processing, artificial intelligence, and resource optimization technologies to form a complete set of UAV countermeasure technology solutions, which not only improves the accuracy and success rate of countermeasures, but also reduces energy consumption, enhances the concealment and reliability of the system, and has important value for improving the air defense security of important areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic flowchart of the method for adaptively generating UAV countermeasure signals with air-ground cooperation according to an embodiment of the present invention; Figure 2 is a comparison chart of the information transfer accuracy rates of different communication schemes under different communication loads; Figure 3 is a comparison chart of the optimized performance. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0020] Figure 1 is a schematic flowchart of the method for adaptively generating UAV countermeasure signals with air-ground cooperation according to an embodiment of the present invention, as Figure 1 shown, the method includes: Obtain the communication signal of the target UAV, extract the signal characteristic parameters and analyze to obtain the type of communication protocol; The signal feature parameters are encoded using a two - level coding structure, where the first level represents the type of interference strategy and the second level represents the specific parameters; a genetic algorithm with a composite fitness function is used to optimize the encoded parameters, and the composite fitness function is used to evaluate the interference effect, power consumption efficiency, and anti - detection ability; according to the type of communication protocol, the corresponding basic template is selected from the interference waveform template library, and the optimized parameters are used for parameter matching and adjustment to generate an interference signal template; The interference signal template is sent to the airborne counter - measure device and the ground counter - measure device; Execute resource optimization allocation for air - ground cooperation, including: constructing an air - ground cooperation strategy based on a game - theory model; using a reinforcement learning algorithm to perform dynamic resource allocation according to the air - ground cooperation strategy, and the reinforcement learning algorithm performs real - time optimization allocation of power resources and computing resources based on device status and target characteristics; According to the resource optimization allocation result, control the airborne and ground counter - measure devices to generate interference signals and transmit them to the target UAV.
[0021] In an alternative embodiment, the signal feature parameters are encoded using a two - level coding structure, where the first level represents the type of interference strategy and the second level represents the specific parameters, including: The type of interference strategy is encoded at the first level using a 4 - bit binary code to obtain a strategy code, and the types of interference strategies include a noise interference strategy, a deception interference strategy, a blocking interference strategy, and a combined interference strategy; The specific parameters are encoded at the second level using a variable - length coding structure to obtain a parameter code. The specific parameters include frequency - domain parameters, time - domain parameters, and power parameters. The frequency - domain parameters include the interference bandwidth ratio, frequency offset, and power spectrum shape factor. The time - domain parameters include the interference duration, time delay, and repetition period. The power parameters include the interference power, power distribution coefficient, and waveform factor; the coding bit - width of the parameter code is optimized based on parameter importance to obtain an optimized code, and the coding bit - width is determined according to the parameter value range and quantization accuracy requirements; Exemplarily, the type of interference strategy is encoded at the first level using a 4 - bit binary code. The types of interference strategies specifically include a noise interference strategy, a deception interference strategy, a blocking interference strategy, and a combined interference strategy. The encodings for each type are as follows: Noise interference strategy: 0001; Deception interference strategy: 0010; Blocking interference strategy: 0100; Combined interference strategy: 1000; When it is necessary to represent a combined interference strategy, for example, a combined strategy of using noise interference and deception interference at the same time, the encoding is: 1001.
[0022] The specific parameters are secondarily encoded using a variable-length coding structure to form parameter codes. The specific parameters include frequency-domain parameters, time-domain parameters, and power parameters.
[0023] For each parameter, its coding bit width determines the numerical range to be represented according to the value range of the parameter, then determines the minimum resolution according to the required quantization accuracy requirement, and finally calculates the minimum number of binary digits required to meet these two conditions.
[0024] For frequency-domain parameters, it includes the interference bandwidth ratio, frequency offset, and power spectral shape factor. The interference bandwidth ratio represents the ratio of the interference signal bandwidth to the target signal bandwidth, with a value range of 0.1 to 10, a quantization accuracy of 0.1, and requires 7 binary digits to represent (can represent the range of 0 to 12.7). The frequency offset represents the offset amount between the center frequency of the interference signal and the center frequency of the target signal, with a value range of -50 MHz to +50 MHz, a quantization accuracy of 1 MHz, and requires 7 binary digits to represent (can represent the range of -64 MHz to +63 MHz). The power spectral shape factor is used to describe the spectral distribution characteristics of the interference signal, with a value range of 0 to 1, a quantization accuracy of 0.01, and requires 7 binary digits to represent (can represent the range of 0 to 1.27).
[0025] For time-domain parameters, it includes the interference duration, time delay, and repetition period. The interference duration represents the duration of a single interference signal, with a value range of 1 ms to 1000 ms, a quantization accuracy of 1 ms, and requires 10 binary digits to represent (can represent the range of 0 to 1023 ms). The time delay represents the delay time of the interference signal relative to the target signal, with a value range of 0 to 100 ms, a quantization accuracy of 0.2 ms, and requires 9 binary digits to represent (can represent the range of 0 to 102.4 ms). The repetition period represents the time interval at which the interference signal repeats, with a value range of 10 ms to 10000 ms, a quantization accuracy of 10 ms, and requires 10 binary digits to represent (can represent the range of 0 to 10230 ms).
[0026] For power parameters, it includes the interference power, power distribution coefficient, and waveform factor. The interference power represents the transmission power of the interference signal, with a value range of 0 to 50 dBm, a quantization accuracy of 0.2 dBm, and requires 8 binary digits to represent (can represent the range of 0 to 51 dBm). The power distribution coefficient represents the power distribution ratio between channels in multi-channel interference, with a value range of 0 to 1, a quantization accuracy of 0.01, and requires 7 binary digits to represent (can represent the range of 0 to 1.27). The waveform factor is used to describe the waveform characteristics of the interference signal, with a value range of 0.5 to 5, a quantization accuracy of 0.1, and requires 6 binary digits to represent (can represent the range of 0 to 6.3).
[0027] Optimize the encoding bit width of parameter encoding based on parameter importance. The parameter importance is sorted from high to low as follows: interference power, interference bandwidth ratio, frequency offset, interference duration, time delay, power spectrum shape factor, repetition period, power distribution coefficient, waveform factor. Keep the original bit width for parameters with high importance and appropriately reduce the bit width for parameters with lower importance. The adjusted bit widths are as follows: Interference power: Keep 8 bits; interference bandwidth ratio: Keep 7 bits; frequency offset: Keep 7 bits; interference duration: Adjust to 8 bits (precision adjusted to about 4 ms); time delay: Adjust to 8 bits (precision adjusted to about 0.4 ms); power spectrum shape factor: Adjust to 6 bits (precision adjusted to about 0.02); repetition period: Adjust to 8 bits (precision adjusted to about 40 ms); power distribution coefficient: Adjust to 5 bits (precision adjusted to about 0.03); waveform factor: Adjust to 5 bits (precision adjusted to about 0.15); Through the above adjustments, the total bit width of parameter encoding is reduced from 71 bits to 62 bits, improving the encoding efficiency.
[0028] Perform adaptive mapping on the strategy encoding and optimized encoding according to the target characteristics to obtain the final characteristic parameter encoding. The adaptive mapping process selects a suitable parameter subset based on the target characteristics and makes adaptive adjustments to the interference requirements in different scenarios. For example, for low-altitude and slow-speed targets, the parameter subset selected by the adaptive mapping includes: interference power, interference bandwidth ratio, frequency offset, and interference duration. The corresponding parameter encoding bit widths are 8 bits, 7 bits, 7 bits, and 8 bits respectively, totaling 30 bits. The strategy encoding remains 4 bits unchanged, and the final characteristic parameter encoding is 34 bits. For high-speed and maneuvering targets, the parameter subset selected by the adaptive mapping includes: interference power, interference bandwidth ratio, frequency offset, interference duration, time delay, and repetition period. The corresponding parameter encoding bit widths are 8 bits, 7 bits, 7 bits, 8 bits, 8 bits, and 8 bits respectively, totaling 46 bits. The strategy encoding remains 4 bits unchanged, and the final characteristic parameter encoding is 50 bits. For targets in complex electromagnetic environments, the parameter subset selected by the adaptive mapping includes: interference power, interference bandwidth ratio, frequency offset, interference duration, time delay, power spectrum shape factor, repetition period, power distribution coefficient, and waveform factor, i.e., all parameters, totaling 62 bits. The strategy encoding remains 4 bits unchanged, and the final characteristic parameter encoding is 66 bits.
[0029] For example, for an X-band radar target, a noise jamming strategy is required with a jamming power of 30 dBm, a jamming bandwidth ratio of 2.5, a frequency offset of -15 MHz, a jamming duration of 200 ms, a time delay of 20 ms, a power spectral shape factor of 0.8, a repetition period of 500 ms, a power distribution coefficient of 0.6, and a waveform factor of 2.0. According to the above parameter values, the strategy code is "0001", and the parameter code is "00111100 0011001 1110001 00110010 00110010 01100101100100 10011 01101". After adaptive mapping, the final characteristic parameter code obtained is "0001 001111000011001 1110001 00110010 00110010 011001 01100100 10011 01101".
[0030] The two-level coding structure and adaptive mapping method proposed by the present invention achieve an efficient digital representation of jamming strategies and parameters; by combining the first-level coding of strategy types and the variable-length second-level coding of parameters, it solves the problems of limited expression ability, coding redundancy, and poor adaptability of traditional coding methods; innovatively optimizes the coding bit width based on parameter importance, reduces the overall coding length while ensuring the accuracy of key parameters, and effectively improves the coding efficiency; the adaptive mapping mechanism can dynamically select a suitable parameter subset according to different target characteristics, making the coding structure flexibly adapt to diverse jamming scenarios.
[0031] In an alternative embodiment, a genetic algorithm with a composite fitness function is used to optimize the encoded parameters. The composite fitness function for evaluating jamming effects, power consumption efficiency, and anti-detection includes: Construct a composite fitness function, which includes: a jamming effect evaluation term for evaluating signal-to-jamming ratio, bit error rate, and jamming coverage; a power consumption efficiency evaluation term for evaluating power efficiency, time efficiency, and energy utilization rate, where the energy utilization rate is determined by the time-domain integral ratio of the effective jamming power to the total jamming power; an anti-detection evaluation term for evaluating detection probability, exposure time, and low probability of intercept characteristics, where the low probability of intercept characteristics is determined by the jamming bandwidth ratio, power fluctuation degree, and transmission pattern characteristics; The composite fitness function uses dynamic weight coefficients to perform weighted combination on the evaluation terms. The dynamic weight coefficients calculate the weight adjustment amount according to the deviation between the target performance and the actual performance. The target performance includes the desired jamming effect, power consumption efficiency, and anti-detection index, and the actual performance includes the real-time measured jamming effect, power consumption efficiency, and anti-detection index; Iteratively optimize the encoding of the strategy and the optimized encoding using a genetic algorithm to obtain an optimal solution, and output the optimized interference strategy and parameter configuration; the genetic algorithm uses an adaptive mutation rate, and the adaptive mutation rate decreases as the fitness value increases.
[0032] Exemplarily, in the interference effect evaluation item, the signal-to-interference ratio (SIR) evaluation is calculated using the ratio of the interference signal power to the useful signal power at the target receiver. For example: divide the measured interference power at the target receiver (such as -45 dBm) by the useful signal power (such as -75 dBm) to obtain an SIR of 30 dB. The bit error rate (BER) evaluation is obtained by monitoring the BER performance of the target communication system. For example, under the action of interference, the BER of the target system increases from 10 -5 to 10 -2 , indicating a significant interference effect. The interference coverage range is verified through actual measurement, and the area of the region where effective interference can be achieved at different distances (such as 100 meters, 200 meters, 500 meters) is recorded.
[0033] In the power consumption efficiency evaluation item, the power efficiency is calculated as the ratio of the effective interference power to the input power. For example: if the input power is 100 W and the measured effective interference power is 85 W, then the power efficiency is 85%. The time efficiency is determined by the ratio of the interference duration to the total task time. For example, if the effective interference time is 1.8 hours in a 2-hour task, then the time efficiency is 90%. The energy utilization rate is calculated by integrating the ratio of the effective interference power to the total interference power over time. Specifically, the system samples the interference power every 10 milliseconds and records the interference power data within 24 hours. Assuming the total transmit power is 50 W and the cumulative effective interference power is 42 W·hours, then the energy utilization rate is 84%.
[0034] In the anti-detection evaluation items, the detection probability is obtained through Monte Carlo simulation. In 1000 simulations, the enemy detection equipment successfully detected the jammer 150 times, so the detection probability is 15%. The exposure time refers to the average time when the jammer is detected by the enemy equipment. The test results show that the average exposure time is 3.5 seconds. The low intercept characteristic is determined by three sub-indicators: the interference bandwidth ratio (the ratio of the interference bandwidth to the target system bandwidth. For example, if the interference bandwidth is 5 MHz and the target system bandwidth is 20 MHz, the interference bandwidth ratio is 0.25); the power fluctuation degree (the ratio of the standard deviation of the interference power to the average power. For example, if the standard deviation is 2 W and the average power is 40 W, the power fluctuation degree is 5%); the characteristics of the radiation pattern include: the main lobe direction (the direction of the maximum radiation power of the antenna); the main lobe width (the angular range when the radiation power drops to half of the maximum value (-3 dB). For example, the main lobe width is 30° (azimuth) × 20° (elevation)); the side lobe level (the ratio of the power in the secondary radiation direction to the maximum power of the main lobe, usually expressed in dB. For example, the maximum side lobe level is -25 dB); the front-to-back ratio (the ratio of the power in the main lobe direction to the power in its opposite direction. For example, the front-to-back ratio is 18 dB); the null depth (the ratio of the minimum power point to the maximum power in the radiation pattern. For example, the null depth is -40 dB); the directivity coefficient (a dimensionless coefficient characterizing the concentrated radiation ability of the antenna. For example, the directivity coefficient is 12 dB). In the evaluation of the low intercept characteristic, an ideal radiation pattern should have a narrow main lobe, low side lobes, and a deep null. For example, a narrow beam directional antenna (main lobe width 10°, side lobe level -30 dB) has better low intercept characteristics than a wide beam antenna (main lobe width 60°, side lobe level -15 dB).
[0035] The composite fitness function uses dynamic weight coefficients to perform weighted combination on each evaluation item. The initial values of the weight coefficients are set as 0.5 for interference effect, 0.3 for power consumption efficiency, and 0.2 for anti-detection. The system calculates the deviation between the target performance and the actual performance every 1 minute and adjusts the weight coefficients. For example, when the actual interference effect (signal-to-interference ratio 25 dB) is lower than the target value (signal-to-interference ratio 30 dB), the system increases the weight of the interference effect from 0.5 to 0.6 and correspondingly reduces the other weights (the weight of power consumption efficiency is reduced to 0.25, and the weight of anti-detection is reduced to 0.15). The amount of weight adjustment is proportional to the performance deviation, and the maximum adjustment amplitude is 20% of the initial weight.
[0036] The optimization process of the genetic algorithm is as follows: Step 1: Initialize the population. Randomly generate 50 individuals, each of which contains the encoded interference strategy and parameters. The encoding of the interference strategy includes the interference mode (e.g., 0 represents noise interference, 1 represents drag interference, 2 represents repetitive interference), the time pattern (e.g., 0 represents continuous, 1 represents periodic, 2 represents random), and the frequency pattern (e.g., 0 represents fixed frequency, 1 represents frequency hopping, 2 represents frequency sweeping). The encoding of the parameters includes the interference power (range 20 - 100W, accuracy 0.1W), the interference frequency (range 2 - 6GHz, accuracy 1MHz), the frequency hopping rate (range 100 - 1000 hops / second, accuracy 10 hops / second), etc.
[0037] Step 2: Evaluate the fitness. Calculate the composite fitness value for each individual. For example, the interference effect score of individual A is 0.85, the power consumption efficiency score is 0.75, and the anti - detection score is 0.60. Using the current weight coefficients (0.6, 0.25, 0.15) to calculate, the composite fitness value is 0.78.
[0038] Step 3: Selection operation. Use the roulette wheel method to select individuals. Individuals with higher fitness have a greater probability of being selected. For example, the individual with the highest fitness (0.88) has a selection probability of 8.8%, while the individual with the lowest fitness (0.35) has a selection probability of only 3.5%.
[0039] Step 4: Crossover operation. Randomly select two parent individuals and perform crossover with an 80% probability. The crossover point is randomly selected. For example, the encoding of parent A is "101|10110", and the encoding of parent B is "010|01001". If the crossover point is after the 3rd bit, then the offspring "10101001" and "01010110" are generated.
[0040] Step 5: Mutation operation. For each gene position of the individual, perform mutation with a mutation rate probability. An adaptive mutation rate is used, and the initial mutation rate is 5%. When the highest fitness of the population reaches 0.8, the mutation rate drops to 3%; when the fitness reaches 0.9, the mutation rate further drops to 1% to ensure convergence stability.
[0041] Step 6: Iterative optimization. Repeat steps 2 to 5 until the termination condition is met: reaching the preset 200 - generation iteration number, or the improvement of the optimal fitness in 20 consecutive generations is less than 0.1%.
[0042] Finally, output the optimal solution. Select the individual with the highest fitness from the final population and decode to obtain the optimized interference strategy and parameter configuration. Specifically: the interference strategy is "repetitive interference (strategy code: 2) + periodic time pattern (code: 1) + frequency hopping pattern (code: 1)"; the key parameter configuration includes an interference power of 75.3 W, a center frequency of 3.5 GHz, a frequency coverage range of 3.2 - 3.8 GHz, a frequency hopping rate of 800 hops per second, a duty cycle of 65%, and a pattern configuration with a main lobe width of 15° and a side lobe level of -28 dB.
[0043] The interference waveform template library contains basic interference templates designed for various mainstream communication protocols, such as WiFi interference templates, Bluetooth interference templates, 4G / 5G interference templates, satellite communication interference templates, etc. First, determine the type of the target communication protocol through the signal recognition module, such as identifying a WiFi signal of the IEEE 802.11ac protocol. Subsequently, retrieve the basic WiFi interference template from the template library. This template presets the basic interference parameters for WiFi signals, including spectral characteristics (center frequency in the 5.2 GHz band, bandwidth 80 MHz), time-domain characteristics (periodic frame structure, frame length of 1 ms), and modulation characteristics (quadrature amplitude modulation format, symbol rate). Match and adjust the parameters optimized by the genetic algorithm with the basic template. For example, apply the optimized interference power of 75.3 W to the template; combine the frequency hopping pattern parameters with the basic template to make the interference signal hop at a rate of 800 hops per second within the 5.15 - 5.35 GHz band; apply the optimized 65% duty cycle parameter to the time-domain structure to set the ratio of the duration to the silent time of the interference signal; apply the pattern configuration (main lobe width of 15°, side lobe level of -28 dB) to the antenna parameters to achieve directional interference. The system further optimizes the template according to the evaluation results of the composite fitness function. When detecting that the WiFi signal uses the dynamic frequency selection mechanism, automatically adjust the interference bandwidth ratio and power distribution of the interference template to enhance the interference intensity on the key channels; when the target system enables forward error correction coding, the system will correspondingly adjust the time characteristics of the interference signal so that the interference duration exceeds the coding interleaving period and reduces the error correction ability. The finally generated interference signal template not only contains waveform parameters but also emission control parameters, forming a complete interference strategy. The system saves this template and transmits it to the signal generation module for real-time generation of optimized interference signals for the target communication system, achieving an efficient and precise interference effect.
[0044] The present invention realizes the intelligent configuration of interference system parameters by constructing a composite fitness function that combines interference effect, power consumption efficiency, and anti-detection performance, and integrating a dynamic weight adjustment mechanism and genetic algorithm optimization; it overcomes the limitations of single-index optimization and manual parameter adjustment in the traditional interference system parameter configuration, and can find the optimal configuration scheme under multi-objective constraints; the dynamic weight coefficient mechanism can adaptively adjust the evaluation focus according to the deviation between the actual performance and the target performance, while the adaptive mutation rate design effectively balances the global search ability and local convergence speed of the algorithm, significantly improving the overall effectiveness and environmental adaptability of the interference system.
[0045] In an alternative embodiment, the determination of the dynamic weight coefficient includes: Establish a performance index system, including the signal-to-interference ratio, bit error rate, and interference coverage in the interference effect index set, the power efficiency, time efficiency, and energy utilization rate in the power consumption efficiency index set, and the detection probability, exposure time, and low intercept characteristics in the anti-detection performance index set, and determine the ideal target value and acceptable range for each index; Use the normalized deviation calculation function δ ij =(T ij -A ij ) / N ij Calculate the performance deviation of each index, where T ij is the target performance value, A ij is the actual measured value, and N ij is the normalization factor; sum the weighted deviations of the same type of indexes to obtain the comprehensive deviation Δ i ; Use the non-linear deviation response function to calculate the weight adjustment amount: ; where ΔW i is the adjustment amount of the weight corresponding to the i-th type of index, α is the adjustment amplitude coefficient, β is the non-linear exponent used to control the curvature of the response function, γ is the sensitivity coefficient used to control the response degree to small deviations, τ min is the minimum response threshold used to filter out noise, τ max is the maximum response threshold used to limit abnormal situations; sign(Δ i ) is the sign function of the deviation, which determines the direction of weight adjustment; Smoothly update and normalize the weights to ensure that the sum of all weights is 1, and the weight values are restricted by upper and lower limits.
[0046] Exemplarily, a complete performance index system is established, including three categories of indexes: interference effect index set, power consumption efficiency index set, and anti-detection index set. The interference effect index set includes signal-to-interference ratio, bit error rate, and interference coverage range; the power consumption efficiency index set includes power efficiency, time efficiency, and energy utilization rate; the anti-detection index set includes detection probability, exposure time, and low intercept feature.
[0047] For each specific index, an ideal target value and an acceptable range are set. For example, for the signal-to-interference ratio index, the ideal target value is set to 6 dB, and the acceptable range is 3 dB to 10 dB; for the bit error rate index, the ideal target value is set to 0.25, and the acceptable range is 0.15 to 0.35; the ideal value of the interference coverage range is set to 95%, and the acceptable range is 85% to 100%. The ideal value of the power efficiency index is set to 0.8, and the acceptable range is 0.6 to 0.9; the ideal value of the time efficiency is set to 0.75, and the acceptable range is 0.6 to 0.85; the ideal value of the energy utilization rate is set to 0.7, and the acceptable range is 0.5 to 0.8. The ideal value of the detection probability is set to 0.15, and the acceptable range is 0.1 to 0.25; the ideal value of the exposure time is set to 30 seconds, and the acceptable range is 20 to 60 seconds; the ideal value of the low intercept feature is set to 0.85, and the acceptable range is 0.7 to 0.95.
[0048] The performance deviation of each index is calculated through a normalization deviation calculation function. This function standardizes the difference between the target performance value and the actual measurement value through a normalization factor. The normalization deviations of the same category of indexes are weighted and summed to obtain the comprehensive deviation of each category of indexes. Assume that the internal weights of the interference effect index set are 0.4 for the signal-to-interference ratio, 0.35 for the bit error rate, and 0.25 for the interference coverage range. If the normalization deviations of the three are 0.57, -0.5, and 0.2 respectively, then the comprehensive deviation of the interference effect index is 0.4×0.57 + 0.35×(-0.5) + 0.25×0.2 = 0.108. Similarly, calculate the comprehensive deviations of the power consumption efficiency index and the anti-detection index.
[0049] The weight adjustment amount is calculated using a non-linear deviation response function. This response function contains multiple parameters: the adjustment amplitude coefficient α is usually set to 0.15, which is used to control the maximum amplitude of each adjustment; the non-linear exponent β is set to 2.0, which controls the curvature of the response function; the sensitivity coefficient γ is set to 0.5, which controls the response degree to small deviations; the minimum response threshold τ min is set to 0.05, which is used to filter out noise; the maximum response threshold τ max is set to 0.85, which is used to limit the influence of abnormal situations. sign(Δ i ) is used to extract the sign of a real number. For any non-zero real number, the sign function returns the positive or negative sign of the number; for zero, it returns zero. In the weight adjustment formula, sign(Δ iThe role of () is to determine the direction of weight adjustment: when the comprehensive deviation Δ i is positive (indicating that the actual performance is lower than the target performance), sign(Δ i ) returns +1, and the weight will increase; when the comprehensive deviation Δ i is negative (indicating that the actual performance exceeds the target performance), sign(Δ i ) returns -1, and the weight will decrease; when the comprehensive deviation Δ i equals zero (indicating that the actual performance exactly equals the target performance), sign(Δ i ) returns 0, and the weight remains unchanged.
[0050] Finally, perform weighted smoothing update and normalization, and at the same time check whether the updated weight exceeds the preset upper and lower limits. If it exceeds, adjust it to the boundary value. For example, if the upper limit of the interference effect index weight is 0.45 and the lower limit is 0.25, then 0.395 is within the reasonable range and does not need to be adjusted.
[0051] During the actual operation process, the above-mentioned weight update process is executed every predetermined time period (such as 5 minutes), so that the system can dynamically adjust the weights of each index according to the performance deviation, thereby optimizing the overall performance of the system in different working environments.
[0052] The present invention realizes the precision and adaptability of the interference system performance evaluation by establishing a comprehensive performance index system and an intelligent dynamic weight adjustment mechanism; adopts normalized deviation calculation and a non-linear response function, enabling the weight adjustment to accurately reflect the difference between the actual performance and the target performance; this method overcomes the limitations of the traditional fixed-weight evaluation method, can automatically optimize the weight ratio of each index according to the changes in the working environment, takes into account the power consumption efficiency and anti-detection performance while ensuring the interference effect, and significantly improves the adaptability, working efficiency and overall combat effectiveness of the interference system in complex electromagnetic environments.
[0053] In an alternative embodiment, constructing an air-ground cooperation strategy based on a game theory model includes: Obtain the state vectors of the air countermeasure equipment and the ground countermeasure equipment, and construct an air-ground countermeasure game theory model. The state vectors include power, position, coverage range and energy state; Take the air countermeasure equipment and the ground countermeasure equipment as game participants, and construct a game strategy space including power selection, interference direction and resource allocation according to the state vectors. The game strategy space is used to constrain the strategy selection range of the game participants; Construct a multi-objective utility function, where the multi-objective utility function includes an interference effect evaluation function obtained by weighted combination of signal-to-interference ratio evaluation, coverage effect evaluation, and directivity evaluation, and a resource consumption evaluation function obtained by weighted combination of power consumption, communication overhead, and maneuvering cost. Combine the multi-objective utility function with a resource penalty factor to construct a game payoff matrix, and solve for the optimal response strategy under the Nash equilibrium based on the game payoff matrix; Calculate the air-ground cooperation superiority degree based on the optimal response strategy, and construct a strategy selection probability distribution, where the air-ground cooperation superiority degree is determined by the ratio of the sum of the joint countermeasure benefits to the individual countermeasure benefits; Establish a state value function according to the strategy selection probability distribution. The state value function calculates the long-term expected payoff of the game payoff matrix based on a discount factor, and obtains the air-ground cooperation strategy according to the gradient of the strategy selection probability distribution with respect to the state value function.
[0054] Exemplarily, the state vector of the airborne countermeasure device includes: transmission power (adjustable from 20W to 200W), geographical location coordinates (latitude, longitude, and altitude, such as N30°11′24″, E120°13′49″, altitude 500 meters), coverage range (a circular area with a radius of 1000 meters), and energy state (remaining battery power 80%). The state vector of the ground countermeasure device includes: transmission power (adjustable from 50W to 300W), geographical location coordinates (such as N30°11′26″, E120°13′52″, altitude 5 meters), coverage range (a sector area with a radius of 500 meters), and energy state (remaining battery power 65%). The state acquisition is updated in real time through a wireless communication link, and the update period is 100 milliseconds.
[0055] Construct an air-ground countermeasure game theory model, with the airborne countermeasure device as game participant A and the ground countermeasure device as game participant B. Construct a game strategy space according to the state vector: power selection strategy set (for airborne device: {40W, 80W, 120W, 160W}, for ground device: {75W, 150W, 225W, 300W}), interference direction strategy set (for airborne device: {omnidirectional, 30° sector, 60° sector, 90° sector}, for ground device: {45° sector, 90° sector, 180° sector, 360°}), resource allocation strategy set (for airborne device: {20%, 40%, 60%, 80%}, for ground device: {25%, 50%, 75%, 100%}). These strategies are combined to form 4×4×4 = 64 airborne countermeasure strategies and 4×4×4 = 64 ground countermeasure strategies, for a total of 64×64 = 4096 joint strategy combinations.
[0056] When constructing the multi-objective utility function, first calculate the interference effect evaluation function. For the signal-to-interference-plus-noise ratio (SINR) evaluation, three parameters, namely signal strength, spectral coverage width, and interference incident angle, are used to calculate the score. For example, when the airborne device selects a power of 120 W, a 60° sector, and a 60% resource allocation, the average SINR in the target area is calculated to be -12 dB, corresponding to a score of 85 points. The coverage effect evaluation considers the coverage ratio of the target area, and under the same conditions, the coverage score is 78 points. The directivity evaluation considers the energy concentration, with a score of 92 points. The total score of the interference effect is calculated as a weighted combination (weights are 0.5, 0.3, and 0.2 respectively) of the three, which is 84.9 points. The resource consumption evaluation function is evaluated by a weighted combination of power consumption, communication overhead, and maneuvering cost. Under the above conditions, the power consumption score is 65 points (high consumption), the communication overhead score is 88 points (low overhead), and the maneuvering cost score is 72 points (medium consumption). The weights of the three are set to 0.6, 0.2, and 0.2 respectively, and the total score of resource consumption is 71.4 points.
[0057] Combining the interference effect evaluation and the resource consumption evaluation with a resource penalty factor (set to 0.7), a game payoff matrix is constructed. The calculation formula is: Payoff value = Interference effect score - Resource penalty factor × Resource consumption score. Under the above conditions, the payoff value of the airborne device acting alone is 84.9 - 0.7×71.4 = 34.92. By iteratively calculating the payoff values of all strategy combinations, a complete 4096×2 payoff matrix is filled.
[0058] Based on the constructed 4096×2 game payoff matrix, the optimal response iteration method is used to solve the Nash equilibrium: Randomly initialize the strategy probability distributions for the airborne device and the ground device, and generate initial probability vectors on their respective 64 strategies; In each iteration, two steps are performed: Fix the current strategy of the ground device and calculate the optimal response strategy of the airborne device, that is, under the existing ground device strategy, calculate the expected payoff value of each possible strategy of the airborne device, and select the strategy with the maximum payoff value as the updated strategy of the airborne device; Fix the updated strategy of the airborne device and calculate the optimal response strategy of the ground device. The convergence condition is set as: The change in strategy between two consecutive iterations is less than the threshold of 0.01 or the maximum number of iterations is reached. In this example, after 23 iterations, the algorithm converges, and the Nash equilibrium strategy combination is obtained: The airborne device selects (80 W power, 60° sector, 40% resource allocation), and the ground device selects (150 W power, 90° sector, 50% resource allocation), with corresponding payoff values of 42.5 and 39.8 respectively. This strategy combination represents the optimal response strategies of both devices in the comprehensive game considering interference effect and resource consumption.
[0059] Calculate the air-ground cooperation advantage degree based on the optimal response strategy. The specific method is to obtain the combined countermeasure benefit (such as 78.6 points) and the individual countermeasure benefits (42.5 points for air equipment and 39.8 points for ground equipment) under Nash equilibrium. The air-ground cooperation advantage degree is obtained by calculating the ratio of the combined benefit to the sum of individual benefits, which is 78.6 / (42.5 + 39.8) = 0.96. This value indicates that the cooperative action has obvious advantages over individual actions. Subsequently, construct a probability distribution of strategy selection based on the air-ground cooperation advantage degree, and use the softmax method to convert the expected benefits of each strategy into selection probabilities. The calculation results show that the selection probability of the optimal strategy (80W power, 60° sector, 40% resource allocation) for air equipment is 0.35, and the probability of the sub-optimal strategy (120W power, 60° sector, 60% resource allocation) is 0.28; the selection probability of the optimal strategy (150W power, 90° sector, 50% resource allocation) for ground equipment is 0.33.
[0060] Directly establish a state value function according to the probability distribution of strategy selection. The method is to multiply the selection probability of each possible strategy by the payoff value corresponding to that strategy and then sum them to obtain the expected value of the current state. For example, the state value calculation for air equipment is: 0.35×42.5 (optimal strategy payoff value) + 0.28×40.2 (sub-optimal strategy payoff value) +... = 41.3. When establishing the state value function, set the discount factor to 0.85, which is used to calculate the long-term expected benefit of the game payoff matrix, that is, the current benefit plus the present value of future benefits. According to the gradient of the state value function with respect to the probability distribution of strategy selection, the system uses the policy gradient method for optimization. Specifically, it increases the selection probability of high-benefit strategies and decreases the selection probability of low-benefit strategies. After 200 iterations of optimization, the final air-ground cooperation strategy is obtained: air equipment selects (80W power, 60° sector, 40% resource allocation) with a 40% probability, (120W power, 60° sector, 60% resource allocation) with a 35% probability, and an alternative strategy with a 25% probability; ground equipment selects (150W power, 90° sector, 50% resource allocation) with a 45% probability, (225W power, 90° sector, 75% resource allocation) with a 30% probability, and an alternative strategy with a 25% probability.
[0061] The prior art mainly adopts fixed preset strategies or simple adaptive methods for countermeasure decision-making, making decisions based on local information respectively, and lacking a systematic cooperation mechanism. The present invention models airborne and ground countermeasure devices as game participants, constructs a multi-dimensional strategy space including power selection, interference direction, and resource allocation; designs a multi-objective utility function combining interference effect and resource consumption; quantifies the cooperation benefit by calculating the air-ground cooperation superiority degree; and optimizes the cooperation decision using the state value function and the policy gradient method. This game theory-based method enables air-ground countermeasure decision-making to be based on a strict mathematical framework, realizing the guidance of theory to practice. The present invention greatly improves the cooperation of air-ground countermeasures, enabling both devices to make optimal responses based on the Nash equilibrium principle, avoiding decision conflicts in traditional methods; significantly enhancing the countermeasure effect, achieving higher interference efficiency under the same resource constraints through the optimization of the multi-objective utility function; and improving the resource utilization efficiency, reducing unnecessary resource consumption while ensuring the countermeasure effect through game equilibrium calculation and air-ground cooperation superiority degree evaluation.
[0062] In an alternative embodiment, the reinforcement learning algorithm performs real-time optimal allocation of power resources and computing resources based on device status and target characteristics, including: Obtain the device status and target characteristics, construct a resource status tensor including platform types and resource types, where the platform types include airborne platforms and ground platforms, and the resource types include power resources and computing resources; Construct a variational autoencoder to encode the resource status tensor to generate latent variables, and decode the latent variables into communication messages between air-ground platforms; Based on the resource status tensor and the communication messages between air-ground platforms, design a two-way deep Q-network with value stream and advantage stream, where the value stream evaluates the state value of resource allocation, the advantage stream evaluates the action advantage of resource allocation, and combine the state value and the action advantage to obtain the Q value of the resource allocation strategy; train the policy network using the proximal policy optimization algorithm, where the policy network includes an actor network and a critic network, the actor network outputs the mean and standard deviation of the resource allocation action distribution that meets the air-ground cooperation requirements according to the constraints of the air-ground cooperation strategy, and the critic network estimates the state value based on the current state and calculates the action advantage; Construct a hierarchical priority experience replay pool, calculate the sample priority based on the temporal difference error, sample the experience samples according to the sample priority; perform policy update using a meta-learning framework, and construct an air-ground consistency evaluation function to impose additional constraints on the training of the policy network; perform real-time optimal allocation of power resources and computing resources for airborne platforms and ground platforms according to the trained policy network.
[0063] Exemplarily, the device status and target feature information are obtained as the input data of the resource status tensor. The device status includes the current battery level, processor load rate, memory occupancy rate, signal strength, etc.; the target features include task priority, computational complexity, time constraint, etc. Taking a certain UAV formation system as an example, the aerial platform includes 5 UAVs of different models, and the ground platform includes 3 computing nodes, and their status information is collected respectively. For example, the battery level of UAV 1 is 85%, the CPU load is 40%, and it is executing a target tracking task with a priority of 3.
[0064] The resource status tensor contains two dimensions: platform type and resource type. The platform type is divided into aerial platform and ground platform; the resource type includes power resources and computing resources. For each aerial platform, record its available power percentage, computing unit occupancy rate, remaining storage space, etc.; for the ground platform, record its power supply status, computing load, network bandwidth occupancy rate, etc. Normalize each index to ensure that the numerical range is between 0 and 1, and form a resource status tensor with a dimension of [8, 10], where 8 represents the number of platforms, and 10 represents the resource status feature dimension of each platform.
[0065] The variational autoencoder consists of an encoder and a decoder. The encoder contains three fully connected networks. The number of nodes in the input layer is 80 (i.e., the dimension after flattening the resource status tensor), the number of nodes in the hidden layers are 64 and 32 respectively, and the output layer generates a 16-dimensional latent variable and its mean and variance. The decoder also uses three fully connected networks to map the latent variable back to the space of the original dimension and generate communication messages between the aerial and ground platforms. The communication messages contain information such as resource demand prediction, task priority adjustment suggestions, cooperation strategy parameters, etc. For example, when UAV 2 has insufficient computing resources, it will send a task offloading request to ground computing node 1, including information such as task type, data size, and expected completion time.
[0066] Design a dual - path deep Q - network based on resource state tensors and communication messages. This network consists of a shared feature extraction layer and separate value and advantage streams. The feature extraction layer is composed of two convolutional layers and one fully - connected layer. The number of filters in the convolutional layers is 16 and 32 respectively, the convolutional kernel size is 3×3, and the number of nodes in the fully - connected layer is 128. The value stream contains two fully - connected networks with 64 and 1 nodes respectively, which are used to evaluate the overall value of the current resource allocation state. The advantage stream also contains two fully - connected networks with 64 and the dimension of the action space (20 in this example, representing optional resource allocation combinations) respectively, which are used to evaluate the advantage of each action relative to the average level. Combine the output of the value stream and the output of the advantage stream to obtain the final Q - value, which represents the expected return of taking each resource allocation action in the current state. The dual - path deep Q - network generates Q - value estimates for all possible resource allocation actions in the current state through forward propagation calculations once every 100 milliseconds. These Q - values are stored in an action - value vector of size 20 and are used to guide the action selection and evaluation process of the policy network subsequently.
[0067] Train the resource allocation policy network using the Proximal Policy Optimization (PPO) algorithm. This policy network directly receives the output Q - values of the dual - path deep Q - network as additional inputs to optimize action selection. The policy network consists of an Actor Network and a Critic Network. The two share the first two feature extraction layers but have different output layers. The Actor Network contains three fully - connected layers with 128, 64, and 40 nodes respectively. Its input is the resource state tensor and the action Q - value vector provided by the dual - path deep Q - network, and its output is the mean vector (20 - dimensional) and standard deviation vector (20 - dimensional) of the action distribution, which jointly define the Gaussian policy distribution of resource allocation. Specifically, the Actor Network designs a fusion layer after its second hidden layer (the 64 - node layer), concatenates the 20 - dimensional Q - value vector of the dual - path deep Q - network with the current hidden - layer representation, and then performs feature fusion through an additional fully - connected layer (64 nodes) to ensure that the Q - value information can directly affect policy generation. The 40 - dimensional final output corresponds to two types of decisions, power and computing resource allocation, on 8 platforms (5 aerial + 3 ground), a total of 16 decision dimensions, plus two types of parameters, mean and variance, for a total of 40 output nodes. The Actor Network ensures that the output resource allocation actions meet the requirements of air - ground cooperation and have high long - term benefits by receiving the constraints of the aforementioned air - ground cooperation strategy (such as through additional terms in the loss function) and combining the action evaluation information provided by the Q - values. For example, ensure that the computing resource allocation ratio between the aerial platform and the ground platform is within a reasonable range (such as normalizing the allocation ratio using the SoftMax function), and the ground platform reserves at least 30% of the computing resources for critical task processing (by setting a minimum threshold).
[0068] The critic network also consists of three fully-connected layers with 128, 64, and 1 nodes respectively. Its input is also the resource state tensor and the Q-value vector of the dual-path deep Q-network, and the output is the estimated state value V, which is used to estimate the long-term reward of the current resource allocation state. In the implementation, the critic network designs a comparison mechanism to compare the state value V estimated by itself with the maximum Q-value estimated by the dual-path deep Q-network, and calculates a consistency loss term to encourage the consistency of the two value estimations, thereby improving the accuracy of evaluation. By comparing the state value estimation of the critic network with the actually obtained reward, the temporal difference error is calculated, and then the action advantage value is calculated. This calculation process refers to the advantage flow design idea of the dual-path deep Q-network, but uses the value estimation of the critic itself. The action advantage value is obtained by subtracting the state value function V from the action-state value function Q, indicating the degree of advantage of a specific action relative to the average value of the state. The PPO algorithm uses this advantage value (while referring to the advantage information provided by the dual-path Q-network) to guide the actor network to adjust the policy direction, and updates the policy parameters by maximizing a specific objective function. This objective function takes the smaller value between the product of the ratio of the new and old policy probabilities and the advantage value, and the product of the ratio after being clipped by the upper and lower limits and the advantage value. The ratio of the new and old policy probabilities refers to the ratio of the action probability under the new policy to the action probability under the old policy. The clipping parameter is set to 0.2 in this example to limit the step size of policy update and ensure training stability.
[0069] During system operation, the dual-path deep Q-network and the policy network form a complementary and collaborative relationship: after each calculation and update of the Q-network, the output Q-value is immediately used by the policy network to guide the next decision; and the actions generated by the policy network and their actual effects are used to update the parameters of the Q-network, forming a closed-loop optimization process.
[0070] A hierarchical priority experience replay pool is constructed to store and sample training data. The capacity of the replay pool is set to 10,000 records, and each record contains five parts: the current state, the executed action, the obtained reward, the next state, and the termination flag. The sample priority is calculated based on the temporal difference error. Samples with larger temporal difference errors indicate that they contain more valuable information and have higher priorities. In actual operation, for a sample where the system performance is significantly improved after a certain resource allocation, the temporal difference error can reach 0.85, which is much higher than the average level of 0.32, so it is given a higher sampling probability.
[0071] The meta - learning framework is adopted for policy update, and the network parameters are adjusted by the gradient descent method. In each training cycle, 128 records are sampled according to the priority from the experience replay pool. The loss functions of the actor network and the critic network are calculated, and the parameters are updated. Meanwhile, an air - ground consistency evaluation function is constructed to impose constraints on the policy network, ensuring that the resource allocation meets the collaborative requirements between the air and ground platforms, such as ensuring the balanced communication bandwidth allocation between the ground command center and the UAVs, and reasonable computing task allocation. After 1000 rounds of training, it is determined that the policy network converges, and the resource optimization allocation is realized according to the trained policy network. For example, when it is detected that UAV 3 has insufficient computing resources for target recognition tasks, the policy network will automatically offload 50% of its computing tasks to ground node 2, and at the same time adjust the power configuration of the UAV, and allocate the saved 25% power to the communication module to improve the data transmission efficiency.
[0072] Figure 2 It is a comparison chart of the information transfer accuracy of different communication schemes under different communication loads. This chart shows the comparison of the information transfer accuracy of three different communication schemes under different communication loads. The horizontal axis represents the air - ground platform communication load (Mbps), ranging from 0 to 60 Mbps; the vertical axis represents the information transfer accuracy (%), ranging from 50% to 99%. The performance of the three communication schemes is as follows: The present invention (VAE communication) is represented by circular markers: It maintains the highest transmission accuracy under all communication load conditions. Even under high - load (60 Mbps) conditions, it can still maintain an accuracy of 92.1%. The accuracy can reach 98.5% under low - load conditions, showing excellent communication performance and stability. Traditional direct communication is represented by diamond markers: The performance of this scheme drops rapidly as the load increases. The accuracy is 91.8% at low load, but drops sharply to 56.3% at high load (60 Mbps), indicating its obvious limitations under high communication pressure. Compressive - sensing - based is represented by square markers: The performance is between the former two, reaching 95.3% at low load and 72.3% at high load, showing a medium level of anti - interference ability and coding efficiency.
[0073] Figure 3 It is a comparison chart of optimized performance. This chart compares the training effects and convergence performances of three different reinforcement - learning optimization algorithms. The horizontal axis represents the number of training rounds (0 - 1200 rounds), and the vertical axis represents the resource allocation efficiency score (0 - 100 points). The present invention (dual - path Q - network + policy network) is represented by circular markers: It shows the fastest convergence speed and the highest final performance. The efficiency score reaches 68.3 after 400 rounds of training and finally converges to a high efficiency score of 97.8 after 1200 rounds. The single DQN network is represented by square markers: It has the slowest convergence speed. The single PPO algorithm is represented by diamond markers: The performance is between the former two, showing a medium level of air - ground collaborative advantage.
[0074] The prior art mainly adopts fixed - rule strategies or simple adaptive algorithms for resource allocation, lacking a unified coordination mechanism. The hybrid architecture integrating variational auto - encoder, dual - path deep Q - network and policy network in the present invention realizes the unified optimization and allocation of air - ground platform resources. The present invention designs a collaborative mechanism between the dual - path Q - network and the policy network, enabling the Q - value to directly participate in policy decision - making; introducing an air - ground consistency evaluation function, incorporating resource balance degree, task completion degree and collaborative efficiency into a unified evaluation framework. At the same time, through hierarchical priority experience replay and meta - learning framework, the adaptability and generalization of the algorithm to dynamic environments are enhanced. The present invention realizes the unified optimization of power resources and computing resources, avoiding resource allocation conflicts between platforms; the system response speed is significantly accelerated, and the resource strategy can be adjusted in real - time under complex and changeable environments; the task completion quality is significantly improved, especially being able to maintain high - efficiency execution under resource - constrained conditions; the overall energy efficiency ratio of the system is improved, reducing energy consumption while ensuring task performance.
[0075] In an alternative embodiment, the air - ground consistency evaluation function includes: The air - ground consistency evaluation function includes a resource balance degree evaluation item, a task completion degree evaluation item and a collaborative efficiency evaluation item. The resource balance degree evaluation item calculates the resource allocation difference between the air platform and the ground platform. The task completion degree evaluation item evaluates the task execution effect based on the platform resource utilization rate. The collaborative efficiency evaluation item is evaluated based on the communication quality and collaborative response time between the air - ground platforms; the compliance degree of the resource allocation action and the air - ground collaborative strategy is calculated according to the weighted combination of the resource balance degree evaluation item, the task completion degree evaluation item and the collaborative efficiency evaluation item.
[0076] Exemplarily, the resource balance degree evaluation item is used to calculate the difference in resource allocation between the aerial platform and the ground platform. The aerial platform set and the ground platform set can be set, and each platform includes four resource types: computing resources, storage resources, energy resources, and communication resources. For each resource type, calculate the variance of the resource allocation between the aerial platform and the ground platform. For example, the allocation ratio of computing resources on the aerial platform is [0.6, 0.7, 0.5], and on the ground platform is [0.8, 0.9, 0.75], then the difference in computing resource allocation is 0.037. Similarly, calculate the difference values of other resource types, and finally, sum up the weighted difference values of the four resource types to obtain the resource balance degree evaluation value. In practical applications, the smaller the resource allocation difference, the higher the balance degree evaluation value, indicating that the resource allocation is more reasonable. Suppose there are 3 aerial platforms and 4 ground platforms, and the weights of the four resource types are 0.3, 0.2, 0.3, and 0.2 respectively. The computing resource allocation rate of the aerial platform is [0.65, 0.72, 0.58], and that of the ground platform is [0.82, 0.85, 0.78, 0.80]; the storage resource allocation rate of the aerial platform is [0.70, 0.75, 0.68], and that of the ground platform is [0.60, 0.65, 0.62, 0.67]; the energy resource allocation rate of the aerial platform is [0.55, 0.60, 0.58], and that of the ground platform is [0.88, 0.92, 0.85, 0.90]; the communication resource allocation rate of the aerial platform is [0.80, 0.85, 0.82], and that of the ground platform is [0.75, 0.78, 0.72, 0.80]. Calculate the variances of each resource type to be 0.026, 0.005, 0.093, and 0.002 respectively. After weighting, the resource balance degree evaluation value is 0.035, indicating that the resource allocation is relatively balanced.
[0077] The task completion degree evaluation item evaluates the execution effect of tasks based on the platform resource utilization rate, collects the resource utilization rate data of all platforms, including CPU usage rate, memory usage rate, energy consumption rate, and bandwidth usage rate. For each task T, calculate its resource utilization rate indicators on each platform. By setting the target completion threshold, determine whether the task meets the completion conditions. For example, a certain task is executed on 5 platforms, and the resource utilization rate of each platform is [0.75, 0.82, 0.68, 0.79, 0.73]. If the set resource utilization rate threshold is 0.65, then the task meets the completion conditions on all platforms. Suppose there are 10 tasks in the system, and the average resource utilization rates of each task on different platforms are [0.78, 0.72, 0.80, 0.65, 0.81, 0.75, 0.69, 0.82, 0.77, 0.70] respectively. Set the completion threshold to 0.65, and the calculated task completion degree is 100%. If the threshold is increased to 0.70, the task completion degree drops to 90%. In addition, the task completion degree evaluation value can also be calculated through weighted tasks according to priorities. For example, the weight of critical tasks is 0.6, and the weight of ordinary tasks is 0.4.
[0078] The collaborative effectiveness evaluation item evaluates based on the communication quality and collaborative response time between air and ground platforms. The communication quality can be quantified by indicators such as data transmission success rate, signal strength, and bit error rate. The collaborative response time is determined by measuring the time interval from the instruction issuance to the execution completion. Monitor the communication link quality between each pair of air and ground platforms. For example, the data transmission success rate between the air platform A and the ground platform B is 95%, the signal strength is -65dBm, and the bit error rate is 0.002. Its communication quality score can be calculated as 0.85. At the same time, record the response time of each collaborative task. For example, the time from the instruction issuance to the completion of the collaborative reconnaissance task is 1.5 seconds. If the preset threshold is 2 seconds, the response time score is 0.75. Suppose the air-ground system contains 3 air platforms and 4 ground platforms, forming 12 communication links. The communication quality scores of each link are [0.88, 0.92, 0.85, 0.90, 0.83, 0.87, 0.91, 0.86, 0.89, 0.84, 0.93, 0.88], and the average communication quality score is 0.88. The collaborative task response time scores are [0.82, 0.78, 0.85, 0.80, 0.83], and the average response time score is 0.82. Set the communication quality weight to 0.6 and the response time weight to 0.4, and the calculated collaborative effectiveness evaluation value is 0.856.
[0079] By weighted combination of the above three evaluation indicators, calculate the compliance degree of the resource allocation action and the air-ground collaboration strategy. Set the resource balance degree weight to 0.3, the task completion degree weight to 0.4, and the collaborative effectiveness weight to 0.3. Then the overall evaluation function value is calculated as the weighted sum of each item.
[0080] In an actual case, assume that the evaluation value of resource balance is 0.85, the evaluation value of task completion is 0.92, and the evaluation value of collaborative efficiency is 0.88. Then the overall air-ground consistency evaluation function value is 0.3×0.85 + 0.4×0.92 + 0.3×0.88 = 0.887. This value indicates that the current resource allocation action has a high degree of compliance with the air-ground collaboration strategy, and the resource allocation is reasonable and effective.
[0081] During the application process of the evaluation function, the weights of each evaluation item can be dynamically adjusted according to specific scenarios. For example, in an environment with limited resources, the weight of resource balance can be increased; in a scenario with high requirements for task timeliness, the weight of task completion can be increased; in a complex collaborative scenario, the weight of collaborative efficiency can be increased. By reasonably setting the weight parameters, the evaluation function can better adapt to the requirements of different application scenarios and improve the overall performance of the air-ground collaboration system.
[0082] The present invention designs an air-ground consistency evaluation function, comprehensively considers three key dimensions of resource balance, task completion, and collaborative efficiency, and realizes the accurate evaluation of the compliance between resource allocation actions and air-ground collaboration strategies. This evaluation function can effectively coordinate the resource allocation differences between the air platform and the ground platform, optimize the task execution effect, improve the communication quality and collaborative response efficiency between platforms. By dynamically adjusting the weights of each evaluation item, the system can adapt to the requirements of different application scenarios, improve the resource utilization rate and task completion quality, and enhance the overall performance and adaptability of the air-ground collaboration system.
[0083] In the second aspect of the embodiment of the present invention, an adaptive generation system for unmanned aerial vehicle countermeasure signals for air-ground collaboration is provided, including: A first unit for acquiring the communication signal of the target unmanned aerial vehicle, extracting signal feature parameters, and analyzing to obtain the communication protocol type; A second unit for encoding the signal feature parameters using a two-level coding structure, where the first level represents the interference strategy type and the second level represents specific parameters; optimizing the encoded parameters using a genetic algorithm with a composite fitness function, and the composite fitness function is used to evaluate interference effect, power consumption efficiency, and anti-detection performance; selecting a corresponding basic template from a preset interference waveform template library according to the communication protocol type, and performing parameter matching and adjustment using the optimized parameters to generate an interference signal template; A third unit for sending the interference signal template to the air countermeasure device and the ground countermeasure device; A fourth unit for performing resource optimization allocation for air-ground collaboration, including: constructing an air-ground collaboration strategy based on a game theory model; using a reinforcement learning algorithm to perform dynamic resource allocation according to the air-ground collaboration strategy, and the reinforcement learning algorithm performs real-time optimization allocation of power resources and computing resources based on device status and target characteristics; The fifth unit is used to control the air and ground countermeasure devices to generate interference signals and transmit them to the target UAV according to the result of optimized resource allocation.
[0084] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0085] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0086] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for performing various aspects of the present invention are loaded.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive generation method for UAV countermeasure signals with air-ground collaboration, characterized in that, Including: Obtain the communication signal of the target UAV, extract the signal feature parameters and analyze to obtain the communication protocol type; Encode the signal feature parameters using a two-level coding structure, where the first level represents the interference strategy type and the second level represents the specific parameters; use a genetic algorithm with a composite fitness function to optimize the encoded parameters, and the composite fitness function is used to evaluate the interference effect, power consumption efficiency and anti-detection performance; select the corresponding basic template from a preset interference waveform template library according to the communication protocol type, and use the optimized parameters for parameter matching and adjustment to generate an interference signal template; Send the interference signal template to the aerial countermeasure device and the ground countermeasure device; Execute the resource optimization allocation for air-ground cooperation, including: constructing an air-ground cooperation strategy based on a game theory model; using a reinforcement learning algorithm to perform dynamic resource allocation according to the air-ground cooperation strategy, and the reinforcement learning algorithm performs real-time optimization allocation of power resources and computing resources based on the device state and target characteristics; According to the resource optimization allocation result, control the aerial and ground countermeasure devices to generate interference signals and transmit them to the target UAV.
2. The method according to claim 1, characterized in that, Encode the signal feature parameters using a two-level coding structure, where the first level represents the interference strategy type and the second level represents the specific parameters, including: Perform first-level encoding on the interference strategy type using a 4-bit binary code to obtain a strategy code, and the interference strategy types include a noise interference strategy, a deception interference strategy, a blocking interference strategy, and a combined interference strategy; Perform second-level encoding on the specific parameters using a variable-length coding structure to obtain a parameter code, and the specific parameters include frequency domain parameters, time domain parameters, and power parameters. The frequency domain parameters include the interference bandwidth ratio, frequency offset, and power spectrum shape factor. The time domain parameters include the interference duration, time delay, and repetition period. The power parameters include the interference power, power distribution coefficient, and waveform factor; optimize the coding bit width of the parameter code based on the parameter importance to obtain an optimized code, and the coding bit width is determined according to the parameter value range and quantization accuracy requirements; Perform adaptive mapping on the strategy code and the optimized code according to the target characteristics to obtain the final feature parameter code.
3. The method according to claim 2, wherein Use a genetic algorithm with a composite fitness function to optimize the encoded parameters, and the composite fitness function is used to evaluate the interference effect, power consumption efficiency and anti-detection performance, including: Construct a composite fitness function, and the composite fitness function includes: an interference effect evaluation item for evaluating the signal-to-interference ratio, bit error rate, and interference coverage; a power consumption efficiency evaluation item for evaluating the power efficiency, time efficiency, and energy utilization rate, and the energy utilization rate is determined by the time domain integral ratio of the effective interference power to the total interference power; an anti-detection performance evaluation item for evaluating the detection probability, exposure time, and low intercept characteristic, and the low intercept characteristic is determined by the interference bandwidth ratio, power fluctuation degree, and transmission pattern characteristics; The composite fitness function uses dynamic weight coefficients to perform weighted combination on evaluation items. The dynamic weight coefficients calculate the weight adjustment amount according to the deviation between the target performance and the actual performance. The target performance includes the expected interference effect, power consumption efficiency, and anti-detection index, and the actual performance includes the interference effect, power consumption efficiency, and anti-detection index measured in real time; The genetic algorithm is used to iteratively optimize the strategy encoding and the optimized encoding to obtain the optimal solution, and output the optimized interference strategy and parameter configuration; the genetic algorithm adopts an adaptive mutation rate, and the adaptive mutation rate decreases as the fitness value increases.
4. The method according to claim 3, wherein The determination of the dynamic weight coefficient includes: Establish a performance index system, including the signal-to-interference ratio, bit error rate, and interference coverage range in the interference effect index set, the power efficiency, time efficiency, and energy utilization rate in the power consumption efficiency index set, and the detection probability, exposure time, and low intercept characteristic in the anti-detection index set, and determine the ideal target value and the acceptable range for each index; Adopt the normalized deviation calculation function δ ij =(T ij -A ij ) / N ij Calculate the performance deviation of each index, where T ij is the target performance value, A ij is the actual measured value, and N ij is the normalization factor; perform weighted summation on the deviations of similar indexes to obtain the comprehensive deviation Δ i ; Use the non-linear deviation response function to calculate the weight adjustment amount: ; Among them, ΔW i is the adjustment amount of the weight corresponding to the i-th type of index, α is the adjustment amplitude coefficient, β is the non-linear exponent used to control the curvature of the response function, γ is the sensitivity coefficient used to control the response degree to small deviations, τ min is the minimum response threshold for filtering noise, τ max is the maximum response threshold for limiting abnormal situations; sign(Δ i ) is the sign function of the deviation, which determines the direction of weight adjustment; Smoothly update and normalize the weights to ensure that the sum of all weights is 1, and the weight values are restricted by the upper and lower limits.
5. The method according to claim 1, characterized in that, Constructing an air-ground cooperation strategy based on the game theory model includes: Obtain the state vectors of the aerial countermeasure equipment and the ground countermeasure equipment, and construct an air-ground countermeasure game theory model. The state vectors include power, position, coverage range, and energy state; Take the aerial countermeasure equipment and the ground countermeasure equipment as game participants, and construct a game strategy space including power selection, interference direction, and resource allocation according to the state vectors. The game strategy space is used to restrict the strategy selection range of the game participants; Construct a multi-objective utility function. The multi-objective utility function includes an interference effect evaluation function obtained by weighted combination of signal-to-interference ratio evaluation, coverage effect evaluation, and directivity evaluation, and a resource consumption evaluation function obtained by weighted combination of power consumption, communication overhead, and maneuver cost. Combine the multi-objective utility function with a resource penalty factor to construct a game payoff matrix, and solve the optimal response strategy under the Nash equilibrium based on the game payoff matrix; Calculate the air-ground cooperation advantage degree based on the optimal response strategy, and construct a strategy selection probability distribution. The air-ground cooperation advantage degree is determined by the ratio of the sum of the joint countermeasure benefits to the separate countermeasure benefits; Establish a state value function according to the strategy selection probability distribution. The state value function calculates the long-term expected payoff of the game payoff matrix based on the discount factor, and obtains the air-ground cooperation strategy according to the gradient of the strategy selection probability distribution with respect to the state value function.
6. The method according to claim 1, characterized in that The reinforcement learning algorithm performs real-time optimal allocation of power resources and computing resources based on the device state and target characteristics, including: Obtain the device state and target characteristics, and construct a resource state tensor including platform types and resource types. The platform types include aerial platforms and ground platforms, and the resource types include power resources and computing resources; Construct a variational autoencoder to encode the resource state tensor to generate latent variables, and decode the latent variables into communication messages between the aerial and ground platforms; Based on the resource state tensor and the communication messages between the aerial and ground platforms, design a dual-path deep Q-network with value stream and advantage stream. The value stream evaluates the state value of resource allocation, and the advantage stream evaluates the action advantage of resource allocation. Combine the state value and action advantage to obtain the Q-value of the resource allocation strategy. Use the proximal policy optimization algorithm to train the policy network, which includes an actor network and a critic network. The actor network outputs the mean and standard deviation of the resource allocation action distribution that meets the requirements of aerial-ground cooperation according to the constraints of the aerial-ground cooperation strategy, and the critic network estimates the state value based on the current state and calculates the action advantage. Construct a hierarchical prioritized experience replay pool, calculate the sample priorities based on the temporal difference error, and sample the experience samples according to the sample priorities. Use the meta-learning framework for policy update, and construct an aerial-ground consistency evaluation function to impose additional constraints on the training of the policy network. According to the trained policy network, perform real-time optimal allocation of power resources and computing resources for the aerial platform and the ground platform.
7. The method according to claim 6, characterized in that, The aerial-ground consistency evaluation function includes: The aerial-ground consistency evaluation function includes a resource balance evaluation item, a task completion evaluation item, and a cooperation efficiency evaluation item. The resource balance evaluation item calculates the difference in resource allocation between the aerial platform and the ground platform. The task completion evaluation item evaluates the task execution effect based on the platform resource utilization rate. The cooperation efficiency evaluation item is evaluated based on the communication quality and cooperation response time between the aerial and ground platforms. Calculate the compliance degree of the resource allocation action and the aerial-ground cooperation strategy according to the weighted combination of the resource balance evaluation item, the task completion evaluation item, and the cooperation efficiency evaluation item.
8. An adaptive generation system for UAV countermeasure signals with air-ground cooperation, which is used to implement the method described in any one of the foregoing claims 1-7, is characterized in that It includes: The first unit is used to obtain the communication signal of the target UAV, extract the signal feature parameters, and analyze to obtain the communication protocol type. The second unit is used to encode the signal feature parameters using a two-level coding structure, where the first level represents the interference strategy type and the second level represents the specific parameters. Use a genetic algorithm with a composite fitness function to optimize the encoded parameters. The composite fitness function is used to evaluate the interference effect, power consumption efficiency, and anti-detection performance. Select the corresponding basic template from the preset interference waveform template library according to the communication protocol type, and use the optimized parameters for parameter matching and adjustment to generate an interference signal template. The third unit is used to send the interference signal template to the aerial countermeasure device and the ground countermeasure device. The fourth unit is used to perform optimal allocation of resources for aerial-ground cooperation, including: constructing an aerial-ground cooperation strategy based on a game theory model; using a reinforcement learning algorithm to perform dynamic resource allocation according to the aerial-ground cooperation strategy. The reinforcement learning algorithm performs real-time optimal allocation of power resources and computing resources based on the device state and target characteristics. The fifth unit is used to control the aerial and ground countermeasure devices to generate interference signals and transmit them to the target UAV according to the results of the optimal resource allocation.
9. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Air-ground cooperative detection networking deployment optimization method
CN116151105A
Unmanned aerial vehicle trajectory optimization method and system based on bionic algorithm and BP neural network
CN116782269A
Unmanned aerial vehicle communication anti-interference decision model and method based on deep reinforcement learning
CN118138173A
Unmanned aerial vehicle covert communication resource optimization method based on intelligent reflecting surface assistance
CN119629614A
Efficient Localization of Transmitters Within Complex Electromagnetic Environments
US20160127931A1
Cited By
Resource adjustment method, device and equipment of system
CN120407199A
Unmanned aerial vehicle instruction link blocking interference method and device based on communication protocol identification
CN121000333A
Unmanned aerial vehicle command link jamming method and device based on communication protocol identification
CN121000333B
Dynamic defense model construction method for unmanned aerial vehicle countering
CN121028568A