Working state switching method and system for equipment adopting adjustable electromagnetic metamaterial
By optimizing the state switching strategy of adjustable electromagnetic metamaterials, the survivability of the target equipment and radar interference effect are improved, and the problem of insufficient adaptability of passive jammer devices in complex electromagnetic environments is solved.
Patent Information
- Application Number
- CN202510745481.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The electromagnetic scattering characteristics of existing passive jammer devices are fixed and cannot adapt to the complex and changeable electromagnetic environment, which limits the survivability of the target equipment and the radar interference effect.
Devices that use adjustable electromagnetic metamaterials, set the search environment and target device parameters of phased array radar to establish an initial working state sequence, use action agents and iterative agents to optimize the state switching strategy, select state replacement actions, calculate action and iterative output values, adjust weights, optimize the agents, and switch the working state of the device.
While reducing the requirements for adjustable electromagnetic metamaterial materials, the interference capability of the radar and the survivability of the target equipment are improved.
Smart Images

Figure CN120254777A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radar jamming, and particularly relates to a method and system for switching the working state of a device using tunable electromagnetic metamaterials. Background Art
[0002] In modern search missions, the survivability of penetration target devices largely depends on their electronic jamming capabilities against radars. Electronic jamming can be divided into active jamming and passive jamming. Among them, passive jamming realizes jamming by reflecting radar waves, and has advantages such as a large jamming airspace, a wide frequency band, strong real-time performance, flexible use, good concealment, and low cost. However, the electromagnetic scattering characteristics of traditional passive jamming devices (such as chaff, corner reflectors, etc.) are fixed and cannot adapt to complex and changeable electromagnetic environments, which limits their application effects. Therefore, tunable electromagnetic metamaterials are introduced into the field of passive jamming to endow traditional passive devices with dynamic jamming capabilities.
[0003] In related technologies, for the survivability of target devices, tunable electromagnetic metamaterials are mainly used on the surface of target devices. The existing technologies mainly focus on the performance optimization of the materials themselves. For example: improving the broadband absorption efficiency in the absorption state, enhancing the anisotropic scattering characteristics in the reflection state, and shortening the state switching response time, etc. By optimizing the materials, the survivability of target devices is improved.
[0004] In view of the above related technologies, when using tunable electromagnetic metamaterials on the surface of target devices, in addition to the research on the materials themselves, the interference effect of the state switching of tunable electromagnetic metamaterials on phased array radars is not considered. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and system for switching the working state of a device using tunable electromagnetic metamaterials, which can improve the interference ability against radars and the survivability of target devices while reducing the material requirements for tunable electromagnetic metamaterials.
[0006] A method for switching the working state of a device using tunable electromagnetic metamaterials includes: Setting the search environment of a phased array radar and target device parameters, where the target device parameters include the number of targets, the target speed, and the time for target state switching; Setting the working state of the target device, and establishing an initial working state sequence according to the working state, where the working state includes an absorption state and a reflection state; Under the search environment and the target device parameters, obtaining the corresponding beam ratio, total tracking time, average search frame period, and first capture time according to the initial working state sequence; Setting an action agent and an iterative agent; An initial reward is obtained based on the beam ratio, the total tracking time, the average search frame period, the first capture time, and the weight ratio. By executing a policy, a state replacement action is selected, and the state replacement action is performed on the initial working state sequence to obtain the next state sequence. The updated state sequence generated when performing the state replacement action is input into the action agent to obtain a number of action output values. The next state sequence is used as the current state sequence, and the reward corresponding to the current state sequence is used as the current reward. The current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence are stored in the experience pool as experience values. A number of target experience values are randomly drawn from the experience pool, and the next state sequence, the state replacement action corresponding to the next state sequence, and the initial reward in the target experience values are input into the iterative agent to obtain an iterative output value. A loss value is calculated based on the maximum action output value and the iterative output value. The weights of the action agent are adjusted according to the loss value to obtain an optimized agent. The next state sequence is input into the optimized agent as the updated current state sequence for continuous iteration to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is the largest. The iteratively updated state sequence with the largest current reward is used as the optimized working state sequence, and the working state of the device is switched according to the optimized working state sequence.
[0007] Optionally, in the step of "by executing a policy, selecting a state replacement action, performing the state replacement action on the initial working state sequence, and obtaining the next state sequence", the execution policy includes a first policy and a second policy, including: A probability threshold and a random probability are obtained. When the random probability is greater than the probability threshold, the first policy is executed. The first policy is: For the initial working state sequence, state replacement actions are performed to obtain all updated working state sequences with different state replacement actions. The action output values of all updated working state sequences are calculated by the action agent, and the target state replacement action corresponding to the maximum action output value is selected. The target state replacement action is performed on the initial working state sequence to obtain the next state sequence. When the random probability is less than or equal to the probability threshold, the second policy is executed. The second policy is: A randomly selected action is performed on the initial working state sequence to obtain the next state sequence.
[0008] Optionally, obtaining the beam ratio includes: Setting an event scheduling period and a beam dwell time, where the event scheduling period is the interval time of the phased array radar scheduling task, and the beam dwell time is the residence time when the phased array radar scans; Calculating the total number of beams based on the event scheduling period and the beam dwell time; Obtaining a confirmation beam; Obtaining the beam ratio based on the confirmation beam and the total number of beams.
[0009] Optionally, obtaining the initial reward according to the beam ratio, the total tracking time, the average search frame period, the first capture time, and the weight ratio includes: Obtaining the beam ratio weight, the total tracking time weight, the average search frame period weight, and the first capture time weight according to the weight ratio; Substituting the beam ratio weight, the total tracking time weight, the average search frame period weight, the first capture time weight, the beam ratio, the total tracking time, the average search frame period, and the first capture time into a reward function to obtain the initial reward; The reward function is: ; Wherein, is the beam ratio weight, is the beam ratio, is the total tracking time weight, is the total tracking time, is the average search frame period weight, is the average search frame period, is the first capture time weight, is the first capture time.
[0010] Optionally, performing a state replacement action on the initial working state sequence to obtain all updated working state sequences performing different state replacement actions, and calculating the action output values of all updated working state sequences by an action agent includes: Setting the initial weight parameters of the action agent; Obtaining an action formula according to the action agent; Performing a state replacement action on the initial working state sequence to obtain all updated working state sequences performing different state replacement actions; Inputting the initial weight parameters, the updated working state sequences, and the state replacement actions into the action formula to obtain the action output values; The action formula is expressed as: ; Among them, is the action output value, is the action agent output value, is the updated working state sequence, is the replacement action for the current state, is the initial weight parameter.
[0011] Optionally, randomly extracting several target experience values from the experience pool, and inputting the next state sequence, the state replacement action corresponding to the next state sequence, and the initial reward in the target experience values into the iterative agent to obtain the iterative output value includes: Obtaining an iterative formula according to the iterative agent; Substituting the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward corresponding to the current state into the iterative formula to obtain the iterative output value; The iterative formula is: ; Among them, is the iterative output value, is the iterative updated reward, is the discount factor, is the iterative agent output value, is the iterative working state sequence, is the replacement action for the next state, , is the initial weight parameter.
[0012] Optionally, calculating the loss value according to the maximum action output value and the iterative output value includes: Obtaining a loss function; Substituting the maximum action output value and the iterative output value into the loss function to calculate the loss value; The loss function is expressed as: ; Among them, is the loss value, is the calculation expectation, is the action output value, is the iterative output value.
[0013] A working state switching system for a device using tunable electromagnetic metamaterials, comprising: A first setting module, configured to set the search environment of the phased array radar and the target device parameters, where the target device parameters include the number of targets, the target speed, and the time of the target switching state; A second setting module, configured to set the working state of the target device, and establish an initial working state sequence according to the working state, where the working state includes an absorbing state and a reflecting state; A first calculation module, configured to obtain corresponding beam ratios, total tracking time, average search frame period, and first capture time according to the initial working state sequence under the search environment and the target device parameters; A third setting module, configured to set an action agent and an iterative agent; A second calculation module, configured to obtain an initial reward according to the beam ratio, total tracking time, average search frame period, first capture time, and weight ratio; An action selection module, configured to select a state replacement action through an execution policy, and perform the state replacement action on the initial working state sequence to obtain a next state sequence; A third calculation module, configured to input the updated state sequence generated when performing the state replacement action into the action agent to obtain a plurality of action output values; Use the next state sequence as the current state sequence, and use the reward corresponding to the current state sequence as the current reward; An iteration module, configured to store the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence into an experience pool as experience values; A fourth calculation module, configured to randomly extract a plurality of target experience values from the experience pool, and input the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the experience values into the iterative agent to obtain an iterative output value; A fifth calculation module, configured to calculate a loss value according to the maximum action output value and the iterative output value; An optimization module, configured to adjust the weights of the action agent according to the loss value to obtain an optimized agent; A switching module, configured to input the next state sequence as an updated current state sequence into the optimized agent for continuous iteration to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is the largest, use the iteratively updated state sequence with the largest current reward as the optimized working state sequence, and switch the working state of the device according to the optimized working state sequence.
[0014] A terminal device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor loads and executes the computer program, it adopts a method for switching the working state of a device using tunable electromagnetic metamaterials.
[0015] A computer-readable storage medium stores a computer program. When the computer program is loaded and executed by a processor, a method for switching the working state of a device using tunable electromagnetic metamaterials is adopted.
[0016] The beneficial effects of the present invention are as follows: Set the search environment of the phased array radar and the target device parameters. Then, in this environment, calculate the initial reward according to the working state sequence. Then calculate the action output values of all updated working state sequences after executing the state replacement action, and select the maximum action output value to execute the state replacement action to obtain the next state sequence. Input the next state sequence, the initial reward, and the state replacement action into the iterative agent to calculate the iterative output value. By calculating the loss values of the action output value and the iterative output value, the action agent is optimized to obtain an optimized agent. Input the next state sequence as the new current state sequence into the optimized agent and iterate continuously to obtain the iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is the largest. Take the iteratively updated state sequence with the largest current reward as the real-time working state sequence, and switch the working state of the device according to the real-time working state sequence. This application can improve the interference ability to the radar and the survival ability of the target device by optimizing the switching strategy while reducing the requirements for the materials of the tunable electromagnetic metamaterials. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the search method of the phased array radar of the present invention; Figure 2 It is a schematic diagram of the agent structure of the present invention; Figure 3 It is a schematic diagram of the agent training process of the present invention; Figure 4 It is a first comparison diagram after the result smoothing process of the switching method of the present invention; Figure 5 It is a second comparison diagram after the result smoothing process of the switching method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] A method for switching the working state of a device using tunable electromagnetic metamaterials, as Figure 1 shown, the present invention includes: S1. Set the search environment of the phased array radar and the target device parameters. The target device parameters include the number of targets, the target speed, and the time for the target to switch states.
[0019] Specifically, the search environment includes search airspace division, waveform arrangement, and search mode design. Search airspace division includes determining the radar's search area, clarifying the airspace range and its boundary conditions. Waveform arrangement: within the search area, arrange and design the waveforms to determine the position and coverage range of each beam direction. Search mode design: stipulate the radar's search mode, including parameters such as beam scanning order and dwell time.
[0020] In this embodiment, the radar's search area range is: azimuth angle from -20° to 20° (north is 0°), elevation angle from 0° to 30°. Assuming the beam bandwidth is 1.7°, the waveform arrangement of the radar within the search area adopts staggered arrangement, as Figure 1 shown. The scanning mode is to search from left to right, the beam dwell time is 2 ms, and the event scheduling period is 100 ms. Therefore, within one event scheduling period, at most 50 waveform events can be executed. Assuming the target enters the radar search area from a place 100 km away from the radar, with an azimuth angle of -19.32° and an elevation angle of 24.94°, and a speed of 2000 m / s, passing through the radar search area from west to east. The switching period of the target's wave absorption / reflection state is 0.5 s, that is, the target decides whether to change the current state every 0.5 s.
[0021] In the normal state, the radar performs the search task and continuously scans the specified airspace. When a potential target is detected within the search area, the radar first confirms and re-illuminates the target, that is, verifies the existence of the target through multiple beam irradiations. When the cumulative confirmation probability of the target reaches the preset threshold, the radar changes the target state from search to tracking.
[0022] The target device parameters are mainly for the target device, and are set according to the target parameters in the search airspace, mainly including the number of target devices, speed, position where the target enters the radar search area, time when the target switches states, etc.
[0023] S2. Set the working state of the target device, and establish an initial working state sequence according to the working state. The working state includes the wave absorption state and the reflection state.
[0024] Specifically, use the adjustable electromagnetic metamaterial wave absorption / reflection state sequence covering the target as the state, which is an n-bit binary sequence, that is According to the parameter setting in this embodiment, the time for the target device to pass through the radar search area is 31s, so the sequence length n=62. According to the total time of the target passing through the radar search area and the time when the target is set to switch state, the length n of the adjustable electromagnetic metamaterial absorption / reflection state sequence is calculated. Each bit in the sequence represents the absorption / reflection state of the target in each switching cycle. If it is "0", it means that the target is in the absorption state. At this time, the RCS (radar cross section) of the target is small, and it is difficult for the radar to detect the target. If it is "1", it means that the target is in the reflection state. At this time, the RCS of the target is large, and the radar can easily detect the target.
[0025] According to the total time of the target passing through the radar search area and the target state switching cycle, the length n of the absorption / reflection state sequence is calculated as follows: .
[0026] S3. Under the search environment and target device parameters, the corresponding beam ratio, total tracking time, average search frame period and first capture time are obtained according to the target action sequence.
[0027] Obtaining beam ratio includes: S31. Setting an event scheduling cycle and a beam dwell time. The event scheduling cycle is the interval time of the phased array radar scheduling tasks, and the beam dwell time is the stay time of the phased array radar during scanning.
[0028] S32. Calculate the total number of beams according to the event scheduling period and the beam dwell time.
[0029] S33. Obtain a confirmation beam.
[0030] S34. Obtain a beam ratio according to the confirmed beam and the total number of beams.
[0031] Specifically, the beam ratio is calculated as follows: .
[0032] .
[0033] in, is the event scheduling cycle, is the beam dwell time, is the total number of beams, To search the beam, To confirm the beam, is the number of tracking beams.
[0034] The search frame period is calculated as .
[0035] in, is the search frame period, and [[ID=]] is the total number of waveform positions.
[0036] In simulation software, such as MATLAB, after setting the search environment and target device parameters and performing the simulation, the beam ratio, total tracking time, average search frame period, and first capture time parameters can be obtained.
[0037] Confirm the beam ratio : Defined as the ratio of the beam resources used by the radar for target confirmation tasks to the total beam resources. The higher the confirmation beam ratio, the more beam resources the radar uses for target confirmation, resulting in a relative reduction in the beam resources for search and tracking tasks.
[0038] Total tracking time : Defined as the cumulative time during which the radar tracks the target within a specified time.
[0039] Search frame period : Defined as the average time required for the radar to complete a full scan of the entire search airspace.
[0040] First capture time : Defined as the time required for the radar to transfer the target from entering the radar surveillance airspace to being first tracked by the radar.
[0041] S4. Set the action agent and the iterative agent.
[0042] Specifically, two deep learning networks with the same structure are constructed: TQ-net (action agent) and TargetTQ-net (iterative agent). Among them, TQ-net is used to estimate the Q value (action output value) in real time and guide the action selection of the agent. Target TQ-net is used to provide a stable target Q value (iterative output value) to reduce fluctuations during the training process. The structures of the two networks are the same, as Figure 2 shown: Input embedding layer: Map the input sequence to a continuous vector representation through the embedding layer to capture the semantic information of the input features. Specifically, the input embedding layer uses a fully connected layer to transform the features of the sequence from one dimension to embedding_dim = 64 dimensions.
[0043] Position encoding: Add position encoding to the embedded vector to introduce the order information of the sequence.
[0044] Feature extraction layer: The input sequence is processed by a Transformer encoder through multiple layers of self-attention mechanism and feed-forward neural network to capture the global dependencies between elements in the sequence. Specifically, the number of heads in the multi-head attention mechanism is 8, the number of layers in the encoder is 4, the dimension of the hidden layer in the feed-forward neural network is 512, and the Dropout probability is 0.1.
[0045] Output layer: The prediction results for each time step are generated through the output layer, that is, the Q-value estimation corresponding to the action. Specifically, the output layer is a fully connected layer that reshapes the sequence from the embedding_dim dimension back to one dimension. The difference between the action agent and the iterative agent lies in the different initial weights and calculation formulas.
[0046] S5. The initial reward is obtained based on the beam ratio, total tracking time, average search frame period, first capture time, and weight ratio.
[0047] Obtaining the initial reward based on the beam ratio, total tracking time, average search frame period, first capture time, and weight ratio includes: S51. Based on the weight ratio, the beam ratio weight, total tracking time weight, average search frame period weight, and first capture time weight are obtained.
[0048] S52. Substitute the beam ratio weight, total tracking time weight, average search frame period weight, first capture time weight, beam ratio, total tracking time, average search frame period, and first capture time into the reward function to obtain the initial reward.
[0049] The reward function is: .
[0050] Where, is the beam ratio weight, is the beam ratio, is the total tracking time weight, is the total tracking time, is the average search frame period weight, is the average search frame period, is the first capture time weight, is the first capture time.
[0051] , , and They are the weight coefficients of 4 indicators respectively, used to balance the influence of different indicators on the system performance. In this embodiment, the values are 600, -0.15, 0.1, and 1 respectively. This reward function comprehensively evaluates the interference effect by quantifying the consumption of radar search resources and the influence on target search and tracking performance, aiming to weaken the radar's ability to search for new targets and at the same time reduce its continuous tracking ability for the tracked targets.
[0052] S6. By executing the policy, select a state replacement action, and perform the state replacement action on the initial working state sequence to obtain the next state sequence.
[0053] By executing the policy, select a state replacement action, and perform the state replacement action on the initial working state sequence to obtain the next state sequence. The execution policy includes a first policy and a second policy, including: S61. Obtain the probability threshold and the random probability; S62. When the random probability is greater than the probability threshold, execute the first policy. The first policy is: Perform the state replacement action on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions. Calculate the action output values of all updated working state sequences through the action agent, select the target state replacement action corresponding to the maximum action output value, and perform the target state replacement action on the initial working state sequence to obtain the next state sequence; S63. When the random probability is less than or equal to the probability threshold, execute the second policy. The second policy is: Perform a randomly selected action on the initial working state sequence to obtain the next state sequence.
[0054] Specifically, here an exhaustive method is adopted. Each state switching period has two states, namely 0 and 1, that is, there are a total of 2 n combinations.
[0055] Calculate the action output values of these n working state sequences through the action agent, select the state replacement action corresponding to the maximum action output value as the target state replacement action, and execute it to obtain the next state sequence. At this time, one iteration is completed, and then the next state sequence is used as the basis, and the same operation is performed again to obtain the next-next state sequence.
[0056] Specifically, in each step size, for an initial adjustable electromagnetic metamaterial absorption / reflection state sequence, use the policy to select an action . Specifically, the policy is expressed as follows: Among them, Q is the output value of the action agent, s is the current state sequence, a is the state replacement action, is the weight of the action agent, is the random action, is the probability of randomly selecting an action.
[0057] When selecting an action each time, a random number is first generated, between 0 and 1, as the probability threshold. When the random probability is greater than the probability threshold ( ), the first strategy is executed. When the random probability is less than or equal to the probability threshold ( ), the second strategy is executed.
[0058] Specifically, the action action is related to the sequence length n. Each action corresponds to a flipping operation on a certain bit in the sequence. Therefore, the size of the action space is n, that is , for example, when , for , the first bit is inverted, and so on.
[0059] It means that there is a probability of choosing the best action according to the current knowledge, which is also called exploitation. There is a probability of randomly choosing an action, which is also called exploration, . In the initial stage of training, it is more inclined to exploration. Therefore is a relatively large value. As the training progresses, will gradually decrease and be more inclined to choose actions by exploiting the action agent. That is to say, after the initial working state sequence is determined, each possible action is determined. For example, if the initial state sequence is [1, 1, 1, 1], then there may be 4 situations for each change, that is, the 1 in the first bit becomes 0 or the 1 in the second bit becomes 0 or the 1 in the third bit becomes 0 or the 1 in the fourth bit becomes 0. By calculating the Q value (action output value) of each state, the replacement action with the largest Q value is selected to execute, and the working state sequence of the next state is obtained. Assuming that the Q value is the largest after the first bit becomes 0, the replacement action is to change the 1 in the first bit to 0, and the obtained sequence is [0, 1, 1, 1].
[0060] S7. Input the updated state sequence generated when executing the state replacement action into the action agent to obtain several action output values; S8. Use the next state sequence as the current state sequence and the reward corresponding to the current state sequence as the current reward.
[0061] S9. Store the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence into the experience pool as experience values.
[0062] Specifically, multiple data can be stored in the experience pool. Usually, a preset quantity is set. When the number of experience values stored in the experience pool is equal to the preset quantity, no more data will be stored, and the preset quantity can be set by oneself.
[0063] Perform state replacement actions on the initial working state sequence to obtain all updated working state sequences after performing different state replacement actions. The action output values of all updated working state sequences are calculated by the action agent, including: S621. Set the initial weight parameters of the action agent.
[0064] S622. Obtain the action formula according to the action agent.
[0065] S623. Perform state replacement actions on the initial working state sequence to obtain all updated working state sequences after performing different state replacement actions.
[0066] S624. Input the initial weight parameters, all updated working state sequences, and the state replacement actions into the action formula to obtain the action output value.
[0067] The action formula is expressed as: .
[0068] Among them, y is the action output value, is the action agent output value, is the updated working state sequence, is the replacement action of the current state, is the initial weight parameter.
[0069] Specifically, randomly draw a small batch of data samples from the experience pool D.
[0070] For each sample, use the current state as the input of the TQ-net to obtain the Q value.
[0071] Specifically, the training process is as Figure 3 shown. Initialize the experience pool D. The experience pool can store N state-action pairs. Initialize the TQ-net with random weights, and the weight parameter is , initialize the Target TQ-net, and the weight parameter is .
[0072] In each time step, for an initial state sequence, use the ε-greedy strategy to select an action .
[0073] Execute the action , observe the new state given by the environment and the reward .
[0074] Store the obtained new data experience in the experience pool D. The preset number of experience values that can be stored in the experience pool can be set by oneself.
[0075] S10. Randomly extract several target experience values from the experience pool, and input the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the target experience values into the iterative agent to obtain an iterative output value.
[0076] Randomly extract several target experience values from the experience pool, and input the next state sequence, the state replacement action corresponding to the next state sequence, and the initial reward into the iterative agent to obtain an iterative output value, including: S101. Obtain an iterative formula according to the iterative agent.
[0077] S102. Substitute the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward corresponding to the current state into the iterative formula to obtain an iterative output value.
[0078] The iterative formula is: .
[0079] Among them,[[]] is the iterative output value,[[]] is the iterative update reward,[[]] is the discount factor,[[]] is the iterative agent output value,[[]] is the iterative working state sequence,[[]] is the replacement action of the next state,[[]] ,[[]] is the initial weight parameter.[[]]
[0080] Specifically, take the next state as the input of the Target TQ-net and output the target Q value.[[]]
[0081] S11. Calculate the loss value according to the action output value and the iterative output value.[[]]
[0082] Calculating the loss value according to the action output value and the iterative output value includes:[[]] S111. Obtain the loss function.[[]]
[0083] S112. Substitute the action output value and the iterative output value into the loss function to calculate the loss value.[[]]
[0084] The loss function is expressed as:[[]] .
[0085] Among them, is the loss value, E is the calculated expectation, is the action output value, is the iterative output value.
[0086] S12. Adjust the weights of the action agent according to the loss value to obtain an optimized agent.
[0087] S13. Take the next state sequence as the new current state sequence and input it into the optimized agent for continuous iteration to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is the largest. Take the iteratively updated state sequence with the largest current reward as the real-time working state sequence, and switch the working state of the device according to the real-time working state sequence.
[0088] Specifically, use gradient update to update the weights of the TQ-net , every C steps, then copy the weights of the TQ-net to the weights of the Target TQ-net in, and update the state as for the next loop.
[0089] It should be noted that each time the state replacement action is executed, a current state sequence and a next state sequence will be generated. Therefore, there will also be an action output value and an iterative output value, and a loss value can be calculated. The weights are optimized through backpropagation. However, because the next state sequence will be continuously generated, the optimization process also exists multiple times. The stopping condition is that the maximum reward converges. Each optimization will change the parameters of the action agent , obtain an optimized agent, and then use the optimized agent as a new optimization object to change the parameters until the iteration ends.
[0090] As Figure 4 and Figure 5 shown, the reward values obtained by the deep reinforcement learning algorithm in each round have large fluctuations. However, after smoothing, we can more easily see the overall upward trend of the reward values, and finally gradually tend to be stable. This shows that as the number of iterations increases, the algorithm gradually adapts to the environment, and the reward value also increases accordingly. This indicates that under the guidance of the reward function, the deep reinforcement learning algorithm continuously learns and improves in the interaction with the environment, and finally learns a near-optimal strategy. From Figure 5 it can be seen that the maximum value of the reward converges at 5.4, while using the unoptimized strategy, the obtained reward is 3.644. This comparison shows that the switching strategy optimized by deep reinforcement learning can significantly improve the radar jamming effect.
[0091] A working state switching system for a device using tunable electromagnetic metamaterials, comprising: A first setting module for setting the search environment of a phased array radar and target device parameters, where the target device parameters include the number of targets, the target speed, and the time of the target switching state; A second setting module for setting the working state of the target device and establishing an initial working state sequence according to the working state, where the working state includes an absorbing state and a reflecting state; A first calculation module for obtaining the corresponding beam ratio, total tracking time, average search frame period, and first capture time according to the initial working state sequence under the search environment and target device parameters; A third setting module for setting an action agent and an iterative agent; A second calculation module for obtaining an initial reward according to the beam ratio, total tracking time, average search frame period, first capture time, and weight ratio; An action selection module for selecting a state replacement action through an execution policy, performing the state replacement action on the initial working state sequence, and obtaining a next state sequence; A third calculation module for inputting the updated state sequence generated when performing the state replacement action into the action agent to obtain a number of action output values; Taking the next state sequence as the current state sequence and the reward corresponding to the current state sequence as the current reward; An iteration module for storing the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence into an experience pool as experience values; A fourth calculation module for randomly extracting a number of target experience values from the experience pool and inputting the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the experience values into the iterative agent to obtain an iterative output value; A fifth calculation module for calculating a loss value according to the maximum action output value and the iterative output value; An optimization module for adjusting the weights of the action agent according to the loss value to obtain an optimized agent; A switching module is used to input the next state sequence as the updated current state sequence into the optimization agent for continuous iteration, obtaining an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized. The iteratively updated state sequence when the current reward is maximized is used as the optimized working state sequence, and the working state of the device is switched according to the optimized working state sequence. An embodiment of the present application also discloses a terminal device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor loads and executes the computer program, a method for switching the working state of a device using tunable electromagnetic metamaterials is adopted.
[0092] Among them, the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server. And the terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may also include input / output devices, network access devices, and a bus, etc.
[0093] Among them, the processor can adopt a central processing unit (CPU). Of course, according to the actual usage situation, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc. The present application does not make any restrictions in this regard.
[0094] Among them, the memory can be an internal storage unit of the terminal device. For example, the hard disk or memory of the terminal device, or it can also be an external storage device of the terminal device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital card (SD), or a flash card (FC), etc. equipped on the terminal device. And the memory can also be a combination of the internal storage unit and the external storage device of the terminal device. The memory is used to store the computer program and other programs and data required by the terminal device. The memory can also be used to temporarily store the data that has been output or will be output. The present application does not make any restrictions in this regard.
[0095] Among them, through this terminal device, the method for switching the working state of a device using tunable electromagnetic metamaterials in the above embodiment is stored in the memory of the terminal device, and is loaded and executed on the processor of the terminal device, which is convenient for use.
[0096] An embodiment of the present application also discloses a computer-readable storage medium. And the computer-readable storage medium stores a computer program. Among them, when the computer program is executed by the processor, the method for switching the working state of a device using tunable electromagnetic metamaterials in the above embodiment is adopted.
[0097] Among them, the computer program can be stored in a computer-readable medium. The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some middleware form, etc. The computer-readable medium includes any entity or device, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer-readable medium includes but is not limited to the above components.
[0098] Among them, through this computer-readable storage medium, the working state switching method of a device adopting tunable electromagnetic metamaterials in the above embodiments is stored in the computer-readable storage medium, and is loaded and executed on a processor to facilitate the storage and application of the above method.
[0099] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the protection scope of the present application is limited to these examples; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments in the present application as above, and they are not provided in detail for the sake of brevity.
[0100] One or more embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the present application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of one or more embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A method for switching the operating state of a device using tunable electromagnetic metamaterials, characterized in that, Including: Set the search environment of the phased array radar and the target device parameters, where the target device parameters include the number of targets, the target speed, and the time of the target switching state; Set the working state of the target device, and establish an initial working state sequence according to the working state, where the working state includes an absorbing state and a reflecting state; Under the search environment and the target device parameters, obtain the corresponding beam ratio, total tracking time, average search frame period, and first capture time according to the initial working state sequence; Set an action agent and an iterative agent; Obtain an initial reward according to the beam ratio, total tracking time, average search frame period, first capture time, and weight ratio; By executing a policy, select a state replacement action, and perform the state replacement action on the initial working state sequence to obtain the next state sequence; Input the updated state sequence generated when performing the state replacement action into the action agent to obtain a number of action output values; Take the next state sequence as the current state sequence, and take the reward corresponding to the current state sequence as the current reward; Take the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence as experience values and store them in the experience pool; Randomly extract a number of target experience values from the experience pool, and input the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the target experience values into the iterative agent to obtain an iterative output value; Calculate a loss value according to the maximum action output value and the iterative output value; Adjust the weight of the action agent according to the loss value to obtain an optimized agent; Input the next state sequence as the updated current state sequence into the optimized agent for continuous iteration to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is the largest. Take the iteratively updated state sequence with the largest current reward as the optimized working state sequence, and switch the working state of the device according to the optimized working state sequence.
2. The method for switching the working state of a device using tunable electromagnetic metamaterials according to claim 1, characterized in that, The step of "by executing a policy, select a state replacement action, and perform the state replacement action on the initial working state sequence to obtain the next state sequence", the execution policy includes a first policy and a second policy, and includes: Obtain a probability threshold and a random probability; When the random probability is greater than the probability threshold, execute the first policy, and the first policy is: Perform the state replacement action on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions. Calculate the action output values of all updated working state sequences through the action agent, select the target state replacement action corresponding to the maximum action output value, and perform the target state replacement action on the initial working state sequence to obtain the next state sequence; When the random probability is less than or equal to the probability threshold, execute the second policy, and the second policy is: Perform a randomly selected action on the initial working state sequence to obtain the next state sequence.
3. The method for switching the working state of a device using tunable electromagnetic metamaterials according to claim 1, characterized in that, Obtaining the beam ratio includes: Set the event scheduling period and the beam dwell time. The event scheduling period is the interval time of the phased array radar scheduling task, and the beam dwell time is the residence time when the phased array radar scans; Calculate the total number of beams based on the event scheduling period and the beam dwell time; Obtain the confirmed beam; Obtain the beam ratio based on the confirmed beam and the total number of beams.
4. The method for switching the working state of a device using tunable electromagnetic metamaterials as described in claim 1, characterized in that, The obtaining of the initial reward according to the beam ratio, the total tracking time, the average search frame period, the first capture time, and the weight ratio includes: Obtain the beam ratio weight, the total tracking time weight, the average search frame period weight, and the first capture time weight according to the weight ratio; Substitute the beam ratio weight, the total tracking time weight, the average search frame period weight, the first capture time weight, the beam ratio, the total tracking time, the average search frame period, and the first capture time into the reward function to obtain the initial reward; The reward function is: ; Among them, is the beam ratio weight, is the beam ratio, is the total tracking time weight, is the total tracking time, is the average search frame period weight, is the average search frame period, is the first acquisition time weight, is the first acquisition time.
5. The method for switching the working state of a device using tunable electromagnetic metamaterials according to claim 2, characterized in that, The performing of state replacement actions on the initial working state sequence to obtain all updated working state sequences performing different state replacement actions, and the calculation of the action output values of all updated working state sequences by the action agent includes: Set the initial weight parameters of the action agent; Obtain the action formula according to the action agent; Perform state replacement actions on the initial working state sequence to obtain all updated working state sequences performing different state replacement actions; Input the initial weight parameters, all the updated working state sequences, and the state replacement actions into the action formula to obtain the action output values; The action formula is expressed as: ; Among them, is the action output value, is the action agent output value, is the updated working status sequence, is the replacement action for the current state, is the initial weight parameter.
6. The method for switching the working state of a device using tunable electromagnetic metamaterials according to claim 1, characterized in that, The randomly extracting a number of target experience values from the experience pool and inputting the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the target experience values into the iterative agent to obtain the iterative output value includes: Obtain the iterative formula according to the iterative agent; Substitute the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward corresponding to the current state into the iterative formula to obtain the iterative output value; The iterative formula is: ; Among them, is the iterative output value, is the iterative updated reward, is the discount factor, is the iterative agent output value, is the iterative working state sequence, is the replacement action for the next state, , is the initial weight parameter.
7. The method for switching the working state of a device using tunable electromagnetic metamaterials according to claim 1, characterized in that, The calculating of the loss value according to the maximum action output value and the iterative output value includes: Obtain the loss function; Substitute the maximum action output value and the iterative output value into the loss function to calculate the loss value; The loss function is expressed as: ; Among them, is the loss value, is the calculated expectation, is the action output value, is the iterative output value.
8. A working state switching system for a device using tunable electromagnetic metamaterials, characterized in that, Includes: The first setting module is used to set the search environment of the phased array radar and the target device parameters. The target device parameters include the number of targets, the target speed, and the time of the target switching state; The second setting module is used to set the working state of the target device, and establish an initial working state sequence according to the working state. The working state includes the wave absorption state and the reflection state; The first calculation module is used to obtain the corresponding beam ratio, total tracking time, average search frame period, and first capture time according to the initial working state sequence under the search environment and the target device parameters; The third setting module is used to set the action agent and the iterative agent; A second calculation module, configured to obtain an initial reward according to the beam ratio, the total tracking time, the average search frame period, the first capture time, and the weight ratio; An action selection module, configured to select a state replacement action by executing a policy, and perform the state replacement action on the initial working state sequence to obtain a next state sequence; A third calculation module, configured to input the updated state sequence generated when performing the state replacement action into the action agent to obtain a plurality of action output values; Use the next state sequence as the current state sequence, and use the reward corresponding to the current state sequence as the current reward; An iteration module, configured to store the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence in an experience pool as experience values; A fourth calculation module, configured to randomly extract a plurality of target experience values from the experience pool, and input the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the target experience values into an iteration agent to obtain an iteration output value; A fifth calculation module, configured to calculate a loss value according to the maximum action output value and the iteration output value; An optimization module, configured to adjust the weights of the action agent according to the loss value to obtain an optimized agent; A switching module, configured to input the next state sequence as an updated current state sequence into the optimized agent for continuous iteration to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is the largest. Use the iteratively updated state sequence with the largest current reward as the optimized working state sequence, and switch the working state of the device according to the optimized working state sequence.
9. A terminal device, comprising a memory and a processor, characterized in that, The memory stores a computer program capable of running on a processor. When the processor loads and executes the computer program, the method described in any one of claims 1 to 7 is adopted.
10. A computer-readable storage medium storing a computer program therein, characterized in that, When the computer program is loaded and executed by the processor, the method described in any one of claims 1 to 7 is adopted.
Citation Information
Patent Citations
Multifunctional radar interference decision-making method
CN118501823A
Intelligent interference decision-making method and system based on prior knowledge embedded LSTM-PPO model
CN118818440A
Scheduling analysis method for integrated avionics system based on reinforcement learning
CN119806783A
Methods and systems for constrained reinforcement learning
US20240265263A1
Decision-making method based on deep reinforcement learning
WO2022083029A1