Method and system for switching operating states of a device using tunable electromagnetic metamaterials
By using adjustable electromagnetic metamaterials on the target device, combining action agents and iterative agents to optimize state switching strategies, the problem that passive jammer devices cannot adapt to complex electromagnetic environments is solved, and radar interference and target equipment survivability are improved.
Patent Information
- Application Number
- CN202510745481.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The electromagnetic scattering characteristics of existing passive jammer devices are fixed and cannot adapt to the complex and changeable electromagnetic environment, which limits the survivability of the target equipment and the radar interference effect.
The adjustable electromagnetic metamaterial is adopted to set the search environment and target device parameters of the phased array radar to establish a working state sequence, and use the action agent and iterative agent to optimize the state switching strategy, select the optimal state to replace the action, and optimize the weight of the agent to improve interference ability and survivability.
On the premise of reducing the requirements for adjustable electromagnetic metamaterial materials, the interference capability of the radar and the survivability of the target equipment are improved.
Smart Images

Figure CN120254777B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radar interference, and in particular relates to a method and system for switching the working state of a device using adjustable electromagnetic metamaterials. Background Art
[0002] In modern search and rescue missions, the survivability of target penetration equipment depends heavily on its ability to electronically jam radar. Electronic jamming can be categorized as active or passive. Passive jamming, which achieves interference by reflecting radar waves, offers advantages such as a large jamming volume, wide bandwidth, strong real-time performance, flexible use, good concealment, and low cost. However, traditional passive jamming devices (such as chaff and corner reflectors) have fixed electromagnetic scattering characteristics and are unable to adapt to complex and changing electromagnetic environments, limiting their effectiveness. Therefore, tunable electromagnetic metamaterials are being introduced into the field of passive jamming, endowing traditional passive devices with dynamic jamming capabilities.
[0003] Related technologies primarily enhance the survivability of target devices by using tunable electromagnetic metamaterials on their surfaces. Existing techniques primarily focus on optimizing the material's performance, such as improving broadband absorption efficiency in the absorbing state, enhancing anisotropic scattering properties in the reflecting state, and shortening state switching response time. By optimizing these materials, the survivability of target devices can be improved.
[0004] Regarding the above-mentioned related technologies, when using tunable electromagnetic metamaterials on the surface of target devices, in addition to the research on the material itself, the interference effect of the state switching of the tunable electromagnetic metamaterials on the phased array radar is not considered. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and system for switching the working state of a device using an adjustable electromagnetic metamaterial, which can reduce the material requirements of the adjustable electromagnetic metamaterial and also improve the interference capability against radar and the survivability of the target device.
[0006] A method for switching the working state of a device using an adjustable electromagnetic metamaterial, comprising:
[0007] Setting the search environment and target device parameters of the phased array radar, wherein the target device parameters include the number of targets, target speed, and target switching time;
[0008] Setting the working state of the target device, and establishing an initial working state sequence according to the working state, wherein the working state includes an absorbing state and a reflecting state;
[0009] Under the search environment and the target device parameters, obtaining a corresponding beam ratio, a total tracking time, an average search frame period, and a first capture time according to the initial working state sequence;
[0010] Set up action agents and iteration agents;
[0011] An initial reward is obtained according to the beam ratio, the total tracking time, the average search frame period, the first capture time, and the weight ratio;
[0012] By executing the strategy, a state replacement action is selected, and the state replacement action is performed on the initial working state sequence to obtain the next state sequence;
[0013] Inputting the updated state sequence generated when executing the state replacement action into the action agent to obtain a number of action output values;
[0014] Taking the next state sequence as the current state sequence, and taking the reward corresponding to the current state sequence as the current reward;
[0015] The current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence are stored in the experience pool as experience values;
[0016] Randomly extract a number of target experience values from the experience pool, and input the next state sequence in the target experience value, the state replacement action corresponding to the next state sequence, and the initial reward into the iterative agent to obtain an iterative output value;
[0017] Calculate the loss value based on the maximum action output value and the iterative output value;
[0018] Adjust the weight of the action agent according to the loss value to obtain an optimized agent;
[0019] The next state sequence is input into the optimization agent as the updated current state sequence and iterated continuously to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized. The iteratively updated state sequence when the current reward is maximized is used as the optimized working state sequence, and the working state of the device is switched according to the optimized working state sequence.
[0020] Optionally, the state replacement action is selected by executing a strategy, and the state replacement action is executed on the initial working state sequence to obtain a next state sequence. The execution strategy includes a first strategy and a second strategy, including:
[0021] Get probability threshold and random probability;
[0022] When the random probability is greater than the probability threshold, the first strategy is executed, wherein the first strategy is:
[0023] The state replacement action is performed on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions, the action output values of all updated working state sequences are calculated by the action agent, the target state replacement action corresponding to the maximum action output value is selected, and the target state replacement action is performed on the initial working state sequence to obtain the next state sequence;
[0024] When the random probability is less than or equal to the probability threshold, the second strategy is executed, which is:
[0025] The next state sequence is obtained by performing randomly selected actions on the initial working state sequence.
[0026] Optionally, obtaining the beam ratio includes:
[0027] Setting the event scheduling period and beam dwell time, where the event scheduling period is the interval between phased array radar scheduling tasks and the beam dwell time is the dwell time during phased array radar scanning;
[0028] Calculating the total number of beams according to the event scheduling period and the beam dwell time;
[0029] Get confirmation beam;
[0030] A beam ratio is obtained according to the confirmed beam and the total number of beams.
[0031] Optionally, obtaining the initial reward according to the beam ratio, total tracking time, average search frame period, first capture time, and weight ratio includes:
[0032] According to the weight proportions, the beam ratio weight, the total tracking time weight, the average search frame period weight, and the first capture time weight are obtained;
[0033] Substituting the beam ratio weight, the total tracking time weight, the average search frame period weight, the first capture time weight, the beam ratio, the total tracking time, the average search frame period, and the first capture time into a reward function to obtain an initial reward;
[0034] The reward function is:
[0035] ;
[0036] in, is the beam proportion weight, is the beam ratio, To track the total time weight, To track total time, is the average search frame period weight, is the average search frame period, is the first capture time weight, The time of first capture.
[0037] Optionally, performing a state replacement action on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions, and calculating the action output values of all updated working state sequences by the action agent includes:
[0038] Set the initial weight parameters of the action agent;
[0039] According to the action agent, an action formula is obtained;
[0040] Performing a state replacement action on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions;
[0041] Inputting the initial weight parameter, the updated working state sequence and the state replacement action into the action formula to obtain an action output value;
[0042] The action formula is expressed as:
[0043] ;
[0044] in, is the action output value, Output value for the action agent, To update the working status sequence, is the replacement action for the current state, is the initial weight parameter.
[0045] Optionally, the step of randomly extracting a number of target experience values from the experience pool and inputting the next state sequence, the state replacement action corresponding to the next state sequence, and the initial reward in the target experience values into the iterative agent to obtain the iterative output value includes:
[0046] Acquiring an iterative formula according to the iterative agent;
[0047] Substituting the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward corresponding to the current state into the iterative formula to obtain an iterative output value;
[0048] The iterative formula is:
[0049] ;
[0050] in, is the iterative output value, To iterate and update rewards, is the discount factor, is the output value of the iterative agent, is the iterative working state sequence, is the replacement action for the next state, , is the initial weight parameter.
[0051] Optionally, calculating the loss value according to the maximum action output value and the iterative output value includes:
[0052] Get the loss function;
[0053] Substituting the maximum action output value and the iterative output value into the loss function to calculate the loss value;
[0054] The loss function is expressed as:
[0055] ;
[0056] in, is the loss value, To calculate the expectation, is the action output value, Output value for the iteration.
[0057] A working state switching system for a device using an adjustable electromagnetic metamaterial, comprising:
[0058] A first setting module is used to set the search environment and target device parameters of the phased array radar, wherein the target device parameters include the number of targets, target speed, and target switching time;
[0059] A second setting module is used to set the working state of the target device and establish an initial working state sequence according to the working state, wherein the working state includes an absorbing state and a reflecting state;
[0060] A first calculation module is configured to obtain, according to the initial working state sequence, a corresponding beam ratio, a total tracking time, an average search frame period, and a first capture time under the search environment and the target device parameters;
[0061] The third setting module is used to set the action agent and the iteration agent;
[0062] A second calculation module is used to obtain an initial reward based on the beam ratio, the total tracking time, the average search frame period, the first capture time and the weight ratio;
[0063] An action selection module is used to select a state replacement action by executing a strategy, and execute the state replacement action on the initial working state sequence to obtain a next state sequence;
[0064] A third calculation module is used to input the updated state sequence generated when executing the state replacement action into the action agent to obtain a number of action output values;
[0065] Taking the next state sequence as the current state sequence, and taking the reward corresponding to the current state sequence as the current reward;
[0066] An iteration module, configured to store the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence as experience values in an experience pool;
[0067] a fourth calculation module, configured to randomly extract a number of target experience values from the experience pool, and input the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the experience values into the iterative agent to obtain an iterative output value;
[0068] a fifth calculation module, configured to calculate a loss value based on the maximum action output value and the iterative output value;
[0069] An optimization module, configured to adjust the weight of the action agent according to the loss value to obtain an optimized agent;
[0070] The switching module is used to input the next state sequence as the updated current state sequence into the optimization agent for continuous iteration to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized, and the iteratively updated state sequence when the current reward is maximized is used as the optimized working state sequence, and the working state of the device is switched according to the optimized working state sequence.
[0071] A terminal device includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, a working state switching method of a device using adjustable electromagnetic metamaterial is adopted.
[0072] A computer-readable storage medium stores a computer program. When the computer program is loaded and executed by a processor, a method for switching the working state of a device using an adjustable electromagnetic metamaterial is adopted.
[0073] The beneficial effects of the present invention are:
[0074] The search environment and target device parameters of the phased array radar are set. Then, in this environment, the initial reward is calculated according to the working state sequence. Then, the action output values of all updated working state sequences after the state replacement action is executed are calculated. The largest action output value is selected to execute the state replacement action to obtain the next state sequence. The next state sequence, the initial reward, and the state replacement action are input into the iterative agent to calculate the iterative output value. By calculating the action output value and the loss value of the iterative output value, the action agent is optimized to obtain an optimized agent. The next state sequence is input into the optimized agent as the new current state sequence and iterated continuously to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized. The iteratively updated state sequence when the current reward is maximized is used as the real-time working state sequence. According to the real-time working state sequence, the working state of the device is switched. This application has the advantages of improving the interference capability to the radar and the survivability of the target device by optimizing the switching strategy while reducing the material requirements for the adjustable electromagnetic metamaterial. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 Schematic diagram of the search method of the phased array radar of the present invention;
[0076] Figure 2 This is a schematic diagram of the intelligent agent structure of the present invention;
[0077] Figure 3 Schematic diagram of the training process of the intelligent agent of the present invention;
[0078] Figure 4 This is a first comparison diagram of the results of the switching method of the present invention after smoothing;
[0079] Figure 5 It is a second comparison diagram of the result after smoothing of the switching method of the present invention. DETAILED DESCRIPTION
[0080] A method for switching the working state of a device using an adjustable electromagnetic metamaterial, such as Figure 1 As shown, the present invention includes:
[0081] S1. Set the search environment and target device parameters of the phased array radar. The target device parameters include the number of targets, target speed, and target switching time.
[0082] Specifically, the search environment includes search airspace division, beam arrangement and search method design. Search airspace division includes determining the radar's search area, clarifying the airspace range and its boundary conditions. Beam arrangement: within the search area, the beam is arranged and designed to determine the position of each beam and its coverage range. Search method design: specifies the radar's search method, including parameters such as beam scanning sequence and dwell time.
[0083] In this embodiment, the radar search area ranges from -20° to 20° in azimuth (0° for north) and from 0° to 30° in elevation. Assuming the beam width is 1.7°, the radar position arrangement within the search area adopts staggered arrangement, such as Figure 1 As shown, the scanning mode is left-to-right, the beam dwell time is 2ms, and the event scheduling period is 100ms. Therefore, a maximum of 50 beam position events can be executed within one event scheduling period. Assume that a target enters the radar search area from a distance of 100km, at an azimuth of -19.32° and an elevation of 24.94°, at a speed of 2000m / s, and moves from west to east. The target's absorption / reflection state switching period is 0.5s, meaning that the target decides whether to change its current state every 0.5 seconds.
[0084] In its normal state, the radar performs a search mission, continuously scanning a designated airspace. When a potential target is detected within the search area, the radar first performs a confirmation re-illumination, verifying its presence by illuminating the target multiple times with the beam. When the cumulative probability of target confirmation reaches a preset threshold, the radar switches from search to tracking.
[0085] Target device parameters are mainly for target devices, and are set when passing through the search airspace. They mainly include the number of target devices, speed, location of entering the radar search area, time for target switching status, etc.
[0086] S2. Set the working state of the target device and establish an initial working state sequence according to the working state. The working state includes an absorbing state and a reflecting state.
[0087] Specifically, the state sequence of the adjustable electromagnetic metamaterial absorption / reflection state covered on the target is used as the state, which is an n-bit binary sequence, that is, According to the parameter settings in this embodiment, the target device takes 31 seconds to pass through the radar search area, so the sequence length n = 62. The length n of the adjustable electromagnetic metamaterial absorption / reflection state sequence is calculated based on the total time the target passes through the radar search area and the time the target switches between different states. Each bit in the sequence represents the absorption / reflection state of the target in each switching cycle. If it is "0", it means the target is in the absorption state, and the RCS (radar cross section) of the target is small, making it difficult for the radar to detect the target. If it is "1", it means the target is in the reflection state, and the RCS of the target is large, making it easy for the radar to detect the target.
[0088] According to the total time the target passes through the radar search area and the target state switching period, the length n of the absorption / reflection state sequence is calculated as follows:
[0089] .
[0090] S3. Under the search environment and target device parameters, the corresponding beam ratio, total tracking time, average search frame period and first capture time are obtained according to the target action sequence.
[0091] Obtaining beam ratio includes:
[0092] S31. Setting an event scheduling period and a beam dwell time. The event scheduling period is the interval time of the phased array radar scheduling tasks, and the beam dwell time is the dwell time of the phased array radar during scanning.
[0093] S32. Calculate the total number of beams according to the event scheduling period and the beam dwell time.
[0094] S33. Acquire a confirmation beam.
[0095] S34. Obtain a beam ratio based on the confirmed beam and the total number of beams.
[0096] Specifically, the beam ratio is calculated as follows:
[0097] .
[0098] .
[0099] in, is the event scheduling cycle, is the beam dwell time, is the total number of beams, To search the beam, To confirm the beam, is the number of tracking beams.
[0100] The search frame period is calculated as
[0101] .
[0102] in, To search for the frame period, is the total number of wave positions.
[0103] In simulation software, such as MATLAB, after setting the search environment and target device parameters and performing simulation, the beam ratio, total tracking time, average search frame period, and first capture time parameters can be obtained.
[0104] Confirm beam ratio : Defined as the ratio of the radar's beam resources used for target confirmation to the total beam resources. A higher confirmation beam ratio means the radar uses more beam resources for target confirmation, resulting in a relative reduction in beam resources for search and tracking missions.
[0105] Track total time : Defined as the cumulative time the radar tracks the target within a specified time.
[0106] Search frame period : Defines the average time required for the radar to complete a complete scan of the entire search airspace.
[0107] First capture time : Defined as the time from the target entering the radar surveillance airspace to the radar first switching to tracking the target.
[0108] S4. Set up action agent and iteration agent.
[0109] Specifically, two deep learning networks with the same structure were constructed: TQ-net (action agent) and TargetTQ-net (iteration agent). TQ-net is used to estimate Q values (action output values) in real time and guide the agent's action selection. Target TQ-net is used to provide stable target Q values (iteration output values) to reduce fluctuations during training. The structures of the two networks are consistent, such as Figure 2 As shown:
[0110] Input Embedding Layer: The input sequence is mapped into a continuous vector representation through the embedding layer to capture the semantic information of the input features. Specifically, the input embedding layer uses a fully connected layer to transform the sequence features from one dimension to embedding_dim = 64 dimensions.
[0111] Positional encoding: Positional encoding is added to the embedding vector to introduce the order information of the sequence.
[0112] Feature extraction layer: The Transformer encoder processes the input sequence using a multi-layer self-attention mechanism and feedforward neural network to capture global dependencies between elements in the sequence. Specifically, the multi-head attention mechanism has 8 heads, the encoder has 4 layers, the hidden layer dimension of the feedforward neural network is 512, and the dropout probability is 0.1.
[0113] Output layer: This layer generates predictions for each time step, i.e., the Q-value estimates for the corresponding actions. Specifically, the output layer is a fully connected layer that converts the sequence from the embedding_dim dimension back to one dimension. The difference between the action agent and the iterative agent lies in the different initial weights and calculation formulas.
[0114] S5. Get the initial reward based on the beam ratio, total tracking time, average search frame period, first capture time, and weight ratio.
[0115] Based on the beam ratio, total tracking time, average search frame period, first capture time and weight ratio, the initial rewards include:
[0116] S51. Obtain the beam ratio weight, the total tracking time weight, the average search frame period weight, and the first capture time weight according to the weight ratio.
[0117] S52. Substitute the beam ratio weight, total tracking time weight, average search frame period weight, first capture time weight, beam ratio, total tracking time, average search frame period, and first capture time into the reward function to obtain the initial reward.
[0118] The reward function is:
[0119] .
[0120] in, is the beam proportion weight, is the beam ratio, To track the total time weight, To track total time, is the average search frame period weight, is the average search frame period, is the first capture time weight, The time of first capture.
[0121] 、 、 as well as The reward function is a weighting factor for each of the four indicators, used to balance their impact on system performance. In this embodiment, the values are 600, -0.15, 0.1, and 1, respectively. This reward function comprehensively evaluates the jamming effect by quantifying the consumption of radar search resources and the impact on target search and tracking performance. It aims to weaken the radar's ability to search for new targets while reducing its ability to continuously track already tracked targets.
[0122] S6. By executing the strategy, a state replacement action is selected, and the state replacement action is performed on the initial working state sequence to obtain the next state sequence.
[0123] By executing the strategy, a state replacement action is selected, and the state replacement action is executed on the initial working state sequence to obtain the next state sequence. The execution strategy includes a first strategy and a second strategy, including:
[0124] S61, obtaining a probability threshold and a random probability;
[0125] S62: When the random probability is greater than the probability threshold, execute the first strategy, which is:
[0126] Execute the state replacement action on the initial working state sequence to obtain all updated working state sequences that execute different state replacement actions. Calculate the action output values of all updated working state sequences through the action agent, select the target state replacement action corresponding to the maximum action output value, and execute the target state replacement action on the initial working state sequence to obtain the next state sequence.
[0127] S63. When the random probability is less than or equal to the probability threshold, execute the second strategy, which is:
[0128] The next state sequence is obtained by performing randomly selected actions on the initial working state sequence.
[0129] Specifically, here we use the traversal method, each state switching cycle has two states, namely 0 and 1, that is, a total of 2 n combinations.
[0130] The action output values of each of these n working state sequences are calculated by the action agent. The state replacement action corresponding to the maximum action output value is selected as the target state replacement action and executed to obtain the next state sequence. At this point, one iteration is completed. The next state sequence is then used as the basis and the same operation is performed again to obtain the next state sequence.
[0131] Specifically, in each step, for an initial sequence of adjustable electromagnetic metamaterial absorption / reflection states, use Strategy selection action Specifically, The strategy is expressed as follows:
[0132]
[0133] Among them, Q is the output value of the action agent, s is the current state sequence, a is the state replacement action, is the action agent weight, For random actions, The probability of choosing a randomly selected action.
[0134] Each time an action is selected, a random number is first generated between 0 and 1 as the probability threshold. When the random probability is greater than the probability threshold ( ), execute the first strategy, when the random probability is less than or equal to the probability threshold ( ), execute the second strategy.
[0135] Specifically, the action is related to the sequence length n. Each action corresponds to a flip operation on a certain bit in the sequence, so the size of the action space is n, that is, , for example when When, right The first bit in the digit is inverted, and so on.
[0136] Description The probability of choosing the best action based on current knowledge is also called utilization, with The probability of randomly selecting an action is also called exploration. In the early stages of training, it is more inclined to explore, so is a larger value. As the training progresses, It will gradually decrease, preferring to select actions through the use of action agents. That is, after the initial working state sequence is determined, the possible actions for each time are determined. For example, if the initial state sequence is [1, 1, 1, 1], then each change may have four situations, namely, the first 1 becomes 0, the second 1 becomes 0, the third 1 becomes 0, or the fourth 1 becomes 0. By calculating the Q value (action output value) of each state, the replacement action with the largest Q value is selected to execute, and the working state sequence for the next state is obtained. Assuming that the Q value is the largest after the first 1 becomes 0, the replacement action is to change the first 1 to 0, and the resulting sequence is [0, 1, 1, 1].
[0137] S7, inputting the updated state sequence generated when executing the state replacement action into the action agent to obtain a number of action output values;
[0138] S8. Take the next state sequence as the current state sequence, and take the reward corresponding to the current state sequence as the current reward.
[0139] S9. The current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence are stored in the experience pool as experience values.
[0140] Specifically, the experience pool can store multiple data, and a preset number is usually set. When the number of experience points stored in the experience pool is equal to the preset number, no more data will be stored. The preset number can be set by yourself.
[0141] Execute the state replacement action on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions. The action output values of all updated working state sequences calculated by the action agent include:
[0142] S621. Set the initial weight parameters of the action agent.
[0143] S622. Obtain an action formula based on the action agent.
[0144] S623: Execute a state replacement action on the initial working state sequence to obtain all updated working state sequences that execute different state replacement actions.
[0145] S624: Input the initial weight parameters, all updated working state sequences, and state replacement actions into the action formula to obtain the action output value.
[0146] The action formula is expressed as:
[0147] .
[0148] Among them, y is the action output value, Output value for the action agent, To update the working status sequence, is the replacement action for the current state, is the initial weight parameter.
[0149] Specifically, small batch data samples are randomly extracted from the experience pool D.
[0150] For each sample, the current state As the input of TQ-net, the Q value is obtained.
[0151] Specifically, the training process is as follows Figure 3 As shown, the experience pool D is initialized, and the experience pool can store N state-action pairs. TQ-net is initialized with random weights, and the weight parameters are , initialize Target TQ-net, the weight parameter is .
[0152] At each step, for an initial state sequence, an action is selected using the ε-greedy strategy .
[0153] Execute an action , observe the new state given by the environment and rewards .
[0154] New data experience will be obtained Stored in the experience pool D, the preset number of experience points that can be stored in the experience pool can be set by yourself.
[0155] S10. Randomly extract several target experience values from the experience pool, and input the next state sequence in the target experience value, the state replacement action corresponding to the next state sequence, and the current reward into the iterative agent to obtain the iterative output value.
[0156] Randomly extract a number of target experience values from the experience pool, and input the next state sequence in the target experience value, the state replacement action corresponding to the next state sequence, and the initial reward into the iterative agent. The iterative output values include:
[0157] S101. Obtain an iterative formula according to an iterative agent.
[0158] S102. Substitute the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward corresponding to the current state into the iterative formula to obtain the iterative output value.
[0159] The iteration formula is:
[0160] .
[0161] in, is the iterative output value, To iterate and update rewards, is the discount factor, is the output value of the iterative agent, is the iterative working state sequence, is the replacement action for the next state, , is the initial weight parameter.
[0162] Specifically, the next state As the input of Target TQ-net, it outputs the target Q value.
[0163] S11. Calculate the loss value based on the action output value and the iteration output value.
[0164] According to the action output value and the iteration output value, the loss value is calculated including:
[0165] S111. Obtain a loss function.
[0166] S112: Substitute the action output value and the iteration output value into the loss function to calculate the loss value.
[0167] The loss function is expressed as:
[0168] .
[0169] in, is the loss value, E is the expected value, is the action output value, Output value for the iteration.
[0170] S12. Adjust the weight of the action agent according to the loss value to obtain the optimized agent.
[0171] S13. Input the next state sequence as the new current state sequence to optimize the intelligent weight and continuously iterate to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized. The iteratively updated state sequence when the current reward is maximized is used as the real-time working state sequence. According to the real-time working state sequence, the working state of the device is switched.
[0172] Specifically, the weights of TQ-net are updated using gradient updates. , every C steps, the weight of TQ-net is Weights copied to Target TQ-net Update status for , in order to proceed to the next cycle.
[0173] It is worth noting that each time a state replacement action is executed, a current state sequence and a next state sequence are generated. Therefore, there will also be an action output value and an iteration output value. A loss value can be calculated and weight optimization can be performed through backpropagation. However, because the next state sequence will continue to be generated, the optimization process will also occur multiple times. The stopping condition is that the reward maximum value converges. Each optimization will change the parameters of the action agent. , get the optimized agent, then use the optimized agent as the new optimization object, change the parameters until the iteration ends.
[0174] like Figure 4 and Figure 5 As shown in the figure, the reward value obtained by the deep reinforcement learning algorithm in each round has large fluctuations, but after smoothing, we can more easily see the overall upward trend of the reward value, and finally gradually stabilize. This shows that as the number of iterations increases, the algorithm gradually adapts to the environment and the reward value also increases. This shows that under the guidance of the reward function, the deep reinforcement learning algorithm learns and improves through continuous interaction with the environment, and finally learns a near-optimal strategy. Figure 5 It can be seen that the maximum value of the reward converges to 5.4, while the reward obtained using the unoptimized strategy is 3.644. This comparison shows that the switching strategy optimized by deep reinforcement learning can significantly improve the radar jamming effect.
[0175] A working state switching system for a device using an adjustable electromagnetic metamaterial, comprising:
[0176] The first setting module is used to set the search environment and target device parameters of the phased array radar, the target device parameters including the number of targets, target speed and target switching time;
[0177] The second setting module is used to set the working state of the target device and establish an initial working state sequence according to the working state, where the working state includes an absorbing state and a reflecting state;
[0178] A first calculation module is used to obtain the corresponding beam ratio, total tracking time, average search frame period and first capture time according to the initial working state sequence under the search environment and target device parameters;
[0179] The third setting module is used to set the action agent and the iteration agent;
[0180] The second calculation module is used to obtain the initial reward based on the beam ratio, total tracking time, average search frame period, first capture time and weight ratio;
[0181] The action selection module is used to select a state replacement action through the execution strategy, and perform the state replacement action on the initial working state sequence to obtain the next state sequence;
[0182] A third calculation module is used to input the updated state sequence generated when executing the state replacement action into the action agent to obtain a number of action output values;
[0183] Take the next state sequence as the current state sequence, and the reward corresponding to the current state sequence as the current reward;
[0184] The iteration module is used to store the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence as experience values in the experience pool;
[0185] The fourth calculation module is used to randomly extract a number of target experience values from the experience pool, and input the next state sequence in the experience value, the state replacement action corresponding to the next state sequence, and the current reward into the iterative agent to obtain the iterative output value;
[0186] A fifth calculation module is used to calculate a loss value based on the maximum action output value and the iterative output value;
[0187] The optimization module is used to adjust the weight of the action agent according to the loss value to obtain the optimized agent;
[0188] The switching module is configured to input the next state sequence as the updated current state sequence into the optimization agent for continuous iteration, obtaining an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized. The iteratively updated state sequence at which the current reward is maximized is then used as the optimized working state sequence, and the device's working state is switched based on the optimized working state sequence. The present application also discloses a terminal device comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, a method for switching the working state of a device using an adjustable electromagnetic metamaterial is employed.
[0189] Among them, the terminal device can be a computer device such as a desktop computer, a laptop computer or a cloud server, and the terminal device includes but is not limited to a processor and a memory. For example, the terminal device can also include input and output devices, network access devices and buses, etc.
[0190] Among them, the processor can adopt a central processing unit (CPU). Of course, according to actual usage, other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc., and this application does not impose any restrictions on this.
[0191] Among them, the memory can be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device, or it can be an external storage device of the terminal device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD) or flash memory card (FC) equipped on the terminal device, etc., and the memory can also be a combination of the internal storage unit and the external storage device of the terminal device. The memory is used to store computer programs and other programs and data required by the terminal device. The memory can also be used to temporarily store data that has been output or is to be output. This application does not impose any restrictions on this.
[0192] Among them, through this terminal device, a working state switching method of a device using adjustable electromagnetic metamaterial in the above embodiment is stored in the memory of the terminal device, and is loaded and executed on the processor of the terminal device for easy use.
[0193] An embodiment of the present application further discloses a computer-readable storage medium, and the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, a method for switching the working state of a device using an adjustable electromagnetic metamaterial in the above embodiment is adopted.
[0194] Among them, the computer program can be stored in a computer-readable medium, the computer program includes computer program code, the computer program code can be in the form of source code, object code, executable file or certain middleware, etc. The computer-readable medium includes any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that computer-readable medium includes but is not limited to the above-mentioned components.
[0195] Among them, through this computer-readable storage medium, a working state switching method of a device using adjustable electromagnetic metamaterials in the above embodiment is stored in a computer-readable storage medium, and is loaded and executed on a processor to facilitate the storage and application of the above method.
[0196] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of protection of the present application is limited to these examples. In line with the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as above, which are not provided in detail for the sake of simplicity.
[0197] The one or more embodiments of this application are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this application should be included in the scope of protection of this application.
Claims
1. A method for switching the working state of a device using an adjustable electromagnetic metamaterial, characterized in that: include: Setting the search environment and target device parameters of the phased array radar, wherein the target device parameters include the number of targets, target speed, and target switching time; Setting the working state of the target device, and establishing an initial working state sequence according to the working state, wherein the working state includes an absorbing state and a reflecting state; Under the search environment and the target device parameters, obtaining a corresponding beam ratio, a total tracking time, an average search frame period, and a first capture time according to the initial working state sequence; Set up action agents and iteration agents; An initial reward is obtained according to the beam ratio, the total tracking time, the average search frame period, the first capture time, and the weight ratio; By executing the strategy, a state replacement action is selected, and the state replacement action is performed on the initial working state sequence to obtain the next state sequence; Inputting the updated state sequence generated when executing the state replacement action into the action agent to obtain a number of action output values; Taking the next state sequence as the current state sequence, and taking the reward corresponding to the current state sequence as the current reward; The current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence are stored in the experience pool as experience values; Randomly extract a number of target experience values from the experience pool, and input the next state sequence in the target experience value, the state replacement action corresponding to the next state sequence, and the current reward into the iterative agent to obtain an iterative output value; Calculate the loss value based on the maximum action output value and the iterative output value; Adjust the weight of the action agent according to the loss value to obtain an optimized agent; The next state sequence is input into the optimization agent as the updated current state sequence and iterated continuously to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized. The iteratively updated state sequence when the current reward is maximized is used as the optimized working state sequence, and the working state of the device is switched according to the optimized working state sequence.
2. The method for switching the working state of a device using an adjustable electromagnetic metamaterial according to claim 1, wherein: The state replacement action is selected by executing the strategy, and the state replacement action is executed on the initial working state sequence to obtain the next state sequence. The execution strategy includes a first strategy and a second strategy, including: Get probability threshold and random probability; When the random probability is greater than the probability threshold, the first strategy is executed, wherein the first strategy is: The state replacement action is performed on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions, the action output values of all updated working state sequences are calculated by the action agent, the target state replacement action corresponding to the maximum action output value is selected, and the target state replacement action is performed on the initial working state sequence to obtain the next state sequence; When the random probability is less than or equal to the probability threshold, the second strategy is executed, which is: The next state sequence is obtained by performing randomly selected actions on the initial working state sequence.
3. The method for switching the working state of a device using an adjustable electromagnetic metamaterial according to claim 1, wherein: Obtaining beam ratio includes: Setting the event scheduling period and beam dwell time, wherein the event scheduling period is the interval time of the phased array radar scheduling tasks, and the beam dwell time is the dwell time of the phased array radar during scanning; Calculating the total number of beams according to the event scheduling period and the beam dwell time; Get confirmation beam; A beam ratio is obtained according to the confirmed beam and the total number of beams.
4. The method for switching the working state of a device using an adjustable electromagnetic metamaterial according to claim 1, wherein: The initial reward obtained according to the beam ratio, total tracking time, average search frame period, first capture time and weight ratio includes: According to the weight proportions, the beam ratio weight, the total tracking time weight, the average search frame period weight, and the first capture time weight are obtained; Substituting the beam ratio weight, the total tracking time weight, the average search frame period weight, the first capture time weight, the beam ratio, the total tracking time, the average search frame period, and the first capture time into a reward function to obtain an initial reward; The reward function is: ; in, is the beam proportion weight, is the beam ratio, To track the total time weight, To track total time, is the average search frame period weight, is the average search frame period, is the first capture time weight, The time of first capture.
5. The method for switching the working state of a device using an adjustable electromagnetic metamaterial according to claim 2, wherein: The performing of the state replacement action on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions, and calculating the action output values of all updated working state sequences by the action agent include: Set the initial weight parameters of the action agent; According to the action agent, an action formula is obtained; Performing a state replacement action on the initial working state sequence to obtain all updated working state sequences that perform different state replacement actions; Inputting the initial weight parameter, all the updated working state sequences and the state replacement action into the action formula to obtain an action output value; The action formula is expressed as: ; in, is the action output value, Output value for the action agent, To update the working status sequence, is the replacement action for the current state, is the initial weight parameter.
6. The method for switching the working state of a device using an adjustable electromagnetic metamaterial according to claim 1, wherein: The step of randomly extracting a number of target experience values from the experience pool and inputting the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward in the target experience value into the iterative agent to obtain the iterative output value includes: Acquire an iterative formula according to the iterative agent; Substituting the next state sequence, the state replacement action corresponding to the next state sequence, and the current reward corresponding to the current state into the iterative formula to obtain an iterative output value; The iterative formula is: ; in, is the iterative output value, To iterate and update rewards, is the discount factor, is the output value of the iterative agent, is the iterative working state sequence, is the replacement action for the next state, , is the initial weight parameter.
7. The method for switching the working state of a device using an adjustable electromagnetic metamaterial according to claim 1, wherein: The calculating the loss value according to the maximum action output value and the iterative output value includes: Get the loss function; Substituting the maximum action output value and the iterative output value into the loss function to calculate the loss value; The loss function is expressed as: ; in, is the loss value, To calculate the expectation, is the action output value, Output value for the iteration.
8. A working state switching system for a device using an adjustable electromagnetic metamaterial, characterized in that: include: A first setting module is used to set the search environment and target device parameters of the phased array radar, wherein the target device parameters include the number of targets, target speed, and target switching time; A second setting module is used to set the working state of the target device and establish an initial working state sequence according to the working state, wherein the working state includes an absorbing state and a reflecting state; A first calculation module is configured to obtain, according to the initial working state sequence, a corresponding beam ratio, a total tracking time, an average search frame period, and a first capture time under the search environment and the target device parameters; The third setting module is used to set the action agent and the iteration agent; A second calculation module is used to obtain an initial reward based on the beam ratio, the total tracking time, the average search frame period, the first capture time and the weight ratio; An action selection module is used to select a state replacement action by executing a strategy, and execute the state replacement action on the initial working state sequence to obtain a next state sequence; A third calculation module is used to input the updated state sequence generated when executing the state replacement action into the action agent to obtain a number of action output values; Taking the next state sequence as the current state sequence, and taking the reward corresponding to the current state sequence as the current reward; An iteration module, configured to store the current state sequence, the state replacement action corresponding to the current state sequence, the initial reward, and the next state sequence corresponding to the current state sequence as experience values in an experience pool; a fourth calculation module, configured to randomly extract a number of target experience values from the experience pool, and input the next state sequence in the target experience values, the state replacement action corresponding to the next state sequence, and the current reward into the iterative agent to obtain an iterative output value; a fifth calculation module, configured to calculate a loss value based on the maximum action output value and the iterative output value; An optimization module, configured to adjust the weight of the action agent according to the loss value to obtain an optimized agent; The switching module is used to input the next state sequence as the updated current state sequence into the optimization agent for continuous iteration to obtain an iteratively updated state sequence until the current reward corresponding to the obtained iteratively updated state sequence is maximized, and the iteratively updated state sequence when the current reward is maximized is used as the optimized working state sequence, and the working state of the device is switched according to the optimized working state sequence.
9. A terminal device comprising a memory and a processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the method according to any one of claims 1 to 7 is adopted.
10. A computer-readable storage medium storing a computer program, wherein: When the computer program is loaded and executed by a processor, the method according to any one of claims 1 to 7 is adopted.
Citation Information
Patent Citations
Multifunctional radar interference decision-making method
CN118501823A
Scheduling analysis method for integrated avionics system based on reinforcement learning
CN119806783A