An electromagnetic most sensitive waveform testing method based on reinforcement learning
Through reinforcement learning-based methods, building state space and action space, using TD3 network and automated closed-loop testing environment, and designing reward functions, the problem of not being able to quickly find the most sensitive waveform in existing EMS tests is solved, and lower sensitivity thresholds and higher testing efficiency are achieved.
Patent Information
- Application Number
- CN202410974310.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-07-19
AI Technical Summary
The existing EMS testing standards use fixed waveforms for frequency sweep testing, which cannot fully capture the specific electromagnetic sensitivity characteristics of electronic devices. The optimization method has too much calculation overhead when scanning parameters in a huge waveform parameter space, making it difficult to quickly find the most sensitive waveform.
Using reinforcement learning-based methods, state space and action space are constructed, and reinforcement learning models are established using dual-delay depth deterministic strategic gradient network (TD3), combining automated closed-loop testing environment and reward functions to quickly search for the most sensitive waveforms.
The most sensitive waveform that makes the subject's sensitivity threshold the lowest in the modulated waveform set quickly finds the most sensitive waveform that reduces the sensitivity threshold and improves the testing efficiency and accuracy.
Smart Images

Figure CN119178938B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to electromagnetic sensitivity testing, and particularly to an electromagnetic most sensitive waveform testing method based on reinforcement learning. Background Art
[0002] Electromagnetic sensitivity (EMS) refers to the characteristic that a device, component or system may cause performance degradation due to electromagnetic interference, and has become a key factor affecting the reliability and integrity of electronic devices. Electromagnetic sensitivity can be obtained by conducting EMS tests on electronic devices. In this test, an interference signal with a certain power is applied to the equipment under test (EUT). When the interference power exceeds a certain threshold and the EUT experiences performance degradation, this threshold is defined as the sensitivity threshold, which is used to reflect the electromagnetic sensitivity of the EUT. As the integration level of devices is getting higher and the electromagnetic environment is getting more complex, revealing the subtle differences in the electromagnetic sensitivity of devices to different interference signals is crucial for ensuring the safe and stable operation of the devices. Therefore, it is very meaningful to find a signal waveform that can stimulate the electromagnetic sensitivity of the EUT with lower power.
[0003] Existing EMS test standards use fixed waveforms for swept-frequency testing to obtain sensitivity thresholds at different frequencies. For example, GJB-151B recommends using a pulse amplitude modulation (PM) signal with a duty cycle of 50, a pulse repetition frequency of 1 kHz, and a modulation depth of 100%. The International Electrotechnical Commission standard IEC 61000-4-3 recommends using a sinusoidal amplitude modulation (AM) waveform with a modulation frequency of 1 kHz and a modulation depth of 80%. These waveforms recommended by these standards come from the interferences that may exist in the real world. The waveform recommended by GJB-151B is similar to the interference caused by a switching power supply, while the waveform recommended by IEC 61000-4-3 may be the interference caused by an AM broadcast signal. In addition, the interference generated by frequency modulation (FM) broadcasts or frequency modulated continuous wave radars may cause electromagnetic sensitivity. Although these fixed waveforms recommended by these standards are easy to implement and promote, they may not be able to fully capture the specific electromagnetic sensitivity characteristics of some EUTs.
[0004] There have been some studies at home and abroad that have tested specific EUTs with complex test waveforms. For example, the application of signal waveforms such as single-carrier frequency division multiple access, quadrature phase shift keying, and filtered noise in the radiation sensitivity testing of avionics systems. In addition, some scholars have studied the electromagnetic sensitivity of CAN networks under electrical fast transient pulse interference, the electromagnetic sensitivity of personal computer systems under double-exponential electromagnetic pulse interference, and the electromagnetic sensitivity of temperature sensors under the interference of basic emission element waveforms. The above studies have considered various complex waveforms that may be generated in the actual environment, but they may not be the most sensitive waveforms of the EUT.
[0005] Some studies have utilized optimization algorithms to determine the parameters of the most sensitive waveforms. For example, Bayesian optimization can be used for the adaptive sensitivity testing of UAV data links. It is used to search for combinations of interference waveform parameters that cause the UAV data link to be sensitive in complex electromagnetic environments and predict the sensitivity threshold. Given the working principle of the EUT, traditional optimization methods can obtain the most sensitive waveforms that excite the electromagnetic sensitivity of the EUT. However, for EUTs with unknown principles, scanning or optimizing parameters in the vast waveform parameter space incurs unaffordable computational overhead.
[0006] The optimization ability of reinforcement learning in high-dimensional spaces and its adaptability to complex environments make it suitable for searching waveform parameters in EMS testing. Reinforcement learning can determine the optimal strategy in complex high-dimensional spaces and demonstrates strong learning and optimization capabilities even in vast parameter spaces. Recent studies have shown that reinforcement learning has great application potential in the field of electromagnetic compatibility. It has been used for the optimization of PCB ground vias layout and pin assignment optimization in system-in-package chips.
[0007] In summary, existing standard test methods cannot obtain the most sensitive waveforms, and existing optimization methods expose some limitations when searching for the most sensitive waveforms. Therefore, there is an urgent need to develop a new intelligent test method to solve these problems. Summary of the Invention
[0008] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an electromagnetic most sensitive waveform test method based on reinforcement learning, which can quickly search for the most sensitive waveform that minimizes the sensitivity threshold of the test item within the defined modulation waveform set.
[0009] The purpose of the present invention is achieved through the following technical solutions: An electromagnetic most sensitive waveform test method based on reinforcement learning, comprising the following steps:
[0010] S1. Establish a reinforcement learning architecture for electromagnetic sensitivity testing, and construct a state space and an action space;
[0011] The step S1 includes:
[0012] Take the RF signal source in the EMS test as the intelligent agent in reinforcement learning, whose goal is to learn how to perform actions in each step to maximize the cumulative reward and minimize the sensitivity threshold of the test item excited by the test waveform;
[0013] Take the test system in the EMS test as the environment in reinforcement learning, which is used to receive the signal waveform output of the RF signal source and return the sensitive state of the test item;
[0014] Take the waveform parameters of the signal source and the sensitive state of the DUT in the EMS test as the states in reinforcement learning. There are a total of 8 waveform parameters and 2 sensitive state parameters, which together constitute the state space in reinforcement learning. Each parameter in the state space is a state, with a total of 10 states. The value of each state is within the range of [-1, 1].
[0015] Take the change in waveform parameters in the EMS test as the actions in reinforcement learning. The purpose is to change the current state of the agent in order to obtain a greater reward. There are a total of 8 actions, corresponding to the change amounts of 8 waveform parameters, namely the change amount of the carrier frequency, the change amount of the modulation type parameter, the change amount of the modulation waveform parameter, the change amount of the modulation frequency, the change amount of the modulation depth, the change amount of the frequency offset, the change amount of the pulse duty cycle, and the change amount of the pulse period. The values of these 8 actions are all within the range of [-1, 1], which constitutes the action space in reinforcement learning.
[0016] The 8 waveform parameters include:
[0017] Carrier frequency: The symbol is s(1), representing the normalized carrier frequency value, with a value range of [-1, 1]. Let the physical range of the true carrier frequency be [f min , f max . According to s(1) and the physical range [f min , f max , the true carrier frequency is:
[0018] The unit is MHz.
[0019] Modulation type: The symbol is s(2), with a value range of [-1, 1]. The value range is divided into four intervals: [-1, -0.5), [-0.5, 0), [0, 0.5), [0.5, 1]. Each interval corresponds to a true modulation type, and the true modulation types include no modulation, amplitude modulation, frequency modulation, and pulse modulation.
[0020] Modulation waveform: The symbol is s(3), with a value range of [-1, 1]. The value range is divided into five intervals: [-1, -0.6), [-0.6, -0.2), [-0.2, 0.2), [0.2, 0.6), [0.6, 1). Each interval corresponds to a true waveform, and the waveforms include sine, triangle, square wave, positive ramp, and negative ramp.
[0021] Modulation frequency: The symbol is s(4), representing the normalized modulation frequency value, with a value range of [-1, 1]. Let the physical range of the true modulation frequency be According to s(4) and the physical range The true modulation frequency is:
[0022] The unit is MHz;
[0023] Modulation depth: The symbol is s(5), representing the normalized modulation depth value, with a value range of [-1, 1]; assuming the physical range of the true modulation depth is [a min , a max , according to s(5) and the physical range [a min , a max , the true modulation depth is:
[0024] The unit is %;
[0025] Frequency offset: The symbol is s(6), representing the normalized frequency offset value, with a value range of [-1, 1]; assuming the physical range of the true frequency offset is According to s(6) and the physical range The true frequency offset is:
[0026] The unit is MHz;
[0027] Pulse duty cycle: The symbol is s(7), representing the normalized pulse duty cycle value, with a value range of [-1, 1]; assuming the physical range of the true pulse duty cycle is [k min , k max , according to s(7) and the physical range [k min , k max , the true pulse duty cycle is:
[0028] The unit is %;
[0029] Pulse period: The symbol is s(8), representing the normalized pulse period value, with a value range of [-1, 1], assuming the physical range of the true pulse period is [T min , T max , according to s(8) and the physical range [T min , T max , the true pulse period is:
[0030] The unit is ms;
[0031] The two sensitive state parameters include:
[0032] Sensitivity, the symbol is s(9), representing whether the test item is sensitive under the current test waveform, with a value range of [-1, 1]; the value range is divided into two intervals: [-1, 0), [0, 1], indicating that the test item is sensitive and not sensitive under the current test waveform respectively;
[0033] The sensitivity threshold, denoted as s(10), represents the normalized sensitivity threshold value, with a value range of [-1, 1]. Let the physical range of the true sensitivity threshold be [P min ,P max . According to s(10) and the physical range [P min ,P max , the true sensitivity threshold is: The unit is dBm.
[0034] S2. Establish a reinforcement learning model based on the Twin Delayed Deep Deterministic Policy Gradient network;
[0035] In step S2, the Twin Delayed Deep Deterministic Policy Gradient network, i.e., the TD3 network, includes an online learning network and a target network. Each network contains an actor network and a Twin Q network inside, and the Twin Q network contains two critic networks;
[0036] Both the actor network and the critic network are fully connected neural networks with two hidden layers, and each hidden layer contains N hidden neurons; the input dimension of the actor network is 10, used to input the state, and the output dimension is 8, used to output the action. The input dimension of the critic network is 18, used to input the state + action, and the output dimension is 1, used to output the Q value; the activation function of the actor network is the hyperbolic tangent function, and the activation function of the critic network is the RELU function.
[0037] S3. Build an automated closed-loop test environment for the sensitivity threshold;
[0038] Step S3 includes:
[0039] The output signal of the radio frequency signal source is amplified by a power amplifier, and the power-amplified signal is connected to a directional coupler. The through-end of the directional coupler is connected to the DUT (Device Under Test) to directly inject the signal power into the DUT. The coupled-end of the directional coupler is connected to a signal receiver through an attenuator to measure the actual injected power, and the power amplifier is powered by a DC power supply. The DUT is fixed by a test bench to improve the repeatability of the experiment;
[0040] The core of the automatic closed-loop test is a control computer. It obtains the sensitive state of the DUT from the microcontroller and monitors the power at the current test frequency through the signal receiver. According to the information obtained, it makes an intelligent decision to control the radio frequency signal source to change the output signal waveform; the sensitivity excited by the new waveform is fed back to the control computer in real time to support the next intelligent decision, thus forming a closed loop to search for the most sensitive waveform that minimizes the sensitivity threshold of the DUT.
[0041] S4. Construct an electromagnetic sensitivity reward function and design a basic reward function and a shaping reward function;
[0042] Step S4 includes:
[0043] The reward function is used to guide the agent to achieve the goal. If the test item shows electromagnetic sensitivity, the agent will receive a reward. Considering that most waveforms do not trigger the electromagnetic sensitivity of the test item, which will lead to the problem of sparse rewards, the reward shaping method is adopted: by dividing the total reward into the basic reward r i b and the shaped reward r i s to achieve;
[0044] The basic reward function r i b will make the agent obtain a larger reward when the current state is close to s min where s min represents the most sensitive waveform state so far;
[0045] The reward mainly considers the carrier frequency s of the most sensitive waveform min (1), and the basic reward function at the i-th step is defined as:
[0046]
[0047] where * represents the non-normalized value, is the sensitivity threshold of the most sensitive waveform so far;
[0048] The state of the most sensitive waveform is unknown in the initial stage. Therefore, the following shaped reward function is designed to guide the agent to find a lower sensitivity threshold:
[0049]
[0050] where the square term helps the agent search for the action that reduces the maximum sensitivity threshold in one step, thus accelerating the search for the most sensitive waveform. When the sensitivity threshold at the i-th step is the lowest so far, the agent will be given a huge reward r i s = r i s + 100;
[0051] In addition, the following shaped reward function is used to accelerate convergence:
[0052] r i s = r i s + 1 if s i+1 (9)==1
[0053] r is = r i s -1 if s i+1 (9) == -1
[0054] r i s = r i s +5 if s i+1 (9) == 1 and s i (9) == -1
[0055] r i s = r i s -10 if s i+1 (9) == -1 and s i (9) == 1
[0056] Finally, the total reward is expressed as:
[0057] r i = r i b + μr i s
[0058] where μ is the reward shaping decay factor, starting from 1 and linearly decaying to 0 according to the total number of steps.
[0059] S5. Start the electromagnetic most sensitive waveform test and reinforcement learning training to obtain the most sensitive waveform.
[0060] The step S5 includes:
[0061] S501. Start the electromagnetic most sensitive waveform test and synchronously train the constructed reinforcement learning model. The training is divided into N ep games, and each game has N s steps. The environment needs to be reset at the beginning of each game: the output of the RF signal source is turned off, the state is reset to a random number uniformly distributed between -1 and 1, and the game end flag B done is reset to 0; within N warm steps is the warm-up stage. During this period, the action is a random number uniformly distributed between -1 and 1, rather than being inferred by the actor network in the TD3 network; in the non-warm-up stage, based on the current state s i use the actor network in the online network to infer the action; add Gaussian noise with a mean of 0 and a variance of E n as exploration noise to the action; update the current waveform parameter state based on the obtained action a i as follows:
[0062] si+1 (1:8) = s i (1:8) + βa i
[0063] where β represents the maximum step size of the action, and the updated waveform state s i+1 (1:8) needs to be clipped to the range of [-1, 1];
[0064] S502. By controlling the RF signal source, output a waveform with a carrier power of P max whose waveform parameters are the state s i+1 (1:8) represents the physical real value. Observe the sensitivity of the EUT after the signal output stays for 2 seconds. If the EUT is not sensitive, set s i+1 (9) = -1 and s i+1 (10) = 1;
[0065] Otherwise, set s i+1 (9) = 1, and perform a binary search within the range of [P min , P max ; The change in carrier power should stay for 2 seconds each time; Use the binary search method to change the carrier power a total of B n times to find the sensitivity threshold, with an accuracy of
[0066] Normalize the sensitivity threshold from [P min , P max to [-1, 1], and record it in s i+1 (10);
[0067] S503. Calculate the reward r i according to step S4. If the sensitivity threshold is the lowest so far, the current state should be recorded in s min ;
[0068] If this game has ended, set B done = 1, otherwise set B done = 0;
[0069] S504. Record s i , a i , r i , s i+1 and B done in the experience replay buffer, and enter the next loop. Until each game has ended, obtain the most sensitive waveform state s min ;
[0070] During the testing process, decay the learning rate l rdelay by l rγ every l r steps, and every N EnStep with γ En Decay the exploration noise variance E n , after T start steps, train the TD3 network every T train steps. The training method includes:
[0071] Obtain a batch of N batch samples s i , a i , r i , s i+1 from the experience replay buffer. Feed s i+1 into the actor network of the target network to get a i+1 . Feed s i+1 and a i+1 into the critic network of the target network to get Q * 1 and Q * 2. Take the minimum value of Q * 1 and Q * 2 as Q. Calculate the target Q value Q t = r i +(~B done )γQ, where ~B done represents the negation of B done , and γ represents the discount factor;
[0072] Feed s i and a i into the twin Q network of the online network to get Q1 and Q2. Calculate MSE Loss1 = Q t - Q1, MSELoss2 = Q t - Q2. The total loss loss = MSE Loss1 + MSE Loss2. Use loss as the loss function and update the parameters of the twin Q network in the online network through the backpropagation algorithm;
[0073] The parameters of the actor network in the online network are updated every T delay steps during training. The update process is as follows: Feed s i into the actor network of the online network to get Feed s i and into the twin Q network of the online network. Only obtain the value of the first critic network among them. Take its negative value as the loss function and update the parameters of the actor network in the online network through the backpropagation algorithm;
[0074] The weight parameters of the target network are updated every T delay steps during training through a soft update process. These parameters are merged with the parameters of the online network in the ratio of (1 - τ):τ.
[0075] The beneficial effects of the present invention are as follows: Compared with the electromagnetic sensitivity standard test method and the optimization test method based on genetic algorithms, the present invention can search for the most sensitive waveform with the lowest sensitivity threshold of the test object in the defined modulation signal set. The present invention also introduces relevant concepts of reinforcement learning into EMS testing, constructs a reinforcement learning framework based on the double-delayed deep deterministic policy gradient network, and designs a reward function for electromagnetic sensitivity testing. The present invention can determine the most sensitive waveform faster, and the sensitivity threshold of the determined most sensitive waveform is lower. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is the flowchart of the method of the present invention;
[0077] Figure 2 is a schematic diagram of the TD3 network framework for electromagnetic sensitivity testing;
[0078] Figure 3 is a layout diagram of the automatic closed-loop test environment for sensitivity threshold;
[0079] Figure 4 is a schematic diagram of the test result of the electromagnetic sensitivity threshold of the standard test method;
[0080] Figure 5 is a schematic diagram of the optimization process of the lowest electromagnetic sensitivity threshold of the test method optimized by GA;
[0081] Figure 6 is a schematic diagram of the optimization process of the lowest electromagnetic sensitivity threshold of the test method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0082] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0083] The present invention introduces relevant concepts of reinforcement learning into EMS testing, and proposes a method for testing the most electromagnetic sensitive waveform based on reinforcement learning, called EMS-RL, which can search for the most sensitive waveform with the lowest sensitivity threshold of the test object in the defined modulation waveform set. Compared with the current standard test method and the most sensitive waveform search method based on genetic algorithms, the present invention can determine the most sensitive waveform faster, and the threshold for stimulating the electromagnetic sensitivity of the test object is lower. Specifically:
[0084] As Figure 1 shown, a method for testing the most electromagnetic sensitive waveform based on reinforcement learning includes the following steps:
[0085] S1. Establish a reinforcement learning architecture for electromagnetic sensitivity testing, and construct a state space and an action space;
[0086] Taking the RF signal source in the EMS test as the agent in reinforcement learning, its goal is to learn how to take actions at each step to maximize the cumulative reward, thereby minimizing the sensitivity threshold of the device under test excited by the test waveform.
[0087] Taking the test system in the EMS test as the environment in reinforcement learning, it is used to receive the signal waveform output of the RF signal source and return the sensitive state of the device under test.
[0088] Taking the waveform parameters of the signal source and the sensitive state of the device under test in the EMS test as the state in reinforcement learning, there are a total of 8 waveform parameters and 2 sensitive state parameters, which constitute the state space in reinforcement learning. All state parameters are normalized to the range of [-1, 1]. The physical meaning and the physical range in the specific embodiment are shown in Table 1.
[0089] Table 1 State space parameter table of reinforcement learning in the present invention
[0090]
[0091] Taking the change of waveform parameters in the EMS test as the action in reinforcement learning, its purpose is to change the current state of the agent in order to obtain a greater reward. There are a total of 8 actions, corresponding to the change amounts of the first 8 states, which are the change amount of the carrier frequency, the change amount of the modulation type parameter, the change amount of the modulation waveform parameter, the change amount of the modulation frequency, the change amount of the modulation depth, the change amount of the frequency offset, the change amount of the pulse duty cycle, and the change amount of the pulse period. These 8 actions constitute the action space in reinforcement learning. All action parameters are normalized to the range of [-1, 1].
[0092] S2. Establish a reinforcement learning model based on the Twin Delayed Deep Deterministic Policy Gradient network;
[0093] Establish a reinforcement learning model based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) network, as Figure 2 shown, where the TD3 network mainly includes the following key networks:
[0094] Online learning network: It actively learns and updates parameters from the online interaction with the environment, which directly affects the next decision-making and learning progress.
[0095] Target network: In the temporal difference (TD) update process, both the actor network and the critic network use it to provide stable targets, which can reduce the risk of overfitting in the learning process and improve the convergence stability.
[0096] Dual Q-network: It uses two independent critic networks to independently estimate the Q-value (the potential reward after taking an action in a certain state), and uses the smaller one, so as to more accurately predict the Q-value.
[0097] Actor Network: Based on the current state, it determines the next action and directly optimizes the policy to obtain higher rewards.
[0098] Critic Network: It evaluates the potential rewards of taking actions in a given state and guides the agent to take actions with greater rewards by accurately estimating the Q-value.
[0099] Both the actor network and the critic network are fully connected neural networks with two hidden layers, and each hidden layer contains N hidden neurons. The input dimension of the actor network is 10 (state), and the output dimension is 8 (action), while the input dimension of the critic network is 18 (state + action), and the output dimension is 1 (Q-value). The activation function of the actor network is the hyperbolic tangent function, and the activation function of the critic network is the RELU function.
[0100] S3. Build an automated closed-loop test environment for sensitivity thresholds;
[0101] Compared with traditional electromagnetic sensitivity test methods, the method proposed in the present invention needs to run in an automated closed-loop test environment for sensitivity thresholds.
[0102] The test layout can be based on the direct power injection method layout or the large current injection method layout. Taking the test layout of the direct power injection method as an example, as Figure 3 shown. The output signal of the RF signal source is amplified by a power amplifier. The power-amplified signal is connected to a directional coupler. The through-end of the directional coupler is connected to the device under test to directly inject the signal power into the device under test. The coupled-end of the directional coupler is connected to a signal receiver through an attenuator to measure the actual injected power. The DC power supply is used to power the power amplifier. The test bench is used to fix the device under test to improve the repeatability of the experiment.
[0103] The core of the automated closed-loop test is a control computer. It obtains the sensitive state of the device under test from the microcontroller and monitors the power at the current test frequency through the signal receiver. According to the information obtained, it makes intelligent decisions using the method proposed in the present invention and controls the RF signal source to change the output signal waveform. The sensitivity excited by the new waveform is fed back to the control computer in real time to support the next intelligent decision, thus forming a closed loop to search for the most sensitive waveform that minimizes the sensitivity threshold of the device under test.
[0104] S4. Construct an electromagnetic sensitivity reward function, and design a basic reward function and a shaping reward function;
[0105] The reward function is used to guide the agent to achieve the goal. If the test article shows electromagnetic sensitivity, the agent will receive a reward. However, most waveforms do not trigger the electromagnetic sensitivity of the test article, which leads to the problem of sparse rewards. Therefore, the reward shaping method needs to be adopted. This can be achieved by dividing the total reward into a basic reward r i b and a shaped reward r i s . The reward shaping method can alleviate the sparse reward problem and accelerate the convergence speed. The design ideas of the basic reward function and the shaped reward function are as follows.
[0106] In the present invention, the purpose of the proposed method is to determine the most sensitive waveform. Therefore, the basic reward function r i b will give the agent a large reward when the current state is close to s min , where s min represents the state of the most sensitive waveform so far. The reward mainly considers the carrier frequency s min (1) of the most sensitive waveform. The basic reward function at the i-th step is defined as:
[0107]
[0108] where * represents the non-normalized value, is the sensitivity threshold of the most sensitive waveform so far.
[0109] However, the agent does not know the state of the most sensitive waveform at the beginning, so the following shaped reward function is designed to guide the agent to find a lower sensitivity threshold.
[0110]
[0111] where the square term helps the agent search for actions that reduce the maximum sensitivity threshold in one step, thus accelerating the search for the most sensitive waveform. When the sensitivity threshold at the i-th step is the lowest so far, a huge reward r i s = r i s + 100 will be given to the agent. In addition, there are some other shaped reward functions to accelerate convergence:
[0112] r i s = r i s + 1 if s i+1 (9)==1,
[0113] r is = r i s -1 if s i+1 (9) == -1,
[0114] r i s = r i s +5 if s i+1 (9) == 1 and s i (9) == -1,
[0115] r i s = r i s -10 if s i+1 (9) == -1 and s i (9) == 1.
[0116] Finally, the total reward is expressed as:
[0117] r i = r i b + μr i s ,
[0118] where μ is the reward shaping decay factor. μ starts from 1 and linearly decays to 0 according to the total number of steps. The decay of reward shaping helps the agent find a way to reduce the sensitivity threshold in the early stage and avoid relying on reward shaping in the later stage.
[0119] S5. Start the electromagnetic most sensitive waveform test and reinforcement learning training to obtain the most sensitive waveform.
[0120] Table 2 Parameter table of the test method based on reinforcement learning proposed in the present invention
[0121]
[0122]
[0123] Start the electromagnetic most sensitive waveform test and synchronously train the constructed reinforcement learning model. The relevant parameters are shown in Table 2. The training is divided into N ep games, and each game has N s steps. The environment needs to be reset at the beginning of each game: the output of the RF signal source is turned off, the state is reset to a random number uniformly distributed between -1 and 1, and the game end flag B done is reset to 0. N warmWithin the first few steps is the warm-up stage. During this period, actions are randomly generated with a uniform distribution, rather than being inferred by the actor network in the TD3 network. In the non-warm-up stage, based on the current state s i Use the actor network in the online network to infer the action. Add Gaussian noise with a mean of 0 and a variance of E n as exploration noise to the action. The current waveform parameter state is updated based on the obtained action a i as follows:
[0124] s i+1 (1:8) = s i (1:8) + βa i
[0125] where β represents the maximum step size of the action. The updated waveform state s i+1 (1:8) needs to be clipped to the range of [-1, 1].
[0126] By controlling the RF signal source, output a waveform with a carrier power of P max , and the waveform parameters are the true values corresponding to the values represented by the state s i+1 (1:8). Observe the sensitivity of the EUT after the signal output stays for 2 seconds. If the EUT is not sensitive, set s i+1 (9) = -1 and s i+1 (10) = 1. Otherwise, set s i+1 (9) = 1, and perform a binary search within the range of [P min , P max , where P min = -50 dBm and P max = -10 dBm. Each change in carrier power should stay for 2 seconds. Use the binary search method to change the carrier power a total of B n times to find the sensitivity threshold, and its accuracy can reach Normalize the sensitivity threshold from [-50, -10] to [-1, 1] and record it in s i+1 (10). Then, calculate the reward r i according to the method detailed in step S4. If the sensitivity threshold is the lowest so far, the current state should be recorded in s min . If this game has ended, set B done = 1, otherwise set B done = 0. Finally, record s i , a i , r i , s i+1 and B done in the experience replay buffer. Finally, enter the next loop until each game has ended. During the test, every l rdelay steps with lrγ Decay learning rate l r , every N En steps with γ En decay the exploration noise variance E n , after T start steps, train the TD3 network every T train steps. The specific training method includes:
[0127] Obtain a batch of N batch samples s i , a i , r i , s i+1 from the experience replay buffer. Feed s i+1 into the actor network of the target network to get a i+1 . Feed s i+1 and a i+1 into the critic network of the target network to get Q * 1 and Q * 2. Take the minimum value of Q * 1 and Q * 2 as Q. Calculate the target Q value Q t = r i +(~B done )γQ, where ~B done represents the negation of B done , and γ represents the discount factor;
[0128] Feed s i and a i into the twin Q network of the online network to get Q1 and Q2. Calculate MSE Loss1 = Q t - Q1, MSELoss2 = Q t - Q2. The total loss loss = MSE Loss1 + MSE Loss2. Use loss as the loss function and update the parameters of the twin Q network in the online network through the backpropagation algorithm;
[0129] The parameters of the actor network in the online network are updated every T delay steps during training. The update process is as follows: Feed s i into the actor network of the online network to get Feed s i and into the twin Q network of the online network, and only obtain the value of the first critic network among them. Take its negative value as the loss function and update the parameters of the actor network in the online network through the backpropagation algorithm;
[0130] The weight parameters of the target network are updated through a soft update process every T delayThese parameters are merged with the parameters of the online network in the ratio (1-τ):τ).
[0131] The technical effects of the present invention are described in detail below in conjunction with actual measurement experiments.
[0132] 1. Test object:
[0133] Using the direct power injection layout, the ADS1110 analog-to-digital conversion (A / D) chip and the OPA2192 operational amplifier (OPA) chip were selected as the test products for testing. The test product chips were installed in a special PCB circuit board and fixed on the test bench. In the circuit board, the test interference signal was injected into the VCC pin of the chip through a 6.8nF isolation capacitor.
[0134] 2. Obtain the status of the test product:
[0135] The AD chip is used in the voltage acquisition circuit, and the microcontroller reads the AD chip control register data through the I2C bus and then forwards it to the control computer. The OPA chip is used in a second-order Butterworth low-pass filter circuit with a gain of 1, and the microcontroller reads the output voltage data and forwards it to the control computer.
[0136] 3. Sensitivity criteria:
[0137] For AD chips, if the read data is inconsistent with the written data, the test product is considered to be sensitive. For OPA chips, if the output DC voltage differs by 0.1V from the output when not disturbed, the test product is considered to be sensitive.
[0138] 4. Test result analysis
[0139] The test methods using continuous wave waveform (CW), GJB 151B CS114 standard test waveform (PM-S) and IEC 61000-4-3 standard test waveform (AM-S) are selected as comparative test methods, and the genetic algorithm (GA) is selected as the comparative optimization method of the reinforcement learning method proposed in the present invention.
[0140] In the experiment using the standard test method, the carrier frequency of the test interference signal was swept from 20MHz to 400MHz by 1MHz. The sensitivity thresholds of the AD chip and OPA chip were measured to change with frequency as shown in the following table. Figure 4 (a) and Figure 4 (b) as shown.
[0141] In the test method optimized by genetic algorithm, the first 8 states (waveform parameters) are used as the optimization variables, and the sensitivity threshold is used as the optimization target value. The population size is 50, the maximum number of iterations is 20, the mutation probability is 0.1, and the selection operator is a tournament operator with a size of 30. The optimization processes of the lowest electromagnetic sensitivity thresholds of the AD chip and the OPA chip are respectively as Figure 5 (a) and (b) shown.
[0142] Using the test method of the present invention, the specific parameters are shown in Table 2. The optimization processes of the lowest electromagnetic sensitivity thresholds of the AD chip and the OPA chip are respectively as Figure 6 (a) and (b) shown.
[0143] The test results show that the lowest sensitivity thresholds of the standard test waveforms are -30dBm and -24dBm respectively. Using genetic algorithm optimization to find the most sensitive waveforms, the obtained lowest sensitivity thresholds are -35.20dBm and -26.37dBm respectively. Using the method proposed in the present invention to find the most sensitive waveforms, the obtained lowest sensitivity thresholds are -35.20dBm and -28.71dBm respectively.
[0144] The test results of the method of the present invention show that for the AD chip, the most sensitive waveform is a square wave amplitude modulation signal with a carrier frequency of 20MHz, a modulation frequency of 34kHz, and a modulation depth of 100%; for the OPA chip, the most sensitive waveform is a sine frequency modulation signal with a carrier frequency of 82.49MHz, a modulation frequency of 19kHz, and a frequency deviation of 175kHz. These most sensitive waveforms are 5.20dB and 4.71dB lower than the lowest sensitivity thresholds of the standard test waveforms respectively, and the result of the OPA chip is 2.34dB lower than that of GA.
[0145] In terms of test efficiency, the present invention found the most sensitive waveforms at the 133rd and 206th steps respectively, while the GA method found the most sensitive waveforms after the 12th and 20th iterations. Considering that there are 20 population sizes in one iteration, converting to the number of test steps is 240 steps and 1000 steps. The number of steps required for testing is much higher than that of the present invention, which proves that the test method of the present invention is more efficient.
[0146] The above is the preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention should be within the protection scope of the appended claims of the present invention.
Claims
1. An electromagnetic most sensitive waveform testing method based on reinforcement learning, characterized in that: It includes the following steps: S1. Establish a reinforcement learning architecture for electromagnetic susceptibility testing, and construct a state space and an action space; The step S1 includes: Regarding the RF signal source in the EMS test as the agent in reinforcement learning, whose goal is to learn how to take actions at each step to maximize the cumulative reward and minimize the sensitivity threshold of the test item excited by the test waveform; Regarding the test system in the EMS test as the environment in reinforcement learning, which is used to receive the signal waveform output of the RF signal source and return the sensitive state of the test item; Regarding the waveform parameters of the signal source and the sensitive state of the test item in the EMS test as the states in reinforcement learning. There are a total of 8 waveform parameters and 2 sensitive state parameters, which together constitute the state space in reinforcement learning. Each parameter in the state space is a state, with a total of 10 states, and the value of each state is within the range of [-1, 1]; Regarding the change in waveform parameters in the EMS test as the action in reinforcement learning, the purpose of which is to change the current state of the agent in order to obtain a greater reward; there are a total of 8 actions, corresponding to the change amounts of 8 waveform parameters, namely the change amount of the carrier frequency, the change amount of the modulation type parameter, the change amount of the modulation waveform parameter, the change amount of the modulation frequency, the change amount of the modulation depth, the change amount of the frequency offset, the change amount of the pulse duty cycle, and the change amount of the pulse period; the values of these 8 actions are all within the range of [-1, 1], constituting the action space in reinforcement learning; S2. Establish a reinforcement learning model based on the twin-delayed deep deterministic policy gradient network; S3. Build an automated closed-loop test environment for the sensitivity threshold; S4. Construct an electromagnetic sensitivity reward function, and design a basic reward function and a shaping reward function; The step S4 includes: The reward function is used to guide the agent to achieve the goal. If the test item exhibits electromagnetic sensitivity, the agent will receive a reward. Considering that most waveforms do not trigger the electromagnetic sensitivity of the test item, which will lead to the problem of sparse rewards, a reward shaping method is adopted: by dividing the total reward into a basic reward function and a shaping reward function to achieve it; Basic reward function will cause the agent to receive a larger reward when the current state is close to s min where s min represents the most sensitive waveform state so far; The reward mainly considers the carrier frequency s of the most sensitive waveform min (1), The basic reward function for the i-th step is defined as: where * represents a non-normalized value, is the sensitivity threshold of the most sensitive waveform so far; The state of the most sensitive waveform is unknown in the initial stage. Therefore, design the following shaping reward function to guide the agent to find a lower sensitivity threshold: The square term helps the agent search for actions that reduce the maximum sensitivity threshold in one step, thus accelerating the search for the most sensitive waveform. When the sensitivity threshold at the i-th step is the lowest so far, a huge reward will be given to the agent Finally, the total reward is expressed as: In the formula, μ is the reward shaping decay factor, which starts from 1 and linearly decays to 0 according to the total number of steps; S5. Start the electromagnetic most sensitive waveform test and reinforcement learning training to obtain the most sensitive waveform.
2. The electromagnetic most sensitive waveform testing method based on reinforcement learning according to claim 1, wherein: The 8 waveform parameters include: Carrier frequency: The symbol is s(1), representing the normalized carrier frequency value, with a value range of [-1, 1]; assuming the physical range of the true carrier frequency is [f min , f max , according to s(1) and the physical range [f min , f max , the true carrier frequency is: The unit is MHz; Modulation type: The symbol is s(2), and the value range is [-1, 1]. The value range is divided into four intervals: [-1, -0.5), [-0.5, 0), [0, 0.5), [0.5, 1]; each interval corresponds to a real modulation type, and the real modulation types include no modulation, amplitude modulation, frequency modulation, and pulse modulation; Modulation waveform: The symbol is s(3), and the value range is [-1, 1]. The value range is divided into five intervals: [-1, -0.6), [-0.6, -0.2), [-0.2, 0.2), [0.2, 0.6), [0.6, 1); each interval corresponds to a real waveform, and the waveforms include sine, triangle, square wave, positive ramp, and negative ramp; Modulation frequency: The symbol is s(4), representing the normalized modulation frequency value, with a value range of [-1, 1]; assuming the physical range of the true modulation frequency is According to s(4) and the physical range The true modulation frequency is: The unit is MHz; Modulation depth: The symbol is s(5), representing the normalized modulation depth value, with a value range of [-1, 1]; assuming the physical range of the true modulation depth is [a min , a max , according to s(5) and the physical range [a min , a max , the true modulation depth is: Unit: %; Frequency offset: The symbol is s(6), representing the normalized frequency offset value, with a value range of [-1, 1]; assume the physical range of the true frequency offset is According to s(6) and the physical range The true frequency offset is: Unit: MHz; Pulse duty cycle: The symbol is s(7), representing the normalized pulse duty cycle value, with a value range of [-1, 1]; assuming the physical range of the true pulse duty cycle is [k min , k max , according to s(7) and the physical range [k min , k max , the true pulse duty cycle is: Unit: %; Pulse period: The symbol is s(8), representing the normalized pulse period value, with a value range of [-1, 1]. Let the physical range of the true pulse period be [T min , T max . According to s(8) and the physical range [T min , T max , the true pulse period is: Unit: ms; The 2 sensitive state parameters include: Sensitivity, denoted by s(9), indicates whether the test item is sensitive under the current test waveform, with a value range of [-1, 1]. The value range is divided into two intervals: [-1, 0) and [0, 1], indicating that the test item is sensitive and not sensitive under the current test waveform, respectively. The sensitivity threshold, denoted as s(10), represents the normalized sensitivity threshold value, with a value range of [-1, 1]. Let the physical range of the true sensitivity threshold be [P min , P max . According to s(10) and the physical range [P min , P max , the true sensitivity threshold is: The unit is dBm.
3. The electromagnetic most sensitive waveform testing method based on reinforcement learning according to claim 1, characterized in that: In step S2, based on the twin-delayed deep deterministic policy gradient network, i.e., the TD3 network, which includes an online learning network and a target network. Each network contains an actor network and a twin Q-network internally, and the twin Q-network contains two critic networks. Both the actor network and the critic network are fully connected neural networks with two hidden layers, and each hidden layer contains N hidden neurons; the input dimension of the actor network is 10 for inputting states, the output dimension is 8 for outputting actions, while the input dimension of the critic network is 18 for inputting states + actions, and the output dimension is 1 for outputting Q-values; the activation function of the actor network is the hyperbolic tangent function, while the activation function of the critic network is the RELU function.
4. The electromagnetic most sensitive waveform testing method based on reinforcement learning according to claim 1, characterized in that: Step S3 includes: The output signal of the radio frequency signal source is amplified by a power amplifier. The power-amplified signal is connected to a directional coupler. The through-end of the directional coupler is connected to the test item to directly inject the signal power into the test item. The coupled-end of the directional coupler is connected to a signal receiver through an attenuator to measure the actual injected power. The power amplifier is powered by a DC power supply, and the test item is fixed by a test bench to improve the repeatability of the experiment. The core of the automatic closed-loop test is a control computer. It obtains the sensitive state of the test item from the microcontroller and monitors the power at the current test frequency through the signal receiver. Based on the information obtained, it makes intelligent decisions to control the radio frequency signal source to change the output signal waveform. The sensitivity excited by the new waveform is fed back to the control computer in real time to support the next intelligent decision, thus forming a closed loop to search for the most sensitive waveform that minimizes the sensitivity threshold of the test item.
5. The electromagnetic most sensitive waveform testing method based on reinforcement learning according to claim 2, wherein: In step S4, there is also the following shaping reward function to accelerate convergence:
6. The electromagnetic most sensitive waveform testing method based on reinforcement learning according to claim 5, characterized in that: Step S5 includes: S501. Start the electromagnetic most sensitive waveform test and synchronously train the constructed reinforcement learning model. The training is divided into N ep rounds, and each round has N s steps. The environment needs to be reset at the beginning of each round: the output of the RF signal source is turned off, the state is reset to a random number uniformly distributed between -1 and 1, and the end-of-round flag B done is reset to 0; within N warm steps is the warm-up phase. During this period, the action is a random number uniformly distributed between -1 and 1, rather than being inferred by the actor network in the TD3 network; in the non-warm-up phase, based on the current state s i use the actor network in the online network to infer the action; add Gaussian noise with a mean of 0 and a variance of E n as exploration noise to the action; update the current waveform parameter state based on the obtained action a i as follows: s i+1 (1:8) = s i (1:8) + βa i where β represents the maximum step size of the action, and the updated waveform state s i+1 (1:8) needs to be clipped to the range of [-1, 1]; S502. Output a waveform with a carrier power of P by controlling the RF signal source. max The waveform parameters are the state s. i+1 (1:8) represents the physical true value. After the signal output stays for 2 seconds, observe the sensitivity of the EUT. If the EUT is not sensitive, set s i+1 (9) = -1 and s i+1 (10) = 1; Otherwise, set s i+1 (9) = 1, and perform a binary search within [P min , P max ; the change in carrier power should stay for 2 seconds each time; use the binary search method to change the carrier power a total of B n times to find the sensitivity threshold, with an accuracy of Normalize the sensitivity threshold from [P min , P max to [-1, 1] and record it in s i+1 (10); S503. Calculate the reward r according to step S4 i , if the sensitivity threshold is the lowest so far, the current state should be recorded in s min ; If this game has ended, set B done to 1, otherwise set B done to 0; S504. Record s i , a i , r i , s i+1 and B done into the experience replay buffer, and enter the next loop. Until each game ends, obtain the most sensitive waveform state s min ; During the test, every l rdelay Step 1 rγ Decay learning rate l r , every N En Step γ En Attenuated exploration noise variance E n , in T start After the step, every T train The TD3 network is trained once.
7. The electromagnetic most sensitive waveform testing method based on reinforcement learning according to claim 6, characterized in that: Every T train The process of training the TD3 network once every train steps includes: Fetch a batch of N from the experience replay buffer batch Sample s i 、a i 、r i 、s i+1 , send s i+1 into the actor network of the target network to obtain a i+1 , send s i+1 and a i+1 into the critic network of the target network to obtain and Take and The minimum value of is used as Q, and the target Q value Q is calculated according to the Bellman formula t = r i +(~B done )γQ, where ~B done represents the negation of B done and γ represents the discount factor; Send s i and a i into the double Q-network of the online network to obtain Q1 and Q2, calculate MSE Loss1 = Q t - Q1, MSE Loss2 = Q t - Q2, the total loss loss = MSE Loss1 + MSE Loss2, use loss as the loss function, and update the parameters of the double Q-network in the online network through the backpropagation algorithm; The parameters of the actor network in the online network are updated every T steps during training. The update process is as follows: Feed s into the actor network of the online network to obtain delay Feed s i into the actor network of the online network to obtain Feed s i and into the double Q-network of the online network, and only obtain the value of the first critic network among them. Use its negative value as the loss function, and update the parameters of the actor network in the online network through the backpropagation algorithm. The weight parameters of the target network are updated once every T steps during training through a soft update process, and these parameters are merged with the parameters of the online network in the ratio of (1 - τ):τ. delay