Sensor hardware Trojan horse detection method and system based on signal injection
Through signal injection and reinforcement learning, the physical signal parameters of the sensor are optimized, combined with 3σ outlier detection, the accuracy and efficiency of sensor hardware Trojan detection are solved, and the safety and reliability of the sensor are ensured.
Patent Information
- Application Number
- CN202510333107.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The prior art is difficult to effectively identify and detect hardware Trojans embedded in sensors, especially in analog signal processing environments, resulting in sensors that may fail or provide false readings, affecting the reliability and efficiency of safety-critical applications.
Using a signal injection method, reinforcement learning is used to optimize the parameters of sound wave, laser and electromagnetic signals, and by detecting the sensor's response characteristics such as node voltage, power consumption and output data, combined with the 3σ outlier detection model, we can judge whether the sensor is implanted in a Trojan in real time.
It realizes efficient and accurate detection of sensor hardware Trojans, reduces false alarm rates, improves detection reliability and efficiency, and is suitable for a variety of sensor products to ensure their safety and functionality in different application scenarios.
Smart Images

Figure CN120337216A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sensor anomaly detection, and particularly to a method and system for detecting sensor hardware trojans based on signal injection. Background Art
[0002] Modern sensor technology and its role in cyber-physical systems (CPS) are becoming increasingly important. With technological advancements, sensors have not only become more sensitive and accurate but also integrated more complex processing capabilities to support intelligent decision-making. However, this complexity has also provided new attack vectors for malicious actors, namely "sensor trojans". A sensor trojan is a malicious entity hidden inside a sensor that can affect the behavior or output of the sensor and may be embedded in any component of the sensor. Once activated, the trojan can cause the sensor to fail, provide false readings, or stop working altogether, posing a serious threat to safety-critical applications that rely on accurate sensing data.
[0003] Since multiple parties involved in the design and manufacturing of modern sensors may have access to design details and production processes, there are opportunities to embed Trojan horses at different stages of production. These trojans may be hardware tampering or hidden code in embedded software and may be triggered under specific conditions, such as when receiving an input signal in a preset pattern. Traditional digital circuit detection methods are not always effective in identifying these Trojan horses in the analog domain because the latter often utilize the characteristics of analog signal processing.
[0004] With the increasing dependence of intelligent devices (such as cars and drones) on high-precision sensing data, ensuring the security and reliability of sensors has become a crucial task. Methods for detecting sensor trojans through physical signals must consider variables in the actual operating environment and be able to distinguish fluctuations under normal operating conditions from potential malicious activities. Such a detection mechanism requires a deep understanding of sensor performance to ensure that even the most hidden trojans are not missed. In addition, this method should also minimize the false alarm rate to maintain the overall reliability and efficiency of the system.
[0005] In summary, it is urgent to solve the huge threat of sensor backdoors and improve the detection mechanism for sensor backdoors. Summary of the Invention
[0006] In view of the above problems, the present invention proposes a method and system for detecting sensor hardware trojans based on signal injection, which focuses on analyzing the real physical characteristics such as the node voltage level, power consumption, output, etc. of the sensor response to discover abnormal behaviors (or signs of hidden trojans). The sensor trojans are covertly triggered by injecting remote physical signals such as sound waves, ultrasonic waves, electromagnetic signals, lasers, etc., and reinforcement learning is used to automatically adjust the physical signal parameters and optimize the transmission signal strategy to detect trojans more efficiently.
[0007] The technical solution proposed by the present invention is as follows:
[0008] In a first aspect, the present invention proposes a method for detecting sensor hardware trojans based on signal injection, including the following steps:
[0009] (1) Place a clean sensor without a trojan in a target control environment, and collect the output characteristic data of the sensor. The output characteristic data includes response time, node voltage, power consumption, and output data.
[0010] (2) Train a policy model using the reinforcement learning method. Use the sensor and the physical signal transmitting device to form an environment. Use the output characteristics of the sensor, the current transmission parameters of the physical signal transmitting device, and the corresponding transmitted physical signal type to form the state space, and use the adjustable parameters of the physical signal transmitting device to form the action space. During the training process, the policy model adjusts the transmission parameters of the physical signal transmitting device according to the current state. The environment generates a new physical signal under the current transmission parameters and acts on the sensor. After the sensor responds, the environment returns a new state and a reward.
[0011] (3) In the test phase, select the physical signal transmitting device, place the sensors to be tested and the clean sensors of the same batch in the target control environment, and give the initial parameters for the swept-frequency test, including: start frequency, cut-off frequency, and adjustable parameters. Use the trained policy network to optimize the adjustable parameters of the signal transmitting device in real time according to the current state and the environment. Based on the output characteristic data of the sensors to be tested and the clean sensors, judge whether the response of the sensor is normal based on the 3σ outlier detection method.
[0012] If it is normal, the sensor to be tested is not implanted with a trojan; otherwise, the sensor to be tested is implanted with a trojan.
[0013] Further, the physical signal transmitting device is two or more of an acoustic wave signal transmitting device, a laser signal transmitting device, and an electromagnetic signal transmitting device. During the test, traverse different physical signal transmitting devices. Only when all are normal, it is judged that the sensor to be tested is not implanted with a trojan.
[0014] Further, the adjustable parameters of the physical signal transmitting device include frequency step size, time step size, amplitude, and signal duration.
[0015] Further, during the reinforcement learning process, when an abnormal output characteristic of the sensor is detected, a positive reward is given. When no abnormal output characteristic of the sensor is detected, or when the sensor is damaged due to unreasonable setting of the transmission signal parameters, a negative reward is given.
[0016] Further, the policy model adopts the DDPG model.
[0017] Further, the start frequency and the cut-off frequency are preset fixed values.
[0018] In a second aspect, the present invention provides a sensor hardware Trojan detection system based on signal injection for implementing the above-mentioned sensor hardware Trojan detection method.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0020] Based on reinforcement learning, the present invention optimizes the parameters of various physical signals such as sound waves, laser, and electromagnetic signals to conduct safety tests on sensors. This multi-dimensional detection method can effectively reveal the responses of sensors under different signal excitations, thereby accurately identifying potential hidden hardware Trojans.
[0021] Based on the 3σ outlier detection model to determine whether the sensor characteristics are abnormal, it is possible to automatically judge the response of the sensor according to the preset standard, thereby reducing the false alarm rate and improving the detection accuracy, ensuring the reliability of the detection results.
[0022] The general and high-efficiency detection process of the present invention makes it applicable to various types of sensor products, and reduces manual intervention through automated testing, which not only improves the detection efficiency but also ensures the safety and functionality of sensors in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a control block diagram of a sensor hardware Trojan detection method based on signal injection proposed by the present invention;
[0024] Figure 2 is a flowchart of a sensor hardware Trojan detection method based on signal injection proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following further elaborates and explains the present invention in conjunction with specific embodiments. The embodiments are only demonstrations of the disclosed content and do not delimit the scope of limitation. Without conflict, the technical features of each embodiment of the present invention can be combined accordingly.
[0026] The drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0027] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all steps. For example, some steps can be further decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0028] The present invention proposes a method for detecting sensor hardware Trojans based on signal injection, focusing on using a neural network constructed by reinforcement learning to optimize physical signal characteristics to detect these potential threats. As Figure 1 shown, the host computer, as the main control unit, receives instructions using a serial port or Ethernet communication protocol and synchronously generates control signals to drive the signal transmitting device. The signal transmitting device includes acoustic wave, laser, and electromagnetic emission modules, which respectively transmit physical signals to the sensor array. The sensor part consists of the sensor to be tested and a reference sensor (clean sensor). Both of them receive the interference of physical signals in a controllable environment, collect the output characteristic data of the sensors, and perform data analysis through collection by the host computer. This method can carefully analyze the electrical signal characteristics generated when the sensor works, including but not limited to voltage level, current mode, frequency response, and noise distribution, etc., so as to capture any abnormal behavior or unexpected signal changes.
[0029] As Figure 2 shown, a method for detecting sensor hardware Trojans based on signal injection proposed by the present invention mainly includes the following steps:
[0030] S1, Data preparation
[0031] Place the sensor to be detected in a controlled environment and connect it to the sensor data acquisition platform to ensure a stable connection between the sensor and the data acquisition platform for accurate data collection.
[0032] Start the data acquisition platform, set initial parameters to match the specifications and expected working environment of the sensor. Collect the response time, node voltage, power consumption, and output data of a clean sensor (without Trojans), and keep them unchanged for subsequent tests. Repeat the collection of 100 groups of data during this process, and these data will be used as the benchmark for subsequent tests for comparison and analysis.
[0033] S2: Construct a reinforcement learning model.
[0034] 1. Environment and policy design
[0035] In the process of detecting sensor Trojans using reinforcement learning, it is first necessary to define the environment, state space, and action space.
[0036] In a specific implementation of the present invention, the environment consists of a sensor and a physical signal transmitting device, and the physical signal transmitting device is any one of an acoustic signal transmitting device, a laser signal transmitting device, and an electromagnetic signal transmitting device; the physical signal transmitting device operates at a set start frequency, cut-off frequency, frequency step, time step, amplitude, and signal duration, from the start frequency to the cut-off frequency; here, the start frequency and the cut-off frequency are preset parameters, and the frequency step, time step, amplitude, and signal duration are adjusted according to the interaction process using a reinforcement learning algorithm. Here, specific amplitude signals and signal durations are more likely to be used as Trojan trigger conditions.
[0037] The state space includes: the output characteristics of the sensor (including response time, node voltage, power consumption, and output data) and the current transmission parameters of the physical signal transmitting device and the corresponding transmitted physical signal type. Among them, in the output characteristics of the sensor, the response time reflects the speed at which the sensor responds to external stimuli, the node voltage is directly related to the internal electrical state of the sensor, the power consumption reflects the energy consumption during the operation of the sensor, and the output data is the intuitive presentation of the physical quantity sensed by the sensor. The transmitted physical signal type is related to the physical signal transmitting device, that is, the acoustic signal type, the laser signal type, and the electromagnetic signal type. In this embodiment, one transmitted physical signal type is fixed each time during testing, such as electromagnetic or laser.
[0038] The action space is the adjustable parameters of the physical signal transmitting device, such as frequency step, time step, amplitude, and signal duration. Among them, changing the frequency step can adjust the frequency change amplitude of the transmitted signal, and the time step controls the fine adjustment degree of the signal in the time dimension; the amplitude determines the intensity of the signal and directly affects the interaction effect between the signal and the sensor; the signal duration specifies the duration of the transmitted signal, and different duration settings may trigger different responses from the sensor.
[0039] The above definitions lay the foundation for subsequent policy learning.
[0040] 2. Reinforcement Learning Algorithm and Reward Mechanism
[0041] The interaction process with the environment: The policy network selects an action (adjusts the transmission parameters) according to the current state, the environment generates a new physical signal according to the action and acts on the sensor, and after the sensor responds, the environment returns a new state and a reward.
[0042] Design positive rewards: Detect abnormal output characteristics of the sensor, such as abnormal response time, node voltage fluctuations, abnormal power consumption, or output data deviating from expectations.
[0043] Design negative rewards: The output features of the sensor are not detected as abnormal, and the transmission signal parameter settings are unreasonable (such as excessive signal strength causing damage to the sensor).
[0044] Selecting a policy network suitable for the continuous action space is one of the key steps. In this embodiment, the policy network adopts the Deep Deterministic Policy Gradient (DDPG) algorithm, which combines the advantages of the Deep Q-Network (DQN) and the policy gradient method.
[0045] Combined with the above reward mechanism, the algorithm can learn how to adjust the parameters of the physical signal transmission device to maximize the probability of detecting a sensor trojan.
[0046] 3. In the training phase, the signal transmission device adjusts the parameters of the physical signal transmission device according to the current policy, observes the output features of the sensor and calculates the reward, and then updates the policy network according to the reward.
[0047] S3: Testing phase
[0048] In the testing phase, tests are randomly conducted under the acoustic wave signal transmission device, the laser signal transmission device, and the electromagnetic signal transmission device. Under each type of test:
[0049] Place the sensors to be tested and clean sensors of the same batch in a controlled environment, connect them to the sensor data acquisition platform, and start the corresponding signal transmission device. Given the initial parameters for the sweep frequency test, including: start frequency, cut-off frequency, frequency step, time step, amplitude, signal duration; where the start frequency and cut-off frequency are fixed values, and the frequency step, time step, amplitude, and signal duration are adjustable parameters for initialization.
[0050] Use the trained policy network to optimize the adjustable parameters of the signal transmission device in real time according to the current state and environment, observe the output features of the sensor, and compare the differences in the output features between the sensors to be tested and the clean sensors. Based on the 3σ outlier detection method, determine whether the response of the sensor to be tested is within the normal range.
[0051] It should be noted that what the policy network optimizes in real time here are the frequency step, time step, amplitude, and signal duration, and the start frequency and cut-off frequency will not be changed. Therefore, the testing process will follow from the start frequency to the cut-off frequency step by step and finally automatically stop the test.
[0052] The described 3σ outlier detection method is based on the normal distribution hypothesis in statistics. By calculating the mean and standard deviation of the data, it determines whether the data points deviate from the normal range. Here, by comparing the data of the sensor under test with that of the clean sensor, abnormal behaviors are identified to prevent misjudgments caused by the vulnerability of the sensor itself. The clean sensor data and the data of the sensor under test are collected under the same test parameters. For the response time, node voltage, power consumption, and output data, it is respectively determined whether they are abnormal. Taking a certain output feature (such as output data) as an example, assume the collected output data is:
[0053] Clean sensor data: X_normal = {x_1, x_2, x_n};
[0054] Test sensor data: Y_test = {y_1, y_2, y_n}.
[0055] Calculate the mean and standard deviation of the clean sensor data. Based on the 3σ principle, the range of normal data can be obtained; determine whether each data point of the sensor under test falls within the normal range. If so, it is determined that the sensor is normal under the output feature of output data. Otherwise, it is judged as abnormal, and it is considered that the sensor is implanted with a Trojan. The judgment of the remaining output features is the same. Only when all output features are normal can the test of the next signal transmitting device be entered.
[0056] If all of the above three signal transmitting devices are normal after traversal, it is determined that the sensor is not implanted with a Trojan; once any abnormality is detected during any process, the detection is stopped and it is determined that the sensor is implanted with a Trojan.
[0057] Take Figure 2 as an example. In this embodiment, the above test process is carried out in the order of acoustic wave test, laser test, and electromagnetic test.
[0058] (1) Acoustic wave test
[0059] Test system settings: Start the acoustic wave sweep test system on the control host. In the system interface, set the initial parameters for the acoustic wave sweep test, including: start frequency, cut-off frequency, frequency step, time step, amplitude, and signal duration.
[0060] Test execution: After clicking the "Start" button, the acoustic wave test system will perform the test according to the preset parameters. During the test process, the system will display the response of the sensor to the acoustic waves with different parameters in real time and collect the test data.
[0061] Outlier Judgment: Analyze the test results and compare the response differences between the sensor under test and the normal sensor. Based on the 3σ outlier detection model, determine whether the response of the sensor under test is within the normal range. If the response of the sensor under test is abnormal at any test point, this may indicate the existence of a hardware Trojan. If the test results show normal, continue with the laser test to ensure the security and reliability of the sensor.
[0062] (2) Laser Test
[0063] Test System Setup: Start the laser test system on the control host. Here we use a laser with a fixed wavelength, and then use the test equipment to generate sine waves of different frequencies, and modulate the current of the laser generator through a modulation circuit to change the amplitude. In the system interface, set the initial parameters for the acoustic frequency sweep test, including: start frequency, cut-off frequency, frequency step, time step, amplitude, signal duration.
[0064] Test Execution: First, align the laser focus emitted by the laser generator with the sensor. After clicking the "Start" button, the laser test system will perform the test according to the preset parameters. During the test, the system will display the response of the sensor to the laser with different parameters in real time and collect the test data.
[0065] Outlier Judgment: Analyze the test results and compare the response differences between the sensor under test and the normal sensor. Based on the 3σ outlier detection model, determine whether the response of the sensor under test is within the normal range. If the response of the sensor under test is abnormal at any test point, this may indicate the existence of a hardware Trojan. If the test results show normal, continue with the electromagnetic test to ensure the security and reliability of the sensor.
[0066] (3) Electromagnetic Test
[0067] Test System Setup: Connect and configure the experimental device to ensure that the antenna is aligned with the sensor under test. Use gnuradio-companion to control the USRP (Universal Software Radio Peripheral) to emit electromagnetic signals of specific frequencies. In the system interface, set the initial key parameters for the electromagnetic test, including: start frequency, cut-off frequency, frequency step, time step, amplitude, signal duration.
[0068] Test Execution: After clicking the "Start" button, the electromagnetic test system will perform the test according to the preset parameters. The system displays the response of the sensor to the electromagnetic signals of different frequencies, amplitudes, and durations in real time and collects the test data.
[0069] Outlier Judgment: Analyze the test results and compare the response differences between the sensor under test and the normal sensor. Based on the 3σ outlier detection model, judge whether the response of the sensor under test is within the normal range. If the response of the sensor under test is abnormal at any test point, this may indicate the existence of a hardware Trojan. If the test results show normal, it can be judged that there is no obvious hardware Trojan in the sensor, thus ensuring the security and reliability of the sensor.
[0070] The present invention provides an efficient and accurate method for detecting hardware Trojans in sensors. Based on reinforcement learning, it optimizes the emitted acoustic waves, laser, and electromagnetic signals, realizing a comprehensive detection of sensors under different physical signals and accurately identifying potential hardware Trojans. This method is based on the 3σ outlier detection algorithm, effectively distinguishing the performance differences between normal sensors and sensors with Trojans, reducing the false alarm rate. At the same time, the automated test process reduces manual intervention and improves the detection efficiency. The present invention is applicable to various types of sensor products, ensuring their security and functionality in different application scenarios, providing a strong guarantee against malicious tampering, and significantly improving the security detection level of sensors.
[0071] Based on the same inventive concept, in this embodiment, a sensor hardware Trojan detection system based on signal injection is also provided, which is used to implement the above embodiment. The following terms such as "module" and "unit" can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.
[0072] In this embodiment, a sensor hardware Trojan detection system based on signal injection includes:
[0073] A data acquisition module, which is used to acquire the output characteristic data of the sensor in the target control environment, and the output characteristic data includes response time, node voltage, power consumption, and output data;
[0074] A reinforcement learning module, which is used to train a policy model by using the reinforcement learning method. The sensor and the physical signal emission device form an environment, the output characteristics of the sensor, the current emission parameters of the physical signal emission device, and the corresponding emission physical signal type form a state space, and the adjustable parameters of the physical signal emission device form an action space; during the training process, the policy model adjusts the emission parameters of the physical signal emission device according to the current state, the environment generates a new physical signal under the current emission parameters and acts on the sensor, and after the sensor responds, the environment returns a new state and a reward;
[0075] The test module is used to test the sensors to be tested and the clean sensors of the same batch under the target control environment of the selected physical signal transmitter, and to give the initial parameters of the frequency sweep test, including: the start frequency, the cutoff frequency, and the adjustable parameters; and to optimize the adjustable parameters of the signal transmitter in real time according to the current state and environment using the trained strategy network;
[0076] The anomaly detection module is used to determine whether the sensor's response is normal based on the output characteristic data of the sensor to be tested and the clean sensor; if normal, the sensor to be tested is not implanted with a Trojan horse; otherwise, the sensor to be tested is implanted with a Trojan horse.
[0077] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0078] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.
[0079] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A method for detecting sensor hardware Trojans based on signal injection, characterized in that, Including the following steps: (1) Place a clean sensor without Trojan in the target control environment, and collect the output characteristic data of the sensor. The output characteristic data includes response time, node voltage, power consumption, and output data; (2) Use the reinforcement learning method to train a policy model. The environment is composed of the sensor and the physical signal transmitting device. The state space is composed of the output characteristics of the sensor, the current transmitting parameters of the physical signal transmitting device, and the corresponding types of transmitted physical signals. The action space is composed of the adjustable parameters of the physical signal transmitting device; During the training process, the policy model adjusts the transmitting parameters of the physical signal transmitting device according to the current state. The environment generates a new physical signal under the current transmitting parameters and acts on the sensor. After the sensor responds, the environment returns a new state and a reward; (3) In the test phase, select the physical signal transmitting device, place the sensors to be tested and the clean sensors of the same batch in the target control environment, and give the initial parameters for the swept-frequency test, including: start frequency, cut-off frequency, and adjustable parameters; Use the trained policy network to optimize the adjustable parameters of the signal transmitting device in real time according to the current state and the environment; Based on the output characteristic data of the sensors to be tested and the clean sensors, judge whether the response of the sensors to be tested is normal based on the 3σ outlier detection method; If it is normal, the sensor to be tested is not implanted with a Trojan; otherwise, the sensor to be tested is implanted with a Trojan.
2. The method for detecting sensor hardware Trojans based on signal injection according to claim 1, wherein The physical signal transmitting device is two or more of an acoustic signal transmitting device, a laser signal transmitting device, and an electromagnetic signal transmitting device; During the test, different physical signal transmitting devices are traversed. Only when all are normal, it is judged that the sensor to be tested is not implanted with a Trojan.
3. A method for detecting sensor hardware Trojans based on signal injection according to claim 1, characterized in that, The adjustable parameters of the physical signal transmitting device include frequency step, time step, amplitude, and signal duration.
4. A sensor hardware Trojan detection method based on signal injection according to claim 1, characterized in that During the reinforcement learning process, when an abnormal output characteristic of the sensor is detected, a positive reward is given. When no abnormal output characteristic of the sensor is detected, or the sensor is damaged due to unreasonable setting of the transmitted signal parameters, a negative reward is given.
5. A method for detecting sensor hardware Trojans based on signal injection according to claim 1, characterized in that, The policy model uses the DDPG model.
6. The method for detecting sensor hardware Trojans based on signal injection according to claim 1, wherein, The start frequency and the cut-off frequency are preset fixed values.
7. A sensor hardware Trojan detection system based on signal injection is used to implement the above-mentioned sensor hardware Trojan detection method; characterized in that, The system includes: A data acquisition module, which is used to collect the output characteristic data of the sensor in the target control environment. The output characteristic data includes response time, node voltage, power consumption, and output data; A reinforcement learning module, which is used to train a policy model by using the reinforcement learning method. The environment is composed of the sensor and the physical signal transmitting device. The state space is composed of the output characteristics of the sensor, the current transmitting parameters of the physical signal transmitting device, and the corresponding types of transmitted physical signals. The action space is composed of the adjustable parameters of the physical signal transmitting device; During the training process, the policy model adjusts the transmitting parameters of the physical signal transmitting device according to the current state. The environment generates a new physical signal under the current transmitting parameters and acts on the sensor. After the sensor responds, the environment returns a new state and a reward; A test module, which is used to test the sensors to be tested and clean sensors of the same batch in the target control environment of a selected physical signal transmitting device, and given the initial parameters of the swept-frequency test, including: start frequency, cut-off frequency, and adjustable parameters; and use the trained policy network to optimize the adjustable parameters of the signal transmitting device in real time according to the current state and environment. An anomaly detection module, which is used to judge whether the response of the sensor to be tested is normal based on the output characteristic data of the sensor to be tested and the clean sensor; and based on the 3σ outlier detection method; if it is normal, the sensor to be tested is not implanted with a Trojan horse; otherwise, the sensor to be tested is implanted with a Trojan horse.
8. The sensor hardware Trojan detection system based on signal injection according to claim 7, characterized in that, The physical signal transmitting device is two or more of an acoustic wave signal transmitting device, a laser signal transmitting device, and an electromagnetic signal transmitting device.
9. The sensor hardware Trojan detection system based on signal injection according to claim 7, wherein The adjustable parameters of the physical signal transmitting device include frequency step size, time step size, amplitude, and signal duration.
10. The sensor hardware Trojan detection system based on signal injection according to claim 7, wherein During the reinforcement learning process, when an abnormal output characteristic of the sensor is detected, a positive reward is given, and when an abnormal output characteristic of the sensor is not detected, or the sensor is damaged due to unreasonable setting of the transmission signal parameters, a negative reward is given.
Citation Information
Patent Citations
Automatic driving system backdoor attack method based on deep reinforcement learning and related device
CN116389041A
Hardware trojan detection using reinforcement learning
US20220188415A1