A sensor hardware Trojan detection method and system based on signal injection

By combining signal injection and reinforcement learning, sensor parameters are optimized using acoustic waves, lasers, and electromagnetic signals. Based on a 3σ outlier detection model, the accuracy and efficiency issues of sensor hardware Trojan detection are solved, ensuring the safety and reliability of the sensor.

CN120337216BActive Publication Date: 2026-01-30ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510333107.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2026-01-30
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and detect hardware Trojans embedded in sensors, especially in analog signal processing environments, where traditional methods suffer from high false alarm rates and low detection efficiency.

Method used

A signal injection-based approach is adopted, which uses remote physical signals such as sound waves, lasers and electromagnetic signals to trigger sensors. Reinforcement learning is combined to optimize signal parameters. By detecting the node voltage, power consumption and output characteristics of the sensor, a 3σ outlier detection model is used to determine whether the sensor has been implanted with a Trojan.

Benefits of technology

It achieves efficient and accurate detection of sensor hardware Trojans, reduces false alarm rate, and improves the reliability and applicability of detection. It is applicable to a variety of sensor products, ensuring their security and functionality in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337216B_ABST
    Figure CN120337216B_ABST
Patent Text Reader

Abstract

This invention discloses a sensor hardware Trojan detection method and system based on signal injection, belonging to the field of sensor anomaly detection. A clean sensor is placed in the target environment, and characteristic data such as response time are collected. Reinforcement learning is used to construct the environment using the sensor and the transmitting device, and a policy model is trained. The state space includes sensor output, transmission parameters, and signal type, while the action space contains adjustable parameters of the transmitting device. During training, the model adjusts parameters according to the state. The environment generates new signals that act on the sensor and provide feedback on the state and reward. In the testing phase, a transmitting device is selected, and a batch of sensors to be tested and clean sensors are placed in it. Initial parameters for frequency sweep are set, and the trained policy network optimizes the transmission parameters in real time. Based on the outputs of the two sensors, the 3σ outlier detection method is used to determine whether the response of the sensor under test is normal, thereby determining whether the sensor has been implanted with a Trojan. This invention utilizes reinforcement learning to optimize transmission signal parameters, enabling more efficient Trojan detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sensor anomaly detection, and more specifically to a sensor hardware Trojan detection method and system based on signal injection. Background Technology

[0002] Modern sensor technology and its role in cyber-physical systems (CPS) are becoming increasingly important. With technological advancements, sensors have not only become more sensitive and accurate but also integrated more sophisticated processing capabilities to support intelligent decision-making. However, this complexity also provides malicious actors with new attack vectors: "sensor Trojans." A sensor Trojan is a malicious entity hidden inside a sensor that can affect its behavior or output, potentially embedding itself in any component of the sensor. Once activated, a Trojan can cause the sensor to malfunction, provide incorrect readings, or cease operation entirely, posing a serious threat to security-critical applications that rely on accurate sensor data.

[0003] Because multiple stakeholders involved in the design and manufacturing of modern sensors may have access to design details and production processes, there is an opportunity to embed Trojan horses at different stages of production. These Trojans may be hardware tampering or hidden code in embedded software, and may be triggered under specific conditions, such as upon receiving input signals of a preset pattern. Traditional digital circuit detection methods are not always effective in identifying these analog-domain Trojans, as the latter often exploit the characteristics of analog signal processing.

[0004] As smart devices (such as cars and drones) increasingly rely on high-precision sensor data, ensuring sensor security and reliability has become paramount. Methods for detecting sensor-based malware using physical signals must consider variables in the actual operating environment and be able to distinguish between fluctuations under normal operating conditions and potential malicious activity. Such detection mechanisms require a deep understanding of sensor performance to ensure that even the most stealthy malware is detected. Furthermore, this approach should minimize false alarms to maintain the overall reliability and efficiency of the system.

[0005] In summary, it is urgent to address the significant threat posed by sensor backdoors and to improve the detection mechanisms for them. Summary of the Invention

[0006] To address the aforementioned issues, this invention proposes a sensor hardware Trojan detection method and system based on signal injection. It focuses on analyzing the actual physical characteristics of sensor responses, such as node voltage levels, power consumption, and output, to detect abnormal behavior (or signs of hidden Trojans). The method uses remote physical signal injection, such as sound waves, ultrasound, electromagnetic signals, and lasers, to covertly trigger sensor Trojans. Reinforcement learning is then used to automatically adjust physical signal parameters and optimize the signal transmission strategy, resulting in more efficient Trojan detection.

[0007] The technical solution proposed in this invention is as follows:

[0008] In a first aspect, the present invention proposes a sensor hardware Trojan detection method based on signal injection, comprising the following steps:

[0009] (1) Place a clean sensor without Trojan horse in the target control environment and collect the output characteristic data of the sensor. The output characteristic data includes response time, node voltage, power consumption and output data.

[0010] (2) A policy model is trained using reinforcement learning. The environment consists of sensors and physical signal transmitting devices. The state space consists of the output characteristics of the sensors, the current transmission parameters of the physical signal transmitting devices, and the corresponding types of transmitted physical signals. The action space consists of the adjustable parameters of the physical signal transmitting devices. During the training process, the policy model adjusts the transmission parameters of the physical signal transmitting devices according to the current state. The environment generates new physical signals under the current transmission parameters and acts on the sensors. After the sensors respond, the environment returns a new state and reward.

[0011] (3) During the testing phase, a physical signal transmitting device is selected, and the sensors to be tested and clean sensors of the same batch are placed in the target control environment. The initial parameters of the frequency sweep test are given, including the start frequency, the cutoff frequency, and the adjustable parameters. The adjustable parameters of the signal transmitting device are optimized in real time according to the current state and environment using the trained policy network. Based on the output characteristic data of the sensors to be tested and clean sensors, the 3σ outlier detection method is used to determine whether the sensor response is normal.

[0012] If normal, the sensor under test has not been infected with a Trojan horse; otherwise, the sensor under test has been infected with a Trojan horse.

[0013] Furthermore, the physical signal transmitting device is one or more of the following: sound wave signal transmitting device, laser signal transmitting device, and electromagnetic signal transmitting device; during testing, different physical signal transmitting devices are traversed, and only when all are normal is it determined that the sensor under test has not been implanted with a Trojan horse.

[0014] Furthermore, the adjustable parameters of the physical signal transmitting device include frequency step size, time step size, amplitude, and signal duration.

[0015] Furthermore, during the reinforcement learning process, a positive reward is given when an abnormality in the sensor's output characteristics is detected, and a negative reward is given when no abnormality in the sensor's output characteristics is detected, or when unreasonable settings of the transmission signal parameters lead to sensor damage.

[0016] Furthermore, the strategy model described adopts the DDPG model.

[0017] Furthermore, the start frequency and the cutoff frequency are preset fixed values.

[0018] Secondly, this invention proposes a sensor hardware Trojan detection system based on signal injection, which is used to implement the aforementioned sensor hardware Trojan detection method.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0020] This invention optimizes the parameters of various physical signals such as sound waves, lasers, and electromagnetic signals using reinforcement learning to perform security testing on sensors. This multi-dimensional detection method can effectively reveal the sensor's response under different signal stimuli, thereby accurately identifying potentially hidden hardware Trojans.

[0021] The 3σ outlier detection model can be used to determine whether sensor features are abnormal. It can automatically judge the sensor response according to preset standards, thereby reducing the false alarm rate and improving the accuracy of detection, and ensuring the reliability of the detection results.

[0022] The versatility and high efficiency of the detection process of this invention make it applicable to a variety of sensor products, and the automation of testing reduces human intervention, which not only improves detection efficiency, but also ensures the safety and functionality of the sensor in different application scenarios. Attached Figure Description

[0023] Figure 1 This is a control block diagram of a sensor hardware Trojan detection method based on signal injection proposed in this invention;

[0024] Figure 2 This is a flowchart of a sensor hardware Trojan detection method based on signal injection proposed in this invention. Detailed Implementation

[0025] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.

[0026] The accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0027] The flowchart shown in the attached diagram is merely an illustrative example and does not necessarily include all steps. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0028] This invention proposes a sensor hardware Trojan detection method based on signal injection, focusing on using a neural network constructed through reinforcement learning to optimize physical signal characteristics to detect these potential threats. For example... Figure 1 As shown, the host computer acts as the main control unit, receiving commands via serial port or Ethernet communication protocol and simultaneously generating control signals to drive the signal transmitting device. The signal transmitting device includes acoustic, laser, and electromagnetic emission modules, which transmit physical signals to the sensor array respectively. The sensor section consists of the sensor under test and a reference sensor (clean sensor), both of which are subjected to physical signal interference in a controlled environment. The host computer collects and analyzes the sensor's output characteristic data. This method can meticulously analyze the electrical signal characteristics generated by the sensors during operation, including but not limited to voltage levels, current patterns, frequency response, and noise distribution, thereby capturing any abnormal behavior or unexpected signal changes.

[0029] like Figure 2 As shown, the present invention proposes a sensor hardware Trojan detection method based on signal injection, which mainly includes the following steps:

[0030] S1, Data Preparation

[0031] Place the sensor to be tested in the control environment and connect it to the sensor data acquisition platform to ensure a stable connection between the sensor and the data acquisition platform in order to collect data accurately.

[0032] Start the data acquisition platform and set initial parameters to match the sensor specifications and expected operating environment. Collect response time, node voltage, power consumption, and output data from a clean sensor (free of malware) and keep these consistent for subsequent tests. Repeat this process to collect 100 sets of data, which will serve as the benchmark for subsequent tests for comparison and analysis.

[0033] S2: Construct a reinforcement learning model.

[0034] 1. Environmental and Strategy Design

[0035] In the process of using reinforcement learning to detect sensor trojans, it is first necessary to define the environment, state space, and action space.

[0036] In one specific embodiment of the present invention, the environment consists of sensors and a physical signal transmitting device, which is any one of an acoustic signal transmitting device, a laser signal transmitting device, or an electromagnetic signal transmitting device. The physical signal transmitting device operates at a set start frequency, cutoff frequency, frequency step size, time step size, amplitude, and signal duration, from the start frequency to the cutoff frequency. Here, the start frequency and cutoff frequency are preset parameters, and the frequency step size, time step size, amplitude, and signal duration are adjusted using a reinforcement learning algorithm based on the interaction process. Here, specific amplitude signals and signal durations are more likely to serve as Trojan trigger conditions.

[0037] The state space includes: the sensor's output characteristics (including response time, node voltage, power consumption, and output data) and the current transmission parameters and corresponding physical signal types of the physical signal transmitting device. Among the sensor's output characteristics, response time reflects how quickly the sensor reacts to external stimuli; node voltage directly relates to the sensor's internal electrical state; power consumption reflects the energy consumption during sensor operation; and output data is a direct representation of the sensed physical quantity. The transmitted physical signal types are related to the physical signal transmitting device, namely, acoustic signal types, laser signal types, and electromagnetic signal types. In this embodiment, one type of transmitted physical signal, such as electromagnetic or laser, is fixed for each test.

[0038] The action space refers to the adjustable parameters of the physical signal transmitting device, such as frequency step size, time step size, amplitude, and signal duration. Among these, changing the frequency step size adjusts the amplitude of the frequency variation of the transmitted signal, while the time step size controls the degree of fine adjustment of the signal in the time dimension. The amplitude determines the strength of the signal, directly affecting the interaction between the signal and the sensor. The signal duration specifies the duration of the transmitted signal, and different duration settings may trigger different responses from the sensor.

[0039] The above definition lays the foundation for subsequent strategy learning.

[0040] 2. Reinforcement Learning Algorithms and Reward Mechanisms

[0041] Interaction process with the environment: The policy network selects an action (adjusts the emission parameters) based on the current state, the environment generates a new physical signal based on the action and applies it to the sensor, and after the sensor responds, the environment returns a new state and reward.

[0042] Positive design reward: Detecting abnormal output characteristics of the sensor, such as abnormal response time, node voltage fluctuations, abnormal power consumption, or output data deviating from expectations.

[0043] Negative design incentives: failure to detect abnormal sensor output characteristics, or unreasonable transmission signal parameter settings (such as excessively high signal strength leading to sensor damage).

[0044] Selecting a policy network suitable for the continuous action space is one of the key steps. In this embodiment, the policy network adopts the Deep Deterministic Policy Gradient (DDPG) algorithm, which combines the advantages of Deep Q-Network (DQN) and policy gradient methods.

[0045] By combining the above reward mechanism, the algorithm can learn how to adjust the parameters of the physical signal transmitting device to maximize the probability of detecting sensor Trojans.

[0046] 3. During the training phase, the signal transmitting device adjusts the parameters of the physical signal transmitting device according to the current strategy, observes the output characteristics of the sensor and calculates the reward, and then updates the policy network based on the reward.

[0047] S3: Testing Phase

[0048] During the testing phase, tests were randomly conducted using acoustic signal transmitting devices, laser signal transmitting devices, and electromagnetic signal transmitting devices. Under each test:

[0049] Place the sensors under test and clean sensors from the same batch in the control environment, connect them to the sensor data acquisition platform, and start the corresponding signal transmission device. Provide the initial parameters for the frequency sweep test, including: start frequency, cutoff frequency, frequency step size, time step size, amplitude, and signal duration; where the start frequency and cutoff frequency are fixed values, and the frequency step size, time step size, amplitude, and signal duration are adjustable parameters initialized from the beginning.

[0050] The trained policy network optimizes the adjustable parameters of the signal transmitting device in real time based on the current state and environment. The output characteristics of the sensor are observed, and the differences in output characteristics between the sensor under test and the clean sensor are compared. Based on the 3σ outlier detection method, it is determined whether the response of the sensor under test is within the normal range.

[0051] It should be noted that the policy network here optimizes the frequency step size, time step size, amplitude, and signal duration in real time, without changing the start frequency or cutoff frequency. Therefore, the test process will follow the step-by-step approach from the start frequency to the cutoff frequency, and finally stop the test automatically.

[0052] The 3σ outlier detection method described above is based on the normal distribution assumption in statistics. It calculates the mean and standard deviation of the data to determine whether data points deviate from the normal range. Here, abnormal behavior is identified by comparing the data from the sensor under test and a clean sensor, preventing misjudgments due to the inherent fragility of the sensor itself. The data from the clean sensor and the sensor under test are collected under the same test parameters. Anomalies are assessed separately for response time, node voltage, power consumption, and output data. Taking a specific output characteristic (such as output data) as an example, suppose the collected output data is as follows:

[0053] Clean sensor data: X_normal = {x_1, x_2, x_n};

[0054] Test sensor data: Y_test={y_1,y_2,y_n}.

[0055] Calculate the mean and standard deviation of the clean sensor data. Based on the 3σ principle, the range of normal data can be obtained. Determine whether each data point of the sensor under test falls within the normal range. If so, the sensor is considered normal in terms of output data. Otherwise, it is considered abnormal, and the sensor is believed to have been infected with malware. The judgment of other output characteristics is similar. Only when all output characteristics are normal can the test proceed to the next signal transmitting device.

[0056] If all three signal transmitting devices function normally, the sensor is determined to be free of malware; if any abnormality is detected in any process, the detection is stopped, and the sensor is determined to be infected with malware.

[0057] by Figure 2 For example, in this embodiment, the above testing process is performed in the order of acoustic wave testing, laser testing, and electromagnetic testing.

[0058] (1) Sound wave test

[0059] Test system setup: Start the acoustic wave frequency sweep test system on the control host. In the system interface, specify the initial parameters for the acoustic wave frequency sweep test, including: start frequency, cutoff frequency, frequency step, time step, amplitude, and signal duration.

[0060] Test Execution: After clicking the "Start" button, the acoustic wave testing system will perform the test according to the preset parameters. During the test, the system will display the sensor's response to the acoustic waves at different parameters in real time and collect test data.

[0061] Outlier Detection: Analyze the test results and compare the response differences between the sensor under test and normal sensors. Use a 3σ outlier detection model to determine if the sensor's response is within the normal range. If the sensor's response is abnormal at any test point, this may indicate the presence of a hardware malware. If the test results are normal, continue with laser testing to ensure the sensor's safety and reliability.

[0062] (2) Laser testing

[0063] Test System Setup: Start the laser test system on the control host. Here we use a fixed wavelength laser and then use the test equipment to generate sine waves of different frequencies. The amplitude is changed by modulating the current of the laser generator through a modulation circuit. In the system interface, set the initial parameters for the acoustic wave frequency sweep test, including: start frequency, cutoff frequency, frequency step, time step, amplitude, and signal duration.

[0064] Test Execution: First, align the laser focus emitted by the laser generator with the sensor. After clicking the "Start" button, the laser testing system will perform the test according to the preset parameters. During the test, the system will display the sensor's response to different laser parameters in real time and collect test data.

[0065] Outlier Detection: Analyze the test results and compare the response differences between the sensor under test (DUT) and normal sensors. Use a 3σ outlier detection model to determine if the DUT's response is within the normal range. If the DUT's response is abnormal at any test point, this may indicate the presence of a hardware malware. If the test results are normal, continue with electromagnetic testing to ensure the sensor's safety and reliability.

[0066] (3) Electromagnetic testing

[0067] Test system setup: Connect and configure the experimental setup, ensuring the antenna is aligned with the sensor under test. Use gnuradio-companion to control the USRP (Universal Software Radio Peripheral) to emit electromagnetic signals at a specific frequency. Within the system interface, set the initial key parameters for the electromagnetic test, including: start frequency, cutoff frequency, frequency step, time step, amplitude, and signal duration.

[0068] Test Execution: After clicking the "Start" button, the electromagnetic testing system will perform the test according to the preset parameters. The system displays the sensor's response to electromagnetic signals of different frequencies, amplitudes, and durations in real time and collects test data.

[0069] Outlier Detection: Analyze the test results and compare the response differences between the sensor under test and normal sensors. Use a 3σ outlier detection model to determine if the sensor's response is within the normal range. If the sensor's response is abnormal at any test point, this may indicate the presence of a hardware malware. If the test results are normal, it can be determined that the sensor does not have obvious hardware malware, thus ensuring the sensor's security and reliability.

[0070] This invention provides a highly efficient and accurate method for detecting hardware Trojans in sensors. Based on reinforcement learning, it optimizes the emitted acoustic waves, lasers, and electromagnetic signals, enabling comprehensive detection of sensors under different physical signals and accurately identifying potential hardware Trojans. This method utilizes a 3σ outlier detection algorithm to effectively distinguish performance differences between normal sensors and those containing Trojans, reducing false alarm rates. Simultaneously, the automated testing process reduces manual intervention and improves detection efficiency. This invention is applicable to various types of sensor products, ensuring their security and functionality in different application scenarios, providing strong protection against malicious tampering, and significantly improving the security detection level of sensors.

[0071] Based on the same inventive concept, this embodiment also provides a sensor hardware Trojan detection system based on signal injection, which is used to implement the above embodiments. The terms "module," "unit," etc., used below can refer to a combination of software and / or hardware that performs a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible.

[0072] In this embodiment, a sensor hardware Trojan detection system based on signal injection includes:

[0073] The data acquisition module is used to acquire the output characteristic data of the sensors under the target control environment. The output characteristic data includes response time, node voltage, power consumption and output data.

[0074] The reinforcement learning module is used to train a policy model using reinforcement learning methods. The environment consists of sensors and physical signal transmitters. The state space consists of the output characteristics of the sensors, the current transmission parameters of the physical signal transmitters, and the corresponding types of transmitted physical signals. The action space consists of the adjustable parameters of the physical signal transmitters. During training, the policy model adjusts the transmission parameters of the physical signal transmitters according to the current state. The environment generates new physical signals under the current transmission parameters and acts on the sensors. After the sensors respond, the environment returns a new state and reward.

[0075] The testing module is used to test the same batch of sensors under test and clean sensors in the target control environment of a selected physical signal transmitting device. It provides initial parameters for the frequency sweep test, including the start frequency, cutoff frequency, and adjustable parameters. The trained policy network optimizes the adjustable parameters of the signal transmitting device in real time based on the current state and environment.

[0076] The anomaly detection module is used to determine whether the sensor response is normal based on the output characteristic data of the sensor under test and the clean sensor, using the 3σ outlier detection method. If it is normal, the sensor under test has not been infected with a Trojan; otherwise, the sensor under test has been infected with a Trojan.

[0077] For the system embodiments, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments; the implementation methods of the remaining modules will not be repeated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0078] The system embodiments of the present invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The system embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution.

[0079] The above examples are merely specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments and many variations are possible. All variations that can be directly derived or conceived by those skilled in the art from the disclosure of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A signal injection based sensor hardware Trojan detection method, characterized in that, The method comprises the following steps: (1) placing a clean sensor without a Trojan horse in a target control environment, collecting output characteristic data of the sensor, wherein the output characteristic data comprises response time, node voltage, power consumption and output data; (2) training a strategy model by using a reinforcement learning method, wherein the environment is formed by the sensor and a physical signal emitting device, the state space is formed by the output characteristic of the sensor, the current emission parameter of the physical signal emitting device and the corresponding emitted physical signal type, and the action space is formed by the adjustable parameter of the physical signal emitting device; During the training process, the strategy model adjusts the emission parameter of the physical signal emitting device according to the current state, the environment generates a new physical signal under the current emission parameter and acts on the sensor, and the environment returns the new state and a reward after the sensor responds; (3) in the test stage, selecting a physical signal emitting device, placing the same batch of sensors to be tested and clean sensors in the target control environment, giving initial parameters of the sweep test, including a start frequency, a stop frequency and adjustable parameters, and using the trained strategy network to optimize the adjustable parameters of the signal emitting device in real time according to the current state and the environment; and judging whether the response of the sensor to be tested is normal or not based on the 3σ outlier detection method according to the output characteristic data of the sensor to be tested and the clean sensor; If the response is normal, the sensor to be tested is not implanted with a Trojan horse; otherwise, the sensor to be tested is implanted with a Trojan horse.

2. The method of claim 1, wherein, The physical signal emitting device is two or more of an acoustic signal emitting device, a laser signal emitting device and an electromagnetic signal emitting device; during the test, different physical signal emitting devices are traversed, and only when all the physical signal emitting devices are normal, it is determined that the sensor to be tested is not implanted with a Trojan horse.

3. The method of claim 1, wherein, The adjustable parameters of the physical signal emitting device include a frequency step, a time step, an amplitude and a signal duration.

4. The signal injection based sensor hardware Trojan detection method of claim 1, wherein, During the reinforcement learning process, a positive reward is given when the output characteristic of the sensor is detected to be abnormal, and a negative reward is given when the output characteristic of the sensor is not detected to be abnormal or the sensor is damaged due to unreasonable emission signal parameter setting.

5. The signal injection based sensor hardware Trojan detection method of claim 1, wherein, The strategy model adopts a DDPG model.

6. The signal injection based sensor hardware Trojan detection method of claim 1, wherein The start frequency and the stop frequency are preset fixed values.

7. A signal injection based sensor hardware Trojan detection system for implementing the sensor hardware Trojan detection method described above; characterized in that, The system comprises: a data collection module for collecting output characteristic data of the sensor in the target control environment, wherein the output characteristic data comprises response time, node voltage, power consumption and output data; a reinforcement learning module for training a strategy model by using a reinforcement learning method, wherein the environment is formed by the sensor and a physical signal emitting device, the state space is formed by the output characteristic of the sensor, the current emission parameter of the physical signal emitting device and the corresponding emitted physical signal type, and the action space is formed by the adjustable parameter of the physical signal emitting device; during the training process, the strategy model adjusts the emission parameter of the physical signal emitting device according to the current state, the environment generates a new physical signal under the current emission parameter and acts on the sensor, and the environment returns the new state and a reward after the sensor responds; The test module is used for testing the to-be-tested sensor and the clean sensor in the same batch under a target control environment of a selected physical signal emission device, and gives initial parameters of a sweep test, including a start frequency, a cut-off frequency, and adjustable parameters; the trained strategy network is used for optimizing the adjustable parameters of the signal emission device in real time according to a current state and an environment; The abnormality detection module is used for judging whether the response of the to-be-tested sensor is normal based on 3σ abnormal value detection method according to output characteristic data of the to-be-tested sensor and the clean sensor; if yes, the to-be-tested sensor is not implanted with a Trojan horse; otherwise, the to-be-tested sensor is implanted with a Trojan horse.

8. The signal injection based sensor hardware Trojan detection system of claim 7, wherein, The physical signal emission device is two or more of an acoustic wave signal emission device, a laser signal emission device and an electromagnetic signal emission device.

9. The signal injection based sensor hardware Trojan detection system of claim 7, wherein, The adjustable parameters of the physical signal emission device include a frequency step, a time step, an amplitude and a signal duration.

10. The signal injection based sensor hardware Trojan detection system of claim 7, wherein, In the reinforcement learning process, a positive reward is given when the output characteristic of the sensor is detected to be abnormal, and a negative reward is given when the output characteristic of the sensor is not detected to be abnormal or the sensor is damaged due to unreasonable emission signal parameters.