An intelligent security control system and method based on a simulation environment

By constructing a simulation environment and using multi-agent reinforcement learning, security strategies are optimized, solving the problems of poor environmental adaptability and isolated sensor information in existing intelligent security systems, and achieving high-precision intrusion detection and effective defense.

CN120526516BActive Publication Date: 2025-12-09GUANGDONG CHENGCUN CONSTRUCTION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510347500.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-12-09
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

Existing intelligent security systems suffer from poor environmental adaptability, lack dynamic learning capabilities, isolated sensor information, and a lack of intelligent simulation testing environments, resulting in insufficient detection accuracy and inadequate effectiveness of defense measures.

Method used

An intelligent security control system based on a simulation environment is constructed. By combining a random generation module, a virtual sensing module, a collaborative simulation module, and an actual monitoring module with multi-agent reinforcement learning, security strategies are optimized to improve the accuracy of intrusion detection and the effectiveness of defense measures.

Benefits of technology

It has achieved intelligent and dynamic optimization of security strategies, improved the intelligence level, adaptability and defense capabilities of security systems, and significantly enhanced the accuracy of intrusion detection and the effectiveness of defense measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526516B_ABST
    Figure CN120526516B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent security control system and method based on a simulation environment, comprising a random generation module, a virtual sensing model, a collaborative simulation module, an actual monitoring module, a parameter comparison module and an optimization module. The random generation module constructs a simulation environment and generates variable parameters, the virtual sensing module outputs simulation sensing data through modeling, the collaborative simulation module constructs intrusion, defense and regulation intelligent agents based on a multi-agent reinforcement learning framework, realizes simulation and defense optimization of various intrusion behaviors and response strategies. The actual monitoring module obtains real environment data, the parameter comparison module analyzes the deviation of simulation and reality data, and the optimization module dynamically adjusts the intelligent agent model based on learning parameters to improve the adaptability and response capability of the security system. The application can dynamically optimize security strategies, improve the accuracy of intrusion detection and the effectiveness of defense measures.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent security technology, and in particular to an intelligent security control system and method based on a simulation environment. BACKGROUND

[0002] With the development of artificial intelligence, Internet of Things and intelligent monitoring technology, security systems are increasingly widely used in smart cities, enterprise parks, residential areas and other scenarios. Current intelligent security systems usually rely on fixed monitoring devices (such as cameras, sensors, etc.) for intrusion detection, and use pre-set rules or traditional machine learning algorithms for alarm and defense. However, these traditional security systems have the following technical limitations:

[0003] 1. Poor environmental adaptability: Existing security systems usually analyze based on fixed rules or static data, making it difficult to cope with complex and changing intrusion behaviors and environmental changes. For example, factors such as camera monitoring angle, light conditions, sensor sensitivity, etc. will affect the accuracy of detection.

[0004] 2. Lack of dynamic learning ability: Most security systems cannot self-optimize according to actual environmental feedback, making it difficult to cope with new or hidden intrusion behaviors. In the face of attackers' evasion strategies, traditional systems often react slowly, lacking continuous learning and dynamic adjustment capabilities.

[0005] 3. Sensory information is isolated: In current security systems, different types of sensors (such as infrared detectors, radars, cameras, etc.) usually operate independently, with low data fusion, making the detection results susceptible to single sensor false positives or false negatives, reducing the overall reliability of the system.

[0006] 4. Lack of intelligent simulation test environment: In actual applications, the testing and optimization of security systems usually rely on real environment intrusion testing, which is costly and risky. Existing simulation test environments often lack realism and cannot accurately reflect various complex intrusion behaviors and abnormal situations.

[0007] To address the above problems, the present application proposes an intelligent security control system based on a simulation environment, which constructs a simulation environment, a virtual sensor system, and combines multi-agent reinforcement learning to realize dynamic simulation and defense optimization of intrusion behaviors. SUMMARY

[0008] To address the deficiencies of the prior art, the present application aims to provide an intelligent security control system and method based on a simulation environment for dynamically optimizing security strategies and improving the accuracy of intrusion detection and the effectiveness of defense measures.

[0009] To achieve the above-mentioned purpose, the present application provides the following technical solution: An intelligent security control system based on a simulation environment, comprising:

[0010] a random generation module configured to construct a simulation environment of the intelligent security system and randomly generate a plurality of change parameters in the simulation environment, the change parameters including simulation environment parameters and sensing anomaly parameters;

[0011] a virtual sensing module configured to model a plurality of types of sensors to obtain a plurality of virtual sensing models, the virtual sensing models being configured to output simulation sensing data and fuse the simulation sensing data to obtain simulation fusion data;

[0012] a collaborative simulation module connected to the random generation module and the virtual sensing module, and configured to construct an intrusion agent, a defense agent and a regulation agent in the simulation environment according to a pre-introduced multi-agent reinforcement learning framework, the intrusion agent being configured to simulate a plurality of simulated intrusion behaviors and simulated countermeasures, the regulation agent being configured to generate a defense strategy according to the simulation fusion data when facing the simulated intrusion behaviors and the simulated countermeasures of the intrusion agent, and the defense agent being configured to adjust a camera angle, an alarm threshold and a light state according to the defense strategy, the simulation environment parameters and the sensing anomaly parameters;

[0013] an actual monitoring module configured to obtain actual environment parameters, actual anomaly parameters, actual intrusion behaviors, actual sensing fusion data and actual countermeasures;

[0014] a parameter comparison module connected to the collaborative simulation module and the actual monitoring module, and configured to compare the actual environment parameters, the actual anomaly parameters, the actual intrusion behaviors, the actual sensing fusion data and the actual countermeasures with the simulation environment parameters, the sensing anomaly parameters, the simulated intrusion behaviors, the simulation sensing data and the simulated countermeasures to obtain a plurality of learning parameters, the learning parameters including environment deviation data, anomaly deviation data, intrusion deviation data, sensing fusion deviation data and countermeasure deviation data;

[0015] an optimization module connected to the parameter comparison module and the collaborative simulation module, and configured to dynamically correct and optimize the intrusion agent, the defense agent and the regulation agent according to the learning parameters.

[0016] Further, the random generation module comprises:

[0017] an environment generation unit configured to randomly match and generate a plurality of simulation environment parameters from a preset environment database according to positioning data and time data at a location where the simulation environment-based intelligent security control system is located, the simulation environment parameters including simulation temperature, simulation humidity, simulation air quality, simulation illumination parameters, simulation electromagnetic interference data and simulation biological interference data;

[0018] anomaly introducing unit, configured to introduce the sensor anomaly parameters in a sensor anomaly state in the simulation environment, the sensor anomaly parameters including simulation sensor noise data, simulation signal interruption data, simulation sensor conflict data, and simulation link anomaly data.

[0019] Further, the virtual sensor module includes:

[0020] a sensor modeling unit, configured to construct a plurality of virtual sensor models for simulating response curves, delays, noises, and errors of various sensors and outputting simulation sensor data including simulation video data, simulation infrared data, simulation audio data, and simulation vibration data.

[0021] a preliminary fusion unit connected to the sensor modeling unit and configured to construct a multi-sensor data fusion algorithm in the simulation environment to sequentially perform time-space alignment, information extraction, and data fusion on different sources of the simulation sensor data to obtain simulation fusion data.

[0022] Further, the data fusion calculation formula of the multi-sensor data fusion algorithm is configured as:

[0023]

[0024]

[0025]

[0026] wherein, represents the simulation fusion data, represents the total time of fusion data calculation, represents a discrete series index, represents a set of n-dimensional real number vectors, represents a Gamma function for controlling the scale change of data, represents a Riemann Zeta function for enhancing the nonlinear mapping of data, represents a first kind of Bessel function for regulating the smoothness of data, represents a complex information filtering function for optimizing data fusion by combining an error function, a logarithmic function, and an elliptic integral, represents different types of sensor data, wherein is simulation video data, is simulation infrared data, is simulation audio data, is simulation vibration data, Weights representing various types of sensor data, used to allocate according to the confidence and contribution of the data, Indicates the index of the sensor data in the spatial domain, is a filter input variable, is an integral variable, is a first type of elliptic integral, used to correct the nonlinear error of the data.

[0027] Further, it also includes a framework modeling module connected to the collaborative simulation module, the framework modeling module is used to build a hierarchical reinforcement learning framework, and includes:

[0028] An outer layer global policy unit is configured to generate a security policy sub-goal according to hierarchical deep reinforcement learning, and pass the security policy sub-goal to the regulation intelligent agent, and the regulation intelligent agent generates the defense strategy based on the security policy sub-goal;

[0029] An inner layer local policy unit is configured to perform fine-grained adjustment for cameras, alarm devices and lights, and train the adjustment mode based on a policy gradient method to optimize the local control task of the defense intelligent agent on the camera angle, the alarm threshold and the light state.

[0030] Further, the hierarchical reinforcement learning framework adopts a hierarchical reward structure, the hierarchical reward structure includes an overall reward, an outer layer reward and an inner layer reward, and the conversion relationship between the overall reward, the outer layer reward and the inner layer reward is configured as:

[0031]

[0032] Wherein, represents the overall reward, represents the outer layer reward, represents the inner layer reward, represents the outer layer weight, represents the inner layer weight, represents the total number of inner layer rewards.

[0033] Further, the actual environment parameters include actual temperature, actual humidity, actual air quality, actual light parameters, actual electromagnetic interference data and actual biological interference data, and the framework modeling module further includes an adaptive weight adjustment unit, the adaptive weight adjustment unit is used to establish a conversion relationship between the actual environment parameters and the outer layer weight, the inner layer weight, and dynamically adjust the outer layer weight, the inner layer weight according to the actual environment parameters, wherein the conversion relationship between the actual environment parameters and the outer layer weight, the inner layer weight is configured as:

[0034]

[0035] wherein, represents the actual temperature, represents the actual humidity, represents the actual air quality, represents the actual light parameter, represents the actual electromagnetic interference data, represents the actual biological interference data, represents the average value of the actual temperature, represents the average value of the actual humidity, represents the average value of the actual air quality, represents the average value of the actual light parameter, represents the average value of the actual electromagnetic interference data, represents the average value of the actual biological interference data, respectively represent the value range of the actual temperature, the actual humidity, the actual air quality, the actual light parameter, the actual electromagnetic interference data and the actual biological interference data.

[0036] Further, the actual abnormal parameter includes actual video data, actual infrared data, actual audio data and actual vibration data, and the framework modeling module further includes a reward adjustment unit, the reward adjustment unit is used for calculating a sensing comprehensive abnormal value according to the actual video data, the actual infrared data, the actual audio data and the actual vibration data, and dynamically adjusting the overall reward according to the sensing comprehensive abnormal value, and the conversion relationship between the sensing comprehensive abnormal value and the overall reward is configured as:

[0037]

[0038]

[0039] wherein, represents the sensing comprehensive abnormal value, represents the actual video data abnormal value,

[0040] represents the actual infrared data abnormal value, represents the actual audio data abnormal value, represents the actual vibration data abnormal value, ,​​​​​​ , respectively represent the mean value of each sensor data, , , , respectively represent the value range of each sensor data, represents a preset basic overall reward, represents an anomaly suppression coefficient, used to control the degree of attenuation of abnormal values on the overall reward, represents a preset anomaly compensation coefficient, used to prevent the reward value from being excessively reduced, for providing a smooth anomaly adjustment effect and preventing mutation.

[0041] Further, in the hierarchical reinforcement learning framework, an edge computing device and a cloud server are included, the edge computing device is used to perform the local control task in real time, and the cloud server is used to continuously update the defense strategy.

[0042] An intelligent security control method based on a simulation environment, applied to the intelligent security control system based on the simulation environment, comprising:

[0043] Step S1, a random generation module constructs a simulation environment of an intelligent security system, and randomly generates a plurality of change parameters in the simulation environment, the change parameters including simulation environment parameters and sensor anomaly parameters;

[0044] Step S2, a virtual sensor module models a plurality of sensors to obtain a plurality of virtual sensor models, the virtual sensor models are used to output simulation sensor data, and simulation fusion data is obtained by fusing each simulation sensor data;

[0045] Step S3, a cooperative simulation module constructs an intrusion agent, a defense agent and a regulation and control agent in the simulation environment according to a pre-introduced multi-agent reinforcement learning framework, the intrusion agent is used to simulate a plurality of simulated intrusion behaviors and simulated countermeasures, the regulation and control agent is used to generate a defense strategy according to the simulation fusion data when facing each simulated intrusion behavior and simulated countermeasure of the intrusion agent, and the defense agent is used to adjust the camera angle, alarm threshold and light state according to the defense strategy, simulation environment parameters and sensor anomaly parameters;

[0046] Step S4, an actual monitoring module acquires actual environment parameters, actual anomaly parameters, actual intrusion behaviors and actual countermeasures;

[0047] Step S5, the parameter comparison module compares the actual environment parameter, the actual anomaly parameter, the actual intrusion behavior and the actual response strategy with the simulation environment parameter, the sensing anomaly parameter, the simulated intrusion behavior and the simulated response strategy respectively to obtain a plurality of learning parameters, the learning parameters including environment deviation data, anomaly deviation data, intrusion deviation data and response deviation data;

[0048] Step S6, the optimization module dynamically corrects and optimizes the intrusion intelligent agent, the defense intelligent agent and the regulation intelligent agent according to the learning parameters.

[0049] The beneficial effects of the present application are:

[0050] The present application introduces a simulation environment, virtual sensing technology and multi-agent reinforcement learning, realizes intelligent and dynamic security strategy optimization, breaks through the limitations of traditional security systems, significantly improves the intelligent level, adaptability and defense capability of the security system, improves the accuracy of intrusion detection and the effectiveness of defense measures, and has a wide application prospect. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 is a structural schematic diagram of the intelligent security control system in the present application;

[0052] Figure 2 is a step flow chart of the intelligent security control method in the present application.

[0053] Reference signs: 1, random generation module; 11, environment generation unit; 12, anomaly introduction unit; 2, virtual sensing module; 21, sensing modeling unit; 22, preliminary fusion unit; 3, cooperative simulation module; 4, actual monitoring module; 5, parameter comparison module; 6, optimization module; 7, framework construction module; 71, outer layer global strategy unit; 72, inner layer local strategy unit; 73, self-adaptive weight adjustment unit; 74, reward adjustment unit. DETAILED DESCRIPTION

[0054] The present application will be further described in detail below in combination with the drawings and embodiments. Identical parts are denoted by identical reference signs in the following description. It should be noted that the words "front", "back", "left", "right", "up" and "down" used in the following description refer to the directions in the drawings, and the words "bottom surface" and "top surface", "inner" and "outer" refer to the directions towards or away from the geometric center of a particular part.

[0055] Embodiment 1, with reference to Figure 1 As the first embodiment of the present application, the embodiment provides an intelligent security control system based on a simulation environment, which can improve the accuracy of intrusion detection and the effectiveness of defense measures, comprising:

[0056] Random generation module 1 is used to construct a simulation environment for an intelligent security system and randomly generate multiple variable parameters in the simulation environment, including simulation environment parameters and sensor anomaly parameters.

[0057] Virtual sensing module 2 is used to model multiple types of sensors to obtain multiple virtual sensing models. The virtual sensing models are used to output simulated sensing data and fuse the simulated sensing data to obtain simulated fused data.

[0058] The collaborative simulation module 3 connects the random generation module 1 and the virtual sensing module 2. It is used to construct an intrusion agent, a defense agent, and a control agent in the simulation environment based on a pre-introduced multi-agent reinforcement learning framework. The intrusion agent is used to simulate and generate various simulated intrusion behaviors and simulated response strategies. The control agent is used to generate a defense strategy based on the simulation fusion data when facing various simulated intrusion behaviors and simulated response strategies of the intrusion agent. The defense agent is used to adjust the camera angle, alarm threshold, and light status based on the defense strategy, simulation environment parameters, and sensor anomaly parameters.

[0059] Actual monitoring module 4 is used to acquire actual environmental parameters, actual abnormal parameters, actual intrusion behavior, and actual response strategies;

[0060] The parameter comparison module 5 connects the co-simulation module 3 and the actual monitoring module 4. It is used to compare the actual environmental parameters, actual abnormal parameters, actual intrusion behavior and actual response strategies with the simulated environmental parameters, sensor abnormal parameters, simulated intrusion behavior and simulated response strategies to obtain multiple learning parameters. The learning parameters include environmental deviation data, abnormal deviation data, intrusion deviation data and strain deviation data.

[0061] The optimization module 6 connects the parameter comparison module 5 and the collaborative simulation module 3, and is used to dynamically correct and optimize the intrusion agent, the defense agent and the control agent according to each learning parameter.

[0062] Working principle of Example 1:

[0063] The random generation module 1 is used to construct an intelligent security simulation environment and randomly generate multiple simulation environment parameters and sensor anomaly parameters within the environment to simulate different external conditions and sensor malfunctions. The simulation environment parameters include: temperature, humidity, air quality, illumination parameters, electromagnetic interference data, and biological interference data; the sensor anomaly parameters include: sensor noise data, signal interruption data, sensor conflict data, and link anomaly data, such as: increased signal noise from cameras in low-light environments; false detections by infrared sensors in high-temperature environments; and data drift caused by electromagnetic interference in radar.

[0064] The goal of the random generation module 1 is to generate diverse simulation data to ensure that the security system can adapt to complex environments. The virtual sensor module 2 models various sensors such as cameras, infrared sensors, pressure sensors, vibration sensors, and radars, and outputs simulated sensor data. This data is derived from the random generation module 1 and simulates the working state of the sensors in various environments. For example: camera: simulates target recognition ability in daytime, nighttime, low light, and strong light; infrared sensor: simulates the infrared characteristics of high temperature, low temperature, and dynamic targets; radar: simulates the echo changes of microwave radar in signal interference conditions.

[0065] All sensor data is processed through fusion to form simulation fusion data, improving the data accuracy of the system. For example: camera detects suspicious target movement trajectory; infrared sensor senses target temperature anomaly; radar analyzes target distance and speed changes. These fusion data will be used as input for subsequent agent learning and optimization.

[0066] In this embodiment, a multi-agent reinforcement learning framework is used to construct intrusion agents, defense agents, and regulation agents in the simulation environment to form a dynamic defense mechanism:

[0067] Intrusion agents: simulate illegal intrusion behaviors such as climbing over walls, blocking cameras, signal interference, password cracking, etc., and optimize their attack strategies to enhance the defense capabilities of the system.

[0068] Defense agents: adjust camera angles, increase alarm sensitivity, start strong light illumination, etc. according to simulation fusion data and defense strategies to respond to intrusion behaviors in real time.

[0069] Regulation agents: optimize alarm thresholds, sensor parameters, camera patrol paths, etc. according to the attack methods of intrusion agents and the response strategies of defense agents to reduce false positives and false negatives and improve system adaptability.

[0070] The actual monitoring module 4 is deployed in the real environment and is responsible for collecting real-time data, including: actual environmental parameters, actual abnormal parameters, actual intrusion behaviors, and actual response strategies.

[0071] These data are used to compare with simulation data to optimize system defense strategies.

[0072] The parameter comparison module 5 is used to compare simulation data with real data and calculate key deviations. Among them, environmental deviation: difference between simulation environment parameters and real environment; abnormal deviation: difference between simulation sensor anomalies and actual sensor false positives; intrusion deviation: matching degree between simulation intrusion behavior and real intrusion event; response deviation: difference in performance of defense agents in simulation and actual scenarios; and sensor fusion deviation data.

[0073] If the simulation system predicts an 80% probability of nighttime intrusion, but the actual monitoring data shows only 50%, the simulation environment needs to be adjusted to be closer to reality.

[0074] If a certain type of false alarm occurs frequently, the anomaly detection algorithm of the optimization system is optimized to improve the response accuracy of the defense agent.

[0075] The optimization module 6 dynamically optimizes the agent based on the data provided by the parameter comparison module 5: reinforcement learning optimization: continuously adjust the agent's strategy to improve the accuracy of identifying intrusion behavior. For example: after multiple training, the defense agent finds that there are more nighttime intrusions, automatically increases the scanning frequency of the nighttime patrol camera; by adjusting the alarm threshold, reduce false alarms, improve the practicality of the system.

[0076] When there is a large difference between the abnormal situation in the real environment (such as foggy weather) and the simulation environment, the optimization module 6 will update the simulation model to be closer to reality.

[0077] Application of this embodiment in an intelligent residential community: A certain community has installed this intelligent security system. One night, a suspicious person tried to climb over the wall to break in, and the system's workflow is as follows:

[0078] The infrared sensor detects temperature anomalies near the wall, and the camera synchronously captures the suspicious target;

[0079] The simulation system analyzes the behavior and finds that it has a high degree of match with previous simulated intrusion behaviors;

[0080] The defense agent automatically adjusts the camera angle to track the target and increases the alarm sensitivity;

[0081] The control agent analyzes the current environment and determines whether to turn on the high-intensity light to deter the target and start the alarm system;

[0082] The actual monitoring module 4 records this event and inputs its data into the parameter comparison module 5 for optimization of subsequent defense strategies.

[0083] This embodiment shows how an intelligent security control system based on a simulation environment optimizes defense strategies in a simulation environment and intelligently defends in the actual environment. Through reinforcement learning and dynamic adjustment, this system can improve the accuracy of intrusion detection, enhance the adaptability and defense capabilities of the security system, and is suitable for intelligent security, unattended monitoring and other scenarios.

[0084] Preferably, the random generation module 1 includes:

[0085] The environment generation unit 11 is configured to randomly match and generate a plurality of simulation environment parameters from a preset environment database according to positioning data and time data at a location where the intelligent security and protection control system based on a simulation environment, the simulation environment parameters including simulation temperature, simulation humidity, simulation air quality, simulation illumination parameters, simulation electromagnetic interference data, and simulation biological interference data.

[0086] The anomaly introduction unit 12 is configured to introduce sensor abnormality parameters in a sensor abnormal state in the simulation environment, the sensor abnormality parameters including simulation sensor noise data, simulation signal interruption data, simulation sensor conflict data, and simulation link abnormality data.

[0087] Specifically, in the embodiment, the environment generation unit 11 matches a plurality of simulation environment parameters from a preset environment database based on positioning data (GPS coordinates) and time data (current time, season, day and night information) of a location where the intelligent security and protection system is located, to simulate different weather, geographical environment and environmental interference factors.

[0088] The simulation temperature is determined according to the geographical location and the time, historical temperature data is queried from a meteorological database, and a fluctuation factor is randomly added to form the simulation temperature, so as to enhance the diversity of the simulation environment; the simulation humidity is randomly set according to the temperature and the weather condition, so as to ensure the matching with the temperature; the simulation air quality is generated from historical air quality index (AQI) data, and a time period (for example, the AQI is higher during morning and evening peak hours) is considered to form the simulation air quality; the simulation illumination parameters are obtained by combining the sunshine angle and the weather (sunny day / overcast day / rainy day); the simulation electromagnetic interference data is formed by simulating random noise interference generated by urban electromagnetic environment (for example, WiFi, 5G base station, industrial equipment); and the simulation biological interference data is used to simulate the interference of small animals and tree swaying on the sensor, so as to test the false alarm rate.

[0089] In order to improve the response capability of the intelligent security and protection system to abnormal conditions, the anomaly introduction unit 12 actively introduces the sensor abnormality state in the simulation environment, simulates the sensor failure or interference condition that may occur in actual deployment.

[0090] Among them, the simulation sensor noise data is used to simulate the noise of camera, infrared, sound and other sensors, such as image noise, audio background noise, temperature drift noise, etc. The simulation sensor noise data is simulated by using a Gaussian noise model; The simulation signal interruption data simulates signal loss, transmission delay and other situations to ensure that the system has the ability to resist network fluctuations. This data simulates the signal loss probability by using the Poisson process; The simulation sensor conflict data: when multiple sensors act on the same area at the same time, data conflicts may occur due to signal interference, for example: camera and infrared sensor data do not match, multiple radar signals interfere with each other, etc. The data uses covariance analysis to calculate the conflict degree; The simulation link abnormal data is used to simulate network fluctuations, data packet loss, bandwidth limitations and other situations to enable the security system to respond to actual network failures. This data is calculated by multiplying the network packet loss rate and bandwidth fluctuation.

[0091] This embodiment shows how the environment generation unit 11 and the anomaly introduction unit 12 work together to dynamically build a high-fidelity simulation environment and actively introduce various sensor anomalies to improve the adaptive ability, anti-interference ability and fault tolerance of the intelligent security system, ensuring stable operation in complex environments.

[0092] Preferably, the virtual sensor module 2 includes:

[0093] The sensor modeling unit 21 is configured to construct a plurality of virtual sensor models, the virtual sensor models being configured to simulate response curves, delays, noises and errors of various sensors, and output simulation sensor data, the simulation sensor data including simulation video data, simulation infrared data, simulation audio data and simulation vibration data.

[0094] The preliminary fusion unit 22 is connected to the sensor modeling unit 21 and is configured to construct a multi-sensor data fusion algorithm in the simulation environment to sequentially perform time-space alignment, information extraction and data fusion on different sources of simulation sensor data to obtain simulation fusion data.

[0095] Specifically, in this embodiment, the sensor modeling unit 21 constructs a plurality of virtual sensor models according to the physical characteristics and response mechanisms of real sensors to simulate the response curves, delays, noises and errors of different sensors. The following sensor types are mainly involved in this system:

[0096] Video sensor (camera): used for visible light environment monitoring, responding to changes in light and generating simulation video data.

[0097] Infrared sensor: used for detecting heat sources, significantly affected by environmental temperature, outputting simulation infrared data.

[0098] Audio sensor (microphone): used for monitoring environmental sound, which may be disturbed by background noise, outputting simulation audio data.

[0099] Vibration sensor: used to detect ground or object vibration, for identifying abnormal motion, outputting simulated vibration data.

[0100] Sensor response curve modeling: the response characteristics of each sensor are described by a response function, the standard response curve of the sensor needs to be modeled according to its physical characteristics:

[0101] The video sensor uses a linear Gamma correction model, the infrared sensor uses a nonlinear temperature response model, the audio sensor uses a frequency spectrum analysis model, and the vibration sensor uses an elastic damping model.

[0102] Noise and error terms are used to simulate the noise interference of real sensors, noise is simulated using Gaussian noise, and error is simulated using signal drift.

[0103] Generation of simulated sensor data: running the sensor model in the simulation environment, generating simulated video data, simulated infrared data, simulated audio data and simulated vibration data to provide complete multi-modal perception information.

[0104] Implementation of the preliminary fusion unit 22:

[0105] (1) Data space-time alignment

[0106] Due to different sampling frequencies and time delays of different sensors, space-time alignment must be performed. Dynamic time warping (DTW) method is used to adjust the time stamps of different sensor data.

[0107] (2) Information extraction

[0108] Extract key features from various sensor data:

[0109] Video data uses target detection (YOLO) and motion trajectory analysis to extract key features; infrared data uses heat source distribution and target temperature change to extract key features; audio data uses voiceprint recognition and abnormal sound source positioning to extract key features; vibration data uses vibration amplitude and frequency pattern to extract key features.

[0110] (3) Data fusion calculation

[0111] Using data fusion calculation formula, combining multi-sensor data, calculating simulated fusion data.

[0112] The embodiment constructs a simulation environment through a virtual sensing module 2, and completes data modeling and fusion by using a sensing modeling unit 21 and a preliminary fusion unit 22. The embodiment effectively improves the authenticity of the simulation environment, enhances the adaptability to complex environments, reduces the false alarm rate of sensors, eliminates the wrong judgment of a single sensor through data fusion, and optimizes the system decision-making ability to provide more accurate input data for intelligent security control strategies.

[0113] The embodiment realizes modeling, space-time alignment and fusion of multi-sensor data, ensures that the intelligent security system can accurately simulate the real situation in the simulation environment, and optimizes the defense strategy.

[0114] Preferably, the data fusion calculation formula of the multi-sensor data fusion algorithm is configured as:

[0115]

[0116]

[0117]

[0118] wherein, represents the simulation fusion data, represents the total time of fusion data calculation, represents the discrete series index, represents a set of real vectors of dimension, represents the Gamma function for controlling the scale change of data, represents the Riemann Zeta function for enhancing the nonlinear mapping of data, represents the first kind of Bessel function for regulating the smoothness of data, represents a complex information filtering function for combining the error function, the logarithmic function and the elliptic integral to optimize data fusion, represents different types of sensing data, wherein is simulation video data, is simulation infrared data, is simulation audio data, is simulation vibration data, represents the weight of each type of sensor data for distribution according to the confidence and contribution of the data, represents the index of the sensing data in the spatial domain, is a filtering input variable, is an integral variable, is the first kind of elliptic integral for correcting the nonlinear error of data.

[0119] Specifically, in the embodiment, the value range of is determined by the integral interval and the distribution of each sensor data affects:

[0120] When the sensor data signal is strong and consistent, the value is larger, indicating that the data fusion quality is higher; when the sensor data noise is larger or the information is not matched, the value is smaller, indicating that the uncertainty of the fused data is higher; the formula combines the Gamma function, the Riemann Zeta function and the Bessel function to ensure the stability and nonlinear mapping effect of data fusion, and uses the error function and elliptic integral to optimize the information, minimizing the error of data fusion. This formula can accurately describe the process of multi-sensor data fusion, and realize nonlinear mapping, noise suppression and optimal data fusion strategy by combining multiple complex mathematical functions, which can be used for simulation data fusion tasks in complex scenarios such as intelligent security, autonomous driving and multi-modal recognition.

[0121] Embodiment 2 is the second embodiment of the present application, which is different from the previous embodiment. The embodiment provides a framework modeling module 7, which can realize efficient strategy optimization and intelligent control of the intelligent security system, which includes a framework modeling module 7 connected with the cooperative simulation module 3. The framework modeling module 7 is used to build a hierarchical reinforcement learning framework, and includes:

[0122] an outer global policy unit 71 for generating security policy sub-targets according to hierarchical deep reinforcement learning, and transmitting the security policy sub-targets to a regulation intelligent agent, and the regulation intelligent agent generates defense strategies based on the security policy sub-targets;

[0123] an inner local policy unit 72 for fine-grained adjustment of cameras, alarm devices and lights, and training the adjustment method based on policy gradient method to optimize the local control task of the defense intelligent agent on camera angle, alarm threshold and light state.

[0124] Preferably, the hierarchical reinforcement learning framework adopts a hierarchical reward structure, which includes overall reward, outer reward and inner reward, and the conversion relationship between the overall reward, the outer reward and the inner reward is configured as:

[0125] wherein, represents the overall reward, representing the defense effect of the entire intelligent security system; represents the outer reward, used to evaluate the long-term security and resource optimization effect of the overall security strategy; represents the inner reward, used to evaluate the real-time effect of the local control strategy (camera, alarm, light) in a short time; represents the outer weight, represents the inner weight, total number of global rewards

[0126] The working principle of embodiment 2 is as follows:

[0127] The framework building module 7 adopts hierarchical deep reinforcement learning, including an outer layer global policy unit 71 and an inner layer local policy unit 72. The outer layer global policy unit 71 is responsible for formulating an overall security policy, generating sub-goals, and passing them to the regulation agent. The outer layer global policy unit 71 is based on hierarchical deep reinforcement learning (HDRL) and adopts an options framework, enabling the outer layer agent to generate optimal sub-goals in different scenarios. Task examples include:

[0128] High alert mode (increasing monitoring intensity when threats are detected), energy-saving monitoring mode (optimizing energy consumption when there are no threats), and full defense mode (increasing the monitoring intensity of the entire region).

[0129] The outer layer global policy unit 71 performs reinforcement learning training based on sensor data fusion results, historical alarm data, security events, and other factors. The outer layer agent passes the generated security policy sub-goals to the regulation agent, which optimizes the defense strategy based on the sub-goals.

[0130] The inner layer local policy unit 72 is responsible for fine-grained adjustment of cameras, alarm devices, and lights to improve local defense efficiency. This unit is trained based on policy gradient methods and optimizes local tasks such as camera angles, alarm thresholds, and light control through agent reinforcement learning.

[0131] Task examples include camera control: adjusting camera angles to maximize monitoring coverage. Alarm optimization: optimizing alarm sensitivity based on sensor data to reduce false positives. Light control: optimizing lighting based on nighttime security needs to ensure visibility.

[0132] Preferably, the actual environmental parameters include actual temperature, actual humidity, actual air quality, actual lighting parameters, actual electromagnetic interference data, and actual biological interference data. The framework building module 7 further includes an adaptive weight adjustment unit 73 for establishing a conversion relationship between the actual environmental parameters and the outer layer weight and the inner layer weight, and dynamically adjusting the outer layer weight and the inner layer weight based on the actual environmental parameters. The conversion relationship between the actual environmental parameters and the outer layer weight and the inner layer weight is configured as follows:

[0133]

[0134] in, Indicates the actual temperature. Indicates actual humidity. Indicates actual air quality. Indicates actual illumination parameters. This represents actual electromagnetic interference data. This represents actual biological interference data. This represents the average actual temperature. This represents the average actual humidity. This represents the average actual air quality. This represents the average value of the actual illumination parameters. This represents the average value of the actual electromagnetic interference data. This represents the average value of the actual biological interference data. , , , , , These represent the value ranges of actual temperature, actual humidity, actual air quality, actual light parameters, actual electromagnetic interference data, and actual biological interference data, respectively.

[0135] Specifically, in this embodiment, the formula uses a standardization method to normalize each environmental parameter to ensure that different physical quantities have the same dimensions, so that they can be processed within the same calculation framework.

[0136] Controlled by the Sigmoid function The value range is between (0,1), which allows for smooth adjustment of the global weights and ensures stability. Depend on Ensure that the sum of the global and local policy weights is 1, allowing for adaptive adjustment.

[0137] The more the external environment deviates from the average (i.e., the more severe the external disturbance), the more likely it is to cause problems. When the value approaches 1, the global strategy becomes more important to ensure system stability; conversely, when the environment is stable, As the value approaches 1, local strategies become more important in order to optimize resource efficiency.

[0138] When the environment is completely stable (all parameters are equal to their average values): the sum of the normalized terms is 0. , =0.5, at this time and balance.

[0139] When the environment fluctuates drastically (a parameter deviates far from its mean): the absolute value of the normalized term increases, and the exponential term tends to reach a maximum or minimum. Approach 1 or 0, so that the global or local strategy dominates.

[0140] The formula can be used in intelligent security systems, enabling the system to adaptively adjust the decision weight according to environmental changes, optimizing the balance between global and local strategies.

[0141] Preferably, the actual abnormal parameters include actual video data, actual infrared data, actual audio data and actual vibration data, and the framework modeling module 7 further includes a reward adjustment unit 74 for calculating a sensor comprehensive abnormal value according to the actual video data, the actual infrared data, the actual audio data and the actual vibration data, and dynamically adjusting the overall reward according to the sensor comprehensive abnormal value. The conversion relationship between the sensor comprehensive abnormal value and the overall reward is configured as:

[0142]

[0143]

[0144] wherein, represents the sensor comprehensive abnormal value, represents the actual video data abnormal value,

[0145] represents the actual infrared data abnormal value, represents the actual audio data abnormal value, represents the actual vibration data abnormal value, , , , respectively represent the mean value of each sensor data, , , , respectively represent the value range of each sensor data, represents the preset basic overall reward, represents an abnormality suppression coefficient for controlling the degree of attenuation of the abnormal value on the overall reward, represents a preset abnormality compensation coefficient for preventing the reward value from being excessively reduced, for providing smooth abnormality adjustment effect and preventing mutation.

[0146] Specifically, in the present embodiment, the sensor comprehensive abnormal value is calculated by adopting a standardization method to normalize each sensor abnormal data and taking the mean value, so as to ensure that different physical quantities are calculated under the same dimension. If a sensor abnormal value deviates from the mean value by a large margin, the overall abnormal value increases.

[0147] Adjusting the overall reward : an exponential decay function is adopted When the anomaly increases, the overall reward gradually decreases, ensuring that the system reduces incentives when the anomaly increases, enhancing the safety policy.

[0148] The hyperbolic tangent function is used To ensure that the reward adjustment is smooth when the anomaly is small, avoiding over-sensitivity of the system to minor anomalies. Parameters and Can be adjusted according to actual needs, for example, in a security system, if the false alarm cost is high, a larger can be set to improve the suppression ability of the anomaly.

[0149] Value range analysis:

[0150] Normal environment (no anomaly): ≈0, then

[0151] It indicates that in the normal state, the system reward maintains the default level.

[0152] Anomaly increases (sensor data deviates from the mean value greatly): Increases, the exponential term Rapidly decays, making the overall reward decrease, indicating that the system enters an abnormal state.

[0153] Extreme anomaly (highly abnormal system): →∞, then It indicates that when the anomaly is very large, the system's reward still retains a certain lower limit, preventing excessive punishment from affecting policy optimization. This formula can be applied to intelligent security systems, intelligent control systems, multi-sensor data fusion, etc. The system can dynamically adjust the reward according to the anomaly situation, ensuring that the reward is maximized in a normal environment and appropriately reducing incentives in an abnormal environment, to enhance the safety and robustness of the system.

[0154] Preferably, in a hierarchical reinforcement learning framework, including edge computing devices and cloud servers, edge computing devices are used for real-time execution of local control tasks, and cloud servers are used for continuous updating of defense strategies.

[0155] Specifically, in this embodiment, the system uses an edge-cloud collaborative computing architecture, which consists of the following:

[0156] Edge computing device:

[0157] Device type: smart camera, edge server (NVIDIA Jetson, Intel Movidius), embedded AI chip, etc.

[0158] Main tasks: Real-time execution of local control tasks, including camera angle adjustment, alarm triggering, light control, etc. Local data processing reduces transmission delay and avoids cloud computing bottlenecks. Preliminary reinforcement learning inference uses local models to quickly calculate the current optimal local control strategy.

[0159] Cloud server:

[0160] Device type: High-performance GPU servers (such as AWS EC2, Google Cloud TPU), AI computing clusters.

[0161] Main tasks: Global defense strategy optimization, training reinforcement learning models based on large amounts of historical data, and generating global optimization strategies. Continuous model updates will distribute the latest learned optimization strategies to edge computing devices to improve intelligence levels. Large-scale data analysis processes massive video, infrared, audio, and vibration sensor data to provide global threat assessment capabilities.

[0162] Workflow:

[0163] When the system is running, edge computing devices and cloud servers work together, with the following division of labor:

[0164] (1) Edge computing devices execute real-time local control tasks

[0165] When an abnormal event is detected in the target area, the edge computing device first analyzes the current environmental data and executes the optimal control strategy using the deployed reinforcement learning model.

[0166] For example: The camera adjusts to the best monitoring angle to improve the tracking ability of suspicious targets. The alarm device adjusts the alarm volume or triggers an emergency notification based on the real-time threat level. The lighting system dynamically adjusts the brightness based on the location of the intrusion event to provide optimal monitoring conditions. Since the calculation is performed locally on the edge side, the response time of the entire process is usually in milliseconds, ensuring high real-time performance of the system.

[0167] (2) Cloud server continuously optimizes defense strategies

[0168] The cloud server collects data from multiple edge devices, analyzes the effectiveness of long-term defense strategies, and continuously optimizes global strategies using reinforcement learning algorithms. The specific optimization process includes:

[0169] Data synchronization: Edge computing devices regularly upload abnormal detection data, local execution results, and sensor information.

[0170] Model training: The cloud server trains global reinforcement learning models using historical data to optimize the long-term effectiveness of strategies.

[0171] Policy delivery: After training, the cloud server sends the latest optimized defense strategy back to the edge device to update its control logic.

[0172] To improve the efficiency and stability of the system, the following optimization strategies are adopted in the edge-cloud computing architecture:

[0173] (1) Edge computing task optimization

[0174] Lightweight AI model: Adopt model pruning, quantization, and other techniques to enable reinforcement learning inference computation to run efficiently on low-power devices such as embedded AI chips.

[0175] Asynchronous inference mechanism: For urgent tasks such as intrusion detection, the system prioritizes local decision-making to avoid delays caused by waiting for cloud computing.

[0176] (2) Cloud optimization strategy

[0177] Federated learning: Avoid uploading all data to the cloud, and the cloud server only synchronizes the model weights learned by each edge device to improve privacy security.

[0178] Adaptive training scheduling: The cloud server dynamically adjusts the policy update frequency based on the task load of the edge device to ensure that the device's computing resources are not overused.

[0179] This embodiment effectively improves the real-time performance, stability, and adaptability of the intelligent security system through the hierarchical reinforcement learning architecture of edge computing devices and cloud servers, ensuring that the system maintains efficient and intelligent security control capabilities in various complex environments.

[0180] An intelligent security control method based on a simulation environment, applied to the intelligent security control system based on a simulation environment described above, with reference to Figure 2 , comprising:

[0181] Step S1, the random generation module 1 constructs a simulation environment of the intelligent security system and randomly generates a plurality of change parameters in the simulation environment, the change parameters including simulation environment parameters and sensor anomaly parameters;

[0182] Step S2, the virtual sensor module 2 models a plurality of sensors to obtain a plurality of virtual sensor models, the virtual sensor models are used to output simulation sensor data, and the simulation sensor data are fused to obtain simulation fusion data;

[0183] Step S3, the cooperative simulation module 3 constructs the intrusion agent, the defense agent and the regulation agent in the simulation environment according to the pre-introduced multi-agent reinforcement learning framework, the intrusion agent is used for simulating generating various simulated intrusion behaviors and simulated response strategies, the regulation agent is used for generating defense strategies according to simulation fusion data when facing various simulated intrusion behaviors and simulated response strategies of the intrusion agent, and the defense agent is used for adjusting the camera angle, the alarm threshold and the light state according to the defense strategy, the simulation environment parameter and the sensing abnormal parameter;

[0184] Step S4, the actual monitoring module 4 acquires actual environment parameters, actual abnormal parameters, actual intrusion behaviors and actual response strategies;

[0185] Step S5, the parameter comparison module 5 compares the actual environment parameters, the actual abnormal parameters, the actual intrusion behaviors and the actual response strategies with the simulation environment parameters, the sensing abnormal parameters, the simulated intrusion behaviors and the simulated response strategies respectively to obtain a plurality of learning parameters, and the learning parameters include environment deviation data, abnormal deviation data, intrusion deviation data and response deviation data;

[0186] Step S6, the optimization module 6 dynamically corrects and optimizes the intrusion agent, the defense agent and the regulation agent according to the learning parameters.

[0187] The above is only the preferred embodiment of the present application, the protection scope of the present application is not limited to the above-mentioned examples only, any technical scheme belonging to the idea of the present application is also within the protection scope of the present application. It should be noted that for ordinary skilled in the art, some improvements and decorations without departing from the principles of the present application are also considered as the protection scope of the present application.

Claims

1. An intelligent security control system based on a simulation environment, characterized in that, The system comprises: a random generation module (1) for constructing a simulation environment of an intelligent security system and randomly generating a plurality of change parameters in the simulation environment, the change parameters including simulation environment parameters and sensor anomaly parameters; a virtual sensor module (2) for modeling a plurality of types of sensors to obtain a plurality of virtual sensor models, the virtual sensor models being used to output simulation sensor data and fuse the simulation sensor data to obtain simulation fusion data; a collaborative simulation module (3) connected to the random generation module (1) and the virtual sensor module (2), for constructing an intrusion agent, a defense agent and a regulation agent in the simulation environment according to a pre-introduced multi-agent reinforcement learning framework, the intrusion agent being used to simulate a plurality of simulated intrusion behaviors and simulated response strategies, the regulation agent being used to generate a defense strategy according to the simulation fusion data when facing the simulated intrusion behaviors and the simulated response strategies of the intrusion agent, and the defense agent being used to adjust a camera angle, an alarm threshold and a light state according to the defense strategy, the simulation environment parameters and the sensor anomaly parameters; an actual monitoring module (4) for acquiring actual environment parameters, actual anomaly parameters, actual intrusion behaviors, actual sensor fusion data and actual response strategies; a parameter comparison module (5) connected to the collaborative simulation module (3) and the actual monitoring module (4), for comparing the actual environment parameters, the actual anomaly parameters, the actual intrusion behaviors, the actual sensor fusion data and the actual response strategies with the simulation environment parameters, the sensor anomaly parameters, the simulated intrusion behaviors, the simulation sensor data and the simulated response strategies to obtain a plurality of learning parameters, the learning parameters including environment deviation data, anomaly deviation data, intrusion deviation data, sensor fusion deviation data and response deviation data; an optimization module (6) connected to the parameter comparison module (5) and the collaborative simulation module (3), for dynamically correcting and optimizing the intrusion agent, the defense agent and the regulation agent according to the learning parameters; and a framework construction module (7) connected to the collaborative simulation module (3), the framework construction module (7) being used to construct a hierarchical reinforcement learning framework and comprising: an outer layer global policy unit (71) for generating a security strategy sub-goal according to hierarchical deep reinforcement learning and transmitting the security strategy sub-goal to the regulation agent, the regulation agent generating the defense strategy based on the security strategy sub-goal; an inner layer local policy unit (72) for fine-grained adjustment of a camera, an alarm device and a light, and training the adjustment mode based on a policy gradient method to optimize the local control task of the defense agent on the camera angle, the alarm threshold and the light state; the hierarchical reinforcement learning framework adopts a hierarchical reward structure, the hierarchical reward structure including an overall reward, an outer layer reward and an inner layer reward, and the conversion relationship between the overall reward, the outer layer reward and the inner layer reward being configured as: ; wherein, represents the overall reward, for reflecting the defense effect of the entire security system, represents the outer-layer reward, for evaluating the long-term security and resource optimization effect of the global security strategy, represents the inner-layer reward, for the user to evaluate the real-time control effect of the local device, represents the outer-layer weight, represents the inner-layer weight, represents the total number of inner-layer rewards.

2. The intelligent security control system based on simulation environment according to claim 1, characterized in that: The random generation module (1) comprises: An environment generation unit (11) configured to generate a plurality of simulation environment parameters from a preset environment database according to positioning data and time data at a location where the simulation environment-based intelligent security control system is located, the simulation environment parameters comprising simulation temperature, simulation humidity, simulation air quality, simulation illumination parameters, simulation electromagnetic interference data, and simulation biological interference data; An anomaly introduction unit (12) configured to introduce the sensor abnormality parameters in a state of sensor abnormality into the simulation environment, the sensor abnormality parameters comprising simulation sensor noise data, simulation signal interruption data, simulation sensor conflict data, and simulation link abnormality data. 3.The intelligent security control system based on simulation environment according to claim 1, characterized in that: The virtual sensor module (2) comprises: A sensor modeling unit (21) configured to construct a plurality of virtual sensor models, the virtual sensor models being configured to simulate response curves, delays, noises, and errors of various sensors and output simulation sensor data, the simulation sensor data comprising simulation video data, simulation infrared data, simulation audio data, and simulation vibration data; A preliminary fusion unit (22) connected to the sensor modeling unit (21) and configured to construct a multi-sensor data fusion algorithm in the simulation environment to sequentially perform time-space alignment, information extraction, and data fusion on different sources of the simulation sensor data to obtain simulation fusion data.

4. The intelligent security control system based on simulation environment according to claim 3, characterized in that: The data fusion calculation formula of the multi-sensor data fusion algorithm is configured as: ; ; ; wherein, represents the simulated fused data, represents the total time of the fused data computation, represents the discrete series index, represents a set of real-valued vectors, represents the Gamma function for controlling the scale variation of data, represents the Riemann Zeta function for enhancing the nonlinear mapping of data, represents the first kind of Bessel function for regulating the smoothness of data, represents the complex information filtering function for combining the error function, logarithmic function, and elliptic integral to optimize data fusion, represents the representative of different types of sensor data, wherein is the simulated video data, is the simulated infrared data, is the simulated audio data, is the simulated vibration data, represents the weight of each type of sensor data for allocation according to the confidence and contribution of the data, represents the index of the sensor data in the spatial domain, is the filtering input variable, is the integral variable, is the first kind of elliptic integral for correcting the nonlinear error of data, represents the filtering adjustment parameter.

5. The intelligent security control system based on simulation environment according to claim 1, characterized in that: The actual environment parameters comprise actual temperature, actual humidity, actual air quality, actual illumination parameters, actual electromagnetic interference data, and actual biological interference data, and the framework construction module (7) further comprises an adaptive weight adjustment unit (73) configured to establish a conversion relationship between the actual environment parameters and the outer layer weight and the inner layer weight, and dynamically adjust the outer layer weight and the inner layer weight according to the actual environment parameters, wherein the conversion relationship between the actual environment parameters and the outer layer weight and the inner layer weight is configured as: ; wherein represents the actual temperature, represents the actual humidity, represents the actual air quality, represents the actual light parameter, represents the actual electromagnetic interference data, represents the actual biological interference data, represents the average of the actual temperature, represents the average of the actual humidity, represents the average of the actual air quality, represents the average of the actual light parameter, represents the average of the actual electromagnetic interference data, represents the average of the actual biological interference data, , , , , , represents the value range of the actual temperature, the actual humidity, the actual air quality, the actual light parameter, the actual electromagnetic interference data and the actual biological interference data, respectively. 6.The intelligent security control system based on simulation environment according to claim 1, characterized in that: The actual abnormality parameters comprise actual video data, actual infrared data, actual audio data, and actual vibration data, and the framework construction module (7) further comprises a reward adjustment unit (74) configured to calculate a sensor comprehensive abnormality value according to the actual video data, the actual infrared data, the actual audio data, and the actual vibration data, and dynamically adjust the overall reward according to the sensor comprehensive abnormality value, wherein a conversion relationship between the sensor comprehensive abnormality value and the overall reward is configured as: ; ; wherein, represents the sensor composite anomaly value, represents the actual video data anomaly value, represents the actual infrared data anomaly value, represents the actual audio data anomaly value, represents the actual vibration data anomaly value, , , , respectively represent the mean value of each sensor data, , , , respectively represent the value range of each sensor data, represents the preset basic overall reward, represents the anomaly suppression coefficient, used to control the attenuation degree of the anomaly value on the overall reward, represents the preset anomaly compensation coefficient, used to prevent the reward value from being excessively reduced, used to provide a smooth anomaly adjustment effect and prevent mutation.

7. The simulation environment based intelligent security control system according to claim 1, wherein: In the hierarchical reinforcement learning framework, an edge computing device is configured to perform the local control task in real time, and a cloud server is configured to continuously update the defense strategy.

8. The intelligent security control method based on the simulation environment, applied to the intelligent security control system based on the simulation environment in any one of claims 1-7, characterized in that, Comprise: Step S1, the random generation module (1) constructs a simulation environment of an intelligent security system, and randomly generates a plurality of change parameters in the simulation environment, the change parameters comprising simulation environment parameters and sensor abnormality parameters; Step S2, the virtual sensor module (2) models multiple types of sensors to obtain multiple virtual sensor models, the virtual sensor models are used to output simulation sensor data, and simulation fusion data is obtained by fusing each simulation sensor data; Step S3, the cooperative simulation module (3) constructs an intrusion agent, a defense agent and a regulation agent in the simulation environment according to a pre-introduced multi-agent reinforcement learning framework, the intrusion agent is used to simulate multiple simulated intrusion behaviors and simulated response strategies, the regulation agent is used to generate a defense strategy according to the simulation fusion data when facing each simulated intrusion behavior and simulated response strategy of the intrusion agent, and the defense agent is used to adjust the camera angle, alarm threshold and light state according to the defense strategy, simulation environment parameters and sensor anomaly parameters; Step S4, the actual monitoring module (4) acquires actual environment parameters, actual anomaly parameters, actual intrusion behaviors and actual response strategies; Step S5, the parameter comparison module (5) compares the actual environment parameters, actual anomaly parameters, actual intrusion behaviors and actual response strategies with the simulation environment parameters, sensor anomaly parameters, simulated intrusion behaviors and simulated response strategies respectively to obtain multiple learning parameters, the learning parameters include environment deviation data, anomaly deviation data, intrusion deviation data and response deviation data; Step S6, the optimization module (6) dynamically corrects and optimizes the intrusion agent, defense agent and regulation agent according to each learning parameter.

Citation Information

Patent Citations

  • Optical fiber sensing based borderline boundary anti-intrusion alarming system and method

    CN108765814A

  • Image recognition system based on machine vision

    CN119399534A