Indoor Smoking Behavior Detection Method and System Based on Transient Change Characteristics
By capturing data streams from multiple sources of sensors to construct a dynamic behavior response model and using a transient feature analysis algorithm to generate an intelligent agent, the problem of high false positive rate and privacy invasion in indoor smoking behavior detection is solved, achieving high accuracy and rapid environmental intervention.
Patent Information
- Application Number
- CN202511521431.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing indoor smoking detection technologies suffer from high false alarm rates, difficulty in distinguishing between e-cigarettes and traditional cigarettes, susceptibility to interference, and privacy violations. Traditional smoke detectors and camera monitoring technologies are ineffective in complex environments.
An indoor smoking behavior detection system based on transient mutation features is adopted. Data streams are captured in real time by multi-source sensors, a dynamic behavior response model is constructed, and a mutation-sensing intelligent agent is generated using a transient feature analysis algorithm to execute environmental intervention commands.
It improves the accuracy and adaptability of smoking behavior detection, reduces the false positive rate, enables rapid and effective environmental intervention, and protects air quality and privacy.
Smart Images

Figure CN120995285B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of indoor behavior detection technology, specifically to a method and system for detecting indoor smoking behavior based on transient change characteristics. Background Technology
[0002] Smoking poses a significant threat to human health, severely endangering not only the smoker's own health but also exposing secondhand smoke to those around them. This harm is particularly pronounced indoors because the relatively enclosed space makes it difficult for smoke and harmful components to dissipate quickly. As societal awareness of health issues continues to rise, the importance of indoor smoking control is becoming increasingly apparent.
[0003] Currently, indoor smoking detection technology faces numerous challenges. Traditional smoke detectors are primarily designed for fire smoke and have significant limitations when detecting e-cigarettes and traditional cigarettes. The smoke emitted by e-cigarettes and traditional cigarettes differs from regular smoke in composition and concentration. Traditional smoke detectors often fail to accurately detect the emissions from e-cigarette devices, making it even more difficult to issue effective alarms, resulting in unsatisfactory detection performance. Some e-cigarette detectors on the market rely solely on a single sensor to detect the concentration of smoke in the air. This single detection method not only struggles to effectively distinguish between e-cigarettes and traditional cigarettes but is also highly susceptible to interference from aerosols, volatile organic compounds from cleaning products, and vapors, leading to false alarms or masking actual smoking behavior in complex environments.
[0004] When camera surveillance technology is applied to detect smoking, significant problems have emerged. When a person is in a blind spot, such as being obstructed or facing away from the camera, the technology cannot accurately analyze smoking behavior, resulting in a significant reduction in detection rate and accuracy. Furthermore, camera surveillance may involve privacy violations, making it difficult to fully protect individual privacy. This is a major factor hindering the widespread application of this technology in places with high privacy requirements, such as private offices and changing rooms. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for detecting indoor smoking behavior based on transient change characteristics, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides an indoor smoking behavior detection system based on transient change characteristics, the system comprising:
[0007] Transient sensing module: Real-time capture of multi-source sensor data streams within the target monitoring space;
[0008] Environmental modeling module: Based on the multi-source sensor data stream, a dynamic behavior response model is constructed, which formalizes the smoking behavior detection task into an environmental state evolution process;
[0009] Behavior decision module: Employs transient feature parsing algorithm as behavior detection engine to generate mutation-aware intelligent agent;
[0010] Execution verification module: Executes environmental intervention instructions based on the behavior judgment results output by the mutation-aware agent.
[0011] Preferably, the dynamic behavior response model comprises a four-element structure:
[0012] The state space is defined as a multidimensional feature vector of the current environment, including the transient feature values, historical behavior markers, and environmental physical parameters of the multi-source sensor data stream;
[0013] Action space is defined as the set of judgment levels for suspected smoking behavior;
[0014] The state transition function maps the current environment state to the behavior decision action and then to the next environment state.
[0015] The reward function calculates the environmental feedback value based on the accuracy rate of the behavior judgment results and the timeliness of the response.
[0016] Preferably, the transient feature parsing algorithm includes:
[0017] The mutation sensing unit: performs parameterized modeling of the multi-source sensing data stream through a transient energy field analysis network. The input vector is the current environmental state and the behavior judgment action, and the output value is the environmental feedback prediction value.
[0018] Behavior Decision Unit: Inputs the current environment state through a dynamic gradient tracking network and outputs the probability distribution of the behavior decision action.
[0019] Preferably, the transient energy field analysis network adopts a dual-network architecture, specifically:
[0020] Two independent feature parsing networks are constructed, and the minimum value between the two is selected as the final environmental feedback prediction value.
[0021] Each feature parsing network comprises a transient feature extraction layer and an energy field aggregation layer. The transient feature extraction layer captures abrupt change patterns in the sensor data through a trainable nonlinear function, and the energy field aggregation layer fuses the spatial correlation of multidimensional features.
[0022] Preferably, the dynamic gradient tracking network comprises:
[0023] The input layer receives the current environment state vector;
[0024] The state evolution layer iteratively updates the behavioral decision path through multi-level dynamic gradient units;
[0025] The output layer generates an optimized probability distribution of the behavior determination action.
[0026] Preferably, the operation flow of the behavior decision module is as follows:
[0027] Initialize the network parameters of the transient feature parsing algorithm;
[0028] The current environmental state is obtained from the dynamic behavior response model.
[0029] The behavior decision unit generates the current behavior determination action.
[0030] The environmental state at the next moment is calculated based on the state transition function.
[0031] Calculate the environmental feedback value based on the reward function;
[0032] The storage environment state transfer experience data is transferred to the circular memory pool.
[0033] Preferably, the update mechanism of the behavior decision module is as follows:
[0034] Periodically sample environmental state transfer experience data from the circulating memory pool;
[0035] The target environment feedback prediction value is calculated through the mutation sensing unit;
[0036] The parameter weights of the transient energy field analysis network are updated based on the difference between the predicted value of the target environment feedback and the actual environment feedback value.
[0037] The policy parameters of the dynamic gradient tracking network are updated by maximizing the environmental feedback prediction values.
[0038] Preferably, the calculation method for the environmental feedback value includes:
[0039] When a sudden change in smoke concentration is detected to occur simultaneously with a change in thermal radiation waveform, a positive environmental feedback value is assigned.
[0040] When a continuous match is detected between human posture features and waving gesture features, a secondary environmental feedback value is assigned;
[0041] When environmental physical parameters exceed the preset safety threshold, a negative environmental feedback value is assigned.
[0042] Preferably, the execution verification module includes:
[0043] Multi-source collaborative verification unit: compares the spatiotemporal consistency between the real-time sensing data stream captured by the transient sensing module and the behavior determination result;
[0044] Instruction distribution unit: When multi-source verification passes, the execution logic of the environmental intervention instruction is triggered.
[0045] Preferably, the present invention also includes an indoor smoking behavior detection method based on transient change characteristics, the method comprising all the modules and method flow of the indoor smoking behavior detection system based on transient change characteristics described above.
[0046] Compared with the prior art, the beneficial effects of the present invention are:
[0047] The system's transient sensing module can capture multi-source sensor data streams within the target monitoring space in real time. Traditional detection methods often rely on a single smoke sensor or camera, making it difficult to comprehensively acquire environmental information. This system, however, integrates multiple sensors to simultaneously collect multi-dimensional data such as temperature, humidity, air quality, and images. This allows for more accurate capture of subtle environmental changes related to smoking behavior. For example, smoking releases heat, causing a temporary rise in local temperature and altering surrounding air quality. The multi-source sensor data stream can capture these changes simultaneously, greatly improving the comprehensiveness and accuracy of capturing smoking-related features.
[0048] The environmental modeling module constructs a dynamic behavior response model based on multi-source sensor data streams, formalizing the smoking behavior detection task as an environmental state evolution process. This innovative method is more adaptable than traditional fixed-model detection. In complex and changing indoor environments, such as offices with frequent personnel movement and equipment interference, traditional models are prone to misjudgment. However, this system's dynamic behavior response model can continuously adjust its understanding of the environmental state based on real-time collected data, automatically adapting to environmental changes. For example, during meetings in a conference room, factors such as personnel activity and heat dissipation from electronic devices cause continuous changes in the environmental state. The model can accurately determine whether smoking behavior exists based on the dynamic changes in multi-source data, effectively avoiding false positives and false negatives caused by environmental changes.
[0049] The behavior decision-making module employs a transient feature analysis algorithm as its behavior detection engine to generate a mutation-aware intelligent agent. This algorithm can keenly identify transient mutation features in multi-source sensor data streams, which are key signals of smoking behavior. Unlike traditional detection algorithms based on fixed thresholds or simple pattern matching, this algorithm can deeply analyze data change trends and feature combinations. For example, in image data, traditional algorithms might only judge based on the presence of smoke, while the transient feature analysis algorithm comprehensively analyzes the instantaneous changes in human movements and object positions within the image, as well as the correlation with other sensor data, such as the synchronicity between a sudden increase in harmful gas concentration in air quality data and the smoking action in the image. This allows for a more accurate determination of smoking behavior and significantly reduces the false positive rate.
[0050] The execution verification module executes environmental intervention commands based on the behavior judgment results output by the mutation-sensing agent. This module achieves seamless integration of detection and intervention. Once smoking is detected, corresponding measures can be taken immediately, such as activating ventilation equipment to accelerate air circulation and reduce the concentration of smoke and harmful components indoors; simultaneously, alarms are triggered to remind relevant personnel, and the monitoring system can be linked to record smoking behavior and on-site conditions, providing a basis for subsequent management. In some places with extremely high air quality requirements, such as hospital operating rooms and precision instrument laboratories, this rapid and effective environmental intervention measure can minimize the harm of smoking to the environment and personnel, ensuring the normal operation of the venue and the health of personnel. Attached Figure Description
[0051] Figure 1 This is a timing diagram of the indoor smoking behavior detection system based on transient change characteristics described in this invention.
[0052] Figure 2 A flowchart of the four-element structure of a dynamic behavior response model;
[0053] Figure 3 This is a graph showing the results of running the dynamic behavior response model;
[0054] Figure 4 A flowchart of the dual-network architecture for transient energy field analysis networks;
[0055] Figure 5 The diagram shows the results of the transient energy field analysis network operation.
[0056] Figure 6 A flowchart for the operation of the behavioral decision-making module. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] Please see Figure 1 This invention provides an indoor smoking behavior detection system based on transient change characteristics. The system includes a transient perception module, an environmental modeling module, a behavior decision-making module, and an execution verification module. Specific implementation details are as follows:
[0059] The transient sensing module deploys a multi-source sensor array, including a gas sensor, an infrared thermal imager, and a millimeter-wave radar, to collect real-time data on smoke concentration, thermal radiation distribution, and human movement trajectories within the target monitoring space. The environmental modeling module converts the raw sensor data stream into a structured environmental state vector, constructing a dynamic behavioral response model to characterize the spatiotemporal evolution of smoking behavior. The behavior decision-making module employs a deep reinforcement learning framework, generating a mutation-sensing agent through a transient feature analysis algorithm to map environmental states to behavioral judgment actions. The execution verification module ensures the reliability of behavior judgment results through a multi-source data cross-validation mechanism and triggers environmental intervention commands such as starting / stopping ventilation equipment or sending alarm signals.
[0060] Example 1: See Figure 2 The dynamic behavioral response model is constructed using a Markov decision process framework. Its implementation utilizes a four-element structure to dynamically map environmental states to behavioral decisions. The state space consists of a multi-dimensional feature vector that fuses features from three types of sensor data in real time: transient gradient features obtained from PM2.5 concentration values collected by a gas sensor array through first-order difference calculation; waveform abrupt change coefficients extracted from local temperature field distribution captured by an infrared thermal imager at a 10Hz sampling frequency; and hand-waving frequency features calculated from human joint movement trajectories tracked by millimeter-wave radar. Historical behavioral markers record the sequence of judgment results over the past 30 seconds using binary encoding, forming a temporally correlated state encoding vector. Environmental physical parameters integrate temperature and humidity sensor data and spatial volume parameters, collectively constituting the fundamental dimensions of the state vector.
[0061] A primary warning corresponds to a situation where the transient gradient of smoke concentration exceeds the baseline value by 50% but no synchronous change in thermal radiation is detected, triggering a low-level response. A medium-level alert requires both a sudden change in PM2.5 gradient lasting more than 3 seconds and a pulse peak in the thermal radiation waveform at the 2.5μm band, initiating moderate intervention measures. Emergency response level activation conditions include high-risk composite characteristics: hand waving frequency exceeding 1.5Hz and a sudden temperature rise of 2°C in the mouth area. The state transition function is implemented through a three-layer fully connected neural network. The input layer receives a feature tensor composed of a 128-dimensional state vector and a 3-dimensional one-hot encoding of actions. The hidden layer uses 256 neurons for nonlinear transformation, and the output layer generates a 128-dimensional predicted value of the state vector for the next time step. The network is trained using a mean squared error loss function, learning environmental evolution patterns through historical state transition data.
[0062] The base reward value is a linear combination of the accuracy coefficient and the response delay coefficient: the accuracy coefficient is dynamically adjusted based on the verification results, assigning a positive value for correct judgments and deducting double the score for incorrect judgments; the response delay coefficient decays exponentially with processing time, turning negative after 2 seconds. Enhancement factors are activated when specific feature combinations occur: when the temporal correlation between the second derivative of smoke concentration and the thermal radiation pulse exceeds 0.85, the base reward value is amplified by 1.8 times; when the time difference between the hand-mouth distance feature and the abrupt change in smoke concentration is less than 500 milliseconds, an additional environmental feedback value is added.
[0063] The mutation sensing unit employs a convolutional long short-term memory hybrid network to process multi-source sensor data streams. The input layer normalizes the raw data and performs sliding window segmentation, with a window length set to 5 seconds. The spatiotemporal feature extraction layer deploys a dual-branch structure: the gas data branch uses five 1D convolutional kernels with a width of 3 for waveform feature extraction, capturing concentration mutation patterns through a max pooling layer; the thermal imaging branch uses a 3×3 2D convolutional kernel group to process the temperature matrix, and spatial gradient features are reduced in dimensionality through global average pooling. The temporal modeling layer concatenates the outputs of the two branches and inputs them into a bidirectional LSTM unit, with the hidden layer state dimension set to 64. Noise interference is filtered through a forget gate mechanism, and the output is connected to a self-attention module to strengthen the weights of key features.
[0064] The input layer receives the state vector and first performs feature scaling. The state evolution layer contains three cascaded gated recurrent units. Each unit deploys a dynamic gradient calculation module: the temporal convolution module uses causal convolution with a dilation factor of 2 to extract multi-scale temporal features; the self-attention module calculates the dependencies between feature dimensions to generate a feature weight matrix. The output layer optimizes the action probability distribution through a policy iteration algorithm and uses a baseline-based dominance function to calculate the gradient update direction. The probability distribution of the three decision levels is normalized using the Softmax function, while an ε-greedy policy is introduced to maintain a 10% exploration probability. During network training, an asynchronous update mechanism is used, performing forward inference every 200 milliseconds and updating parameters every 10 seconds.
[0065] During the initialization phase, random initial values are set for the neural network parameters, and the dynamic behavior response model updates the environmental state at a fixed frequency. The current state vector is constructed using a sliding window mechanism. Data within the window undergoes Fast Fourier Transform to extract frequency domain features, which are then concatenated with real-time sampled values to form a 128-dimensional input vector. The behavior decision unit performs Monte Carlo tree search to generate a candidate action set, uses a confidence interval upper bound algorithm to balance exploration and utilization, and finally outputs the action selection result. When calculating the state transition function, target network technology is enabled, using a parameter-frozen replica network to predict the next state to enhance stability. After the environmental feedback value is calculated, the state transition experience data is sorted by timestamp and stored in a circular memory pool with a capacity of 1000 records, employing a first-in, first-out (FIFO) management strategy.
[0066] Gas sensor data streams are filtered by a Butterworth filter to remove low-frequency noise, and transient fluctuation characteristics are obtained by calculating the moving standard deviation. Infrared thermal imaging data uses background subtraction to separate human thermal radiation, and a dynamic window of interest is set in the mouth and nose area to extract temperature change curves. Millimeter-wave radar point cloud data is segmented into human targets using the DBSCAN clustering algorithm, and hand movement trajectories are calculated based on a skeletal joint kinematic model. The three types of feature data are synchronized at the millisecond level through a timestamp alignment module and finally fused into a unified state description vector.
[0067] The system maintains a 30-entry queue of decision results, with each record containing a timestamp, action level, and confidence score. The encoder maps discrete action levels to 8-dimensional vectors through an embedding layer, which are then positionally encoded and input into the Transformer encoder layer. The output generates a fixed-dimensional historical feature vector through max pooling. This vector is concatenated with real-time sensing features to form a complete state description, providing temporal context information for behavioral decisions.
[0068] Batch data is randomly sampled from the recurrent memory pool, containing a quadruple of current state, executed action, next state, and environmental feedback value. The neural network updates its parameters by minimizing the mean squared error between the predicted and actual states, and uses the Adam optimizer to adaptively adjust the learning rate. A target network delayed update mechanism is introduced during training, synchronizing the parameters of the main network and the target network every 100 iterations to reduce correlation interference during training.
[0069] The base reward weight matrix is automatically adjusted based on historical accuracy: the accuracy coefficient weight is increased when 10 consecutive judgments are correct; a penalty mechanism is activated when a misjudgment occurs, deducting points and reducing the response latency coefficient ratio. The activation conditions for reinforcement factors are double-verified: time-series correlation detection uses Pearson correlation coefficient calculation, and feature synchronicity verification uses a dynamic time warping algorithm to align feature curves. Environmental feedback values are compressed to the [-1,1] interval using a Sigmoid function before output to avoid gradient explosion during training.
[0070] The policy optimization of the action decision unit employs a proximal policy optimization algorithm, controlling the policy update magnitude through the importance sampling ratio. The value function estimator uses a generalized advantage estimation method to balance bias and variance. An action entropy regularization term is added to the network output layer to maintain policy diversity and prevent premature convergence. Gradient pruning is implemented during training to limit parameter updates within a preset threshold, ensuring the stability of the learning process. The decision network performs policy evaluation every 10 seconds, calculating the policy gradient and updating network parameters using sampled trajectory data.
[0071] See Figure 3The results of the dynamic behavior response model are shown in the figure above. The top figure displays transient characteristic data collected by multiple sensors, including PM2.5 concentration gradient characteristics (red curve), thermal radiation waveform abrupt change coefficient (blue curve), and hand waving frequency characteristics (green curve). These characteristic data were collected in real time by gas sensors, infrared thermal imagers, and millimeter-wave radar, and processed through first-order difference calculations and waveform analysis. Obvious characteristic peaks were observed in specific time intervals (e.g., 10-15 seconds, 25-30 seconds, and 40-45 seconds), corresponding to possible smoking behaviors. The bottom figure shows the probability distribution of the three-level judgment levels output by the behavior decision module, including the probability values of three actions: primary warning (red curve), intermediate alarm (blue curve), and emergency response (green curve). The system dynamically adjusts the probability distribution of each action based on a comprehensive analysis of sensor characteristics, increasing the probability of a higher-level alarm when multiple characteristics show anomalies simultaneously. This dynamic behavior response model based on Markov decision processes can effectively characterize the spatiotemporal evolution of smoking behavior and achieve a precise mapping between environmental states and behavioral judgment actions.
[0072] Example 2: See Figure 4 The transient energy field analysis network adopts a dual-network architecture design, achieving robust estimation of environmental feedback predictions through parallel computing paths. Two independent feature extraction networks share the same topology but employ differentiated initialization parameters, each constructing a complete signal processing link. The network input layer receives preprocessed multi-source sensor data streams, including gas concentration time-series sequences, thermal imaging spatial matrices, and radar point cloud feature vectors. In the feature extraction path of the first network, the gas data branch employs a five-layer one-dimensional convolutional structure, with the kernel width decreasing from 5 to 3 layer by layer. Each layer's output is passed through the LeakyReLU activation function, and batch normalization operations are inserted between layers to suppress internal covariate shifts. The thermal imaging data processing branch deploys three-dimensional convolutional operations, using 3×3 convolutional kernels in the spatial dimension and a sliding window mechanism in the temporal dimension to capture dynamic changes in the temperature field. The radar feature branch models the data using a graph neural network, representing human joints as vertices in a graph structure, with joint motion trajectories as edge attributes, and employing a graph attention mechanism to calculate the interaction weights between nodes.
[0073] The second network introduces structural variations within the same functional layers. In the gas data branch, dilated convolutions replace standard convolutions, with the dilation factor gradually increasing from 1 to 4 to construct a multi-scale receptive field. The thermal imaging branch uses separable convolutions to reduce the number of parameters, and depthwise convolutions process spatial features before fusion of channel information via point convolutions. Radar data processing introduces dynamic edge weight calculations, adjusting the graph structure connection strength in real-time based on the joint movement speed. Intermediate features from both networks are spatially aligned in the energy field aggregation layer, establishing a feature mapping relationship through a cross-attention mechanism. The output of the first network serves as the query vector, and the output of the second network serves as the key-value pair. The aggregated feature vectors are then dimensionality-reduced by a fully connected layer, outputting independent environmental feedback prediction values. A minimum selector compares the outputs of the two networks, selecting the smaller value as the final prediction output. This mechanism effectively suppresses prediction bias caused by network overconfidence.
[0074] The input layer standardizes the state vector, and missing values are filled in using nearest neighbor interpolation. The first-level dynamic gradient unit (RTU) consists of a parallel structure of a temporal convolution module and a self-attention module. The temporal convolution module uses dilated convolution with causal constraints to ensure that temporal feature extraction does not introduce future information leakage. The output of the dilated convolution is adjusted by a gating mechanism and then weighted and fused with the calculation results of the self-attention module. The self-attention module calculates the correlation matrix between feature dimensions and assigns feature weights by scaling the dot product attention. The second-level RTU introduces a residual connection structure, concatenating the previous stage output with the original input features before inputting it into the gated recurrent unit. Internal state updates use a dynamic routing algorithm to adjust the information transmission path according to the importance of input features. The third-level unit deploys a feature pyramid structure, capturing multi-scale temporal patterns through pooling operations of different granularities. Features at each scale are adaptively weighted and then fused.
[0075] The probability distribution generation network comprises two parallel paths: a random path generates a candidate action set through Monte Carlo sampling, with each action corresponding to a probability density estimate; the deterministic path directly outputs the action preference distribution through a policy network. The outputs of the two paths are fused using a Boltzmann distribution function to form the final action probability distribution for the behavior decision. During training, a dual optimization objective is employed, simultaneously updating the sampling policy of the random path and the network parameters of the deterministic path.
[0076] The Dynamic Gradient Tracking Network maintains two parameter sets: an online network and a target network. The online network is responsible for real-time decision generation, while the target network is used to compute a stable optimization objective. After every 100 forward inference iterations, the parameters of the online network are slowly synchronized to the target network via a soft update method, with the update coefficient set to 0.01. This gradual parameter transfer method effectively smooths policy fluctuations during training. Gradient calculation employs truncated backpropagation, limiting the propagation of error signals to the most recent five time steps, thus avoiding gradient vanishing or exploding problems caused by long-range dependencies.
[0077] The transient energy field analysis network employs a time-sharing training mechanism for its two subnetworks. In odd-numbered epochs, the parameters of the first network are updated while the parameters of the second network are frozen; conversely, in even-numbered epochs, the second network is frozen. Training data is sampled from a recurrent memory pool according to priority, which is dynamically adjusted based on historical prediction errors. The loss function is designed as a combination of Huber loss and cosine similarity; the former constrains the absolute error of the predictions, while the latter maintains consistency in the feature space. After each parameter update, a weight pruning operation is performed to limit the L2 norm of the network parameters to a preset threshold.
[0078] In the initial training phase, the dimensionality of the state vector is limited, and policy learning is performed using only gas concentration and baseline temperature features. As training epochs increase, more complex inputs, such as thermal radiation waveform features and human posture features, are gradually introduced. Transitional conditions in the training phase are automatically triggered based on the network's performance on the validation set; when the reward value increases below a threshold after three consecutive evaluations, the dimensionality of the input features is expanded. This progressive training method effectively alleviates the exploration difficulties brought about by high-dimensional state spaces.
[0079] The input data first undergoes validity verification, eliminating abnormal readings that exceed the sensor's range. The intermediate feature layer performs gradient pruning to limit the parameter update magnitude during backpropagation. The output layer deploys an uncertainty estimation module, which assesses the confidence level of the prediction results by calculating the Monte Carlo Dropout sampling variance. When the difference between the outputs of the two networks exceeds a threshold, the system automatically triggers a recalculation process to prevent occasional errors from affecting decision quality.
[0080] The spatial distribution information of gas sensor nodes is encoded as a graph structure, with each node containing coordinates and sensor type attributes. Graph convolution operations iteratively update node features, and inter-layer transfer functions fuse neighbor node information with the node's own features. A projection relationship is established between the spatial grid of thermal imaging data and the gas sensor nodes, and resolution alignment is achieved through bilinear interpolation. Radar point cloud features are mapped to a unified world coordinate system through coordinate transformation, forming spatial relationships with other sensor data.
[0081] During the initial training phase, a high probability of random exploration is set, which is gradually reduced as network performance improves. The generation of exploration actions not only involves completely random sampling but also incorporates a heuristic search based on state similarity: similar historical states are retrieved from the memory pool, and their corresponding actions are subjected to noise perturbation before being used as candidate exploration actions. This guided exploration strategy accelerates the convergence process of policy optimization.
[0082] Feature extraction and policy generation are performed on separate computation threads, with data exchange facilitated by a circular buffer. Time-sensitive feature processing operations are deployed in a cache, while computational tasks on non-critical paths are scheduled using low-priority threads. The network inference process implements operator fusion optimization, merging consecutive convolutions and activation functions into a single computational kernel to reduce memory access overhead.
[0083] The minimum selector compares not only the numerical values of the predicted values but also considers the output confidence scores of the two networks. When the main network's predicted value is smaller but its confidence score is lower than that of the backup network, a verification calculation process is initiated. The verification calculation introduces a temporal smoothing constraint, requiring the current predicted value to maintain continuity with historical trends. The final output includes a confidence score for subsequent validation modules. The network maintains two sets of state buffers, storing the temporary state at the current time step and the persistent state across time steps, respectively. A gating mechanism controls the information exchange ratio between the two types of states, dynamically adjusting the memory retention strength based on input features. Normalization is implemented during the state update process to prevent numerical drift during iterative calculations. The adjustment of the action probability distribution not only considers maximizing immediate rewards but also introduces action smoothness constraints to avoid drastic jumps in judgment results between adjacent time steps. The weights of the constraint terms dynamically decay during training, initially strengthening constraints to stabilize training and gradually relaxing them later to pursue higher performance. This adaptive constraint mechanism balances the stability and flexibility of the policy.
[0084] See Figure 5 The results of the transient energy field analysis network are shown. The top figure displays the environmental feedback predictions of the dual-network architecture, including the predictions from Network 1 (solid red line), Network 2 (dashed blue line), and the final selected minimum prediction (solid green line). The two independent feature parsing networks have the same topology but use differentiated initialization parameters, each constructing a complete signal processing link. The minimum value selector compares the outputs of the two networks and selects the smaller value as the final prediction output. This mechanism effectively suppresses prediction bias caused by network overconfidence. The bottom figure shows the loss changes of the two networks during training, using a logarithmic scale to illustrate the decreasing trend of the loss value. Both the loss of Network 1 (red curve) and the loss of Network 2 (blue curve) gradually decrease with increasing training time, indicating that the network parameters are continuously optimized and the prediction performance is continuously improved. This dual-network architecture achieves robust estimation of environmental feedback predictions through parallel computing paths, improving the accuracy and reliability of the system in detecting smoking behavior. The decreasing trend of the loss function reflects the convergence of the network learning process, ensuring the stability of the system in practical applications.
[0085] Example 3: See Figure 6The behavioral decision-making module's operation is built upon a temporal difference learning framework. Its implementation involves multi-stage data processing and model interaction to achieve continuous evolution of the environmental state. During system initialization, a parameter pre-configuration strategy is employed. The neural network weight matrix W is generated through a truncated normal distribution, with its standard deviation set as the reciprocal of the square root of the input dimension. The dynamic behavioral response model updates the environmental state vector every 200 milliseconds, a time interval determined by balancing sensor sampling frequency and computational resource consumption. The current state vector is constructed using a sliding window mechanism with a window length of 10 seconds and a 50% overlap rate to ensure the continuity of temporal features. The raw data stream within the window undergoes three levels of preprocessing: gas sensor data is filtered using moving average to eliminate high-frequency noise; thermal imaging data is processed using a background subtraction algorithm to extract human thermal radiation features; and millimeter-wave radar point clouds are segmented using DBSCAN clustering to divide independent moving targets. The preprocessed data is then input into the feature extraction pipeline. Temporal features are converted to frequency domain representation using Fast Fourier Transform, and the power spectral density is calculated by dividing the frequency band energy distribution within the 0.1-5Hz range into 20 equally wide intervals.
[0086] Monte Carlo Tree Search generates a set of candidate actions within the action decision unit. Its implementation consists of four phases: The selection phase starts from the root node and traverses to leaf nodes using a tree strategy. This strategy, combined with confidence interval upper bound algorithms and prior knowledge, guides the search direction. The expansion phase, upon encountering an unexpanded node, generates an initial action probability distribution through a policy network to create new child nodes. The simulation phase executes sampled actions in a virtual environment and uses a simplified model to predict subsequent state changes. The backtracking phase propagates value estimates along the search path, updating node visit counts and action values. During the search process, an exploration coefficient c is maintained, whose value decays exponentially with search depth.
[0087]
[0088] in: Indicates the intensity of basic exploration. The attenuation coefficient is... This represents the current search depth. This mechanism ensures that the action space is broadly explored in the early stages of the search, and that high-value areas are gradually focused on later. The final action decision is generated through value-weighted voting, with the weight of each candidate action proportional to its number of visits in the search tree.
[0089] The main network receives the current state vector. With action coding As input, state transition features are calculated through three fully connected layers, with the hidden layer activation function employing the Swish nonlinear transform. The target network serves as a parameter-delayed replica of the main network, with an identical structure but its weights are synchronized every 100 iterations. State prediction values. The calculation formula is:
[0090]
[0091] in: This represents the target network's frozen parameters. Prediction errors are handled through a dual-checking mechanism: when... When the threshold is exceeded, a recalculation process is triggered and the abnormal event is recorded.
[0092] Basic feedback components It consists of a linear combination of accuracy and response time. The accuracy term is handled using a sign function, and a positive value is assigned when a correct judgment is made. Deduct when misjudgment The timeliness factor is designed as a piecewise linear function; a response delay of less than 1 second receives a full reward, while a delay exceeding 2 seconds incurs a penalty. (Feature: Synchronicity feedback component) Detecting the spatiotemporal correlation of multi-source signals, when the time difference between a sudden change in smoke concentration and a thermal radiation pulse... Activate synchronization reward when less than 500ms ,in The decay time constant. Physical constraint feedback component. Monitor environmental safety parameters and impose a fixed penalty when any parameter exceeds a safety threshold. The final environmental feedback value is the weighted sum of all components. The weighting coefficients are dynamically adjusted according to the operational phase.
[0093] The experience data is divided into two regions based on the positive and negative feedback values. The positive experience pool stores records with reward values greater than 0.5, while the regular pool stores the remaining data. During sampling, positive and negative samples are mixed in a 7:3 ratio, which is gradually adjusted as the training progresses. Each experience data point contains a quintuple. ,in Let be the environmental feedback value at time t. Metadata tags are used to record the sensor type and timestamp information of the data source. The memory pool implements an automatic cleanup mechanism, prioritizing the removal of low-priority samples when the storage capacity reaches its limit. The priority is calculated based on the access frequency and the time decay factor.
[0094] The behavior decision module maintains an independent parameter server and multiple worker threads. The worker threads periodically sample batch data from the memory pool and calculate gradients. The parameter server aggregates the gradients and then performs an update. The optimizer uses a variant of Adam with a warm-up phase, and the initial learning rate is set to... The first 1000 steps linearly increase to the peak value. Gradient updates undergo global pruning, limiting the L2 norm of parameter updates to within 0.5. Policy evaluation is performed after each parameter update, and the average reward is calculated on the validation set to guide hyperparameter tuning.
[0095] Value estimation considers not only immediate rewards but also the discounted summation of n future returns. For each time step... The formula for calculating the n-step return is: ,in As a discount factor, This represents the state-value function. In actual training, the value of n is dynamically adjusted. Initially, a smaller n value is used to accelerate convergence, and later it is gradually increased to 5 steps to improve estimation accuracy. The value target calculation employs importance sampling techniques, and bias correction is performed on data from different strategies.
[0096] After feature extraction, the raw sensor data forms a multidimensional feature tensor, and the correlation between features is calculated through a multi-head attention layer. Key-value pairs are generated from historical state sequences, and the query vector comes from the current state features. The attention weight matrix is normalized by Softmax and used to weight and aggregate contextual information. A feature pyramid structure is deployed at the encoder output, capturing multi-granularity spatiotemporal patterns through pooling operations at different scales, and finally concatenating them to form a 128-dimensional state vector.
[0097] Base exploration rate The initial value was set to 0.3, which linearly decayed to 0.05 with each training step. The generation of exploration actions not only involved uniform random sampling but also introduced state cluster-based targeted exploration: historical state vectors were clustered using the k-means algorithm, and when a new state arrived, the nearest neighbor cluster center was matched, and exploration actions were sampled from the action distribution corresponding to that cluster. This semantic similarity-based exploration strategy improved exploration efficiency.
[0098] Input data undergoes range validation, rejecting abnormal readings exceeding the sensor's range. Network intermediate layer activation value monitoring uses quantile statistics, triggering an alarm when a feature value deviates from its historical distribution by three standard deviations. The state prediction module implements consistency checks, automatically rolling back to the previous stable version when three consecutive prediction errors exceed a threshold. A heartbeat detection mechanism is maintained during system runtime; any component failing to respond within a timeout will trigger an automatic restart process.
[0099] Before deployment, the neural network model undergoes operator fusion optimization, merging consecutive matrix multiplications and activation functions into a single computational unit. Convolution operations are implemented using Winograd transformation to reduce computational complexity, and recurrent neural network layers employ a caching mechanism to avoid redundant computations. Memory management utilizes a pre-allocation strategy, reserving buffer areas based on the maximum possible input size to reduce runtime dynamic allocation overhead.
[0100] Each sensor data stream carries a hardware timestamp, achieving microsecond-level clock synchronization via the PTP protocol. The data processing pipeline maintains an adaptive delay buffer to compensate for transmission delays from different sensors. When clock drift exceeds a threshold, a timestamp resynchronization process is triggered, using a linear correction model fitted based on the least squares method.
[0101] Controlled noise perturbation is applied to the original empirical data to generate derived samples, with the noise amplitude determined based on the sensor's accuracy characteristics. Time-series data is randomly sliced and spliced to generate new sequences, maintaining the integrity of causal relationships. Interpolation enhancement is applied to the feature space to linearly generate intermediate samples between similar state vectors. These techniques effectively alleviate the problem of insufficient training data. The newly trained network parameters are first run in shadow mode, processing the same input in parallel with the existing model but without affecting the results. After performance improvement is confirmed through validation set testing, the network is gradually migrated to the new model through incremental traffic switching. A rollback mechanism is automatically triggered when performance degradation exceeds a threshold, restoring the parameters to the previous stable version. For composite events spanning multiple time steps, the total feedback value is allocated to each decision point according to its contribution. The weight allocation is calculated using an attention mechanism, considering the causal relationship between the action and subsequent state changes. Delayed feedback events establish a temporary cache queue, and backtracking updates are performed after the causal relationship is clarified. This refined credit allocation improves the accuracy of policy optimization.
[0102] Example 4: The update mechanism of the behavior decision-making module achieves continuous policy improvement through stratified sampling and incremental optimization. The recurrent memory pool organizes experience data using a three-dimensional storage structure, dividing it into high-reward, medium-reward, and low-reward zones according to the range of environmental feedback values. Within each zone, a secondary index is established based on sensor type. Stratified random sampling is implemented during batch construction: 40% of samples are drawn from the high-reward zone, 40% from the medium-reward zone, and 20% from the low-reward zone. This ratio is dynamically adjusted according to the training phase. Each sample contains a complete state transition quintuple: environmental state vector, action, immediate reward, subsequent state, and metadata tag. The metadata records the environmental context information during data collection, including temperature range, humidity range, and personnel density level.
[0103] 50% of the samples in the current batch are randomly selected for prediction by the main network, while the remaining 50% are processed by the target network. The outputs of the two networks undergo consistency verification; if the difference exceeds a preset threshold, a review process is triggered. The review calculation uses a historical moving average technique, taking the median of the most recent 10 predictions as the final value. A soft-constraint strategy is implemented for network parameter updates, limiting the magnitude of each gradient update to within a parameter space sphere, the radius of which gradually shrinks with each training iteration.
[0104] The direction of network weight updates is determined by the current gradient and the exponential moving average of historical gradients, with the momentum coefficient linearly increasing from an initial value of 0.5 to 0.9. A phased adjustment strategy is used for the learning rate: a fixed value is maintained for the first 1000 updates, and then it decays by 5% every 500 updates thereafter. Gradient calculation employs layer-by-layer normalization, scaling the gradient norm of each hidden layer to the same scale. Random noise perturbation is added to the network output layer; the noise intensity is proportional to the prediction uncertainty, preventing the training process from getting trapped in local optima.
[0105] The policy update objective function consists of three components: a term maximizing expected reward that drives the policy towards higher reward regions, an action entropy regularization term that maintains exploration capability, and a policy similarity constraint term that limits the update magnitude. The weights of these three components are dynamically adjusted based on the network's current performance. Regularization weights are increased when recent reward variance is large, and similarity constraints are gradually weakened in the later stages of policy convergence. A random deactivation technique is implemented in the network's hidden layers, randomly disabling 15% of neurons during each forward propagation to enhance the model's robustness.
[0106] The system monitors the triggering of three types of feature combinations, along with the corresponding feedback weight and duration for each combination. When a sudden change in smoke concentration occurs simultaneously with a thermal radiation waveform, the primary feature channel is activated, assigned a base weight of 1.0, and a duration window of 3 seconds. When human posture and waving gestures match continuously, the secondary feature channel is activated, with a base weight of 0.6 and a duration window extended to 5 seconds. When environmental physical parameters exceed a safety threshold, a negative feedback channel is triggered, with a fixed weight of -1.5, until the parameters return to normal. The real-time weights of each channel are dynamically adjusted based on the environmental context. In densely populated scenarios, the weight of the primary feature channel is increased by 20%, while in night mode, the weight of the secondary feature channel is decreased by 30%. See Table 1.
[0107] Table 1: Parameter configuration for calculating environmental feedback values.
[0108] Feature combination type Basic weights Duration window Intensive personnel adjustment Night mode adjustment Smoke and thermal radiation synchronized 1.0 3 seconds +20% constant Posture and gesture matching 0.6 5 seconds constant -30% Physical parameters exceed limits -1.5 continued constant constant
[0109] Each active feature channel continuously generates feedback pulses within a time window, with the pulse intensity decreasing linearly over the remaining time. The integrator calculates the weighted sum of each channel in real time, outputting the final environmental feedback value. When multiple channels are activated simultaneously, a nonlinear superposition process is implemented, establishing an inhibition relationship between positive and negative feedback channels to avoid fractional cancellation effects.
[0110] A central server maintains a copy of the global network parameters, while multiple worker nodes sample different batches of data from a memory pool and compute gradients in parallel. When aggregating gradients, the server applies time-decay weighting, assigning higher weights to recently computed gradients. Parameter update messages are broadcast via a publish-subscribe pattern, and worker nodes periodically synchronize the latest parameters. This architecture supports horizontal scaling; when computational resources are sufficient, the number of worker nodes can be increased to accelerate training.
[0111] Each piece of empirical data has its priority score updated when it is sampled and used. The score is determined by both the base access frequency and the time-series decay factor. When storage space reaches its limit, the 20% of records with the lowest scores are evicted first. A replication protection mechanism is implemented for high-value samples. Records with a reward value exceeding 1.0 are automatically retained in three copies, distributed in different memory areas to prevent concentrated loss.
[0112] The raw input data first undergoes sensor validity verification to eliminate abnormal readings caused by hardware failures. During the feature extraction stage, the statistical characteristics of intermediate results are monitored, and values exceeding the historical distribution range of 3σ trigger an alarm. Activation value pruning is implemented during network forward propagation to limit the output amplitude of any neuron to the [-5,5] interval. Gradient anomalies are dynamically detected during training, and automatic switching to a safe computation mode is initiated when numerical overflow is detected.
[0113] The initial stage only enables basic smoke detection features, and the policy network structure is simplified to a single hidden layer. The intermediate stage introduces thermal radiation waveform analysis, expanding the network to a dual-hidden-layer architecture. The advanced stage integrates all multi-source sensor features, enabling a complete network topology. Stage transitions are based on validation set performance evaluation; an upgrade is triggered when the reward growth rate exceeds a threshold after three consecutive evaluations. The dashboard displays historical curves for core parameters such as instantaneous reward values, policy entropy, and memory pool sample distribution. The monitoring agent continuously checks the system's health status, including operational metrics such as CPU / memory usage, inference latency, and data throughput. Alarms are triggered when any metric exceeds the normal range, and severe anomalies automatically initiate fault isolation procedures.
[0114] Hot updates of network parameters are implemented using a version control strategy. A parameter snapshot is created before each major update, saving the complete network state to the repository. After deployment, a new version enters an observation period, during which the old and new versions are run in parallel to compare the output differences. A rollback mechanism is automatically triggered when performance metrics drop by more than 10%, restoring to the previous stable version. The version management system records performance changes for each update, providing a basis for decision-making and subsequent optimization. Data streams from each sensor carry precise timestamps, and a unified timeline is reconstructed using interpolation algorithms. The data processing pipeline maintains an adaptive latency window, the size of which is automatically adjusted based on network jitter. The clock synchronization service periodically calibrates the time base of each node, with a maximum allowable deviation set to 50 milliseconds.
[0115] For complex behaviors with long durations, the system records the timestamp sequence of key events. The feedback integrator assigns weights based on the causal relationships between events, with the primary triggering event receiving 60% of the base score, and subsequent related events sharing the remaining 40%. This allocation method more accurately reflects the actual contribution of each action. The historical state vector is divided into 200 clusters through unsupervised learning, with each cluster maintaining a corresponding distribution of successful actions. When a new state arrives, it is matched against the most similar cluster, and the exploration direction is sampled from the action distribution of that cluster. Gaussian noise perturbation is applied to the exploration actions, and the noise variance gradually decreases as training progresses.
[0116] The system continuously monitors statistical changes in the distribution of input data and initiates incremental training when a significant drift is detected. This mode retains the core parameters unchanged, adjusting only the network weights of the last two layers, focusing on optimizing the ability to adapt to new environments using recent data. Incremental training implements resource isolation to ensure the stable operation of the main model is not affected. For high-value empirical data, neighbor-sample pairing is implemented, and new samples are generated by linear interpolation of two state vectors in the feature space. Interpolation coefficients are randomly sampled from a uniform distribution to ensure reasonable coverage of intermediate states. The action labels of the generated samples are re-predicted through the policy network, and the reward value is calculated weighted according to the interpolation ratio. Continuous linear transformations in the neural network's forward propagation path are merged into composite operators to reduce intermediate result storage overhead. The memory allocator implements pooled management, pre-allocating the maximum buffer area required for computation, eliminating the overhead of dynamic allocation at runtime. Hardware-accelerated instruction sets are enabled for intensive operations such as matrix multiplication. Minor anomalies trigger an automatic retry mechanism; after three failed retries, the system degrades to a simplified model. Serious errors initiate an isolation recovery process, pausing the current task and resetting the computing environment. Detailed context information is recorded for all abnormal events for subsequent root cause analysis and system improvement.
[0117] Example 5: The execution verification module achieves reliable execution of behavior judgment results through a distributed verification architecture. Its implementation is based on the spatiotemporal consistency verification of multi-source sensor data. The multi-source collaborative verification unit constructs a three-dimensional spatial alignment engine, mapping the concentration change curves of the gas sensor, the temperature distribution matrix of the thermal imager, and the point cloud trajectory data of the millimeter-wave radar to a unified world coordinate system. The origin of the coordinate system is set at the geometric center of the monitoring area, the X-axis points due north, the Y-axis is vertically upward, and the Z-axis completes the construction of a right-handed coordinate system. The physical location of the gas sensor nodes is calibrated using a laser rangefinder, and the location information is encoded as a four-dimensional vector containing spatial coordinates and installation height. The thermal imager data is projected onto the world coordinate system through a perspective transformation matrix, with each pixel mapped to a spatial temperature sampling point. The millimeter-wave radar point cloud data undergoes coordinate rotation transformation to eliminate equipment installation angle deviations, generating a three-dimensional trajectory point set.
[0118] Upon system startup, initial registration is performed: five reference reflectors are deployed in the monitoring area, their precise locations are acquired through radar scanning, a thermal imager captures the thermal characteristics of the reflectors, and gas sensors record the concentration values at the reference points. The registration algorithm calculates the transformation parameters between the coordinate systems of each sensor and stores them as an initial transformation matrix. During operation, online correction is implemented: using a human target as a dynamic reference point, as the target moves within the monitoring area, the coordinate differences reported by multiple sensors are compared in real time, and parameters are smoothed and corrected using a Kalman filter. Time synchronization employs a hardware clock compensation mechanism; each sensor data packet carries a precise timestamp, and the central processing unit achieves microsecond-level alignment through linear interpolation.
[0119] For each behavior determination result, the support probability of the three types of sensors is calculated independently: the gas sensor probability is based on the joint distribution function of the concentration change amplitude and duration; the thermal imaging probability is based on the waveform matching degree of the temperature change curve in the mouth and nose area; and the radar probability is based on the similarity between the hand movement trajectory and the preset smoking gesture template. The spatial consistency probability is calculated by the target location overlap, using the radar positioning point as a reference to detect the spatial offset between the thermal imaging target area and the gas concentration peak. The final verification result is calculated by the posterior probability using a Bayesian fusion formula. When the combined probability exceeds 0.85 and all three types of sensors provide positive evidence, a verification pass flag is triggered.
[0120] The initial warning response rule activates the audible and visual alarm device, controlling the LED warning lights to flash amber light at a 1Hz frequency, while simultaneously triggering a low-frequency buzzer to emit intermittent alert sounds. The intermediate alarm response rule activates the area ventilation system, calculates the optimal fan speed based on a smoke diffusion model, and sends structured alarm SMS messages to a pre-set contact list, containing the event time, location coordinates, and confidence score. The emergency response rule activates a composite response: it opens local nozzles of the fire sprinkler system, broadcasts evacuation instructions via audio, sends an emergency event code to the safety management platform, and simultaneously locks access routes to the relevant area. After all response actions are executed, the system automatically enters a status monitoring loop.
[0121] After the command is executed, a timer is started, collecting environmental parameters every 500 milliseconds. Gas sensors monitor the PM2.5 concentration decay curve, setting the first-stage target as the concentration dropping to 120% of the baseline value. Thermal imagers track the temperature field equalization process, requiring the temperature difference between the mouth / nose area and the background to be less than 0.5℃. Radar continuously monitors personnel movement, identifying evacuation behavior or abnormal lingering. The three monitoring channels operate independently; a ready signal is sent when any channel reaches the recovery standard. Once all channels are ready and maintain a stable state for 10 seconds, the system automatically de-alarms and generates an event report. If any channel fails to meet the standard within the preset time limit, the response level is escalated, and the handling procedure is retried.
[0122] The large monitoring space is divided into multiple sub-regions, each deploying a local verification node. Each node contains a dedicated processor that runs a simplified verification algorithm, and the local decision results are uploaded to a central coordinator. The coordinator implements a voting mechanism; when more than 60% of the nodes support a certain decision, the central system adopts that conclusion. Nodes communicate with each other via redundant network links, and a heartbeat detection mechanism monitors node status. If any node fails, a neighboring node automatically expands its monitoring range and takes over its responsibilities.
[0123] When sensor data is abnormal, a reliability assessment is initiated: if a certain type of sensor samples three times consecutively outside the reasonable range, its decision weight is automatically reduced; if two types of sensors fail simultaneously, an expert review mode is triggered to suspend automatic decision-making. In the event of communication interruption, a local emergency strategy is activated: the verification node executes a preset safety response based on the last valid state. In the event of hardware failure, device redundancy switching is activated, and the backup sensor group automatically takes over the data acquisition task. Detailed diagnostic information for all abnormal events is recorded for subsequent maintenance and analysis.
[0124] Verification tasks are processed in a priority-based manner: high-confidence alarms are allocated 90% of computing resources, while low-priority tasks are placed in a latency queue. Processor cores implement a task binding strategy, with time-sensitive computations consistently allocated to dedicated cores. Memory management employs a prefetch mechanism, ensuring that frequently accessed data resides in the cache. Network transmissions undergo traffic shaping, and critical control commands are marked as highest priority to ensure immediate delivery.
[0125] The data transmission channel employs end-to-end encryption, and control commands are accompanied by digital signatures to verify their origin. Device authentication utilizes a two-way certificate mechanism, preventing unauthorized devices from accessing the system. Operation logs are stored on a blockchain to prevent post-event tampering. Integrity checks are performed upon system startup, with hash values of critical code segments calculated and compared to benchmark values. The intrusion detection system monitors abnormal access patterns and automatically blocks source IPs upon identifying brute-force attacks. The report header includes basic information such as event number, occurrence time, and duration. The main body records a summary of raw data from multiple sensors, key parameters of the verification process, a list of executed commands, and environmental recovery curves. Attachments store complete data snapshots, including high-frequency sampling data from three minutes before and after the event. Report output uses a standardized template, supporting export in both JSON and XML formats to adapt to the data access requirements of different management platforms.
[0126] The hardware health monitoring system tracks sensor sensitivity degradation curves and generates replacement recommendations when a performance drop exceeds 15%. Software components are deployed with continuous integration, and security patches and algorithm updates are pushed out via a canary release strategy. Maintenance windows are intelligently scheduled, selecting low-risk periods for upgrades based on historical event frequency. A remote diagnostic interface allows technical experts to access and analyze complex faults.
[0127] Degradation operation strategy ensures basic functionality under extreme conditions. When the main computing unit fails, the backup control board automatically activates simplified decision-making logic: performing basic alarm functions based solely on gas sensor data. Local storage buffering is enabled during network interruptions, synchronizing event data once connection is restored. In the event of a power failure, the system switches to UPS power supply mode, prioritizing the operation of core sensors and alarm devices, maintaining basic monitoring capabilities for at least 30 minutes.
[0128] The console displays real-time environmental parameter curves, overlaid with behavior judgment result markers. A 3D spatial view renders the multi-source data fusion results, dynamically displaying smoke diffusion simulation and thermal distribution. An event timeline clearly presents the time markers for each stage of detection, verification, and execution. The management interface supports customized response rules, allowing adjustment of weight parameters and threshold settings for each verification channel.
[0129] Gas sensors acquire baseline values in a smoke-free environment, thermal imagers calibrate temperature readings against a constant-temperature reference source, and the radar system verifies ranging accuracy using a standard reflector. Calibration data automatically generates a deviation compensation table, dynamically correcting the original readings during subsequent data processing. The calibration process automatically starts during idle periods in the monitored area to avoid interfering with normal monitoring tasks.
[0130] After the new software version passed verification in the test environment, it was run in parallel with the production network in shadow mode: receiving the same input but with different output results. After 72 hours of consistency comparison to confirm that there were no errors, traffic was gradually switched to the new version. The rollback mechanism remains available at all times, automatically reverting to the previous stable version when performance metrics drop below a preset threshold. Detailed version change records are archived, including descriptions of feature changes and performance benchmark test results.
[0131] The temperature compensation module adjusts the sensor sensitivity parameters based on changes in ambient temperature and humidity, eliminating the impact of climate factors on detection accuracy. The personnel density estimation algorithm analyzes the number of moving targets and dynamically corrects the group behavior detection threshold. The illumination condition monitor adjusts the thermal imaging exposure parameters to ensure stable image quality under different lighting conditions. These adaptive mechanisms maintain the system's detection reliability under various environmental conditions.
[0132] The operational roles are divided into three levels: system administrator, security supervisor, and general monitor. Administrators can modify core algorithm parameters and network configurations, security supervisors can adjust response rules and contact lists, and monitors only have status viewing and event confirmation permissions. Critical operations are subject to a dual-review mechanism, and sensitive configuration changes require authorization from both levels of permissions. Complete auditing of permission change records ensures traceability to the specific operator and time point.
[0133] Real-time operational data is synchronously replicated to a backup site, automatically switching business traffic within 10 seconds in the event of a primary site failure. A full system snapshot backup is performed daily at midnight, preserving the complete system state for the past 7 days. Recovery drills are conducted quarterly to test the complete process of rebuilding the system from backup data and verify the achievement of recovery time targets. Backup data is stored offline to prevent data loss due to ransomware attacks.
[0134] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0135] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A system for detecting indoor smoking behavior based on transient mutation characteristics, characterized in that, The application relates to a dynamic behavior response model for a target monitoring space, comprising: a transient perception module for capturing a multi-source sensing data stream in the target monitoring space in real time; an environment modeling module for constructing a dynamic behavior response model based on the multi-source sensing data stream, and formalizing a smoking behavior detection task into an environment state evolution process; a behavior decision module for generating a mutation perception agent by adopting a transient feature analysis algorithm as a behavior detection engine; an execution verification module for executing an environment intervention instruction according to a behavior judgment result output by the mutation perception agent; the dynamic behavior response model comprises a four-element structure: a state space defined as a multi-dimensional feature vector of a current environment, including transient feature values of the multi-source sensing data stream, historical behavior labels and environment physical parameters; an action space defined as a set of judgment levels of suspected smoking behaviors; a state transition function for mapping a current environment state and a behavior judgment action to a next environment state; a reward function for calculating an environment feedback value according to a correctness rate and a response timeliness of the behavior judgment result; the transient feature analysis algorithm comprises: a mutation perception unit for parameterizing modeling the multi-source sensing data stream by a transient energy field analysis network, input vectors being the current environment state and the behavior judgment action, and output values being environment feedback prediction values; a behavior decision unit for inputting the current environment state into a dynamic gradient tracking network, and outputting a distribution probability of the behavior judgment action; a running process of the behavior decision module is as follows: initializing network parameters of the transient feature analysis algorithm; acquiring a current environment state in the dynamic behavior response model; generating a current behavior judgment action by the behavior decision unit; calculating a next environment state based on the state transition function; calculating an environment feedback value according to the reward function; storing environment state transition experience data into a recurrent memory pool.
2. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 1, characterized in that, the transient energy field analysis network adopts a double-network architecture, specifically: two independent feature analysis networks are constructed, and the minimum value of the two is selected as a final environment feedback prediction value; each feature analysis network comprises a transient feature extraction layer and an energy field aggregation layer, the transient feature extraction layer captures a sensing data mutation mode by a trainable nonlinear function, and the energy field aggregation layer fuses spatial correlation of multi-dimensional features.
3. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 1, characterized in that, the dynamic gradient tracking network comprises: an input layer for receiving a current environment state vector; a state evolution layer for iteratively updating a behavior decision path by a multi-level dynamic gradient unit; an output layer for generating an optimized probability distribution of the behavior judgment action.
4. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 1, characterized in that, an updating mechanism of the behavior decision module is as follows: periodically sampling environment state transition experience data from the recurrent memory pool; calculating a target environment feedback prediction value by the mutation perception unit; updating parameter weights of the transient energy field analysis network based on a difference between the target environment feedback prediction value and an actual environment feedback value; updating policy parameters of the dynamic gradient tracking network by maximizing the environment feedback prediction value.
5. The indoor smoking behavior detection system based on transient mutation features of claim 4, wherein, a calculation mode of the environment feedback value comprises: when a smoke concentration mutation feature and a thermal radiation waveform feature are detected to synchronously occur, a positive environment feedback value is given; when a human posture feature and a waving gesture feature are detected to continuously match, a secondary environment feedback value is given. When the environmental physical quantity exceeds the preset safety threshold, a negative environmental feedback value is given.
6. The indoor smoking behavior detection system based on transient mutation features of claim 1, wherein, The execution verification module comprises: Multi-source collaborative verification unit: compare the real-time sensing data stream captured by the transient perception module with the spatiotemporal consistency of the behavior judgment result; Instruction distribution unit: when multi-source verification passes, trigger the execution logic of the environmental intervention instruction.
7. A method for detecting indoor smoking behavior based on transient mutation characteristics, characterized in that, All modules and method processes of the indoor smoking behavior detection system based on transient mutation characteristics according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for testing smoke alarm
CN118629184A
Multi-equipment cooperative control method and system for coal mine work
CN120406370A