Indoor smoking behavior detection method and system based on transient mutation characteristics

The indoor smoking behavior detection system, which integrates multi-source sensors and a dynamic behavior response model, solves the problems of high false positive rate and privacy invasion in existing technologies. It achieves accurate detection and rapid intervention, adapts to complex environmental changes, and ensures indoor air quality and health.

CN120995285AActive Publication Date: 2025-11-21SHANGHAI LINGZE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511521431.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing indoor smoking detection technologies suffer from problems such as high false alarm rates, difficulty in distinguishing between e-cigarettes and traditional cigarettes, susceptibility to interference, and privacy violations, especially in complex environments where accurate detection and intervention are difficult.

Method used

An indoor smoking behavior detection system based on transient change characteristics is adopted. It integrates multiple source sensors to collect temperature, humidity, air quality and image data. Through dynamic behavior response model and transient feature analysis algorithm, a change perception intelligent agent is generated to achieve accurate detection and environmental intervention.

Benefits of technology

It improves the comprehensiveness and accuracy of smoking behavior detection, reduces the false judgment rate, can automatically adapt to changes in complex environments, achieve rapid and effective environmental intervention, and protect air quality and human health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995285A_ABST
    Figure CN120995285A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of indoor behavior detection, and discloses an indoor smoking behavior detection method and system based on transient mutation characteristics. The system comprises a transient sensing module which captures multi-source sensing data streams such as temperature, humidity, air quality and images in a target space in real time and accurately captures environment changes related to smoking; the environment modeling module constructs a dynamic behavior response model based on multi-source data, converts smoking detection into an environment state evolution process, can adapt to a complex indoor environment, and avoids false detection and missing detection; the behavior decision-making module takes a transient feature analysis algorithm as an engine, generates a sudden change sensing agent, identifies a smoking behavior by analyzing a data trend and feature combination, and reduces a misjudgment rate; the execution verification module rapidly executes intervention instructions such as ventilation equipment starting and alarm triggering according to the judgment result, seamless joint of detection and intervention is achieved, the harm of smoking is reduced, and an efficient scheme is provided for indoor smoke control.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of indoor behavior detection, in particular to an indoor smoking behavior detection method and system based on transient mutation characteristics. BACKGROUND

[0002] Smoking is extremely harmful to human health, not only seriously threatening the health of smokers themselves, but also posing a potential threat to the surrounding population. In an indoor environment, this harm is particularly prominent, as the space is relatively closed and smoke and harmful components are difficult to dissipate quickly. With the increasing attention of society to health problems, the importance of indoor smoking control is becoming increasingly prominent. Currently, indoor smoking behavior detection technology faces many challenges. Traditional smoke detectors are mainly designed for fire smoke, and have significant limitations in detecting electronic cigarettes and cigarettes. The smoke emitted by electronic cigarettes and cigarettes differs in composition and concentration from ordinary smoke, and traditional smoke detectors often cannot accurately detect the emissions of electronic cigarette devices, making it more difficult to issue effective alarms, and the detection effect is not satisfactory. Some electronic cigarette detectors on the market rely solely on a single sensor to detect the concentration of smoke in the air. This single detection method not only makes it difficult to effectively distinguish between electronic cigarettes and traditional cigarettes, but also is easily disturbed by aerosols, volatile substances from cleaning supplies, and vapors, etc., leading to false positives or masking the real smoking behavior in complex environments.

[0003] Camera monitoring technology also exposes obvious problems when applied to smoking behavior detection. When the human body is in the detection blind area, such as being blocked or facing away from the camera, the technology cannot accurately analyze the smoking behavior, resulting in a significant reduction in detection rate and accuracy. In addition, camera monitoring may involve the invasion of citizens' privacy, making it difficult to fully protect personal privacy and safety, which in some places with high requirements for privacy protection, such as private offices, changing rooms, etc., becomes an important factor hindering the widespread application of this technology. SUMMARY

[0004] The purpose of the present application is to provide an indoor smoking behavior detection method and system based on transient mutation characteristics to solve the problems raised in the background art.

[0005] To achieve the above purpose, the present application provides an indoor smoking behavior detection system based on transient mutation characteristics, which comprises:

[0006] a transient perception module: capturing real-time multi-source sensor data streams within the target monitoring space;

[0007] an environment modeling module: constructing a dynamic behavior response model based on the multi-source sensor data streams, and formalizing the smoking behavior detection task as an environmental state evolution process;

[0008] Behavior decision module: adopt transient feature analysis algorithm as behavior detection engine, generate mutation perception agent;

[0009] Execution verification module: execute environmental intervention instructions according to the behavior judgment result output by the mutation perception agent.

[0010] Preferably, the dynamic behavior response model comprises a four-element structure:

[0011] State space, defined as the multi-dimensional feature vector of the current environment, including the transient feature value of the multi-source sensor data stream, the historical behavior label and the environmental physical parameter;

[0012] Action space, defined as a set of judgment levels for suspected smoking behavior;

[0013] State transition function, mapping the current environment state and behavior judgment action to the next environment state;

[0014] Reward function, calculating the environmental feedback value according to the correctness rate and response timeliness of the behavior judgment result.

[0015] Preferably, the transient feature analysis algorithm comprises:

[0016] Mutation perception unit: parameterize modeling of the multi-source sensor data stream through transient energy field analysis network, input vector is current environment state and behavior judgment action, output value is environmental feedback prediction value;

[0017] Behavior decision unit: input current environment state through dynamic gradient tracking network, output distribution probability of the behavior judgment action.

[0018] Preferably, the transient energy field analysis network adopts a double network architecture, specifically:

[0019] Two independent feature analysis networks are constructed, and the minimum value of the two is selected as the final environmental feedback prediction value;

[0020] Each feature analysis network comprises a transient feature extraction layer and an energy field aggregation layer, the transient feature extraction layer captures the mutation pattern of sensor data through a trainable nonlinear function, and the energy field aggregation layer fuses the spatial correlation of multi-dimensional features.

[0021] Preferably, the dynamic gradient tracking network comprises:

[0022] The input layer receives the current environment state vector;

[0023] The state evolution layer iteratively updates the behavior decision path through multiple dynamic gradient units;

[0024] The output layer generates the optimized probability distribution of the behavior judgment action.

[0025] Preferably, the operation flow of the behavior decision module is as follows:

[0026] Initialize the network parameters of the transient feature analysis algorithm;

[0027] Obtain the current time environment state in the dynamic behavior response model;

[0028] Generate the current behavior decision action through the behavior decision unit;

[0029] Calculate the next time environment state based on the state transition function;

[0030] Calculate the environment feedback value based on the reward function;

[0031] Store the environment state transition experience data in the recurrent memory pool.

[0032] Preferably, the update mechanism of the behavior decision module is as follows:

[0033] Periodically sample the environment state transition experience data from the recurrent memory pool;

[0034] Calculate the target environment feedback prediction value through the mutation perception unit;

[0035] Update the parameter weight of the transient energy field analysis network based on the difference between the target environment feedback prediction value and the actual environment feedback value;

[0036] Update the policy parameter of the dynamic gradient tracking network by maximizing the environment feedback prediction value.

[0037] Preferably, the calculation method of the environment feedback value includes:

[0038] When the smoke concentration mutation feature and the thermal radiation waveform feature are detected to occur synchronously, a positive environment feedback value is given;

[0039] When the human posture feature and the waving gesture feature are continuously matched, a secondary environment feedback value is given;

[0040] When the environment physical parameters exceed the preset safety threshold, a negative environment feedback value is given.

[0041] Preferably, the execution verification module includes:

[0042] Multi-source collaborative verification unit: compare the real-time sensing data stream captured by the transient perception module with the spatiotemporal consistency of the behavior decision result;

[0043] Instruction distribution unit: when the multi-source verification passes, trigger the execution logic of the environment intervention instruction.

[0044] Preferably, the present application also includes a method for detecting indoor smoking behavior based on transient mutation characteristics, which comprises all the modules and method processes of the indoor smoking behavior detection system based on transient mutation characteristics as described above.

[0045] Compared with the prior art, the present application has the following beneficial effects:

[0046] The transient perception module of the system can capture multi-source sensing data streams in the target monitoring space in real time. Traditional detection methods often rely only on a single smoke sensor or camera, making it difficult to obtain comprehensive environmental information. However, by integrating multiple sensors, the system can simultaneously collect multi-dimensional data such as temperature, humidity, air quality, and images. In this way, it can more accurately capture subtle environmental changes related to smoking behavior. For example, smoking releases heat, causing a temporary increase in local temperature, and changes the surrounding air quality. The multi-source sensing data stream can collect these changes together, greatly improving the comprehensiveness and accuracy of capturing features related to smoking behavior.

[0047] The environmental modeling module constructs a dynamic behavior response model based on multi-source sensing data streams, formalizing the smoking behavior detection task as an environmental state evolution process. This innovative approach has greater adaptability compared to traditional fixed model detection. In complex and variable indoor environments, such as offices with frequent personnel flow and equipment operation interference, traditional models are prone to misjudgment. However, the dynamic behavior response model of the system can continuously adjust its understanding of the environmental state based on real-time data collected, automatically adapting to environmental changes. For example, during a meeting in a conference room, factors such as personnel activity and electronic device heat dissipation cause the environmental state to continuously change. The model can accurately determine whether there is smoking behavior based on the dynamic changes in multi-source data, effectively avoiding false positives and false negatives caused by environmental changes.

[0048] The behavior decision module uses a transient feature analysis algorithm as the behavior detection engine to generate a mutation perception agent. The transient feature analysis algorithm can sensitively identify transient mutation features in multi-source sensing data streams, which are key signals of smoking behavior. Unlike traditional detection algorithms based on fixed thresholds or simple pattern matching, this algorithm can deeply analyze data trends and feature combinations. For example, traditional algorithms may only rely on the presence of smoke images to make a judgment, while the transient feature analysis algorithm can analyze the instantaneous changes in human actions and object positions in the image, as well as the correlation with other sensor data, such as the sudden increase in harmful gas concentration in air quality data and the synchronization of smoking actions in the image, to more accurately determine smoking behavior and significantly reduce the false positive rate.

[0049] The execution verification module executes the environmental intervention instruction according to the behavior judgment result output by the mutation-aware intelligent agent. This module realizes the seamless connection of detection and intervention. Once the smoking behavior is detected, corresponding measures can be taken immediately, such as starting the ventilation equipment to accelerate air circulation, reducing the concentration of smoke and harmful components in the room; at the same time, an alarm is triggered to remind the relevant personnel, and the monitoring system can also be linked to record the smoking behavior and the scene, providing a basis for subsequent management. In some places with extremely high requirements for air quality, such as hospital operating rooms, precision instrument laboratories, etc., this kind of fast and effective environmental intervention measure can minimize the harm of smoking behavior to the environment and personnel, and ensure the normal operation of the place and the health of the personnel. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 Timing diagram of the indoor smoking behavior detection system based on transient mutation characteristics according to the present application;

[0051] Figure 2 Flow chart of the four-element structure of the dynamic behavior response model;

[0052] Figure 3 Graph of the running results of the dynamic behavior response model;

[0053] Figure 4 Flow chart of the double-network architecture of the transient energy field analysis network;

[0054] Figure 5 Graph of the running results of the transient energy field analysis network;

[0055] Figure 6 Flow chart of the running of the behavior decision module. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0057] Please refer to Figure 1 The present application provides an indoor smoking behavior detection system based on transient mutation characteristics, which comprises a transient perception module, an environmental modeling module, a behavior decision module and an execution verification module. The specific implementation is as follows:

[0058] The transient perception module deploys a multi-source sensor array, including gas sensors, infrared thermal imagers, and millimeter wave radars, to collect real-time data on smoke concentration, thermal radiation distribution, and human motion trajectory in the target monitoring space. The environmental modeling module converts the raw sensor data stream into a structured environmental state vector, and constructs a dynamic behavior response model to represent the spatio-temporal evolution law of smoking behavior. The behavior decision-making module uses a deep reinforcement learning framework to generate a mutation perception agent through a transient feature analysis algorithm, realizing the mapping relationship between environmental state and behavior judgment action. The execution verification module ensures the reliability of the behavior judgment result through a multi-source data cross-validation mechanism, and triggers environmental intervention instructions such as starting or stopping the ventilation equipment or sending alarm signals.

[0059] Embodiment 1: refer to Figure 2 The dynamic behavior response model is constructed as a Markov decision process framework, and its implementation process realizes the dynamic mapping of environmental state and behavior decision through a four-element structure. The state space is composed of a multi-dimensional feature vector, which real-time integrates three types of sensor data features: the PM2.5 concentration value collected by the gas sensor array is calculated by first-order difference to obtain the transient gradient feature, the infrared thermal imager captures the local temperature field distribution at a sampling frequency of 10Hz and extracts the waveform mutation coefficient, and the millimeter wave radar tracks the human joint motion trajectory and calculates the hand waving frequency feature. The historical behavior label uses binary encoding to record the judgment result sequence in the past 30 seconds, forming a state encoding vector with time sequence correlation. The environmental physical parameters integrate the temperature and humidity sensor data and the space volume parameters to form the basic dimensions of the state vector.

[0060] The primary warning corresponds to the situation that the transient gradient of smoke concentration exceeds the baseline value by 50% but no synchronous change in thermal radiation is detected, triggering a low-level response; the intermediate alarm requires simultaneous satisfaction of PM2.5 gradient mutation lasting more than 3 seconds and pulse peak value of thermal radiation waveform at 2.5μm band, starting moderate intervention measures; the emergency disposal level activates conditions include high-risk composite features: hand waving frequency above 1.5Hz and mouth area temperature sudden rise of 2℃. The state transition function is realized by a three-layer fully connected neural network, the input layer receives a feature tensor spliced by a 128-dimensional state vector and a 3-dimensional action one-hot encoding, the hidden layer uses 256 neurons for non-linear transformation, and the output layer generates a 128-dimensional prediction value of the next state vector. The network training uses mean square error loss function to learn the environmental evolution law through historical state transition data.

[0061] The base reward value is linearly combined by the accuracy coefficient and the response delay coefficient: the accuracy coefficient is dynamically adjusted according to the verification result, and a positive value is given when the judgment is correct, and double the score is deducted when the judgment is wrong; the response delay coefficient exponentially decays with the processing time, and turns to negative value after more than 2 seconds. The reinforcement factor is activated when a specific feature combination appears: when the time sequence correlation between the second derivative of smoke concentration and the heat radiation pulse exceeds 0.85, the base reward value is amplified by 1.8 times; when the hand-mouth distance feature and the smoke concentration mutation time difference are less than 500 milliseconds, the environmental feedback value is additionally increased.

[0062] The mutation perception unit uses a convolutional long short-term memory hybrid network to process multi-source sensor data streams. The input layer normalizes the original data and divides it into sliding windows, with a window length of 5 seconds. The spatiotemporal feature extraction layer deploys a double-branch structure: the gas data branch uses 5 one-dimensional convolution kernels with a width of 3 to extract waveform features, and captures the concentration mutation pattern through the max pooling layer; the thermal imaging branch uses a group of 3x3 two-dimensional convolution kernels to process the temperature matrix, and the spatial gradient features are reduced in dimension through global average pooling. The time series modeling layer inputs the concatenated outputs of the double branches into a bidirectional LSTM unit, with a hidden layer state dimension of 64, which filters noise interference through the forget gate mechanism, and the output end is connected to a self-attention module to strengthen the key feature weights.

[0063] The input layer receives the state vector and first performs feature scaling, and the state evolution layer includes 3 cascaded gated recurrent units. Each unit internally deploys a dynamic gradient calculation module: the time convolution module uses a causal convolution with a dilation factor of 2 to extract multi-scale temporal features; the self-attention module calculates the dependency between feature dimensions to generate a feature weight matrix. The output layer optimizes the action probability distribution through the policy iteration algorithm, and uses the advantage function with baseline to calculate the gradient update direction. The probability distribution of the three-level judgment level is normalized by the Softmax function, and an epsilon-greedy strategy is introduced to maintain a 10% exploration probability. During network training, an asynchronous update mechanism is used, with forward inference performed every 200 milliseconds and parameter update performed every 10 seconds.

[0064] The initialization phase sets random initial values for the neural network parameters, and the dynamic behavior response model updates the environment state at a fixed frequency. The current state vector is constructed through a sliding window mechanism, and the windowed data is extracted after fast Fourier transform to extract frequency domain features, then concatenated with real-time sampling values to form a 128-dimensional input vector. The behavior decision unit performs Monte Carlo tree search to generate a candidate action set, and uses the upper bound algorithm of the confidence interval to balance exploration and utilization, and finally outputs the action selection result. The state transition function is enabled when the target network technology is used, and a parameter-frozen copy network is used to predict the next state to enhance stability. After the environmental feedback value is calculated, the state transition experience data is sorted by timestamp and stored in the circular memory pool, with a pool capacity of 1000 records and a first-in-first-out management strategy.

[0065] After the gas sensor data stream passes through the Butterworth filter to remove low-frequency noise, the transient fluctuation characteristics are obtained by moving standard deviation calculation. The infrared thermal imaging data adopts the background difference method to separate the human body thermal radiation, and sets a dynamic region of interest window in the oral and nasal area to extract the temperature change curve. The millimeter wave radar point cloud data is segmented by the DBSCAN clustering algorithm to separate the human body target, and the hand movement trajectory is calculated based on the kinematic model of the skeletal joint. The three types of feature data are synchronized to the millisecond level through the time stamp alignment module, and finally fused into a unified state description vector.

[0066] The system maintains a decision result queue with a length of 30, and each record contains a timestamp, an action level, and a confidence. The encoder maps the discrete action level to an 8-dimensional vector through an embedding layer, and after position encoding, it inputs into the Transformer encoder layer. The output end generates a fixed-dimensional historical feature vector through max-pooling. This vector is concatenated with real-time sensing features to form a complete state description, providing temporal context information for behavior decision-making.

[0067] Batch data is randomly sampled from the recurrent memory pool, containing four tuples of current state, executed action, next state, and environment feedback value. The neural network updates the parameters by minimizing the mean square error between the predicted state and the actual state, and uses the Adam optimizer to adaptively adjust the learning rate. During training, the target network delay update mechanism is introduced, which synchronizes the parameters of the main network and the target network once every 100 iterations, reducing the correlation interference in the training process.

[0068] The basic reward weight matrix is automatically adjusted according to the historical judgment accuracy: when the judgment is correct for 10 consecutive times, the accuracy coefficient weight is increased; when a mistake is made, a penalty mechanism is started, and the score is deducted while the response delay coefficient proportion is reduced. The reinforcement factor activation condition is set with double verification: the time correlation detection uses the Pearson correlation coefficient calculation, and the feature synchronization verification uses the dynamic time warping algorithm to align the feature curve. The environment feedback value is compressed to the [-1, 1] interval by the Sigmoid function before output, to avoid gradient explosion during training.

[0069] The strategy optimization of the behavior decision unit uses the proximal policy optimization algorithm, controls the policy update amplitude through the importance sampling ratio, and uses the generalized advantage estimation method for value function estimation to balance bias and variance. The action entropy regularization term is added to the output layer of the network to maintain policy diversity and prevent premature convergence. Gradient clipping technology is implemented during training to limit the parameter update within a preset threshold, ensuring the stability of the learning process. The decision network performs policy evaluation every 10 seconds, calculates the policy gradient through sampled trajectory data, and updates the network parameters.

[0070] Referring to Figure 3, the upper figure shows the transient characteristic data collected by multi-source sensors, including PM2.5 concentration gradient characteristics (red curve), thermal radiation waveform mutation coefficient (blue curve) and hand waving frequency characteristics (green curve). These characteristic data are collected in real time by gas sensors, infrared thermal imagers and millimeter wave radars, and are processed through first-order difference calculation and waveform analysis. It can be observed that there are obvious characteristic peaks in certain time intervals (such as 10-15 seconds, 25-30 seconds and 40-45 seconds), which correspond to possible smoking behavior. The lower figure shows the probability distribution of three-level decision grades output by the behavior decision module, including primary warning (red curve), intermediate alarm (blue curve) and emergency disposal (green curve). The system dynamically adjusts the probability distribution of each action according to the comprehensive analysis of sensor characteristics, and increases the probability of high-level alarm when multiple characteristics appear abnormal at the same time. This dynamic behavior response model based on Markov decision process can effectively represent the spatio-temporal evolution law of smoking behavior and realize accurate mapping of environmental state and behavior decision action.

[0071] Example 2: see Figure 4 , the transient energy field analysis network adopts a double network architecture design, and realizes robust estimation of environmental feedback prediction value through parallel computing paths. Two independent feature analysis networks have the same topological structure but use different initialization parameters, and each constructs a complete signal processing link. The network input layer receives preprocessed multi-source sensor data stream, including gas concentration time sequence, thermal imaging space matrix and radar point cloud feature vector. In the feature extraction path of the first network, the gas data branch adopts a five-layer one-dimensional convolution structure, and the convolution kernel width decreases from 5 to 3 layer by layer. Each layer output is transmitted through LeakyReLU activation function, and batch normalization operation is inserted between layers to suppress internal covariate shift. The thermal imaging data processing branch deploys three-dimensional convolution operation, using 3x3 convolution kernel in spatial dimension and sliding window mechanism in time dimension to capture the dynamic changes of temperature field. The radar feature branch models the human body joint as a vertex in the graph structure, and the joint motion trajectory as an edge attribute, and uses graph attention mechanism to calculate the interaction weight between nodes.

[0072] The second network introduces structural variations in the same functional layer, and the gas data branch uses dilated convolution instead of standard convolution, with the dilation factor gradually increasing from 1 to 4 to build multi-scale receptive fields. The thermal imaging branch uses separable convolution to reduce the number of parameters, and the spatial features are processed by deep convolution followed by point convolution to fuse channel information. The radar data processing introduces dynamic edge weight calculation, which adjusts the connection strength of the graph structure in real time according to the motion speed of the key node. The intermediate features of the two networks are spatially aligned in the energy field aggregation layer, and the feature mapping relationship is established through the cross-attention mechanism, with the output of the first network as the query vector and the output of the second network as the key-value pair. After dimension reduction by the fully connected layer, the aggregated feature vector outputs independent environmental feedback prediction values. The minimum selector compares the output results of the two networks and selects the smaller value as the final prediction output, which effectively suppresses the prediction bias caused by excessive confidence of the network.

[0073] The input layer normalizes the state vector, and the missing values are completed by the nearest neighbor interpolation method. The first level dynamic gradient unit contains a parallel structure of time convolution module and self-attention module. The time convolution module uses dilated convolution with causal constraints to ensure that the time series feature extraction does not introduce future information leakage. The output of the dilated convolution is adjusted by the gating mechanism and then weighted and fused with the calculation result of the self-attention module. The self-attention module calculates the correlation matrix between feature dimensions and assigns feature weights through scaled dot product attention. The second level dynamic gradient unit introduces a residual connection structure, concatenates the output of the previous stage with the original input features, and inputs them into the gated recurrent unit. The internal state update uses a dynamic routing algorithm to adjust the information transmission path according to the importance of the input features. The third level unit deploys a feature pyramid structure to capture multi-scale temporal patterns through different granularity pooling operations, and the features of each scale are fused after adaptive weighting.

[0074] The probability distribution generation network contains two parallel paths: the random path generates a candidate action set through Monte Carlo sampling, and each action corresponds to a generated probability density estimate value; the deterministic path directly outputs the action preference distribution through the policy network. The outputs of the two paths are fused through the Boltzmann distribution function to form the final action probability distribution for behavior judgment. During training, a double optimization target is used to simultaneously update the sampling strategy of the random path and the network parameters of the deterministic path.

[0075] The dynamic gradient tracking network maintains two sets of parameter sets, online network and target network. The online network is responsible for real-time decision generation, and the target network is used to calculate stable optimization targets. After completing 100 forward inferences, the parameters of the online network are slowly synchronized to the target network through soft update, with an update coefficient of 0.01. This gradual parameter transfer method effectively smooths the strategy fluctuations during training. The gradient calculation uses the truncated backpropagation-through-time technique to limit the error signal to propagate within the last 5 time steps, avoiding the problem of gradient vanishing or explosion caused by long-range dependencies.

[0076] The two sub-networks of the transient energy field analysis network are trained in time, and the second network is frozen when the first network parameters are updated in odd rounds, and vice versa in even rounds. Training data is sampled from the circular memory pool according to priority, and the priority is dynamically adjusted according to the historical prediction error. The loss function is designed as a combination of Huber loss and cosine similarity, the former constrains the absolute error of the predicted value, and the latter maintains the consistency of the feature space. After each parameter update, a weight clipping operation is performed to limit the L2 norm of the network parameters within a preset threshold.

[0077] The dimension of the state vector is limited at the beginning of training, and only gas concentration and basic temperature features are used for strategy learning. As the training round increases, complex inputs such as thermal radiation waveform features and human posture features are gradually introduced. The course phase transition condition is automatically triggered according to the network performance on the validation set, and when the reward value growth of three consecutive evaluations is less than the threshold, the input feature dimension is expanded. This progressive training method effectively alleviates the exploration difficulty brought by high-dimensional state space.

[0078] The input data is first subjected to validity check, and abnormal readings exceeding the sensor range are removed. Gradient clipping operation is implemented in the intermediate feature layer to limit the parameter update amplitude in the back propagation process. The uncertainty estimation module is deployed in the output layer to evaluate the confidence of the prediction results by calculating the Monte Carlo Dropout sampling variance. When the difference between the outputs of the two networks exceeds the threshold, the system automatically triggers the recalculation process to avoid accidental errors affecting the decision quality.

[0079] The spatial distribution information of the gas sensor nodes is encoded into a graph structure, each node contains coordinate position and sensor type attributes. Graph convolution operation iteratively updates node features, and the interlayer transfer function fuses neighbor node information and its own features. The spatial grid of thermal imaging data establishes a projection relationship with the gas sensor nodes, and the resolution alignment is realized through bilinear interpolation. The radar point cloud features are mapped to a unified world coordinate system through coordinate transformation, and are spatially related to other sensor data.

[0080] A higher random exploration probability is set in the initial training stage, and the exploration rate is gradually reduced as the network performance improves. The generation of exploration actions not only includes completely random sampling, but also introduces heuristic search based on state similarity: similar historical states are retrieved from the memory pool, and the corresponding actions are added to the noise disturbance as candidate exploration actions. This guided exploration strategy speeds up the convergence process of policy optimization.

[0081] Feature extraction and policy generation belong to different computing threads, and data exchange is realized through a ring buffer. Time-sensitive feature processing operations are deployed in the cache, and non-critical path computing tasks use low-priority thread scheduling. The network inference process implements operator fusion optimization, combining consecutive convolution and activation functions into a single computing kernel to reduce memory access overhead.

[0082] The minimum selector not only compares the numerical value of the predicted value, but also considers the output confidence of the two networks. When the predicted value of the main network is small but the confidence is lower than that of the backup network, the review calculation process is started. The review calculation introduces a timing smoothing constraint, requiring the current predicted value to maintain continuity with the historical trend. The final output result is attached with a confidence score for the subsequent verification module to reference. The network maintains two sets of state buffers, respectively storing the temporary state of the current time step and the persistent state across time steps. The gating mechanism controls the exchange ratio of the two types of state information, dynamically adjusting the memory retention strength according to the input features. The state update process implements normalization processing to prevent numerical drift problems in the cyclic calculation process. The adjustment of the action probability distribution not only considers the maximization of immediate rewards, but also introduces an action smoothness constraint to avoid drastic jumps in the decision results of adjacent time steps. The weight of the constraint term is dynamically decayed during the training process, with the constraint being strengthened in the early stage to stabilize the training, and gradually relaxed in the later stage to pursue higher performance. This adaptive constraint mechanism balances the stability and flexibility of the policy.

[0083] Referring to Figure 5 , the running results of the transient energy field analysis network are shown. The upper graph shows the environment feedback predicted value of the dual network architecture, including the network 1 predicted value (red solid line), network 2 predicted value (blue dashed line), and the final selected minimum predicted value (green solid line). The two independent feature analysis networks have the same topological structure but use different initialization parameters, each building a complete signal processing link. The minimum selector compares the output results of the two networks and selects the smaller value as the final prediction output. This mechanism effectively suppresses the prediction deviation caused by excessive confidence of the network. The lower graph shows the loss change of the two networks during the training process, with the loss value shown in logarithmic scale. The network 1 loss (red curve) and network 2 loss (blue curve) both gradually decrease with the increase of training time, indicating that the network parameters are continuously optimized and the prediction performance is continuously improved. This dual network architecture realizes robust estimation of environment feedback predicted value through parallel computing paths, improving the accuracy and reliability of the system for detecting smoking behavior. The downward trend of the loss function reflects the convergence of the network learning process, ensuring the stability of the system in practical applications.

[0084] Example 3: Referring to Figure 6, the running process of the behavior decision module is built based on the time difference learning framework, and the implementation process realizes the continuous evolution of the environment state through multi-stage data processing and model interaction. In the system initialization stage, the parameter pre-configuration strategy is adopted, the neural network weight matrix W is generated by truncated normal distribution, and the standard deviation is set to the inverse of the square root of the input dimension. The dynamic behavior response model updates the environment state vector every 200 milliseconds, and this time interval is determined according to the balance between sensor sampling frequency and computing resource consumption. The construction of the state vector at the current time adopts the sliding window mechanism, the window length is set to 10 seconds and the overlap rate is configured to 50%, which ensures the continuity of the time sequence characteristics. The original data stream in the window is preprocessed in three stages: the moving average filter is performed on the gas sensor data to eliminate high-frequency noise, the background subtraction algorithm is used on the thermal imaging data to extract human thermal radiation features, and the DBSCAN clustering segmentation is performed on the millimeter wave radar point cloud to separate independent moving targets. The preprocessed data is input into the feature extraction pipeline, the time domain features are converted into frequency domain representation by fast Fourier transform, and the frequency band energy distribution is divided into 20 equal-width intervals in the range of 0.1-5Hz to calculate the power spectral density.

[0085] Monte Carlo tree search generates a candidate action set inside the behavior decision unit, which includes four stages: the selection stage starts from the root node, traverses to the leaf node through the tree strategy, and the tree strategy combines the upper bound algorithm of the confidence interval and the prior knowledge to guide the search direction. The expansion stage generates an initial action probability distribution to create a new child node when encountering an unexpanded node. The simulation stage executes the sampling action in the virtual environment and uses a simplified model to predict the subsequent state change. The backtracking stage propagates the value estimation along the search path and updates the node access count and action value. An exploration coefficient c is maintained during the search process, which decays exponentially with the search depth:

[0086]

[0087] wherein: represents the basic exploration strength, is the decay coefficient, represents the current search depth. This mechanism ensures extensive exploration of the action space in the early stage of search and gradually focuses on high-value areas in the later stage. The final output behavior decision action is generated through value weighted voting, and the weight of each candidate action is proportional to its access count in the search tree.

[0088] The main network receives the current state vector and action encoding as input, calculates the state transition features through three fully connected layers, and the hidden layer activation function adopts Swish nonlinear transformation. The target network is a parameter delayed copy of the main network, which has the same structure but the weights are synchronized every 100 updates. The state prediction value is calculated as:

[0089]

[0090] wherein: denotes the target network frozen parameters. The prediction error is processed by a double-check mechanism: when exceeds the threshold, a re-computation procedure is triggered and an abnormal event is recorded.

[0091] the basic feedback component is composed of a linear combination of the decision accuracy and response timeliness, the accuracy term is processed by a sign function, giving a positive value for correct decision and deducting for misjudgment; the timeliness term is designed as a piecewise linear function, giving full reward for response delay within 1 second and turning to penalty when exceeding 2 seconds. The feature synchronization feedback component detects the spatio-temporal correlation of multi-source signals, when the time difference between smoke concentration mutation and thermal radiation pulse is less than 500 ms, the synchronization reward is activated , where is the decay time constant. The physical constraint feedback component monitors the environmental safety parameters, imposing a fixed penalty when any parameter exceeds the safety threshold . The final environmental feedback value is the weighted sum of each component , and the weight coefficients are dynamically adjusted according to the running stage.

[0092] The experience data is divided into two regions according to the positive and negative feedback values, the positive experience pool stores records with reward values greater than 0.5, and the regular pool saves the rest of the data. When sampling, positive and negative samples are mixed in a 7:3 ratio, and this ratio is gradually adjusted as the training progresses. Each piece of experience data contains a five-tuple , where is the environmental feedback value at time t, is the metadata label, recording the data source sensor type and timestamp information. The memory pool implements an automatic cleaning mechanism, when the storage capacity reaches the upper limit, low-priority samples are preferentially eliminated, and the priority is calculated according to the access frequency and time decay factor.

[0093] The behavior decision module maintains an independent parameter server and multiple worker threads, the worker threads periodically sample batches of data from the memory pool and calculate gradients, and the parameter server aggregates the gradients and performs updates. The optimizer uses an Adam variant with a warm-up phase, the initial learning rate is set to , and increases linearly to the peak value in the first 1000 steps. Gradient updates implement global clipping, limiting the L2 norm of parameter updates to within 0.5. After each parameter update, a policy evaluation is performed to calculate the average return value on the validation set to guide hyperparameter adjustment.

[0094] The value estimation not only considers the immediate reward, but also introduces the discounted accumulation of n-step future rewards. For time step The n-step reward calculation formula is where is the discount factor, represents the state value function. In actual training, the value of n is adjusted dynamically. A small n value is used in the early stage to accelerate convergence, and then gradually increased to 5 steps to improve estimation accuracy. The importance sampling technique is used to calculate the value target, and bias correction is performed on the off-policy data.

[0095] After feature extraction, the original sensor data forms a multi-dimensional feature tensor, and the correlation between features is calculated through a multi-head attention layer. The key-value pair is generated from the historical state sequence, and the query vector comes from the current state feature. After the attention weight matrix is normalized by Softmax, it is used to weight and aggregate the context information. A feature pyramid structure is deployed at the output end of the encoder, which captures multi-granularity spatio-temporal patterns through different scale pooling operations, and finally forms a 128-dimensional state vector.

[0096] The basic exploration rate is initially set to 0.3 and linearly decays to 0.05 with the number of training steps. The generation of exploration actions not only includes uniform random sampling, but also introduces directed exploration based on state clusters: through the k-means algorithm, the historical state vectors are clustered, and when a new state arrives, it is matched with the nearest neighbor cluster center, and an exploration action is sampled from the action distribution corresponding to the cluster. This exploration strategy based on semantic similarity improves the efficiency of exploration.

[0097] The input data implements range checking and rejects abnormal readings that exceed the sensor range. The activation values in the middle layer of the network are monitored using quantile statistics, and an alarm is triggered when a feature value deviates from the historical distribution by more than 3 standard deviations. The state prediction module deploys a consistency check, which automatically rolls back to the last stable version when the prediction error exceeds the threshold for three consecutive times. A heartbeat detection mechanism is maintained during system operation, and any component that does not respond within a timeout will trigger an automatic restart process.

[0098] Before deploying the neural network model, operator fusion optimization is performed to combine continuous matrix multiplication and activation functions into a single calculation unit. The convolution operation implements Winograd transformation to reduce computational complexity, and the recurrent neural network layer uses a caching mechanism to avoid repeated calculations. Memory management uses a pre-allocation strategy, which reserves a buffer area based on the maximum possible input size, reducing the overhead of dynamic allocation at runtime.

[0099] Each sensor data stream carries a hardware timestamp, and microsecond-level clock synchronization is achieved through the PTP protocol. The data processing pipeline maintains an adaptive delay buffer to compensate for the transmission delay of different sensors. When clock drift is detected to exceed the threshold, a timestamp resynchronization process is triggered, and a linear correction model is fitted based on the least squares method.

[0100] Controllable noise is added to the raw experience data to generate derived samples, with the noise amplitude determined according to the sensor accuracy characteristics. Time series data is processed by random slicing and splicing to generate new sequences, maintaining the integrity of the causal relationship. Interpolation is implemented in the feature space to enhance the generation of intermediate samples between similar state vectors. These techniques effectively alleviate the problem of insufficient training data. The newly trained network parameters are first run in shadow mode, processing the same input as the existing model but with no effective results. After the performance improvement is confirmed by the validation set test, the new model is gradually migrated through progressive traffic switching. The rollback mechanism is automatically triggered when the performance drops below a threshold, restoring the parameters of the last stable version. For composite events that span multiple time steps, the total feedback value is allocated to each decision point based on its contribution. The allocation weights are calculated using an attention mechanism, taking into account the causal relationship between actions and subsequent state changes. Delayed feedback events establish a temporary buffer queue, and backtracking updates are performed after the causal relationship is clear. This refined credit allocation improves the accuracy of policy optimization.

[0101] Embodiment 4: The update mechanism of the behavior decision module realizes continuous improvement of the strategy through hierarchical sampling and progressive optimization. The recurrent memory pool uses a three-dimensional storage structure to organize experience data, which is divided into high reward, medium reward, and low reward regions according to the range of environmental feedback values. Within each region, a secondary index is established based on sensor type. Hierarchical random sampling is implemented when constructing the sampling batch, with 40% of samples taken from the high reward region, 40% from the medium reward region, and 20% from the low reward region. The proportion is dynamically adjusted during the training phase. Each sample contains a complete state transition quintuple: environmental state vector, executed action, immediate reward, subsequent state, and metadata label. Metadata records the environmental context information during data collection, including temperature range, humidity interval, and personnel density level.

[0102] From the current batch, 50% of the samples are randomly selected for prediction by the main network, and the remaining 50% are processed by the target network. The output results of the two networks are subjected to consistency verification, and if the difference exceeds the preset threshold, a review process is triggered. The review calculation uses a historical moving average technique, taking the median of the last 10 prediction results as the final value. Network parameter updates implement a soft constraint strategy, limiting the magnitude of each gradient update within the parameter space sphere, with the sphere radius gradually shrinking with each training round.

[0103] The direction of network weight updates is determined by the current gradient and the exponentially moving average of historical gradients, with the momentum coefficient increasing linearly from 0.5 to 0.9. The learning rate is adjusted in stages, maintaining a fixed value for the first 1000 updates, then decaying by 5% every 500 updates. Gradient calculation implements layer-by-layer normalization, scaling the gradient norm of each hidden layer to the same scale. Random noise is added to the network output layer, with the noise intensity proportional to the prediction uncertainty, preventing the training process from falling into local optima.

[0104] The policy update objective function consists of three components: the expected return maximization term drives the policy to move towards high reward regions, the action entropy regularization term maintains the exploration ability, and the policy similarity constraint term limits the update amplitude. The weight coefficients of the three terms are dynamically adjusted according to the current performance of the network. When the recent return variance is large, the regularization weight is increased, and the similarity constraint is gradually weakened in the later stage of policy convergence. The random inactivation technique is implemented in the network hidden layer, and 15% of the neurons are randomly shielded during each forward propagation to enhance the robustness of the model.

[0105] The system monitors the triggering conditions of three types of feature combinations, and each combination corresponds to a feedback weight and duration. When the smoke concentration mutation occurs synchronously with the thermal radiation waveform, the main feature channel is activated, with a base weight of 1.0 and a duration window set to 3 seconds. When the human posture and waving gesture are continuously matched, the secondary feature channel is activated, with a base weight of 0.6 and a duration window extended to 5 seconds. When the environmental physical parameters exceed the safety threshold, the negative feedback channel is triggered, with a fixed weight of -1.5 until the parameters return to normal. The real-time weight of each channel is dynamically adjusted according to the environmental context, with the main feature channel weight increased by 20% in crowded scenes and the secondary feature channel weight reduced by 30% in night mode. See Table 1.

[0106] Table 1: Environmental feedback value calculation parameter configuration.

[0107] Feature combination type Base weight Duration window People density adjustment Night mode adjustment Smoke and heat radiation synchronization 1.0 3 seconds +20% Invariable Pose and gesture matching 0.6 5 seconds Invariable -30% Physical quantity overrun -1.5 Persistent Invariable Invariable

[0108] Each active feature channel continuously generates feedback pulses within the time window, with the pulse intensity linearly decaying with the remaining time. The integrator calculates the weighted sum of each channel in real time as the final environmental feedback value output. When multiple channels are activated simultaneously, nonlinear superposition processing is implemented, and an inhibition relationship is established between positive and negative feedback channels to avoid fractional cancellation effects.

[0109] The central server maintains a copy of the global network parameters, and multiple worker nodes sample different batches of data from the memory pool and calculate gradients in parallel. The server aggregates the gradients with time decay weighting, giving higher weights to recent calculations. Parameter update messages are broadcast through the publish-subscribe mode, and each worker node synchronizes the latest parameters regularly. This architecture supports horizontal expansion, and when computing resources are sufficient, the number of worker nodes can be increased to speed up training.

[0110] The priority score of each experience data is updated when it is sampled for use, and the score is determined by the base access frequency and the time decay factor. When the storage space reaches the upper limit, the 20% of the records with the lowest score are preferentially eliminated. A copy protection mechanism is implemented for high-value samples, and records with reward values exceeding 1.0 are automatically retained in three copies distributed in different memory areas to prevent concentrated loss.

[0111] The original input data is first checked for sensor validity to eliminate abnormal readings caused by hardware failures. In the feature extraction stage, the statistical properties of the intermediate results are monitored, and values exceeding the historical distribution by 3σ trigger an alarm. During the forward propagation of the network, activation value clipping is implemented to limit the output amplitude of any neuron to the [-5, 5] interval. During training, gradient anomalies are dynamically detected, and when numerical overflow is found, the system automatically switches to a safe computing mode.

[0112] In the initial stage, only basic smoke detection features are enabled, and the policy network structure is simplified to a single hidden layer. In the intermediate stage, thermal radiation waveform analysis is introduced, and the network is expanded to a two-hidden-layer architecture. In the advanced stage, all multi-source sensor features are integrated, and the complete network topology structure is enabled. The stage transition condition is based on performance evaluation on the validation set, and when the return growth rate of three consecutive evaluations exceeds the threshold, the upgrade is triggered. The dashboard displays the historical curves of key parameters such as instantaneous return value, strategy entropy, and memory pool sample distribution. The monitoring agent continuously checks the system health status, including CPU / memory usage, inference delay, data throughput, and other operating indicators. Any indicator exceeding the normal range triggers an alarm, and serious anomalies automatically start the fault isolation program.

[0113] Hot updates of network parameters implement version control strategies. Before each major update, a parameter snapshot is created, and the complete network state is saved to the version library. After deploying the new version, an observation period is entered, during which the new and old versions are run in parallel to compare output differences. The rollback mechanism is automatically triggered when the performance indicator decreases by more than 10%, restoring to the previous stable version. The version management system records the performance changes of each update, forming a decision basis for subsequent optimization reference. Each sensor data stream carries an accurate time stamp, which is reconstructed into a unified time axis through interpolation algorithms. The data processing pipeline maintains an adaptive delay window, and the window size is automatically adjusted according to network jitter. The clock synchronization service periodically calibrates the time reference of each node, with a maximum allowed deviation of 50 milliseconds.

[0114] For complex behaviors with longer durations, the system records the timestamp sequence of key events. The feedback integrator assigns weights based on the causal relationship between events, with the main trigger event receiving 60% of the base score and subsequent associated events sharing the remaining 40%. This allocation method more accurately reflects the actual contribution of each action. The historical state vector is divided into 200 clusters through unsupervised learning, and each cluster maintains the corresponding successful action distribution. When a new state arrives, it is matched to the most similar cluster, and the exploration direction is sampled from the action distribution of that cluster. The exploration action is disturbed by Gaussian noise, and the noise variance is gradually reduced as training progresses.

[0115] The system continuously monitors the statistical changes of the input data distribution and initiates an incremental training mode when significant drift is detected. This mode keeps the core parameters unchanged and only adjusts the network weights of the last two layers, focusing on optimizing the ability to adapt to the new environment using recent data. Incremental training implements resource isolation to ensure that it does not affect the stable operation of the main model. High-value experience data is implemented with adjacent sample pairing, and new samples are generated by linear interpolation of two state vectors in the feature space. The interpolation coefficient is randomly sampled from a uniform distribution to ensure that the intermediate state is covered. The action label of the generated sample is predicted by the policy network, and the reward value is calculated by weighting the interpolation proportion. The continuous linear transformations in the forward propagation path of the neural network are combined into a composite operator to reduce the storage overhead of intermediate results. The memory allocator implements pool management to pre-allocate the maximum buffer area required for calculation, eliminating the overhead of dynamic allocation at runtime. Dense operations such as matrix multiplication enable hardware acceleration instruction sets. Slight abnormalities trigger an automatic retry mechanism, and after three failed retries, the system downgrades to a simplified model. Serious errors initiate an isolated recovery process, suspending the current task and resetting the computing environment. All abnormal events are recorded with detailed context information for subsequent root cause analysis and system improvement.

[0116] In embodiment 5, the execution verification module implements reliable execution of behavior judgment results through a distributed verification architecture. The implementation process is based on the spatio-temporal consistency verification of multi-source sensing data. The multi-source collaborative verification unit constructs a three-dimensional space alignment engine to map the concentration change curve of the gas sensor, the temperature distribution matrix of the thermal imager, and the point cloud trajectory data of the millimeter wave radar to a unified world coordinate system. The origin of the coordinate system is set at the geometric center of the monitoring area, the X-axis points to the north, the Y-axis is perpendicular to the upper, and the Z-axis completes the right-handed coordinate system construction. The physical position of the gas sensor node is calibrated by a laser range finder, and the position information is encoded as a four-dimensional vector containing spatial coordinates and installation height. The thermal imager data is projected to the world coordinate system through the perspective transformation matrix, and each pixel point is mapped to a spatial temperature sampling point. After eliminating the installation angle deviation of the device, the millimeter wave radar point cloud data is transformed through coordinate rotation to generate a three-dimensional trajectory point set.

[0117] The system performs initial registration when starting: five reference reflectors are arranged in the monitoring area, the accurate positions are obtained by radar scanning, the thermal characteristics of the reflectors are captured by the thermal imager, and the concentration values of the reference points are recorded by the gas sensor. The registration algorithm calculates the transformation parameters between the coordinate systems of each sensor and stores them as the initial transformation matrix. During operation, online correction is implemented: a human target is used as a dynamic reference point, and when the target moves in the monitoring area, the coordinate differences reported by multiple sensors are compared in real time, and the parameters are corrected by a Kalman filter. Time synchronization uses a hardware clock compensation mechanism, and each sensor data packet carries an accurate timestamp, and the central processor realizes microsecond-level alignment through linear interpolation.

[0118] For each behavior judgment result, the support probability of three types of sensors is independently calculated: the gas sensor probability is based on the joint distribution function of concentration mutation amplitude and duration; the thermal imaging probability is determined by the waveform matching degree of the temperature change curve of the mouth and nose area; and the radar probability is determined by the similarity of the hand movement trajectory and the preset smoking gesture template. The spatial consistency probability is calculated by the target position coincidence degree, taking the radar positioning point as the reference, and detecting the spatial offset of the thermal imaging target area and the gas concentration peak. The final verification result is calculated by the posterior probability of the Bayesian fusion formula. When the comprehensive probability exceeds 0.85 and all three types of sensors provide positive evidence, the verification pass flag is triggered.

[0119] The primary early warning response rule activates the sound and light alarm device, controls the LED warning light to flash amber light at a frequency of 1 Hz, and simultaneously triggers the low-frequency buzzer to emit intermittent prompt sound. The medium alarm response rule starts the regional ventilation system, calculates the optimal fan speed according to the smoke diffusion model, sends a structured alarm SMS to the preset contact list, and the SMS content includes the event time, location coordinates and confidence score. The emergency disposal rule activates the composite response: opens the local nozzles of the fire sprinkler system, plays the evacuation instruction through audio broadcast, sends the emergency event code to the safety management platform, and locks the access channels of the relevant areas. After all the response actions are executed, the system automatically enters the state monitoring cycle.

[0120] After the instruction is executed, a timer is started, and environmental parameters are collected every 500 milliseconds. The gas sensor monitors the PM2.5 concentration decay curve, and sets the concentration to 120% of the baseline value as the first stage target. The thermal imager tracks the temperature field equalization process, requiring the temperature difference between the mouth and nose area and the background to be less than 0.5°C. The radar continuously detects the movement state of personnel and identifies evacuation behavior or abnormal retention. The three monitoring channels run independently, and any channel that meets the recovery standard sends a ready signal. When all channels are ready and maintain a stable state for 10 seconds, the system automatically removes the alarm state and generates an event report. If any channel does not meet the standard within the preset time limit, the response level is upgraded and the disposal process is triggered again.

[0121] The large monitoring space is divided into multiple sub-regions, and each sub-region is deployed with a local verification node. The node contains a dedicated processor running a simplified verification algorithm, and the local decision result is uploaded to the central coordinator. The coordinator implements a voting mechanism, and when more than 60% of the nodes support a certain judgment result, the central system adopts the conclusion. The nodes communicate through redundant network links, and use a heartbeat detection mechanism to monitor node status. When any node fails, the adjacent node automatically expands the monitoring range to take over its responsibilities.

[0122] Start credibility assessment when sensor data is abnormal: when a certain type of sensor exceeds the reasonable range for three consecutive samples, automatically reduce its decision weight; when two types of sensors fail at the same time, trigger expert review mode to suspend automatic decision. Enable local emergency strategy in case of communication interruption: the verification node executes the preset safety response according to the last valid state. Activate device redundancy switching when hardware fails, and the standby sensor group automatically takes over the data acquisition task. All abnormal events record detailed diagnostic information for subsequent maintenance analysis.

[0123] Verification tasks are processed according to priority: high-confidence alerts are allocated 90% of computing resources, and low-priority tasks are placed in a delay queue. Processor cores implement task binding strategies, with time-sensitive computations fixed to dedicated cores. Memory management uses a prefetch mechanism, with frequently accessed data resident in cache areas. Network transmission implements traffic shaping, with critical control instructions marked as the highest priority to ensure immediate delivery.

[0124] Data transmission channels use end-to-end encryption, and control instructions are attached with digital signatures to verify the source. Device authentication uses a two-way certificate mechanism, and unauthorized devices cannot access the system. Operation logs are stored using a blockchain, preventing tampering after the fact. The system performs integrity checks when starting, with critical code segments calculating hash values compared to baseline values. The intrusion detection system monitors abnormal access patterns and identifies brute force attacks to automatically block the source IP. The report header contains basic information such as event number, occurrence time, and duration. The main part records multi-sensor raw data summaries, verification process key parameters, execution instruction lists, and environment recovery curves. Attachments store complete data snapshots, including three minutes of high-frequency sampling data before and after the event. Report output uses standardized templates, supporting JSON and XML formats for export, adapting to the data access needs of different management platforms.

[0125] The hardware health monitoring system tracks sensor sensitivity decay curves and generates replacement recommendations when performance degradation exceeds 15%. The software components implement continuous integration and deployment, with security patches and algorithm updates pushed through a gray release strategy. Maintenance windows are intelligently scheduled, with low-risk periods selected based on historical event frequency to perform upgrade operations. The remote diagnostic interface supports technical experts to access and analyze complex faults.

[0126] Degraded operation strategies ensure basic functionality in extreme situations. When the main computing unit fails, the backup control board automatically activates simplified decision logic: only basic alarm functions are executed based on gas sensor data. Enable local storage buffering when the network is interrupted, and synchronize event data after the connection is restored. Switch to UPS power mode in case of power failure, prioritizing the operation of core sensors and alarm devices, maintaining basic monitoring capabilities for at least 30 minutes.

[0127] The console displays real-time environmental parameter curves, superimposed with behavior judgment result markers. The three-dimensional space view renders multi-source data fusion results, dynamically displaying smoke diffusion simulation and thermal distribution. The event timeline clearly presents time markers for detection, verification, and execution stages. The management interface supports customizing response rules, allowing adjustment of weight parameters and threshold settings for each verification channel.

[0128] The gas sensor collects baseline values in a smoke-free environment, the thermal imager aligns with a constant temperature reference source to calibrate temperature readings, and the radar system checks ranging accuracy through a standard reflector plate. Calibration data automatically generates a deviation compensation table, which dynamically corrects original readings in subsequent data processing. The calibration process automatically starts during idle periods in the monitored area to avoid interfering with normal monitoring tasks.

[0129] After the new version of software passes the test environment verification, it runs in shadow mode parallel to the live network: it receives the same input but the output results are not effective. After 72 hours of consistency comparison, it is confirmed to be error-free, and gradually switches traffic to the new version. The rollback mechanism remains available at any time, and automatically reverts to the previous stable version when performance indicators drop below the preset threshold. Version change records are archived in detail, including function change explanations and performance benchmark test results.

[0130] The temperature compensation module adjusts sensor sensitivity parameters based on environmental temperature and humidity changes, eliminating the impact of climate factors on detection accuracy. The personnel density estimation algorithm analyzes the number of moving targets and dynamically adjusts the group behavior detection threshold. The light condition monitor adjusts thermal imaging exposure parameters to ensure stable imaging quality under different lighting conditions. These adaptive mechanisms maintain the detection reliability of the system under various environmental conditions.

[0131] The operating roles are divided into three levels: system administrator, security supervisor, and ordinary monitor. Administrators can modify core algorithm parameters and network configurations, security supervisors adjust response rules and contact lists, and monitors only have state viewing and event confirmation permissions. Key operations implement a two-person review mechanism, and sensitive configuration changes require joint authorization from two levels of permissions. Permission changes are fully audited and can be traced back to specific operators and time nodes.

[0132] Real-time running data is synchronously replicated to the standby site, and the main site automatically switches business traffic within 10 seconds in case of failure. Daily snapshot backups are performed at midnight, preserving the complete system state for the past 7 days. Recovery drills are performed once every quarter to test the complete system reconstruction process from backup data and verify the achievement of recovery time objectives. Backup data is stored offline to prevent data loss due to ransomware attacks.

[0133] It is to be understood that the terminology used herein such as first and second, and the like, is only used to distinguish one entity or action from another entity or action, and does not necessarily require or imply any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0134] While embodiments of the present application have been shown and described with reference to particular embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application. The scope of the application is defined by the appended claims and their equivalents.

Claims

1. A system for detecting indoor smoking behavior based on transient mutation characteristics, characterized in that, The application relates to a dynamic behavior response model for monitoring a target monitoring space, and comprises the following modules: a transient perception module for capturing a multi-source sensing data stream in the target monitoring space in real time; an environment modeling module for constructing a dynamic behavior response model based on the multi-source sensing data stream, and formalizing a smoking behavior detection task into an environment state evolution process; a behavior decision module for generating a mutation perception intelligent agent by adopting a transient feature analysis algorithm as a behavior detection engine; an execution verification module for executing an environment intervention instruction according to a behavior judgment result output by the mutation perception intelligent agent; the dynamic behavior response model comprises a four-element structure: a state space defined as a multi-dimensional feature vector of a current environment, including transient feature values of the multi-source sensing data stream, historical behavior markers and environment physical parameters; an action space defined as a set of judgment levels of suspected smoking behaviors; a state transition function for mapping a current environment state and a behavior judgment action to a next environment state; a reward function for calculating an environment feedback value according to a correctness rate and a response timeliness of the behavior judgment result.

2. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 1, characterized in that, The transient feature analysis algorithm comprises: a mutation perception unit for parameterizing modeling the multi-source sensing data stream by a transient energy field analysis network, input vectors being the current environment state and the behavior judgment action, and output values being environment feedback prediction values; a behavior decision unit for inputting the current environment state into a dynamic gradient tracking network, and outputting a distribution probability of the behavior judgment action.

3. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 2, characterized in that, The transient energy field analysis network adopts a double-network architecture, specifically: two independent feature analysis networks are constructed, and the minimum value of the two networks is selected as the final environment feedback prediction value; each feature analysis network comprises a transient feature extraction layer and an energy field aggregation layer, the transient feature extraction layer captures a sensing data mutation mode by a trainable nonlinear function, and the energy field aggregation layer fuses the spatial correlation of multi-dimensional features.

4. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 2, characterized in that, The dynamic gradient tracking network comprises: an input layer for receiving a current environment state vector; a state evolution layer for iteratively updating a behavior decision path by a multi-level dynamic gradient unit; an output layer for generating an optimized probability distribution of the behavior judgment action.

5. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 3, characterized in that, The running process of the behavior decision module is as follows: network parameters of the transient feature analysis algorithm are initialized; an environment state at a current time is obtained in the dynamic behavior response model; a current behavior judgment action is generated by the behavior decision unit; a next time environment state is calculated based on the state transition function; an environment feedback value is calculated based on the reward function; environment state transition experience data is stored in a recurrent memory pool.

6. The indoor smoking behavior detection system based on transient mutation features according to claim 5, characterized in that, The updating mechanism of the behavior decision module is as follows: environment state transition experience data is periodically sampled from the recurrent memory pool; a target environment feedback prediction value is calculated by the mutation perception unit; parameter weights of the transient energy field analysis network are updated based on a difference between the target environment feedback prediction value and an actual environment feedback value; policy parameters of the dynamic gradient tracking network are updated by maximizing the environment feedback prediction value.

7. The indoor smoking behavior detection system based on transient mutation features of claim 6, wherein, The calculation method of the environment feedback value comprises: when a smoke concentration mutation feature and a thermal radiation waveform feature are detected to synchronously occur, a positive environment feedback value is given; when a human posture feature and a waving gesture feature are detected to continuously match, a secondary environment feedback value is given. When the environmental physical quantity exceeds the preset safety threshold, a negative environmental feedback value is given.

8. The indoor smoking behavior detection system based on transient mutation characteristics according to claim 1, characterized in that, The execution verification module comprises: Multi-source collaborative verification unit: compare the real-time sensing data stream captured by the transient perception module with the spatiotemporal consistency of the behavior judgment result; Instruction distribution unit: when multi-source verification passes, trigger the execution logic of the environmental intervention instruction. 9.A method for detecting indoor smoking behavior based on transient mutation characteristics, characterized in that, All modules and method processes of the indoor smoking behavior detection system based on transient mutation characteristics according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for testing smoke alarm

    CN118629184A

  • Intelligent power distribution room safety monitoring management system

    CN119399908A

  • Multi-equipment cooperative control method and system for coal mine work

    CN120406370A

  • Household intelligent alarm system

    CN120656272A

  • Novel Aureobasidium pullulans strain and use thereof

    KR1020220036790A

Cited By

  • Preservation method for water content of preserved fruit based on multi-sensor data driving

    CN122065252A