Space-time dual attention enhanced behavior recognition system and method
Patent Information
- Application Number
- CN202610676721.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-18
AI Technical Summary
[0007]为了解决上述技术问题,本发明提供时空双注意力增强的行为识别系统及方法,以解决现有WiFi-CSI行为识别方案集中式架构覆盖范围有限、穿墙场景信号衰减与抗干扰能力差、低资源边缘设备部署适配性不足、固定阈值链路筛选环境适应性弱、多径干扰下CSI数据纯度低、相似动作特征提取不充分导致误判率高的问题
本发明中,通过设有边缘环境自适应多链路清洗与调度模块,构建了子载波空间一致性与信噪比联合评估、非对称EMA动态阈值追踪、带迟滞区间的链路信誉状态机调度的完整链路优化体系,从信号源头解决了传统固定阈值筛选适应性差、伪高质链路无法识别、多径干扰下CSI数据纯度不足的问题,大幅提升了复杂环境下的信号质量与系统抗干扰能力。
Smart Images

Figure CN122595191A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of WiFi wireless sensing and human behavior recognition technology, and more specifically, it relates to a behavior recognition system and method with spatiotemporal dual attention enhancement. Background Technology
[0002] With the rapid development of smart homes, smart elderly care, and indoor security, human behavior recognition technology has become one of the core technologies for indoor intelligent sensing. Currently, mainstream human behavior recognition technologies are mainly divided into three categories: wearable sensor recognition, visual image recognition, and WiFi channel state information (CSI) wireless sensing recognition. Among them, WiFi-CSI behavior recognition technology requires no user-worn devices, eliminates the risk of visual privacy leakage, is unaffected by environmental factors such as light and smoke, and can reuse existing WiFi infrastructure, resulting in low deployment costs. Therefore, it has become a research hotspot and application focus in the field of indoor human behavior recognition.
[0003] Currently, WiFi-CSI-based human behavior recognition technology has achieved certain research results, but in actual engineering deployment, the following technical pain points still urgently need to be addressed: Centralized data acquisition architectures have inherent drawbacks. Existing solutions mostly adopt a centralized architecture of "single router + terminal". WiFi sensing signals need to penetrate obstacles such as walls and furniture to complete detection, which easily leads to signal attenuation and severe multipath interference, causing a sharp drop in detection accuracy in scenarios with walls in between. At the same time, centralized architectures only support serial data processing in a single area, and the coverage is limited by the communication distance of a single node, making it impossible to achieve large-scale parallel data acquisition in multiple areas. The system has poor scalability and insufficient real-time transmission performance.
[0004] Low-resource AIoT devices have poor deployment adaptability. Existing high-precision identification solutions mostly concentrate data collection, feature extraction, and model inference on high-computing-power terminals, which cannot adapt to the deployment requirements of low-resource edge AIoT devices such as ESP32. On the other hand, lightweight edge solutions, in order to adapt to low-computing-power devices, greatly simplify the algorithm and collection process, resulting in a significant decrease in recognition accuracy. They cannot achieve a balance between collection deployment flexibility and recognition accuracy.
[0005] The existing solution suffers from weak anti-interference capability in complex environments and poor practical implementation results. In complex deployment environments such as homes and nursing homes, the signal-to-noise ratio of CSI data is significantly reduced due to factors such as wall obstruction, furniture reflection, and crosstalk between signals in multiple areas. Effective motion features are submerged by noise, resulting in large fluctuations in recognition accuracy and failing to meet the stable detection requirements of real-world scenarios.
[0006] The feature extraction of behavior recognition algorithms is insufficient, resulting in a high misjudgment rate for similar actions. Most current mainstream recognition algorithms employ models such as one-way LSTM, random forest, and HMM. One-way LSTM can only capture historical temporal information and cannot utilize future contextual information, leading to incomplete temporal feature extraction. Single attention mechanisms can only achieve weight allocation in the temporal dimension and cannot mine highly sensitive spatial features strongly correlated with actions. For behaviors that are highly similar to falling, sitting, bending over, or squatting, misjudgment and missed judgment are very likely. Especially in safety-sensitive scenarios such as elderly care, misjudgment and missed judgment of fall behavior can pose serious safety hazards and fail to meet the requirements of high-reliability applications. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a spatiotemporal dual attention-enhanced behavior recognition system and method to solve the problems of limited coverage of the centralized architecture in existing WiFi-CSI behavior recognition solutions, poor signal attenuation and anti-interference capabilities in wall-penetrating scenarios, insufficient adaptability to deployment of low-resource edge devices, weak environmental adaptability of fixed threshold link screening, low purity of CSI data under multipath interference, and high misjudgment rate due to insufficient extraction of similar action features.
[0008] The present invention provides a spatiotemporal dual attention-enhanced WiFi-CSI behavior recognition system, comprising an ESP32-Mesh distributed data acquisition module and a high-resource-processing terminal recognition module; The ESP32-Mesh distributed data acquisition module includes several regional acquisition units set up for multiple independent monitoring areas, a distributed Mesh transmission network built based on the ESP-Mesh protocol, and an edge environment adaptive multi-link cleaning and scheduling module deployed in each regional acquisition unit. Each of the aforementioned regional acquisition units includes one regional aggregation ANN node operating in AP mode and at least three sensing acquisition PCN nodes operating in STA mode. The ANN node and the corresponding PCN node are deployed in the same independent monitoring area to construct a closed-loop WiFi signal field within the area. Each PCN node and ANN node form three intersecting WiFi links in the vertical, horizontal, and diagonal directions. The PCN node is used to collect WiFi CSI amplitude data caused by human behavior within the area. Each ANN node connects to the distributed Mesh transmission network to aggregate and perform lightweight preprocessing of data collected by PCN nodes within its region. After adding a triple-structured unique identifier to each preprocessed data entry, the data is uploaded through the distributed Mesh transmission network. The format of the triple structured unique identifier is [Region ID][64-bit global synchronization hardware timestamp][Node ID]; The edge environment adaptive multi-link cleaning and scheduling module is deployed within the ANN node to complete joint link quality assessment, dynamic threshold adaptive update, link health state machine scheduling, and multi-link data weighted fusion, outputting clean CSI data with high signal-to-noise ratio. The high-resource processing terminal identification module includes a data preprocessing unit, a DAB-LSTM behavior recognition unit with spatiotemporal feature decoupling calibration, and a multi-region parallel inference unit. The data preprocessing unit is used to parse and standardize the received labeled CSI data and output a standardized time series tensor. The DAB-LSTM behavior recognition unit with spatiotemporal feature decoupling calibration adopts a network architecture of bidirectional LSTM layer combined with spatiotemporal dual attention enhancement layer and spatiotemporal feature decoupling calibration layer, which is used to extract, enhance and calibrate features of standardized temporal tensor and output discriminative feature vector. The multi-region parallel inference unit is used to allocate independent inference threads to each monitoring region, complete behavior classification and recognition based on discriminative feature vectors, and output behavior recognition results bound to region ID, hardware timestamp, and node ID.
[0009] This invention also provides a spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition method, implemented based on the aforementioned spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system, comprising the following steps: S1 Distributed Closed-Loop Acquisition and Adaptive Link Scheduling: Deploy one ANN node operating in AP mode and at least three PCN nodes operating in STA mode in each independent monitoring area to construct a closed-loop WiFi signal field and multiple cross WiFi links within the area. Collect WiFi CSI amplitude data caused by human behavior in the area through the PCN nodes. Based on the ESP-Mesh protocol, ANN nodes in all regions are assembled into a distributed Mesh transmission network to complete global time synchronization of all nodes. The ANN nodes complete the joint evaluation of link quality, dynamic threshold update, link state machine scheduling and multi-link data weighted fusion through the edge environment adaptive multi-link cleaning and scheduling module. Then, lightweight preprocessing is performed in sequence, and a triple structured unique identifier is added to each data before it is uploaded to the high-resource processing terminal through the distributed Mesh transmission network. S2 data standardization preprocessing: The processing terminal parses and classifies the received labeled CSI data, performs PCA dimensionality reduction and adaptive activity segmentation in sequence, and outputs a standardized temporal tensor of the input dimension of the matching recognition model. S3 Spatiotemporal Dual Attention Enhancement and Decoupling Calibration Behavior Recognition: Standardized temporal tensors are input into a pre-trained DAB-LSTM network model with spatiotemporal feature decoupling calibration. First, basic spatiotemporal features are extracted through a bidirectional LSTM layer. Then, feature weighting enhancement is performed through a spatiotemporal dual attention layer. Subsequently, intrinsic feature decoupling, adaptive calibration, and residual fusion are completed through a spatiotemporal feature decoupling calibration layer. Finally, the behavior category probability distribution is output through a fully connected layer and a Softmax activation function to complete behavior classification and recognition. S4 Multi-Region Parallel Inference Output: Allocates independent inference threads to each monitoring area, performs behavior recognition inference of CSI data in multiple areas in parallel, and outputs behavior recognition results bound to area ID, hardware timestamp, and node ID, triggering alarm linkage for high-risk behaviors such as falls.
[0010] Preferably, the edge environment adaptive multi-link cleaning and scheduling module has a built-in subcarrier spatial consistency and signal-to-noise ratio joint weighted evaluation submodule, which is used to complete the two-dimensional quantitative evaluation of link quality and candidate link screening, specifically: Let H be the CSI amplitude matrix collected by the k-th WiFi link within a preset time window T. k ∈R N×T Where N=52 is the number of subcarriers, and T is the number of sampling frames within the time window; First, calculate the mean vector of all subcarriers within the time window: ,in Let be the timing vector of the i-th subcarrier; Next, calculate the subcarrier spatial consistency index of the k-th link. The formula is: ,in, σ is the covariance calculation function, and σ is the standard deviation calculation function; Define the joint quality factor of the k-th link. Where SNRk is the real-time signal-to-noise ratio of the link; Set dual threshold filtering conditions: only when SNRk≥20dB and When the value is ≥0.7, the link enters the valid candidate pool Ω; For all links within the candidate pool Ω, CSI data fusion is performed using a normalized joint weighting formula, which is: ,in, = , where is the fusion weight of the k-th link.
[0011] Preferably, the edge environment adaptive multi-link cleaning and scheduling module also incorporates a dynamic threshold adaptive submodule based on asymmetric exponential weighted moving average, used to track environmental noise in real time and update the dynamic threshold for link selection, specifically: set up The instantaneous environmental noise energy during the silent period of frame t is used to update the mean environmental noise in real time using the asymmetric EMA algorithm. The formula is: ; in, , This is the noise rise smoothing coefficient. This is the noise reduction smoothing coefficient; Let the mean absolute noise floor of the initial calibration environment be . Construct a dynamic signal-to-noise ratio threshold With amplitude stationarity threshold The formula is: ; in, The reference signal-to-noise ratio threshold This is the signal-to-noise ratio threshold adjustment factor. As the benchmark stationarity threshold, This is the stability threshold adjustment coefficient.
[0012] Preferably, the edge environment adaptive multi-link cleaning and scheduling module also has a built-in link reputation state machine scheduling submodule with hysteresis interval, used to achieve seamless link switching and anti-ping-pong effect control, specifically: For each WiFi link k, maintain a link health reputation value ranging from [0,1]. At each time step t, the reputation value is updated according to a dynamic threshold, using the following formula: Where ΔS(t) is the reputation value step update, defined as: ; Where R is the positive reward step size, P is the negative penalty step size, and R <P; Set the activation state A for link k k ∈{0,1}, where 1 represents the active state and 0 represents the hot standby state. State switching is completed based on hysteresis dual thresholds, and the state transition equation is: ; in, This is the threshold for link disconnection. This is the link activation threshold, and ; When the primary link goes offline, the ANN node immediately selects the link with the highest health status from the backup links to complete a seamless activation switch.
[0013] Preferably, the distributed Mesh transmission network uses 1 / 6 / 11 interference-free channels in the 2.4GHz band, and a single network supports a maximum of 50 nodes to access. The network is equipped with multi-level relay nodes to relay and forward the uploaded data of ANN nodes in remote areas that exceed the transmission range of a single node, thereby extending the reliable transmission range of the system to more than 200 meters. The triple structured unique identifier is simultaneously embedded in the filename and the first 32 bytes of the file header of the CSI data file. The region ID is a 2-digit globally unique decimal number, the 64-bit timestamp is generated based on the global software clock of the Mesh network with a synchronization error of ≤10 milliseconds, and the node ID is a unique decimal number of the PCN node in the same region.
[0014] Preferably, the lightweight preprocessing of the ANN node specifically includes: using an 8th-order Butterworth low-pass filter to filter high-frequency noise and signal burst interference in the environment, using a 1-D linear interpolation algorithm to repair missing data caused by packet loss during signal transmission, and using a min-max normalization algorithm to map the CSI amplitude data of each subcarrier to the [-1,1] interval; The standardization process of the data preprocessing unit specifically includes: compressing the 52-dimensional subcarrier CSI data to 30 dimensions with a cumulative contribution rate ≥95% through PCA principal component analysis; extracting effective action segments using an adaptive threshold segmentation algorithm based on moving variance; uniformly processing the effective segments to a length of 120 time steps; and outputting a standardized time series tensor with the shape [number of samples, 120, 30].
[0015] Preferably, the network architecture of the DAB-LSTM behavior recognition unit with spatiotemporal feature decoupling calibration consists of, in sequence, a bidirectional LSTM basic feature extraction layer, a spatiotemporal dual attention feature enhancement layer, a spatiotemporal feature decoupling calibration layer, and a fully connected classification output layer; The bidirectional LSTM basic feature extraction layer consists of two bidirectional LSTM networks, each with 200 hidden nodes in each direction. It is used to simultaneously learn the forward and reverse temporal dependencies of the time-series data. The forward hidden state output by the forward LSTM and the reverse hidden state output by the reverse LSTM at the same time are concatenated as vectors to obtain a basic spatiotemporal feature matrix X∈R with dimensions of 120×400. 120×400 ; Furthermore, each unidirectional LSTM unit in the bidirectional LSTM network consists of a forget gate, an input gate, a cell state update unit, and an output gate. This gate mechanism addresses the vanishing gradient problem in long-term time-dependent data, accurately capturing the long-range dynamic features of human behavior in CSI time-series signals. The core computation process of the unidirectional LSTM unit is as follows: Forget gate calculation is used to selectively discard redundant noise information from the cell state at the previous time step. The calculation formula is as follows: ; The input gate calculation is used to filter effective behavioral information from the input features at the current time step and update the candidate cell state. The calculation formula is as follows: ; ; Cell state updates are performed by fusing the forget gate selection results with the input gate candidate features to generate the current global cell state. The calculation formula is as follows: ; Output gate calculation, based on the cell state, generates the current hidden state output. The calculation formula is as follows: ; ; In the formula, Let t be the CSI time series feature vector input at time t. This is the hidden state from the previous moment. This represents the cell state at the previous moment; For each gate weight matrix, σ represents the corresponding bias vector; σ is the Sigmoid activation function, tanh is the hyperbolic tangent activation function, and ⊙ represents the element-wise multiplication operation.
[0016] In this invention, the forward LSTM learns historical temporal dependencies by propagating forward in time, while the backward LSTM learns future contextual dependencies by propagating backward in time. The positive hidden state at the same time With reverse hidden state Perform dimensional concatenation to obtain the fused features at that moment. After being stacked with two layers of bidirectional LSTM, the final output is a basic spatiotemporal feature matrix X∈R120×400 with a dimension of 120×400, which provides a complete temporal feature base for subsequent spatiotemporal attention enhancement; The spatiotemporal dual attention feature enhancement layer includes a temporal attention branch, a spatial attention branch, and a feature fusion module; The time attention branch is used to assign normalized weights to each time step. The formula for calculating the time attention weights is: ; in, ∈R 400×1 , ∈R 1 For time-attention, trainable parameters, Let be the normalized weight at time step t, and let the sum of the weights at all time steps be 1. The spatial attention branch is used to assign normalized weights to each feature dimension. The formula for calculating the spatial attention weights is: ; in, ∈R 120×1 , ∈R 1 For trainable parameters of spatial attention, Let be the normalized weight of the d-th feature, and the sum of the weights of all feature dimensions is 1; The feature fusion module performs an outer product operation on the temporal weights and spatial weights to obtain the spatiotemporal joint attention weight matrix A∈R. 120×400 The weight matrix and the basic spatiotemporal feature matrix are then subjected to an element-wise Hadamard product to obtain the weighted and enhanced spatiotemporal joint feature matrix. ∈R 120×400 .
[0017] Preferably, the spatiotemporal feature decoupling calibration layer is a self-developed core module used to complete temporal-spatial intrinsic feature decoupling, adaptive calibration, and residual fusion of the spatiotemporal joint feature matrix, specifically as follows: Let the input spatiotemporal joint feature matrix be... ∈R T×D Where T=120 is the time step length and D=400 is the feature dimension; Step 1: Decoupling intrinsic features, extracting temporal eigenvectors and spatial eigenvectors respectively: Temporal eigenvectors ∈R T The calculation formula is: ; Spatial eigenvectors ∈R D The calculation formula is: ; Step 2: Two-branch adaptive calibration, performing gated calibration on temporal and spatial intrinsic features separately: Timing calibration gating = Timing characteristics after calibration = ; Space calibration gating = Spatial features after calibration ; in, ∈R T×T , ∈R T Trainable parameters for time-series calibration. ∈R D×D , ∈R D represents the trainable parameters for spatial calibration, and ⊙ represents the element-wise multiplication operation; Step 3: Decouple feature refusion and residual connection. Reconstruct the outer product of the calibrated bi-branch features, and then perform residual fusion with the original features. The formula is: ; ; in, For outer product operation, This is the output calibrated discriminant feature matrix.
[0018] Preferably, the training process of the DAB-LSTM network model in step S3 adopts a self-developed contrastive regularization loss function with inter-class distance constraints, and the total loss function formula is: ; in, Here, λ is the standard cross-entropy loss function, and λ is the regularization coefficient. The regularization term for the inter-class distance constraint is calculated using the following formula: ; In the formula, C is the number of behavior categories, B is the training batch size, and m is the minimum inter-class distance threshold. Let i be the class center feature vector of the i-th class. Let be the global feature vector of the b-th training sample. The true label for the b-th training sample; During training, class center feature vectors The method uses a moving average to update in real time, which forcibly increases the feature distance between similar actions and compresses the feature distribution within the class, thus solving the problem of misjudgment of similar actions from the perspective of loss function.
[0019] Compared with the prior art, the present invention has the following beneficial effects: In this invention, by incorporating an edge environment adaptive multi-link cleaning and scheduling module, a complete link optimization system is constructed, which includes joint evaluation of subcarrier spatial consistency and signal-to-noise ratio, dynamic threshold tracking of asymmetric EMA, and link reputation state machine scheduling with hysteresis interval. This solves the problems of poor adaptability of traditional fixed threshold screening, inability to identify pseudo-high-quality links, and insufficient purity of CSI data under multipath interference from the signal source, and significantly improves signal quality and system anti-interference capability in complex environments.
[0020] In this invention, an ANN+PCN node architecture with closed-loop deployment in the same area and ESP-Mesh distributed networking are used to achieve WiFi signal transmission without wall penetration within the area. This solves the problem of signal attenuation and sharp drop in recognition accuracy caused by wall penetration in traditional centralized architecture. With the help of multi-level relay transmission, the reliable coverage range of the system is extended to more than 200 meters. At the same time, it is adapted to the deployment requirements of ESP32 low-resource edge devices, taking into account both deployment flexibility and recognition accuracy.
[0021] In this invention, a triple structured unique identifier mechanism is used to bind a region ID, a globally synchronized hardware timestamp, and a node ID to each CSI data entry. This enables end-to-end traceability, automatic classification, and time-series alignment of multi-region, multi-node data, avoiding cross-contamination of multi-source data and providing reliable data support for large-scale multi-node system deployment.
[0022] In this invention, a DAB-LSTM network architecture with a spatiotemporal feature decoupling calibration layer is used to achieve decoupling of temporal and spatial intrinsic features, adaptive gating calibration and residual fusion based on bidirectional LSTM and spatiotemporal dual attention feature enhancement. This further strengthens discriminative features strongly correlated with behavior, suppresses redundant noise information, and solves the problem of insufficient feature extraction in traditional models.
[0023] In this invention, by setting a contrastive regularization loss function with inter-class distance constraints, the inter-class feature distance between falling and similar actions such as sitting, bending over, and squatting is forcibly increased during model training, and the intra-class feature distribution is compressed. This solves the industry pain point of high misjudgment rate of similar actions from the root of training. The fall recognition accuracy can reach 99%, which is fully adapted to the application needs of high safety and sensitive scenarios such as elderly care monitoring.
[0024] In this invention, by setting up a multi-region parallel inference unit and allocating an independent inference thread to each monitoring region, parallel acquisition, transmission and identification inference of multi-region data are realized, with an end-to-end transmission latency of ≤50 milliseconds, which greatly improves the real-time performance and large-scale deployment adaptability of the system.
[0025] In this invention, by employing a preprocessing architecture with end-edge collaboration, lightweight preprocessing steps such as filtering, interpolation, and normalization are moved to the edge of the ESP32, reducing transmission bandwidth usage by more than 30%. At the same time, high-precision model inference is concentrated at the terminal, achieving a balance between low-resource device deployment and high-precision recognition.
[0026] In this invention, by deploying three cross WiFi links and using a weighted fusion mechanism, blind-spot-free coverage of the monitoring area is achieved, ensuring that human movements at any location can be detected simultaneously by at least two links, further improving the system's detection stability and blind-spot-free coverage capability. Attached Figure Description
[0027] Figure 1 This is a schematic diagram comparing the recognition performance of different CNN models in this invention; Figure 2 These are schematic diagrams illustrating the classification performance in different scenarios of this invention; Figure 3 This is a schematic diagram illustrating the accuracy of distinguishing actions similar to falling in this invention; Figure 4 This is a flowchart illustrating the present invention; Figure 5 This is a diagram of the ESP-Mesh distributed network architecture of the present invention; Figure 6 This is a schematic diagram of the single-region ANN+PCN node deployment of the present invention; Figure 7 This is a schematic diagram of the multi-regional large-scale deployment of nursing homes according to the present invention; Figure 8 This is a schematic diagram of cross-regional long-distance transmission according to the present invention; Figure 9 This is a diagram of the DAB-LSTM algorithm architecture of the present invention. Detailed Implementation
[0028] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.
[0029] This invention provides a spatiotemporal dual attention-enhanced WiFi-CSI behavior recognition system and method, which can be applied to indoor human behavior recognition scenarios such as smart senior living apartments, smart homes, and indoor security. It is particularly designed for the high-reliability recognition of fall behavior in senior living scenarios. Those skilled in the art can fully reproduce the technical solution of this invention and achieve the corresponding technical effects based on the following implementation details.
[0030] The system in this embodiment is divided into two core parts: the ESP32-Mesh distributed data acquisition terminal and the high-resource processing terminal. The two achieve low-latency and high-reliability data communication through the ESP-Mesh self-organizing network, and build an end-edge collaborative architecture of "lightweight edge acquisition and preprocessing - high-precision terminal recognition and inference".
[0031] I. Hardware Platform and Basic Parameter Configuration: Edge node hardware configuration (ANN node, PCN node, relay node): Core hardware: All nodes use the Espressif ESP32-WROOM-32D development board, which integrates a dual-core 32-bit LX6 microprocessor with a fixed frequency of 240MHz, built-in 448KBROM and 520KB / sRAM, and an integrated 2.4GHz WiFi RF module, compatible with IEEE802.11b / g / n protocols, fully meeting the computing power requirements of lightweight computing and network transmission at the edge.
[0032] Fixed hardware parameter configuration (uniform across all nodes to ensure data consistency): WiFi operating frequency band: 2.4GHz, channel bandwidth 40MHz, fixed selection of three non-overlapping interference-free channels 1, 6 and 11 to avoid crosstalk of signals on the same frequency; RF transmit power: fixed at 20dBm, single-node unobstructed communication distance ≥50 meters, communication distance ≥20 meters under 120mm concrete wall obstruction; CSI data acquisition configuration: A fixed 52-channel subcarrier CSI amplitude data is acquired at a single sampling point, with a fixed sampling rate of 100 frames / second (100Hz) to ensure accurate capture of subtle dynamic changes in CSI signals caused by human movements; Power supply: 5V / 2A Type-C DC power supply is preferred, which is suitable for general indoor power supply conditions; in the absence of fixed power supply, two 18650 lithium batteries (3.7V / 2500mAh) are connected in series for power supply, with a battery life of ≥72 hours.
[0033] High-resource processing terminal hardware and software environment: Hardware configuration: Industrial control computer, equipped with Intel Core i7-12700H processor (14 cores and 20 threads), 16GB DDR4 3200MHz memory, NVIDIA RTX 3060 6GB dedicated graphics card, and 512GB NVMe SSD solid-state drive; Software environment: Ubuntu 20.04 LTS operating system, Python 3.8 programming language, PyTorch 1.12.1 deep learning framework, ESP-IDFv4.4 development environment (for edge node program development), and MQTT protocol for remote push of alarm information.
[0034] II. ESP32-Mesh Distributed Data Acquisition Terminal: Closed-loop deployment of regional data acquisition units: Deployment rules: For each independent room (independent monitoring area, 15-20㎡), a closed-loop deployment architecture of 1 ANN node + 3 PCN nodes in the same area is adopted. All nodes are deployed inside the monitoring area. The WiFi sensing signal does not need to penetrate the wall, which solves the problem of signal attenuation and sharp drop in recognition accuracy when passing through walls in the traditional centralized architecture from the root.
[0035] Specific deployment locations: The ANN node is deployed in the center of the room ceiling and operates in WiFiAP mode; the three PCN nodes are deployed in the three opposite corners of the room and operate in WiFiSTA mode. They only establish dedicated WiFi connections with the ANN nodes in the same area. The three PCN nodes and the ANN nodes form three intersecting WiFi links in a vertical, horizontal, and diagonal fashion, ensuring that human movements at any location within the monitoring area can be detected simultaneously by at least two links, with no blind spots.
[0036] Relay node deployment: For remote monitoring areas at the end of corridors and across floors, deploy 1-2 levels of relay nodes at corridor corners. The relay nodes use the same model ESP32-WROOM-32D development board, which is only responsible for data forwarding and does not participate in CSI data acquisition. Through multi-level relays, the reliable transmission range of the system is extended to 220 meters.
[0037] ESP-Mesh distributed networking: Network architecture: All ANN nodes in the monitoring area are used as child nodes of the Mesh network. One ANN node near the duty room is selected and configured as the Mesh root node. The root node and the processing terminal establish a stable connection through wired Ethernet. The remaining ANN nodes are connected to the Mesh network through 2.4GHz 1 / 6 / 11 interference-free channels. A single Mesh network supports a maximum of 50 nodes, which can be adapted to large-scale deployment of up to 50 independent monitoring areas.
[0038] Global time synchronization: When all ANN nodes are connected to the Mesh network, global hardware time synchronization is completed through the built-in time synchronization protocol of ESP-Mesh. The synchronization error is ≤1 millisecond, providing a unified time reference for data time alignment in multiple regions.
[0039] Parallel transmission mechanism: The Mesh network allocates an independent MAC layer data transmission channel and a 1MB buffer queue to each monitoring area. Data acquisition, preprocessing, and transmission in each area are executed in parallel without data crossover or confusion. The link disconnection and reconnection timeout threshold is set to 500 milliseconds to ensure the stability of the transmission link.
[0040] Edge environment adaptive multi-link cleaning and scheduling module: This module is deployed at the application layer of the ANN node, with an execution cycle of 100ms. It is the core module of the data acquisition end of this invention, and is specifically divided into 4 sub-modules executed sequentially: (1) Submodule for joint weighted evaluation of subcarrier spatial consistency and signal-to-noise ratio: Input: CSI amplitude matrix collected for each WiFi link within a 200ms time window T (20 frames). ∈R 52×20 , where k is the link number; Step 1: Calculate the global mean subcarrier vector ,in This is the 20-frame timing vector of the i-th subcarrier; Step 2: Calculate the subcarrier spatial consistency index of the k-th link. The formula is: ; in, Here, σ is the covariance calculation function, and σ is the standard deviation calculation function. In this embodiment, the dsps_math.h mathematical library built into ESP-IDF is used for implementation, reducing the computational complexity from O(N) to O(N). 2 The computational power is reduced to O(N), fully adapting to the low computing power characteristics of ESP32; Step 3: Define the joint quality factor SQFk for the k-th link = SNRk is read in real time via the ESP32's built-in WiFi radio frequency register; Step 4: Dual-threshold hard screening: The link is only included in the valid candidate pool Ω if SNRk≥20dB and Corrk≥0.7; Step 5: Normalized Weighted Fusion: For links in the candidate pool, calculate the fusion weights according to the formula and complete the CSI data fusion. ; ; The final output is fused CSI data with a high signal-to-noise ratio.
[0041] (2) Dynamic threshold adaptive submodule based on asymmetric EMA: Quiet period determination: When the moving variance of 100 consecutive frames of CSI data is <0.1, it is determined to be an environmental quiet period, and instantaneous environmental noise energy is collected. ; Asymmetric EMA parameter configuration: In this embodiment, the noise rise smoothing coefficient... =0.2, noise reduction smoothing coefficient =0.8, ensuring that the system responds quickly to sudden noise increases and smoothly to noise decreases, avoiding frequent fluctuations in the threshold; Initial noise floor calibration: After the system is powered on, collect 30 seconds of static data in an environment with no human activity, and calculate the initial mean absolute noise floor. ; Dynamic threshold construction: In this embodiment, the baseline signal-to-noise ratio threshold... =18dB, signal-to-noise ratio adjustment factor =2; Reference stationarity threshold =0.5, stability adjustment coefficient =0.3, the dynamic threshold formula is: ; ; The link filtering threshold is adaptively adjusted according to the ambient noise level, thus solving the problem of false alarms / missed alarms with fixed thresholds.
[0042] (3) Link reputation state machine scheduling submodule with hysteresis interval: Link health and reputation value update: Maintain a health and reputation value for each link k, with a value ranging from [0,1]. Each time step updates according to the following rules: The reputation value step update ΔS(t) is defined as follows: ; Hysteresis state switching logic: Set link offline threshold =0.3, Link Activation Threshold =0.7, forming a hysteresis interval of 0.4, completely avoiding the ping-pong switching effect at the threshold edge, the state transition equation is: ; Seamless switching implementation: When the primary link goes offline, the ANN node immediately interrupts the current data frame transmission and selects the link with the highest health status from the backup links to complete the activation switch. The switching delay is ≤10ms and there is no data packet loss.
[0043] Triple structured unique identification mechanism: The identifier has a fixed format: [Region ID][64-bit global synchronization hardware timestamp][Node ID]. For example:
[01] [1714523658000][1], which represents monitoring region 1, timestamp 1714523658000, and PCN node 1. Identifier binding rule: When packaging each preprocessed CSI data, the identifier is simultaneously bound to the first 32 bytes of the data file name and the file header. This double binding prevents the identifier from being lost during transmission. The flag is clearly defined: Region ID: A 2-digit decimal number, globally unique, with 01-50 corresponding to different monitoring regions, enabling traceability of data source regions; 64-bit hardware timestamp: generated based on the GPTimer built into the ANN node, with a precision of 1 millisecond, and globally synchronized through the Mesh network with a synchronization error of ≤1 millisecond, ensuring the consistency of data timing across multiple regions; Node ID: A 1-digit decimal number, which is a unique number for PCN nodes in the same area. 1-3 correspond to 3 PCN nodes and are used for link fault diagnosis and multi-link data tracing.
[0044] Lightweight preprocessing at the edge: The merged CSI data undergoes three lightweight preprocessing steps on the ANN node, all implemented using the ESP-IDF built-in math library, making it suitable for low-computing-power hardware environments. Low-pass filtering noise reduction: An 8th-order Butterworth low-pass filter is used with a fixed cutoff frequency of 0.5Hz and zero-phase filtering to filter high-frequency electromagnetic noise and signal burst interference in the environment, while retaining the low-frequency fluctuation characteristics of CSI within 0.5Hz caused by human behavior. Missing data repair: A 1-D linear interpolation algorithm is used to fill missing segments with ≤5 consecutive packet loss by linear interpolation using the CSI amplitude of the adjacent valid frames; invalid data segments with >5 consecutive packet loss are directly removed to avoid invalid data interfering with subsequent identification. min-max normalization: The min-max normalization algorithm is used to uniformly map the CSI amplitude data of each subcarrier to the [-1,1] interval, eliminating the differences in hardware gain and dimensions between different subcarriers. The calculation formula is as follows: ,in, This represents the original CSI amplitude value of the i-th sampling point and the j-th subcarrier. , These are the minimum and maximum CSI amplitude values of the j-th subcarrier within the current data segment, respectively.
[0045] Data Packaging and Uploading: After preprocessing, the ANN node packages the data into units of 1000 frames, adds a triple structured unique identifier, and uploads it to the processing terminal through the Mesh network. Compared with transmitting the original data, this process reduces the transmission bandwidth usage by 35%, and the end-to-end transmission latency is ≤40ms.
[0046] High-resource processing terminal: Terminal data standardization preprocessing: After receiving the data uploaded by the acquisition terminal, the processing terminal first parses the triple unique identifier, classifies the data by region ID, aligns the time sequence by timestamp, and then performs standardized preprocessing: PCA Principal Component Analysis Dimensionality Reduction: For 52-dimensional subcarrier CSI data, the covariance matrix of the normalized CSI data is calculated, the covariance matrix is decomposed into eigenvalues, and the eigenvalues are sorted from largest to smallest. The top 30 principal components with a cumulative contribution rate of ≥95% are selected, and the original 52-dimensional CSI data is compressed to 30-dimensional, reducing the amount of computation while retaining the core features related to human movement. Adaptive Activity Segment Segmentation: An adaptive threshold segmentation algorithm based on moving variance is adopted. The sliding window length is set to 20 frames and the step size is 1 frame. The moving variance of CSI data within each sliding window is calculated sequentially, and the moving variance sequence is processed by first-order difference and Gaussian smoothing. When the moving variance exceeds the adaptive threshold (5 times the mean of the moving variance of 100 consecutive frames of static data), it is determined as the start point of the action. When the moving variance falls back to below the threshold and continues for more than 50 frames, it is determined as the end point of the action. Valid action segments are accurately extracted, and static redundant data is removed. Uniform temporal length: The extracted valid action segments are uniformly trimmed or padded with zeros to a length of 120 time steps; for segments with a duration of more than 120 frames, the middle 120 frames of the core action change are extracted; for segments with a duration of less than 120 frames, zeros are padded at the end of the segment to a length of 120 frames; the final output is a standardized temporal tensor with the shape of [number of samples, 120, 30], which fully matches the input dimension requirements of the subsequent recognition model.
[0047] DAB-LSTM Behavior Recognition Unit with Spatiotemporal Feature Decoupling Calibration: The core recognition network in this embodiment is named DAB-LSTM-DC. Its overall architecture is as follows: Bidirectional LSTM basic feature extraction layer → Spatiotemporal dual attention feature enhancement layer → Spatiotemporal feature decoupling calibration layer → Global average pooling layer → Fully connected classification output layer. Specific implementation details are as follows: (1) Input / output dimension configuration: Input dimensions: [batch_size, 120, 30], batch_size is fixed at 32 during training and batch_size=1 during inference; 120 is the time step length and 30 is the feature dimension after PCA dimensionality reduction; Output dimension: [batch_size,C], where C=10 is the number of behavior categories, corresponding to: standing, walking, bending over, squatting, drinking water, cleaning, running, sitting down, lying down, and falling down.
[0048] (2) Bidirectional LSTM basic feature extraction layer: Network structure: a two-layer stacked bidirectional LSTM network, with a fixed number of 200 hidden nodes in each direction, and a Dropout rate of 0.2 to prevent overfitting. Temporal feature learning: The forward LSTM propagates along the forward time sequence (t=1→t=120) to learn the temporal dependencies from the historical moment to the current moment; the backward LSTM propagates along the reverse time sequence (t=120→t=1) to learn the temporal dependencies from the future moment to the current moment, thus fully capturing the temporal context information of human behavior. Bidirectional Feature Fusion: The forward and backward hidden states at the same time step are concatenated as vectors to obtain the output features of the bidirectional LSTM at that time step. The concatenation formula is as follows: Where ⊕ represents the vector concatenation operation, In a positive hidden state, This is the reverse hidden state; after extraction by two layers of bidirectional LSTM, the final output is a basic spatiotemporal feature matrix X∈R with a dimension of 120×400. 120×400 .
[0049] (3) Spatiotemporal dual attention feature enhancement layer: Temporal attention branch: Normalized weights are assigned to each time step to amplify features from key time steps that contribute more to action recognition and suppress noise interference from irrelevant time steps. The weight calculation formula is as follows: ,in, ∈R 400×1 , ∈R 1 For time-attention, trainable parameters, Let be the normalized weight at time step t, and let the sum of the weights at all time steps be 1. Spatial attention branch: Assigns normalized weights to each feature dimension, amplifies highly sensitive feature dimensions that are strongly correlated with human behavior, and suppresses interference from redundant features. The weight calculation formula is as follows: ,in, ∈R 120×1 , ∈R 1 For trainable parameters of spatial attention, Let be the normalized weight of the d-th feature, and the sum of the weights of all feature dimensions is 1; Spatiotemporal joint feature weighting: The temporal weights and spatial weights are multiplied by an outer product to obtain a 120×400 spatiotemporal joint attention weight matrix A. Then, an element-wise Hadamard product is performed with the basic spatiotemporal feature matrix to obtain the weighted and enhanced spatiotemporal joint feature matrix. ∈R 120×400 The formula is: ; where ⊙ represents element-wise multiplication.
[0050] (4) Spatiotemporal feature decoupling calibration layer: This layer decouples and purifies redundant information in the spatiotemporal coupling features, with the weighted and enhanced feature matrix as input. ∈R 120×400 The process involves three steps: Step 1: Decoupling intrinsic features, extracting global intrinsic features in both temporal and spatial dimensions: Temporal intrinsic vector. ∈R 120 The calculation formula is: ; Spatial eigenvectors ∈R 400 The calculation formula is: .
[0051] Step 2: Two-branch adaptive gating calibration, performing adaptive weighted calibration on temporal and spatial intrinsic features respectively: Timing calibration gating = Timing characteristics after calibration = ; Space calibration gating = Spatial features after calibration = ; in, ∈R 120×120 , ∈R 120 Trainable parameters for time-series calibration. ∈R 400×400 , ∈R 400 Trainable parameters for spatial calibration; Step 3: Outer product reconstruction and residual fusion. The calibrated bi-branch features are reconstructed into a feature matrix, and then residual fusion is performed with the original input to avoid feature degradation. ,in, For outer product operation, The output is the calibrated discriminant feature matrix, which still has a dimension of 120×400.
[0052] (5) Fully connected classification output layer: Global average pooling: for the calibrated feature matrix Global average pooling is performed along the time dimension to compress the 120×400 feature matrix into a 400-dimensional global feature vector. The calculation formula is: ; Classification Output: The 400-dimensional global feature vector is input into the fully connected layer, mapped to a 10-dimensional category output. Then, the predicted probability of each behavior category is output through the Softmax activation function, as shown in the formula: ,in, ∈R 400×10 , ∈R 10 y represents the trainable parameters of the fully connected layer, y represents the predicted probability of the corresponding category, and finally the category with the highest predicted probability is selected as the behavior recognition result of the model.
[0053] Model training and loss function: (1) Dataset construction: Self-built dataset: In three typical scenarios, namely a laboratory, a bedroom in a senior living apartment, and a living room, CSI data of 10 types of human behavior were collected from 20 volunteers (10 men and 10 women, aged 22-65 years). 50 samples were collected from each volunteer for each type of behavior, for a total of 10,000 samples. The dataset covers people of different heights, weights, and ages, as well as different environmental interferences, to ensure the diversity of the dataset. Public datasets: 10,000 additional samples were added to the HAR-2 and WiAR public WiFi-CSI behavior recognition datasets to expand the dataset size and improve the model's generalization ability; Dataset partitioning: The total dataset is divided into training set, validation set and test set in a ratio of 7:2:1. The training set is used for model parameter updates, the validation set is used for hyperparameter tuning and early stopping judgment, and the test set is used for model performance evaluation.
[0054] (2) Configuration of training hyperparameters and loss function: Optimizer: Adam optimizer, initial learning rate 0.001, weight decay 1e-4; Learning rate scheduling: The learning rate decays to 0.5 of its original value every 10 epochs; Loss function: A self-developed contrastive regularization loss function with inter-class distance constraints is adopted. The total loss formula is as follows: ,in, The standard cross-entropy classification loss is used, with a regularization coefficient λ = 0.1. The regularization term for the inter-class distance constraint is calculated using the following formula: In the formula, C=10 is the number of behavior categories, B is the training batch size, and the minimum inter-class distance threshold m=10. Let i be the class center feature vector of the i-th class. Let be the global feature vector of the b-th training sample. The true label for the b-th training sample; Training configuration: 100 total epochs, early stopping strategy is to stop training if the validation set loss does not decrease for 5 consecutive epochs to prevent model overfitting; Class center update: After each epoch, the class center feature vectors of each class are updated using a moving average method. The update formula is as follows: , where zb is the global feature vector of all samples of class i in the current epoch. By updating the class center, the inter-class feature distance of similar actions is forcibly increased, and the intra-class feature distribution is compressed, thus solving the problem of misjudgment of similar actions from the root of training.
[0055] Multi-region parallel inference unit: Multi-process architecture: It adopts Python's multiprocessing library to allocate an independent inference process for each monitoring area. Each process is bound to the receiving buffer, preprocessing subprocess, and model inference subprocess of the corresponding area. Data is transferred between processes through shared memory. The inference process of each area is completely parallel and does not interfere with each other. It can support real-time behavior recognition of up to 50 monitoring areas at the same time. Real-time inference performance: Single-sample inference time ≤15ms, meeting real-time recognition requirements; Results output and alarm linkage: The fixed format of the recognition result is [Region ID][Hardware Timestamp][Node ID]_[Behavior Category], for example: "01_1714523658000_1_falldown"; When a high-risk behavior such as a fall is detected, the local sound and light alarm is immediately triggered on the terminal, and the alarm information is pushed to the duty room management terminal and the caregiver's mobile APP via the MQTT protocol, with an alarm delay of ≤200ms.
[0056] A spatiotemporal dual-attention-enhanced WiFi-CSI behavior recognition method: This method is implemented based on the above system, and the specific execution steps are as follows: Step S1: Distributed Closed-Loop Acquisition and Adaptive Link Scheduling S1.1 Node Deployment and Networking: For each independent monitoring area, one ANN node is deployed in the center of the ceiling, and one PCN node is deployed in each of the three opposite corners to build a three-way cross WiFi link; based on the ESP-Mesh protocol, all ANN nodes are used to form a distributed Mesh network to complete global hardware time synchronization with a synchronization error of ≤1ms; S1.2 Real-time CSI data acquisition: Each PCN node continuously acquires WiFi CSI amplitude data caused by human behavior in its area at a sampling rate of 100Hz and uploads it to the ANN node in the same area in real time. S1.3 Adaptive Link Cleaning and Scheduling: The ANN node performs a link quality assessment every 100ms, filters effective links by jointly evaluating subcarrier spatial consistency and signal-to-noise ratio, updates the dynamic filtering threshold based on the asymmetric EMA algorithm, completes seamless switching between primary and backup links through a link reputation state machine with hysteresis interval, and performs weighted fusion of effective link data to obtain high signal-to-noise ratio CSI data. S1.4 Edge Preprocessing and Data Upload: The ANN node sequentially performs 8th-order Butterworth low-pass filtering, 1-D linear interpolation missing data repair, and min-max normalization on the fused CSI data. A triple structured unique identifier is added to each preprocessed data. After being packaged in units of 1000 frames, the data is uploaded to the high-resource processing terminal through the Mesh network.
[0057] Step S2: Data standardization preprocessing S2.1 Data Parsing and Classification: After receiving the data, the processing terminal parses the triple structured unique identifier, classifies the data by region ID, and aligns the time sequence by timestamp; S2.2PCA Dimensionality Reduction: Using the PCA principal component analysis algorithm, the 52-dimensional CSI data is compressed to 30 dimensions with a cumulative contribution rate of ≥95%, and redundant information is removed; S2.3 Effective Action Segment Extraction: An adaptive threshold segmentation algorithm based on moving variance is used to extract effective action segments from the continuous CSI time-series stream and remove static redundant data; S2.4 Time Length Unification: Unify the effective action segments by trimming / padding them with zeros to a time step length of 120, and output a standardized time tensor with the shape of [number of samples, 120, 30].
[0058] Step S3: Spatiotemporal dual attention enhancement and decoupling calibration for behavior recognition: S3.1 Bidirectional LSTM Basic Feature Extraction: The standardized temporal tensor is input into the pre-trained DAB-LSTM-DC network model. The forward and reverse temporal dependencies of the temporal data are learned synchronously through a two-layer bidirectional LSTM network, and a basic spatiotemporal feature matrix of 120×400 is output. S3.2 Spatiotemporal Dual Attention Feature Enhancement: By using temporal attention branches and spatial attention branches, normalized weights for the temporal and spatial dimensions are calculated respectively to generate a spatiotemporal joint attention weight matrix. The basic spatiotemporal feature matrix is then enhanced element-wise to obtain the enhanced spatiotemporal joint feature matrix. S3.3 Spatiotemporal Feature Decoupling Calibration: The enhanced feature matrix is input into the spatiotemporal feature decoupling calibration layer to complete temporal-spatial intrinsic feature decoupling, adaptive gating calibration, outer product reconstruction and residual fusion, and output the final discriminative feature matrix; S3.4 Classification and Recognition: Global average pooling is performed on the calibrated feature matrix to obtain a 400-dimensional global feature vector. The predicted probability of each behavior category is output through a fully connected layer and a Softmax activation function. The category with the highest probability is selected as the final behavior recognition result.
[0059] Step S4 Multi-region Parallel Inference Output: The processing terminal allocates an independent inference process to each monitoring area, and completes the full-process processing and behavior recognition inference of CSI data from multiple areas in parallel, outputting the recognition result bound with a triple unique identifier; when high-risk behaviors such as falls are detected, it immediately triggers an audible and visual alarm and remote information push.
[0060] This embodiment uses multiple sets of control experiments to quantitatively verify the technical effect of the present invention, as detailed below: Comparison experiment on the discrimination of similar actions: For five easily confused similar actions—bending over, squatting, sitting, lying down, and falling—this invention was compared with existing mainstream algorithms, and the recognition accuracy results are as follows:
[0061] Experimental results show that the accuracy of the present invention in recognizing similar actions is significantly better than that of existing mainstream algorithms. In particular, for falling behavior, the recognition accuracy reaches 99%, effectively solving the core pain point of high misjudgment rate of falling and similar actions such as sitting down and bending over in traditional solutions.
[0062] Experiment on anti-interference capability in complex environments: The recognition accuracy of this invention was tested in three typical scenarios: a laboratory, a living room, and a bedroom with complex furniture obstruction. The results were: 97.2% in the laboratory environment, 95.1% in the living room environment, and 91.3% in the bedroom environment. Even in the bedroom environment with many furniture obstructions and complex multipath effects, it still maintained a high recognition accuracy of over 90%, demonstrating excellent environmental anti-interference ability and adaptability to local deployment.
[0063] System core performance indicators: Reliable transmission range: With two levels of relay nodes, the system's reliable transmission range can reach 220 meters, far exceeding the coverage range of a single node in a traditional centralized architecture; End-to-end latency: The total end-to-end latency from CSI data acquisition to behavior recognition result output is ≤200ms, which meets the real-time recognition requirements; Scalable deployment capability: A single Mesh network supports up to 50 nodes, which can adapt to the scalable deployment needs of up to 50 independent monitoring areas.
[0064] In other embodiments of the present invention, the SNR threshold and Corr threshold can be adaptively adjusted according to the actual deployment environment (such as the degree of obstruction, transmission distance, and interference intensity). For example, in strong interference and long-distance scenarios, the SNR threshold can be adjusted to 15dB and the Corr threshold can be adjusted to 0.6 to ensure the normal operation of the system in complex environments.
[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0066] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A behavior recognition system enhanced with spatiotemporal dual attention, characterized in that: Includes the ESP32-Mesh distributed data acquisition module and the high-resource processing terminal identification module; The ESP32-Mesh distributed data acquisition module includes several regional acquisition units set up for multiple independent monitoring areas, a distributed Mesh transmission network built based on the ESP-Mesh protocol, and an edge environment adaptive multi-link cleaning and scheduling module deployed in each regional acquisition unit. Each of the aforementioned regional acquisition units includes one regional aggregation ANN node operating in AP mode and at least three sensing acquisition PCN nodes operating in STA mode. The ANN node and the corresponding PCN node are deployed in the same independent monitoring area to construct a closed-loop WiFi signal field within the area. Each PCN node and ANN node form three intersecting WiFi links in the vertical, horizontal, and diagonal directions. The PCN node is used to collect WiFi CSI amplitude data caused by human behavior within the area. Each ANN node connects to the distributed Mesh transmission network to aggregate and perform lightweight preprocessing of data collected by PCN nodes within its region. After adding a triple-structured unique identifier to each preprocessed data entry, the data is uploaded through the distributed Mesh transmission network. The format of the triple structured unique identifier is [Region ID][64-bit global synchronization hardware timestamp][Node ID]; The edge environment adaptive multi-link cleaning and scheduling module is deployed within the ANN node to complete joint link quality assessment, dynamic threshold adaptive update, link health state machine scheduling, and multi-link data weighted fusion, outputting clean CSI data with high signal-to-noise ratio. The high-resource processing terminal identification module includes a data preprocessing unit, a DAB-LSTM behavior recognition unit with spatiotemporal feature decoupling calibration, and a multi-region parallel inference unit. The data preprocessing unit is used to parse and standardize the received labeled CSI data and output a standardized time series tensor. The DAB-LSTM behavior recognition unit with spatiotemporal feature decoupling calibration adopts a network architecture of bidirectional LSTM layer combined with spatiotemporal dual attention enhancement layer and spatiotemporal feature decoupling calibration layer, which is used to extract, enhance and calibrate features of standardized temporal tensor and output discriminative feature vector. The multi-region parallel inference unit is used to allocate independent inference threads to each monitoring region, complete behavior classification and recognition based on discriminative feature vectors, and output behavior recognition results bound to region ID, hardware timestamp, and node ID.
2. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system according to claim 1, characterized in that, The edge environment adaptive multi-link cleaning and scheduling module has a built-in sub-module for joint weighted evaluation of subcarrier spatial consistency and signal-to-noise ratio, which is used to complete the two-dimensional quantitative evaluation of link quality and candidate link selection, specifically: Let H be the CSI amplitude matrix collected by the k-th WiFi link within a preset time window T. k ∈R N×T Where N=52 is the number of subcarriers, and T is the number of sampling frames within the time window; First, calculate the mean vector of all subcarriers within the time window: ,in Let be the timing vector of the i-th subcarrier; Next, calculate the subcarrier spatial consistency index of the k-th link. The formula is: ,in, σ is the covariance calculation function, and σ is the standard deviation calculation function; Define the joint quality factor of the k-th link. Where SNRk is the real-time signal-to-noise ratio of the link; Set dual threshold filtering conditions: only when SNRk≥20dB and When the value is ≥0.7, the link enters the valid candidate pool Ω; For all links within the candidate pool Ω, CSI data fusion is performed using a normalized joint weighting formula, which is: ,in, = , where is the fusion weight of the k-th link.
3. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system according to claim 2, characterized in that, The edge environment adaptive multi-link cleaning and scheduling module also incorporates a dynamic threshold adaptive submodule based on asymmetric exponential weighted moving average, used to track environmental noise in real time and update the dynamic threshold for link selection, specifically: set up The instantaneous environmental noise energy during the silent period of frame t is used to update the mean environmental noise in real time using the asymmetric EMA algorithm. The formula is: ; in, , This is the noise rise smoothing coefficient. This is the noise reduction smoothing coefficient; Let the mean absolute noise floor of the initial calibration environment be . Construct a dynamic signal-to-noise ratio threshold With amplitude stationarity threshold The formula is: ; in, The reference signal-to-noise ratio threshold This is the signal-to-noise ratio threshold adjustment factor. As the benchmark stationarity threshold, This is the stability threshold adjustment coefficient.
4. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system according to claim 3, characterized in that, The edge environment adaptive multi-link cleaning and scheduling module also has a built-in link reputation state machine scheduling submodule with hysteresis intervals, used to achieve seamless link switching and anti-ping-pong effect control, specifically: For each WiFi link k, maintain a link health reputation value ranging from [0,1]. At each time step t, the reputation value is updated according to a dynamic threshold, using the following formula: Where ΔS(t) is the reputation value step update, defined as: ; Where R is the positive reward step size, P is the negative penalty step size, and R <P; Set the activation state A for link k k ∈{0,1}, where 1 represents the active state and 0 represents the hot standby state. State switching is completed based on hysteresis dual thresholds, and the state transition equation is: ; in, This is the threshold for link disconnection. This is the link activation threshold, and ; When the primary link goes offline, the ANN node immediately selects the link with the highest health status from the backup links to complete a seamless activation switch.
5. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system according to claim 1, characterized in that, The distributed Mesh transmission network uses 1 / 6 / 11 interference-free channels in the 2.4GHz band. A single network supports a maximum of 50 nodes. The network is equipped with multi-level relay nodes to relay and forward uploaded data from ANN nodes in remote areas that exceed the transmission range of a single node, extending the reliable transmission range of the system to more than 200 meters. The triple structured unique identifier is simultaneously embedded in the filename and the first 32 bytes of the file header of the CSI data file. The region ID is a 2-digit globally unique decimal number, the 64-bit timestamp is generated based on the global software clock of the Mesh network with a synchronization error of ≤10 milliseconds, and the node ID is a unique decimal number of the PCN node in the same region.
6. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system according to claim 1, characterized in that, The lightweight preprocessing of the ANN node specifically includes: using an 8th-order Butterworth low-pass filter to filter high-frequency noise and signal burst interference in the environment; using a 1-D linear interpolation algorithm to repair missing data caused by packet loss during signal transmission; and using a min-max normalization algorithm to map the CSI amplitude data of each subcarrier to the [-1,1] interval. The standardization process of the data preprocessing unit specifically includes: compressing the 52-dimensional subcarrier CSI data to 30 dimensions with a cumulative contribution rate ≥95% through PCA principal component analysis; extracting effective action segments using an adaptive threshold segmentation algorithm based on moving variance; uniformly processing the effective segments to a length of 120 time steps; and outputting a standardized time series tensor with the shape [number of samples, 120, 30].
7. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system according to claim 1, characterized in that, The network architecture of the DAB-LSTM behavior recognition unit with spatiotemporal feature decoupling calibration consists of, in order, a bidirectional LSTM basic feature extraction layer, a spatiotemporal dual attention feature enhancement layer, a spatiotemporal feature decoupling calibration layer, and a fully connected classification output layer. The bidirectional LSTM basic feature extraction layer consists of two bidirectional LSTM networks, each with 200 hidden nodes in each direction. It is used to simultaneously learn the forward and reverse temporal dependencies of the time-series data. The forward hidden state output by the forward LSTM and the reverse hidden state output by the reverse LSTM at the same time are concatenated as vectors to obtain a basic spatiotemporal feature matrix X∈R with dimensions of 120×400. 120×400 ; The spatiotemporal dual attention feature enhancement layer includes a temporal attention branch, a spatial attention branch, and a feature fusion module; The time attention branch is used to assign normalized weights to each time step. The formula for calculating the time attention weights is: ; in, ∈R 400×1 , ∈R 1 For time-attention, trainable parameters, Let be the normalized weight at time step t, and let the sum of the weights at all time steps be 1. The spatial attention branch is used to assign normalized weights to each feature dimension. The formula for calculating the spatial attention weights is: ; in, ∈R 120×1 , ∈R 1 For trainable parameters of spatial attention, Let be the normalized weight of the d-th feature, and the sum of the weights of all feature dimensions is 1; The feature fusion module performs an outer product operation on the temporal weights and spatial weights to obtain the spatiotemporal joint attention weight matrix A∈R. 120×400 The weight matrix and the basic spatiotemporal feature matrix are then subjected to an element-wise Hadamard product to obtain the weighted and enhanced spatiotemporal joint feature matrix. ∈R 120×400 .
8. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition system according to claim 7, characterized in that, The spatiotemporal feature decoupling and calibration layer is a self-developed core module used to complete temporal-spatial intrinsic feature decoupling, adaptive calibration, and residual fusion of the spatiotemporal joint feature matrix. Specifically: Let the input spatiotemporal joint feature matrix be... ∈R T×D Where T=120 is the time step length and D=400 is the feature dimension; Step 1: Decoupling intrinsic features, extracting temporal eigenvectors and spatial eigenvectors respectively: Temporal eigenvectors ∈R T The calculation formula is: ; Spatial eigenvectors ∈R D The calculation formula is: ; Step 2: Two-branch adaptive calibration, performing gated calibration on temporal and spatial intrinsic features separately: Timing calibration gating = Timing characteristics after calibration = ; Space calibration gating = Spatial features after calibration ; in, ∈R T×T , ∈R T Trainable parameters for time-series calibration. ∈R D×D , ∈R D represents the trainable parameters for spatial calibration, and ⊙ represents the element-wise multiplication operation; Step 3: Decouple feature refusion and residual connection. Reconstruct the outer product of the calibrated bi-branch features, and then perform residual fusion with the original features. The formula is: ; ; in, For outer product operation, This is the output calibrated discriminant feature matrix.
9. A spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition method, characterized in that, The system implementation based on any one of claims 1-8 includes the following steps: S1 Distributed Closed-Loop Acquisition and Adaptive Link Scheduling: Deploy one ANN node operating in AP mode and at least three PCN nodes operating in STA mode in each independent monitoring area to construct a closed-loop WiFi signal field and multiple cross WiFi links within the area. Collect WiFi CSI amplitude data caused by human behavior in the area through the PCN nodes. Based on the ESP-Mesh protocol, a distributed mesh transmission network is built from the ANN nodes in all regions to complete the global time synchronization of all nodes; ANN nodes use an edge environment adaptive multi-link cleaning and scheduling module to complete joint link quality assessment, dynamic threshold update, link state machine scheduling and multi-link data weighted fusion. Then, they perform lightweight preprocessing, add a triple structured unique identifier to each data, and upload it to the high-resource processing terminal through a distributed Mesh transmission network. S2 data standardization preprocessing: The processing terminal parses and classifies the received labeled CSI data, performs PCA dimensionality reduction and adaptive activity segmentation in sequence, and outputs a standardized temporal tensor of the input dimension of the matching recognition model; S3 Spatiotemporal Dual Attention Enhancement and Decoupling Calibration Behavior Recognition: Standardized temporal tensors are input into a pre-trained DAB-LSTM network model with spatiotemporal feature decoupling calibration. First, basic spatiotemporal features are extracted through a bidirectional LSTM layer. Then, feature weighting enhancement is performed through a spatiotemporal dual attention layer. Subsequently, intrinsic feature decoupling, adaptive calibration, and residual fusion are completed through a spatiotemporal feature decoupling calibration layer. Finally, the behavior category probability distribution is output through a fully connected layer and a Softmax activation function to complete behavior classification and recognition. S4 Multi-Region Parallel Inference Output: Allocates independent inference threads to each monitoring area, performs behavior recognition inference of CSI data in multiple areas in parallel, and outputs behavior recognition results bound to area ID, hardware timestamp, and node ID, triggering alarm linkage for high-risk behaviors such as falls.
10. The spatiotemporal dual-attention enhanced WiFi-CSI behavior recognition method according to claim 9, characterized in that, The training process of the DAB-LSTM network model described in step S3 uses a self-developed contrastive regularization loss function with inter-class distance constraints. The total loss function formula is as follows: ; in, Here, λ is the standard cross-entropy loss function, and λ is the regularization coefficient. The regularization term for the inter-class distance constraint is calculated using the following formula: ; In the formula, C is the number of behavior categories, B is the training batch size, and m is the minimum inter-class distance threshold. Let i be the class center feature vector of the i-th class. Let be the global feature vector of the b-th training sample. The true label for the b-th training sample; During training, class center feature vectors The method uses a moving average to update in real time, which forcibly increases the feature distance between similar actions and compresses the feature distribution within the class, thus solving the problem of misjudgment of similar actions from the perspective of loss function.