A human behavior recognition method based on multi-sensing time-frequency information enhancement

CN122528077BActive Publication Date: 2026-09-22WUHAN TEXTILE UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611026322.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-22
Estimated Expiration
2046-07-10

AI Technical Summary

Technical Problem

但是该方法主要依赖于单一刚性惯性特征的提取和加强,无法充分利用肢体柔顺形变等其他模态的特征,在区分重力分量相似的微动态动作时极易产生混淆

Benefits of technology

本发明适配多源异构传感数据特性,通过非对称流形投影匹配惯性与柔顺数据的信息熵差异,避免低维稀疏数据的过度参数化映射,有效降低模型参数量,减少边缘终端计算负载,提升模型在树莓派、微控制器等轻量化硬件上的部署可行性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528077B_ABST
    Figure CN122528077B_ABST
Patent Text Reader

Abstract

The application discloses a human behavior recognition method based on multi-sensing time-frequency information enhancement, relates to the cross field of artificial intelligence, edge computing and wearable Internet of Things, and comprises the following steps: based on a distributed wearable sensing network, rigid inertia pose data and compliant deformation data corresponding to human actions are synchronously collected, the collected original data is subjected to time sequence slicing and standardization processing, and a standardized space-time tensor is generated; the standardized space-time tensor is subjected to channel orthogonal decoupling, and inertia submanifold and compliant submanifold are obtained; and according to the Shannon information entropy difference of the two submanifolds, an asymmetric projection network is constructed, and the two submanifolds are respectively mapped to corresponding dimension hidden feature spaces. The application designs an asymmetric initial hidden space mapping through asymmetric manifold scale adaptive matching, avoids excessive parameterization mapping of low-dimensional sparse data from the root, and realizes extreme compression of model parameter quantity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technologies of artificial intelligence, edge computing and wearable Internet of Things, specifically a human behavior recognition method based on multi-sensor time-frequency information enhancement. Background Technology

[0002] Human Activity Recognition (HAR) is the core engine for building digital health monitoring, immersive human-computer interaction, and sports rehabilitation assessment systems. Existing mainstream HAR systems heavily rely on triaxial acceleration and angular velocity data collected by rigid inertial measurement units (IMUs). However, when IMU sensors identify quasi-static or micro-dynamic movements (such as sitting versus standing) with highly similar spatial displacements but different limb morphologies, severe category confusion often occurs due to the high overlap of gravitational acceleration mapping components. In recent years, the development of flexible wearable electronics technology has made it possible to continuously acquire stress deformation signals from the skin or joint surfaces. Flexible sensors can directly physically map the flexion and extension angles of joints, forming a perfect modal complement to IMUs, which excel at monitoring relative spatial displacement. However, achieving deep integration of rigid IMU data and compliant deformable data faces significant technical barriers: First, manifold heterogeneity. IMU data is rich in high-frequency transient oscillations, while compliant data is mostly characterized by low-frequency continuous manifold boundaries. Moreover, the feature dimensions of the two differ by orders of magnitude. Using traditional symmetric network architectures will cause severe "feature collapse" and information overload. Second, edge computing power is limited. Deploying complex multimodal networks on resource-constrained edge terminals such as Raspberry Pi or microcontrollers faces the dual challenges of power consumption and latency.

[0003] In the prior art, Chinese patent CN111582361A discloses a "human behavior recognition method based on inertial sensors." This method constructs a deep neural network to extract spatial and temporal features of IMU signals for action classification, thereby improving the insufficient feature representation capabilities of traditional shallow models. However, this method mainly relies on the extraction and enhancement of single rigid inertial features and cannot fully utilize features of other modalities such as limb compliance deformation. It is also prone to confusion when distinguishing micro-dynamic actions with similar gravitational components.

[0004] To address this, we propose a human behavior recognition method based on enhanced multi-sensor time-frequency information. Summary of the Invention

[0005] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a human behavior recognition method based on multi-sensor time-frequency information enhancement, which can effectively solve the problems of the existing technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions; This invention discloses a human behavior recognition method based on multi-sensor time-frequency information enhancement, comprising: Step 1: Based on a distributed wearable sensor network, synchronously collect rigid inertial pose data and compliant deformation data corresponding to human body movements, perform temporal slicing and standardization processing on the collected raw data, and generate a standardized spatiotemporal tensor. Step 2: Perform channel orthogonal decoupling on the standardized spatiotemporal tensor to obtain inertial submanifolds and compliant submanifolds. Based on the difference in Shannon information entropy between the two submanifolds, construct an asymmetric projection network to map the two submanifolds to the latent feature space of the corresponding dimension. Step 3: Construct a temporal analysis branch. Based on the multi-granularity extended temporally separable operator, perform multi-branch parallel temporal feature extraction and aggregation on the two latent features to obtain temporal motion features. Step 4: Construct a frequency domain analysis branch, perform adaptive wavelet frequency domain decomposition on the two latent features respectively, construct a cross-manifold cooperative gating mechanism based on the frequency domain features of the compliant submanifold, perform nonlinear residual modulation on the frequency domain features of the inertial submanifold, and obtain the fused frequency domain features after noise filtering. Step 5: Based on the multi-order expectation-covariance collaborative aggregation mechanism, multi-order statistical feature extraction and aggregation are performed on the fused frequency domain features to obtain frequency domain spectral features; Step 6: Input the temporal motion features and frequency spectral features into the joint decision space, perform feature fusion and behavior classification, and output the human behavior recognition results.

[0007] Furthermore, step 1 includes the following sub-steps: S11: Configure the sampling clock for the distributed wearable sensor network, to 9-axis rigid inertial pose data were acquired synchronously at high frequency. and 2-channel compliant deformation data This forms the original observation matrix. ; S12: Construct a time-series sliding window operator to perform time-series slicing on the original observation matrix and generate a standardized spatiotemporal tensor. The calculation formula for the standardized spatiotemporal tensor is as follows: ; in, Denotes the τ-th normalized spacetime tensor; This indicates the zero-mean-unit-variance normalization operation performed on each independent channel; This represents the multi-source heterogeneous raw observation matrix obtained synchronously by wearable sensor networks; The index represents the starting time step index of the τth time series sliding window on the time axis corresponding to the original observation matrix; L represents the fixed length of the preset time series sliding window; and T represents the matrix transpose operator.

[0008] Furthermore, step 2 includes the following sub-steps: S21: Standardize the spacetime tensor Decoupled into an inertial submanifold in the orthogonal dimensions of the channel. and compliant manifold ; S22: Construct an initial feature projection layer with an asymmetric receptive field, respectively for the inertial submanifold and compliant manifold Perform convolutional manifold mapping; where the inertial submanifold Using high-dimensional mapping tensors Mapped to dimension Latent feature space; compliant submanifold Using low-dimensional mapping tensors Mapped to dimension The latent feature space; The mapping dimension parameter satisfies the inequality > .

[0009] Furthermore, in step 3, the computation process of the multi-granularity expanded temporally separable operator includes: S31: For the input original feature tensor, construct three parallel temporal topology branches, namely the local transient motion capture branch, the long-range temporal dependency branch, and the extreme value significance preservation branch. S32: Temporal feature extraction for the corresponding dimension is completed through three parallel branches, and the output features of each branch are obtained; S33: The output features of the three branches are aggregated by an adaptive channel activation operator with a cross-layer residual structure to obtain the final temporal motion features.

[0010] Furthermore, the local transient motion capture branch constructs a difference compensation sequence by extracting the first-order time derivative of the feature tensor. After merging this sequence with the original features, it is fed into the depthwise separable convolution module for processing to obtain the output features of this branch. The calculation formula is as follows: ; In the formula: This represents the output characteristics of the local transient motion capture branch; This represents the Swish activation function; This indicates a one-dimensional batch normalization operation; This represents a 1×1 one-dimensional pointwise convolution operation; This represents a one-dimensional depthwise convolution operation with kernel size K=3 and dilation rate d=1. ⊕ represents the original feature tensor of the input; ⊕ represents the summation and compensation of features in the channel dimension; λ is the differential compensation weight that the network can learn; · represents the scalar multiplication operation; This represents the discrete first-order difference derivative sequence calculated over the time dimension; The long-range temporal dependency branch generates a local energy envelope mask by calculating the squared magnitude of the signal within a local temporal window. It then performs nonlinear residual modulation on the result of a one-dimensional depthwise convolution with a kernel size of 5 and an expansion rate of 2 to obtain the output features of this branch. The calculation formula for this branch is as follows: ; In the formula: This represents the output characteristics of long-term time-dependent branches; This represents the element-wise product of Hadamah; This represents a one-dimensional dilated depthwise convolution operation with kernel size K=5 and dilation rate d=2; Sigmoid This represents the Sigmoid activation gate function; This represents a one-dimensional average pooling operation with a kernel size of 5. The extreme value significance preservation branch uses a local signal-to-noise ratio self-calibration relative peak operator to calculate the relative deviation between the maximum value and the mean within a local window, thus obtaining the output feature of this branch. The calculation formula is as follows: ; In the formula: The output characteristics of the branch that preserves the significance of extreme values; This represents a one-dimensional max-pooling operation with a kernel size of 3; tanh This represents the Tanh hyperbolic tangent activation function; This is an adaptive scaling factor; This represents a one-dimensional average pooling operation with a kernel size of 3. To prevent extremely small constant biases when dividing by zero; The final aggregated output feature of the multi-granularity extended temporally separable operator The calculation formula is: ; In the formula: Channel cascading operations for characteristic manifolds; Convolution of the channel's dimension-reduced projection points; It is an adaptive channel activation operator, which includes average pooling, a fully connected layer, a ReLU activation function, a fully connected layer, and a Sigmoid activation function in sequence.

[0011] Furthermore, in step 4, the adaptive wavelet frequency domain decomposition operation follows the following rules: The high-dimensional latent features corresponding to the inertial submanifold after mapping by the asymmetric projection network are input into the Haar wavelet for frequency domain decomposition, and the low-dimensional latent features corresponding to the compliant submanifold are input into the Dobersey wavelet for frequency domain decomposition. The features after the two frequency domain decompositions are respectively sent to the feature extraction module to complete the initial frequency domain feature extraction. The feature extraction module consists of a one-dimensional depthwise convolution with Kernel=3, a normalization layer, a ReLU activation function, a 1×1 one-dimensional pointwise convolution, a normalization layer, and a ReLU activation function.

[0012] Furthermore, in step 4, the cross-manifold cooperative gating mechanism includes: S41: Define the shallow compliant frequency domain features directly output by Dobessie wavelet decomposition as follows: The deep rigid frequency domain features extracted by the feature extraction module are: ; Applying shallow compliant frequency domain features A discrete nonlinear Tegor-Kaiser motion energy operator is introduced along the time dimension to calculate the transient deformation energy matrix. The corresponding calculation formula is: ; In the formula: Represents the transient deformation energy matrix; Indicates compliant frequency domain characteristics; This represents the element-wise product of Hadamah; and These represent shift operators that translate one time step to the left and one time step to the right along the time dimension, respectively. S42: The transient deformation energy matrix is ​​reconstructed using one-dimensional local moving average pooling and latent space features to generate a dynamic cooperative gating matrix that is strictly aligned with the time step. The corresponding calculation formula is: ; In the formula: Represents a dynamic collaborative gating matrix; This represents a local average pooling operation that preserves temporal resolution. and These are respectively the reduction and enhancement of the latent space. Orthogonal transformation operation; It is the ReLU activation function; Use the Sigmoid activation function; S43: The deep rigid frequency domain features are modulated step-by-step using a dynamic collaborative gating matrix, and then fused into frequency domain features via point convolution transformation and cross-layer residual connections. The corresponding calculation formula is: ; In the formula: Indicates the fused frequency domain characteristics; Represents rigid frequency domain characteristics; Indicates the use of feature reconstruction One-dimensional convolution operation; This indicates the summation of characteristic residuals across layers.

[0013] Furthermore, in step 5, the multi-order expectation-covariance collaborative aggregation mechanism follows the following: S51: For the fused frequency domain features of the input, calculate the first-order global expectation vector in the channel dimension; S52: Divide the fused frequency domain features into multiple orthogonal subspaces in the channel dimension, and calculate the regularized second-order unbiased covariance matrix of the features in each orthogonal subspace respectively; S53: Extract the upper triangular independent elements of the covariance matrix of each subspace based on the manifold semi-vectorization operator, and concatenate them with the first-order global expectation vector in the latent distribution space to output the final frequency domain spectral features.

[0014] Furthermore, the calculation process of the first-order global expectation vector follows: Let the input fused frequency domain features be a high-dimensional spectral state feature matrix. The first-order global expectation vector of the channel dimension is calculated. The calculation formula is: ; In the formula: Represents the first-order global expectation vector along the channel dimension; Represents the high-dimensional spectral state characteristic matrix; Indicates the length of the feature sequence; Indicates size is A column vector of all 1s; The calculation process of the regularized second-order unbiased covariance matrix is ​​as follows: The high-dimensional spectral characteristic matrix is... Divide the channel dimension into K orthogonal subspaces. For the κth group of features Calculate its regularized second-order unbiased covariance matrix. The corresponding calculation formula is: ; In the formula: Indicates the first The regularized second-order unbiased covariance matrix of the group feature subspace; Indicates the first Characteristic matrices of a set of orthogonal subspaces; express The identity matrix of order 1; Indicates length is A vector of all 1s; Represents the matrix transpose operator; To maintain the positive definiteness of the matrix, a minimal perturbation factor; The dimension is The identity matrix; The final calculation formula for the frequency domain spectral characteristics is: ; In the formula: This represents the multidimensional spectral characterization of the final output; Represents the semi-vectorization operator of a manifold; Let represent the covariance matrix of each orthogonal subspace; This represents the cascading operation of eigenvectors in the latent space.

[0015] Furthermore, the joint decision space in step 6 includes, in sequence, a feature flattening layer, a feature fully connected mapping layer, a layer normalization layer, a nonlinear activation layer, a random deactivation layer, and an output classification layer; After feature flattening, temporal motion features and frequency spectral features are fused through fully connected mapping, and finally the human behavior recognition result is output through the output classification layer.

[0016] Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects: This invention adapts to the characteristics of multi-source heterogeneous sensing data. By matching the information entropy difference between inertial and compliant data through asymmetric manifold projection, it avoids excessive parameterization mapping of low-dimensional sparse data, effectively reduces the number of model parameters, reduces the computing load on edge terminals, and improves the feasibility of deploying the model on lightweight hardware such as Raspberry Pi and microcontrollers.

[0017] This invention employs a multi-granularity extended temporally separable operator to extract local transient, long-range temporal, and extreme value features in parallel. Combined with differential compensation and signal-to-noise ratio self-calibration mechanisms, it filters out high-frequency glitches and jitter noise from sensors, enhances the anti-disturbance capability of temporal feature extraction, and improves the representation accuracy of motion features. This invention is based on adaptive wavelet decomposition and cross-manifold collaborative gating mechanism to achieve targeted fusion and noise suppression of rigid and flexible frequency domain features, preserve signal temporal resolution, accurately capture frequency domain details of limb deformation and motion, and improve the problem of feature smoothing distortion in traditional frequency domain processing. This invention uses multi-order expectation-covariance collaborative aggregation to simultaneously extract static posture bias and dynamic action texture features, reducing the confusion rate of recognition of highly similar micro-dynamic behaviors such as sitting and standing, and filling the accuracy gap of conventional methods in micro-action recognition. This invention shortens the single-frame inference latency by classifying time-frequency dual-branch features through a lightweight decision space, thereby meeting the real-time human behavior recognition needs of wearable IoT devices and balancing recognition accuracy with edge computing efficiency. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0019] Figure 1 This is a flowchart illustrating a human behavior recognition method based on multi-sensor time-frequency information enhancement. Figure 2 This is a schematic diagram of an asymmetric manifold decoupling network in a human behavior recognition method based on multi-sensor time-frequency information enhancement; Figure 3 This is a schematic diagram of a temporal-domain flow multi-granularity extended temporally separable operator module in a human behavior recognition method based on multi-sensor time-frequency information enhancement; Figure 4 This is a schematic diagram of the mid-frequency domain flow cross-manifold collaborative gating mechanism and multi-stage pooling aggregation module in a human behavior recognition method based on multi-sensor time-frequency information enhancement. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0021] The present invention will be further described below with reference to embodiments.

[0022] Example 1: This embodiment presents a human behavior recognition method based on multi-sensor time-frequency information enhancement, such as... Figures 1-4 As shown, it includes: Step 1: Based on a distributed wearable sensor network, synchronously collect rigid inertial pose data and compliant deformation data corresponding to human body movements, perform temporal slicing and standardization processing on the collected raw data, and generate a standardized spatiotemporal tensor. Step 1 includes the following sub-steps: S11: Configure the sampling clock for the distributed wearable sensor network, to 9-axis rigid inertial pose data were acquired synchronously at high frequency. and 2-channel compliant deformation data This forms the original observation matrix. ; S12: Construct a time-series sliding window operator with length L=100 and step size S=50 to perform time-series slicing on the original observation matrix and generate a normalized spatiotemporal tensor. The calculation formula is as follows: ; The original multi-source heterogeneous sensing data is extracted by a time-series sliding window with a fixed length and step size. Zero-mean-unit variance normalization is performed on each sensor channel to eliminate the deviation of physical dimensions. At the same time, the transpose matrix is ​​adapted to the network input dimension and normalized to generate a standardized spatiotemporal tensor, which lays a unified and high-quality input foundation for subsequent heterogeneous feature decoupling. in, Denotes the τ-th normalized spacetime tensor; This indicates the zero-mean-unit-variance normalization operation performed on each independent channel to eliminate the dimensional shifts of physical quantities between different sensors; This represents the multi-source heterogeneous raw observation matrix obtained synchronously by wearable sensor networks; The index of the starting time step of the τ-th temporal sliding window on the time axis corresponding to the original observation matrix is ​​represented; L represents the fixed length of the preset temporal sliding window; T represents the matrix transpose operator, which is used to transform the rows and columns of the truncated two-dimensional submatrix to adapt to the dimension format requirements of the input tensor of the downstream network layer. Step 2: Perform channel orthogonal decoupling on the standardized spatiotemporal tensor to obtain inertial submanifolds and compliant submanifolds. Based on the difference in Shannon information entropy between the two submanifolds, construct an asymmetric projection network to map the two submanifolds to the latent feature space of the corresponding dimension. Step 2 includes the following sub-steps: S21: Standardize the spacetime tensor Decoupled into an inertial submanifold in the orthogonal dimensions of the channel. and compliant manifold ; S22: Construct an initial feature projection layer with an asymmetric receptive field, respectively for... and Perform convolutional manifold mapping; for rigid inertial pose data with higher information entropy, the corresponding inertial submanifold... Using high-dimensional mapping tensors Map it to dimension The latent feature space; the compliant submanifold corresponding to compliant deformable data with relatively sparse information entropy. Using low-dimensional mapping tensors Map it to dimension The latent feature space; The mapping dimension parameter satisfies the inequality > To achieve Pareto optimal allocation of edge computing power and modal feature representation capability; Step 3: Construct a temporal analysis branch. Based on the multi-granularity extended temporally separable operator, perform multi-branch parallel temporal feature extraction and aggregation on the two latent features to obtain temporal motion features. In step 3, the computation process of the multi-granularity extended temporally separable operator includes: S31: For the input original feature tensor Three parallel temporal topology branches are constructed: a local transient motion capture branch, a long-range temporal dependency branch, and an extreme value significance preservation branch. S32: Temporal feature extraction for the corresponding dimension is completed through three parallel branches, and the output features of each branch are obtained; S33: By using an adaptive channel activation operator with a cross-layer residual structure, the output features of the three branches are aggregated to obtain the final temporal motion features; The local transient motion capture branch constructs a difference-compensated sequence by extracting the first-order time derivative of the feature tensor. This sequence is then combined with the original features and fed into a depthwise separable convolutional module consisting of a 1D depthwise convolution with Kernel=3, a normalization layer, a Swish activation function, a 1×1 1D pointwise convolution, a normalization layer, and another Swish activation function. Its output... The calculation formula is: ; This formula first extracts the first-order temporal difference of the feature tensor to construct a compensation sequence and fuses it with the original features. Then, it extracts local transient action features through small kernel depth convolution, normalization, Swish activation and pointwise convolution, accurately capturing the details of instantaneous changes in action and making up for the shortcomings of traditional convolution in representing transient actions. In the formula: This represents the output characteristics of the local transient motion capture branch; This represents the Swish activation function; This indicates a one-dimensional batch normalization operation; This represents a 1×1 one-dimensional pointwise convolution operation; This represents a one-dimensional depthwise convolution operation with kernel size K=3 and dilation rate d=1. ⊕ represents the original feature tensor of the input; ⊕ represents the summation and compensation of features in the channel dimension; λ is the differential compensation weight that the network can learn; · represents the scalar multiplication operation; This represents the discrete first-order difference derivative sequence calculated over the time dimension; Long-range temporal dependent branches generate local energy envelope masks by calculating the squared magnitude of the signal within a local temporal window. These masks are then used to perform nonlinear residual modulation on the one-dimensional depthwise convolution result with a kernel size of 5 and an dilation rate of 2. The output of this method... The calculation formula is: ; This formula calculates the squared magnitude of the local window of the feature tensor to generate an energy envelope mask, performs nonlinear residual modulation on the result of the large kernel dilated depth convolution, and then performs normalization, activation and pointwise convolution processing to effectively capture long-range temporal dependencies, avoid interference from dilated convolution grids, and make the extraction of long-term action features more stable. In the formula: This represents the output characteristics of long-term time-dependent branches; This represents the element-wise product of Hadamah; This represents a one-dimensional dilated depthwise convolution operation with kernel size K=5 and dilation rate d=2; Sigmoid represents the Sigmoid activation gate function. This represents a one-dimensional average pooling operation with a kernel size of 5, used to extract the local motion energy envelope; The extreme value significance preservation branch employs a relative peak operator with local signal-to-noise ratio self-calibration to calculate the relative deviation between the maximum value and the mean within a local window, filtering out glitches in the sensor data. Its output... The calculation formula is: ; This formula calculates the relative deviation of features through local window max pooling and average pooling, and combines adaptive scaling factor and hyperbolic tangent activation to filter data spikes, retain the extreme significance information of actions, remove high-frequency noise from sensors, and improve feature purity and recognizability. In the formula: The output characteristics of the branch that preserves the significance of extreme values; This represents a one-dimensional max-pooling operation with a kernel size of 3; tanh This represents the Tanh hyperbolic tangent activation function; This is an adaptive scaling factor; This represents a one-dimensional average pooling operation with a kernel size of 3. To prevent extremely small constant biases when dividing by zero; Temporal motion characteristics of the final aggregated output of multi-granularity extended temporally separable operators The calculation formula is: ; The temporal feature channels of the three parallel branches are concatenated, and after dimensionality reduction and projection convolution, they are fed into an adaptive channel activation operator. Then, the cross-layer residuals are fused with the original features to adaptively enhance the effective feature channels and suppress redundant information, making the aggregation of multi-granular temporal features more efficient. In the formula: Channel cascading operations for characteristic manifolds; Convolution of the channel's dimension-reduced projection points; It is an adaptive channel activation operator, which includes average pooling, a fully connected layer, a ReLU activation function, a fully connected layer, and a Sigmoid activation function in sequence; Step 4: Construct a frequency domain analysis branch, perform adaptive wavelet frequency domain decomposition on the two latent features respectively, construct a cross-manifold cooperative gating mechanism based on the frequency domain features of the compliant submanifold, perform nonlinear residual modulation on the frequency domain features of the inertial submanifold, and obtain the fused frequency domain features after noise filtering. In step 4, the operation of adaptive wavelet frequency domain decomposition follows the following rules: The high-dimensional latent features corresponding to the inertial submanifold after mapping by the asymmetric projection network are input into the Haar wavelet for frequency domain decomposition, and the low-dimensional latent features corresponding to the compliant submanifold are input into the Dobersey wavelet for frequency domain decomposition. The features after the two frequency domain decompositions are respectively fed into the feature extraction module to complete the initial frequency domain feature extraction. The feature extraction module consists of a one-dimensional depthwise convolution with Kernel=3, a normalization layer, a ReLU activation function, a 1×1 one-dimensional pointwise convolution, a normalization layer, and a ReLU activation function. In step 4, the cross-manifold cooperative gating mechanism includes: S41: Define the shallow compliant frequency domain features directly output by Dobessie wavelet decomposition as follows: The deep rigid frequency domain features extracted by the feature extraction module are: ; Applying shallow compliant frequency domain features A discrete nonlinear Tegor-Kaiser motion energy operator is introduced along the time dimension to calculate the transient deformation energy matrix. The corresponding calculation formula is: ; Based on the compliant frequency domain characteristics, the above formula calculates transient deformation energy through element-wise multiplication and time dimension shift. It can accurately capture the transient energy changes of limb deformation without complex preprocessing, providing a reliable energy reference for cross-manifold gating. In the formula: Represents the transient deformation energy matrix; Indicates compliant frequency domain characteristics; and These represent shift operators that translate one time step to the left and one time step to the right along the time dimension, respectively, with zero padding at the boundaries; S42: The transient deformation energy matrix is ​​reconstructed using one-dimensional local moving average pooling and latent space features to generate a dynamic cooperative gating matrix that is strictly aligned with the time step. The corresponding calculation formula is: ; This formula processes the transient deformation energy matrix through local moving average pooling and latent space orthogonal transformation, and generates a dynamic cooperative gating matrix through Sigmoid activation. It retains the time resolution throughout the process and generates a gating vector that is strictly aligned with the time step, thus completing the precise temporal modulation of rigid features. In the formula: Represents a dynamic collaborative gating matrix; This represents a local average pooling operation that preserves temporal resolution. and These are respectively the reduction and enhancement of the latent space. Orthogonal transformation operation; It is the ReLU activation function; S43: The deep rigid frequency domain features are modulated step-by-step using a dynamic collaborative gating matrix, followed by point convolution transformation and cross-layer residual connection to output the fused frequency domain features after targeted noise filtering. The corresponding calculation formula is: ; The above formula uses a dynamic collaborative gating matrix to modulate the rigid frequency domain features step by step and element by element. After 1×1 convolution reconstruction and addition with the cross-layer residual, the high-frequency jitter noise of the rigid features is targeted and filtered out. The rigid-flexible heterogeneous frequency domain features are fused, so that the fused features are more in line with the physical characteristics of the action. In the formula: This indicates the fusion characteristics after targeted modulation; Represents rigid frequency domain characteristics; Indicates the use of feature reconstruction One-dimensional convolution operation; This indicates the summation of characteristic cross-layer residuals; Step 5: Based on the multi-order expectation-covariance collaborative aggregation mechanism, multi-order statistical feature extraction and aggregation are performed on the fused frequency domain features to obtain frequency domain spectral features; In step 5, the multi-order expectation-covariance co-aggregation mechanism follows the following: S51: For the fused frequency domain features of the input, calculate the first-order global expectation vector in the channel dimension to represent the static bias field affected by gravity; S52: Divide the fused frequency domain features into multiple orthogonal subspaces in the channel dimension, and calculate the regularized second-order unbiased covariance matrix of the features in each orthogonal subspace to characterize the texture correlation between non-steady frequency components. S53: Extract the upper triangular independent elements of the covariance matrix of each subspace based on the manifold semi-vectorization operator, and concatenate them with the first-order global expectation vector in the latent distribution space to output the final frequency domain spectral features. The calculation process of the first-order global expectation vector follows: Let the input fused frequency domain features be a high-dimensional spectral state feature matrix. The first-order global expectation vector of the channel dimension is calculated. The calculation formula is: ; This formula performs a global average calculation of the channel dimension of the high-dimensional spectral feature matrix, extracts the first-order global expectation vector to represent the static posture bias under the influence of gravity, accurately distinguishes highly similar static micro-movements such as sitting and standing, and solves the problem of fuzzy static posture recognition in existing methods. In the formula: Represents the first-order global expectation vector along the channel dimension; Represents the high-dimensional spectral state characteristic matrix; Indicates the length of the feature sequence; Indicates size is A column vector of all 1s; The calculation process for the regularized second-order unbiased covariance matrix is ​​as follows: The high-dimensional spectral characteristic matrix is... Divide the channel dimension into K orthogonal subspaces. For the κth group of features Calculate its regularized second-order unbiased covariance matrix. The corresponding calculation formula is: ; The above formula divides the spectral features into orthogonal subspaces and then calculates the regularized second-order unbiased covariance matrix. A minimal perturbation factor is added to ensure the positive definiteness of the matrix, capture the texture correlation of non-steady-state frequency components, mine the frequency domain details of dynamic actions, and make up for the lack of dynamic texture representation by first-order features. In the formula: Indicates the first The regularized second-order unbiased covariance matrix of the group feature subspace; Indicates the first Characteristic matrices of a set of orthogonal subspaces; express The identity matrix of order 1; Indicates length is A vector of all 1s; Represents the matrix transpose operator; To maintain the positive definiteness of the matrix, a minimal perturbation factor; The dimension is The identity matrix; The final formula for calculating the frequency domain spectral characteristics is: ; This formula extracts the upper triangular independent elements of the covariance matrix of each subspace and performs semi-vectorization processing. It is then concatenated with the first-order global expectation vector in the latent space to integrate the frequency domain information of static pose and dynamic texture, generating a multi-dimensional spectral representation and providing a comprehensive frequency domain basis for subsequent fusion classification. In the formula: This represents the multidimensional spectral characterization of the final output; Represents the semi-vectorization operator of a manifold; Let represent the covariance matrix of each orthogonal subspace; This represents the concatenation (merging) operation of feature vectors in the latent space; Step 6: Input the temporal motion features and frequency spectral features into the joint decision space, perform feature fusion and behavior classification, and output the human behavior recognition results; The joint decision space in step 6 includes, in sequence, a feature flattening layer, a feature fully connected mapping layer, a layer normalization layer, a nonlinear activation layer, a random deactivation layer, and an output classification layer; After feature flattening, temporal motion features and frequency spectral features are fused through fully connected mapping, and finally the human behavior recognition result is output through the output classification layer.

[0023] The methods described in the above embodiments can accurately adapt to the heterogeneous data characteristics of high-information-entropy inertial sensing and low-information-entropy flexible sensing in specific implementation scenarios, avoiding excessive modeling of low-dimensional data, significantly compressing model computation, and adapting to the computing power limitations of edge devices. Simultaneously, they effectively filter out high-frequency glitches and jitter noise in the sensing signals, significantly improving signal anti-interference capabilities, and can accurately distinguish highly similar micro-dynamic postures such as sitting and standing, eliminating recognition blind spots. Through the coordinated use of time-frequency information, human behavior recognition on wearable devices becomes more accurate and stable, meeting the needs of real-time monitoring and efficient discrimination.

[0024] It should be noted that: The distributed wearable sensor network adopts a distributed wearing architecture of the human torso and limbs. A 9-axis rigid inertial pose sensor, using an industrial-grade MEMS inertial measurement unit, is fixedly worn at three key attitude nodes: the chest, wrist, and ankle, to collect rigid inertial pose data of three-axis acceleration, three-axis angular velocity, and three-axis magnetic field strength. A 2-channel compliant deformation sensor, using a flexible capacitive deformation sensor, is fitted to the curved surfaces of the elbow and knee joints to collect compliant deformation data generated by joint flexion and extension. The sensor network employs a dual synchronization mechanism of hardware clock synchronization and software timestamp calibration, using a unified sampling frequency of 50Hz. Synchronous triggering and acquisition of data from multiple sensors is achieved via an I2C communication bus, avoiding timing misalignment. In the time-series slicing of the original observation matrix, the sliding window length L=100 and step size S=50 are the optimal parameters to adapt to the short-term characteristics of human behavior, which can cover the complete time-series cycle of a single action. The zero-mean-unit variance normalization operation is performed frame by frame for each independent channel of each sensor to eliminate the numerical offset caused by different sensor ranges and physical dimensions, and ensure the characteristic consistency of the standardized spatiotemporal tensor.

[0025] After channel orthogonal decoupling of the standardized spatiotemporal tensor, the Shannon information entropy of the inertial and compliant submanifolds is calculated using the discrete information entropy formula. The calculated information entropy of the inertial submanifold is approximately 4.2 times that of the compliant submanifold. Based on this, the mapping dimension parameter of the asymmetric projection network is determined. =64、 =16, which satisfies the condition. > The core constraint is a convolutional kernel size of K=3, balancing feature extraction efficiency with edge computing power consumption. The convolutional manifold mapping of the asymmetric projection network is implemented using one-dimensional depthwise separable convolution, avoiding parameter redundancy caused by fully connected layers. This fundamentally avoids over-parameterization of low-dimensional sparse and smooth data, achieving extreme compression of model parameters and adapting to the computing power limitations of edge terminals.

[0026] In the multi-granularity extended temporally separable operator, the learnable parameters are initialized according to the following rules: the differential compensation weight λ is initialized to 0.1, the adaptive scaling factor α is initialized to 1.0, and the zero bias is prevented. The value is 10^-6; the output channels of the local transient action capture branch, the long-range temporal dependency branch, and the extreme value saliency preservation branch are all 16, and the total number of channels after channel concatenation is 48, which is then compressed to 32 channels through channel dimensionality reduction projection point convolution. In the adaptive channel activation operator, global average pooling is used, the hidden layer dimension of the fully connected layer is 16, and the adaptive allocation of channel weights is achieved through ReLU and Sigmoid activation. The cross-layer residual structure directly superimposes the input features and aggregated features, which preserves the original features while enhancing effective temporal information, and accurately captures local transient actions, long-range temporal dependencies, and extreme value saliency features.

[0027] In the adaptive wavelet frequency domain decomposition, the high-dimensional latent features of the inertial submanifold are decomposed using Haar wavelets in a two-level frequency domain to adapt to the high-frequency transient features of the inertial data; the low-dimensional latent features of the compliant submanifold are decomposed using Dobessi wavelets (db4) in a three-level frequency domain to adapt to the low-frequency continuous deformation features of the compliant data. In the cross-manifold collaborative gating mechanism, the kernel size of the local moving average pooling is set to 5 to preserve the complete time resolution; the latent space order reduction transform Υdesc and the order increase transform Υasc are both implemented using 1×1 one-dimensional orthogonal convolution with no parameter redundancy. The Tegor-Kaiser motion energy operator directly acts on the shallow compliant frequency domain features without additional smoothing preprocessing, accurately calculating the transient deformation energy. The generated dynamic collaborative gating matrix is ​​aligned with the rigid frequency domain features step-by-step and element-by-element, completing targeted noise filtering and nonlinear residual modulation, and fusing the rigid-flexible heterogeneous frequency domain features.

[0028] In the multi-order expectation-covariance collaborative aggregation mechanism, the number of orthogonal subspaces K=4, uniformly dividing the fused frequency domain features into 4 orthogonal subspaces; the minimal perturbation factor η, which maintains the positive definiteness of the matrix, takes a value of 10^-5, ensuring that the covariance matrix is ​​invertible and stable. (Manifold semi-vectorization operator) This method is used to extract the independent elements of the upper triangular region of the covariance matrix, remove redundant symmetric information, and cascade the second-order covariance features with the first-order global expectation vector in the latent distribution space. It simultaneously represents the static gravity posture bias and dynamic unsteady action texture, solving the problem of blind spots in the recognition of highly similar micro-dynamic postures such as sitting and standing.

[0029] In the joint decision space, the feature flattening layer flattens the temporal motion features and frequency domain spectral features into a one-dimensional vector. The fully connected mapping layer has a dimension of 128. The layer normalization layer uses sample-by-sample normalization. The nonlinear activation layer uses the Swish function. The dropout rate of the random deactivation layer is set to 0.3. The output classification layer uses the Softmax activation function to output the probability of human behavior category. The model training uses the cross-entropy loss function, the optimizer is AdamW, the initial learning rate is set to 10^-3, the batch size is 32, and the number of training epochs is 50. Real-time inference can be achieved on edge devices Raspberry Pi 4B and STM32H7 microcontrollers. The single-frame inference latency is less than 20ms, and the number of parameters is compressed to 1 / 8 of that of traditional multimodal models, balancing recognition accuracy and edge deployment feasibility.

[0030] In summary, the methods described in the above embodiments, targeting the heterogeneous data characteristics of a 9-channel high-information-entropy IMU and a 2-channel low-information-entropy flexible sensor, construct three core innovative technology systems. First, by using asymmetric manifold scale adaptive matching to design an asymmetric initial latent space mapping, the excessive parameterization mapping of low-dimensional sparse data is avoided from the root, achieving extreme compression of model parameters. Second, a unique physical mechanism-driven temporal capture network is created, completely overturning the traditional feature extraction paradigm. In the temporal branch, kinematic velocity / acceleration constraints are implicitly introduced through first-order differential compensation. Energy envelope gating constructed using the squared magnitude of local signals overcomes the grid interference effect of long-term dilated convolution. At the same time, a dynamic signal-to-noise ratio self-calibrating pooling operator is proposed. The system effectively filters out the inherent high-frequency micro-spur noise of flexible sensors, achieving a significant improvement in the model's robustness against disturbances. Finally, addressing the limitation of conventional global pooling in losing temporal resolution, a one-dimensional local sliding temporal gating mechanism is pioneered in the frequency domain branch. This mechanism directly applies the nonlinear Teger-Kaiser energy operator to the shallow wavelet features that are not disrupted by convolutional smoothing, keenly capturing transient burst energy and generating a dynamic mask matrix aligned one-to-one with the time steps. This enables targeted residual modulation step-by-step. Simultaneously, a multi-order expectation-covariance collaborative aggregation mechanism is combined to simultaneously capture the first-order global expectation representing static gravity attitude and the second-order covariance representing high-frequency non-steady-state motion texture, thus largely solving the problem of blind spots in the recognition of highly similar micro-dynamic attitudes.

[0031] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for human behavior recognition based on multi-sensor time-frequency information enhancement, characterized in that, include: Step 1: Based on a distributed wearable sensor network, synchronously collect rigid inertial pose data and compliant deformation data corresponding to human body movements, perform temporal slicing and standardization processing on the collected raw data, and generate a standardized spatiotemporal tensor. Step 2: Perform channel orthogonal decoupling on the standardized spatiotemporal tensor to obtain inertial submanifolds and compliant submanifolds. Based on the difference in Shannon information entropy between the two submanifolds, construct an asymmetric projection network to map the two submanifolds to the latent feature space of the corresponding dimension. Step 3: Construct a temporal analysis branch. Based on the multi-granularity extended temporally separable operator, perform multi-branch parallel temporal feature extraction and aggregation on the two latent features to obtain temporal motion features. Step 4: Construct a frequency domain analysis branch, perform adaptive wavelet frequency domain decomposition on the two latent features respectively, construct a cross-manifold cooperative gating mechanism based on the frequency domain features of the compliant submanifold, perform nonlinear residual modulation on the frequency domain features of the inertial submanifold, and obtain the fused frequency domain features after noise filtering. Step 5: Based on the multi-order expectation-covariance collaborative aggregation mechanism, multi-order statistical feature extraction and aggregation are performed on the fused frequency domain features to obtain frequency domain spectral features; Step 6: Input the temporal motion features and frequency spectral features into the joint decision space, perform feature fusion and behavior classification, and output the human behavior recognition results.

2. The human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 1, characterized in that, Step 1 includes the following sub-steps: S11: Configure the sampling clock for the distributed wearable sensor network, to 9-axis rigid inertial pose data were acquired synchronously at high frequency. and 2-channel compliant deformation data This forms the original observation matrix. ; S12: Construct a time-series sliding window operator to perform time-series slicing on the original observation matrix and generate a standardized spatiotemporal tensor. The calculation formula for the standardized spatiotemporal tensor is as follows: ; in, Let τ be the τ-th normalized spacetime tensor; This indicates the zero-mean-unit-variance normalization operation performed on each independent channel; This represents the multi-source heterogeneous raw observation matrix obtained synchronously by wearable sensor networks; The index represents the starting time step index of the τth time series sliding window on the time axis corresponding to the original observation matrix; L represents the fixed length of the preset time series sliding window; and T represents the matrix transpose operator.

3. The human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 1, characterized in that, Step 2 includes the following sub-steps: S21: Standardize the spacetime tensor Decoupled into an inertial submanifold in the orthogonal dimensions of the channel. and compliant manifold ; S22: Construct an initial feature projection layer with an asymmetric receptive field, respectively for the inertial submanifold and compliant manifold Perform convolutional manifold mapping; where the inertial submanifold Using high-dimensional mapping tensors Mapped to dimension Latent feature space; compliant submanifold Using low-dimensional mapping tensors Mapped to dimension The latent feature space; The mapping dimension parameter satisfies the inequality > .

4. The human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 1, characterized in that, In step 3, the computation process of the multi-granularity expanded temporally separable operator includes: S31: For the input original feature tensor, construct three parallel temporal topology branches, namely the local transient motion capture branch, the long-range temporal dependency branch, and the extreme value significance preservation branch. S32: Temporal feature extraction for the corresponding dimension is completed through three parallel branches, and the output features of each branch are obtained; S33: The output features of the three branches are aggregated by an adaptive channel activation operator with a cross-layer residual structure to obtain the final temporal motion features.

5. The human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 4, characterized in that, The local transient motion capture branch constructs a difference compensation sequence by extracting the first-order time derivative of the feature tensor. After merging this sequence with the original features, it is fed into the depthwise separable convolution module for processing to obtain the output features of this branch. The calculation formula is as follows: ; In the formula: This represents the output characteristics of the local transient motion capture branch; This represents the Swish activation function; This indicates a one-dimensional batch normalization operation; This represents a 1×1 one-dimensional pointwise convolution operation; This represents a one-dimensional depthwise convolution operation with kernel size K=3 and dilation rate d=1. ⊕ represents the original feature tensor of the input; ⊕ represents the summation and compensation of features in the channel dimension; λ is the differential compensation weight that the network can learn; · represents the scalar multiplication operation; This represents the discrete first-order difference derivative sequence calculated over the time dimension; The long-range temporal-dependent branch generates a local energy envelope mask by calculating the squared magnitude of the signal within a local temporal window. It then performs nonlinear residual modulation on the result of a one-dimensional depthwise convolution with a kernel size of 5 and an expansion rate of 2 to obtain the output features of this branch. The calculation formula for this branch is as follows: ; In the formula: This represents the output characteristics of long-term time-dependent branches; This represents the element-wise product of Hadamah; This represents a one-dimensional dilated depthwise convolution operation with kernel size K=5 and dilation rate d=2; Sigmoid This represents the Sigmoid activation gate function; This represents a one-dimensional average pooling operation with a kernel size of 5. The extreme value significance preservation branch uses a local signal-to-noise ratio self-calibration relative peak operator to calculate the relative deviation between the maximum value and the mean within a local window, thus obtaining the output feature of this branch. The calculation formula is as follows: ; In the formula: The output characteristics of the branch that preserves the significance of extreme values; This represents a one-dimensional max-pooling operation with a kernel size of 3; tanh This represents the Tanh hyperbolic tangent activation function; This is an adaptive scaling factor; This represents a one-dimensional average pooling operation with a kernel size of 3. To prevent extremely small constant biases when dividing by zero; The temporal motion characteristics of the final aggregated output of the multi-granularity extended temporally separable operator The calculation formula is: ; In the formula: Channel cascading operations for characteristic manifolds; Convolution of the channel's dimension-reduced projection points; It is an adaptive channel activation operator, which includes average pooling, a fully connected layer, a ReLU activation function, a fully connected layer, and a Sigmoid activation function in sequence.

6. The human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 1, characterized in that, In step 4, the operation of adaptive wavelet frequency domain decomposition follows the following rules: The high-dimensional latent features corresponding to the inertial submanifold after mapping by the asymmetric projection network are input into the Haar wavelet for frequency domain decomposition, and the low-dimensional latent features corresponding to the compliant submanifold are input into the Dobersey wavelet for frequency domain decomposition. The features after the two frequency domain decompositions are respectively sent to the feature extraction module to complete the initial frequency domain feature extraction. The feature extraction module consists of a one-dimensional depthwise convolution with Kernel=3, a normalization layer, a ReLU activation function, a 1×1 one-dimensional pointwise convolution, a normalization layer, and a ReLU activation function.

7. The human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 6, characterized in that, In step 4, the cross-manifold cooperative gating mechanism includes: S41: Define the shallow compliant frequency domain features directly output by Dobessie wavelet decomposition as follows: The deep rigid frequency domain features extracted by the feature extraction module are: ; Applying shallow compliant frequency domain features A discrete nonlinear Tegor-Kaiser motion energy operator is introduced along the time dimension to calculate the transient deformation energy matrix. The corresponding calculation formula is: ; In the formula: Represents the transient deformation energy matrix; Indicates compliant frequency domain characteristics; This represents the element-wise product of Hadamah; and These represent shift operators that translate one time step to the left and one time step to the right along the time dimension, respectively. S42: The transient deformation energy matrix is ​​reconstructed using one-dimensional local moving average pooling and latent space features to generate a dynamic cooperative gating matrix that is strictly aligned with the time step. The corresponding calculation formula is: ; In the formula: Represents a dynamic collaborative gating matrix; This represents a local average pooling operation that preserves temporal resolution. and These are respectively the reduction and enhancement of the latent space. Orthogonal transformation operation; It is the ReLU activation function; Use the Sigmoid activation function; S43: The deep rigid frequency domain features are modulated step-by-step using a dynamic collaborative gating matrix, and then fused into frequency domain features via point convolution transformation and cross-layer residual connections. The corresponding calculation formula is: ; In the formula: Indicates the fused frequency domain characteristics; Represents rigid frequency domain characteristics; Indicates the use of feature reconstruction One-dimensional convolution operation; This indicates the sum of residuals across layers.

8. The human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 1, characterized in that, In step 5, the multi-order expectation-covariance collaborative aggregation mechanism follows the following: S51: For the fused frequency domain features of the input, calculate the first-order global expectation vector in the channel dimension; S52: Divide the fused frequency domain features into multiple orthogonal subspaces in the channel dimension, and calculate the regularized second-order unbiased covariance matrix of the features in each orthogonal subspace respectively; S53: Extract the upper triangular independent elements of the covariance matrix of each subspace based on the manifold semi-vectorization operator, and concatenate them with the first-order global expectation vector in the latent distribution space to output the final frequency domain spectral features.

9. A human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 8, characterized in that, The calculation process of the first-order global expectation vector follows: Let the input fused frequency domain features be a high-dimensional spectral state feature matrix. The first-order global expectation vector of the channel dimension is calculated. The calculation formula is: ; In the formula: Represents the first-order global expectation vector along the channel dimension; Represents the high-dimensional spectral state characteristic matrix; Indicates the length of the feature sequence; Indicates size is A column vector of all 1s; The calculation process of the regularized second-order unbiased covariance matrix is ​​as follows: The high-dimensional spectral characteristic matrix is... Divide the channel dimension into K orthogonal subspaces. For the κth group of features Calculate its regularized second-order unbiased covariance matrix. The corresponding calculation formula is: ; In the formula: Indicates the first The regularized second-order unbiased covariance matrix of the group feature subspace; Indicates the first Characteristic matrices of a set of orthogonal subspaces; express The identity matrix of order 1; Indicates length is A vector of all 1s; Represents the matrix transpose operator; To maintain the positive definiteness of the matrix, a minimal perturbation factor; The dimension is The identity matrix; The final calculation formula for the frequency domain spectral characteristics is: ; In the formula: This represents the multidimensional spectral characterization of the final output; Represents the semi-vectorization operator of a manifold; Let represent the covariance matrix of each orthogonal subspace; This represents the cascading operation of eigenvectors in the latent space.

10. A human behavior recognition method based on multi-sensor time-frequency information enhancement according to claim 1, characterized in that, The joint decision space in step 6 includes, in sequence, a feature flattening layer, a feature fully connected mapping layer, a layer normalization layer, a nonlinear activation layer, a random deactivation layer, and an output classification layer. After feature flattening, temporal motion features and frequency spectral features are fused through fully connected mapping, and finally the human behavior recognition result is output through the output classification layer.

Citation Information

Patent Citations

  • Human body behavior recognition method based on inertial sensor

    CN111582361A

  • Main and auxiliary shaft hoisting visual identification method based on digital twinning

    CN122347611A

  • Methods and apparatus for preserving information between layers within a neural network

    US20200117975A1