Human body activity identification method and system for effectively capturing time-space relationships of variables in sensors and between sensors

Through technical means such as modal specific embedding and self-attention mechanism, the spatiotemporal relationship between variables within and between sensors is effectively captured, and the problem of low human activity recognition performance in the prior art is solved, achieving higher recognition accuracy and computing efficiency.

CN120164142APending Publication Date: 2025-06-17CHONGQING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510190553.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the spatiotemporal relationship between variables within and between sensors, resulting in a degradation of human activity recognition performance.

Method used

The sensor data is converted into high-dimensional data through modal specific embedding processing, local time feature extraction and cross-channel and cross-variable fusion, different sensor features are integrated with the self-attention mechanism, and finally input a fully connected linear classifier for activity recognition.

Benefits of technology

It improves data integration capabilities and activity recognition accuracy, enhances the generalization capabilities of the model, optimizes computing efficiency, and supports personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164142A_ABST
    Figure CN120164142A_ABST
Patent Text Reader

Abstract

The invention discloses a human body activity identification method and system capable of effectively capturing variable time-space relations in sensors and between the sensors. The method comprises the following steps: 1) carrying out cross-variable fusion on local cross-channel fusion features; 2) performing global time aggregation on the local cross-variable fusion features to obtain sensor feature data; 3) integrating different sensor feature data by using a self-attention mechanism to generate human body activity features; and 4) inputting the human body activity features into a full-connection linear classifier to obtain human body activity categories. The system comprises N wearable motion sensors, a data conversion module, a local time feature extraction module, a cross-channel fusion module, a cross-variable fusion module, a global time aggregation module, a human body activity feature extraction module and a human body activity classification module. By capturing the space-time relationship between the interior of the sensor and the sensors, the frame can provide more systematic information, so that the model is more accurate during analysis and prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human activity recognition, and specifically to a human activity recognition method and system that can effectively capture the spatio-temporal relationships of intra-sensor and inter-sensor variables. Background Art

[0002] Human activity recognition (HAR) methods can be classified into three categories: vision-based, environment-based, and wearable sensor-based. Vision-based HAR relies on video data but is limited by lighting conditions and camera coverage. Environment-based HAR uses environmental sensors such as sound or WiFi, but it is restricted in terms of location and often performs poorly in terms of accuracy.

[0003] However, WHAR (wearable human activity recognition) involves attaching sensors directly to the human body. This method provides continuous and direct motion measurement, is less affected by environmental factors, and has strong adaptability in different scenarios. It also benefits from the use of off-the-shelf devices such as smartphones and watches, making it practical in the fields of healthcare and sports. WHAR can be classified as a multivariate time series classification problem. The multivariate nature comes from the modalities and axes of different sensors. A typical motion sensor may include a 3-axis accelerometer, a 3-axis gyroscope, and a 3-axis magnetometer, generating a time series data sample containing 9 variables.

[0004] The first key challenge of WHAR lies in how to effectively capture the local and global temporal features of multiple variables in each sensor ( Figure 1 intra-sensor modal variables in

[0005] ), while handling the complex relationships between these diverse variables. Previous SOTA (state-of-the-art) models such as DeepConvLSTM and AttendandDiscriminate use shared convolutional kernels to capture the local temporal patterns in each variable, while LSTM is used to integrate the variables and capture the global temporal dependencies. Recent research has used Conv1D multimodal fusion before temporal extraction, which may lead to the loss of temporal information in each variable. These methods often struggle to retain the inherent high-level features of each variable and perform inadequately in modeling complex cross-variable interactions, which may limit their full utilization of the multimodal intra-sensor data of the method and result in a decline in recognition performance. Figure 1(inter-modal variables between sensors). For example, during running, there is a strong correlation between the forward swing of one arm and the forward movement of the opposite leg. Similarly, recognizing the relationship between the stationary hand and the moving leg during cycling can significantly improve activity recognition performance. However, the above methods usually fuse all variables of all sensors indiscriminately using dense layers or RNN layers. This simple fusion method cannot capture the spatial relationship between sensors and lacks interpretability. To address this issue, DynamicWHAR separates sensors placed at different positions, performs multi-modal fusion and temporal feature extraction within each sensor respectively, and then uses a graph convolutional network (GCN) to capture the dynamic associations between sensors. However, GCN faces limitations due to its reliance on predefined graph structures, which restricts its ability to model dynamic relationships and is difficult to effectively scale in large graphs. Another significant limitation is that GCN calculates the associations between sensors symmetrically, potentially overlooking the directionality and asymmetry of the interactions between sensors. In addition, due to the message passing mechanism, GCN is less efficient in parallel computing. Summary of the Invention

[0006] The object of the present invention is to provide a human activity recognition method that can effectively capture the spatio-temporal relationships of intra-sensor and inter-sensor variables, including the following steps:

[0007] 1) Obtain human activity data of N wearable motion sensors, denoted as x = {x1, x2, …, x N} ∈ R N×M×L ; x n ∈ R M×L represents the data of the nth sensor at L time steps; M is the number of variables of each sensor;

[0008] 2) Perform modality-specific embedding processing on the human activity data of the wearable motion sensors to convert the human activity data of the sensors into high-dimensional sensing data;

[0009] 3) Perform local temporal feature extraction on each high-dimensional sensing data to obtain local temporal features;

[0010] 4) Perform cross-channel fusion on the local temporal features in sequence to obtain local cross-channel fusion features;

[0011] 5) Perform cross-variable fusion on the local cross-channel fusion features to obtain local cross-variable fusion features;

[0012] 6) Perform global temporal aggregation on the local cross-variable fusion features to obtain sensor feature data;

[0013] 7) Use the self-attention mechanism to integrate different sensor feature data to generate human activity features;

[0014] 8) Input the human activity features into the fully-connected linear classifier to obtain the human activity category.

[0015] Further, the wearable motion sensor includes accelerometers, gyroscopes, and magnetometers worn on various parts of the human body. The human body parts include the wrist, waist, thigh, and head.

[0016] Further, in step 2), the high-dimensional sensing data is as follows:

[0017] X emb = Conv1D(x n , P, S, D)(1)

[0018] In the formula, P is the convolutional kernel size, S is the stride, and D is the number of output channels; X emb ∈R N×M×D×T ; is the new length of the embedded sequence; Conv1D represents the 1D convolutional operation.

[0019] Further, in step 3), the steps for extracting local temporal features from each high-dimensional sensing data include:

[0020] 3.1) Reshape the high-dimensional sensing data X emb ∈R N×M×D×T to obtain the reshaped sensing data X emb ’ ∈ R N ×(M×D)×T ; N is the number of wearable motion sensors;

[0021] 2) Use the deep convolutional model to perform a deep convolutional operation on the reshaped sensing data X emb ’ to obtain the local temporal feature X dw , that is:

[0022] X dw = DWConv1D(X emb ’, K dw , G dw = M × D)(2)

[0023] Among them, K dw is the convolutional kernel size, G dw is the number of groups of the deep convolution; D is the number of output channels; DWConv1D represents the deep convolutional operation.

[0024] Further, the local cross-channel fusion feature is as follows:

[0025] X ccf = PWConv1D(X dw , G ccf = M)(3)

[0026] Among them, Gccf is the number of groups of point convolutions; X dw is the local temporal feature; PWConv1D is the point convolution operation; X ccf ∈R N ×(M×D)×T is the local cross-channel fusion feature.

[0027] Furthermore, the local cross-variable fusion feature is as follows:

[0028] X cvf = PWConv1D(X ccf , G cvf = D) (4)

[0029] where G cvf is the number of groups of point convolutions; X cvf ∈R N×(D×M)×T is the local cross-variable fusion feature.

[0030] Furthermore, in step 6), the steps of performing global temporal aggregation on the local cross-variable fusion feature include:

[0031] 6.1) Perform global average pooling on the local cross-channel fusion feature X ccf to obtain the average pooling feature X gap , that is:

[0032]

[0033] where n is the sensor index, d is the channel index, and t is the time step; X gap ∈R N×D×T ;

[0034] 6.2) Reshape the average pooling feature X gap into the feature X stack ∈R (N×D)×T ;

[0035] 6.3) Use the selective state space model to perform linear projection and convolution operations on the feature X stack to obtain the sensor feature data X mb , that is:

[0036]

[0037] where σ is the activation function, represents element-wise multiplication.

[0038] Furthermore, in step 7), the steps of integrating different sensor feature data using the self-attention mechanism include:

[0039] 7.1) Concatenate the sensor feature data X mbReshaped into feature X = {X1, X2, …, X N} ∈ R N×(D×T) ; X i ∈ R D×T represents the feature of the i-th sensor;

[0040] 7.2) Calculate the attention score A i,i′ , that is:

[0041]

[0042] where Q is the query function that projects the sensor into the query space, and K is the key function that projects the sensor into the key space; X i′ is the feature of the i'-th sensor;

[0043] 7.3) Based on the attention score A i,i′ , generate the self-attention feature map O i for each sensor, that is:

[0044]

[0045] where W is the linear embedding with learnable weights, and V is the value function that projects the sensor into the value space;

[0046] 7.4) Combine the feature map O i with the wearable motion sensor data x i through a residual connection to generate the human activity feature X csi .

[0047] Furthermore, the fully connected linear classifier is as follows:

[0048]

[0049] where, represents the output, and c is the number of human activity categories;

[0050] The loss function of the fully connected linear classifier is as follows:

[0051]

[0052] where y i , is the true label, is the predicted probability.

[0053] A system for a human activity recognition method based on the spatio-temporal relationship of variables within and between effective capture sensors, comprising N wearable motion sensors, a data conversion module, a local time feature extraction module, a cross-channel fusion module, a cross-variable fusion module, a global time aggregation module, a human activity feature extraction module, and a human activity classification module;

[0054] The N wearable motion sensors are attached to the human body for acquiring human activity data and transmitting it to the data conversion module;

[0055] The data conversion module performs modality-specific embedding processing on the human activity data, converts the human activity data of the sensors into high-dimensional sensing data, and transmits it to the local time feature extraction module;

[0056] The local time feature extraction module extracts local time features from each high-dimensional sensing data, obtains local time features, and transmits them to the cross-channel fusion module;

[0057] The cross-channel fusion module sequentially performs cross-channel fusion on the local time features, obtains local cross-channel fusion features, and transmits them to the cross-variable fusion module;

[0058] The cross-variable fusion module performs cross-variable fusion on the local cross-channel fusion features, obtains local cross-variable fusion features, and transmits them to the global time aggregation module;

[0059] The global time aggregation module performs global time aggregation on the local cross-variable fusion features, obtains sensor feature data, and transmits it to the human activity feature extraction module;

[0060] The human activity feature extraction module uses the self-attention mechanism to integrate different sensor feature data, generates human activity features, and transmits them to the human activity classification module;

[0061] The human activity classification module stores a fully connected linear classifier;

[0062] The human activity classification module inputs the human activity features into the fully connected linear classifier to obtain the human activity category.

[0063] The human activity category includes basic actions, daily behaviors, and special exercises;

[0064] The basic actions include walking, running, and jumping; the daily behaviors include sitting still, standing, drinking water, and going up and down stairs; the special exercises include cycling, weightlifting, and swimming, and the specific recognition types can be dynamically adjusted according to the application scenario.

[0065] The technical effect of the present invention is beyond doubt, and the beneficial effects of the present invention are as follows:

[0066] 1. Improve data integration capabilities: The new framework can effectively integrate data from different sensors, enhancing the comprehensive understanding of user activities. By capturing the spatio-temporal relationships within and between sensors, the framework can provide more comprehensive information, making the model more accurate in analysis and prediction.

[0067] 2. Enhance activity recognition accuracy: By deeply exploring spatio-temporal relationships, the framework can identify complex human activity patterns, thus improving the accuracy of activity recognition. This is crucial for application scenarios such as health monitoring, fitness tracking, and smart homes.

[0068] 3. Improve model generalization ability: The framework is designed to adapt to a variety of wearable devices and different types of sensor data, enabling the model to maintain high performance when facing new or unseen activities, thus possessing better generalization ability.

[0069] 4. Optimize computational efficiency: When capturing spatio-temporal relationships, the new framework can effectively reduce redundant calculations and improve processing speed. This is particularly important for real-time monitoring and feedback systems, helping to quickly respond to user needs.

[0070] 5. Facilitate multimodal fusion: The framework can effectively fuse data from different modalities (such as acceleration, gyroscope, and physiological signals, etc.), providing more comprehensive analysis and decision-making support by integrating information from different data sources.

[0071] 6. Support personalized services: By deeply understanding the user's activity patterns and habits, the framework can provide more targeted health management and exercise guidance programs for individuals, realizing personalized services and improving the user experience. Description of the Drawings

[0072] Figure 1 Variables within and between sensors in WHAR;

[0073] Figure 2 For the human activity recognition process. Detailed Implementation Modes

[0074] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject scope of the present invention is limited to the following embodiments. Without departing from the above-mentioned technical idea of the present invention, various substitutions and changes should be included within the protection scope of the present invention according to the common general knowledge and conventional means in the art.

[0075] Embodiment 1:

[0076] See Figures 1 to 2 , a method for human activity recognition that effectively captures the spatio-temporal relationships of variables within and between sensors, including the following steps:

[0077] 1) Obtain the human activity data of N wearable motion sensors, denoted as x = {x1, x2, …, x N} ∈ R N×M×L ; x n ∈ R M×L represents the data of the nth sensor at L time steps; M is the number of variables of each sensor;

[0078] 2) Perform modality-specific embedding processing on the human activity data of the wearable motion sensors to convert the human activity data of the sensors into high-dimensional sensing data;

[0079] 3) Extract local time features from each high-dimensional sensing data to obtain local time features;

[0080] 4) Perform cross-channel fusion on the local time features in sequence to obtain local cross-channel fusion features;

[0081] 5) Perform cross-variable fusion on the local cross-channel fusion features to obtain local cross-variable fusion features;

[0082] 6) Perform global time aggregation on the local cross-variable fusion features to obtain sensor feature data;

[0083] 7) Use the self-attention mechanism to integrate different sensor feature data to generate human activity features;

[0084] 8) Input the human activity features into a fully connected linear classifier to obtain the human activity categories.

[0085] The wearable motion sensors include accelerometers, gyroscopes, and magnetometers worn on various parts of the human body. The human body parts include the wrist, waist, thigh, and head.

[0086] In step 2), the high-dimensional sensing data is as follows:

[0087] X emb = Conv1D(x n , P, S, D)(1)

[0088] where P is the convolution kernel size, S is the stride, and D is the number of output channels; X emb ∈ R N×M×D×T ; is the new length of the embedded sequence; Conv1D represents the 1D convolution operation.

[0089] In step 3), the steps for extracting local time features from each high-dimensional sensing data include:

[0090] 3.1) Reshape the high-dimensional sensing data X emb ∈ R N×M×D×T to obtain the reshaped sensing data Xemb ’ ∈ R N ×(M×D)×T ; N is the number of wearable motion sensors;

[0091] 2) Use the deep convolutional model to perform a deep convolutional operation on the reshaped sensing data X emb ’ to obtain the local temporal feature X dw , that is:

[0092] X dw = DWConv1D(X emb ’, K dw , G dw = M × D) (2)

[0093] Among them, K dw is the convolutional kernel size, G dw is the number of groups of depth convolution; D is the number of output channels; DWConv1D represents the depth convolution operation.

[0094] The local cross-channel fusion feature is as follows:

[0095] X ccf = PWConv1D(X dw , G ccf = M) (3)

[0096] Among them, G ccf is the number of groups of point convolution; X dw is the local temporal feature; PWConv1D is the point convolution operation; X ccf ∈ R N ×(M×D)×T is the local cross-channel fusion feature.

[0097] The local cross-variable fusion feature is as follows:

[0098] X cvf = PWConv1D(X ccf , G cvf = D) (4)

[0099] Among them, G cvf is the number of groups of point convolution; X cvf ∈ R N×(D×M)×T is the local cross-variable fusion feature.

[0100] In step 6), the steps for global temporal aggregation of the local cross-variable fusion feature include:

[0101] 6.1) Perform global average pooling on the local cross-channel fusion feature X ccf to obtain the average pooling feature X gap , that is:

[0102]

[0103] where n is the sensor index, d is the channel index, and t is the time step; X gap ∈R N×D×T ;

[0104] 6.2) Reshape the average pooling feature X gap into feature X stack ∈R (N×D)×T ;

[0105] 6.3) Use the selective state space model to perform linear projection and convolution operations on feature X stack to obtain the sensor feature data X mb , that is:

[0106]

[0107]

[0108] where σ is the activation function, represents element-wise multiplication.

[0109] In step 7), the steps of integrating different sensor feature data using the self-attention mechanism include:

[0110] 7.1) Reshape the sensor feature data X mb into feature X = {X1, X2,..., X N} ∈ R N×(D×T) ; X i ∈ R D×T represents the feature of the i-th sensor;

[0111] 7.2) Calculate the attention score A i,i′ , that is:

[0112]

[0113] where Q is the query function that projects the sensor into the query space, and K is the key function that projects the sensor into the key space; X i′ is the feature of the i'-th sensor;

[0114] 7.3) Based on the attention score A i,i′ , generate the self-attention feature map O i for each sensor, that is:

[0115]

[0116] where W is the linear embedding with learnable weights, and V is the value function that projects the sensor into the value space;

[0117] 7.4) Combine the feature map o i with the wearable motion sensor data x i to generate the human activity feature X csi .

[0118] The fully connected linear classifier is as follows:

[0119]

[0120] where, represents the output, and C is the number of human activity categories;

[0121] The loss function of the fully connected linear classifier is as follows:

[0122]

[0123] where, y i , is the true label, is the predicted probability.

[0124] Example 2:

[0125] A human activity recognition method for effectively capturing the spatio-temporal relationships of variables within and between sensors, comprising the following steps:

[0126] 1) Obtain the human activity data of N wearable motion sensors, denoted as x = {x1, x2,..., x N} ∈ R N×M×L ; c n ∈ R M×L represents the data of the nth sensor at L time steps; M is the number of variables of each sensor;

[0127] 2) Perform modality-specific embedding processing on the human activity data of the wearable motion sensors to convert the human activity data of the sensors into high-dimensional sensing data;

[0128] 3) Extract local time features from each high-dimensional sensing data to obtain local time features;

[0129] 4) Perform cross-channel fusion on the local time features in sequence to obtain local cross-channel fusion features;

[0130] 5) Perform cross-variable fusion on the local cross-channel fusion features to obtain local cross-variable fusion features;

[0131] 6) Perform global time aggregation on the local cross-variable fusion features to obtain sensor feature data;

[0132] 7) Integrate the feature data of different sensors using the self-attention mechanism to generate human activity features;

[0133] 8) Input the human activity features into a fully connected linear classifier to obtain the human activity categories.

[0134] Example 3:

[0135] A human activity recognition method that effectively captures the spatio-temporal relationships of variables within and between sensors. The technical content is the same as that of Example 2. Further, the wearable motion sensors include accelerometers, gyroscopes, and magnetometers worn on various parts of the human body. The human body parts include the wrist, waist, thigh, and head.

[0136] Example 4:

[0137] A human activity recognition method that effectively captures the spatio-temporal relationships of variables within and between sensors. The technical content is the same as any one of Examples 2 - 3. Further, in step 2), the high-dimensional sensing data is as follows:

[0138] X emb = Conv1D(x n , P, S, D)(1)

[0139] In the formula, P is the convolution kernel size, S is the stride, D is the number of output channels; X emb ∈R N×M×D×T ; is the new length of the embedded sequence; Conv1D represents the 1D convolution operation.

[0140] Example 5:

[0141] A human activity recognition method that effectively captures the spatio-temporal relationships of variables within and between sensors. The technical content is the same as any one of Examples 2 - 4. Further, in step 3), the steps for extracting local time features from each high-dimensional sensing data include:

[0142] 3.1) Reshape the high-dimensional sensing data X emb ∈R N×M×D×T to obtain the reshaped sensing data X emb ’∈R N ×(M×D)×T ;

[0143] 2) Perform a depth convolution operation on the reshaped sensing data X emb ’ using a deep convolution model to obtain the local time feature X dw , that is:

[0144] X dw = DWConv1D(X emb ’, K dw , Gdw = M × D) (2)

[0145] Among them, K dw is the size of the convolution kernel, G dw is the number of groups of depthwise convolution; D is the number of output channels; DWConv1D represents the depthwise convolution operation.

[0146] Example 6:

[0147] A human activity recognition method for effectively capturing the spatio-temporal relationships of intra-sensor and inter-sensor variables, the technical content is the same as any one of Examples 2-5. Further, the local cross-channel fusion features are as follows:

[0148] X ccf = PWConv1D(X dw , G ccf = M) (3)

[0149] Among them, G ccf is the number of groups of pointwise convolution; X dw is the local temporal feature; PWConv1D is the pointwise convolution operation; X ccf ∈ R N ×(M×D)×T is the local cross-channel fusion feature.

[0150] Example 7:

[0151] A human activity recognition method for effectively capturing the spatio-temporal relationships of intra-sensor and inter-sensor variables, the technical content is the same as any one of Examples 2-6. Further, the local cross-variable fusion features are as follows:

[0152] X cvf = PWConv1D(X ccf , G cvf = D) (4)

[0153] Among them, G cvf is the number of groups of pointwise convolution; X cvf ∈ R N×(D×M)×T is the local cross-variable fusion feature.

[0154] Example 8:

[0155] A human activity recognition method for effectively capturing the spatio-temporal relationships of intra-sensor and inter-sensor variables, the technical content is the same as any one of Examples 2-7. Further, in step 6), the steps of performing global temporal aggregation on the local cross-variable fusion features include:

[0156] 6.1) Perform global average pooling on the local cross-channel fusion feature X ccf to obtain the average pooling feature X gap , that is:

[0157]

[0158] where n is the sensor index, d is the channel index, and t is the time step; X gap ∈R N×D×T ;

[0159] 6.2) Reshape the average pooling feature X gap into the feature X stack ∈R (N×D)×T ;

[0160] 6.3) Use the selective state space model to perform linear projection and convolution operations on the feature X stack to obtain the sensor feature data X mb , that is:

[0161]

[0162]

[0163] where σ is the activation function, represents element-wise multiplication.

[0164] Example 9:

[0165] A human activity recognition method for effectively capturing spatio-temporal relationships of intra-sensor and inter-sensor variables, the technical content is the same as any one of Examples 2-8. Further, in step 7), the steps of integrating different sensor feature data using the self-attention mechanism include:

[0166] 7.1) Reshape the sensor feature data X mb into the feature X = {X1, X2,..., X N}∈R N×(D×T) ; X i ∈R D×T represents the feature of the i-th sensor;

[0167] 7.2) Calculate the attention score A i,i′ , that is:

[0168]

[0169] where Q is the query function that projects the sensor into the query space, and K is the key function that projects the sensor into the key space; X i′ is the feature of the i'-th sensor;

[0170] 7.3) Based on the attention score A i,i′ , generate the self-attention feature map O i for each sensor, that is:

[0171]

[0172] Among them, W is a linear embedding with learnable weights, and V is a value function that projects the sensor into the value space;

[0173] 7.4) Combine the feature map O i with the wearable motion sensor data x i to generate the human activity feature X csi .

[0174] Example 10:

[0175] A human activity recognition method for effectively capturing the spatio-temporal relationships of variables within and between sensors, the technical content is the same as any one of Examples 2-9. Further, the fully connected linear classifier is as follows:

[0176]

[0177] Among them, represents the output, and C is the number of human activity categories;

[0178] The loss function of the fully connected linear classifier is as follows:

[0179] The predicted activity is compared with the true label (y) through the cross-entropy loss function:

[0180]

[0181] Among them, y i , is the true label, is the predicted probability.

[0182] Example 11:

[0183] A human activity recognition method for effectively capturing the spatio-temporal relationships of variables within and between sensors, the technical content is the same as any one of Examples 2-10. Further, the human activity categories include basic actions, daily behaviors, and special exercises;

[0184] The basic actions include walking, running, and jumping; the daily behaviors include sitting still, standing, drinking water, and going up and down stairs; the special exercises include cycling, weightlifting, and swimming, and the specific recognition types can be dynamically adjusted according to the application scenario.

[0185] Example 12:

[0186] A system for a human activity recognition method based on effectively capturing the spatio-temporal relationships of intra-sensor and inter-sensor variables according to any one of Embodiments 1-11, comprising N wearable motion sensors, a data conversion module, a local time feature extraction module, a cross-channel fusion module, a cross-variable fusion module, a global time aggregation module, a human activity feature extraction module, and a human activity classification module;

[0187] The N wearable motion sensors are attached to the human body for acquiring human activity data and transmitting it to the data conversion module;

[0188] The data conversion module performs modality-specific embedding processing on the human activity data, converting the human activity data of the sensors into high-dimensional sensing data and transmitting it to the local time feature extraction module;

[0189] The local time feature extraction module extracts local time features from each high-dimensional sensing data, obtaining local time features and transmitting them to the cross-channel fusion module;

[0190] The cross-channel fusion module sequentially performs cross-channel fusion on the local time features, obtaining local cross-channel fusion features and transmitting them to the cross-variable fusion module;

[0191] The cross-variable fusion module performs cross-variable fusion on the local cross-channel fusion features, obtaining local cross-variable fusion features and transmitting them to the global time aggregation module;

[0192] The global time aggregation module performs global time aggregation on the local cross-variable fusion features, obtaining sensor feature data and transmitting it to the human activity feature extraction module;

[0193] The human activity feature extraction module uses a self-attention mechanism to integrate different sensor feature data, generating human activity features and transmitting them to the human activity classification module;

[0194] The human activity classification module stores a fully-connected linear classifier;

[0195] The human activity classification module inputs the human activity features into the fully-connected linear classifier to obtain the human activity category.

[0196] The fully-connected linear classifier is trained through historical data, with the input being human activity features and the output being the human activity category.

[0197] The human activity categories include basic actions, daily behaviors, and special exercises;

[0198] The basic actions include walking, running, and jumping; the daily behaviors include sitting still, standing, drinking water, and going up and down stairs; the special exercises include cycling, weightlifting, and swimming, and the specific recognition types can be dynamically adjusted according to the application scenario.

[0199] Example 13:

[0200] A human activity recognition method that effectively captures the spatio-temporal relationships of variables within and between sensors, which includes a modal perception signal decomposition stage and a hierarchical interaction fusion stage.

[0201] Modal perception signal decomposition first separates different sensors and focuses on the modal interactions within the sensors. Inspired by modern pure convolutional time series structures, through a modality-specific embedding layer and a local time extraction module, deep convolution is used to independently extract the implicit high-level temporal features of the modal variables within each sensor. This ensures the independence at the sensor level and variable level, and can capture the detailed temporal features of each modality. In the hierarchical interaction fusion stage, the decomposed channel and modal variables are integrated through pointwise convolution to effectively capture the complex dependencies between these within-sensor modal variables. Through global time aggregation and cross-sensor interaction, using the Mamba attention module, the model integrates the features across the entire time dimension and dynamically captures the spatio-temporal correlations between sensors.

[0202] Specifically, given (N) wearable motion sensors, each sensor having (M) different modalities of variables (such as accelerometer, gyroscope, magnetometer data), the task is to recognize human activities from these multi-sensor data streams. The data can be represented as (X = {X1, X2, …, X N} ∈ R N×M×L ), where (X i ∈ R M×L ) represents the data of the (i)-th sensor at (L) time steps. For example, if each sensor records 3-axis accelerometer ((a x , a y , a z ))、gyroscope ((g x , g y , g z )) and magnetometer ((m x , m y , m z )) data, then (M = 9).

[0203] The goal is to develop a model (F(θ)) to predict the activity (F(X)) from the input (X). The goal of the model is to minimize the difference between the predicted activity and the true label (y). The ultimate goal is to ensure that the predicted label matches the actual activity label (y) as closely as possible.

[0204] In the processing of multivariate time series data, various convolution techniques can capture spatial and temporal correlations across channels. The Conv2D method uses shared convolutional kernels to learn cross-variable features across all channels. Although it is effective for cross-variable correlations, it cannot capture patterns specific to individual variables.

[0205] Inspired by depthwise separable convolutions, Depth-WiseConv1D applies independent 1D convolutional kernels to each channel separately. This method focuses on extracting features specific to individual variables and has fewer parameters and faster computational speed compared to Conv2D, thus reducing computational complexity.

[0206] After Depth-WiseConv1D, Point-WiseConv1D uses an operation with a kernel size of 1 to fuse cross-variable information. It integrates the features extracted by depthwise convolution into a unified representation while linearly transforming the depth dimension. Point-WiseConv1D further reduces the number of parameters and improves computational efficiency, optimizing the multivariate data processing process.

[0207] The goal is to independently extract the temporal features of each modal variable in each sensor, avoiding interference from other variables. This stage ensures independence at the sensor level and variable level. First, each sensor is isolated, and then modal-specific embedding (MSE) is performed on the variables of each sensor to transform the variables of each sensor into high-dimensional vectors. Next, local time extraction (LTE) is applied to independently convolve each channel of these high-dimensional vectors to extract temporal features from each channel. Therefore, this stage involves decomposition at the sensor level, variable level, and channel level, preserving the unique features of each modality.

[0208] MSE is designed to transform the original multi-sensor time series data into a high-dimensional representation, allowing the temporal dynamics of each modal variable to be captured independently before further processing.

[0209] Given input data from (N) wearable motion sensors, each sensor having (M) variables and (L) time steps, the goal is to independently embed each variable sequence into a high-dimensional space. The input data is represented as a tensor (X ∈ R N×M×L ), where (X n ∈ R M×L ) represents the data from the (n)-th sensor. The embedding process uses independent 1D convolution to transform each variable sequence:

[0210] X emb = Conv1D(X n , P, S, D),

[0211] where (P) is the convolution kernel size, (S) is the stride, and (D) is the number of output channels. This operation is performed along the time dimension of the variable sequence to obtain the embedded tensor (X emb ∈R N×M×D×T ), where is the new length of the embedded sequence, which helps to shorten the time length for more efficient subsequent calculations. The embedding process preserves the unique features of each variable by processing them independently and avoiding interference from other variables.

[0212] Local time extraction captures local time features while maintaining the independence of different modality variables. Each variable channel undergoes a convolution operation separately, preserving their unique features. Given the embedded tensor (X emb ∈R N×M×D×T ), where (N) represents the number of sensors, (M) represents the number of variables, (D) represents the number of channels, and (T) represents the time length, the tensor is first reshaped into (R N×(M×D)×T ). Depth convolution is used as local time extraction to capture the local time dependencies of each variable channel.

[0213] Depth convolution is defined as:

[0214] X dw =DWConv1D(X emb ,K dw ,G dw =M×D),

[0215] where (K dw ) is the convolution kernel size and (G dw ) is the number of groups of depth convolution. This operation is performed separately for each variable channel, ensuring the preservation of unique features.

[0216] In the hierarchical interaction fusion stage, the decomposed features are integrated at the channel, variable, and sensor levels to capture the spatio-temporal relationships within and between sensors. This method first combines the features within each sensor and then extends to cross-sensor interactions. The process is further refined by a global time aggregation module that integrates the features across the entire time dimension, enabling the model to effectively capture long-range dependencies.

[0217] Cross-channel fusion (CCF)

[0218] Cross-channel fusion combines the features of different sensor channels to capture the inter-channel dependencies within each variable. Starting from the tensor output (R N×(M×D)×T ) of depthwise separable convolution, the following pointwise convolution operation is performed:

[0219] X ccf =PWConv1D(X dw ,G ccf= M),

[0220] where (G ccf ) is the number of groups for point convolution. This operation fuses the information between channels and then combines and restores the dimensions by reshaping and performing two point convolutions.

[0221] Cross-Variable Fusion (CVF)

[0222] After cross-channel fusion, the intra-channel relationships of each variable are integrated. However, the interaction between different modality variables has not been addressed. To capture the cross-variable dependencies within the same sensor, a method similar to CCF is adopted. The input tensor (X ccf ∈ R N×(M×D)×T ) is reshaped to (X ccf ∈ R N×(D×M)×T ), and the number of groups is changed to (D). Then point convolution is applied:

[0223] X cvf = PWConv1D(X ccf , G cvf = D),

[0224] where (G cvf ) is the number of groups for point convolution. This operation fuses the information between variables and captures the cross-variable dependencies within each sensor.

[0225] Global Time Aggregation (GTA)

[0226] The decomposition stage extracts the modality-specific local temporal features but fails to fully capture the temporal context of the entire time series. To address this issue, a Global Time Aggregation (GTA) module is introduced, which integrates the information of all time steps to capture long-range dependencies and the overall temporal dynamics.

[0227] First, global average pooling (GAP) is applied to reduce the spatial dimension of each feature map to a scalar by taking the average over the variable dimension (M). For the input tensor (X ccf ∈ R N×(D×M)×T ), the application of GAP is as follows:

[0228]

[0229] where (n) is the sensor index, (d) is the channel index, and (t) is the time step. The result is (X gap ∈ R N×D×T ), which is then reshaped to (X stack ∈ R (N×D)×T ).

[0230] Next, the MambaBlock processes (X stack ∈ R (N×D)×T)To capture critical temporal information. The MambaBlock has a selective state space model (SSM) that uses linear projection and convolution to extract local features and selectively retains or discards information. Its operation is as follows:

[0231]

[0232] where (σ) is the activation function, denotes element-wise multiplication. The MambaBlock utilizes the selective SSM to effectively capture long-range dependencies, ensuring comprehensive temporal feature representation for accurate human activity recognition.

[0233] Cross-Sensor Interaction (CSI)

[0234] After completing the fusion at the channel, variable, and time levels, the next step is to integrate data from different sensors. The second challenge is to manage the spatial correlation between sensors. Inspired by the self-attention mechanism, which effectively captures the relationships between tokens, this method is used to facilitate interaction between sensors.

[0235] Given the output tensor (X mb ∈ R (N×D)×T ) from the MambaBlock, reshape it to (X = {X1, X2, …, X N} ∈ R N×(D×T) ), where (X i ∈ R D×T ) represents the data of the (i)-th sensor. Each (X i ) serves as a token for the attention layer.

[0236] The self-attention mechanism calculates the response of each sensor by attending to the representations of all sensors. The normalized correlation between all sensor data pairs (X i ) and (X i′ ) is calculated using an embedded Gaussian function. The attention score (A i,i′ ) measures the correlation of the data from sensor (i′) in optimizing the representation of sensor (i) and is calculated as follows:

[0237]

[0238] where (Q) is the query function that projects the sensors into the query space and (K) is the key function that projects the sensors into the key space.

[0239] Then, these correlations are used to generate the self-attention feature map (O i ) for each sensor:

[0240]

[0241] Among them, (W) is a linear embedding with learnable weights, and (V) is a value function that projects the sensors into the value space. The feature map (O) is combined with the original sensor data through a residual connection to generate a refined feature representation (X csi ), enabling the adaptive integration or exclusion of relevant information.

[0242] By using the CSI module, the model captures the interactions between different sensors and encodes these correlations through self-attention weights. During the inference process, these learned correlations are used to enhance the prediction, providing a robust method for integrating information from multiple sensors.

[0243] Fully connected linear classifier

[0244] After completing the CSI module through the self-attention mechanism, the output tensor (X csi ∈R N×(D×T) ) is obtained. This tensor is then reshaped and classified through a fully connected (FC) layer:

[0245]

[0246] Here, represents the final output, where (C) is the number of human activity categories. The predicted activity is compared with the true label (y) through the cross-entropy loss function:

[0247]

[0248] where (y i ) is the true label, is the predicted probability of the (i)-th class.

[0249] Example 14:

[0250] Verification of a human activity recognition method that effectively captures the spatio-temporal relationships of intra-sensor and inter-sensor variables is as follows:

[0251] To verify the effectiveness and generality of the proposed model, experiments were conducted on three benchmark datasets widely recognized in the WHAR community. These datasets are known for their complexity and diversity.

[0252] Opportunity: This dataset involves data collected from 4 users during kitchen activities. The focus is on data collected by 5 IMUs placed on the back, upper right arm, lower right arm, upper left arm, and lower left arm. Each IMU provides data from a 3-axis accelerometer, gyroscope, and magnetometer, with a sampling frequency of 30 Hz.

[0253] Realdisp: This dataset collected data on 33 fitness activities performed by 17 users, using 9 IMUs placed on different parts of the body, including the arms, calves, thighs, and back. Each IMU provides data from a 3-axis accelerometer, gyroscope, and magnetometer, with a sampling frequency of 50Hz. Due to incomplete data, recordings from 10 users were used.

[0254] Skoda: This dataset captured data on a user performing car maintenance activities, using 10 accelerometers placed on both hands, with a sampling frequency of 98Hz.

[0255] Experiments were conducted using the PyTorch framework on an RTX4090 GPU. The Adam optimizer was used to optimize the model parameters. The learning rate for the Opportunity dataset was set to 0.001, and the learning rates for the Realdisp and Skoda datasets were set to 0.0001. The maximum number of training epochs was set to 80. The batch size was set to 64 for the Opportunity and Skoda datasets, and 128 for the Readisp dataset. The number of Mamba blocks and attention layers was both set to 1, and the number of attention heads was set to 8. The output dimension (D) of the MSE was set to 64.

[0256] For the Opportunity and Realdisp datasets, a method similar to leave-one-user cross-validation was adopted. In this method, the data of a certain user was used as the test set, and the data of the remaining users was used for training. This process was repeated until each user had been used as a test subject once, and the average performance of all iterations was reported. Since the Skoda dataset only contains data from one user, a holdout evaluation method was adopted, with 80% of the data used for training, 10% for validation, and the remaining 10% for testing, to evaluate the performance of the model on the Skoda dataset. The evaluation metrics were accuracy and macro-F1 score.

[0257] The latest models in the WHAR field are introduced here and compared with the proposed model. The baseline models include DeepConvLSTM and its variants: DeepConvLSTM, Att.Model, AttendandDiscriminate. GCN-based models: GraphConvLSTM, HAR-PBD, DynamicWHAR. Transformer-based models: IFConvTransformer. Mamba-based models: HARMamba.

[0258] Table 1 compares the performance of the method with other state-of-the-art models in terms of accuracy and F1-score. The model is consistently superior to previous methods on the Opportunity, Realdisp, and Skoda datasets, significantly outperforming DynamicWHAR and other early models in both metrics.

[0259]

[0260]

[0261] Table 1 Comparison of Model Results

[0262] In WHAR applications, due to the limitations of embedded devices, it is crucial to manage computational resources and energy consumption while achieving accurate recognition. The figure evaluates efficiency by comparing model parameters, inference latency, and energy consumption. The FLOPs of DecomposeWHAR, HARMamba, and DynamicWHAR are all below 600M, while those of other models exceed 6000M. The model achieves comparable computational efficiency to the best models through depthwise separable convolutions, reducing parameters and accelerating the calculation speed. In addition, the selection mechanism of the Mamba block, the hardware-aware algorithm, and the efficient self-attention layer further enhance the overall efficiency.

Claims

1. A human activity recognition method that effectively captures the temporal and spatial relationships of intra-sensor and inter-sensor variables, characterized in that: The following steps are involved: 1) Obtain human activity data from N wearable motion sensors, denoted as x = {x1, x2, …, x N }∈R N×M×L ;x n ∈R M ×L Represents the data of the nth sensor at L time steps; M is the number of variables for each sensor. 2) Perform modality-specific embedding processing on the human activity data of wearable motion sensors to convert the human activity data of the sensors into high-dimensional sensing data; 3) Extract local time features from each high-dimensional sensor data to obtain local time features; 4) The local time features are sequentially fused across channels to obtain local cross-channel fusion features; 5) Perform cross-variable fusion on the local cross-channel fusion features to obtain local cross-variable fusion features; 6) Perform global temporal aggregation on the local cross-variable fusion features to obtain sensor feature data; 7) Use the self-attention mechanism to integrate the feature data of different sensors to generate human activity features; 8) Input the human activity features into the fully connected linear classifier to obtain the human activity category.

2. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: The wearable motion sensor includes an accelerometer, a gyroscope, and a magnetometer worn on various parts of the human body, including the wrist, waist, thigh, and head.

3. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: In step 2), the high-dimensional sensor data is as follows: X emb =Conv1D(x n ,P,S,D)(1) Where P is the convolution kernel size, S is the step size, and D is the number of output channels; X emb ∈R N×M×D×T ; T is the new length of the embedding sequence; Conv1D represents a 1D convolution operation.

4. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: In step 3), the step of extracting local time features from each high-dimensional sensor data includes: 3.1) For high-dimensional sensor data X emb ∈R N×M×D×T Reshape to obtain the reshaped sensor data X emb '∈R N×(M×D)×T ; N is the number of wearable motion sensors; 2) Use deep convolutional models to reshape sensor data X emb 'Perform deep convolution operation to obtain local time feature X dw ,Right now: X dw =DWConv1D(X emb ’,K dw ,G dw =M×D)(2) Among them, K dw is the convolution kernel size, G dw is the number of groups of depthwise convolution; D is the number of output channels; DWConv1D represents the depthwise convolution operation.

5. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: The local cross-channel fusion features are as follows: X ccf =PWConv1D(X dw ,G ccf =M)(3) Among them, G ccf is the number of point convolution groups; X dw is the local time feature; PWConv1D is the point convolution operation; X ccf ∈R N ×(M×D)×T It is the local cross-channel fusion feature.

6. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: The local cross-variable fusion features are as follows: X cvf =PWConv1D(X ccf ,G cvf =D)(4) Among them, G cvf is the number of point convolution groups; X cvf ∈R N×(D×M)×T It is a local cross-variable fusion feature.

7. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: In step 6), the step of performing global temporal aggregation on the local cross-variable fusion features includes: 6.1) Fusion of local cross-variable features X cvf Perform global average pooling to obtain the average pooling feature X gap ,Right now: Where n is the sensor index, d is the channel index, and t is the time step; X gap ∈R N×D×T ; 6.2) Average pooling feature X gap Reshape to feature X stack ∈R (N×D)×T ; 6.3) Using the Selective State Space Model to Model the Feature X stack Perform linear projection and convolution operations to obtain sensor feature data X mb ,Right now: Where σ is the activation function, Represents element-wise multiplication.

8. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: In step 7), the steps of integrating feature data of different sensors using the self-attention mechanism include: 7.1) Transform the sensor characteristic data X mb Reshape into features X = {X1, X2, …, X N }∈R N×(D×T) ;X i ∈R D×T represents the characteristics of the i-th sensor; 7.2) Calculate the attention score A i,i′ ,Right now: Where Q is the query function that projects the sensor into the query space, K is the key function that projects the sensor into the key space; X i′ is the characteristic of the i'th sensor; 7.3) Based on the attention score A i,i′ , generate the self-attention feature map O of each sensor i ,Right now: Where W is a linear embedding with learnable weights and V is the value function that projects the sensor into the value space; 7.4) Connect the feature map O through residual connection i Wearable motion sensor data i Combine to generate human activity feature X csi .

9. A human activity recognition method for effectively capturing the temporal and spatial relationships of variables within and between sensors according to claim 1, characterized in that: The fully connected linear classifier looks like this: in, represents the output, C is the number of human activity categories; Loss function for a fully connected linear classifier As shown below: Among them, y i , is the true label, is the predicted probability.

10. A system for human activity recognition based on the method for effectively capturing the temporal and spatial relationship of variables within and between sensors as claimed in any one of claims 1 to 9, characterized in that: It includes N wearable motion sensors, data conversion module, local time feature extraction module, cross-channel fusion module, cross-variable fusion module, global time aggregation module, human activity feature extraction module, and human activity classification module; The N wearable motion sensors are attached to the human body to obtain human activity data and transmit them to the data conversion module; The data conversion module performs modality-specific embedding processing on the human activity data, converts the human activity data of the sensor into high-dimensional sensing data, and transmits it to the local time feature extraction module; The local time feature extraction module extracts local time features from each high-dimensional sensor data to obtain local time features, and transmits them to the cross-channel fusion module; The cross-channel fusion module sequentially performs cross-channel fusion on the local time features to obtain local cross-channel fusion features, and transmits them to the cross-variable fusion module; The cross-variable fusion module performs cross-variable fusion on the local cross-channel fusion features to obtain local cross-variable fusion features, and transmits them to the global time aggregation module; The global time aggregation module performs global time aggregation on the local cross-variable fusion features to obtain sensor feature data, and transmits the data to the human activity feature extraction module; The human activity feature extraction module integrates the feature data of different sensors using a self-attention mechanism to generate human activity features and transmits them to the human activity classification module; The human activity classification module stores a fully connected linear classifier; The human activity classification module inputs the human activity features into a fully connected linear classifier to obtain a human activity category; Human activity categories include basic movements, daily behaviors, and special sports; Basic movements include walking, running, and jumping; daily behaviors include sitting, standing, drinking water, and going up and down stairs; special sports include cycling, weightlifting, and swimming. The specific recognition type can be dynamically adjusted according to the application scenario.

Citation Information

Cited By

  • Human body gravity center track estimation and balance capability detection method and device

    CN120616463A