A general human activity recognition method in multimodal environments based on contrastive learning

By introducing contrast learning and multimodal feature fusion technology into the deep learning human activity recognition method, the problems of insufficient feature extraction and insufficient model generalization capabilities in multimodal environments are solved, and higher recognition accuracy and adaptability are achieved.

CN119046792BActive Publication Date: 2025-06-06NANTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411128411.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-06-06
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

Existing deep learning based on human activity recognition methods has problems such as insufficient feature extraction and insufficient generalization capabilities in multimodal environments, especially in terms of data scarcity and adaptability to new data.

Method used

The general human activity recognition method in a multimodal environment based on contrast learning is adopted to enrich the data distribution through time and frequency domain data augmentation technology, and more valuable feature representations are extracted using the weighted multimodal feature fusion method, and the discernibility of the features is enhanced through contrast learning, and finally the features are sent to the classifier for activity recognition.

Benefits of technology

It significantly improves the generalization ability and accuracy of the model, maintains high recognition accuracy under different environments and conditions, and enhances the adaptability and robustness to new data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119046792B_ABST
    Figure CN119046792B_ABST
Patent Text Reader

Abstract

The present invention provides a universal human activity recognition method in a multimodal environment based on contrastive learning, which belongs to the field of deep learning technology. The technical problem of the generalization performance of the model when the test domain data is inaccessible is solved. The technical solution is: comprising the following steps: S1, data enhancement in the time domain and the frequency domain; S2, multimodal feature fusion; S3, contrastive learning. The beneficial effect of the present invention is: while improving the generalization ability of the unknown data domain, the accuracy of activity recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and specifically to a daily activity recognition method based on deep learning. The method can be applied to a human activity recognition system in a multimodal environment, and in particular to a general human activity recognition method in a multimodal environment based on contrastive learning. Background Art

[0002] HAR (Human Activity Recognition) is a research direction in computer science and artificial intelligence that focuses on developing algorithms and techniques to identify, classify, and understand human activities through data collected from various sensors. The goal of HAR is to automatically identify and interpret human actions and behaviors using data from sensors such as accelerometers, gyroscopes, magnetometers, and even cameras. Existing HAR methods are mainly divided into two categories: vision-based HAR and sensor-based HAR. In addition, due to the rapid development of wearable technology and the Internet of Things, sensor-based HAR has become a research focus. This approach has advantages such as reduced dependence on ambient light, enhanced privacy protection capabilities, and reduced computational overhead.

[0003] Sensor-based HAR can be divided into two categories: HAR based on environmental perception sensors uses sensors distributed in the environment to monitor and identify human activities, while HAR methods based on wearable sensors use data collected by sensors placed in different parts of the human body to identify activities. The increase in computing power in recent years has promoted the development of deep learning, thus giving HAR greater room for performance improvement.

[0004] Deep learning uses stacked convolutional network layers and activation functions to extract features from raw time series data and finally implements activity classification through a fully connected layer. Although the HAR method based on deep learning is superior to traditional machine learning technology, it requires a large amount of data for training, and the label cost of time series is too high, which affects the accuracy of the model. In addition, different people have different activity styles due to age, habits, etc., which reduces the performance of the model when processing new data.

[0005] Deep learning uses nonlinear information processing layers to extract features and classify activities. These layers form a hierarchical network of dense layers that extract features from raw sensor data and automatically classify activities. The HAR method based on deep learning outperforms traditional machine learning techniques and achieves high prediction and classification accuracy. However, some existing deep learning methods rarely consider the dependencies between the variable dimension and the time dimension at the same time, resulting in the extracted features not being able to perform activity recognition well.

[0006] Data augmentation technology generates new training samples by performing a series of transformations and expansions on the original data to increase the diversity and quantity of data, thereby improving the generalization ability and performance of the model. When data augmentation technology is applied to time series, the difficulty lies in how to effectively utilize its characteristics to design appropriate augmentation methods.

[0007] Multimodal fusion technology plays an increasingly important role in the field of human activity recognition based on wearable sensors. When processing sensor data, the difficulty of multimodal fusion technology lies in how to divide high-dimensional time series data into different modes, because the modes in multimodality usually refer to different data types.

[0008] Contrastive learning learns meaningful representations by comparing the similarities and differences between different samples in the data. When contrastive learning is used to process sequence data, the difficulty lies in how to select appropriate positive and negative samples, because existing contrastive learning usually only uses the enhanced view of the original data as positive samples and the rest as negative samples. Summary of the invention

[0009] The purpose of the present invention is to provide a universal human activity recognition method in a multimodal environment based on contrastive learning. First, data enhancement technology is used in the time domain and frequency domain to enrich the distribution of original data; then a weighted multimodal feature fusion method is used to extract more valuable feature representations from high-dimensional time series data; then a contrastive learning method is used to enhance the discriminability of features; finally, the extracted features are sent to a classifier for activity recognition.

[0010] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is specifically: a universal human activity recognition method in a multimodal environment based on contrastive learning named CLEAR, comprising the following steps:

[0011] S1. Data enhancement in time domain and frequency domain: In order to enrich the distribution of label data that can be used for training, data enhancement is performed from two perspectives: time domain and frequency domain. The data enhancement methods in time domain include rotation, arrangement, time warping, scaling, jittering and random sampling. The frequency domain part perturbs the spectrum of time series data by manipulating various parts of the frequency domain such as amplitude, phase, frequency, etc.

[0012] S2. Multimodal feature fusion: Considering the high dimensionality and multimodal characteristics of data collected by multiple sensors, we first use the convolutional neural network to extract the features of each modality separately; then introduce the attention mechanism and use the learnable weight parameters to adaptively fuse these features;

[0013] S3. Contrastive learning: To improve the distinguishability of the features finally learned, contrastive learning is used to optimize the feature representation by clustering samples of the same category in the feature space and dispersing samples of different categories, so that the similarities between samples of the same category are more significant, and the differences between samples of different categories are clearer.

[0014] Further, as a preferred technical solution of the present invention, the S1 specifically comprises the following steps:

[0015] First, the original time series data is divided into several samples using appropriate window size and overlap size, and then the corresponding data enhancement is performed; the time domain enhancement techniques include rotation, arrangement, time warping, scaling, jittering and random sampling; the rotation operation can change the position of the time series data on the time axis, thereby providing different time observation perspectives, which helps the model learn the features in different time periods and increase the model's generalization ability for time series data:

[0016] x' r =Rx (1)

[0017] x' r represents the result vector obtained after the rotation transformation, where R is the linear transformation matrix acting on the vector x;

[0018] The permutation operation provides a new way to organize data by rearranging the order of elements in time series data, thereby helping the model discover different patterns and structures in the data and improving the model's adaptability to different data arrangements:

[0019] x' p =(x π(1) ,x π(2) ,...x π(n) ) (2)

[0020] x' p represents the result vector after permutation transformation, x π(n) represents the nth element selected from vector x according to the permutation function π;

[0021] The time warping operation changes the time interval between data points in the time series data. It can speed up or slow down the sampling rate of the time series data, which helps the model learn at different time scales:

[0022] x t ' w =(x f(1) ,x f(2) ...x f(n) ) (3)

[0023] x t ' wrepresents the result vector after time warping transformation, x f ( n) represents the nth element of the result of applying function f to vector x;

[0024] Scaling operations increase or decrease the magnitude of the data by changing the amplitude or range of the time series data, thereby adapting the model to different data ranges:

[0025] x' s =(c 1 x 1 ,c 2 x 2 ...c n x n ) (4)

[0026] x' s represents the result vector after scaling transformation, c n is a constant used to scale the elements of vector x;

[0027] The jitter operation adds random noise to time series data to simulate uncertainty and randomness in the data, thereby helping the model better understand the true nature of the data and increasing the model's robustness to data noise:

[0028] x' j =(x 1 +ε 1 ,x 2 +ε 2 ...x n +ε n ) (5)

[0029] x' j represents the result vector after the jitter transformation, ε n represents the random noise added to the nth element;

[0030] Random sampling creates new samples by randomly selecting subsequences or data points in the time series data, which helps introduce more changes and diversity in the training process, thereby improving the robustness and generalization ability of the model and reducing the risk of overfitting:

[0031] x' rs =(x i1 ,x i2 ...x im ) (6)

[0032] x' rs represents the result vector obtained after random sampling transformation, x im represents the mth element selected from the vector x, where im represents the index of the selected element;

[0033] In order to enhance the time series in the frequency domain, the time series signal needs to be converted to the frequency domain through Fourier transform:

[0034]

[0035] Where u and v are indices, and H and W are the height and width respectively.

[0036] Use phase offset technology to adjust the phase of the spectrum. The original number is segmented according to a fixed period, and the start and end time of the sample is adjusted by appropriate phase offset, which helps to better understand each activity stage:

[0037]

[0038] Where R(x) and I(x) represent the real and imaginary parts of F(x) respectively; φ offset is the phase shift.

[0039] In the frequency domain, frequency represents periodicity, while amplitude represents the intensity of activity. We increase the data distribution by increasing and decreasing these parameters, as shown in equations (9) and (10):

[0040]

[0041] The parameters α, β, and γ are used to adjust the amplitude and frequency.

[0042] Since a small perturbation of the time series in the frequency domain may have a significant impact on the temporal pattern in the time domain, in order to ensure that the time series after the perturbation operation remains similar to the original sample in the time domain, we introduce a parameter μ to control the frequency perturbation, which represents the number of manipulated frequency components.

[0043] Further as a preferred technical solution of the present invention, S2 specifically includes the following steps:

[0044] A single convolutional subnetwork is used to extract features from each sensor. The subnetwork includes multiple stacked convolutional layers and a fully connected layer. Each convolutional layer consists of a convolution operation (Equation 11), a ReLU activation function (Equation 12), and a maximum pooling (Equation 13) operation. Maximum pooling is used to reduce the spatial size of the feature map while retaining important features. The ReLU function introduces nonlinearity to each neuron in the neural network, enabling the network to learn complex nonlinear mapping relationships. In addition, it has simple derivative calculations, which avoids the gradient vanishing problem during back propagation. (11)

[0046]

[0047] (X*C)[i, j] represents the element in the output feature map, h and w are the height and width of the convolution kernel;

[0048] ReLU: f(x) = max(0, x) (12)

[0049] Y i,j =max(X i:i+h-1,j:j+w-1 ) (13)

[0050] Each element Y in the output feature map Y i,j It is given by the maximum value of the input feature map X on the window [i,i+h-1]×[j,j+w-1];

[0051] The features extracted by the convolutional sub-network will be adaptively fused through the attention mechanism;

[0052] First, use a single-layer perceptron to transform the feature vector F k Operate to obtain the intermediate vector α k :

[0053] α k =tanh(W k F k +a k ) (14)

[0054] W k Representation and feature matrix F k The associated weight matrix, a k represents the bias term;

[0055] Then calculate the weight vector δ of each mode relative to the overall feature:

[0056]

[0057] Finally, the features of each modality are fused based on these weights to obtain the final feature vector V:

[0058]

[0059] Further, as a preferred technical solution of the present invention, S3 specifically includes the following steps:

[0060] Use the cosine of the angle between two vectors to evaluate their similarity:

[0061]

[0062] A i and B i are the components of vectors A and B respectively;

[0063] Then, we use similarity to select appropriate positive and negative samples. We need to pay special attention to samples that are close to or even overlap with other classes in the feature space. These samples occupy the boundary area in the feature space and depict the boundary between various class activities. To this end, we select positive samples from samples that have the same label as the anchor but with low similarity, and select negative samples from samples that have different labels but with high similarity to the anchor.

[0064]

[0065] v is an element in the feature set V, v i is the feature representation of the anchor point, v p is the positive sample set, v n is the negative sample set, threshold pos 、threshold neg It is the threshold for screening positive and negative samples; this improves the distinguishability of features while also reducing the difficulty of model convergence;

[0066] Then, we use the loss function L cl The loss function enhances the distinguishability of features by encouraging features from the same class to aggregate while separating features from different classes:

[0067]

[0068] The obtained features are then fed into a classifier for activity recognition. The classifier is a linear layer whose task is to map the features to a one-dimensional vector whose size matches the number of activities. The model is then trained using the cross entropy loss:

[0069]

[0070] Where N represents the number of samples, C represents the number of categories, represents the output prediction probability distribution, y represents the true category label, represents the predicted probability that the i-th sample belongs to class c, y ic is an indicator function. When the true label of the i-th sample belongs to class c, y ic is equal to 1, otherwise it is equal to 0; finally, the Adam gradient descent optimization algorithm is used to train the model; after obtaining the optimal parameters of the model, the model is applied to the test data set for verification.

[0071] Compared with the prior art, the present invention has the following beneficial effects:

[0072] (1) The present invention greatly enriches the distribution of source data through multi-field data enhancement technology, and this innovation brings significant technical advantages. In traditional activity recognition methods, the scarcity of training data often leads to poor performance of the model when faced with new data. Through multi-field data enhancement technology, the present invention can generate more diverse data samples in different fields and scenarios, thereby effectively alleviating the problem of data scarcity. This not only improves the diversity of the model in the training stage, but also enhances the adaptability and robustness of the model to new data in practical applications. Therefore, when faced with new data in different environments and under different conditions, the model can still maintain a high accuracy and stability, significantly improving the generalization ability of the model.

[0073] (2) The present invention integrates the attention mechanism with the convolutional neural network (CNN) subnetwork, so that the spatiotemporal dependencies in multimodal sensor signals can be effectively captured. The attention mechanism can dynamically assign importance weights to different time steps and different sensors, allowing the model to focus on more critical feature areas. Combined with the powerful feature extraction capability of CNN, this collaborative approach can effectively capture important feature information in high-dimensional time series, which is of great significance for accurate classification. This not only improves the accuracy of the model, but also enhances the model's ability to recognize complex data patterns, ensuring efficient recognition of human activities.

[0074] (3) The present invention introduces an enhanced contrastive learning method to improve the discriminability of features by strategically selecting contrast samples near classification boundaries. This innovation greatly improves the accuracy and generalization ability of the model technically. Traditional contrastive learning methods often only focus on global features and ignore key samples near classification boundaries. By selecting contrast samples in a targeted manner, the present invention can form clearer classification boundaries in the feature space, thereby significantly improving the model's ability to distinguish between different categories. This not only improves the accuracy of the model on training data, but also enhances its performance when facing unknown data in a real environment, allowing the model to more reliably complete various complex human activity recognition tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0076] Figure 1 This is an overall flow chart of a general human activity recognition method in a multimodal environment based on contrastive learning provided by the present invention.

[0077] Figure 2 Schematic diagram comparing the test results of the CLEAR model in the present invention and other methods on three public datasets.

[0078] Figure 3 Schematic diagram of the classification confusion matrix after testing on the DSADS dataset in the present invention.

[0079] Figure 4 This is a schematic diagram of the classification confusion matrix after testing on the USC-HAD dataset in the present invention.

[0080] Figure 5 Schematic diagram of the change of test results when the training data on the DSADS data set in the present invention changes.

[0081] Figure 6 Schematic diagram of the change of test results when the training data on the USC-HAD dataset in the present invention changes. DETAILED DESCRIPTION

[0082] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Of course, the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0083] Example 1

[0084] See also Figures 1 to 4 This embodiment proposes a universal human activity recognition method in a multimodal environment based on contrastive learning. In this embodiment, the experiment of the present invention is carried out on pytorch, and the RTX-3070ti GPU is used for model training and testing. During the experiment, the model is trained and evaluated on the DSADS public dataset. This dataset is used for comparison with other methods.

[0085] A general human activity recognition method in a multimodal environment based on contrastive learning comprises the following steps:

[0086] S1. Data enhancement in time domain and frequency domain: In order to enrich the distribution of labeled data available for training, we use two methods:

[0087] The data enhancement methods in the time domain include rotation, arrangement, time warping, scaling, jittering and

[0088] Random Sampling

[0089] ; The frequency domain part perturbs the spectrum of the time series data by manipulating the frequency components and their amplitudes, such as adding or deleting them; S2, multimodal feature fusion: Considering the high dimensionality and multimodal characteristics of multi-sensor data, we first use the convolutional neural network

[0090] The network extracts the features of each modality respectively; then the attention mechanism is introduced to adaptively fuse these features using learnable weight parameters;

[0091] S3. Contrastive learning: To improve the distinguishability of the features learned, contrastive learning is used to separate samples of the same category in the features.

[0092] The feature representation is optimized by clustering samples in the feature space and dispersing samples of different categories, making the similarities between samples of the same category more significant and the differences between samples of different categories more clear.

[0093] S1 specifically includes the following steps: First, the original time series data is divided into several

[0094] Take a sample and then perform corresponding data enhancement; the time domain enhancement techniques include rotation, permutation, time warping, scaling, jittering and random sampling;

[0095] The rotation operation can change the position of time series data on the time axis, thereby providing different time observation perspectives, helping the model learn the features of different time periods, and increasing the model's generalization ability for time series data:

[0096] x' r =Rx (1)

[0097] x' r represents the result vector obtained after the rotation transformation, where R is the linear transformation matrix acting on the vector x;

[0098] The permutation operation provides a new way to organize data by rearranging the order of elements in time series data, thereby helping the model discover different patterns and structures in the data and improving the model's adaptability to different data arrangements:

[0099] x' p =(x π(1) ,x π(2) ,...x π(n) ) (2)

[0100] x' p represents the result vector after permutation transformation, x π(n) represents the nth element selected from vector x according to the permutation function π;

[0101] The time warping operation changes the time interval between data points in the time series data. It can speed up or slow down the sampling rate of the time series data, which helps the model learn at different time scales:

[0102] x t ' w =(x f(1) ,x f(2) ...x f(n) ) (3)

[0103] x t ' w represents the result vector after time warping transformation, x f(n) represents the nth element of the result of applying function f to vector x;

[0104] Scaling operations increase or decrease the magnitude of the data by changing the amplitude or range of the time series data, thereby adapting the model to different data ranges:

[0105] x' s =(c 1 x 1 ,c 2 x 2 ...c n x n ) (4)

[0106] x' s represents the result vector after scaling transformation, c n is a constant used to scale the elements of vector x;

[0107] The jitter operation adds random noise to time series data to simulate uncertainty and randomness in the data, thereby helping the model better understand the true nature of the data and increasing the model's robustness to data noise:

[0108] x' j =(x 1 +ε 1 ,x 2 +ε 2 ...x n +ε n ) (5)

[0109] x' j represents the result vector obtained after the jitter transformation, ε n represents the random noise added to the nth element;

[0110] Random sampling creates new samples by randomly selecting subsequences or data points in the time series data, which helps introduce more changes and diversity in the training process, thereby improving the robustness and generalization ability of the model and reducing the risk of overfitting:

[0111] x' rs =(x i1 ,x i2 ...x im ) (6)

[0112] x' rs represents the result vector obtained after random sampling transformation, x imrepresents the mth element selected from the vector x, where im represents the index of the selected element;

[0113] In order to enhance the time series in the frequency domain, the time series signal needs to be converted to the frequency domain through Fourier transform:

[0114]

[0115] Where u and v are indices, and H and W are the height and width respectively.

[0116] We use phase offset technology to adjust the phase of the spectrum. The original number is segmented according to a fixed period, and the start and end times of the samples are adjusted by appropriate phase offset, which helps to understand each activity stage more deeply:

[0117]

[0118] Where R(x) and I(x) represent the real and imaginary parts of F(x) respectively; φ offset is the phase shift.

[0119] In the frequency domain, frequency represents periodicity, while amplitude represents the intensity of activity. We increase the data distribution by increasing and decreasing these parameters, as shown in equations (9) and (10):

[0120]

[0121] The parameters α, β, and γ are used to adjust the amplitude and frequency.

[0122] Since a small perturbation of the time series in the frequency domain may have a significant impact on the temporal pattern in the time domain, in order to ensure that the time series after the perturbation operation remains similar to the original sample in the time domain, we introduce a parameter μ to control the frequency perturbation, which represents the number of manipulated frequency components.

[0123] S2 specifically includes the following steps: extracting features from each sensor using a single convolutional subnetwork, the subnetwork includes multiple stacked convolutional layers and a fully connected layer, each convolutional layer consists of a convolution operation (Formula 10), a ReLU activation function (Formula 11), and a maximum pooling (Formula 12) operation; the maximum pooling is used to reduce the spatial size of the feature map while retaining important features, and the ReLU function introduces nonlinearity on each neuron of the neural network, so that the network can learn complex nonlinear mapping relationships. In addition, it has a simple derivative calculation, which avoids the gradient vanishing problem during back propagation;

[0124]

[0125] (X*C)[i, j] represents the element in the output feature map, h and w are the height and width of the convolution kernel;

[0126] ReLU: f(x) = max(0, x) (12)

[0127] Y i,j =max(X i:i+h-1,j:j+w-1 ) (13)

[0128] Each element Y in the output feature map Y i,j It is given by the maximum value of the input feature map X on the window [i,i+h-1]×[j,j+w-1]; the features extracted by the convolutional sub-network will be adaptively fused through the attention mechanism; first, a single-layer perceptron is used to ensemble the feature vector F k Perform the operation to obtain the intermediate vector α k :

[0129] α k =tanh(W k F k +a k ) (14)

[0130] W k Representation and feature matrix F k The associated weight matrix, a k represents the bias term;

[0131] Then calculate the weight vector δ of each mode relative to the overall feature:

[0132]

[0133] Finally, the features of each modality are fused based on these weights to obtain the final feature vector V:

[0134]

[0135] S3 specifically includes the following steps: Use the cosine value of the angle between two vectors to evaluate their similarity:

[0136]

[0137] A i and B i are the components of vectors A and B respectively;

[0138] Then, we use similarity to select appropriate positive and negative samples. We need to pay special attention to samples that are close to or even overlap with other classes in the feature space. These samples occupy the boundary area in the feature space and depict the boundary between various class activities. To this end, we select positive samples from samples that have the same label as the anchor but with low similarity, and select negative samples from samples that have different labels but with high similarity to the anchor.

[0139]

[0140] v is an element in the feature set V, v i is the feature representation of the anchor point, v p is the positive sample set, v n is the negative sample set, threshold pos 、threshold neg It is the threshold for screening positive and negative samples; this improves the distinguishability of features while also reducing the difficulty of model convergence;

[0141] Then, we use the loss function L cl The loss function enhances the distinguishability of features by encouraging features from the same class to aggregate while separating features from different classes:

[0142]

[0143] The obtained features are then fed into a classifier for activity recognition. The classifier is a linear layer whose task is to map the features to a one-dimensional vector whose size matches the number of activities. The model is then trained using the cross entropy loss:

[0144]

[0145] Where N represents the number of samples, C represents the number of categories, represents the output prediction probability distribution, y represents the true category label, represents the predicted probability that the i-th sample belongs to class c, y ic is an indicator function. When the true label of the i-th sample belongs to class c, y ic is equal to 1, otherwise it is equal to 0; finally, the Adam gradient descent optimization algorithm is used to train the model; after obtaining the optimal parameters of the model, the model is applied to the test data set for verification.

[0146] Comparative test:

[0147] DSADS dataset. This dataset includes sensor data collected from 8 individuals, evenly distributed between 4 females and 4 males, aged between 20 and 30 years old. The data was collected by placing five Xsens MTx units strategically on the torso, arms and legs of the participants. Each participant performed 19 activities ranging from daily activities to personalized exercise regimens, each lasting 5 minutes. Data acquisition was performed at a sampling frequency of 25 Hz.

[0148] The models that also use this dataset are selected for recognition accuracy comparison, namely RSC model, Mixup model, MLDG model, SimCLR model, Fish model, and DDlearn model. Figure 2 It can be seen that CLEAR has the best results. Figure 4 It can be seen that CLEAR is highly robust to changes in the amount of training data u, proving the generalization ability of the model.

[0149] Example 2

[0150] Based on Example 1, the same operating steps were implemented to test the USC-HAD dataset of Example 2, which involved 14 participants, including 7 males and 7 females, for data collection. The dataset covers 12 different behaviors. During the data collection process, a separate MotionNode device was securely placed in a standard-sized smartphone pocket, and the participant wore it on the right hip throughout the experiment. Each participant performed each behavior 5 times, each lasting about 24 seconds. Throughout the experiment, the MotionNode device recorded data from its three-axis accelerometer and three-axis gyroscope at a sampling frequency of 100 Hz. We used 60 percent of the data for training and 40 percent for validation and testing, respectively. Figure 5 This is the confusion matrix of the USC-HAD dataset test results. Figure 6 The change in the test accuracy of the model when the amount of training data changes.

[0151] The present invention provides a deep learning method for daily activity recognition, a training phase and an inference phase. In the training phase, firstly, the data enhancement technology in the time domain and frequency domain is used to enrich the diversity and quantity of the training data, which helps the model to be better generalized to various scenarios and environments; then, the multimodal features of the sensor data are used, and a multimodal fusion method is used to obtain a more representative feature representation, and by merging data from different sensors, richer and more comprehensive information is extracted, thereby improving the performance and generalization ability of the model; then, we apply contrastive learning to further optimize the feature learning process, by clustering samples of the same category together and separating samples of different categories, the discriminability of the feature representation is enhanced, thereby improving the accuracy and robustness of the model in the classification task; then we use an activity classifier to predict the label of the sample and use the accuracy as the evaluation criterion; finally, the optimal parameters obtained in the training phase are fixed, and the performance of the model is evaluated on the test set, and the experimental results prove that CLEAR has a higher accuracy rate than the existing models.

[0152] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A general human activity recognition method in a multimodal environment based on contrastive learning, characterized in that: The following steps are involved: S1. Data enhancement in time domain and frequency domain. Data enhancement is performed from two perspectives: time domain and frequency domain. The data enhancement methods in time domain include rotation, arrangement, time warping, scaling, jittering and random sampling. The frequency domain part disturbs the spectrum of time series data by manipulating various parts of the frequency domain. The step S1 comprises the following steps: First, the original time series data is divided into several samples using appropriate window size and overlap size, and then the corresponding data enhancement is performed; the time domain enhancement techniques include rotation, arrangement, time warping, scaling, jittering and random sampling; The rotation operation changes the position of the time series data on the time axis, thereby providing different time observation perspectives, helping the model learn the features of different time periods, and increasing the model's generalization ability for time series data: x' r =Rx (1) x' r Represents the result vector obtained after the rotation transformation, where R is the linear transformation matrix acting on the vector x; the permutation operation provides a new way of organizing data by rearranging the order of elements of the time series data, thereby helping the model discover different patterns and structures in the data and improving the model's adaptability to different data arrangements: x' p =(x π(1) ,x π(2) ,...x π(n) ) (2) x' p represents the result vector after permutation transformation, x π(n) represents the nth element selected from vector x according to the permutation function π; The time warping operation changes the time interval between data points in the time series data, speeding up or slowing down the sampling rate of the time series data, which helps the model learn at different time scales: x′ tw =(x f(1) ,x f(2) ...x f(n) ) (3) x′ tw represents the result vector after time warping transformation, x f(n) represents the nth element of the result of applying function f to vector x; Scaling operations increase or decrease the magnitude of the data by changing the amplitude or range of the time series data, thereby adapting the model to different data ranges: x' s =(c1x1,c2x2...c n x n ) (4) x' s represents the result vector after scaling transformation, c n is a constant used to scale the elements of vector x; the jitter operation adds random noise to the time series data, simulating the uncertainty and randomness in the data, thereby helping the model better understand the true nature of the data and increasing the model's robustness to data noise: x' j =(x1+ε1,x2+ε2...x n +e n ) (5) x' j represents the result vector after the jitter transformation, ε n Represents the random noise added to the nth element; Random sampling creates new samples by randomly selecting subsequences or data points in time series data, which helps introduce more changes and diversity in the training process, thereby improving the robustness and generalization ability of the model and reducing the risk of overfitting: x' rs =(x i1 ,x i2 ...x im ) (6) x' rs represents the result vector obtained after random sampling transformation, x im represents the mth element selected from the vector x, where im represents the index of the selected element; In order to enhance the time series in the frequency domain, the time series signal is converted to the frequency domain through Fourier transform: Where u and v are indices, H and W are height and width respectively; The phase offset technique is used to adjust the phase of the spectrum. The original number is segmented according to a fixed period. The start and end times of the samples are adjusted through appropriate phase offsets, which helps to gain a deeper understanding of each activity stage: Where R(x) and I(x) represent the real and imaginary parts of F(x) respectively; φ offset is the phase shift; In the frequency domain, frequency represents periodicity, while amplitude represents the intensity of activity. The data distribution is increased by increasing and decreasing these parameters, as shown in equations (9) and (10): Among them, parameters α, β, and γ are used to adjust the amplitude and frequency; Since a small perturbation of the time series in the frequency domain can have a significant impact on the temporal pattern in the time domain, in order to ensure that the time series after the perturbation operation remains similar to the original sample in the time domain, a parameter μ is introduced to control the frequency perturbation, which represents the number of manipulated frequency components; S2. Multimodal feature fusion: Considering the high dimensionality and multimodal characteristics of data collected by multiple sensors, we first use the convolutional neural network to extract the features of each modality separately; then introduce the attention mechanism and use the learnable weight parameters to adaptively fuse these features; S3. Contrastive learning: Contrastive learning is used to optimize feature representation by clustering samples of the same category in the feature space and dispersing samples of different categories, so that the similarities between samples of the same category are significant and the differences between samples of different categories are clear.

2. The general human activity recognition method in a multimodal environment based on contrastive learning according to claim 1 is characterized in that: The S2 specifically includes the following steps: A single convolutional subnetwork is used to extract features from each sensor. The subnetwork includes multiple stacked convolutional layers and a fully connected layer. Formula 11 represents each convolutional layer consisting of a convolution operation, Formula 12 represents a ReLU activation function, and Formula 13 represents a maximum pooling operation. Maximum pooling is used to reduce the spatial size of the feature map while retaining important features. The ReLU function introduces nonlinearity to each neuron in the neural network, enabling the network to learn complex nonlinear mapping relationships. In addition, it has a simple derivative calculation, which avoids the gradient vanishing problem during back propagation. (X*C)[i, j] represents the element in the output feature map, h and w are the height and width of the convolution kernel; ReLU: f(x) = max(0, x) (12) AND i,j =max(X i:i+h-1,j:j+w-1 ) (13) Each element Y in the output feature map Y i,j It is given by the maximum value of the input feature map X on the window [i,i+h-1]×[j,j+w-1]; The features extracted by the convolutional sub-network will be adaptively fused through the attention mechanism; First, use a single-layer perceptron to transform the feature vector F k Operate to obtain the intermediate vector α k : α k =tanh(W k F k +a k ) (14) W k Representation and feature matrix F k The associated weight matrix, a k represents the bias term; Then calculate the weight vector δ of each mode relative to the overall feature: Finally, the features of each modality are fused based on these weights to obtain the final feature vector V:

3. The universal human activity recognition method in a multimodal environment based on contrastive learning according to claim 2 is characterized in that: The S3 specifically includes the following steps: Use the cosine of the angle between two vectors to evaluate their similarity: A i and B i are the components of vectors A and B respectively; Use similarity to select appropriate positive and negative samples; samples that are close to or even overlap with other classes in the feature space need special attention; these samples occupy the boundary area in the feature space and depict the boundaries between various class activities; To do this, positive samples are selected from samples that have the same label as the anchor but with lower similarity, while negative samples are selected from samples that have different labels but with higher similarity to the anchor; v is an element in the feature set V, v i is the feature representation of the anchor point, v p is the positive sample set, v n is the negative sample set, threshold pos 、threshold neg It is the threshold for screening positive and negative samples; it improves the distinguishability of features while also reducing the difficulty of model convergence; Then, we use the loss function L cl The loss function enhances the distinguishability of features by encouraging features from the same class to aggregate and separating features from different classes: The obtained features are then fed into a classifier for activity recognition. The classifier is a linear layer whose task is to map the features to a one-dimensional vector whose size matches the number of activities. The model is then trained using the cross entropy loss: Where N represents the number of samples, C represents the number of categories, represents the output prediction probability distribution, y represents the true category label, represents the predicted probability that the i-th sample belongs to class c, y ic is an indicator function. When the true label of the i-th sample belongs to class c, y ic is equal to 1, otherwise it is equal to 0; finally, the Adam gradient descent optimization algorithm is used to train the model; after obtaining the optimal parameters of the model, the model is applied to the test data set for verification.

Citation Information

Patent Citations

  • Intelligent emotion recognition method based on data enhancement and cross-modal feature fusion

    CN116011457A

  • Method for multimodal emotion classification based on modal space assimilation and contrastive learning

    US20240119716A1