Activity Recognition Method and Device Based on Representation-Guided Multimodal Diffusion Model

By characterizing the guided multimodal diffusion model, using self-supervised models and diffusion models to generate multimodal activity data, the overfitting problem of multimodal data in wearable devices is solved, the accuracy of activity recognition is improved, and a wide range of health care and motion monitoring applications are provided.

CN119066475BActive Publication Date: 2025-07-18ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411556232.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-07-18
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

The existing wearable human activity recognition model is prone to overfitting problems on multimodal data. Traditional data augmentation methods destroy data information, insufficient diversity in GAN generation data, and the diffusion model-based method loses local details, resulting in low accuracy of activity recognition.

Method used

The characterization-guided multimodal diffusion model is adopted, and multimodal representation and global representation are extracted through the self-supervised model, and the diffusion model of the modal layer and the global layer is introduced. The multimodal activity data is generated using labeled data, and the activity recognition model is trained in combination with labeled data to improve accuracy.

Benefits of technology

Complex multimodal data with modal details and consistent activity information are generated, overfitting problems are solved, and the accuracy of activity recognition is improved, suitable for healthcare and motion monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119066475B_ABST
    Figure CN119066475B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for activity recognition based on a representation-guided multimodal diffusion model, including: obtaining original multimodal activity data and performing preprocessing to construct an unlabeled data set and a labeled data set; training a multimodal self-supervised model with the unlabeled data set so that it extracts multimodal representations containing complex patterns and global representations containing structural information from the data; introducing a modality layer guided by multimodal representations and a global layer guided by global representations in the diffusion model, and training the diffusion model with the unlabeled data set so that it generates multimodal activity data; training an activity recognition model using the labeled data set and the multimodal activity data generated by the trained diffusion model based on the labeled data set, which can solve the overfitting problem of the activity recognition model, improve the accuracy of activity recognition, and has broad application prospects in the fields of healthcare, sports monitoring, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wearable human activity recognition, and particularly relates to an activity recognition method and device based on a representation-guided multi-modal diffusion model. Background Art

[0002] Wearable Human Activity Recognition (WHAR) infers human activities using wearable sensor data and has extensive application values in scenarios such as sports monitoring, health support, and work assistance. Nowadays, more and more deep learning-based activity recognition models use multi-modal data, such as data from accelerometers, gyroscopes, and magnetometers, to obtain complementary activity information, thereby achieving higher accuracy than using single-modal data.

[0003] Since sensor data must be labeled during acquisition, obtaining a large amount of diverse labeled data is costly. Therefore, human activity recognition models are prone to overfitting problems, and in the multi-modal case, due to the increase in model parameters, the overfitting problem is particularly serious. The overfitting problem can usually be solved by using data augmentation methods to increase labeled data. In the field of wearable human activity recognition, existing data augmentation methods can be mainly divided into three types, namely traditional data augmentation methods, data augmentation methods based on Generative Adversarial Network (GAN), and data augmentation methods based on Diffusion Model.

[0004] Traditional data augmentation methods, such as rotation, permutation, scaling, etc., rely on human experience and may destroy key information contained in the data, such as temporal context information, sensor attitude, signal change intensity, etc.

[0005] The data augmentation method based on GAN solves the above problems, but the task of the discriminator in GAN is to judge the authenticity of samples and does not judge the diversity of the generated data, resulting in the generator not learning the complete data distribution. Therefore, GAN has the problem of insufficient diversity of the generated data.

[0006] The data augmentation method based on the diffusion model has an optimization objective of minimizing the negative log-likelihood loss, which can learn the complete data distribution without the problem of insufficient diversity in generated data. In the field of wearable human activity recognition, the data augmentation method based on the diffusion model utilizes statistical features such as mean and variance to guide the diffusion model, enabling the use of a large amount of unlabeled data to augment a small amount of labeled data. However, statistical features mainly describe the overall data, losing local detailed information and resulting in the loss of some complex patterns and structures in the augmented data. In addition, this method is designed only for unimodal data and will lose modal details when directly extended to multimodal data. Summary of the Invention

[0007] In view of the above, the object of the present invention is to provide a method and device for activity recognition based on a representation-guided multimodal diffusion model to solve the problem of how to use the diffusion model to improve the accuracy of activity recognition models for multimodal data activity recognition.

[0008] To achieve the above object of the invention, an activity recognition method based on a representation-guided multimodal diffusion model provided by an embodiment includes the following steps:

[0009] Data preprocessing stage: Obtain the original multimodal activity data, preprocess it, and construct an unlabeled data set and a labeled data set:

[0010] Self-supervised model training stage: Use the unlabeled data set to train a multimodal self-supervised model to extract multimodal representations containing complex patterns and global representations containing structural information from the data;

[0011] Diffusion model training stage: Introduce a modality layer guided by multimodal representations and a global layer guided by global representations in the diffusion model, and use the unlabeled data set to train the diffusion model to generate multimodal activity data;

[0012] Activity recognition model training stage: Use the labeled data set and the multimodal activity data generated by the trained diffusion model based on the labeled data set to train the activity recognition model, and the trained activity recognition model is used for activity recognition.

[0013] Preferably, using the unlabeled data set to train a multimodal self-supervised model to extract multimodal representations containing complex patterns and global representations containing structural information from the data includes:

[0014] Construct a self-supervised model including multiple independent modality encoders, attention networks, and multiple independent modality decoders;

[0015] Select unlabeled data samples from the unlabeled dataset, perform random masking to obtain masked data, and use multiple independent modality encoders to encode each modality in the masked data to obtain multi-modal representations containing complex patterns;

[0016] After concatenating the multi-modal representations, use the attention network to obtain a global representation containing structural information, and use multiple independent modality decoders to decode each modality in the global representation to obtain reconstructed data;

[0017] Construct a masking reconstruction loss based on the reconstructed data and the unlabeled data samples, and use the masking reconstruction loss to train the self-supervised model to update the parameters of the self-supervised model.

[0018] The masking reconstruction loss uses the mean squared error, which is expressed as:

[0019] ;

[0020] where, represents the masking reconstruction loss, and represent the unlabeled data sample and the reconstructed data respectively, represents the unlabeled data sample from the distribution that the unlabeled dataset follows , represents the expectation, represents the square of the L2 norm.

[0021] Preferably, introduce a modality layer guided by multi-modal representation and a global layer guided by global representation in the diffusion model, and use the unlabeled dataset to train the diffusion model to generate multi-modal activity data, including:

[0022] (1) Introduce multiple modality layers and global layers in the denoising network of the diffusion model. Among them, each modality layer and global layer includes a continuous number of downsampling modules and multiple upsampling modules. The global layer also includes a fusion module. Among them, the number of downsampling modules and upsampling modules in each modality layer is equal, and a skip connection is formed between the downsampling module and its corresponding upsampling module; the number of downsampling modules and upsampling modules in the global layer is also equal. The output of the last downsampling module is directly connected to the input of the corresponding upsampling module, and a skip connection is formed between the remaining downsampling modules and their corresponding upsampling modules;

[0023] (2) Select unlabeled data samples from the unlabeled dataset, and use the trained self-supervised model to extract multi-modal representations and global representations based on the unlabeled data samples. After taking the average of the multi-modal representations and global representations along a fixed length, the multi-modal representation and global representation used for guidance are obtained;

[0024] (3)For each time step, the following process is performed:

[0025] (3-1)Add noise to the input data to obtain the noisy data, and use two encoding networks to encode the time step respectively to obtain the first time step encoding and the second time step encoding;

[0026] (3-2)Use the noisy data as the input of the denoising network. Each modality layer processes the noisy data corresponding to one modality. Utilize multiple downsampling modules in each modality layer, and use the first time step encoding and the modality representation corresponding to each modality extracted from the multi-modal representation as guiding conditions to extract the modality features of each modality from the input noisy data;

[0027] (3-3)Use the fusion module of the global layer to fuse the modality features of all modalities to obtain the fused features. Utilize multiple downsampling modules and upsampling modules in the global layer, and use the second time step encoding and the global representation as guiding conditions to extract the global features from the input fused features;

[0028] (3-4)Utilize multiple upsampling modules in each modality layer, and use the first time step encoding and the modality representation corresponding to each modality extracted from the multi-modal representation as guiding conditions to extract the denoised data of each modality from the input global features. The denoised data of all modalities constitute the generated multi-modal activity data;

[0029] (4)Construct a reconstruction loss based on the unlabeled data samples and the generated multi-modal activity data, and use the reconstruction loss to train the diffusion model to update the denoising network parameters in the diffusion model.

[0030] Preferably, the same processing process adopted by each downsampling module is as follows: calculate the first scaling parameter and the first translation parameter based on the input guiding condition, and then use the first scaling parameter and the first translation parameter to scale and translate the input data or input features of each downsampling module, and then obtain the output features after convolution and activation processing; for each modality layer, the input data of the first downsampling module is the noisy data, the input features of other downsampling modules are the output features of the previous downsampling module in each modality layer, and the output of the last downsampling module is the modality feature; for the global layer, the input feature of the first downsampling module is the fused feature, and the input features of other downsampling modules are the output features of the previous downsampling module in the global layer;

[0031] The same processing procedure adopted by each upsampling module is as follows: calculate the second scaling parameter and the second translation parameter based on the input guiding condition, then scale and translate the input features of each upsampling module using the second scaling parameter and the second translation parameter, and then obtain the output features after deconvolution and activation processing; for the global layer, the input features of the first upsampling module are the output features of the last downsampling module in the global layer, the input features of other upsampling modules are the output features of the previous upsampling module in the global layer and the output features of the same-layer downsampling module through skip connection, and the output features of the last upsampling module are the global features; for each modality layer, the input features of the first upsampling module are the global features output by the last upsampling module in the global layer and the output features of the same-layer downsampling module through skip connection, the input features of other upsampling modules are the output features of the previous upsampling module in the modality layer and the output features of the same-layer downsampling module through skip connection, and the output of the last upsampling module is the denoised data.

[0032] Preferably, use the fusion module of the global layer to fuse the modality features of all modalities to obtain the fused features, including:

[0033] Concatenate the modality features of all modalities to obtain the concatenated features, and input the concatenated features into a multi-layer convolutional layer to obtain the modality weight matrix. At the same time, input the concatenated matrix into two convolutional layers respectively to obtain the first fusion parameter and the second fusion parameter. Use the first fusion parameter to process the modality weight matrix to obtain the modality weights, and use the second fusion parameter to process the modality weights to obtain the fused features.

[0034] Preferably, the reconstruction loss adopts the mean square error, expressed as:

[0035] ;

[0036] Where, represents the reconstruction loss, and respectively represent the unlabeled data samples and the generated multi-modal activity data, represents the unlabeled data samples from the distribution that the unlabeled data set follows , represents the noise during the noise addition process from the normal distribution with a mean of 0 and a variance of , represents the time step, represents the square of the L2 norm.

[0037] Preferably, use the labeled data set and the multi-modal activity data generated by the trained diffusion model based on the labeled data set to train the activity recognition model, including:

[0038] Select labeled data samples from the labeled dataset, and use the trained self-supervised model to generate multimodal representations and global representations based on the labeled data samples, and obtain the multimodal representations and global representations for guidance after taking the average along a fixed length;

[0039] Use the trained diffusion model to perform multi-step denoising and adding noise to the initial random noise, and introduce multimodal representations and global representations as guiding conditions during the denoising process to generate multimodal activity data;

[0040] Input the labeled data samples selected from the labeled dataset and the generated modal activity data into the activity recognition model to obtain classification results, and construct a classification loss based on the classification results and the true labels, and use the classification loss to train the activity recognition model to update the parameters of the activity recognition model.

[0041] Preferably, the classification loss adopts cross-entropy loss, which is expressed as:

[0042] ;

[0043] wherein, represents the classification loss, represents the true label, represents the classification result corresponding to the labeled data sample, represents the classification result corresponding to the generated modal activity data, represents calculating the cross-entropy loss, represents the expectation, represents the labeled data sample from the distribution that the labeled dataset follows ,[[]]END]] represents the noise in the noise addition process from a normal distribution with a mean of 0 and a variance of

[0044] To achieve the above object of the invention, an embodiment of the present invention further provides an activity recognition device based on a representation-guided multimodal diffusion model, including:

[0045] A data preprocessing unit, which is used to obtain the original multimodal activity data and perform preprocessing to construct an unlabeled dataset and a labeled dataset:

[0046] A self-supervised model training unit, which is used to train a multimodal self-supervised model using the unlabeled dataset so that it extracts multimodal representations containing complex patterns and global representations containing structural information from the data;

[0047] ​A diffusion model training unit, which is used to introduce a modality layer guided by multi-modal representations and a global layer guided by global representations into the diffusion model, and train the diffusion model using an unlabeled dataset to generate multi-modal activity data;

[0048] An activity recognition model training unit, which is used to train an activity recognition model using a labeled dataset and multi-modal activity data generated by the trained diffusion model based on the labeled dataset, and the trained activity recognition model is used for activity recognition.

[0049] Compared with the prior art, the beneficial effects of the present invention at least include:

[0050] The present invention introduces a modality layer guided by multi-modal representations and a global layer guided by global representations into the diffusion model, which can use a large amount of unlabeled data to augment a small amount of labeled data on complex multi-modal activity data, generate complex multi-modal activity data with modal details and consistent activity information, solve the overfitting problem of the activity recognition model, improve the accuracy of activity recognition, and has broad application prospects in the fields of healthcare, sports monitoring, etc. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 is a flowchart of the activity recognition method based on the representation-guided multi-modal diffusion model provided by the embodiment;

[0053] Figure 2 is a flowchart of the activity recognition method based on the representation-guided multi-modal diffusion model provided by the embodiment;

[0054] Figure 3 is a schematic structural diagram of the self-supervised model provided by the embodiment;

[0055] Figure 4 is a schematic structural diagram of the denoising network in the diffusion model provided by the embodiment;

[0056] Figure 5 is a schematic structural diagram of the upsampling module and the downsampling module provided by the embodiment;

[0057] Figure 6 is a schematic structural diagram of the activity recognition device based on the representation-guided multi-modal diffusion model provided by the embodiment. Detailed Embodiments

[0058] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0059] As Figure 1 and Figure 2 shown, an activity recognition method based on a representation-guided multimodal diffusion model provided by an embodiment includes the following steps:

[0060] S1. Data preprocessing stage: Obtain the original multimodal activity data, perform preprocessing, and construct an unlabeled data set and a labeled data set.

[0061] The specific process of the data preprocessing stage in the embodiment includes:

[0062] S1-1. Obtain the multimodal time-series data of different users performing different activities collected by wearable devices, which contains types of multimodal activity data collected by sensors.

[0063] S1-2. Perform outlier elimination processing and normalization processing on the collected multimodal activity data. Use a sliding window to divide the processed data to obtain data samples, and divide them into an unlabeled data set and a labeled data set according to whether they contain labels. The data samples of the unlabeled data set are , , represents the single-modal activity data corresponding to sensor , represents the number of sensors, represents the length of the sliding window, represents the number of channels of sensor , represents the data space, and the data samples of the labeled data set are , where represents the data sample with a true label , , represents the set of activity category labels.

[0064] Specifically, the preprocessing process is as follows:

[0065] a) Perform outlier elimination processing on the multimodal activity data, replace the values outside the normal range with the maximum or minimum value within the normal range, and replace the missing values with 0;

[0066] b) Perform normalization processing on the processed data, and the formula is as follows:

[0067] ;

[0068] Wherein is the original value, is the minimum value of the channel where the original value is located, is the maximum value of the channel where the original value is located, is the value after normalization;

[0069] c) Manually set the size of the time window according to experience , the overlapping degree of the sliding window is 50%, and the normalized data is divided by the sliding window to obtain data samples ;

[0070] d) Based on each data sample, an unlabeled data set is constructed, and the labeled data samples and their corresponding labels are used as samples to be data-augmented to form a labeled data set.

[0071] S1-3, batch the unlabeled data set and the labeled data set in a fixed size.

[0072] In this step, the total number of batches of the unlabeled data set is , and the total number of batches of the labeled data set is , and the formulas are as follows:

[0073] ;

[0074] ;

[0075] Wherein, represents the total number of samples in the unlabeled data set, represents the batch size of the unlabeled data set, which is manually set according to experience, represents the total number of samples in the labeled data set, represents the batch size of the labeled data set, which is manually set according to experience.

[0076] S2, self-supervised model training stage: Use the unlabeled data set to train a multi-modal self-supervised model so that it can extract multi-modal representations containing complex patterns and global representations containing structural information from the data.

[0077] In the embodiment, as Figure 3 shown, the self-supervised model includes multiple independent modal encoders, an attention network, and multiple independent modal decoders. Based on the self-supervised model, multi-modal representations containing complex patterns and global representations containing structural information can be extracted from the input multi-modal activity data. The specific training process of this self-supervised model is as follows:

[0078] S2-1. Select a batch of unlabeled samples from the unlabeled dataset, and for each unlabeled data sample in the selected batch , repeat step S2-2.

[0079] S2-2. After performing random masking on the unlabeled data sample to obtain masked data , calculate the reconstructed data through the self-supervised model;

[0080] During random masking, each value in the unlabeled data sample has a 20% probability of being replaced by 0 to achieve the masking process;

[0081] Adopt K independent modality encoders to encode each modality in the masked data to obtain a multimodal representation containing complex patterns k , where , represents the modality feature corresponding to a single modality , and k represents the feature dimension; among them, each modality encoder has the same structure and is composed of two one-dimensional convolutional layers, where the convolutional operation is performed along the time dimension; After concatenating the multimodal representation

[0082] , use the attention network to obtain a global representation containing structural information , specifically, concatenate all the single-modality representations in the multimodal representation to obtain the global representation ; ;

[0083] Use K independent modality decoders to decode each modality in the global representation to obtain the reconstructed data ; among them, each modality decoder has the same structure and is composed of a fully connected layer.

[0084] S2-3. For all unlabeled data samples in the current batch , calculate the masked reconstruction loss between the reconstructed data and the unlabeled data sample , and use the masked reconstruction loss to train the self-supervised model to update the self-supervised model parameters. When training, the goal is to minimize the masked reconstruction loss , Adopt the mean square error, expressed as:

[0085] ;

[0086] Among them, represents an unlabeled data sample from the distribution that the unlabeled data set follows , represents the expectation, represents the square of the L2 norm.

[0087] S2-4. When the training iteration number is not reached, that is, when the training is not completed, repeat steps S2-1 to S2-3 until all batches of the unlabeled data set participate in the training. When the specified training iteration number is reached, the training ends.

[0088] S3, Diffusion model training stage: Introduce a modality layer guided by multi-modal representation and a global layer guided by global representation in the diffusion model, and use the unlabeled data set to train the diffusion model to generate multi-modal activity data.

[0089] In the embodiment, as Figure 4 shown, introduce multiple modality layers and global layers in the denoising network of the diffusion model. Among them, each modality layer and global layer includes a plurality of consecutive downsampling modules and a plurality of upsampling modules. The global layer also includes a fusion module. Among them, the number of downsampling modules and upsampling modules in each modality layer is equal, and a skip connection is formed between the downsampling module and the corresponding upsampling module; the number of downsampling modules and upsampling modules in the global layer is also equal. The output of the last downsampling module is directly connected to the input of the corresponding upsampling module, and a skip connection is formed between the remaining downsampling modules and the corresponding upsampling modules. The specific training process of this diffusion model is as follows:

[0090] S3-1, Select a batch of unlabeled data samples from the unlabeled data set, and for each unlabeled data sample in the batch , repeat steps S3-2 to S3-3.

[0091] S3-2, Fix the self-supervised model, and use the trained self-supervised model to extract multi-modal representation and global representation based on the unlabeled data sample; specifically, input the unlabeled data sample into the trained self-supervised model to extract the multi-modal representation and global representation , and then the multi-modal representation and global representation are averaged along a fixed length (such as the sliding window size l ) to obtain the multi-modal representation and global representation used as guidance, where represents the modality k corresponding modality representation.

[0092] S3-3. Use the multimodal representation and the global representation as guiding conditions, and train the diffusion model on unlabeled data samples. Specifically, the training process of the diffusion model includes: randomly sampling time steps , where represents the maximum number of steps of the diffusion model. For each time step , perform the following process;

[0093] (a) Add noise to the input data to obtain the noise-added data. Use two encoding networks to encode the time steps respectively to obtain the first time step encoding and the second time step encoding;

[0094] Specifically, add noise to the input data to obtain the noise-added data , where represents the modality k at the time step t corresponding noise-added data. The specific input data and the random noise sampled from the normal distribution are added proportionally according to the time step to achieve the noise-adding process.

[0095] For each time step t , use two encoding networks to encode the time step t respectively to obtain the first time step encoding and the second time step encoding , where the encoding network can use a fully connected layer.

[0096] (b) Use the noise-added data as the input of the denoising network. Each modality layer processes the noise-added data k corresponding to one modality . Utilize multiple downsampling modules in each modality layer, and use the first time step encoding and the modality representation extracted from the multimodal representation k corresponding to each modality as guiding conditions. Extract the modality features of each modality k in the input noise-added data t at the time step , represents the length of the downsampled features. The modality features of all modalities of the noise-added data are .

[0097] For example Figure 5As shown, each downsampling module in each modality layer adopts the same processing procedure, specifically: based on the input guiding condition, the first-time step encoding and the modality representation calculate the first scaling parameter and the first translation parameter , for example, using the SiLU activation function (SiLU) and the fully connected layer (Linear) based on and calculate the first scaling parameter and the first translation parameter , then use and to scale and translate the input data or input features of each downsampling module, and then after processing by one-dimensional convolution (Conv1D) and the SiLU activation function (SiLU), the output features are obtained, expressed as:

[0098] ;

[0099] ;

[0100] Among them, represents the input data or input features for modality k. For the input data of the first downsampling module, it is the noisy data , and the input features of other downsampling modules are the output features of the previous downsampling module in each modality layer. The output of the last downsampling module is the modality feature k corresponding to modality , represents the output features of each downsampling module in each modality layer.

[0101] (c) Use the fusion module of the global layer to fuse the modality features of all modalities to obtain the fusion feature , use multiple downsampling modules and upsampling modules in the global layer, and use the second-time step encoding and the global representation as the guiding conditions to extract the global feature from the input fusion feature ;

[0102] In the embodiment, the specific process of fusing the modality features of all modalities to obtain the fusion feature is as follows: After concatenating the modality features of each modality for the modality features of all modalities, the concatenated feature is obtained, and the concatenated feature is input into the multi-layer convolutional layer to obtain the modality weight matrix , while splicing the matrix They are respectively input into two convolutional layers to obtain the first fusion parameter and the second fusion parameter . Using the first fusion parameter to process the modal weight matrix to obtain the modal weight . Using the second fusion parameter to process the modal weight to obtain the fused feature . The formula is:

[0103] ;

[0104] ;

[0105] As Figure 5 shown, each downsampling module in the global layer adopts the same processing process as the downsampling module in the modal layer. Specifically: based on the input guided conditional second time step encoding and the global representation calculate the first scaling parameter and the first translation parameter . For example, using the SiLU activation function (SiLU) and the fully connected layer (Linear) based on and to calculate and . Then, using and to scale and translate the input data or input features of each downsampling module, and then after processing by one-dimensional convolution (Conv1D) and the SiLU activation function (SiLU), the output features are obtained, expressed as:

[0106] ;

[0107] ;

[0108] Among them, represents the input data or input features of each downsampling module in the global layer. For the input data of the first downsampling module, it is the fused feature , and the input features of other downsampling modules are the output features of the previous downsampling module in the global layer, represents the output features of each downsampling module in the global layer.

[0109] As Figure 5 shown, the specific implementation process of each upsampling module in the global layer is basically the same as that of the downsampling module. The difference is that one-dimensional convolution replaces layer one-dimensional transposed convolution. The specific process is: based on the input guided conditional second time step encoding and the global representation Calculate the second scaling parameter and the second translation parameter , for example, using the SiLU activation function (SiLU) and the fully connected layer (Linear) based on and calculate and , then use and to scale and translate the input features of each upsampling module, and then after processing by one-dimensional transposed convolution (DeConv1D) and the SiLU activation function (SiLU), the output features are obtained, expressed as:

[0110] ;

[0111] ;

[0112] wherein, represents the input features of each upsampling module in the global layer, wherein the input features of the first upsampling module are the output features of the last downsampling module in the global layer, and the input features of other upsampling modules are the output features of the previous upsampling module in the global layer and the output features of the same-layer downsampling module through the skip connection, and the output features of the last upsampling module are the global features , represents the output features of each upsampling module in the global layer.

[0113] (d) Utilize multiple upsampling modules in each modality layer and encode with the first time step and each modality extracted from the multi-modal representation k corresponding modality representation as the guiding condition, extract the denoised data of each modality k from the input global features , and use the denoised data of all modalities as the generated multi-modal activity data.

[0114] As Figure 5 shown, the upsampling modules in each modality layer adopt the same processing procedure as the upsampling modules in the global layer, specifically: based on the input guiding condition first time step encoding and the modality representation calculate the second scaling parameter and the second translation parameter , for example, using the SiLU activation function (SiLU) and the fully connected layer (Linear) based on and Calculate and , and then use and to scale and translate the input features of each upsampling module, and then obtain the output features after being processed by one-dimensional transposed convolution (DeConv1D) and SiLU activation function (SiLU), expressed as:

[0115] ;

[0116] ;

[0117] wherein, represents the input features for modality k , and the input features of the first upsampling module are the global features output by the last upsampling module of the global layer and the output features of the downsampling module of the same layer through skip connection. The input features of other upsampling modules are the output features of the previous upsampling module in the modality layer and the output features of the downsampling module of the same layer through skip connection. The output of the last upsampling module is the denoised data of modality k , , represents the output features of each upsampling module for modality k .

[0118] For all unlabeled data samples in the current batch, based on the unlabeled data samples and the generated multi-modal activity data construct the reconstruction loss , and use the reconstruction loss to train the diffusion model to update the denoising network parameters in the diffusion model. When training, the goal is to minimize the reconstruction loss . Among them, the reconstruction loss adopts the mean square error, expressed as:

[0119] ;

[0120] wherein, represents the noise during the noise addition process coming from a normal distribution with a mean of 0 and a variance of .

[0121] S3-4. When the training iteration count is not reached, i.e., the training is not completed, repeat steps S3-1 to S3-3 until all batches of the unlabeled dataset have participated in the training. When the specified training iteration count is reached, the training ends. S4. Activity recognition model training phase: Use the labeled dataset and the multimodal activity data generated by the diffusion model based on the labeled dataset after training to train the activity recognition model. The trained activity recognition model is used for activity recognition.

[0122] The specific process of training the activity recognition model using the labeled dataset is as follows:

[0123] S4-1. Select a batch of labeled data samples from the labeled dataset, and for each labeled data sample in the batch , repeat steps S4-2 to S4-4.

[0124] S4-2. Fix the self-supervised model and use the trained self-supervised model to extract multimodal representations and global representations based on the labeled data samples; specifically, input the labeled data samples into the trained self-supervised model to extract multimodal representations and global representations through forward inference, and obtain the multimodal representations and global representations used as guidance after taking the average along the fixed length.

[0125] S4-3. Use the multimodal representations and global representations as guidance conditions to generate multimodal activity data using the diffusion model; specifically, use the trained diffusion model to perform multi-step denoising and adding noise on the initial random noise, and introduce the multimodal representations and global representations as guidance conditions to generate multimodal activity data . In each time step, after denoising through the denoising network and then adding noise, it is used as the input for the next time step, and so on, cycling through multiple time steps to generate multimodal activity data .

[0126] S4-4. Use the real data and the generated data to train the activity recognition model. Specifically, the labeled data samples selected from the labeled dataset and the generated modal activity data are input into the activity recognition model to obtain classification results and . For all labeled data samples in the current batch, based on the classification results and and the true labels construct a classification loss , and use the classification loss The activity recognition model is trained to update the parameters of the activity recognition model. During training, the classification loss is minimized, where the classification loss uses the cross-entropy loss, which is expressed as:

[0127] The classification loss uses the cross-entropy loss, which is expressed as:

[0128] ;

[0129] where represents calculating the cross-entropy loss, represents the labeled data samples from the distribution that the labeled data set follows , represents the noise of the noise addition process from the normal distribution with a mean of 0 and a variance of .

[0130] S4-5. When the training iteration number is not reached, that is, when the training is not completed, repeat steps S4-1 to S4-4 until all batches of the unlabeled data set participate in the training. When the specified training iteration number is reached, the training ends.

[0131] The activity recognition model after training is used for activity recognition. Specifically, during activity recognition, the multi-modal activity data to be recognized is input into the activity recognition model, and the recognized classification result is obtained through calculation.

[0132] As Figure 6 shown, the embodiment also provides an activity recognition device 60 based on a representation-guided multi-modal diffusion model, including: a data preprocessing unit 61, a self-supervised model training unit 62, a diffusion model training unit 63, and an activity recognition model training unit 64. Among them, the data preprocessing unit 61 is used to obtain the original multi-modal activity data, preprocess it, and construct an unlabeled data set and a labeled data set: the self-supervised model training unit 62 is used to train the multi-modal self-supervised model with the unlabeled data set so that it extracts multi-modal representations containing complex patterns and global representations containing structural information from the data; the diffusion model training unit 63 is used to introduce a modality layer guided by multi-modal representations and a global layer guided by global representations in the diffusion model, and train the diffusion model with the unlabeled data set so that it generates multi-modal activity data; the activity recognition model training unit 64 is used to train the activity recognition model using the labeled data set and the multi-modal activity data generated by the trained diffusion model based on the labeled data set. The trained activity recognition model is used for activity recognition.

[0133] It should be noted that when the activity recognition device based on the representation-guided multi-modal diffusion model provided in the above embodiments performs activity recognition, the above examples are given based on the division of each functional module. The above functions can be allocated to different functional units or modules according to needs, that is, the internal structure of the terminal or server is divided into different functional units or modules to complete all or part of the functions described above. In addition, the activity recognition device based on the representation-guided multi-modal diffusion model provided in the above embodiments and the embodiments of the activity recognition method based on the representation-guided multi-modal diffusion model belong to the same concept. For the specific implementation process, please refer to the embodiments of the activity recognition method based on the representation-guided multi-modal diffusion model, which will not be elaborated here.

[0134] In the above activity recognition method and device, by introducing the modality layer guided by multi-modal representation and the global layer guided by global representation into the diffusion model, a large amount of unlabeled data can be used to augment a small amount of labeled data on complex multi-modal data, generate complex multi-modal data with modality details and consistent activity information, solve the overfitting problem of the activity recognition model, improve the accuracy of activity recognition, and have broad application prospects in the fields of healthcare, sports monitoring, etc.

[0135] The above specific embodiments have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above are only the most preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An activity recognition method based on a representation-guided multi-modal diffusion model, characterized in that, It includes the following steps: Data preprocessing stage: Obtain the original multi-modal activity data of different users performing different activities collected by wearable devices, and after preprocessing, construct an unlabeled dataset and a labeled dataset: Self-supervised model training stage: Use the unlabeled dataset to train a multi-modal self-supervised model so that it can extract multi-modal representations containing complex patterns and global representations containing structural information from the data; Diffusion model training stage: Introduce a modality layer guided by multi-modal representations and a global layer guided by global representations into the denoising network of the diffusion model. Each modality layer and global layer include a continuous number of downsampling modules and upsampling modules. The global layer also includes a fusion module. Among them, the number of downsampling modules and upsampling modules in each modality layer is equal, and a skip connection is formed between the downsampling module and its corresponding upsampling module; the number of downsampling modules and upsampling modules in the global layer is also equal. The output of the last downsampling module is directly connected to the input of the corresponding upsampling module, and a skip connection is formed between the remaining downsampling modules and their corresponding upsampling modules; and use the unlabeled dataset to train the diffusion model to generate multi-modal activity data, including the following steps: (1) Select unlabeled data samples from the unlabeled dataset, and use the trained self-supervised model to extract multi-modal representations and global representations based on the unlabeled data samples. After taking the average value along a fixed length, the multi-modal representations and global representations used as guidance are obtained; (2) For each time step, the following process is carried out: (2-1) Add noise to the input data to obtain the noise-added data, and use two encoding networks to encode the time step respectively to obtain the first time step encoding and the second time step encoding; (2-2) Use the noise-added data as the input of the denoising network. Each modality layer processes the noise-added data corresponding to one modality, and use multiple downsampling modules in each modality layer, and use the first time step encoding and the modality representations corresponding to each modality extracted from the multi-modal representations as guidance conditions to extract the modality features of each modality from the input noise-added data; (2-3) Use the fusion module of the global layer to fuse the modality features of all modalities to obtain the fused features, and use multiple downsampling modules and upsampling modules in the global layer, and use the second time step encoding and the global representation as guidance conditions to extract the global features from the input fused features; (2-4) Use multiple upsampling modules in each modality layer, and use the first time step encoding and the modality representations corresponding to each modality extracted from the multi-modal representations as guidance conditions to extract the denoised data of each modality from the input global features. The denoised data of all modalities constitute the generated multi-modal activity data; (3) Construct a reconstruction loss based on the unlabeled data samples and the generated multi-modal activity data, and use the reconstruction loss to train the diffusion model to update the parameters of the denoising network in the diffusion model; Activity recognition model training stage: An activity recognition model is trained using a labeled dataset and multimodal activity data generated by a trained diffusion model based on the labeled dataset. The trained activity recognition model is used for activity recognition.

2. The activity recognition method based on a representation-guided multi-modal diffusion model according to claim 1, wherein A multimodal self-supervised model is trained using an unlabeled dataset to enable it to extract multimodal representations containing complex patterns and global representations containing structural information from the data, including: Construct a self-supervised model that includes multiple independent modality encoders, an attention network, and multiple independent modality decoders; Select unlabeled data samples from the unlabeled dataset for random masking to obtain masked data, and use multiple independent modality encoders to encode each modality in the masked data to obtain multimodal representations containing complex patterns; After concatenating the multimodal representations, use the attention network to obtain a global representation containing structural information, and use multiple independent modality decoders to decode each modality in the global representation to obtain reconstructed data; Construct a masked reconstruction loss based on the reconstructed data and the unlabeled data samples, and use the masked reconstruction loss to train the self-supervised model to update the parameters of the self-supervised model.

3. The activity recognition method based on a representation-guided multimodal diffusion model according to claim 2, wherein The masked reconstruction loss uses the mean squared error, expressed as: ; Among them, represents the mask reconstruction loss, and represent unlabeled data samples and reconstructed data respectively, represents unlabeled data samples from the distribution that the unlabeled data set follows , represents the expectation, represents the square of the L2 norm.

4. The activity recognition method based on a representation-guided multimodal diffusion model according to claim 1, wherein The same processing procedure for each downsampling module is: Calculate the first scaling parameter and the first translation parameter based on the input guidance condition, and then use the first scaling parameter and the first translation parameter to scale and translate the input data or input features of each downsampling module, and then obtain the output features after convolution and activation processing; For each modality layer, the input data of the first downsampling module is the noisy data, the input features of other downsampling modules are the output features of the previous downsampling module in each modality layer, and the output of the last downsampling module is the modality feature; For the global layer, the input feature of the first downsampling module is the fused feature, and the input features of other downsampling modules are the output features of the previous downsampling module in the global layer; The same processing procedure for each upsampling module is: Calculate the second scaling parameter and the second translation parameter based on the input guidance condition, and then use the second scaling parameter and the second translation parameter to scale and translate the input features of each upsampling module, and then obtain the output features after transposed convolution and activation processing; For the global layer, the input feature of the first upsampling module is the output feature of the last downsampling module in the global layer, the input features of other upsampling modules are the output features of the previous upsampling module in the global layer and the output features of the same-layer downsampling module through skip connections, and the output feature of the last upsampling module is the global feature; For each modality layer, the input feature of the first upsampling module is the global feature output by the last upsampling module in the global layer and the output features of the same-layer downsampling module through skip connections, the input features of other upsampling modules are the output features of the previous upsampling module in the modality layer and the output features of the same-layer downsampling module through skip connections, and the output of the last upsampling module is the denoised data.

5. The activity recognition method based on a representation-guided multimodal diffusion model according to claim 1, wherein Use the fusion module in the global layer to fuse the modality features of all modalities to obtain a fused feature, including: After splicing the modal features of all modalities to obtain the spliced features, the spliced features are input into a multi-layer convolutional layer to obtain the modal weight matrix. At the same time, the splicing matrix is input into two convolutional layers respectively to obtain the first fusion parameter and the second fusion parameter. The modal weight matrix is processed using the first fusion parameter to obtain the modal weights, and the modal weights are processed using the second fusion parameter to obtain the fused features.

6. The activity recognition method based on a representation-guided multi-modal diffusion model according to claim 1, wherein The reconstruction loss uses the mean squared error, expressed as: ; Among them, represents the reconstruction loss, and represent unlabeled data samples and the generated multimodal activity data respectively, represents the unlabeled data samples from the distribution that the unlabeled data set follows , represents the noise in the noise addition process from a normal distribution with a mean of 0 and a variance of . represents the time step, represents the square of the L2 norm.

7. The activity recognition method based on a representation-guided multimodal diffusion model according to claim 1, characterized in that Training the activity recognition model using the labeled dataset and the multi-modal activity data generated by the trained diffusion model based on the labeled dataset, including: Selecting labeled data samples from the labeled dataset, using the trained self-supervised model to generate multi-modal representations and global representations based on the labeled data samples, and obtaining the multi-modal representations and global representations used as guidance after taking the average along a fixed length; Using the trained diffusion model to perform multi-step denoising and adding noise to the initial random noise, and introducing the multi-modal representations and global representations as guidance conditions during the denoising process to generate multi-modal activity data; Inputting the labeled data samples selected from the labeled dataset and the generated modal activity data into the activity recognition model to obtain the classification results, constructing the classification loss based on the classification results and the true labels, and using the classification loss to train the activity recognition model to update the parameters of the activity recognition model.

8. The activity recognition method based on a representation-guided multi-modal diffusion model according to claim 7, wherein The classification loss uses the cross-entropy loss, expressed as: ; Among them, represents the classification loss, represents the true label, represents the classification result corresponding to the labeled data sample, represents the classification result corresponding to the generated modal activity data, represents calculating the cross-entropy loss, represents the expectation, represents the labeled data sample from the distribution that the labeled data set follows , represents the noise of the noise addition process from a normal distribution with a mean of 0 and a variance of the normal distribution.

9. An activity recognition device based on a representation-guided multi-modal diffusion model, characterized in that, Including: A data preprocessing unit, which is used to obtain the original multi-modal activity data of different users performing different activities collected by wearable devices and perform preprocessing to construct an unlabeled dataset and a labeled dataset: A self-supervised model training unit, which is used to train a multi-modal self-supervised model using the unlabeled dataset so that it extracts multi-modal representations containing complex patterns and global representations containing structural information from the data; A diffusion model training unit, which is used to introduce a modality layer guided by multi-modal representations and a global layer guided by global representations into the denoising network of the diffusion model. Among them, each modality layer and global layer include a continuous plurality of downsampling modules and a plurality of upsampling modules. The global layer also includes a fusion module. Among them, the number of downsampling modules and upsampling modules in each modality layer is equal, and a skip connection is formed between the downsampling module and its corresponding upsampling module; the number of downsampling modules and upsampling modules in the global layer is also equal. The output of the last downsampling module is directly connected to the input of the corresponding upsampling module, and a skip connection is formed between the remaining downsampling modules and their corresponding upsampling modules; and using the unlabeled dataset to train the diffusion model to generate multi-modal activity data, including the following steps: (1) Selecting unlabeled data samples from the unlabeled dataset, and using the trained self-supervised model to extract multi-modal representations and global representations based on the unlabeled data samples. The multi-modal representations and global representations are obtained as the multi-modal representations and global representations used as guidance after taking the average along a fixed length; (2) For each time step, the following process is performed: (2-1) Add noise to the input data to obtain the noisy data, and use two encoding networks to encode the time steps respectively to obtain the first time step encoding and the second time step encoding; (2-2) Use the noisy data as the input of the denoising network. Each modality layer processes the noisy data corresponding to one modality. Utilize multiple downsampling modules in each modality layer, and use the first time step encoding and the modality representation corresponding to each modality extracted from the multi-modal representation as guiding conditions to extract the modality features of each modality from the input noisy data; (2-3) Use the fusion module in the global layer to fuse the modality features of all modalities to obtain the fused features. Utilize multiple downsampling modules and upsampling modules in the global layer, and use the second time step encoding and the global representation as guiding conditions to extract the global features from the input fused features; (2-4) Use multiple upsampling modules in each modality layer, and use the first time step encoding and the modality representation corresponding to each modality extracted from the multi-modal representation as guiding conditions to extract the denoised data of each modality from the input global features. The denoised data of all modalities constitute the generated multi-modal activity data; (3) Construct a reconstruction loss based on the unlabeled data samples and the generated multi-modal activity data, and use the reconstruction loss to train the diffusion model to update the denoising network parameters in the diffusion model; An activity recognition model training unit, which is used to train an activity recognition model using a labeled data set and multi-modal activity data generated by the trained diffusion model based on the labeled data set. The trained activity recognition model is used for activity recognition.

Citation Information

Patent Citations

  • Multi-modal human activity recognition method based on generative adversarial network

    CN110309861A

  • Multimodal bacterial colony sample classification and identification method and system based on diffusion model

    CN117292195A