Model construction method and device for identifying motion and daily activities
By processing data through adaptive filtering algorithms and multi-sensor fusion systems, and combining deep convolutional neural networks and convolutional block attention mechanisms, the problems of complex data preprocessing and insufficient feature extraction in existing systems are solved, achieving efficient and accurate recognition of motion and daily activities.
Patent Information
- Application Number
- CN202510732018.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-10-28
AI Technical Summary
Existing motion and daily activity recognition systems suffer from problems such as complex data preprocessing, difficulty in completely eliminating noise, insufficient feature extraction, and low model training efficiency, resulting in insufficient recognition accuracy and reliability.
Adaptive filtering algorithms and multi-sensor fusion systems are used for data processing. Expert annotation and machine annotation are combined. An improved deep convolutional neural network and convolutional block attention mechanism are used to enhance feature extraction through channel and spatial attention modules. Sliding window technology and batch normalization layers are introduced. Cross-entropy loss function and Adam optimizer are used for model training.
It significantly improves the accuracy and robustness of the data, enhances the model's recognition precision and generalization ability, and can efficiently identify different types of motion and daily activities.
Smart Images

Figure CN120851089A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for constructing a model that recognizes motion and daily activities. Background Technology
[0002] Existing motion and daily activity recognition systems are widely used in health monitoring, sports training, smart homes, and other fields. These systems typically rely on a variety of sensor devices, such as accelerometers, gyroscopes, and magnetometers, to collect multimodal time-series data. Accelerometers measure the linear acceleration of an object along three axes, gyroscopes measure the angular velocity of an object around three axes, and magnetometers measure the components of the Earth's magnetic field along three axes.
[0003] These sensor devices are typically integrated into smart wearable devices, smartphones, or dedicated motion capture systems. To ensure data accuracy and reliability, each sensor requires calibration before use, including determining the sensor's bias, scaling factor, and non-orthogonal error. Furthermore, during data acquisition, the sensors continuously record data at a fixed sampling frequency, the choice of which depends on the application requirements; typically, accelerometers and gyroscopes use higher sampling frequencies, while magnetometers use relatively lower frequencies. In addition, to ensure data consistency and synchronization, data acquisition from all sensors must be performed on the same time base, usually achieved through hardware synchronization signals or software timestamps.
[0004] Although existing technologies can identify motion and daily activities, current motion and daily activity recognition systems often have the following problems and shortcomings:
[0005] First, the data preprocessing process is quite complex, requiring the removal of noise and outliers, including low-pass filtering, high-pass filtering, and median filtering. Low-pass filtering removes high-frequency noise, high-pass filtering eliminates low-frequency drift, and median filtering effectively removes impulse noise. Furthermore, data normalization is necessary to convert data from different sensors to the same numerical range for subsequent analysis and processing. Although these preprocessing methods improve data quality to some extent, they still cannot completely eliminate noise and outliers in practical applications, affecting the accuracy and reliability of the data.
[0006] Secondly, existing feature extraction methods mainly rely on time-domain features, frequency-domain features, and time-frequency-domain features. While these features can describe the characteristics of motion data, they often fail to fully capture key features in complex scenarios, resulting in low model recognition accuracy. Furthermore, traditional machine learning methods also have limitations in training efficiency and generalization ability when processing large-scale multimodal time series data, making it difficult to meet the needs of practical applications. Summary of the Invention
[0007] This application discloses a method and apparatus for building a model that identifies motion and daily activities.
[0008] In a first aspect, this application discloses a model building method for recognizing motion and daily activities, the method comprising:
[0009] The user's physiological signal data is collected using a pre-set wearable device and / or sensor, and the physiological signal data includes at least one of heart rate, gait and acceleration.
[0010] The physiological signal data is processed according to preset data processing rules; the data processing rules include: sequentially performing filtering and noise reduction, missing value filling, smoothing, and standardization or normalization on the physiological signal data;
[0011] This method combines expert annotation and automatic annotation by machine annotation models to annotate physiological signal data after data processing. The expert annotation is performed manually by professional medical personnel and / or sports scientists, while the machine annotation model automatically generates new annotated data based on historical annotated data using machine learning algorithms. The machine annotation model uses support vector machines or decision tree classification models to identify key features in a large amount of unannotated data and generate annotation results.
[0012] The labeled physiological signal data is used as input data to an improved deep convolutional neural network to extract the spatiotemporal features corresponding to motion and daily activities; the deep convolutional neural network is a deep network formed by a stacked structure of multiple convolutional layers and pooling layers under the deep learning framework.
[0013] The feature map, after being processed by multiple convolutions, pooling, and residual connections, is fed into a fully connected layer to map the high-dimensional spatiotemporal features to a low-dimensional space, generating a low-dimensional feature vector.
[0014] The convolutional block attention mechanism is used to enhance key features of low-dimensional feature vectors. The feature vectors processed by the convolutional block attention module are then input into a deep learning model for training, thereby building a model to recognize motion and daily activities. The deep learning model consists of multiple convolutional layers, pooling layers, fully connected layers, and activation functions. The convolutional block attention module consists of two independent sub-modules: a channel attention module and a spatial attention module. The channel attention module weights the feature map by calculating the importance weight of each channel.
[0015] Optionally, the step of inputting the labeled physiological signal data as input data into an improved deep convolutional neural network to extract the spatiotemporal features corresponding to motion and daily activities includes:
[0016] After preprocessing, the input data is fed into the convolutional layer of the convolutional neural network. Multiple convolutional kernels sequentially perform convolution operations on the input data according to the following expression, and generate multiple feature maps F.
[0017] F = σ(W*X + b)
[0018] Where W represents the convolution kernel corresponding to the convolutional layer in the convolutional neural network, X represents the input data, b represents the bias term, * represents the convolution operation, and σ represents the activation function.
[0019] Optionally, after generating multiple feature maps, the process may also include:
[0020] The bidirectional gated loop module performs forward and reverse loop calculations on the multiple feature maps one by one and then splices or fuses them. The spliced or fused feature maps are then input into the pooling layer for downsampling to reduce the dimensionality of the feature maps while retaining the main feature information. The pooling operation includes at least one of max pooling and average pooling.
[0021] Optionally, it also includes: using batch normalization technology to normalize each batch of input data according to the following expression;
[0022]
[0023] Where x represents the input data, μ B This represents the mean of the data in this batch. Let represent the variance of this batch of data, ∈ represent a small constant used to prevent division by zero errors, and γ and β represent learnable parameters.
[0024] Optionally, it also includes: introducing residual connections on top of the deep convolutional neural network, and the output H(x) of the residual block is represented as:
[0025] H(x)=F(x, W) i )+x
[0026] Where F(x, W) i ) represents the forward propagation function inside the residual block, and x represents the input data.
[0027] Optionally, the step of weighting the feature map by calculating the importance weight of each channel in the channel attention module includes:
[0028] The channel attention module compresses the feature map into two one-dimensional vectors in the spatial dimension through max pooling and / or average pooling operations. and The data is fed into a multilayer perceptron network and two weight vectors are obtained. and The multilayer perceptron network includes two fully connected layers. The first fully connected layer reduces the dimension of the input vector to C / r, and the second fully connected layer restores the dimension-reduced input vector to the C dimension.
[0029] The two weight vectors and Channel attention weights are generated by summing elements one by one and passing them through a Sigmoid activation function, according to the following expression.
[0030] M c =σ(MLP(F) max )+MLP(F avg ))
[0031] in, Let represent the given input feature map, C represent the number of channels, H represent the height of the feature map, and W represent the width of the feature map; r represents a hyperparameter used to control the dimensionality reduction ratio, and σ represents the Sigmoid activation function.
[0032] Optionally, it also includes:
[0033] The cross-entropy loss function is used as the optimization objective, and the difference between the predicted probability distribution of the constructed model and the true label is measured according to the following expression; by minimizing the cross-entropy loss function, the parameters of the constructed model are adjusted to make the prediction results closer to the true label.
[0034]
[0035] Where N represents the number of samples, C represents the number of categories, and y ij p represents the true label of the i-th sample belonging to the j-th category. ij This represents the probability that the i-th sample, as predicted by the constructed model, belongs to the j-th category.
[0036] Secondly, this application discloses a model building apparatus for recognizing motion and daily activities, the apparatus comprising:
[0037] A data acquisition unit is used to collect physiological signal data of a user using a preset wearable device and / or sensor, wherein the physiological signal data includes at least one of heart rate, gait and acceleration.
[0038] The data processing unit is used to process the physiological signal data according to preset data processing rules; the data processing rules include: sequentially performing filtering and noise reduction, missing value filling, smoothing, and standardization or normalization on the physiological signal data;
[0039] The data annotation unit is used to annotate physiological signal data after data processing by combining expert annotation and automatic annotation based on machine annotation models. The expert annotation is performed manually by professional medical personnel and / or sports scientists, and the machine annotation model automatically generates new annotated data based on historical annotated data through machine learning algorithms. The machine annotation model uses support vector machines or decision tree classification models to identify key features in a large amount of unannotated data and generate annotation results.
[0040] The convolutional neural network processing unit is used to input labeled physiological signal data into an improved deep convolutional neural network to extract spatiotemporal features corresponding to motion and daily activities; the deep convolutional neural network is a deep network formed by a stacked structure of multiple convolutional layers and pooling layers under the deep learning framework.
[0041] The feature vector dimensionality reduction processing unit is used to feed the feature map after multiple layers of convolution, pooling and residual connection into the fully connected layer, mapping the high-dimensional spatiotemporal features to the low-dimensional space and generating low-dimensional feature vectors.
[0042] The model building unit is used to perform key feature enhancement processing on low-dimensional feature vectors based on the convolutional block attention mechanism, and input the feature vectors processed by the convolutional block attention module into the deep learning model for model training to realize the model construction for recognizing motion and daily activities. The deep learning model consists of multiple convolutional layers, pooling layers, fully connected layers and activation functions. The convolutional block attention module consists of two independent sub-modules: a channel attention module and a spatial attention module. The channel attention module weights the feature map by calculating the importance weight of each channel.
[0043] Thirdly, this application discloses an electronic device comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to perform the method as described in any of the preceding aspects.
[0044] Fourthly, this application discloses a non-transitory computer-readable storage medium in which, when the instructions in the storage medium are executed by a processor of an electronic device, enable the electronic device to perform the methods described in any of the preceding aspects.
[0045] Fifthly, this application discloses a computer program product in which, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the method described in any of the preceding aspects.
[0046] The technical solution provided in this application may include the following beneficial effects:
[0047] (1) Improve the accuracy and reliability of data.
[0048] This application presents an adaptive filtering algorithm, including denoising (low-pass, high-pass, band-pass, and band-stop filters), missing value imputation (mean, median, mode, and interpolation imputation), smoothing (moving average, exponential smoothing, and Savitzky-Golay filter), and standardization or normalization (Z-score and Min-Max standardization). This algorithm processes high-quality raw motion data collected from multiple sensor devices, significantly improving data quality and ensuring it meets the requirements for subsequent analysis. Through a multi-sensor fusion system and the adaptive filtering algorithm, the accuracy and purity of the data are ensured.
[0049] (2) Enhance segmentation accuracy and robustness.
[0050] This application introduces a Convolutional Block Attention (CNN-CBAM) mechanism into convolutional neural networks. It calculates the weights of feature maps through two sub-modules: channel attention and spatial attention, thereby achieving channel-wise and position-wise weighted processing of the input feature maps, significantly improving the model's ability to identify key features. A sliding window technique is also introduced, dividing the continuous data stream into fixed-length time windows. Data within each window is treated as an independent sample, improving data robustness and model training efficiency.
[0051] (3) Improve the training speed and generalization ability of the model.
[0052] This application adds batch normalization and dropout layers between multiple convolutional and pooling layers. By standardizing each batch of input and randomly dropping some neurons, the training process is accelerated and the model's generalization ability is improved, enabling the deep learning model to efficiently learn and recognize different types of motion and daily activities. The cross-entropy loss function and Adam optimizer ensure the model's efficient learning and recognition of different types of motion and daily activities, improving the model's accuracy and robustness. The use of convolutional neural networks and convolutional block attention (CNN-CBAM) effectively extracts spatiotemporal features related to motion and daily activities, significantly improving the model's recognition accuracy and generalization ability. Attached Figure Description
[0053] Figure 1 A flowchart illustrating the model construction method for recognizing motion and daily activities provided in this application.
[0054] Figure 2 A flowchart illustrating the overall algorithm for constructing the model for recognizing motion and daily activities provided in this application.
[0055] Figure 3Channel attention map for the convolutional block attention mechanism provided in this application.
[0056] Figure 4 Spatial attention map for the convolutional block attention mechanism provided in this application.
[0057] Figure 5 A flowchart of the convolutional block attention mechanism provided in this application.
[0058] Figure 6 A structural diagram of the model building device for recognizing motion and daily activities provided in this application.
[0059] Figure 7 A block diagram of an electronic device provided in this application.
[0060] Figure 8 A block diagram of another electronic device provided in this application. Detailed Implementation
[0061] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0062] As can be seen from the foregoing, existing motion and daily activity recognition systems often have the following problems and defects: (1) The data preprocessing process is relatively complex, requiring the removal of noise and outliers, including low-pass filtering, high-pass filtering, and median filtering. (2) Existing feature extraction methods mainly rely on time-domain features, frequency-domain features, and time-frequency-domain features. Although these features can describe the characteristics of motion data, they often fail to fully capture key features in complex scenarios, resulting in low recognition accuracy of the model. (3) Traditional machine learning methods also have limitations in training efficiency and generalization ability when processing large-scale multimodal time series data, making it difficult to meet the needs of practical applications.
[0063] In view of this, this application provides a method and apparatus for building a model for recognizing motion and daily activities based on a combination of algorithms including CNN (Convolutional Neural Networks), BIGRU (Bidirectional Gated Recurrent Unit), and CBAM (Convolutional Block Attention Module). Given an intermediate feature map, the module provided in this application sequentially infers an attention map along two independent dimensions (channel and space), and then multiplies the attention map by the input feature map for adaptive feature enhancement. The following section combines... Figures 1 to 5 The method for constructing a model to identify motion and daily activities provided in this application is described.
[0064] Example 1
[0065] Please refer to Figure 1 The above is a flowchart of a model construction method for recognizing motion and daily activities provided in this application. This method can be applied to electronic devices, and specifically includes the following steps:
[0066] In step S101, the user's physiological signal data is collected using a preset wearable device and / or sensor, the physiological signal data including at least one of heart rate, gait and acceleration.
[0067] Specifically, heart rate data is acquired using a photoplethysmography (PPG) sensor, while gait and acceleration data are acquired using an inertial measurement unit (IMU) sensor, which integrates a triaxial accelerometer, a triaxial gyroscope, and a triaxial magnetometer. To improve the efficiency and quality of data acquisition, this application provides a multi-sensor fusion system capable of simultaneously acquiring data from multiple sensors and ensuring temporal consistency of data from each sensor through timestamp synchronization technology.
[0068] In step S102, the physiological signal data is processed according to preset data processing rules; the data processing rules include: sequentially performing filtering and noise reduction, missing value filling, smoothing, standardization or normalization on the physiological signal data.
[0069] Understandably, in order to ensure that the data quality meets the requirements of subsequent analysis, a series of detailed processing operations are required on the collected data.
[0070] First, filtering and noise reduction is an essential step. By removing noise signals, the purity of the data can be significantly improved. Filtering methods include low-pass filters, high-pass filters, band-pass filters, and band-stop filters. Low-pass filters are used to remove high-frequency noise and retain low-frequency signals; high-pass filters, on the other hand, are used to remove low-frequency noise and retain high-frequency signals; band-pass filters allow signals within a specific frequency range to pass through, while band-stop filters block signals within a specific frequency range from passing through. These filters can convert the time-domain signal to the frequency-domain signal using Fourier transform, then perform filtering operations in the frequency domain, and finally convert the signal back to the time domain using inverse Fourier transform.
[0071] Secondly, missing value imputation is also a crucial step in data preprocessing. The presence of missing values can affect the accuracy of subsequent analysis, thus requiring appropriate imputation methods. Common missing value imputation methods include mean imputation, median imputation, mode imputation, and interpolation imputation. Mean imputation uses the average of all non-missing values for a given feature to fill in the missing value; median imputation uses the median of all non-missing values for that feature; mode imputation uses the most frequent value among all non-missing values for that feature; and interpolation imputation estimates the missing value based on known data points using methods such as linear interpolation or spline interpolation. Furthermore, the K-nearest neighbor algorithm or decision tree algorithm can be used to predict missing values based on the values of other features.
[0072] Next, smoothing is another crucial step in data preprocessing. Smoothing reduces random fluctuations in the data, making it smoother and easier for subsequent analysis. Methods include moving averages, exponential smoothing, and the Savitzky-Golay filter. Moving averages smooth data by calculating the average of data points within a certain window; the window size affects the smoothing effect. Exponential smoothing assigns different weights to data points at different times, with data closer to the current time point receiving greater weight; this method better captures data trends. The Savitzky-Golay filter is a polynomial smoothing method that smooths data by fitting local polynomials, effectively preserving the data's characteristic information.
[0073] After completing the filtering, denoising, missing value imputation, and smoothing processes described above, the data also needs to be standardized or normalized to ensure consistency in the dimensions of different features and to avoid adverse effects on model training due to excessively large or small numerical ranges for certain features. Common standardization methods include Z-score standardization and Min-Max standardization.
[0074] Z-score standardization is achieved by transforming the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. The formula is:
[0075]
[0076] Min-Max normalization is achieved by scaling the data to a range of 0 to 1, as shown in the formula:
[0077]
[0078] Where x represents the original data, μ represents the mean, and σ represents the standard deviation.
[0079] Furthermore, to further improve the robustness of the data, a sliding window technique is employed. This technique divides the continuous data stream into fixed-length time windows, with each window's data treated as an independent sample. The size and step size of the sliding window can be adjusted according to specific application requirements, typically ranging from a few seconds to tens of seconds, with a step size approximately half the window size. This approach not only helps reduce the amount of data but also improves the model's training efficiency.
[0080] In summary, by performing filtering and noise reduction, missing value imputation, smoothing, and standardization or normalization on the collected data, the quality of the data can be significantly improved, ensuring that it meets the requirements of subsequent analysis.
[0081] In step S103, the physiological signal data after data processing is labeled using a combination of expert labeling and machine labeling model automatic labeling. The expert labeling is performed manually by professional medical personnel and / or sports scientists, and the machine labeling model automatically generates new labeled data based on historical labeled data using machine learning algorithms. The machine labeling model uses support vector machine or decision tree classification model to identify key features in a large amount of unlabeled data and generate labeling results.
[0082] In the data annotation phase, a combination of expert annotation and automatic annotation is employed to ensure the accuracy and authority of the annotations. Expert annotation is completed by professional medical personnel or sports scientists, while automatic annotation uses machine learning algorithms to generate new annotation results based on existing annotated data. It employs classification models such as Support Vector Machines (SVM) or decision trees to quickly identify key features from large amounts of unlabeled data, generating reliable annotation results. Mathematically, we assume the data segment is x. i Its corresponding label is y i The annotation process can then be formalized as labeling each data fragment x. i Mapped to the corresponding label y i , that is, y i =f(x) i ), where f is the labeling function. After labeling, the resulting dataset is D = {(x1, y1), (x2, y2), ..., (x... n ,yn )} will be used for training the supervised learning model.
[0083] In step S104, the labeled physiological signal data is used as input data to the improved deep convolutional neural network to extract the spatiotemporal features corresponding to motion and daily activities; the deep convolutional neural network is a deep network formed by a stacked structure of multiple convolutional layers and pooling layers under the deep learning framework.
[0084] It should be noted that, in order to further enhance the expressive power of features, this application introduces the above-mentioned stacked structure, which is composed of multiple convolutional layers and multiple pooling layers.
[0085] In the feature extraction stage, the input data is preprocessed and then fed into the convolutional layer of the convolutional neural network. Multiple convolutional kernels sequentially perform convolution operations on the input data according to the following expression, and generate multiple feature maps F.
[0086] F = σ(W*X + b)
[0087] Where W represents the convolution kernel corresponding to the convolutional layer in the convolutional neural network, X represents the input data, b represents the bias term, * represents the convolution operation, and σ represents the activation function, usually the ReLU function.
[0088] See Figure 2 After generating multiple feature maps, the process further includes: a bidirectional gated loop module performing forward and reverse loop calculations on each of the multiple feature maps and then concatenating or fusing them. The concatenated or fused feature maps are then input into a pooling layer for downsampling to reduce the dimensionality of the feature maps while retaining the main feature information. The pooling operation includes at least one of max pooling and average pooling. The pooling operation can be expressed as P = f(F), where P represents the pooled feature map and f represents the pooling function.
[0089] In step S105, the feature map after being processed by multiple convolutions, pooling and residual connections is fed into a fully connected layer to map the high-dimensional spatiotemporal features to a low-dimensional space and generate a low-dimensional feature vector.
[0090] Specifically, the output of the fully connected layer can be expressed as O = σ(W f ·F+b f ), where O represents the output of the fully connected layer, W f Let F represent the weight matrix of the fully connected layer, and b represent the feature map. f σ represents the bias term, and σ represents the activation function.
[0091] Channel attention weights Mc are multiplied channel-wise with the original feature map F to obtain the channel-attention-enhanced feature map F. c The formula is
[0092] The spatial attention module further optimizes the feature map by calculating the importance weight of each location. The spatial attention module first processes the channel-attention-enhanced feature map F... c Max pooling and average pooling operations are performed along the channel dimension to generate two two-dimensional feature maps. and These two feature maps are stacked together along the channel dimension to form a three-dimensional feature map.
[0093] Then, this 3D feature map is fed into a 7×7 convolutional layer, which is used to generate spatial attention weights. The formula is M s ∈σ(Conv(F concat )), where Conv represents the convolution operation and σ represents the Sigmoid activation function.
[0094] Finally, the spatial attention weight M s Feature map F after channel attention c Perform position-by-position multiplication to obtain the final spatial attention-based feature map F. out The formula is Through the above steps, CBAM can effectively enhance the model's ability to capture key features and improve the model's performance in various tasks.
[0095] In step S106, key feature enhancement processing is performed on the low-dimensional feature vector based on the convolutional block attention mechanism, and the feature vector processed by the convolutional block attention module is input into the deep learning model for model training to realize the model construction for recognizing motion and daily activities; the deep learning model consists of multiple convolutional layers, pooling layers, fully connected layers and activation functions, and the convolutional block attention module consists of two independent sub-modules: a channel attention module and a spatial attention module. The channel attention module weights the feature map by calculating the importance weight of each channel.
[0096] Features processed by CBAM are input into a deep learning model for training. CBAM enhances the expressive power of feature maps in both spatial and channel dimensions, thereby improving feature discriminativity and robustness. These enhanced features are fed into a deep neural network model consisting of multiple convolutional layers, pooling layers, fully connected layers, and activation functions. Convolutional layers extract local features, pooling layers reduce feature dimensionality, fully connected layers integrate global information, and activation functions introduce non-linearity, enabling the model to learn more complex patterns.
[0097] To ensure that the model can accurately identify different movements and daily activities, the cross-entropy loss function is used as the optimization objective. The cross-entropy loss function is used to measure the difference between the predicted probability distribution of the constructed model and the true label. By minimizing the cross-entropy loss function, the parameters of the constructed model are adjusted to make the prediction results closer to the true label.
[0098]
[0099] Where N represents the number of samples, C represents the number of categories, and y ij p represents the true label of the i-th sample belonging to the j-th category. ij This represents the probability that the i-th sample, as predicted by the constructed model, belongs to the j-th category.
[0100] To efficiently update parameters, the Adam optimizer is chosen. The Adam optimizer combines the advantages of momentum and RMSProp, and performs well in cases of sparse gradients and non-stationary objective functions. The update rule of the Adam optimizer is as follows:
[0101] m t =β1m t-1 +(1-β1)g t
[0102]
[0103] Where, m t and v t These are the first-order moment estimate and the second-order moment estimate of the gradient, g. t α is the gradient at the current time step, β1 and β2 are the decay rates, α is the learning rate, and ∈ is a small constant to prevent division by zero.
[0104] Furthermore, batch normalization is used to normalize each batch of input data according to the following expression to accelerate the training process and improve the model's generalization ability.
[0105]
[0106] Where x represents the input data, μ B This represents the mean of the data in this batch. Let represent the variance of this batch of data, ∈ represent a small constant used to prevent division by zero errors, and γ and β represent learnable parameters.
[0107] In one scenario, residual connections are introduced into deep convolutional neural networks, which skip direct connections in certain layers to address the vanishing gradient problem and improve the network's training performance. The output H(x) of the residual block is expressed as:
[0108] H(x)=F(x, W)i )+x
[0109] Where F(x, W) i ) represents the forward propagation function inside the residual block, and x represents the input data.
[0110] Specifically, the feature map can be weighted by calculating the importance weight of each channel using the following steps:
[0111] The channel attention module compresses the feature map into two one-dimensional vectors in the spatial dimension through max pooling and / or average pooling operations. and The data is fed into a multilayer perceptron network and two weight vectors are obtained. and The multilayer perceptron network includes two fully connected layers. The first fully connected layer reduces the dimension of the input vector to C / r, and the second fully connected layer restores the dimension-reduced input vector to the C dimension.
[0112] The two weight vectors and Channel attention weights are generated by summing elements one by one and passing them through a Sigmoid activation function, according to the following expression.
[0113] M c =σ(MLP(F) max )+MLP(F avg ))
[0114] in, Let represent the given input feature map, C represent the number of channels, H represent the height of the feature map, and W represent the width of the feature map; r represents a hyperparameter used to control the dimensionality reduction ratio, and σ represents the Sigmoid activation function.
[0115] Through these steps, the Adam optimizer effectively tunes model parameters, accelerates the convergence process, and improves the model's generalization ability. Batch normalization is also employed throughout the training process to reduce internal covariate bias, speeding up training and improving model performance. Batch normalization normalizes the mean and variance of each mini-batch of data, resulting in features with zero mean and unit variance, thus stabilizing the training process and reducing gradient vanishing and exploding problems. Furthermore, to avoid overfitting, dropout regularization is used. Dropout randomly deactivates some neurons, reducing the model's dependence on specific input features and enhancing its generalization ability. Through these comprehensive measures, the deep learning model can efficiently learn and recognize different types of motion and daily activities, ensuring the model's accuracy and robustness.
[0116] In summary, applying the solution provided in this application to identify daily motion activities can achieve the following beneficial effects: (1) Improve the accuracy and reliability of the data. The adaptive filtering algorithm of this application includes filtering and denoising (low-pass, high-pass, band-pass and band-stop filters), missing value filling (mean, median, mode and interpolation filling), smoothing processing (moving average method, exponential smoothing method and Savitzky-Golay filter) and standardization or normalization processing (Z-score and Min-Max standardization), which processes high-quality raw motion data collected from multiple sensor devices, significantly improving the quality of the data and ensuring that it meets the requirements of subsequent analysis. Through the multi-sensor fusion system and adaptive filtering algorithm, the accuracy and purity of the data are ensured. (2) Enhance the segmentation accuracy and robustness. This application introduces the convolutional block attention mechanism (CNN-CBAM) in the convolutional neural network, and calculates the weights of the feature map through the two sub-modules of channel attention and spatial attention, thereby realizing the weighted processing of the input feature map channel by channel and position by position, which significantly improves the model's ability to identify key features. A sliding window technique is introduced to divide the continuous data stream into fixed-length time windows. The data in each window is regarded as an independent sample, which improves the robustness of the data and the training efficiency of the model. (3) Improve the training speed and generalization ability of the model. This application adds batch normalization and dropout layers between multiple convolutional and pooling layers. By standardizing the input of each batch and randomly dropping some neurons, the training process is accelerated and the generalization ability of the model is improved, enabling the deep learning model to efficiently learn and recognize different types of motion and daily activities. Through the cross-entropy loss function and Adam optimizer, it is ensured that the model can efficiently learn and recognize different types of motion and daily activities, improving the accuracy and robustness of the model. Convolutional neural networks and convolutional block attention mechanism (CNN-CBAM) are used to effectively extract the spatiotemporal features related to motion and daily activities, which significantly improves the recognition accuracy and generalization ability of the model.
[0117] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by this application.
[0118] Example 2
[0119] Reference Figure 6This is a structural diagram of a model building device for recognizing motion and daily activities provided in this application. The device includes:
[0120] Data acquisition unit 210 is used to acquire physiological signal data of a user using a preset wearable device and / or sensor, wherein the physiological signal data includes at least one of heart rate, gait and acceleration;
[0121] The data processing unit 220 is used to process the physiological signal data according to preset data processing rules; the data processing rules include: sequentially performing filtering and noise reduction, missing value filling, smoothing, and standardization or normalization on the physiological signal data;
[0122] The data annotation unit 230 is used to annotate physiological signal data after data processing by combining expert annotation and automatic annotation based on machine annotation models. The expert annotation is performed manually by professional medical personnel and / or sports scientists. The machine annotation model automatically generates new annotated data based on historical annotation data through machine learning algorithms. The machine annotation model uses support vector machines or decision tree classification models to identify key features in a large amount of unannotated data and generate annotation results.
[0123] The convolutional neural network processing unit 240 is used to input labeled physiological signal data into an improved deep convolutional neural network to extract spatiotemporal features corresponding to motion and daily activities; the deep convolutional neural network is a deep network formed by a stacked structure composed of multiple convolutional layers and pooling layers under the deep learning framework.
[0124] The feature vector dimensionality reduction processing unit 250 is used to feed the feature map after multi-layer convolution, pooling and residual connection processing into the fully connected layer, mapping the high-dimensional spatiotemporal features to the low-dimensional space and generating low-dimensional feature vectors.
[0125] The model building unit 260 is used to perform key feature enhancement processing on low-dimensional feature vectors based on the convolutional block attention mechanism, and input the feature vectors processed by the convolutional block attention module into the deep learning model for model training, thereby realizing the model construction for recognizing motion and daily activities. The deep learning model consists of multiple convolutional layers, pooling layers, fully connected layers, and activation functions. The convolutional block attention module consists of two independent sub-modules: a channel attention module and a spatial attention module. The channel attention module weights the feature map by calculating the importance weight of each channel. The technical solution provided by this application can include the following beneficial effects:
[0126] (1) Improve the accuracy and reliability of data.
[0127] This application presents an adaptive filtering algorithm, including denoising (low-pass, high-pass, band-pass, and band-stop filters), missing value imputation (mean, median, mode, and interpolation imputation), smoothing (moving average, exponential smoothing, and Savitzky-Golay filter), and standardization or normalization (Z-score and Min-Max standardization). This algorithm processes high-quality raw motion data collected from multiple sensor devices, significantly improving data quality and ensuring it meets the requirements for subsequent analysis. Through a multi-sensor fusion system and the adaptive filtering algorithm, the accuracy and purity of the data are ensured.
[0128] (2) Enhance segmentation accuracy and robustness.
[0129] This application introduces a Convolutional Block Attention (CNN-CBAM) mechanism into convolutional neural networks. It calculates the weights of feature maps through two sub-modules: channel attention and spatial attention, thereby achieving channel-wise and position-wise weighted processing of the input feature maps, significantly improving the model's ability to identify key features. A sliding window technique is also introduced, dividing the continuous data stream into fixed-length time windows. Data within each window is treated as an independent sample, improving data robustness and model training efficiency.
[0130] (3) Improve the training speed and generalization ability of the model.
[0131] This application adds batch normalization and dropout layers between multiple convolutional and pooling layers. By standardizing each batch of input and randomly dropping some neurons, the training process is accelerated and the model's generalization ability is improved, enabling the deep learning model to efficiently learn and recognize different types of motion and daily activities. The cross-entropy loss function and Adam optimizer ensure the model's efficient learning and recognition of different types of motion and daily activities, improving the model's accuracy and robustness. The use of convolutional neural networks and convolutional block attention (CNN-CBAM) effectively extracts spatiotemporal features related to motion and daily activities, significantly improving the model's recognition accuracy and generalization ability.
[0132] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0133] Example 3
[0134] Optionally, this application also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0135] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0136] Figure 7 This application provides a block diagram of an electronic device 800. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0137] Reference Figure 7 The electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0138] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0139] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, images, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0140] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0141] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0142] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0143] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0144] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0145] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast operation information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0146] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0147] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0148] Example 4
[0149] Figure 8 A block diagram of another electronic device 1900 provided for this application. For example, electronic device 1900 may be provided as a server.
[0150] Reference Figure 8 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0151] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output (I / O) interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0152] Example 5
[0153] Fifthly, this application discloses a computer program product in which, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the method described in any of the preceding aspects.
[0154] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0156] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
[0157] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0158] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0159] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0161] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0162] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0163] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for constructing a model to identify movement and daily activities, characterized in that, The method includes: The user's physiological signal data is collected using a pre-set wearable device and / or sensor, and the physiological signal data includes at least one of heart rate, gait and acceleration. The physiological signal data is processed according to preset data processing rules; the data processing rules include: sequentially performing filtering and noise reduction, missing value filling, smoothing, and standardization or normalization on the physiological signal data; This method combines expert annotation and automatic annotation by machine annotation models to annotate physiological signal data after data processing. The expert annotation is performed manually by professional medical personnel and / or sports scientists, while the machine annotation model automatically generates new annotated data based on historical annotated data using machine learning algorithms. The machine annotation model uses support vector machines or decision tree classification models to identify key features in a large amount of unannotated data and generate annotation results. The labeled physiological signal data is used as input data to an improved deep convolutional neural network to extract the spatiotemporal features corresponding to motion and daily activities; the deep convolutional neural network is a deep network formed by a stacked structure of multiple convolutional layers and pooling layers under the deep learning framework. The feature map, after being processed by multiple convolutions, pooling, and residual connections, is fed into a fully connected layer to map the high-dimensional spatiotemporal features to a low-dimensional space, generating a low-dimensional feature vector. The convolutional block attention mechanism is used to enhance key features of low-dimensional feature vectors. The feature vectors processed by the convolutional block attention module are then input into a deep learning model for training, thereby building a model to recognize motion and daily activities. The deep learning model consists of multiple convolutional layers, pooling layers, fully connected layers, and activation functions. The convolutional block attention module consists of two independent sub-modules: a channel attention module and a spatial attention module. The channel attention module weights the feature map by calculating the importance weight of each channel.
2. The method for constructing a model to identify motion and daily activities according to claim 1, characterized in that, The step of inputting labeled physiological signal data as input data into an improved deep convolutional neural network to extract spatiotemporal features corresponding to motion and daily activities includes: After preprocessing, the input data is fed into the convolutional layer of the convolutional neural network. Multiple convolutional kernels sequentially perform convolution operations on the input data according to the following expression, and generate multiple feature maps F. F = σ(W*X + b) Where W represents the convolution kernel corresponding to the convolutional layer in the convolutional neural network, X represents the input data, b represents the bias term, * represents the convolution operation, and σ represents the activation function.
3. The method for constructing a model to identify motion and daily activities according to claim 2, characterized in that, After generating multiple feature maps, the process also includes: The bidirectional gated loop module performs forward and reverse loop calculations on the multiple feature maps one by one and then splices or fuses them. The spliced or fused feature maps are then input into the pooling layer for downsampling to reduce the dimensionality of the feature maps while retaining the main feature information. The pooling operation includes at least one of max pooling and average pooling.
4. The method for constructing a model to identify motion and daily activities according to claim 1, characterized in that, Also includes: Using batch normalization technology, normalize the input data for each batch according to the following expression; Where x represents the input data, μ B This represents the mean of the data in this batch. Let represent the variance of this batch of data, ∈ represent a small constant used to prevent division by zero errors, and γ and β represent learnable parameters.
5. The method for constructing a model to identify motion and daily activities according to claim 1, characterized in that, Also includes: Residual connections are introduced on top of deep convolutional neural networks, and the output H(x) of the residual block is expressed as: H(x)=F(x,W i )+x Where F(x, W) i ) represents the forward propagation function inside the residual block, and x represents the input data.
6. The method for constructing a model to identify motion and daily activities according to claim 1, characterized in that, The channel attention module weights the feature map by calculating the importance weight of each channel, including the following steps: The channel attention module compresses the feature map into two one-dimensional vectors in the spatial dimension through max pooling and / or average pooling operations. and The data is fed into a multilayer perceptron network and two weight vectors are obtained. and The multilayer perceptron network includes two fully connected layers. The first fully connected layer reduces the dimension of the input vector to C / r, and the second fully connected layer restores the dimension-reduced input vector to the C dimension. The two weight vectors and Channel attention weights are generated by summing elements one by one and passing them through a Sigmoid activation function, according to the following expression. M c <σ(MLP(F max )+MLP(F avg )) in, Let represent the given input feature map, C represent the number of channels, H represent the height of the feature map, and W represent the width of the feature map; r represents a hyperparameter used to control the dimensionality reduction ratio, and σ represents the Sigmoid activation function.
7. The method for constructing a model to identify motion and daily activities according to claim 1, characterized in that, Also includes: The cross-entropy loss function is used as the optimization objective, and the difference between the predicted probability distribution of the constructed model and the true label is measured according to the following expression; by minimizing the cross-entropy loss function, the parameters of the constructed model are adjusted to make the prediction results closer to the true label. Where N represents the number of samples, C represents the number of categories, and y ij p represents the true label of the i-th sample belonging to the j-th category, for example, it can be represented by 0 or 1. ij This represents the probability that the i-th sample, as predicted by the constructed model, belongs to the j-th category.
8. A model building device for recognizing movement and daily activities, characterized in that, The device includes: A data acquisition unit is used to collect physiological signal data of a user using a preset wearable device and / or sensor, wherein the physiological signal data includes at least one of heart rate, gait and acceleration. The data processing unit is used to process the physiological signal data according to preset data processing rules; the data processing rules include: sequentially performing filtering and noise reduction, missing value filling, smoothing, and standardization or normalization on the physiological signal data; The data annotation unit is used to annotate physiological signal data after data processing by combining expert annotation and automatic annotation based on machine annotation models. The expert annotation is performed manually by professional medical personnel and / or sports scientists, and the machine annotation model automatically generates new annotated data based on historical annotated data through machine learning algorithms. The machine annotation model uses support vector machines or decision tree classification models to identify key features in a large amount of unannotated data and generate annotation results. The convolutional neural network processing unit is used to input labeled physiological signal data into an improved deep convolutional neural network to extract spatiotemporal features corresponding to motion and daily activities; the deep convolutional neural network is a deep network formed by a stacked structure of multiple convolutional layers and pooling layers under the deep learning framework. The feature vector dimensionality reduction processing unit is used to feed the feature map after multiple layers of convolution, pooling and residual connection into the fully connected layer, mapping the high-dimensional spatiotemporal features to the low-dimensional space and generating low-dimensional feature vectors. The model building unit is used to perform key feature enhancement processing on low-dimensional feature vectors based on the convolutional block attention mechanism, and input the feature vectors processed by the convolutional block attention module into the deep learning model for model training to realize the model construction for recognizing motion and daily activities. The deep learning model consists of multiple convolutional layers, pooling layers, fully connected layers and activation functions. The convolutional block attention module consists of two independent sub-modules: a channel attention module and a spatial attention module. The channel attention module weights the feature map by calculating the importance weight of each channel.
9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.