Wearable human body behavior activity identification method based on Hash space optimization
By constructing a compact hash feature representation and introducing MMD and self-updating center loss optimization mechanisms, the problems of insufficient class separation and loose intra-class distribution in human activity recognition by hash methods are solved, thereby improving recognition accuracy and generalization performance. It is suitable for resource-constrained wearable devices and edge computing terminals.
Patent Information
- Application Number
- CN202510800098.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing hashing methods suffer from problems such as insufficient class separation, loose distribution of intra-class samples, and difficulty in dealing with class imbalance in human activity recognition, resulting in insufficient recognition accuracy and generalization performance.
By constructing a compact hash feature representation, the maximum mean difference (MMD) method is introduced to optimize the distribution difference between categories, and the self-updating center loss is used to optimize the hash vector of the same category. The MMD loss and center loss are combined for joint optimization.
It significantly improves the model's recognition accuracy and generalization performance, especially in classification performance in multi-class and complex environments, and is suitable for resource-constrained wearable devices and edge computing terminals.
Smart Images

Figure CN120910602A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human activity recognition, and particularly relates to a wearable human behavior activity recognition method based on hash space optimization. BACKGROUND
[0002] With the rapid development of artificial intelligence, deep learning and Internet of Things technology, human activity recognition has gradually become one of the important research hotspots in the field of intelligent perception and wisdom interaction. Human activity recognition is to automatically realize the recognition, understanding and classification of human motion and behavior by analyzing sensor data, video image data or other types of time series data. In recent years, this technology has been widely applied in intelligent security monitoring, smart home, human-computer interaction, medical health monitoring, rehabilitation training and sports analysis and other fields. For example, in the field of medical health, human activity recognition technology can monitor the physical state of patients in real time, assist in diagnosis and rehabilitation training; in the field of smart home, it can actively perceive the behavior habits of users and provide more personalized and convenient life services; and in the field of security monitoring, it can effectively improve the security protection capability by automatically identifying abnormal behaviors.
[0003] Traditional activity recognition methods, such as support vector machine SVM, logistic regression and random forest, usually rely on feature engineering, that is, a large number of features are designed and extracted by human, and then the human activity is recognized and classified based on these features. However, this manual feature extraction method not only needs to consume a lot of manpower and time cost, but also the extracted features often have strong scene dependence, which is difficult to flexibly adapt to the dynamic changes in the actual environment. Therefore, in recent years, with the rise of deep learning technologies such as convolutional neural network CNN and recurrent neural network RNN, automatic feature learning methods have begun to attract widespread attention and have shown significant advantages in generalization performance and feature representation ability. However, CNN is very effective in extracting spatial features, but it is usually difficult to capture long-term temporal dependencies in time series data; and RNN model is prone to gradient vanishing or gradient explosion phenomenon, which seriously limits its modeling ability for long-term dependencies.
[0004] In view of the limitations of the above traditional and deep learning methods, researchers need to explore alternative technical solutions that take into account efficiency and accuracy. In this context, the method based on hash coding has gradually become one of the important research directions in the field of activity recognition due to its efficient computing performance, low storage overhead and good scalability. Hash method can efficiently compress and map high-dimensional original feature data to compact binary representation, significantly improving the storage efficiency and retrieval speed of the model, and is particularly suitable for edge computing scenarios with limited computing resources and storage space.
[0005] However, the existing hash method still faces several challenges in human activity recognition tasks, mainly in the following aspects: (1) The traditional hash method usually ignores the distribution difference between classes, resulting in insufficient separation degree between classes in the hash space, which is easy to cause class confusion, thereby reducing the accuracy of activity classification. (2) The existing method does not fully consider the dispersion of samples in the same class, and the hash coding of intra-class samples is often distributed loosely in the hash space, resulting in large hash difference between samples of the same class, which further affects the recognition performance. (3) In practical application scenarios, data often shows obvious class imbalance characteristics, and the existing hash method is difficult to effectively cope with this imbalance, resulting in poor generalization performance and robustness of the model on the minority class SUMMARY
[0006] The present application aims to solve the above-mentioned problems in human activity recognition, and proposes a wearable human behavior activity recognition method based on hash space optimization. The method improves the efficiency and accuracy of the model in the multi-class human activity recognition task by constructing a compact hash feature representation and introducing a joint optimization strategy.
[0007] In order to achieve the above-mentioned application purposes, the technical scheme adopted by the present application is as follows: A wearable human behavior activity recognition method based on hash space optimization, comprising the following steps:
[0008] S1, data preprocessing, processing the human activity recognition dataset, such as OPPORTUNITY, PAMAP2 and WISDM, etc. The original sensor signal data into time series data with better robustness;
[0009] S2, feature extraction and hash embedding, the processed human activity time series data is extracted by neural network method, and embedded into hash space using tanh function;
[0010] S3, MMD optimization, for different hash vectors in the hash space of different classes, use MMD method to maximize the distribution difference between different classes;
[0011] S4, Center optimization, for different hash vectors of the same class in the hash space, use self-updating center loss optimization to reduce the dispersion of the same class in the hash space.
[0012] S5, loss calculation, combine MMD loss and center loss to jointly optimize the hash vector in the hash space.
[0013] S6, use the optimized hash vector in the hash space to predict different classes of human activity recognition.
[0014] Further preferred embodiments of the present application are as follows:
[0015] S101, data set denoising: because the sensor is worn on the human body, sudden large changes or device shaking will generate a lot of noise, WT method is used for denoising, and the wavelet transform retains the information of time domain and frequency domain at the same time, thereby minimizing the loss of key information. The key wavelet transform formula is as shown in the formula below:
[0016] x(t) = ∫T(W(x), T) ψ(t) (1)
[0017] Where x represents the original signal, T(W(x), T) represents the threshold function, and ψ(t) represents the wavelet function.
[0018] S102, data standardization: in order to improve the efficiency and stability of subsequent model training, the z-score method is used to standardize the data after denoising, and the original sensor signal is standardized to a unified scale, so that the model can be more stable. The mean of the standardized data is 0 and the standard deviation is 1. The formula is as follows:
[0019]
[0020] Where μ represents the mean of the original data, and σ represents the standard deviation of the original data.
[0021] S103, data division: the data set is divided into training set, validation set and test set according to the ratio of 7:2:1.
[0022] Further preferred schemes of the application are that the step S2 comprises the following steps:
[0023] S201, feature extraction: the preprocessed human activity recognition data Where T represents the time step, and D represents the feature dimension. Through the method of neural network, a series of nonlinear transformations are performed to map high-dimensional feature representation Where K represents the dimension of the feature space, and the formula is as follows:
[0024] f(x) = σ(Wx+b) (3)
[0025] Where W is the weight matrix, b is the bias term, and σ is the activation function.
[0026] S202, hash embedding: in order to map the dimension of the feature space to the low-dimensional hash space, considering the smoothness of gradient update, a nonlinear function tanh is used to map to [-1, 1] to generate a hash vector in the hash space. The formula is as follows:
[0027] H = tanh(f(x)) (4)
[0028] where H represents a hash vector in a hash space.
[0029] It is further preferred in the present application that the step S3 comprises the following steps:
[0030] S301, in order to maximize the distribution difference between different categories in the hash space, the MMD method is introduced to constrain it. For the hash vector in the space First, it is sampled, and since the calculation of MMD needs two different categories, different categories need to be sampled respectively. For each category y i ∈{1,2,...,C}, the hash vector set of the cth category is defined as follows:
[0031]
[0032] A fixed number of hash vectors are randomly sampled from each category set. For the method of sampling K samples, when the number of samples of a certain category is less than K, we use repeated sampling or padding randomly selected to ensure the number of samples is consistent, and the formula is as follows:
[0033]
[0034] Where the Sample function encapsulates various logics of the above sampling.
[0035] S302, for the set after sampling, it is applied to the calculation of MMD. In order to keep the training smooth and stable, the Gaussian kernel function is first used as the similarity measure index, and its formula definition is as follows:
[0036]
[0037] Where σ is the kernel bandwidth parameter, which controls the smoothness of the kernel function.
[0038] For any category c1 and c2, the sample set is And The unbiased estimate of MMD can be calculated as:
[0039]
[0040] Finally, in order to integrate the distribution difference between all hash pairs, we average the MMD unbiased estimates of all different category pairs to get the overall MMD loss as follows:
[0041]
[0042] Where C is the total number of categories, which is used for normalization to prevent the loss amplitude from being affected by the change of the number of categories.
[0043] A further preferred embodiment of the present application is that the step S4 comprises the following steps:
[0044] S401, self-updating class center establishment and update: in order to optimize the aggregation of the same category samples in the hash space, a learnable center vector is introduced for each active identified category to represent the center position of the category. Initially, each center vector is initialized according to the standard normal distribution:
[0045]
[0046] In each batch b, in order to prevent the interference of class imbalance on center update, a weighted update strategy is adopted. For each hash vector h i , whose category is y i , its weight is defined according to the number of samples in the category: wherein represents the number of samples of the category y i in the current batch, and ∈ represents a minimum constant. Then, the weighted mean of the category in the current batch is obtained by weighted aggregation of all samples:
[0047]
[0048] The momentum update method is adopted, that is, interpolation is performed between the original center c y and the weighted mean value obtained in the current batch , wherein α is a smoothing factor, and the value is close to 1:
[0049]
[0050] S402, calculation of center loss:
[0051] After obtaining the center vector of each category, the center loss is used to make the hash vectors of the same category close to the corresponding category center, so as to improve the intra-class aggregation. For each batch b, the definition of the center loss is as follows:
[0052]
[0053] A further preferred embodiment of the present application is that the step S5 is as follows:
[0054] S5, in order to make the human activity recognition data mapped into the hash space to obtain hash vectors with stronger discriminability and compactness, the hash vectors are constrained by jointly optimizing two objective functions of MMD loss and center loss. As shown below:
[0055]
[0056] wherein λ1 and λ2 represent weight factors of MMD loss and center loss, respectively.
[0057] It is further preferred in the present application that the step S6 is as follows:
[0058] S6, the hash vector in the hash space after optimization, the last optimized hash vector is used for final classification when predicting human activity, as follows:
[0059]
[0060] The present application is based on a wearable human behavior activity recognition method based on hash space optimization, which has the following advantages compared with the prior art:
[0061] (1) The present application abstracts and extracts the original high-dimensional sensor features through a neural network, and then further maps them to a hash space for discrete quantization expression, thereby compressing the continuous high-dimensional vector into a low-dimensional approximate binary hash vector. Compared with the floating-point representation method in the traditional feature space, this method greatly reduces the storage cost and memory occupation of the features, and is especially suitable for resource-limited wearable devices or edge computing terminals.
[0062] (2) The present application introduces the maximum mean difference (MMD) method to finely model and optimize the distribution of different categories of hash vectors. This mechanism effectively expands the distribution distance between classes by comparing the distribution centers and kernel mean differences between classes, and enhances the distinguishability of classes in the hash space. This improves the model's ability to recognize the boundaries between similar behaviors and significantly improves the overall recognition accuracy, especially in multi-class or complex environment classification performance.
[0063] (3) In view of the problem of vector dispersion and inconsistent representation of the same category in the hash space, the present application introduces a center loss optimization method with a self-updating mechanism. By dynamically calculating the center vector of each category during training and guiding the samples of the same category to gradually aggregate around the center, the spatial difference between the intra-class samples is effectively reduced. This mechanism not only improves the organization and consistency of the hash space, but also enhances the sensitivity of the model to changes in action details, making it particularly suitable for micro-motion or continuous action classification tasks in complex human behavior recognition. BRIEF DESCRIPTION OF DRAWINGS
[0064] Figure 1 The flowchart of the present application.
[0065] Figure 2 The overall framework diagram of the present application.
[0066] Figure 3is the comparison of the method of the present application with other methods in four data sets.
[0067] Figure 4 is the confusion matrix graph of the present application in four data sets.
[0068] Figure 5 is the loss and accuracy curve of the present application. DETAILED DESCRIPTION
[0069] In order to make the purpose and embodiment of the present application more clear, the present application is further explained in detail below according to the drawings, and the reference examples are only used to explain the present application, not as the limitation of the present application.
[0070] Example 1
[0071] As shown in Figure 1 and Figure 2 A wearable human activity recognition method based on hash space optimization includes the following steps:
[0072] S1, data preprocessing, processing the original sensor signal data of human activity recognition data set, such as OPPORTUNITY, PAMAP2 and WISDM, into time series data with better robustness;
[0073] S2, feature extraction and hash embedding, extracting features of the processed human activity time series data through neural network method, and embedding into hash space using tanh function;
[0074] S3, MMD optimization, using MMD method to maximize the distribution difference between different categories in hash space for different hash vectors of different categories in hash space;
[0075] S4, Center optimization, using self-updating center loss optimization to reduce the dispersion of the same category in hash space for different hash vectors of the same category in hash space.
[0076] S5, loss calculation, combining MMD loss and center loss to jointly optimize hash vectors in hash space.
[0077] S6, using the optimized hash vectors in hash space to predict different categories of human activity recognition.
[0078] In step S1, the following steps are included:
[0079] S101, data set denoising: because the sensor is worn on the human body, sudden large changes or device shaking will produce a lot of noise, WT method is used for denoising, wavelet transform retains the information of time domain and frequency domain at the same time, thereby minimizing the loss of key information. The key wavelet transform formula is as shown in the formula below:
[0080] x(t) = T(W(x), T) ψ(t) (1)
[0081] Where x represents the original signal, T(W(x), T) represents the threshold function, and ψ(t) represents the wavelet function.
[0082] S102, data standardization: in order to improve the efficiency and stability of subsequent model training, the z-score method is used to standardize the data after denoising, and the original sensor signal is standardized to a unified scale, so that the model can be more stable. The mean of the standardized data is 0 and the standard deviation is 1. The formula is as follows:
[0083]
[0084] Where μ represents the mean of the original data, and σ represents the standard deviation of the original data.
[0085] S103, data division: the data set is divided into training set, validation set and test set according to the ratio of 7:2:1.
[0086] In step S2, the following steps are included:
[0087] S201, feature extraction: the preprocessed human activity recognition data Where T represents the time step, and D represents the feature dimension. Through the method of neural network, a series of nonlinear transformations are performed to map high-dimensional feature representation Where K represents the dimension of the feature space, and the formula is as follows:
[0088] f(x) = σ(Wx + b) (3)
[0089] Where W is the weight matrix, b is the bias term, and σ is the activation function.
[0090] S202, hash embedding: in order to map the dimension of the feature space to the low-dimensional hash space, considering the smoothness of gradient update, a nonlinear function tanh is used to map to [-1, 1] to generate a hash vector in the hash space. The formula is as follows:
[0091] H = tanh(f(x)) (4)
[0092] Where H represents the hash vector in the hash space.
[0093] Step S3 includes the following steps:
[0094] S301, in order to maximize the difference between the distribution of different categories in the hash space, the MMD method is introduced to constrain it. For the hash vector in space First of all, it is sampled, because the calculation of MMD needs two different categories, it needs to be sampled respectively for different categories. For each category y i ∈{1,2,…,C}, the hash vector set of the cth category is defined as follows:
[0095]
[0096] We randomly sample a fixed number of hash vectors from each category set. For the method of sampling K samples, when the number of samples of a certain category is less than K, we use repeated sampling or padding randomly selected to ensure the number of samples is consistent, the formula is as follows:
[0097]
[0098] Where the Sample function encapsulates various logics of the above sampling.
[0099] S302, for the set after sampling, it is applied to the calculation of MMD, in order to keep the training smooth and stable, first of all, the Gaussian kernel function is used as the similarity measure index, its formula definition is as follows:
[0100]
[0101] Where σ is the kernel bandwidth parameter, which controls the smoothness of the kernel function.
[0102] For any category c1 and c2, its sample set is respectively And The unbiased estimate of MMD can be calculated as:
[0103]
[0104] Finally, in order to integrate the distribution difference between all hash pairs, we average the MMD unbiased estimates of all different category pairs to get the overall MMD loss as follows:
[0105]
[0106] Where C is the total number of categories, used for normalization to prevent the change of the number of categories from affecting the loss amplitude.
[0107] Step S4 includes the following steps:
[0108] S401, self-updating class center establishment and update: in order to optimize the aggregation of the same class samples in the hash space, a learnable center vector is introduced for each active identified class to represent the center position of the class. Initially, each center vector is initialized according to the standard normal distribution:
[0109]
[0110] In each batch b, in order to prevent the interference of class imbalance on center update, a weighted update strategy is adopted. For each hash vector h i , whose class is y i , its weight is defined according to the number of samples in the class: wherein represents the number of samples of class y i in the current batch, and ∈ represents a small constant. Then, the weighted aggregation of all samples is obtained to obtain the weighted mean of the class in the current batch:
[0111]
[0112] The momentum update method is adopted, that is, interpolation is performed between the original center c y and the weighted mean obtained in the current batch , wherein α is a smoothing factor, and the value is close to 1:
[0113]
[0114] S402, calculation of center loss:
[0115] After obtaining the center vector of each class, the center loss is used to make the hash vectors of the same class close to the corresponding class center, so as to improve the intra-class aggregation. For each batch b, the definition of the center loss is as follows:
[0116]
[0117] In step S5, the following steps are included:
[0118] S5, in order to make the human activity recognition data mapped to the hash space to obtain hash vectors with stronger discriminability and compactness, the hash vectors are constrained by jointly optimizing the MMD loss and the center loss two objective functions. As follows:
[0119]
[0120] Wherein λ1 and λ2 represent the weight factors of MMD loss and center loss respectively.
[0121] In step S6, the following steps are included:
[0122] S6, the hash vector in the hash space after optimization, the last optimized hash vector is used for final classification when predicting human activity, as follows:
[0123]
[0124] The application provides a human activity recognition method based on hash space optimization, which can significantly improve the discrimination ability between different categories, effectively reduce the distribution dispersion of samples in the same category in the hash space, and thus improve the overall recognition accuracy and generalization performance of the model. By introducing the joint optimization mechanism of maximum mean difference (MMD) and center loss, the model has stronger stability and robustness in complex human activity feature extraction and discrimination.
[0125] Embodiment 2
[0126] In another specific example, in order to verify the performance of the proposed wearable human behavior activity recognition method based on hash space optimization in different data environments, four widely used public human activity recognition datasets are selected as evaluation benchmarks, namely: OPPORTUNITY, PAMAP2, WISDM and UniMiB_SHAR. The above datasets cover different collection environments, sensor types and behavior categories, have strong representativeness, and can effectively verify the universality and robustness of the proposed method.
[0127] In this embodiment, according to the complete processing flow described in embodiment 1, first, the unified data preprocessing is performed on each dataset, including signal denoising, standardization processing and training set division, etc., to ensure the consistency of data quality. Then, the neural network model is constructed to extract features from the time series data, and the high-dimensional feature vector is embedded into the low-dimensional hash space by using the tanh function, so as to significantly reduce the subsequent computational complexity and improve the vector compactness. In order to further improve the recognition effect, the MMD optimization mechanism is introduced to pull apart the distribution distance between different categories, and the center loss mechanism is combined to make the hash vectors of the same category as close to the category center as possible. Finally, the optimized hash vector is classified and predicted to realize efficient and accurate human activity recognition.
[0128] As Figure 3 shown, the accuracy comparison chart of the method of the application and the mainstream recognition model on the four datasets. This embodiment compares three representative deep learning models in current human activity recognition: CNN, DanHAR and GRU-INC. Under the same experimental conditions and data input, the method of the application achieves the best performance on the above datasets.
[0129] The method proposed in this invention exhibits good universality in terms of accuracy, maintaining high accuracy across different types of datasets. Compared to other mainstream human activity recognition methods, the method of this invention also demonstrates superior accuracy. It achieves 94.97% accuracy on the OPPORTUNITY dataset, 94.11% accuracy on the PAMAP2 dataset, 98.93% accuracy on the WISDM dataset, and 92.3% accuracy on the UniMiB_SHAR dataset.
[0130] Furthermore, such as Figure 4 As shown in the figure, the performance of the method of this invention in the identification of specific behavior categories on various datasets is illustrated by using a confusion matrix visualization. Specifically, the OPPORTUNITY dataset contains 17 behavior categories, PAMAP2 contains 12 categories, WISDM contains 6 categories, and UniMiB_SHAR also contains 17 categories. The figure shows that the method can achieve high recognition accuracy in most categories.
[0131] Example 3
[0132] Depend on Figure 5 As shown, this embodiment was designed to further evaluate the impact of the proposed hash space-optimized wearable human behavior activity recognition method on the model training process. The core of this embodiment is to observe the changes in the loss function and the evolution of prediction accuracy during the training and validation processes.
[0133] like Figure 5 As shown, the MMDC method proposed in this invention exhibits extremely high optimization efficiency in the early stages of training. Within the first few epochs of training, the overall loss decreases rapidly and then gradually stabilizes, indicating a significant advantage in convergence speed, and no drastic oscillations or significant overfitting occur during training. Furthermore, the validation set accuracy curve shows that this method achieves high prediction accuracy early in training and completes high-precision convergence in the middle stages, demonstrating that the method effectively improves training efficiency while maintaining generalization ability.
[0134] By introducing a joint mechanism of MMD optimization and center optimization, the method of this invention can more effectively widen the distribution differences between different categories, while aggregating similar samples, thereby forming a more discriminative and highly aggregated vector distribution in the hash space. This allows the method to capture the hash distribution of each human behavior well in the early stages of training and achieve convergence more quickly.
[0135] The above embodiments explain the technical concept, implementation path and potential advantages of the present application in detail, and the content is only the preferred examples of the present application, and cannot be regarded as the limitation of the technical scope of the present application. For ordinary skilled in the art, various forms of equivalent adjustment and technical improvement made without departing from the core idea and essential characteristics of the present application, should be regarded as included in the protection scope of the present application.
Claims
1. A wearable human behavior activity recognition method based on hash space optimization, characterized in that, The method comprises the following steps: S1, data preprocessing, processing the original sensor signal in the wearable human activity behavior data set, removing noise, standardizing the data format, and converting into input data with stable timing structure and expression ability; S2, feature extraction and hash embedding, the preprocessed human activity timing data is extracted by a neural network method, and the features are mapped and embedded into a hash space by a hyperbolic tangent activation function to obtain a compact binary approximate vector representation; S3, MMD optimization, for different hash vectors in the hash space of different categories, the maximum mean difference method is used to measure and enhance the distribution distinguishability of different category hash vectors; S4, Center optimization, for different hash vectors in the hash space of the same category, a self-updating center loss optimization is used to reduce the dispersion of the same category in the hash space; S5, loss calculation, the MMD loss and the center loss are combined to jointly optimize the hash vectors in the hash space; S6, using the optimized hash vectors in the hash space to predict different categories of human activity recognition.
2. The method of claim 1, wherein, The step S1 comprises the following steps: S101, data set denoising: since the sensor is worn on the human body, WT method is used for denoising, wavelet transform retains the information of time domain and frequency domain at the same time, and the key wavelet transform formula is as shown in the formula below: x(t)=∫T(W(x),T)ψ(t) (1) Where x represents the original signal, T(W(x),T) represents the threshold function, and ψ(t) represents the wavelet function; S102, data standardization: the z-score method is used to standardize the data after denoising, and the original sensor signal is standardized to a unified scale, and the standardized data has a mean value of 0 and a standard deviation of 1, and the formula is as follows: Where μ represents the mean value of the original data, and σ represents the standard deviation of the original data; S103, data division: the data set is divided into training set, validation set and test set according to the proportion of 7:2:
1. 3.The method of claim 2, wherein, The step S2 comprises the following steps: S201、Feature extraction: the pre-processed human activity recognition data wherein T represents a time step, D represents a feature dimension, and the method of neural network is mapped to a high-dimensional feature representation through a series of nonlinear transformations wherein K represents the dimension of the feature space, and the formula is as follows: f(x)=σ(Wx+b) (3) Where W is the weight matrix, b is the bias term, and σ is the activation function; S202, hash embedding: using a nonlinear function tanh to map to [-1, 1] to generate hash vectors in the hash space, and the formula is as follows: H=tanh(f(x)) (4) Where H represents the hash vector in the hash space.
4. The method of claim 3, wherein, The step S3 comprises the following steps: S301、In the hash space, the distribution difference between different categories is maximized, and the MMD method is introduced to constrain it. For the hash vector in the space First, it is sampled. Since the calculation of MMD requires two different categories, different categories need to be sampled respectively, and for each category y i ∈{1,2,...,C}, the hash vector set of the cth category is defined as follows: A fixed number of hash vectors are randomly sampled from each category set, and for the method of sampling K samples, when the number of samples of a certain category is less than K, repeated sampling or filling random selection is adopted to ensure the consistency of the number of samples, and the formula is as follows: Where the Sample function encapsulates various logics of the above sampling; S302, for the set after sampling, it is applied to the calculation of MMD, and a Gaussian kernel function is used as a similarity measurement index, and its formula definition is as follows: Where σ is the kernel bandwidth parameter, which controls the smoothness of the kernel function; For any classes c1 and c2, whose sample subsets are respectively and An unbiased estimate of the MMD is computed as: The MMD unbiased estimates of all different category pairs are averaged to obtain the overall MMD loss as follows: Where C is the total number of categories, and is used for normalization to prevent the loss magnitude from being affected by the number of categories.
5. The method of claim 4, wherein, The step S4 includes the following steps: S401, self-updating class center establishment and update: a learnable center vector is introduced for each class identified for each activity The center vector is initialized according to a standard normal distribution at the beginning, and is used to represent the center position of the class. In each batch b, a weighted update strategy is adopted for each hash vector h. i Its category is y i The weights are defined based on the number of samples within each category. in Indicates category y in the current batch i The number of samples is given, where ∈ represents the minimum constant. Then, all samples are weighted and aggregated to obtain the weighted mean of that category in the current batch: The momentum updating method is adopted, that is, the original center c y is interpolated between the weighted mean value obtained in the current batch , where a is a smoothing factor, and the value is close to 1. S402, calculation of the center loss: After obtaining the center vector of each category, the center loss is used to make the hash vectors of the same category close to the corresponding category center. For each batch b, the definition of the center loss is as follows:
6. The method of claim 5, wherein, The step S5 includes the following steps: S5, constraint the hash vector by jointly optimizing the MMD loss and the center loss two objective functions as follows: Where λ1 and λ2 represent the weight factors of the MMD loss and the center loss respectively.
7. The method of claim 6, wherein, The step S6 includes the following steps: S6, the hash vector in the hash space after optimization, the last optimized hash vector is used for final classification when predicting human activities as follows:
Citation Information
Patent Citations
Cross-platform burying-point-free data acquisition and intelligent circle selection rule generation method and system
CN119537158A
Indoor non-contact human activity recognition method and system
US20230055065A1