Human body activity identification method, system and equipment
Through multi-sensor fusion technology, multimodal fusion of data level, feature level and decision-making level is carried out, solving the problem of low accuracy in human body activity recognition and achieving more efficient and accurate recognition effects.
Patent Information
- Application Number
- CN202510096879.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-19
- Filing Date
- 2025-01-21
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the accuracy of human body activity recognition is low. Especially in the face of different human body postures, complex motion patterns and environmental interference, the single-modal recognition method is poor in effect and is difficult to meet the practical application needs.
Multi-sensor fusion technology is adopted to obtain human joint motion data, joint relative motion data and magnetic source excitation data to perform multimodal fusion at data level, feature level and decision-making level, including data splicing, feature threading and decision-making fusion, and use D-S evidence theory, weight averaging or voting methods to improve identification accuracy.
It improves the accuracy and flexibility of human activity recognition, reduces calculation complexity and energy consumption, reduces response time, and achieves more efficient and accurate recognition.
Smart Images

Figure CN120279592A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent titled "A Human Activity Recognition Method, System and Device" with the application number 202411660955.X and filed with the Chinese Patent Office on November 19, 2024. The entire content thereof is incorporated herein by reference. Technical Field
[0002] The present invention relates to the field of terminals and communication technologies, and particularly to a human activity recognition method, system and device. Background Art
[0003] With the rapid development of intelligent sensing and wearable technologies, human activity recognition (HAR) has gradually become an important research field in modern intelligent technologies. By real-time monitoring and analyzing the motion data of the human body, human activity recognition technology has shown great application potential in fields such as health monitoring, sports monitoring, rehabilitation training, and human-computer interaction. Especially in the context of an aging society and the growing demand for health management, technologies that can accurately and efficiently identify and judge human activities are of great significance for the construction of intelligent health management systems.
[0004] The technology of human activity recognition can reveal the activity status and physiological information of users, providing many conveniences for daily life and work. For example, in the field of medical health, since certain diseases are associated with specific human movements or functional changes (such as Parkinson's disease, trauma resuscitation, etc.), human activity recognition technology can identify the regularity of specific actions of patients (including movement trajectories, action amplitudes, and frequencies), thus providing guidance for the disease rehabilitation process.
[0005] In addition, human activity recognition technology can reveal the patterns and regularities of the human body in a normal state. When the human body is in an abnormal state, this technology can capture these abnormal states, so as to take timely measures to prevent the occurrence of diseases. In the field of physical exercise, human activity recognition technology can track the exercise duration, intensity, and current physical function status of users, providing health-related reports and suggestions for users, and helping to formulate appropriate exercise plans. In the field of smart home, users can directly control smart devices through gestures or actions, and at the same time, smart devices can identify users through behavioral characteristics, and then match personalized working modes for different users to improve the product experience and quality of life.
[0006] Therefore, there is an urgent need to provide a reliable human activity recognition solution. Summary of the Invention
[0007] The purpose of the present invention is to provide a human activity recognition method, system and device, which is used to solve the problem of low accuracy of human activity recognition in the prior art.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] In a first aspect, the present invention provides a human activity recognition method, the method comprising:
[0010] Obtaining raw data; the raw data at least includes human joint motion data, human joint relative motion data, and magnetic source excitation data;
[0011] Preprocessing the raw data to obtain preprocessed data, and extracting features from the preprocessed data to obtain target features;
[0012] Performing data-level fusion on the raw data, performing feature-level fusion on the target features, and performing decision-level fusion on the feature fusion model to obtain fusion information;
[0013] Performing human activity recognition based on the fusion information.
[0014] Optionally, the preprocessing at least includes data conversion, noise filtering, registration, compression, interpolation, and normalization; the preprocessed data includes first preprocessed data and second preprocessed data; the target features include a first target feature set and a second target feature set;
[0015] Extracting features from the preprocessed data to obtain target features, specifically including:
[0016] Extracting features from the first preprocessed data to obtain a first target feature set; the first target feature set is the skeletal joint motion feature of the target action;
[0017] Extracting features from the second preprocessed data to obtain a second target feature set; the second target feature set is the skeletal joint relative motion feature of the target action.
[0018] Optionally, extracting features from the preprocessed data to obtain target features, specifically including:
[0019] Performing IMU attitude solution, IMU data feature extraction, and magnetic field data feature extraction on the preprocessed data;
[0020] Performing data dimensionality reduction on the extracted feature data.
[0021] Optionally, performing IMU attitude solution, IMU data feature extraction, and magnetic field data feature extraction on the preprocessed data, specifically including:
[0022] Performing IMU attitude solution on the preprocessed data and expressing it in the form of quaternions and Euler angles;
[0023] Extract IMU data features from the human joint motion data in the preprocessed data to obtain the triaxial angular velocity modulus, the difference between the triaxial acceleration modulus and the gravitational acceleration modulus, and the magnetic field intensity;
[0024] Extract IMU data features from the human joint relative motion data in the preprocessed data, and extract features from the magnetic field data; the magnetic field data is used to represent the relative position relationship with the magnetic source, and the second norm of the triaxial magnetometer data is used to represent the magnitude of the magnetic field intensity;
[0025] The data reduction of the extracted feature data specifically includes:
[0026] Perform PCA data reduction on the linear data in the extracted feature data. Through linear transformation, project the transformed data into a new coordinate system;
[0027] Perform T-SNE data reduction on the non-linear data in the extracted feature data, project similar points in the high-dimensional space into the low-dimensional space, and maintain the local structure.
[0028] Optionally, the data-level fusion of the original data specifically includes:
[0029] Stitch and integrate the first preprocessed data and the second preprocessed data in the first sequence stitching method, the first structure stitching method, and the first pixel stitching method;
[0030] Send the integrated data into the training model according to the sliding window length for data-level fusion;
[0031] The first sequence stitching method includes:
[0032] Input the serial data of the first preprocessed data and the second preprocessed data into the model. The stitching methods of multiple groups of different joint data include intra-group stitching and inter-group stitching, and finally input in the form of a 1*9 sequence, and the sequence length is the sum of the sequence lengths of the first preprocessed data and the second preprocessed data; different joint data includes 6 groups;
[0033] The first structure stitching method includes:
[0034] Stitch the first preprocessed data and the second preprocessed data of the same group into a 3*3 matrix, and stitch the matrices of different groups to form a 3*3*6 structure;
[0035] The first pixel stitching method includes:
[0036] Concatenate the first preprocessed data of all groups with the second preprocessed data into a 6*9 pixel matrix, and expand it to a 9*9 pixel matrix by zero-padding according to the convolution requirement.
[0037] Optionally, the feature-level fusion of the target features specifically includes:
[0038] Concatenate and integrate the first target feature set and the second target feature set in a second sequence concatenation manner, a second structure concatenation manner, and a second pixel concatenation manner for feature-level fusion;
[0039] The second sequence concatenation manner includes:
[0040] Perform a second sequence concatenation on the first target feature set and the second target feature set. The concatenation manner of the second sequence concatenation includes intra-group concatenation and inter-group concatenation, and finally input in the form of a 1*M sequence, where M is the sum of the data lengths of the first preprocessed data and the second preprocessed data;
[0041] The second structure concatenation manner includes:
[0042] Concatenate the joint feature data of the same group into a 3*3 matrix, and then concatenate the matrices of different groups to form a structure of N*P*6. Reshape M into N*P, and fill the vacancies with zeros;
[0043] The second pixel concatenation manner includes:
[0044] Concatenate the joint feature data of all groups into an M*M pixel matrix, and fill the vacancies with zeros.
[0045] Optionally, the decision-level fusion of the feature fusion model specifically includes:
[0046] Train a first human activity recognition training model based on joint motion according to the first preprocessed data, the first target feature set, and the action labels after One-hot encoding;
[0047] Determine the prediction result of the first human activity recognition training model as the first decision,
[0048] Train a second human activity recognition training model based on joint motion according to the second preprocessed data, the second target feature set, and the action labels after One-hot encoding;
[0049] Determine the prediction result of the second human activity recognition training model as the second decision,
[0050] Obtain the final decision through training fusion of the first decision and the second decision by a preset method; the preset method at least includes: D-S evidence theory, weighted average, or voting;
[0051] According to the first decision and the second decision, assign basic probabilities to each hypothesis, reflecting the degree of support of the data for the corresponding hypothesis;
[0052] Synthesize the evidence of all decisions and calculate the final basic probability assignment using Dempster's rule;
[0053] Based on the fused evidence, calculate the confidence or likelihood of each hypothesis;
[0054] Obtain the target result according to the maximum confidence or maximum likelihood.
[0055] Compared with the prior art, a human activity recognition method provided by the present invention. By acquiring original data; preprocessing the original data to obtain preprocessed data, and extracting features from the preprocessed data to obtain target features; performing data-level fusion on the original data, performing feature-level fusion on the target features, and performing decision-level fusion on the feature fusion model to obtain fusion information; performing human activity recognition based on the fusion information. The present invention uses multiple fusion models to process the obtained multi-class data, and performs data-level fusion, feature-level fusion and decision-level fusion respectively. The input data is integrated by means of sequence splicing, structure splicing and pixel splicing of the preprocessed data and the feature data sequences. The decision-level fusion model further fuses the first decision result obtained by training and predicting with the first preprocessed data and the first decision data, and the second decision result obtained by training and predicting with the second preprocessed data and the second decision data to obtain the final decision, which can improve the flexibility, anti-interference ability and prediction accuracy of the prediction model, and can more efficiently and accurately perform human activity recognition on the target user; moreover, the small throughput reduces the computational complexity and the occupied hardware resources of the human activity recognition algorithm, thereby reducing the energy consumption and response time of the human activity recognition method.
[0056] In a second aspect, the present invention provides a human activity recognition system, the system comprising:
[0057] An inertial measurement device, a magnetic field measurement device, an electromagnetic drive device and a terminal device;
[0058] The inertial measurement device is integrated with the magnetic field measurement device and establishes a communication connection with the terminal device; the inertial measurement device and the magnetic field measurement device are placed on the target part of the human body; the electromagnetic drive device is a passive device; the electromagnetic drive device is arranged on the back of the human body and is spaced from the inertial measurement device and the magnetic field measurement device by a preset distance;
[0059] The terminal device is used to receive the original data; the original data at least includes human joint movement data, human joint relative movement data and magnetic source excitation data;
[0060] Preprocess the original data to obtain preprocessed data, and extract features from the preprocessed data to obtain target features;
[0061] Perform data-level fusion on the original data, perform feature-level fusion on the target features, and perform decision-level fusion on the feature fusion model to obtain fusion information;
[0062] Perform human activity recognition based on the fusion information.
[0063] In a third aspect, the present invention provides a human activity recognition device, which includes:
[0064] A memory, a processor, and a communication interface coupled to the processor; a computer program that can be run by the processor is stored on the memory; when the processor runs the computer program, it executes the above-mentioned human activity recognition method.
[0065] The technical effects achieved by the system-level solution provided in the second aspect and the device-level solution provided in the third aspect are the same as those of the method-level solution provided in the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0067] Figure 1 is a schematic flow chart of a human activity recognition method provided by the present invention;
[0068] Figure 2 is a schematic diagram of the overall implementation process of a human activity recognition method provided by the present invention;
[0069] Figure 3 is a schematic diagram of data-level fusion provided by the present invention;
[0070] Figure 4 is a schematic diagram of feature-level fusion provided by the present invention;
[0071] Figure 5 is a schematic diagram of decision-level fusion provided by the present invention;
[0072] Figure 6 is a schematic diagram of the structure of a human activity recognition system provided by the present invention;
[0073] Figure 7 is a schematic diagram of a human activity recognition device provided by the present invention.
[0074] Reference Signs:
[0075] 1- Inertial measurement equipment, 2- Magnetic field measurement equipment, 3- Electromagnetic transmission equipment, 4- Terminal equipment. DETAILED DESCRIPTION
[0076] In order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and their order is not limited. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0077] It should be noted that, in the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0078] In the present invention, "at least one" means one or more, and "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b, c can be single or multiple.
[0079] Traditional human activity recognition methods mainly rely on a single type of sensor or a single-modal data source, such as an accelerometer or gyroscope. Although these methods perform well in static and simple action recognition, their recognition accuracy and reliability are often limited for dynamic activities involving complex skeletal joint movements. Especially in the face of different human postures, complex movement patterns, and environmental interference, the effect of single-modal recognition methods is poor and it is difficult to meet the needs of practical applications.
[0080] In recent years, with the development of inertial sensor technology, multi-sensor fusion technology has gradually become an effective means to solve this problem. By fusing multiple sensor data, especially inertial sensors such as accelerometers, gyroscopes and magnetometers, it is possible to capture the multi-dimensional information of human motion more comprehensively and improve the accuracy of recognition. However, the existing multi-sensor fusion methods still have certain limitations in feature extraction and multimodal data fusion, and it is difficult to fully reflect the skeletal joint motion characteristics and relative motion characteristics between joints in complex human activities. Therefore, how to deeply mine the motion information of skeletal joints on the basis of multimodal data fusion, and then improve the accuracy of human activity recognition, is still a technical problem that needs to be solved urgently in this field.
[0081] In order to solve the defects in the prior art, the present invention provides a method, system and device for human activity recognition. Next, the scheme provided by the embodiment of this specification is described in conjunction with the accompanying drawings:
[0082] Example 1
[0083] Figure 1 A flow chart of a human activity recognition method provided by the present invention is shown in FIG. Figure 1 As shown, the execution subject of the process may be a terminal device, which obtains measurement data of each measurement device in the system and performs human activity recognition based on the obtained data. Specifically, the following steps may be included:
[0084] Step 110: Obtain original data.
[0085] The raw data may include at least human joint motion data, human joint relative motion data and magnetic source excitation data. These data may be data collected by the system in Embodiment 2 of the present invention.
[0086] Step 120: preprocessing the original data to obtain preprocessed data, and extracting features from the preprocessed data to obtain target features;
[0087] Preprocessing can include raw data conversion, noise filtering, registration, compression, interpolation, normalization and other processing.
[0088] The preprocessed data includes first preprocessed data and second preprocessed data;
[0089] Feature extraction can include various feature collection and data dimensionality reduction. Feature collection includes IMU attitude solution, IMU data feature extraction, magnetic field data feature extraction, etc.; data dimensionality reduction is the process of extracting effective features from the above features for human activity recognition. The target features include a first target feature set and a second target feature set;
[0090] When performing feature extraction, the first preprocessed data can be used for feature extraction to obtain a first target feature set; the first target feature set is the skeletal joint movement feature of the target action.
[0091] Perform feature extraction on the second preprocessed data features to obtain a second target feature set; the second target feature set is the relative skeletal joint movement feature of the target action.
[0092] Step 130: Perform data-level fusion on the original data, perform feature-level fusion on the target features, and perform decision-level fusion on the feature fusion model to obtain fusion information.
[0093] In this step, more comprehensive and multi-dimensional multi-modal fusion features are fused from the data level, feature level, and decision level to achieve efficient multi-modal recognition of human activities.
[0094] Step 140: Perform human activity recognition based on the fusion information.
[0095] Figure 1 The method in [description] obtains the original data; preprocesses the original data to obtain preprocessed data, and performs feature extraction on the preprocessed data to obtain target features; performs data-level fusion on the original data, performs feature-level fusion on the target features, and performs decision-level fusion on the feature fusion model to obtain fusion information; performs human activity recognition based on the fusion information. The present invention uses multiple fusion models to process the obtained multi-class data, and performs data-level fusion, feature-level fusion, and decision-level fusion respectively. The input data is integrated by means of sequence splicing, structure splicing, and pixel splicing of the preprocessed data and feature data. The decision-level fusion model further fuses the first decision result obtained by training and predicting with the first preprocessed data and the first decision data, and the second decision result obtained by training and predicting with the second preprocessed data and the second decision data to obtain the final decision, which can improve the flexibility, anti-interference ability, and prediction accuracy of the prediction model, and can more efficiently and accurately perform human activity recognition on the target user; moreover, the small throughput reduces the computational complexity and occupied hardware resources of the human activity recognition algorithm, thereby reducing the energy consumption and response time of the human activity recognition method.
[0096] Based on Figure 1 the method of [description], the embodiments of this specification also provide some specific implementation manners of this method, which will be described below.
[0097] Figure 1 When the solution in [description] is specifically implemented, the implementation process is as Figure 2 shown, and may include the following specific implementation steps:
[0098] The extracted features include joint motion information and relative joint motion information. The original information of the relative motion depends on the magnetic vector information captured by the magnetometer, such as Figure 2 the magnetic source excitation in Figure 2 ; data preprocessing is performed on the extracted data, feature extraction is performed on the preprocessed data, then data-level fusion is performed on the preprocessed data, feature-level fusion is performed on the extracted features, and decision-level fusion is performed on the feature fusion model corresponding to the features.
[0099] Further, in step 120, the feature extraction may include:
[0100] Performing IMU attitude solution, IMU data feature extraction, and magnetic field data feature extraction on the preprocessed data;
[0101] Performing data dimensionality reduction on the extracted feature data.
[0102] More specifically, the performing IMU attitude solution, IMU data feature extraction, and magnetic field data feature extraction on the preprocessed data may specifically include:
[0103] Performing IMU attitude solution on the preprocessed data and expressing it in the form of quaternions and Euler angles;
[0104] Performing IMU data feature extraction on the human joint motion data in the preprocessed data to obtain the triaxial angular velocity modulus, the difference between the triaxial acceleration modulus and the gravitational acceleration modulus, and the magnetic field intensity;
[0105] Performing IMU data feature extraction on the human joint relative motion data in the preprocessed data and performing feature extraction on the magnetic field data; the magnetic field data is used to represent the relative position relationship with the magnetic source, and the second norm of the triaxial magnetometer data is used to represent the magnitude of the magnetic field intensity;
[0106] The performing data dimensionality reduction on the extracted feature data specifically includes:
[0107] Performing PCA dimensionality reduction on the linear data in the extracted feature data, and through linear transformation, projecting the transformed data into a new coordinate system;
[0108] Performing T-SNE dimensionality reduction on the non-linear data in the extracted feature data, projecting the similar points in the high-dimensional space into the low-dimensional space, and maintaining the local structure.
[0109] Among them, linear data is suitable for dimensionality reduction using PCA (Principal Component Analysis), and non-linear data uses T-SNE (T-Distributed Stochastic Neighbor Embedding) to achieve the mapping of data from a high-dimensional space to a low-dimensional space.
[0110] In the above steps, feature extraction may include various types of feature collection and data dimensionality reduction.
[0111] Feature collection may include IMU attitude solution, IMU data feature extraction, magnetic field data feature extraction, etc.;
[0112] In this step, for the first type of preprocessed data, attitude information at the joint can be obtained by resolving acceleration and angular velocity data; IMU attitude solution may include Euler angles and quaternions. Euler angles include three rotations, and the orientation of a rigid body is specified based on these three rotations. These three rotations are around the x-axis, y-axis, and z-axis respectively, namely Roll, Pitch, and Yaw.
[0113] IMU data feature extraction includes features such as the modulus of the three-axis angular velocity, the difference between the modulus of the three-axis acceleration and the modulus of the gravitational acceleration, and the magnetic field strength; the modulus of the three-axis angular velocity represents the change in the overall rotation amplitude of the joint; while the difference between the second-order norm of the three-axis acceleration and the modulus of the gravitational acceleration represents the overall movement amplitude of the joint: for the relative movement information of skeletal joints, the magnetic field data represents the relative position relationship with the magnetic source, and the second-order norm of the three-axis magnetometer data can intuitively represent the magnitude of the magnetic field strength; the above features can directly represent the motion features of the human body. More data features such as: mean, median, root mean square, quartile, inter-axis correlation coefficient, Pearson correlation coefficient, wavelet transform, zero crossing rate, skewness, kurtosis, spectral energy, standard deviation, median absolute deviation, sample entropy, and autoregressive coefficient, etc. The calculation formulas for these data features are all prior arts, and this specification does not give specific descriptions thereof; Next, the calculation of the key data features adopted by the present invention will be described:
[0114] The calculation of Euler angles from acceleration is as described in Equation Group (1):
[0115]
[0116] where a x represents the acceleration value output on the x-axis, a y represents the acceleration value output on the y-axis, a z represents the acceleration value output on the z-axis. The formula for calculating Euler angles from angular velocity is as Equation Group (2):
[0117]
[0118] Where, n represents the Euler angles at the nth moment, Δt represents the minimum sampling time, dr / dt represents the derivative of the roll angle with respect to time, dp / dt represents the derivative of the pitch angle with respect to time, and dy / dt represents the derivative of the yaw angle with respect to time.
[0119] The Euler angles after the fusion of acceleration and gyroscope information data are as shown in Equation Group (3):
[0120]
[0121] Where, roll acc represents the roll angle calculated from the acceleration data, roll gyro represents the roll angle calculated from the angular velocity data, pitch acc represents the pitch angle obtained from the acceleration data, pitch gyro represents the pitch angle obtained from the angular velocity data, yaw gyro represents the yaw angle obtained from the angular velocity data.
[0122] Quaternions can represent the rotation of an object about any vector axis, including 1 real part and 3 imaginary parts. They are different representations of Euler angles and have different ways of calculating and updating the attitude. Equation Group (4) is the conversion formula between quaternions and Euler angles:
[0123]
[0124] Q(q i , q j , q k , q r ) = q i + q j + q k j + q r k (5)
[0125] Where q i , q j , q k , q r are real numbers, which means that quaternions are imaginary numbers with three imaginary parts and can be regarded as four-dimensional complex spaces. The above quaternions can be written as Equation (6), where s represents real numbers and v represents vectors.
[0126]
[0127] The magnitude of the three-axis angular velocity |gyr| represents the overall change in the rotation amplitude of the joint; while the difference between the second norm of the three-axis acceleration and the magnitude of the gravitational acceleration represents the overall movement amplitude of the joint, as shown in Equation (7):
[0128] ||g|| = ||a acc || — gG | (7)
[0129] Among them, g represents the overall motion amplitude of the joint, and a acc represents the modulus of the acceleration reading, and g G represents the acceleration due to gravity.
[0130] For the second preprocessed data, the magnetic field data expresses the relative position relationship with the magnetic source, and the second-order norm of the three-axis magnetometer data can intuitively express the magnitude of the magnetic field strength, as shown in (8);
[0131]
[0132] Among them, ||mag|| represents the magnetic field strength, and mag x represents the output on the x-axis of the magnetometer, and mag y represents the output on the y-axis of the magnetometer, and mag z represents the output on the z-axis of the magnetometer.
[0133] The above features can directly express the motion features of the human body.
[0134] The purpose of data dimensionality reduction is to extract the most effective feature set from numerous features for human activity recognition. Linear data is suitable for dimensionality reduction using PCA (Principal Component Analysis). Through linear transformation, the data is projected into a new coordinate system. The principal components are sorted according to the variance of the data, and a larger variance represents more original data information. By selecting the first few principal components, dimensionality reduction can be achieved; for non-linear data, T-SNE (T-Distributed Stochastic Neighbor Embedding) is used. It projects similar points in the high-dimensional space into the low-dimensional space, maintaining the local structure while avoiding data overlap after dimensionality reduction.
[0135] After the above steps, the original data, the first preprocessed data, the first target feature set, the second preprocessed data, and the second target feature set are obtained. These data achieve the classification and recognition task of human activities through data fusion.
[0136] Data fusion can be divided into data-level fusion, feature-level fusion, and decision-level fusion. Next, the fusion at these three levels will be explained separately:
[0137] Data-level fusion. Data-level fusion is the fusion of the first preprocessed data and the second preprocessed data at the model input data level. The first preprocessed data and the second preprocessed data are spliced and integrated in the form of sequences, structures, and pixels. Next, Figure 3 the splicing methods of sequences, structures, and pixels at the data-level fusion level will be explained respectively, asFigure 3 As shown below:
[0138] The first sequence splicing method: The serial data of the first preprocessed data and the second preprocessed data are input into the model. The splicing method of 6 groups of different joint data involves in-group splicing and inter-group splicing, and finally is input in the form of a 1*9 sequence. The sequence length is the sum of the sequence lengths of the first preprocessed data and the second preprocessed data.
[0139] The first structure splicing method: The first preprocessed data and the second preprocessed data in the same group are spliced into a 3*3 matrix, and then the matrices of different groups are spliced together to form a 3*3*6 structure;
[0140] The first pixel splicing method: The first preprocessed data and the second preprocessed data of all groups are spliced into a 6*9 pixel matrix, and are padded with zeros to 9*9 according to the convolution requirements.
[0141] The data-level fusion process is to integrate the first preprocessed data and the second preprocessed data together by splicing, and send them into the training model according to the sliding window length. The time series models involve NN\LSTM\Transformer\GRU, etc., and the structure and pixel splicing models involve CNN\GANs. Each action label is in the form of One-hot encoding as the output of the training model.
[0142] As an optional implementation, the sliding window time can be set to 2-10s, 3s for recognition tasks of daily activities such as walking and running, and 5s for complex or long-lasting activities such as standing up and sitting down, going up and down stairs, etc.;
[0143] As an optional implementation, the sliding window step size can be set to 1 / 4-1 / 2 of the window.
[0144] Feature-level fusion is the fusion of the first target feature and the second target feature at the model input data level. The first target feature and the second target feature are spliced and integrated in the ways of sequence, structure, and pixel. Next, combined with Figure 4 The splicing methods of sequence, structure, and pixel at the feature-level fusion level are described respectively, as Figure 4 shown below:
[0145] The second sequence splicing method: The first target feature and the second target feature. The splicing method of 6 groups of different joint feature data involves in-group splicing and inter-group splicing, and finally is input in the form of a 1*M sequence, where M is the sum of the data lengths of the first preprocessed data and the second preprocessed data;
[0146] The second structure splicing method is to splice the joint feature data of the same group into a 3*3 matrix, and then splice the matrices of different groups together to form an N*P*6 structure, which reshapes M into N*P and fills the vacancies with zeros.
[0147] The third pixel splicing method is to splice the joint feature data of all groups into an M*M pixel matrix and fill the vacancies with zeros.
[0148] As an alternative implementation, the feature-level fusion process integrates the first target feature and the second target feature by splicing and feeds them into the training model according to the sliding window length. The time series model involves NN, LSTM, Transformer, GRU, etc., and the structure and pixel splicing models involve CNN and GANs. Each action label is in the form of One-hot encoding as the output of the training model.
[0149] As an alternative implementation, the sliding window time and step size are the same as described above.
[0150] For the decision-level fusion process, it can be combined with Figure 5 for illustration:
[0151] The first preprocessed data / first target feature and the action label after One-hot encoding are used to train the human activity recognition training model based on joint motion, and the prediction result is denoted as the first decision;
[0152] The second preprocessed data / second target feature and the action label after One-hot encoding are used to train the human activity recognition training model based on joint motion, and the prediction result is denoted as the second decision;
[0153] The first decision data and the second decision data are further trained and fused through methods such as D-S evidence theory, weighted average, and voting to obtain the final decision.
[0154] Preferably, the steps of the D-S evidence reasoning method may include:
[0155] Assume that the space Ω is the set of all possible classification or target states (such as motion states: walking, running, standing, etc.);
[0156] According to the first decision and the second decision data, assign the basic probability m(A) to each hypothesis, reflecting the degree of support of this data for hypothesis A;
[0157] Synthesize the evidence of all decisions and use Dempster's rule to calculate the final basic probability assignment;
[0158] Based on the fused evidence, calculate the confidence (Bel) or likelihood (Pl) of each hypothesis, and then make a final judgment according to the maximum confidence or maximum likelihood.
[0159] Among them, the D-S evidence theory (Dempster-Shafer) allows evidence from multiple sources to be aggregated through a combination method to improve the accuracy of decision-making. Different from probability theory, it can handle partial confidence and can deal with no knowledge or conflicting evidence. The specific process is as follows:
[0160] First, it is necessary to define the confidence of each piece of evidence (each information source, such as each sensor data). The basic probability of a hypothesis A is m(A), which satisfies formula (9):
[0161]
[0162] Among them, m(φ) represents that the probability of the empty set is 0, and Ω is the hypothesis space, representing the set of all possible conclusions, that is, all human activity states: walking, running, standing, etc.
[0163] The belief function is used to quantify the credibility of a set A, representing the sum of all subsets that can fully support A. Its definition is formula (10):
[0164]
[0165] Among them, Bel(A) represents the belief function of event A, and m(B) represents that the basic probability of hypothesis B is m(B).
[0166] The likelihood function represents the possibility of a set A, quantifying the evidence that does not conflict with A, and is defined as:
[0167]
[0168] Among them, Pl(A) represents the likelihood function of A, and Bel(A′) represents the belief function of A′.
[0169] For the first decision-making and second decision-making data, evidence regarding the hypothesis space Ω is given respectively, and their basic probability assignments are called m1 and m2. These evidences are fused in the following way:
[0170]
[0171] Among them, K represents the degree of conflict, and is defined as
[0172]
[0173] This formula (13) is used to calculate the combined result of the two basic probability assignments m1 and m2 given by the sensors. If K = 1, it means that the evidence given by the two sensors is completely conflicting and cannot be fused.
[0174] For more complex hybrid fusion, the above prediction results are further used to obtain the final decision through methods such as voting, confidence, and weighted averaging.
[0175] Embodiment 2
[0176] Based on the same idea, the present invention also provides a human activity recognition system, as Figure 6 shown, the system may include:
[0177] Inertial measurement device 1, magnetic field measurement device 2, electromagnetic drive device 3, and terminal device 4;
[0178] The inertial measurement device 1 is integrated with the magnetic field measurement device 2 and establishes a communication connection with the terminal device 4; the inertial measurement device 1 and the magnetic field measurement device 2 are placed on the target parts of the human body; the electromagnetic drive device 3 is a passive device; the electromagnetic drive device 3 is arranged on the human back and is preset with a distance from the inertial measurement device 1 and the magnetic field measurement device 2;
[0179] The terminal device 4 is used to receive the original data; the original data at least includes human joint movement data, human joint relative movement data, and magnetic source excitation data;
[0180] Preprocess the original data to obtain preprocessed data, and extract features from the preprocessed data to obtain target features;
[0181] Perform data-level fusion on the original data, perform feature-level fusion on the target features, and perform decision-level fusion on the feature fusion model to obtain fusion information;
[0182] Perform human activity recognition based on the fusion information, as Figure 6 shown, and further illustrate the system with examples:
[0183] For example: The inertial measurement device 1 and the magnetic field measurement device 2 in the above system can be integrated together, a total of 6 groups, and establish a communication connection with the terminal device 4; the electromagnetic drive device 3 is a passive device.
[0184] The inertial measurement device 1 and the magnetic field measurement device 2 are placed on key parts of the human body such as the wrist, heel, head, and back.
[0185] The electromagnetic drive device 3 is arranged on the back and maintains a distance of at least 10 cm from the inertial measurement device 1 and the magnetic field measurement device 2 on the back.
[0186] The terminal device 4 is used to receive the original data.
[0187] Preferably, the inertial measurement device 1 can be an accelerometer and a gyroscope; the magnetic field measurement device 2 is a magnetometer;
[0188] Preferably, the electromagnetic drive device 3 is an electromagnetic coil or a permanent magnet. For the material of the electromagnetic drive device 3, copper is preferred for the coil, and rubidium magnet is preferred for the permanent magnet.
[0189] The terminal device 4 is a host computer or a terminal device 4 such as a smart phone.
[0190] For the communication among the inertial measurement device 1, the magnetic field measurement device 2 and the terminal device 4, the hardware system transmits signals through the wireless Bluetooth technology to meet the activity requirements of users indoors and outdoors. The inertial measurement device 1, the magnetic field measurement device 2 and the terminal device 4 work synchronously, collect and upload human activity data in real time, and then perform calculations through the cloud platform to determine the activity state in real time. Specifically, the accelerometer, gyroscope and magnetometer in this step are collectively referred to as the MARG sensing unit. The hardware system transmits signals through the wireless Bluetooth technology to meet the activity requirements of users indoors and outdoors. Multiple MARG sensing units work synchronously with a host computer or a terminal device 4 such as a smart phone, collect and upload human activity data in real time, and then perform calculations through the cloud platform to determine the activity state in real time.
[0191] The first data received by the terminal device 4 is inertial data, including three-axis acceleration data and three-axis gyroscope data, and the second data is three-axis magnetic data, which can reflect the motion information of different parts of the human body, and further determine the motion trajectory through the attitude resolution determination module. By analyzing the motion characteristics of different parts, different human actions can be distinguished; the magnetic data reflects the relative position change between the magnetometer and the magnetic source, representing the position relationship of the human hand, foot, head and back relative to the magnetic source (located on the back). This information is also applicable to the recognition of human activities. Combining these two characteristics can process and distinguish target actions more precisely.
[0192] The human activity recognition method provided by the present invention integrates inertial data and magnetic data, which are collected by the aforementioned hardware system. The inertial data includes three-axis acceleration data and three-axis gyroscope data, which can reflect the motion information of different parts of the human body, and further determine the motion trajectory through the attitude resolution determination module. By analyzing the motion characteristics of different parts, different human actions can be distinguished; the magnetic data represents the position relationship of the human hand, foot, head and back relative to the magnetic source. This information is also applicable to the recognition of human activities. Combining these two characteristics can process and distinguish target actions more precisely.
[0193] In the technical solution provided by the present invention, an inertial sensor module integrating an accelerometer, a gyroscope, and a magnetometer is adopted, and supplemented by a magnetic source device, which jointly act on the accurate capture of human activities. Specifically, this method involves feature extraction at two levels: one is to conduct a detailed analysis of the motion characteristics of the bones and joints themselves in the target action; the other is to deeply explore the relative motion characteristics between the bones and joints. Among them, the original motion information of the bones and joints covers the acceleration and angular velocity data of the joints, while the original information of the relative motion depends on the magnetic vector information captured by the magnetometer, and this information is derived from the magnetic sources ingeniously arranged on the body. The present invention deeply fuses the above two types of features, aiming to construct a more comprehensive and multi-dimensional multi-modal fusion feature, which can accurately determine the recognition result corresponding to the target action. By cleverly integrating the direct motion data and relative motion data of the bones and joints, it realizes the efficient multi-modal recognition of human activities, effectively overcomes the problem of poor recognition effect in traditional single-modal recognition methods, and thus significantly improves the accuracy and reliability of human activity recognition.
[0194] Based on the same idea, the embodiment of this specification also provides a human activity recognition device. As Figure 7 shown, this device includes:
[0195] A memory, a processor, and a communication interface coupled to the processor; a computer program that can be run by the processor is stored on the memory; when the processor runs the computer program, it executes the aforementioned human activity recognition method.
[0196] As Figure 7 shown, the above-mentioned processor can be a general central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. The above-mentioned communication interface can be one or more. The communication interface can use any transceiver-like system for communicating with other devices or communication networks.
[0197] As Figure 7 shown, the above-mentioned terminal device 4 may further include a communication line. The communication line may include a path for transmitting information between the above-mentioned components.
[0198] Optionally, as Figure 7 shown, the above-mentioned terminal device 4 may further include a memory. A computer program that can be run by the processor is stored on the memory; when the processor runs the computer program, the method provided by the embodiment of the present invention is thus realized.
[0199] As Figure 7As shown in the figure, the memory can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited to this. The memory can exist independently and be connected to the processor through a communication line. The memory can also be integrated with the processor.
[0200] Optionally, the computer-executable instructions in the embodiments of the present invention can also be referred to as application program code, and the embodiments of the present invention do not make specific limitations on this.
[0201] In a specific implementation, as an embodiment, as Figure 7 shown, the processor can include one or more CPUs, such as Figure 7 CPU0 and CPU1 in
[0202] In a specific implementation, as an embodiment, as Figure 7 shown, the terminal device 4 can include multiple processors, such as Figure 7 the processors in
[0203] Optionally, the computer-executable instructions in the embodiments of the present invention can also be referred to as application program code, and the embodiments of the present invention do not make specific limitations on this.
[0204] The method disclosed in the embodiments of the present invention above can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0205] Although the present invention has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0206] Although the present invention has been described in connection with specific features and their embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of the present invention. Accordingly, the present specification and the drawings are merely exemplary illustrations of the present invention as defined by the appended claims, and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A method for human activity recognition, characterized in that, The method includes: Obtaining original data; The original data at least includes human joint motion data, human joint relative motion data, and magnetic source excitation data; Performing preprocessing on the original data to obtain preprocessed data, and performing feature extraction on the preprocessed data to obtain target features; Performing data-level fusion on the original data, performing feature-level fusion on the target features, and performing decision-level fusion on the feature fusion model to obtain fusion information; Performing human activity recognition based on the fusion information.
2. The human activity recognition method according to claim 1, characterized in that, The preprocessing at least includes data conversion, noise filtering, registration, compression, interpolation, and normalization; the preprocessed data includes first preprocessed data and second preprocessed data; The target features include a first target feature set and a second target feature set; Performing feature extraction on the preprocessed data to obtain target features, specifically including: Performing feature extraction on the first preprocessed data to obtain a first target feature set; the first target feature set is the skeletal joint motion feature of the target action; Performing feature extraction on the second preprocessed data feature to obtain a second target feature set; the second target feature set is the skeletal joint relative motion feature of the target action.
3. The method for human activity recognition according to claim 1, characterized in that Performing feature extraction on the preprocessed data to obtain target features, specifically including: Performing IMU attitude solution, IMU data feature extraction, and magnetic field data feature extraction on the preprocessed data; Performing data dimensionality reduction on the extracted feature data.
4. A human activity recognition method according to claim 3, characterized in that, The performing IMU attitude solution, IMU data feature extraction, and magnetic field data feature extraction on the preprocessed data specifically includes: Performing IMU attitude solution on the preprocessed data and expressing it in the form of quaternions and Euler angles; Performing IMU data feature extraction on the human joint motion data in the preprocessed data to obtain the triaxial angular velocity modulus, the difference between the triaxial acceleration modulus and the gravitational acceleration modulus, and the magnetic field intensity; Performing IMU data feature extraction on the human joint relative motion data in the preprocessed data and performing feature extraction on the magnetic field data; the magnetic field data is used to represent the relative position relationship with the magnetic source, and the second-order norm of the triaxial magnetometer data is used to represent the magnitude of the magnetic field intensity; The performing data dimensionality reduction on the extracted feature data specifically includes: Performing PCA dimensionality reduction on the linear data in the extracted feature data, and projecting the transformed data into a new coordinate system through linear transformation; Performing T-SNE dimensionality reduction on the non-linear data in the extracted feature data, and projecting the similar points in the high-dimensional space into the low-dimensional space to maintain the local structure.
5. A human activity recognition method according to claim 2, characterized in that, The performing data-level fusion on the original data specifically includes: Performing splicing and integration on the first preprocessed data and the second preprocessed data in a first sequence splicing method, a first structure splicing method, and a first pixel splicing method; Feeding the integrated data into a training model according to the sliding window length for data-level fusion; The first sequence splicing method includes: Input the serial data of the first preprocessed data and the second preprocessed data into the model. The splicing methods for multiple groups of different joint data include intra-group splicing and inter-group splicing, and finally input in the form of a 1*9 sequence, where the sequence length is the sum of the sequence lengths of the first preprocessed data and the second preprocessed data; the different joint data includes 6 groups. The splicing method of the first structure includes: Splice the first preprocessed data and the second preprocessed data of the same group into a 3*3 matrix, and splice the matrices of different groups to form a 3*3*6 structure. The splicing method of the first pixel includes: Splice the first preprocessed data and the second preprocessed data of all groups into a 6*9 pixel matrix, and expand it to a 9*9 pixel matrix by zero-padding according to the convolution requirement.
6. The human activity recognition method according to claim 2, characterized in that, The feature-level fusion of the target feature specifically includes: Splice and integrate the first target feature set and the second target feature set in the second sequence splicing method, the second structure splicing method, and the second pixel splicing method for feature-level fusion. The second sequence splicing method includes: Perform second sequence splicing on the first target feature set and the second target feature set. The splicing method of the second sequence splicing includes intra-group splicing and inter-group splicing, and input in the form of a 1*M sequence, where M is the sum of the data lengths of the first preprocessed data and the second preprocessed data. The second structure splicing method includes: Splice the joint feature data of the same group into a 3*3 matrix, splice the matrices of different groups to form an N*P*6 structure, reshape M into N*P, and fill the vacancies with zeros. The second pixel splicing method includes: Splice the joint feature data of all groups into an M*M pixel matrix, and fill the vacancies with zeros.
7. A human activity recognition method according to claim 1, characterized in that The decision-level fusion of the feature fusion model specifically includes: Train the first human activity recognition training model based on joint movement according to the first preprocessed data, the first target feature set, and the action label after One-hot encoding. Determine the prediction result of the first human activity recognition training model as the first decision. Train the second human activity recognition training model based on joint movement according to the second preprocessed data, the second target feature set, and the action label after One-hot encoding. Determine the prediction result of the second human activity recognition training model as the second decision. Obtain the final decision through training fusion of the first decision and the second decision by a preset method; the preset method at least includes: D-S evidence theory, weighted average, or voting. Allocate basic probabilities to each hypothesis according to the first decision and the second decision, reflecting the degree of support of the data for the corresponding hypothesis. Synthesize the evidence of all decisions, and use Dempster's rule to calculate the final basic probability assignment. Based on the fused evidence, calculate the confidence or likelihood of each hypothesis. Obtain the target result according to the maximum confidence or maximum likelihood.
8. A human activity recognition system, characterized in that, The system includes: Inertial measurement device, magnetic field measurement device, electromagnetic drive device, and terminal device. The inertial measurement device is integrated with the magnetic field measurement device and establishes a communication connection with the terminal device; the inertial measurement device and the magnetic field measurement device are placed on the target part of the human body; the electromagnetic drive device is a passive device; the electromagnetic drive device is arranged on the human back and is spaced from the inertial measurement device and the magnetic field measurement device by a preset distance; The terminal device is used to receive the original data; the original data at least includes human joint movement data, human joint relative movement data, and magnetic source excitation data; Preprocess the original data to obtain preprocessed data, and extract features from the preprocessed data to obtain target features; Perform data-level fusion on the original data, perform feature-level fusion on the target features, and perform decision-level fusion on the feature fusion model to obtain fusion information; Perform human activity recognition based on the fusion information.
9. The human activity recognition system according to claim 8, wherein The communication between the inertial measurement device, the magnetic field measurement device and the terminal device is carried out by wireless Bluetooth technology for signal transmission; The inertial measurement device, the magnetic field measurement device and the terminal device work synchronously, collect and upload human activity data in real time, and calculate through the cloud platform to identify human activities.
10. A human activity recognition device, characterized in that the device Including: A memory, a processor, and a communication interface coupled to the processor; A computer program that can be run by the processor is stored on the memory; When the processor runs the computer program, it executes a human activity recognition method according to any one of claims 1 to 7.
Citation Information
Cited By
Method and device for determining state of pet, electronic equipment and storage medium
CN121502385A