A method for classifying construction worker actions in a hierarchical integrated network structure
By integrating a hierarchical network structure and multiple IMU configuration schemes, and combining artificial neural networks and tree breadth-first search algorithms, the problems of accuracy and real-time performance in classifying construction worker actions are solved, achieving high-accuracy and fast-response action recognition, which is suitable for wearable exoskeleton devices.
Patent Information
- Application Number
- CN202411349693.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-09-26
AI Technical Summary
Existing methods for classifying the movements of construction workers cannot simultaneously satisfy both accuracy and real-time performance, and traditional methods are not effective for real-time control on resource-constrained wearable exoskeleton devices.
A hierarchical integrated network structure was designed. By constructing a directed label tree and various IMU configuration schemes, and combining artificial neural networks and tree breadth-first search algorithms, the location and number of IMU sensors, the sliding time window size, and feature extraction were optimized to achieve accuracy and real-time performance in action classification.
It improves the accuracy and real-time performance of motion classification, is applicable to different wearable devices, and ensures that the exoskeleton device can accurately and in real-time assist construction workers in completing a variety of tasks.
Smart Images

Figure CN119312195B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of construction worker action classification, and particularly relates to a construction worker action classification method based on a hierarchical integrated network structure. BACKGROUND
[0002] Wearable devices have great potential for construction workers to prevent work-related musculoskeletal diseases. Advanced research shows that exoskeletons have the potential to reduce muscle activity and range of motion in various body movements, such as tasks involving walking (flat and inclined), carrying, lifting (including lifting, lowering, and repeated lifting), pushing and pulling, and kneeling on the ground and roof. The widespread use of exoskeletons in a variety of tasks significantly improves the working environment and efficiency of construction workers. As the application tasks expand, action classification technology is essential for exoskeletons to provide adaptive assistance. At present, most wearable robots use rule-based methods to classify movement patterns, such as finite state machines, fuzzy logic, or decision trees. However, the above methods require detailed descriptions of each phase to design rules. As the number of actions increases, rule design will become more complex, time-consuming, and lack of certain adaptability.
[0003] Machine learning-based methods can perform action classification tasks without relying too much on tedious rule design. For example, traditional machine learning based on support vector machines, linear discriminant analysis, Bayesian networks, Markov models, and artificial neural networks (ANN) is widely used for action classification. However, due to the complexity of input signal data (containing spatial and temporal information), the above methods are difficult to accurately classify body movements. In addition, the above methods rely on feature engineering to improve accuracy, and the feature extraction strategy in feature engineering requires human intervention, resulting in increased time consumption.
[0004] In contrast, deep learning models have the advantage of automatically extracting complex abstract features from raw data, thereby improving decision-making and enabling classification of data with complex spatial and temporal information. For example, convolutional neural networks (CNN), long short-term memory networks (LSTM), and further improved action classification accuracy through algorithm adjustment and / or model structure. However, the above methods usually classify at a high specificity level and represent the transition between high specificity actions as a single transition occurring at a single time point, without considering the hierarchical relationship between multiple actions, i.e., without concatenating multiple outputs to jointly determine the final action category. In addition, while the above methods improve accuracy, they also increase computational complexity and spatial complexity, which is not conducive to real-time control of resource-constrained exoskeletons.
[0005] In addition, advanced research usually provides a single solution based on certain sensor combinations. For example, numerous studies discuss the influence of feature extraction, IMU configuration, and sliding time window size on classification accuracy. However, the above studies do not provide multiple alternative configuration schemes, and a single solution can have an indirect impact on different assistive devices (algorithms using foot IMUs are not suitable for bilateral portable hip exoskeletons or unilateral knee exoskeletons).
[0006] The invention patent with application number 201811328308.3 discloses a human action recognition method based on a TP-STG framework. The method includes: taking video information as input, adding prior knowledge to the SVM classifier, and proposing a posteriori discriminant criterion to remove non-person targets; segmenting the person target by target positioning and detection algorithm, and outputting in the form of target frame and coordinate information to provide input data for human key point detection; using an improved posture recognition algorithm for body part positioning and correlation analysis to extract all human key point information and form a key point sequence; constructing a space-time graph on the key point sequence through an action recognition algorithm, applying it to multi-layer space-time graph convolution operation, and classifying the action by a Soft-max classifier to realize human action recognition in a complex scene. The method of the above invention first combines the actual scene of the offshore platform, and the proposed TP-STG framework first attempts to use target detection, posture recognition, and space-time graph convolution to recognize the actions of workers on the offshore drilling platform. However, the above patent needs to convert video and picture information into point sequence information, which is complex and time-consuming in feature engineering, and is not conducive to real-time control of resource-constrained exoskeletons; the multi-layer space-time convolution operation does not consider the hierarchical relationship between multiple actions or calculate the output probability based on multiple Soft-max classifiers, thereby affecting the accuracy of action classification. SUMMARY
[0007] To solve the technical problem that existing construction worker action classification methods are difficult to meet both accuracy and real-time performance, the present application proposes a construction worker action classification method with a hierarchical integrated network structure, designs six configurations with high accuracy, strong real-time performance, and IMU sensors worn on different body parts, and promotes multiple wearable devices to accurately and timely assist construction workers in completing various work tasks.
[0008] To achieve the above purpose, the technical scheme of the present application is as follows: a construction worker action classification method with a hierarchical integrated network structure, comprising the following steps:
[0009] S1: Obtain action data of a construction worker wearing an IMU performing multiple actions, preprocess the action data, and establish an action data set;
[0010] S2: select different IMU quantity and installation position, sliding time window size and feature extraction data combination to obtain multiple IMU configuration schemes; a directed label tree is established based on the specificity between multiple actions, a hierarchical integrated network is constructed based on the directed label tree, the hierarchical integrated network is trained by using the corresponding action data in the action data set, and a trained action classification model is obtained;
[0011] S3: input the action data of the complete action cycle of each IMU configuration scheme into the trained action classification model, evaluate the performance of each IMU configuration scheme according to four accuracy performance indicators and two real-time performance indicators, and select the IMU configuration scheme meeting the preset condition as the selectable scheme of the action of the construction worker.
[0012] Preferably, the method for obtaining action data in step S1 is: one IMU is installed on the center of the subject construction worker's two front thighs, two front shanks and torso respectively; the subject construction worker performs multiple actions of walking, uphill, downhill, ascending stairs, descending stairs, running, standing, squatting, kneeling, squatting up and kneeling up respectively, and the action data collected by each IMU is wirelessly transmitted to the PC end for data preprocessing; 3D acceleration and 3D angular velocity of the IMU sensor are selected as the signal channel for feature extraction.
[0013] The data preprocessing includes: suppressing action data noise by Kalman filtering, cleaning action data by data visualization method, and obtaining time series X of a complete action cycle.
[0014] Preferably, the Z axis of all IMUs points to the front of the subject construction worker, and the Y axis is perpendicular to the sagittal plane of the subject construction worker.
[0015] The action of the subject construction worker in a complete action cycle is: 1) standing for 30s, 2) static half-kneeling for 30s, 3) static deep squatting for 30s, 4) walking on flat ground at a self-selected speed for 30s, 5) running on flat ground at a self-selected speed for 30s, 6) ascending stairs at a self-selected speed for 30s, 7) descending stairs at a self-selected speed for 30s, 8) ascending a 20m slope at a self-selected speed, 9) descending a 20m slope at a self-selected speed, 10) repeating squatting 5 times at a 4s rhythm, and 11) repeating kneeling 5 times at a 4s rhythm.
[0016] Preferably, the method for constructing a directed label tree is: dividing the plurality of actions into static actions and dynamic actions according to different action attributes, the left child node and the right child node of the root node are static actions and dynamic actions respectively, the left child node and the right child node of the static action are bending the knee and standing still respectively, the left child node and the right child node of the bending the knee are static squatting and static kneeling respectively; the left child node and the right child node of the dynamic action are moving around and moving in place respectively, the left child node and the right child node of the moving in place are squatting up and kneeling up respectively, the left child node and the right child node of the moving around are flat ground and uneven ground respectively, the left child node and the right child node of the flat ground are walking and running respectively, the left child node and the right child node of the uneven ground are stairs and slope respectively, the left child node and the right child node of the stairs are climbing stairs and descending stairs respectively, and the left child node and the right child node of the slope are ascending and descending respectively.
[0017] Preferably, the directed label tree of the 11 actions includes 6 layers of the root node.
[0018] The encoding method of the binary label of the plurality of actions in the directed label tree is: in the directed label tree, the left child node and the right child node are marked as "1" and "0" respectively; the child nodes of all the actions to be classified are represented as an N-1 dimensional vector; if the child node is not in the Nth layer, the difference M between N and the layer level of the current child node is calculated, and then the last (N-M) bit components are filled with "0".
[0019] Preferably, the method for constructing a hierarchical integrated network based on the directed label tree is: the directed label tree has 10 parent nodes and 11 child nodes, the 11 child nodes represent 11 actions respectively, and the 10 parent nodes include a root node, static actions, dynamic actions, bending the knee, moving around, moving in place, flat ground, uneven ground, stairs and slope; an ANN network is constructed for each parent node to classify the actions, and the ANN classifier of each parent node is trained separately.
[0020] The input of each ANN network is the input variable obtained by preprocessing and signal processing of the action data of all child nodes under the parent node as the root node, and the output of each ANN network is the action represented by the left child node and the right child node.
[0021] Preferably, each ANN network is composed of an input layer, an encoder layer, a decoder layer and an output layer, wherein the size of the input layer is determined by the input vector V, the encoder layer, the decoder layer and the output layer are 8x1, 4x1 and 2x1 respectively, and a Soft-max layer is arranged after the output layer, and the ANN classifier outputs a binary label vector to represent the action category.
[0022] In the action classification, the binary label of standing is [0 1 0 0 0], the binary label of squatting is [0 0 0 0 0], the binary label of kneeling is [0 0 1 0 0], the binary label of standing up is [1 1 0 0 0], the binary label of kneeling up is [1 1 1 0 0], the binary label of walking is [1 0 0 0 0], the binary label of running is [1 0 0 1 0], the binary label of going upstairs is [1 0 1 0 0], the binary label of going downstairs is [1 0 1 0 1], the binary label of going uphill is [1 0 1 1 1], and the binary label of going downhill is [1 0 1 1 0];
[0023] The action classification starts from the sub-network of the root node, and the output action node of the sub-network of the root node will be used as the sub-network of the next level node, and so on, until the output node has no child node;
[0024] The loss function of the training hierarchical integrated network adopts a classification cross-entropy, and the loss function is:
[0025]
[0026] Wherein, m represents the output action number of a single sub-network in the hierarchical integrated network, c m represents the binary index of the real label of the mth action, p m represents the classification probability of the mth action of the sub-network of a single node in the hierarchical integrated network.
[0027] Preferably, the different IMU numbers and installation positions include: 3 cases of installing 1 IMU, respectively installed on the torso, the thigh and the calf; 3 cases of installing 2 IMUs, respectively installed on two thighs, two calves and a thigh & a calf; 2 cases of installing 3 IMUs, respectively installed on the torso & two thighs and the torso & two calves; 1 case of installing 4 IMUs, installed on two thighs and two calves; and 1 case of installing 5 IMUs, respectively installed on the torso, two thighs and two calves.
[0028] The sliding time window size includes 100 ms, 150 ms and 200 ms;
[0029] The feature extraction data includes 3D acceleration collected by the accelerometer, 3D angular velocity collected by the gyroscope, and 3D acceleration collected by the accelerometer & 3D angular velocity collected by the gyroscope;
[0030] The different IMU numbers and installation positions, the sliding time window size and the feature selection data are combined with each other, and there are 90 IMU configuration solutions in total.
[0031] The action data collected by the plurality of IMU configuration schemes and the action data in the action data set are subjected to signal processing before being input into the hierarchical integrated network, and the method of signal processing is as follows: a sliding time window is used to extract features including maximum value, minimum value, average value, standard deviation, first signal value, middle signal value and last signal value from each signal channel of the feature extraction data to form an input vector V; and the input vector V is input into the hierarchical integrated network.
[0032] Preferably, a group Soft-max method is used to traverse the directed label tree nodes in breadth-first order, obtain path encoding labels and calculate the Soft-max probability output of each node on the path, and then multiply the Soft-max probability output of each node by the probability output of its parent node to calculate the path probability output.
[0033] By integrating the probability outputs of the sub-networks of multiple parent nodes, a unique path representing a sub-node leading to the current action category is indexed out, so as to determine the action category.
[0034] Preferably, each action node on the path corresponds to a classification probability p m , wherein the classification probability of the root node and other nodes filled with "0" is set to 1.
[0035] The label of each action is represented by a path y k = (v1,..., v E ), wherein v1 and v E are the root node and the terminal node of the path respectively, and E represents the number of nodes on the path; the path probability P(y k ) of the kth action node is calculated by using a group Soft-max method as , wherein v j-1 represents the j-1th node on the path.
[0036] The four accuracy performance indicators include the accuracy, precision, recall and F1 score of the action classification model; and the two real-time performance indicators include the parameter quantity and inference time of the action classification model.
[0037] The preset condition is that the scheme with an accuracy higher than 93% and a proper number of IMUs is selected as the optional scheme for the action classification of the construction worker.
[0038] Compared with the prior art, the application has the beneficial effects: constructing a specific directed label tree structure to classify multiple actions hierarchically (from top to bottom), i.e., first classifying low-specificity patterns, and gradually classifying higher-specificity patterns as the level of the tree increases, thereby reducing inference time; correlating the correlation between multiple actions, i.e., assigning a classifier to each parent node and integrating the outputs of multiple classifiers to jointly index the unique path leading to the current action category leaf node, thereby improving accuracy; developing 6 configurations of wearing IMU sensors on different body parts, including the position and number of IMUs, sliding time window size and feature extraction, which can be applied to different wearable devices. The application combines artificial intelligence training, integrates artificial neural networks and tree breadth-first search algorithms to solve the problems of insufficient classification accuracy and poor real-time performance of existing wearable auxiliary device action classification strategies, and the possible indirect adverse effects of single sensor combination solutions on different auxiliary devices. At the same time, the application has the characteristics of a large number of classified actions, high accuracy, strong real-time performance and the provision of multiple IMU configuration schemes, which can ensure that the wearable exoskeleton device accurately and timely assists construction workers to complete various work tasks. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0040] Figure 1 Flowchart of the present application.
[0041] Figure 2 IMU installation method diagram for action data collection of the present application, wherein (a) is IMU configuration and data collection, and (b) is the subject performing multiple actions.
[0042] Figure 3 Basic action diagram of action data collection of the present application.
[0043] Figure 4 Structure diagram of the hierarchical integrated network of the present application.
[0044] Figure 5 Experimental results diagram of all 90 IMU configuration schemes (including 6 optional IMU configuration schemes) of the present application in terms of accuracy, recall rate, accuracy and F1 score.
[0045] Figure 6 Accuracy performance experimental results diagram of the 6 optional solutions of the present application.
[0046] Figure 7 Fig. 6 is a diagram of real-time performance experimental results of the six optional solutions of the present application. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0048] As shown in Fig. 1, a construction worker action classification method of a hierarchical integrated network structure comprises the following steps: Figure 1 Step S1: acquiring action data of a construction worker wearing an IMU performing a plurality of actions, transmitting the action data to a PC end through wireless transmission for data preprocessing, and establishing an action data set.
[0049] The installation scheme method of the action data acquisition IMU is shown in Fig. 2(a). The subject is installed with a total of 5 IMUs (LPMS-B2, Alubi) on the torso and legs. The action data collected by the IMU is transmitted to the PC end through wireless transmission for data preprocessing, which can bring higher data accuracy and comprehensive analysis capability. The data preprocessing process specifically comprises: processing through Kalman filtering to effectively suppress action data noise, thereby obtaining smoother and more reliable state estimation; and cleaning data through a data visualization method, thereby obtaining a time sequence X of a complete action cycle.
[0050] Figure 2 The 5 IMUs are respectively installed on the front segments of the subject's two thighs (2 cm above the knee joint), the front segments of the subject's two calves (2 cm above the ankle joint), and the center of the subject's torso. The 4 IMUs installed on the legs are used to record the motion data of the respective parts, and the 1 IMU installed on the torso is used to monitor the overall body movement of the subject. The IMU installation method will help to obtain data synchronization and accurate motion analysis. The Z axis of all IMUs points to the front of the subject, and the Y axis is perpendicular to the sagittal plane of the subject, which ensures that all IMUs are oriented consistently, which helps to more easily align and integrate data when data fusion or analysis is performed between multiple IMUs. The IMU records data at a sampling rate of 100 Hz and transmits the data to the PC end through wireless transmission for data preprocessing. As shown in Fig. 2(b), the plurality of actions include walking, uphill, downhill, ascending stairs, descending stairs, running, standing still, squatting still, kneeling still, standing up from squatting, and standing up from kneeling, a total of 11 actions.
[0051] The 5 IMUs are respectively installed on the front segments of the subject's two thighs (2 cm above the knee joint), the front segments of the subject's two calves (2 cm above the ankle joint), and the center of the subject's torso. The 4 IMUs installed on the legs are used to record the motion data of the respective parts, and the 1 IMU installed on the torso is used to monitor the overall body movement of the subject. The IMU installation method will help to obtain data synchronization and accurate motion analysis. The Z axis of all IMUs points to the front of the subject, and the Y axis is perpendicular to the sagittal plane of the subject, which ensures that all IMUs are oriented consistently, which helps to more easily align and integrate data when data fusion or analysis is performed between multiple IMUs. The IMU records data at a sampling rate of 100 Hz and transmits the data to the PC end through wireless transmission for data preprocessing. As shown in Fig. 2(b), the plurality of actions include walking, uphill, downhill, ascending stairs, descending stairs, running, standing still, squatting still, kneeling still, standing up from squatting, and standing up from kneeling, a total of 11 actions. Figure 2
[0052] After the IMU sensors were installed, the subjects completed a variety of motions of construction workers according to the specified operation and procedure for compiling the data set of the training hierarchical ensemble network, and the motion data acquisition protocol is shown in FIG. 1. Figure 3 The motions of the construction workers were as follows: 1) standing still for 30 s, 2) static half-kneeling for 30 s, 3) static deep squatting for 30 s, 4) walking on the flat ground at a self-selected speed for 30 s, 5) running on the flat ground at a self-selected speed for 30 s, 6) climbing stairs at a self-selected speed for 30 s, 7) descending stairs at a self-selected speed for 30 s, 8) ascending a slope (30° slope) at a self-selected speed for 20 m, 9) descending a slope (30° slope) at a self-selected speed for 20 m, 10) repeating squatting for 5 times at a 4 s rhythm, and 11) repeating kneeling for 5 times at a 4 s rhythm. The rest time between static motions was 1 min, and the rest time between dynamic motions was 5 min. The rest time between static and dynamic motions was set to help reduce the influence of fatigue of the subjects on data acquisition and ensure the stability of data collection for each motion.
[0053] The subjects included 10 people, 5 males and 5 females, with an average age of 21.9 ± 2.3 years, an average height of 172.0 ± 7.0 cm, and an average weight of 69.2 ± 8.3 kg. All the subjects gave informed consent and completed various motions of construction workers under the guidance of safety personnel according to the motion data acquisition protocol in FIG. 1. The motion data acquisition protocol has been reviewed and approved by the Institutional Review Board (IRB: xxxb202404-1X) of Beijing University of Technology. Figure 3
[0054] Step S2: Select different IMU quantities and installation positions, sliding time window sizes, and feature extraction data combinations to obtain various IMU configuration schemes; establish a directed label tree based on the specificity between various motions, construct a hierarchical ensemble network based on the directed label tree, and train the parent nodes of the hierarchical ensemble network using the corresponding motion data in the motion data set of step one to obtain a trained motion classification model.
[0055] The hierarchical ensemble network structure is shown in FIG. 2. The entire structure diagram is composed of three parts, namely signal processing, hierarchical ensemble network, and motion category. Figure 4
[0056] Signal processing: process the motion data set obtained in step one to obtain the input vector and input label of the hierarchical ensemble network.
[0057] The signal processing is to process the data signal, that is, to extract features from the signal channels of the motion data set of step one by using the sliding time window method to reduce the influence of data noise and improve the robustness and accuracy of feature extraction. Since the Euler angle data will drift over time during walking, the 3D acceleration and 3D angular velocity of the IMU sensor are selected as the signal channels for feature extraction.
[0058] Each of the signal channels extracts 7 features, respectively: maximum value, minimum value, average value, standard deviation, first, middle and last, which together constitute the input vector V. The 7 features can more comprehensively and stably describe the basic statistical characteristics of the signal, which helps to improve the effect of feature extraction and improve the performance of subsequent data analysis or machine learning model. Figure 4 In the embodiment, the size of the input vector V needs to be determined according to the specific IMU configuration scheme. In the present application, z kinds of IMU configuration schemes can be combined by selecting different IMU numbers, sliding time window sizes and feature extraction data. For example, in the zth IMU configuration scheme Sz, there are 2 IMUs, the sliding time window size is fixed at 100 ms, and the feature extraction data selects 3D acceleration and 3D angular velocity (6 signal channels), that is, the size of the input vector V is 84 (2x1x6x7).
[0059] The IMU configuration scheme covers all 10 cases of capturing lower limb motion data, with IMU numbers ranging from 1 to 5:
[0060] There are 3 cases of installing 1 IMU, which are installed on the torso, the thigh and the calf respectively;
[0061] There are 3 cases of installing 2 IMUs, which are installed on two thighs, two calves and a thigh & a calf (single side) respectively;
[0062] There are 2 cases of installing 3 IMUs, which are installed on the torso & two thighs and the torso & two calves respectively;
[0063] There is 1 case of installing 4 IMUs, which is installed on two thighs and two calves;
[0064] There is 1 case of installing 5 IMUs, which is installed on the torso, two thighs and two calves respectively.
[0065] The sliding time window size of the present application selects 3 cases, which are 100 ms, 150 ms and 200 ms respectively.
[0066] The feature extraction aspect of the present application selects 3 cases, which are only accelerometer (3D acceleration), only gyroscope (3D angular velocity) and accelerometer & gyroscope (3D acceleration and 3D angular velocity) respectively.
[0067] The above-mentioned IMU configuration, sliding time window size and feature selection cases are combined with each other, and there are 90 kinds of IMU configuration solutions, that is, z=90.
[0068] A directed label tree is constructed for hierarchical modeling of multiple action labels. The rules for hierarchical modeling are based on the specificity between different actions, where specificity refers to different action attributes. For example, squatting and standing are static attribute actions, while walking on flat ground and walking on stairs are dynamic attribute actions. More specifically, walking on stairs can be divided into going up stairs and going down stairs based on the direction attribute. (See attached image) Figure 4 As shown, this invention constructs a directed label tree for 11 types of lower limb movements of construction workers. The number of levels in the directed label tree is represented by N, and N is set to 6 (including the root node). The left and right child nodes of the root node are static and dynamic movements, respectively. The left and right child nodes of static movements are bent knee and standing still, respectively; the left and right child nodes of bent knee are squatting and kneeling, respectively. The left and right child nodes of dynamic movements are moving around and moving in place, respectively; the left and right child nodes of moving in place are squatting and kneeling, respectively; the left and right child nodes of moving around are flat ground and uneven ground, respectively; the left and right child nodes of flat ground are walking and running, respectively; the left and right child nodes of uneven ground are stairs and slopes, respectively; the left and right child nodes of stairs are going up stairs and going down stairs, respectively; and the left and right child nodes of slopes are going uphill and going downhill, respectively.
[0069] In the directed label tree, the left and right child nodes are labeled as "1" and "0" respectively. All leaf nodes of actions to be classified are represented as N-1 dimensional vectors. If a leaf node is not at level N, the difference M between N and the level of the current leaf node is calculated, and then the last (NM) bits are filled with "0". For example, the tree path of the level 4 leaf node "kneeling up" is represented in binary multi-label format as [1 1 1 0 0], where the 4th and 5th bits are filled with "0"; the tree path of the level 6 leaf node "going down the stairs" is represented in binary multi-label format as [1 0 1 0 1].
[0070] The binary multi-label format represents paths that must end with a leaf node, and all paths must have the same length and be represented by a 5-dimensional vector. This invention applies to time series X = {x1,...,x...} t Classify the data, where i = 1, 2, ..., t, and the input labels are Y = {y1, ..., y2}. n}, where k = 1, 2, ..., n. x t y represents a sample in a time series. n The path label represents the action to be identified, t represents the time step, and n represents the number of actions. This invention uses a fixed-length sliding time window to classify the time series X. Since classification must be performed in real-time, the sample x... i Label prediction must depend on samples recorded before the current time step. The window slides over time, updating its contents and reclassifying the samples.
[0071] Hierarchical integrated network: based on the directed label tree, a hierarchical integrated network is constructed.
[0072] The directed label tree has 10 parent nodes and 11 leaf nodes, and the 11 leaf nodes respectively represent 11 actions, and the 10 parent nodes include a root node, a static action, a dynamic action, a bent knee, a moving everywhere, a moving in place, a flat ground, an uneven ground, a staircase and an inclined slope. In order to fully utilize the hierarchical structure and avoid inconsistent action label prediction, the present application adopts a local classifier method to classify the actions, that is, an ANN classifier is constructed for each parent node and is separately trained. The networks of the 10 parent nodes are respectively: the root node is Net0, the static action is Net2, the dynamic action is Net1, the bent knee is Net5, the moving everywhere is Net4, the moving in place is Net3, the flat ground is Net7, the uneven ground is Net6, the staircase is Net9, and the inclined slope is Net8. The input of each ANN network is the action data of all leaf nodes under the parent node network as the root node, and the output is the action represented by the left and right leaf nodes. The action data is based on the action data set established in step one. For example, the training data of Net 6 is four groups of action data: "going upstairs", "going downstairs", "going uphill" and "going downhill". The training data of Net 7 is two groups of action data: "walking" and "running".
[0073] Each of the ANN classifiers consists of an input layer, an encoder layer, a decoder layer and an output layer. Among them, the size of the input layer is determined by the input vector V, and the encoder layer, the decoder layer and the output layer are all fixed sizes: 8x1, 4x1 and 2x1 respectively. The fixed size structure makes the parameter quantity and the calculation complexity of the model controllable, which helps to avoid overfitting and improve the efficiency of training and inference. In addition, it also simplifies the design and implementation process of the model. After the output layer, there is a Soft-max layer, and the ANN classifier outputs a vector to represent the action category. Each action is assigned a unique identifier: [1, 0] and [0, 1] represent the actions on the left and right child nodes respectively. Therefore, in the action classification, the binary label of static standing is [0 1 0 0 0], the binary label of static squatting is [0 0 0 0 0], the binary label of static kneeling is [0 0 1 0 0], the binary label of squatting up is [1 1 0 0 0], the binary label of kneeling up is [1 1 1 0 0], the binary label of walking is [1 0 0 0 0], the binary label of running is [1 0 0 1 0], the binary label of going upstairs is [1 0 1 0 0], the binary label of going downstairs is [1 0 1 0 1], the binary label of going uphill is [1 0 1 1 1], and the binary label of going downhill is [1 0 1 1 0].
[0074] The training process of action classification aims to minimize the loss function, and the loss function of the present application adopts classification cross-entropy, as shown in formula (1)
[0075]
[0076] Wherein, m represents the output action number of a single subnetwork in the hierarchical integrated network, c m represents the binary index of the true label of the mth action, p m represents the prediction probability of the mth action of the subnetwork of a single node in the hierarchical integrated network.
[0077] The hierarchical classification starts from the subnetwork of the root node, and the output action node of the root node subnetwork will be used as the subnetwork of the next level node, and so on, until there is no subnode of the output node. Consider that each action label is represented by a path y k =(v1,...,v E ), wherein v1 and v E are the starting node (root node) and the terminal node of the path respectively, and E represents the number of nodes on the path, and in the present application, E=5. The present application adopts the grouped Soft-max method to calculate the path probability P(y k ) of the action label. The path probability P(y k ) represents the path calculation probability of the kth action node. Each action node on the path corresponds to a classification probability p m . Among them, the classification probability p of the root node and other nodes filled with "0" is set to 1 to maintain the consistency of path calculation, so that the output probability distribution of the final classification node is between 0-1, so the path probability P(v1)=1. The root node is the starting node of the search path, which is not a node on the path obtained by classification calculation probability; the zero padding node does not exist in reality, so it is also not a node on the path obtained by classification calculation probability. The path probability calculation is shown in formula (2).
[0078]
[0079] Wherein, j=1,2…E, v j-1 represents the j-1th node on the path.
[0080] The Softmax layer of each subnetwork mentioned above is combined for calculation, and the grouped Soft-max is to traverse the directed label tree node in breadth-first order, obtain the path encoding label and calculate the Soft-max probability output of each node on the path, and then multiply the grouped Soft-max output of each node by the output of its parent node to calculate the probability output of the path.
[0081] Action category: according to the hierarchical probability calculation result, output the final classification target.
[0082] The application determines the action category by integrating the probability outputs of multiple sub-networks and indexing a unique path leading to the current action category leaf node. Figure 4 The red arrow route in FIG. Figure 4 The red box in FIG.
[0083] Step S3: After feature extraction of the time sequence of the complete action cycle obtained by each IMU configuration scheme, input the trained action classification model, and evaluate the performance of each IMU configuration scheme according to four accuracy performance indicators (accuracy, precision, recall, and F1-score) and two real-time performance indicators (parameters and inference time), and select six IMU configuration schemes meeting the preset conditions as optional schemes for action classification of construction workers.
[0084] The calculation of the four accuracy performance indicators is shown in formulas (3)-(6):
[0085]
[0086] The four accuracy performance indicators accuracy, precision, recall, and F1-score represent the accuracy, precision, recall, and F1-score of the action classification model prediction, respectively. Among them, TP, TN, FP, and FN are the four basic elements of the "confusion matrix", which represent true positive, true negative, false positive, and false negative, respectively. True positive TP is the number of samples correctly predicted as positive by the action classification model in the positive class samples. True negative TN is the number of samples correctly predicted as negative by the action classification model in the negative class samples. False positive FP is the number of samples incorrectly predicted as positive by the action classification model in the negative class samples. False negative FN is the number of samples incorrectly predicted as negative by the action classification model in the positive class samples.
[0087] In addition, the application also uses the number of parameters (parameters) and the inference time (inference time) to evaluate the real-time performance of the action classification model.
[0088] The number of parameters reflects the size and complexity of the model, and is usually related to the representation ability of the model. The larger the parameter scale, the better the model can fit the complex data distribution. However, a large number of parameters may require more computing resources for training and inference, so the number of parameters may be an important consideration in resource-limited situations.
[0089] The inference time directly affects the real-time performance and efficiency of the model in practical applications. In general, a model with shorter inference time is more suitable for mobile devices, embedded systems or real-time application scenarios. In the present application, in order to synchronize the ANN classifier with the IMU sensor (sampling rate of 100 Hz), the inference time must be less than 10 ms. The calculation method of the inference time in the present application is as follows: first, classify 10,000 randomly generated input samples to measure the inference time. Then take the average of the results after repeating the measurement 100 times, and finally divide by 10,000.
[0090] The preset condition is to select the scheme with accuracy higher than 93% and appropriate number of IMUs as the action classification scheme for construction workers. The experimental results of all 90 IMU configuration schemes in terms of classification accuracy performance are shown in FIG. 8. Figure 5 Among all the configuration schemes, the configuration scheme with 5 IMUs installed has higher results in accuracy, precision, recall and F1-score than other schemes, but has no obvious advantage compared with the configuration scheme with 4 IMUs installed. In addition, the configuration scheme with 1 IMU installed also has unsatisfactory results in accuracy, precision, recall and F1-score, and the accuracy is much lower than 93%. Therefore, the present application does not regard the configuration scheme with 5 IMUs or 1 IMU installed as a selectable scheme. FIG. 8 shows the experimental results of all 90 IMU configuration schemes in terms of classification accuracy performance. Figure 5 The six selectable IMU configuration schemes that meet the preset condition are marked by the green dashed boxes in FIG. 8.
[0091] The six selectable IMU configuration schemes that meet the preset condition are as follows:
[0092] Scheme 1 (S1): The accuracy of this scheme is 95.71%, and a total of 4 IMUs are installed, of which 2 IMUs are installed on the thighs and the other 2 IMUs are installed on the calves. The time window size is set to 200 ms. Both the accelerometer and the gyroscope are used for feature extraction.
[0093] Scheme 2 (S2): The accuracy of this scheme is 95.54%, and a total of 3 IMUs are installed, of which 1 IMU is installed on the torso and the other 2 IMUs are installed on the thighs. The time window size is set to 200 ms. Both the accelerometer and the gyroscope are used for feature extraction.
[0094] Scheme 3 (S3): The accuracy of this scheme is 95.21%, and a total of 3 IMUs are installed, of which 1 IMU is installed on the torso, and the other 2 IMUs are installed on the lower leg. The time window size is set to 150 ms. Both the accelerometer and the gyroscope are used for feature extraction.
[0095] Scheme 4 (S4): The accuracy of this scheme is 94.87%, and a total of 2 IMUs are installed, both of which are installed on the thigh. The time window size is set to 150 ms. Both the accelerometer and the gyroscope are used for feature extraction.
[0096] Scheme 5 (S5): The accuracy of this scheme is 93.17%, and a total of 2 IMUs are installed, both of which are installed on the lower leg. The time window size is set to 200 ms. Both the accelerometer and the gyroscope are used for feature extraction.
[0097] Scheme 6 (S6): The accuracy of this scheme is 93.34%, and a total of 2 IMUs are installed, which are installed on the thigh and the lower leg (single side) respectively. The time window size is set to 200 ms. Both the accelerometer and the gyroscope are used for feature extraction.
[0098] The 6 optional IMU configuration schemes that meet the preset conditions are compared with a single ANN (baseline model), a CNN and an LSTM model in terms of accuracy performance and real-time performance as follows:
[0099] Accuracy performance: Appendix Figure 6 The experimental results of the 6 optional schemes based on the hierarchical integrated network structure in terms of accuracy performance are shown. The results show that the model of the present application is significantly better than the single ANN model in all schemes, and achieves similar performance to the single CNN and LSTM model. In particular, in scheme S3, the model of the present application shows the highest performance in all indicators. The single ANN model has the lowest performance, with accuracy, precision, recall and F1-score being 0.9052, 0.9240, 0.9052 and 0.8996 respectively. The single LSTM model has the second lowest performance, with accuracy, precision, recall and F1-score being 0.9428, 0.9505, 0.9428 and 0.9389 respectively. Compared with the single LSTM model, the model of the present application improves accuracy, precision, recall and F1-score by 0.99%, 1.48%, 0.99% and 1.04% respectively. In scheme S6, although the accuracy of the model of the present application is slightly lower than that of the single CNN model, the precision, recall and F1-score are increased by 0.21%, 0.22% and 0.05% respectively compared with the single CNN model.
[0100] Real-time performance: Appendix Figure 7 The experimental results of the real-time performance of the six optional schemes based on the hierarchical integrated network structure and the single ANN, CNN and LSTM models are shown. The results show that the model of the present application performs best in the number of parameters and inference time. Among them, the models with the most parameters and the longest inference time are single CNN and LSTM models, respectively. The single ANN model performs the second best in the number of parameters and inference time. The inference time of the single CNN and LSTM models is more than 10 ms, the inference time of the single ANN model is between 2-4 ms, and the model of the present application is maintained within 2 ms. In particular, in S5, the number of parameters of the model of the present application is 2.733k, which is reduced by 73.57%, 93.93% and 72.38% compared with the single ANN, CNN and LSTM models, respectively. The inference time is 1.4 ms, which is reduced by 46.15%, 88.30% and 94.62% compared with the single ANN, CNN and LSTM models, respectively.
[0101] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for classifying a construction worker's action in a hierarchical integrated network structure, characterized by, The method comprises the following steps: S1: acquiring action data of a plurality of actions performed by a construction worker wearing an IMU, preprocessing the action data, and establishing an action data set; S2: selecting different IMU numbers and installation positions, sliding time window sizes, and feature extraction data combinations to obtain a plurality of IMU configuration schemes; a directed label tree is established based on the specificity between the plurality of actions, a hierarchical integrated network is constructed based on the directed label tree, the hierarchical integrated network is trained using corresponding action data in the action data set, and a trained action classification model is obtained; the directed label tree of the 11 actions comprises 6 layers of root nodes; an encoding method of binary labels of the plurality of actions in the directed label tree is that, in the directed label tree, left child nodes and right child nodes are marked as "1" and "0", respectively; all child nodes of the action to be classified are expressed as an N-1-dimensional vector; if a child node is not in the Nth layer, a difference M between N and a layer level at which the child node is located is calculated, and "0" is filled in the last (N-M) bit components; a method for constructing the hierarchical integrated network based on the directed label tree is that: the directed label tree has 10 parent nodes and 11 child nodes, the 11 child nodes represent 11 actions, and the 10 parent nodes include a root node, static actions, dynamic actions, bent knees, everywhere movement, in-place movement, flat ground, uneven ground, stairs, and slopes; an ANN network is constructed for each parent node to classify actions, and the ANN classifier of each parent node is trained separately; an input of each ANN network is action data of all child nodes under the parent node as the root node, and an input variable obtained through preprocessing and signal processing of the action data, and an output of each ANN network is an action represented by the left child node and the right child node; each ANN network comprises an input layer, an encoder layer, a decoder layer, and an output layer, wherein the size of the input layer is determined by an input vector V, the encoder layer, the decoder layer, and the output layer are 8x1, 4x1, and 2x1, respectively, a Soft-max layer is arranged after the output layer, and the ANN classifier outputs a binary label vector to represent an action category; in action classification, the binary label of static standing is [0 1 0 0 0], the binary label of static squatting is [0 0 0 0 0], the binary label of static kneeling is [0 0 1 0 0], the binary label of squatting up is [1 1 0 0 0], the binary label of kneeling up is [1 1 1 0 0], the binary label of walking is [1 0 0 0 0], the binary label of running is [1 0 0 1 0], the binary label of climbing stairs is [1 0 1 0 0], the binary label of descending stairs is [1 0 1 0 1], the binary label of climbing uphill is [1 0 1 1 1], and the binary label of descending downhill is [1 0 1 1 0]; action classification starts from a subnetwork of a root node, an output action node of the subnetwork of the root node is used as a subnetwork of a next level node, and this is repeated until there is no child node of an output node; S3: input the action data of the complete action cycle of each IMU configuration scheme into the trained action classification model, evaluate the performance of each IMU configuration scheme according to the four accuracy performance indicators and the two real-time performance indicators, and select the IMU configuration scheme meeting the preset condition as the optional scheme for the action classification of the construction worker.
2. The layered integrated network structure construction worker action classification method of claim 1, wherein, The method for obtaining the action data in the step S1 is: one IMU is installed on the center of the front segment of each thigh, the front segment of each lower leg and the torso of the subject construction worker; the subject construction worker performs multiple actions of walking, uphill, downhill, ascending stairs, descending stairs, running, standing still, squatting still, kneeling still, squatting up and kneeling up, respectively, and the action data collected by each IMU is wirelessly transmitted to the PC end for data preprocessing; The data preprocessing includes: inhibiting the action data noise by Kalman filtering processing, and cleaning the action data by a data visualization method to obtain the time sequence X of a complete action cycle.
3. The layered integrated network structure builder action classification method of claim 2, wherein, The Z axis of all IMUs points to the front of the subject construction worker, and the Y axis is perpendicular to the sagittal plane of the subject construction worker; The actions of the subject construction worker in one complete action cycle are in sequence: 1) standing still for 30s, 2) static half-kneeling for 30s, 3) static deep squatting for 30s, 4) walking on flat ground at a self-selected speed for 30s, 5) running on flat ground at a self-selected speed for 30s, 6) ascending stairs at a self-selected speed for 30s, 7) descending stairs at a self-selected speed for 30s, 8) ascending a slope of 20m at a self-selected speed, 9) descending a slope of 20m at a self-selected speed, 10) repeating squatting up 5 times at a 4s rhythm, and 11) repeating kneeling up 5 times at a 4s rhythm.
4. The layered integrated network structure builder action classification method according to any one of claims 1-3, characterized in that, The method for constructing the directed label tree is: multiple actions are divided into static actions and dynamic actions according to different action attributes, the left child node and the right child node of the root node are static actions and dynamic actions respectively, the left child node and the right child node of the static action are bent knees and standing still respectively, the left child node and the right child node of the bent knees are static squatting and static kneeling respectively; the left child node and the right child node of the dynamic action are everywhere movement and in-place movement respectively, the left child node and the right child node of the in-place movement are squatting up and kneeling up respectively, the left child node and the right child node of the everywhere movement are flat ground and uneven ground respectively, the left child node and the right child node of the flat ground are walking and running respectively, the left child node and the right child node of the uneven ground are stairs and slope respectively, the left child node and the right child node of the stairs are ascending stairs and descending stairs respectively, and the left child node and the right child node of the slope are ascending and descending respectively.
5. The layered integrated network structure builder action classification method of claim 4, wherein, The loss function for training the hierarchical ensemble network adopts a classification cross-entropy, and the loss function is: where m denotes the number of output actions of a single subnetwork in the hierarchical ensemble network, c m denotes the binary indicator of the mth action true label, p m denotes the classification probability of the mth action by the subnetwork of a single node in the hierarchical ensemble network.
6. The layered integrated network structure construction worker action classification method according to any one of claims 1 to 3, 5, characterized by, The different IMU numbers and installation positions include: 3 cases of installing 1 IMU, respectively installed on the torso, the thigh and the calf; 3 cases of installing 2 IMUs, respectively installed on two thighs, two calves and the thigh & the calf; 2 cases of installing 3 IMUs, respectively installed on the torso & two thighs and the torso & two calves; 1 case of installing 4 IMUs, installed on two thighs and two calves; and 1 case of installing 5 IMUs, respectively installed on the torso, two thighs and two calves; The sliding time window size includes 100 ms, 150 ms and 200 ms; The feature extraction data includes 3D acceleration collected by the accelerometer, 3D angular velocity collected by the gyroscope and 3D acceleration collected by the accelerometer & 3D angular velocity collected by the gyroscope; The different IMU numbers and installation positions, the sliding time window size and the feature extraction data are combined with each other, and there are 90 IMU configuration solutions in total; The action data collected by the multiple IMU configuration solutions and the action data in the action data set are subjected to signal processing before being input into the hierarchical integrated network, and the signal processing method is: extracting the features including the maximum value, the minimum value, the average value, the standard deviation, the first signal value, the middle signal value and the last signal value from each signal channel of the feature extraction data by using the sliding time window to form an input vector V; and inputting the input vector V into the hierarchical integrated network.
7. The layered integrated network structure builder action classification method of claim 6, wherein, The grouping Soft-max method is used to traverse the directed label tree nodes in the breadth-first order, obtain the path encoding label, calculate the Soft-max probability output of each node on the path, and then multiply the Soft-max probability output of each node by the probability output of its parent node to calculate the path probability output; By integrating the probability outputs of the sub-networks of multiple parent nodes, the unique path capable of representing the child node leading to the current action category is indexed out, so as to determine the action category.
8. The layered integrated network structure builder action classification method of claim 7, wherein, Each action node on the path corresponds to a classification probability p m where the root node and other nodes filled with "0" have a classification probability of 1; The label of each action is represented by a path y k = (v1,..., v E ), where v1 and v E are the root node and the terminal node of the path respectively, and E represents the number of nodes on the path; the path probability P(y k ) of the kth action node is calculated by using the group Soft-max method as where v j-1 represents the j-1th node on the path; The 4 accuracy performance indicators include the accuracy, the precision, the recall and the F1 score of the action classification model; and the 2 real-time performance indicators include the parameter number and the inference time of the action classification model. The preset condition is to select the scheme with an accuracy higher than 93% and a suitable IMU number as the action classification scheme of the construction worker. The different IMU numbers and installation positions include: 3 cases of installing 1 IMU, respectively installed on the torso, the thigh and the calf; 3 cases of installing 2 IMUs, respectively installed on two thighs, two calves and the thigh & the calf; 2 cases of installing 3 IMUs, respectively installed on the torso & two thighs and the torso & two calves; 1 case of installing 4 IMUs, installed on two thighs and two calves; and 1 case of installing 5 IMUs, respectively installed on the torso, two thighs and two calves; The sliding time window size includes 100 ms, 150 ms and 200 ms; The feature extraction data includes 3D acceleration collected by the accelerometer, 3D angular velocity collected by the gyroscope and 3D acceleration collected by the accelerometer & 3D angular velocity collected by the gyroscope; The different IMU numbers and installation positions, the sliding time window size and the feature extraction data are combined with each other, and there are 90 IMU configuration solutions in total; The action data collected by the multiple IMU configuration solutions and the action data in the action data set are subjected to signal processing before being input into the hierarchical integrated network, and the signal processing method is: extracting the features including the maximum value, the minimum value, the average value, the standard deviation, the first signal value, the middle signal value and the last signal value from each signal channel of the feature extraction data by using the sliding time window to form an input vector V; and inputting the input vector V into the hierarchical integrated network. The grouping Soft-max method is used to traverse the directed label tree nodes in the breadth-first order, obtain the path encoding label, calculate the Soft-max probability output of each node on the path, and then multiply the Soft-max probability output of each node by the probability output of its parent node to calculate the path probability output; By integrating the probability outputs of the sub-networks of multiple parent nodes, the unique path capable of representing the child node leading to the current action category is indexed out, so as to determine the action category. The 4 accuracy performance indicators include the accuracy, the precision, the recall and the F1 score of the action classification model; and the 2 real-time performance indicators include the parameter number and the inference time of the action classification model; The preset condition is to select the scheme with an accuracy higher than 93% and a suitable IMU number as the action classification scheme of the construction worker.
Citation Information
Patent Citations
Human body action recognition method based on a TP-STG framework
CN109492581A
Method, device and system for retrieving keywords
CN103544281A
3D posture estimation method based on multi-view deep sensor frame
CN108389227A