Automatic plantar pressure classification method and system based on multi-task learning framework

By adopting a multi-task learning framework and deep learning technology in the plantar pressure data processing, combined with DBSCAN clustering and Gaussian filtering processing, the problem of plantar pressure data identification and classification in the existing technology is solved, and high-accurate left and right foot recognition and completeness judgment are achieved, which is suitable for barefoot and shoe-wearing data.

CN120046053AActive Publication Date: 2025-05-27CHENGDU UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202411832917.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-27
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify and classify different types of plantar pressure data, especially in the shoe state, and the uneven sample results in the model biasing towards data with a larger sample size.

Method used

Using a multi-task learning framework method, the sole pressure data is extracted through DBSCAN clustering algorithm and Gaussian filtering processing, combined with the Inception block and channel and spatial attention module of GoogLeNet, an Inception-CBAM block is established, a multi-task learning model is constructed, and the sample loss weight is adjusted using the focal loss function during the training process.

Benefits of technology

The recognition accuracy of left and right feet on barefoot and shoe data reached 99.78% and 98.59% respectively, and the reliability of the method was verified through cross-population testing, solving the problem of sample imbalance and improving the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046053A_ABST
    Figure CN120046053A_ABST
Patent Text Reader

Abstract

The invention discloses a plantar pressure automatic classification method and system based on a multi-task learning framework, and belongs to the field of plantar pressure data research. The method comprises the following steps: S1, collecting plantar pressure data; s2, preprocessing the plantar pressure data: a plantar pressure extraction algorithm: segmenting the plantar pressure data through a DBSCAN clustering algorithm and extracting to generate a footprint sequence; footprint image representation: accumulating and summing the extracted plantar pressure sequence to reduce the length, and performing noise suppression processing on the image through Gaussian filtering; carrying out data enhancement; s3, a channel and a space attention module are introduced into an Inception block of the GoogLeNet, an Inception-CBAM block is established, and a multi-task learning model is constructed; the multi-task learning model is used for carrying out classification processing on the footprint images; and S4, obtaining plantar pressure classification data. The method provided by the invention is not only suitable for barefoot data, but also suitable for complex footwear data in a pressure area, and lays a data preparation foundation for automatic gait analysis and recognition based on plantar pressure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of plantar pressure data research, and particularly to an automatic classification method and system for plantar pressure based on a multi-task learning framework. Background Art

[0002] Plantar pressure refers to the pressure field acting on the sole of the foot and the supporting surface during daily activities of a person. A plantar pressure distribution measurement system can accurately record the dynamic pressure distribution of the sole of the foot with high spatio-temporal resolution, thereby reflecting the characteristics of a person's gait. The kinetic characteristics of different anatomical regions of the sole of the foot can be obtained through plantar pressure data.

[0003] Due to different foot shapes, differences in the texture structure of the shoe sole, and different walking abilities of the experimenters, the footprint patterns formed by plantar pressure vary greatly, and continuous and discontinuous footprints may appear. It is a relatively difficult problem to accurately identify the left and right feet of the experimenter and the integrity. In addition, since the shoe-wearing samples may carry interference information, filtering processing needs to be performed on them before identifying the left and right feet and integrity. Since the sample sizes of 60 experimenters collected are not equal and there are also differences in their step angles, in order to avoid the trained model being too biased towards the footprint patterns of experimenters with larger sample sizes due to unbalanced footprint data, it is necessary to increase the diversity of the experimenters' footprint samples. There are some footprints in the footprint samples where the heel and the sole are disconnected, and the footprint patterns of high-arched feet, flat feet are different from those of normal feet, and it is difficult even for manual discrimination.

[0004] In addition, most current work has studied plantar pressure samples in the case of bare feet or socks on feet, and it is not clear how these technologies will be applicable to different types of footwear. Before this research, only MACDONALD et al. used plantar pressure data collected in the shoe-wearing state for left and right foot classification. However, the footprints collected by MACDONALD are relatively complete, without taking into account the situation that there may be incomplete footprints in daily walking activities. And the plantar pressure morphology after the experimenter wears shoes may be involved in these applications, which is more complex in spatial morphology than the barefoot pressure data, and this may bring additional difficulties to simple comparison-based algorithm annotation (such as: the angle between the two feet, the number of pixels in different parts of the foot, the similarity between the footprint and the template). Summary of the Invention

[0005] The purpose of the present invention is to overcome one or more deficiencies of the prior art and provide an automatic classification method and system for plantar pressure based on a multi-task learning framework.

[0006] The purpose of the present invention is achieved by the following technical solutions:

[0007] An automatic classification method for plantar pressure based on a multi-task learning framework, the steps of the method include:

[0008] S1. Collect plantar pressure data;

[0009] S2. Preprocess the plantar pressure data:

[0010] Plantar pressure extraction algorithm: Segment and extract the plantar pressure data through the DBSCAN clustering algorithm to generate a plantar pressure sequence;

[0011] Footprint image representation: Accumulate and sum the extracted plantar pressure sequence to reduce its length, and represent it as a single-frame image;

[0012] Suppress noise in the image through Gaussian filtering; then perform data augmentation;

[0013] S3. Introduce channel and spatial attention modules into the Inception block of GoogLeNet to establish an Inception-CBAM block and construct a multi-task learning model; the multi-task learning model is used to classify the footprint image;

[0014] S4. Obtain plantar pressure classification data.

[0015] Furthermore, in step S2, the plantar pressure data is extracted through the DBSCAN algorithm and the contour merging algorithm. The metric formula of the DBSCAN algorithm is:

[0016] ;

[0017] Among them, is the Euclidean distance matrix, is the Manhattan distance matrix, and α, β are weight parameters.

[0018] Furthermore, in step S2, footprint image representation: Accumulate and sum the extracted plantar pressure sequence to reduce its length. By reducing the length of the segmented pressure sequence, the model convergence speed can be accelerated; represent a single footprint sequence s i as a single-frame cumulative pressure image , which is expressed as:

[0019] ;

[0020] Among them, , , and respectively correspond to the pressure values at their respective positions in the pressure matrix sequence; is the cumulative sum of the pressure values at the same position in the pressure matrix sequence, and its expression is as follows:

[0021] AP nm = ∑ F nm [ t ] , ( 0 ≤ t ≤ T max ) ;

[0022] Among them, is the maximum value of the effective sampling of the effective plantar pressure sequence, is the pressure received by the sensor at the nth row and mth column in the sensor matrix at time t.

[0023] Furthermore, in step S2, the image is processed by Gaussian filtering to suppress noise, which is expressed as:

[0024] ;

[0025] Among them, k is 1, σ is 1, i and j both take values from -1 to 1. After calculation, the template needs to be processed in one step: normalize all coefficients of the obtained Gaussian template, that is, divide each coefficient by the sum of the template coefficients.

[0026] Furthermore, in step S3, a multi-task learning model is established: modify the single classification head of the two auxiliary classifiers of GoogLeNet into a dual classification head, defined as the left and right foot classification head and the integrity classification head, which are used to process left and right foot prediction and integrity prediction respectively; modify the single classification head structure of the main classifier into a dual classification head, and the rest of GoogLeNet is used as a hard sharing parameter structure.

[0027] Furthermore, in step S3, the focal loss function is added in the training of the multi-task learning model to adjust the sample loss weight in the model training process. The formula of the focal loss function is:

[0028] ;

[0029] ;

[0030] Among them, is the focusing parameter, is the modulation factor, is the prediction probability, , in the left and right foot prediction, the left foot is 1 and the right foot is 0; in the integrity prediction, the complete is 1 and the incomplete is 0.

[0031] Furthermore, in step S3, the total loss of the two types of tasks is calculated by the focal loss function and then weighted. The two types of tasks are the left and right foot recognition task and the integrity recognition task, which is expressed as:

[0032] ;

[0033] Among them, is the weight parameter of the left and right foot recognition task, is the weight parameter of the integrity recognition task, is the loss for the left and right foot recognition prediction task, is the loss for the integrity recognition prediction task. For both tasks, the corresponding losses are calculated by the focal loss function.

[0034] A plantar pressure automatic classification system based on a multi-task learning framework is provided for collecting, processing, classifying, and analyzing plantar pressure.

[0035] The beneficial effects of the present invention are:

[0036] (1) The custom distance metric not only considers the spatial distance between force sensors but also the relationship in the force application time series; compared with the existing methods for extracting and locating the left and right feet from array sensors, it is applicable not only to barefoot data but also to shod data with complex pressure regions; ablation studies show that the features extracted by combining shallow, middle, and deep Inception blocks have the strongest expression ability, the convolutional attention module can improve the model recognition performance, and the focal loss adjusts the loss weights of simple samples and sparse difficult samples, improving the model generalization ability;

[0037] (2) It is proved by existing methods and deep learning methods that barefoot data does not require filtering, while shod data requires filtering; the left and right foot recognition accuracies on barefoot and shod data reach 99.78% and 98.59% respectively, and the reliability is verified through cross-population tests;

[0038] (3) By collecting more plantar pressure data of patients, combining this solution, optimizing the existing algorithm, and integrating it into an automated evaluation software, it is convenient for further research. Brief Description of the Drawings

[0039] Figure 1 is the step flow chart of the method;

[0040] Figure 2 is the difficult sample diagram;

[0041] Figure 3 is the schematic diagram for obtaining experimental data, (a) is for barefoot walking acquisition, and (b) is for shod walking acquisition;

[0042] Figure 4 is the data preprocessing structure diagram;

[0043] Figure 5 is the footprint sequence diagram;

[0044] Figure 6 is the multi-task learning model structure diagram;

[0045] Figure 7 is the schematic diagram of the Inception-CBAM block structure;

[0046] Figure 8 Schematic diagram of the basic structural unit of the network, (a) is the basic convolution block, (b) is the main classifier and (c) is the auxiliary classifier. DETAILED DESCRIPTION

[0047] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0048] See also Figure 1 , provides an automatic classification method for plantar pressure based on a multi-task learning framework, characterized in that the method steps include:

[0049] S1. Collect plantar pressure data;

[0050] S2. Preprocessing of plantar pressure data:

[0051] Plantar pressure extraction algorithm: The plantar pressure data is segmented and extracted using the DBSCAN clustering algorithm to generate a plantar pressure sequence;

[0052] Footprint image representation: The extracted plantar pressure sequence is accumulated and summed to reduce the length and represented as a single frame image;

[0053] The image is processed to suppress noise through Gaussian filtering, and then data enhancement is performed;

[0054] S3. Introduce channel and spatial attention modules into the Inception block of GoogLeNet, establish the Inception-CBAM block, and build a multi-task learning model; the multi-task learning model is used to classify footprint images;

[0055] S4. Obtain plantar pressure classification data.

[0056] 45 male and 15 female subjects were recruited for data collection experiments. The subjects were mainly students (age 23.17±1.3 years, height 169.075±9.19 cm, weight 62.93±11.76 kg, shoe size 39.3±2.29). The subjects' shoes included basketball shoes, skateboard shoes, casual shoes, sports shoes, and Martin boots. The pressure platform parameters are shown in Table 1:

[0057] Table 1

[0058]

[0059] See alsoFigure 3 , the experimenter adopted the Mid-step protocol, that is, after each experimenter walked at least three steps, stepped on the pressure test board, and then walked at least 3 more steps and stopped. The data collected in this way was obtained in the most stable state of walking because the data obtained by the same person at different times had better repeatability. For each experimenter, complete and incomplete data with shoes on and complete and incomplete data without shoes were obtained. The experimental data acquisition method is as Figure 3 shown.

[0060] Data preprocessing:

[0061] Refer to Figure 2 , Figure 4 , in this study, the distance metric standard of DBSCAN was improved according to the temporal and spatial relationship of the force points of the sensors. After clustering into pressure regions, they were automatically merged into the same class of footprints to achieve automatic sample extraction. At the same time, in this embodiment, GoogLeNet was modified into a multi-task network and CBAM was embedded into it, and focal loss was used in the training stage to adjust the loss weights of difficult-to-recognize samples and easy-to-recognize samples. By collecting the complete and incomplete plantar pressure data of 60 people and cross-population testing, the effectiveness of the proposed method was verified, and the results showed that the recognition accuracy was much higher than the recent methods. The data processing framework of this embodiment is as Figure 4 shown. In the data preprocessing stage, the plantar pressure data collected by the Footscan system was first saved, and the left and right footprints and their integrity were labeled based on the self-developed software. For the extraction of footprints, an improved DBSCAN algorithm was used to automatically extract footprints. Considering that the plantar pressure images collected when wearing shoes contain more high-frequency information, and this information may not serve well for footprint classification and the judgment of its integrity, a Gaussian filtering algorithm was adopted to process the high-frequency information. Since the number of samples in each category collected was unbalanced, this embodiment also adopted changing the stepping direction angle to achieve data augmentation.

[0062] Plantar pressure extraction algorithm:

[0063] For the pressure data collected without shoes, due to different foot shapes, the pressure data will show different shapes. For example, the forefoot and hindfoot of a high-arched foot may be separated; for the plantar pressure data collected when wearing shoes, due to the differences in shoe shapes, the footprints present more complex patterns, and the formed pressure regions may be continuous or discontinuous. In addition, there are differences in the step lengths of different people, and the starting points of each person's steps when collecting data are different, making it difficult to complete footprint positioning by presetting the number of footprints.

[0064] In this embodiment, the pressure areas that are relatively close to each other are merged together through a clustering algorithm to form a footprint, and an automatic extraction sequence is realized according to the footprint position. Based on the above considerations, an improved DBSCAN algorithm and a contour merging algorithm are used to complete this part of the work.

[0065] First, according to the three-dimensional plantar pressure matrix Determine the feature array, that is, the effective position index matrix of the footprint on the sensor matrix , and calculate the Euclidean distance matrix and The Manhattan distance matrix of the initial acting force time in . At the same time, for and , multiply them by the weight coefficients α and β respectively, and set the initial values of the two parameters to 0.1 and 0.1 respectively. Then search for the optimal parameters, and set the search step size to 0.01. The optimal α and β obtained through the search are 0.95 and 0.05 respectively. Calculate the DBSCAN metric standard by operating the two matrices according to the formula:

[0066] ;

[0067] The clustering and extraction process is as follows:

[0068] Algorithm 1: Plantar pressure extraction;

[0069] Input , the pressure point scanning radius eps: 0.2, the minimum number of pressure points min_samples: 3, the distance matrix Output the segmented footprint sequence:

[0070] 1) Mark the determined in as unvisited;

[0071] 2) Do;

[0072] 3) Randomly select an unvisited ;

[0073] 4) Mark as visited;

[0074] 5) If the number of pressure points within the eps of

[0075] 6) Create a new cluster , and at the same time add and the points within eps to ;

[0076] 7) Add the points within the eps of to ;

[0077] 8) Traverse (1 ≤ k ≤ max), ;

[0078] 9) if is unvisited;

[0079] 10) Mark as visited;

[0080] 11) if the number of pressure points within the eps of ≥ min_samples, add these pressure points to ;

[0081] 12) if has not been assigned to any cluster, then assign it to the current cluster ;

[0082] 13) End the traversal;

[0083] 14) else Mark as noise;

[0084] 15) until there are no pressure points marked as unvisited;

[0085] 16) Use the labels_ method of DBSCAN to extract the cluster labels after clustering;

[0086] 17) for each cluster;

[0087] 18) Calculate the cluster center label category and merge the clusters belonging to the same label;

[0088] 19) end for;

[0089] 20) Search for the footprint position information on the clustered footprint array;

[0090] 21) Determine the position of the footprint in the dynamic sequence according to the array footprint position;

[0091] 22) Calculate the rectangle with the shortest width and height that encloses the footprint;

[0092] 23) Determine the index of the footprint in the sequence according to the length, width and coordinates to complete the segmentation;

[0093] 24) Output the segmented footprint sequence;

[0094] Footprint image representation:

[0095] After extracting the footprint data using the proposed plantar pressure extraction algorithm, a representative footprint sequence is visualized as Figure 5 shown. There are still many similar frames in the time dimension. If the sequence is input into the network, a large number of image features need to be extracted for each training, which will slow down the convergence speed of the network. By reducing the length of the segmented pressure sequence, the convergence speed of the model can be accelerated. In this embodiment, a single footprint sequence s i is represented as a single-frame cumulative pressure image .

[0096] ;

[0097] wherein, , , and respectively correspond to the pressure values at their respective positions in the pressure matrix sequence; is the sum of the pressure values at the same position in the pressure matrix sequence, and its expression is as follows:

[0098] AP nm = ∑ F nm [ t ] , ( 0 ≤ t ≤ T max ) ;

[0099] wherein, is the maximum value of the effective sampling of the effective plantar pressure sequence, is the pressure received by the sensor at the nth row and mth column of the sensor matrix at time t. After this step, a footprint sequence of dozens of frames is represented as one frame.

[0100] Gaussian filtering processing:

[0101] For the plantar pressure data of wearing shoes, the texture of the sole and sundries will cause the existence of high-frequency information. After a single footprint sequence s is represented as a single-frame pressure image img, the high-frequency information in the sequence still exists, which is not conducive to model classification and integrity judgment. Therefore, it is necessary to filter the single-frame pressure image. The Gaussian filter is a linear filter that can effectively suppress noise and smooth the image. Its function and principle are similar to those of the mean filter, which is to take the mean value of the pixels within the filter window as the output. Before processing the single-frame cumulative pressure image, it is necessary to discretize the continuous Gaussian function to obtain a Gaussian filter template. In this embodiment, a 3×3 template is generated through parameters, and sampling is performed with the center point of this template as the origin. The value obtained by substituting the value of each position into the Gaussian function is the coefficient of the Gaussian template. The calculation method of the Gaussian template in this embodiment is as follows:

[0102] ;

[0103] Here, k is 1, σ is 1, and i and j both take values from -1 to 1. After calculation, the template needs to be further processed: normalize all the coefficients of the obtained Gaussian template, that is, divide each coefficient by the sum of the template coefficients.

[0104] Data augmentation: In this embodiment, aiming at the problem of extremely unbalanced sample numbers among experimenters, a data augmentation method based on the step direction angle is proposed to improve the robustness of the model for footprint pattern recognition. Under normal circumstances, there are natural differences in the step direction angles among experimenters due to physiological characteristics and walking habits. To avoid the model training being overly biased towards experimenters with a large number of samples, this embodiment first identifies the maximum and minimum step direction angles of the left and right feet in the plantar pressure image set of each experimenter (denoted as , and , ). Subsequently, each left-foot and right-foot sample is rotated by an angle equal to the original step direction angle plus a random angle α (generated between and ) to simulate the natural variation of the actual step direction angle. If α is greater than the current angle, the step direction angle is increased; otherwise, it is decreased to ensure that the enhanced data fits the real sample distribution.

[0105] Table 2 Sample distribution table

[0106]

[0107] Table 2 shows the sample distribution before and after data augmentation. This method effectively alleviates the problem of sample imbalance and improves the generalization ability of the model. The sample enhancement algorithm process is as follows:

[0108] Algorithm 2: Sample enhancement:

[0109] Input experimenter samples:

[0110] Output newly generated samples:

[0111] 1) for in :;

[0112] 2) Search pressure area contour;

[0113] 3) if ==1:

[0114] 4) Calculate the rectangular angle of the minimum bounding rectangle of the contour;

[0115] 5) else: Merge ;

[0116] 6) The rectangular angle of the minimum bounding rectangle for calculating the contour;

[0117] 7) end if;

[0118] 8) if : Add to ;

[0119] 9) else : Add to ;

[0120] 10) end if;

[0121] 11) end for;

[0122] 12) Calculate the maximum and minimum angles of the left and right feet;

[0123] 13) Randomly generate an angle value between the maximum and minimum values;

[0124] 14) Calculate the center coordinates of the image;

[0125] 15) Generate a rotation matrix based on the angle and the center coordinates;

[0126] 16) Affine transform the image;

[0127] Multi-task learning model:

[0128] In this embodiment, the left and right foot classification and integrity classification of plantar pressure are targeted (denoted as Task1 and Task2 respectively). For any , since the left and right foot attributes and integrity attributes exist simultaneously, there is a correlation between the two types of tasks, and the two recognition tasks can be parallel. Compared with the single-task model, the multi-task model can utilize the correlation between the left and right foot classification tasks and the integrity classification tasks, learn more general and rich feature representations, and this shared learning mechanism makes the model not overly dependent on the specific features of a certain task, reducing the risk of overfitting.

[0129] The Inception block of GoogLeNet can perform multi-scale processing on feature maps, but it can only handle single tasks. Training separate models to complete Task1 and Task2 not only requires setting hyperparameters for the corresponding recognition tasks, but also takes longer training time. Therefore, in this embodiment, its structure is modified to adapt to multi-tasks. First, the single classification heads of the two auxiliary classifiers of GoogLeNet are respectively modified to dual classification heads, that is, the left and right foot classification heads and the integrity classification heads, and then the single classification head structure of the main classifier is modified to a dual classification head, and the rest of the model network is used as a hard-sharing parameter structure. As Figure 6As shown in the figure, the entire model network consists of a basic convolutional block (Basic convblock), an Inception-CBAM block (Inception-CBAM block), a main classifier (Mainclassifier), and an auxiliary classifier (Auxiliary classifier). The convolutional attention mechanism CBAM is introduced into the Inception block to extract more important spatial features. In terms of the implementation of each module, the basic convolutional block obtains the basic feature map through alternating convolutional pooling operations. The Inception block uses multiple convolutional kernels to extract features and pays more attention to some complex features after passing through CBAM. The multi-task learning model learns the features of two types of tasks through a shared layer (hard sharing) during the training phase. Finally, the main classifier predicts the sample categories of the test set. The auxiliary classifier enables the intermediate layer of the network to also have a classification effect, preventing gradient disappearance and regularizing the network. The Inception-CBAM block is as shown in Figure 7 , and the basic structural units of the network (basic convolutional block, main classifier, auxiliary classifier) are as shown in Figure 8 . Although the Inception block has a strong ability to capture multi-scale features, the recognition task in this embodiment is relatively difficult. To further strengthen important features, CBAM is introduced into the Inception block in the network. This attention mechanism includes a spatial attention module and a channel attention module. The channel attention module generates two descriptors for each channel through global average pooling and global max pooling. Then, the descriptors are input into a multi-layer shared perceptron to capture the relationships between channels. Next, the outputs of the multi-layer perceptron are combined and passed through an activation function to generate channel attention weights. Finally, these weights are multiplied by the original feature map to emphasize or suppress certain channels. The spatial attention module uses channel average pooling and channel max pooling to generate two two-dimensional feature maps through the cross-channel dimension, merges the feature maps, and then uses a two-dimensional convolution to generate a spatial attention feature map. After passing through the activation function, it is multiplied by the original feature to obtain the emphasized or suppressed spatial positions.

[0130] Focal loss parameter settings:

[0131] However, in the left and right foot classification task, due to the fact that some barefoot and shod samples only have pressure on the toes or heels, and the existence of a large number of relatively complete left and right foot samples, it is easy for the network gradient to deviate towards samples with large partial vectors and easy to identify during the training process. In the integrity classification task, due to the structure of the sole, it is difficult to accurately identify the integrity of some shod samples. Inspired by the fact that focal loss can weight the losses of two types of samples to improve the model performance in the case of unbalanced easy and hard samples, focal loss is migrated to this field. In each round of iteration of traversing the training set, among the 128 samples input into the model each time, there are more easy-to-identify samples, and it is easy to make mistakes when identifying samples with fewer pressure areas and larger shape differences. Focal loss will calculate the loss of each sample. When the sample is misidentified, a larger loss value will be obtained. By weighting this loss to increase the contribution of the misidentified sample, the sample loss weight during the training process is adjusted to improve generalization. This will make the network pay more attention to the misidentified samples during the training process while not ignoring the simple samples. The focal loss formula is as follows:

[0132] ;

[0133] where γ is the focusing parameter, is the modulation factor.

[0134] ;

[0135] where p ∈ [0, 1] of, is the probability distribution of the predicted sample categories by the multi-task model during the training phase, y ∈ {0, 1}, with the left foot being 1 and the right foot being 0 when predicting the left and right feet. When predicting integrity, complete is 1 and incomplete is 0. When tends to 1 and is an easy-to-identify simple sample, at this time calculate , reducing the contribution of simple samples to the loss. When approaches 0, and the sample is misclassified, obtains a larger weight for the hard sample, increasing the loss value of this sample. After calculating the loss weights of simple samples and hard samples through the predicted probability of the sample and the focusing parameter γ, then multiply the weight by its cross-entropy loss, so that the model pays attention to hard samples during the training phase and avoids the network being biased towards a large number of simple samples.

[0136] In this embodiment, focal loss is introduced as the loss function to balance the loss weights of hard samples and simple samples of the multi-task model during the training phase, improving the model performance. After the focal loss function balances the loss weights of hard samples and simple samples, a weighted sum is performed to obtain the total loss, expressed as:

[0137] ;

[0138] Among them, weight1 is the weight parameter for the left and right foot recognition task, and weight2 is the weight parameter for the integrity recognition task. is the loss of the left and right foot recognition task. is the loss of the integrity recognition task. The corresponding losses of both tasks are calculated by focal loss. In this embodiment, the initial values of the gamma parameter and the weights of the two tasks are set to 0.5, 0.1, and 0.9 respectively. Gamma and the task weights go through 7 rounds and 9 rounds of iteration respectively. Gamma adds a step size of 0.5 in the new round of iteration, and the task weights add and subtract a step size of 0.1 respectively in the new round of iteration. The task weights do not include 0 and 1 because the model needs to complete two recognition tasks simultaneously. When any one of the weights is 0 or 1, the multi-task model degrades to a single-task model.

[0139] When modeling barefoot data, the weight parameters of Task1 and Task2 are 0.9 and 0.1 respectively, and when the gamma parameter of focal loss is set to 3, the model performs best, with the recognition rate of the left and right feet reaching 99.78% and the recognition rate of integrity reaching 97.79%. When modeling shod data, the weight parameters of Task1 and Task2 are 0.5 and 0.5 respectively, and when the gamma parameter of focal loss is set to 1, the model performs best on shod data, with the recognition rate of the left and right feet reaching 98.59% and the recognition rate of integrity reaching 94.35%.

[0140] Comparison method:

[0141] This embodiment is compared with the latest left and right foot classification methods, including the comparison based on the center of plantar pressure, the similarity comparison between the footprint and the template (COP End Points (EP), COP Dynamic Time Warping (DTW), P100 Template Matching (TM)), and the deep learning-based methods (P100 Convolutional Neural Network (CNN), COP Temporal Convolutional Networks (TCN)). In addition, this embodiment also uses image classification models including AlexNet, DenseNet, and Vision Transformer (VIT) as baseline models for comparison. Among them, the AlexNet network uses the LRN normalization layer, enhancing the generalization ability and robustness of the model. The dense connection of DenseNet can alleviate the vanishing gradient problem, strengthen feature propagation, promote feature reuse, and greatly reduce the number of parameters. Different from convolution, VIT directly slices the image and then encodes each slice, introducing the position information of each slice into the model through position encoding.

[0142] Results and Discussion:

[0143] Experimental Setup:

[0144] The method of subject-wise data division is adopted in the article. The data of 48 people is used to train the model, and the data of 12 people is used to test the model. The ratio of the number of the two is 8:2. At the initial stage, the learning rate is set to 2e-5, the learning rate decay factor is 1 / 10, and the learning rate is adjusted once every 4 rounds of iteration of the model. And Adam is used as the parameter optimizer of the model, the weight decay parameter is set to 1e-3, the batch_size is set to 128, the sample size is (128, 128), and the loss function is the weighted loss sum of Task1 and Task2. The loss functions of the deep learning-based comparison methods are all cross-entropy losses, the optimizer is Adam, the learning rate is 2e-5, the weight decay is 1e-3, the learning rate decay factor is 0.1, and the batch size is 128. The three methods of COP EP, COP DTW, and P100 T.M. are neither machine learning methods nor deep learning methods, so no hyperparameters need to be set.

[0145] Experimental Results and Analysis:

[0146] Comparative Analysis with Baseline Methods:

[0147] The results comparison of this embodiment with several image classification models is shown in Tables 3 and 4. Compared with the image classification method, the deep learning method proposed in this paper has achieved the highest accuracy under the subject-wise division. Under this division, the recognition rates of the left and right feet data of wearing shoes and bare feet are 98.59% and 99.78% respectively, and the recognition accuracies of integrity are 94.35% and 97.79% respectively. In addition, under the above division, the recognition rate of barefoot data mostly decreases after filtering, while the recognition rate of wearing shoes data mostly increases after filtering.

[0148] Through the comparison results of the image classification model before and after barefoot filtering, it is found that except for VIT, the recognition accuracies of the left and right feet of AlexNet, DenseNet and the model proposed in this embodiment are all above 93%, and the accuracies of these four models in recognizing integrity basically reach above 96%, indicating the reliability of the deep learning method in image recognition. Through the comparison results of the image classification model before and after wearing shoes filtering, it is found that the results of the four models in recognizing the left and right feet are lower than those of bare feet, indicating that the footprint pattern of wearing shoes is more complex and difficult to recognize than that of bare feet. The recognition effect of VIT on the left and right feet is still lower than that of AlexNet, DenseNet and the model proposed in this embodiment, because after VIT cuts the image into blocks, the continuity between the original adjacent pressure blocks is no longer retained, and the continuity of the plantar pressure points is the key to recognizing the left and right feet, and this is also the limitation of VIT itself, that is, it lacks inductive bias compared with convolution. While AlexNet, DenseNet and the model proposed in this embodiment all use two-dimensional convolution to extract image features, can utilize the continuity between pixels, and are well applicable to the tasks of this embodiment.

[0149] Another discovery is that except for VIT, most of the recognition accuracies of the left and right feet are higher than those of integrity recognition, which may be because integrity recognition is more complex than left and right feet recognition.

[0150] The influence of filtering on recognition:

[0151] Table 3 Comparison of the effects of barefoot pressure data before and after filtering (unit: %)

[0152]

[0153] Table 4 Comparison of the effects of wearing shoes pressure data before and after filtering (unit: %)

[0154]

[0155] As can be seen from Table 3, after filtering the barefoot sole pressure data, the performance of identifying the left and right feet and the integrity judgment decreases in most cases. On the model proposed in this embodiment, the recognition rate of the left and right feet decreases by 0.22% after filtering. This may be because the interference information in the original barefoot data is relatively small, and filtering such samples will lead to a reduction in the details of the original image, thereby reducing the discrimination degree and resulting in misclassification. As can be seen from Table 4, after filtering the shod sole pressure data, the accuracy of identifying the left and right feet and the integrity judgment increases in most cases. On the model proposed in this embodiment, the integrity judgment is improved by 0.23% after filtering. This shows that it is necessary to filter the shod foot pressure data because the structure of the sole pattern and the material of the sole have an impact on the sole pressure pattern, and this impact may bring more high-frequency information. Filtering can weaken this impact, thereby improving the recognition rate of the model.

[0156] Comparison with the recent methods:

[0157] The comparison between the model proposed in this embodiment and the recent classification methods is shown in Table 5. Whether it is the data collected in the barefoot situation or the data collected in the shod situation, the method of this embodiment has advantages in identifying the left and right feet and the integrity judgment.

[0158] Table 5 Comparison with the recent classification methods (unit: %)

[0159]

[0160] Recently, MACDONALD et al. proposed COP EP, COP DTW, and COP TCN to distinguish the sole pressure data of the left and right feet. These three methods classify the left and right feet through the Anterior-Posterior (AP) and Medio-Lateral (ML) sequences of COP. By comparing the results of barefoot and shod data, it can be seen that COP EP and COP DTW perform similarly on the two types of data samples. The former judges left and right by comparing the sizes of the heel and toe in the ML direction, and the latter distinguishes the left and right feet by calculating the similarity between the DTW templates of the left and right feet and the samples. Due to the incomplete situation of the collected samples, for incomplete time series samples, since COP EP only judges by comparing the values at the endpoints of the sequence, this is likely to lead to misjudgment of incomplete footprints in the left and right foot recognition tasks. And COP DTW needs to find a representative template, and it contains incomplete samples. After averaging the COP samples of the training set, this may cause a large difference between the COP time series template and the COP sequence obtained from normal footprint data. This is the reason for the poor recognition effect of COP EP and COP DTW.

[0161] COP TCN is a causal convolutional neural network for time series, that is, the current inference result only depends on the time steps before the input sequence and does not rely on the subsequent time steps. The performance of TCN is better than that of COP EP and COP DTW. This is because TCN introduces residual connections and causal convolutions. Residual connections enable the network to have the effect of identity mapping and ensure that the learning effect will not deteriorate due to deepening the number of network layers during training. Causal convolutions are more suitable for processing time series data than ordinary convolutional layers and can extract more representative features. However, the effect of this method is not as good as that of the method in this embodiment, which may be caused by the difference in length between the incomplete sequence and the complete sequence. The incomplete COP sequence needs to be zero-padded before being input into the model for training, which may make it difficult to learn discriminative features from the incomplete footprint COP data. The pressure sequence representation method used in this embodiment avoids the problem of inconsistent sequence lengths.

[0162] P100 CNN classifies the left and right feet and integrity by extracting image features. P100 CNN directly uses the pressure image as the input, extracts features through two-dimensional convolution, and then uses a fully connected layer for classification. However, its recognition effect is not as good as the model proposed in this embodiment. This may be because the two convolutional layers can only extract relatively shallow features and fail to learn more advanced image features, resulting in insufficient feature expression ability. P100 CNN has a good effect in recognizing integrity because it is more accurate to judge integrity through images. Since there are some similar left and right footprints in the walking habits of the experimenters, it is difficult to distinguish them only using the extracted shallow image features, which is the reason for the poor recognition effect of the left and right feet.

[0163] The performance of template matching (P100 TM) is not good because there are a large number of incomplete samples in the collected samples. Before calculating the SAD, this algorithm needs to find the best alignment direction. The direction angles calculated from the incomplete samples are very different from the direction angles of the templates, and template matching is sensitive to the direction angles of the samples. When the direction angles differ too much, the error of SAD will also increase, resulting in recognition errors. In addition, the performance of template matching depends on the construction of representative templates, and it is unrealistic to construct such templates in the case of a large dataset.

[0164] Among the existing classification algorithms, the comparison-based methods (COP EP, COP DTW, P100 TM) are not as accurate as the deep learning-based methods (COP TCN, P100 CNN) in terms of recognition accuracy; the existing deep learning methods are not as good as the method proposed in this embodiment. This is because the existing comparison-based methods do not consider the situation of incomplete footprint data; the time and space feature expression abilities extracted by the competitive deep learning methods COP TCN and P100 CNN are limited, so the accuracy of the model is still much lower than the method proposed in this embodiment.

[0165] Cross-population testing:

[0166] To verify the performance of the plantar pressure extraction algorithm proposed in this embodiment on other datasets, the dataset collected by YI is used in this embodiment. This dataset includes the plantar pressure data of 29 healthy and 21 hemiplegic middle-aged and elderly people. Since the walking ability of middle-aged and elderly people is weak and they are prone to falling, only the data in the shod state is collected in this dataset. The overall situation of the dataset is shown in Table 6. When processing the samples of healthy middle-aged and elderly people, the hyperparameters eps and min_samples of DBSCAN are set to 0.2 and 3 respectively, and the hyperparameters α and β of the distance matrix are 0.95 and 0.05 respectively. Under these parameter settings, the samples are completely correctly segmented. When processing the data of hemiplegic young, middle-aged and elderly people, the hyperparameters eps and min_samples of DBSCAN are 0.12 and 3 respectively, and the hyperparameters of the distance matrix are the same as those of healthy middle-aged and elderly people. The number of mis-segmented footprints obtained with such parameter settings is 9.

[0167] Table 6 Elderly experimental subjects

[0168]

[0169] The segmentation method proposed in this embodiment has a worse effect on the data of hemiplegic experimental subjects than that of healthy experimental subjects. This is because the footprints of hemiplegic experimental subjects are smaller than the stride of healthy people, and there may even be a state where two consecutive footprints overlap. The algorithm directly groups them together, resulting in the segmented footprints being abnormal footprints, while the footprint spacing of healthy experimental subjects is larger than that of hemiplegic experimental subjects, making it easier to segment.

[0170] To verify the reliability of the proposed multi-task model in plantar pressure recognition, the author conducts recognition task tests on the samples with correct segmentation in the above-mentioned middle-aged and elderly people dataset. The test results are shown in Table 7.

[0171] Table 7 Cross-population testing (unit: %)

[0172]

[0173] As can be seen from Table 7, the multi-task model trained on the shod data of young students has good generalization on the data of healthy middle-aged and elderly people, and the left and right foot recognition rates and integrity recognition rates reach 98.12% and 92.79% respectively; the recognition rates on the hemiplegic patient dataset are relatively poor, which are 95.62% and 87.23% respectively. This is because students are all normal people and their footprint patterns are more similar, while the generalization on hemiplegic patients is weak because the footprint patterns of hemiplegic patients are quite different from those of students, and the footprints left by some hemiplegic patients under normal walking conditions are similar to the incomplete footprints of normal people, resulting in misjudgment by the model.

[0174] Comparison between multitask and single-task:

[0175] In this embodiment, the multitask model is modified into a single-task model, and models 1 and 2 are trained on the shoe-wearing data of students for the two tasks respectively. The parameter settings during the model training are the same as those in 3.1. First, the best gamma parameters are found for the two models respectively, and the gamma parameters found are both 4. And the recognition accuracy of model 1 for left and right feet is 1.65% lower than that of the multitask model, and the recognition accuracy of model 2 for integrity is 0.47% lower than that of the multitask model. To observe the generalization difference between the single-task model and the multitask model on the data collected by YI, the test results of Task1 of model 1 on the data of healthy and hemiplegic middle-aged and elderly people are 97.49% and 92.34% respectively, and the test results of Task2 of model 2 on the data of healthy and hemiplegic middle-aged and elderly people are 89.03% and 81.39% respectively, both lower than the recognition rate of the multitask model. And in the case of the same model size, the single-task model can only complete one recognition task. In this embodiment, a multitask learning model is used, with hard sharing of parameters between tasks. Because the tasks are correlated, and the two tasks are weighted to reduce the possibility of parameter tearing. In addition, the two types of tasks can play a role of regularization for each other through hard sharing of parameters. By learning the parameters of another task, it can prevent the model from being biased towards a single task and reduce the risk of overfitting.

[0176] Ablation experiment:

[0177] To quantify the impact of different levels in the Inception module on the overall performance of the model, we gradually conducted ablation experiments. The original GoogleNet has 3 Inception layer structures, namely the shallow layer (Inception 3a, Inception 3b), the middle layer (Inception 4a, Inception 4b, Inception 4c, Inception 4d, Inception 4e), and the deep layer (Inception 5a, Inception 5b). By removing or retaining different levels, we analyzed the contributions of these three-layer modules and CBAM to the classification performance of the model. Because the shoe-wearing data is more in line with the real scenario, the ablation of different levels of the network is only carried out on the shoe-wearing data.

[0178] When only the shallow blocks are retained, the model performance is poor. When the shallow and middle-level modules are combined, the performance is improved. This indicates that on the basis of supplementing the shallow features, the middle-level features can effectively improve the model's ability to capture higher semantics. When the shallow and deep modules are combined, the accuracy of the model is further improved, which shows that the deep features are crucial for the extraction of high-order semantic information. When only the middle-level modules are retained, the recognition rate is second only to the combination of shallow, middle-level, and deep modules, indicating that the middle-level features play a key role in capturing complex semantics. When the middle-level and deep modules are combined, the model performance is slightly lower than that when only the middle-level modules are used, probably because the middle-level features are already sufficient to capture most of the important semantics. When all of the shallow, middle-level, and deep modules are involved, the model has the best effect, and the accuracy rates of the two types of tasks are increased by 0.71% and 0.7% respectively after adding CBAM.

[0179] Through ablation experiments, we found that the role of the shallow module is limited. It can only capture low-level features and performs poorly when used alone. The middle-level module contributes significantly to the model performance and can independently capture discriminative features. Although the deep module is important for the extraction of high-order semantic information, its independent effect is not as good as that of the middle-level module. The optimal performance is in the combination of shallow + middle-level + deep, indicating that features at different levels complement each other in the model and form the most effective feature representation.

[0180] In this embodiment, DBSCAN is improved. The custom distance metric standard not only considers the spatial distance between force sensors but also the relationship in force application time series. Compared with the existing methods for extracting and locating the left and right feet from array sensors, the method proposed in this embodiment is applicable not only to barefoot data but also to shod data with complex pressure regions. Ablation studies show that the feature expression ability extracted by combining shallow, middle-level, and deep Inception blocks is the strongest. The convolutional attention module can improve the model recognition performance. Focal loss adjusts the loss weights of simple samples and sparse difficult samples, improving the model generalization ability. It is proved by existing methods and deep learning methods that barefoot data does not require filtering, while shod data requires filtering. The method proposed in this embodiment achieves left and right foot recognition accuracy rates of 99.78% and 98.59% respectively on barefoot and shod data, and the reliability is verified through cross-population tests.

[0181] This embodiment improves DBSCAN. The custom distance metric standard not only considers the spatial distance between force sensors, but also the relationship in force application time series. Compared with the existing method of extracting and positioning the left and right feet from array sensors, it is applicable not only to barefoot data, but also to shod data with complex pressure areas. Ablation studies have shown that the feature expression ability is the strongest when combining shallow, middle, and deep Inception blocks. The convolutional attention module can improve the model recognition performance. Focal loss adjusts the loss weights of simple samples and sparse difficult samples, improving the model generalization ability. It is proved by existing methods and deep learning methods that barefoot data does not require filtering, while shod data does. The left and right foot recognition accuracies on barefoot and shod data reach 99.78% and 98.59% respectively, and the reliability is verified through cross-population tests, which is better than other proposed methods. More plantar pressure data of patients can be collected, the existing algorithm can be optimized, and it can be integrated into an automated evaluation software for further research.

[0182] The dataset of this embodiment is established based on normal people and a small number of hemiplegic patients. Therefore, the extraction of footprints of plantar pressure, the recognition of left and right feet, and the integrity can still be verified on a larger dataset. In future work, it is planned to collect more plantar pressure data of patients to optimize the existing algorithm and integrate it into an automated evaluation software.

[0183] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for automatic plantar pressure classification based on a multi-task learning framework, characterized in that: The method steps include: S1. Collect plantar pressure data; S2. Preprocessing of plantar pressure data: Plantar pressure extraction algorithm: The plantar pressure data is segmented and the plantar pressure sequence is extracted using the DBSCAN clustering algorithm; Footprint image representation: The extracted plantar pressure sequence is accumulated and summed to reduce the length and represented as a single frame image; The image is processed to suppress noise through Gaussian filtering, and then data enhancement is performed; S3. Introduce channel and spatial attention modules into the Inception block of GoogLeNet to build a multi-task learning model; the multi-task learning model is used to classify footprint images; S4. Obtain plantar pressure classification data.

2. The method for automatic plantar pressure classification based on a multi-task learning framework according to claim 1, characterized in that: In step S2, plantar pressure data is extracted by using the DBSCAN algorithm and the contour merging algorithm. The measurement formula of the DBSCAN algorithm is: ; in, is the Euclidean distance matrix, is the Manhattan distance matrix, α and β are weight parameters.

3. The method for automatic plantar pressure classification based on a multi-task learning framework according to claim 1, characterized in that: In step S2, the footprint image is represented by: the extracted plantar pressure sequence is accumulated and summed to reduce the length, and the single footprint sequence s i Represented as a single frame accumulated pressure image , expressed as: ; in, , , and Then they are respectively expressed as the pressure values ​​at their respective positions in the pressure matrix sequence; is the same in the pressure matrix sequence The cumulative sum of the pressure values ​​at a location is expressed as follows: ; in, is the maximum value of the valid sampling of the valid plantar pressure sequence, It is the pressure exerted on the sensor in the nth row and mth column in the sensor matrix at time t.

4. The method for automatic plantar pressure classification based on a multi-task learning framework according to claim 1, characterized in that: In step S2, the image is subjected to noise suppression processing by Gaussian filtering, which is expressed as: ; Among them, k is 1, σ is 1, i and j both take values ​​from -1 to 1. After calculation, the template needs to be processed in one step: all coefficients of the obtained Gaussian template are normalized.

5. The method for automatic plantar pressure classification based on a multi-task learning framework according to claim 1, characterized in that: In step S3, a multi-task learning model is established: the single classification heads of the two auxiliary classifiers of GoogLeNet are modified into dual classification heads, which are defined as left and right foot classification heads and integrity classification heads, and are used to process left and right foot prediction and integrity prediction respectively; the single classification head structure of the main classifier is modified into a dual classification head, and the rest of GoogLeNet is used as a hard shared parameter structure.

6. The method for automatic plantar pressure classification based on a multi-task learning framework according to claim 1, characterized in that: In step S3, a focal loss function is added to the multi-task learning model training to adjust the sample loss weight during the model training process. The focal loss function formula is: ; ; in, is the focusing parameter, is the modulation factor, is the predicted probability, When predicting the left and right feet, the left foot is 1 and the right foot is 0; when predicting the completeness, the complete is 1 and the incomplete is 0.

7. The method for automatic plantar pressure classification based on a multi-task learning framework according to claim 1, characterized in that: In step S3, the total loss of the task is calculated by the focal loss function, which is expressed as: ; in, is the weight parameter of the left and right foot recognition task, Identify task weight parameters for completeness, The prediction task loss for left and right foot recognition, To completely identify the prediction task loss, the tasks are all calculated using the focal loss function to calculate the corresponding loss.

8. An automatic plantar pressure classification system based on a multi-task learning framework, characterized in that: The system uses an automatic plantar pressure classification method based on a multi-task learning framework as described in claims 1 to 7.

Citation Information

Patent Citations

  • Left and right foot dynamic recognition method based on plantar pressure distribution information

    CN104434128A

  • Terrain classification device and method based on surface electromyographic signals and plantar force

    CN111053555A

  • Indoor elder walking health detection method and system, storage medium and terminal

    CN112684430A

  • Plantar pressure image processing method, plantar pressure image recognition method and gait analysis system

    CN112766142A

  • Egg freshness detection method based on Inception module and Attention mechanism

    CN113012244A

Cited By

  • Shuttlecock holding action intelligent identification and scoring method based on cross-feature interaction

    CN121388706A