A Gesture Recognition Method and System Based on Multimodal Optimal Data Selection and Enhancement

Through multimodal optimal data selection and enhanced gesture recognition methods, the problem of low data acquisition burden and recognition accuracy in the prior art is solved, and gesture recognition with high accuracy under single acquisition is achieved, which is suitable for patients with motor dysfunction in rehabilitation training.

CN116956063BActive Publication Date: 2025-08-05BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310840617.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-10
Publication Date
2025-08-05
Estimated Expiration
2043-07-10

AI Technical Summary

Technical Problem

The existing gesture recognition algorithm based on sEMG requires multiple data collection in rehabilitation training, which increases the burden on patients. Transfer learning and signal enhancement methods have problems with low recognition accuracy, especially when facing patients with motor dysfunction, there is a large difference between the EMG signal domains and insufficient sample diversity.

Method used

The gesture recognition method of multimodal optimal data selection and enhancement is adopted. By establishing a multimodal signal database, the best matching signal screening and signal enhancement are performed. The long and short-term timing Transformer multimodal gesture recognition model is used, combining joint angle regression and upper limb pose estimation to reduce differences between users and improve recognition accuracy.

Benefits of technology

Under a single acquisition of each gesture, the recognition accuracy is equivalent to training of sufficient surface electromyography signal data sets, reducing the burden of multiple acquisitions, and improving the accuracy and generalization of gesture recognition in small samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116956063B_ABST
    Figure CN116956063B_ABST
Patent Text Reader

Abstract

The present invention discloses a gesture recognition method based on multimodal optimal data selection and enhancement, including establishing and training a gesture recognition model and performing online gesture recognition based on the gesture recognition model. Establishing and training the gesture recognition model includes: establishing a multimodal signal database, collecting and forming a new user calibration gesture dataset; performing optimal matching signal screening on new user calibration gesture multimodal signal samples in the new user calibration gesture dataset based on multimodal signal samples in the multimodal signal database to obtain the optimal matching signal to form a first training set; performing signal enhancement based on the optimal matching signal to form a second training set; training a long-short time series Transformer multimodal gesture recognition model based on the first and second training sets; and performing online gesture recognition based on the gesture recognition model includes inputting the online collected new user multimodal gesture signals into the model for recognition and obtaining gesture categories. The present invention also discloses a corresponding system, electronic device, and computer-readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of gesture recognition processing and rehabilitation training technology, and in particular to a gesture recognition method and system based on multimodal optimal data selection and enhancement. Background Art

[0002] Current sEMG-based gesture recognition algorithms are based on traditional recognition algorithms used in common scenarios. These algorithms require patients to perform the intended movements multiple times before rehabilitation training to collect sufficient data for model training. These patients often experience one or more movement disorders, such as bradykinesia, decreased muscle strength, and fatigue. The cumbersome data collection process increases the data collection burden, imposes significant physical and time costs on patients, and impacts their rehabilitation experience.

[0003] To address the above issues, existing technologies propose to achieve small-sample gesture recognition through transfer learning and signal enhancement. Transfer learning uses a small amount of labeled or unlabeled data from the target domain to calibrate the model of the source domain user to achieve cross-domain recognition. This algorithm usually needs to take into account the individual differences between different users. It can transfer learning between different users and train the gesture recognition model with a small amount of data, effectively solving the problem that the recognition model based on individual surface electromyography signals requires multiple signal acquisitions for each patient. Signal enhancement is a method to solve the problem of small-sample gesture recognition from the perspective of expanding signal samples. When faced with patients with movement disorders, the number of available signal samples is very limited. Signal enhancement uses the information of existing signal samples to generate new and more signal samples to expand the training set and improve model performance. However, there are still the following problems in transfer learning and signal enhancement when constructing small-sample gesture recognition models: (1) There are large differences between domains of electromyographic signal samples. Transfer learning methods can only reduce the differences between domains to a certain extent, which is prone to negative transfer and affects recognition accuracy; (2) The samples obtained in the GAN-based signal enhancement method are generated through adversarial training. The generator will generate samples that are as similar to the real dataset as possible, making it impossible for the discriminator to distinguish between true and false. Therefore, the diversity of generated samples will be limited, and it is impossible to simulate the diversity of electromyographic signals under the same gesture. Summary of the Invention

[0004] To address the problems existing in the prior art and reduce the burden of data collection, the present invention provides a gesture recognition method based on multimodal optimal data selection and enhancement. Based on a small sample of data collected for each gesture, hand motion information is introduced to increase the universality of gesture features across users, thereby resolving the negative transfer problem caused by large differences between EMG signal domains and the problem of insufficient sample diversity generated by signal enhancement methods. The method proposes optimal matching signal screening and signal enhancement based on signal similarity calculation. Based on small sample multimodal physiological signals, when facing a new user (patient), each gesture only needs to be collected once to achieve a recognition accuracy comparable to that of a traditional recognition model trained with sufficient surface EMG signal data sets, thus avoiding the burden of multiple collections on the patient.

[0005] The method of the present invention is based on the theoretical foundation that data from different physiological sensors and devices can be combined in recognition tasks such as joint angle regression and upper limb posture estimation. It adopts multimodal data fusion to obtain more comprehensive, accurate and reliable physiological information to improve the recognition accuracy and generalization of the model. The present invention applies the multi-source data fusion method to small-sample gesture recognition, thereby supplementing the universal physiological information of gesture movement, reducing differences between user domains, and improving the accuracy of small-sample gesture recognition.

[0006] On one hand, the present invention proposes a gesture recognition method based on multimodal optimal data selection and enhancement, comprising establishing and training a gesture recognition model and performing online gesture recognition based on the gesture recognition model; wherein the gesture recognition model is a long-short time series Transformer multimodal gesture recognition model;

[0007] The establishing and training of the gesture recognition model includes: S1, establishing a multimodal signal database based on existing data, collecting multimodal signal samples of new user calibration gestures to form a new user calibration gesture data set, and performing best matching signal screening on the new user calibration gesture multimodal signal samples in the new user calibration gesture data set based on the multimodal signal samples in the multimodal signal database to obtain the best matching signal to form a first training set; S2, performing signal enhancement based on the best matching signal to form a second training set; S3, training the long-short time series Transformer multimodal gesture recognition model based on the first training set and the second training set; performing online gesture recognition based on the gesture recognition model includes: S1', collecting new user multimodal gesture signals online; S2', inputting the new user multimodal gesture signals into the trained long-short time series Transformer multimodal gesture recognition model to obtain a gesture category after recognition.

[0008] Preferably, the S1 includes: S11, collecting multimodal signal samples of a new user's single repeated gesture to form a new user calibration gesture data set; S12, accessing a database; the database stores a variety of gesture signals collected historically; S13, calculating the similarity between the calibration gesture signal and the variety of gesture signals in the new user calibration gesture data set, thereby determining the signal similarity between the new user's calibration gesture and the variety of gesture signals stored in the database; S14, determining the best matching signal as the first training set based on the multimodal signal adaptive selection strategy and the signal similarity, the first training set being a partial training set for training the long-short time series Transformer multimodal gesture recognition model.

[0009] Preferably, the similarity calculation method in S13 comprehensively describes the similarity between the data in the database and the new user signal from the time domain and frequency domain perspectives according to the number of gesture repetitions; the time domain similarity calculation part adopts the time warping algorithm (Dynamic Time Warping, DTW), and the frequency domain part adopts the mean square error (MSE) to calculate the Fast Fourier Transform (Fast Fourier Transform). Transform, FFT); including: (1) forming a database of existing modal data according to the number of gesture repetitions, and calculating the signal similarity according to the gesture type, first calculating the first type of gesture signal similarity; (2) under the first type of gesture, selecting the single modal data of the first repeated gesture for similarity calculation; (3) calculating the time domain similarity between the first repeated gesture and the calibration gesture of the new user based on the DTW algorithm; (4) calculating the frequency domain similarity between the first repeated gesture and the calibration gesture of the new user based on the mean square error of FFT; (5) time-frequency domain similarity scaling: in the multimodal signal adaptive selection module, the time domain similarity and frequency domain similarity values of the two signals are uniformly scaled to the same scale; (6) establishing a similarity graph: repeating step (1) to obtain the time domain similarity set D of each gesture in the database and the calibration gesture TD and frequency domain similarity set D FD ; Take the time domain similarity of the signal as the horizontal axis and the frequency domain similarity of the signal as the vertical axis. TD and frequency domain similarity set D FD Forming a time-frequency domain pair Form a similarity graph; (7) Calculate signal similarity, and calculate the distance between each point in the graph and the calibration gesture as the signal similarity value; repeat step (3) until the similarity value of all data of this type of gesture in the database is obtained, and sort them in ascending order to indicate the similarity size; after the calculation is completed, if there are other gestures in the database whose similarity has not been calculated, repeat steps (2) to (7) to perform the next round of signal similarity calculation; if not, sort all signals in the database by similarity; (8) Sort by database signal similarity: After calculating the signal time domain and frequency domain similarity, a unique similarity value is obtained as a measurement standard to comprehensively consider the signal similarity of each modal signal in the database in the time domain and frequency domain, and sort the similarity between the signals in the database and the calibration gesture.

[0010] Preferably, the multimodal signal adaptive selection strategy includes: (1) selecting and combining the first N similar data: obtaining the similarity ranking of each modal signal according to the signal similarity calculation module, setting the N value, increasing N from 1, 2, 3, ... to L in sequence, selecting data on each modality, cutting them to the same sequence length for combination, and making a training set, where L represents the total number of modal signals; (2) model training: training the data set obtained in step (1) under the network LST-EMG-Net to verify the recognition accuracy; (3) implementing the early stopping model training strategy: if N = n, the recognition accuracy Accuracy n >Accuracy n+1 Accuracy n >Accuracy n+2 When n is the optimal value, the first n similar data are the best matching data.

[0011] Preferably, the S2 includes: S21, generating a plurality of expanded new user signal samples and / or best matching signal samples using a variational autoencoder network structure; S22, calculating the difference between the generated new user signal samples and / or best matching signal samples and the original samples of the new user calibration gesture and / or best matching signal based on a time series differentiable loss function; S23, repeating steps S21-S22 after updating the parameters of the variational autoencoder network structure based on the difference and back propagation, thereby obtaining a second training set required for training the long-short time series Transformer multimodal gesture recognition model, wherein the second training set is another part of the training set used to train the long-short time series Transformer multimodal gesture recognition model.

[0012] Preferably, the variational autoencoder network structure comprises three parts: an encoder, a latent variable generation part, and a decoder, wherein the encoder part is used to calculate the low-dimensional mean μ and variance σ of each input electromyographic window from the signal window; the latent variable generation part performs mathematical operations on the added noise and the mean μ and variance σ to obtain a probability density function Z = {Z1, Z..., Zn}; the decoder part performs sample reconstruction through the probability density function and generates a new signal window through the decoder.

[0013] Preferably, the time series differentiable loss function is based on a soft-DTW algorithm to minimize the reconstruction error partial loss, measuring the difference between the input sample and the generated sample; the soft-DTW algorithm uses a continuous and smooth path to represent the alignment between the two time series and calculates the distance gradient.

[0014] The second aspect of the present invention is to provide a gesture recognition system based on multimodal optimal data selection and enhancement, comprising: a model building module for building and training a gesture recognition model; and a gesture recognition module for performing online gesture recognition based on the gesture recognition model; wherein the gesture recognition model is a long-short time series Transformer multimodal gesture recognition model; wherein the model building module comprises: a best matching signal screening submodule for building a multimodal signal database based on existing data, collecting multimodal signal samples of new user calibration gestures to form a new user calibration gesture dataset, and performing online gesture recognition on the new user calibration gesture dataset based on the multimodal signal samples in the multimodal signal database. The user calibration gesture multimodal signal samples are screened for the best matching signal to obtain the best matching signal to form a first training set; a signal enhancement submodule is used to perform signal enhancement based on the best matching signal to form a second training set; a model training submodule is used to train the long-short time series Transformer multimodal gesture recognition model based on the first training set and the second training set; the gesture recognition module includes: a gesture signal acquisition submodule, used to collect new user multimodal gesture signals online; a gesture category label generation submodule, used to input the new user multimodal gesture signal into the trained long-short time series Transformer multimodal gesture recognition model to obtain the gesture category after recognition.

[0015] A third aspect of the present invention provides an electronic device, comprising a processor and a memory, wherein the memory stores a plurality of instructions, and the processor is configured to read the instructions and execute the method described in the first aspect.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a plurality of instructions, and the plurality of instructions can be read by a processor to execute the method described in the first aspect.

[0017] The gesture recognition method, system, electronic device, and computer-readable storage medium based on multimodal optimal data selection and enhancement provided by the present invention have the following beneficial technical effects:

[0018] A gesture recognition method based on multimodal optimal data selection and enhancement uses a small sample size for each gesture. It introduces hand motion information to increase the universality of gesture features across users, thereby addressing the negative transfer problem caused by large differences between EMG signal domains and the insufficient sample diversity produced by signal enhancement methods. The method proposes optimal matching signal selection and signal enhancement based on signal similarity calculation. Based on a small sample size of multimodal physiological signals, for new users (patients), only a single acquisition of each gesture is required to achieve recognition accuracy comparable to traditional recognition models trained with sufficient surface EMG datasets, avoiding the burden of multiple acquisitions on patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flowchart of the gesture recognition method based on multimodal optimal data selection and enhancement according to the present invention.

[0020] Figure 2 This is a flow chart of the best matching signal screening method described in the present invention.

[0021] Figure 3 This is a flow chart of the signal similarity calculation method described in the present invention.

[0022] Figure 4 This is a flow chart of the similarity calculation method of the present invention; Figure 4 (a) shows the flow chart of the Euclidean distance similarity calculation method; Figure 4 (b) shows the flow chart of the DTW similarity calculation method.

[0023] Figure 5 This is a schematic diagram of the variational autoencoder network structure and loss function calculation principle described in the present invention.

[0024] FIG6( a ) is a schematic diagram of the LST-EMG-Net vector sorting under the ideal state described in the present invention; FIG6( b ) is a schematic diagram of the LST-EMG-Net vector sorting disorder described in the present invention.

[0025] Figure 7 This is a schematic diagram of the linear transformation module structure of the long-short time series Transformer multimodal model described in the present invention.

[0026] Figure 8 Schematic diagram of the structure of the gesture recognition system based on multimodal optimal data selection and enhancement according to the present invention.

[0027] Figure 9This is a structural diagram of the electronic device described in the present invention. DETAILED DESCRIPTION

[0028] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0029] Example 1

[0030] like Figure 1 As shown, this embodiment proposes a gesture recognition method based on multimodal optimal data selection and enhancement, including establishing and training a gesture recognition model and performing online gesture recognition based on the gesture recognition model; wherein the gesture recognition model is a long-short time series Transformer multimodal gesture recognition model.

[0031] The establishment and training of the gesture recognition model includes: Figure 2 As shown, S1, a multimodal signal database is established based on existing data, multimodal signal samples of new user calibration gestures are collected to form a new user calibration gesture data set, and the multimodal signal samples of new user calibration gestures in the new user calibration gesture data set are screened for the best matching signals based on the multimodal signal samples in the multimodal signal database to obtain the best matching signals to form a first training set.

[0032] In gesture recognition tasks, transfer learning can alleviate data scarcity and reduce the cost and time of data collection. By learning shared features and knowledge in the source domain and transferring them to the target domain, models can be trained more quickly and accurately. However, the similarity between the source and target domains is crucial for transfer learning to be effective. If the differences between the source and target domains are too great, negative transfer may occur, whereby knowledge from the source domain negatively impacts performance in the target domain.

[0033] In small-sample gesture recognition tasks, to avoid negative transfer, when performing transfer learning between different users, user data with a similar distribution to the target domain data should be selected as the source domain. However, since it is difficult to directly determine the similarity of electromyographic signals through direct observation, if two people have very different muscle activity patterns, their electromyographic signals will often differ significantly. In this case, using two people's data as the source and target domains, respectively, can easily lead to negative transfer.

[0034] Therefore, step S1 adopts the best matching signal screening method, which consists of two parts: similarity quantification and multimodal signal adaptive selection. Figure 2As shown in the figure. Signal similarity is used to calculate the signal similarity between the new user's calibration gesture and the gestures in the database. A multimodal signal adaptive selection strategy is then used to select the best matching signal, avoiding the problem of negative transfer. This method first constructs a multimodal signal database (user1, 2, ..., n) based on existing user data. Signal similarity is then used to calculate the time-frequency domain similarity between the calibration gesture's modal signals and the signals in the database using the signal similarity. Multimodal signal adaptive selection is then used to select the top N similar modal data, which are then combined. An early stopping model training strategy is used to reduce the time required to determine the optimal N value, and these are used as the best matching data to construct the first training set.

[0035] As a preferred embodiment, the S1 includes: S11, collecting multimodal signal samples of a single repeated gesture of a new user to form a new user calibration gesture data set; S12, accessing a database; the database stores a variety of gesture signals collected historically; see Figure 3 , S13, calculating the similarity between the calibration gesture signal and the multiple gesture signals, thereby determining the signal similarity between the new user calibration gesture and the multiple gesture signals stored in the database.

[0036] The signal similarity calculation method in this embodiment comprehensively describes the similarity between the data in the database and the new user's signal from the time domain and frequency domain perspectives according to the number of gesture repetitions. However, most of the methods in the prior art calculate the similarity between samples based on signal windows. The use of signal window samples can only capture local information of the data, while ignoring the global information in the entire action time series, which may lead to inaccurate similarity calculations. This embodiment uses a time warping algorithm (Dynamic Time Warping, DTW) in the time domain similarity calculation part to solve the problem that similarity calculation methods such as Euclidean distance and Pearson correlation coefficient cannot adapt to time offset: the frequency domain part uses the mean square error (MSE) to calculate the signal amplitude difference after the fast Fourier transform (Fast Fourier Transform, FFT).

[0037] The flow chart of the signal similarity calculation method in this embodiment is as follows: Figure 3As shown, the specific steps are described as follows: (1) The existing modal data are organized into a database according to the number of gesture repetitions, and the signal similarity is calculated according to the gesture type in turn. The first type of gesture signal similarity calculation is performed first. (2) Under the first type of gesture, the single modal data of the first repeated gesture is selected for similarity calculation. (3) The time domain similarity between the first repeated gesture and the new user calibration gesture is calculated based on the DTW algorithm. The calculation formula is as follows (time domain similarity calculation based on DTW): Time domain similarity calculation is adopted. The traditional time domain similarity calculation is the point-to-point Euclidean distance. The time domain similarity calculation method based on DTW in this embodiment can effectively avoid the problem that time delay and amplitude inconsistency between time series affect the description of signal time domain similarity. When using DTW to calculate signal time domain similarity, in actual applications, problems such as time delay and shape inconsistency between two time series are often encountered. For example, for the electromyographic signal sequence of the same gesture, since the limb movement speed of each person is different, even if the behavior pattern of the two movements is exactly the same, there is the same amplitude change trend in the signal. However, the waveform on the signal may be very different, resulting in the inability to correspond one to one at the time point. At this point, the Euclidean distance cannot determine the true similarity of two time series, such as Figure 4 (a) shows that in order to calculate the distance of the path, we need to first calculate the distance between each data point. Finally, the path with the smallest distance among all possible DTW paths is used as the similarity measure between the two time series, as shown in Figure 4 (b) In the form of time series, the points are no longer calculated one-to-one, but "one-to-many" and "many-to-one" correspondences are performed, eliminating the influence of signal time delay on time domain similarity.

[0038] The steps for DTW to calculate the time domain similarity of EMG signals are as follows:

[0039] ① The multi-channel electromyographic signal sequences of two complete gestures are recorded as X = {A1, A2, ..., A H}, Y={B1,B2,...,B H}, where H is the number of channels.

[0040] ② Select the first channel signal subsequence A1 of each of the two sequences X and Y = {a1, a2, ..., a n}, B1={b1, b2, ..., b m}, where n and m are the signal lengths of the two subsequences respectively.

[0041] ③ Calculate the distance matrix. Use the defined distance metric to calculate the distance P between all points in the two time series = {p1, ..., p s ,...,p k}, where p s=(i s ,j s ) represents the distance matrix.

[0042] ④Calculate the cumulative distance matrix Each element of the cumulative distance matrix represents the distance to that location, starting from the lower left corner and ending at the upper right corner. Each element in the cumulative distance matrix is calculated according to the following rules: the cumulative distance of the initial position is the distance to the corresponding position in the distance matrix. The calculation formula (1) of the cumulative distance matrix is as follows:

[0043]

[0044] Where δ(p s ) means i s With j s Distance, w s >0, indicating weighting coefficient.

[0045] ⑤Calculate the minimum distance path At the lower right corner of the cumulative distance matrix, the minimum cumulative distance to that location is obtained. Starting from the upper right corner and tracing back to the lower left corner, the path representing the minimum distance path is obtained. Output The minimum distance path is the result of the DTW algorithm, as shown in formula (2).

[0046]

[0047] ⑥ Repeat steps ② to ⑤ until the minimum distance matrix D of each channel is calculated p , and add them together to get the minimum distance matrix sum The value of is taken as the time series similarity of X and Y.

[0048] (4) The frequency domain similarity between the first repeated gesture and the new user's calibration gesture is calculated based on the mean square error of FFT. The calculation formula is as follows (frequency domain similarity calculation based on FFT):

[0049] The Fast Fourier Transform (FFT) is a computational method used to convert signals from the time domain to the frequency domain. The FFT decomposes time-domain signals into sine and cosine waves of varying frequencies. Therefore, to measure the frequency-domain similarity between two sequence signals, the mean squared error (MSE) between the signal FFT amplitudes can be used to calculate the difference between each sampling point in the frequency domain. The larger the MSE, the smaller the frequency-domain similarity.

[0050] First, the two signal sequences x are subjected to fast Fourier transform to obtain the frequency domain sequence X after FFT, and the other set of signals y are subjected to fast Fourier transform FFT to obtain Y. The following formula (3) calculates the mean square error (MSE) of the two signals as the signal frequency domain similarity:

[0051]

[0052] Where N is the length of the signal, X n and Y n Respectively represent the amplitude values of the two signals at the nth frequency in FFT. If x and y are real signals, then the amplitude spectrum in DFT is symmetrical, so only the first half of the frequency, that is, n = 0 to If x and y are complex signals, all frequencies need to be calculated.

[0053] (5) Time-frequency domain similarity scaling. In the multimodal signal adaptive selection module, the time-domain similarity and frequency-domain similarity values of the two signals are uniformly scaled to the same scale. After scaling, the time-domain similarity of the i-th gesture is (TimeDomain, TD) and frequency domain similarity (Frequency Domain, FD) is shown in formula (4) and formula (5):

[0054]

[0055]

[0056] in, is the time domain similarity between each gesture data in the database and the calibration gesture, is the frequency domain similarity between each gesture data in the database and the calibration gesture, σ is the scaling factor, and L is the total number of repetitions of each gesture in the database. For example, if there are 8 user data in the database and each gesture has 6 gesture data collected, then L = 48 gestures can be screened.

[0057] (6) Establish a similarity graph. Repeat step (1) to obtain the time domain similarity set D between each gesture in the database and the calibration gesture. TD and frequency domain similarity set D FD , as shown in equations (6) and (7):

[0058]

[0059]

[0060] The time domain similarity of the signal is the horizontal axis, and the frequency domain similarity of the signal is the vertical axis. The above set is formed into the form of time-frequency domain pairs Form a similarity graph.

[0061] (7) Signal similarity calculation: The distance between each point in the graph and the calibration gesture is calculated as the signal similarity value (SSV). The similarity of the first gesture repetition in the dataset is shown in the following formula (8):

[0062]

[0063] The larger the SSV value, the smaller the similarity. Repeat step (3) until the similarity values of all gesture data of this type in the database are obtained, and sort them in ascending order to represent the similarity.

[0064] After the calculation is completed, if there are other gesture repetitions in the database whose similarity has not been calculated, repeat steps (2) to (7) to perform the next round of signal similarity calculation; if there are no other gesture repetitions in the database, then sort the similarity of all signals in the database.

[0065] (8) Database signal similarity ranking. After calculating the similarity of the signals in the time domain and frequency domain, a unique similarity value is obtained as a measurement standard to comprehensively consider the signal similarity of each modal signal in the database in the time domain and frequency domain, and to achieve the ranking of the similarity between the signals in the database and the calibration gesture.

[0066] S14: Determine the best matching signal as a first training set based on the multimodal signal adaptive selection strategy and the signal similarity, where the first training set is a partial training set for training a long-short time series Transformer multimodal gesture recognition model.

[0067] In step S13, a similarity ranking of each modal signal in the database is constructed. However, due to the influence of the data quality in the database, the best signal selection cannot be performed based solely on the average or median similarity as a threshold. When the behavior pattern of the patient in the database is similar to that of the new user, there may be more similar data in the database. On the contrary, there is less similar data. Therefore, in order to adaptively realize the selection of the top N similar data, reduce the domain difference between the training set and the new user, solve the problem of negative transfer and improve the recognition accuracy. The present invention proposes an adaptive selection method for the best signal, which can adaptively and automatically select the signal with high similarity in the database as the best matching signal to determine the optimal N value.

[0068] The specific method is to calculate the signal similarity and get the similarity ranking of the data in the database. Then, put the signals of each modality into the training set in descending order according to their similarity to verify the accuracy change of the gesture recognition model. When the model accuracy is the best, the first N similarity signals are selected as high similarity signals, that is, the best matching signals. N increases in turn. The model accuracy is compared under the early stopping strategy. When Accuracy is the best,n >Accuracy n+1 Accuracy n >Accuracy n+2 The first n items are the best matching data, and L is the total number of gesture repetitions in the database. The design steps of the adaptive selection method of the above best signal are as follows:

[0069] (1) Select and combine the top N similar data. According to the signal similarity calculation module, obtain the similarity ranking of each modal signal, set the N value, and increase N to 1, 2, 3, ... L in sequence. Select data from each modality, cut them to the same sequence length, and combine them to create a training set.

[0070] (2) Model training: The dataset obtained in step (1) is trained on the network LST-EMG-Net to verify the recognition accuracy.

[0071] (3) Early stopping model training strategy. In order to reduce the model training time to determine the optimal N value, if N = n, the recognition accuracy Accuracy n >Accuracy n+1 Accuracy n >Accuracy n+2 When n is the optimal value, the first n similar data are the best matching data.

[0072] From the above steps, we can see that adaptive selection of multimodal signals can filter out data in the database that is helpful for recognition accuracy and put it into the training set as the best matching data, thereby avoiding the negative transfer phenomenon caused by the large difference between the source domain and the target domain signals.

[0073] S2, performing signal enhancement based on the best matching signal, including: S21, generating multiple expanded new user signal samples and / or best matching signal samples using a variational autoencoder network structure; the new user signal samples and / or best matching signal samples can be expanded as needed, wherein each sample can be expanded with a single modality signal or a multimodal signal, wherein the preferred solution includes only enhancing the electromyographic signal. S22, calculating the difference between the generated new user signal samples and / or best matching signal samples and the original samples of the new user calibration gesture and / or best matching signal based on a time series differentiable loss function; S23, updating the parameters of the variational autoencoder network structure based on the difference and backpropagation, and then repeating steps S21-S22, thereby obtaining a second training set required for training the long-short time series Transformer multimodal gesture recognition model, wherein the second training set is another part of the training set used to train the long-short time series Transformer multimodal gesture recognition model.

[0074] In recent years, data augmentation methods have been proven to be an effective way to solve small sample recognition tasks. Therefore, a large number of researchers have studied how to use deep learning to automatically learn the features of small samples and generate new enhanced samples.

[0075] Currently, data augmentation networks can be divided into generative adversarial networks (GANs), which do not require a manually designed loss function, and encoder-decoder networks, which do. GANs do not require a designed loss function. Instead, they autonomously learn to find the optimal strategy to improve the target metric. This allows for a higher level of automation. For example, Anicet Zanini R proposed a method based on a deep convolutional generative adversarial network (DCGAN) to enhance Parkinson's disease (PD) electromyography (EMG) signal data. However, with small sample sizes, DCGAN training can suffer from pattern collapse. The generator can only generate a small number of samples that are excessive, repetitive, or non-representative of the true data distribution, but cannot generate more data patterns, resulting in insufficient signal diversity. Furthermore, GANs primarily focus on the time domain of the signal, making it difficult to recover features such as high-frequency components and harmonics, thus limiting the quality of signal enhancement.

[0076] Unlike GANs, the variational autoencoder (VAE) of the encoder-decoder network learns the distribution of latent variables such as the signal's mean and variance, and by adding random noise, ensures that each generated sample is both similar and different from the original. VAE training is performed by minimizing the sum of the reconstruction error and the KL divergence, where the KL divergence improves the diversity of generated samples. However, the reconstruction error minimization loss is often designed for image data, typically using the MSE loss to compare pixel-by-pixel differences. For signal enhancement, this method makes it difficult to effectively utilize the temporal nature of the signal. Therefore, strengthening the VAE network is essential to minimize the temporal reconstruction error of the signal.

[0077] Therefore, the present invention uses a variational autoencoder (VAE) to design its loss function based on the signal similarity calculation method, in order to better utilize the temporal nature of the signal and achieve better signal enhancement. Based on this, the present invention proposes a signal enhancement method based on signal similarity calculation. The following details the variational autoencoder network (VAE) and the loss function based on time series differentiability.

[0078] The autoencoder network VAE mainly consists of three parts: encoder, latent variable generation part, and decoder. Figure 5As shown in the figure, the encoder calculates the low-dimensional mean μ and variance σ of each input EMG window. The latent variable generator performs mathematical operations on the added noise and the mean μ and variance σ to obtain a probability density function Z = {Z1, Z..., Zn}. The decoder reconstructs samples using this probability density function and generates new EMG windows through the decoder.

[0079] In order to generate more diverse signal samples, the latent variable generation part of the variational autoencoder no longer simply reconstructs samples from the features generated by the encoder, but instead reconstructs samples from a probability density function distribution. This distribution reconstruction requires a mean μ, a variance σ, and a randomly sampled e from a normal distribution, as shown in Equation (9). This avoids the autoencoder simply reconstructing the original features, enhancing the model's generalization ability. In addition, the randomly sampled e can be controlled to generate the expected EMG window.

[0080] Z=μ+exp(σ)×e (9)

[0081] Where μ is the mean, σ is the variance, e is a random sample from a normal distribution, and Z = {Z1, Z..., Zn} is the probability density function.

[0082] At the same time, in order to enhance VAE's utilization of the temporal characteristics of the signal, in the loss function part, this paper adopts a DTW algorithm called Soft-DTW that can calculate the distance gradient, replacing the original VAE network's use of the mean square error (MSE) as the reconstruction error loss to measure the difference between the input sample and the generated sample. Soft-DTW uses a continuous and smooth path to represent the alignment between two time series and can calculate the distance gradient, thereby strengthening the learning of time domain similarity and optimizing the enhanced signal quality. The Soft-DTW formula is shown in (10):

[0083]

[0084] Among them, β>0 is a parameter that controls the strength of Soft-DTW. Represents the cost function value of the original DTW algorithm given the path Y. SoftPlus is a smooth ReLU function that ensures the function is differentiable.

[0085] The normal distribution loss uses KL divergence to measure the difference between the general normal distribution and the standard normal distribution. Given a latent variable space dimension of n, with known mean μ and variance σ, KL is defined as formula (11):

[0086]

[0087] In summary, the similarity quantification loss function LOSS is formula (12):

[0088] LOSS=α*SoftDTW+β*KL(12)

[0089] Where α and β are the coefficients of the SoftDTW loss term and the KL loss term, respectively.

[0090] S3: Training the long-short time series Transformer multimodal gesture recognition model based on the first training set and the second training set.

[0091] For the multimodal signal input task of the present invention, when the transformer linear transformation module is extended to multimodal signals, the vector structure after the linear transformation module is shown in Figure 6(a). However, when added to the multimodal gesture recognition task, the encoder can only distinguish the vector's origin, not the modality. To the encoder, it may see the vector order disorder shown in Figure 6(b), which cannot guarantee the continuity of the information between the modalities and is not conducive to information exchange between the modalities.

[0092] Therefore, this embodiment adds a modal marker vector based on the above vectors according to the different modalities between signal segments, as follows Figure 7 As shown in the modal marker vector, “1” indicates that the slice vector comes from the electromyographic signal, and “2” indicates that the slice vector comes from the motion signal.

[0093] This module embeds multiple modalities into different vector spaces of the same dimension, allowing their vector representations to correspond to different types, facilitating temporal learning for the encoder. The final model concatenates the modality-labeled embedding vector with the slice vector and position-labeled vector to produce a comprehensive representation. This comprehensive representation incorporates both positional information and the modality information of the slice, interacting with other modalities to identify gesture types.

[0094] The online gesture recognition based on the gesture recognition model includes: S1', online collection of multimodal gesture signals of the new user; in this embodiment, the multimodal gesture signals of the new user are collected online based on a collection device worn by the new user; S2', inputting the multimodal gesture signals of the new user into the trained long-short time series Transformer multimodal gesture recognition model to obtain a gesture category after recognition.

[0095] In summary, in this embodiment, the application process of the gesture recognition algorithm based on multimodal optimal data selection and enhancement is divided into a model training phase and an online gesture recognition phase: In the model training phase, a single new user gesture signal is first collected as a calibration multimodal signal. This new user calibration gesture multimodal signal and the multimodal signal database are input into the best matching signal screening module, which selects the best matching signal as part of the training set. Next, the new user calibration gesture multimodal signal is input into the signal enhancement method module based on signal similarity calculation. The variational autoencoder sample generation network is used to expand the new user calibration gesture multimodal signal samples as another part of the training set. The training set samples are input into the long-short time series Transformer multimodal recognition model for training. In the online gesture recognition phase, the user wears a collection device to collect online multimodal signals in real time. The signals are input into the trained Transformer multimodal recognition model to obtain gesture category labels.

[0096] This example evaluates the effectiveness of the proposed algorithm on self-collected data from stroke patients and Ninapro DB5 healthy subjects, and analyzes and summarizes the experimental results.

[0097] (1) Multimodal Datasets and Computer Development Environment

[0098] The proposed gesture recognition method based on multimodal optimal data selection and enhancement is evaluated using the Ninapro DB5 public dataset.

[0099] 1. Ninapro DB5 multimodal dataset

[0100] The Ninapro DB5 Exercise C dataset features seven hand gestures that fully stimulate muscle activity and facilitate muscle recovery training. The DB5 dataset was collected from 10 healthy subjects using two Myo (Thalmic Labs) wristbands and a data glove. Each Myo wristband has eight single-channel electrodes, collecting a total of 16 channels of myoelectric signals and 22 channels of palm and finger joint motion information. Each electrode has a sampling rate of 200Hz. Each gesture in the Ninapro DB5 Exercise C is repeated six times, with 5 seconds of activity signal collected each time, and a 3-second interval between each acquisition.

[0101] 2. Computer development environment

[0102] Model training and testing were performed using a deep learning framework on a computer platform. The computer hardware configuration used was an Intel Core i7-10700K CPU processor (64GB of RAM), a GeForce GTX 3090 GPU (24GB of video memory), and an Ubuntu 18.04.5LTS operating system. The network model was constructed, trained, and validated using the Python 3.6.5 programming language and the PyTorch 1.8.0 deep learning framework. The computer development environment for the gesture recognition algorithm based on multimodal optimal data selection and enhancement is shown in Table 1.

[0103] Table 1

[0104] Hardware environment Software Environment CPU: Intel(R)Core(TM)i7-10700K CPU 3.8GHz Programming language: Python 3.6.5 Memory: 64.00GB Deep learning framework: Pytorch 1.8.0 System type: Ubuntu 18.04.5LTS Development tool: JetBrains PyCharm Graphics card model: NVIDIA GeForce GTX 3090

[0105] (2) Evaluation indicators

[0106] In order to verify the computer effectiveness of the gesture recognition algorithm based on multimodal optimal data selection and enhancement, when verifying the algorithm, each patient in the dataset was taken as a new user, and the data of other users in the dataset was used as the database. This process was repeated to obtain the recognition accuracy of each patient as a new user, and the average recognition accuracy of all patients in the dataset was used to verify the performance of the algorithm.

[0107] (3) Verification and analysis

[0108] Ablation experiments are performed on the Ninapro DB5 multimodal public dataset to evaluate the effectiveness of the algorithm.

[0109] In the ablation experiment section, to verify whether the gesture recognition algorithm based on multimodal optimal data selection and enhancement can achieve effective recognition of new users while reducing data collection, this example verifies the recognition accuracy of new users in two ways: not using new user data in the training set (experiments 1-4) and using a small sample of new user calibration gesture (CG) data (experiments 5-8). In both methods, the optimal matching signal screening method (OMSS), the long-short time series Transformer multimodal recognition model (Multimodal-LSTEMGNet, MM-LSTEMGNet), and the signal enhancement method based on the variational autoencoder (VAE) are gradually added to prove their effectiveness on the Ninapro DB5 multimodal dataset. The ablation experiment of the gesture recognition method based on multimodal optimal data selection and enhancement is shown in Table 2.

[0110] Table 2

[0111]

[0112]

[0113] As shown in Table 2, the method of the present invention achieved accuracies of 93.76% and 98.56% on the Ninapro DB5 public dataset in Experiment 4, which did not use new user data, and Experiment 8, which used a small sample of new user calibration gesture (CG) data, respectively. This is a significant improvement over traditional recognition methods.

[0114] Compared to Experiment 1, which used the entire database as the training set, Experiment 2 employed the Optimal Matching Signal Screening (OMSS) method to filter out data similar to the new user from the database, effectively avoiding negative transfer and improving accuracy by 18.77% on the dataset. VAE data augmentation effectively leverages the temporal nature of signals for enhancement.

[0115] The accuracy of the model trained using only the best matching data reached 93.76% on the dataset, which can meet basic rehabilitation needs; when only using the calibration gesture collected once, it can achieve an accuracy of 98.56% on the above dataset, which greatly reduces the burden of data collection and achieves the same effect as the accuracy of the personal data training model, making smart rehabilitation equipment more patient-friendly and conducive to the implementation of rehabilitation assistive devices.

[0116] Example 2

[0117] See also Figure 8 This embodiment provides a gesture recognition system based on multimodal optimal data selection and enhancement, including: a model building module 1, used to establish and train a gesture recognition model; and a gesture recognition module 2, used to perform online gesture recognition based on the gesture recognition model: wherein the gesture recognition model is a long-short time series Transformer multimodal gesture recognition model.

[0118] Among them, the model establishment module 1 includes: a best matching signal screening submodule 101, which is used to establish a multimodal signal database based on existing data, collect multimodal signal samples of new user calibration gestures to form a new user calibration gesture data set, and perform best matching signal screening on the new user calibration gesture multimodal signal samples in the new user calibration gesture data set based on the multimodal signal samples in the multimodal signal database to obtain the best matching signal to form a first training set; a signal enhancement submodule 102, which is used to perform signal enhancement based on the best matching signal to form a second training set; a model training submodule 103, which is used to train the long-short time series Transformer multimodal gesture recognition model based on the first training set and the second training set.

[0119] The gesture recognition module 2 includes: a gesture signal collection submodule 201 for online collection of multimodal gesture signals of new users; a gesture category label generation submodule 202 for inputting the multimodal gesture signals of new users into the trained long-short time series Transformer multimodal gesture recognition model to obtain gesture categories after recognition.

[0120] The present invention also provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method described in the first embodiment.

[0121] like Figure 9 As shown, the present invention also provides an electronic device, including a processor 301 and a memory 302 connected to the processor 301, wherein the memory 302 stores multiple instructions, which can be loaded and executed by the processor to enable the processor to execute the method described in Example 1.

[0122] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.

Claims

1. A gesture recognition method based on multimodal optimal data selection and enhancement, characterized in that: The method includes establishing and training a gesture recognition model and performing online gesture recognition based on the gesture recognition model; The gesture recognition model is a long-short time series Transformer multimodal gesture recognition model; The establishing and training of the gesture recognition model includes: S1, establishing a multimodal signal database based on existing data, collecting multimodal signal samples of a single calibration gesture of a new user to form a new user calibration gesture dataset, and performing best matching signal screening on the multimodal signal samples of the new user calibration gesture in the new user calibration gesture dataset based on the multimodal signal samples in the multimodal signal database to obtain the best matching signal to form a first training set; S2, performing signal enhancement based on the best matching signal to form a second training set; S3, training the long-short time series Transformer multimodal gesture recognition model based on the first training set and the second training set; The performing online gesture recognition based on the gesture recognition model includes: S1', online collection of multimodal gesture signals of new users; S2', inputting the multimodal gesture signal of the new user into the trained long-short time series Transformer multimodal gesture recognition model to obtain the gesture category after recognition; Said S1 comprises: S11, collecting multimodal signal samples of a new user's single repeated gesture to form a new user calibration gesture dataset; S12, accessing a database; the database stores a variety of gesture signals collected historically; S13, calculating similarities between the multiple gesture signals and the multiple gesture signals in the new user calibration gesture dataset, thereby determining signal similarities between the new user calibration gesture and the multiple gesture signals stored in the database; S14, determining the best matching signal as a first training set based on the multimodal signal adaptive selection strategy and the signal similarity, wherein the first training set is a partial training set for training the long-short time series Transformer multimodal gesture recognition model; The similarity calculation method described in S13 comprehensively describes the similarity between the data in the database and the new user signal from the time domain and frequency domain perspectives according to the number of gesture repetitions; the time domain similarity calculation part adopts the time warping algorithm DTW, and the frequency domain part uses the mean square error MSE to calculate the signal amplitude difference after fast Fourier transform FFT.

2. The gesture recognition method based on multimodal optimal data selection and enhancement according to claim 1, characterized in that: The similarity calculation method in S13 comprehensively describes the similarity between the data in the database and the new user signal from the perspectives of time domain and frequency domain according to the number of gesture repetitions, including the following steps: (1) The existing modal data are organized into a database according to the number of gesture repetitions, and the signal similarity is calculated according to the gesture type. The first type of gesture signal similarity is calculated first; (2) Under the first type of gesture, the single modal data of the first repeated gesture is selected for similarity calculation; (3) Calculate the time domain similarity between the first repeated gesture and the calibration gesture of the new user based on the DTW algorithm; (4) Calculate the frequency domain similarity between the first repeated gesture and the new user's calibration gesture based on the mean square error of FFT; (5) Time-frequency domain similarity scaling, including: scaling the time-domain similarity and frequency-domain similarity values of two signals to the same scale in the multimodal signal adaptive selection module; (6) Build a similarity graph: Repeat step (1) to obtain the time domain similarity set of each gesture in the database and the calibration gesture and frequency domain similarity set ; With the time domain similarity of the signal as the horizontal axis and the frequency domain similarity of the signal as the vertical axis, the time domain similarity set and frequency domain similarity set Forming a time-frequency domain pair ,… , , forming a similarity graph; (7) Signal similarity calculation: calculate the distance between each point in the graph and the calibration gesture as the signal similarity value; repeat step (3) until the similarity value of all data of this type of gesture in the database is obtained, and sort them in ascending order to represent the similarity size; after the calculation is completed, if there are other gestures in the database whose similarity has not been calculated, repeat steps (2) to (7) to perform the next round of signal similarity calculation; if there are no similarities, then sort all signals in the database by similarity; (8) Database signal similarity ranking: After calculating the similarity of the signals in the time domain and frequency domain, a unique similarity value is obtained as a measurement standard to comprehensively consider the signal similarity of the time domain and frequency domain of each modal signal in the database, and the similarity between the signals in the database and the calibration gesture is ranked.

3. The gesture recognition method based on multimodal optimal data selection and enhancement according to claim 1, characterized in that: The multimodal signal adaptive selection strategy includes the following steps: (1) Selection and combination of the first N similar data: According to the signal similarity calculation module, the similarity ranking of each modal signal is obtained, and the N value is set, increasing N from 1, 2, 3, ... to L in sequence. Data is selected on each modality, cut to the same sequence length, and combined to make a training set, where L represents the total number of modal signals; (2) Model training: The dataset obtained in step (1) is trained under the network LST-EMG-Net to verify the recognition accuracy; (3) Implement early stopping model training strategy: If N=n, recognition accuracy > and > When n is the optimal value, the first n similar data are the best matching data.

4. The gesture recognition method based on multimodal optimal data selection and enhancement according to claim 3, characterized in that: The S2 includes: S21, generating a plurality of expanded new user signal samples and / or best matching signal samples using a variational autoencoder network structure; S22, calculating the difference between the new user signal sample and / or the best matching signal sample generated based on a time series differentiable loss function and the original sample of the new user calibration gesture and / or the best matching signal; S23, after updating the parameters of the variational autoencoder network structure based on the difference and backpropagation, repeat steps S21-S22 to obtain a second training set required for training the long-short time series Transformer multimodal gesture recognition model, where the second training set is another part of the training set for training the long-short time series Transformer multimodal gesture recognition model.

5. The gesture recognition method based on multimodal optimal data selection and enhancement according to claim 4, characterized in that: The variational autoencoder network structure consists of three parts: encoder, latent variable generator and decoder, where the encoder part is used to calculate the low-dimensional mean μ and variance of each input electromyographic window from the signal window. ; Latent variables are generated by adding noise with mean μ and variance Perform mathematical operations to obtain a probability density function Z={Z1, Z ..., Zn}, use the probability density function to reconstruct samples, and generate a new signal window through the decoder.

6. The gesture recognition method based on multimodal optimal data selection and enhancement according to claim 5, characterized in that: The time series differentiable loss function minimizes the reconstruction error loss based on the soft-DTW algorithm, measuring the difference between the input sample and the generated sample; the soft-DTW algorithm uses a continuous and smooth path to represent the alignment between the two time series and calculates the distance gradient.

7. A gesture recognition system based on multimodal optimal data selection and enhancement, used to implement the method according to any one of claims 1 to 6, characterized in that: include: Model building module (1), used to build and train the gesture recognition model; as well as A gesture recognition module (2), configured to perform online gesture recognition based on the gesture recognition model; The gesture recognition model is a long-short time series Transformer multimodal gesture recognition model; Wherein, the model building module (1) includes: A best matching signal screening submodule (101) is used to establish a multimodal signal database based on existing data, collect multimodal signal samples of a new user calibration gesture to form a new user calibration gesture data set, and perform best matching signal screening on the multimodal signal samples of the new user calibration gesture in the new user calibration gesture data set based on the multimodal signal samples in the multimodal signal database to obtain the best matching signal to form a first training set; A signal enhancement submodule (102), configured to perform signal enhancement based on the best matching signal to form a second training set; A model training submodule (103), configured to train the long-short temporal sequence Transformer multimodal gesture recognition model based on the first training set and the second training set; The gesture recognition module (2) comprises: A gesture signal collection submodule (201) is used to collect multimodal gesture signals of new users online; The gesture category label generation submodule (202) is used to input the multimodal gesture signal of the new user into the trained long-short time series Transformer multimodal gesture recognition model to obtain the gesture category after recognition.

8. An electronic device comprising a processor and a memory, wherein the memory stores a plurality of instructions, characterized in that: The processor is configured to read the instruction and execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The plurality of instructions may be read by a processor and used to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal progressive hierarchical fusion method for natural gesture recognition

    CN116028889A

  • Gesture recognition method and system based on multi-modal feature fusion and small sample learning

    CN116343261A