Gesture motion tracking system and method based on surface muscle electricity

By using the technology of generative enhancement modules and cross-crowd adjustment modules on sparse sEMG devices, high cost and complex calibration problems are solved, and a gesture motion tracking system with high accuracy and compatibility is achieved.

CN120045071APending Publication Date: 2025-05-27BEIJING INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191847.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing gesture motion tracking technology based on surface muscle electricity faces the problems of high cost, difficulty in practical application, and complex calibration across populations and across days.

Method used

The generative enhancement module and the cross-crowd adjustment module are adopted. The generative enhancement module realizes signal space enhancement on sparse sEMG devices. The cross-crowd adjustment module completes muscle function partitioning through two joint gestures, simplifying the calibration process.

Benefits of technology

Reduces deployment costs, improves the accuracy of finger tracking, simplifies the calibration process, and improves system compatibility and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045071A_ABST
    Figure CN120045071A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of signal representation and classification, and discloses a gesture motion tracking system and method based on surface muscle electricity. The system comprises a data acquisition module used for acquiring an sEMG vector; the pre-filling module is used for constructing an sEMG matrix I containing an sEMG vector; the generative enhancement module is used for receiving the sEMG matrix I and then enhancing the sEMG matrix I into an independent gesture representation graph; the cross-crowd adjustment module is used for decomposing the myoelectricity panoramic representation image into a corresponding number of function partition mask images; the generative enhancement module and the cross-crowd adjustment module both need to be pre-trained; the generative enhancement module and the cross-crowd adjustment module both need to be pre-trained, the judgment module is used for judging whether the cross-crowd adjustment module needs to be calibrated or not, and finger-muscle relationships are determined in different modes; the method is used for determining the biokinematics relation between the muscle and the finger of the subject, the deployment cost is reduced, richer muscle information is obtained, and the accuracy of gesture motion tracking is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of signal characterization and classification, and in particular, to a gesture motion tracking system and method based on surface electromyography. Background Art

[0002] Motion tracking based on surface electromyography (sEMG) includes tracking of several parts, such as the hand, leg, foot, abdomen, etc. Taking the realization of hand motion tracking by an electromyography bracelet as an example, the research in this area has always faced two challenges: one is the balance between hardware cost and performance. Although a sparse sEMG bracelet is easy to deploy and has a low cost, its coverage area is small, and it cannot obtain more muscle information in space, resulting in inaccurate motion tracking. While the high-density electromyography (HD-sEMG) solution can cover a larger area of the arm, it is not easy to be deployed daily and has a high cost. The other is the calibration problem across different people and days. Traditional calibration methods often require complex processes, which is not convenient for daily use.

[0003] The key to continuous finger motion tracking is to establish an arm model, that is, a biomechanical model in which the bending and extension of each finger are affected by muscle traction in different regions. However, differences in arm circumference, skin impedance, muscle strength, and muscle distribution among different individuals easily lead to the failure of the "finger-muscle model". At the same time, some muscles, such as the extensor digitorum, can trigger different muscle fibers of its own to control 4 fingers, that is, one muscle controls 4 joints. Therefore, when the muscle functional regions of different subjects are distinguished, an adaptive finger-muscle model can be made for each individual. In order to more accurately distinguish surface muscles and improve the performance of finger motion tracking, the prior art has conducted research in three aspects: electromyogram partitioning, electromyogram signal quality and channel utilization, and electromyogram correction. However, most of these studies cannot identify the muscle distribution of a single individual, or although they can identify the muscle distribution of a single individual, individual differences and electrode offsets will still lead to a decline in model performance.

[0004] Therefore, there is an urgent need to propose a gesture motion tracking method that can solve problems such as high cost, difficulty in practical application, and complex calibration across different people and days. Summary of the Invention

[0005] The object of the present invention is to provide a gesture motion tracking system and method based on surface electromyogram, which adopts a generative enhancement module and a cross-population adjustment module. Among them, the generative enhancement module can realize signal space enhancement on a sparse sEMG device with low cost and easy to wear, which provides richer information for the functional partition of muscles. With this information, in the calibration process, the cross-population adjustment module only needs two combined gestures to divide the relevant muscle areas that independently drive finger movement, and can well solve the problems existing in the prior art, such as high cost, difficult to be actually applied, and complex cross-population and cross-day calibration.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides a gesture motion tracking system based on surface electromyogram, which is used to determine the kinematic relationship between the muscles and fingers of the current subject. The gesture motion tracking system includes:

[0008] A data acquisition module, which is used to output the sEMG vector of the current subject;

[0009] A pre-filling module, which is used to construct an sEMG matrix I containing the sEMG vector;

[0010] A generative enhancement module, which is used to enhance the sEMG matrix I into an independent gesture representation map after receiving it;

[0011] A cross-population adjustment module, which is used to decompose the electromyogram panoramic representation map into a corresponding number of functional partition mask maps;

[0012] Both the generative enhancement module and the cross-population adjustment module need to be pre-trained;

[0013] A judgment module, which is used to judge whether the cross-population adjustment module needs to be calibrated; if so, after the measured representation maps of K independent gestures of the current subject are constructed and filled through sequence vectors, and then an electromyogram panoramic representation map is generated through the cross-population adjustment module, and then N functional partition mask maps are decomposed, and the relevant channels in the independent gesture representation map are filtered by the N functional partition mask maps of the current subject to determine the kinematic relationship between the muscles and fingers, that is, the finger-muscle relationship; otherwise, N functional partition mask maps are generated by using the pre-trained electromyogram panoramic maps of N independent gestures to filter the relevant channels in the measured representation map of the independent gesture to determine the kinematic relationship between the muscles and fingers, that is, the finger-muscle relationship; K is greater than or equal to 2 and less than N, and N is greater than or equal to 6.

[0014] As a possible implementation, the generative enhancement module is pre-trained through the following steps:

[0015] S10. Vector construction: Extract one row of the EMG panoramic representation map of an independent gesture from a standard or measured dataset to generate a one-dimensional sEMG vector corresponding to the independent gesture;

[0016] The EMG panoramic representation map of the independent gesture is the ground truth;

[0017] S11. Prefilling: Prefill the one-dimensional sEMG vector into an sEMG matrix I;

[0018] The sEMG matrix I is a two-dimensional matrix;

[0019] S12. Encoder-decoder training: Train through a GPAT encoder and a GPAT decoder to obtain the EMG panoramic representation map corresponding to the independent gesture after training;

[0020] S13. Compare the EMG panoramic representation map corresponding to the independent gesture after training with the ground truth, and adjust the parameters of the GPAT encoder and the GPAT decoder according to the comparison result;

[0021] Continuously repeat S10 to S13, train and adjust the parameters of the GPAT encoder and the GPAT decoder until the absolute value of the correlation coefficient or the similarity between the EMG panoramic representation map after training and the ground truth is greater than or equal to the GPAT threshold, and stop the pre-training.

[0022] Specifically, when the GPAT threshold is the correlation coefficient, if the absolute value of the correlation coefficient is greater than or equal to 0.8, it is considered that the EMG panoramic representation map after training is very close to the ground truth; or when the GPAT threshold is the similarity, if the similarity is greater than or equal to 0.7 or 0.8, it is considered that the EMG panoramic representation map after training is very close to the ground truth; stop the pre-training.

[0023] As a possible implementation, pre-train the cross-population adjustment module through the following steps:

[0024] S20. Convert each of the N EMG panoramic representation maps into a one-dimensional sEMG vector and then splice them into a two-dimensional matrix to form the ground truth;

[0025] S21. Sequence vector construction: Select K from the N EMG panoramic representation maps, and then generate K one-dimensional sEMG vectors;

[0026] The N independent gesture reference representation maps in S20 and S21 are from a standard or measured dataset;

[0027] S22. Sequence filling: Fill the K one-dimensional sEMG vectors to generate a two-dimensional matrix with the same dimension as the ground truth;

[0028] S23. Encoding and decoding training: The two-dimensional matrix filled in S22 is trained through an AT encoder and an AT decoder to obtain N reconstructed independent gesture EMG panoramic representation maps;

[0029] S24. Compare the N reconstructed independent gesture EMG panoramic representation maps with the ground truth, and adjust the parameters of the AT encoder and the AT decoder according to the comparison results;

[0030] Continuously repeat S20 to S24 to train and adjust the parameters of the AT encoder and the AT decoder until the absolute value of the correlation coefficient or the similarity between the trained reconstructed EMG panoramic representation map and the ground truth is greater than or equal to the AT threshold, and stop the pre-training;

[0031] Specifically, when the AT threshold is the correlation coefficient, if the absolute value of the correlation coefficient is greater than or equal to 0.8, it is considered that the trained EMG panoramic representation map is very close to the ground truth; or when the AT threshold is the similarity, if the similarity is greater than or equal to 0.7 or 0.8, it is considered that the trained EMG panoramic representation map is very close to the ground truth; stop the pre-training.

[0032] S25. Decompose N functional partition mask maps from the N pre-trained independent gesture EMG panoramic representation maps.

[0033] As a possible implementation, the cross-population adjustment module is calibrated by the following method:

[0034] S30. The current subject performs K standard actions to obtain K independent gesture measured representation maps;

[0035] S31. Generate K sEMG vectors from the K independent gesture measured representation maps;

[0036] S32. Sequence filling: Construct a two-dimensional matrix Ⅲ containing K sEMG vectors through sequence filling;

[0037] S33. Encoding and decoding training: Reconstruct the EMG panoramic representation maps of N independent gestures through an AT encoder and an AT decoder;

[0038] S34. Decompose N functional partition mask maps of the current subject from the EMG panoramic representation maps of N independent gestures.

[0039] The generative enhancement module is pre-trained based on the Transformer architecture, including a GPAT encoder and a GPAT decoder, and is used to implement sequence-to-sequence signal conversion; both the GPAT encoder and the GPAT decoder internally include multiple GPAT encoding layers and GPAT decoding layers respectively to increase the scale of the weight parameters of the pre-trained Transformer architecture.

[0040] Each GPAT encoding layer includes at least a multi-head attention layer and a position-wise feed-forward network layer;

[0041] Among them, the multi-head attention layer has six heads, and each head includes Q, K, and V matrices. The self-attention mechanism is applied to perform the inner product of Q and the transpose of K to obtain a correlation matrix. The value A obtained by cross-multiplying the correlation matrix and the V matrix represents the influence degree of other vectors on the current vector; the value A is superimposed on the input value and then imported into the position-wise feed-forward network layer to obtain the output value O;

[0042] The multi-head attention layer included in each GPAT encoding layer is masked with a lower triangular matrix mask to ensure sequential generation; each GPAT encoding layer also includes a multi-head attention module, which uses the output of the multi-head attention module of the previous layer as the index value, the output of the GPAT encoder as the key and value, and the feed-forward network layer of the GPAT encoding layer as the output layer.

[0043] The cross-population adjustment module is pre-trained based on the Transformer architecture and includes an AT encoder and an AT decoder; the AT decoder does not include a mask layer, that is, during operation, the output of the AT encoder is not used as the next index value to complete iterative generation; the loss value between the output of the AT decoder and the true value will be backpropagated to the AT encoder of the cross-population adjustment module to update the parameters.

[0044] As another aspect of the present invention, a gesture motion tracking method based on surface electromyography relies on a generative enhancement module and a cross-population adjustment module, and includes the following steps:

[0045] Collect the independent gestures of the current subject and obtain a one-dimensional sEMG vector;

[0046] Fill and construct a two-dimensional matrix Ⅰ containing the one-dimensional sEMG vector;

[0047] After receiving the two-dimensional matrix Ⅰ, enhance it to an actual measurement representation map of the independent gesture;

[0048] Pre-train the generative enhancement module to obtain an independent gesture representation map corresponding to the independent gesture;

[0049] Pre-train the cross-population adjustment module to obtain N functional partition mask maps decomposed from the myoelectric panoramic representation map based on N independent gestures;

[0050] Determine whether the cross-population adjustment module needs to be calibrated; if so, use the K independent gesture measured representation maps of the current subject to construct and fill in the sequence vector, then generate the electromyographic panoramic representation map through the cross-population adjustment module, and then decompose N functional partition mask maps, and use the N functional partition mask maps of the current subject to filter the relevant channels in the independent gesture representation map to determine the biokinematic relationship between muscles and fingers, that is, the finger-muscle relationship; otherwise, use the pre-trained N independent gesture electromyographic panoramic maps to generate masks to obtain N functional partition mask maps to filter the relevant channels in the independent gesture representation map to determine the biokinematic relationship between muscles and fingers, that is, the finger-muscle relationship, K is greater than or equal to 2 and less than N, and N is greater than or equal to 6.

[0051] Fill and construct a two-dimensional matrix I containing a one-dimensional sEMG vector, and after receiving the two-dimensional matrix I, enhance it into an independent gesture measured representation diagram, specifically: generate each row of sEMG data on the two-dimensional matrix I row by row through an iterative method to fully consider the one-dimensional sEMG vector.

[0052] Beneficial effects:

[0053] The present invention proposes a gesture motion tracking system and method based on surface muscle electricity, which has the following beneficial effects compared with the prior art:

[0054] 1. The surface electromyography-based gesture motion tracking system provided by the present invention includes a generative enhancement module that can spatially expand sparse sEMG signals into HD-sEMG on a low-cost and easy-to-wear sparse sEMG device, thereby obtaining richer muscle information, reducing costs while improving the accuracy of finger tracking.

[0055] 2. The surface electromyography-based gesture motion tracking system provided by the present invention includes a cross-population adjustment module which does not need to use a mask to achieve autoregressive serialized output, but outputs the result in a one-time manner in a non-autoregressive manner, which can significantly improve the computing efficiency.

[0056] 3. The surface electromyography-based gesture motion tracking system provided by the present invention has a cross-population adjustment module that can complete the partitioning of 8 muscle functions through only two calibration actions. Each time an individual user wears it, the user does not need to adjust all the parameters of the gesture motion tracking system again, but only needs to train the cross-population adjustment module again to complete the calibration, which saves time and improves compatibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0058] Figure 1 This is a schematic structural diagram of the gesture motion tracking system based on surface electromyography in the embodiments of the present invention;

[0059] Figure 2 This is a schematic diagram of the hardware and data reconstruction of the Hyser dataset in the embodiments of the present invention;

[0060] Figure 3 This is an information enhancement comparison diagram of various movements of Subject 1 among 15 subjects in the embodiments of the present invention;

[0061] Figure 4 This is a bar chart of the similarity SSIM of 15 subjects in the embodiments of the present invention;

[0062] Figure 5 This is the effect of the sEMG diagrams of multiple actions of Subject 17 among another 5 subjects before and after GPAT processing and the bar chart of SSIM of another 5 subjects in the embodiments of the present invention;

[0063] Figure 6 This is the muscle segmentation effect diagram after Subject 1 transferred the electrodes in the embodiments of the present invention.

[0064] Reference numerals:

[0065] 1 - Data acquisition module, 2 - Prefilling module, 3 - Generative enhancement module, 30 - GPAT encoder, 30 - GPAT decoder, 4 - Cross - population adjustment module, 40 - AT encoder, 41 - AT decoder, 5 - Judgment module. Detailed implementation manners

[0066] In order to clearly describe the technical solutions in the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and no limitation is imposed on their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.

[0067] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0068] In the present invention, "at least one" means one or more, and "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The following at least one (item) or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b, and c, where a, b, and c may be single or multiple.

[0069] The embodiments of the present invention provide a gesture motion tracking system and method based on surface electromyography. The goal of the gesture motion tracking system is to establish the kinematic relationship between muscles and fingers. The system includes: a data acquisition module for acquiring sEMG vectors; a pre-filling module for constructing an sEMG matrix I containing sEMG vectors; a generative enhancement module for enhancing the sEMG matrix I into an independent gesture representation map after receiving it; a cross-population adjustment module for decomposing the electromyography panoramic representation map into a corresponding number of functional partition mask maps; both the generative enhancement module and the cross-population adjustment module require pre-training; both the generative enhancement module and the cross-population adjustment module require pre-training, and a judgment module is used to judge whether the cross-population adjustment module needs to be calibrated and determine the finger-muscle relationship in different ways; the present invention is used to determine the kinematic relationship between the muscles and fingers of a subject, reduce the deployment cost, obtain richer muscle information, and improve the accuracy of gesture motion tracking.

[0070] In order to obtain richer muscle information, the prior art usually adopts HD-sEMG or iEMG solutions, but these solutions are not easy to deploy in practical applications or may cause physiological pain. The gesture motion tracking system proposed in this embodiment adopts a generative enhancement module (Generative Pre-trained Augmented Transformer, GPAT) and a cross-population adjustment module (Adapter, AT). The generative enhancement module can perform signal space enhancement on a sparse sEMG device with a lower cost and easy to wear, which provides richer information for the functional partition of muscles. With this information, in the calibration process, the cross-population adjustment module only needs two joint gestures to divide the relevant muscle areas that independently drive finger movement. It can well solve the problems existing in the prior art, such as high cost, difficulty in practical application, and complexity of cross-population and cross-day calibration.

[0071] In a first aspect, the present invention provides a gesture motion tracking system based on surface electromyogram, which is used to determine the kinematic relationship between the muscles and fingers of the current subject. Specifically, the muscles of the subject refer to the muscles of their wrist.

[0072] See Figure 1 , the gesture motion tracking system includes: a data acquisition module 1, a pre-filling module 2, a generative enhancement module 3, a cross-population adjustment module 4, and a judgment module 5. The functions of each module are as follows:

[0073] The data acquisition module 1 is used to output the sEMG vector of the current subject;

[0074] The pre-filling module 2 is used to construct an sEMG matrix I containing the sEMG vector;

[0075] As an example, assuming that the sEMG vector of the current subject output by the data acquisition module 1 is a one-dimensional sEMG vector of 1×16, then the sEMG matrix I constructed by the pre-filling module is a two-dimensional matrix of 16×16. Specifically, the pre-filling module 2 uses blank data for shape filling. The input sEMG data is 16 vectors, that is, the dimension of each vector is 16, and only the first vector contains sEMG data, and the other vectors are empty.

[0076] The generative enhancement module 3 is used to enhance the sEMG matrix I into an independent gesture representation map after receiving it;

[0077] As a possible implementation manner, the generative enhancement module 3 is pre-trained based on the Transformer architecture; in this embodiment, the generative enhancement module 3 uses extended sparse sEMG to obtain an sEMG panoramic map with richer spatial information, and can achieve signal spatial enhancement on a sparse sEMG device with lower cost and easy to wear.

[0078] See Figure 1 , the generative enhancement module 3 includes a GPAT encoder 30 and a GPAT decoder 31, which are used to implement sequence-to-sequence signal conversion; both the GPAT encoder 30 and the GPAT decoder 31 respectively include multiple GPAT encoding layers and GPAT decoding layers inside to increase the scale of the weight parameters of the pre-trained Transformer architecture.

[0079] As a possible implementation manner, each GPAT encoding layer includes at least a multi-head attention layer and a position-wise feed-forward network layer;

[0080] Among them, the multi-head attention layer is set with six heads, and each head includes Q, K, and V matrices. The self-attention mechanism is applied to perform the inner product of Q and the transpose of K to obtain the relevance matrix, and the value A obtained by cross-multiplying the relevance matrix and the V matrix represents the influence degree of other vectors on the current vector; the value A is superimposed on the input value and then imported into the position-wise feed-forward network layer to obtain the output value O.

[0081] As an example, the Q, K, and V matrices are generated by a linear network of the input vector via W q 、W k 、W v . Q is an array composed of the index values q n of all input vectors. K and V are arrays of the keys k n and values v n of all input vectors respectively. The self-attention mechanism obtains the relevance matrix by performing the inner product of Q and the transpose of K, which is used to establish the correlation between the Q, K, and V matrices. The standard transformer represents the influence degree of other vectors on the current vector. The value obtained by superimposing A on the input is imported into the position-wise feed-forward network and other modules to obtain the output value O = V × A. When generating the sEMG data of each row, the generative enhancement module generates row by row in an iterative manner, which can fully consider the information of the previous signals.

[0082] Each multi-head attention layer included in each GPAT encoding layer is masked with a lower triangular matrix mask to ensure sequential generation; each GPAT encoding layer also includes a multi-head attention module, which uses the output of the previous multi-head attention module as the index value, the output of the GPAT encoder as the key and value, and the feed-forward network layer included in the GPAT encoding layer as the output layer. The multi-head attention module is masked with a lower triangular matrix mask to ensure sequential generation and provide dimensionality-raising information for the cross-population adjustment module. During training, the data does not require manual annotation, and only the sEMG map and the sEMG panoramic map after panoramic map downsampling need to be input as the input and the ground truth. Then, the output value and the ground truth are fed back to the network through the evaluation of the loss function for parameter update. This semi-supervised training method saves labor.

[0083] The cross-population adjustment module 4 is used to decompose the myoelectric panoramic representation map into the corresponding number of functional partition mask maps;

[0084] As a possible implementation, the cross-population adjustment module 4 is pre-trained based on the Transformer architecture; the cross-population adjustment module 4 includes an AT encoder 40 and an AT decoder 41; the AT decoder 41 does not include a mask layer, that is, during operation, the output of the AT encoder 40 is not used as the next index value to complete iterative generation; the loss value between the output of the AT decoder 41 and the ground truth will be backpropagated to the cross-population adjustment module to achieve parameter update.

[0085] In this embodiment, the cross-crowd adjustment module is different from the generative enhancement module in that it needs to use a mask to achieve the serialized output of autoregression. Instead, it outputs the result in a one-time manner in a non-autoregressive manner, which can significantly improve the computing efficiency. In addition, in order to simplify the calibration steps and improve the adaptability across populations, this embodiment separates the sEMG active areas when the four fingers are stretched and bent independently through the combined action of stretching and bending the four fingers (index finger, middle finger, ring finger and little finger) at the same time. When the user wears it each time, the user does not need to adjust all the parameters of the gesture motion tracking system again, but only needs to train the cross-crowd adjustment module again to complete the calibration.

[0086] Both the generative enhancement module 3 and the cross-population adjustment module 4 require pre-training;

[0087] As a possible implementation, the generative enhancement module 3 is pre-trained by the following steps:

[0088] S10. Vector construction: extract a row of the panoramic EMG representation of an independent gesture from the standard or measured data set to generate a one-dimensional sEMG vector corresponding to the independent gesture; the panoramic EMG representation of the independent gesture is the groundtruth;

[0089] S11. Pre-filling: pre-filling the one-dimensional sEMG vector into the sEMG matrix I; the sEMG matrix I is a two-dimensional matrix;

[0090] S12. Encoder and decoder training: training is performed through the GPAT encoder and the GPAT decoder to obtain a myoelectric panoramic representation corresponding to the independent gesture after training;

[0091] S13. Compare the trained EMG panoramic representation corresponding to the independent gesture with the groundtruth, and adjust the GPAT encoder and GPAT decoder parameters according to the comparison results;

[0092] Repeat S10 to S13 continuously, train and adjust the parameters of the GPAT encoder and GPAT decoder, until the correlation coefficient or similarity between the trained EMG panoramic representation and the groundtruth is greater than the GPAT threshold, and stop pre-training;

[0093] In specific implementation, when the GPAT threshold is the correlation coefficient, if the absolute value of the correlation coefficient is greater than or equal to 0.8, it is considered that the trained EMG panoramic representation is very close to the groundtruth; or when the GPAT threshold is the similarity, if the similarity is greater than or equal to 0.7 or 0.8, it is considered that the trained EMG panoramic representation is very close to the groundtruth; stop pre-training.

[0094] As a possible implementation, the cross-population adjustment module 4 is pre-trained through the following steps:

[0095] S20. Convert each of the N myoelectric panoramic representation maps into a one-dimensional sEMG vector and then splice them into a two-dimensional matrix to form the ground truth;

[0096] S21. Sequence vector construction: Take K from the N myoelectric panoramic representation maps and then generate K one-dimensional sEMG vectors;

[0097] The N independent gesture reference representation maps in S20 and S21 are from a standard or measured data set; N is greater than K; K is greater than or equal to 2;

[0098] S22. Sequence padding: Pad the K one-dimensional sEMG vectors to generate a two-dimensional matrix with the same dimension as the ground truth;

[0099] S23. Encoder-decoder training: Train the two-dimensional matrix padded in S22 through an AT encoder and an AT decoder to obtain the reconstructed N independent gesture myoelectric panoramic representation maps;

[0100] S24. Compare the reconstructed N independent gesture myoelectric panoramic representation maps with the ground truth and adjust the parameters of the AT encoder and the AT decoder according to the comparison results;

[0101] Continuously repeat S20 to S24 to train and adjust the parameters of the AT encoder and the AT decoder until the correlation coefficient or similarity between the trained reconstructed myoelectric panoramic representation map and the ground truth is greater than or equal to the AT threshold, and stop the pre-training;

[0102] Specifically, when the AT threshold is the correlation coefficient, if the absolute value of the correlation coefficient is greater than or equal to 0.8, it is considered that the trained myoelectric panoramic representation map is very close to the ground truth; or when the AT threshold is the similarity, if the similarity is greater than or equal to 0.7 or 0.8, it is considered that the trained myoelectric panoramic representation map is very close to the ground truth; stop the pre-training.

[0103] S25. Decompose N functional partition mask maps from the pre-trained N independent gesture myoelectric panoramic representation maps.

[0104] A judgment module is used to judge whether the cross-population adjustment module needs to be calibrated; if so, after the measured representation maps of K independent gestures of the current subject are constructed and filled through sequence vectors, an electromyogram panoramic representation map is generated through the cross-population adjustment module, and then N functional partition mask maps are decomposed. The relevant channels in the independent gesture representation map are filtered by the N functional partition mask maps of the current subject to determine the biokinematic relationship between the muscle and the finger, that is, the finger-muscle relationship; otherwise, N functional partition mask maps are generated by applying the pre-trained electromyogram panoramic maps of N independent gestures to filter the relevant channels in the measured representation map of the independent gesture to determine the biokinematic relationship between the muscle and the finger, that is, the finger-muscle relationship; K is greater than or equal to 2 and less than N, and N is greater than or equal to 6.

[0105] Different from the traditional method of extracting time-domain, frequency-domain, and time-frequency-domain features from sEMG signals to obtain more muscle information, and also different from the interpolation method of earlier methods, the GPAT module performs spatial generative expansion on sparse sEMG signals. Through a pre-trained Transformer model, GPAT expands the one-dimensional sEMG information (vector: 1×16) into a two-dimensional sEMG panoramic map (matrix: 16×16), thereby obtaining additional sEMG information in space. Compared with sparse one-dimensional data, the sEMG panoramic map is more convenient for locating the region where the electromyogram is located when distinguishing muscle functions.

[0106] As a possible implementation method, the cross-population adjustment module is calibrated by the following method:

[0107] S30. The current subject performs K standard actions to obtain K measured representation maps of independent gestures; K is greater than or equal to 2;

[0108] S31. Generate K sEMG vectors from the K measured representation maps of independent gestures;

[0109] Specifically, when S30 and S31 are executed, K is greater than or equal to 2, that is, the K independent gestures are multi-finger gesture movements of at least two items, and are used to obtain at least two sEMG vectors;

[0110] S32. Sequence filling: Construct a two-dimensional matrix Ⅲ containing K sEMG vectors through sequence filling;

[0111] S33. Encoding and decoding training: Reconstruct the electromyogram panoramic representation maps of N independent gestures through an AT encoder and an AT decoder; N is greater than K;

[0112] S34. Decompose N functional partition mask maps of the current subject from the electromyogram panoramic representation maps of N independent gestures.

[0113] The technical solution of the present invention will be further elaborated below in conjunction with the accompanying drawings and specific experimental examples.

[0114] The goal of the gesture motion tracking system provided in this embodiment is to establish the biomechanical relationship between muscles and fingers. During specific implementation, a dataset needs to be constructed first. Here, the Hyser dataset released in 2021 is selected; the Hyser dataset used a hardware platform with 256 sEMG channels to collect gesture information for 20 subjects in 5 sub-datasets: Sub-dataset 1 collected sEMG data of 34 common gestures; Sub-dataset 2 collected the maximum voluntary muscle contraction (MVC) data of each finger; Sub-dataset 3 recorded the force change of each finger with a single degree of freedom (1-DoF); Sub-dataset 4 saved the force change of the combination of multiple fingers (n-DoF); Sub-dataset 5 collected the force state during random finger movement; This embodiment uses Sub-datasets 3 and 4, and the generative enhancement module and cross-population adjustment module transform and reorganize the Hyser dataset. Since the Hyser dataset is based on 4 8×8 arrays of HD-sEMG units for signal acquisition, its channel index numbers are as Figure 2 shown. As can be seen from Figure 2 (a), the index numbers cannot be directly spliced into a 16×16 sEMG panoramic image in sequence. Therefore, data reorganization is required. Each frame of data (1×256) is combined into a sEMG panoramic image by the splicing method in Figure 2 (b), which is more conducive to decomposing the activation areas of different muscles.

[0115] Since each group of original data in Hyser records 3 finger movements (extension-flexion-extension) of the subject within 10 s, and the movement holding time is 2 s, a data packet contains 256 channels, and each channel has 51,200 data points. The experiment only selected the middle-section data of each movement time period (abduction: 16,501 - 17,000 and 36,501 - 37,000, adduction: 26,001 - 27,000) to make sEMG characterization maps of different actions.

[0116] The gesture motion tracking system provided in this application realizes the "seeing more from less" and "seeing parts from the whole" of sEMG signals. It is necessary to pre-train the generative enhancement module and the cross-population adjustment module respectively; it is necessary to collect different public datasets and evaluate their usability; because the generative enhancement module needs to finally generate a sEMG panoramic image, a dataset with more sEMG channels should be selected; in addition, the cross-population adjustment module needs to decompose independent finger movements from combined finger movements, so the dataset should also have data collection of various gesture actions. The construction of the dataset is elaborated in detail below:

[0117] (1) Construction of the dataset required for the generative enhancement module: Based on the input sparse sEMG signals, the generative enhancement module generates a dense sEMG panoramic map in the spatial dimension. However, the preparation of the dataset is the opposite. For all sEMG panoramic maps in sub-datasets 3 and 4, inverse downsampling is performed, that is, only the information of the first row of the 16×16 matrix is retained, and other spatial information is set to zero. The processed data is used as the input data for model training, while the original sEMG panoramic map is used as the ground truth for comparison with the output value.

[0118] (2) Construction of the dataset required for the cross-population adjustment module: The input of the cross-population adjustment module is the sEMG characterization map of joint finger movements, and the output is the sEMG panoramic maps of 8 independent finger movements. Therefore, in the experiment, the 9th joint action (i.e., simultaneous extension and flexion of 4 fingers) in sub-dataset 4 was selected as the input data, and the sEMG feature maps of the independent movements of the 2nd to 5th fingers in sub-dataset 3 were used as the ground truth for calculating the loss function.

[0119] The generative enhancement module of this system adopts a scheme to expand the sparse sEMG to obtain a more spatially rich sEMG panoramic map. Further, to address the issue of individual differences, this scheme designs an AT regulator based on the Transformer architecture, which can adjust its own model parameters to avoid the parameter update of the overall framework of the system, saving the training time in practical applications. At the same time, to simplify the user's calibration actions, the AT can generate 8 sEMG masks (extension and flexion) of 4 fingers (index finger, middle finger, ring finger, little finger) using only 2 gestures.

[0120] Next, two experiments are conducted to demonstrate the sEMG information space expansion ability of the generative enhancement module.

[0121] Experiment 1: In this experiment, 15 out of 20 subjects (subject 1 to subject 15) were selected for model training. 90% of their data in Hyser sub-datasets 3 and 4 was used as the training set, and 10% of the data was used as the test set. As Figure 3 shown, taking 10 finger movements of subject 1 as an example, the differences between the sEMG characterization map processed by the generative enhancement module and the ground truth are visually compared. It can be seen from the left column of sEMG maps that the muscle activation areas for extending the fingers are concentrated on the back of the arm, while in the right column of maps, it is more uniform, and the muscle activation state of the inner side of the arm is obvious. However, it does not rule out that some subjects may unconsciously generate muscle linkage during hand movements. Further, the similarity SSIM between the two is quantitatively evaluated. The experiment uses SSIM as the evaluation index, as shown in equations (1) to (6):

[0122]

[0123] c 1=(k 1 L) 2 ,c 2 =(k 2 L) 2 (6)

[0124] x i ,y i respectively represent the pixel values of each of the two images whose similarities are to be compared, μ x , μ y is the mean of x and y, σ x , σ y is the variance of x and y. The constants c 1 and c 2 are to avoid the situation where the denominator is zero. L is the pixel dynamic range (0 - 255) of the sEMG characterization map, and k 1 and k 2 are 0.01 and 0.03 respectively.

[0125] See Figure 4 . For the SSIM barrel charts after spatial enhancement of various hand movements of 15 subjects in the sEMG space, from left to right are subjects No. 1 to No. 15. It can be seen from the figure that the average fitting degree reaches 76.28%. This result shows that the GPAT unit can effectively perform spatial enhancement on sparse information and generate the sEMG panoramic characterization map.

[0126] Experiment 2: In this experiment, data of another 5 subjects out of 20 were selected to test the enhancement effect of the data of subjects not participating in the model training in the generative enhancement module.

[0127] See Figure 5 . Among them, Figure 5 (a) shows the effects of the sEMG maps of multiple movements of Subject No. 17 outside the pre - training dataset before and after being processed by the generative enhancement module. It can be seen from the figure that the similarity between the sEMG panoramic map after information enhancement and the muscle activation area of the ground truth is high. This shows that the pre - trained GPAT unit still has a certain degree of compatibility for new subjects. Figure 5 (b) Statistics were made on new subjects No. 16 to No. 20 outside the dataset. Although the performance of these 5 new subjects in the pre - trained model is lower than that of the old subjects, the average SSIM still reaches 75.61%.

[0128] In addition, in practice, different users need to complete the actions of simultaneous extension and flexion of 4 fingers according to the instructions during calibration to obtain the combined sEMG map. 90% of the four - finger linkage data of 20 subjects was used as the input of the training set, and the corresponding data of the independent movement of 4 fingers was used as the ground truth of the training set. The remaining 10% was used as the test set to demonstrate the information decomposition ability of the cross - population adjustment module.

[0129] By selecting representative examples (index finger extension and little finger flexion), it is found that the muscle partition mask produced by the combined movement of Subject 1 after being processed by the cross-population adjustment module is very similar to the actual muscle partition. Through quantitative calculation of the similarity, specifically in this embodiment, N = 8, K = 6; both the GPAT threshold and the AT threshold are greater than or equal to 0.7 or 0.8; the average SSIM of the muscle partitions of 20 subjects reaches 74.87%. For the detailed statistical results, see Table 1:

[0130] Table 1 SSIM Statistical Table of 8 Partitions of 20 Subjects

[0131]

[0132] In addition to the decomposition effect on different individuals, effective positioning of the position offset during wearing is also required in actual calibration. In the experiment, the sEMG characterization map is cyclically shifted by columns to simulate the situation of electrode deflection, that is, the last column is moved to the front column, and the other columns are shifted backward column by column. Figure 6 The muscle partition effect when the electrode of Subject 1 is offset by 3 columns, where Figure 6 (a) is the sEMG characterization map of four-finger flexion before and after AT calibration, Figure 6 (b) is the sEMG characterization map of four-finger abduction before and after AT calibration, Figure 6 (c.1) is the abduction mask of the index finger before and after AT calibration, Figure 6 (c.2) is the flexion mask of the middle finger, Figure 6 (c.3) is the abduction mask of the ring finger. It can be seen that the newly generated partition mask of the pre-trained model also shifts correspondingly with the movement of the input information. Therefore, the pre-trained model of AT proposed in this embodiment can also play a role in calibrating the offset problem.

[0133] In the second aspect, the present invention provides a gesture motion tracking method based on surface electromyography, relying on the generative enhancement module and the cross-population adjustment module. The gesture motion tracking method includes the following steps:

[0134] Collect the independent gestures of the current subject and obtain a one-dimensional sEMG vector;

[0135] Fill and construct a two-dimensional matrix Ⅰ containing the one-dimensional sEMG vector;

[0136] After receiving the two-dimensional matrix Ⅰ, enhance it to the actual measurement characterization map of the independent gesture;

[0137] Pre-train the generative enhancement module to obtain the independent gesture characterization map corresponding to the independent gesture;

[0138] Pre-train the cross-population adjustment module to obtain N functional partition mask maps decomposed from the myoelectric panoramic characterization map based on N independent gestures, where N is an even number greater than or equal to 6;

[0139] Determine whether the cross-population adjustment module needs to be calibrated; if so, after the measured representation maps of K independent gestures of the current subject are constructed and filled through sequence vectors, generate an electromyogram panoramic representation map through the cross-population adjustment module, and then decompose it into N functional partition mask maps. Use the N functional partition mask maps of the current subject to filter the relevant channels in the independent gesture representation map to determine the biokinematic relationship between the muscle and the finger, that is, the finger-muscle relationship; otherwise, use the N pre-trained electromyogram panoramic maps of independent gestures to generate masks to obtain N functional partition mask maps to filter the relevant channels in the independent gesture representation map to determine the biokinematic relationship between the muscle and the finger, that is, the finger-muscle relationship, where K is greater than or equal to 2 and less than N.

[0140] Fill and construct a two-dimensional matrix Ⅰ containing one-dimensional sEMG vectors, and after receiving the two-dimensional matrix Ⅰ, enhance it to a measured representation map of independent gestures. Specifically: generate row by row in an iterative manner, and generate each row of sEMG data on the two-dimensional matrix Ⅰ to fully consider the one-dimensional sEMG vectors.

[0141] Although the present invention has been described in connection with various embodiments, however, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure content, and the like. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude the case of multiple. A single processor or other unit can implement several functions listed in the specification. Certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0142] Although the present invention has been described in connection with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present invention. Accordingly, the present specification and the drawings are merely exemplary descriptions of the present invention and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A gesture motion tracking system based on surface electromyography for determining the biokinematic relationship between the muscles and fingers of a current subject, characterized in that: The gesture motion tracking system comprises: A data acquisition module, used to output the sEMG vector of the current subject; A pre-filling module for constructing a sEMG matrix I containing sEMG vectors; A generative enhancement module, configured to receive the sEMG matrix I and enhance it into an independent gesture representation graph; A cross-population adjustment module is used to decompose the EMG panoramic representation map into a corresponding number of functional partition mask maps; The generative enhancement module and the cross-population adjustment module both require pre-training; A judgment module is used to judge whether the cross-population adjustment module needs to be calibrated; if so, the K independent gesture measured representation maps of the current subject are used to construct and fill in the sequence vector, and then the electromyographic panoramic representation map is generated through the cross-population adjustment module, and then N functional partition mask maps are decomposed, and the N functional partition mask maps of the current subject are used to filter the relevant channels in the independent gesture representation map to determine the biokinematic relationship between muscles and fingers, that is, the finger-muscle relationship; otherwise, the pre-trained N independent gesture electromyographic panoramic maps are used to generate masks to obtain N functional partition mask maps to filter the relevant channels in the independent gesture measured representation map to determine the biokinematic relationship between muscles and fingers, that is, the finger-muscle relationship; the K is greater than or equal to 2 and less than N.

2. The gesture motion tracking system based on surface muscle electricity according to claim 1, characterized in that: The generative enhancement module is pre-trained by the following steps: S10. Vector construction: extract a row of an independent gesture electromyographic panoramic representation image from a standard or measured data set to generate a one-dimensional sEMG vector corresponding to the independent gesture; The myoelectric panoramic representation of the independent gesture is groundtruth; S11. Pre-filling: pre-filling the one-dimensional sEMG vector into the sEMG matrix I; The sEMG matrix I is a two-dimensional matrix; S12. Encoder and decoder training: training is performed through the GPAT encoder and the GPAT decoder to obtain a myoelectric panoramic representation corresponding to the independent gesture after training; S13. Compare the trained EMG panoramic representation corresponding to the independent gesture with the groundtruth, and adjust the GPAT encoder and GPAT decoder parameters according to the comparison results; Repeat S10 to S13 continuously to train and adjust the parameters of the GPAT encoder and GPAT decoder until the absolute value of the correlation coefficient or similarity between the trained EMG panoramic representation and the groundtruth is greater than or equal to the GPAT threshold, and then stop pre-training.

3. The gesture motion tracking system based on surface muscle electricity according to claim 1, characterized in that: The cross-population adjustment module is pre-trained by the following steps: S20. Convert each of the N panoramic electromyographic representation images into a one-dimensional sEMG vector and then concatenate them into a two-dimensional matrix to form a groundtruth; S21. Sequence vector construction: Take K from N EMG panoramic representations and generate K one-dimensional sEMG vectors; The N independent gesture reference representation images in S20 and S21 are derived from standard or measured data sets; S22.Sequence filling: fill the K one-dimensional sEMG vectors to generate a two-dimensional matrix with the same dimension as the groundtruth; S23. Encoder and decoder training: The two-dimensional matrix filled by S22 is trained through an AT encoder and an AT decoder to obtain a reconstructed N independent gesture electromyography panoramic representation map; S24. Compare the reconstructed N independent gesture electromyography panoramic representation images with the groundtruth, and adjust the AT encoder and AT decoder parameters according to the comparison results; Repeat S20 to S24 continuously, train and adjust the parameters of the AT encoder and AT decoder, until the absolute value of the correlation coefficient or similarity between the reconstructed EMG panoramic representation image after training and the groundtruth is greater than or equal to the AT threshold, and stop pre-training; S25. Decompose N functional partition mask images from the N pre-trained independent gesture electromyography panoramic representation images.

4. The gesture motion tracking system based on surface muscle electricity according to claim 1, characterized in that: The cross-population adjustment module is calibrated as follows: S30. The current subject performs K standard actions to obtain K independent gesture measured representation graphs; S31. Generate K sEMG vectors from K independent gesture measured representation images; S32.Sequence filling: Construct a two-dimensional matrix containing K sEMG vectors by sequence filling III; S33. Encoder-decoder training: Reconstruct the EMG panoramic representation of N independent gestures through AT encoder and AT decoder; S34. Decompose N functional partition mask images of the current subject from the electromyographic panoramic representation images of N independent gestures.

5. The gesture motion tracking system based on surface muscle electricity according to claim 1, characterized in that: The generative enhancement module is obtained based on the pre-training of the Transformer architecture; The generative enhancement module includes a GPAT encoder and a GPAT decoder, which are used to realize sequence-to-sequence signal conversion; the GPAT encoder and the GPAT decoder each include multiple GPAT encoding layers and GPAT decoding layers to increase the scale of the pre-trained Transformer architecture weight parameters.

6. The gesture motion tracking system based on surface muscle electricity according to claim 5, characterized in that: Each of the GPAT encoding layers at least includes a multi-head attention layer and a position-by-position feedforward network layer; The multi-head attention layer is provided with six heads, each of which includes Q, K, and V matrices. The self-attention mechanism is applied to perform inner product on the transpose of Q and K to obtain a correlation matrix. The value A obtained by cross-multiplying the correlation matrix by the V matrix represents the degree of influence of other vectors on the current vector. The value A is superimposed on the input value and then introduced into the position-by-position feedforward network layer to obtain the output value O. The multi-head attention layer included in each of the GPAT encoding layers is masked with a lower triangular matrix to ensure sequential generation; each of the GPAT encoding layers also includes a multi-head attention module, the output of the multi-head attention module in the previous layer is used as the index value, the output of the GPAT encoder is used as the key and value, and the feedforward network layer of the GPAT encoding layer is used as the output layer.

7. The gesture motion tracking system based on surface muscle electricity according to claim 1, characterized in that: The cross-population adjustment module is obtained based on Transformer architecture pre-training; The cross-crowd adjustment module AT includes an AT encoder and an AT decoder; the AT decoder does not include a mask layer, that is, the AT encoder output is not used as the next index value to complete the iterative generation during the operation; the loss value between the output of the AT decoder and the true value will be back-propagated to the AT encoder of the cross-crowd adjustment module to achieve parameter update.

8. The gesture motion tracking system based on surface muscle electricity according to claim 1, characterized in that: The N is greater than or equal to 6.

9. A gesture motion tracking method based on surface muscle electricity, characterized in that: Relying on the generative enhancement module and the cross-crowd adjustment module, the gesture motion tracking method includes the following steps: Collect the current subject's independent gestures and obtain a one-dimensional sEMG vector; Fill and construct a two-dimensional matrix I containing the one-dimensional sEMG vector; After receiving the two-dimensional matrix I, it is enhanced into an independent gesture measured representation diagram; Pre-training the generative enhancement module to obtain independent gesture representation graphs corresponding to independent gestures; Pre-training the cross-population adjustment module to obtain N functional partition mask images based on the electromyographic panoramic representation image of N independent gestures; Determine whether the cross-crowd adjustment module needs calibration; if so, use the K independent gesture measured representations of the current subject to construct and fill the sequence vectors, then generate the electromyographic panoramic representation through the cross-crowd adjustment module, and then decompose N functional partition mask images, and use the N functional partition mask images of the current subject to filter the relevant channels in the independent gesture representation image to determine the biokinematic relationship between muscles and fingers, that is, the finger-muscle relationship; Otherwise, the pre-trained N independent gesture electromyography panoramic images are used to generate masks to obtain N functional partition mask images to filter the relevant channels in the independent gesture representation image to determine the biokinematic relationship between muscles and fingers, that is, the finger-muscle relationship, where K is greater than or equal to 2 and less than N, and N is an even number greater than or equal to 6.

10. The method for tracking gesture motion based on surface muscle electricity according to claim 9, characterized in that: The two-dimensional matrix I containing the one-dimensional sEMG vector is filled and constructed, and after receiving the two-dimensional matrix I, it is enhanced into an independent gesture measured representation diagram, specifically: each row of sEMG data is generated on the two-dimensional matrix I row by row through iteration to fully consider the one-dimensional sEMG vector.