Hand action recognition method and system
Through the EWT and EEMD of the boundary adaptive backfill mechanism, the hand sensor signals are processed, important IMF components are screened and combined with the improved CNN network, the boundary effect and noise influence in hand motion recognition are solved, and the accuracy and robustness of recognition are improved.
Patent Information
- Application Number
- CN202510839002.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When faced with the fusion of complex movements and multi-sensor data, the existing hand movement recognition methods have problems such as obvious boundary effects, great noise impact, and difficulty in selecting features, resulting in inaccurate identification.
The hand sensor signals are processed using the boundary adaptive backfill mechanism, and the IMF components with large influence factors of multi-scale nonlinear relationships are screened, and the hand motion recognition is combined with time-frequency domain feature extraction and improved convolutional neural network (CNN).
Effectively reduce noise and boundary effects, improve the accuracy and robustness of hand movement recognition, especially in complex environments, which can better capture signal characteristics and improve recognition accuracy.
Smart Images

Figure CN120372407A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of flexible sensors, and particularly relates to a hand gesture recognition method and system. Background Art
[0002] With the rapid development of artificial intelligence and intelligent wearable devices, hand gesture recognition, as an important human-computer interaction method, has received extensive attention. Hand gesture recognition technology can sense hand gestures through sensors and perform analysis, thus providing effective support for fields such as virtual reality, intelligent healthcare, and robot control. However, the dynamics, complexity, and high-precision requirements of hand gestures pose many challenges to hand gesture recognition. Especially in the process of sensor data processing, how to efficiently denoise, reduce signal interference, extract key features, and how to process multi-dimensional heterogeneous sensor data have become important problems in this field.
[0003] Existing hand gesture recognition methods mostly adopt technologies combining traditional signal processing and deep learning. Traditional methods extract signal features through time-domain or frequency-domain analysis. Although they can complete basic tasks to a certain extent, they often encounter problems such as obvious boundary effects, large noise influence, and difficult feature selection when facing complex gestures and multi-sensor data fusion. Therefore, existing methods cannot meet the actual application requirements of high precision and high robustness. Summary of the Invention
[0004] The present invention provides a hand gesture recognition method and system for solving the technical problem of inaccurate hand gesture recognition caused by obvious boundary effects, large noise influence, difficult feature selection, and the like.
[0005] In a first aspect, the present invention provides a hand gesture recognition method, including: Obtaining a strain signal of a hand sensor, and performing EWT transformation on the strain signal according to a boundary adaptive backfill mechanism to obtain a target strain signal; Decomposing the target strain signal according to EEMD to obtain at least one IMF component, calculating a multi-scale nonlinear relationship influence factor of the at least one IMF component, and selecting at least one target IMF component whose multi-scale nonlinear relationship influence factor is greater than a preset threshold; Performing feature extraction on the at least one target IMF component, combining the extracted time-domain features into a time-domain feature vector, and combining the extracted frequency-domain features into a frequency-domain feature vector, and screening the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector; Normalize the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector; Input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand gesture recognition result corresponding to the strain signal.
[0006] In a second aspect, the present invention provides a hand gesture recognition system, including: An acquisition module configured to acquire the strain signal of a hand sensor and perform EWT transformation on the strain signal according to a boundary adaptive backfill mechanism to obtain a target strain signal; A calculation module configured to decompose the target strain signal according to EEMD to obtain at least one IMF component, calculate the multi-scale non-linear relationship influence factor of the at least one IMF component, and select at least one target IMF component whose multi-scale non-linear relationship influence factor is greater than a preset threshold; A screening module configured to extract features from the at least one target IMF component, combine the extracted time-domain features into a time-domain feature vector, and combine the extracted frequency-domain features into a frequency-domain feature vector, and screen the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector; A fusion module configured to normalize the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector; An output module configured to input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand gesture recognition result corresponding to the strain signal.
[0007] In a third aspect, there is provided an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the hand gesture recognition method according to any embodiment of the present invention.
[0008] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program instructions are executed by a processor, the processor executes the steps of the hand gesture recognition method according to any embodiment of the present invention.
[0009] The hand gesture recognition method and system of the present application Brief Description of the Drawings
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] Figure 1 It is a flowchart of a hand motion recognition method provided by an embodiment of the present invention; Figure 2 It is a structural block diagram of a hand motion recognition system provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0013] Please refer to Figure 1 , which shows a flowchart of a hand motion recognition method of the present application.
[0014] As Figure 1 shown, the hand motion recognition method specifically includes the following steps: Step S101, obtain the strain signal of the hand sensor, and perform EWT transformation on the strain signal according to the boundary adaptive backfill mechanism to obtain the target strain signal.
[0015] In this step, the functional expression of the boundary adaptive backfill mechanism is: , wherein, is the boundary weighting function, is the time point of the left boundary of the signal, is the time, is the total time length of the signal, , are both exponents for controlling Gaussian decay, , are both rates for controlling signal decay, , All are factors affecting the polynomial, , , , are all the orders of the polynomial, is the time point of the right boundary of the signal, represents the exponential operation.
[0016] Specifically, according to the boundary adaptive backfilling mechanism, the EWT transform (Empirical Wavelet Transform) is performed on the strain signal to obtain the target strain signal, including: According to the boundary adaptive backfilling mechanism, boundary weighting processing is performed on the strain signal to obtain the strain signal after boundary backfilling, and the expression is: , In the formula, is the strain signal after boundary backfilling, is the strain signal; Perform the EWT transform on the strain signal after boundary backfilling to obtain the target strain signal, and the expression is: , In the formula, is the target strain signal of the th frequency band, is the inverse Fourier transform, is 's inverse Fourier transform, is the th wavelet filtering function of the frequency band.
[0017] Step S102: Decompose the target strain signal according to EEMD to obtain at least one IMF component, calculate the multi-scale nonlinear relationship influence factor of the at least one IMF component, and select at least one target IMF component whose multi-scale nonlinear relationship influence factor is greater than a preset threshold.
[0018] In this step, according to EEMD (Ensemble Empirical Mode Decomposition), the target strain signal is decomposed to obtain at least one IMF component, and the expression is: , In the formula, is the th decomposition result obtained by decomposing the th frequency band signal through EEMD, is the th frequency band signal's IMF components is the residual term obtained by EEMD decomposition of the th band signal, is the total number of IMFs of each band signal; Calculate the multi-scale non-linear relationship influence factor of the at least one IMF component, and the expression is: , , , , In the formula, represents the multi-scale non-linear relationship influence factor of the th IMF of the th band signal, represents and non-linear mutual information between, represents a non-linear function, represents and weighting function of, represents the th correlation coefficient of the th IMF of the represents the th correlation coefficient of the th IMF of the represents and similarity in the frequency domain, represents a weighted time-domain correlation function, represents the th value of the th IMF of the th band signal at time represents the th value of the th IMF of the th band signal at time represents a similarity function of the spectrum, represents frequency, represents polynomial coefficients, represents the order of the polynomial, represents the integral from to , represents a weighted factor of frequency, represents an exponent for controlling the distance metric, , are different time points, is the th value of the th IMF of the th band signal at frequency is the th value of the th IMF of the th band signal at frequency is the total number of frequency bands, is the correlation weighting factor, is the total number of IMFs obtained by decomposing one band signal, is the correlation difference index; Select at least one target IMF component whose multi-scale non-linear relationship influence factor is greater than a preset threshold.
[0019] Step S103: Extract features from the at least one target IMF component, combine the extracted time-domain features into a time-domain feature vector, and combine the extracted frequency-domain features into a frequency-domain feature vector, and screen the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector.
[0020] In this step, features are extracted from the at least one target IMF component to obtain time-domain features, and the time-domain features at multiple scales are combined into a time-domain feature vector. The expression is: , , where is the vector obtained by combining the time-domain features extracted from the screened IMFs, , , are the first-order, second-order, and th value of the th IMF of the th band signal, is the th value of the th IMF of the th band signal, is the time length of the IMF, is the translation of the IMF by scale factors, represents the th value of the th IMF of the th band signal at time is the non-linear index, Indicates scaling the scale factor, which is a constant; Extract features from the at least one target IMF component to obtain frequency-domain features, and combine each frequency-domain feature with the weighted second moment of each frequency-domain feature to obtain a frequency-domain feature vector, with the expression: , , , wherein, is the frequency-domain feature vector of the th IMF of the th band signal, is the frequency-domain structure compression ratio of the th IMF of the th band signal, is the total number of frequency bands, is the frequency, is the th band signal of the th IMF at the frequency , is the th band signal of the th IMF at the th frequency, is the smoothing factor, represents performing a Fourier transform on the th IMF of the th band signal, represents the value of the th IMF of the th band signal at time , is the Fourier transform; Calculate the first feature score of each time-domain feature vector and the second feature score of each frequency-domain feature vector according to a preset feature scoring function, where the expression of the feature scoring function is: , wherein, is the scoring function of the th IMF of the th band signal for the feature , is the dispersion of the feature , is the th band signal of the th IMF for the time-domain feature, is the The frequency domain features of the th IMF of a frequency band signal, is the mean value of the features in all IMFs, and are both skewness penalty control parameters; At least one target time-domain feature vector with the first feature score greater than the first score threshold and at least one target frequency-domain feature vector with the second feature score greater than the second score threshold are respectively selected.
[0021] Step S104: Normalize the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector.
[0022] In this step, the normalization processing of the target time-domain feature vector and the target frequency-domain feature vector is performed, and the expression is: , , , wherein, is the time-domain feature of the th IMF of the th frequency band signal after normalization, is the frequency-domain feature of the th IMF of the th frequency band signal after normalization, is a non-linear normalization mapping function, is the time-domain feature of the th IMF of the th frequency band signal, is the frequency-domain feature of the th IMF of the th frequency band signal, is the standard deviation of the feature set , is the mean value of the feature set , is an adjustment parameter, is a constant; Cross-weighted fusion is performed on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector, and the expression is: , , , , In the formula, is the th composite feature vector of the th IMF of the th band signal, is the th weighted coefficient of the time-domain and frequency-domain features of the th IMF of the th band signal, is the th weighted coefficient of the time-domain feature of the th IMF of the th band signal, is the th weighted coefficient of the frequency-domain feature of the th IMF of the th band signal, is the th mutual information between and the target variable
[0023] Step S105, input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand gesture recognition result corresponding to the strain signal.
[0024] In this step, multi-dimensional influence factors are added before the convolutional layer of the CNN network, including a time influence factor to dynamically adjust the feature importance at each time point according to the time-domain part of the input composite feature vector, and a frequency-domain influence factor to measure the importance of different frequency components to the final feature in the frequency domain. The expression of the time influence factor is: , In the formula, represents the time influence factor of the th IMF of the th band signal at the time point , represents the time-domain feature in the composite feature vector, represents the time-domain change adjustment factor, represents the time-domain smooth transition adjustment factor, represents the instantaneous change rate of the time-domain signal at the time point , represents the time-domain balance term to prevent over-response to extreme changes, represents the total step length of the time point; The expression of the domain influence factor is: , wherein, represents the th frequency domain influence factor of the th IMF of the th band signal, represents the frequency domain feature in the composite feature vector, represents the frequency domain change adjustment factor, represents the frequency domain smooth transition adjustment factor, represents the instantaneous change rate of the frequency component, represents the frequency domain balance term, is the total number of frequency bands; Combines the time influence factor and the frequency domain influence factor and applies them to the features in the convolutional layer. The expression is: , wherein, represents the adjusted composite feature, represents the input signal, which is the adjusted composite feature vector here, is the composite feature vector of the th IMF of the
[0025] After the convolutional layer, apply the improved non-linear activation function to adjust the composite feature vector adjusted by the influence factor, so that the model can learn more complex non-linear relationships. , wherein, , , respectively represent different activation functions, , , respectively represent the adjustable parameters of the corresponding activation functions, controlling the contribution of each activation function, represents the improved non-linear activation function; Add an attention mechanism after the output of the convolutional layer, including a time domain attention mechanism and a frequency domain attention mechanism, so that the network can adaptively focus on the most informative features and improve the network's learning ability for key features.
[0026] The expression of the time domain attention mechanism is: , wherein, represents the th time domain attention weight of the th IMF of the th band signal at the time point represents the current time point, represents the total number of time steps, represents the number of time steps in the time window, represents the length of the time window, represents the th th IMF of the time-domain feature after being processed by the activation function at the time point represents in the time window near the time point the cumulative sum of the time-domain features of the past time points, represents the adjustment parameter of the time-domain attention weight, represents the time-domain adjustment term parameter; The expression of the frequency-domain attention mechanism is: , In the formula, represents the th th IMF of the frequency-domain attention weight at the frequency point represents the adjustment factor of the frequency-domain attention, represents the th th IMF of the frequency-domain feature after being processed by the activation function at the frequency point represents near the frequency the cumulative sum of the frequency-domain features of the adjacent represents the size of the sliding window, represents an index used to select the adjacent frequency points near the current frequency point, represents the frequency-domain adjustment term parameter.
[0027] Perform a weighting operation on the features after non-linear activation, and the expression is: , , In the formula, represents the time-domain feature weighted by the time-domain attention mechanism, represents element-wise multiplication, represents the time-domain feature after being processed by the activation function, represents the frequency-domain feature weighted by the frequency-domain attention mechanism, represents the frequency-domain feature after being processed by the activation function.
[0028] In summary, for the method of the present application, the original strain signal is subjected to EWT transformation, and a boundary adaptive backfilling mechanism is added to decompose the original strain signal into components of different frequency bands to remove noise and clutter. By introducing the boundary adaptive backfilling mechanism, this method can effectively solve the boundary effect generated when using EWT to decompose signals. Traditional backfilling methods (such as zero padding) will produce mutations at the boundaries of the signal, affecting subsequent feature extraction and model training. After using the boundary adaptive backfilling mechanism, the boundary part of the signal will be smoothly transitioned according to the local mean and gradient, making the signal boundary more consistent with the internal features, thus avoiding the negative impact of the boundary effect on the analysis results.
[0029] The signal after EWT transformation is subjected to EEMD decomposition to screen important IMF components and improve the quality of IMF components: Through EEMD decomposition, a complex signal can be decomposed into multiple IMF components with different frequency components. By screening the most important IMF components, useful information is retained, and irrelevant or low-quality components are removed.
[0030] Dynamically screen important IMFs: By applying the non-linear relationship influence factor, the most representative IMF components can be automatically screened out, thus improving the efficiency of feature selection and avoiding the interference of redundant information.
[0031] Time-domain and frequency-domain feature extraction and scoring function screening Time-domain and frequency-domain feature extraction: In this step, a new type of time-frequency domain feature extraction method is proposed, which is different from the traditional time-frequency domain feature extraction methods. The newly proposed time-frequency domain features can more accurately characterize the time-domain and frequency-domain characteristics of the signal and capture more complex change patterns of the signal in the time-frequency domain.
[0032] Scoring function screening: A custom scoring function is used to screen the extracted time-frequency domain features, and the most important features are selected to obtain the final time-frequency domain feature set.
[0033] Innovative time-frequency domain features: The newly proposed time-frequency domain features are more refined and can capture complex patterns of the signal in the time-frequency domain. These features can provide more information than traditional features and contribute to improving the accuracy of signal analysis.
[0034] Feature screening: By screening the time-frequency features through the scoring function, irrelevant or redundant features can be effectively removed, and only the feature part that best represents the signal is retained, thus improving the efficiency and accuracy of the model.
[0035] Normalization, weighted fusion of time-frequency domain features, and input to the improved CNN network Normalization and weighted fusion of time-frequency domain features: By normalizing and weighted fusing the selected time-frequency domain features, the time domain and frequency domain features can be combined more balancedly, eliminating the differences between different feature dimensions. Through weighted fusion, the expression ability of time-frequency features is enhanced. Balancing time-frequency features: Through normalization and weighted fusion, the balance of time domain and frequency domain features when inputting into the network is ensured, enabling the network to fairly utilize time-frequency feature information and avoiding overemphasis on features of a certain dimension.
[0036] The introduction of multi-dimensional influence factors (TIF and FIF) and the combination of time domain and frequency domain attention mechanisms enable the time domain and frequency domain features of the signal to more accurately reflect their contributions to the task, thereby improving the expression ability of features and the learning ability of the model. Especially in complex signals, it can highlight the most informative part and enhance the accuracy of hand gesture recognition. It optimizes the model's processing ability for time-frequency features, avoids the limitations of single time domain or frequency domain features, enables the network to perform feature selection and weighting in a higher dimension, and thus enhances the recognition accuracy of the model for hand gestures, especially the robustness in complex scenarios.
[0037] The technical effects are specifically as follows: Improve signal quality: Effectively remove noise and clutter, reduce the influence of boundary effects, and ensure signal stability; Enhance feature selection and expression ability: By dynamically screening the most important IMF components and time-frequency features, the feature extraction ability and accuracy of the model are improved; Optimize feature weighted fusion: Through the adaptive weighted fusion of time-frequency domain features, the model can better capture important information in the signal; Improve classification accuracy: Input the optimized features into the improved CNN network to improve the accuracy and robustness of hand gesture recognition, especially in complex environments.
[0038] Please refer to Figure 2 , which shows the structural block diagram of a hand gesture recognition system of the present application.
[0039] As Figure 2 shown, the hand gesture recognition system 200 includes an acquisition module 210, a calculation module 220, a screening module 230, a fusion module 240, and an output module 250.
[0040] Among them, the acquisition module 210 is configured to acquire the strain signal of the hand sensor and perform EWT transformation on the strain signal according to the boundary adaptive backfill mechanism to obtain the target strain signal; A calculation module 220, configured to decompose the target strain signal according to EEMD to obtain at least one IMF component, calculate a multi-scale non-linear relationship influence factor of the at least one IMF component, and select at least one target IMF component whose multi-scale non-linear relationship influence factor is greater than a preset threshold; A screening module 230, configured to extract features from the at least one target IMF component, combine the extracted time-domain features into a time-domain feature vector, and combine the extracted frequency-domain features into a frequency-domain feature vector, and screen the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector; A fusion module 240, configured to perform normalization processing on the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector; An output module 250, configured to input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand motion recognition result corresponding to the strain signal.
[0041] It should be understood that Figure 2 The modules described in Figure 1 correspond to the respective steps in the method described in reference Figure 2 Thus, the operations, features, and corresponding technical effects described above for the method also apply to the modules in
[0042] In some other embodiments, the embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the program instructions are executed by a processor, the processor executes the hand motion recognition method in any of the above method embodiments; As an implementation manner, the computer-readable storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as: Obtain the strain signal of the hand sensor, and perform EWT transformation on the strain signal according to the boundary adaptive backfill mechanism to obtain a target strain signal; Decompose the target strain signal according to EEMD to obtain at least one IMF component, calculate a multi-scale non-linear relationship influence factor of the at least one IMF component, and select at least one target IMF component whose multi-scale non-linear relationship influence factor is greater than a preset threshold; Extract features from the at least one target IMF component, combine the extracted time-domain features into a time-domain feature vector, and combine the extracted frequency-domain features into a frequency-domain feature vector, and screen the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector; Perform normalization processing on the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector; Input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand action recognition result corresponding to the strain signal.
[0043] A computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area may store an operating system and application programs required for at least one function; the storage data area may store data created according to the use of the hand action recognition system, etc. In addition, the computer-readable storage medium may include high-speed random access memory, and may also include memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the hand action recognition system through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0044] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 3 shown, the device includes: a processor 310 and a memory 320. The electronic device may further include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means, Figure 3 taking connection through a bus as an example. The memory 320 is the above-mentioned computer-readable storage medium. The processor 310 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 320, that is, implements the hand action recognition method in the above method embodiment. The input device 330 may receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the hand action recognition system. The output device 340 may include a display device such as a display screen.
[0045] The above electronic device can execute the method provided by the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiments of the present invention.
[0046] As an implementation manner, the above electronic device is applied to a hand movement recognition system and is used for a client, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtain the strain signal of the hand sensor, and perform EWT transformation on the strain signal according to the boundary adaptive backfilling mechanism to obtain a target strain signal; Decompose the target strain signal according to EEMD to obtain at least one IMF component, calculate the multi-scale non-linear relationship influence factor of the at least one IMF component, and select at least one target IMF component whose multi-scale non-linear relationship influence factor is greater than a preset threshold; Extract features from the at least one target IMF component, combine the extracted time-domain features into a time-domain feature vector, and combine the extracted frequency-domain features into a frequency-domain feature vector, and screen the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector; Perform normalization processing on the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector; Input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand movement recognition result corresponding to the strain signal.
[0047] Through the description of the above implementation manners, those skilled in the art can clearly understand that each implementation manner can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hand gesture recognition method, characterized in that, Including: Obtain the strain signal of the hand sensor, and perform EWT transformation on the strain signal according to the boundary adaptive backfilling mechanism to obtain the target strain signal; Decompose the target strain signal according to EEMD to obtain at least one IMF component, calculate the multi-scale nonlinear relationship influence factor of the at least one IMF component, and select at least one target IMF component whose multi-scale nonlinear relationship influence factor is greater than a preset threshold; Extract features from the at least one target IMF component, combine the extracted time-domain features into a time-domain feature vector, and combine the extracted frequency-domain features into a frequency-domain feature vector, and screen the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector; Perform normalization processing on the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector; Input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand action recognition result corresponding to the strain signal.
2. The hand motion recognition method according to claim 1, characterized in that The functional expression of the boundary adaptive backfilling mechanism is: , In the formula, is the boundary weighting function, is the time point of the left boundary of the signal, is the time, is the total time length of the signal, and are both exponents for controlling Gaussian decay, and are both rates for controlling signal decay, and are both factors affecting the polynomial, and and and are all orders of the polynomial, is the time point of the right boundary of the signal, represents the exponentiation operation.
3. A hand movement recognition method according to claim 1, characterized in that The step of decomposing the target strain signal according to EEMD to obtain at least one IMF component, calculating the multi-scale nonlinear relationship influence factor of the at least one IMF component, and selecting at least one target IMF component whose multi-scale nonlinear relationship influence factor is greater than a preset threshold includes: Decompose the target strain signal according to EEMD to obtain at least one IMF component, and the expression is: , Wherein, is the -th decomposition result obtained by EEMD decomposition of the -th band signal, is the -th IMF component of the -th band signal, is the -th residue obtained by EEMD decomposition of the -th band signal, is the total number of IMFs of each band signal; Calculate the multi-scale nonlinear relationship influence factor of the at least one IMF component, and the expression is: , , , , In the formula, represents the th multi-scale non-linear relationship influence factor of the th IMF of the th band signal, represents the non-linear mutual information between and represents a non-linear function, and weighting function of represents the th correlation coefficient of the th IMF of the th band signal, represents the th correlation coefficient of the th IMF of the th band signal, represents the weighted time-domain correlation function, represents the th value of the th IMF of the th band signal at time represents the th value of the th IMF of the th band signal at time represents the spectral similarity function, represents frequency, represents polynomial coefficients, represents the order of the polynomial, represents the integral from to represents the frequency weighting factor, represents the exponent that controls the distance metric, and are respectively different time points, is the th value of the th IMF of the th band signal at frequency is the th value of the th IMF of the th band signal at frequency is the total number of frequency bands, is the correlation weighting factor, is the total number of IMFs decomposed from one band signal, is the correlation difference index; Select at least one target IMF component whose multi-scale nonlinear relationship influence factor is greater than a preset threshold.
4. A hand movement recognition method according to claim 1, characterized in that, The step of extracting features from the at least one target IMF component, combining the extracted time-domain features into a time-domain feature vector, and combining the extracted frequency-domain features into a frequency-domain feature vector, and screening the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector includes: Extract features from the at least one target IMF component to obtain time-domain features, and combine the time-domain features at multiple scales into a time-domain feature vector, and the expression is: , , In the formula, is a vector for combining the time-domain features extracted from the IMF after screening, , , are respectively the first-order, second-order, and -order dynamic fluctuation indices of the -th IMF of the -th band signal, is the -order dynamic fluctuation index of the -th IMF of the -th band signal, is the time length of the IMF, is the translation of the IMF by scale factors, represents the value of the -th IMF of the -th band signal at time , is the non-linear index, represents the scaling of the scale factor, is a constant; Extract features from the at least one target IMF component to obtain frequency-domain features, and combine each frequency-domain feature with the weighted second moment of each frequency-domain feature to obtain a frequency-domain feature vector, and the expression is: , , , In the formula, is the th frequency domain feature vector of the th IMF of the th band signal, is the th frequency domain structure compression rate of the th band signal, is the total number of frequency bands, is the th frequency, is the th spectral power of the th IMF of the th band signal at frequency is the th spectral power of the th IMF of the th band signal at the th frequency, is the smoothing factor, represents the th Fourier transform of the th IMF of the th band signal, is the Fourier transform; Calculate the first feature score of each time-domain feature vector and the second feature score of each frequency-domain feature vector according to a preset feature scoring function, where the expression of the feature scoring function is: , Wherein, is the th score function of the th IMF of the th band signal, is the dispersion of the feature , is the th time domain feature of the th IMF of the th band signal, is the th frequency domain feature of the th IMF of the th band signal, is the median of the means of all features, , are both skew penalty control parameters; Select at least one target time-domain feature vector whose first feature score is greater than the first score threshold, and select at least one target frequency-domain feature vector whose second feature score is greater than the second score threshold.
5. A hand movement recognition method according to claim 1, characterized in that Normalizing the target time-domain feature vector and the target frequency-domain feature vector, and performing cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector, including: Normalizing the target time-domain feature vector and the target frequency-domain feature vector, and the expression is: , , , Wherein, is the time-domain feature of the th IMF of the th band signal after normalization, is the frequency-domain feature of the th IMF of the th band signal after normalization, is the non-linear normalization mapping function, is the time-domain feature of the th IMF of the th band signal, is the frequency-domain feature of the th IMF of the th band signal, is the standard deviation of the feature set , is the mean of the feature set , is the adjustment parameter, is a constant; Performing cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector, and the expression is: , , , , Wherein, is the composite feature vector of the th IMF of the th band signal, is the weighted coefficient of the time-domain and frequency-domain features of the th IMF of the th band signal, is the weighted coefficient of the time-domain feature of the th IMF of the th band signal, is the weighted coefficient of the frequency-domain feature of the th IMF of the th band signal, is the mutual information between and the target variable , is the target variable, is the mutual information between and the target variable .
6. The hand motion recognition method according to claim 1, wherein The improved CNN model includes: A convolutional layer including multi-dimensional influence factors, where the multi-dimensional influence factors include a time influence factor and a frequency-domain influence factor, and the expression of the time influence factor is: , wherein, represents the th IMF of the th band signal at the time point represents the time domain feature in the composite feature vector, represents the time domain change adjustment factor, represents the time domain smooth transition adjustment factor, represents the instantaneous change rate of the time domain signal at the time point ; represents the time domain balance term to prevent over - response to extreme changes, represents the total step size of the time point; The expression of the domain influence factor is: , In the formula, represents the th IMF of the th band signal in the frequency domain influence factor at frequency represents the frequency domain feature in the composite feature vector, represents the frequency domain change adjustment factor, represents the frequency domain smooth transition adjustment factor, represents the instantaneous change rate of the frequency component, represents the frequency domain balance term, is the total number of frequency bands; A non-linear activation function connected to the convolutional layer, where the expression of the non-linear activation function is: , wherein, , , respectively represent different activation functions, , , respectively represent the adjustable parameters of the corresponding activation functions, controlling the contribution of each activation function, represents the improved non - linear activation function, represents the input signal; An attention mechanism connected to the non-linear activation function, and the attention mechanism includes a time-domain attention mechanism and a frequency-domain attention mechanism, and the expression of the time-domain attention mechanism is: , In the formula, represents the -th IMF of the -th band signal at the time point denotes the current time point, represents the total number of time steps, represents the number of time steps in the time window, represents the length of the time window, represents the -th IMF of the -th band signal at the time point after being processed by the activation function, denotes the cumulative sum of the time domain features of the past time points in the time window near the time point represents the adjustment parameter of the time domain attention weight, represents the time domain adjustment term parameter; The expression of the frequency-domain attention mechanism is: , In the formula, represents the -th frequency domain attention weight of the -th IMF of the -th band signal at the frequency point represents the adjustment factor of the frequency domain attention, represents the -th frequency domain feature of the -th IMF of the -th band signal after being processed by the activation function at the frequency point represents the cumulative sum of the frequency domain features of adjacent frequency points near the frequency represents the size of the sliding window, represents an index used to select adjacent frequency points near the current frequency point represents the frequency domain adjustment term parameter.
7. A hand gesture recognition system, characterized in that, Including: An acquisition module configured to acquire the strain signal of the hand sensor and perform EWT transformation on the strain signal according to the boundary adaptive backfill mechanism to obtain a target strain signal; A calculation module configured to decompose the target strain signal according to EEMD to obtain at least one IMF component, calculate the multi-scale non-linear relationship influence factor of the at least one IMF component, and select at least one target IMF component whose multi-scale non-linear relationship influence factor is greater than a preset threshold; A screening module configured to perform feature extraction on the at least one target IMF component, combine the extracted time-domain features into a time-domain feature vector, and combine the extracted frequency-domain features into a frequency-domain feature vector, and screen the time-domain feature vector and the frequency-domain feature vector according to a preset feature scoring function to obtain a target time-domain feature vector and a target frequency-domain feature vector; A fusion module configured to normalize the target time-domain feature vector and the target frequency-domain feature vector, and perform cross-weighted fusion on the normalized target time-domain feature vector and the normalized target frequency-domain feature vector to obtain a composite feature vector; An output module configured to input the composite feature vector into a preset improved CNN model, and the improved CNN model outputs a hand action recognition result corresponding to the strain signal.
8. An electronic device, characterized in that, Including: At least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Surface electromyography (SEMG)-based human hand interior action identification method
CN110604578A
Gesture recognition method based on variational mode decomposition and support vector machine
CN114220164A
Dynamic gesture recognition method based on hand key point and double-layer bidirectional LSTM network
CN117576783A
Partial discharge type identification method and system based on multi-feature extraction and fusion
CN118709095A
Bionic manipulator control method, system and equipment based on sEMG signal and storage medium
CN119175715A