Lower limb motion intention recognition method and system based on multi-source information fusion and deep learning optimization
By using multi-source information fusion and deep learning optimization, a lower limb movement intention recognition model is constructed using LSTM and CNN, which solves the problems of low recognition accuracy and poor adaptability in existing technologies, and achieves efficient and accurate recognition of lower limb movement intention.
Patent Information
- Application Number
- CN202510922415.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-21
AI Technical Summary
Existing methods for recognizing lower limb movement intentions suffer from insufficient signal processing and information fusion, as well as inadequate model construction and optimization, resulting in low recognition accuracy, poor adaptability, and difficulty in meeting real-time and portability requirements.
By employing a multi-source information fusion and deep learning optimization method, surface electromyography signals and posture information are collected, and an intention recognition model is constructed using LSTM combined with CNN. By deeply fusing time series features and spatial features, the accurate recognition of lower limb movement intentions is achieved.
It improves the accuracy and adaptability of lower limb movement intention recognition, enabling accurate recognition in complex movement patterns and meeting the application requirements of real-time performance and portability.
Smart Images

Figure CN120822002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of motion recognition technology, and more specifically to a method and system for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization. Background Art
[0002] Lower-limb movement intention recognition technology plays a crucial role in modern medical rehabilitation and intelligent assistive devices. With the aging of society and the increasing number of people suffering from lower-limb motor dysfunction due to illness and accidents, the demand for devices that can accurately identify lower-limb movement intentions and provide corresponding assistance, such as intelligent prostheses and lower-limb exoskeleton robots, is becoming increasingly urgent. These devices can help people with limited mobility regain their ability to move independently and improve their quality of life, showing broad application prospects in a variety of fields, including medical rehabilitation, military, and industry.
[0003] Currently, a variety of technologies have been applied to the recognition of lower limb movement intentions. For example, surface electromyography (sEMG) signals have become a common data source for driving lower limb exoskeleton robots to move synchronously with the wearer because they can reflect the state of muscle activity and have mature acquisition technology and rich information. By analyzing and processing surface electromyography signals and combining them with artificial intelligence technologies such as BP neural networks, the recognition of lower limb movement intentions can be achieved to a certain extent. Inertial measurement unit (IMU) signals are also often used to obtain human posture information. They have the characteristics of high accuracy, fast speed, good flexibility, and high-frequency response. In addition, plantar pressure sensors can collect plantar pressure data for analysis of information such as human gait phase.
[0004] However, existing methods for lower limb movement intention recognition still have many defects. On the one hand, in terms of signal processing and information fusion, existing technologies often rely only on a single signal source or a simple superposition of multiple signal sources, and are unable to fully tap the potential correlation information between different signals. For example, although surface electromyography signals can reflect muscle activity, they are easily interfered by noise, and signal feature extraction is complex; although IMU signals can provide posture information, they have the problem of poor stability; plantar pressure signals have limited information acquisition in certain situations, such as the swing phase. At the same time, the fusion methods between different signal sources are mostly relatively crude, failing to effectively integrate the advantages of each signal, resulting in low information utilization and difficulty in achieving high-precision movement intention recognition.
[0005] On the other hand, in terms of model construction and optimization, traditional recognition methods based on manual feature engineering rely heavily on the manual selection and extraction of features, resulting in high subjectivity and low efficiency. While models using deep learning techniques, such as common neural network models, have overcome the shortcomings of feature engineering to some extent, they lack generalization capabilities for complex lower limb movement patterns. Accuracy can drop by up to 20%-30% in cross-individual tests (e.g., muscle strength and movement habits) and in varying scenarios (e.g., flat ground versus climbing stairs), and feature extraction is time-consuming. They also exhibit poor adaptability to complex lower limb movement patterns, with accuracy fluctuating by up to 15%-25% across different movement scenarios, making them unable to effectively adapt to the variability across individuals and movement scenarios. Furthermore, existing models lack in-depth exploration and effective modeling of intrinsic relationships when processing multi-source information, resulting in poor accuracy and robustness. Furthermore, current recognition models generally suffer from long training times and high computational resource consumption, making them difficult to meet the real-time and portability requirements of practical applications.
[0006] Therefore, how to provide a lower limb movement intention recognition method and system based on multi-source information fusion and deep learning optimization to improve the recognition accuracy is an urgent problem that needs to be solved by technical personnel in this field. Summary of the Invention
[0007] In view of this, the present invention provides a method and system for lower limb movement intention recognition based on multi-source information fusion and deep learning optimization. By collecting the surface electromyographic signals and posture information of the lower limb muscles, and then using LSTM combined with CNN to construct an intention recognition model, the advantages of these two deep learning networks are fully utilized, and the time series characteristics of the surface electromyographic signals and the spatial characteristics of the posture information are deeply integrated, thereby realizing accurate recognition of lower limb movement intention.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] On the one hand, the present invention provides a method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization, comprising:
[0010] Collect surface electromyographic signals and posture information of lower limb muscles;
[0011] Preprocessing the surface electromyography signal and posture information respectively;
[0012] An intention recognition model was constructed using LSTM combined with CNN. The pre-processed surface electromyography signals and posture information were recognized using the trained intention recognition model to obtain the lower limb movement intention.
[0013] Preferably, preprocessing the surface electromyography signal includes:
[0014] Converting the surface electromyography signal into a time-frequency graph by wavelet transform;
[0015] Suppressing noise in the time-frequency graph by binarization;
[0016] The noise-suppressed time-frequency diagram is scaled using bicubic interpolation;
[0017] Normalize the scaled time-frequency diagram.
[0018] Preferably, preprocessing the posture information includes:
[0019] Align the posture data by timestamp;
[0020] Use compensation algorithm to fuse the aligned posture information and calculate the angle information;
[0021] performing low-pass filtering on the angle information;
[0022] The filtered angle information is normalized.
[0023] Preferably, the pre-processed surface electromyographic signals and posture information are identified using a trained intention recognition model to obtain lower limb movement intention, including:
[0024] Extracting the spatial feature vector of the pre-processed posture information based on the convolutional neural network (CNN) in the intention recognition model;
[0025] Extracting the time series feature vector of the preprocessed surface electromyography signal based on the long short-term memory network (LSTM) in the intention recognition model;
[0026] fusing the spatial feature vector and the time series feature vector in chronological order to obtain a spatiotemporal feature vector;
[0027] The spatiotemporal feature vector is mapped to the intention recognition task to obtain the lower limb intention recognition result.
[0028] Preferably, fusing the spatial feature vector and the time series feature vector in chronological order to obtain the spatiotemporal feature vector includes:
[0029] Aligning the spatial feature vector and the time series feature vector according to a timestamp;
[0030] In chronological order, the aligned spatial feature vector and time series feature vector are concatenated in series to form a spatiotemporal feature vector.
[0031] An attention mechanism is used to optimize the spatiotemporal feature vector.
[0032] Preferably, extracting the spatial feature vector of the preprocessed posture information based on the convolutional neural network (CNN) in the intention recognition model includes:
[0033] The preprocessed posture information is input into the convolutional neural network (CNN), and the local features are extracted using the convolution operation.
[0034] After the convolution operation, nonlinearity is introduced through the activation function to enhance the local features;
[0035] Perform pooling operation on the enhanced local features to reduce the dimension of the features;
[0036] The output of the last convolutional layer or pooling layer is flattened to form a one-dimensional feature vector, which is input into the fully connected layer to generate a spatial feature vector.
[0037] Preferably, extracting the time series feature vector of the preprocessed surface electromyography signal based on the long short-term memory network LSTM in the intention recognition model includes:
[0038] Extracting a time series corresponding to the spatial feature vector in each time step of the surface electromyography signal;
[0039] The extracted spatial feature vector is input into the long short-term memory network LSTM according to the time series to learn the evolution law and dependency relationship of lower limb movement with the time series, and obtain the time series feature vector corresponding to the spatial feature vector.
[0040] On the other hand, the present invention provides a lower limb movement intention recognition system based on multi-source information fusion and deep learning optimization, comprising:
[0041] Acquisition module, used to collect surface electromyographic signals and posture information of lower limb muscles;
[0042] A preprocessing module, used for preprocessing the surface electromyography signal and posture information respectively;
[0043] The intention recognition module is used to build an intention recognition model using LSTM combined with CNN, and use the trained intention recognition model to identify the preprocessed surface electromyography signals and posture information to obtain the lower limb movement intention.
[0044] The above technical solution shows that compared with the prior art, the present invention provides a method and system for lower limb movement intention recognition based on multi-source information fusion and deep learning optimization. By fusing two multi-source information sources, surface electromyography (EMG) signals and posture information, the system more comprehensively captures lower limb movement characteristics. Surface electromyography (EMG) signals reflect muscle electrophysiological activity, while posture information provides the spatial position and movement state of the lower limbs. The combination of the two avoids the limitations of a single information source. The LSTM is combined with the CNN to construct an intention recognition model, giving full play to the advantages of both. The CNN extracts the spatial features of posture information, while the LSTM processes the surface electromyography signal time series data, learning the temporal dependency and dynamic evolution laws. The deep fusion of spatiotemporal features enables accurate recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0046] Figure 1 A schematic diagram of the process provided by the present invention.
[0047] Figure 2 Flowchart for lower limb movement intention recognition based on multi-source information fusion.
[0048] Figure 3 This is the structure diagram of intent recognition combining CNN and LSTM.
[0049] Figure 4 This is a schematic diagram of the system structure provided by the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0051] The embodiment of the present invention discloses a method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization, such as Figure 1-2 Shown, including:
[0052] Collect surface electromyographic (SEM) signals and posture information from lower limb muscles. Surface EMG signals reflect the electrophysiological activity of lower limb muscles and are a direct reflection of muscle contraction and movement intent. Posture information provides the spatial position and motion state of the lower limb, helping to fully capture the movement characteristics of the lower limb. To ensure data accuracy and reliability, SEM signals can be collected using surface or implanted electrodes, while posture information can be acquired using devices such as inertial measurement units (posture information).
[0053] In this embodiment, during exercise, a three-channel surface electromyography signal acquisition device with a frequency of 1000Hz is used to synchronously collect the tibialis anterior and other three leg muscles. At the same time, two nine-axis posture information sensors with a frequency of 200Hz are used as angle sensors to collect the movement angle information of the knee joint and ankle joint respectively. The collected data is transmitted to the computer via Bluetooth, reducing the signal interference of the line during the collection process. Before the experiment, 75% medical alcohol is used to clean the skin to reduce impedance. During the experiment, attention should be paid to the simple environment of the collection location to avoid interference from other equipment and lines.
[0054] Preprocess the surface electromyography signal and posture information respectively;
[0055] An intent recognition model is constructed using LSTM combined with CNN. The trained model is then used to identify preprocessed surface EMG signals and posture information to determine lower-limb movement intentions. CNN effectively extracts spatial features from posture information, capturing local spatial patterns and structures. LSTM excels at processing time series data and can learn the temporal evolution patterns and dependencies of surface EMG signals. By combining these two networks, the time series features of surface EMG signals and the spatial features of posture information can be deeply integrated, enabling accurate recognition of lower-limb movement intentions.
[0056] In practical applications, when a user uses the lower limb movement intention recognition system based on the present invention to walk up or down stairs, surface electromyographic (SEM) signals and posture information of the lower limb muscles are first acquired. For example, when walking up stairs, the contraction pattern and strength of the leg muscles differ significantly from walking on level ground. The muscle activation sequence, changes in muscle contraction intensity, and the frequency and time domain characteristics of the SEM signals reflect characteristics similar to walking on level ground. Furthermore, the angle changes of the knee and ankle joints in the posture information are more frequent and complex.
[0057] By decomposing the time-frequency graph of the surface electromyographic signal during stair walking through wavelet transform, the high-frequency characteristics (100-300 Hz) of the tibialis anterior muscle during the buffering phase can be identified, while the posture information simultaneously captures the sudden change in angular velocity of the ankle dorsiflexion (>100° / s).
[0058] When climbing stairs, CNN extracts the spatial features of the knee joint angle from 0° to 60°, and LSTM captures the rising edge timing of the rectus femoris electromyographic signal within 0-200ms. After fusing the two, it can accurately distinguish between "climbing stairs" and "walking fast on flat ground."
[0059] Since the peak value of the surface electromyographic signal and the sudden change of the joint angle of the posture information at the moment of starting on the stairs contribute most to intention recognition, the attention mechanism automatically assigns a higher weight to the features of this stage (weight coefficient α = 0.25).
[0060] After the fusion of spatiotemporal features, the model can capture the rule that "the electromyographic signal lags behind the posture change" when going downstairs (for example, the ankle joint angle reaches 30° first, and the electromyographic signal of the tibialis anterior muscle peaks 50ms later), avoiding misjudgment caused by timing misalignment in traditional models.
[0061] Traditional intention recognition models typically rely on a single signal source, such as surface electromyography (EMG) or posture information, making it difficult to fully capture the multi-dimensional characteristics of complex motion patterns such as stair climbing. This new approach, by fusing multi-source information and leveraging the advantages of surface electromyography (EMG) in reflecting muscle activity and posture information in reflecting the spatial position and motion of the lower limbs, significantly improves the recognition accuracy of complex movements such as stair climbing.
[0062] Furthermore, the surface electromyography signal is preprocessed including:
[0063] The surface electromyography signal is converted into a time-frequency graph through wavelet transform; wavelet transform is a signal processing tool that can convert the signal from the time domain to the time-frequency domain, thereby retaining the time and frequency information of the signal at the same time. This is particularly important for analyzing non-stationary signals (such as surface electromyography signals) because it can reveal the frequency characteristics of the signal at different time points and provide richer information for subsequent feature extraction. Specifically, the wavelet transform decomposes the surface electromyography signal into coefficients in different frequency bands by selecting appropriate wavelet basis functions and decomposition scales, and then generates a time-frequency graph. The time-frequency graph shows the time-frequency distribution characteristics of the surface electromyography signal in the form of a two-dimensional image, providing a more intuitive and comprehensive signal representation for subsequent processing. The original surface electromyography signal s(t) is converted into a time-frequency graph through the wavelet basis function ψ:
[0064]
[0065] Among them, a is the scale factor, b is the translation factor, ψ * is the complex conjugate of the wavelet basis function.
[0066] The noise of the time-frequency graph is suppressed by binarization; noise is one of the important factors affecting signal quality, especially in complex surface electromyography signal acquisition environments, where various high-frequency or low-frequency noises may be mixed in. Binarization converts the grayscale image into a binary image. By setting an appropriate threshold, points with pixel values greater than the threshold are set as the foreground (usually white, with a pixel value of 255), and points less than or equal to the threshold are set as the background (usually black, with a pixel value of 0). Through repeated experiments and optimization, the binarization threshold for the surface electromyography signal time-frequency graph is determined, thereby effectively suppressing noise interference in the time-frequency graph, highlighting the main features of the signal, and laying a good foundation for subsequent feature extraction and analysis. Noise is suppressed by the threshold T = μ + kσ:
[0067]
[0068] The noise-suppressed time-frequency graph is scaled using bicubic interpolation. Bicubic interpolation calculates the value of the new pixel by considering the values of the surrounding 16 pixels (i.e., the pixels within a 3×3 neighborhood), thereby achieving smooth scaling of the image. Compared with simple methods such as nearest neighbor interpolation, bicubic interpolation can better preserve image details and edge information and reduce distortion during scaling. In practice, the scaling ratio is determined based on the specific requirements of the deep learning model for the input image size, and the bicubic interpolation algorithm is used to adjust the time-frequency graph to a uniform size to ensure data consistency and standardization, while also providing higher-quality input data for model training and recognition. The target pixel (x, y) is obtained by weighting the 4×4 pixels in the neighborhood:
[0069]
[0070] Among them, W is the bicubic kernel function, s x , s y is the scaling factor.
[0071] The scaled time-frequency graph is normalized. The purpose of normalization is to map the value range of the data to a specific interval (usually [0,1] or [-1,1]) to improve the stability and convergence speed of the model training. In an embodiment of the present invention, a linear normalization method is adopted, and the value of each pixel point in the time-frequency graph is subtracted from the minimum value of the data set, and then divided by the range of the data set (i.e., the difference between the maximum and minimum values), so as to normalize the pixel value to the [0,1] interval. The normalized data can not only eliminate the interference caused by the amplitude difference between different surface electromyography signals, but also enable the model to learn the inherent laws of the data more quickly during the training process, thereby improving the generalization ability and recognition performance of the model.
[0072] Furthermore, the posture information is preprocessed, including:
[0073] The posture data is aligned by timestamps; due to differences in sampling frequency and data transmission delay of different sensors, the collected posture data may not be synchronized in time. Therefore, it is necessary to use the timestamps added to the data by each sensor to perform alignment operations. First, determine the time series of the main reference sensor (such as the inertial measurement unit with the highest sampling frequency and the best stability) as the benchmark. Then, for the posture data collected by other sensors, use linear interpolation or spline interpolation algorithms to map the data to the benchmark time series according to the timestamp. For example, if the data collected by a certain sensor is missing between time points t1 and t2, the intermediate data points in the time period can be calculated by linear interpolation to ensure that all posture data correspond one to one in the time dimension and achieve precise alignment. Specifically, the posture data P(t) is aligned to the unified time series t by linear interpolation. k :
[0074]
[0075] Among them, t i ≤t k ≤t i+1 .
[0076] The compensation algorithm is used to fuse the aligned posture information to calculate the angle information; the aligned posture data may still have inaccuracies due to factors such as sensor errors (such as zero bias, drift), installation position deviation, etc. For this reason, the compensation algorithm is used to process the data. Taking the inertial measurement unit as an example, it contains an accelerometer, a gyroscope, and a magnetometer. The complementary filtering or Kalman filtering algorithm can be used to fuse the data of these three sensors. Complementary filtering combines the stable posture information of the accelerometer in the low frequency band with the fast response information of the gyroscope in the high frequency band based on the performance advantages of the sensors at different frequencies; Kalman filtering establishes a state space model, uses the dynamic model and measurement model of the system, and takes the minimum mean square error as the criterion to make the best estimate of the posture data, converting the original posture data into accurate angle information, such as the rotation angle of the joint, the tilt angle of the limb, etc., to provide reliable data for subsequent analysis. The angle θ is calculated by fusing the accelerometer c and gyroscope ω data:
[0077] θ(t)=α·(θ(t-1)+ω(t)·Δt)+(1-α)·arctan 2(c y ,c x )
[0078] Among them, α is the fusion coefficient and Δt is the sampling time interval.
[0079] The angle information is low-pass filtered; the calculated angle information is often mixed with high-frequency noise, which may be caused by factors such as sensor measurement errors and environmental interference, and will affect subsequent feature extraction and analysis. Therefore, low-pass filtering is used to process the angle information. Commonly used low-pass filtering methods include Butterworth low-pass filter, Chebyshev low-pass filter, etc. In an embodiment of the present invention, a Butterworth low-pass filter is used to determine the cutoff frequency and order of the filter according to the actual application scenario. If the posture information mainly reflects the low-frequency movement of the lower limbs of the human body (such as walking, standing, etc., the frequency is generally 0.5-5Hz), the cutoff frequency can be set to 5Hz, and the angle information is processed by the filter to retain the real angle change information of the low frequency and filter out the high-frequency noise, so that the angle information is smoother and more stable.
[0080] Normalize the filtered angle information. There are differences in limb size and joint range of motion among different individuals, and the measurement range of the sensor may also be different, which will lead to inconsistent numerical ranges of the angle information. In order to eliminate the impact of these differences on subsequent model training and recognition results, the filtered angle information needs to be normalized. Common normalization methods include minimum-maximum normalization and Z-score normalization. This embodiment adopts Z-score normalization, and the data is standardized to a distribution with a mean of 0 and a standard deviation of 1 according to the formula x′=(x-μ) / σ, where μ is the mean of the angle information and σ is the standard deviation. Through normalization, the angle information has a unified scale, which improves the training efficiency and recognition accuracy of the model.
[0081] Furthermore, the trained intention recognition model is used to identify the preprocessed surface electromyographic signals and posture information to obtain the lower limb movement intention, including:
[0082] The convolutional neural network (CNN) in the intention recognition model extracts the spatial feature vector of the pre-processed posture information;
[0083] The time series feature vector of the pre-processed surface electromyography signal is extracted based on the long short-term memory network (LSTM) in the intention recognition model;
[0084] Fusion of spatial feature vectors and time series feature vectors in chronological order to obtain spatiotemporal feature vectors;
[0085] The spatiotemporal feature vector is mapped to the intention recognition task to obtain the lower limb intention recognition result. Specifically, the spatiotemporal feature vector is input into the fully connected layer. After nonlinear transformation by the activation function, the probability values corresponding to each type of movement intention are obtained. The movement intention corresponding to the highest probability is the recognition result. Simultaneously, the softmax classifier maps the feature vector to multiple movement intention categories and outputs a probability distribution, clearly defining the category corresponding to the highest probability as the final lower limb movement intention.
[0086] Furthermore, Figure 3 As shown in Figure 2, the spatial feature vector and the time series feature vector are fused in time order to obtain the spatiotemporal feature vector, including:
[0087] The spatial and time series feature vectors are aligned based on their timestamps. During the feature extraction phase, the CNN and LSTM networks process the timestamped data separately, and the output spatial and time series feature vectors also retain the timestamp information. By comparing and adjusting the timestamps of the feature vectors, feature vectors from different modalities are mapped to the same time frame, ensuring that the extracted spatial and time series feature vectors accurately reflect the lower limb movement state at the same moment.
[0088] In chronological order, the aligned spatial feature vector and time series feature vector are concatenated in series to form a spatiotemporal feature vector.
[0089] The attention mechanism is used to optimize the spatiotemporal feature vector. The attention mechanism improves the important features of lower limb movement intention recognition by learning the importance weights of different parts of the feature vector. Specifically, by collecting surface electromyographic signals and posture signals from different individuals and different scenarios, similar historical spatiotemporal feature vectors are obtained through clustering. The similar historical spatiotemporal feature vectors, spatial feature vectors, and time series features are used as inputs to the spatiotemporal attention module. In the spatiotemporal attention module, the similar historical spatiotemporal feature vectors and spatiotemporal feature vectors are processed by the gated residual network to obtain similar historical spatiotemporal features and spatiotemporal features respectively. The similar historical spatiotemporal features are activated by the ELU function and multiplied element-wise with the spatial features, and then passed through the multi-layer perceptron to obtain the spatial vector. The spatial vector is multiplied and summed with each component of the spatial feature to obtain the key vector and value vector of the spatial attention mechanism. The spatiotemporal feature vector is passed through the gated residual network to obtain the spatiotemporal feature. The spatiotemporal feature is used as the query vector, and the spatial attention weight vector is obtained based on the key vector and value vector. The spatial attention weight vector is sequentially passed through the gated linear unit, layer normalization, and fully connected layer to obtain the spatial attention feature vector.
[0090] Similarly, the temporal attention feature vector is obtained, the spatial attention feature vector is added to the temporal attention feature vector, and then normalized through the gated linear unit layer and the fully connected layer in sequence to obtain the spatiotemporal features.
[0091] Furthermore, the convolutional neural network (CNN) in the intent recognition model extracts the spatial feature vectors of the preprocessed posture information, including:
[0092] The preprocessed posture information is input into a convolutional neural network (CNN), where local features are extracted using convolution operations. In this embodiment, the convolution kernels are designed to capture spatial patterns in the posture information, such as changes in joint angles and the direction of limb movement. Each convolution kernel is responsible for detecting a specific feature, and multiple convolution kernels work together to extract a variety of local features from the posture information. These local features are crucial for understanding the movement patterns of the lower limbs because they can reflect the position and movement trend of the limbs in space.
[0093] After the convolution operation, nonlinearity is introduced through the activation function to enhance the local features. Common activation functions include ReLU (Rectified Linear Unit), Sigmoid, and Tanh. In this embodiment, the ReLU activation function is used to set all negative values to zero while keeping the positive values unchanged, so that the network can better capture the complex patterns in the posture information. The feature map processed by the activation function not only retains the original local feature information, but also highlights those features that are more critical to motion intention recognition. The calculation of the l-th layer convolution feature map Fl is:
[0094]
[0095] Among them, Wl is the convolution kernel and σ is the activation function.
[0096] The enhanced local features are pooled to reduce the dimension of the features; the pooling operation is a downsampling technique that aims to reduce the size of the feature map while retaining the most important feature information. Commonly used pooling methods include maximum pooling and average pooling. In this embodiment, maximum pooling is used to slide a pooling window on the feature map and take the maximum value in the window as the output, thereby reducing the spatial resolution of the feature map. The pooling operation not only reduces the amount of calculation and the number of parameters, but also improves the model's translation invariance to the input data, making the model less sensitive to small position changes in the posture information, thereby improving the model's generalization ability. Using maximum pooling to reduce feature dimensions, F pool (i,j)=max m,n∈R F l (i·s+m,j·s+n), where R is the pooling area and s is the stride.
[0097] The output of the last convolutional or pooling layer is flattened to form a one-dimensional feature vector, which is then input into the fully connected layer to generate a spatial feature vector. Flattening is the process of converting a multidimensional feature map into a one-dimensional vector. This is done to integrate the local features extracted by the convolutional layer into a global feature representation. The flattened feature vector contains the key spatial features of the posture information, and these feature vectors are then input into the fully connected layer. Each neuron in the fully connected layer is connected to all neurons in the previous layer. It performs linear transformation and nonlinear activation on the input feature vector, further integrating and abstracting the feature information, and ultimately generating a fixed-length spatial feature vector. This spatial feature vector condenses the key features of the posture information and can effectively characterize the motion state and spatial position of the lower limbs, providing a high-quality feature representation for subsequent motion intention recognition.
[0098] Specifically, the time series feature vector of the preprocessed surface electromyography signal is extracted based on the long short-term memory network (LSTM) in the intention recognition model, including:
[0099] Extract the time series corresponding to the spatial feature vector in each time step of the surface electromyography signal;
[0100] The extracted spatial feature vectors are input into the long short-term memory network (LSTM) according to the time series to learn the evolution law and dependency of lower limb movement along the time series, and obtain the time series feature vector corresponding to the spatial feature vector.
[0101] In an LSTM, the input at each time step passes through a series of gating mechanisms (input gate, forget gate, and output gate) to control the flow of information. The input gate determines the extent to which new information should be stored in the cell state; the forget gate determines the extent to which old information should be forgotten; and the output gate determines which information in the cell state should be output. These gating mechanisms enable the LSTM to flexibly adjust the retention and updating of information, effectively capturing long-term dependencies in surface EMG signals.
[0102] The input gate is fed by the input x t and the hidden state h at the previous moment t-1 Determine, the calculation formula is:
[0103] i t =σ(W i [h t-1 ,x t ]+b i )
[0104] Among them, σ is the Sigmoid activation function, W i is the weight matrix of the input gate, b i is the bias term.
[0105] The calculation formula of the forget gate is:
[0106] f t =σ(W f [h t-1 ,x t ]+b f )
[0107] Among them, W f is the weight matrix of the forget gate, b f is the bias term.
[0108] The calculation formula of the output gate is:
[0109] o t =σ(W o [h t-1 ,x t ]+b o )
[0110] Among them, W o is the weight matrix of the output gate, b o is the bias term.
[0111] The aligned surface electromyography signal segments E s (i) As input sequence X = {E s (1),E s (2),…,E s (n)} is input into the LSTM network, where n represents the total number of time steps.
[0112] Within the LSTM, the input feature vector at each time step is combined with the hidden state from the previous time step. Through nonlinear transformations and gating mechanisms, the cell state and hidden state for the current time step are updated. The hidden state is then passed to the next time step, enabling the flow and accumulation of information over time. By processing multiple time steps, the LSTM learns the temporal evolution of surface EMG signals and the dependencies between different time steps.
[0113] Cell state c t The update formula is:
[0114] c t =f t c t-1 +i t tanh(W c [h t-1 ,x t ]+b c )
[0115] Among them, tanh is the hyperbolic tangent activation function, W c is the weight matrix for unit state update, bc is the bias term.
[0116] Hidden state h t The update formula is:
[0117] h t =o t tanh(c t )
[0118] Through the above updating process, the LSTM network can learn the dependencies between lower limb movements at different time steps.
[0119] Finally, the time series feature vector output by the LSTM represents a high-level summary of the temporal characteristics of the surface EMG signal. This feature vector not only captures the local features of the surface EMG signal at each time step but also reflects the temporal trends and interrelationships of these features. In this way, the LSTM effectively captures the dynamic characteristics of lower-limb movement, providing a critical time series feature vector for subsequent intention recognition. Together with the previously extracted spatial feature vector, this time series feature vector forms the basis of the spatiotemporal feature vector, providing a comprehensive feature representation for accurate recognition of lower-limb movement intentions.
[0120] On the other hand, the present invention provides a lower limb movement intention recognition system based on multi-source information fusion and deep learning optimization, such as Figure 4 Shown, including:
[0121] Acquisition module, used to collect surface electromyographic signals and posture information of lower limb muscles;
[0122] A preprocessing module, used to preprocess the surface electromyography signal and posture information respectively;
[0123] The intention recognition module is used to build an intention recognition model using LSTM combined with CNN, and use the trained intention recognition model to identify the preprocessed surface electromyography signals and posture information to obtain the lower limb movement intention.
[0124] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0125] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for lower limb movement intention recognition based on multi-source information fusion and deep learning optimization, characterized in that: include: Collect surface electromyographic signals and posture information of lower limb muscles; Preprocessing the surface electromyography signal and posture information respectively; An intent recognition model is constructed based on LSTM combined with CNN. CNN is used to extract the spatial features of the posture information, and LSTM is used to extract the time series features of the surface electromyography signal. The two are fused into a spatiotemporal feature vector. The feature weights are optimized using the attention mechanism to obtain an optimized intent recognition model. The optimized intention recognition model is used to identify the preprocessed surface electromyography signals and posture information to obtain the lower limb movement intention.
2. The method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization according to claim 1 is characterized in that: Preprocessing the surface electromyography signal includes: Converting the surface electromyography signal into a time-frequency graph by wavelet transform; Suppressing noise in the time-frequency graph by binarization; The noise-suppressed time-frequency diagram is scaled using bicubic interpolation; Normalize the scaled time-frequency diagram.
3. The method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization according to claim 1 is characterized in that: Preprocessing the posture information includes: Align the posture data by timestamp; Use compensation algorithm to fuse the aligned posture information and calculate the angle information; performing low-pass filtering on the angle information; The filtered angle information is normalized.
4. The method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization according to claim 1, characterized in that: The trained intention recognition model is used to identify the pre-processed surface electromyographic signals and posture information to obtain the lower limb movement intention, including: Extracting the spatial feature vector of the pre-processed posture information based on the convolutional neural network (CNN) in the intention recognition model; Extracting the time series feature vector of the preprocessed surface electromyography signal based on the long short-term memory network (LSTM) in the intention recognition model; fusing the spatial feature vector and the time series feature vector in chronological order to obtain a spatiotemporal feature vector; The spatiotemporal feature vector is mapped to the intention recognition task to obtain the lower limb intention recognition result.
5. The method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization according to claim 4 is characterized in that: The spatial feature vector and the time series feature vector are fused in time order to obtain a spatiotemporal feature vector, including: Aligning the spatial feature vector and the time series feature vector according to a timestamp; In chronological order, the aligned spatial feature vector and time series feature vector are concatenated in series to form a spatiotemporal feature vector. An attention mechanism is used to optimize the spatiotemporal feature vector.
6. The method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization according to claim 4, characterized in that: Extracting the spatial feature vector of the pre-processed posture information based on the convolutional neural network (CNN) in the intention recognition model includes: The preprocessed posture information is input into the convolutional neural network (CNN), and the local features are extracted using the convolution operation. After the convolution operation, nonlinearity is introduced through the activation function to enhance the local features; Perform pooling operation on the enhanced local features to reduce the dimension of the features; The output of the last convolutional layer or pooling layer is flattened to form a one-dimensional feature vector, which is input into the fully connected layer to generate a spatial feature vector.
7. The method for recognizing lower limb movement intention based on multi-source information fusion and deep learning optimization according to claim 4, characterized in that: The time series feature vector of the preprocessed surface electromyography signal is extracted based on the long short-term memory network (LSTM) in the intention recognition model, including: Extracting a time series corresponding to the spatial feature vector in each time step of the surface electromyography signal; The extracted spatial feature vector is input into the long short-term memory network LSTM according to the time series to learn the evolution law and dependency relationship of lower limb movement with the time series, and obtain the time series feature vector corresponding to the spatial feature vector.
8. A lower limb movement intention recognition system based on multi-source information fusion and deep learning optimization, characterized by: include: Acquisition module, used to collect surface electromyographic signals and posture information of lower limb muscles; A preprocessing module, used for preprocessing the surface electromyography signal and posture information respectively; The intention recognition module is used to construct an intention recognition model based on LSTM combined with CNN, use CNN to extract the spatial features of the posture information, use LSTM to extract the time series features of the surface electromyography signal, and fuse them into a spatiotemporal feature vector. The attention mechanism is used to optimize the feature weights to obtain an optimized intention recognition model, and the optimized intention recognition model is used to recognize the preprocessed surface electromyography signal and posture information to obtain the lower limb movement intention.