Mechanical arm control method based on electroencephalogram function signal and artificial neural network

By combining a cascaded sparse autoencoder neural network and a Softmax classifier, the spatiotemporal features of EEG signals are automatically extracted, solving the problem of step-by-step optimization of feature extraction and classifier training in existing technologies, and realizing efficient and robust EEG signal decoding and robotic arm control.

CN121870745APending Publication Date: 2026-04-17ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ACADEMY OF MILITARY MEDICAL SCIENCES
Filing Date
2025-12-29
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing EEG signal decoding methods, feature extraction and classifier training are optimized step by step and independently. This means that the extracted features may not be the most favorable for the final classification decision. Furthermore, they rely on manual feature engineering, which is cumbersome and cannot fully capture complex and subtle nonlinear patterns, especially information related to fine motor intentions.

Method used

A cascaded sparse autoencoder neural network is used to extract and fuse spatiotemporal features of EEG signals, including a temporal sparse autoencoder and a spatial sparse autoencoder. Robust spatiotemporal features are automatically learned and decoded using a Softmax classifier to generate robotic arm control commands.

Benefits of technology

It enables automatic, layer-by-layer learning to transform raw EEG signals into high-level abstract features, improving decoding performance, enhancing cross-individual adaptability, supporting rapid calibration and instant control feedback, and forming stable and natural human-machine integrated control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121870745A_ABST
    Figure CN121870745A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of brain-computer interfaces and neural engineering, in particular to a mechanical arm control method based on electroencephalogram function signals and an artificial neural network. The method comprises the following steps: collecting a multi-channel electroencephalogram signal during motor imagery of a user; performing step-by-step spatial-temporal feature extraction and fusion on the signal by using a pre-trained cascade sparse self-encoding neural network to obtain a feature code; a motion intention is decoded through a Softmax classifier; and a control instruction is generated according to the decoding result to drive the mechanical arm to execute actions. Wherein the neural network training adopts sparsity constraint based on KL divergence to learn robust features; the system is also integrated with a graphical user interface, provides real-time visual feedback and can adaptively adjust task difficulty according to user performance to form closed-loop interaction. According to the method, high-precision and self-adaptive decoding of electroencephalogram signals and natural and smooth mechanical arm control are achieved, and the method is suitable for rehabilitation training and man-machine cooperative operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of brain-computer interface and neural engineering technology, specifically to a robotic arm control method based on electroencephalogram (EEG) functional signals and artificial neural networks. Background Technology

[0002] Brain-computer interface (BCI) technology establishes a direct communication pathway between the brain and external devices, showing great promise in fields such as neurorehabilitation, assisted control, and human-computer interaction. Among these, non-invasive BCI based on scalp EEG signals has become a mainstream research direction due to its high safety and good temporal resolution. A typical EEG decoding process includes signal preprocessing, feature extraction, and classification. Feature extraction usually relies on manually designed algorithms to extract specific features from the time, frequency, or spatial domains, such as event-related desynchronization / synchronization power and common spatial patterns. These features are then fed into a classifier for intent recognition.

[0003] However, existing technical solutions have significant limitations: in traditional methods, feature extraction and classifier training are optimized step-by-step and independently, and their objective functions are not entirely consistent. This means that the extracted features may not be the most favorable for the final classification decision, thus limiting the overall decoding performance of the system. Manual feature engineering relies heavily on researchers' prior knowledge and experience, is tedious and time-consuming, and the extracted features may not be able to fully capture the complex and subtle nonlinear patterns in EEG signals, especially information related to fine motor intentions.

[0004] Therefore, there is an urgent need in this field for a high-performance brain-controlled robotic arm method and system that can automatically learn robust spatiotemporal features in EEG signals, reduce reliance on artificial feature engineering, have good cross-individual adaptability, and be deeply integrated with immersive feedback environments. Summary of the Invention

[0005] In view of this, embodiments of this application aim to provide a robotic arm control method based on electroencephalogram (EEG) functional signals and artificial neural networks.

[0006] The first aspect of this application provides a robotic arm control method based on brain functional signals and artificial neural networks, comprising the following steps:

[0007] S1: Collect multi-channel EEG signals from users when performing specific motor imagery tasks; S3: The multi-channel EEG signal is input into a pre-trained cascaded sparse autoencoder neural network for decoding; the cascaded sparse autoencoder neural network includes at least a temporal sparse autoencoder and a spatial sparse autoencoder connected in sequence, used to extract and fuse the spatiotemporal features of the EEG signal step by step, and output feature codes; S3: Input the feature encoding into the classifier to obtain the decoding result of the user's motion intention; S4: Generate control commands based on the decoding results and send them to the robotic arm control system to drive the robotic arm to perform corresponding actions.

[0008] Optionally, before step S3, a step of training the cascaded sparse autoencoder neural network is further included, the training step comprising: S31: Obtain the training dataset, which includes multi-channel EEG signals from multiple trials and their corresponding motor imagery category labels; S32: Train the temporal sparse autoencoder, whose loss function includes a reconstruction error term between the input signal and the reconstructed signal, a regularization term for the network weights, and a sparsity penalty term for the activation values ​​of the hidden layer neurons. S33: Encode the EEG signals of the training dataset using a trained temporal sparse autoencoder to obtain temporal features; S34: Train the spatial domain sparse autoencoder using the temporal features; S35: Encode the temporal features using a trained spatial domain sparse autoencoder to obtain the final feature code for classification.

[0009] Optionally, in step S31, for each motor imagery category, the average response of the EEG signals of all training trials under that category is calculated to form an average response matrix; the training input data of the time-domain sparse autoencoder is composed of the average response matrices of different categories.

[0010] Optionally, the sparsity penalty term is constructed based on Kullback-Leibler divergence to constrain the average activation value of hidden layer neurons to be close to a preset small sparsity parameter ρ.

[0011] Optionally, in step S3, the classifier is a Softmax classifier, whose output is the probability that the decoding result belongs to each motion image category, and the category corresponding to the maximum probability is taken as the final decoding result.

[0012] Optionally, in step S1, while collecting EEG signals, the user is provided with visual feedback of virtual reality scenes or robotic arm movements corresponding to the motor imagery task through a graphical user interface to form closed-loop training or control.

[0013] Optionally, the graphical user interface can adaptively adjust the difficulty of the motor imagery task based on the user's EEG signal decoding results or task completion status.

[0014] In a second aspect, the present invention provides a robotic arm control system based on brain functional signals and artificial neural networks for implementing the method described in the first aspect, comprising: The signal acquisition module is used to acquire the user's multi-channel EEG signals; The signal processing and decoding module integrates the pre-trained cascaded sparse autoencoder neural network and classifier, which is used to receive the EEG signals, decode them, and output the user's movement intention. The control command generation module is used to generate control commands for the robotic arm based on the user's movement intention; The robotic arm drive module is used to receive the control commands and drive the robotic arm to perform actions.

[0015] Optionally, it also includes: A feedback presentation module is used to provide a graphical user interface to the user, the graphical user interface being used at least to present a virtual reality scene or the state of a robotic arm to provide visual feedback.

[0016] Optionally, the system further includes a model training module for performing the training process of the cascaded sparse autoencoder neural network as described above.

[0017] Compared with existing technologies, the technical solution of this invention achieves automatic, layer-by-layer learning from raw EEG signals to high-level abstract features through the design of cascaded sparse autoencoders. By minimizing a loss function containing sparse constraints, the network can autonomously discover the most effective combination of spatiotemporal features for decoding motion intentions, completely avoiding the subjectivity and limitations of manual feature engineering, thereby achieving superior decoding performance.

[0018] Furthermore, the training strategy proposed in this invention, particularly the use of multi-subject data (group level) for pre-training or fine-tuning of the network, enables the model to learn common neural representations related to motor imagery that transcend individual specificity. This greatly enhances the model's adaptability to new users, making it possible to achieve practical BCI systems with rapid calibration or even "zero training." By deeply integrating a high-performance decoder with an immersive virtual reality feedback environment, this invention provides users with intuitive and immediate control feedback. This not only accelerates the process of users learning to regulate their own EEG signals but also maintains user engagement and training effectiveness through an adaptive difficulty adjustment mechanism, ultimately achieving more stable and natural human-computer fusion control. This invention covers a complete technology chain from signal acquisition, intelligent decoding, visual feedback to physical device control, forming a mature system solution applicable to rehabilitation training, assisted operation, and other scenarios, possessing high practical value and promising prospects for promotion. Attached Figure Description

[0019] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 This is a flowchart illustrating a robotic arm control method based on brain functional signals and artificial neural networks provided in one embodiment of this application. Figure 2 This is a deep learning architecture provided in one embodiment of this application; Figure 3 This is a partial flowchart of a robotic arm control based on brain functional signals and artificial neural networks provided in one embodiment of this application; Figure 4 This is a schematic diagram of the structure of a robotic arm control system based on brain functional signals and artificial neural networks provided in one embodiment of this application.

[0021] Figure 5 This is a schematic diagram of an electronic device structure provided in one embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Before introducing the method provided in this application, the overall architecture of the platform provided in this application will first be explained: The overall architecture of the platform provided in this application is still divided into three main parts: a virtual reality system, an EEG acquisition module, and an EEG analysis module. The EEG acquisition module is used to acquire EEG signals; the EEG analysis module is used to analyze EEG signals; the virtual reality system includes a virtual system and / or a real system; it is used to respond to the user's EEG signals; for example, the virtual reality system may include a display screen and a robotic arm; the display screen is equipped with multiple buttons and a movable dot; when a button is covered by the dot, it controls the robotic arm to perform a preset operation; the user controls the movable dot to move through EEG, and when the dot moves to a button, the robotic arm performs the operation corresponding to that button.

[0024] Reference Figure 1 The robotic arm control method based on brain functional signals and artificial neural networks provided in this application includes the following steps: S1: Collect multi-channel EEG signals from users when performing specific motor imagery tasks; This step marks the starting point for the interaction and data acquisition between the invention and the user. In practice, the system guides the user into a structured control or training task. First, a clear task prompt is presented to the user through a graphical user interface (GUI) or virtual reality (VR) environment. For example, displaying an arrow pointing to the left on the screen, or highlighting several buttons with different functions, constitutes a specific motor imagery task: "Imagine a ball moving downwards to control a robotic arm to move to the left and grasp." Simultaneously with or immediately after the task prompt, the system begins recording brainwave activity using a multi-channel EEG acquisition device worn on the user's head (such as a 64-lead EEG cap following the international 10-20 system setup). Following the prompts, the user actively and continuously imagines the movement of a ball on a pre-set display screen. The acquisition module synchronously records multi-channel raw EEG signals from the brain's motor sensory cortex and related areas at a high sampling rate (e.g., 1000 Hz). These signals contain characteristic neural oscillation patterns related to motor preparation and execution, such as event-correlated desynchronization of sensorimotor rhythms. The acquisition process typically lasts several seconds (e.g., the visualization phase lasts 2-4 seconds), forming a complete "trial". The system performs multiple trials consecutively to accumulate sufficient training or real-time control data. To ensure signal quality, the acquisition process is usually conducted in an electromagnetically shielded room, and signal impedance and noise levels can be monitored in real time via software.

[0025] S2: The multi-channel EEG signal is input into a pre-trained cascaded sparse autoencoder neural network for decoding; the cascaded sparse autoencoder neural network includes at least a temporal sparse autoencoder and a spatial sparse autoencoder connected in sequence, used to extract and fuse the spatiotemporal features of the EEG signal step by step, and output feature codes; This step is the core of the invention for achieving high-precision decoding. After necessary preprocessing (such as bandpass filtering to remove high-frequency electromyography and low-frequency drift, noise reduction, etc.), the acquired raw multi-channel EEG signals are fed into a cascaded sparse autoencoder neural network that is pre-trained offline or continuously optimized online.

[0026] This network is a deep learning architecture specifically designed for the spatiotemporal characteristics of EEG signals. For example... Figure 2As shown, it employs a cascaded (sequentially connected) structure, with the first stage being a temporal sparse autoencoder. This encoder is trained to learn robust temporal representations for a single channel or all channels jointly. It takes a trial of multi-point time-series data as input, compresses and encodes it through a bottleneck layer (with fewer hidden neurons than the input layer), and then reconstructs it back to the original temporal dimension. Its training objective is not only to minimize reconstruction error, but more importantly, to force the network to activate only a few neurons in the hidden layers to represent the most salient and essential temporal patterns in the input data by introducing a sparsity penalty term (such as a constraint based on KL divergence), thereby automatically filtering out noise and extracting key temporal features.

[0027] After the first level of encoding, the EEG signal is transformed from a high-dimensional "channel × time point" raw space into a low-dimensional temporal feature encoding vector. This vector is then input into the second level—a spatial domain sparse autoencoder. This encoder is designed to learn functional connectivity and coordination patterns between different spatial locations in the brain (i.e., different acquisition channels). It takes the temporal feature encodings from all channels as input and further extracts and fuses cross-channel spatial information through an encoder-decoder process with sparse constraints. This level of encoding can capture co-activation patterns associated with specific motor images, distributed in specific brain regions (e.g., left-hand motor images primarily activate the right hemisphere sensorimotor cortex).

[0028] Through the step-by-step processing of these two sparse autoencoders, the complex, high-dimensional spatiotemporal information in the original EEG signal is automatically and abstractly extracted layer by layer into a highly condensed and information-rich final feature code. This code integrates the discriminative features of the signal in two key dimensions, time and space, laying a solid foundation for subsequent accurate classification.

[0029] S3: Input the feature encoding into the classifier to obtain the decoding result of the user's motion intention; After obtaining the high-dimensional feature encoding output by the cascaded sparse autoencoder neural network, it needs to be mapped to specific, understandable motion intent categories. This mapping function is performed by a classifier. In a preferred embodiment, the classifier is a Softmax classifier.

[0030] Specifically, the final feature encoding vector obtained from step S3 is input into the trained Softmax classifier. The Softmax classifier is essentially a linear layer plus a normalized exponential function. It performs a linear weighted summation of the input feature vectors and calculates a raw score for each possible motion imagery category (e.g., "the ball moves up," "the ball moves down," "the ball moves left," "the ball moves right"). These raw scores are then transformed into a probability distribution using the Softmax function. This probability distribution visually represents the likelihood that the current EEG signal belongs to each preset category, with the sum of the probabilities of all categories being 1.

[0031] The decoding result is the category with the highest probability. For example, if the system pre-defines four categories corresponding to "the ball moves in four directions," and the probability vector output by the Softmax classifier is [0.7, 0.05, 0.12, 0.13], then the system determines the current user's movement intention as "the ball moves upward," corresponding to 0.7. This step explicitly translates the abstract mathematical features extracted in the previous step into concrete, semantic user commands, completing the crucial transformation from neural signals to control semantics.

[0032] Specifically, in practical applications, the display screen is equipped with multiple buttons. The user controls the ball to move to the preset button position, and then controls the robotic arm to perform the operation corresponding to the button.

[0033] S4: Generate control commands based on the decoding results and send them to the robotic arm control system to drive the robotic arm to perform corresponding actions.

[0034] After obtaining the decoded motion intent (such as "the ball moves upward"), the system needs to translate it into a specific action that the robotic arm can perform in the physical or virtual world. This step is completed by the control command generation module.

[0035] First, the system maintains an instruction mapping table or a set of motion planning algorithms. This mapping table associates discrete decoding categories with specific, parameterized robotic arm motion instructions.

[0036] In some embodiments, for example, "the ball moves upward" may be mapped to an instruction: "Control the end effector of the robotic arm to move in a straight line at a speed V in the direction of three-dimensional spatial coordinates (X1, Y1, Z1); In some embodiments, "moving the ball upward" may involve controlling the ball on the display screen to move so that the ball reaches the button area of ​​the display screen, and then controlling the robotic arm to execute the command corresponding to the button.

[0037] It should be noted that the logic for generating control instructions for the specific decoding results can be set based on actual needs, and this application does not limit it in this regard.

[0038] Specifically, the control command generation module queries the mapping table or calls the planning algorithm based on the decoding results to generate corresponding low-level control commands that conform to the robotic arm communication protocol (such as ROS topic messages and Modbus commands). These commands are sent in real time to the connected physical robotic arm or the simulated robotic arm model in the virtual reality environment through the robotic arm drive module (such as ROS controller or PLC).

[0039] Ultimately, the robotic arm's servo drive receives these instructions and precisely controls the rotation of each joint motor, thereby enabling the end effector (such as a gripper) to complete the expected movement, grasping, and placement actions.

[0040] In some embodiments, refer to Figure 2 , Figure 3 Before step S3, the method further includes a step of training the cascaded sparse autoencoder neural network, the training step including: S31: Obtain the training dataset, which includes multi-channel EEG signals from multiple trials and their corresponding motor imagery category labels; Training begins with the construction of the dataset. To train a general model, it is typically necessary to collect EEG data from multiple subjects performing a standardized motor imagery task to capture group-level neural response patterns. Each "trial" represents a complete data recording unit: the system provides explicit task prompts (e.g., the screen displays "imagine a ball moving downwards"), and the subject performs the corresponding motor imagery within a specified time (e.g., 3 seconds), while a multi-channel EEG device simultaneously records the voltage changes of all channels during that time period. This process is repeated dozens to hundreds of times, covering all categories of motor imagery to be identified (the ball moving in various directions). Each recording is precisely labeled with its corresponding imagery category, forming "signal-label" pairs. The data from all trials are preprocessed (e.g., 0.5-40 Hz bandpass filtering, artifact removal, baseline correction) to form the original training dataset and label set. To enhance the model's robustness to individual differences, this dataset preferably includes data from subjects of different ages, genders, and physiological states to achieve group-level modeling.

[0041] S32: Train the temporal sparse autoencoder, whose loss function includes a reconstruction error term between the input signal and the reconstructed signal, a regularization term for the network weights, and a sparsity penalty term for the activation values ​​of the hidden layer neurons. The first step is to train the first level—the temporal sparse autoencoder. Its purpose is to learn to extract robust, low-dimensional temporal representations from the raw high-dimensional temporal signal. For initial training, for each motor imagery category, the average EEG response across all training trials is calculated, resulting in an average response matrix representing the typical spatiotemporal pattern of that category. The average matrices for all categories are then concatenated column-wise to form the initial training input for the temporal SAE.

[0042] In actual training, the smaller the difference between the reconstructed signal (the reconstructed signal obtained by decoding the time-domain features and reconstructing the original signal) and the original signal, the better.

[0043] S33: Encode the EEG signals of the training dataset using a trained temporal sparse autoencoder to obtain temporal features; Once the time-domain autoencoder is trained, it is treated as a fixed feature extraction tool. The raw EEG signals from each trial in the entire training dataset are input into this network. The focus shifts from the network's final output signal to extracting its internal compressed, low-dimensional "code." This "code" is the essential summary of the EEG signal in the time dimension for that trial, which we call the "temporal feature." Through this step, all the original high-dimensional, lengthy time-series data is uniformly transformed into a series of more compact and representative feature vectors.

[0044] S34: Train the spatial domain sparse autoencoder using the temporal features; Next, the second autoencoder is trained using all the temporal features obtained in the previous step. The training objective of this new network is to learn the relationships between features from different brain channels (spatial locations). The temporal features are then further compressed to obtain spatiotemporal features. For example, it needs to learn to recognize a specific pattern of coordinated changes in the feature values ​​of certain brain channels when the ball is imagined moving downwards. When training this network, a loss function that includes reconstruction requirements, weight constraints, and sparsity penalties is used, forcing it to encode only a few key combinations to summarize the spatial cooperative relationships between all channels.

[0045] Specifically, the smaller the difference between the reconstructed signal and the original time-domain signal, the better. This can be achieved by using the corresponding decoder (here, the decoder is used to decode the spatiotemporal features and reconstruct the time-domain signal).

[0046] Correspondingly, its loss function includes a reconstruction error term between the input signal (time domain signal) and the reconstructed signal (reconstructed time domain signal), a regularization term for the network weights, and a sparsity penalty term for the activation values ​​of the hidden layer neurons. S35: Encode the temporal features using a trained spatial domain sparse autoencoder to obtain the final feature code for classification.

[0047] After the spatial domain autoencoder is trained, it is also used as a feature extractor. The temporal feature matrix obtained in step S33 is input into this network, and its compressed "code" is extracted. This final "code" is the ultimate product of the original EEG signal after being refined in the temporal dimension and fused in the spatial dimension. It contains dual key information: "when" the signal appears in what pattern and "where" this pattern is distributed in the brain. This highly refined "final feature code" concentrates all the discriminative information most relevant to motor intention and will then be directly used to train a simple classifier (such as a Softmax classifier) ​​to complete the final mapping from neural signals to explicit commands. Through this series of progressively abstract training steps, the system achieves automatic and efficient parsing of the deep spatiotemporal structure of EEG signals.

[0048] In some preferred training embodiments, to initialize the temporal sparse autoencoder more efficiently and stably, enabling it to quickly capture essential temporal patterns related to motion imagery categories, we employ a data preparation method based on category-averaged responses. This method is a key preliminary step in building a high-performance decoding model.

[0049] Specifically, for each motor imagery category, the average response of the EEG signals of all training trials under that category is calculated to form an average response matrix; the training input data of the time-domain sparse autoencoder is composed of the average response matrices of different categories.

[0050] After obtaining the original training dataset containing multiple trials and labels, we do not directly feed all the raw signals from all trials into the network for training. Instead, we first classify and refine the data. Specifically, the system groups all training trials according to the motion imagery category labels (e.g., "Category A: The ball moves upward" and "Category B: The ball moves downward").

[0051] For each motor imagery category, the system performs the following operation: averaging the multi-channel EEG signals from all training trials within that category over time. In other words, for each EEG channel, at each identical time point, the system calculates the average signal amplitude of all trials belonging to that category at that point. This process is equivalent to extracting a "template signal" or "typical neural response pattern" for that category from a large number of single trials containing random noise and individual transient fluctuations.

[0052] Through the above calculations, a unique average response matrix will be generated for each category of motor imagery. The dimension of this matrix is ​​"number of channels × number of time points", which centrally represents the most stable and common spatiotemporal patterns of electrical activity generated in multiple brain regions when performing this specific type of motor imagery.

[0053] Subsequently, these average response matrices from different categories are concatenated column-wise to form a larger matrix. This concatenated matrix constitutes the training input data for the first stage of the temporal sparse autoencoder.

[0054] The advantage of this method lies in its ability to effectively suppress irrelevant random noise and irrelevant EEG fluctuations in individual trials by averaging trials of the same type. This allows the network to directly learn the core temporal features with the highest signal-to-noise ratio that are most relevant to the task. The input data is directly composed of "typical templates" for each category, forcing the autoencoder to learn how to distinguish and reconstruct these templates with different patterns during encoding. This allows its hidden layer features to better capture discriminative information between categories. Using averaged templates as initial input, compared to using a large amount of messy single-trial data, provides a clearer and more consistent starting point for network training, helping the model converge faster and reducing the risk of getting trapped in local optima.

[0055] This training data preparation strategy, which uses the average response of each category as a "guide," is a crucial foundation for the SSAE model of this invention to effectively learn group-level neural representations and achieve efficient spatiotemporal feature decoupling. It provides a high-quality input starting point rich in discriminative information for subsequent sparse coding and cascaded training processes.

[0056] In some key embodiments, to enable the sparse autoencoder to truly learn the essential, non-redundant feature representations behind the data, this invention introduces a sparsity penalty term based on Kullback-Leibler (KL) divergence into the loss function during training the autoencoder (including the temporal and spatial domains). This mechanism is a crucial engine driving the network to perform efficient feature learning.

[0057] During training, we aim for “sparse” neuronal activity in the autoencoder’s hidden layers (i.e., the coding layers). This means that for any given input sample, we expect only a few specific neurons to be significantly activated, while the majority of neurons remain silent or in a near-zero activation state. This mimics the sparse coding characteristic of information processing in biological neural systems, helping the network learn more representative and interpretable features from the data, rather than simply memorizing the input.

[0058] To achieve this goal, we set a global, desired low activation level target for the entire hidden layer, denoted as the sparsity parameter ρ (e.g., it can be set to 0.05, 0.1, or 0.2). ρ is a small positive number that indicates we expect each neuron to be active for only ρ proportions of the time on average.

[0059] In each round of training, we calculate the actual average activation of each hidden layer neuron across all training samples in the current batch, denoted as . (For the j-th neuron). Then, we use the Kullback-Leibler divergence as a mathematical tool to accurately measure the actual average activation of that neuron. The “difference” or “distance” between us and our expected sparse target ρ.

[0060] The larger the KL divergence value, the greater the deviation between the actual activation pattern and the expected sparse pattern. We sum the KL divergences of all hidden layer neurons to form the total sparsity penalty term, multiply it by a coefficient (sparse penalty weight η), and then add it to the overall loss function of the autoencoder.

[0061] Achieving sparsity through process optimization: The goal of network training (such as through gradient descent) is to minimize the total loss function. Therefore, to minimize the loss including this term, the optimization algorithm automatically adjusts the network's connection weights, driving the actual average activation of each hidden layer neuron. To get as close as possible to our preset minimum target value ρ.

[0062] The technical effect of this mechanism is that the network is forced to allocate only a small number of "active" neurons to encode the core information for each input. This forces these active neurons to learn to capture the most salient and recurring patterns in the input data (such as specific spatiotemporal EEG patterns associated with specific motor images), thereby achieving automatic feature selection and abstraction. Sparse representation reduces the co-adaptability between neurons, reduces the network's dependence on random noise or irrelevant details in the training data, and helps to learn more robust and generalizable features. This is crucial for processing EEG signals with large individual variability and significant noise. Sparse activation makes the ultimately learned features (corresponding to activated neurons) often associated with certain specific and meaningful components in the input data, improving the interpretability of the model.

[0063] In summary, through this sparsity penalty constraint based on KL divergence, the cascaded autoencoder of this invention can spontaneously form an efficient and sparse neural code during training without human intervention. This is one of the core mechanisms that enables it to extract robust and discriminative spatiotemporal features from complex EEG signals, thereby achieving high-precision decoding.

[0064] In some key embodiments, to complete the final mapping from high-level neural features to specific control commands, the present invention employs a Softmax classifier as the final step in decoding in step S3. This classifier receives the final feature encoding extracted by a cascaded sparse autoencoder neural network and outputs a decision with a clear probabilistic interpretation to determine the user's motion intention.

[0065] The role and processing flow of the Softmax classifier are as follows: The final feature encoding vector is then fed into a pre-trained Softmax classifier. This classifier first performs a linear transformation, mapping the input feature vector to a "score" space equal to the number of motion image categories using a weight matrix. The score for each category initially reflects the degree to which the input features match that category.

[0066] These raw scores are then processed by the Softmax function. The Softmax function is a mathematical normalization function whose core function is to transform any set of real numbers (i.e., scores for each category) into a set of probability values ​​that sum to 1. Specifically, it amplifies the differences between scores through exponential operations and then normalizes them so that the output value for each category falls between 0 and 1, and the sum of the output values ​​for all categories is strictly equal to 1. In this way, the classifier's output is no longer an abstract score, but a clear probability distribution vector.

[0067] For example, in a scenario where a ball is controlled to move up, down, left, and right, the Softmax classifier outputs probability vectors of [0.92, 0.05, 0.01, 0.03] corresponding to the four categories of "the ball moves in four directions." This can be intuitively interpreted as: the system determines that the current user's brainwave signal has a 92% probability of corresponding to an "up" movement intention, a 5% probability of corresponding to a "down" movement intention, a 1% probability of corresponding to a "left" movement intention, and a 3% probability of corresponding to a "right" movement intention. The system employs the maximum probability decision rule. That is, it selects the category with the highest probability value from the probability distribution output by the Softmax function and uses it as the final and unique output result for this decoding. In the example above, the category corresponding to the maximum probability of 0.92 is "the ball moves upwards," so the system determines "the ball moves upwards" as the user's movement intention and passes this result to the subsequent control command generation module.

[0068] For example, in a scenario where a robotic arm is controlled to move left and right, two predefined motion intention categories are "left" and "right". The output of a feature encoding vector by the Softmax classifier might be a probability vector such as [0.92, 0.08]. This can be intuitively interpreted as: the system determines that the current user's EEG signal has a 92% probability of corresponding to a "left" motion intention and an 8% probability of corresponding to a "right" motion intention. The system uses the maximum probability decision rule. That is, it selects the category with the highest probability value from the probability distribution output by the Softmax function and uses it as the final and unique output result of this decoding. In the example above, the category corresponding to the maximum probability of 0.92 is "left", so the system determines "left" as the user's motion intention and passes this result to the subsequent control command generation module.

[0069] The technical advantage of using the Softmax classifier lies in the fact that the probabilistic output provides confidence information for the decision, which not only makes the results easier to understand but also provides additional judgment criteria for subsequent systems (such as risk control or adaptive adjustment modules). For example, when the highest probability value is low (e.g., 0.55), the system can determine that the confidence of this decoding is insufficient and take strategies such as ignoring it or requesting the user to repeat the operation. This structure naturally supports multi-class motion imagery decoding tasks of two or more classes, requiring only the corresponding increase in the output layer dimension, providing a convenient architectural foundation for the expansion of system functions. The Softmax classifier can be seamlessly integrated with the aforementioned sparse autoencoder to form an end-to-end trainable network (although they are trained step-by-step in one training phase of this invention), facilitating future global fine-tuning. Therefore, by transforming abstract spatiotemporal feature encoding into class probabilities with clear statistical significance through the Softmax classifier and making decisions based on the principle of maximum probability, this invention completes a crucial and reliable step in EEG signal decoding from continuous feature space to discrete control semantics.

[0070] In some embodiments, in step S1, while collecting EEG signals, the user is provided with visual feedback of virtual reality scenes or robotic arm movements corresponding to the motor imagery task through a graphical user interface, so as to form a closed-loop training or control.

[0071] When a user puts on the EEG device and prepares to start a task, the system launches a graphical user interface. This interface is not a static command screen, but a dynamic virtual reality scene or robotic arm motion simulation interface that is tightly coupled with the content of the motor imagery task.

[0072] For example, to control a robotic arm to pick up a target object, the user first uses brainwaves to control the movement of a small ball, then uses the ball to click buttons, which in turn control the robotic arm to perform different actions to pick up the target object. During this process, the user can see the movement of both the ball and the robotic arm in real time. It should be noted that the robotic arm can be positioned on one side of the display screen, moving based on commands. Alternatively, a virtual robotic arm can be constructed using a virtual scene, moving based on commands and displayed on a preset screen for the user to view.

[0073] This process constitutes a complete biofeedback closed loop of "perception-decoding-feedback-regulation": The user issues neural commands.

[0074] The system decodes and presents the signals: It collects EEG signals, decodes the motor intention, and immediately converts the intention into a visual action.

[0075] User reception and evaluation feedback: Users can see the control effects produced by their "thoughts" in real time.

[0076] Neural strategy adaptive modulation: Based on visual feedback, the user's brain subconsciously adjusts its motor imagery strategies or focus in an attempt to achieve better control. This modulation alters the patterns of the generated brain electrical signals.

[0077] The closed-loop system has significant technical advantages and benefits: Instant feedback greatly shortens the "learning curve" for users to understand how to effectively modulate their own brain signals, enabling users to quickly master the "tricks" of cooperating with the system, which is crucial for the practical application of brain-computer interfaces.

[0078] Vivid visual feedback transforms abstract EEG signal training into concrete, goal-oriented tasks, effectively maintaining the user's attention and reducing fatigue or distraction caused by tedious tasks, thereby collecting higher quality and more stable EEG data.

[0079] Closed-loop control allows the system to continuously adapt to changes in the user's mental state during operation. When slight deviations occur in decoding, the user can observe the feedback and proactively adjust, spontaneously compensating for system errors, thereby achieving more stable and reliable control performance overall.

[0080] The system can dynamically adjust the difficulty of tasks (such as changing the size or distance of the target or adding distractions) based on the user's performance in the closed loop (such as task success rate and reaction time), so that the training is always in a personalized "challenge zone", thereby maximizing the rehabilitation or training effect.

[0081] Therefore, by introducing immersive real-time visual feedback at the signal acquisition source, this invention not only achieves simple "brain control" functions, but also constructs a dynamic closed-loop ecosystem that promotes neuroplasticity and enables human-machine intelligent collaborative evolution. This is one of the core design concepts that enables this system to move from the laboratory to practical applications.

[0082] In some more intelligent and interactive embodiments, the graphical user interface of the present invention has an adaptive difficulty adjustment function. This function enables the system to dynamically adjust the complexity of the motor imagery task based on the user's real-time performance or long-term training progress, thereby providing the user with a personalized and optimized training or control experience. This is a key intelligent feature for achieving efficient brain-computer interface interaction.

[0083] The signal decoding process of the brain-controlled robotic arm employs a cascaded sequential sparse autoencoder (SSAE) structure to illustrate the training and testing process of the aforementioned EEG signal decoding model. This SSAE structure consists of cascaded time-domain / spatial-domain sparse autoencoders and a softmax decoder, as shown in the diagram. Figure 4 As shown. Wherein: 1) Training of time-domain sparse autoencoders The temporal sparse autoencoder has 1200 artificial neurons in both its input and output layers, with only 1 neuron in each hidden layer. During training, the EEG signals from each trial of the training set for each motor cortex are z-normalized and then divided into two classes according to the direction of motor imagery. The average response signal for each class of training set data is then obtained. The average response matrix of the multichannel EEG signals for each directional category of spatial motor imagery is defined as follows: (Matrix size: OK, (list), among which The number of neural response pathways in the motor cortex. The length of the time-domain data of the motor cortex EEG signal. The direction of motion is categorized. Subsequently, the two average response data matrices are concatenated column-wise to serve as the training input data matrix for the time-domain sparse autoencoder. The mean square error of the network's input and output vectors is used to measure the network's encoding and decoding performance. The loss function is constructed according to formula (1) to adjust and update the connection weights between network nodes.

[0084] (1) in The term is a sparsity penalty term, designed to restrict some hidden layer neurons from being inactive. This is the coefficient of the penalty term. Input data vector Feature encoding representation in the hidden layer This represents the number of neurons in the hidden layer. For network connectivity matrix, This is the regularization parameter. For any hidden neuron... Its sparsity is defined by the Kullback-Leibler divergence, which is expressed as follows: middle, This represents the average activation value of the corresponding hidden layer neurons across all training trials. Given a sparsity parameter (set to 0.2 in this study), the model uses KL distance to measure the difference between the two. After the sparse autoencoder is trained, the training set data matrix of neural signals from each motor region ( By inputting the data into this network, the encoded result of feature dimensionality reduction can be obtained. This is used for the next step of training the spatially sparse autoencoder.

[0085] 2) Spatial Domain Sparse Autoencoder Training Subsequently, the data matrix It was used for training the second part of this SSAE architecture, namely the spatial domain SAE. The number of nodes in the input and output layers of this network architecture is 1. This is equal to the number of EEG signal channels in the motor cortex and the number of hidden layer nodes. After the network training is completed, Input the network and obtain the activation value matrix of the hidden layer nodes. This means completing the entire spatiotemporal feature extraction process.

[0086] 3) Training the softmax classifier Data Matrix The training set serves as the training set for the softmax classifier, the third part of the SSAE architecture. After training, this classifier provides a test set. The probability of each spatial motion stimulus angle corresponding to each sample in each trial.

[0087] The system embodiments of this application can be used to execute the method embodiments of this application. For details not disclosed in the system embodiments of this application, please refer to the method embodiments of this application.

[0088] Reference Figure 4 The robotic arm control system based on brain functional signals and artificial neural networks provided in this application includes: Signal acquisition module 51 is used to acquire multi-channel EEG signals from the user; The signal processing and decoding module 52 integrates the pre-trained cascaded sparse autoencoder neural network and classifier, which is used to receive the EEG signal and decode it to output the user's movement intention. The control instruction generation module 53 is used to generate robotic arm control instructions based on the user's movement intention. The robotic arm drive module 54 is used to receive the control commands and drive the robotic arm to perform actions.

[0089] In some embodiments, it also includes: A feedback presentation module is used to provide a graphical user interface to the user, the graphical user interface being used at least to present a virtual reality scene or the state of a robotic arm to provide visual feedback.

[0090] In some embodiments, the system further includes a model training module for performing the training process of the cascaded sparse autoencoder neural network described above.

[0091] Below, for reference Figure 5 This describes an electronic device according to embodiments of the present application. Figure 5 A block diagram of an electronic device according to an embodiment of this application is illustrated.

[0092] like Figure 5 As shown, the electronic device 600 includes one or more processors 610 and memory 620.

[0093] The processor 610 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 600 to perform desired functions.

[0094] The memory 620 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 610 may execute the program instructions to implement the robotic arm control methods based on brain functional signals and artificial neural networks described in the various embodiments of this application above, and / or other desired functions. Various contents, such as category correspondences, may also be stored in the computer-readable storage medium.

[0095] In one example, the electronic device 600 may also include an input device 630 and an output device 640, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0096] In addition, the input device 630 may also include, for example, a keyboard, mouse, interface, etc. The output device 640 can output various information to the outside, including analysis results, etc. The output device 640 may include, for example, a display, speaker, printer, and communication network and its connected remote output devices, etc.

[0097] Of course, for the sake of simplicity, Figure 5 Only some of the components of the electronic device relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device may include any other suitable components depending on the specific application.

[0098] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the robotic arm control method based on brain functional signals and artificial neural networks according to various embodiments of this application as described in the "Exemplary Methods" section above.

[0099] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0100] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the robotic arm control method based on brain functional signals and artificial neural networks according to various embodiments of this application as described in the "Exemplary Methods" section above.

[0101] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0102] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A robotic arm control method based on brain functional signals and artificial neural networks, characterized in that, Includes the following steps: S1: Collect multi-channel EEG signals from users when performing specific motor imagery tasks; S3: The multi-channel EEG signal is input into a pre-trained cascaded sparse autoencoder neural network for decoding; the cascaded sparse autoencoder neural network includes at least a temporal sparse autoencoder and a spatial sparse autoencoder connected in sequence, used to extract and fuse the spatiotemporal features of the EEG signal step by step, and output feature codes; S3: Input the feature encoding into the classifier to obtain the decoding result of the user's motion intention; S4: Generate control commands based on the decoding results and send them to the robotic arm control system to drive the robotic arm to perform corresponding actions.

2. The robotic arm control method according to claim 1, characterized in that, Before step S3, the method further includes a step of training the cascaded sparse autoencoder neural network, the training step including: S31: Obtain the training dataset, which includes multi-channel EEG signals from multiple trials and their corresponding motor imagery category labels; S32: Train the temporal sparse autoencoder, whose loss function includes a reconstruction error term between the input signal and the reconstructed signal, a regularization term for the network weights, and a sparsity penalty term for the activation values ​​of the hidden layer neurons. S33: Encode the EEG signals of the training dataset using a trained temporal sparse autoencoder to obtain temporal features; S34: Train the spatial domain sparse autoencoder using the temporal features; S35: Encode the temporal features using a trained spatial domain sparse autoencoder to obtain the final feature code for classification.

3. The robotic arm control method according to claim 2, characterized in that, In step S31, for each motor imagery category, the average response of the EEG signals of all training trials under that category is calculated to form an average response matrix; the training input data of the time-domain sparse autoencoder is composed of the average response matrices of different categories.

4. The robotic arm control method according to claim 2, characterized in that, The sparsity penalty term is constructed based on Kullback-Leibler divergence and is used to constrain the average activation value of hidden layer neurons to be close to a preset small sparsity parameter ρ.

5. The robotic arm control method according to claim 1, characterized in that, In step S3, the classifier is a Softmax classifier, whose output is the probability that the decoding result belongs to each motion image category, and the category corresponding to the maximum probability is taken as the final decoding result.

6. The robotic arm control method according to any one of claims 1 to 5, characterized in that, In step S1, while collecting EEG signals, the user is provided with visual feedback of virtual reality scenes or robotic arm movements corresponding to the motor imagery task through a graphical user interface, so as to form a closed-loop training or control.

7. The robotic arm control method according to claim 6, characterized in that, The graphical user interface can adaptively adjust the difficulty of the motor imagery task based on the user's EEG signal decoding results or task completion status.

8. A robotic arm control system based on brain functional signals and artificial neural networks for implementing the method of any one of claims 1 to 7, characterized in that, include: The signal acquisition module is used to acquire the user's multi-channel EEG signals; The signal processing and decoding module integrates the pre-trained cascaded sparse autoencoder neural network and classifier, which is used to receive the EEG signals, decode them, and output the user's movement intention. The control command generation module is used to generate robotic arm control commands based on the user's movement intention; The robotic arm drive module is used to receive the control commands and drive the robotic arm to perform actions.

9. The robotic arm control system according to claim 8, characterized in that, Also includes: A feedback presentation module is used to provide a graphical user interface to the user, the graphical user interface being used at least to present a virtual reality scene or the state of a robotic arm to provide visual feedback.

10. The robotic arm control system according to claim 8 or 9, characterized in that, The system further includes a model training module for performing the training process of the cascaded sparse autoencoder neural network as described in any one of claims 2 to 4.