Robot control system and device and storage medium

Through mixed reality interactive equipment and brain-computer interface technology, combined with PPO algorithm to optimize motion strategies, the problems of EEG signal noise interference and low control accuracy are solved, efficient and accurate robot control is achieved, and immersive interactive experience is provided.

CN120335607APending Publication Date: 2025-07-18EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510413242.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the direct use of brain-computer interfaces to control robots has problems such as large interference in EEG signal noise and low control accuracy, which is difficult to meet the flexible operation needs in complex environments.

Method used

Mixed reality interactive equipment is used to collect robot environment information, combine the brain-computer interface module to identify user intentions using the trained EEG classification model, generate motion strategies through the learning module, and execute actions by the robot control module, and optimize motion strategies using the PPO algorithm to provide an immersive interactive experience.

Benefits of technology

It realizes efficient and accurate robot control, and users can intuitively observe the environment and respond quickly to commands, avoiding the problems of noise interference and low control accuracy, and providing an immersive interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335607A_ABST
    Figure CN120335607A_ABST
Patent Text Reader

Abstract

The invention relates to a robot control system and device and a storage medium, the robot control system comprises a mixed reality interaction device, a brain-computer interface module, a learning module and a robot control module, and the mixed reality interaction device is used for collecting current environment information and a real-time state of a robot in reality; mapping the current environment information and the real-time state in a virtual interface of mixed reality interaction equipment; the brain-computer interface module is used for collecting a current electroencephalogram signal of a user and identifying the current electroencephalogram signal by using a trained electroencephalogram classification model to obtain a control intention of the user; the learning module is used for generating a motion strategy according to the current environment information, the real-time state and the reward function; and the robot control module is used for generating a motion instruction of the robot according to the control intention and the motion strategy, and controlling the robot to execute a corresponding action. Therefore, the robot can quickly respond to the instruction of the user, and efficient and accurate robot control is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application mainly relates to the field of control systems, and particularly to a robot control system, device, and storage medium. Background Art

[0002] Quadruped robots have high mobility and flexibility and have broad application prospects in complex terrains. Traditional control methods for quadruped robots mainly include remote control and pre-programmed control, but these methods have certain limitations and are difficult to meet the flexible operation requirements in complex environments. With the continuous development of technology, brain-computer interface technology has been continuously matured. The development of brain-computer interface technology provides new possibilities for the interaction between humans and robots. By capturing the electroencephalogram (EEG) signals of humans, the mind control of robots can be realized. However, due to the noise interference of EEG signals and the control accuracy problem, directly using brain-computer interfaces to control robots still faces challenges. Summary of the Invention

[0003] An object of this application is to provide a robot control system, device, and storage medium to solve the problems of large noise interference of EEG signals and low control accuracy caused by directly using brain-computer interfaces to control robots in the prior art.

[0004] According to one aspect of this application, a robot control system is provided. The control system includes:

[0005] A mixed reality interaction device, a brain-computer interface module, a learning module, and a robot control module. Among them,

[0006] The mixed reality interaction device is configured to collect the current environmental information and real-time state of the robot in reality, and map the current environmental information and real-time state in the virtual interface of the mixed reality interaction device. Among them, the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies for the user;

[0007] The brain-computer interface module is configured to collect the current EEG signals of the user, and use a trained EEG classification model to identify the current EEG signals to obtain the control intention of the user. Among them, the EEG classification model is obtained by classifying and training the measured EEG signals of the user at different stimulation frequencies;

[0008] The learning module is configured to generate a motion strategy according to the current environmental information, real-time state, and reward function;

[0009] The robot control module is configured to generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions.

[0010] Optionally, the SSVEP virtual interface provides flashing squares representing different instructions. The mixed reality interaction device is configured such that when a user trigger signal is received, the squares on the SSVEP virtual interface start to flash and continue for a specified time.

[0011] The brain-computer interface module is used to collect electroencephalogram (EEG) signals consistent with the flashing frequency of the squares.

[0012] Optionally, the mixed reality interaction device is used to make a specified mark on the flashing squares corresponding to the instruction according to the received classification result, so as to prompt the user that the instruction has been successfully recognized and executed.

[0013] Optionally, the training process of the EEG classification model is as follows:

[0014] Classify the measured EEG signals of the user at different stimulation frequencies to obtain the SSVEP classification result.

[0015] Extract the motor imagery features of the motor imagery task, and fuse the SSVEP classification result with the motor imagery features.

[0016] Use the features after feature fusion to train a deep learning-based classifier to obtain the EEG classification model.

[0017] Optionally, the calculation process of the SSVEP classification result is as follows:

[0018] Extract the measured EEG signals of the user at each stimulation frequency, calculate the correlation coefficient between the template signal corresponding to each stimulation frequency and the measured EEG signal, match the target frequency according to the correlation coefficient, and use the target frequency as the SSVEP classification result.

[0019] Optionally, the brain-computer interface module is used to calculate the quality index of the current EEG signal or the measured EEG signal based on a sliding window. When the quality index is less than the threshold, the corresponding EEG signal is marked as invalid.

[0020] Optionally, the mixed reality interaction device includes a spatial positioning module, an interaction logic module, and a multi-modal information display module. Among them, the spatial positioning module is used to position the mixed reality interaction device and display the positioning information.

[0021] The interaction logic module is used to control the button layout on the virtual interface and set an effective fixation area for displaying the stimulation frequency for the user.

[0022] The multi-modal information display module is used to provide the user with the navigation path information of the robot and the status display information of the robot.

[0023] Optionally, the learning module is used to construct a simulation motion environment for the robot according to the current environment information, configure relevant parameters of PPO and set a multi-objective reward function, and perform learning and training in the simulation motion environment to generate a motion strategy.

[0024] Optionally, the multi-objective reward function satisfies the following formula:

[0025] R = w1v forward - w2‖τ‖ 2 + w3h body - w4||w body || 2 - w5∑‖q - q neutral ‖ - w6F impact + w7I goal ;

[0026] where τ represents a regularization term, v forward represents the forward speed, h body represents the torso height, w body represents the angular velocity stability, q neutral represents the neutral state, F impact represents the foot end impact force, I goal represents the performance index of task achievement, and w1, w2... w7 represent the weight coefficients of the corresponding items.

[0027] According to another aspect of the present application, there is also provided a device for robot control, the device includes:

[0028] One or more processors; and

[0029] A memory storing computer-readable instructions, the computer-readable instructions, when executed, cause the processor to perform the following steps when executed:

[0030] Collect the current environment information and real-time state of the robot in reality, and map the current environment information and real-time state in the virtual interface of the mixed reality interaction device, where the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies to the user;

[0031] Collect the current electroencephalogram signal of the user, use the trained electroencephalogram classification model to identify the current electroencephalogram signal, and obtain the control intention of the user, where the electroencephalogram classification model is obtained by classifying and training the measured electroencephalogram signals of the user at different stimulation frequencies;

[0032] Generate a motion strategy according to the current environment information, real-time state and reward function;

[0033] Generate a motion instruction for the robot according to the control intention and the motion strategy, and control the robot to execute the corresponding action.

[0034] According to another aspect of the present application, there is also provided a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the following steps are implemented:

[0035] Collect the current environmental information and real-time state of the robot in reality, and map the current environmental information and real-time state in the virtual interface of the mixed reality interaction device, where the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies to the user;

[0036] Collect the current electroencephalogram (EEG) signal of the user, and use the trained EEG classification model to identify the current EEG signal to obtain the control intention of the user, where the EEG classification model is obtained by classifying and training the measured EEG signals of the user under different stimulation frequencies;

[0037] Generate a motion strategy according to the current environmental information, real-time state, and reward function;

[0038] Generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions.

[0039] Compared with the prior art, the present application provides a robot control system, which includes: a mixed reality interaction device, a brain-computer interface module, a learning module, and a robot control module. Among them, the mixed reality interaction device is used to collect the current environmental information and real-time state of the robot in reality, and map the current environmental information and real-time state in the virtual interface of the mixed reality interaction device, where the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies to the user; the brain-computer interface module is used to collect the current EEG signal of the user, and use the trained EEG classification model to identify the current EEG signal to obtain the control intention of the user, where the EEG classification model is obtained by classifying and training the measured EEG signals of the user under different stimulation frequencies; the learning module is used to generate a motion strategy according to the current environmental information, real-time state, and reward function; the robot control module is used to generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions. Thus, an immersive interaction experience is provided for the user. The user can more intuitively observe the environment where the robot is located and send control instructions. Through the motion strategy, the robot can quickly respond to the user's instructions, thereby realizing efficient and precise robot control. Description of the Drawings

[0040] To make the above objects, features, and advantages of the present application more obvious and understandable, the following provides a detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings, where:

[0041] Figure 1 Schematic diagram showing the structure of a robot control system provided according to one aspect of the present application;

[0042] Figure 2 Signal classification flowchart of the brain-computer interface module in an embodiment of the present application;

[0043] Figure 3 Schematic diagram of the interface of the mixed reality interaction device in an embodiment of the present application;

[0044] Figure 4 Schematic diagram of the process of training with the PPO algorithm in an embodiment of the present application;

[0045] Figure 5 Schematic diagram of the hierarchical distributed architecture of the control system in an embodiment of the present application;

[0046] Figure 6 Schematic diagram of the framework of a device for robot control provided according to another aspect of the present application.

[0047] Identical or similar reference numerals in the drawings represent identical or similar components. Detailed implementation manners

[0048] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific implementation manners of the present application is provided in conjunction with the accompanying drawings.

[0049] In the following description, many specific details are set forth to facilitate a thorough understanding of the present application. However, the present application may be implemented in other ways different from those described herein. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0050] As shown in the present application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0051] The present application provides a robot control system that uses brain-computer interface technology to capture the user's control intention, combines mixed reality technology to provide an immersive interaction experience, and at the same time uses the PPO algorithm to optimize the robot's motion strategy to achieve efficient and precise robot control. This system has the advantages of innovation, immersive experience, high efficiency, ease of use, etc., and has broad application prospects. The specific implementation steps are as follows:

[0052] Figure 1The structural schematic diagram of a robot control system provided according to an aspect of the present application is shown. The control system includes: a mixed reality interaction device 100, a brain-computer interface module 200, a learning module 300, and a robot control module 400. Among them, the mixed reality interaction device 100 is used to collect the current environmental information and real-time status of the robot in reality, and map the current environmental information and real-time status to the virtual interface of the mixed reality interaction device. Among them, the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies to the user; the brain-computer interface 200 module is used to collect the current electroencephalogram signal of the user, and use the trained electroencephalogram classification model to identify the current electroencephalogram signal to obtain the control intention of the user. Among them, the electroencephalogram classification model is obtained by classifying and training the measured electroencephalogram signals of the user at different stimulation frequencies; the learning module 300 is used to generate a motion strategy according to the current environmental information, real-time status, and reward function; the robot control module 400 is used to generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions.

[0053] The mixed reality (MR) interaction device 100 provides an immersive interaction environment for the user, establishes an interaction connection between the user and the robot, and displays environmental information, robot status information, etc. through the interface of the MR device. The user can more intuitively observe the environment where the robot is located and the robot status, and send control instructions. The Quest3 headset can be selected as the mixed reality interaction device, and this device supports gesture recognition. On the virtual interface of MR, a three-dimensional interface display panel is designed, and the user can see the real-time status of the robot in reality and the information of the current environment where the robot is located. A mixed reality interaction interface is constructed using the Quest3 headset, and the interface is mapped to the user's mixed reality interface using the mixed reality function of the Quest3. The user can observe the environment where the robot is located through the headset and send control instructions. The user can send control instructions in the mixed reality interface by means of gazing, gestures, etc., such as selecting the motion mode of the robot, setting the target position, etc.

[0054] In the embodiment of the present application, the virtual interface provides an electroencephalogram control mode, in which the robot autonomously determines a motion strategy based on the electroencephalogram intention. In addition, other types of control modes can also be provided on the virtual interface for the user to choose, such as a manual control mode. In the manual control mode, the user can directly control the motion of the robot through gestures or voice.

[0055] The virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies to the user. Among them, SSVEP (Steady-State Visual Evoked Potential) is a brain electrical activity triggered by visual stimulation. In the SSVEP virtual interface, visual stimulation is simulated to provide different stimulation frequencies to the user. When the user gazes at a certain stimulation frequency, the EEG signal will generate corresponding brain wave responses, and the system identifies the user's selection by performing frequency domain analysis on these signals.

[0056] The brain-computer interface (BCI) module 200 is responsible for capturing the user's neural signals, extracting the features of the captured neural signals, inputting them into a deep learning-based classifier, using the measured EEG signals of the user at different stimulation frequencies as training samples to train the classifier, obtaining an EEG classification model, and then using the trained EEG classification model to identify the current EEG signal of the current user collected by the brain-computer interface module 200 and converting it into a control instruction. In the embodiment of the present application, a non-invasive EEG (electroencephalogram) device can be used to capture the user's EEG signals. After the user wears the EEG helmet, the signals are collected, and then the signals are processed and identified to recognize the user's control intention, such as "forward", "backward", "left turn", "right turn", etc.

[0057] The learning module 300 is responsible for optimizing the motion strategy of the robot. In the embodiment of the present application, the Proximal Policy Optimization (PPO) algorithm can be used. This algorithm is an efficient policy gradient-based reinforcement learning method suitable for control tasks in continuous action spaces. The PPO algorithm continuously updates the policy network through interaction with the environment to maximize the cumulative reward. The design of the reward function is the key to reinforcement learning. Here, the present application designs a multi-objective reward function, which includes a comprehensive evaluation of indicators such as motion speed, stability, and energy consumption. During the training process, a simulation environment is used for offline training to accelerate the training speed and reduce the dependence on actual hardware. After training, the trained motion strategy is deployed to the actual robot, and online fine-tuning is performed through a mixed reality interface to adapt to changes in the actual environment. Finally, the robot control module 400 generates specific motion instructions for the robot according to the received user's control intention and motion strategy, and controls the robot to execute corresponding actions. Thus, the user can control the robot without going through complex training and also avoids directly using the brain-computer interface module, thereby solving the problems of noise and low control accuracy.

[0058] In a specific embodiment of the present application, the robot is preferably a quadruped robot, which is equipped with an IMU, a force sensor, a lidar, and a depth camera; combined with a Quest3 headset and a brain-computer interface module, an instruction system for controlling the quadruped robot based on a mixed reality brain-computer interface is constructed; the user can enter the mixed reality environment through the Quest3 headset, observe the environment where the quadruped robot is located, and send control instructions through the brain-computer interface module; the system uses the PPO algorithm to optimize the motion strategy of the quadruped robot, enabling the quadruped robot to autonomously complete tasks according to the user's intention.

[0059] Specifically, in the initialization stage: after the user wears the EEG helmet and the Quest3 device, the system performs calibration, including EEG signal calibration and spatial mapping of the mixed reality interaction device, and the quadruped robot enters the standby state, waiting for control instructions. In the control stage: the brain-computer interface module continuously collects the user's EEG signals and converts them into control instructions. The mixed reality interaction device provides a virtual interface, and the user can select the control mode through gestures or voice and view the real-time state of the robot. The learning module generates the optimal motion strategy according to the current environmental state and the reward function; in the said control system, the environment is defined as the motion task of the quadruped robot in a specific scenario, such as walking on complex terrain or avoiding obstacles in a dynamic environment. The state information of the quadruped robot includes position, speed, joint angle, etc., and the action space includes torque or speed commands for each joint. The control module realizes the control of the quadruped robot, generates specific joint motion commands according to the received control instructions and the motion strategy, and controls the robot to execute the corresponding actions. In the feedback stage: the mixed reality interaction device displays the motion state and environmental information of the quadruped robot in real time, providing visual feedback; the brain-computer interface module adjusts the control instructions according to the user's neural feedback to achieve closed-loop control.

[0060] In some embodiments of the present application, the brain-computer interface module 200 is used to preprocess the collected electroencephalogram (EEG) signals of the user to extract the target signals, and extract the time-domain features and frequency-domain features of the target signals to obtain the feature signals for identification. Here, an EEG helmet is used, which has 64 electrodes and can capture EEG signals in regions such as the frontal lobe, parietal lobe, and temporal lobe. To improve the signal quality, the present application adopts a filtering technique to remove noise, including 50Hz power frequency interference and electrooculogram (EOG) interference. During the signal processing, a band-pass filter is used, and the filtering range is set between 0.5Hz and 30Hz to remove low-frequency drift and high-frequency noise. In terms of feature extraction, the present application adopts a method combining time domain and frequency domain. Among them, the time-domain features include the mean, standard deviation, and kurtosis of the signal; the frequency-domain features include the power spectral density (PSD) of each frequency band, especially the power changes of the alpha wave, beta wave, and gamma wave. These features are input into a deep learning-based classifier to identify the user's control intention. The system uses the brain-computer interface module to capture the user's EEG signals and uses machine learning algorithms to classify the EEG signals to identify the user's control intention.

[0061] Specifically, the brain-computer interface module is used to collect EEG signal data of the user under different control intentions, such as "forward", "backward", "left turn", "right turn", etc. The collected EEG signals are preprocessed, such as filtering and denoising, to extract useful information; time-domain, frequency-domain and other methods are used to extract the features of the EEG signals, such as power spectrum, waveform features, etc.; the features are classified to train an EEG signal classification model; the system continuously collects the current EEG signals of the user and uses the trained EEG classification model to classify the current EEG signals to identify the user's control intention.

[0062] In some embodiments of the present application, the SSVEP virtual interface provides flashing squares representing different instructions. The mixed reality interaction device 100 is used to make the squares on the SSVEP virtual interface start to flash and last for a specified time when receiving the user's trigger signal; the brain-computer interface module is used to collect EEG signals consistent with the flashing frequency of the squares. Here, multiple flashing squares representing different instructions are provided in the SSVEP virtual interface, and each square corresponds to a specific action instruction, such as forward, backward, left turn, right turn, etc. The user issues a control command by gazing at a specific square. To enhance the sense of space, the arrangement of the squares adopts a 3D layout, and the position, size, and depth of the squares are finely controlled to ensure the vivid display of the interface elements in 3D space. Each square flashes periodically at a different frequency to simulate visual stimuli, allowing each square to quickly switch states within a specified time, thus generating a frequency-specific flashing effect. When the user gazes at a certain square, the EEG signal will generate a corresponding brain wave response, and the system identifies the user's selection by performing frequency-domain analysis on these signals.

[0063] In a specific embodiment of the present application, when the user gazes at the flashing square on the screen, the brain generates brain waves that match the flashing frequency. The EEG device is used to capture these brain electrical signals in real time and extract fluctuations of different frequencies. The frequency of each flashing square is set to a specific value between 8 Hz and 15 Hz. Through the EEG response, the system can identify the square that the user is focusing on. The data collected by the EEG is transmitted to the data recognition module (such as the PC side) in real time. The PC side algorithm analyzes the brain electrical signals and determines the user's selection, and then feedbacks the corresponding control instructions to the mixed reality interaction device and the quadruped robot through the TCP protocol.

[0064] In some embodiments of the present application, the mixed reality interaction device is used to perform a designated marking on the square corresponding to the instruction according to the received classification result, so as to prompt the user that the instruction has been successfully recognized and executed. Here, during the SSVEP flashing, the background of the square flashes. When the control intention of the user is recognized, the elements in the square are marked with a designated color, such as turning red, thereby helping the user intuitively select the instruction.

[0065] When collecting and recognizing the user's brain electrical signals, it is completed through the following steps: First, the user clenches their teeth as a trigger signal. When the system recognizes the electromyogram signal of the user clenching their teeth, the SSVEP interface in the mixed reality interaction device starts to flash for 4 seconds. This flashing process guides the user's attention through visual stimulation and also provides conditions for the brain to generate brain electrical signals consistent with the flashing frequency, facilitating subsequent brain electrical signal recognition. Then, the system enters the waiting stage, and the PC side starts to analyze and classify the collected brain electrical signals. Through algorithmic frequency domain analysis and pattern recognition of the EEG signals, the PC side can identify the user's intention, that is, the instruction (such as forward, backward, etc.) selected by the user in the SSVEP interface. After the classification is completed, the PC side transmits the classification result to the MR helmet through the TCP protocol. After receiving the signal, the MR interface will change the module corresponding to the instruction to red for 1 second as a visual feedback to prompt the user that the instruction has been successfully recognized and executed. Finally, the system enters the standby state, waiting for the next electromyogram signal trigger to restart the entire operation process. This timing process, through the combination of electromyogram signals, brain electrical signals, and the MR interface, provides a smooth, biometric signal-based control method that can achieve precise device interaction and instruction execution.

[0066] In some embodiments of the present application, the training process of the electroencephalogram classification model is as follows: Classify the measured electroencephalogram signals of the user at different stimulation frequencies to obtain the SSVEP classification result; Extract the motor imagery features from the motor imagery task, and fuse the SSVEP classification result with the motor imagery features; Use the features after feature fusion to train a deep learning-based classifier to obtain the electroencephalogram classification model. Here, when training the model, identify and train by collecting the electroencephalogram signals of the user at different stimulation frequencies, extract the electroencephalogram components related to specific visual stimulation frequencies to obtain the SSVEP classification result; Extract the motor imagery features from the motor imagery task, where the motor imagery refers to the behavior of imagining a specific action but not actually performing the action, and the task includes tasks of limb movement, such as waving the right hand, flicking the fingers, waving the left hand, etc. Fuse the SSVEP classification result with the motor imagery features, and use the fused features as training samples to input into a deep learning-based classifier to train the electroencephalogram classification model. Deep learning-based classifiers include, for example, convolutional neural networks (CNNs), recurrent neural networks (RNNs), etc. Thus, when collecting the real-time electroencephalogram signals of the user, the trained model can be used to more accurately identify the control intention of the user.

[0067] As Figure 2 shown in the signal classification flowchart of the brain-computer interface module, the SSVEP and motor imagery hybrid paradigm is adopted, and millisecond-level real-time processing is achieved on the FPGA+GPU heterogeneous computing platform; In the feature fusion stage, event-related desynchronization features (ERD) or event-related synchronization (ERS) features are extracted for the motor imagery task, that is, calculate the event-related desynchronization index in the μ rhythm (8-12Hz) and β rhythm (18-26Hz) frequency bands, for example where, P baseline represents the baseline power, and P task represents the power after performing the task. Integrate the SSVEP classification result and the motor imagery features through decision-level fusion to improve the instruction recognition accuracy, and the accuracy can reach 98.7% in actual measurements.

[0068] In some embodiments of the present application, the calculation process of the SSVEP classification result is as follows: Extract the measured electroencephalogram signals of the user at each stimulation frequency, calculate the correlation coefficient between the template signal corresponding to each stimulation frequency and the measured electroencephalogram signal, match the target frequency according to the correlation coefficient, and use the target frequency as the SSVEP classification result. Here, continue to refer to Figure 2, Signal acquisition stage: Start the SSVEP visual stimulation interface, which is rendered using the Unity engine with a stimulation frequency encoded at 8 - 15 Hz; synchronously trigger the EEG device to start signal acquisition; use canonical correlation analysis to calculate the correlation coefficients between the template signals corresponding to each stimulation frequency and the measured signals, and then use the correlation coefficients to match the most suitable frequency, thereby identifying the user's control intention. Among them, the template signals can refer to the sine-cosine basis function group, and the following formula is used when calculating the correlation coefficients:

[0069]

[0070] Among them, X is the EEG signal matrix, Y is the reference template signal matrix, and w x and w y are the linear transformation weight vectors of the template signal and the measured signal respectively.

[0071] In the embodiment of the present application, the EEG signal classification model is trained based on the Task-Related Component Analysis (TRCA) algorithm in SSVEP. SSVEP is the EEG signal induced by visual stimuli of specific frequencies (such as flickering lights). The goal is to extract the components related to the stimulation frequency from multi-channel EEG signals. TRCA extracts the EEG components related to the stimulation frequency by maximizing the consistency between trials while suppressing noise. Obtain the SSVEP data corresponding to the stimulation in the virtual interface. This data includes the number of channels, the number of sampling points, and the number of trials for repeated acquisition of the same stimulation during the calibration stage; then perform filter bank analysis to decompose the original data into multiple sub-bands, and subsequently calculate the covariance matrix. The calculation formula is as follows:

[0072] Among them, S n represents the cross-covariance matrix, Q n represents the auto-covariance matrix, represents the data of the i-th trial of the n-th stimulation, and N t represents the number of trials for repeated acquisition of the same stimulation during the calibration stage.

[0073] TRCA extracts the task-related components by maximizing the consistency between trials, calculates the spatial filter, and uses the obtained spatial filter for classification to determine the stimulation frequency at which the subject is gazing, and calculates the correlation coefficient between the test data and the training data. Among them, the formula is as follows:

[0074]

[0075] Among them, X (b) is the test data (single-trial data) of the b-th sub-band, is the training data of the nth visual stimulus, represents the spatial filter calculated for the bth sub-band and the nth visual stimulus. Calculate the weighted feature γ n , and comprehensively integrate the information of different sub-bands by weighted fusion of the correlation coefficients of each sub-band, so as to improve the recognition accuracy of the target stimulus. The formula for the weighted feature is as follows:

[0076]

[0077] where c(b) is the weight of each sub-band, and usually c(b) = b -1.25 +0.25. Finally, perform target frequency recognition. By comparing the correlation coefficients of all possible frequencies, select the most matching frequency as the final recognition result. In the SSVEP task, TRCA extracts the EEG components related to the specific visual stimulus frequency by maximizing the consistency between trials, and is applicable to the SSVEP-based brain-computer interface system.

[0078] In an embodiment of the present application, the brain-computer interface module 200 is used to calculate the quality index of the current EEG signal or the measured EEG signal based on a sliding window. When the quality index is less than the threshold, the corresponding EEG signal is marked as invalid. Here, the present application provides an adaptive calibration mechanism to update the TRCA spatial filter online, such as recalculating the covariance matrix of the task-related components every 5 minutes; and can also dynamically adjust the classification threshold, calculate the signal command index (SQI) based on a sliding window (for example, with a length of 2s). When SQI < 0.6, automatically switch to the redundant control channel, that is, consider the signal at this time to be invalid and no longer need to be recognized.

[0079] It should be noted that when calculating the covariance matrix, the following formula is used:

[0080]

[0081] where S b represents the between-class covariance matrix, S w represents the within-class covariance matrix, N t represents the total number of categories, X i represents the mean of the ith class, represents the overall mean of all samples, X i (t) is the sample value in the ith class, is the mean of the samples in the ith class.

[0082] In some embodiments of the present application, the mixed reality interaction device 100 includes a spatial positioning module 101, an interaction logic module 102, and a multi-modal information display module 103. Among them, the spatial positioning module 101 is used to position the mixed reality interaction device 100, realize the display of positioning information, and align the virtual control panel in the virtual interface with the physical environment; the interaction logic module 102 is used to control the button layout on the virtual interface, set an effective gaze area for displaying the stimulation frequency for the user, and calculate the gaze heat map of the user; the multi-modal information display module 103 is used to provide the user with the navigation path information of the robot and the status display information of the robot. Here, as Figure 3 shown in the interface schematic diagram of the mixed reality interaction device, this interface adopts a hierarchical information display architecture and includes the following core components:

[0083] (1) Spatial positioning module. This spatial positioning module realizes sub-millimeter-level spatial positioning (accuracy 0.3 mm), positions the MR device, and can prevent the user wearing it from colliding with obstacles; aligns the virtual control panel with the physical environment through the iterative closest point (ICP) algorithm. And depth perception compensation is adopted to dynamically adjust the transparency of virtual objects according to the time-of-flight (ToF) data. For example, when the depth difference > 0.5 mm, the transparency rises to 70%.

[0084] (2) Interaction logic module. The Fitts' law is adopted to optimize the control button layout, set the effective gaze area within the range of ±15° from the center of the viewing angle, and calculate the gaze heat map, so as to more intuitively determine which flash block the user's stimulation response is to and more accurately obtain the user's electroencephalogram signal.

[0085] (3) Multi-modal information display module. It is used to display the relevant information of the mixed reality interaction device and the robot body; the path navigation of the mixed reality interaction device. This path navigation adopts cubic B-spline curve planning, projects a rainbow-colored light band in the user's field of view, and the width of this light band changes with the confidence level, adjustable from 0.2 to 0.5 m; displays the body status of the robot, and fixedly displays the robot joint temperature (color warning threshold 65 °C), battery remaining capacity (accuracy 1%) and network latency (color coding: green < 50 ms, red > 200 ms) in the lower right corner of the field of view.

[0086] In some embodiments of the present application, the learning module 300 is used to construct a simulation motion environment for the robot, configure the relevant parameters of PPO and set a multi-objective reward function, and perform learning and training in the simulation motion environment to generate a motion strategy. Among them, the multi-objective reward function satisfies the following formula:

[0087] R = w1v forward - w2‖τ‖ 2 + w3hbody -w4||w body || 2 -w5∑‖q - q neutral ‖-w6F impact +w7I goal ;

[0088] Among them, τ represents the regularization term, v forward represents the forward speed, h body represents the torso height, w body represents the angular velocity stability, q neutral represents the neutral state, F impact represents the foot-end impact force, I goal represents the performance index for task accomplishment, and w1, w2... w7 represent the weight coefficients of the corresponding terms.

[0089] In the embodiment of the present application, for the control of the quadruped robot, w1 = 1.0 (forward speed reward), w2 = 0.01 (joint torque penalty), w3 = 2.0 (torso height maintenance), w4 = 0.05 (angular velocity stability term), w5 = 0.1 (joint neutral position deviation penalty), w6 = 10.0 (foot-end impact force penalty), and w7 = 100.0 (task completion reward) can be set.

[0090] The PPO algorithm aims to ensure the stability and efficiency of training by restricting the amplitude of policy updates. Its core idea is to limit the amplitude of policy updates through a clipping mechanism to avoid unstable training caused by excessive policy updates. The main features of PPO include: Stability: Restrict the amplitude of policy updates through the clipping mechanism to prevent policy collapse; Efficiency: Support multiple small-batch updates to improve data utilization; Universality: Applicable to continuous action spaces and high-dimensional state spaces.

[0091] In the embodiment of the present application, when performing PPO algorithm training, adopt as Figure 4For the process shown, the construction of the simulation environment is carried out first, including the construction of the terrain library and obstacles. Among them, the terrain library includes twelve types of complex terrains such as gravel, slopes (0 - 40°), stairs (step height 0 - 0.2m), etc., and Perlin noise is used to generate surface irregularity (roughness coefficient 0.1 - 0.9); the construction of dynamic obstacles includes introducing moving cylinders (speed 0 - 1.5m / s) and pendulums (period 2 - 5s) to form unstructured threats. Then, the PPO hyperparameters are configured, including the calculation of the advantage function, the update of the policy network, and experience replay; the Generalized Advantage Estimation (GAE, λ = 0.95) is used to estimate the advantage value in the calculation of the advantage function, and the discount factor γ = 0.99; when updating the policy network, the KL divergence threshold δ = 0.01 is set, and the learning rate is automatically reduced when KL > δ; the Proximal Policy Optimization - Clip (PPO - Clip) algorithm (ε = 0.2) is used for policy improvement in the experience replay process, with a batch size of 2048, and 10 rounds of partial incremental (mini - batch) updates are performed in each iteration. The Sim - to - Real transfer includes domain randomization and online adaptation. Among them, domain randomization randomizes the ground friction coefficient (0.4 - 1.2), motor delay (0 - 20ms), and sensor noise (IMU white noise σ = 0.05) in the simulation; online adaptation uses the Meta - SAC algorithm in the deployment stage, and the bias terms of the hidden layer of the policy network are updated every 30 seconds to compensate for the Sim - to - Real gap.

[0092] The trained PPO algorithm is used to optimize the motion strategy of the robot, enabling the robot to autonomously complete tasks according to the user's control intentions; the specific implementation includes the following six aspects:

[0093] (1) State representation, the state of the robot includes information such as the current position, speed, joint angles, etc., and the feedback from the environment includes whether the target position is reached, whether an obstacle is encountered, etc.

[0094] (2) Action space, the actions of the robot include the angular changes of each joint, constituting a continuous action space.

[0095] (3) Reward function, the reward function is designed to obtain a positive reward when the target position is reached, and a negative reward when an obstacle is encountered or the robot falls.

[0096] (4) Policy network, a policy network is constructed using a deep neural network, which takes the state information of the robot as input and outputs action instructions.

[0097] (5) Training process, in the simulation environment, the PPO algorithm is used to train the policy network so that the robot can learn the optimal motion strategy.

[0098] (6) Online control. In practical applications, the system uses the trained policy network to control the robot in real time, outputs action instructions according to the current state, and continuously adjusts the policy based on the environmental feedback.

[0099] The control system described in this application includes a brain-computer interface module, a mixed reality interaction device, a learning module, and a robot control module. The modules interact with each other through a data communication protocol to ensure the real-time performance and stability of the system. As Figure 5 shown, this system adopts a hierarchical distributed architecture, which consists of a perception layer, a decision-making layer, an execution layer, and a human-computer interaction layer. The modules achieve data synchronization through ROS middleware.

[0100] Specifically, the perception layer uses the brain-computer interface module, which includes a neural signal acquisition unit and an environmental perception unit. Among them, the neural signal acquisition unit uses a helmet with 64 channels and a sampling rate of 1000Hz to collect electroencephalogram signals such as frontal / parietal lobe through dry electrodes, and completes 50Hz power frequency notch filtering (Q = 30) and 0.5 - 40Hz band-pass filtering through an embedded device; the environmental perception unit exists in the quadruped robot and constructs a dense point cloud map through the SLAM algorithm.

[0101] The decision-making layer uses the learning module, which includes a brain-computer interface decoding engine and motion planning. The brain-computer interface decoding engine adopts a TRCA-SSVEP hybrid classification model and runs real-time decoding on a remote PC (delay < 150ms). The core of motion planning is a reinforcement learning policy network (Actor-Critic structure, hidden layer dimension 256×256) based on PPO that receives decoding instructions and environmental states and outputs a 12-dimensional joint space trajectory (update frequency 100Hz).

[0102] The execution layer uses the robot control module of the quadruped robot, which includes a driver module and a safety monitoring unit. Among them, the driver module uses an open-source motor controller and adopts an impedance control mode, with a stiffness coefficient of 1200 Nm / rad and a damping ratio of 0.7; the safety monitoring unit uses a six-axis force sensor (Range ±200N) and an IMU (BMI088) to form a whole-body dynamics observation system, and the trigger emergency stop threshold is that the foot-end contact force mutation rate > 500N / s.

[0103] The human-computer interaction layer uses a mixed reality interaction device, which realizes MR rendering and multi-modal input functions. Quest3 receives point cloud data through a wireless access point (5.8GHz frequency band) and realizes dynamic light and shadow rendering (guaranteed frame rate of 90fps) in Unity HDRP. For multi-modal input, gesture recognition (accuracy 0.01mm) and voice parsing are integrated to form a "brain electricity + gesture + voice" triple redundant control channel.

[0104] The following method steps are implemented through the above-mentioned layers: collect the EGG raw signal, perform band-pass filtering on the raw signal, extract features using TRCA, and use CPS spatial filtering in the frequency band of 8 - 30 Hz for feature extraction; perform LSTM time series classification on the data after feature extraction to generate motion primitives, use the PPO policy network to process the body information of the robot and the environmental feature vector, and output joint torque commands based on the motion primitives. The joint torque commands can be displayed through a 12-channel PWM waveform; finally, the robot executes the commands and transmits the real-time state back to the MR device interface through the RTPS protocol.

[0105] This system realizes the control of a quadruped robot based on a hybrid reality brain-computer interface and reinforcement learning. Users can control the robot to complete complex tasks through simple thoughts, and it has broad application prospects, such as in the field of medical rehabilitation. The system combines brain-computer interface technology, hybrid reality technology, and reinforcement learning algorithms to achieve efficient and precise control of the quadruped robot and provide an immersive interaction experience.

[0106] Figure 6 The frame schematic diagram of a device for robot control provided according to another aspect of the present application is shown. The device at least includes a processor 601 and a memory 602.

[0107] The processor 601 may include one or more processing cores, such as: a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0108] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory, as well as non-volatile memory, such as one or more magnetic disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one instruction for being executed by the processor 601 to implement the following steps:

[0109] Collect the current environmental information and real-time state of the robot in reality, and map the current environmental information and real-time state in the virtual interface of the mixed reality interaction device, wherein the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies for the user;

[0110] Collect the current electroencephalogram signal of the user, and use the trained electroencephalogram classification model to identify the current electroencephalogram signal to obtain the control intention of the user, wherein the electroencephalogram classification model is obtained by classifying and training the measured electroencephalogram signals of the user at different stimulation frequencies;

[0111] Generate a motion strategy according to the current environmental information, real-time state, and reward function;

[0112] Generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions.

[0113] In some embodiments, the device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface may be connected through a bus or signal line. Each peripheral device may be connected to the peripheral device interface through a bus, signal line, or circuit board. Schematically, the peripheral devices include but are not limited to: radio frequency circuits, touch display screens, audio circuits, and power supplies, etc.

[0114] Of course, the device may also include fewer or more components, and this embodiment does not limit this.

[0115] According to another aspect of the present application, there is also provided a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the following steps are implemented:

[0116] Collect the current environmental information and real-time state of the robot in reality, and map the current environmental information and real-time state in the virtual interface of the mixed reality interaction device, wherein the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies for the user;

[0117] Collect the user's current EEG signal, and use the trained EEG classification model to identify the current EEG signal to obtain the user's control intention, where the EEG classification model is obtained by classifying and training the measured EEG signals of the user at different stimulation frequencies;

[0118] Generate a motion strategy according to the current environmental information, real-time state, and reward function;

[0119] Generate a motion instruction for the robot according to the control intention and the motion strategy, and control the robot to execute the corresponding action.

[0120] When the method for robot control is implemented as a computer program, it can also be stored in a computer-readable storage medium as an article of manufacture. For example, the computer-readable storage medium may include, but is not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs)), smart cards, and flash memory devices (e.g., electrically erasable programmable read-only memories (EPROMs), cards, sticks, key drives). In addition, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media (and / or storage media) that can store, contain, and / or carry code and / or instructions and / or data.

[0121] It should be understood that the above-described embodiments are illustrative only. The embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or any combination thereof. For a hardware implementation, the processor can be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and / or other electronic units designed to perform the functions described herein, or in combination therewith.

[0122] Some aspects of the present application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above-mentioned hardware or software can all be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". The processor can be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors or combinations thereof. In addition, aspects of the present application may be embodied as a computer product located in one or more computer-readable media, which includes computer-readable program code. For example, computer-readable media may include, but are not limited to, magnetic storage devices (such as hard disks, floppy disks, magnetic tapes...), optical discs (such as compact discs CD, digital versatile discs DVD...), smart cards, and flash memory devices (such as cards, sticks, key drives...).

[0123] A computer-readable medium may contain a propagated data signal having computer program code embodied therein, for example, on a baseband or as part of a carrier wave. The propagated signal may take various forms, including electromagnetic form, optical form, etc., or a suitable combination thereof. A computer-readable medium can be any computer-readable medium other than a computer-readable storage medium, which can communicate, propagate, or transport a program for use by being connected to an instruction execution system, apparatus, or device. The program code located on the computer-readable medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, radio frequency signal, or similar media, or any combination of the above media.

[0124] The basic concepts have been described above. Obviously, for those skilled in the art, the above invention disclosure is only an example and does not constitute a limitation to the present application. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to the present application. Such modifications, improvements, and corrections are proposed in the present application, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of the present application.

[0125] At the same time, the present application uses specific terms to describe the embodiments of the present application. Such as "one embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of the present application. Therefore, it should be emphasized and noted that the "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the present application can be appropriately combined.

[0126] In some embodiments, numbers are used to describe components and the quantity of attributes. It should be understood that such numbers used in the description of embodiments are modified by the modifiers "about", "approximate" or "substantially" in some examples. Unless otherwise specified, "about", "approximate" or "substantially" indicate that the said numbers allow a variation of ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, and such approximate values may change according to the characteristics required by individual embodiments. In some embodiments, the numerical parameters should consider the specified significant digits and adopt the method of retaining the general number of digits. Although the numerical ranges and parameters used in some embodiments of the present application to confirm the breadth of their scope are approximate values, in specific embodiments, such numerical settings are as precise as possible within the feasible range.

Claims

1. A robot control system, characterized in that, The control system includes: a mixed reality interaction device, a brain-computer interface module, a learning module, and a robot control module. Among them, the mixed reality interaction device is used to collect the current environmental information and real-time status of the robot in reality, and map the current environmental information and real-time status in the virtual interface of the mixed reality interaction device. Among them, the virtual interface includes an SSVEP virtual interface, which is used to provide different stimulation frequencies for the user; the brain-computer interface module is used to collect the current electroencephalogram (EEG) signal of the user, and use the trained EEG classification model to identify the current EEG signal to obtain the control intention of the user. Among them, the EEG classification model is obtained by classifying and training the measured EEG signals of the user at different stimulation frequencies; the learning module is used to generate a motion strategy according to the current environmental information, real-time status, and reward function; the robot control module is used to generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions.

2. The control system according to claim 1, wherein The SSVEP virtual interface provides flashing squares representing different instructions. The mixed reality interaction device is used to make the squares on the SSVEP virtual interface start to flash and last for a specified time when a trigger signal from the user is received; the brain-computer interface module is used to collect the EEG signal consistent with the flashing frequency of the square.

3. The control system according to claim 2, wherein The mixed reality interaction device is used to make a specified mark on the flashing square corresponding to the instruction according to the received classification result to prompt the user that the instruction has been successfully recognized and executed.

4. The control system according to claim 1, wherein The training process of the EEG classification model is as follows: Classify the measured EEG signals of the user at different stimulation frequencies to obtain the SSVEP classification result; Extract the motor imagery features of the motor imagery task, and fuse the SSVEP classification result with the motor imagery features; Use the features after feature fusion to train a deep learning-based classifier to obtain the EEG classification model.

5. The control system according to claim 4, characterized in that The calculation process of the SSVEP classification result is as follows: Extract the measured EEG signals of the user at each stimulation frequency, calculate the correlation coefficient between the template signal corresponding to each stimulation frequency and the measured EEG signal, match the target frequency according to the correlation coefficient, and use the target frequency as the SSVEP classification result.

6. The control system according to claim 1, wherein, The brain-computer interface module is used to calculate the quality index of the current EEG signal or the measured EEG signal based on a sliding window. When the quality index is less than the threshold, the corresponding EEG signal is marked as invalid.

7. The control system according to claim 1, wherein The mixed reality interaction device includes a spatial positioning module, an interaction logic module, and a multi-modal information display module. Among them, the spatial positioning module is used to position the mixed reality interaction device and display the positioning information; the interaction logic module is used to control the button layout on the virtual interface and set an effective fixation area for displaying the stimulation frequency for the user; the multi-modal information display module is used to provide the navigation path information of the robot and the status display information of the robot for the user.

8. The control system according to claim 1, characterized in that, The learning module is used to construct a simulation motion environment for the robot according to the current environment information, configure the relevant parameters of PPO and set a multi-objective reward function, and perform learning and training in the simulation motion environment to generate a motion strategy.

9. The control system according to claim 8, wherein The multi-objective reward function satisfies the following formula: R = w1v forward - w2‖τ‖ 2 + w3h body - w4||w body || 2 - w5∑‖q - q neutral ‖- w6F impact + w7I goal ; Among them, τ represents the regularization term, v forward represents the forward speed, h body represents the trunk height, W body represents the angular velocity stability, q neutral represents the neutral state, F impact represents the foot-end impact force, I goal represents the performance index for task achievement, and w1, w2... w7 represent the weight coefficients of the corresponding items.

10. A device for robot control, characterized in that, The device includes: one or more processors; and a memory storing computer-readable instructions, which when executed by the processor, perform the following steps: Collect the current environment information and real-time state of the robot in reality, and map the current environment information and real-time state in the virtual interface of the mixed reality interaction device, where the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies to the user; Collect the current electroencephalogram (EEG) signal of the user, and use the trained EEG classification model to identify the current EEG signal to obtain the control intention of the user, where the EEG classification model is obtained by classifying and training the measured EEG signals of the user at different stimulation frequencies; Generate a motion strategy according to the current environment information, real-time state, and reward function; Generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions.

11. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instructions are executed by the processor, the following steps are performed: Collect the current environment information and real-time state of the robot in reality, and map the current environment information and real-time state in the virtual interface of the mixed reality interaction device, where the virtual interface includes an SSVEP virtual interface for providing different stimulation frequencies to the user; Collect the current electroencephalogram (EEG) signal of the user, and use the trained EEG classification model to identify the current EEG signal to obtain the control intention of the user, where the EEG classification model is obtained by classifying and training the measured EEG signals of the user at different stimulation frequencies; Generate a motion strategy according to the current environment information, real-time state, and reward function; Generate a motion instruction for the robot according to the control intention and motion strategy, and control the robot to execute corresponding actions.