Real-time control system and method for humanoid robots integrating electromyography and monocular vision
By integrating electroencephalography (EEG) and monocular vision into a control system, the problems of poor real-time performance and synchronization in humanoid robot control have been solved. This enables precise control of healthy and mobility-impaired users, making it applicable to multiple application fields and enhancing the accuracy of hand movement control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV CHINA
- Filing Date
- 2023-12-06
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies for humanoid robot control suffer from problems such as low real-time performance, poor synchronization, high economic costs, inaccurate hand motion capture, and difficulty in adapting to users with mobility impairments, which limits their application, especially in the healthcare and home service industries.
A control system integrating electromyography (EMG) and monocular vision is employed. Through a posture data, EMG data, and EEG data acquisition system, combined with a computer system and a humanoid robot, real-time multi-modal control is achieved. The system includes a monocular camera, EMG sensors, EEG sensors, a data transmission module, and a humanoid robot. It utilizes 2D human posture estimation algorithms, EMG signal classification algorithms, and EEG signal classification algorithms to control the humanoid robot's movements.
It enables real-time and precise control for different users, is simple and lightweight, and is suitable for both healthy and mobility-impaired users. It is widely used in industrial, educational, entertainment, elderly care, and medical fields, enhancing the accuracy of hand movement control and providing control opportunities for mobility-impaired users.
Smart Images

Figure CN117532609B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal humanoid robot control technology, and in particular to a real-time control system and method for a humanoid robot that integrates electromyography (EMG) and monocular vision. Background Technology
[0002] In recent years, artificial intelligence (AI) technology has been continuously developing, from initial manual programming to sensor-based adaptive control technology, and then to deep learning and neural network technologies. AI technology is increasingly being applied to robot control, making robot control technology more intelligent and integrated, and increasingly sophisticated and mature. Humanoid robots are a type of biomimetic robot, resembling humans in appearance, capable of performing various tasks, and widely used. For example, in real life, especially when there are elderly or mobility-impaired users at home, humanoid robots can learn human behavior and movements through visual sensors or electromyography (EMG) sensors, and are used in the home service industry. In intelligent factories, humanoid robots can learn user movements through visual sensors or EMG sensors, enabling the robot to acquire the same motor skills as the user and complete high-difficulty, high-intensity tasks.
[0003] Humanoid robots typically consist of motors, sensors, and computers. Their movements and behaviors can be controlled through programming. Using traditional control methods, humanoid robots can perform corresponding actions based on instructions, which is a mechanical form of movement. In contrast, mimicking human actions to create corresponding movements is a natural form of movement. In recent years, researchers have continuously innovated in the development of humanoid robot control technology. Analyzing human movements and simultaneously controlling the movements of humanoid robots is one of the important research directions. Existing control methods have the following problems:
[0004] 1) The traditional method mainly involves programming the angles of the servos corresponding to multiple joints one by one to achieve control of different joints. This not only lacks good real-time performance, resulting in low efficiency, but also greatly affects the synchronization between joints. Existing patents (Chinese patent publication number CN112571446A) only focus on the control of humanoid robot arms and do not pay attention to the control of the robot torso.
[0005] 2) Existing patents (Chinese Patent Announcement No. CN113305830B) use multiple gyroscopes for attitude measurement, which increases economic costs and restricts user activities. Controlling the humanoid robot's movements using human movements requires the person being imitated to wear corresponding sensors to capture signals from human skeletal points and convert them into control signals. This not only requires a large number of sensors, leading to increased economic costs, but also restricts human activities due to prolonged wearing of sensor devices.
[0006] 3) Visual sensors are used to detect key points of the human body. Existing algorithms cannot fit the movement of key points of a human body to the simulated human movement of a humanoid robot. Furthermore, due to phenomena such as self-occlusion of the human body and occlusion by clothing, the capture of hand details is not accurate, making it difficult to control hand movements.
[0007] 4) For users with mobility impairments, relying on visual sensors and electromyography sensors cannot collect user information, making it difficult to promote in the medical and health and home service industries. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to address the shortcomings of the prior art by providing a real-time control system and method for humanoid robots that integrates electromyography (EMG) and monocular vision, thereby realizing real-time control of humanoid robots.
[0009] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0010] On one hand, this invention provides a real-time control system for a humanoid robot that integrates electromyography (EMG) and monocular vision, including a posture data acquisition system, an EMG data acquisition system, an EMG data acquisition system, a computer system, a data transmission module, and a humanoid robot. The posture data acquisition system is used to acquire two-dimensional motion videos of the user's whole body, preprocess them, and then transmit them to the computer system through the data transmission module. The EMG data acquisition system is used to acquire EMG signals of the user's arm movements, process them, and then transmit them to the computer system through the data transmission module. The EMG data acquisition system acquires the user's motor imagery EMG signals and transmits them to the computer system through the data transmission module. The computer system converts the received two-dimensional motion videos of the user's whole body, EMG signals of arm movements, and EMG signals of motor imagery into corresponding control commands to control the humanoid robot to perform corresponding actions.
[0011] Preferably, the posture data acquisition system includes a monocular camera and a video processing unit. The monocular camera is used to acquire two-dimensional motion video of the user's whole body and transmit the motion video to the video processing unit. The video processing unit sets the video transmission settings for the motion video images.
[0012] The electromyography (EMG) data acquisition system includes an EMG sensor and an EMG signal processing unit. The EMG sensor is used to collect EMG signals from the user's arm movements and transmit the EMG signals to the EMG signal processing unit. The EMG signal processing unit preprocesses the EMG signals and then transmits them to the computer system through the data transmission module.
[0013] The EEG data acquisition system includes an EEG signal sensor and an EEG signal transmission unit. The EEG signal sensor is used to collect the user's motor imagery EEG signals, and the EEG signal transmission unit transmits the EEG signals to the computer system.
[0014] Preferably, the computer system includes a visual interface and humanoid robot control software with deployed algorithms. The visual interface visualizes the two-dimensional human pose estimation skeleton after processing the human pose estimation model, and sets two control modes: manual control and automatic control. At the same time, buttons are set for each type of action in the manual mode, so that users can select the mode and the control method.
[0015] The humanoid robot control software allows for manual parameter setting of the servos in each joint of the humanoid robot. The algorithms deployed in the humanoid robot control software include a 2D human posture estimation algorithm, an electromyography signal classification algorithm, and an electroencephalogram (EEG) signal classification algorithm.
[0016] Preferably, the control mode of the humanoid robot includes two modes: manual and automatic. In manual mode, the computer system realizes manual control of the humanoid robot according to the actions selected by the user. In automatic mode, the computer system performs automatic motion control of the humanoid robot based on the user's visual, electromyographic, or electroencephalographic signals, including visual control, electromyographic control, and electroencephalographic control.
[0017] In manual control mode, the corresponding actions of the humanoid robot have been fixed and set to the corresponding control signals. Users only need to select the corresponding actions in the visual interface.
[0018] In the automatic control mode, the user moves freely in front of the monocular camera, ensuring that the whole body appears within the camera's range. The monocular camera transmits video data to the computer system. At the same time, the computer system uses a 2D human posture estimation model to perform human posture recognition, converting the recognized posture information into control signals, so that the humanoid robot follows the user to perform corresponding actions.
[0019] In automatic control mode, the user wears electromyography (EMG) sensors on both hands and transmits EMG signals to the computer system via Bluetooth. At the same time, the computer system uses an EMG signal classification model to classify the EMG signals into actions, and converts the classified actions into control signals so that the humanoid robot can follow the user to perform corresponding actions.
[0020] In the automatic control mode, the user wears an EEG sensor on their brain and uses a Bluetooth module to transmit the EEG signals to the computer system. At the same time, the computer system uses a motor imagery EEG classification model to classify the EEG signals into actions, and converts the classified actions into control signals so that the humanoid robot can perform the corresponding actions according to the user's intention.
[0021] In automatic control mode, you can choose one of the three control methods, or you can choose two or three control methods.
[0022] The 2D human posture estimation algorithm includes a human posture estimation algorithm based on YOLOv7-POSE and a joint information and control signal fitting algorithm. The electromyography (EMG) signal classification algorithm uses a hybrid classification action model based on surface EMG signals for motion classification. The electroencephalogram (EEG) signal classification algorithm uses a motion imagery EEG classification model based on a CNN-LSTM feature fusion network for motion classification.
[0023] Preferably, the data transmission module includes a network card module and a Bluetooth communication module. The network card communication module is used for data transmission between the monocular camera and the computer system, and the Bluetooth communication module is used for data transmission between the electromyography sensor, the electroencephalogram (EEG) sensor, and the computer system.
[0024] Preferably, the humanoid robot includes a humanoid robot body and a signal receiving module. The signal receiving module is used to receive control signals sent by the computer system and transmit the control signals to the servos at each joint of the humanoid robot to control the movement of the humanoid robot.
[0025] On the other hand, the present invention also provides a real-time control method for a humanoid robot that integrates electromyography and monocular vision, comprising the following steps:
[0026] Step 1: Label each servo motor of the humanoid robot body, and use the humanoid robot control software in the computer system to adjust and record the control signal magnitude of the initial position of each servo motor in the humanoid robot body offline.
[0027] Where n is the servo motor number, N is the total number of servo motors in the humanoid robot body, and θ is the control signal magnitude of each servo motor. The magnitude of the control signal for the servo motor numbered n at its initial position e;
[0028] Step 2: Use a computer system to adjust the robot body to its maximum range of motion and record the maximum value of the servo motor control signal. and minimum value
[0029] in, Let n be the maximum value of the control signal for the servo motor with the number n within its maximum range of motion. This represents the minimum value of the control signal for the servo motor numbered n within its maximum range of motion.
[0030] Step 3: Connect the monocular camera to the computer system using the network card module and fix it in a suitable position so that it can fully capture the user's full-body movements;
[0031] Place the monocular camera and the computer system on the same network segment, so that the computer system can call the monocular camera through the RTSP protocol to capture the user's movement process and transmit the video to the computer system.
[0032] Step 4: Connect the electromyography (EMG) sensor to the computer system using a Bluetooth module, so that the computer system can receive the EMG signals collected by the EMG sensor;
[0033] Step 5: Connect the EEG signal sensor to the computer system using a Bluetooth module, so that the computer system can receive the EEG signals collected by the EEG signal sensor;
[0034] Step 6: Position the user within the field of view of the monocular camera, open the computer system's visual interface, and select the robot control mode.
[0035] Step 7: The computer system uses a 2D human posture estimation model to process the two-dimensional posture data of the whole body, calculates the motion angles of each skeleton in the upper limbs, and converts the user's motion angles into control signals to control the servo motor.
[0036] Step 7.1: The 2D human pose recognition model deployed in the computer system uses YOLOv7-POSE to predict the positions of human key points, achieving end-to-end detection of key points and returning the two-dimensional coordinate data of 17 key points throughout the body. After numbering the key points, a data set {(o h ,u h )|h=0,1,2,……,16}, where, o h ,u h Here are the two-dimensional coordinates of the h-th joint.
[0037] Step 7.2: Calculate the joint vectors using the two-dimensional coordinate data of human body key points:
[0038] β cz =(o z -o c ,u z -u c ),c,z∈{0,1,2,……,16},c≠z
[0039] Where the subscripts c and z are the key point numbers, β czIt is the joint vector pointing from keypoint c to keypoint z;
[0040] Step 7.3: Select h′ key points and calculate the joint angles at these key points as the main motion angles δ of the human torso. h′ , where h′∈h;
[0041] Step 7.4: Enable the 2D human pose recognition model The camera measures the main movement angle δ of the human torso by having the user move within its field of view. h′ The maximum value δ of (h′=5,6,7,8,12) h′f With minimum value δ h′g ;
[0042] Step 7.5: Match the corresponding servo motor numbers and motion angles for controlling the movement of the humanoid robot's torso.
[0043] Step 7.6: Using linear function fitting, establish the relationship between the control signal of the humanoid robot and the angle of human torso movement offline, as follows:
[0044]
[0045] in, This represents the maximum value of the servo control signal numbered n. Let δ be the minimum value of the servo control signal numbered n. h′f δ represents the maximum angle of the human body key point numbered h. h′g Let h be the minimum angle of the human body key point numbered h, where ψ and ζ are both coefficients;
[0046] Step 7.7: The human posture skeleton model is displayed on the computer system's visualization interface, enabling simultaneous observation and timely adjustment of the human posture skeleton and humanoid robot movements;
[0047] Step 8: The computer system uses a hybrid classification model based on a self-updating algorithm to process the electromyography data, perform real-time classification, and update the data in real time, converting different categories of hand movements into control signals to control the servo motor.
[0048] Step 8.1: The electromyography (EMG) signal classification model deployed within the computer system applies a hybrid classification action model based on surface EMG signals. This model integrates a Gaussian mixture model (GMM) to eliminate external motion interference and a multilinear discriminant analysis (LDA) model to classify target motion data, achieving real-time, high-precision classification of hand movements and returning the hand movement category. Where K is the number of action categories that the hybrid classification action model can recognize, l represents the left hand, and r represents the right hand;
[0049] Step 8.2: Based on the hand movements, adjust the corresponding servo control signals using the computer system. Where ξ is related to the category of hand movements The servo motor numbers of the relevant robot joints enable the humanoid robot to respond to control signals. Able to perform user hand movements To perform real-time control of the robot's hand movements;
[0050] When action In this case, the above action needs to be defined as a. K+1 As a new target class, the LDA model is updated to achieve incremental recognition of new actions;
[0051] Step 9: The computer system uses a motion imagery EEG classification model based on a CNN-LSTM feature fusion network to process the EEG data, perform real-time classification, and convert different categories of hand movements into control signals to control the servo motor.
[0052] Step 9.1: The EEG classification model deployed within the computer system applies a motor imagery EEG classification model based on a CNN-LSTM feature fusion network. This model integrates a convolutional neural network (CNN) and a long short-term memory network (LSTM). CNN extracts spatial features, and LSTM extracts temporal features. A flattening layer is added after the convolutional layers for feature fusion, improving the accuracy of motor imagery EEG classification and returning the classification of motor imagery EEG. in This represents the number of action categories that the model can recognize;
[0053] Step 9.2: Based on the hand movements, adjust the corresponding servo control signals using the computer system. Where τ is related to the category of motor imagery EEG movements. The servo motor numbers of the relevant robot joints enable the humanoid robot to... Able to perform the actions imagined by the user. To perform real-time robot motion control;
[0054] Step 10: The humanoid robot receives the servo control signals from steps 7 to 9 in real time based on the human's movements, and drives the servo motors to start moving.
[0055] The beneficial effects of adopting the above technical solution are as follows: The real-time control system and method for humanoid robots that integrates electromyography and monocular vision provided by the present invention are as follows:
[0056] (1) It adopts a multi-mode and multi-modal control method and can be applied to different groups of users, including healthy users and users with mobility impairments such as hemiplegia. Users can choose the appropriate mode according to their actual needs and their own requirements.
[0057] (2) It is simple and lightweight, with a good user visualization and control interface, which allows users to operate it conveniently and easily.
[0058] (3) It can provide real-time control, ensure high efficiency, and can be widely used in industrial, educational, entertainment, elderly care and medical fields.
[0059] (4) For healthy users, the main limb movements of the whole body are captured non-contactly using monocular vision. Users do not need to wear gyroscopes or other inertial measurement units, and it does not hinder the user's movement.
[0060] (5) Modular control of the actuator (servo motor) means that it can be selected according to the servo motor number, without the need for separate programming control of each servo motor.
[0061] (6) The needle uses electromyography signals to control the hand movements of the humanoid robot, which enhances the accuracy of hand movement control.
[0062] (7) Use EEG signal classification to perform motor imagery EEG control, providing opportunities for users with hemiplegia or other mobility impairments to control robots.
[0063] (8) By integrating multiple types of information, the humanoid robot’s main joint movements can be controlled using monocular visual information, and the humanoid robot’s hand and wrist movements can be precisely controlled using electromyographic information. This enables users with hemiplegia or other mobility impairments to control the humanoid robot’s movements using electroencephalography (EEG). Attached Figure Description
[0064] Figure 1 This is a schematic diagram of the operation of the real-time control system for a humanoid robot that integrates electromyography and monocular vision, provided in an embodiment of the present invention.
[0065] Figure 2 This is a structural block diagram of a real-time control system for a humanoid robot that integrates electromyography and monocular vision, provided in an embodiment of the present invention.
[0066] Figure 3 A schematic diagram of the structure of a real-time control system for a humanoid robot that integrates electromyography and monocular vision, provided in an embodiment of the present invention.
[0067] Figure 4 A schematic diagram of the electromyography sensor and EEG cap worn according to an embodiment of the present invention;
[0068] Figure 5 A flowchart of a real-time control method for a humanoid robot that integrates electromyography and monocular vision, provided in an embodiment of the present invention;
[0069] Figure 6This is a framework diagram of a 2D human pose estimation model based on YOLOv7-POSE provided in an embodiment of the present invention;
[0070] Figure 7 This is a schematic diagram of key points for two-dimensional human body pose estimation provided in an embodiment of the present invention;
[0071] Figure 8 This is a framework diagram of the hybrid electromyography classification model provided in an embodiment of the present invention;
[0072] Figure 9 A schematic diagram of the CNN-LSTM feature fusion network structure for EEG classification provided in an embodiment of the present invention.
[0073] In the picture, 0 is the nose; 1 is the left eye; 2 is the right eye; 3 is the left ear; 4 is the right ear; 5 is the left shoulder; 6 is the right shoulder; 7 is the left elbow; 8 is the right elbow; 9 is the left wrist; 10 is the right wrist; 11 is the left hip; 12 is the right hip; 13 is the left knee; 14 is the right knee; 15 is the left foot; and 16 is the right foot. Detailed Implementation
[0074] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0075] In this embodiment, a real-time control system for a humanoid robot that integrates electromyography (EMG) and monocular vision is used, such as... Figure 1-3 As shown, the system includes a posture data acquisition system, an electromyography (EMG) data acquisition system, an electroencephalography (EEG) data acquisition system, a computer system, a data transmission module, and a humanoid robot. The posture data acquisition system acquires two-dimensional motion videos of the user's entire body, preprocesses them, and then transmits them to the computer system via the data transmission module. The EMG data acquisition system acquires EMG signals of the user's arm movements, processes them, and then transmits them to the computer system via the data transmission module. The EEG data acquisition system acquires EEG signals of the user's motor imagery and transmits them to the computer system via the data transmission module. The computer system converts the received two-dimensional motion videos of the user's entire body, EMG signals of arm movements, and EEG signals of motor imagery into corresponding control commands to control the humanoid robot to perform corresponding actions.
[0076] In this embodiment, the posture data acquisition system includes a HIKVISION monocular camera and a video processing unit. The monocular camera is used to acquire two-dimensional motion video of the user's entire body and transmit the motion video to the video processing unit. The video processing unit configures the video transmission settings for the motion video images. Simultaneously, the HIKVISION monocular camera can also transmit video streams to the computer system via a network card module.
[0077] The electromyography (EMG) data acquisition system includes an MYO EMG sensor and an EMG signal processing unit. The MYO EMG sensor is used to collect EMG signals from the user's arm movements and transmits the EMG signals to the EMG signal processing unit. The EMG signal processing unit preprocesses the EMG signals and then transmits them to the computer system through the data transmission module.
[0078] The EEG data acquisition system includes an EEG signal sensor and an EEG signal transmission unit. The EEG signal sensor is used to collect the user's motor imagery EEG signals, and the EEG signal transmission unit transmits the EEG signals to the computer system. In this embodiment, the EEG signal sensor uses an EEG cap to collect the user's motor imagery EEG signals. The wearing of the electromyography (EMG) sensor and the EEG cap is as follows: Figure 4 As shown, the MYO electromyography (EMG) sensor is worn in the middle of the user's left and right forearms, and the EEG cap is worn on the user's head.
[0079] The computer system includes a visual interface and humanoid robot control software with deployed algorithms. The visual interface is used to select the control mode of the humanoid robot and visualize the 2D human posture estimation. The humanoid robot control software can manually set the parameters of the servos in each joint of the humanoid robot. The algorithms deployed in the humanoid robot control software include 2D human posture estimation algorithm, electromyography signal classification algorithm and electroencephalography signal classification algorithm.
[0080] The humanoid robot has two control modes: manual and automatic. In manual mode, the computer system controls the humanoid robot manually based on the actions selected by the user. In automatic mode, the computer system automatically controls the humanoid robot's movements based on the user's visual, electromyographic, or electroencephalographic signals.
[0081] The 2D human pose estimation algorithm includes a human pose estimation algorithm based on YOLOv7-POSE and a joint information and control signal fitting algorithm. The electromyography (EMG) signal classification algorithm uses a hybrid classification action model based on surface EMG signals for motion classification. The electroencephalogram (EEG) signal classification algorithm uses a motion imagery EEG classification model based on a CNN-LSTM feature fusion network for motion classification.
[0082] In this embodiment, the computer system's visualization interface is designed using the tkinter library in Python. It visualizes the two-dimensional human posture estimation skeleton after processing the human posture estimation model and sets up two control modes: manual control and automatic control (visual control, electromyography control, or electroencephalography control). At the same time, buttons are set for each type of action in the manual mode, allowing users to select the mode and control method.
[0083] Robot control modes include: manual control and automatic control, where automatic control includes visual control, electromyographic control and electroencephalographic control;
[0084] In manual control mode, the corresponding actions of the humanoid robot have been fixed and set to the corresponding control signals. Users only need to select the corresponding actions in the visual interface.
[0085] In the automatic control mode, the user moves freely in front of the HIKVISION monocular camera, ensuring that the whole body is within the camera's range. The HIKVISION monocular camera transmits video data to the computer system. At the same time, the computer system uses a 2D human posture estimation model to perform human posture recognition, converting the recognized posture information into control signals, so that the humanoid robot follows the user to perform corresponding actions.
[0086] In automatic control mode, the user wears MYO electromyography sensors on both hands and transmits electromyography signals to the computer system via Bluetooth. At the same time, the computer system uses an electromyography signal classification model to classify the electromyography signals into actions, and converts the classified actions into control signals so that the humanoid robot can follow the user to perform corresponding actions.
[0087] In the automatic control mode, the user wears an EEG cap on their head and uses a Bluetooth module to transmit EEG signals to the computer system. At the same time, the computer system uses a motor imagery EEG classification model to classify the EEG signals into actions, and converts the classified actions into control signals so that the humanoid robot can perform corresponding actions according to the user's intentions.
[0088] In automatic control mode, you can choose one of the three control methods, or you can choose two or three control methods.
[0089] The human pose estimation algorithm based on YOLOv7-POSE not only has a fast inference speed but also enables long-distance recognition, avoiding the impact of changes in the distance between the user and the HIKVISION monocular camera. Furthermore, it can predict and complete key points of the limbs in areas that the HIKVISION monocular camera cannot capture, avoiding problems caused by occlusion by the user or other objects.
[0090] Currently, most electromyography (EMG) recognition models can only recognize a limited number of target movements, failing to meet users' needs for new movements. Furthermore, the recognition performance of these models is reduced by external movement interference. In contrast, hybrid classification movement models based on surface EMG signals can overcome external movement interference, and their recognition ability can grow online. Moreover, these models do not require pre-trained movements, effectively reducing the user's training burden and improving the robustness of the system.
[0091] The data transmission module includes a network card module and a Bluetooth communication module. The network card communication module is used for data transmission between the HIKVISION monocular camera and the computer system, and the Bluetooth communication module is used for data transmission between the MYO electromyography sensor, electroencephalography signal sensor and the computer system.
[0092] The humanoid robot consists of the humanoid robot body and a signal receiving module. The signal receiving module is used to receive control signals sent by the computer system and transmit the control signals to the servos at each joint of the humanoid robot to control the movement of the humanoid robot.
[0093] In this embodiment, the humanoid robot body includes a power supply, an emergency stop switch, and a servo module; the power switch is used to control the power on and off of the entire system; the emergency stop switch immediately stops the running humanoid robot when the control system malfunctions; the servo module includes multiple servos that control the movement of each joint of the humanoid robot.
[0094] In this embodiment, a real-time control method for a humanoid robot that integrates electromyography (EMG) and monocular vision is described, such as... Figure 5 As shown, it includes the following steps:
[0095] Step 1: Label each servo motor of the humanoid robot body, and use the humanoid robot control software in the computer system to adjust and record the control signal magnitude of the initial position of each servo motor in the humanoid robot body offline. Where n is the servo motor number, N is the total number of servo motors in the humanoid robot body, and θ is the control signal magnitude of each servo motor. The magnitude of the control signal for the servo motor numbered n at its initial position e;
[0096] In this embodiment, based on the human body structure, the total number of servo motors in the humanoid robot body is set to N=28, and the initial state of the humanoid robot is set to the state of arms hanging naturally with hands open, corresponding to the state of the user standing relaxed with hands open.
[0097] Step 2: Use a computer system to adjust the robot body to its maximum range of motion and record the maximum value of the servo motor control signal. and minimum value
[0098] in, Let n be the maximum value of the control signal for the servo motor with the number n within its maximum range of motion. This represents the minimum value of the control signal for the servo motor numbered n within its maximum range of motion.
[0099] Step 3: Connect the HIKVISION monocular camera to the computer system using the network card module, and fix it in a suitable position so that it can fully capture the user's full-body movements;
[0100] The HIKVISION monocular camera and the computer system are placed on the same network segment, so that the computer system can call the HIKVISION monocular camera through the RTSP (Real Time Streaming Protocol) to capture the user's movement process and transmit the video to the computer system;
[0101] Step 4: Connect the MYO electromyography sensor to the computer system using a Bluetooth module, so that the computer system can receive the electromyography signals collected by the MYO electromyography sensor;
[0102] Step 5: Connect the EEG signal sensor to the computer system using a Bluetooth module, so that the computer system can receive the EEG signals collected by the EEG signal sensor;
[0103] Step 6: Position the user within the field of view of the monocular camera, open the computer system's visual interface, and select the robot control mode.
[0104] Step 7: The computer system uses a 2D human posture estimation model to process the two-dimensional posture data of the whole body, calculates the motion angles of each skeleton in the upper limbs, and converts the user's motion angles into control signals to control the servo motor.
[0105] Step 7.1: The 2D human pose recognition model deployed in the computer system uses the extended model of the YOLO series—YOLOv7-POSE. The YOLOv7-POSE model is based on a deep convolutional neural network architecture, consisting of multiple convolutional layers, max pooling layers, and fully connected layers. The model network receives the input image and generates feature maps to predict the positions of human key points, achieving end-to-end detection of human key points and returning two-dimensional coordinate data of 17 key points throughout the body. After numbering the key points, a data set {(o h ,u h )|h=0,1,2,……,16}, where, o h ,u h Here are the two-dimensional coordinates of the h-th joint.
[0106] In this embodiment, the YOLOv7-POSE model simultaneously implements bounding box detection and keypoint detection based on the YOLOv7 framework. Figure 6As shown, it mainly consists of four parts: the Input layer, the Backbone network, the Neck network, and the Head network. First, the image (640×640 pixels) enters the Input layer and is pre-trained in the Backbone network to extract features. Then, the Neck network fuses features from various feature layers, combining positional and semantic information. Finally, the Head network classifies the image, outputting three feature maps of different sizes. The number of channels is then adjusted to output the human bounding box and keypoint prediction results.
[0107] Step 7.2: Calculate the joint vectors using the two-dimensional coordinate data of human body key points:
[0108] β cz =(o z -o c ,u z -u c ),c,z∈{0,1,2,……,16},c≠z
[0109] Where the subscripts c and z are the key point numbers, β cz It is the joint vector pointing from keypoint c to keypoint z;
[0110] Step 7.3: Select h′ key points and calculate the joint angles at these key points as the main motion angles δ of the human torso. h′ , h′ <h;
[0111] In this embodiment, key points numbered 5, 6, 7, 8, and 12 are selected, and the joint angles at these key points are calculated as the main motion angles of the human torso, as follows:
[0112]
[0113] Where h′ is 5, 6, 7, 8, 12, and when h′=5, the key points q and s are 7 and 11 respectively; when h′=6, the key points q and s are 8 and 12 respectively; when h′=7, the key points q and s are 5 and 9 respectively; when h′=8, the key points q and s are 6 and 10 respectively; when h′=12, the key points q and s are 5 and 11 respectively.
[0114] In this embodiment, the key points for human two-dimensional pose estimation are as follows: Figure 7 As shown, the angles of human joints can be calculated using the relationships between them. The specific construction method is as follows:
[0115] Taking the left elbow 7 angle construction method as an example, with the three joint points left shoulder 5, left elbow 7, and left wrist 9 as the vertices of the triangle, the joint vector is calculated using the two-dimensional coordinate data of the key points of the human body in step 7.2, and β is obtained. 57 ,β 79,β 95 That is, a triangle with three vertices. Using the law of cosines, the angle between the three points on the left elbow and the seven key points on the body can be calculated. The specific formula is:
[0116]
[0117] The angles of other key points on the human body can be achieved using the above construction method.
[0118] Step 7.4: Enable the 2D human pose recognition model to allow the user to move within the camera's field of view and measure the main motion angles δ of the human torso. h′ The maximum value δ of (h′=5,6,7,8,12) h′f With minimum value δ h′g ;
[0119] Step 7.5: Match the corresponding servo motor numbers and motion angles for controlling the movement of the humanoid robot's torso.
[0120] In this embodiment, the corresponding servo motor number and motion angle of the humanoid robot controlling the corresponding human torso movement are matched as (h′,n)∈{(5,10),(6,20),(7,7),(8,17),(12,27)};
[0121] Step 7.6: Using linear function fitting, establish the relationship between the control signal of the humanoid robot and the angle of human torso movement offline, as follows:
[0122]
[0123] in, This represents the maximum value of the servo control signal numbered n. Let δ be the minimum value of the servo control signal numbered n. h′f δ represents the maximum angle of the human body key point numbered h. h′g Let h be the minimum angle of the human body key point numbered h, where ψ and ζ are both coefficients, obtained by solving the equation;
[0124] Step 7.7: The human posture skeleton model is displayed on the computer system's visualization interface, enabling simultaneous observation and timely adjustment of the human posture skeleton and humanoid robot movements;
[0125] Step 8: The computer system uses a hybrid classification model based on a self-updating algorithm to process the electromyography data, perform real-time classification, and update the data in real time, converting different categories of hand movements into control signals to control the servo motor.
[0126] Step 8.1: The electromyography (EMG) signal classification model deployed within the computer system applies a hybrid classification action model based on surface EMG signals. This model integrates a Gaussian mixture model (GMM) to eliminate external motion interference and a multilinear discriminant analysis (LDA) model to classify target motion data, achieving real-time, high-precision classification of hand movements and returning the hand movement category. Where K is the number of action categories that the hybrid classification action model can recognize, l represents the left hand, and r represents the right hand;
[0127] In this embodiment, the framework of the hybrid electromyography classification model is as follows: Figure 8 As shown, to avoid interference from external data, a general hybrid model framework that integrates a single-class classifier and a multi-class classifier is used. The single-class classifier determines whether a sample belongs to an external class, while the multi-class classifier assigns non-external class samples to a specific target class. The single-class classifier employs a Gaussian Mixture Model (GMM), and the multi-class classifier uses Multilinear Discriminant Analysis (LDA) to achieve real-time, high-precision classification of actions.
[0128] In this embodiment, the electromyography (EMG) sensor collects EMG data, filters and denoises it, and extracts features to obtain EMG samples.
[0129] A Gaussian Mixture Model (GMM) consists of K GMMs, corresponding to K target classes, used to distinguish target samples from external samples. For each target class a k (k = 1, 2, ..., K), using sample set X k Build GMM offline:
[0130]
[0131]
[0132] Where p(·) denotes the probability density function of the GMM, p(x; a k ) represents the probability that the xk feature vector in the sample belongs to class ak. It is the eigenvector x k The probability of belonging to the j-th probability density function in a GMM; D i The number of mixed components can be determined by minimizing Akaike information; γ j Is it satisfying γ j >0 and The mixing coefficient, C j It is the covariance matrix, v j It is a d-dimensional mean vector, |C j | represents C j The determinant value, {γ j ,v j Cj It can be determined using the Expectation Maximization (EM) algorithm.
[0133] The core of multilinear discriminant analysis (LDA) is to find a projection matrix W that projects the original samples into a low-dimensional space, maximizing the distance between each class in the projected space while minimizing the distance between data within each class itself, thus making the samples from each class separable. The specific algorithm flow and calculation method are as follows:
[0134]
[0135]
[0136]
[0137]
[0138] Using an offline sEMG sample set {X1, X2, ... X} of K target classes K Training the LDA classifier. First, calculate the mean m of each target action sample. k And the overall sample mean m, then calculate the within-class covariance matrix S. w and the inter-class covariance matrix S b The optimal projection matrix W is given by the matrix corresponding to S. w -1 S b The eigenvectors are composed of the q largest eigenvalues, where the largest eigenvalue is the largest. Using trained LDA, the target sample z is classified. First, the distance from the center of the target sample to the sample point in the projection space is calculated, and the sample with the smallest distance is the target class.
[0139]
[0140] The MYO electromyography sensor can detect either the left or right hand depending on its placement, with left-hand movements represented as... Right-hand movements are represented as
[0141] Step 8.2: Based on the hand movements, adjust the corresponding servo control signals using the computer system. Where ξ is related to the category of hand movements The servo motor numbers of the relevant robot joints enable the humanoid robot to respond to control signals. Able to perform user hand movements To perform real-time control of the robot's hand movements;
[0142] In this embodiment, the left-hand servo control number is ζ. l ∈{1,2,3,4,5}, left hand as The corresponding servo control signal is:
[0143]
[0144] Similarly, the right-hand servo control number is ζ. r ∈{11,12,13,14,15}, right hand as The corresponding servo control signal is:
[0145]
[0146] When action In this case, the above action needs to be defined as a. K+1 As a new target class, and using the samples in it to update the in-class covariance matrix S. w and the inter-class covariance matrix S b Update the LDA model to achieve incremental recognition of new actions;
[0147] In this embodiment, the sample mean, sum of squares, and covariance matrix of the new target class are first calculated:
[0148]
[0149]
[0150] S K+1 =G K+1 -Mm K+1 (m K+1 ) T
[0151]
[0152]
[0153]
[0154] The superscript “~” indicates the updated value.
[0155] After the update is complete, calculate The matrix formed by the eigenvectors corresponding to the first q largest eigenvalues. By updating the projection matrix W, the updated LDA can be obtained.
[0156] Step 9: The computer system uses a motion imagery EEG classification model based on a CNN-LSTM feature fusion network to process the EEG data, perform real-time classification, and convert different categories of hand movements into control signals to control the servo motor.
[0157] Step 9.1: The EEG classification model deployed within the computer system applies a motor imagery EEG classification model based on a CNN-LSTM feature fusion network. This model integrates a convolutional neural network (CNN) and a long short-term memory network (LSTM). CNN extracts spatial features, and LSTM extracts temporal features. A flattening layer is added after the convolutional layer for feature fusion to achieve accurate motor imagery EEG classification and return the classification of motor imagery EEG. in This represents the number of action categories that the model can recognize;
[0158] In this embodiment, the CNN-LSTM feature fusion network for EEG signal classification is as follows: Figure 9 As shown, the CNN network consists of an input layer, a 1-D convolutional layer, a separable convolutional layer, and two flat layers; the LSTM network consists of an input layer, an LSTM layer, and a flat layer; finally, the two networks are classified into fully connected layers.
[0159] Step 9.2: Based on the hand movements, adjust the corresponding servo control signals using the computer system. Where τ is related to the category of motor imagery EEG movements. The servo motor numbers of the relevant robot joints enable the humanoid robot to... Able to perform the actions imagined by the user. To perform real-time robot motion control;
[0160] Step 10: The humanoid robot receives the servo control signals from steps 7 to 9 in real time based on the human's movements, and drives the servo motors to start moving.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A real-time control system for a humanoid robot integrating electromyography (EMG) and monocular vision, characterized in that: The system includes a posture data acquisition system, an electromyography (EMG) data acquisition system, an electroencephalography (EEG) data acquisition system, a computer system, a data transmission module, and a humanoid robot. The posture data acquisition system acquires two-dimensional motion videos of the user's entire body, preprocesses them, and then transmits them to the computer system via the data transmission module. The EMG data acquisition system acquires EMG signals from the user's arm movements, processes them, and then transmits them to the computer system via the data transmission module. The EEG data acquisition system acquires the user's motor imagery EEG signals and transmits them to the computer system via the data transmission module. The computer system converts the received two-dimensional motion videos of the user's entire body, EMG signals from arm movements, and EEG signals from motor imagery into corresponding control commands to control the humanoid robot to perform corresponding actions. The computer system includes a visual interface and humanoid robot control software with deployed algorithms. The visual interface visualizes the two-dimensional human posture estimation skeleton after processing the human posture estimation model and sets two control modes: manual control and automatic control. In manual mode, buttons are provided for various types of actions, allowing users to select the mode and control method. In manual mode, the computer system manually controls the humanoid robot based on the user's selected actions. In automatic mode, the computer system automatically controls the humanoid robot's movements based on the user's visual, electromyographic, or electroencephalographic signals, including visual control, electromyographic control, and electroencephalographic control. The humanoid robot control software can manually set the parameters of the servo motors in each joint of the humanoid robot. The algorithms deployed in the humanoid robot control software include 2D human posture estimation algorithm, electromyography signal classification algorithm and electroencephalography signal classification algorithm. In manual control mode, the corresponding actions of the humanoid robot have been fixed and set to the corresponding control signals. Users only need to select the corresponding actions in the visual interface. In the automatic control mode, the user moves freely in front of the monocular camera, ensuring that the whole body appears within the camera's range. The monocular camera transmits video data to the computer system. At the same time, the computer system uses a 2D human posture estimation model to perform human posture recognition, converting the recognized posture information into control signals, so that the humanoid robot follows the user to perform corresponding actions. In automatic control mode, the user wears electromyography (EMG) sensors on both hands and transmits EMG signals to the computer system via Bluetooth. At the same time, the computer system uses an EMG signal classification model to classify the EMG signals into actions, and converts the classified actions into control signals so that the humanoid robot can follow the user to perform corresponding actions. In the automatic control mode, the user wears an EEG sensor on their brain and uses a Bluetooth module to transmit the EEG signals to the computer system. At the same time, the computer system uses a motor imagery EEG classification model to classify the EEG signals into actions, and converts the classified actions into control signals so that the humanoid robot can perform the corresponding actions according to the user's intention. In automatic control mode, you can choose one of the three control methods, or you can choose two or three control methods. The 2D human posture estimation algorithm includes a human posture estimation algorithm based on YOLOv7-POSE and a joint information and control signal fitting algorithm. The electromyography (EMG) signal classification algorithm uses a hybrid classification action model based on surface EMG signals for motion classification. The electroencephalogram (EEG) signal classification algorithm uses a motion imagery EEG classification model based on a CNN-LSTM feature fusion network for motion classification.
2. The real-time control system for a humanoid robot integrating electromyography and monocular vision according to claim 1, characterized in that: The posture data acquisition system includes a monocular camera and a video processing unit. The monocular camera is used to acquire two-dimensional motion video of the user's whole body and transmit the motion video to the video processing unit. The video processing unit sets the video transmission settings for the motion video images. The electromyography (EMG) data acquisition system includes an EMG sensor and an EMG signal processing unit. The EMG sensor is used to collect EMG signals from the user's arm movements and transmit the EMG signals to the EMG signal processing unit. The EMG signal processing unit preprocesses the EMG signals and then transmits them to the computer system through the data transmission module. The EEG data acquisition system includes an EEG signal sensor and an EEG signal transmission unit. The EEG signal sensor is used to collect the user's motor imagery EEG signals, and the EEG signal transmission unit transmits the EEG signals to the computer system.
3. The real-time control system for a humanoid robot integrating electromyography and monocular vision according to claim 2, characterized in that: The data transmission module includes a network card module and a Bluetooth communication module. The network card communication module is used for data transmission between the monocular camera and the computer system, and the Bluetooth communication module is used for data transmission between the electromyography sensor, the electroencephalography signal sensor and the computer system.
4. The real-time control system for a humanoid robot integrating electromyography and monocular vision according to claim 1, characterized in that: The humanoid robot includes a humanoid robot body and a signal receiving module. The signal receiving module is used to receive control signals sent by the computer system and transmit the control signals to the servos at each joint of the humanoid robot to control the movement of the humanoid robot.
5. A real-time control method for a humanoid robot integrating electromyography (EMG) and monocular vision, implemented based on the system described in claim 1, characterized in that: Includes the following steps: Step 1: Label each servo motor of the humanoid robot body, and use the humanoid robot control software in the computer system to adjust and record the control signal magnitude of the initial position of each servo motor in the humanoid robot body offline. ; in, Let N be the number of the servo motor, and N be the total number of servo motors in the humanoid robot body. The magnitude of the control signal for each servo motor. For the number The servo motor controls the signal magnitude at the initial position e; Step 2: Use a computer system to adjust the robot body to its maximum range of motion and record the maximum value of the servo motor control signal. and minimum value ; in, For the number The maximum value of the control signal within the maximum range of motion of the servo motor. For the number The minimum value of the control signal within the maximum range of motion of the servo motor; Step 3: Connect the monocular camera to the computer system using the network card module and fix it in a suitable position so that it can fully capture the user's full-body movements; Place the monocular camera and the computer system on the same network segment, so that the computer system can call the monocular camera through the RTSP protocol to capture the user's movement process and transmit the video to the computer system. Step 4: Connect the electromyography (EMG) sensor to the computer system using a Bluetooth module, so that the computer system can receive the EMG signals collected by the EMG sensor; Step 5: Connect the EEG signal sensor to the computer system using a Bluetooth module, so that the computer system can receive the EEG signals collected by the EEG signal sensor; Step 6: Position the user within the field of view of the monocular camera, open the computer system's visual interface, and select the robot control mode. Step 7: The computer system uses a 2D human posture estimation model to process the two-dimensional posture data of the whole body, calculates the motion angles of each skeleton in the upper limbs, and converts the user's motion angles into control signals to control the servo motor. Step 8: The computer system uses a hybrid classification model based on a self-updating algorithm to process the electromyography data, perform real-time classification, and update the data in real time, converting different categories of hand movements into control signals to control the servo motor. Step 9: The computer system uses a motion imagery EEG classification model based on a CNN-LSTM feature fusion network to process the EEG data, perform real-time classification, and convert different categories of hand movements into control signals to control the servo motor. Step 10: The humanoid robot receives the servo control signals from steps 7 to 9 in real time based on the human's movements, and drives the servo motors to start moving.
6. The real-time control method for humanoid robots integrating electromyography and monocular vision according to claim 5, characterized in that: The specific method for step 7 is as follows: Step 7.1: The 2D human pose recognition model deployed in the computer system uses YOLOv7-POSE to predict the positions of human key points, achieving end-to-end detection of key points and returning the two-dimensional coordinate data of 17 key points throughout the body. After numbering the key points, a dataset is formed. ,in, For the first Two-dimensional coordinates of each joint point; Step 7.2: Calculate the joint vectors using the two-dimensional coordinate data of human body key points: ; Among them, subscript These are the key point numbers. It is the joint vector pointing from keypoint c to keypoint z; Step 7.3: Select Calculate the joint angles at these key points as the main motion angles of the human torso. ; Step 7.4: Enable 2D human posture recognition model, allowing users to move within the camera's field of view and measuring the main movement angles of the human torso. maximum value and minimum value ; Step 7.5: Match the corresponding servo motor numbers and motion angles for controlling the movement of the humanoid robot's torso. Step 7.6: Using linear function fitting, establish the relationship between the control signal of the humanoid robot and the angle of human torso movement offline, as follows: ; in, This represents the maximum value of the servo control signal numbered n. Let n be the minimum value of the servo control signal numbered n. The maximum angle of the human body key point numbered h. Let h be the minimum angle of the key human body point. and All are coefficients; Step 7.7: The human body posture skeleton model is displayed on the computer system's visualization interface, enabling simultaneous observation and timely adjustment of the human body posture skeleton and humanoid robot movements.
7. The real-time control method for humanoid robots integrating electromyography and monocular vision according to claim 6, characterized in that: The specific method for step 8 is as follows: Step 8.1: The electromyography (EMG) signal classification model deployed within the computer system applies a hybrid classification action model based on surface EMG signals. This model integrates a Gaussian mixture model (GMM) to eliminate external motion interference and a multilinear discriminant analysis (LDA) model to classify target motion data, achieving real-time, high-precision classification of hand movements and returning the hand movement category. Where K is the number of action categories that the hybrid classification action model can recognize, l represents the left hand, and r represents the right hand; Step 8.2: Based on the hand movements, adjust the corresponding servo control signals using the computer system. ,in, It is related to the category of hand movements The servo motor numbers of the relevant robot joints enable the humanoid robot to respond to control signals. Able to perform user hand movements To perform real-time control of the robot's hand movements; When action At that time, the above actions need to be defined as As a new target class, the LDA model is updated to achieve incremental recognition of new actions.
8. The real-time control method for humanoid robots integrating electromyography and monocular vision according to claim 7, characterized in that: The specific method for step 9 is as follows: Step 9.1: The EEG classification model deployed within the computer system applies a motor imagery EEG classification model based on a CNN-LSTM feature fusion network. This model integrates a convolutional neural network (CNN) and a long short-term memory network (LSTM). CNN extracts spatial features, and LSTM extracts temporal features. A flattening layer is added after the convolutional layers for feature fusion, improving the accuracy of motor imagery EEG classification and returning the classification of motor imagery EEG. ,in This represents the number of action categories that the model can recognize; Step 9.2: Based on the hand movements, adjust the corresponding servo control signals using the computer system. ,in It is related to the category of motor imagery EEG movements. The servo motor numbers of the relevant robot joints enable the humanoid robot to... Able to perform the actions imagined by the user. To perform real-time robot motion control.