Exoskeleton control method and apparatus based on reinforcement learning
By using a reinforcement learning-based exoskeleton control method, the joint torques of the exoskeleton and the rendering parameters of the virtual reality scene are dynamically adjusted, which solves the problem of low safety in traditional exoskeleton control methods and improves the safety of matching the user's physiological state and responding to real-time anomalies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-05-18
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional exoskeleton control methods cannot dynamically adjust according to the user's real-time physiological state, resulting in a disconnect between torque commands and the user's actual movement intentions and physiological tolerance. The lack of a closed-loop feedback mechanism reduces the safety of users during training.
An exoskeleton control method based on reinforcement learning is adopted. By acquiring the user's physiological and motion signal set, preprocessing and feature extraction are performed to construct a state feature vector. The reinforcement learning agent is used to generate control parameters and dynamically adjust the exoskeleton joint torque and virtual reality scene rendering parameters to achieve real-time response and improve safety.
It improves safety during user training, ensures that torque commands match the user's physiological state, responds to abnormal situations in real time, and avoids visual dizziness and safety hazards.
Smart Images

Figure CN122440433A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a method and apparatus for controlling exoskeletons based on reinforcement learning. Background Technology
[0002] With the integration of rehabilitation engineering and artificial intelligence technology, the rehabilitation training model combining lower limb exoskeleton robots and virtual reality (VR) technology is becoming increasingly mature. Currently, in lower limb rehabilitation training for stroke patients, the common approach is to drive the exoskeleton to perform preset movements by recognizing motor intentions in electroencephalogram (EEG) or electromyogram (EMG) signals.
[0003] However, when using the above methods to control exoskeletons, the following technical problems often arise: Traditional methods cannot dynamically adjust based on the user's real-time physiological state, causing the generated torque commands to become disconnected from the user's actual movement intentions and physiological tolerance. At the same time, traditional methods lack a closed-loop feedback mechanism and cannot respond to abnormal situations during training in real time, resulting in reduced user safety during training.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure propose reinforcement learning-based exoskeleton control methods, devices, electronic devices, and computer-readable media to address one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide an exoskeleton control method based on reinforcement learning. The method includes: in response to receiving a control request from a target user for the exoskeleton, acquiring a set of physiological and motion signals of the target user and a current virtual reality scene complexity level; preprocessing and extracting features from the physiological and motion signal set to obtain electroencephalogram (EEG) features, electromyogram (EMG) features, and joint motion features; constructing a state feature vector based on the current virtual reality scene complexity level, the EEG features, the EMG features, and the joint motion features; inputting the state feature vector into a preset reinforcement learning agent to obtain a set of control parameters, wherein the control parameters in the set include an auxiliary torque gain and a target virtual reality scene complexity level; generating rendering parameters based on the target virtual reality scene complexity level, and controlling the virtual reality device to perform scene rendering based on the rendering parameters; generating exoskeleton joint torque commands based on the auxiliary torque gain, and controlling the exoskeleton to perform corresponding torque adjustment operations based on the exoskeleton joint torque commands.
[0008] Secondly, some embodiments of this disclosure provide an exoskeleton control device based on reinforcement learning. The device includes: an acquisition unit configured to, in response to receiving a control request from a target user for the exoskeleton, acquire a set of physiological and motion signals of the target user and a current virtual reality scene complexity level; a preprocessing and feature extraction unit configured to preprocess and extract features from the set of physiological and motion signals to obtain electroencephalogram (EEG) features, electromyogram (EMG) features, and joint motion features; and a construction unit configured to construct a state based on the current virtual reality scene complexity level, the EEG features, the EMG features, and the joint motion features. The system comprises: a feature vector; an input unit configured to input the aforementioned state feature vector into a preset reinforcement learning agent to obtain a set of control parameters, wherein the control parameters in the set include an auxiliary torque gain and a target virtual reality scene complexity level; a first generation unit configured to generate rendering parameters based on the target virtual reality scene complexity level, and to control the virtual reality device to perform scene rendering based on the rendering parameters; and a second generation unit configured to generate exoskeleton joint torque commands based on the auxiliary torque gain, and to control the exoskeleton to perform corresponding torque adjustment operations based on the exoskeleton joint torque commands.
[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0011] The above embodiments of this disclosure have the following beneficial effects: the reinforcement learning-based exoskeleton control method of some embodiments of this disclosure can improve user safety during training. Specifically, the reason for the reduced user safety during training is that traditional solutions cannot dynamically adjust according to the user's real-time physiological state, causing the generated torque commands to be out of sync with the user's actual movement intentions and physiological tolerance. Furthermore, traditional solutions lack a closed-loop feedback mechanism and cannot respond to abnormal situations during training in real time, leading to reduced user safety during training. Based on this, the reinforcement learning-based exoskeleton control method of some embodiments of this disclosure first, in response to receiving a control request from the target user for the exoskeleton, obtains the target user's physiological and motion signal set and the current virtual reality scene complexity level. This allows for the acquisition of basic data for controlling exoskeleton movement. Second, the physiological and motion signal set is preprocessed and feature extracted to obtain EEG features, EMG features, and joint motion features. This removes environmental noise and interference from the original data. Third, based on the current virtual reality scene complexity level, the EEG features, the EMG features, and the joint motion features, a state feature vector is constructed. Therefore, multi-dimensional discrete features can be fused into a unified-dimensional state space representation. Next, the aforementioned state feature vector is input into a preset reinforcement learning agent to obtain a set of control parameters, including auxiliary torque gain and the complexity level of the target virtual reality scene. This yields device control parameters that match the user's current situation. Then, based on the aforementioned complexity level of the target virtual reality scene, rendering parameters are generated, and the virtual reality device is controlled to render the scene according to these parameters. This allows for the output of a suitable virtual image to the user while avoiding visual dizziness caused by sudden changes in the image. Finally, based on the aforementioned auxiliary torque gain, exoskeleton joint torque commands are generated, and the exoskeleton is controlled to perform corresponding torque adjustment operations according to these commands. Ultimately, this improves user safety during training. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0013] Figure 1This is a flowchart of some embodiments of the reinforcement learning-based exoskeleton control method according to the present disclosure; Figure 2 This is a schematic diagram of the structure of some embodiments of the reinforcement learning-based exoskeleton control device according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] Figure 1 A flow 100 of some embodiments of a reinforcement learning-based exoskeleton control method according to the present disclosure is shown. The reinforcement learning-based exoskeleton control method includes the following steps: Step 101: In response to receiving a control request from the target user for the exoskeleton, obtain the target user's physiological and motion signal set and the current virtual reality scene complexity level.
[0021] In some embodiments, the execution entity (e.g., a server) of the reinforcement learning-based exoskeleton control method can, in response to receiving a control request from a target user for the exoskeleton, acquire the target user's physiological and motion signal set and the current virtual reality scene complexity level. The reinforcement learning-based exoskeleton control method can be applied to an exoskeleton control system. The exoskeleton control system includes an exoskeleton, a virtual reality device, and a multi-source acquisition unit that are communicatively connected. The exoskeleton can be a robotic device worn on the lower limbs, integrating a drive unit and a motion unit. The exoskeleton can be used to provide adhesion or guidance to a patient's lower limbs during rehabilitation training. The drive unit can be a device for generating auxiliary torque gain, such as a hydraulic rod. The motion unit can be a device for transmitting power and guiding motion, such as a linkage mechanism and joint bearings. The virtual reality device can be a head-mounted display capable of generating visual training scenes. The multi-source acquisition unit can include an EEG electrode cap, a surface electromyography sensor, and a joint encoder. The EEG electrode cap can be a wearable device for acquiring electroencephalogram (EEG) signals generated by the cerebral cortex. The aforementioned surface electromyography (EMG) sensor can be an electrode pad attached to the skin surface of the main muscle groups (such as the quadriceps and hamstrings) of the target user's lower limbs, used to collect the weak electrical signals (surface EMG signals) generated during muscle contraction. The aforementioned joint encoder can be a position sensor installed at each joint of the exoskeleton. The aforementioned joint encoder can measure the joint angle and angular velocity in real time. The aforementioned current virtual reality scene complexity level can be used to define the difficulty level of the current training task. The aforementioned current virtual reality scene complexity level can be a classification label, ranging from level 1 to level 10. The higher the level of the aforementioned current virtual reality scene complexity level, the greater the mass of the virtual objects in the virtual reality scene, the greater the friction, and the faster the speed of the moving target.
[0022] In practice, the aforementioned executing entity can use the aforementioned multi-source acquisition device to obtain the physiological and motion signal set of the target user and the complexity level of the current virtual reality scene.
[0023] Step 102: Preprocess and extract features from the physiological and motor signal set to obtain EEG features, EMG features, and joint movement features.
[0024] In some embodiments, the aforementioned executing entity may preprocess and extract features from the aforementioned physiological and motor signal set to obtain electroencephalogram (EEG) features, electromyogram (EMG) features, and joint motion features. The physiological and motor signals in the aforementioned physiological and motor signal set include EEG signals, surface EMG signals, and joint motion data.
[0025] In some optional implementations of certain embodiments, the aforementioned execution entity may perform preprocessing and feature extraction on the aforementioned physiological and motor signal set through the following steps to obtain electroencephalogram (EEG) features, electromyogram (EMG) features, and joint motion features: Step one involves bandpass filtering and artifact removal of the aforementioned EEG signals to obtain processed EEG signals. In practice, the executing entity can use a bandpass filter to perform bandpass filtering on the EEG signals to obtain filtered EEG signals. Next, an adaptive filtering algorithm can be used to remove artifacts from the filtered EEG signals to obtain processed EEG signals. For example, using the acquired signal from the electrooculography (EOG) channel as a reference input, artifact signals generated by eye movements are removed from the EEG signals of each motor cortex (e.g., C3, C4, Cz) to obtain processed EEG signals. The aforementioned adaptive filtering algorithm can be: .
[0026] in, This represents a discrete-time index. This indicates ocular artifact signals. In discrete time The transpose of the weight coefficient vector. Representing discrete time The reference signal vector at that time. In discrete time The weight coefficient vector. This indicates the transpose operation.
[0027] Step two involves feature extraction from the processed EEG signal to obtain EEG features. In practice, the executing entity can use the short-time Fourier transform method to determine the signal power of the processed EEG signal in a specific frequency band, which can then be used as the EEG features.
[0028] Step three involves filtering and rectifying the surface electromyography (EMG) signal to obtain a processed EMG signal. In practice, the executing entity can perform bandpass filtering on the EMG signal to remove motion artifacts and high-frequency noise, resulting in a filtered EMG signal. Then, the filtered EMG signal is subjected to full-wave rectification to obtain the processed EMG signal. For example, the absolute value of the negative portion of the filtered EMG signal can be taken.
[0029] Step four: Determine the time-frequency characteristics of the processed surface electromyography (EMG) signal to obtain an EMG time-frequency feature set. In practice, the execution entity can determine the root mean square (RMS) and median frequency (MF) of the processed surface EMG signal within each sliding window as EMG time-frequency features to obtain the EMG time-frequency feature set. The sliding window can be a 150-millisecond time window, with each sliding step being 50 milliseconds.
[0030] Step 5: Extract features from the aforementioned electromyographic time-frequency feature set to obtain electromyographic features. In practice, the executing entity can determine the rate of change of the median frequency in the aforementioned electromyographic time-frequency feature set relative to the baseline reference value as a muscle fatigue factor. Then, the muscle fatigue factor and the aforementioned electromyographic time-frequency feature set are combined to form the electromyographic features. The aforementioned baseline reference value can be the root mean square value measured when the target user completes one maximum voluntary contraction (MVC) action in the initial stage of wearing the device. For example, it could be 0.8 mV.
[0031] Step six involves feature extraction from the aforementioned joint motion data to obtain joint motion features. In practice, the executing entity can extract the joint angle values at each moment in the joint motion data to obtain a joint angle value sequence, and then perform numerical differentiation on the joint angle value sequence to obtain a joint angular velocity value sequence. Subsequently, the joint angle value sequence and the joint angular velocity value sequence are identified as the joint motion features.
[0032] Step 103: Construct a state feature vector based on the current virtual reality scene complexity level, EEG characteristics, EMG characteristics, and joint motion characteristics.
[0033] In some embodiments, the aforementioned execution entity may construct a state feature vector based on the aforementioned current virtual reality scene complexity level, the aforementioned electroencephalogram (EEG) features, the aforementioned electromyogram (EMG) features, and the aforementioned joint motion features.
[0034] In some optional implementations of certain embodiments, the aforementioned execution entity may construct a state feature vector based on the aforementioned current virtual reality scene complexity level, the aforementioned electroencephalogram (EEG) features, the aforementioned electromyogram (EMG) features, and the aforementioned joint motion features through the following steps: Step one: Based on the aforementioned EEG features, generate an accuracy scalar. In practice, the executing entity can input these EEG features into a pre-trained classifier to obtain the matching probability between the EEG features and the target intent, which serves as the standard accuracy scalar. The pre-trained classifier can be trained using a large amount of EEG data and is used to represent the target intent through EEG feature representation. For example, the classifier could be a Support Vector Machine (SVM). The target intent could be a specific limb movement that the user expects to perform during rehabilitation training, generated through mental imagery or motor planning. For example, it could be "swing the left leg forward."
[0035] Step two: Based on the aforementioned electromyographic time-frequency feature set, determine the scalar value of the muscle activation level. In practice, the executing entity can extract the root mean square (RMS) value of the aforementioned EMG time-frequency feature set. Then, determine the percentage of the aforementioned RMS value relative to the RMS value at the target user's maximum voluntary contraction to obtain the scalar value of the muscle activation level.
[0036] Step 3: Based on the above electromyographic characteristics, muscle fatigue factors are generated. In practice, the aforementioned executing entity can extract muscle fatigue factors from the above electromyographic characteristics.
[0037] Step four: Based on the aforementioned joint motion characteristics, generate a motion smoothness index. In practice, the executing entity can determine the joint angular velocity in the aforementioned joint motion characteristics and the angular velocity difference between adjacent sampling points, obtaining an angular velocity difference sequence. Then, the absolute value of the angular velocity difference sequence is taken and averaged using a sliding window to obtain the angular velocity fluctuation amplitude. Finally, the angular velocity fluctuation amplitude is normalized to obtain the motion smoothness index.
[0038] Step 5: Based on the aforementioned current virtual reality scene complexity level, muscle activation level scalar, muscle fatigue factor, motion stability index, and accuracy scalar, a standardized feature set is generated. In practice, the executing entity can perform max-min normalization on the aforementioned muscle activation level scalar, muscle fatigue factor, motion stability index, and accuracy scalar to obtain a normalized feature set. Next, the current virtual reality scene complexity level is normalized to obtain normalized complexity features. Then, the normalized feature set and the normalized complexity features are merged into a standardized feature set. For example, if the current virtual reality scene complexity level is level 6, dividing the current virtual reality scene complexity level by 10 yields the normalized complexity features.
[0039] Step six: Concatenate the above standardized feature set to obtain the state feature vector. In practice, the executing entity can sort the standardized features in the above standardized feature set according to a preset order to obtain the state feature vector. As an example, the state feature vector can be obtained using the following formula: .
[0040] in, This represents the state feature vector. Represents a scalar measure of accuracy. This indicates the scalar level of muscle activation and the muscle fatigue factor. This represents a sequence of joint angle values. This represents a sequence of joint angular velocity values.
[0041] Step 104: Input the state feature vector into the preset reinforcement learning agent to obtain the control parameter set.
[0042] In some embodiments, the execution entity can input the state feature vector into a preset reinforcement learning agent to obtain a set of control parameters. The control parameters in the set include an auxiliary torque gain and a target virtual reality scene complexity level. The target virtual reality scene complexity level can be a classification label used to characterize the difficulty of the training task. The auxiliary torque gain can be a proportional amplification factor for the reinforcement learning agent to dynamically adjust the output torque of the exoskeleton motors.
[0043] The aforementioned pre-defined reinforcement learning agent can be a deep neural network model trained using a proximal policy optimization algorithm, used to output control parameters based on the user's real-time state during rehabilitation training. This pre-defined reinforcement learning agent can include a policy network and a value function network. The policy network takes the aforementioned state feature vector as input and the control parameter set as output. The policy network can consist of input layers, hidden layers, and an output layer. The input layer takes the aforementioned state feature vector as input. The hidden layer can consist of multiple fully connected layers, connected by non-linear activation functions (such as ReLU or Tanh). The hidden layer extracts higher-order non-linear relationships from the aforementioned state feature vector. The output layer outputs the control parameter set, including the sigmoid function and the softmax activation function. The output layer uses the sigmoid function to limit the auxiliary torque gain to between 0 and 1. The output layer uses the softmax activation function to output the probability distribution of various virtual reality scene complexity levels. The aforementioned value network can be used to provide attitude guidance for policy network updates during training and to measure the expected cumulative reward of the current state. This value network can include an input layer, hidden layers, and an output layer. The input layer of the value network can take the state feature vector as input. The hidden layers of the value network can consist of fully connected layers and activation functions, used to extract the intrinsic value representation of the state feature vector. The output layer of the value network outputs a value estimate of the current state. This value estimate can be a scalar numerical value.
[0044] Optionally, the parameters of the policy network can be updated using the loss function of the near-end policy optimization algorithm. This can be achieved using the following formula: .
[0045] in, This indicates the network parameters (such as weights and biases) that need to be optimized for the network using the above strategy. This represents the loss function. Indicates in At each time step, the average loss is calculated on the collected empirical data (training samples in the buffer). This indicates that the minimum value is returned. Indicates in The ratio of the probability that the new policy outputs the same action to the old policy at any given time (e.g., the ratio of the probability that the current policy network outputs the same control action to the probability that the policy network before the update outputs the same control action). Indicates in The estimated value of the dominance function at time 1. This function accepts three arguments. If the first argument is the maximum value among the three arguments, the third argument is returned. If the first argument is the minimum value among the three arguments, the second argument is returned. Otherwise, the first argument is returned. This represents the pruning boundary threshold, used to limit the ratio of the probabilities of the new and old strategies, and can be a constant value (such as 0.1 or 0.2). Indicates a time index.
[0046] Step 105: Generate rendering parameters based on the complexity level of the target virtual reality scene, and control the virtual reality device to render the scene based on the rendering parameters.
[0047] In some embodiments, the execution entity may generate rendering parameters based on the complexity level of the target virtual reality scene, and control the virtual reality device to perform scene rendering based on the rendering parameters.
[0048] In addressing the aforementioned technical problems in the process of adopting technical solutions, the application scenario—rehabilitation training for first-time users or patients prone to dizziness in a virtual reality environment—often presents the following technical challenges: During the generation of the virtual reality scene, the lack of real-time perception of the user's physiological tolerance leads to a disconnect between scene parameters and the user's physiological tolerance state, thereby reducing the user's safety during training. Considering the following requirements for this application scenario—high security, dynamic adjustment, and adaptability—we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity may generate rendering parameters based on the complexity level of the target virtual reality scene through the following steps, and control the virtual reality device to perform scene rendering based on the rendering parameters: Step one: Based on the aforementioned target virtual reality scene complexity level, generate a basic rendering parameter set. In practice, the executing entity can look up the parameters corresponding to the target virtual reality scene complexity level in a preset level parameter mapping table to obtain the basic rendering parameter set. This preset level parameter mapping table can be a mapping table used to characterize the correspondence between virtual reality scene complexity levels and specific physical rendering parameters. The basic rendering parameter set can be a set of values used to define the physical properties and interaction logic of objects in the virtual reality scene. The basic rendering parameters in the basic rendering parameter set may include, but are not limited to, angular velocity parameters, acceleration parameters, and field of view angle change rate.
[0049] Step two involves performing anomaly assessment on the aforementioned basic rendering parameter set to obtain an abnormal rendering parameter set. In practice, the execution entity can determine the basic rendering parameters as abnormal rendering parameters and obtain an abnormal rendering parameter set if the basic rendering parameters in the aforementioned basic rendering parameter set exceed a preset motion sickness induction threshold. The aforementioned preset motion sickness induction threshold can be a safety limit value used to determine whether virtual reality scene parameters will induce motion sickness in the user. The aforementioned preset motion sickness induction threshold may include an angular velocity threshold, an acceleration threshold, and a field of view angle change rate threshold.
[0050] Step three involves threshold correction and smoothing of the aforementioned abnormal rendering parameter set to obtain the processed rendering parameter set. In practice, the executing entity can correct each abnormal rendering parameter in the abnormal rendering parameter set to its maximum or minimum value within a safe threshold range, thus obtaining the corrected rendering parameter set. For example, if the angular velocity parameter is 25 degrees / second and the safe angular velocity range is [0, 20 degrees / second], the angular velocity parameter can be corrected to 20 degrees / second. Then, using an interpolation algorithm, each corrected rendering parameter in the corrected rendering parameter set is interpolated with the corresponding rendering parameters from the previous rendering cycle to obtain the processed rendering parameter set.
[0051] Step four: Based on the processed rendering parameter set and the target virtual reality scene complexity level, control the virtual reality device to perform scene rendering. In practice, the executing entity can send the processed rendering parameter set to the virtual reality device to generate the corresponding 3D image. Furthermore, based on the target virtual reality scene complexity level, generate corresponding task prompts and interactive objects in the virtual reality device.
[0052] Step 5: Acquire the physiological feedback signals of the target user in real time, and determine the physiological feedback level based on these signals. In practice, the executing entity can acquire the target user's EEG and EMG signals in real time as physiological feedback signals. Then, determine the frequency values of specific frequency bands (such as Theta waves) of the EEG signals and the root mean square (RMS) and median frequency of the EMG signals. Next, normalize the frequency values using the sigmoid function to obtain a cognitive stress score. Quantify the RMS and median frequency of the EMG signals to obtain a muscle risk score. For example, the percentage of the RMS value relative to the RMS value generated during the user's maximum voluntary contraction can be determined as the muscle activation rate, and the percentage of the median frequency relative to the median frequency at the beginning of training can be determined as the fatigue decline rate. The muscle activation rate and fatigue decline rate are then weighted and summed (weights can be set to 0.5 each) to obtain the muscle risk score. Finally, the cognitive stress score and the muscle risk score are weighted and summed to obtain the physiological feedback level. As an example, the weights of the cognitive stress score and the muscle risk score mentioned above can each be set to 0.5.
[0053] Step six: Based on the aforementioned physiological feedback level, dynamically fine-tune the processed rendering parameter set to obtain a fine-tuned rendering parameter set. Then, control the virtual reality device to perform scene rendering based on this fine-tuned rendering parameter set. In practice, the executing entity can dynamically fine-tune the processed rendering parameter set in response to the physiological feedback level exceeding a safety level (e.g., 0.7), obtaining a fine-tuned rendering parameter set, and then control the virtual reality device to perform scene rendering based on this fine-tuned rendering parameter set. As an example, if angular velocity is a potential factor causing motion sickness, the angular velocity can be reduced by 10% of the excess.
[0054] Steps one through six and related content described above constitute an inventive point of this disclosure, solving the technical problem of "the lack of real-time perception of user physiological tolerance during the generation of virtual reality scenes, leading to a disconnect between scene parameters and user physiological tolerance, and consequently reducing user safety during training." The reason for this reduced user safety during training is the lack of real-time perception of user physiological tolerance during the generation of virtual reality scenes, resulting in a disconnect between scene parameters and user physiological tolerance, thus reducing user safety during training. Solving this problem would resolve the issue of reduced user safety during training. To achieve this, the first step involves generating a basic rendering parameter set based on the aforementioned target virtual reality scene complexity level. This provides users at different training stages with an initial training environment that meets basic ability expectations, avoiding direct physiological shock due to excessively high starting difficulty. The second step involves anomaly evaluation of the basic rendering parameter set to obtain an abnormal rendering parameter set. This accurately identifies dangerous parameters that meet the difficulty level but induce vestibular conflict or visual fatigue. The third step involves threshold correction and smoothing transition processing of the abnormal rendering parameter set to obtain a processed rendering parameter set. This ensures that changes generated by the system remain within the user's physiological limits, while preventing dizziness caused by sudden visual changes. The fourth step involves controlling the virtual reality device to render the scene based on the processed rendering parameter set and the complexity level of the target virtual reality scene. This presents the user with a virtual reality interactive screen that meets training requirements and is within a safe range. The fifth step involves acquiring the physiological feedback signals of the target user in real time and determining the physiological feedback level based on these signals. This quantifies the user's current cognitive stress and muscle load. The sixth step involves dynamically fine-tuning the processed rendering parameter set based on the physiological feedback level to obtain a fine-tuned rendering parameter set. The virtual reality device is then controlled to render the scene according to this fine-tuned parameter set. This allows for timely adjustments to the screen when the user experiences discomfort. Ultimately, this ensures the user's safety during training.
[0055] Step 106: Generate exoskeleton joint torque commands based on the auxiliary torque gain, and control the exoskeleton to perform corresponding torque adjustment operations based on the exoskeleton joint torque commands.
[0056] In some embodiments, the aforementioned execution entity may generate exoskeleton joint torque commands based on the aforementioned auxiliary torque gain, and control the exoskeleton to perform corresponding torque adjustment operations based on the aforementioned exoskeleton joint torque commands.
[0057] In addressing the technical problems mentioned above, the application scenario—lower limb rehabilitation training for specific patients (such as those with sudden spasticity or motor incoordination)—often presents the following technical challenges: During exoskeleton-assisted movement, if the user experiences sudden involuntary tonic spasticity or loss of joint control, the control commands generated based on the spasticity signature will directly clash with the sudden abnormal torque generated by the user's muscles, posing a significant safety hazard during training. Considering the following requirements for this application scenario—high security, millisecond-level response, and multimodal anomaly fusion—we have decided to adopt the following solution: In some optional implementations of certain embodiments, the execution entity may generate exoskeleton joint torque commands based on the aforementioned auxiliary torque gain through the following steps, and control the exoskeleton to perform corresponding torque adjustment operations based on the aforementioned exoskeleton joint torque commands: Step one: Determine the basic joint torque command based on the aforementioned auxiliary torque gain and preset reference torque. In practice, the executing entity can determine the basic joint torque command by multiplying the aforementioned auxiliary torque gain and preset reference torque. The aforementioned preset reference torque can be the maximum reference torsional threshold output when the exoskeleton's auxiliary torque gain is 1. For example, it could be 30 N·m.
[0058] Step two: Based on the aforementioned basic joint torque command, acquire real-time joint motion data. In practice, the executing entity can acquire real-time joint motion data by sending the aforementioned basic joint torque command to the joint encoder of the multi-source acquisition device. This real-time joint motion data includes joint angles and joint angular velocities.
[0059] Step 3: Based on the aforementioned joint angles and angular velocities, determine the target joint torque. The target joint torque can be calculated as stiffness coefficient × joint angle + damping coefficient × joint angular velocity. The stiffness coefficient can be a preset value, for example, 100 N·m / radian. The damping coefficient can be 5 N·m·s / radian.
[0060] Step four: Based on the aforementioned real-time joint motion data, generate motion abnormality markers. In practice, the executing entity can generate a genuine motion abnormality marker if the joint angle of the real-time joint motion data is greater than 120 degrees or the joint angular velocity of the real-time joint motion is greater than 180 degrees / second. Conversely, it can generate a false motion abnormality marker if the joint angle of the real-time joint motion data is not less than 120 degrees or the joint angular velocity of the real-time joint motion is not greater than 180 degrees / second.
[0061] Step 5: Based on the acquired real-time electromyography (EMG) features and the aforementioned muscle fatigue factor, generate a muscle abnormality marker. In practice, the executing entity can generate a true muscle abnormality marker in response to the root mean square value of the real-time EMG features exceeding a preset spasticity threshold or the aforementioned muscle fatigue factor exceeding a preset fatigue value. Conversely, it can generate a false muscle abnormality marker in response to the root mean square value of the real-time EMG features not exceeding a preset spasticity threshold or the aforementioned muscle fatigue factor not exceeding a preset fatigue value. The preset spasticity threshold can be three times the root mean square value of the real-time EMG features from the previous second. The preset fatigue value can be 0.8.
[0062] Step six: Based on the aforementioned motion abnormality markers and muscle abnormality markers, generate a safety correction factor. In practice, the executing entity may set the safety correction factor to 0.5 in response to the aforementioned motion abnormality marker being true or the aforementioned muscle abnormality marker being true. In response to the aforementioned motion abnormality marker being false or the aforementioned muscle abnormality marker being false, the safety correction factor may be set to 1. The aforementioned safety correction factor may be a coefficient used to adjust the target joint torque, with a value range of [0, 1].
[0063] Step seven: Based on the aforementioned safety correction factor, dynamically adjust the target joint torque to obtain the exoskeleton joint torque command. In practice, the executing entity can determine the exoskeleton joint torque as the product of the target joint torque and the aforementioned safety correction factor, and then encapsulate the exoskeleton joint torque into a JSON-formatted exoskeleton joint torque command.
[0064] Step eight involves dynamically saturating the exoskeleton joint torque commands to obtain the final exoskeleton joint torque commands. The dynamic saturation constraint can be an operation used to limit the exoskeleton torque commands within the maximum output range allowed by the hardware (e.g., the exoskeleton).
[0065] Step nine: Based on the aforementioned final exoskeleton joint torque command, control the exoskeleton to perform the corresponding torque adjustment operation. In practice, the executing entity can send the aforementioned final exoskeleton joint torque command to the exoskeleton's drive motor controller to control the exoskeleton to perform the corresponding torque adjustment operation.
[0066] Steps one through nine above, and their related content, constitute an inventive point of this disclosure, solving the technical problem that "during exoskeleton-assisted movement, if the user suddenly experiences involuntary spasticity or loss of joint control, the control commands generated based on the spasticity signature's state characteristics will directly clash with the sudden abnormal torque generated by the user's muscles, leading to significant safety hazards during training." The reason for these safety hazards is that during exoskeleton-assisted movement, if the user suddenly experiences involuntary spasticity or loss of joint control, the control commands generated based on the spasticity signature's state characteristics will directly clash with the sudden abnormal torque generated by the user's muscles, resulting in significant safety hazards during training. Solving these factors can reduce the safety hazards during training. To achieve this, the first step is to determine the basic joint torque command based on the aforementioned auxiliary torque gain and preset baseline torque. This ensures that the initial output is within a safe range, avoiding excessive microadenomas at the start of training and laying a reliable foundation for subsequent safety corrections. The second step involves acquiring real-time joint motion data, including joint angles and angular velocities, based on the aforementioned basic joint torque commands. This allows for real-time monitoring of the user's limb movement changes. The third step determines the target joint torque based on the joint angles and angular velocities. This reduces the resistance forces generated during human-machine interaction. The fourth step generates motion abnormality markers based on the real-time joint motion data. This prevents safety hazards caused by the exoskeleton blindly following commands. The fifth step generates muscle abnormality markers based on the acquired real-time electromyographic characteristics and the aforementioned muscle fatigue factors. This allows for early identification of potential muscle strain risks through kinematic data analysis. The sixth step generates safety correction factors based on the motion abnormality markers and muscle abnormality markers. This enables multimodal information fusion decision-making, avoiding misjudgments and training interruptions. The seventh step dynamically adjusts the target joint torque based on the safety correction factors to obtain exoskeleton joint torque commands. This quickly releases the exoskeleton's restraining force on the limbs, reducing strain injuries to spastic muscles. Step 8: Dynamically saturate and limit the exoskeleton joint torque commands to obtain the final exoskeleton joint torque commands. This ensures that the exoskeleton moves within the maximum torque range allowed by the actuators. Step 9: Based on the final exoskeleton joint torque commands, control the exoskeleton to perform corresponding torque adjustment operations. This reduces safety hazards for users during training.
[0067] Optionally, after step 106, the above method further includes: Step one: Obtain the second set of physiological and motor signals of the target user. This second set of physiological and motor signals can be the data set of physiological responses and motor states collected by the multi-source data acquisition device after the target user has completed the action of the previous control cycle.
[0068] Step two: Based on the aforementioned second set of physiological and motor signals, determine the comprehensive reward signal.
[0069]
[0070] in, This indicates a comprehensive reward signal. Represents a scalar measure of accuracy. It indicates the characteristics of joint movement. This indicates a muscle fatigue factor. This indicates an index of motion smoothness. This represents the preset weighting coefficient corresponding to the accuracy scalar. This represents the preset weighting coefficients corresponding to the joint motion characteristics. This represents the preset weighting coefficient corresponding to the muscle fatigue factor. This represents the preset weighting coefficient corresponding to the motion smoothness index. Represents a logarithmic function. This represents the activation function. This represents an exponential function.
[0071] Step 3: Based on the optimization algorithm, the aforementioned comprehensive reward signal, and the aforementioned control parameter set, the aforementioned preset reinforcement learning agent is updated. In practice, the executing agent can combine the aforementioned comprehensive reward signal and the aforementioned control parameter set into training samples and place the training samples into a buffer. Subsequently, in response to the number of training samples in the aforementioned buffer reaching a preset batch size, the aforementioned preset reinforcement learning agent is updated through the following steps: First, the probability ratio between the new and old policies is determined using the aforementioned proximal policy optimization algorithm, and the parameters of the aforementioned policy network and the aforementioned value network are optimized using the aforementioned comprehensive reward signal.
[0072] The above embodiments of this disclosure have the following beneficial effects: the reinforcement learning-based exoskeleton control method of some embodiments of this disclosure can improve user safety during training. Specifically, the reason for the reduced user safety during training is that traditional solutions cannot dynamically adjust according to the user's real-time physiological state, causing the generated torque commands to be out of sync with the user's actual movement intentions and physiological tolerance. Furthermore, traditional solutions lack a closed-loop feedback mechanism and cannot respond to abnormal situations during training in real time, leading to reduced user safety during training. Based on this, the reinforcement learning-based exoskeleton control method of some embodiments of this disclosure first, in response to receiving a control request from the target user for the exoskeleton, obtains the target user's physiological and motion signal set and the current virtual reality scene complexity level. This allows for the acquisition of basic data for controlling exoskeleton movement. Second, the physiological and motion signal set is preprocessed and feature extracted to obtain EEG features, EMG features, and joint motion features. This removes environmental noise and interference from the original data. Third, based on the current virtual reality scene complexity level, the EEG features, the EMG features, and the joint motion features, a state feature vector is constructed. Therefore, multi-dimensional discrete features can be fused into a unified-dimensional state space representation. Next, the aforementioned state feature vector is input into a preset reinforcement learning agent to obtain a set of control parameters, including auxiliary torque gain and the complexity level of the target virtual reality scene. This yields device control parameters that match the user's current situation. Then, based on the aforementioned complexity level of the target virtual reality scene, rendering parameters are generated, and the virtual reality device is controlled to render the scene according to these parameters. This allows for the output of a suitable virtual image to the user while avoiding visual dizziness caused by sudden changes in the image. Finally, based on the aforementioned auxiliary torque gain, exoskeleton joint torque commands are generated, and the exoskeleton is controlled to perform corresponding torque adjustment operations according to these commands. Ultimately, this improves user safety during training.
[0073] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an exoskeleton control device based on reinforcement learning. These device embodiments are similar to... Figure 2 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0074] like Figure 2As shown, some embodiments of the reinforcement learning-based exoskeleton control device 200 include: an acquisition unit 201, a preprocessing and feature extraction unit 202, a construction unit 203, an input unit 204, a first generation unit 205, and a second generation unit 206. The acquisition unit 201 is configured to, in response to receiving a control request from a target user for the exoskeleton, acquire the target user's physiological and motion signal set and the current virtual reality scene complexity level; the preprocessing and feature extraction unit 202 is configured to preprocess and extract features from the physiological and motion signal set to obtain EEG features, EMG features, and joint motion features; the construction unit 203 is configured to construct a state feature vector based on the current virtual reality scene complexity level, the EEG features, the EMG features, and the joint motion features; the input unit 204 is configured to... The aforementioned state feature vector is input to a preset reinforcement learning agent to obtain a set of control parameters, wherein the control parameters in the set include an auxiliary torque gain and a target virtual reality scene complexity level; the first generation unit 205 is configured to generate rendering parameters based on the target virtual reality scene complexity level, and control the virtual reality device to perform scene rendering based on the rendering parameters; the second generation unit 206 is configured to generate exoskeleton joint torque commands based on the auxiliary torque gain, and control the exoskeleton to perform corresponding torque adjustment operations based on the exoskeleton joint torque commands.
[0075] It is understandable that the units described in the device 200 are related to the reference. Figure 2 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.
[0076] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0077] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0078] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0079] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 309, or installed from a storage device 308, or installed from a ROM 302. When the computer program is executed by the processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0080] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0081] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0082] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently without being assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: in response to receiving a control request from a target user for the exoskeleton, acquire the target user's physiological and motion signal set and the current virtual reality scene complexity level; preprocess and extract features from the physiological and motion signal set to obtain EEG features, EMG features, and joint motion features; construct a state feature vector based on the current virtual reality scene complexity level, the EEG features, the EMG features, and the joint motion features; input the state feature vector to a preset reinforcement learning agent to obtain a control parameter set, wherein the control parameters in the control parameter set include an auxiliary torque gain and a target virtual reality scene complexity level; generate rendering parameters based on the target virtual reality scene complexity level, and control the virtual reality device to perform scene rendering based on the rendering parameters; generate exoskeleton joint torque commands based on the auxiliary torque gain, and control the exoskeleton to perform corresponding torque adjustment operations based on the exoskeleton joint torque commands.
[0083] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0085] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a preprocessing and feature extraction unit, a construction unit, an input unit, a first generation unit, and a second generation unit. The names of these units do not necessarily limit the specific unit; for example, the acquisition unit may be described as "a unit that, in response to receiving a control request from a target user for the exoskeleton, acquires the target user's physiological and motion signal set and the current virtual reality scene complexity level."
[0086] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0087] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A reinforcement learning-based exoskeleton control method, applied to an exoskeleton control system, the exoskeleton control system comprising an exoskeleton and a virtual reality device interconnected, the method comprising: In response to receiving a control request from a target user for the exoskeleton, the system acquires the target user's physiological and motion signal set and the current virtual reality scene complexity level. The physiological and motor signal set is preprocessed and feature extracted to obtain EEG features, EMG features, and joint movement features. Based on the current virtual reality scene complexity level, the electroencephalogram (EEG) features, the electromyogram (EMG) features, and the joint motion features, a state feature vector is constructed; The state feature vector is input into a preset reinforcement learning agent to obtain a set of control parameters, wherein the control parameters in the set of control parameters include the auxiliary torque gain and the complexity level of the target virtual reality scene; Based on the complexity level of the target virtual reality scene, rendering parameters are generated, and based on the rendering parameters, the virtual reality device is controlled to perform scene rendering. Based on the auxiliary torque gain, exoskeleton joint torque commands are generated, and based on the exoskeleton joint torque commands, the exoskeleton is controlled to perform corresponding torque adjustment operations.
2. The method according to claim 1, wherein, The method further includes: Obtain the second set of physiological and motor signals of the target user; Based on the second set of physiological and motor signals, a comprehensive reward signal is determined; The preset reinforcement learning agent is updated based on the optimization algorithm, the comprehensive reward signal, and the control parameter set.
3. The method according to claim 1, wherein, The physiological and motor signals in the set of physiological and motor signals include electroencephalogram (EEG) signals, surface electromyography (EMG) signals, and joint motion data.
4. The method according to claim 3, wherein, The preprocessing and feature extraction of the physiological and motor signal set yields EEG features, EMG features, and joint movement features, including: The EEG signal is subjected to bandpass filtering and artifact removal to obtain the processed EEG signal; Feature extraction is performed on the processed EEG signal to obtain EEG features; The surface electromyography (EMG) signal is filtered and rectified to obtain the processed surface EMG signal. The time-frequency characteristics of the processed surface electromyography signal are determined to obtain the electromyography time-frequency feature set; Feature extraction is performed on the electromyographic time-frequency feature set to obtain electromyographic features; Feature extraction is performed on the joint motion data to obtain joint motion features.
5. The method according to claim 4, wherein, The step of constructing a state feature vector based on the current virtual reality scene complexity level, the electroencephalogram (EEG) features, the electromyogram (EMG) features, and the joint motion features includes: Based on the aforementioned EEG characteristics, an accuracy scalar is generated; Based on the electromyographic time-frequency feature set, a scalar value for muscle activation level is determined; Based on the electromyographic characteristics, muscle fatigue factors are generated. Based on the joint motion characteristics, a motion stability index is generated; A standardized feature set is generated based on the current virtual reality scene complexity level, the muscle activation level scalar, the muscle fatigue factor, the motion stability index, and the accuracy scalar. The standardized feature set is concatenated to obtain the state feature vector.
6. An exoskeleton control device based on reinforcement learning, comprising: The acquisition unit is configured to acquire, in response to receiving a control request from a target user for the exoskeleton, the target user’s set of physiological and motion signals and the current virtual reality scene complexity level; The preprocessing and feature extraction unit is configured to preprocess and extract features from the physiological and motor signal set to obtain electroencephalogram (EEG) features, electromyogram (EMG) features, and joint motion features. The construction unit is configured to construct a state feature vector based on the current virtual reality scene complexity level, the electroencephalogram (EEG) features, the electromyogram (EMG) features, and the joint motion features; The input unit is configured to input the state feature vector into a preset reinforcement learning agent to obtain a set of control parameters, wherein the control parameters in the set of control parameters include the auxiliary torque gain and the complexity level of the target virtual reality scene; The first generation unit is configured to generate rendering parameters based on the complexity level of the target virtual reality scene, and control the virtual reality device to perform scene rendering based on the rendering parameters. The second generation unit is configured to generate exoskeleton joint torque commands based on the auxiliary torque gain, and control the exoskeleton to perform corresponding torque adjustment operations based on the exoskeleton joint torque commands.
7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.