Robot control method and device, electronic equipment and storage medium
By using a fast-slow system for whole-body control, the first system decodes user instructions to generate target motion parameters and combines them with the robot's current position to generate control commands. This solves the problem of inconvenient interaction in complex scenarios in traditional robot control methods, and enables intuitive interaction and precise operation of the robot in different work scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional robot control methods struggle to achieve intuitive and convenient human-computer interaction in complex work scenarios, especially in full-body motion control, where they cannot simultaneously meet the requirements for coordination and fine operation of complex movements.
A whole-body control method based on a fast-slow system is adopted. The first system decodes the user's instruction information to generate target motion parameters, and combines the robot's current position to generate control commands. The training dataset and privileged information are used to train the model to improve the interaction performance.
It enables intuitive interactive control of robots in different work scenarios, improves the robot's ability to understand user intentions, and meets the needs of coordinating complex movements and performing precise operations.
Smart Images

Figure CN121777142A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, and more particularly to a robot control method, apparatus, electronic device, and storage medium. Background Technology
[0002] Humanoid robots possess a significant advantage in their anthropomorphic appearance, enabling them to perform human-like operations in human-dominated environments, thereby enhancing productivity. Traditional solutions rely on remote control for full-body robot control or single-action following, resulting in inconvenient and unintuitive interaction that struggles to adapt to complex robot work scenarios. Summary of the Invention
[0003] This disclosure provides a robot control method, apparatus, electronic device, and storage medium, which can improve the interactive performance of robots to adapt to different robot working scenarios.
[0004] A first aspect of this disclosure provides a robot control method, comprising: generating target motion parameters for at least one active part of the robot via a first system based on acquired user instruction information and at least one historical motion parameter of the robot; generating control commands via a second system based on the target motion parameters and the current position of at least one active part, wherein the operating frequency of the first system is lower than the operating frequency of the second system; and controlling the robot according to the control commands to adjust at least one active part of the robot according to the target motion parameters. In some embodiments of this disclosure, generating target motion parameters for at least one active part of the robot via a first system based on acquired user instruction information and at least one historical motion parameter of the robot includes: performing intent analysis based on the user instruction information and at least one historical motion parameter to determine a target vector, the target vector representing the action the user expects the robot to perform; and decoding the target vector to obtain the target motion parameters for at least one active part.
[0005] In some embodiments of this disclosure, intention analysis is performed on user instruction information and at least one historical motion parameter to determine a target vector, including: converting the user instruction information into a semantic vector; encoding at least one historical motion parameter to obtain at least one historical action vector corresponding to at least one historical motion parameter; and, based on the target codebook of the first system, determining the vector among the multiple vectors included in the target codebook that satisfies the first similarity condition with the semantic vector and the historical action vector as the target vector.
[0006] In some embodiments of this disclosure, generating control commands through a second system based on target motion parameters and the current position of at least one moving part includes: determining the current state of the robot based on the current position of at least one moving part; and generating control commands based on the offset between the current state of the robot and the state of the robot indicated by the target motion parameters.
[0007] In some embodiments of this disclosure, generating control commands based on the offset between the robot's current state and the robot's state indicated by the target motion parameters includes: determining the control operation to be performed for the robot to move from its current state to the robot's state indicated by the target motion parameters based on the robot's current state, the target mapping relationship, and the offset between the robot's current state, the target mapping relationship, and the robot's state indicated by the target motion parameters; and generating control commands as needed for the control operation to be performed.
[0008] In some embodiments of this disclosure, the method further includes: acquiring a training dataset, the training dataset including multiple first training data, the first training data including motion parameters corresponding to training actions and semantic information corresponding to training actions; updating at least one of the preset codebook, encoder and decoder of the first initial system according to the training dataset until the first initial system meets the first preset condition, thereby obtaining the first system.
[0009] In some embodiments of this disclosure, updating at least one of the preset codebook, encoder, and decoder of the first initial system based on the training dataset until the first initial system meets a first preset condition to obtain the first system includes: encoding the first training data using the encoder of the first initial system to obtain a first vector corresponding to the first training data; identifying a second vector among multiple vectors included in the preset codebook using the first initial system that satisfies a second similarity condition; decoding the second vector using the decoder of the first initial system to obtain reconstructed training data corresponding to the second vector; and updating at least one of the vectors included in the preset codebook of the first initial system, the encoding weight parameters of the encoder, and the decoding weight parameters of the decoder based on the difference between the first training data and the reconstructed training data until the first initial system meets the first preset condition to obtain the first system.
[0010] In some embodiments of this disclosure, the method further includes: acquiring a training dataset, the training dataset including multiple second training datasets, the second training datasets including motion parameters corresponding to training actions and semantic information corresponding to training actions; acquiring privileged information, the privileged information including at least one of the physical parameters of the robot's environment, interference information in a future time period, obstacle location information, and internal sensor data of the robot; and training a first model and a second model based on the training dataset and the privileged information to obtain a second system.
[0011] In some embodiments of this disclosure, training a first model and a second model to obtain a second system based on a training dataset and privileged information includes: performing reinforcement learning training on the first model using multiple second training data included in the training dataset and privileged information until the first model meets a second preset condition; training the second model based on multiple second training data included in the training dataset and simulation data of the trained first model until the second model meets a third preset condition to obtain the second system, wherein the simulation data of the trained first model is the mapping relationship between the state of the robot and the control commands output by the first model during the simulation process.
[0012] In some embodiments of this disclosure, obtaining the training dataset includes: collecting multiple raw motion capture data through a motion capture system; determining multiple third training data based on the differences between the raw motion capture data and the learning data obtained by learning from the raw motion capture data, wherein the multiple third training data includes multiple first training data and / or multiple second training data; and generating a training dataset based on the multiple third training data.
[0013] In some embodiments of this disclosure, determining multiple third training data based on the difference between the original motion capture data and the learning data obtained by learning from the original motion capture data includes: learning the original motion capture data through a reinforcement learning algorithm to obtain the learning data corresponding to the original motion capture data; determining multiple original motion capture data whose difference from the corresponding learning data is less than or equal to a preset threshold as multiple first motion capture data; and extracting the dynamic data of the multiple first motion capture data as multiple third training data.
[0014] A second aspect of this disclosure provides a robot control device, comprising: a processing module, configured to generate target motion parameters for at least one movable part of the robot via a first system based on acquired user instruction information and at least one historical motion parameter of the robot; generate control commands via a second system based on the target motion parameters and the current position of at least one movable part, wherein the operating frequency of the first system is less than the operating frequency of the second system; and control the robot according to the control commands so that at least one movable part of the robot adjusts according to the target motion parameters.
[0015] In some embodiments of this disclosure, the processing module is further configured to: perform intent analysis based on user instruction information and at least one historical motion parameter to determine a target vector, the target vector being used to represent the action that the user expects the robot to perform; and decode the target vector to obtain target motion parameters for at least one active part.
[0016] In some embodiments of this disclosure, the processing module is further configured to: convert user instruction information into semantic vectors; encode at least one historical motion parameter to obtain at least one historical action vector corresponding to at least one historical motion parameter; and, based on the target codebook of the first system, determine the vector among the multiple vectors included in the target codebook that satisfies the first similarity condition with the semantic vector and the historical action vector as the target vector.
[0017] In some embodiments of this disclosure, the processing module is further configured to: determine the current state of the robot based on the current position of at least one moving part; and generate control commands based on the offset between the current state of the robot and the state of the robot indicated by the target motion parameters.
[0018] In some embodiments of this disclosure, the processing module is further configured to: determine the control operation to be performed for the robot to move from its current state to the state indicated by the target motion parameters, based on the robot's current state, the target mapping relationship, and the offset between the robot's current state, the target mapping relationship, and the state indicated by the target motion parameters; and generate control commands as required by the control operation to be performed.
[0019] In some embodiments of this disclosure, the processing module is further configured to: acquire a training dataset, the training dataset including multiple first training data, the first training data including motion parameters corresponding to training actions and semantic information corresponding to training actions; update at least one of the preset codebook, encoder and decoder of the first initial system according to the training dataset until the first initial system meets the first preset condition, thereby obtaining the first system.
[0020] In some embodiments of this disclosure, the processing module is further configured to: encode the first training data using the encoder of the first initial system to obtain a first vector corresponding to the first training data; among the multiple vectors included in the preset codebook, determine the vector that satisfies the second similarity condition with the first vector as a second vector using the first initial system; decode the second vector using the decoder of the first initial system to obtain the reconstructed training data corresponding to the second vector; and update at least one of the vectors included in the preset codebook of the first initial system, the encoding weight parameters of the encoder, and the decoding weight parameters of the decoder according to the difference between the first training data and the reconstructed training data, until the first initial system satisfies the first preset condition to obtain the first system.
[0021] In some embodiments of this disclosure, the processing module is further configured to: acquire a training dataset, the training dataset including multiple second training datasets, the second training datasets including motion parameters corresponding to training actions and semantic information corresponding to training actions; acquire privileged information, the privileged information including at least one of the following: physical parameters of the robot's environment, interference information in a future time period, location information of obstacles, and internal sensor data of the robot; and train a first model and a second model based on the training dataset and the privileged information to obtain a second system.
[0022] In some embodiments of this disclosure, the processing module is further configured to: perform reinforcement learning training on the first model using multiple second training data and privileged information included in the training dataset until the first model meets a second preset condition; train the second model according to the multiple second training data included in the training dataset and the simulation data of the trained first model until the second model meets a third preset condition to obtain a second system, wherein the simulation data of the trained first model is the mapping relationship between the state of the robot and the control commands output by the first model during the simulation process.
[0023] In some embodiments of this disclosure, the processing module is further configured to: collect multiple raw motion capture data through a motion capture system; determine multiple third training data based on the differences between the raw motion capture data and the learning data obtained by learning from the raw motion capture data, wherein the multiple third training data includes multiple first training data and / or multiple second training data; and generate a training dataset based on the multiple third training data.
[0024] In some embodiments of this disclosure, the processing module is further configured to: learn the original motion capture data through a reinforcement learning algorithm to obtain the learning data corresponding to the original motion capture data; determine multiple original motion capture data whose difference from the corresponding learning data is less than or equal to a preset threshold as multiple first motion capture data; and extract the dynamic data of the multiple first motion capture data as multiple third training data.
[0025] A third aspect of this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described in the first aspect of this disclosure.
[0026] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the first aspect of this disclosure.
[0027] In summary, the robot control method proposed in this disclosure, based on the acquired user instruction information and multiple historical motion parameters of the robot, transforms (or "translates") the user-input instruction information through a first system, translating it into target motion parameters for at least one moving part of the robot. These target motion parameters reflect the state the user expects the robot to achieve, enabling the robot to understand the user's intention through the first system. The second system, based on the target motion parameters and the current position of at least one moving part, determines what control commands need to be issued to the robot at its current position to achieve the state indicated by the target motion parameters. Therefore, the above-mentioned solution, through the cooperation of the first and second systems, can adjust the robot's spatial posture according to the user's instruction information, improving the robot's human-machine interaction performance and meeting the needs of different robot working scenarios.
[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0030] Figure 1 A flowchart illustrating a robot control method provided in this embodiment of the present disclosure. Figure 1 ; Figure 2 A flowchart illustrating a robot control method provided in this embodiment of the present disclosure. Figure 2 ; Figure 3 A flowchart illustrating a robot control method provided in this embodiment of the present disclosure. Figure 3 ; Figure 4AA flowchart illustrating a robot whole-body control method based on a fast-slow system provided in this disclosure embodiment; Figure 4B This is a schematic diagram illustrating a training dataset construction process provided in an embodiment of the present disclosure; Figure 4C This is a schematic diagram of the training process of a slow system provided in an embodiment of the present disclosure; Figure 4D This is a schematic diagram illustrating the application process of a slow system provided in an embodiment of the present disclosure; Figure 4E This is a schematic diagram of the training process of a fast system provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of a robot control device provided in an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0031] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0032] Humanoid robots possess a significant advantage in their anthropomorphic appearance, enabling them to perform human-like operations in human-dominated environments, thereby improving productivity. Current methods for controlling the whole-body motion of humanoid robots include model-based control, which combines traditional methods with model prediction control (MPC) and model optimization; and reinforcement learning (RL)-based whole-body motion control strategies, which use reward functions to enable the robot to explore and acquire the ability to maintain its balance and perform prescribed actions.
[0033] Robot whole-body control can be divided into upper and lower limb decoupled control or whole-body unified control. In terms of control effect, upper and lower limb decoupled control has a faster response and more precise upper body operation, but it is difficult to achieve coordinated movements of complex whole-body motions. Whole-body unified control can meet the needs of coordinated whole-body motions, but the movements are more difficult and precise operations are hard to achieve. According to traditional robot control methods, it is usually easier to achieve the following of a single robot movement, but it is difficult to adapt to the complex working scenarios of robots.
[0034] In terms of control methods, robot mapping control is generally achieved through motion capture teleoperation, or robot movement is controlled by the speed of the remote control joystick. Full-body control via teleoperation can be implemented using traditional model-based algorithms, or reinforcement learning algorithms, with decoupled control of upper and lower limbs, or joint control of upper and lower limbs, such as in multi-action training. Single-action following involves collecting motion data through motion capture, redirecting it, and then setting rewards in reinforcement learning to make the robot follow this single sequence of actions. This allows the robot to complete pre-set actions while maintaining balance. However, these two methods have relatively complex interaction operations and are not convenient or intuitive enough. Therefore, a new, human-computer interaction-friendly full-body motion control training and deployment framework is needed to achieve intuitive and precise full-body robot control, meeting the needs of complex scenarios such as daily life or factory operations.
[0035] Therefore, to solve the above problems, this disclosure proposes a robot control method based on a fast-slow system for whole-body control. This method enables different levels of tasks to be calculated at different frequencies, ensuring that the robot can complete the most complex action instructions at the highest level, rather than simply following single actions, provided that computing power allows. This gives the robot strong interactive capabilities. After training, the robot itself can perform various actions and has strong generalization ability. Through the user's voice and text input, the robot can more easily achieve interactive control to complete various actions.
[0036] The specific details of this method are as follows.
[0037] Figure 1 A flowchart illustrating a robot control method provided in this embodiment of the present disclosure. Figure 1 .like Figure 1 As shown, the method may include the following steps.
[0038] Step 101: Based on the obtained user instruction information and at least one historical motion parameter of the robot, the first system generates target motion parameters for at least one active part of the robot.
[0039] In some embodiments, user instruction information can be user commands, such as voice commands, text commands, etc., which can be used to instruct the robot to perform corresponding actions, such as walking, running, raising its hand, etc. Optionally, the robot can receive user instruction information; for example, the robot can have an audio acquisition module and can use user voice commands.
[0040] In some embodiments, at least one historical motion parameter of the robot may be at least one historical sequence of actions of the robot, or parameters of at least one action historically performed by the robot, such as motion parameters of 20 historical frames. That is, motion parameters that can predict the next action to be performed by the robot based on user instructions and historical sequence of actions. The motion parameter may be a parameter of the movement of the robot's moving parts, wherein the moving parts may be, for example, adjustable components when the robot adjusts its spatial posture, such as at least one joint of a humanoid robot.
[0041] In some embodiments, the target motion parameters of at least one moving part of the robot can be the motion parameters of the next action that the robot needs to perform, such as the movement angle, movement direction, movement distance, etc. of the joints of a humanoid robot.
[0042] In some embodiments, generating target motion parameters for at least one active part of the robot through a first system based on the acquired user instruction information and at least one historical motion parameter of the robot includes: performing intent analysis based on the user instruction information and at least one historical motion parameter to determine a target vector, the target vector being used to represent the action that the user expects the robot to perform; and decoding the target vector to obtain the target motion parameters for at least one active part.
[0043] In some embodiments, intent analysis based on user instruction information and at least one historical motion parameter is performed to determine a target vector, including: converting the user instruction information into a semantic vector; encoding at least one historical motion parameter to obtain at least one historical action vector corresponding to the at least one historical motion parameter; and, based on the target codebook of the first system, determining the vector among the multiple vectors included in the target codebook that satisfies a first similarity condition with the semantic vector and the historical action vector as the target vector. In some embodiments, intent analysis based on user instruction information and at least one historical motion parameter may involve determining the state that the user expects the robot to achieve through the instruction information. The target motion parameter may be used to describe the state that the user expects the robot to achieve through the instruction information. Intent analysis may involve determining the action that the user expects the robot to perform. Optionally, if the target vector can be used to describe the action that the user expects the robot to perform, then the target motion parameter may be the state that the robot can achieve after performing the action corresponding to the target vector. For example, if the action corresponding to the target vector is raising the arm, then the target motion parameter may be that the robot's arm node is located at a first coordinate, the current angle of the robot's arm node is 35°, etc.
[0044] In some embodiments, the first system can be used to perform language processing on user instruction information. For example, when the user instruction information is a voice command, the system can perform text recognition, keyword extraction, word segmentation, semantic analysis, user intent analysis, etc. That is, the first system can analyze and understand the user's knowledge information, and after determining it, can convert the language-processed user instruction information into a semantic vector.
[0045] In some embodiments, the first system may include a target encoder, a target decoder, and a target codebook. The target encoder may encode multiple historical motion parameters to obtain historical motion vectors corresponding to each of the multiple historical parameters. The multiple historical motion parameters may be motion parameters corresponding to multiple historical actions performed by the robot. Each historical action may correspond to multiple motion parameters. For example, each action may correspond to multiple motion parameters of at least one moving part of the robot. That is, each historical motion parameter may be a set of historical motion parameters. A set of historical motion parameters may include motion parameters of at least one moving part of the robot.
[0046] In some embodiments, after obtaining the semantic vector and the historical action vector, the vector that satisfies the first similarity condition with the semantic vector and the historical action vector can be determined as the target vector from the target codebook of the first system. Then, the target decoder of the first system can decode the target vector to determine the target motion parameters corresponding to the target vector. Since the target vector has the highest similarity with the semantic vector and the historical action vector, the features of the target vector are most similar to the user instruction information and most related to the historical action sequence. Therefore, the motion parameters obtained by decoding the target vector best meet the user's expectations. Thus, the target motion parameters determined by the first system based on the user instruction information and multiple historical motion parameters can be the motion parameters of the action that the user expects the robot to perform, or in other words, the motion parameters that can realize the action instructed by the user.
[0047] Optionally, the first similarity condition can be the highest similarity, or the similarity greater than or equal to a similarity threshold, etc.
[0048] In some embodiments, optionally, determining the target vector based on the target codebook may involve determining action tokens based on the target codebook. The action tokens may also be an index sequence [i], where the action tokens may be identifiers of the target vector, such as the index value of the target vector. Then, the first system can determine the target vector based on the action tokens, and the target vector is then processed.
[0049] Step 102: Based on the target motion parameters and the current position of at least one moving part, generate control commands through the second system.
[0050] In some embodiments, the operating frequency of the first system is lower than that of the second system. For example, the first system may operate at a frequency of 10 Hz, and the second system may operate at a frequency of 50 Hz. In some embodiments, the target motion parameters may be parameter information that the robot cannot directly understand. The second system can generate control instructions that the robot can understand based on the target motion parameters, and then send the control instructions to the robot's underlying controller to control the robot to perform actions.
[0051] In some embodiments, generating control commands through the second system based on target motion parameters and the current position of at least one moving part includes: determining the current state of the robot based on the current position of at least one moving part; and generating control commands based on the offset between the current state of the robot and the state of the robot indicated by the target motion parameters.
[0052] In some embodiments, generating control commands based on the offset between the robot's current state and the robot's state indicated by target motion parameters includes: determining the control operation to be performed for the robot to move from its current state to the robot's state indicated by the target motion parameters, based on the robot's current state, the target mapping relationship, and the offset between the robot's current state, the target mapping relationship, and the robot's state indicated by the target motion parameters; and generating control commands as required by the control operation. In some embodiments, the robot's current state can be determined based on the current position of at least one moving part, in other words, the robot's current posture can be determined. Then, control commands can be generated based on the target motion parameters, the current state, and the target mapping relationship. Optionally, a second system can be trained to learn the target mapping relationship. Then, the second system can determine the corresponding control commands based on the robot's current state and the target motion parameters. The target mapping relationship can also be a state-action pair, where the state can be the robot's current state, and the action can be a decision output by the second system, i.e., a control command.
[0053] In some embodiments, the target motion parameters can indicate a target state that the robot is expected to reach, and the current position of at least one moving part of the robot can be used to indicate the current state of the robot. After determining the target state and the current state, it can be determined how to control the robot so that the robot can transition from the current state to the target state. The control operation to be performed is determined to move the robot from the current state to the state indicated by the target motion parameters. The control operation can be, for example, planning a path for the robot, such as planning a path for at least one joint of the robot. Then, control instructions can be generated based on the planned path.
[0054] In some embodiments, the method further includes: acquiring a training dataset, the training dataset including multiple second training datasets, the second training datasets including motion parameters corresponding to training actions and semantic information corresponding to training actions; acquiring privileged information, the privileged information including at least one of the physical parameters of the robot's environment, interference information in a future time period, obstacle location information, and internal sensor data of the robot; and training a first model and a second model based on the training dataset and the privileged information to obtain a second system.
[0055] In some embodiments, training a first model and a second model to obtain a second system based on a training dataset and privileged information includes: performing reinforcement learning training on the first model using multiple second training data included in the training dataset and privileged information until the first model meets a second preset condition; and training the second model based on multiple second training data included in the training dataset and simulation data of the trained first model until the second model meets a third preset condition to obtain the second system. The simulation data of the trained first model represents the mapping relationship between the robot's state and the control commands output by the first model during simulation. In some embodiments, optionally, a teacher-student model can be used to train the second system. In this case, the first model can be a teacher model, the second model can be a student model, and the trained second system can be a student model. During training, the teacher model can be trained first. Optionally, a reinforcement learning method can be used to train the teacher model.
[0056] In some embodiments, during training, the training data for the teacher model may include the aforementioned training dataset and privileged information. The training data may be data collected on training actions, as well as semantic information corresponding to those actions. Speech information can be used to describe the training actions, such as semantic information indicating actions like raising a hand or walking. Privileged information refers to information available during the training phase (in simulation) but which the actual robot cannot directly perceive or obtain in the real world. Examples include precise physical parameters of the environment (physical parameters of the robot's environment / physical parameters of the environment of the first model), interference predictions for a short period in the future (interference information within the future time period), the location of hidden objects (location information of obstacles), and internal force sensor readings (internal sensor data of the robot). Privileged information enables the teacher model to learn better-performing strategies in simulation.
[0057] In some embodiments, the training data of the first model may also include the self-information of the first model, or may include the self-information of the robot on which the first model is deployed. The self-information is information that the robot's actual body sensors can measure, such as joint angles, joint velocities, motor currents, inertial measurement unit (IMU) data (e.g., acceleration, angular velocity), foot contact force, etc.
[0058] In some embodiments, the first model can predict future keyframe actions and the actions to be tracked in the next frame based on the training dataset, and update the first model based on the difference between the predicted future keyframe actions and the actions to be tracked in the next frame and the preset future keyframe actions and the actions to be tracked in the next frame. Alternatively, the training dataset may include future keyframe actions and the actions to be tracked in the next frame, and the first model can be trained using the future keyframe actions and the actions to be tracked in the next frame in the training dataset.
[0059] In some embodiments, when the first model meets the second preset condition, the training of the first model can be terminated to obtain the trained first model. The second preset condition may be, for example, the number of falls and blows being greater than or equal to a preset number, or the performance of the first model meeting a preset target, etc., and there is no limitation on the disclosure of this.
[0060] In some embodiments, after the first model is trained, the second model can be trained using the first model, that is, the student model can be trained using the trained teacher model. Optionally, the second model can be trained based on the training dataset and the simulation data of the trained first model, wherein the training data can be the training dataset used to train the first model, or it can be a new training dataset. Similarly, the second model can predict future keyframe actions and the actions to be tracked in the next frame based on the training dataset, and update the second model based on the difference between the predicted future keyframe actions and the actions to be tracked in the next frame and the preset future keyframe actions and the actions to be tracked in the next frame. Alternatively, the training dataset may include future keyframe actions and the actions to be tracked in the next frame, and the second model can be trained using the future keyframe actions and the actions to be tracked in the next frame in the training dataset.
[0061] In some embodiments, the second model can be trained based on its own information, or the information of the robot with the second model deployed, and historical motion parameters. Optionally, the training process of the student model is essentially behavioral cloning. This involves collecting a large number of "state-action" pairs (state = prophetic history sequence + motion target, action = action output by the teacher's policy) when the teacher model executes the policy in simulation. This data is then used to train the student model's policy (e.g., using a neural network), enabling it to reproduce the teacher's actions based solely on the prophetic history sequence and the motion target. In other words, training the student model based on the robot's state data and corresponding output decision data during the simulation process can improve the training efficiency of the student model.
[0062] In some embodiments, when the second model meets a third preset condition, the training of the second model can be terminated to obtain the first system. The third preset condition may be, for example, the number of falls and blows being greater than or equal to a preset number, or the performance of the second model meeting a preset target, etc., and there is no limitation on the disclosure of this.
[0063] Step 103: Control the robot according to the control instructions so that at least one moving part of the robot is adjusted according to the target motion parameters.
[0064] In some embodiments, control commands can be sent to a lower-level controller, which can then control at least one moving part of the robot to adjust according to target motion parameters in accordance with the control commands, thereby adjusting the robot's spatial posture.
[0065] In summary, the embodiments of this disclosure, based on the acquired user instruction information and multiple historical motion parameters of the robot, transform (or "translate") the user-input instruction information through a first system, translating the user-input instruction information into target motion parameters for at least one moving part of the robot. These target motion parameters can reflect the state the user instruction information expects the robot to achieve. This allows the robot to understand the user's intention through the first system. The second system, based on the target motion parameters and the current position of at least one moving part, can determine what control commands need to be issued to the robot at its current position to control the robot and enable it to reach the state indicated by the target motion parameters. Therefore, the above-described solution of this disclosure, through the cooperation of the first and second systems, can adjust the robot's spatial posture according to the user's instruction information, improving the robot's human-computer interaction performance and meeting the needs of different robot working scenarios.
[0066] Figure 2 A flowchart illustrating a robot control method provided in this embodiment of the present disclosure. Figure 2.like Figure 2 As shown, based on Figure 1 The illustrated embodiment shows that the method includes the following steps.
[0067] Step 201: Obtain the training dataset.
[0068] In some embodiments, the training dataset includes multiple first training data, which include motion parameters corresponding to the training actions and semantic information corresponding to the training actions. Raw motion capture data can be obtained from the motion capture system, and then the raw motion capture data can be processed to obtain training data. The training dataset can then be determined based on the training data.
[0069] Step 202: Update at least one of the preset codebook, encoder, and decoder of the first initial system according to the training dataset until the first initial system meets the first preset condition, thus obtaining the first system. In some embodiments, the first initial system may adopt a large language model design, such as a Transformer-based architecture. The first initial system may include an encoder, a decoder, and a preset codebook. The preset codebook may include multiple vectors, and the multiple vectors included in the preset codebook may be preset.
[0070] In some embodiments, updating at least one of the preset codebook, encoder, and decoder of the first initial system based on the training dataset until the first initial system meets a first preset condition to obtain the first system includes: encoding the first training data using the encoder of the first initial system to obtain a first vector corresponding to the first training data; identifying a second vector whose similarity to the first vector satisfies a second similarity condition among a plurality of vectors included in the preset codebook using the first initial system; decoding the second vector using the decoder of the first initial system to obtain reconstructed training data corresponding to the second vector; and updating at least one of the vectors included in the preset codebook of the first initial system, the encoding weight parameters of the encoder, and the decoding weight parameters of the decoder based on the difference between the first training data and the reconstructed training data until the first initial system meets the first preset condition to obtain the first system.
[0071] In some embodiments, when training the first system, the encoder of the first initial system can encode the training data to obtain a first vector corresponding to the training data.
[0072] In some embodiments, the first initial system can determine the similarity between the first vector and each of the multiple vectors included in the preset codebook, and then determine the vector whose similarity to the first vector satisfies the second similarity condition as the second vector.
[0073] Optionally, the first similarity condition can be the highest similarity, similarity greater than or equal to a similarity threshold, etc., in which case the action patterns of the second vector and the first vector have a high similarity.
[0074] In some embodiments, the decoder of the first initial system can decode the second vector. In other words, the decoder of the first initial system can use the second vector to attempt to reconstruct the training data, or to attempt to reconstruct the actions corresponding to the training data, to obtain the reconstructed training data corresponding to the second vector.
[0075] In some embodiments, the difference between training data and reconstructed training data can be determined. Then, based on the difference between the training data and the reconstructed training data, at least one of the vectors included in the preset codebook of the first initial system, the encoding weight parameters of the encoder, and the decoding weight parameters of the decoder can be updated until the first initial system satisfies the first preset condition to obtain the first system. After updating the vectors included in the preset codebook, the target codebook can be obtained. After updating the encoding weight parameters of the encoder and the decoding weight parameters of the decoder, the trained encoder and the trained decoder can be obtained.
[0076] In other words, during training, the encoder of the first initial system compresses a short action segment into a continuous latent vector z, essentially compressing training data into a first vector. It then searches the codebook for the closest prototype vector e_i to z and uses its index i as the action token. The decoder of the first initial system attempts to reconstruct the original action segment using the selected e_i. The entire process is trained through reconstruction loss and a codebook update mechanism. The goal is to learn a high-quality codebook and encoder / decoder that reconstructs the action using the discrete index sequence [i] by searching the codebook and using the decoder, making it as close as possible to the original action. Thus, the discrete token sequence (i.e., action tokens) [tokens] faithfully represents the original action.
[0077] In summary, the above embodiments of this application can train the first initial system based on the training dataset to obtain the first system. The first system can generate target motion parameters of at least one active part of the robot based on the acquired user instruction information and multiple historical motion parameters of the robot. This enables the robot to understand the user instruction information through the first system and convert the user instruction information into motion parameters of the robot's active parts.
[0078] Figure 3 A flowchart illustrating a robot control method provided in this embodiment of the present disclosure. Figure 3 .like Figure 3 As shown, based on Figure 1 The illustrated embodiment shows that the method includes the following steps.
[0079] Step 301: Collect multiple raw motion capture data using a motion capture system.
[0080] In some embodiments, multiple raw motion capture data can be collected by a motion capture system. For example, raw motion capture data can be collected when a worker wears a motion capture suit or optical markers and performs a target action (such as walking, grabbing, or jumping) in a specified scene.
[0081] Step 302: Based on the differences between the original motion capture data and the learning data obtained by learning from the original motion capture data, determine multiple third training data.
[0082] In some embodiments, the plurality of third training data includes a plurality of first training data and / or a plurality of second training data.
[0083] In some embodiments, determining multiple third training data based on the difference between the original motion capture data and the learning data obtained by learning from the original motion capture data includes: learning the original motion capture data through a reinforcement learning algorithm to obtain the learning data corresponding to the original motion capture data; determining multiple original motion capture data whose difference from the corresponding learning data is less than or equal to a preset threshold as multiple first motion capture data; and extracting the dynamic data of the multiple first motion capture data as multiple third training data. In some embodiments, the collected original motion capture data has high noise and contains some actions that cannot be completed under dynamic constraints (such as climbing, jumping, etc.). To address this issue, an original motion capture data filtering process is designed. Optionally, a filtering model can be used to filter the original motion capture data. For example, the filtering model can be a Physics-based Constrained Handling (PHC) model, that is, for a certain action, the PHC model is used for training and evaluation, and if the evaluation index is found to be too high, the action is removed.
[0084] In some embodiments, the filtering model can use reinforcement learning to learn each original motion capture data to obtain the learning data corresponding to each original motion capture data, that is, it can be trained and evaluated through the PHC model.
[0085] In some embodiments, the difference between each original motion capture data and the corresponding learning data can be determined. Optionally, the difference between each original motion capture data and the corresponding learning data can be determined as the above-mentioned evaluation index. That is, when the difference between each original motion capture data and the corresponding learning data is large, it can be determined that the original motion capture data does not meet the dynamic constraints or has too much noise. In this case, the original motion capture data can be removed.
[0086] In some embodiments, the first motion capture data may be data with low noise, data that satisfies dynamic constraints, or data that satisfies other conditions. That is, the original motion capture data with low noise and that satisfies dynamic constraints can be determined as the first motion capture data, which can achieve filtering of the original motion capture data.
[0087] In some embodiments, after filtering the original motion capture data to obtain multiple first motion capture data, dynamic data of the multiple first motion capture data can be extracted. Dynamic data includes contact force and torque, trajectory data, momentum and energy, etc. Then, the dynamic data of the multiple first motion capture data can be used as multiple training data, and a training dataset can be generated based on the multiple training data.
[0088] In other words, multiple qualified first motion capture data can be determined from the collected raw motion capture data. Then, in the simulation environment, the dynamic properties related to the multiple first motion capture data, such as the time the feet are in the air and the position of the force application, can be extracted to provide a relatively clean supervision signal for the subsequent training of the first and second systems.
[0089] Step 303: Generate a training dataset based on multiple third-party training data.
[0090] In some embodiments, after determining multiple third training data, a training dataset can be generated based on the multiple third training data. For example, all the determined third training data can be combined into a training dataset, or the third training data with higher training performance among the multiple third training data can be used to generate a training dataset, etc.
[0091] In summary, the above embodiments of this disclosure can filter the original motion capture data and determine the training dataset based on the multiple first motion capture data obtained from the filtering. Through the above process, a feasible training dataset under robot dynamics or structural constraints can be finally obtained, ensuring the quality of subsequent training and learning.
[0092] The technical solutions of this disclosure will be further described in detail below with reference to specific application embodiments.
[0093] The following is a robot whole-body control method based on a fast-slow system provided by an embodiment of this disclosure. Current methods for controlling the whole body of a robot have two main shortcomings: 1. Insufficient generalization ability: Current robot control methods usually rely on reinforcement learning to reinforce a single action. When a new action is encountered, the robot needs to be retrained. Our proposed solution can complete multiple actions with a single strategy, and does not require retraining when a new action is needed.
[0094] 2. Weak interactive teaching capability: Current methods for controlling the whole body of robots usually require operators to wear motion capture suits to remotely control the robot in practical applications, which is time-consuming and labor-intensive; our proposed solution supports users to directly control the robot's movement through voice and text input.
[0095] Therefore, the solution in this example is mainly used for full-body robot control, controlling the robot to perform corresponding actions through user voice and text input. The entire system is as follows: Figure 4A As shown: The system includes a slow system for processing user commands, a fast system for processing motion markers, and the robot body (including the robot's underlying controller).
[0096] The user inputs relevant commands into the slow system, which operates at a frequency of 10 Hz. After inference, the slow system generates corresponding action tokens, which are then fed into the fast system. The fast system operates at a frequency of 50 Hz and generates the specific joint angles to be executed by the robot, which are ultimately sent to the robot's underlying controller. Finally, the robot can move according to the user's commands at the top level. Through the combination of the fast and slow systems, stable control of the robot's entire body is achieved. The proposed robot system will be introduced in three parts: training dataset construction, slow system training, and fast system training. 1. Dataset construction system.
[0097] The training dataset construction process is as follows Figure 4B As shown: (1) The raw training data is obtained through a motion capture system. The raw motion capture data has a lot of noise and contains some actions that cannot be completed under dynamic constraints (such as climbing and jumping). To address this issue, a Physics-based Constrained Handling (PHC) model filtering process was designed. That is, for a certain action, the PHC model is used for training and evaluation. If the evaluation index is found to be too high, the action is removed.
[0098] (2) For qualified actions, relevant dynamic attributes such as the time the feet are in the air and the position of force are extracted in the simulation environment to provide a relatively clean supervision signal for subsequent model training.
[0099] After the above process, a feasible training dataset under robot dynamics or structural constraints is finally obtained, ensuring the quality of subsequent training and learning.
[0100] 2. Slow system The training process of a slow system is as follows Figure 4C As shown: The slow system mainly draws on the design of the large language model and adopts a Transformer-based architecture, which mainly consists of two parts: an action tokenizer (which discretizes continuous actions) and an action generation model.
[0101] The action tokenizer model consists of an encoder, a decoder, and a codebook. The encoder compresses a short action segment into a continuous latent vector z. Vector Quantization (VQ) finds the closest prototype vector e_i (the most similar action pattern) to z in the codebook and uses its index i as the token. The decoder attempts to reconstruct the original action segment using the selected e_i. The entire process is trained using reconstruction loss and a codebook update mechanism. The goal is to learn a high-quality codebook and encoder / decoder that reconstructs the action using the discrete index sequence [i] by looking up the codebook and using the decoder, making it as close as possible to the original action. In this way, the discrete token sequence [tokens] faithfully represents the original action.
[0102] Action generation models, such as Figure 4D As shown, this model takes the user's text commands and historical action sequences as input and uses an autoregressive approach (i.e., predicting the next tag one by one) to generate the corresponding humanoid action tags.
[0103] The data source for the action marker and action generation model is the aforementioned dynamically feasible action library.
[0104] 3. Fast System The training process of a fast system is as follows Figure 4E As shown: The fast system is constructed using the teacher-student training framework commonly used in robot reinforcement training. During training, the teacher model receives the following inputs: 1. Privileged information: This refers to information available during the training phase (in simulation) but which the actual robot cannot directly perceive or acquire in the real world. Examples include: precise physical parameters of the environment, predictions of disturbances in the near future, the positions of hidden objects, and internal force sensor readings. This information allows the teacher to learn better-performing strategies in the simulation. 2. Self-information: This is information that the robot's actual body sensors can measure, such as joint angles, joint velocities, motor currents, IMU (Inertial Measurement Unit) data (acceleration, angular velocity), and foot contact forces. 3. Future keyframe actions and the actions to be tracked in the next frame. During training, the teacher model undergoes reinforcement training.
[0105] The training inputs for the student model include: 1. Self-information; 2. Future keyframe actions and the action to be tracked in the next frame; 3. Observations from the previous 20 frames. The process of training the student policy is essentially behavioral cloning: collecting a large number of "state-action" pairs (state = proprioceptive history sequence + moving target, action = action output by the teacher's policy) when the teacher executes the policy in the simulation, and then using this data to train the student policy (e.g., using a neural network) so that it can reproduce the teacher's actions based solely on the proprioceptive history sequence and the moving target.
[0106] In summary, the examples disclosed above can be used for full-body control of robots. Robots can perform a variety of complex movements with limited training datasets. Human-computer interaction can be achieved by directly controlling robots through language and text, resulting in better human-computer interaction performance. Full-body interactive control of robots can be achieved by combining different frequency systems, saving computing power and control costs.
[0107] Figure 5 This is a schematic diagram of the structure of a robot control device 500 provided in an embodiment of this disclosure. Figure 5 As shown, the device includes: a processing module 510, configured to generate target motion parameters for at least one moving part of the robot through a first system based on the acquired user instruction information and at least one historical motion parameter of the robot; generate control commands through a second system based on the target motion parameters and the current position of at least one moving part, wherein the operating frequency of the first system is lower than the operating frequency of the second system; and control the robot according to the control commands so that at least one moving part of the robot is adjusted according to the target motion parameters.
[0108] In some embodiments, the processing module is further configured to: perform intent analysis based on user instruction information and at least one historical motion parameter to determine a target vector, the target vector being used to represent the action that the user expects the robot to perform; and decode the target vector to obtain target motion parameters for at least one active part.
[0109] In some embodiments, the processing module is further configured to: convert user instruction information into semantic vectors; encode at least one historical motion parameter to obtain at least one historical action vector corresponding to at least one historical motion parameter; and, based on the target codebook of the first system, determine the vector among the multiple vectors included in the target codebook that satisfies the first similarity condition with the semantic vector and the historical action vector as the target vector.
[0110] In some embodiments, the processing module is further configured to: determine the current state of the robot based on the current position of at least one moving part; and generate control commands based on the offset between the current state of the robot and the state of the robot indicated by the target motion parameters.
[0111] In some embodiments, the processing module is further configured to: determine the control operation to be performed for the robot to move from its current state to the state indicated by the target motion parameters, based on the robot's current state, the target mapping relationship, and the offset between the robot's current state, the target mapping relationship, and the state indicated by the target motion parameters; and generate control commands as required by the control operation.
[0112] In some embodiments, the processing module is further configured to: acquire a training dataset, the training dataset including multiple first training data, the first training data including motion parameters corresponding to training actions and semantic information corresponding to training actions; update at least one of the preset codebook, encoder and decoder of the first initial system according to the training dataset until the first initial system meets the first preset condition, thereby obtaining the first system.
[0113] In some embodiments, the processing module is further configured to: encode the first training data using the encoder of the first initial system to obtain a first vector corresponding to the first training data; among the multiple vectors included in the preset codebook, determine the vector that satisfies the second similarity condition with the first vector as a second vector using the first initial system; decode the second vector using the decoder of the first initial system to obtain the reconstructed training data corresponding to the second vector; and update at least one of the vectors included in the preset codebook of the first initial system, the encoding weight parameters of the encoder, and the decoding weight parameters of the decoder according to the difference between the first training data and the reconstructed training data, until the first initial system satisfies the first preset condition to obtain the first system.
[0114] In some embodiments, the processing module is further configured to: acquire a training dataset, the training dataset including multiple second training datasets, the second training datasets including motion parameters corresponding to training actions and semantic information corresponding to training actions; acquire privileged information, the privileged information including at least one of the physical parameters of the robot's environment, interference information in a future time period, obstacle location information, and internal sensor data of the robot; and train the first model and the second model according to the training dataset and the privileged information to obtain the second system.
[0115] In some embodiments, the processing module is further configured to: perform reinforcement learning training on the first model using multiple second training data and privileged information included in the training dataset until the first model meets the second preset condition; and train the second model according to the multiple second training data included in the training dataset and the simulation data of the trained first model until the second model meets the third preset condition to obtain the second system, wherein the simulation data of the trained first model is the mapping relationship between the state of the robot and the control commands output by the first model during the simulation process.
[0116] In some embodiments, the processing module is further configured to: acquire multiple raw motion capture data through the motion capture system; determine multiple third training data based on the differences between the raw motion capture data and the learning data obtained by learning from the raw motion capture data, wherein the multiple third training data includes multiple first training data and / or multiple second training data; and generate a training dataset based on the multiple third training data.
[0117] In some embodiments, the processing module is further configured to: learn the original motion capture data through a reinforcement learning algorithm to obtain the learning data corresponding to the original motion capture data; determine multiple original motion capture data whose difference from the corresponding learning data is less than or equal to a preset threshold as multiple first motion capture data; and extract the dynamic data of the multiple first motion capture data as multiple third training data.
[0118] In summary, the robot control device 500 can, based on the acquired user instruction information and multiple historical motion parameters of the robot, transform (or "translate") the user-input instruction information through a first system, translating the user-input instruction information into target motion parameters for at least one moving part of the robot. These target motion parameters can reflect the state the user instruction information expects the robot to achieve. This allows the robot to understand the user's intention through the first system. The second system, based on the target motion parameters and the current position of at least one moving part, can determine what control commands need to be issued to the robot at its current position to control the robot and achieve the state indicated by the target motion parameters. Therefore, the above-mentioned solution of this disclosure, through the cooperation of the first and second systems, can adjust the robot's spatial posture according to the user's instruction information, improving the robot's human-machine interaction performance and meeting different robot working scenarios.
[0119] The methods and apparatus provided in the embodiments of this application have been described above. To implement the functions of the methods provided in the embodiments of this application, the electronic device may include a hardware structure and software modules, and may implement the above functions in the form of a hardware structure, software modules, or a hardware structure plus software modules. One of the above functions may be executed in the form of a hardware structure, software modules, or a hardware structure plus software modules.
[0120] Figure 6 This is a block diagram illustrating an electronic device 600 for implementing the above-described method according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, computer, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0121] In embodiments of this disclosure, the electronic device may be a robot, or the electronic device may be a controller deployed on the robot body, or the electronic device may be a server, such as a cloud server, an edge computing unit, etc.
[0122] Reference Figure 6 The electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0123] Processing component 602 typically controls the overall operation of electronic device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.
[0124] Memory 604 is configured to store various types of data to support the operation of electronic device 600. Examples of such data include instructions for any application or method operating on electronic device 600, contact data, phonebook data, messages, pictures, videos, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0125] Power supply component 606 provides power to various components of electronic device 600. Power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 600.
[0126] Multimedia component 608 includes a screen that provides an output interface between electronic device 600 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When electronic device 600 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0127] Audio component 610 is configured to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) configured to receive external audio signals when electronic device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.
[0128] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0129] Sensor assembly 614 includes one or more sensors for providing state assessments of various aspects of electronic device 600. For example, sensor assembly 614 may detect the on / off state of electronic device 600, the relative positioning of components such as the display and keypad of electronic device 600, changes in position of electronic device 600 or a component of electronic device 600, the presence or absence of user contact with electronic device 600, orientation or acceleration / deceleration of electronic device 600, and temperature changes of electronic device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0130] Communication component 616 is configured to facilitate wired or wireless communication between electronic device 600 and other devices. Electronic device 600 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio), or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0131] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0132] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of an electronic device 600 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0133] Embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the above embodiments of this disclosure.
[0134] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0135] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in at least one embodiment or example.
[0136] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0137] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having at least one wiring (control method), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0138] It should be understood that various parts of the embodiments of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0139] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0140] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc.
[0141] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A robot control method, characterized in that, The method includes: Based on the acquired user instruction information and at least one historical motion parameter of the robot, the first system generates target motion parameters for at least one moving part of the robot. Based on the target motion parameters and the current position of the at least one active part, a control command is generated by the second system, wherein the operating frequency of the first system is lower than the operating frequency of the second system. The robot is controlled according to the control instructions so that at least one moving part of the robot is adjusted according to the target motion parameters.
2. The method according to claim 1, characterized in that, The step of generating target motion parameters for at least one moving part of the robot through the first system based on the acquired user instruction information and at least one historical motion parameter of the robot includes: Based on the user instruction information and the at least one historical motion parameter, intent analysis is performed to determine a target vector, which represents the action that the user expects the robot to perform. The target vector is decoded to obtain the target motion parameters of the at least one moving part.
3. The method according to claim 2, characterized in that, The step of performing intent analysis based on the user instruction information and the at least one historical motion parameter to determine the target vector includes: Convert the user instruction information into a semantic vector; Encode the at least one historical motion parameter to obtain at least one historical action vector corresponding to the at least one historical motion parameter; Based on the target codebook of the first system, the vector among the multiple vectors included in the target codebook that satisfies the first similarity condition with the semantic vector and the historical action vector is determined as the target vector.
4. The method according to claim 1, characterized in that, The step of generating control commands through the second system based on the target motion parameters and the current position of at least one moving part includes: The current state of the robot is determined based on the current position of the at least one active part; The control command is generated based on the offset between the robot's current state and the robot's state indicated by the target motion parameters.
5. The method according to claim 4, characterized in that, The step of generating the control command based on the offset between the current state of the robot and the state of the robot indicated by the target motion parameters includes: determining the control operation to be performed when the robot moves from its current state to the state of the robot indicated by the target motion parameters based on the current state of the robot, the target mapping relationship, and the offset between the current state of the robot, the target mapping relationship, and the state of the robot indicated by the target motion parameters; The control instructions are generated based on the control operations that need to be performed.
6. The method according to claim 2, characterized in that, The method further includes: Obtain a training dataset, which includes multiple first training datasets, each of which includes motion parameters corresponding to a training action and semantic information corresponding to the training action. The first initial system is updated based on the training dataset, including at least one of the preset codebook, encoder, and decoder, until the first initial system meets the first preset condition, thus obtaining the first system.
7. The method according to claim 6, characterized in that, The step of updating at least one of the preset codebook, encoder, and decoder of the first initial system according to the training dataset until the first initial system satisfies the first preset condition to obtain the first system includes: The encoder of the first initial system encodes the first training data to obtain the first vector corresponding to the first training data; Among the multiple vectors included in the preset codebook, the first initial system determines the vector that satisfies the second similarity condition with the first vector as the second vector; The second vector is decoded using the decoder of the first initial system to obtain the reconstructed training data corresponding to the second vector; Based on the difference between the first training data and the reconstructed training data, at least one of the vectors included in the preset codebook of the first initial system, the encoding weight parameters of the encoder, and the decoding weight parameters of the decoder is updated until the first initial system meets the first preset condition, thus obtaining the first system.
8. The method according to claim 4, characterized in that, The method further includes: Obtain a training dataset, which includes multiple second training datasets, each of which includes motion parameters corresponding to a training action and semantic information corresponding to the training action. Obtain privileged information, which includes at least one of the following: physical parameters of the environment in which the robot is located, interference information in the future time period, location information of obstacles, and internal sensor data of the robot; The first model and the second model are trained based on the training dataset and the privileged information to obtain the second system.
9. The method according to claim 8, characterized in that, The step of training the first model and the second model based on the training dataset and the privileged information to obtain the second system includes: The first model is trained using reinforcement learning using multiple second training data included in the training dataset and the privileged information until the first model meets the second preset condition; The second model is trained based on multiple second training data included in the training dataset and the simulation data of the trained first model until the second model meets the third preset condition, thus obtaining the second system. The simulation data of the trained first model is the mapping relationship between the state of the robot and the control commands output by the first model during the simulation process.
10. The method according to any one of claims 6 to 9, characterized in that, The acquisition of the training dataset includes: Multiple raw motion capture data are collected through a motion capture system; Based on the difference between the original motion capture data and the learning data obtained by learning from the original motion capture data, a plurality of third training data are determined, wherein the plurality of third training data includes a plurality of first training data and / or a plurality of second training data. The training dataset is generated based on the plurality of third training data.
11. The method according to claim 10, characterized in that, The step of determining multiple third training data based on the difference between the original motion capture data and the learning data obtained by learning from the original motion capture data includes: learning the original motion capture data through a reinforcement learning algorithm to obtain the learning data corresponding to the original motion capture data; Among the multiple original motion capture data, the multiple original motion capture data whose difference from the corresponding learning data is less than or equal to a preset threshold are determined as multiple first motion capture data; The dynamic data of the multiple first motion capture data are extracted as the multiple third training data.
12. A robot control device, characterized in that, include: Processing module, used for: Based on the acquired user instruction information and at least one historical motion parameter of the robot, the first system generates target motion parameters for at least one moving part of the robot. Based on the target motion parameters and the current position of the at least one active part, a control command is generated by the second system, wherein the operating frequency of the first system is lower than the operating frequency of the second system. The robot is controlled according to the control instructions so that at least one moving part of the robot is adjusted according to the target motion parameters.
13. The apparatus according to claim 12, characterized in that, The processing module is also used for: Based on the user instruction information and the at least one historical motion parameter, intent analysis is performed to determine a target vector, which represents the action that the user expects the robot to perform. The target vector is decoded to obtain the target motion parameters of the at least one moving part.
14. The apparatus according to claim 13, characterized in that, The processing module is also used for: Convert the user instruction information into a semantic vector; Encode the at least one historical motion parameter to obtain at least one historical action vector corresponding to the at least one historical motion parameter; Based on the target codebook of the first system, the vector among the multiple vectors included in the target codebook that satisfies the first similarity condition with the semantic vector and the historical action vector is determined as the target vector.
15. The apparatus according to claim 12, characterized in that, The processing module is also used for: Obtain a training dataset, which includes multiple first training datasets, each of which includes motion parameters corresponding to a training action and semantic information corresponding to the training action. The first initial system is updated based on the training dataset, including at least one of the preset codebook, encoder, and decoder, until the first initial system meets the first preset condition, thus obtaining the first system.
16. The apparatus according to claim 15, characterized in that, The processing module is also used for: The encoder of the first initial system encodes the first training data to obtain the first vector corresponding to the first training data; Among the multiple vectors included in the preset codebook, the first initial system determines the vector that satisfies the second similarity condition with the first vector as the second vector; The second vector is decoded using the decoder of the first initial system to obtain the reconstructed training data corresponding to the second vector; Based on the difference between the first training data and the reconstructed training data, at least one of the vectors included in the preset codebook of the first initial system, the encoding weight parameters of the encoder, and the decoding weight parameters of the decoder is updated until the first initial system meets the first preset condition, thus obtaining the first system.
17. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
18. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.