Machine animal control method, controller and machine animal

By setting up a camera on the machine animal to obtain scene information, using speed and posture models to predict and plan joint data, and combining reinforcement learning algorithms, the problem of insufficient decision-making ability of the machine animal in natural scenes is solved, achieving higher decision-making ability and flexibility.

CN115327917BActive Publication Date: 2025-07-29ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211055024.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-07-29
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

In the prior art, motor learning of robotic animals can only achieve low-coupled motion control tasks and cannot show intelligent decision-making capabilities similar to animals in natural scenarios.

Method used

By setting up a camera on the machine animal to obtain scene information, using the speed model and pose model to predict, planning the joint data of the machine animal to achieve its decision-making and motion control in natural scenarios, and training in combination with reinforcement learning algorithms to improve the adaptability and robustness of the model.

Benefits of technology

It improves the decision-making ability and flexibility of machine animals in various scenarios and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115327917B_ABST
    Figure CN115327917B_ABST
Patent Text Reader

Abstract

The present application provides a method for controlling a robotic animal, a controller, and a robotic animal. The method includes: the controller obtains first scene information through a camera disposed on the robotic animal. The controller may input the first scene information and device information into the speed model respectively, and predict the desired speed of the robotic animal. The controller inputs the desired speed into the pose model to implement the planning of predicted joint data when the robotic animal moves to the desired speed. The predicted joint data includes the angle data of each joint of the robotic animal at the next moment. The controller may control each joint of the robotic animal to move to the angle corresponding to the predicted joint data at the next moment, so as to realize its overall movement. The method of the present application improves the decision-making ability of the robotic animal in various scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular, to a method for controlling a robotic animal, a controller, and a robotic animal. Background Art

[0002] With the rapid development of artificial intelligence technology, significant progress has also been made in decision-making control methods implemented through artificial intelligence. Decision-making control is an important step for robotic animals to achieve intelligence.

[0003] Currently, the decision-making control of robotic animals is mainly achieved through reinforcement learning algorithms. In terms of improving the motion learning ability of robotic animals, reinforcement learning algorithms have broken through the barrier that traditional control methods require long-term experience accumulation, and improved the adaptability and robustness of the algorithms.

[0004] However, the motion learning of robotic animals can only achieve low-coupling motion control tasks, and cannot achieve intelligent decision-making of robotic animals in natural scenarios, resulting in poor decision-making ability. Summary of the Invention

[0005] The present application provides a method for controlling a robotic animal, a controller, and a robotic animal to solve the problem of poor decision-making ability of robotic animals in the prior art.

[0006] In a first aspect, the present application provides a method for controlling a robotic animal, including:

[0007] Obtaining first scene information through a camera disposed on the robotic animal;

[0008] Using a speed model, predicting an expected speed according to the first scene information and the device information of the robotic animal;

[0009] Planning predicted joint data according to the pose model of the robotic animal and the expected speed;

[0010] Controlling each joint of the robotic animal according to the predicted joint data to make the robotic animal move.

[0011] Optionally, the step of using a speed model to predict an expected speed according to the first scene information and the device information of the robotic animal specifically includes:

[0012] Inputting the first scene information into a visual extraction module of the speed model for feature extraction to obtain visual features;

[0013] Inputting the device information into a device information extraction module of the speed model for data processing to generate device features;

[0014] Combining the visual features and the device features to obtain a first fusion feature;

[0015] Input the first fusion feature into the speed prediction module of the speed model to predict the desired speed.

[0016] Optionally, the method further includes:

[0017] Obtain the second scene information and the real speed of the real animal during movement through a camera and sensors arranged on the head of the real animal;

[0018] Use the second scene information and the real speed to train the speed model of the robotic animal, and the speed model is used to predict the desired speed according to the first scene information obtained by the camera of the robotic animal.

[0019] Optionally, the method further includes:

[0020] Obtain the real speed and real joint data of the real animal during movement through motion sensors arranged at the joints of the real animal;

[0021] Use the real speed and the real joint data to train the pose model of the robotic animal, and the pose model is used to generate the predicted joint data of the robotic animal according to the predicted desired speed.

[0022] Optionally, the step of using the real speed and the real joint data to train the pose model of the robotic animal specifically includes:

[0023] Map the real joint data according to the dynamic mapping to obtain the first feature;

[0024] Perform feature fusion on the first feature and the real speed to obtain the second feature;

[0025] Input the second feature and the real speed into the pose model to plan and obtain the predicted joint data.

[0026] In a second aspect, the present application provides a control device for a robotic animal, including:

[0027] An acquisition module, configured to obtain first scene information through a camera arranged on the robotic animal;

[0028] A processing module, configured to use a speed model to predict a desired speed according to the first scene information and the device information of the robotic animal; plan and obtain predicted joint data according to the pose model and the desired speed of the robotic animal;

[0029] A control module, configured to control each joint of the robotic animal according to the predicted joint data so that the robotic animal moves.

[0030] Optionally, the processing module is specifically configured to:

[0031] Input the first scenario information into the visual extraction module of the speed model for feature extraction to obtain visual features;

[0032] Input the device information into the device information extraction module of the speed model for data processing to generate device features;

[0033] Merge the visual features and the device features to obtain a first fusion feature;

[0034] Input the first fusion feature into the speed prediction module of the speed model to predict and obtain the expected speed.

[0035] Optionally, the device further includes:

[0036] A model training module, configured to obtain second scenario information and the real speed of a real animal during movement through a camera and sensors arranged on the head of the real animal;

[0037] Use the second scenario information and the real speed to train the speed model of the robotic animal, where the speed model is used to predict the expected speed according to the first scenario information obtained by the camera of the robotic animal.

[0038] Optionally, the model training module is further configured to:

[0039] Obtain the real speed and real joint data of a real animal during movement through motion sensors arranged at the joints of the real animal;

[0040] Use the real speed and the real joint data to train the pose model of the robotic animal, where the pose model is used to generate the predicted joint data of the robotic animal according to the predicted expected speed.

[0041] Optionally, the model training module is specifically configured to:

[0042] Map the real joint data according to the kinetic mapping to obtain a first feature;

[0043] Perform feature fusion on the first feature and the real speed to obtain a second feature;

[0044] Input the second feature and the real speed into the pose model to plan and obtain the predicted joint data.

[0045] In a third aspect, the present application provides a controller, including: a memory and a processor;

[0046] The memory is used to store a computer program; the processor is used to execute the machine animal control method in the first aspect and any possible design of the first aspect according to the computer program stored in the memory.

[0047] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When at least one processor of a controller executes the computer program, the controller executes the machine animal control method in the first aspect and any possible design of the first aspect.

[0048] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When at least one processor of a controller executes the computer program, the controller executes the machine animal control method in the first aspect and any possible design of the first aspect.

[0049] The machine animal control method, controller, and machine animal provided by the present application obtain first scene information through a camera disposed on the machine animal; input the first scene information and device information into the speed model respectively, and predict the expected speed of the machine animal; input the expected speed into the attitude model to implement the planning of predicted joint data when the machine animal moves to the expected speed. The predicted joint data includes the angle data of each joint of the machine animal at the next moment; according to the predicted joint data, control each joint of the machine animal to move to the angle corresponding to the predicted joint data at the next moment, so as to achieve the overall movement means, improve the decision-making ability of the machine animal in various scenarios, improve the flexibility of the machine animal in the interaction process, and improve the user experience. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a system block diagram of a control system for a machine animal provided by an embodiment of the present application;

[0052] Figure 2 It is a flowchart of a machine animal control method provided by an embodiment of the present application;

[0053] Figure 3 It is a schematic structural diagram of a speed model provided by an embodiment of the present application;

[0054] Figure 4Flowchart of a method for controlling a robotic animal provided by an embodiment of the present application;

[0055] Figure 5 Schematic diagram of the installation of a sampling sensor provided by an embodiment of the present application;

[0056] Figure 6 Schematic diagram of the training process of a speed model provided by an embodiment of the present application;

[0057] Figure 7 Schematic diagram of the architecture of an attitude model provided by an embodiment of the present application;

[0058] Figure 8 Schematic diagram of the training process of an attitude model provided by an embodiment of the present application;

[0059] Figure 9 Schematic diagram of model training provided by an embodiment of the present application;

[0060] Figure 10 Schematic diagram of the structure of a robotic animal control device provided by an embodiment of the present application;

[0061] Figure 11 Schematic diagram of the hardware structure of a controller provided by an embodiment of the present application. Detailed implementation manners

[0062] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0063] The terms "first", "second", "third", "fourth", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances. For example, without departing from the scope of this text, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information.

[0064] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0065] Furthermore, as used in this text, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context indicates otherwise.

[0066] It should be further understood that the terms "comprising" and "including" indicate the presence of features, steps, operations, elements, components, items, kinds, and / or groups, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups.

[0067] As used herein, the terms "or" and "and / or" are interpreted inclusively and mean any one or any combination. Thus, "A, B, or C" or "A, B, and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B, and C". An exception to this definition occurs only when the combination of elements, functions, steps, or operations is inherently mutually exclusive in some manner.

[0068] In recent years, with the rapid development of artificial intelligence technology, artificial intelligence algorithms such as machine learning, deep learning, and reinforcement learning have made great progress and achieved remarkable application results in fields such as finance, gaming, consumer electronics, industrial control, autonomous driving, and robotic mobile platforms. Among them, the field of robotic mobile platforms is an important carrier for humans to aspire to general intelligence. With the rapid development of artificial intelligence, the general mobile platform has also been greatly improved in terms of intelligence. The general mobile platform can include the movement of robots, the movement of autonomous driving vehicles, etc. A robot is an intelligent machine that can work semi-autonomously or fully autonomously. This application mainly focuses on the movement control of quadruped robots such as robotic cats and robotic dogs. This quadruped robot is the machine animal used in this application. Currently, the field of the mobile platform of machine animals is an important branch in recent years. Especially after leading research institutions and companies such as Boston Dynamics and MIT have made great progress in motion control, a new wave of robot research has swept the globe. Recently, some companies or institutions have also begun to pay attention to using reinforcement learning algorithms to solve related technical problems, such as quadruped motion control, end-to-end motion planning control, etc. In the field of robot movement control, decision-making control is an important step towards realizing intelligence.

[0069] At present, the decision-making control of robotic animals is mainly achieved through reinforcement learning algorithms. The basic mechanism of reinforcement learning algorithms is as follows: when an agent interacts with the environment through actions, the environment returns the current reward to the agent, and the agent then evaluates the actions taken based on the current reward. Among these actions, the actions that are beneficial to achieving the goal are retained, and the actions that are not conducive to achieving the goal are attenuated. Among them, the agent can be a robot that can work semi-autonomously or fully autonomously. This agent also includes the main robotic animal in this application. This working mechanism is similar to the brain activities of animals or humans. Scientists at the British artificial intelligence laboratory DeepMind believe that intelligence and its related capabilities are not generated by forming and solving complex problems, but rather stem from the long-term adherence to the principle of "reward maximization". The implementation of reinforcement learning is precisely based on the principle of reward maximization. Therefore, reinforcement learning can effectively improve the motion learning ability of robotic animals in decision-making control, break through the barrier that traditional decision-making control methods require long-term experience accumulation, improve the adaptability and robustness of the algorithm, and provide strong technical support for the large-scale application of robots.

[0070] The above technological developments have established an effective foundation for the general artificial intelligence of robots moving across platforms. However, along with the increasing demand for robots to move across platforms or intelligent companion assistants, consumers' requirements for the intelligence level of robots moving across platforms are constantly rising. How to make robots moving across platforms exhibit simple intelligence similar to that of animals in conventional scenarios is a feasible and realistic exploration, and it is also an important technical support for enhancing product competitiveness and promoting the large-scale application of general mobile platforms. However, in the existing technology, most current research and invention patents focus on aspects such as motion control, simple point-to-point motion planning, and reducing the gap between simulation and real environments. A thinking model for robots has not been established, and a brain-like learning model for robot motion planning and control in conventional scenarios has not been completed. The motion learning of robotic animals can only achieve low-coupling motion control tasks. Fortunately, in terms of simple quadruped motion control algorithms, the existing technology has already had good accumulation, and the simple walking of robotic animals is no longer a very difficult problem. Whether through traditional optimization methods or the emerging reinforcement learning methods in recent years, robotic animals can achieve the expected quadruped motion control effect. However, how to imitate the thinking of animals, establish a thinking model for robotic animals, and improve the intelligent decision-making ability of robotic animals in natural scenarios is an urgent problem to be solved.

[0071] In view of the above problems, the present application proposes a method for controlling a robotic animal. Based on previous research and solutions, the present application uses the robotic animal as a specific implementation carrier, adopts a reinforcement learning algorithm, integrates the motion control and scene recognition of the robot, and establishes a brain-like learning model for the robot, enabling the robot to exhibit behavior patterns similar to those of animals in conventional scenarios in daily life. For example, when the robotic animal observes the owner throwing a frisbee, it can run over to chase the frisbee. Moreover, the model set in the present application has an online learning function, that is, it can continuously expand the types of adaptable scenarios, and finally achieve the effect of being able to complete most of the life scenarios and having the thinking ability of an ordinary animal. In addition, this solution can also be applied to robots on other general mobile platforms.

[0072] In the present application, the controller establishes a complete brain-like learning model framework. This framework can input the collected and processed scene data and animal action data into the model for training together. After the training is completed, the model can be deployed to the real machine of the controller of the robotic animal. When the robotic animal is actually running, it can make decisions on relevant behaviors based on the observed scene data. The scene data used in the present application is specifically image data. The present application uses these scene data for the training of reinforcement learning. By using an optimized traditional reinforcement learning algorithm framework, the present application enables the image data to be used as the input for guiding behaviors, and this kind of guidance is not reflected in the traditional obstacle avoidance guidance, but in the extraction of scene information to guide the motion gait and motion direction of the robot. The present application divides the entire brain-like learning model framework into two parts: a motion strategy and a speed planning strategy. During the model training process, the models corresponding to these two strategies can be trained separately to reduce the difficulty of model training and improve the training speed. The training framework of the present application is different from most of the integrated training frameworks, but decomposes the training task into two small training modules. The division of this module greatly improves the stability and training speed of the training, and provides support for quickly training model parameters when expanding scenarios later. The use of this brain-like learning model not only focuses on the algorithm design of motion control, enabling the robotic animal to have a thinking model similar to that of animals, but also can have an intimate interaction with humans similar to that of animals, and even can assist humans in completing more complex tasks, achieving the effects of accompanying humans and deeply participating in life scenarios.

[0073] The technical solution of the present application will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0074] Figure 1 The system block diagram of a control system for a robotic animal provided by an embodiment of the present application is shown. As Figure 1As shown in the figure, the control system of the robotic animal in this application includes two main strategies: a motion strategy and a speed planning strategy. Among them, the motion strategy can obtain the joint data that changes over time during the movement of a real animal by installing relevant sensors on the real animal. The joint data can be leg joint data. The leg joint data can include hip joint data, thigh joint data, calf joint data, etc. The movement of the real animal can include walking, running, jumping, etc. The controller can use the joint data of the real animal during movement to train the pose model in the motion strategy. The controller can use the trained pose model to generate predicted joint data. The controller can use the predicted joint data to control the movement of the robotic animal so that the robotic animal has a gait similar to that of a real animal during movement. Among them, after obtaining the scene data, the speed planning strategy can combine the scene data with the device information of the robotic animal to obtain the expected speed of the robotic animal's movement in the current scene. After obtaining the expected speed, the controller can input the expected speed into the motion strategy to achieve the behavior of the robotic animal guided in a specific scene and to guide the robotic animal to reach the expected speed. The scene data can be the image information of the scene. For example, the scene can be a person throwing a frisbee. The expected speed can include the longitudinal speed and the yaw angle. In addition, this application can continuously expand its corresponding expected speed according to the scene data, so that the brain-like model of the robot is getting closer and closer to the real animal model and can adapt to most conventional scenes.

[0075] In this application, with the controller as the execution entity, the control method of the robotic animal in the following embodiments is executed. Specifically, the execution entity can be the hardware device of the controller, or the software application in the controller that implements the following embodiments, or the computer-readable storage medium installed with the software application that implements the following embodiments, or the code of the software application that implements the following embodiments.

[0076] Figure 2 The flowchart of a control method for a robotic animal provided by an embodiment of this application is shown. In Figure 1 Based on the embodiment shown, as Figure 2 shown, with the controller set in the robotic animal as the execution entity, after obtaining the completed training model, the controller can control the robotic animal to execute corresponding actions according to the current scene. The method of this embodiment can include the following steps:

[0077] S101. Obtain the first scene information through the camera set on the robotic animal.

[0078] In this embodiment, a camera may be provided on the head of the robotic animal, and the camera may collect first scene information. The camera may send the first scene information to the controller of the robotic animal. Wherein, the first scene may be image information or video information of the current scene captured by the camera of the robotic animal.

[0079] S102. Use the speed model to predict the expected speed according to the first scene information and the device information of the robotic animal.

[0080] In this embodiment, a speed model that has been trained may be stored in the controller. The speed model is used to predict the expected speed according to the first scene information and the device information of the robotic animal. After obtaining the first scene information, the controller may input the first scene information and the device information into the speed model respectively. The speed model may predict the expected speed of the robotic animal in the first scene. Wherein, for different first scenes, the speed model may predict different expected speeds. For example, the expected speeds predicted in the frisbee-throwing scene and the walking scene are different. For different device information, the speed model may predict different expected speeds. For example, when the preset maximum speeds in the device information are 20 m / s and 30 m / s, different expected speeds may be predicted in the frisbee-throwing scene to enable the movement of the robotic animal when catching the frisbee.

[0081] In one example, the expected speed may include an expected horizontal and vertical speed and an expected angular speed.

[0082] In one example, the model architecture of the speed prediction model may be as Figure 3 shown, and the process of the controller using the speed module to predict the expected speed may specifically include the following steps:

[0083] Step 1. The controller inputs the first scene information into the visual extraction module of the speed model for feature extraction to obtain visual features. Before the controller inputs the first scene information into the speed model, the controller may also preprocess the first scene information to extract the depth image information in the first scene information. Alternatively, the first scene information obtained by the controller may be image information including depth information acquired by a depth camera. The controller may input the depth image information corresponding to the first scene information into the visual extraction module. The visual extraction module may specifically be a trained Convolutional Neural Networks (CNN) model. The controller may use the CNN model to perform feature extraction on the first scene information to obtain visual features.

[0084] Step 2: The controller inputs the device information into the device information extraction module of the speed model for data processing to generate device features. Among them, the device information may include an Inertial Measurement Unit (IMU), joint angles, and the previous control action. The controller can input the IMU, joint angles, and the previous control action in the device information into the device information extraction module of the speed model. The device information extraction module can specifically be a Multi-Layer Perception (MLP) neural network model. The multi-layer fully connected neural network model is a multi-layer perceptron neural network. The controller can use the MLP network of the device information extraction module to process the device information to obtain device features. Among them, the IMU in the device information specifically includes information in four dimensions: yaw angle, pitch angle, yaw angular velocity, and pitch angular velocity. Among them, the joint angles include a multi-dimensional data composed of the angle data of each joint. For example, when the robotic animal includes 12 joints (the robotic animal includes four limbs, and each limb includes 3 joints), the joint angles can be a 12-dimensional data. The previous control action includes the angle data of each key part of the robotic animal during the previous movement. The data dimension of the previous control action is the same as the data dimension of the joint angles. For example, when the robotic animal includes 12 joints, the previous control action is a 12-dimensional data. The device information input by the controller may include three sets of device information at three moments that have been obtained. For example, the device information input by the controller can be an 84-dimensional data. The dimension of the device features obtained after the device information is processed by the MLP network can be determined according to a preset dimension.

[0085] Step 3: The controller combines the visual features and the device features to obtain a first fusion feature. The combination process can specifically include connecting the feature vectors of the device features and the visual features to obtain a new feature vector. For example, when the dimensions of both the device features and the visual features are 64 dimensions, the controller can combine them to obtain a 128-dimensional first fusion feature.

[0086] Step 4: The controller inputs the first fusion feature into the speed prediction module of the speed model to predict the desired speed. The speed prediction module can be another MLP network in the speed model. The MLP network of the speed prediction module and the MLP network of the device information extraction module may have different weights and parameters. The server can input the combined first fusion feature into the MLP network of the speed prediction module. The MLP network of the speed prediction module can process the first fusion feature to obtain the desired speed. The desired speed can be the running speed of the robotic animal within the allowable range of the device information in the first scenario information.

[0087] S103. Plan and obtain predicted joint data based on the pose model and desired speed of the robotic animal.

[0088] In this embodiment, the controller may store a pose model that has been trained. After obtaining the desired speed, the controller may input the desired speed into the pose model. The pose model may plan the angle data of each joint of the robotic animal when it moves to reach the desired speed according to the desired speed. The angle data of all joints of the robotic animal constitutes the predicted joint data predicted by the controller. For example, when the robotic animal includes 12 joints, the predicted joint data is a 12-dimensional data.

[0089] S104. Control the robotic animal to obtain each joint according to the predicted joint data, so that the robotic animal moves.

[0090] In this embodiment, the predicted joint data obtained by the controller includes the angles of each joint of the robotic animal at the next moment. The controller may control each joint of the robotic animal to move to the angle corresponding to the predicted joint data at the next moment according to the predicted joint data. The robotic animal may realize its overall movement through the movement of these joints.

[0091] For the robotic animal control method provided in this application, the controller obtains the first scene information through a camera disposed on the robotic animal. The controller may input the first scene information and device information into the speed model respectively, and predict the desired speed of the robotic animal. The controller inputs the desired speed into the pose model to realize the planning of the predicted joint data when the robotic animal moves to reach the desired speed. The predicted joint data includes the angle data of each joint of the robotic animal at the next moment. The controller may control each joint of the robotic animal to move to the angle corresponding to the predicted joint data at the next moment, so as to realize its overall movement. In this application, by using the speed model and the pose model, the prediction of the angle data of each key joint of the robotic animal at the next moment is realized, so that the robotic animal can execute corresponding decisions according to the current scene, improving the decision-making ability of the robotic animal in various scenes, improving the flexibility of the robotic animal in the interaction process, and improving the user experience.

[0092] Figure 4 shows a flowchart of a robotic animal control method provided in an embodiment of this application. Based on Figures 1 to 3 the embodiment shown, this embodiment also needs to train the speed model and the pose model. As Figure 4 shown, with the controller as the execution subject, the method of this embodiment may include the following steps:

[0093] S201. Obtain the second scene information and the real speed of the real animal during movement through the cameras and sensors arranged on the head of the real animal.

[0094] In this embodiment, the controller can obtain training data through the cameras and sensors arranged on various parts of the real animal's body. Specifically, the installation positions of the cameras and sensors can be as Figure 5 shown. Among them, the camera can be arranged on the head of the real animal. The centroid motion state sensor can be arranged at the centroid of the real animal's body. For example, the centroid of the body can be at the waist and abdomen of a dog. The centroid motion sensor can be strapped to the waist and abdomen of the dog through a binding strap to obtain the motion state. The joint sensors can be arranged at the limb joints of the real animal. For example, as Figure 5 shown are the installation positions of three joint sensors on one leg of a dog. The key sensors can be respectively arranged at 12 joints of the four limbs of the dog.

[0095] Since the speed model is mainly used to determine its expected speed according to the current scene. Therefore, the training data of the speed model can include the second scene information and the real speed of the real animal during movement. Among them, the second scene information can be obtained through the camera arranged on the head of the real animal. The real speed can be obtained through the centroid motion state sensor arranged at the centroid of the real animal's body. For example, in the scenario of throwing a frisbee, the second scene information can include the image information of the frisbee appearance recorded by the camera. Since the process of the dog discovering and chasing the frisbee is a continuous process, the second scene information can be the image information at any moment during the entire process from the moment the camera captures the frisbee to the moment the dog touches the frisbee. During the process of catching the frisbee, the controller can continuously obtain multiple pieces of second scene information in time series, and the real speed corresponding to each piece of second scene information.

[0096] S202. Use the second scene information and the real speed to train the speed model of the robotic animal, and the speed model is used to predict the expected speed according to the first scene information obtained by the camera of the robotic animal.

[0097] In this embodiment, after the controller obtains enough second scene information and the real speed corresponding to each piece of second scene information, it can use the second scene information and the real speed to train the speed model. During the training process of the speed model, the real speed can be used as the target speed of the training model. The target speed is used to calculate the reward function for training the speed model.

[0098] The training process of the speed model can be specifically as Figure 6As shown in the figure. After the controller collects sufficient data of the second scenario, it can use a convolutional neural network to extract visual features. The controller can also fuse the preset device information of the robotic animal with the visual features. Among them, the movement speed of the robotic animal is affected by the device itself. Therefore, the controller also needs to obtain the device information of the robotic animal to which the speed model will be applied. And for a robotic animal, the device information is the fixed parameter information of the robotic animal. Therefore, after the controller determines the robotic animal that needs to use the speed model, the controller can directly obtain the device information. The controller can use the visual features fused with the device information for prediction to obtain the desired speed. The desired speed can be the prediction result processed by the controller using the MLP network. The controller can input the prediction result and the target speed corresponding to the second scenario information into the reward function to calculate the reward and punishment result. The controller can use the reward and punishment result to reverse-optimize the speed model. After the controller completes the preliminary training of the speed model, it can use simulation data to perform simulation training on the speed model. The simulation training can include a large amount of simulation data, so as to realize the training of the speed model in big data and improve the model accuracy. After the controller completes the simulation training, it can deploy the speed model to the robotic animal and realize the prediction of the desired speed of the robotic animal.

[0099] The architecture of the speed model can be as Figure 3 shown. The data input into the speed model can include device information and second scenario information. Among them, the device information can include IMU, joint information, and the previous action information. The IMU is 4-dimensional data, including the yaw angle, pitch angle, yaw angular velocity, and pitch angular velocity at the current moment. The joint information contains the joint angles of each joint of the robotic animal. For example, in the robotic dog shown in Figure 1 the figure, it can include 12-dimensional data composed of 12 joint angles. The previous action information includes the joint angles of each joint at the previous moment. At the final input, the controller can use the device data of three moments as the final device features and input them into the speed model. That is, the device features can be 84-dimensional data. Among them, the second scenario information is collected by a camera installed on the top of the animal's head. The image size of the second scenario information can be 64X64. The desired speed finally output by the speed model can include the desired horizontal and vertical speeds and the angular velocity.

[0100] During the model training process, the reward function can be determined according to the difference between the real speed of the real animal and the predicted desired speed. Its specific calculation formula is similar to the motion strategy. When designing the reward function, considering that the speed curve planned by the robotic animal during movement cannot change too violently, it is necessary to make it as smooth as possible. The reward function can be shown as the following formula:

[0101]

[0102] where r mt is the reward function of the speed model. This reward function consists of two parts. One is the difference between the predicted speed and the target speed. The other is and ω t the generated change. The weights corresponding to these two parts are α and β respectively. The magnitudes of these two weights will determine which part is more important in the final result. Generally, the value of β is 0.1 times that of α. For example, when α is 10, β is 1. During the actual training process, the values of α and β can be adjusted according to the actual debugging and testing situations.

[0103] During the training process of this speed model, the overall network framework uses the Proximal Policy Optimization (PPO) algorithm of reinforcement learning, and the Generalized Advantage Estimator (GAE) technology can also be integrated into this network framework to make the network training process of this speed model as stable as possible. This speed model can also use the same neural network architecture for the Actor and the Critic, adopt a 2-layer MLP network to process the perceptual information of device information, and adopt a 3-layer CNN network to process visual information. This speed model is also used to finally process the fusion information using a 2-layer MLP and convert it into the desired speed.

[0104] S203. Obtain the real speed and real joint data of the real animal during movement through motion sensors arranged at the joints of the real animal.

[0105] In this embodiment, the controller can also obtain the real joint data of the real animal during movement through joint sensors. These joint sensors can be installed at each joint of the real animal. For example, Figure 5 shows the installation positions of three joint sensors on a leg of a dog. These key sensors can be respectively arranged at 12 joints of the four limbs of the dog. These joint sensors can be used to obtain the angle of the joint at each moment of the real animal during movement. The data obtained and uploaded by these joint sensors are angle values that continuously change in time series. The controller can also obtain the real speed of the real animal during movement through the centroid motion state sensor. This real speed is the same as the real joint data and is a speed value that continuously changes in time series.

[0106] S204. Use the real speed and real joint data to train the posture model of the robotic animal. The posture model is used to generate the predicted joint data of the robotic animal according to the predicted desired speed.

[0107] In this embodiment, the controller can use the real speed and real joint data to train the posture model after obtaining the real speed and real joint data. The architecture of the posture model can be as follows Figure 7 shown.

[0108] The specific process may include:

[0109] Step 1: Map the real joint data according to the dynamic mapping to obtain the first feature.

[0110] In this step, after acquiring the posture data of the real animal, the controller can use a dynamic mapping method to map the posture data of the real animal to the posture data of the robot animal. Specifically, the posture data can include real joint data. The controller can construct a calculation formula based on the real joint data. The controller can also add the action difference between the real animal and the robot animal to the calculation formula to obtain the calculation formula for the posture data of the robot animal. In this process, the controller can calculate the posture data of the robot animal based on the real joint data. t , calculate the robot animal's imitation state reward r st This imitation state is the first feature. The difference in motion between the robotic animal and the real animal can be continuously optimized as the model is trained. The calculation formula that incorporates the motion difference is the first strategy. This first strategy is used to plan the joint data for each joint.

[0111] Step 2: Fuse the first feature and the true speed to obtain the second feature.

[0112] In this step, Figure 7 The target speed shown is not affected by the actual speed obtained in S203. When the posture model is used for prediction, the target speed can also be the expected speed. The target speed can be expressed as reward r gt The controller can represent the first feature reward r obtained in step 1 st With the real speed reward r gt Fused together and converted into the second feature r t The calculation formula for this feature fusion can be:

[0113] r t =ω g r gt +ω s r st

[0114] where ω g is the weight of the target reward function, ω s is the weight of the state reward function. This design of the reward function takes into account both the effect of reaching the target speed and simulating the animal's movement posture.

[0115] Step 3: Input the second feature and the true speed into the pose model to plan and obtain predicted joint data.

[0116] In this step, the controller can input the second feature and the true speed into the second policy of the pose model to realize the action learning of the robotic animal. The action learning of the robotic animal is a Markov Decision Process (MDP), which can be expressed as (S, A, f, r t , p0, γ). Among them, S is used to represent the state space. A is used to represent the action space. f(s, a) is the system dynamics representation function. r t (s, a, s′) is the reward function. p0 is the initial state distribution. γ is the discount factor. In this pose model, the goal of reinforcement learning is to find the optimal parameter θ of the policy π θ : S → A, so that the expected return is maximized. Among them, T is the total time of each episode.

[0117] In this application, the pose model is the motion policy. The pose model can be divided into two sub-policies. Among them, the first policy is used to imitate the animal pose. The second policy is used to track the desired speed. In this pose model, there is also a state reward function. The state reward function is used to track the animal motion pose at each moment. The animal motion pose is the joint data. The joint data can be expressed as (q′0, q′1, q′2,..., q′ t ) in time series.

[0118] Among them, the specific expression of r st can be:

[0119]

[0120] Among them, is used to minimize the difference in the angles of the actions of the robotic animal and the real animal. ω p is the weight corresponding to . ω v is the weight corresponding to . is the reward function for the joint linear velocity.

[0121] Among them, the specific expression of can be shown as follows:

[0122]

[0123] Among them, and They are the joint angles of a real animal and a robotic animal respectively.

[0124] Among them, The specific expression of can be as follows:

[0125]

[0126] Among them, and They are the joint linear velocities of a real animal and a robotic animal respectively.

[0127] It can be seen from the above formula that the reward function is mainly used for the difference between the pose model trained and the actions of real animals. The closer the actions predicted by the pose model are to the action parameters of real animals, the closer the difference is to 1. Otherwise, the more the actions predicted by the pose model deviate from the poses of real animals, the closer the difference is to -1. During the model training process, the controller will import the dynamic model file of the robotic animal and the URDF file describing the mechanical properties of the entire robotic animal. These two files are used to obtain the action parameters of the robot in the simulation environment.

[0128] The reward function is used to track the speed at time t Among them, is the longitudinal and lateral speed in the forward direction of the desired speed. ω t is the global yaw angle in the desired speed. The specific formula of the reward function can be as follows:

[0129]

[0130] Among them, and ω t ′ are the linear velocity and angular velocity in the desired speed.

[0131] The entire training network of this pose model also uses the PPO architecture in the field of reinforcement learning. Among them, both the Actor and Critic networks adopt the MLP fully connected network. This MLP network includes two hidden layers. These two hidden layers have 512 nodes and 256 nodes respectively. The activation function of this MLP network is Relu.

[0132] The entire training process of this pose model can be specifically as Figure 8As shown in the figure. The controller can collect the motion state data of the animal. The motion state data may include the real joint data and real speed of the real animal. The controller can input the real joint data into the imitation learning PPO strategy to complete the generation of the calculation formula for the imitation action of the robotic animal. The controller can also obtain the target speed. The target speed can be the real joint data or the desired speed. The controller can use the simulation data to train the pose model. After the training is completed, the controller can deploy the pose model to the robotic animal.

[0133] It should be noted that the speed model and the pose model can be trained synchronously. When the two models are trained synchronously, the training process can be as Figure 9 shown. Among them, the target speed used by the two models can use the same data source. Or, the target speeds of the two models can also come from different data sources. The controller can also perform sequential training. That is, the controller first trains the speed model alone. After the speed model is trained, the controller can use the speed model to continue training the pose model, so as to achieve the overall training effect. The controller can use the two models completed by synchronous training as the initial models for sequential training. The use of the initial models can effectively improve the overall training speed. Among them, the training data required by the pose model needs to be collected separately and is not fused with the scene information. The training data collected in the pose model can include dynamic data such as walking and running under normal conditions. In the subsequent optimization training process of the model, using this dynamic data for separate training can also effectively reduce the training difficulty of the model.

[0134] In the robotic animal control method provided by this application, the controller can obtain the second scene information and the real speed of the real animal during movement through the camera and sensors set on the head of the real animal. The controller can use the second scene information and the real speed to train the speed model of the robotic animal. The speed model is used to predict the desired speed according to the first scene information obtained by the camera of the robotic animal. The controller can also obtain the real speed and real joint data of the real animal during movement through the motion sensors set at the joints of the real animal. The controller can use the real speed and real joint data to train the pose model of the robotic animal. The pose model is used to generate the predicted joint data of the robotic animal according to the predicted desired speed. In this application, by parallel training the speed model and the pose model, the training speed of the two models is improved. This application can also perform sequential training on the two models after the training of the two models is completed, further improving the model accuracy and training speed.

[0135] Figure 10 shows a schematic structural diagram of a robotic animal control device provided by an embodiment of this application, as Figure 10As shown, the machine animal control device 10 of this embodiment is used to implement the operations corresponding to the controller in any of the above method embodiments. The machine animal control device 10 of this embodiment includes:

[0136] An acquisition module 11, configured to acquire first scene information through a camera disposed on the machine animal.

[0137] A processing module 12, configured to use a speed model to predict an expected speed according to the first scene information and the device information of the machine animal. According to the posture model and the expected speed of the machine animal, predicted joint data is planned.

[0138] A control module 13, configured to control each joint of the machine animal according to the predicted joint data so that the machine animal moves.

[0139] In one example, the processing module 12 is specifically configured to:

[0140] Input the first scene information into the visual extraction module of the speed model for feature extraction to obtain visual features.

[0141] Input the device information into the device information extraction module of the speed model for data processing to generate device features.

[0142] Merge the visual features and the device features to obtain a first fusion feature.

[0143] Input the first fusion feature into the speed prediction module of the speed model to predict the expected speed.

[0144] In one example, the device further includes:

[0145] A model training module 14, configured to acquire second scene information and real speed of a real animal during movement through a camera and sensors disposed on the head of the real animal. Use the second scene information and the real speed to train the speed model of the machine animal, and the speed model is used to predict the expected speed according to the first scene information acquired by the camera of the machine animal.

[0146] In one example, the model training module 14 is further configured to:

[0147] Acquire the real speed and real joint data of the real animal during movement through motion sensors disposed at the joints of the real animal.

[0148] Use the real speed and the real joint data to train the posture model of the machine animal, and the posture model is used to generate the predicted joint data of the machine animal according to the predicted expected speed.

[0149] The machine animal control device 10 provided by the embodiment of the present application can execute the above method embodiment. For its specific implementation principle and technical effect, please refer to the above method embodiment, which will not be elaborated here in this embodiment.

[0150] Figure 11 The figure shows a schematic hardware structure diagram of a controller provided by the embodiment of the present application. As Figure 11 shown, the controller 20 is used to implement the operations corresponding to the controller in any of the above method embodiments. The controller 20 in this embodiment may include: a memory 21 and a processor 22.

[0151] The memory 21 is used to store computer programs. The memory 21 may include a high-speed random access memory (Random Access Memory, RAM), and may also include non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disc, etc.

[0152] The processor 22 is used to execute the computer program stored in the memory to implement the machine animal control method in the above embodiment. Specifically, please refer to the relevant descriptions in the foregoing method embodiment. The processor 22 may be a central processing unit (Central Processing Unit, CPU), and may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0153] Optionally, the memory 21 may be either independent or integrated with the processor 22.

[0154] When the memory 21 is a device independent of the processor 22, the controller 20 may further include a bus 23. The bus 23 is used to connect the memory 21 and the processor 22. The bus 23 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0155] The controller provided in this embodiment can be used to execute the above-mentioned machine animal control method, and its implementation manner and technical effect are similar, which will not be elaborated here in this embodiment.

[0156] This application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it is used to implement the methods provided by the above various embodiments.

[0157] Among them, the computer-readable storage medium may be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the computer-readable storage medium is coupled to the processor, so that the processor can read information from the computer-readable storage medium and write information to the computer-readable storage medium. Of course, the computer-readable storage medium may also be a component of the processor. The processor and the computer-readable storage medium may be located in an Application Specific Integrated Circuits (ASIC). In addition, the ASIC may be located in a user device. Of course, the processor and the computer-readable storage medium may also exist as discrete components in a communication device.

[0158] Specifically, the computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0159] The present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device can read the computer program from the computer-readable storage medium, and the execution of the computer program by at least one processor causes the device to implement the methods provided by the above various embodiments.

[0160] The embodiments of the present application also provide a chip, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the device installed with the chip executes the methods in the above various possible embodiments.

[0161] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in an electrical, mechanical, or other form.

[0162] Among them, each module can be physically separated, for example, installed at different positions of a device, or installed on different devices, or distributed to multiple network units, or distributed to multiple processors. Each module can also be integrated together, for example, installed in the same device, or integrated in a set of code. Each module can exist in the form of hardware, or can also exist in the form of software, or can also be implemented in the form of software plus hardware. This application can select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] When the integrated modules are implemented in the form of software function modules, they can be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods of various embodiments of this application.

[0164] It should be understood that although the steps in the flowcharts in the above embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and they can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times, and their execution order does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of various embodiments of this application.

Claims

1. A method for controlling a robotic animal, characterized in that, The method includes: Obtaining first scene information through a camera disposed on the robotic animal, where the first scene is image information or video information of the current scene captured by the camera of the robotic animal; Inputting the first scene information into the visual extraction module of the speed model for feature extraction to obtain visual features; Inputting the device information of the robotic animal into the device information extraction module of the speed model for data processing to generate device features; the device information includes an inertial measurement unit, joint angles, and the previous control action; Combining the visual features and the device features to obtain a first fusion feature; Inputting the first fusion feature into the speed prediction module of the speed model to predict an expected speed; Planning predicted joint data according to the pose model of the robotic animal and the expected speed; the pose model is obtained by training with real speed and real joint data; Controlling each joint of the robotic animal according to the predicted joint data to enable the robotic animal to move.

2. The method according to claim 1, wherein The method further includes: Obtaining second scene information and real speed of a real animal during movement through a camera and sensors disposed on the head of the real animal; Training the speed model of the robotic animal using the second scene information and the real speed, where the speed model is used to predict an expected speed according to the first scene information obtained by the camera of the robotic animal.

3. The method according to claim 1 or 2, characterized in that: The method further includes: Obtaining the real speed and real joint data of a real animal during movement through motion sensors disposed at the joints of the real animal; Training the pose model of the robotic animal using the real speed and the real joint data, where the pose model is used to generate the predicted joint data of the robotic animal according to the predicted expected speed.

4. The method according to claim 3, wherein The training of the pose model of the robotic animal using the real speed and the real joint data specifically includes: Mapping the real joint data according to a dynamic mapping to obtain a first feature; Performing feature fusion on the first feature and the real speed to obtain a second feature; Inputting the second feature and the real speed into the pose model to plan predicted joint data.

5. A machine animal control device, characterized in that, The device includes: An acquisition module, configured to obtain first scene information through a camera disposed on the robotic animal, where the first scene is image information or video information of the current scene captured by the camera of the robotic animal; a processing module configured to input the first scene information into a visual extraction module of a velocity model for feature extraction to obtain visual features; input the device information of the robot into the device information extraction module of the velocity model for data processing to generate device features; the device information includes an inertial measurement unit, joint angles, and a previous control action; merge the visual features and the device features to obtain a first fused feature; input the first fused feature into a velocity prediction module of the velocity model to predict an expected velocity; and plan and obtain predicted joint data based on the posture model of the robot and the expected velocity; the posture model is obtained by model training using real velocity and real joint data; The control module is used to control the robot animal to obtain each joint according to the predicted joint data so as to make the robot animal move.

6. A controller, characterized in that: The controller includes: a memory and a processor; the memory is used to store a computer program; the processor is used to implement the robot animal control method as described in any one of claims 1 to 4 according to the computer program stored in the memory.

7. A robotic animal, characterized in that, The robot animal is provided with a plurality of joints, a camera and the controller as claimed in claim 6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, is used to implement the robot animal control method according to any one of claims 1 to 4.

9. A computer program product, characterized in that, The computer program product comprises a computer program, and when the computer program is executed by a processor, the method for controlling a robot animal according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Time and position control method of four-footed bionic robot

    CN102591344A

  • Robot gait autonomous learning method and device, electronic equipment and storage medium

    CN114660947A