Multilegged walking robot and walking method

The multi-legged walking robot uses reinforcement learning to stabilize leg control and optimize torque based on virtual environment models, ensuring stability and efficient crop collection on uneven terrain.

JP7831805B1Active Publication Date: 2026-03-17D SPIRIT CO LTD
View PDF 37 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing multi-legged mobility devices struggle to maintain stability and control on uneven terrain with steep slopes or numerous obstacles, leading to potential premature deterioration of legs and inadequate destination reach.

Method used

A multi-legged walking robot with a control unit that utilizes reinforcement learning to construct a walking learning model, controlling legs based on virtual environment information and leg characteristics, and includes a center of gravity measurement unit to optimize leg torque and position, allowing independent control of an arm unit for crop collection.

Benefits of technology

Enables stable walking on varied terrain by optimizing leg control and suppressing center of gravity fluctuations, while allowing efficient crop collection by minimizing interference from leg movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007831805000001_ABST
    Figure 0007831805000001_ABST
Patent Text Reader

Abstract

This invention provides a multi-legged walking robot and a walking method that enable control suitable for various environments. [Solution] A multi-legged walking robot 1 with four or more legs that walks on a walking surface 9 including uneven ground, comprising a torso 10, a plurality of legs 20 that are in contact with the walking surface 9 and support the torso 10, and a control unit 30 that controls the plurality of legs 20 based on a pre-constructed walking learning model WM. The walking learning model WM is characterized by being constructed by reinforcement learning targeting a virtual robot on a virtual walking surface that mimics the walking surface 9. For example, the walking learning model WM is characterized by being constructed by reinforcement learning using virtual environment information relating to the features of the virtual walking surface, virtual leg information relating to the features of the plurality of legs in relation to the virtual environment information, and a reward set for the result of the virtual robot executing the virtual leg information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-legged walking robot and a walking method.

Background Art

[0002] Conventionally, as a moving device used to achieve stable movement on uneven ground such as mountains and unpaved areas, for example, an uneven ground moving device disclosed in Patent Document 1 has been proposed.

[0003] In Patent Document 1, it is composed of a vehicle body facing the ground, a vehicle body mounted on the vehicle body, and six sets of movable legs that support the vehicle body. Three sets of movable legs are arranged in the front and rear on each side of the vehicle body. Each individual movable leg has its root part, middle part, and tip part integrated. The root part is swingable with respect to the vehicle body, the middle part is attached to the tip of the root part, and the tip part is attached to the tip of the middle part. The middle part is arranged to extend obliquely downward. An uneven ground moving device is disclosed. By having six sets of movable legs, the uneven ground moving device can ensure stability. In addition, by arranging the middle part at an inclination, the load on the movable legs is reduced. Further, the uneven ground moving device can ensure stability even on uneven ground by introducing a control method of always keeping at least four sets of movable legs in contact with the ground.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Here, the off-road mobility device disclosed in Patent Document 1 is based on the premise that at least four sets of movable legs are constantly in contact with the ground in order to ensure stability even on uneven terrain. However, in environments such as ground with a steep slope or ground with many obstacles, there are concerns that simply keeping the four sets of movable legs in contact with the ground may not be sufficient to reach the destination, or that excessive load may lead to premature deterioration of the movable legs. For this reason, it is difficult to control the off-road mobility device disclosed in Patent Document 1 in a way that is suitable for various environments.

[0006] Therefore, the present invention was devised in view of the above-mentioned problems, and its objective is to provide a multi-legged walking robot and a walking method that can be controlled to suit various environments. [Means for solving the problem]

[0007] The multi-legged walking robot according to the first invention is a multi-legged walking robot with four or more legs that walks on a walking surface including uneven ground, comprising: a torso; a plurality of legs that contact the walking surface and support the torso; and a control unit that controls the plurality of legs based on a pre-constructed walking learning model, wherein the plurality of legs are driven at different timings, and the walking learning model is constructed by reinforcement learning targeting a virtual robot on a virtual walking surface that mimics the walking surface. The walking learning model is constructed by reinforcement learning using virtual environment information relating to the characteristics of the virtual walking surface, virtual leg information relating to the characteristics of a plurality of legs in relation to the virtual environment information, and a reward set for the result of the virtual robot executing the virtual leg information. The virtual leg information includes a relationship between a virtual support region enclosed by the position where each of the plurality of virtual legs of the virtual robot touches the virtual walking surface, and a virtual center of gravity position projected onto the virtual walking surface. The reward is set based on whether or not the virtual center of gravity position enters the virtual support region. The robot further includes a center of gravity measurement unit provided in the torso for measuring the center of gravity. Each leg includes a support portion that contacts the walking surface and a motor portion connected to the torso for driving the support portion. The control unit uses the walking learning model to derive the torque of the motor portion corresponding to the center of gravity measured by the center of gravity measurement unit and controls the motor portion. It is characterized by the following:

[0011] The 2 The multi-legged walking robot according to the invention is the 1 In the invention, provided on the torso, The system further includes an altitude measuring unit that measures the height position of the torso relative to the walking surface, and the control The unit measures the height position measured by the height measurement unit and the characteristics of the walking surface. The boundary information is used as input information for the walking learning model, and a plurality of legs are used in response to the input information. The control conditions of the unit are derived as output information of the walking learning model, and based on the output information The system is characterized by controlling the aforementioned leg portion.

[0012] The 3 The multi-legged walking robot according to the invention further comprises an arm unit connected to the torso for collecting crops, the control unit controls the arm unit based on a harvesting learning model pre-constructed by reinforcement learning targeting a virtual crop that mimics the crops and a virtual arm unit that mimics the arm unit, the harvesting learning model is constructed by reinforcement learning using virtual crop information relating to the characteristics of the virtual crops, virtual arm unit information relating to the characteristics of the virtual arm unit in relation to the virtual crop information, and harvesting rewards set for the results of the virtual arm unit executing the virtual arm unit information. The control unit uses the centroid as an explanatory variable for the harvesting learning model, acquires control information for the arm as an objective variable, and controls the arm based on the control information. It is characterized by the following:

[0013] The 4 The multi-legged walking robot according to the invention is the 3 In the invention, the control unit stops the plurality of legs, and then the arm Multiple joints included Control The arm portion includes a cutter portion and a gripper portion, the cutter portion includes a cutting portion for cutting the support portion that supports the crop and a rotating portion linked to the cutting portion, the gripper portion includes a holding portion for holding the crop after the support portion has been cut by the cutter portion and a movable portion for adjusting the position of the holding portion, the movable portion is driven in conjunction with the rotation of the rotating portion after the driving of the multiple joint portions has stopped, and changes the position of the holding portion. It is characterized by the following:

[0014] The 5 The walking method according to the invention is the first invention to the second invention. 4 A walking method for any of the inventions of a multi-legged walking robot, comprising: a walking step of controlling a plurality of the legs and walking on the walking surface; an acquisition step of acquiring the characteristics of the uneven ground in contact with the legs; and a identification step of referring to the walking learning model and identifying the control content of the legs corresponding to the characteristics of the uneven ground. [Effects of the Invention]

[0015] First Invention ~ 5 According to the invention, the control unit controls multiple legs based on a pre-constructed walking learning model. The walking learning model is a model constructed by reinforcement learning targeting a virtual robot on a virtual walking surface that mimics a walking surface. Therefore, by using a walking learning model that takes into account the usage environment of the multi-legged walking robot, walking can be performed according to the purpose. This enables control that is suitable for various environments.

[0016] Furthermore, the 1st to 5th inventions According to the invention, the walking learning model represents a model constructed by reinforcement learning using virtual environment information, virtual leg information, and rewards. Therefore, when using the walking learning model in a real environment, it is possible to perform the most desirable control of the legs with respect to the characteristics of the walking surface.

[0017] Furthermore, the 1st to 5th inventions According to the invention, the virtual leg information includes the relationship between the virtual support area and the virtual center-of-gravity position. Therefore, when using the walking learning model, it is possible to realize the control of the legs considering the position of the center of gravity. Thereby, it becomes possible to improve the stability of walking.

[0018] Furthermore, the 1st to 5th inventions According to the invention, the control unit uses the walking learning model to derive the torque of the motor unit according to the center of gravity and controls the motor unit. Therefore, for example, even on uneven ground such as a mountain with many changes in gradient, it is possible to proceed with walking while suppressing fluctuations in the center of gravity. Thereby, it becomes possible to prevent the possibility of falling.

[0019] In particular, the 2 According to the invention, the control unit uses environmental information regarding the height position and the characteristics of the walking surface as input information for the walking learning model, derives control conditions for a plurality of legs with respect to the input information as output information of the walking learning model, and controls the legs based on the output information. Therefore, in addition to the environmental information, it is possible to perform optimal control of the legs according to the state of the torso. Thereby, it becomes possible to suppress the possibility of falling due to the height of the torso.

[0020] In particular, the 3 According to the invention, the control unit controls the arm unit based on the harvesting learning model. Therefore, it is possible to perform the control of the arm unit independently of the control conditions of the legs. Thereby, it becomes possible to collect agricultural crops in a state where the influence of the walking surface is suppressed.

[0021] In particular, the 4According to the invention, after stopping the plurality of legs, the control unit controls the arm unit. Therefore, when controlling the arm unit, it is possible to suppress the influence of vibrations and the like caused by the driving of the legs. As a result, it becomes possible to easily collect agricultural crops.

[0022] In particular, the 5 According to the invention, the specific step refers to the walking learning model and specifies the control content of the legs corresponding to the characteristics of the rough ground. As a result, it becomes possible to realize control suitable for various environments.

Brief Description of the Drawings

[0023] [Figure 1] FIG. 1 is a schematic perspective view showing an example of a multi-legged walking robot in an embodiment. [Figure 2] FIG. 2 is a schematic top view showing an example of a multi-legged walking robot in an embodiment. [Figure 3] FIGS. 3(a) to 3(d) are schematic perspective views showing examples of legs. [Figure 4] FIGS. 4(a) to 4(c) are schematic views showing examples of the center of gravity of a multi-legged walking robot. [Figure 5] FIG. 5 is a flowchart showing an example of a walking method in an embodiment. [Figure 6] FIG. 6 is a schematic perspective view showing an example of an arm unit. [Figure 7] FIGS. 7(a) and 7(b) are schematic views showing an example of a collection unit. [Figure 8] FIGS. 8(a) and 8(b) are schematic views showing an example of a cutter unit and a gripper unit.

Embodiments for Carrying Out the Invention

[0024] Hereinafter, an example of a multi-legged walking robot and a walking method in the present invention will be described in detail.

[0025] (Embodiment: Multi-legged walking robot 1) Figure 1 is a schematic perspective view showing an example of the multi-legged walking robot 1 in this embodiment. Figure 2 is a schematic top view showing an example of the multi-legged walking robot 1 in this embodiment.

[0026] The multi-legged walking robot 1 refers to a robot with four or more legs that walks on a walking surface 9 that includes uneven terrain. The multi-legged walking robot 1 can automatically walk to a destination set on uneven terrain such as mountainous areas or unpaved areas. The multi-legged walking robot 1 can be used as a substitute for actions that require human intervention, such as collecting crops grown at a destination, transporting goods, or monitoring the area around the walking path.

[0027] The multi-legged walking robot 1 comprises, for example, a torso 10, a plurality of legs 20, and a control unit 30, as shown in Figures 1 and 2, and may also include at least one of a sensor unit 40 and an arm unit 50. The multi-legged walking robot 1 may be a known quadruped walking robot using, for example, a torso 10 and four legs 20, with the control unit 30 and the like mounted on it.

[0028] <Body section 10> The torso 10 represents the core part of the multi-legged walking robot 1. The torso 10 may have a shape such as a rectangular parallelepiped or a plate, and any shape can be used depending on the application. Steel can be used as the material for the torso 10, and any material that can withstand outdoor use can be used.

[0029] The torso section 10 is, for example, a housing that houses the control unit 30 and electronic equipment such as batteries for driving each structure. The torso section 10 has dimensions of, for example, 500 mm to 2000 mm on each side, and can be arbitrarily set according to the application.

[0030] <Legs 20> Multiple leg sections 20 contact the walking surface 9 and support the torso section 10. Four or more leg sections 20 are provided, and any number can be provided depending on the application. Each leg section 20 has, for example, one or more joints. By using joints, it is possible to support the torso section 10 in a posture suitable for the characteristics of the walking surface 9.

[0031] The leg portion 20 includes, for example, a support portion 21, a motor portion 22, a protective portion 23, and a connecting portion 24, as shown in Figure 3(a). For the leg portion 20, for example, the legs of a known walking robot may be used.

[0032] The support portion 21 includes a contact portion 21a that contacts the walking surface 9. The motor portion 22 is connected to the torso portion 10 and drives the support portion 21. The protective portion 23 is provided between the support portion 21 and the motor portion 22 and has, for example, a cavity. The connecting portion 24 is connected to the support portion 21 via a joint portion 25, extends into the cavity of the protective portion 23 and is connected to the motor portion 22. The connecting portion 24 drives the support portion 21 via the motor portion 22.

[0033] The motor unit 22 includes, for example, three motors 22a, 22b, and 22c. The first motor 22a can provide a first rotation r about a first axis X along the walking direction to the support unit 21 and the protective unit 23, as shown in, for example, Figure 3(b). This allows, for example, the height position and degree of parallelism of the torso unit 10 with respect to the walking surface 9 to be changed.

[0034] The second motor 22b can impart a second rotation p1 to the support portion 21 and the protective portion 23, for example, as shown in Figure 3(c), around the second axis Y1 in a direction intersecting the walking direction. This allows, for example, the height position and degree of parallelism of the torso portion 10 relative to the walking surface 9 to be changed.

[0035] The third motor 22c can impart a third rotation p2 to the support portion 21, for example, as shown in Figure 3(d), around the third axis Y2 which passes through the joint portion 25 and is parallel to the second axis Y1. This enables, for example, the multi-legged walking robot 1 to perform walking.

[0036] <Control Unit 30> The control unit 30 controls multiple legs 20 based on a pre-built walking learning model WM. The control unit 30 is housed, for example, in the torso 10.

[0037] The control unit 30 includes, for example, a central control unit, a storage unit, and a communication unit. The central control unit represents a processor such as a CPU (Central Processing Unit) and controls, for example, each component of the multi-legged walking robot 1. The storage unit represents, for example, volatile memory such as RAM (Random Access Memory) or non-volatile memory such as ROM (Read Only Memory), and stores various data and programs such as a walking learning model WM. The communication unit represents known connection equipment for performing wireless communication such as Wi-Fi (registered trademark).

[0038] Furthermore, known control techniques can be used as means for the control unit 30 to control the multi-legged walking robot 1. For example, the central control unit can control the multi-legged walking robot 1 by using the non-volatile memory as a work area and executing a program stored in the non-volatile memory.

[0039] Furthermore, the control unit 30 may control each component (for example, the motor unit 22) using sensor information obtained from imaging devices and sensors attached to the torso unit 10, etc., depending on the application. In this case, the control unit 30 uses sensor information as explanatory variables for the walking learning model WM and acquires control information for each component as the objective variable. The control unit 30 controls each component using the acquired control information.

[0040] The control unit 30 may perform efficient route selection based on sensor information obtained from, for example, the sensor unit 40. For route selection, for example, A * Algorithms and known techniques such as Dijkstra's algorithm can be used. In this case, the control unit 30 can select a route using elevation data and topographic information (slope, presence or absence of obstacles, etc.) that have been stored in advance in the storage unit or the like.

[0041] <Sensor unit 40> The sensor unit 40 is provided, for example, on the torso unit 10, and can also be provided in any location depending on the application, and the number of units provided is arbitrary. The sensor unit 40 includes known sensors that can measure the state of the multi-legged walking robot 1, such as imaging devices such as cameras, position measuring devices, inertial measurement units (IMUs), and obstacle detection sensors (LiDAR (Light Detection and Ranging), ultrasonic sensors, etc.). As position measuring devices, for example, RTK (Real-Time Kinematic) GPS (Global Positioning System) or satellite systems (GLONASS (Global Navigation Satellite System), Galileo, etc.) can be used. The sensor unit 40 includes, for example, at least one of a center of gravity measuring unit and an altitude measuring unit.

[0042] <<Center of Gravity Measurement Unit>> The sensor unit 40 includes, for example, a center of gravity measurement unit. The center of gravity measurement unit is provided on the torso unit 10 and measures the center of gravity G of the multi-legged walking robot 1. When the sensor unit 40 includes a center of gravity measurement unit, the control unit 30 uses a walking learning model WM to derive the torque of the motor unit 22 according to the center of gravity G and controls the motor unit 22.

[0043] For example, as shown in Figure 4(a), the control unit 30 derives a support region R1 enclosed by the positions where each of the multiple contact parts 21a contacts the walking surface 9, and identifies the position of the center of gravity G relative to the support region R1. The support region R1 may be identified, for example, by the degree of torque acting on each motor unit 22, or based on the results of imaging the contact positions between the walking surface 9 and the contact parts 21a using an imaging device or the like. The position of the center of gravity G indicates, for example, the position obtained by projecting the center of gravity G onto the walking surface 9. For example, a range R2 included in the support region R1 may be set in advance, and it may be identified whether the position of the center of gravity G falls within the range R2.

[0044] For example, as shown in Figure 4(b), if the position of the center of gravity G does not fall within the range R2, the control unit 30 controls the motor unit 22 so that the position of the center of gravity G falls within the range R2 (for example, Figure 4(c)). This makes it possible to control the leg unit 20 while taking the position of the center of gravity G into consideration.

[0045] <<Altitude Measurement Unit>> The sensor unit 40 includes, for example, an altitude measurement unit. The altitude measurement unit is provided on the torso unit 10 and measures the height position of the torso unit 10 with respect to the walking surface 9. When the sensor unit 40 includes an altitude measurement unit, the control unit 30 uses a walking learning model WM to derive the torque of the motor unit 22 according to the height position and controls the motor unit 22.

[0046] For example, the control unit 30 compares a preset threshold with the height position. If the height position exceeds the threshold, the control unit 30 controls the motor unit 22 so that the height position becomes below the threshold. This enables optimal control of the leg unit 20 for the state of the torso unit 10.

[0047] <Walking Learning Model WM> The walking learning model WM represents a model constructed by reinforcement learning targeting a virtual robot on a virtual walking surface that mimics the walking surface 9. The virtual robot represents a virtual model with functions equivalent to the multi-legged walking robot 1 in the virtual environment. The virtual walking surface represents the surface on which the virtual robot walks in the virtual environment, and any virtual uneven terrain can be set depending on the application. The walking learning model WM is constructed using known electronic equipment capable of performing reinforcement learning.

[0048] Reinforcement learning is a known learning method that uses a virtual environment to build a model, and is implemented using, for example, Isaac Lab or Isaac Sim (both from NVIDIA). Reinforcement learning derives the Q function and policy using, for example, pre-prepared virtual robot and virtual walking surface data. The Q function is derived using, for example, an offline reinforcement learning algorithm such as CQL (Conservative Q-Learning). The policy provides guidelines (functions) for the virtual robot to select the optimal action in the virtual environment.

[0049] In reinforcement learning used to construct a walking learning model WM, a policy, virtual environment information VA, virtual leg information VB, and reward Cw are used. The learning method involves using the aforementioned policy to determine virtual leg information VB that is suitable for the virtual environment information VA. Then, based on the virtual leg information VB, the virtual robot performs the movement of its virtual legs. A reward Cw is then set for the result of the virtual leg movement, and the policy is updated based on the reward Cw. By repeating the above, the policy is optimized, and the walking learning model WM is constructed.

[0050] <<Virtual Environment Information VA>> Virtual environment information (VA) provides information about the characteristics of the virtual walking surface. For example, virtual environment information (VA) may include conditions such as the slope, surface roughness, and hardness of the virtual walking surface, as well as obstacles such as stones, wood, and puddles, and environmental conditions such as sunny, rainy, and snowy weather.

[0051] <<Virtual Leg Information VB>> The virtual leg information VB provides information about the characteristics of multiple virtual legs relative to the virtual environment information VA. The virtual leg information VB includes, for example, the torque of the virtual motor used to drive each virtual leg. The torque may include, for example, the magnitude of the force, as well as, for example, the change in the magnitude of the force and the reaction force.

[0052] The virtual leg information VB includes a walking position that indicates the degree to which the virtual robot has walked, for example, by driving the virtual legs. The walking position indicates, for example, the difference between the virtual robot's position before driving the virtual legs and the virtual robot's position after driving the virtual legs (current position). In addition to the above, the walking position may also indicate the difference between the target trajectory for heading to a pre-set destination and the current position. As the walking position, data is used that is expected to be calculated based on sensor information obtained from, for example, GPS, inertial measurement units, obstacle detection sensors, etc., and is represented, for example, as a vector.

[0053] The virtual leg information VB includes, for example, the positional relationship between the virtual support region and the virtual center of gravity. The virtual support region represents the area enclosed by the points where each of the multiple virtual legs of the virtual robot touches the virtual walking surface. The virtual center of gravity represents the position obtained by projecting the center of gravity of the virtual robot onto the virtual walking surface.

[0054] The virtual leg information VB may include state information that indicates the state of the virtual robot, such as its angular velocity and tilt. The state information is assumed to be data such as sensor information such as angular velocity (e.g., roll rate, pitch rate, etc.) and tilt acquired using an inertial measurement unit.

[0055] The virtual leg information VB may include, for example, the ground contact pattern of the virtual leg (e.g., a pattern where the front leg is lifted and the rear leg is in contact). As the ground contact pattern, data that is expected to be identified by, for example, the degree of torque controlling the motor section 22 of each leg section 20 may be used, or data that is expected to be identified based on the results of imaging the contact position between the walking surface 9 and the contact section 21a using an imaging device or the like may be used.

[0056] <<Reward Cw>> The reward Cw is set based on the results of the virtual robot's execution of virtual leg information VB. For example, if the result is close to ideal, a high reward Cw is set, and if the result deviates from ideal, a penalty Cw is set.

[0057] The reward Cw is used when updating the policy. The policy is updated, for example, based on multiple reward Cws, and constructed as the optimal function.

[0058] For example, if the virtual leg information VB includes the torque of a virtual motor, a reward Cw may be set based on a comparison between the torque and a pre-set threshold. For instance, if the torque does not exceed the threshold, a high reward C is set, and if the torque exceeds the threshold, a penalty is set as the reward Cw.

[0059] For example, if the virtual leg information VB includes the walking position, the reward Cw is set based on the result of calculating the dot product of the position vector and the target vector. In addition to the above, the reward Cw may also be set based on the result of calculating the difference between the target trajectory and the current position. For example, if the calculated difference is less than or equal to a predetermined threshold, a high reward Cw is set to achieve emphasis on the amount of progress (e.g., increased torque), and if the calculated difference exceeds the threshold, a penalty is set as the reward Cw.

[0060] For example, if the virtual leg information VB includes positional relationships, the reward Cw is set based on whether the virtual center of gravity position falls within the virtual support area (or a pre-defined virtual range included within the virtual support area). For example, if the virtual center of gravity position falls within the virtual support area, a high reward Cw is set, and if the virtual center of gravity position does not fall within the virtual support area, a penalty is set as the reward Cw.

[0061] For example, if the virtual leg information VB includes state information, the reward Cw may be set based on the result of comparing the state information with a pre-set threshold. For example, if the state information does not exceed the threshold, a high reward Cw is set, and if the state information exceeds the threshold, a penalty is set as the reward Cw.

[0062] For example, if the virtual leg information VB includes a ground contact pattern, the reward Cw may be set based on the result of comparing the ground contact pattern with a pre-set reference ground contact pattern. For example, if the ground contact pattern matches or is similar to the reference ground contact pattern, a high reward Cw is set, and if the ground contact pattern is different from the reference ground contact pattern, a penalty Cw is set as the reward Cw for an improper contact state.

[0063] <Arm section 50> The arm 50 is connected to the body 10 and collects crops. The arm 50 is controlled by the control unit 30 based on a harvesting learning model HM that has been constructed in advance using reinforcement learning. Known control techniques can be used as means for the control unit 30 to control the arm 50 and the like.

[0064] Agricultural crops include fruit trees and horticultural crops such as vegetables. For example, agricultural crops refer to fruit trees and the fruits that grow on them, such as oranges and apples.

[0065] For example, a known robot arm can be used as the arm portion 50, and a robot arm of any size and shape can be used depending on the application. The arm portion 50 includes a configuration corresponding to, for example, the "arm portion," "support mechanism," "imaging unit," and "cutting mechanism" described in Japanese Patent Application Publication No. 2021-185758.

[0066] The arm portion 50 includes, for example, a plurality of joints 111 and a collection unit 120, as shown in Figure 6. Each of the plurality of joints 111 is driven independently, and the position and orientation of the tip portion 112 of the arm portion 50 can be changed. The collection unit 120 is provided at the tip portion 112. As the arm portion 50, for example, a configuration in which the collection unit 120 is attached to a known robot arm may be used.

[0067] <Collection Section 120> The collection unit 120 is used to collect agricultural products 105, for example, as shown in Figures 7(a) and 7(b). The collection unit 120 includes, for example, a cutter unit 121 and a gripper unit 122. Alternatively, a known end effector for collecting agricultural products 105 may be used as the collection unit 120.

[0068] <Cutter section 121> The cutter section 121 is used to cut the supporting parts of the crop 105, such as fruit stalks and branches. The cutter section 121 includes a cutting section 121a, a cutting assist section 121b, and a rotating section 121c, as shown in Figures 8(a) and 8(b), for example.

[0069] The cutting portion 121a has a cutting edge for cutting the support portion. The cutting portion 121a may have any shape, for example, according to the characteristics of the crop 105.

[0070] The cutting assist portion 121b is used to suppress fluctuations in the support portion and to facilitate cutting of the support portion by the cutting portion 121a. The cutting assist portion 121b may have, for example, a slit shape that clamps the side surface of the support portion, or it may have any shape depending on the object.

[0071] The rotating part 121c is linked to the cutting part 121a and is used to drive the cutting part 121a. For example, a shaft-shaped member may be used as the rotating part 121c, and the rotation of the rotating part 121c drives the cutting part 121a. The rotating part 121c may also be linked to, for example, the gripper part 122 and used to drive both the cutting part 121a and the gripper part 122.

[0072] <Gripper section 122> The gripper section 122 is used to hold the crops 105. The gripper section 122 includes, for example, a holding section 122a and a movable section 122b.

[0073] The holding portion 122a holds the crop 105 whose support portion has been cut by, for example, the cutter portion 121. The holding portion 122a may have, for example, a flat plate shape, or it may have a shape that is easy to hold depending on the shape of the crop 105.

[0074] The movable part 122b is used to adjust the position of the holding part 122a. The movable part 122b is driven, for example, in conjunction with the rotation of the rotating part 121c, and changes the position of the holding part 122a. For example, when the rotating part 121c rotates to cut the support part, the movable part 122b changes the position of the holding part 122a in a direction that brings it closer to the crop 105. This makes it easier to hold the crop 105 with the holding part 122a when the support part is cut.

[0075] For example, when the control unit 30 controls the arm unit 50, it stops the driving of the multiple joint units 111 before driving the cutter unit 121 and the gripper unit 122. In this case, when collecting crops 105 using the cutter unit 121 and the gripper unit 122, damage to the crops 105 caused by the driving of the joint units 111 can be prevented.

[0076] <Imaging Unit 140> For example, the multi-legged walking robot 1 may be equipped with an imaging unit 140. The imaging unit 140 is provided at the tip 112 and images the crops 105. The imaging unit 140 is mounted from the tip 112 toward the cutter 121, as shown in Figures 7(a) and 7(b). In this case, it becomes easier to image the behavior of the crops 105 in response to the drive of the collection unit 120. As the imaging unit 140, for example, a known imaging device may be used, and for example, an imaging device capable of imaging in the infrared region as well as the visible light region may be used.

[0077] For example, if crop 105 represents a fruit tree and the fruit growing on it, the imaging unit 140 will image the branches and leaves of the fruit tree, the fruit, and the environment surrounding the fruit tree, and generate imaging information. For example, if crop 105 represents a root vegetable and its roots, the imaging unit 140 will image the root vegetable, its roots, and the soil, and generate imaging information.

[0078] For example, there are many obstacles around the fruit, such as branches and leaves of the fruit tree and nets to prevent the fruit from falling. In this case, the imaging unit 140 images the crop 105 and generates imaging information. The control unit 30 then derives the positional relationship between the branches and leaves of the crop 105 and the fruit based on the captured imaging information (e.g., an image). The control unit 30 also uses the above positional relationship as an explanatory variable for the harvesting learning model HM and acquires control information for multiple joints 111 as the objective variable. Using the acquired control information, the control unit 30 controls the multiple joints 111 to bring the arm 50 closer to the fruit. In this way, by using images, the arm 50 can be controlled without the collection unit 120 coming into contact with the obstacles, even when there are many obstacles such as branches and leaves. Furthermore, even if the imaging unit 140 images an obstacle such as a net to prevent the fruit from falling, the arm 50 can be controlled without the collection unit 120 coming into contact with the obstacle.

[0079] Furthermore, the control unit 30 may control each component, such as the arm unit 50, using imaging information obtained from the imaging unit 140, depending on the application. In this case, multiple control commands corresponding to the content of the imaging information may be stored in the storage unit, and a control command may be selected according to the content of the imaging information to control each component. The control unit 30 may also control each component, such as the arm unit 50, based on sensor information obtained from the sensor unit 40.

[0080] <Harvesting Learning Model HM> The harvesting learning model HM represents a model constructed through reinforcement learning targeting a virtual crop modeled after a crop 105 and a virtual arm modeled after an arm 50. The virtual arm represents a virtual model with functions equivalent to the arm 50 in a virtual environment. The virtual crop represents the target that the virtual arm harvests in the virtual environment, and any target can be set depending on the application. The harvesting learning model HM is constructed using known electronic equipment capable of performing reinforcement learning.

[0081] Reinforcement learning employs the same learning methods as when constructing the walking learning model WM described above, and is performed using, for example, Isaac Lab or Isaac Sim (both from NVIDIA). Reinforcement learning derives the Q function and policy using, for example, pre-prepared virtual arm and virtual crop data. The Q function is derived using, for example, an offline reinforcement learning algorithm such as CQL (Conservative Q-Learning). The policy provides guidelines (functions) for the virtual arm to select the optimal action for the virtual crop.

[0082] In reinforcement learning used to construct the harvesting learning model HM, a policy, virtual crop information VH, virtual arm information VM, and reward Ch are used. The learning method involves determining the virtual arm information VM, which is suitable for the virtual crop information VH, using the policy described above. Then, the virtual arm is driven based on the virtual arm information VM. A reward Ch is set for the result of the virtual arm driving, and the policy is updated based on the reward Ch. By repeating the above, the policy is optimized and the harvesting learning model HM is constructed.

[0083] <<Virtual Crop Information VH>> Virtual crop information VH indicates information about the characteristics of a virtual crop. Virtual crop information VH may include, for example, the shape, type, and location of obstacles during harvesting of the virtual crop, as well as environmental information such as sunny, rainy, or snowy conditions. Virtual crop information VH may also include virtual position information indicating the positional relationship between virtual branches and leaves, which are modeled after the branches and leaves of a fruit tree, and virtual fruits, which are modeled after fruits.

[0084] <<Virtual Arm Information VM>> The virtual arm information VM provides information about the characteristics of the virtual arm in relation to the virtual crop information VH. The virtual arm information VM includes, for example, the degree of drive of the virtual joint for moving the tip of the virtual arm (e.g., motor torque). Torque may include, for example, the magnitude of the force, as well as, for example, the change in the magnitude of the force or the reaction force.

[0085] For example, the virtual arm information VM may include parameters for a virtual harvesting unit that is driven to collect virtual crops. The virtual arm information VM includes, for example, driving conditions for bringing the tip of the virtual arm close to a virtual fruit, based on virtual position information. The driving conditions include information for controlling multiple virtual joints so that, for example, the tip of the virtual arm passes through a path that avoids contact with virtual branches and leaves.

[0086] <<Reward Ch>> The reward Ch is set based on the result of the virtual arm executing the virtual arm information VM. For example, if the result is close to ideal, a high reward Ch is set, and if the result deviates from ideal, a penalty is set as the reward Ch.

[0087] The reward Ch is used when updating the policy. The policy is updated, for example, based on multiple reward Chs, and constructed as an optimal function.

[0088] For example, if the virtual arm information VM includes the torque of the virtual joint, a reward Ch may be set based on a comparison between the torque and a pre-set threshold. For instance, if the torque does not exceed the threshold, a high reward Ch is set, and if the torque exceeds the threshold, a penalty is set as the reward Ch.

[0089] For example, if the virtual arm information VM includes driving conditions, the reward Ch is set based on whether or not the position of the tip of the virtual arm enters the harvestable area of ​​the virtual crop. For example, if the position of the tip of the virtual arm enters the harvestable area, a high reward Ch is set, and if the position of the tip of the virtual arm does not enter the harvestable area, a penalty is set as the reward Ch.

[0090] In the constructed harvesting learning model HM, for example, imaging information generated from the imaging unit 140 and sensor information obtained from the sensor unit 40 can be used as explanatory variables corresponding to the virtual crop information VH. In addition, control information of the arm unit 50 corresponding to the virtual arm unit information VM can be obtained by inputting explanatory variables into the harvesting learning model HM.

[0091] (Embodiment: Walking method) Next, an example of a walking method in this embodiment will be described. Figure 5 is a flowchart of an example of a walking method. The walking method can be implemented using the multi-legged walking robot 1 described above. The walking method comprises a walking step S110, an acquisition step S120, and a identification step S130.

[0092] <Walking step S110> The walking step S110 controls multiple legs 20, allowing the multi-legged walking robot 1 to walk on the walking surface 9. In the walking step S110, for example, each leg 20 is driven at a different timing. Therefore, compared to, for example, when two or more legs are driven at the same time, it is easier to maintain balance when walking on uneven ground. This makes it possible to improve stability.

[0093] <Acquisition Step S120> Acquisition step S120 acquires the characteristics of the uneven ground in contact with the leg portion 20. The characteristics of the uneven ground may be identified, for example, based on the reaction force acting on the motor portion 22 of the leg portion 20, or based on images captured by the imaging device included in the sensor portion 40.

[0094] <Specific step S130> In step S130, the walking learning model WM is referenced to identify the control content of the leg 20 corresponding to the characteristics of uneven terrain.

[0095] For example, the walking method may include an arm control step that controls the arm portion 50 after the specific step S130. In the arm control step, for example, the control unit 30 controls the arm portion 50 after stopping the multiple leg portions 20.

[0096] After performing each of the steps described above, the walking method is completed. Note that the acquisition step S120 and the identification step S130 may be repeated multiple times.

[0097] According to this embodiment, the control unit 30 controls the multiple legs 20 based on a pre-constructed walking learning model WM. The walking learning model WM represents a model constructed by reinforcement learning targeting a virtual robot on a virtual walking surface that mimics the walking surface 9. Therefore, by using a walking learning model WM that takes into account the usage environment of the multi-legged walking robot 1, walking can be performed according to the purpose. This enables control that is suitable for various environments.

[0098] Furthermore, according to this embodiment, the walking learning model WM is a model constructed by reinforcement learning using virtual environment information VA, virtual leg information VB, and reward C. Therefore, when the walking learning model WM is used in a real environment, it is possible to perform control of the leg 20 that is most expected to work with the features of the walking surface 9.

[0099] Furthermore, according to this embodiment, the virtual leg information VB includes the relationship between the virtual support area and the virtual center of gravity position. Therefore, when using the walking learning model WM, it is possible to control the leg 20 while considering the position of the center of gravity G. This makes it possible to improve walking stability.

[0100] Furthermore, according to this embodiment, the control unit 30 uses a walking learning model WM to derive the torque of the motor unit 22 according to the center of gravity G and controls the motor unit 22. Therefore, even on uneven terrain such as mountainous areas with many changes in gradient, walking can be carried out while suppressing fluctuations in the center of gravity G. This makes it possible to prevent the possibility of falling.

[0101] Furthermore, according to this embodiment, the control unit 30 uses environmental information regarding height position and characteristics of the walking surface 9 as input information for the walking learning model WM, derives control conditions for multiple legs 20 in relation to the input information as output information for the walking learning model WM, and controls the legs 20 based on the output information. Therefore, in addition to environmental information, it is possible to perform optimal control of the legs 20 in relation to the state of the torso 10. This makes it possible to suppress the possibility of falling due to the height of the torso 10.

[0102] Furthermore, according to this embodiment, the control unit 30 controls the arm unit 50 based on the harvesting learning model HM. Therefore, the control of the arm unit 50 can be performed independently of the control conditions of the leg unit 20. This makes it possible to collect crops 105 while suppressing the influence of the walking surface 9.

[0103] Furthermore, according to this embodiment, the control unit 30 controls the arm unit 50 after stopping the multiple leg units 20. Therefore, when controlling the arm unit 50, the effects of vibrations and other noises caused by the driving of the leg units 20 can be suppressed. This makes it possible to easily collect agricultural products.

[0104] Furthermore, according to this embodiment, the specific step S130 refers to the walking learning model WM and identifies the control content of the leg 20 corresponding to the characteristics of uneven terrain. This makes it possible to realize control suitable for various environments.

[0105] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of Symbols]

[0106] 1: Multilegged walking robot 10: Torso 20: Legs 21: Support part 21a: Contact part 22: Motor section 22a: First motor 22b: Second motor 22c: Third motor 23:Protective part 24: Connection part 25: Joints 30: Control Unit 40: Sensor section 50: Arm section 9: Walking surface S110: Walking step S120: Acquisition Step S130: Specific Step

Claims

1. A multi-legged walking robot with four or more legs that walks on walking surfaces including uneven terrain, The torso and, Multiple legs that contact the walking surface and support the torso, A control unit that controls multiple legs based on a pre-constructed walking learning model, Equipped with, The multiple legs are driven at different timings. The aforementioned walking learning model is constructed by reinforcement learning targeting a virtual robot on a virtual walking surface that mimics the aforementioned walking surface. The aforementioned walking learning model is Virtual environment information relating to the characteristics of the virtual walking surface, Virtual leg information relating to the characteristics of multiple legs for the virtual environment information, and The reward set for the result of the virtual robot executing the virtual leg information Constructed by the aforementioned reinforcement learning using The aforementioned virtual leg information is, Each of the multiple virtual legs in the virtual robot is enclosed by a virtual support region with the position in contact with the virtual walking surface as its vertex, The center of gravity of the virtual robot is the virtual center of gravity position projected onto the virtual walking surface and Including the relationship, The reward is set based on whether or not the virtual center of gravity position falls within the virtual support area. The torso is provided with a center of gravity measuring unit for measuring the center of gravity, The aforementioned leg portion is The support portion that contacts the walking surface, A motor unit connected to the aforementioned body and driving the aforementioned support unit Includes, The control unit uses the walking learning model to derive the torque of the motor unit corresponding to the center of gravity measured by the center of gravity measurement unit, and controls the motor unit. A multi-legged walking robot characterized by [feature].

2. The torso is provided with an altitude measuring unit that measures the height position of the torso with respect to the walking surface, The control unit, The height position measured by the altitude measurement unit and the environmental information relating to the characteristics of the walking surface are used as input information for the walking learning model. Multiple control conditions for the legs in response to the input information are derived as output information of the walking learning model, and the legs are controlled based on the output information. A multi-legged walking robot according to claim 1, characterized by the above.

3. The torso is further equipped with an arm section connected to the aforementioned body section for collecting agricultural products, The control unit controls the arm based on a harvesting learning model that has been pre-constructed by reinforcement learning targeting a virtual crop that mimics the crop and a virtual arm that mimics the arm. The aforementioned learning model for harvesting, Virtual crop information relating to the characteristics of the aforementioned virtual crop, Virtual arm information relating to the characteristics of the virtual arm with respect to the virtual crop information, and The harvest reward set based on the result of the virtual arm unit executing the virtual arm unit information. Constructed by the aforementioned reinforcement learning using The control unit uses the centroid as an explanatory variable of the harvesting learning model, acquires control information of the arm as an objective variable, and controls the arm based on the control information. A multi-legged walking robot according to claim 1, characterized by the above.

4. After stopping the multiple legs, the control unit controls the multiple joints included in the arm, The aforementioned arm portion includes a cutter portion and a gripper portion, The aforementioned cutter section is A cutting section for cutting the support portion that supports the crops, A rotating part linked to the cutting part, Includes, The aforementioned gripper portion is A holding part for holding the crop whose support portion has been cut by the cutter part, A movable part for adjusting the position of the holding part, Includes, The movable part is driven in conjunction with the rotation of the rotating part after the driving of the multiple joints has stopped, thereby changing the position of the holding part. A multi-legged walking robot according to claim 3, characterized by the above.

5. A walking method for a multi-legged walking robot according to any one of claims 1 to 4, A walking step that controls multiple legs and walks on the walking surface, An acquisition step to acquire the characteristics of the uneven ground in contact with the leg portion, A selection step involves referring to the aforementioned walking learning model to identify the control content of the legs corresponding to the characteristics of the uneven terrain, To be prepared A walking method characterized by the following.

Citation Information

Patent Citations

  • Self-propelled self-balancing picking robot

    CN110537419A

  • Lightweight electric four-foot robot

    CN110588828A

  • Quadruped robot gait control method based on reinforcement learning and CPG controller

    CN111208822A

  • Leg mechanism of leg-foot type robot and leg-foot type robot

    CN111232088A

  • Remotely-controlled multifunctional quadruped robot and operation method

    CN111301556A