Actuators and actuator design methodologies

Vision-based machine learning models enhance robotic navigation by creating a three-dimensional occupancy network, addressing sensor complexity and accuracy issues, enabling efficient and accurate obstacle detection and manipulation tasks.

JP2025535012APending Publication Date: 2025-10-22TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025518544
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-28
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing robotic systems face challenges in navigating complex environments due to limitations in sensor-based hardware complexity and accuracy, particularly in detecting dynamic and temporary obstacles using vision systems, which can lead to increased power consumption and interference from additional sensors like radar and lidar.

Method used

The use of vision-based machine learning models that process image data to create a robot-centric three-dimensional occupancy network, reducing sensor complexity by relying on image sensors and enhancing navigation through a projection onto a three-dimensional model, allowing for improved obstacle detection and manipulation tasks.

Benefits of technology

This approach enables more accurate and efficient autonomous navigation and manipulation by reducing sensor complexity, minimizing power consumption, and eliminating interference from additional sensors, thereby enhancing the robot's operational capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535012000001_ABST
    Figure 2025535012000001_ABST
Patent Text Reader

Abstract

A system or methodology for controlling the movement of a robot (600) using actuators, the system may include one or more first type actuators (1002) positioned at the robot's torso, shoulder, and hip positions, one or more second type actuators (1004) positioned at the robot's wrist positions, one or more third type actuators (1006) positioned at the robot's wrist positions, one or more fourth type actuators (1008) positioned at the robot's elbow and ankle positions, one or more fifth type actuators (1010) positioned at the robot's torso and hip positions, and one or more sixth type actuators (1012) positioned at the robot's knee and hip positions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 378,000, filed September 30, 2022, the entire contents of which are incorporated by reference in their entirety for all purposes.

[0002] The present disclosure relates to robots, and more particularly to actuator designs and methodologies. [Background technology]

[0003] Neural networks are relied upon in a variety of applications and increasingly form the basis of technology. For example, neural networks can be utilized to perform object classification on images acquired via a user device (e.g., a smartphone) or a camera system. In this example, the neural network may represent a convolutional neural network that applies convolutional layers, pooling layers, and one or more fully connected layers to classify objects depicted in the images.

[0004] Complex neural networks are further being used to enable autonomous or semi-autonomous driving capabilities for vehicles and the like. For example, unmanned aerial vehicles may utilize neural networks in part to enable navigation around real-world areas. In this example, the unmanned aerial vehicle may utilize sensors to detect approaching objects and navigate around the objects. As another example, a robot may implement a neural network to navigate around a real-world area. A robot's legs or arms may be formed from multiple connectors or members. The movement of each member may contribute to the robot's navigation around a real-world area and may be independently controlled by one or more actuators. The design and characteristics of these actuators may affect the robot's performance characteristics.

[0005] Embodiments of the present disclosure and their advantages are best understood by reference to the following detailed description: It should be understood that like reference numerals have been used to identify like elements shown in one or more of the figures, and that the designations therein are for the purpose of illustrating embodiments of the present disclosure and are not intended to limit the disclosure. Summary of the Invention [Means for solving the problem]

[0006] One aspect is directed to a system for motion control of a robot using actuators, the system may include one or more first type actuators positioned at the robot's torso, shoulder, and hip positions, one or more second type actuators positioned at the robot's wrist positions, one or more third type actuators positioned at the robot's wrist positions, one or more fourth type actuators positioned at the robot's elbow and ankle positions, one or more fifth type actuators positioned at the robot's torso and hip positions, and one or more sixth type actuators positioned at the robot's knee and hip positions.

[0007] Variations on the above embodiment further include one or more motors configured to cause movement of the one or more actuators.

[0008] A variation of the above embodiment further includes one or more batteries located in the body of the robot and connected to the one or more motors.

[0009] A variation of the above embodiment further includes a communication backbone communicatively connected to the one or more actuators and the one or more motors.

[0010] A variation on the above aspect further includes a processor communicatively connected to the communication backbone.

[0011] A variation of the above embodiment is for one or more batteries to be connected to a communications backbone.

[0012] A variation of the above aspect is that the communication backbone is configured to enable communication between sensors on the processor, motors, and actuators.

[0013] A variation of the above embodiment is that the processor is configured to control one or more motors and receive information from sensors on the actuators via a communications backbone.

[0014] A variation of the above embodiment is that one or more of the first type, second type, third type, fourth type, fifth type, and sixth type actuators include a rotary actuator.

[0015] A variation of the above embodiment is that at least one of the rotary actuators comprises a mechanical latch.

[0016] A variation of the above embodiment is that one or more of the first type, second type, third type, fourth type, fifth type, and sixth type actuators include a linear actuator.

[0017] A variation of the above embodiment is that at least one of the linear actuators comprises a planetary roller.

[0018] Another aspect is directed to a method for controlling movement of a robot using actuators, the method may include controlling a torso, shoulder, and hip position of the robot with one or more first type actuators, controlling a wrist position of the robot with one or more second type actuators, controlling the wrist position of the robot with one or more third type actuators, controlling elbow and ankle positions of the robot with one or more fourth type actuators, controlling the torso and hip positions of the robot with one or more fifth type actuators, and controlling knee and hip positions of the robot with one or more sixth type actuators.

[0019] A variation of the above embodiment further includes using one or more motors to control one or more actuators.

[0020] A variation of the above embodiment further includes powering the one or more motors with one or more batteries located in the torso of the robot.

[0021] A variation of the above embodiment further includes communicatively connecting the one or more actuators and the one or more motors over a communication backbone.

[0022] A variation on the above aspect further includes communicatively connecting a communications backbone to the processor.

[0023] A variation of the above embodiment further includes powering the communications backbone with one or more batteries.

[0024] A variation of the above embodiment further includes using a processor to control one or more motors and receive information from sensors on the actuators via a communications backbone.

[0025] A variation of the above embodiment is that one or more of the first type, second type, third type, fourth type, fifth type, and sixth type actuators include a rotary actuator, and one or more of the first type, second type, third type, fourth type, fifth type, and sixth type actuators include a linear actuator. [Brief explanation of the drawings]

[0026] The present invention will now be described with reference to the accompanying drawings, in which like reference numerals refer to like elements and in which:

[0027] [Figure 1A] FIG. 1 is a block diagram illustrating an example autonomous robot including multiple image sensors and an example processor system.

[0028] [Figure 1B] 1 is a block diagram illustrating an exemplary processor system for determining object / signal information based on received image information from an exemplary image sensor.

[0029] [Figure 1C] FIG. 10 shows an example of the resulting degree of robot vision.

[0030] [Figure 2] FIG. 1 is a block diagram of an example vision-based machine learning model including at least one processing network.

[0031] [Figure 3] FIG. 1 is a block diagram of a processing network.

[0032] [Figure 4] FIG. 1 is a block diagram of a modeling network.

[0033] [Figure 5A] FIG. 1 is a block diagram of a robot.

[0034] [Figure 5B] 1 illustrates an exemplary arrangement of battery packs and actuators on a robot according to the present disclosure.

[0035] [Figure 5C] FIG. 5C illustrates an exemplary actuator implemented in the robot of FIG. 5B according to the present disclosure.

[0036] [Figure 5D] FIG. 10 illustrates the battery pack installed in the robot, showing the battery back oriented vertically with a protective cover.

[0037] [Figure 5E] FIG. 1 illustrates the internal components of an exemplary rotary actuator according to the present disclosure.

[0038] [Figure 5F] FIG. 1 illustrates the internal components of an exemplary linear actuator according to the present disclosure.

[0039] [Figure 5G] FIG. 5F shows other internal components of the rotary actuator of FIG. 5E.

[0040] [Figure 5H] FIG. 5C illustrates other internal components of the linear actuator of FIG. 5F.

[0041] [Figure 5I] FIG. 5F shows other internal components of the rotary actuator of FIG. 5E.

[0042] [Figure 5J] 5F shows other internal components of the linear actuator of FIG. 5F.

[0043] [Figure 6]FIG. 5C illustrates the goals and methodology for the exemplary actuator motion (e.g., left hip yaw) shown in FIGS. 5B and 5C, including actuator torque and speed.

[0044] [Figure 7] FIG. 7 illustrates a single association between the goals and methodology shown in FIG. 6 and system cost and actuator mass.

[0045] [Figure 8] FIG. 8 is similar to FIG. 7, but includes multiple associations for use in selecting an optimized design for an actuator.

[0046] [Figure 9] FIG. 7 illustrates performance aspects in terms of torque and speed for actuator motion (e.g., left hip yaw) of the joint of FIG. 6.

[0047] [Figure 10] FIG. 1 illustrates the association between the performance at each position of the robot and the corresponding actuator type implemented at each position of the robot. DETAILED DESCRIPTION OF THE INVENTION

[0048] One or more aspects of the present application relate to improved techniques for autonomous or semi-autonomous (collectively referred to herein as autonomous) operation of machines, generally referred to as robots or robotic machines. In one or more embodiments, the robot may be configured or optimized to perform one or more tasks typically performed by human effort. In some applications, the robot may be humanoid in appearance, at least partially physically resembling a human or human effort. In other applications, the robot may not be constrained by the characteristic of being humanoid in appearance.

[0049] Thus, robots can navigate around real-world environments using vision-based sensor information. As can be appreciated, humans are capable of navigating within various environments and performing detailed tasks using vision and a deep understanding of their real-world surroundings. For example, humans can quickly identify objects (e.g., walls, boxes, machines, tools, etc.) and use these objects to inform navigation / movement (e.g., walking, running, avoiding collisions, etc.) and perform manipulation tasks (e.g., picking up objects, using machines / tools, moving between defined locations, etc.).

[0050] One or more aspects of the present application describe vision-based machine learning models that rely on increasing software complexity to enable reduced sensor-based hardware complexity while improving accuracy. For example, in some embodiments, only image sensors may be used. Through the use of image sensors or vision systems, such as one or more cameras, the described models accommodate refinement and improvement of vision-based robotic movement and task completion. As described below, the machine learning models can acquire images from image sensors and combine (e.g., stitch or fuse) the information contained therein. For example, the information can be combined into a vector space, which is then further processed by the machine learning model to extract objects, signals associated with the objects, etc.

[0051] Additionally, as described below, information may be projected based on a common virtual camera or virtual viewpoint to limit object occlusion and ensure a substantial range of object visibility. To simplify the explanation, objects may be positioned in a vector space according to their positions as seen by a common virtual camera. For example, the virtual camera may be set at a particular height above a robot having a vision system. In this example, the vector space may depict objects proximate to the robot as seen by a camera at that height (e.g., facing forward or diagonally forward). Another exemplary technique may rely on a bird's-eye view of objects positioned around the robot. For example, the bird's-eye view may allow objects to be viewed as positioned based on a virtual camera facing downward at a substantial height.

[0052] In some embodiments or scenarios, processing actual image data or processed virtual camera image data may have limitations or deficiencies in identifying objects that may be in the travel path. In some embodiments, objects may not be fully detectable based on the underlying image data. For example, environmental conditions may cause the image data to be inconsistent or otherwise incomplete. In other aspects, the trained machine-learned model may not be configured to detect certain objects, such as physical objects that may not have been typical objects found in a movement model (e.g., a kinematic model) for the robot, but may present some form of physical obstacle in a particular scenario. For example, temporary equipment or dynamically changing payloads may not be the types of objects the vision system is trained to recognize.

[0053] Thus, in some embodiments, the machine learning models described herein may be relied upon to detect, determine, or identify objects or physical obstacles based on a projection or mapping of vision system data (real or virtual) onto a robot-centric three-dimensional model. Illustratively, the three-dimensional model corresponds to a representation of the physical space within the robot's defined range, subdivided into a grid (e.g., blocks) of three-dimensional volume. Such three-dimensional blocks may individually be referred to as voxels. One or more machine-learned models may then characterize whether individual voxels within the grid of voxels are occluded in response to a query of visual information (e.g., image space). Illustratively, the characteristics of each voxel may be considered binary for purposes of occupancy networking. Furthermore, in some embodiments, the three-dimensional model may associate attribute or metric data, such as speed, orientation, type, associated object, etc., with each voxel. Thus, the processing results from each analysis of image data relate to an occupancy network in which one or more surrounding regions are associated with a prediction / estimation of obstacles / objects. Such occupancy network processing results may be independent of additional vision system processing techniques in which individual objects may be detected and characterized from vision system data.

[0054] The three-dimensional model and associated data may be referred to as an occupancy network. The occupancy network may be relied upon for static object detection, determination, identification, etc. The output of the described model and occupancy network may be used by a planning and / or navigation model or engine, for example, to achieve autonomous or semi-autonomous movement or manipulation. Additional description regarding bird's-eye view networks is included in U.S. Provisional Patent Application No. 63 / 260,439, which is incorporated herein by reference. In such applications, one or more aspects may be applied in the context of a robot.

[0055] The machine learning models described herein may include heterogeneous elements that may be trained end-to-end in some embodiments. As described below, images from an image sensor may be provided to respective backbone networks. In some embodiments, these backbone networks may be convolutional neural networks that output feature maps for subsequent use in the network. A transformer network, such as a self-attention network, may receive the feature maps and transform the information into an output vector space. For example, the output vector space may be associated with virtual cameras at various heights. The processed vector space may then be mapped to a three-dimensional model to provide an organization of one or more regions surrounding the robot. In one embodiment, the three-dimensional model corresponds to a grid of voxels (e.g., a three-dimensional box) that cumulatively models the region surrounding the robot. Each voxel is characterized by a prediction or probability that the mapped image data depicts an object / obstacle (e.g., the occupancy of the individual voxel). Additionally, each voxel may be associated with additional semantic data, such as velocity data, grouping criteria, or object type data. The semantic data may be provided as part of the occupancy network. Still further, in some embodiments, the individual voxel dimensions, commonly referred to as voxel offsets, can be further refined to distinguish or remove potential objects / obstacles that may be depicted in the image data but do not represent actual objects / obstacles utilized in navigation or motion control. For example, changes to the voxel offsets or individual voxel dimensions can account for environmental objects (e.g., dirt on the road) that may be depicted in the image data but are not modeled as obstacles in the occupancy network.

[0056] Still further, in some other embodiments, processing results (e.g., occupancy networks) can be further refined by utilizing historical information within an established time frame so that the occupancy network can be updated. In some embodiments, the update can identify potential discrepancies in the occupancy data, such as inconsistencies in occupancy or non-occupancy characteristics within the time frame. Such discrepancies can be based on errors or limitations in the vision system data. In another aspect, the update can identify verification or validation checks in successive occupancy networks, such as to increase confidence values ​​based on consistent processing results.

[0057] Thus, the disclosed technology enables the enhancement of autonomous movement or manipulation models while reducing sensor complexity. For example, other sensors (e.g., radar, lidar, etc.) may be removed during operation of the robot described herein. As can be appreciated, radar may introduce interference during robot operation that may lead to virtual objects being detected. Furthermore, lidar may introduce errors in certain environmental conditions and significantly complicate the manufacture of the robot. Furthermore, additional detection systems may increase the power consumption of the robot, which may limit its usability and functionality.

[0058] Although descriptions of autonomous robotic locomotion and operation are included herein, it will be understood that the present technology may be applied to other autonomous machines. For example, the machine learning models described herein may be used, in part, to autonomously operate unmanned ground vehicles, unmanned aerial vehicles, and the like. Furthermore, in some embodiments, reference to robots may not be limited to any particular type of environment, such as built environments, factories, commercial facilities, domestic facilities, public areas, safety and protection environments, and the like. Block Diagram - Robotic Handling System

[0059] 1A is a block diagram illustrating an example autonomous robot 600 including multiple image sensors 102A-102F and an example processor system 120. The image sensors may include cameras positioned around the robot 600. For example, the cameras may enable a substantially 360-degree view of the robot 600's surroundings.

[0060] The image sensor can acquire images that are used by the processor system 120 to determine information associated with at least objects located proximate to the vehicle 100. The images can be acquired at a particular frequency, such as 30 Hz, 36 Hz, 60 Hz, or 65 Hz. In some embodiments, certain image sensors can acquire images more quickly than other image sensors. As described below, these images can be processed by the processor system 120 based on the vision-based machine learning models described herein. For illustrative purposes, the processing system 120 is shown as being located in a region of the robot 600 that resembles a human head. However, such an arrangement is merely exemplary in nature and is not required.

[0061] In one embodiment, the first image sensor A includes three image sensors that are laterally offset from one another. For example, the camera housing may include three image sensors that face forward. In this example, a first of the image sensors may have a wide-angle (e.g., fisheye) lens. A second of the image sensors may have a normal or standard lens (e.g., a 35mm equivalent focal length, 50mm equivalent, etc.). A third of the image sensors may have a zoom or narrow field of view lens. In this manner, three images of various focal lengths may be acquired by the robot 600 in the forward direction.

[0062] The second image sensor may be side-facing or rear-facing and located on the left side of the robot 600. Similarly, the third image sensor may also be side-facing or rear-facing and located on the right side of the robot 600. The fourth image sensor may face behind the robot 600 and be positioned to capture images in a direction behind the robot 600 (e.g., assuming the robot 600 is moving forward).

[0063] Although the illustrated embodiment includes an image sensor, it will be appreciated that additional or fewer image sensors may be used and still be included within the techniques described herein.

[0064] The processor system 120 can use the vision-based machine learning models described herein to acquire images from an image sensor and detect objects and signals associated with the objects. Based on the objects, the processor system 120 can adjust one or more position or operational characteristics or tasks. For example, the processor system 120 can cause the robot 600 to turn, slow down, perform a predetermined task, avoid a collision, select a position path, generate an alert, etc. Although not described herein, as can be understood, the processor system 120 can execute one or more planning and / or navigation engines or models that implement autonomous driving using output from the vision-based machine learning models.

[0065] In some embodiments, processor system 120 may include one or more matrix processors configured to rapidly process information associated with a machine learning model. In some embodiments, processor system 120 may be used to perform convolutions associated with a forward pass through a convolutional neural network. For example, input data may be convolved with weight data. Processor system 120 may include multiple multiply-accumulate units to perform the convolutions. As an example, the matrix processor may use input data and weight data that are organized or formatted to facilitate larger convolution operations.

[0066] For example, the input data may be in the form of a three-dimensional matrix or tensor (e.g., two-dimensional data across multiple input channels). In this example, the output data may be across multiple output channels. Thus, the processor system 120 may process larger input data by merging or flattening each two-dimensional output channel into a vector, so that the entire channel, or a significant portion thereof, may be processed by the processor system 120. As another example, data may be efficiently reused, such that weight data may be shared across convolutions. With respect to an output channel, the weight data 106 may represent the weight data (e.g., kernel) used to calculate that output channel.

[0067] Additional exemplary descriptions of processor systems that may use one or more matrix processors are included in U.S. Pat. Nos. 11,157,287, 11,409,692, and 11,157,441, which are incorporated by reference in their entireties and form part of this disclosure as if set forth herein.

[0068] FIG. 1B is a block diagram illustrating an example processor system 120 that determines object / signal information 124 based on received image information 122 from an example image sensor.

[0069] Image information 122 includes images from image sensors positioned around a robot (e.g., robot 600). In the illustrated example of FIG. 1A, there are eight image sensors, and therefore eight images are represented in FIG. 1B. For example, the top row of image information 122 includes three images from forward-facing image sensors. As described above, image information 122 may be received at a particular frequency such that the images shown represent particular timestamps of the images. In some embodiments, image information 122 may represent high dynamic range (HDR) images. For example, different exposures may be combined to form an HDR image. As another example, images from the image sensors may be preprocessed (e.g., using a machine learning model) to convert them into HDR images.

[0070] In some embodiments, each image sensor may acquire multiple exposures, each with a different shutter speed or integration time. For example, the different integration times may be greater than a threshold time difference apart. In this example, in some embodiments, there may be three integration times spaced approximately an order of magnitude apart in time. The processor system 120, or a different processor, may select one of the exposures based on a measure of clipping associated with the image. In some embodiments, the processor system 120, or a different processor, may form an image based on a combination of the multiple exposures. For example, each pixel of the formed image may be selected from one of the multiple exposures based on pixels that do not contain values ​​(e.g., red, green, blue) that are clipped (e.g., exceed a threshold pixel value).

[0071] The processor system 120 can execute a vision-based machine learning model engine 126 to process the image information 122. An example of a vision-based machine learning model is described in more detail below. As described herein, the vision-based machine learning model can combine information contained in images. For example, each image can be provided to a specific backbone network. In some embodiments, the backbone network can represent a convolutional neural network. The outputs of these backbone networks can then, in some embodiments, be combined (e.g., formed into a tensor) or provided as separate tensors to one or more further portions of the model. In some embodiments, an attention network (e.g., cross-attention) can receive the combination or input tensors associated with each image sensor. The combined output can then be provided for analysis, illustratively for determining object detection within the processed image data, as described below.

[0072] 1B, the vision-based machine learning model engine 126 may output object / signal information 124. This information 124 may represent information identifying the object depicted in the image information 122. For example, the information 122 may include one or more of the object's location (e.g., information associated with a cuboid about the object), the object's velocity, the object's acceleration, the object's type or classification, whether the automobile object has a door open, etc. Examples of the object / signal information 124 are described below with respect to FIG.

[0073] With respect to the cuboid, exemplary information 122 may include position information (e.g., with respect to a common virtual or vector space), size information, shape information, etc. For example, the cuboid may be three-dimensional. Exemplary information 122 may further include whether the object is encroaching on an intended path of travel for robot 600. An example of the resulting degree of visibility for robot 600 is shown in exemplary form in FIG. 1C.

[0074] Additionally, as described below, the vision-based machine learning model engine 126 can process multiple images spread over time. For example, a video module can be used to analyze images (e.g., their feature maps generated by a backbone network or subsequently by a vision-based machine learning model) selected from within a previous threshold amount of time (e.g., 3 seconds, 5 seconds, 15 seconds, an adjustable amount of time, etc.). In this way, objects can be tracked over time to monitor their location even when the processor system 120 is temporarily occluded.

[0075] In some embodiments, the vision-based machine learning model engine 126 can output information forming one or more images. Each image can encode specific information, such as the location of an object. For example, a bounding box of an object positioned around the autonomous robot can be formed in the image. In some embodiments, the projections 322 and 422 in FIGS. 3B and 4B can be images generated by a vision-based machine learning model.

[0076] 2 is a block diagram of an exemplary vision-based machine learning model that includes at least one processing network 210. The exemplary model may be executed by an autonomous robot, such as robot 600. Accordingly, the actions of the model may be understood to be executed by a processor system (e.g., system 120) included in robot 600.

[0077] In the illustrated example, images 202A-202F are received by a vision-based machine learning model. These images 202A-202F may be acquired from image sensors positioned around the robot, such as the image sensors of FIG. 1A. The vision-based machine learning model includes a backbone network 200 that receives each image as an input. Thus, backbone network 200 processes the raw pixels contained in images 202A-202F. In some embodiments, backbone network 200 may be a convolutional neural network. For example, there may be 5, 10, 15, etc. convolutional layers in each backbone network.

[0078] In some embodiments, backbone network 200 may include a residual block, a residual network regulated by a recurrent neural network, etc. Additionally, backbone network 200 may include a weighted bidirectional feature pyramid network (BiFPN). The output of the BiFPN may represent multi-scale features determined based on images 202A-202H. In some embodiments, Gaussian blurring may be applied to portions of the images during training and / or inference. For example, road edges may be peak-like in that they are sharply defined in the image. In this example, Gaussian blurring may be applied to the road edges to enable bleeding of visual information, which may allow the road edges to be detectable by a convolutional neural network.

[0079] Additionally, some of the backbone networks 200 may pre-process the images, such as performing rectification, cropping, etc. With respect to cropping, the image 202C from the fisheye's forward-facing lens may be cropped vertically to remove certain elements included based on the curvature of the robot 600 (e.g., curvature or protective features associated with the robot body).

[0080] With respect to rectification, the robot 600 described herein may be an example of a robot that may be implemented in a variety of environments and implementations. Due to manufacturing tolerances and / or differences in use of the robot 600, the image sensors within the robot may be angled or otherwise positioned slightly differently (e.g., differences in roll, pitch, and / or yaw). Furthermore, different models of the robot 600 may run the same vision-based machine learning model. These different models may have image sensors positioned and / or angled differently. The vision-based machine learning models described herein may be trained, at least in part, using information aggregated from a robot fleet. Thus, differences in image perspectives may be apparent due to slight differences between the angles or positions of the image sensors within the robots 600 included in the robot fleet.

[0081] Thus, to address these differences, rectification may be performed through backbone network 200. For example, a transformation (e.g., an affine transformation) may be applied to images 202A-202F, or portions thereof, to normalize the images. In this example, the transformation may be based on camera parameters associated with the image sensor (e.g., image sensor), such as extrinsic and / or intrinsic parameters. In some embodiments, the image sensor may undergo an initial, and optionally repeated, calibration step. For example, as the robot performs a position or manipulation task, the camera may be calibrated to ascertain camera parameters that may be used in the rectification process. In this example, specific markings (e.g., path symbols) may be used to signal the calibration. Rectification may optionally represent one or more layers of backbone network 200, where values ​​for the transformation are learned based on training data.

[0082] Thus, backbone network 200 can output feature maps (e.g., tensors) that are used by processing network 210. In some embodiments, the outputs from backbone network 200 can be combined into matrices or tensors. In some embodiments, the outputs can be provided to processing network 210 as multiple tensors (e.g., eight tensors in the illustrated example). In the illustrated example, the outputs are referred to as visual information 204 that is input to network 210.

[0083] The output tensors from backbone network 200 may be combined (e.g., fused) together into respective virtual camera spaces (e.g., vector spaces) via processing network 210. Image sensors positioned around robot 600 may be at different heights on the robot. For example, left and right image sensors may be positioned higher than front and rear image sensors. Thus, virtual camera spaces may be used to enable consistent views of objects positioned around robot 600. As described above, processing network 210 may use more virtual camera spaces.

[0084] For certain information determined by the vision-based machine learning model, kinematic information 206 of the autonomous robot may be used. Exemplary kinematic information 206 may include the velocity, acceleration, yaw rate, etc. of the robot 600. In some embodiments, the images 202A-202F may be associated with kinematic information 206 determined for the time the images 202A-202F were acquired or a similar time. For example, the kinematic information 206, such as velocity, yaw rate, acceleration, etc., may be encoded (e.g., embedded in a latent space) and associated with the image.

[0085] Thus, with respect to determining the velocity of one or more objects, such as a robot, the vision-based machine learning model may use the velocity of the autonomous robot itself when determining the relative velocity of the objects. Additionally, processing network 210 may process images at a particular frame rate. Thus, consecutive images may be acquired that are the same, or substantially the same, time delta apart. Based on this information, processing network 210 may be trained to estimate the relative velocity of the objects.

[0086] Exemplary outputs from processing network 210 may represent information associated with an object, such as position (e.g., position in virtual camera space), depth, etc. For example, the information may relate to a cuboid associated with an object positioned around robot 600. The outputs may also represent signals utilized by the processor system to autonomously drive robot 600. Exemplary signals may include indications of portions of the processed visual signal that can be characterized as depicting an object or not depicting an object.

[0087] The output may be generated via a forward pass through the network 210. In some embodiments, the forward pass may be calculated at a particular frequency (e.g., 24 Hz, 30 Hz, etc.). In some embodiments, the output may be used, for example, via a planning engine. As an example, the planning engine may determine driving actions (e.g., accelerating, turning, braking, etc.) to be performed by the autonomous robot based on periscope and panoramic views of the real-world environment.

[0088] As further shown in FIG. 2 , the output of the processing network is utilized as an input to the modeling network 220. The output of the modeling network 220 can then correspond to a processing result corresponding to a modeled occupancy network. As previously described, the modeled occupancy network corresponds to a mapped three-dimensional model, resulting in the organization of one or more regions surrounding the robot. In one embodiment, the three-dimensional model corresponds to a grid of voxels (e.g., a three-dimensional box) that cumulatively models the region surrounding the robot 600. Each voxel is characterized by a prediction or probability that the mapped image data depicts an object / obstacle (e.g., the occupancy of the individual voxel). Additionally, each voxel may be associated with additional semantic data, such as velocity data, grouping criteria, object type data, etc. The semantic data may be provided as part of the occupancy network. Still further, in some embodiments, the individual voxel dimensions, commonly referred to as voxel offsets, can be further refined to distinguish or remove potential objects / obstacles that may be depicted in the image data but do not represent actual objects / obstacles utilized for navigation or motion control. For example, changes to voxel offsets or individual voxel dimensions can account for environmental objects (e.g., dirt in a manufacturing site) that may be depicted in the image data but are not modeled as obstacles in the occupancy network.

[0089] Still further, in some other embodiments, processing results (e.g., occupancy networks) can be further refined by utilizing historical information within an established time frame so that the occupancy network can be updated. In some embodiments, the update can identify potential discrepancies in the occupancy data, such as inconsistencies in occupancy or non-occupancy characteristics within the time frame. Such discrepancies can be based on errors or limitations in the vision system data. In another aspect, the update can identify verification or validation checks in successive occupancy networks, such as to increase confidence values ​​based on consistent processing results.

[0090] 3 is a block diagram of processing network 210. As described in FIG. 2, processing network 210 may be used to generate vision image data from one or more cameras or vision systems. In the illustrated example, visual information 204 from a backbone network (e.g., network 200) is provided as input to fixed projection engine 302.

[0091] The fixed projection engine 302 can project information into a virtual camera space associated with a virtual camera. As described above, the virtual camera can be positioned 1 meter, 1.5 meters, 2.5 meters, etc. above the autonomous robot running the vision-based machine learning model. Without being bound by theory, it can be understood that pixels of an input image can be mapped to the virtual camera space. For example, a lookup table can be used in combination with external and internal camera parameters associated with an image sensor (e.g., the image sensor of FIG. 1A ).

[0092] As an example, each pixel may be associated with a depth within the virtual camera space. Each pixel may represent a ray from the image, with the ray extending into the virtual camera space. For a given pixel, a depth may be assumed or otherwise identified. With respect to the ray, the fixed projection engine 302 may identify two different depths along the ray from a given pixel. In some embodiments, these depths may be 5 meters and 50 meters. In other embodiments, the depths may be 3 meters, 7 meters, 45 meters, 52 meters, etc. The processor system 120 may then form the virtual camera space based on a combination of these ray values ​​relative to the pixels of the image. As can be appreciated, the location of a pixel in the input image may substantially correspond to a location within one or more tensors forming the visual information 204.

[0093] In some embodiments, the vector space may be warped by the processing network 210 so that a portion of the three-dimensional vector space is magnified. For example, an object depicted in a view of a real-world environment seen by a camera positioned 1.5 meters, 2 meters, etc. may be warped by the processing network 210. The vector space may be warped so that portions of interest may be magnified or otherwise made more prominent. For example, the width and height dimensions may be warped to elongate the object. To accomplish this warping, training data may be used whose labeled output is the object position adjusted according to the warping. Additionally, the fixed projection engine 302 may warp, for example, the height dimension, to ensure that the object is magnified in at least one dimension.

[0094] The output from the fixed projection engine 302 is provided as input to the frame selector engine 304. To ensure that objects can be tracked over time, even while temporarily occluded, the vision-based machine learning model may utilize multiple frames during a forward pass through the model. For example, each frame may be associated with a time, or a short time range, at which an image sensor is triggered to acquire an image. Thus, the frame selector engine 304 may select visual information 204 corresponding to images taken at different times within a threshold amount of time prior.

[0095] For example, the visual information 204 may be output by the processor system 120 at a particular frame rate (e.g., 20 Hz, 24 Hz, 30 Hz). The visual information 204 may then be queued or otherwise stored by the processor system 120, followed by the fixed projection engine 302. For example, the visual information 204 may be temporally indexed. Thus, the frame selector engine 304 may retrieve the visual information from a queue or other data storage element. In some embodiments, the frame selector engine 304 may retrieve 12, 14, 16, etc. frames spread over the previous 3, 5, 7, 9 seconds (e.g., visual information associated with the 12, 14, or 16 timestamps at which the images were captured). In some embodiments, these frames may be evenly spaced in time over the previous time period. While the discussion of frames is included herein, it will be understood that feature maps associated with image frames captured at a particular time or within a short time range may be selected by the frame selector engine 304.

[0096] The output from the frame selector engine 304 may, in some embodiments, represent a combination of the above frames 306A-N. For example, the output may be combined to form a tensor that is subsequently processed by the rest of the processing network 210.

[0097] For example, the outputs 306A-N (temporally indexed features) may be provided to multiple video modules. In the illustrated example, two video modules 308A-B are used. The video modules 308A-B may represent convolutional neural networks that enable the processor system 120 to perform three-dimensional convolutions. For example, the convolutions may result in a blending of spatial and temporal dimensions. In this manner, the video modules 308A-B may enable tracking of motion and objects over time. In some embodiments, the video modules may represent attention networks (e.g., spatial attention).

[0098] With respect to video module 308A, kinematic information 206 associated with an autonomous robot executing a vision-based machine learning model may be input to module 308A. As described above, kinematic information 206 may represent one or more of acceleration, velocity, yaw rate, turning information, braking information, etc. Kinematic information 206 may be further associated with each of frames 306A-N selected by frame selector engine 304. Thus, video module 308A may, by way of example, encode this kinematic information 206 for use in determining the velocity of objects in the robot's 600 periphery. With respect to processing network 210, velocity may represent allocentric velocity.

[0099] Processing network 210 includes heads 310, 312 for determining different information associated with an object. For example, as shown in Figure 2, head 310 can determine velocity associated with the object, while head 312 can determine position information, etc.

[0100] In general, the vision-based machine learning model described herein may include multiple trunks or heads. As known to those skilled in the art, these trunks or heads (collectively referred to herein as heads) extend from a common portion of the neural network and may be trained as experts to determine specific information. For example, a first head may be trained to output the velocity of each of the objects positioned around the robot. As another example, a second head may be trained to output a specific signal describing a feature or information associated with the object. An exemplary signal may include whether a nearby door is open or ajar.

[0101] In addition to being experts in specific information, the separation into different heads allows for staged training to quickly incorporate new training data. As new training information is acquired, the portions of the machine learning model that would benefit most from the training information can be quickly updated. In this example, the training information may represent images or video clips of specific real-world scenarios collected by the robot in real-world operation. Thus, a specific head or heads can be trained, and the weights included in these portions of the network can be updated. For example, other portions (e.g., earlier portions of the network) may not have updated weights to reduce training time and the time to update the robot 600.

[0102] In some embodiments, training data targeted to one or more of the heads may be adjusted to focus on those heads. For example, an image may be masked (e.g., loss masked) so that only certain pixels of the image are monitored and not others. In this example, certain pixels may be assigned a value of 0, while other pixels may maintain their value or be assigned a value of 1. Thus, if a training image depicts a rarely seen object (e.g., an object with a relatively novel form) or signal (e.g., a known object with an irregular shape), the training image may optionally be masked to focus on that object or signal. During training, the generated errors may be used to train the labeler on the loss of pixels associated with the object or signal. Thus, only heads associated with this type of object or signal may be updated.

[0103] To ensure sufficient training data is acquired, the robot 600 can optionally run a classifier that is triggered to acquire images that meet certain conditions. For example, a robot operated by an end user can automatically acquire training images depicting, for example, tire spray, rain conditions, snow, fog, fire smoke, etc. Further description regarding the use of classifiers is provided in U.S. Patent Publication No. 2021 / 0271259, which is incorporated herein by reference in its entirety as if set forth herein.

[0104] 4 is a block diagram of modeling network 220. As described in FIG. 2, modeling network 220 can be used to generate a three-dimensional model from input visual information, such as processing network 210. Mapping engine 402 can be utilized to map feature data from input image data and project the input data onto the three-dimensional model. As described above, the modeled occupancy network corresponds to the mapped three-dimensional model and provides an organization of one or more regions surrounding the robot. In one embodiment, the three-dimensional model corresponds to a grid of voxels (e.g., a three-dimensional box) that cumulatively model the region surrounding robot 600.

[0105] The output from the mapping engine 402 is provided as input to a query engine 404. For each voxel in the grid of voxels, the query engine 404 illustratively queries the image data to determine whether an object / obstacle is depicted or detected in the image data. As previously mentioned, the determination of an object or obstacle (e.g., occupancy) is a binary determination. For example, any portion of a voxel associated with an object / obstacle may indicate that the voxel is "occupied," regardless of whether the entire voxel encompasses or fills the voxel.

[0106] Output from the query engine 404 may be provided to the processing engine 406 for additional processing. In one aspect, the processing engine 406 may implement voxel offsets or dimension changes to the modeled grid. Illustratively, while voxel occupancy determinations are considered binary, the processing engine 406 may be configured to adjust the dimensions of individual voxels so that the occupancy network provided for navigation or control information may better approximate the details or contours of an object or obstacle. This is as opposed to a more geometric approximation that occurs when voxel dimensions remain static. In another aspect, the processing engine 406 may also create semantics that can group sets of voxels or associate voxels characterized as being part of the same object. Illustratively, an object / obstacle may span multiple modeled voxel spaces such that each voxel is considered individually "occupied." The model may further include organizational information or type identifiers that enable command and control components to consider voxels associated with the same object for decisions.

[0107] As further shown in FIG. 4, the metric engine can also receive and calculate various metrics or additional semantic data for each voxel. Such metric or semantic data may include, but is not limited to, velocity, orientation, type, associated object, and individual voxels. Illustratively, the generation of a three-dimensional model does not need to distinguish between objects that are essentially static and objects that are essentially dynamic. Rather, for each time instance of the vision data, the three-dimensional model can consider a voxel to be either occupied or unoccupied. However, command and control mechanisms may wish to account for the kinematics of the object (e.g., static vs. dynamic). Thus, an occupancy network (unrelated to dynamics) can associate voxel data with additional metrics / semantics that can facilitate the use of occupancy network processing results.

[0108] Thus, the processing results for each analysis of image data relate to an occupancy network in which one or more surrounding regions are associated with obstacle / object predictions / estimations. Such occupancy network processing results may be independent of additional vision system processing techniques in which individual objects may be detected and characterized from the vision system data. Robot Block Diagram and Architecture

[0109] 5A shows a block diagram of a robot 600. The robot 600 may include one or more motors 602 that cause movement of one or more actuators or actuating joints 604. One or more actuators 604 may be associated with each joint of each appendage or limb of the robot 600. For example, in certain embodiments, each limb may include multiple joints or links, and each joint includes multiple actuators (e.g., one or more rotary actuators and one or more linear actuators). In certain embodiments, one or more rotary actuators enable rotation of a link about the joint axis of an adjacent link. In certain embodiments, one or more linear actuators enable translation between the links.

[0110] In certain embodiments, one or more limbs (e.g., arms, legs) each include a series of rotational actuators. In certain embodiments, a first rotational actuator is associated with the shoulder or hip, a second rotational actuator is associated with the elbow or knee, and a third rotational actuator is associated with the wrist or ankle. Of course, more or less than three rotational actuators may be used within a single limb and still fall within the scope of the present disclosure.

[0111] In certain embodiments, one or more limbs (e.g., arms, legs) further comprise a series of linear actuators. In certain embodiments, a first linear rotary actuator is associated with the shoulder or hip, a second linear actuator is associated with the elbow or knee, and a third linear actuator is associated with the wrist or ankle. Of course, more or less than three linear actuators may be used within a single limb and still fall within the scope of the present disclosure.

[0112] Illustrative examples of actuators or manipulated joints 604 are shown in Figures 5B and 5C. In certain embodiments, the robot 600 includes 28 actuators 604. Of course, more or fewer than 28 actuators may be used by the robot 600 and still fall within the scope of the present disclosure. In certain embodiments, the number and location of actuators 604 in any given limb are selected such that the limb achieves six degrees of freedom (e.g., forward / backward, up / down, left / right, yaw, pitch, roll).

[0113] In certain embodiments, the one or more motors 602 may be electric, pneumatic, or hydraulic. Electric motors may include, for example, induction motors, permanent magnet motors, etc. In certain embodiments, the one or more motors 602 drive one or more of the actuators 604. In certain embodiments, a motor 602 is associated with each actuator 604.

[0114] Exemplary embodiments of rotary actuator 500 are shown in Figures 5E, 5G, and 51. In some embodiments, rotary actuator 500 may include a mechanical clutch 502 and an angular contact ball bearing 504 coupled to shaft 506 and integrated into a high-speed side 510 of rotary actuator 500. In some embodiments, rotary actuator 500 may include a cross-roller bearing 512 on a low-speed side 520 of rotary actuator 500. In some embodiments, rotary actuator 500 may include a strain wave gearing 514 disposed between high-speed side 510 and low-speed side 520. In some embodiments, rotary actuator 500 may further include a magnet 516 coupled to an outer surface of rotor 513. In some embodiments, the rotary actuator 500 may also include one or more sensors, such as, for example, an input position sensor 522 configured to detect the angular position of a fast side 510 of the rotary actuator 500 and an output position sensor 524 configured to detect the angular position of a slow side 520 of the rotary actuator 500, as well as a non-contact torque sensor 518 configured to monitor the output torque of the rotary actuator 500.

[0115] Exemplary embodiments of a linear actuator 550 are shown in Figures 5F, 5H, and 5J. In some embodiments, the linear actuator 550 may include a planetary roller 552 disposed on a slow (or linear) side 560 of the linear actuator 550 and disposed between an actuator shaft 561 on the slow side 560 and a rotor 574 on a fast (or rotating) side 570 to provide stability. In some embodiments, the linear actuator 550 may include an inverted roller screw 554 that functions as a gear train between the slow side 560 and the fast side 570 of the linear actuator 550 to enable efficiency and durability. In some embodiments, the linear actuator 550 may further include a ball bearing 562 proximate one end of the fast side 570 of the linear actuator 550 and a four-point contact bearing 564 proximate the other end of the fast side 570 of the linear actuator 550, both of which are disposed between the rotor 574 on the fast side 570 and an enclosure 575 of the linear actuator 550. In some embodiments, linear actuator 550 may also include a stator 572 coupled to an enclosure 575. In some embodiments, linear actuator 550 may include a magnet 566 coupled to an outer surface of rotor 574. In some embodiments, linear actuator 550 may also include one or more sensors, such as, for example, a force sensor 567 attached to a main shaft 580 of linear actuator 550 and configured to monitor a force on main shaft 580, and a position sensor 568 attached to enclosure 575 and configured to detect an angular position of rotor 574.

[0116] As known to those skilled in the art, the battery 606 may include one or more battery packs, each with multiple cells, and may be used to power an electric motor. Figure 5D shows the battery pack installed in the robot 600, showing the battery pack oriented vertically with a protective cover.

[0117] The robot 600 further includes a communication backbone 608 configured to provide communication capabilities between the processor 610, the motors 602, the actuators 604, the battery components 606, the sensors, etc. The communication backbone 608 may illustratively be configured to directly connect the components via a common backbone channel that may form one or more communication loops. Such interaction allows for redundancy and failure of individual components without further disruption to the ability of other components to communicate.

[0118] Additionally, the robot includes a processor system 120 that processes data, such as images received from image sensors positioned around the robot 600. The processor system 120 can additionally output information and receive information (e.g., user input).

[0119] Figure 6 shows the goals and methodology for the motion of the example actuator 604 shown in Figures 5B and 5C (e.g., left hip yaw), including actuator torque and speed. Figure 7 shows a single association between the goals and methodology shown in Figure 6 and system cost and actuator mass. Figure 8 is similar to Figure 7 but includes multiple associations for use in selecting an optimized design for the actuator. Figure 9 shows performance aspects in terms of torque and speed for the actuator motion of the joint of Figure 6 (e.g., left hip yaw).

[0120] In one aspect, a method for selecting an actuator is disclosed herein. As shown in FIGS. 6-9 , various analyses can be performed for each type of motion at each position or joint of the robot 600 to determine which actuators will be used at each position of the robot 600. As shown in FIG. 10 , a performance graph (e.g., a system cost graph) can be created for each type of motion at each position (e.g., right shoulder yaw, right shoulder roll, or right shoulder pitch). In some embodiments, the method for selecting an actuator can include creating such performance graphs for multiple types of motion at multiple positions of the robot 600 and then grouping the motion types at the various positions by their commonality. In some embodiments, the performance graphs for multiple types of motion at multiple positions can be grouped into six types, each corresponding to a different actuator. For example, as shown in FIG. 10, the actuator system disclosed herein may include one or more first type actuators 1002 arranged at the robot's torso, shoulder, and hip positions, one or more second type actuators 1004 arranged at the robot's wrist positions, one or more third type actuators 1006 arranged at the robot's wrist positions, one or more fourth type actuators 1008 arranged at the robot's elbow and ankle positions, one or more fifth type actuators 1010 arranged at the robot's torso and hip positions, and one or more sixth type actuators 1012 arranged at the robot's knee and hip positions.

[0121] All of the processes described herein may be embodied and fully automated via software code modules executed by a computing system including one or more computers or processors. The code modules may be stored on any type of non-transitory computer-readable medium or other computer storage device. Some or all of the methods may be embodied in dedicated computer hardware.

[0122] Many variations beyond those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain operations, events, or functions of any of the algorithms described herein may be performed in a different order, or may be added, merged, or omitted entirely (e.g., not all acts or events described may be necessary to implement an algorithm). Furthermore, in certain embodiments, operations or events may be performed simultaneously rather than sequentially, for example, via multithreading, interrupt processing, or multiple processors or processor cores, or on other parallel architectures. Additionally, different tasks or processes may be performed by different machines and / or computing systems that can function together.

[0123] The various illustrative logic blocks, modules, and engines described in connection with the embodiments disclosed herein may be implemented or executed by a machine such as a processing unit or processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be a microprocessor, but in alternative examples, the processor may be a controller, microcontroller, or state machine, combinations thereof, etc. The processor may include electrical circuitry configured to process computer-executable instructions. In another embodiment, the processor includes an FPGA or other programmable device that performs logical operations without processing computer-executable instructions. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. While described herein primarily with reference to digital technology, a processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented with analog circuitry or mixed analog and digital circuitry. The computing environment may include any type of computer system, including, but not limited to, a computer system based on a computational engine within a microprocessor, mainframe computer, digital signal processor, portable computing device, device controller, or appliance, to name a few.

[0124] Conditional language, particularly "can," "could," "might," or "may," is understood in its commonly used context to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless otherwise specified. Thus, such conditional language is not intended to generally imply that features, elements, and / or steps are somehow required in one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps should be included in or performed in any particular embodiment, with or without user input or prompting.

[0125] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is understood in its commonly used context, unless otherwise indicated, to indicate that an item, term, etc. can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that a particular embodiment requires at least one of X, at least one of Y, or at least one of Z to each be present.

[0126] Any process descriptions, elements, or blocks in the flow diagrams described herein and / or shown in the accompanying figures should be understood as potentially representing modules, segments, or portions of code that comprise one or more executable instructions for implementing a particular logical function or element in the process. As will be appreciated by those skilled in the art, alternative implementations in which elements or functions may be omitted, performed, or described in a different order than that shown, including substantially simultaneously or in reverse order, depending on the functionality involved, are included within the scope of the embodiments described herein.

[0127] Unless otherwise specified, articles such as "a" or "an" should generally be construed to include one or more listed items. Thus, a phrase such as "a device configured to" is intended to include one or more listed devices. Such one or more listed devices may also be collectively configured to perform the stated enumeration. For example, "a processor configured to perform enumerations A, B, and C" may include a first processor configured to perform enumeration A working in conjunction with a second processor configured to perform enumerations B and C.

[0128] It should be emphasized that many variations and modifications can be made to the above-described embodiments, and that the elements thereof are to be understood as being among other acceptable examples, and all such modifications and variations are intended to be included herein within the scope of the present disclosure.

Claims

1. A system for controlling the motion of a robot using an actuator, comprising: one or more first type actuators located at the torso, shoulder, and hip locations of the robot; one or more second type actuators located at the wrist of the robot; one or more third type actuators disposed at the wrist of the robot; one or more fourth type actuators located at the elbows and ankles of the robot; one or more fifth type actuators arranged at the torso position and the hip joint position of the robot; one or more sixth type actuators arranged at knee positions and at hip positions of the robot; A system comprising:

2. The system of claim 1 , further comprising one or more motors configured to cause movement of the one or more actuators.

3. The system of claim 2 , further comprising one or more batteries disposed at the torso location of the robot and connected to the one or more motors.

4. The system of claim 3 , further comprising a communication backbone communicatively connected to the one or more actuators and the one or more motors.

5. The system of claim 4 further comprising a processor communicatively connected to the communication backbone.

6. The system of claim 5 , wherein the one or more batteries are connected to the communications backbone.

7. The system of claim 5 , wherein the communication backbone is configured to enable communication between sensors on the processor, the motor, and the actuator.

8. The system of claim 7 , wherein the processor is configured to control the one or more motors and receive information from the sensors on the actuators via the communication backbone.

9. 10. The system of claim 1, wherein one or more of the first type, the second type, the third type, the fourth type, the fifth type, and the sixth type actuators comprise a rotary actuator.

10. The system of claim 9 , wherein at least one of the rotary actuators comprises a mechanical latch.

11. 10. The system of claim 1, wherein one or more of the first type, the second type, the third type, the fourth type, the fifth type, and the sixth type actuators comprise a linear actuator.

12. The system of claim 11 , wherein at least one of the linear actuators comprises a planetary roller.

13. 1. A method for controlling the motion of a robot using actuators, comprising: controlling the torso, shoulder, and hip positions of the robot using one or more actuators of a first type; controlling a wrist position of the robot using one or more second type actuators; controlling the wrist position of the robot using one or more third type actuators; controlling elbow and ankle positions of the robot using one or more fourth type actuators; controlling the torso position and the hip position of the robot using one or more fifth type actuators; controlling knee and hip positions of the robot using one or more sixth type actuators; A method comprising:

14. The method of claim 13 , further comprising controlling the one or more actuators with one or more motors.

15. The method of claim 14 , further comprising powering the one or more motors with one or more batteries located at the torso location of the robot.

16. The method of claim 15 , further comprising communicatively connecting the one or more actuators and the one or more motors over a communication backbone.

17. The method of claim 16 , further comprising communicatively connecting the communication backbone to a processor.

18. 20. The method of claim 17, further comprising powering the communications backbone with the one or more batteries.

19. 20. The method of claim 19, further comprising using the processor to control the one or more motors and receive information from the sensors on the actuators via the communications backbone.

20. 14. The method of claim 13, wherein one or more of the first type, the second type, the third type, the fourth type, the fifth type, and the sixth type actuators comprise rotary actuators, and one or more of the first type, the second type, the third type, the fourth type, the fifth type, and the sixth type actuators comprise linear actuators.