Generating robot control plans
Patent Information
- Application Number
- CN202180092435.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-01
- Filing Date
- 2021-10-13
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2041-10-13
AI Technical Summary
以避免机器人组件之间的碰撞同时最小化完成任务的时间的方式手动编程这些移动是困难的,因为6D坐标系中的搜索空间非常大,并且不能在合理的时间量内穷尽地搜索
[0008]可以实现本说明书中描述的主题的特定实施例,以便实现一个或多个以下优点。
Smart Images

Figure CN116829314B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Application No. 17 / 108,761, filed December 1, 2020. The disclosure of the aforementioned application is incorporated herein by reference in its entirety. Technical Field
[0003] This manual relates to robots, and in particular to planning robot movement. Background Technology
[0004] Robot planning refers to the sequencing of the physical movements of robot components to perform a task. For example, an industrial robot that builds cars can be programmed to first pick up car parts and then weld them onto the car's frame. Each of these actions can consist of dozens or hundreds of individual movements via robot motors and actuators.
[0005] Robot planning traditionally requires extensive manual programming to precisely instruct robot components on how to move to accomplish specific tasks. Manual programming is tedious, time-consuming, and error-prone. Furthermore, a plan manually generated for one robot operating environment is often incompatible with others. In this specification, the robot operating environment is the physical environment in which robot components will operate. A robot operating environment has specific physical properties, such as physical dimensions, that impose limitations on how robot components can move within it. Therefore, a manually programmed plan for one robot operating environment may be incompatible with robot operating environments of different physical dimensions.
[0006] Robotic operating environments typically contain more than one robot. For example, a robotic operating environment might have multiple robot components, each simultaneously welding different automotive parts to the car's frame. In these cases, the planning process could involve assigning tasks to specific robot components and planning all movements for each robot component. Manually programming these movements to avoid collisions between robot components while minimizing the time to complete the task is difficult because the search space in a 6D coordinate system is very large and cannot be exhaustively searched within a reasonable amount of time. Summary of the Invention
[0007] This specification generally describes how the system can acquire instruction data, which represents instructions for assembling multiple assembly components, such as instructions for assembling a piece of furniture. The system can then generate robot control plans for one or more robot components based on the instruction data to complete the assembly task in the robot's operating environment. In some embodiments, the system may use data from a manual representing the assembly task (e.g., using images from the manual) to generate the instruction data. In other embodiments, the system may identify instruction data provided by an external system, such as instruction data provided by the manufacturer of the assembly components.
[0008] Specific embodiments of the subject matter described in this specification may be implemented in order to achieve one or more of the following advantages.
[0009] Using the techniques described in this specification, the system can automatically generate robot control plans to complete new assembly tasks for which the system has never previously generated robot control plans. For example, the system can automatically generate a robot control plan to assemble a piece of furniture for which the robot components have never assembled before. In some implementations, the user can provide an image of an instruction manual depicting the furniture as input to a machine learning model that has been trained using instruction manuals for assembling other pieces of furniture. The machine learning model can automatically parse the new instruction manual to generate instruction data that the system can use to generate the robot control plan. Therefore, the user can simply capture an image of the instruction manual for any new piece of furniture, and the system can provide a robot control plan for assembling that new piece of furniture.
[0010] Using the techniques described in this specification, the system can generate robot control plans specifically for “temporary” robot operating environments, i.e., environments where robot components will perform only one or a few tasks and / or where robot components will be disassembled or removed after a short period of time (e.g., a day or a week). For example, the robot operating environment could be in a user’s home, such as in a garage, and robot components could be delivered to the user’s home to perform specific tasks, such as assembling furniture. Therefore, the techniques described in this specification enable robust and reliable assembly of complex objects in a fully automated manner and within temporary robot operating environments.
[0011] As a specific example, a user can purchase an unassembled piece of furniture, such as a table, from a store. The store may also send one or more robot components along with the packaged table assembly parts to the user's home.
[0012] Using the technology described in this manual, users can set up robot components in their homes to create a temporary robotic operating environment for assembling a table, such as in their garage. After the user sets up the robot operating environment, a robot planning system can automatically generate a robot control plan for assembling the table. The robot planning system can then provide this plan to a robot control system, which can instruct the robot components to assemble the table. Once assembled, the user can send the robot components back to the store. In this way, stores can enable customers to automatically assemble purchased furniture in a new environment in a time-saving and cost-effective manner.
[0013] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Further features, aspects, and advantages of the subject matter will become apparent from the specification and the accompanying drawings. Attached Figure Description
[0014] Figure 1 This is a diagram of the example system.
[0015] Figures 2A-2D A sample user interface for capturing instruction manual data is shown.
[0016] Figure 3 This is a flowchart of an example process for generating robot control plans.
[0017] The same reference numerals and names in different figures indicate the same elements. Detailed Implementation
[0018] Figure 1 This is a diagram illustrating an example system 100. System 100 is an example of a system that can implement the techniques described in this specification.
[0019] System 100 includes a robot operating environment 102 and a robot planning system 110. The robot operating environment 102 includes a robot control system 150. The robot planning system includes an assembly instruction system 170, a robot component data storage device (store) 180, and a planner 190. Each of these components can be implemented as a computer program installed on one or more computers in one or more locations, these computers being coupled to each other via any suitable communication network (e.g., an intranet or the Internet, or a combination of networks).
[0020] The robot operating environment 102 includes N robot components 160a-n. The robot control system 150 is configured to control robot components 160a-n. The overall objective of the planner 190 of the robot planning system 110 is to generate a robot control plan 192 that allows the robot control system 150 to perform one or more tasks within the robot operating environment 102. Tasks in the robot control plan 192 may include assembly tasks, whereby the robot components 160a-n manipulate one or more assembly components to assemble a final assembled product. Specifically, the robot control system 150 can execute the robot control plan 192 by issuing commands 152 to the robot components 160a-n to drive the movement of the robot components 160a-n.
[0021] The robot operating environment 102 may also include M sensor components 162a-m. The sensor components 162a-m can be any type of sensor capable of measuring the current state of the robot operating environment 102, such as one or more cameras, one or more lidar sensors, one or more ultrasonic sensors, and / or one or more microphones. In embodiments where multiple sensor components 162a-m are present, the different sensor components 162a-m can be of different types, located at different positions in the robot operating environment 102, and / or configured differently compared to other sensor components 162a-m in the robot operating environment 102.
[0022] Sensor components 162a-m can capture sensor data 154 before and / or during the execution of robot control planning 192, wherein sensor data 154 characterizes the robot operating environment 102. Sensor components 162a-m can send sensor data 154 to robot control system 150. Robot control system 150 can use sensor data 154 to execute robot control planning 192. For example, robot control system 150 can use the sensor data 154 captured by sensor components 162a-m to generate commands 152 issued to robot components 160a-n, such as by generating commands to ensure that robot components 160a-n avoid specific obstacles identified in sensor data 154.
[0023] In some embodiments, the robot control system 150 may also issue commands 152 to the sensor assemblies 162a-m. For example, the robot control system 150 may issue command 152 indicating a specific time at which one or more sensor assemblies 162a-m should capture a specific desired observation of the robot operating environment 102. In some embodiments, the sensor assemblies 162a-m may be movable within the robot operating environment 102. For example, the sensor assemblies 162a-m may be attached to a robot arm that can move the sensor assemblies 162a-m to different locations within the robot operating environment 102 to capture the desired observation. In these embodiments, the robot control system 150 may issue command 152 specifying the orientation and / or position of the sensor assemblies 162a-m for each desired observation.
[0024] The robot control system 150 can be configured, for example, by training one or more machine learning models of the robot control system 150, to perform robot control planning 192 in many different types of robot operating environments 102. For example, the robot control system 150 can be configured to operate robot components 160a-n under many different lighting conditions. As a specific example, the robot control system 150 can use sensor data 154 captured by sensor components 162a-m to operate robot components 160a-n in dimly lit operating environments (e.g., in temporary robot operating environments 102 such as garages or attics without industrial-quality lighting).
[0025] Specifically, the planner 190 is configured to generate robot control plan 192 using i) instruction data 172 and ii) assembly component data 174, both of which are provided by the assembly instruction system 170.
[0026] Instruction data 172 is data representing a sequence of subtasks for an assembly task to be performed by robot components 160a-n. In some embodiments, one or more subtasks in the subtask sequence may be performed in parallel. The data representing each subtask may identify the assembly components involved in the subtask, and one or more operations (e.g., torque insertion operation, vertical insertion operation, rotation operation, etc.) that robot components 160a-n must perform to complete the subtask. Optionally, the data representing each subtask may identify one or more robot components 160a-n that should perform one or more actions in robot control planning 192.
[0027] Assembly component data 174 characterizes one or more assembly components required for the final assembled product of the assembly task of the assembly robot control plan 192. For example, assembly component data 174 may define the dimensions of each assembly component, for example, using CAD or STL files. Assembly component data 174 may also include other characteristics of each assembly component, such as the tensile strength of the material of the assembly component.
[0028] In some implementations, the assembly instruction system 170 obtains instruction data 172 and / or assembly component data 174 from an external system. For example, the assembly instruction system 170 may obtain instruction data 172 and / or assembly component data 174 from an external system of an assembly component manufacturer. That is, the manufacturer of the assembly component (e.g., an assembly component of furniture to be assembled) can provide all the information required to complete the assembly task, such as instructions to assemble the assembly component, the material specifications of the assembly component, and the robot component specifications that identify the robot components required for the assembly.
[0029] In some such implementations, the assembly instruction system 170 is configured to receive an identifier of an assembly task from the user equipment 120 of the system 100, for example, by receiving an identifier of a product ready for assembly and sold to the user. For example, the user equipment 120 may obtain user input identifying the assembly task, such as through voice commands or text input. As another example, the manufacturer may include a visual label, such as a barcode or QR code, on the packaging of the product ready for assembly and sold to the user, which the user can scan and provide to the assembly instruction system 170. The assembly instruction system 170 can then obtain instruction data 172 and / or assembly component data 174 based on the user input, for example, from a database provided by the manufacturer.
[0030] In some other implementations, for example, when information about the assembly task is not available from the manufacturer, the assembly instruction system can use data provided by user equipment 120 to generate instruction data 172 and / or assembly component data 174. User equipment 120 can be any suitable device, such as a mobile phone, tablet computer, laptop computer, or desktop computer.
[0031] User equipment 120 can be configured to provide instruction manual data 122 and raw assembly component data 124 to assembly instruction system 170. Instruction manual data 122 represents an instruction manual used to complete the assembly task, such as an instruction manual provided to the user by the manufacturer of the product to be assembled. Raw assembly component data 124 represents the assembly components for the assembly task.
[0032] User equipment 120 may include a sensor system 130 for capturing instruction manual data 122 and raw assembly component data 124. Sensor system 130 may include one or more sensors configured to capture sensor data from the instruction manual and / or the assembled component. The sensors in sensor system 130 may be any suitable type of sensor, such as a camera, lidar sensor, ultrasonic sensor, or microphone. In some embodiments, sensor system 130 may have multiple sensors, wherein different sensors are of different types and / or configured differently from other sensors in sensor system 130.
[0033] For example, sensor system 130 may include a camera that a user can use to capture images of the instruction manual. As a specific example, user equipment 120 may have an application installed that prompts the user to capture an image of each page in the instruction manual. The following references... Figure 2A , Figure 2B , Figure 2C and Figure 2D An example user interface for such an application is described. Images from the instruction manual can then be included in the instruction manual data 122 and provided to the assembly instruction system 170.
[0034] The user can also use the camera of the sensor system 130 to capture one or more images of each assembly component. These images can be included in the raw assembly component data 124. The user can also identify multiple identical copies of the available assembly component for each assembly component; for example, the user might identify multiple copies of the same type of screw available for the assembly task.
[0035] As another example, sensor system 130 can capture measurements of each assembly component and generate geometric data characterizing the geometry of the assembly components, for example, by generating a stereolithography (STL) file for each assembly component. As a specific example, sensor system 130 can capture LiDAR data for each assembly component and use the LiDAR data to generate geometric data. The geometric data can be included in the original assembly component data 124.
[0036] After receiving instruction manual data 122 and raw assembly component data 124, the assembly instruction system 170 can process instruction manual data 122 to generate instruction data 172, and process raw assembly component data 124 to generate assembly component data 174.
[0037] For example, if instruction manual data 122 includes images of instruction manuals, assembly instruction system 170 can use a machine learning model (e.g., using a convolutional neural network) to process each image to generate instruction data 172.
[0038] In some implementations, the assembly instruction system 170 may use multiple different machine learning models to process images corresponding to different types of pages in the instruction manual.
[0039] For example, instruction manuals typically have one or more "Component Identification" pages at the beginning, listing the assembly components required to complete the assembly task. These pages may list each component provided in the product to be assembled. Typically, the Component Identification pages include illustrations and / or text descriptions of the assembly components. A Component Identification machine learning model can process the images on these pages to identify the assembly components required for the assembly task.
[0040] As another example, the instruction manual may have one or more "subtask" pages, each "subtask" page identifying one or more subtasks of the assembly task. That is, for each subtask, the corresponding subtask page of the instruction manual identifies i) the assembly components required to complete the subtask and ii) the actions that must be performed to complete the subtask. Typically, the subtask page includes illustrations and / or text descriptions of the actions required to complete the subtask. The "subtask" machine learning model can process the images of the subtask pages to generate instruction data 172 for each page corresponding to one or more subtasks of the assembly task identified on the page. As a specific example, if each subtask page of the instruction manual illustrates a subtask of the assembly task, the subtask machine learning model can process the images of each subtask page to generate instruction data 172 characterizing the corresponding subtask.
[0041] In some such implementations, the instruction manual data 122 provided by user equipment 120 to assembly instruction system 170 identifies the page type depicted by each image. For example, each image in the instruction manual data 122 may be assigned to one of a plurality of classes corresponding to the page type of the instruction manual. As a particular example, user equipment 120 may prompt the user for user input identifying the class for each image; for example, user equipment 120 may display a list of multiple classes for each image, allowing the user to select one of the classes to which the image should be assigned.
[0042] In some other implementations, the manual data 122 does not identify the class of each image. For example, the assembly instruction system 170 may use a classification machine learning model to process each image to generate a predicted class for the image, and then provide the image to other appropriate machine learning models based on the predicted class.
[0043] Assembly instruction system 170 can use a component identification machine learning model to process images of the component identification page of a manual to determine a representation of the assembly component for each assembly component. For example, the representation could be a feature vector, matrix, or tensor with representative values of the component. In some embodiments, the representation is an embedding. In this specification, an embedding is an ordered set of numerical values representing inputs in a particular embedding space; for example, an embedding could be a vector of floating-point or other numerical values with a fixed dimension. For example, the embedding for each assembly component could be machine-generated. That is, each embedding could represent a corresponding assembly component in a machine learning embedding space determined through training the component identification machine learning model. In some such embodiments, assembly instruction system 170 can determine a different embedding for each unique assembly component; that is, for each assembly component with multiple copies available (e.g., multiple screws of the same type), assembly instruction system 170 can determine a single embedding.
[0044] For example, the component identification machine learning model may include a first machine learning model that processes each component identification page to detect one or more assembly components depicted on the component identification page. That is, because each component identification page may include depictions of multiple different assembly components, the first machine learning model can detect and isolate the corresponding depiction of each assembly component depicted on the component identification page. For example, the first machine learning model may receive an image of the component identification page as input, and the output of the first machine learning model may be data defining the location of one or more depictions of the corresponding assembly component within the input image. As a specific example, the data may define one or more bounding boxes, where each bounding box encloses a region in the input image depicting the corresponding assembly component. For each detected assembly component, the assembly instruction system 170 may then extract the portion of the input image depicting the assembly component (e.g., by extracting pixels within the defined bounding box) and process that portion of the input image with a second machine learning model to determine a representation of the assembly component.
[0045] In some implementations, the assembly instruction system 170 may obtain predetermined representations of one or more assembly components, for example, from a library 140 of the assembly instruction system 170. For example, the library 140 may include representations of common assembly components used in multiple assembly tasks, such as screws or panels of a specific size and type. For convenience, an assembly component having a predetermined representation stored in the library 140 is referred to as a “predetermined assembly component.”
[0046] As a specific example, for each predetermined assembly component, library 140 may include a text description of the predetermined assembly component. Assembly instruction system 170 can then determine whether the same or similar text description appears on the component identification page. If so, assembly instruction system 170 can determine that the assembly component associated with the text description on the component identification page is the same as the predetermined assembly component.
[0047] As another specific example, the assembly instruction system 170 can use a classification machine learning model to process the depiction of each specific assembly component shown on the component identification page (e.g., by processing a portion of the image identified as a component identification by the first machine learning model described above) to generate a prediction of whether a specific assembly component matches a predetermined assembly component. For example, for each predetermined assembly component, the classification machine learning model can generate a confidence value representing the predicted confidence level of a match between the specific assembly component and the predetermined assembly component. If the confidence value corresponding to the corresponding predetermined assembly component exceeds a threshold, such as 0.9 or 0.95, the assembly instruction system 170 can determine that the specific assembly component is the same as the predetermined assembly component.
[0048] As another specific example, instruction manual data 122 or original assembly component data 124 may include identifiers of one or more predetermined assembly components that are assembly components for the current assembly task. For example, instruction manual data 122 or original assembly component data 124 may include images or identifiers of visual markers, such as barcodes or QR codes, corresponding to predetermined assembly components stored in library 140. The visual markers may be printed in instruction manuals, printed on packaging of products ready for assembly sold to users, or printed directly on physical assembly components; users can then capture images or scans of the visual markers and provide them to assembly instruction system 170.
[0049] For assembly components that do not have a predetermined representation, such as an assembly component that is unique to the current assembly task, the component identification machine learning model can generate a corresponding representation.
[0050] To generate representations of assembled components, the component identification machine learning model can use an image processing model, such as a convolutional neural network, to process one or more illustrations of the assembled components. For example, the component identification machine learning model can use a first sub-network to process an image of the component identification page to identify illustrations of the corresponding assembled components, for example, by generating corresponding bounding boxes around each illustration. The component identification machine learning model can then use a second sub-network to process the corresponding illustration of the assembled component for each assembled component (e.g., by extracting pixels within the corresponding bounding box) to generate a representation of the assembled component. Alternatively or additionally, the component identification machine learning model can use a text processing model, such as a recurrent neural network, to process the text description of each assembled component.
[0051] After determining the representation of each assembly component, the assembly instruction system can use a subtask machine learning model to process i) the representation of the assembly component and ii) the image of the subtask page to generate instruction data 172.
[0052] Subtask pages typically include a list of assembly components required for the corresponding subtask; for example, a portion of the subtask page may show the assembly components required by the subtask. A subtask machine learning model can use a first sub-network to process the image on the subtask page to identify the assembly components required by the subtask.
[0053] For example, the first sub-network can identify a diagram of the required assembly component (e.g., by generating a bounding box around the diagram), and the sub-task machine learning model can compare the identified diagram with representations of all assembly components. The sub-task machine learning model can then determine the assembly component most similar to its representation for each diagram. As a specific example, for each diagram and each representation of a corresponding assembly component, the sub-task machine learning model can, for example, use an autoencoder neural network to process i) the representation and ii) the diagram to generate a similarity score. The sub-task machine learning model can then determine the assembly component with the highest corresponding similarity score for each diagram.
[0054] As another example, the first sub-network can identify the text description of each required assembly component, and the sub-task machine learning model can compare the identified text description with the representations of all assembly components. The sub-task machine learning model can then determine, for each text description, the assembly component whose representation is most similar to that text description. As a specific example, for each text description and each representation of a corresponding assembly component, the sub-task machine learning model can process i) the representation and ii) the text description to generate a similarity score. The sub-task machine learning model can then determine, for each text description, the assembly component with the highest corresponding similarity score.
[0055] After determining the required assembly components for each subtask, the subtask machine learning model can process the image of the corresponding subtask page in the instruction manual to determine one or more actions for that subtask, such as one or more manipulations of the required assembly components. The subtask page may include i) a textual description of one or more actions of the subtask, ii) a diagram of one or more actions of the subtask, or iii) both. The subtask machine learning model can therefore process i) the textual description of one or more actions, ii) the diagram of one or more actions, or iii) both to generate instruction data 172 corresponding to the subtask.
[0056] For example, a subtask machine learning model can use a recurrent neural network (e.g., a long short-term memory (LSTM) neural network) to process the text description of one or more actions of the subtask to generate instruction data 172 for the subtask. As a specific example, the text description could be read as "insert part A into the slot in part B", and the recurrent neural network could generate a network output that represents the subtask in the language of instruction data 172, such as "Action: Insert; Part: A; Target: Slot B".
[0057] As another example, a subtask machine learning model can use a convolutional neural network to process a graph of one or more actions of a subtask to generate instruction data for the subtask.
[0058] In some implementations, for one or more subtasks, the component identification machine learning model can generate representations of updated assembled components assembled during the subtask. That is, the subtask includes instructions to assemble two or more assembled components into a single updated assembled component. As described above, the component identification machine learning model can then generate representations of the updated assembled components. These updated representations can then be used to identify the updated assembled components on subsequent subtask pages.
[0059] In some implementations, one or more machine learning models of the assembly instruction system 170 can be trained using instruction manuals for other assembly tasks. That is, the training system can obtain multiple training examples, each including i) instruction manual data corresponding to the respective instruction manual, and ii) a ground-truth model output that should be generated from the instruction manual data. The training system can then process the training examples using one or more machine learning models to determine, for example, updates to the parameters of the machine learning models using backpropagation.
[0060] For example, for each depiction of an assembly component on the component identification page of the training example, the benchmark true model output may include data representing i) the location and / or size of the depiction (e.g., a bounding box defining the description) and ii) a label identifying the depicted assembly component for the current assembly task. In some implementations, as described above, the benchmark true instruction data may also include data identifying the corresponding predetermined assembly component stored in library 140.
[0061] As another example, the benchmark true model output may include benchmark true instruction data representing instruction data 172, which should be generated in response to the instruction manual data of the training examples. The benchmark true instruction manual data may have the same format as instruction data 172, for example, it may be represented in the same language, as described in more detail below. The training system can then use the benchmark true instruction data to train a subtask machine learning model to process images of subtask pages of the instruction manual to extract visual cues from the subtask pages representing the instructions used to complete the subtask. Visual cues may include, for example, arrows pointing to features, depictions of inserted motion, etc.
[0062] In some implementations, one or more machine learning models of the assembly instruction system 170 can be trained using simulated data representing a simulation of a robot control plan generated based on training instruction data 172 generated by the system 170 during training time. That is, the training system can use one or more machine learning models to process training examples including instruction manual data and / or training raw assembly component data to generate training instruction data 172. The simulation system can then simulate the execution of a robot control plan 192 generated based on the instruction data 172 (e.g., a robot control plan 192 generated by the planner 190 in response to processing the training instruction data 172) and generate simulated data representing the simulated execution. For example, the simulated data may include data representing the outcome of the robot control plan and / or one or more intermediate states of the execution of the robot control plan. The training system can then use the simulated data to determine updates to the parameters of the machine learning models. For example, the training system can use the simulated data to train one or more machine learning models using reinforcement learning. As a particular example, the training system can use the simulated data to determine a metric of the feasibility of executing the robot control plan corresponding to the training instruction data 172. The training system can then use the feasibility metric as a reward signal during reinforcement learning. As another specific example, the training system can use simulated data to determine a measure of the correctness of the final assembled product, which will be assembled according to the robot control plan corresponding to training instruction data 172. The training system can then use the measure of correctness as a reward signal during reinforcement learning.
[0063] In some implementations, the training system can train different machine learning models corresponding to the respective manufacturers of the assembled components. Different manufacturers may have different instruction manual formats and practices; for example, a particular manufacturer may use a unique common visual language in each instruction manual. Therefore, for each of the multiple manufacturers, the training system can use training examples corresponding to that manufacturer to train the corresponding machine learning model to recognize the manufacturer's common visual language in the manufacturer's instruction manual.
[0064] Assembly instruction system 170 can process raw assembly component data 124 to generate assembly component data 174. Assembly component data 174 includes corresponding data representing each assembly component of the assembly task. As described above, subtasks of the assembly task may include assembling two or more assembly components to generate an "updated assembly component," which is a combination product of the two or more assembly components. In some embodiments, assembly component data 174 may include corresponding data representing each updated assembly component generated during a corresponding subtask of the assembly task. In some embodiments, assembly component data 174 includes data representing the final assembled product of the assembly task, i.e., the final updated assembly component including each original assembly component of the assembly task.
[0065] For each assembly component, the assembly component data 174 includes i) data defining the dimensions of each assembly component, for example, using a CAD or STL file, and / or ii) one or more characteristics of each assembly component, such as the tensile strength of the material of the assembly component.
[0066] In some implementations, if the original assembly component data 124 includes images of the assembly components, the assembly instruction system 170 can process the images to generate a corresponding 3D model of each assembly component, such as represented by a CAD or STL file. In some implementations where the assembly component data 174 includes data representing updated assembly components, the assembly instruction system 170 can generate updated 3D models of the assembly components. For example, the assembly instruction system 170 can combine the corresponding 3D models of the assembled assembly components to generate an updated assembly component.
[0067] In some embodiments where the raw assembly component data 124 includes images of assembly components, the raw assembly component data 124 provided by user equipment 120 to assembly instruction system 170 identifies the assembly component depicted in the image for each image. For example, for each image of the raw assembly component data 124, the user of user equipment 120 can provide user input identifying the depicted assembly component. As a particular example, user equipment 120 can prompt the user to provide text input for the name of the assembly component for each image. As another particular example, after assembly instruction system 170 processes instruction manual data 122 to identify each requested assembly component, assembly instruction system 170 can provide user equipment 120 with a list of all requested assembly components. User equipment 120 can then prompt the user to provide the identification of the corresponding assembly component for each image; for example, user equipment 120 can display a list of assembly components for each image, allowing the user to select the depicted assembly component from the list.
[0068] In some other implementations, the raw assembly component data 124 does not include the identifiers of the assembly components depicted in each image. For example, the assembly instruction system 170 may use a classification machine learning model to process each image to generate predicted identifiers for the corresponding assembly components. As a specific example, for each image, the classification machine learning model may process i) the image and ii) the corresponding representation of each assembly component determined using the instruction manual data 122 to generate a corresponding similarity score between the image and each representation, for example, using an autoencoder neural network. The classification machine learning model may then determine the representation with the highest similarity score for each image and determine that the image depicts the assembly component corresponding to that representation.
[0069] After identifying the assembly components depicted in each image, the assembly instruction system can determine the assembly component data 174 corresponding to each assembly component. The assembly component data 174 can be interpreted by the planner 190 and includes the characteristics of each assembly component, which will help the planner 190 generate a robot control plan 192 for manipulating the assembly components.
[0070] In some implementations, the assembly instruction system 170 may obtain, for example, predetermined assembly component data 174 of one or more assembly components from a library 140 of the assembly instruction system 170. For example, the library 140 may include assembly component data 174 of common assembly components used in multiple assembly tasks.
[0071] For assembly components that do not have predetermined assembly component data 174, such as assembly components that are unique to the current assembly task, the assembly instruction system 170 can use an assembly component machine learning model to process one or more images describing the assembly component to generate assembly component data 174 corresponding to the assembly component.
[0072] For example, the assembly component machine learning model may include a convolutional neural network configured to process one or more images of the corresponding assembly component to generate a network output representing assembly component data 174 corresponding to the assembly component.
[0073] Assembly component data 174 for a specific assembly component may include one or more of the following characteristics: the material of the assembly component; material ductility specifications (e.g., the degree to which the assembly component can be bent); the maximum force that can be applied to the assembly component; the texture of the assembly component; the weight of the assembly component; the density of the assembly component; the center of mass of the assembly component; or one or more preferred or required contact points for robot manipulation (i.e., identification of corresponding points on the assembly component at which the assembly component can be contacted, grasped, picked up, etc. by the robot component). As mentioned above, assembly component data 174 may also include updated features of the assembly component as an intermediate product of the assembly task, such as one or more of the features listed above.
[0074] In some implementations, one or more machine learning models of the assembly instruction system 170 can be trained using raw assembly component data 124 corresponding to other assembly tasks. That is, the training system can obtain multiple training examples, each including i) raw assembly component data corresponding to the respective assembly task, and ii) baseline real assembly component data 174 that should be generated from the raw assembly component data. For example, the baseline real assembly component data 174 corresponding to each assembly component of each training example may include the actual value of each feature of the assembly component, such as each of the one or more features mentioned above. The training system can then process the training examples using one or more machine learning models to determine, for example, updates to the parameters of the machine learning models using backpropagation.
[0075] Instruction data 172 can be interpreted by planner 190. In some implementations, instruction data 172 is represented using a common computer language across all assembly tasks of the robot planning system 110. That is, instruction data 172 corresponding to any assembly task can be represented using the same computer language, regardless of, for example, the manufacturer of the assembly components for the assembly task.
[0076] As a specific example, instruction data 172 may include the following:
[0077] Phase 1
[0078] Components: 4 components
[0079] Component type: Wooden panel type C
[0080] Visual marker: small hole
[0081] Components (pieces): 16 components
[0082] Part type: Wooden stud type 2
[0083] Number of subtasks: 16
[0084] Move: Stud 2 into panel C
[0085] Skill type: Lateral insertion
[0086] Force threshold: 0.5 lbs
[0087] Operation order: not important
[0088] Success 1: 4 identical objects
[0089] Success 2: Evenly insert studs less than 1 meter long from the bottom of panel C.
[0090] In this example, the assembly task comprises multiple stages, each stage including one or more sub-tasks; instruction data 172 for the first stage is provided. Furthermore, in this example, the robot planning system 110 represents each assembly component as one of two types: “parts,” which are larger components to be assembled into the final assembled product (e.g., shelves, table legs, wooden panels, etc.), and “parts,” which are smaller components used to assemble the parts together (e.g., nails, screws, studs, etc.).
[0091] In this example, the first phase of this exemplary assembly task comprises 16 subtasks, each of which involves placing a "wooden stud type 2" into a corresponding "wooden panel type C". Specifically, robot components 160a-n insert the wooden studs into visually distinct holes in the wooden panels. Instruction data 172 can identify the "skill type" of the subtask, representing the action that the corresponding robot component 160a-n must be able to perform to complete the subtask in robot control planning 192; in this example, the robot component must be able to perform lateral insertion. Instruction data 172 can identify a "force threshold," which is the upper limit of the amount of force that can be applied to the assembly component of the subtask; in this example, the assembly component can withstand a force of up to 0.5 pounds. Instruction data 172 can identify the required or preferred order of operations for the subtasks of the assembly task; in this example, the subtasks can be completed in any order. Instruction data 172 can identify one or more criteria by which the success of a subtask can be determined. In this example, the first measure of success is whether the wooden panels and studs have been assembled into four identical objects, while the second measure of success is whether the studs have been evenly inserted into specific sections of the wooden panels.
[0092] After receiving instruction data 172, the planner 190 can convert instruction data 172 into a robot control plan 192 that can be executed by the robot control system 150.
[0093] In some implementations, planner 190 can generate robot control plan 192 by performing one or more optimization simulations that identify the most efficient robot movement sequences for successfully completing the assembly task. For example, planner 190 can perform thousands, millions, or billions of such simulations to fine-tune robot control plan 192.
[0094] In some implementations, planner 190 may use a "planner" machine learning model to process instruction data 172 and / or assembly component data 174 in order to generate robot control plan 192.
[0095] The planner machine learning model can be configured by training with training examples, each including corresponding instruction data and / or assembly component data for other assembly tasks. By training on a large amount of training data, the planner machine learning model can learn how different types of assembly components are typically assembled together to perform different types of assembly tasks.
[0096] In some implementations, the training system can train a planner machine learning model by generating a training robot control plan using training examples and simulating the execution of the training robot control plan using a simulation system to determine metrics of the training robot control plan's quality. For example, for each of one or more subtasks completed during the training robot control plan, the simulation system can simulate the manipulation of the assembly components of the subtask in a simulated robot operating environment based on the training robot control plan. The training system can then determine the success of the subtask based on the simulation results, such as whether the assembly components of the subtask were successfully assembled to generate an updated assembly component.
[0097] As a specific example, the simulation system can use a 3D model of the assembled component to simulate the manipulation of the assembled component, such as a 3D model defined by the corresponding CAD file. The training system can then evaluate the updated 3D model of the assembled component generated during the simulation (e.g., defined by a combined CAD file generated using the corresponding CAD file of the assembled component) to determine whether the updated assembled component was correctly assembled during the simulation.
[0098] The training system can then use reinforcement learning from the simulation results to train the planner machine learning model. The training system can use one or more of the following to determine the reinforcement learning reward signal: as described above, a measure of the correctness of the updated assembly components; a measure of the future feasibility of the final assembled product using the updated assembly components (i.e., whether future subtasks of the assembly task can be correctly performed using the updated assembly components); or the time required to complete the subtasks.
[0099] In some implementations, the training system can train different planner machine learning models corresponding to different manufacturers. Different manufacturers may have different conventions regarding how to assemble components. Different manufacturers may also have different standard assembly components or different standard instruction sequences. Therefore, for each of the multiple manufacturers, the training system can use training examples corresponding to that manufacturer to train the corresponding planner machine learning model to identify the typical assembly process of that manufacturer.
[0100] In some implementations, the planner 190 can compose a robot control plan from multiple different composable modules. For example, one or more modules may correspond to a specific assembly component or a class of assembly components. Alternatively or additionally, one or more modules may correspond to a corresponding action or a class of actions, such as modules corresponding to insertion operations, rotation operations, etc. The planner 190 can maintain a library of these modules and access the relevant modules when generating a specific robot control plan 192.
[0101] In some implementations, planner 190 may obtain an initial robot control plan already generated by an external system, as well as an updated initial robot control plan, to generate a final robot control plan 192. In some such implementations, the initial robot control plan partially completes the assembly task, but the final assembled product does not meet one or more requirements, or the initial robot control plan includes one or more errors that hinder the execution of the plan. In some other such implementations, the initial robot control plan successfully completes the assembly task, but in an inefficient or suboptimal manner in some respects. For example, the initial robot control plan may be able to execute successfully in a robot operating environment different from robot operating environment 102, but not in robot operating environment 102 (or can execute, but only in a suboptimal manner). Planner 190 may then refine the initial robot control plan, for example, in the manner described above, to generate a final robot control plan 192 optimized for robot operating environment 102. For example, the initial robot control plan may be provided by the manufacturer of the product to be assembled, and / or may be manually programmed by an engineer.
[0102] In some implementations, planner 190 may receive data characterizing one or more humans performing assembly tasks using robotic components 160a-n. For example, planner 190 may receive sensor data (e.g., image data, video data, LiDAR data, tactile sensing data, force sensing data, or motion sensing data) captured during human assembly of the components to generate the final assembled product. Planner 190 may then process this data to generate robot control plan 192, which simulates the movements demonstrated by humans using robotic components 160a-n. This process is sometimes referred to as "learning by demonstration."
[0103] As a specific example, planner 190 can access a predetermined robot control plan corresponding to a specific assembled product (e.g., a specific pre-assembly rack). The manufacturer of the assembled product can issue an updated model of the assembled product, which is the same product except that a specific assembly component has been replaced (e.g., a specific connector has been redesigned). In this example, planner 190 can obtain the predetermined robot control plan corresponding to the previous model of the assembled product, remove the modules corresponding to the replaced assembly components, and insert the modules corresponding to the new assembly components to generate a final robot control plan 192. Therefore, planner 190 can utilize a composability framework to improve efficiency and avoid regenerating a new robot control plan 192 from scratch for each assembled product.
[0104] In some implementations, to generate robot control plan 192, planner 190 also obtains robot component data 182 from robot component data storage device 180. Robot component data 182 characterizes the capabilities of robot components 160a-n. For example, robot component data 182 may include one or more of the following: design files (e.g., CAD files) for each robot component 160a-n, technical specifications for each robot component 160a-n (e.g., payload capacity, range, speed, accuracy threshold, etc.), robot control simulation (RCS) data (e.g., modeled robot motion trajectories), or APIs for interacting with robot control system 150. APIs may include one or more sensor APIs (e.g., sensors for measuring force, torque, motion, vision, gravity, etc.) or data management interfaces (e.g., product lifecycle (PLC), product lifecycle management (PLM), or manufacturing execution system (MES) APIs). Robot component data 182 may also include skill types for each robot component 160a-n, which identify actions that the robot component can perform.
[0105] After generating robot control plan 192, planner 190 can provide robot control plan 192 to robot control system 150. As described above, robot control system 150 can then execute robot control plan 192 by issuing command 152 to robot components 160a-n to drive the movement of robot components 160a-n.
[0106] In some implementations, planner 190 is an online planner. That is, robot control system 150 can receive robot control plan 192 and begin execution, then provide feedback on the execution to planner 190 during execution. Planner 190 can then generate a new robot control plan in response to the feedback. In some other implementations, planner 190 is an offline planner. That is, planner 190 can provide robot control plan 192 to robot control system 150 before robot control system 150 performs any operation, and planner 190 does not receive any direct feedback from robot control system 150.
[0107] In some implementations, the robot planning system 110 resides within the robot's operating environment; that is, the robot control plan 192 can be generated by the on-site planner 190. In other implementations, the robot planning system 110 is hosted in an off-site data center, which can be a distributed computing system with hundreds or thousands of computers in one or more locations. As a specific example, the robot operating environment 102 can be a temporary robot operating environment provided by the user, for example, in the user's home.
[0108] Figures 2A-2D Example user interfaces 210-290 for capturing instruction manual data are shown. User interfaces 210-290 can be displayed to a user device (e.g., Figure 1 The user device 120 depicted is used to capture instruction manual data for generating robot control plans to perform assembly tasks. The instruction manual data characterizes the instruction manual for the assembly task. For example... Figures 2A-2D The assembly task described is to assemble a piece of furniture that is ready to be assembled.
[0109] Figures 2A-2D The user interface shown is for illustrative purposes only. In some embodiments, one or more illustrated user interfaces are not presented to the user. In some embodiments, one or more additional user interfaces are presented to the user. In some embodiments, a user interface with the same or similar functionality as the illustrated user interface may be presented to the user, but with a different design. For example, prompts may be worded differently, colors may be different, layouts may be different, and the interfaces may be presented to the user in a different order, etc.
[0110] refer to Figure 2A In the first user interface 210, the user device prompts the user to obtain the instruction manual for the assembly task.
[0111] In the second user interface 220, the user device prompts the user to capture an image of the component identification page of the instruction manual, that is, a page listing the assembly components required to complete the assembly task.
[0112] In the third user interface 230, the user device prompts the user to capture an image of a subtask page from the instruction manual, that is, a page describing how to complete the assembly task.
[0113] refer to Figure 2B In the fourth user interface 240, the user device prompts the user to capture an image of the final subtask page of the instruction manual, which shows the product ready for assembly after the assembly task has been completed.
[0114] Captured images of the corresponding pages of the instruction manual can be obtained from the assembly instruction system (e.g., Figure 1 The assembly instruction system 170 described is processed to generate instruction data for generating robot control planning.
[0115] refer to Figure 2C In the fifth user interface 250, the user device prompts the user to capture one or more images of the robot operating environment in which robot control planning will be performed. The robot operating environment can be a temporary robot operating environment. For example... Figure 2C The described operating environment for the robot is a user's home garage. For example... Figure 1 The robot planning system 110 described herein can use images of the robot's operating environment to generate robot control plans.
[0116] In the sixth user interface 260, the user device prompts the user to capture one or more videos of the robot's operating environment. The robot planning system can use this video to generate robot control plans.
[0117] In the seventh user interface 270, the user device prompts the user to identify one or more items in the robot's operating environment (such as those depicted in images captured in the fifth user interface 250 and / or videos captured in the sixth user interface 260). For example, the robot planning system may use one or more machine learning models to identify items in the images and / or videos. The robot planning system may send data representing the identified items to the user device. The user device may prompt the user in the seventh user interface 270 to confirm the predicted classification of the items (e.g., such as...). Figure 2C The described item (identified as a refrigerator) is confirmed to be a refrigerator, or the item is assigned a category. In some implementations, the user equipment may also prompt the user whether the identified item can be removed from the robot's operating environment, for example, to free up more space for robot control planning. The robot planning system can use this information to generate a robot control plan.
[0118] refer to Figure 2DIn the eighth user interface 280, the user equipment prompts the user to select one or more robot components to perform robot control planning. For example, the robot planning system may send data representing multiple different candidate robot components that can be used to perform robot control planning, and the user equipment may provide the user with a list of candidate robot components. Alternatively or additionally, such as Figure 2D As shown, users can search a list of candidate robot components. As a specific example, a user can select a robot component that has already been delivered to their home for use in completing an assembly task, such as a robot component delivered by the user from a store where they purchased the product for assembly. The robot planning system can generate a robot control plan that allows the selected robot component to be used to execute the robot control plan. For example, the robot planning system can identify the capabilities of the selected robot component and, for each instruction in the instruction manual, determine whether the identified capabilities meet the requirements of the instruction.
[0119] In the ninth user interface 290, the user equipment notifies the user that a robot control plan has been generated and provides the user with a list of options. The first option allows the user to request the execution of a simulation of the generated robot control plan, for example, to determine the estimated time required to complete the assembly task using the robot control plan.
[0120] The second option allows users to request that the generated robot control plan be sent to an integrator, such as a third-party integrator, so that the integrator can process the robot control plan to estimate integration costs, service considerations, etc.
[0121] The third option allows the user to select a user device, such as a mobile phone or tablet, that will assist the user when setting up the robot's operating environment in their home. For example, the user device could run an application that uses augmented reality to identify the placement of each assembly component and / or robot component within the robot's operating environment. As a specific example, the user device could capture live video from its camera and display it on the user device's screen, where the locations where specific components in the environment should be placed, as depicted in the video, are highlighted, outlined, or otherwise emphasized.
[0122] Figure 3 This is a flowchart of an example process 300 for generating robot control planning. Process 300 can be implemented by one or more computer programs installed on one or more computers and programmed according to this specification. For example, process 300 can be executed by a robot planning system, such as... Figure 1 The robot planning system 110 is depicted in the diagram. For convenience, process 300 will be described as being executed by a system of one or more computers.
[0123] The system acquires an image depicting an instruction manual for assembling multiple assembly components (step 310). The system may acquire the image from an external system, such as a user device.
[0124] The system uses a machine learning model to process images depicting instruction manuals to generate instruction data (step 320). The instruction data represents a sequence of instructions used to assemble the assembly components. The machine learning model can be trained and configured to process images depicting instruction manuals and generate instruction data that characterizes the sequence of instructions identified in the instruction manual.
[0125] In some implementations, the images may include images depicting a portion of the instruction manual, such as one or more component identification pages of the manual, which identify each assembly component. A machine learning model can use this portion of the instruction manual to generate a representation for each assembly component. The machine learning model can then use the generated representation of each assembly component to identify descriptions of the assembly components in subsequent sections of the instruction manual, such as on one or more subtask pages.
[0126] In some implementations, the machine learning model may correspond to the manufacturer assembling the components. That is, the machine learning model can be trained using training examples corresponding to the manufacturer, for example, using training examples that include other instruction manuals already produced by the manufacturer.
[0127] In some implementations, a common computer language can be used to represent instruction data, which can be used to represent instruction manuals produced by multiple different manufacturers or any manufacturer.
[0128] Optionally, the system obtains an image depicting the assembled components (step 330). The system may obtain the image from an external system, such as a user device.
[0129] Optionally, the system processes an image depicting the assembled components to obtain assembled component data (step 340). The assembled component data may identify one or more attributes of the assembled component for each component. Attributes may include one or more of the following: the material of the assembled component, the weight of the assembled component, the density of the assembled component, the center of mass of the assembled component, the strength of the assembled component, or the flexibility of the assembled component.
[0130] In some implementations, the system processes images depicting one or more assembly components to identify the assembly components. The system can then retrieve the images from a data storage device (e.g., Figure 1 The library 140 described in the text obtains pre-defined assembly component data for one or more assembly components.
[0131] In some implementations, the system uses a second machine learning model to process images depicting one or more assembly components to generate assembly component data of the one or more assembly components. The second machine learning model can be trained and configured to process images depicting assembly components and generate assembly component data characterizing one or more attributes of the assembly components.
[0132] The system processes instruction data and, optionally, assembly component data to generate a robot control plan (step 350). The robot control plan can identify one or more robot components that can execute the robot control plan to assemble the assembly components.
[0133] The system provides a robot control plan to the robot control system for execution (step 360). The robot control system can then execute the robot control plan using one or more robot components. For example, the robot components can execute the robot control plan in a temporary robot operating environment, such as in a user's home.
[0134] The robot functions described in this specification can be implemented using a hardware-agnostic software stack, or, for brevity, only a software stack that is at least partially hardware-agnostic. In other words, the software stack can accept commands generated by the planning process described above as input, without requiring the commands to specifically relate to a particular robot model or component. For example, the software stack can be at least partially implemented using... Figure 1 The robot control system 150 is implemented.
[0135] A software stack can include multiple levels that increase hardware specialization in one direction and software abstraction in another. At the lowest level of the software stack are robot components, including devices that perform low-level actions and sensors that report low-level states. For example, robot components can include a variety of lower-level components, including motors, encoders, cameras, actuators, grippers, specialized sensors, linear or rotary position sensors, and other peripherals. As an example, a motor can receive a command indicating the amount of torque that should be applied. In response to receiving this command, the motor can, for example, use an encoder to report the current position of the robot joints to higher levels of the software stack.
[0136] Each next highest level in the software stack can implement interfaces that support multiple different underlying implementations. Generally, each interface between levels provides status messages from lower to higher levels and commands from higher to lower levels.
[0137] Typically, commands and status messages are generated cyclically during each control cycle; for example, one status message and one command per control cycle. Lower levels of the software stack generally have more stringent real-time requirements than higher levels. For example, at the lowest level of the software stack, a control cycle may have actual real-time requirements. In this specification, real-time means that within a specific control cycle, commands received at a level of the software stack must be executed, and optionally, status messages are provided back to higher levels of the software stack. If this real-time requirement is not met, the robot can be configured to enter a fault state, for example, by freezing all operations.
[0138] At the next highest level, the software stack can include software abstractions of specific components, which will be referred to as motor feedback controllers. A motor feedback controller can be a software abstraction of any suitable low-level component, not just a literal motor. Therefore, the motor feedback controller receives state from lower-level hardware components via an interface and, based on higher-level commands received from higher levels in the stack, sends commands back to lower-level hardware components via the same interface. The motor feedback controller can have any suitable control rules that determine how higher-level commands should be interpreted and translated into lower-level commands. For example, the motor feedback controller can use anything from simple logic rules to more advanced machine learning techniques to translate higher-level commands into lower-level commands. Similarly, the motor feedback controller can use any suitable fault rules to determine when a fault state is reached. For example, if the motor feedback controller receives a higher-level command within a specific part of the control cycle but does not receive a lower-level state, it can cause the robot to enter a fault state where all operations cease.
[0139] At the next highest level, the software stack can include an actuator feedback controller. The actuator feedback controller can include control logic for controlling multiple robot components via corresponding motor feedback controllers for multiple robot components. For example, some robot components, such as articulated arms, can actually be controlled by multiple motors. Therefore, the actuator feedback controller can provide software abstraction for the articulated arm by sending commands to the motor feedback controllers of the multiple motors using its control logic.
[0140] At the next highest level, the software stack can include joint feedback controllers. Joint feedback controllers can represent joints mapped to logical degrees of freedom in the robot. Thus, for example, while a robot's wrist might be controlled by a complex network of actuators, a joint feedback controller can abstract this complexity and expose that degree of freedom as a single joint. Therefore, each joint feedback controller can control an arbitrarily complex network of actuator feedback controllers. As an example, a six-DOF robot could be controlled by six different joint feedback controllers, each controlling an independent network of actual feedback controllers.
[0141] Each level of the software stack can also enforce level-specific constraints. For example, if a specific torque value received by the actuator feedback controller is outside the acceptable range, the actuator feedback controller can modify it to be within the range or enter a fault state.
[0142] To drive the inputs to the joint feedback controller, the software stack can use command vectors that include command parameters for each component at lower levels, such as the positive values, torque, and speed of each motor in the system. To expose the state from the joint feedback controller, the software stack can use state vectors that include state information for each component at lower levels, such as the position, speed, and torque of each motor in the system. In some implementations, the command vectors also include constraint information related to constraints to be enforced by the controller at lower levels.
[0143] At the next highest level, the software stack can include a joint collection controller. The joint collection controller can handle the issuance of commands and state vectors exposed as an abstraction of a set of components. Each component can include, for example, a kinematic model for performing inverse kinematics calculations, constraint information, and joint state vectors and joint command vectors. For example, a single joint collection controller can be used to apply different sets of policies to different subsystems at lower levels. The joint collection controller can effectively decouple the relationship between the physical representation of the motors and the control policies and how these components are associated. Thus, for example, if the robotic arm has a movable base, the joint collection controller can be used to implement one set of constraint policies for the arm's movement and a different set of constraint policies for the movable base's movement.
[0144] At the next highest level, the software stack can include a joint selection controller. The joint selection controller is responsible for dynamically selecting between commands issued from different sources. In other words, the joint selection controller can receive multiple commands during a control cycle and select one of them to execute. This ability to dynamically select from multiple commands during a real-time control cycle allows for a significant increase in the control flexibility of traditional robot control systems.
[0145] At the next highest level, the software stack can include joint position controllers. Joint position controllers can receive target parameters and dynamically calculate the commands required to achieve those parameters. For example, a joint position controller can receive a position target and calculate the setpoints used to achieve that target.
[0146] At the next highest level, the software stack can include a Cartesian position controller and a Cartesian selection controller. The Cartesian position controller receives the input target in Cartesian space and uses an inverse kinematics solver to compute the output in joint position space. Then, before passing the computed result in joint position space to the next lowest level joint position controller in the stack, the Cartesian selection controller can impose constraints on the result computed by the Cartesian position controller. For example, the Cartesian position controller can be given three independent target states in Cartesian coordinates x, y, and z. In some degrees, the target state can be position, while in others it can be the desired velocity.
[0147] Therefore, these functionalities provided by the software stack offer extensive flexibility in control instructions, allowing them to be easily expressed as target states in a way that naturally aligns with the aforementioned high-level planning techniques. In other words, when the planning process uses process definition diagrams to generate specific actions to be taken, it is not necessary to specify actions in the low-level commands for individual robot components. Instead, they can be expressed as high-level objectives accepted by the software stack, which are transformed through various levels until they ultimately become low-level commands. Furthermore, actions generated through the planning process can be specified in Cartesian space in a way that makes them understandable to human operators, making debugging and analysis of schedules easier, faster, and more intuitive. Moreover, actions generated through the planning process do not need to be tightly coupled to any specific robot model or low-level command format. Instead, the same actions generated during the planning process can actually be executed by different robot models, as long as they support the same degrees of freedom and have already implemented the appropriate level of control in the software stack.
[0148] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed in this specification and their structural equivalents), or combinations thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium, for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals (e.g., machine-generated electrical, optical, or electromagnetic signals) generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus.
[0149] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of devices, apparatuses, and machines used for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus may also be or may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.
[0150] A computer program, also referred to or described as a program, software, software application, application, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not need to, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer or on multiple computers located in one location or distributed across multiple locations and interconnected via a data communication network.
[0151] For a computer system to be configured to perform a specific operation or action, it means that the system has software, firmware, hardware, or a combination thereof installed thereon, which, in operation, cause the system to perform the operation or action. For a computer program to be configured to perform a specific operation or action, it means that the program includes instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0152] As used herein, "engine" or "software engine" refers to a software-implemented input / output system that provides outputs different from the inputs. An engine can be a coded functional block, such as a library, platform, software development kit ("SDK"), or object. Each engine can be implemented on any suitable type of computing device, including one or more processors and computer-readable media, such as a server, mobile phone, tablet computer, laptop computer, music player, e-book reader, laptop or desktop computer, PDA, smartphone, or other fixed or portable device. Furthermore, two or more engines can be implemented on the same computing device or on different computing devices.
[0153] The processes and logic flows described in this specification can be executed by one or more programmable computers, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by special-purpose logic circuitry (e.g., FPGA or ASIC), or by a combination of special-purpose logic circuitry and one or more programmable computers.
[0154] A computer suitable for executing computer programs can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented or incorporated therein by special-purpose logic circuitry. Generally, a computer will also include or be operatively coupled to one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, to receive data from or deliver data to, or both, these mass storage devices. However, a computer does not need to have such devices. Furthermore, a computer can be embedded in other devices (e.g., mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, GPS receivers, or portable storage devices (e.g., Universal Serial Bus (USB) flash drives, to name just a few).
[0155] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0156] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse, trackball, or presence-sensitive display or other surface through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's device in response to a request received from a web browser. Additionally, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone), running a messaging application, and then receiving response messages from the user.
[0157] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server; or middleware components, such as an application server; or frontend components, such as a client computer with a graphical user interface, a web browser, or an application that allows users to interact with embodiments of the subject matter described in this specification; or any combination of one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0158] A computing system may include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship arises from computer programs running on their respective computers, and they have a client-server relationship. In some embodiments, the server sends data, such as HTML pages, to a user device, for example, with the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from that user. Data generated on the user device, such as the results of user interactions, may be received at the server from the device.
[0159] In addition to the embodiments described above, the following embodiments are also innovative:
[0160] Example 1 is a method comprising:
[0161] Image data is obtained from the user equipment, which depicts instruction manuals for assembling multiple assembly components;
[0162] Image data is processed using a machine learning model to generate instruction data representing a sequence of instructions for assembling multiple assembly components. The machine learning model has been trained and configured to process images depicting instruction manuals and generate instruction data representing the sequence of instructions identified in the instruction manuals.
[0163] Process instruction data to generate a robot control plan, which will be executed by one or more robot components for assembling multiple assembly components; and
[0164] Provide robot control planning to the robot control system for executing the robot control planning using one or more robot components.
[0165] Example 2 is the method of Example 1, wherein generating the robot control plan includes:
[0166] Obtain second image data depicting multiple assembled components from the user equipment;
[0167] For one or more assembly components, obtain assembly component data that characterizes one or more attributes of the assembly components; and
[0168] Process i) instruction data and ii) assembly component data to generate robot control plans.
[0169] Example 3 is the method of Example 2, wherein obtaining assembly component data for a specific assembly component includes:
[0170] A second machine learning model is used to process second image data depicting a specific assembly component to generate assembly component data for that specific assembly component, wherein the second machine learning model has been trained and configured to process images depicting the assembly component and generate assembly component data characterizing one or more attributes of the assembly component.
[0171] Example 4 is a method of either Example 2 or 3, wherein obtaining assembly component data for a specific assembly component includes:
[0172] Identify specific assembly components in the second image data; and
[0173] Obtain predetermined assembly component data for a specific assembly component from a data storage device.
[0174] Example 5 is a method of any of Examples 2-4, wherein the assembly component data includes data identifying one or more of the following for one or more of a plurality of assembly components:
[0175] Materials for assembling components
[0176] The weight of the assembled components,
[0177] Density of assembled components,
[0178] The center of mass of the assembled components
[0179] The strength of the assembled components,
[0180] The flexibility of the assembled components, or
[0181] One or more preferred or required contact points for assembling components.
[0182] Example 6 is a method of any one of Examples 1-5, wherein multiple assembly components have been manufactured by a specific manufacturer, and wherein a machine learning model has been trained using training examples corresponding to the specific manufacturer.
[0183] Example 7 is the method of Example 6, wherein the instruction data is represented using a computer language that can be used to represent instruction manuals produced by multiple different manufacturers.
[0184] Example 8 is a method of any one of Examples 1-7, wherein one or more robot components perform robot control planning in a temporary robot operating environment.
[0185] Example 9 is a method of any one of Examples 1-8, wherein:
[0186] The image data includes images depicting a portion of the instruction manual, which identify each of the multiple assembly components; and
[0187] The generated instruction data includes generating a representation of the assembly for each assembly component, and using the generated representation to identify the corresponding depiction of the assembly component in one or more other images in the image data corresponding to the corresponding other parts of the instruction manual.
[0188] Example 10 is a system comprising: one or more computers and one or more storage devices storing operable instructions that, when executed by the one or more computers, cause the one or more computers to perform the method of any one of Examples 1 to 9.
[0189] Example 11 is one or more non-transitory computer storage media encoded with a computer program, the program including operable instructions that, when executed by a data processing device, cause the data processing device to perform the method of any one of Examples 1 to 9.
[0190] Although this specification contains numerous details of specific embodiments, these should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features characteristic of particular embodiments of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0191] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all shown operations to be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0192] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions stated in the claims can be performed in a different order and the desired result can still be obtained. As an example, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for generating robot control plans, comprising: Image data including multiple images of an instruction manual is obtained from a user equipment, wherein the instruction manual has one or more sub-task pages for an assembly task, each sub-task page identifying multiple assembly components required to complete the sub-task and actions that must be performed to complete the sub-task; For each image, a machine learning model is used to process the image data to generate instruction data representing a sequence of instructions for assembling multiple assembly components. The machine learning model has been trained and configured to process the image data depicting the instruction manual and generate instruction data characterizing the sequence of instructions identified in the instruction manual. The instruction data represents a sequence of sub-tasks representing an assembly task to be performed by one or more robot components. The machine learning model is trained using multiple training examples, each including instruction manual data corresponding to the instruction manual and a baseline real model output that should be generated from the instruction manual data. The baseline real model output includes data representing the depicted position and / or size and data representing labels identifying the depicted assembly components for the current assembly task. Process instruction data to generate a robot control plan, which will be executed by one or more robot components for assembling multiple assembly components; and Provide robot control planning to the robot control system for execution using one or more robot components. The assembly component data for obtaining the assembly components includes: A second machine learning model is used to process the second image data depicting the assembled components to generate assembly component data, wherein the second machine learning model has been trained and configured to process the images depicting the assembled components and generate assembly component data characterizing one or more attributes of the assembled components. The assembly component data for obtaining the assembly components also includes: Identify the assembled components in the second image data; and Obtain the predetermined assembly component data of the assembly component from the data storage device.
2. The method according to claim 1, wherein, Generative robot control planning includes: Obtain second image data depicting multiple assembled components from the user equipment; For one or more assembly components, obtain assembly component data that characterizes one or more attributes of the assembly components; and Process i) instruction data and ii) assembly component data to generate robot control plans.
3. The method according to claim 2, wherein, The assembly component data includes data identifying one or more of the following for one or more of the multiple assembly components: Materials for assembling components The weight of the assembled components, Density of assembled components, The center of mass of the assembled components The strength of the assembled components, The flexibility of the assembled components, or One or more required contact points for assembling components.
4. The method according to claim 1, wherein, The plurality of assembly components have been manufactured by a specific manufacturer, and the machine learning model has been trained using training examples corresponding to the specific manufacturer.
5. The method according to claim 4, wherein, The instruction data is represented using a computer language that can be used to represent instruction manuals produced by multiple different manufacturers.
6. The method according to claim 1, wherein, The one or more robot components perform robot control planning in a temporary robot operating environment.
7. The method according to claim 1, wherein: The image data includes images depicting a portion of the instruction manual, which identify each of the plurality of assembly components; as well as The generation instruction data includes generating a representation of the assembly for each assembly component, and using the generated representation to identify the corresponding depiction of the assembly component in one or more other images in the image data corresponding to the corresponding other parts of the instruction manual.
8. A system for operating a robot, comprising one or more computers and one or more storage devices storing operable instructions, which, when executed by the one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 7.
9. A non-transitory computer storage medium storing one or more instructions, which, when executed by one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Systems and Methods of using a Hieroglyphic Machine Interface Language for Communication with Auxiliary Robotics in Rapid Fabrication Environments
US20140277679A1