METHOD AND SYSTEM FOR PROGRAMMING INDUSTRIAL ROBOTIC SYSTEMS COMPOSED OF ONE OR MORE AGENTS AND ONE OR MORE FEEDBACK SYSTEMS
Patent Information
- Application Number
- IT102024000017422
- Authority / Receiving Office
- IT · IT
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2044-07-26
AI Technical Summary
Scheduling and programming of robotic systems is difficult and non-intuitive, requiring experienced staff, limiting accessibility to non-specialist users.
A method and system that utilizes the 'learning from demonstration' paradigm to program complex industrial robotic systems, incorporating agents and feedback systems, allowing users to demonstrate tasks, learn from the data, and enable autonomous execution.
Enables non-specialist users to program and operate robotic systems autonomously, adapting to various contexts and environments, enhancing accessibility and flexibility.
Description
Method and system for systems programming industrial robotics composed of one or more agents and one or more feedback systems ------ On behalf of: Italian Institute of Technology Foundation (72%), University of Pisa (28%) Inventors: Gianluca LENTINI, Antonio BICCHI, Manuel Giuseppe CATALANO, Giorgio GRIOLI, Elisa STEFANINI ------ The present invention relates to a method and a system for programming robotic systems industrialists composed of one or more agents and one or more feedback systems. State of the art Programmable robotic systems are known to perform of tasks. However, their scheduling is difficult and non-intuitive, and therefore only possible experienced staff. Purpose and object of the invention The purpose of the present invention is to provide a method and system for systems programming industrial robotics composed of one or more agents and one or more feedback systems, which solve problems and overcome the drawbacks of the prior art. The object of the present invention is a method and a system according to the attached claims. Jacobacci & Partners / ANP Detailed description of examples of implementation of the invention List of figures The invention will now be described by way of example. illustrative but not limiting, with particular reference to the drawings of the attached figures, in which: Figure 1 shows a robotic system according to the invention in a warehouse application vertical; Figure 2 shows a hand-eye procedure calibration for a robotic arm; Figure 3 shows an example of an area of interest (ROI) in a system calibration procedure robotic according to the invention. It is specified here that elements of forms of Different realizations can be combined together to provide further embodiments without limits respecting the technical concept of the invention, as the average person skilled in the art intends it without problems as described. This description also refers to the technique known for its implementation, regarding the undescribed detail features, such as example elements of minor importance usually used in the prior art in solutions of the same type. When introducing an element it is always understood that it can be “at least one” or “one or more”. Jacobacci & Partners / ANP When listing a list of items or features in this description are meant to be the found according to the invention “comprises” or alternatively “is composed of” such elements. When listing features in the scope of the same sentence or bullet point, one or more of the individual features may be included in the invention without connection with the others features of the list. Two or more of the parts (elements, devices, systems) described above can be associated freely and considered as kits of parts according to the invention. Embodiments The present invention describes a method for program complex industrial systems. Inspired by to the “learning from demonstration” paradigm (learning by demonstration), this approach elevates industrial systems to a level of intelligence superior, allowing them to assimilate insights implicitly from human demonstrations. These systems, which may include single agents or multiple, incorporate different feedback mechanisms, with the main goal of expanding accessibility of programming to non-specialist users. The term "Agent" refers to any entity physical or virtual whose position and state (set of values of its internal variables) can be controlled and monitored. This includes a wide range Jacobacci & Partners / ANP of devices, such as robotic arms, platforms furniture, grippers, drones, humanoid robots and even common peripherals such as mice and keyboards interfaced with a computer system. “Feedback mechanisms” are processes or mechanisms by which the output of a system is evaluated and used to further regulate or influence the input or behavior of the system itself or of a other system / device. These can take over the form of sensors. Furthermore, devices such as display screens computers and actuators can function as control systems feedback within this framework. The software can be installed on any PC commercially available and managed through various interfaces (smartphone, tablet, screen). The software allows setup, programming, execution and management of a complex robotic system. As for the setup, a robotic system can be composed of robotic arms, final actuators, sensors, mobile bases and related sensors. All these parts must be calibrated and compatible with each other to operate efficiently and safely. Compatibility hardware (HW) is achieved through the first layer of the software architecture (ROS-based), which enables low-level communication between different HW types (robotic arms, mobile bases, actuators final drives, force sensors / torque meters, etc.). As far as programming goes, this key problem is addressed by technology Jacobacci & Partners / ANP present invention through three phases: Demonstration, Learning and, subsequently, Autonomous Execution. Demonstration According to a general embodiment, in the demonstration phase, are recorded by the said calculator a set of agent data from said at least one respective agent sensor and a set of feedback data from said at least one respective feedback sensor, while the system robotic is controlled by a user to perform a first series of tasks. According to one aspect of the present invention, during the Demonstration phase, users actively drive the agent in performing tasks, using various methodologies adapted to the specific robotic system involved. For example, when orchestrating movements of a robotic arm, users can opt for the teleoperation or use teaching methods kinesthetic. Similarly, when you maneuver a mobile base, control can be facilitated via the manipulation of a joystick, while the interactions with a computer interface typically involve the use of conventional input devices like mouse and keyboard. During this interactive phase, a significant volume of data, coming from both feedback mechanisms that are integrated into the agents industrial system. This involves not only the capture of physical movements, but also of the perceptions of Jacobacci & Partners / ANP system. This large dataset provides a rich information repository, forming a solid foundation for further processing and analysis. Each feedback mechanism can be associated to one or more agents depending on the application cases. Furthermore, the versatility of the method facilitates the segmentation of the task demonstration process into manageable subtasks. This segmentation can occur automatically or through intervention of the user using dedicated input mechanisms like buttons or controls. Learning The Learning process consists of the following phases. Each task is represented as a graph whose nodes they are the events and whose arcs are the actions. A graph node encapsulates the features perceptual derived from both feedback mechanisms and from the principal agent relevant to the task in question. Illustrative examples may include a point cloud cluster that outlines the object to be manipulate, obtained by subjecting the image of the RGB-D camera with advanced clustering algorithms such as RANSAC, Euclidean Clustering, or Region Growing. In alternatively, for navigation-oriented tasks, the node can integrate point cloud data coming from a lidar sensor. In scenarios that involving a PC, the node could encapsulate a Jacobacci & Partners / ANP screenshot capturing the relevant monitor's display to the execution of the current task. An edge in the graph denotes the action primitives designated for execution by the agent, facilitating the transition of the perceived state from the one outlined in the parent node to that of the node next child. These action primitives include trajectories, defined by points in a space of the phases which includes positions, and / or orientations and / or feature-related actuation signals of the agent (agent state path). You can have spatial trajectories for a robotic entity, Planar trajectories for a moving base or 2D movements on the screen for a mouse, or trigger events of suckers or grippers, all annotated with markers temporal. These trajectories are paths of states that can be adequately represented through parameters exploited by interpolation algorithms advanced, such as spline, Gaussian mixture models (Gaussian Mixture Models) or dynamic motion primitives (Dynamic Motion Primitives). The salient action-event (examples may include: mouse click on a button, hand grip / release, application of force, stopping the mobile base after a trajectory) outlines the start and end points for the activation of action segments. According to one embodiment, the phase of learning is performed by an expert system on the said computer, in which the expert system learns the carrying out a first series of tasks on the basis Jacobacci & Partners / ANP of the agent dataset and the dataset of feedback, associating, for each time instant (or at predefined time instants) of the said first series of tasks and for each agent, a subset of said feedback data to a subset of said data agent. Such subsets can be extracted between the phase demonstration and data learning phase detected during the latter, on the basis of a criterion of predefined relevance (see below for some examples). The expert system can be an algorithm of artificial intelligence, machine learning or Boolean or other. Autonomous Execution Autonomous Execution: once the execution phase is completed learning, the robot is programmed to replicate the task even in conditions different from those encountered during the demonstration phase. This capability is obtained by comparing current data provided by the robotic system with those stored during the learning phase. Consequently, the robot can adapt and reproduce the task in various contexts, demonstrating an advanced level of flexibility and autonomy. The execution process may include: following stages. When the system starts, the feedback mechanisms are used to capture a snapshot of the data Jacobacci & Partners / ANP sensory related to the operating environment in which the agent is currently busy. The captured photograph is analyzed to recover the perceptual features embedded in the nodes of the learned task. Feature detection salient within a work environment can be made through various methodologies, depending on the nature of the stored perception. For example, when these are point clouds, the techniques can involve setting a threshold for the distance Bhattacharyya's point clouds and the use of particle filters for localization, or use of a Siamese neural network. If the perceptual characteristics corresponding to the initial node of a task are identified with success, the task execution begins. Upon on the contrary, if the perceptual characteristics are not detected, the system stops, inviting the operator to provide further instructions. When a task arc is executed, the agent is controlled to perform the action as directed by the primitive of the arch, generalizing the movement between the parent and child node states to configurations actual departure and arrival of the environment. According to a general embodiment of the invention, in the autonomous execution phase, the expert system generates commands and sends them to said means of movement and / or operation to perform, by of said one or more agents, a second set of tasks based on a comparison, for each instant Jacobacci & Partners / ANP temporal, of a set of current agent data provided by said at least one respective agent sensor and a set of current feedback data provided by said at least one respective feedback sensor with said a subset of said feedback data and a subset of said agent data. As for the system management, the upper level of the software consists of a web application based on a centralized system with a client / server architecture. All clients connected communicate with a server, which manages the robotic system configurations, the interfaces of control for robot movement and connection between hardware components. Operation example - Automation of phases loading of a Vertical Warehouse. The purpose of this application of the invention is to replace the loading and unloading procedures of a vertical warehouse currently carried out by operators humans with a complex robotic system capable of perform the same task independently. The following description is structured as follows: 1) Definition of a vertical warehouse and its operation; 2) Loading phase: Description of the warehouse loading procedure performed by a human operator; Description of the warehouse loading procedure performed by the complex robotic system; and Jacobacci & Partners / ANP Detailed operation of the software and therefore in what do the setup phases consist of, system programming and management. As regards the definition and functioning of a vertical automatic warehouse, it is a system of automatic storage which, unlike a shelf conventional storage, develops vertically upwards. This means that parts, tools and goods of any kind can be stored in the automatic vertical warehouse in a particularly efficient and space-saving. Automatic vertical warehouse systems are controlled and managed by a specific software, which allows for optimal inventory control. In the loading phase, as normally performed by a human operator, the following sub-phases can be distinguished: 1) Preparing the Articles: Before starting the loading phase, operators must prepare the items that will be placed in the warehouse. This could include breaking down items into batches, labelling and verification of correspondence between the goods and the information in the system of inventory management. 2) Access to the Warehouse: the operator accesses the area of loading of the vertical warehouse, where the entrance is located to insert articles. 3) Identification of the compartments: the operators, via inventory management systems or supporting devices, such as mobile terminals or labels, they know the compartment code Jacobacci & Partners / ANP inside the vertical warehouse where they will have to be the articles are placed. 4) Call to the Warehouse: the operators approach the warehouse management PC and enter the code of the compartment in the warehouse management software. After that they wait for the trolley to be placed the items leave the warehouse. 5) Article Placement: The Operators manually place items into the compartment. 6) Warehouse Shipping: After placing all the items in the right compartment, the operators via warehouse management software they send the compartment back into the warehouse vertical. In case the loading phase is carried out with a complex robotic system, such system includes as agents: a mobile base; a robotic arm; a gripper; an automatic warehouse PC + mouse + keyboard; and as feedback mechanisms: a camera; one or more lidars; a warehouse PC monitor. Associating feedback mechanisms with agents it is as follows: the camera is associated with the robotic arm, the gripper and to the PC; Jacobacci & Partners / ANP the lidars are associated with the mobile base; the PC warning to the warehouse PC. Agents have respective agent sensors that provide the following data: Robotic Arm: Reading the current 3D pose of the robot Gripper: Reading the drive signal Mobile Base: Reading the current 2D pose of the robot Warehouse PC (keyboard and mouse): Coordinate reading mouse pixels + reading mouse buttons + reading keyboard buttons. The feedback means have respective sensors feedback that provides the following data: Camera: Point cloud 3D; Lidars: 2D Point Cloud; Warehouse PC Monitor: Screenshot. The sequence of operations described below represents the functioning of the robotic system a once the programming phase is completed. This The process is illustrated below: 1) Article preparation: before starting the phase loading, the items are placed on a tray on a mobile base positioned near a specially designed loading station. 2) Access to the Warehouse: the mobile base navigates independently to position themselves in front of the warehouse, where a robotic arm will be ready to act. Jacobacci & Partners / ANP 3) Call to warehouse: the PC issues a request to the warehouse for the arrival of a dedicated drawer to the allocation of objects. 4) Positioning of items: using a gripper, the robotic arm picks up the objects one at a time identified objects in the mobile base tray through a camera and arranges them precisely inside the warehouse drawer, continuing until when the items are completely sold out. 5) Warehouse Shipping: After placing all the items in their respective compartments, the PC of the warehouse sends the drawer back to its location original inside the warehouse. 6) Mobile Base Shipping: The mobile base returns to the loading point of the objects, thus completing the cycle operating. In these phases, the system is managed by a web application available on Tablet. The system setup first consists of connection of the robotic arm, the warehouse PC, and from the room via Ethernet to a router. The router it is connected to a mini-PC (nuc) where the software according to the invention. The tablet and the mobile base they are connected via Wi-Fi to the router. As for the software setup, it integrates drivers for communication with the hardware parts. For when it comes to communication with the warehouse PC, a small computer has been installed on the warehouse PC program that allows the recording of inputs mouse and keyboard and translates them into machine language Jacobacci & Partners / ANP used by software that operates through ROS. software according to the invention also has already integrated a part that manages communication with the web application. Robotic system programming management will take place entirely via web application. Before proceed with the actual programming, the operator is required to configure the robotic system. This implies inform the software about the agents present and the connected feedback systems, as well as establishing the associations between feedback systems and agents. In case of the robotic arm, the operator must also specify which gripper is connected to the arm. This operation is managed via the web application on a tablet. The operator accesses the web app on the screen of configuration, where it will proceed to add one for one the agents and the feedback systems needed. Once the configuration phase is completed, it is necessary to proceed with the calibration of the sensors according to the their presence in the system. For example, in the case where there is a vision system and a robotic arm, It is important to calibrate the camera compared to the robotic arm. There are two types of calibration: - HAND EYE CALIBRATION: calibration is crucial for inform the robotic arm about the position of the camera with respect to its reference system. This process involves the use of a marker that It is applied to the robot using a specific tool. Jacobacci & Partners / ANP Next, the camera captures several images of the marker in various positions, then allowing to proceed with the calibration and then find the transformation that exists between the chamber and the base of the robot. Figure 2 shows the various transforms between the different frames of the robot, camera and marker. - ROI CALIBRATION: this calibration is necessary for delimit the working area of the room. It happens through two markers that identify the limits of the FoV region of the camera. In Fig. 3 you can see the measurements provided by the camera, image and point cloud, with the final ROI highlighted by a green square obtained by detecting the two markers in the figure. a) DEMONSTRATION: The operator is now ready to start the demonstration phase of the task that the system will have to perform then perform it independently. Using the application, will have to access the dedicated page and start a new task. In the application, the robot control page, which includes a joystick to manage the mobile base, a "free jog" button for control the robotic arm and buttons to command the gripper. To control the mobile base, the joystick will allow the user to move the base in forward / backward, right / left and rotate clockwise clockwise / counterclockwise. For arm control robotic, a "free jog" button will allow the user to manually move the arm. By activating this button, an arm controller will activate robotic called "gravity compensation", which facilitates the manual movement of the robot without effort, Jacobacci & Partners / ANP because it automatically compensates for the weight of the arm. gripper control, the application puts to arrangement of buttons that regulate the strength and the acceleration of the gripper during opening or closing closing, as well as the suction / release force in the case where the gripper is a suction cup. In a START DEMONSTRATION phase, after that the operator has started recording the task by pressing the “start teaching” button on the screen robot control, the operator controls the base mobile and can move it until it reaches the warehouse vertical where the robotic arm is positioned. Next, the operator approaches the PC of the warehouse and enter the compartment code in where you must deposit the objects. Once the cart has left the warehouse, the operator performs the ROI type calibration of the camera, delimiting the camera's field of view (FOV) to the tray only on the mobile base. At this point, the operator can proceed to demonstrate to the system robotic how to pick up items from the tray mobile base and deposit them inside the warehouse compartment. This process is repeated for each object once. that the operator has activated the "free jog" button in the control screen of the robotic arm to move it manually: Move your arm to reach the object. Activate the gripper closure via the control interface. Jacobacci & Partners / ANP Move your arm to reach the correct warehouse compartment. Activate the gripper opening via the control interface. Move the arm to return to the starting position home. Once you have completed managing all the objects, the operator can approach the warehouse, order the shipping the cart and then moving the mobile base to bring her back to the starting point. DATA RECORDED DURING THE DEMONSTRATION PHASE: During the registration phase, data is recorded the following signals: Streaming of the mobile base position, including x, y coordinates and streaming of the location of the mobile base, including coordinates x, y and the orientation (yaw). Streaming of laser Lidar measurements; Streaming mouse position from PC warehouse of the warehouse, with x and y coordinates on the screen. Streaming of command inputs from the keyboard. Streaming of chamber measurements, including RGB images and a Point Cloud. Streaming of the gripper position. Streaming activation / deactivation commands of the gripper. Jacobacci & Partners / ANP In the learning phase, the data is processed differently depending on whether they relate to a agent or to a feedback system. In the case of agent, the streaming data related to the agent's position are encoded through a series of weights and parameters, following the method of dynamic movement primitives (DMP) for the generalization of movement. More specifically, the use of DMPs is an approach used in the field of robotics for learning and the generation of movements. Key components of the DMP are: 1. Dynamic linear system: A system of linear differential equations that define a basic trajectory to a target point. 2. Nonlinear forcing term: A term additional that modifies the basic trajectory for adapt to more complex movements. This term is typically a Gaussian function of the time step. 3. Time phase: A variable that evolves in the time and control the progression of the movement. The DMP equations can be written as: ̇ 𝜏 𝑦 = 𝑧 ̇ 𝜏 𝑧 = 𝛼 ( 𝛽 ( 𝑔 − 𝑦 ) − 𝑧 )+ 𝑓 ( 𝑥 ) 𝑧 𝑧 Where 𝑦 is the position, 𝑧 is the speed, 𝑔 is the goal (point of arrival), 𝛼 , 𝛽 are parameters that determine the convergence, 𝑧 𝑧 Jacobacci & Partners / ANP 𝜏 is a time constant that scales the duration of the movement, and 𝑓 ( 𝑥 ) is the forcing term that models the non- linear motion. The above weights are the result of learning DMPs. In fact, the weights are the crucial components of the nonlinear "forcing term" 𝑓 ( 𝑥 ), which allows DMPs to generate complex movements and adaptable. The forcing term 𝑓 ( 𝑥 ) is a nonlinear function that adds complexity to the movement generated by the system linear dynamic. This term allows to model movements that cannot be simply described from a linear trajectory towards an arrival point. The forcing term is typically represented as a linear combination of weighted basis functions: O ∑ ( ) 𝜓 𝑥 𝜔 𝑖 𝑖 𝑖 = 1 ( ) 𝑓 𝑥 = 𝑥 O ∑ 𝜓 ( 𝑥 ) 𝑖 = 1 𝑖 Where: 𝜓 ( 𝑥 ) are the basic functions, often chosen as 𝑖 Gaussian functions centered at different points of the domain of x; 𝜔 are the weights associated with the basis functions; 𝑖 x is a phase variable that evolves over time and check the progression of the movement. The weights determine the influence of each function based on the final trajectory. These are the parameters to be learned during the DMP training phase. During the execution phase, in this form implementation, once the action to be taken has been identified Jacobacci & Partners / ANP generalize, the new weights of the DMPs are calculated to adapt the learned action to the current situation. In the case of a feedback system, for each system the feedback measurement is stored relevant under specific rules, in order to reproduce the task independently. This is about extract from the complex of feedback information (but also from those of the agent) the data that are considered relevant for the reproduction of the task, so as to reduce the amount of data to be processed. This extraction can be performed by neural networks for tasks such as classification, clustering, segmentation, voice recognition, or any other model flexible and adaptive computation. For example, the lidar measurement is stored at the beginning and end of each task. In the case of withdrawal of objects with the robotic arm and the camera, are record the image and point cloud of the object grabbed by the robot. In the case of keyboard input, all recorded inputs are considered important. The final outcome of this phase is manifested through the visualization of sub-tasks within the web application, which are reflected in specific folders on the mini PC that contain all the data learned. Thanks to the division into sub-tasks, our The software also allows for offline programming, allowing the operator to modify and add new sub-tasks. For example, during the demonstration, the operator instructed the arm Jacobacci & Partners / ANP robotic to pick up three types of objects from the tray of the mobile base and place them in the trolley warehouse. In the future, it will be possible to demonstrate to the place another type of object and add the sub-task corresponding to the initial task of loading the warehouse. In case of autonomous execution, to start autonomous execution the user must select the task desired and click "play" in the application. at this moment, the system starts with the first one below learned task, the one relating to the mobile base, comparing the current laser measurement with that stored. This allows the system to determine the current position of the mobile base in the environment compared to where it was during the phase of demonstration. Once this has been identified correspondence, the mobile base starts to move for reach the user-specified location during the demonstration. This is made possible, for example, by the generalization of movement based on weights and stored parameters. Once the mobile base has reached its destination, the warehouse computer plays faithfully reproduces mouse movements and keyboard inputs performed during the demonstration phase. When the warehouse compartment opens, the room acquires an image of the mobile base tray, and the software of the invention analyzes the data for Jacobacci & Partners / ANP identify objects previously deposited in the warehouse during the demonstration. Next, the robotic arm goes into action, heading to collect each type of object previously learned during the phase of learning. The objects are then deposited inside the warehouse until stocks run out typologies. Once the loading is complete, the computer of the warehouse re-activates cart shipping, and the mobile base re-compares the laser measurement with the one stored, allowing the system to locate the mobile base and start the next phase of navigation. The above phases are better specified below described. During the demonstration, Task 1 is performed which consists of two sub-tasks: - Sub-Task 1: The operator, through the use of joystick, drives the mobile base, equipped with tray with objects to be loaded, from STATION 1 (Initial positioning of the robot) towards the robotic arm placed in front of the warehouse, in STATION 2. - Sub-task 2: The operator approaches the PC of the warehouse, enters the code via keyboard drawer to call, move the mouse over the button called the chest of drawers and pressed the left mouse button. (This operation activates the warehouse, which after a few seconds presents the chest of drawers in the bay (area Jacobacci & Partners / ANP in which the chosen chest of drawers is presented) by warehouse). Task 2 is also carried out, which consists of three subtasks: - Sub-Task 1: The operator manually drives the robotic arm so that the suction cup rests on the surface of the object A present in the tray movable base and activates the suction cup that sucks up the object. - Sub-task 2: The operator manually drives the robotic arm so that the suction cup is above the correct drawer compartment and operate the suction cup that releases object A in the above compartment. - Sub-task 3: The operator manually drives the robotic arm to position it in a configuration of home. Task 3 is also carried out, which consists of three subtasks: - Sub-Task 1: The operator manually drives the robotic arm so that the suction cup rests on the surface of object B present in the tray movable base and activates the suction cup that sucks up the object. - Sub-task 2: The operator manually drives the robotic arm so that the suction cup is above the correct drawer compartment and operate the suction cup that releases object B in the above compartment. - Sub-task 3: The operator manually drives the robotic arm to position it in a configuration of home. Jacobacci & Partners / ANP Task 4 is also performed, which consists of three subtasks: - Sub-Task 1: The operator manually drives the robotic arm so that the suction cup rests on the surface of the object C present in the tray movable base and activates the suction cup that sucks up the object. - Sub-task 2: The operator manually drives the robotic arm so that the suction cup is above the correct drawer compartment and operate the suction cup that releases the object C in the above compartment. - Sub-task 3: The operator manually drives the robotic arm to position it in a configuration of home. Task 5 is also performed, which consists of two subtasks: - Sub-Task 1: The operator moves the mouse to reach the chest shipping button inside the warehouse and presses the left mouse button mouse. - Sub-task 2: The operator drives the mobile base via a joystick from STATION 2 to a new STATION 3. The data recorded during the demonstration are from a part of the data provided by the agent sensors: Robotic Arm: Reading the current 3D pose of the robot; Gripper: Reading the actuation signal; Mobile Base: Reading the current 2D pose of the robot; Jacobacci & Partners / ANP Warehouse PC (keyboard and mouse): Coordinate reading mouse pixels + reading mouse buttons + reading keyboard buttons. and on the other hand those provided by the feedback sensors: Camera: Point cloud 3D; Lidars: 2D Point Cloud; Warehouse PC Monitor: Screenshot. Next, an extraction phase is performed data for each subtask of each task. In this phase, for each sub-task we are going to extract from the Only data relevant to the execution is recorded of the task. For task 1 and subtask 1, since in this sub-task the only agent to be involved it is the mobile base, only its data and those of the feedback system associated with it are important for execution. The extracted data are: Mobile Base: Reading the current 2D pose of the robot Lidars: 2D Point Cloud. For task 1 and subtask 2, since in this sub-task the only agent to be involved in the sub-task is the PC, only its data and those of the feedback system associated with it are important for execution. The extracted data are: Warehouse PC (keyboard and mouse): Reading mouse pixel coordinates + button readings mouse + keyboard button reading Warehouse PC Monitor: Screenshot Camera: Point cloud 3D. Jacobacci & Partners / ANP For tasks 2, 3 and 4 the subtasks differ only for the grasped object. For subtask 1, since the only ones agents to be involved in the sub-task are the robotic arm and the gripper, only their data and those of the feedback system associated with them are important for execution. The extracted data are: Robotic Arm: Reading the current 3D pose of the robot; Gripper: Reading the actuation signal; Camera: Point cloud 3D. For subtask 2, since the only ones agents to be involved in the sub-task are the robotic arm and the gripper, only their data and those of the feedback system associated with them are important for execution. The extracted data are: Robotic Arm: Reading the current 3D pose of the robot; Gripper: Reading the actuation signal; Camera: Point cloud 3D. For subtask 3, since the only agent the arm involved in the subtask is robotic, only its data and those of the system feedback associated with it are important for the execution. The extracted data are: Robotic Arm: Reading the current 3D pose of the robot Camera: Point cloud 3D. Jacobacci & Partners / ANP Regarding task 5, subtask 1, since that the only agent to be involved in the sub-task it's the PC, only its data and those of the system feedback associated with it are important for the execution. The extracted data are: Warehouse PC (keyboard and mouse): Coordinate reading mouse pixels + reading mouse buttons + reading keyboard buttons Warehouse PC Monitor: Screenshot Camera: Point cloud 3D. Regarding task 5, subtask 2, since that the only agent to be involved in the sub-task it is the mobile base, only its data and those of the feedback system associated with it are important for execution. The extracted data are: Mobile Base: Reading the current 2D pose of the robot Lidars: 2D Point Cloud. In the next learning phase, for each task and each subtask we are going to learn and build the graph starting from the relevant data extracted for the execution of the subtask. From the relevant data as an agent we learn the primitives, from the relevant data We learn perception through feedback. Depending on the type of Feedback System involved in the sub-task, It is possible that the System learns only a subset of the relevant feedback data extracted in the phase of extraction. For example, in task 1 subtask 1, the data learned are: 1, 1 - 𝐴 , primitive 2D trajectory; Jacobacci & Partners / ANP 1, 1 - 𝑃 , 2D Point cloud perception. 𝑖 , 𝑗 In general, 𝐴 indicates the set of actions of the agents involved in the i-th task and j-th 𝑖 , 𝑗 under task, and with 𝑃 we indicate the set of perceptions of feedback systems associated with agents involved in the i-th task and j-th sub-task. In task 1 subtask 2, the data learned are: 1, 2 𝐴 , primitive 2D trajectory; 1, 2 𝑃 , screenshot full tray. When the camera's cloud point is paired with the PC, This is processed to return information regarding the presence or absence of objects in the tray (Boolean true / false, full / empty). The graph of task 1 is represented in Fig. 4. In task 2,3,4 subtask 1, the data learned are: 𝑖 , 1 𝐴 with i=2,3,4, primitive 2D trajectory and primitive trigger signal; 𝑖 , 1 𝑃 with i = 2,3,4, point cloud object A. When it comes to a manipulation subtask in which the robotic arm grasps an object, the learned perception, subset of relevant data of feedback, is the cluster of the object's point cloud closer to the gripper at the end of the sub-task. In task 2,3,4 subtask 2, the data learned are: 𝑖 , 2 𝐴 with i=2,3,4, primitive 2D trajectory and primitive trigger signal; 𝑖 , 2 𝑃 with i=2,3,4, empty, the robot did not grasp objects. Jacobacci & Partners / ANP When it comes to a manipulation subtask in which the robotic arm grasps an object, the learned perception, subset of relevant data of feedback, is the cluster of the object's point cloud closer to the gripper at the end of the sub-task. In task 2,3,4 subtask 3, the data learned are: 𝑖 , 3 𝐴 with i=2,3,4, primitive 3D trajectory; 𝑖 , 3 𝑃 with i = 2,3,4, empty, the robot did not grasp objects. When it comes to a manipulation subtask in which the robotic arm grasps an object, the learned perception, subset of relevant data of feedback, is the cluster of the object's point cloud closer to the gripper at the end of the sub-task. In task 5 subtask 1, the data learned are: 5, 1 𝐴 , primitive 3D trajectory; 5, 1 𝑃 , screenshot, empty tray. In task 5 subtask 2, the data learned are: 5, 2 𝐴 , primitive 3D trajectory; 5, 2 𝑃 , 2D point cloud perception. The graph of task 5 is represented in Fig. 6. The learned tasks are then linked by user or by a specific program. The connection is the graph overall of Fig. 7. In the autonomous execution phase, starting from the first node from the left in the graph, the following steps are performed operations: 1. Check which feedback systems have participated in the perceptions stored in the 1, 1 node 𝑃 (in the lidar example); Jacobacci & Partners / ANP 2. Data acquisition at the current instant of execution of only the feedback systems that are 1, 1 it is verified to participate in node 𝑃 (in the example 2D point cloud); 3. Processing of current data just acquired by feedback systems alone that participate in the first node to have the same type of perception that can be compared with the one stored (in the example point cloud 2D); 4. Comparison of current elaborated perceptions with those stored (in the example method ICP for 2D point cloud comparison). In the case of a single perception stored in the node: the current perception is compared with that stored. The comparison is successful if the two perceptions are similar. In the case of multiple stored perceptions, they are compare current perceptions with those stored according to the type of perception, e.g. point cloud 2D lidar with point cloud 2D lidar, PC screenshot with PC screenshots. The comparison is successful if each the compared perception is similar. In the case of outcome positive of the comparison, the following are carried out operations: 1, 1 6. the primitives of the arc 𝐴 are taken 1, 1 connected to the node 𝑃 and calculate the trajectory and / or the trigger signal of the agents whose data have been stored Jacobacci & Partners / ANP primitives in the arc with respect to the environment current (in the example 2D primitive). 7. Commands are sent and actions are repeated 1, 2 previous ones for the next node 𝑃 . In case of a negative outcome of the comparison, there are no the same perceptions and therefore the system does not perform Nothing. These steps are repeated for the second node of the graph. We then move on to the third END node of the graph. Here we carry out the following steps: 1. Check which feedback systems have participated in the perceptions stored in ALL nodes connected to the end node (PC monitor and room). 2. Data acquisition at the current instant of execution of only the feedback systems that are are verified to participate in ALL nodes connected to the end node (screenshot and point cloud 3D). 3. Processing of current data just acquired by feedback systems alone that participate in ALL connected nodes to have the same type of perception that can compare with the stored ones (screenshot, boolean tray full or empty, and object clusters). 4. Comparison of current elaborated perceptions with those stored in each node. (Boolean comparison, image comparison and compare cluster). Jacobacci & Partners / ANP In this case, perceptions are compared currents acquired with each set of perceptions stored. The comparison is successful if the set of acquired perceptions are similar to a set of perceptions stored in a node. In this example: camera identifies object A, or B or C or, tray is empty and screenshot monitor similar to that memorized. In case of a positive outcome of the comparison, yes perform the following operations: 6. the primitives of the connected arc are taken to the node stored with the most perceptions similar to the current ones and calculation of the trajectory and / or trigger signal of the agents whose primitives have been stored in the arc compared to the current environment. 7. Send the commands and proceed with the node along the branch identified as most similar to the one stored. In case of a negative outcome of the comparison, there are no the same perceptions and therefore the system does not perform Nothing. The program ends when it reaches a node of END without children. In the above the favorites have been described embodiments and some have been suggested variations of the present invention, but it is to be understood that experts in the field will be able to contribute modifications and changes without thereby departing from the Jacobacci & Partners / ANP relative scope of protection, as defined by the claims attached. Jacobacci & Partners / ANP
Claims
1. Computer-based method for the programming of composite industrial robotic systems by one or more agents each equipped with means of movement and / or actuation and of at least one respective agent sensor configured to detect a respective agent state and one or more associated feedback systems to said one or more agents, each feedback system being provided with at least one respective sensor feedback configured to detect a configuration environmental associated with at least one of the said one or more agents, including the execution of the following phases: a demonstration phase, in which a set of data is recorded by the said computer of agent coming from said at least one respective agent sensor and a feedback data set coming from said at least one respective sensor of feedback, while the robotic system is controlled from a user to carry out a first series of assignments; a learning phase carried out by a system expert on said computer, in which the system expert learns the performance of the said first series of tasks based on the said set of agent data and said data set feedback, associating, for each instant in time of the said first series of tasks and for each agent, a subset of said feedback data to a subset of said agent data; Jacobacci & Partners / ANP an autonomous execution phase, in which the system expert generates commands and sends them to said means of movement and / or operation to perform, from part of said one or more agents, a second series of tasks based on a comparison, for each instant in time, of a set of data agent current information provided by said at least one respective agent sensor and a set of current feedback data provided by said at least one respective feedback sensor with said a subset of said feedback data and a subset of said agent data.
2. Method according to claim 1, wherein the one or more agents include a mobile base, an arm robotic and a gripper, and the one or more systems of feedback includes a lidar and a camera.
3. Method according to claim 1 or 2, wherein a representation of the first set of tasks is built as a graph in the learning phase, in which each node of the graph is associated with the subset of said feedback data and the subset of said agent data at a respective instant in time, and each arc of the graph is associated with a trajectory defining a path of states that said means of movement and / or operation must be performed, in where a start event and an end event of the trajectory are associated respectively with two connected nodes, at graph being thus associated with a set of trajectories, Jacobacci & Partners / ANP where said comparison is made by adapting the movement along the trajectory associated with two nodes graph connections to an environmental configuration initial and a current final environmental configuration determined on the basis of the current data set of agent and the current feedback data set.
4. Method according to one of claims 1 to 3, where a data extraction phase is performed between said demonstration phase and said demonstration phase learning, in which, for each agent and each feedback system, the subset of said feedback data and the subset of said data of agent on the basis of relevance criterion default.
5. Method according to one of claims 1 to 4, where the said expert system is a machine learning algorithm learning.
6. Method according to one of claims 1 to 4, where the said expert system is an algorithm based on dynamic movement primitives primitives, DMP.
7. Computer program comprising means to code configured to run on a computer electronic the phases of the method according to one of the claims 1 to 6. Jacobacci & Partners / ANP 8. System for programming robotic systems industrial, including: one or more agents each equipped with means of movement and / or actuation; at least one respective configured agent sensor to detect a respective agent state one or more feedback systems associated with said one or more agents, each feedback system being equipped with at least one respective feedback sensor configured to detect a configuration environmental associated with at least one of the said one or more agents, - a computer on which the program is installed claim 7.
9. System according to claim 8, wherein the one or more agents comprise a mobile base, an arm robotic and a gripper, and the one or more systems of feedback includes a lidar and a camera. Jacobacci & Partners / ANP