Reinforced learning method for a system for individually handling items
The reinforcement learning method optimizes individual item handling systems in sorting facilities by addressing fly-out, collision, and multiple gripping errors, enhancing processing efficiency and quality while utilizing existing equipment.
Patent Information
- Application Number
- PCT/EP2024/078973
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-10-15
- Publication Date
- 2025-06-26
AI Technical Summary
Existing individual item handling systems in sorting facilities face challenges such as fly-out errors, collisions, and multiple gripping, which affect the quality and efficiency of item processing. These issues are typically addressed independently, making it difficult to optimize overall system performance.
A reinforcement learning method is implemented to optimize the operation of individual item handling systems. This method simulates the system's operation, generates digital images and depth maps, analyzes images to detect items, determines optimal grasping and movement strategies, and adjusts operational parameters based on rewards for successful actions.
The reinforcement learning method enables the system to manage risks of fly-out, collision, and multiple picking, thereby optimizing individual item processing and improving overall throughput and quality of processing without requiring significant changes to existing sorting line equipment.
Smart Images

Figure EP2024078973_26062025_PF_FP_ABST
Abstract
Description
Reinforced learning process for an individual item handling system TECHNICAL FIELD OF THE INVENTION
[0001] The invention relates to a learning method for optimizing the operation of an individual item handling system, for example for picking up and moving these items during a picking operation in a sorting facility. TECHNOLOGICAL BACKGROUND
[0002] The sorting of items such as postal parcels is carried out within a sorting center comprising sorting lines intended to carry out automated sorting of items, each handling a flow of items. The items are handled along the lines by means of collective conveyor systems but require, for example for individual picking or placement, to be picked individually during an operation known as "pick and place", or "pick and place" in English terminology, implemented by means of a robotic manipulator arm, as described for example in document FR 3 019 068 B1. An example of a manipulator arm using suction cups as a device for gripping items is described in document FR 3 031 046 A1. Of course, in order to make sorting lines profitable, they must handle as large a flow of items as possible and the speed of execution of the "pick and place" operation is essential.
[0003] The function of the manipulator arm is to pick up an item arriving at a so-called picking zone, for example forming one end of a conveyor belt of items where they arrive in bulk, move this item to a predetermined location, then release it. This operation must be fast while remaining reliable in order to limit handling errors.
[0004] A first example of handling error is the "fly-out" in English terminology, which consists of an unexpected release of the article while it is being moved at high speed by the arm and can result in an ejection of the article out of the sorting line. The quality of gripping of an article can depend on the nature of its packaging: the geometry (flat or irregular) and the material (flexible or rigid) packaging the article have a great influence on the risk of fly-out, for example when the gripping is carried out by means of suction cups. More specifically, the quality of gripping of a rigid box with a flat and regular face by the suction cups of a manipulator arm is better than the quality of gripping a flexible packaging, so that the acceleration imposed on the box by the manipulator arm without causing a fly-out can be higher than that imposed on the flexible packaging by the same manipulator arm.Thus, it is understood that for a given level of reliability of the manipulation of an article by the manipulator arm, it will be able to impose on a rigid box a stronger acceleration and therefore a higher speed than that which it will be able to impose on a flexible packaging. Optimizing the overall throughput of processing of the articles therefore requires finding a compromise between the quality of the gripping of the articles and the speed of movement of the arm, this compromise having to take into account the nature of the packaging of each article.
[0005] A second example of handling error is the collision between a picked item and an element of its environment, for example another item or a stack of items located on the path followed by the robotic arm to move it from its picking point to its destination point. Collisions can lead to items being dropped or damaged. Avoiding this type of error requires taking into account the positions of each of the items in the picking area as well as the relative positions of the items to each other, and possibly the characteristics of the handling device which could also collide with items.
[0006] A third example of error is that of the simultaneous gripping of more than one item by the manipulator arm, this is called "multiple gripping". It happens that the gripping head of the manipulator arm accidentally grips more than one item. This is the case, for example, when a gripping head uses a plurality of suction cups to grip items, and some of the suction cups attach to a first item and others attach to a second item. Such an event is accidental and preferably avoided because of its consequences, ranging from a sorting error to a jamming of a sorting machine receiving more than one item simultaneously, and sometimes even the breakage of a component of this sorting machine.
[0007] Each of the above-mentioned errors affects the quality of processing, potentially leading to
[0008] We see that the individual handling of articles arriving in bulk at a picking area requires anticipating one by one distinct problems relating to the consideration of equally distinct parameters. Conventionally, the control of the manipulator arm takes these problems into account, but each one is treated independently of the others with a view to obtaining a compromise between the quality and the speed of handling of the articles. Reference may be made to French patent applications FR2301782, FR2305495 and FR2309639. It remains difficult to optimize the independent treatment of each of these problems while optimizing the overall treatment of the flow of articles arriving to be treated.
[0009] Patent document US 2023 / 0040623 A1 relates to a method of reinforced learning in a pick-and-place system.
[0010] Patent document US 2020 / 086485 A1 relates to a method for manipulating breakable and deformable material.
[0011] INFORMATICS (INDIN), IEEE, July 18, 2023 (2023-07-18), pages 1-7, XP034406020, DOI: 10.1109 / INDIN51400.2023.10217964, is about the integration of manipulator robots for industrial automation.
[0012] Patent document US 2022 / 397887 A1 relates to a computer-implemented method of controlling a logistics system.
[0013] The independent handling of each of the problems that may arise during the individual handling of the items can be described as open-loop operation, i.e. the effects of actions taken to address one of the problems are not taken into account in the resolution of the other problems or for the overall objective of approaching an optimal compromise between the speed and quality of individual item handling.
[0014] An object of the invention is a reinforcement learning method for an individual item handling system, intended to teach an individual item handling system to pick an item arriving at a picking area, grasp it and move it to an arrival point.
[0015] To achieve this aim, one aspect of the invention is a method for reinforcement learning of a system for individual handling of articles, the system comprising a unit for collectively moving the articles to a picking area, a manipulator arm configured to individually grasp and move articles located in the picking area, a sensor configured to acquire images representative of the picking area, the method comprising implementing the following steps by means of a computer system: simulation of the system in a state; generation of a digital color image and a depth map of the virtual picking area, as they would be perceived by the sensor in the system for individual handling of articles, and representative of the picking area of the simulated system;analysis by segmentation of the digital color image and by exploiting the depth map so as to detect and isolate simulated articles visible in the generated digital image and calculation of representative characteristics of the simulated articles visible in the generated image; on the basis of the characteristics calculated in the analysis step, determination of an article to be grasped among the simulated articles and of the trajectories of the simulated manipulator arm to grasp this article and bring it to a drop-off zone; simulation of an effect on the state of the simulated system of the manipulation of the article to be grasped by the manipulator arm in response to the determination; establishment of an indicative report (i) of a representative duration for the manipulator arm to grasp the article to be grasped and bring it to the drop-off zone and (ii) of problems possibly encountered during the simulated manipulation; generation of a reward in the form of a number in response to the report;and adjusting operational parameters of the system, including at least operational parameters of the handling arm of the step of determining in response the reward.;
[0016] Once it has been trained, the system is able to manage the risks of fly-out, collision or even multiple picking, in order to optimize the individual processing of items in a flow of items arriving at the picking area.
[0017] Furthermore, the pre-existing equipment of a sorting line may be sufficient to implement the process, thus allowing an improvement in the quality and processing throughput performance of the system.
[0018] According to additional non-limiting characteristics of the first aspect of the invention, considered individually or in any technically feasible combination:
[0019] - the operational parameters of the manipulator arm can be selected from a choice of an item to be grasped from among the visible simulated items and a configuration of a gripping head of the manipulator arm;
[0020] - the method may further comprise a step of determining a strategy for supplying articles to the picking zone of the simulated system by means of the unit for collective movement of the articles, the step of adjusting the operational parameters of the system further comprising an adjustment of operational parameters of the unit (CONV) for collective movement of the articles;
[0021] - the operational parameters of the unit for collective movement of the articles are chosen from among parameters for controlling the starting and stopping of a conveyor, for modulating a speed of conveying the articles, and for sorting the articles;
[0022] - the report may be representative of the time required to carry out the actions of the manipulator arm, the length of the paths of the manipulator arm, the success of the manipulation of an article and any problems encountered: multiple grips, collisions, and / or fly-outs;
[0023] - the problems possibly encountered can be chosen from multiple picking up of articles by the handling arm, fly-out of an article being handled by the handling arm and collision of an article being moved by the manipulator arm;
[0024] - the step of determining the trajectories of the simulated manipulator arm can be carried out by means of a path search algorithm such as the Dijkstra algorithm.
[0025] The invention extends to a method for individual handling of articles by means of the system having followed the learning method according to the invention, comprising the implementation of the following steps by means of a computer system: capturing a digital image and forming a topographical map of the picking area by means of a sensor; capturing a digital image and generating a topographical map of the picking area by means of a sensor; analyzing the digital image and the topographical map so as to detect and isolate articles visible in the digital image, and calculating individual characteristics of each of these articles; defining a next article to be handled on the basis of the individual characteristics of the articles and the operational parameters; and in response to the definition of the next article to be handled, manipulating the next article to be handled by means of the manipulator arm.
[0026] The invention also extends to a system for implementing the method for individual handling of articles according to the invention, comprising a picking zone; a sensor configured to acquire images representative of the picking zone; a first unit for collective movement of articles configured to bring articles to the picking zone; a manipulator arm configured to individually manipulate articles located in the picking zone; and a computer system for implementing a logic application provided for controlling the manipulator arm and the first unit for collective movement of articles. BRIEF DESCRIPTION OF THE FIGURES
[0027] Other characteristics and advantages of the invention will emerge from the detailed description of the invention which follows with reference to the appended figures in which:
[0028] It represents an overall view of an installation in which the method according to the invention can be applied;
[0029] Illustrates a system to which the learning method according to the invention can be applied, integrated into the installation of the;
[0030] Illustrates the logic modules implemented in a manipulation method to which the learning according to the invention applies;
[0031] This is a diagram illustrating an implementation of the learning method according to the invention;
[0032] It schematically represents a gripping head equipped with suction cups; and
[0033] This is a diagram illustrating an implementation of the method for handling articles according to the invention. DETAILED DESCRIPTION OF THE INVENTION
[0034] Method of implementation
[0035] Figures 1 to 4 illustrate an embodiment of a reinforcement learning method for an individual item handling system, applied to the particular case taken as an example of a system integrated into a sorting installation for these items.
[0036] The illustrates an installation FAC for sorting articles, comprising a PickZ picking zone in which IT articles to be sorted brought by a truck Trck are brought in bulk by means of a first unit CONV for collective movement of articles. A robotic manipulator arm MA, constituting an individual article handling unit, is configured to pick up the IT articles one by one and place them on a second unit INS for collective movement of articles. The second unit for collective movement of articles is in the present embodiment an inserter INS configured to insert articles into a sorting machine SORT. This sorting machine may for example comprise a circular conveyor belt CCB receiving the articles from the inserter INS and sorting outlets Out (three outlets shown here) towards which articles are directed according to their final destination.The CONV and INS collective item movement units are systems for collective item movement and can independently consist of one or more conveyors such as conveyor belts or rollers. The installation can, for example, be designed to sort postal parcels and divide them into distribution rounds, each corresponding to one of the Out sorting outlets.
[0037] More specifically, it illustrates a SYS system for individual item handling integrated into the FAC sorting installation, with the first CONV unit for collective item movement CONV with its PickZ picking zone, the second INS unit for collective item movement with its DZ depositing zone, and the MA manipulator arm for individual item handling with its GH gripping head.
[0038] The first CONV unit for collective movement of articles moves the IT articles from a bulk unloading area of the truck Trck to the PickZ picking area in a direction Dir1represented by an arrow in the. The GH gripping head can be equipped with suction cups connected to a pumping system, each suction cup can be controlled individually so as to apply a vacuum to the surface of an article to be grasped in the PickZ picking area. The illustrates a GF gripping head equipped with six SUC suction cups. The number and distribution of suction cups to be used to grasp a particular article are part of the operational control parameters of the robotic arm, which must be controlled to adapt to the configuration (dimensions, orientation, etc.) specific to each article to be moved.
[0039] A CONT unit for controlling the sorting line, and therefore the SYS system for individual handling of articles, comprises a digital camera, such as a camera, as an image sensor CAM1, and a computer system INF comprising an electronic computer such as a processor and a computer memory, and being functionally connected to the sensor CAM1, to the manipulator arm MA and to the CONV unit for collective movement of articles.
[0040] The CAM1 sensor is configured to acquire images each representing a general view of the PickZ sampling area. The first CAM1 sensor can be adapted to the acquisition of an RGB color image and a pair of digital images capable of forming a depth map of said scene according to known photogrammetry principles. The depth map could be obtained by other techniques, such as laser triangulation or the time-of-flight method.
[0041] In the present embodiment, the articles IT brought into the picking zone PickZ by the first collective article movement unit, which are then manipulated by the manipulator arm MA so as to make them pass from a disordered state resulting from the bulk unloading of the truck Trck on the first article movement unit CONV to an ordered state on the second collective article movement unit INS. The manipulator arm MA has the function of individually manipulating these articles by grasping them one by one to deposit them one by one in a deposit zone DZ of the second movement unit INS.
[0042] The disordered state is characterized by an irregular distribution of the articles, possibly with articles piled on top of each other or simply placed next to each other. The ordered state is characterized, for example, by placing the articles one by one on a continuously rotating conveyor belt, in such a way that each of the articles is placed in the middle of the conveyor belt, preferably at a constant pitch P and defined as the distance separating two of the articles placed consecutively on the conveyor belt by the manipulator arm.
[0043] The second unit INS for collectively moving articles is configured to move the articles placed therein to an entrance of the sorting machine SORT, in a direction Dir2 which is defined as going from an area DZ for depositing the articles by the manipulator arm MA on the second unit for moving articles INS to an entrance of the sorting machine SORT. The articles, entrusted in an orderly manner to the second unit for moving articles INS, can be correctly processed by the sorting machine SORT.
[0044] The ambition of the present invention is, in the present embodiment, a method for optimizing the processing of articles by the system illustrated by the, in particular the control of the manipulator arm MA, but also the control of the first article movement unit CONV. This optimization method is based on reinforcement learning of manipulation of articles by the system SYS, applying firstly to the control of the manipulator arm MA.
[0045] The illustration shows three logic blocks ALIM, VIS and PICK included in an APP logic application of the INF computer system, which can be used to describe the control of the SYS system by the CONT control unit. The APP application can be implemented by the INF computer system. The INF computer system also includes a SIM simulation block configured to simulate the SYS system and its operation.
[0046] The first ALIM supply logic block is designed to control the supply of articles to the picking area. The ALIM block controls the first CONV unit for collective movement of articles: starting and stopping a conveyor belt and / or modulating the conveying speed of the articles, for example. This logic block is also responsible, if necessary, for the separation of the articles in order to aerate the bulk and facilitate the picking of the articles. The separation can be implemented by any conventional means equipping the first CONV movement unit.
[0047] The second VIS visualization logic block is designed to analyze the images of the picking area captured by the image sensor CAM1 and a topographic map derived from these images, according to conventional methods. These images include the articles brought from the picking area in response to the control of the first ALIM logic block. The VIS block in particular performs a segmentation of the images so as to detect and isolate each of the visible articles, and calculates the individual characteristics of each of these isolated articles, such as their positions (heights, locations occupied in the plane formed by the picking area), their orientations in an orthogonal reference associated with the picking area (yaw, pitch, roll), and their colors and / or textures, the latter being indicative of the nature of the packaging of the articles (typically, rigid boxes such as cardboard boxes or flexible packaging such as plastic films).Patent applications FR2305495 and FR2301782 detail implementations of such functions.
[0048] The third logic block PICK is designed to determine a strategy for picking the articles identified by the second logic block VIS: choice of the next article to be picked up by the manipulator arm, approach trajectory of the latter and movement trajectory of the picked article from its picking point to its destination point, here the drop-off zone DZ. The picking strategy is determined based on information on the location of the drop-off point DZ, the location of each of the articles, any overlaps between articles, the height of the articles, the configurations of the manipulator arm and the gripping head, possibly the choices of the gripping suction cups to be activated (implementation of the suction for each suction cup independently of the others) or, where appropriate, the choice of a drop-off zone, for example located on one of several inserters.
[0049] The SIM simulation block is based on a three-dimensional digital twin of the SYS system, which integrates a dynamic engine and automation that reproduce, in particular, realistic configurations of bulk items on the CONV conveyor unit and in the PickZ picking area. Each simulated item is applied with a texture extracted from an image of a real item.
[0050] The three logic blocks ALIM, VIS and PICK can first be pre-trained by classical statistical approaches or heuristic methods determining operational parameters of these logic blocks optimizing the system response to a given situation, defined by the configuration of the items to be processed and the configuration of the system itself, the manipulator arm in particular. Only subsequently, these blocks can benefit from reinforcement learning as described below. An advantage of this two-step learning is to reduce the convergence time of the learning algorithm. Alternatively, these logic blocks can be directly trained by means of reinforcement learning in order to define their operational parameters.
[0051] In all cases, the operational parameters of the logic blocks must be determined in such a way that the system as a whole responds in the most appropriate way possible to a given situation, that is to say in such a way as to maximize the processing rate of the articles while minimizing the occurrence of problems such as collisions, fly-outs or multiple takes. The occurrences of multiple takes and fly-outs can for example be determined by checking the number of articles deposited on the inserter INS. The occurrences of collisions can be determined by trajectory calculations, the topography of the environment and the characteristics of the articles and the manipulator arm being known.
[0052] Illustrates a reinforcement learning process for logic blocks.
[0053] Principle
[0054] The principle of reinforcement learning is to train an intelligent agent to perform actions in a simulation so that it is able to perform them in real conditions once the learning is done. Reinforcement learning has two main elements: the agent and the environment.
[0055] The agent is the algorithm that will seek to learn, and which is expected to eventually be capable of controlling the manipulator arm and the first collective movement unit for supplying articles to the picking area. In this embodiment, it consists of the logic application APP, which controls the manipulator arm MA and the first unit CONV for collective movement of articles. It is this application whose operation we seek to optimize by means of reinforcement learning.
[0056] The environment is defined by the individual characteristics of each of these items (nature, positioning) of the PickZ picking area, the characteristics of the handling arm, and those of the conveyor unit. It is in this environment that the agent is located and it is on this environment that he acts.
[0057] Three types of information are to be considered in the space defined by the agent and the environment: actions, states and rewards.
[0058] An action a is an action on the environment carried out by the agent following a decision on its part. This particularly concerns the choice of an item to be handled by the manipulator arm and the way in which the manipulator arm carries out this manipulation (path taken, etc.). It can also concern the control of the conveyor unit.
[0059] A state s is a situation at a time t of the agent in the environment. The state is updated after each action performed by the agent.
[0060] A reward r is a score indicating the quality of an action applied to a given state of the environment, good (for example: article handled adequately, i.e. brought to the drop-off zone DZ) or bad (for example: occurrence of a collision, a multiple grip, a fly-out, etc.). A reward can be a positive number for good actions, and negative for bad actions.
[0061] Reinforcement learning involves teaching the agent to act in such a way as to maximize the sum of rewards generated. Thus, during learning, it will seek to identify the actions that will allow it to maximize the rewards generated in a given situation, i.e., a given state of the environment.
[0062] Q-learning algorithm
[0063] A reinforcement learning algorithm is Q-learning, which uses the concept of quality (or value) of a state-action tuple, i.e. the total reward generated by an action a in a state s. An action is defined by the operational parameters of the element (manipulator arm and its gripping head, the first unit of collective movement of articles) which implements it.
[0064] The formal definition of quality can be defined as follows for Q-learning:
[0065] Eq. 1
[0066] Eq. 1 mathematically defines the Q-value of a state-action tuple, r(s,a) defining the reward of an action a of the agent applied to a state s of the environment, representing a devaluation factor between 0 and 1, and representing the maximum possible reward of action a in a state s'. States s and s' belong to a state space S. action a belongs to an action space A.
[0067] These algorithms "learn" the quality Q(s, a) of each state-action tuple as they explore the space A of actions a for each state s. The Q-values are used to choose which action a is performed for each state s. At the beginning of the learning process, the Q-values are randomly initialized. As the learning process progresses, the learned Q-values will be increasingly exploited.
[0068] Implementation of Q-learning
[0069] An implementation 100 of the Q-learning algorithm using the INF computer system and the APP application is illustrated by the. This implementation is a learning phase of the APP logic application for controlling the SYS system.
[0070] The environment used for learning is, at least initially, a simulation of the SYS system, in order to avoid damaging the real system.
[0071] At a step S10, a SIMUL simulation of a state S SYS of the SYS system is generated by the SIM simulation module, based on operational parameters Param.ALIM for controlling the CONV article movement unit, and therefore the supply of articles to the picking area. These parameters control the starting and stopping of a conveyor belt, the modulation of the article conveying speed, or the separation of articles.
[0072] At a step S20, the application APP generates a digital image IMG.Virt in color and a depth map of the virtual sampling area, as they would be perceived by the sensor CAM1 in the real system. The rendering of the generated image is preferably photorealistic so as to bring the simulation closer to the operating conditions of the real system SYS.
[0073] In a step S30, the VIS logic block analyzes the digital image by segmenting the digital color image and exploiting the depth map so as to detect and isolate each of the visible articles, and calculates the individual characteristics CAR of each of these articles: positions (heights, locations occupied in the plane formed by the sampling area), orientations, colors and / or textures, as indicated above in the case of the analysis of real system images. The analysis is carried out according to operational parameters Param.VIS of the VIS block, predetermined by the practitioner for the implementation of the image analysis.
[0074] At a step S40, the logic block PICK performs a first task, which is to define a strategy STR for handling the articles by defining the next article IT to be handled on the basis of (i) the individual characteristics CAR of the articles calculated at step S30 and (ii) the characteristics of the manipulator arm and, where applicable, its gripping head.
[0075] The Param.PICK operational parameters are control parameters of the MA manipulator arm, and are representative of the next item that has been decided to be picked up and the characteristics of this item: position, orientations, dimensions, or even nature of the packaging.
[0076] At this stage, the determination of the trajectories of the manipulator arm to pick up the next item to be handled and bring it to the drop-off area can be done by reinforcement learning, but, in this embodiment, they are determined by classic pathfinding algorithms, such as the Dijkstra algorithm. This determination is based on the topological data of the scene obtained by means of the CAM1 sensor: RGB color image, depth map and the geometry and dimensions of the manipulator arm.
[0077] On the other hand, taking into account problem situations (multiple grips, nature of the item's packaging, collisions, fly-outs) is done by reinforcement learning.
[0078] Thus, among the Param.PICK operational parameters for controlling the manipulator arm, parameters of a first set of parameters which are dedicated to defining the trajectories of the manipulator arm can be determined by a path finding algorithm, while parameters of a second set of parameters which are dedicated to taking into account problem situations are determined by reinforcement learning.
[0079] At a step S50, the simulation block SIM simulates the application of the manipulation strategy defined at step S40, which corresponds, in terms of Q-learning, to an action a applied by the agent, which is here the manipulator arm, to a current state s of the environment to arrive at a new state s: movement of the manipulator arm to the article that it has been decided to grasp, grasping of this article, then transport and deposit of the article to its deposit point.
[0080] In a step S60, the supply logic block ALIM performs a second task ALIM?, which is to determine a strategy for managing the supply of articles to the picking zone. The supply is managed by means of the movement unit CONV. The supply management may consist of starting the movement unit CONV for a variable duration, determining a conveying speed of the articles by the conveying unit CONV, or applying a bulk article sorting operation by the conveying unit. The determination is made on the basis of the new configuration of the articles simulated by the simulation block SIM, possibly processed in the same way as in steps S20 and S30.The inability of the manipulator arm to pick up an item may also be a signal indicating the need for the CONV conveyor unit to empty the picking area of items, for example by tipping the items at the end of the conveyor belt into a collection bin.
[0081] At a step S70, the application establishes a REP report representing the time required to carry out the actions of the manipulator arm, the length of the paths of the manipulator arm, the success of the manipulation of an article (actually deposited or not in the deposit zone) and any problems encountered: multiple gripping, collision, and / or fly-out.
[0082] At a step S80, the application APP calculates the reward REC for the actions performed at step S50 by the manipulator arm MA on the basis of the report established at step S70. The principle is to generate a reward that is all the higher when the manipulation of the article takes place without incident, that is to say that potential problems such as multiple gripping and collision do not occur, and that the supply is fluid, that is to say allows a high flow rate of manipulated articles (information deriving from the time taken for the actions to be carried out by the manipulator arm for moving each article).
[0083] The following table summarizes one way to generate rewards within the Q-learning algorithm. Rewards are generated in response to the REP report that indicates the overall system performance with respect to item movement and system throughput, without individually evaluating each of the actions (listed in Action Space A leading to this overall performance, and which correspond to the actions carried out by the system to fulfill the functions of steps S40 and S60 (first task and second task), implemented respectively by the PICK logic block and the ALIM logic block. The State Space S column lists the parameters taken into account in determining the actions to be performed (each defined by operational parameters of the mechanisms responsible for implementing them), which belong to the Action Space.TasksState space SEaction space ARewardsChoice of the item to be handledConfiguration of the manipulator arm at the start of the sequenceGrabs each item in the picking areaEmpty the scene (no item grabbing is possible)Configuration of the gripping head+1 for each item placed in the dropping area-1 for each occurrence of a problem (multiple grabs, collision, fly-out)Supply of items to the picking areaConfiguration of items in the picking areaConfiguration of items on the CONV unitHeight of the chosen itemPosition of the dropping areaStartup of the CONV movement unit, yes or noDuration of start-up of the CONV movement unitMovement speedNeed to fill the scene.
[0084] At step S90, in response to the reward, the parameters Param.ALIM and Param.PICK are adjusted; the adjustment of the parameters Param.VIS can also be considered during this step. Steps S10 to S90 are repeated until the parameters converge sufficiently, according to a criterion established by the system operator SYS, or until an expected performance level of the system is reached.
[0085] Once the parameters are adjusted, the system is ready to be used for sorting items in the actual system. The handling method 200 is illustrated by the.
[0086] In step S210, in the SYS system, IT articles are brought to the picking area by the article movement unit CONV, based on the adjusted Param.ALIM parameters.
[0087] In a step S220, the image sensor CAM1 captures a digital image IMG.Real of the picking area and a topographic map of this picking area is generated, including the articles which were brought to this area in step S210, possibly by means of a pair of digital images, this image and this map are stored in a memory of the computer system INF.
[0088] At a step S230, the logic block VIS analyzes the digital image IMG.Real and the topographic map so as to detect and isolate each of the visible articles, and calculates the individual characteristics CAR of each of these articles, in a similar manner to step S30 of the learning phase with the difference that real data and not virtual or simulation data are analyzed this time.
[0089] In step S240, the PICK logic block defines a strategy STR for handling the articles by defining the next article IT to be handled on the basis of the individual characteristics CAR of the articles calculated in step S230 and the operational parameters Param.PICK.
[0090] At a step S250, the strategy defined at step S240 is implemented: the manipulator arm will grasp the article that it was decided to grasp and moves it to its destination on the inserter INS.
[0091] In a step S260, the logic block ALIM determines a strategy for managing the supply of articles to the picking area, possibly on the basis of new data describing the picking area obtained by means of the sensor.
[0092] The sequence of steps S210 to S260 is repeated cyclically as long as there are articles to be processed. Steps S230, S240, and S260 of implementing the real system are similar to steps S30, S40, and S60 of the learning phase, the only difference being the origin of the data used: simulations (image, topographic map) for learning 100 and data representative of a real situation for implementing the method 200 of handling articles.
[0093] The figures in this document are not necessarily to scale. Some features and components may be shown exaggerated in relation to other components or in a somewhat schematic form, and some details of conventional items may not be shown in the interest of clarity and conciseness.
[0094] Of course, the invention is not limited to the method of implementation described and variant embodiments can be made without departing from the scope of the invention as defined by the claims.
Claims
Method (100) for reinforcement learning of a system (SYS) for individual handling of articles, the system (SYS) comprising a unit (CONV) for collectively moving the articles (IT) to a picking zone (PickZ), a manipulator arm (MA) configured to individually grasp and move articles (IT) located in the picking zone (PickZ), a sensor (CAM1) configured to acquire images representative of the picking zone (PickZ), the method comprising the implementation of the following steps by means of a computer system (INF): - simulation (S10) of the system (SYS) in a state (S SYS) ;- generation (S20) of a digital image (IMG.Virt) in color and a depth map of the virtual picking area, as they would be perceived by the sensor (CAM1) in the individual item handling system (SYS), and representative of the picking area of the simulated system;- analysis (S30) by segmentation of the digital image in color and by exploiting the depth map so as to detect and isolate simulated items visible in the generated digital image and calculation of characteristics (CAR) representative of the simulated items visible in the generated image;- on the basis of the characteristics calculated in the analysis step (S30), determination (S40) of an item to be picked up from among the simulated items and the trajectories of the simulated manipulator arm to pick up this item and bring it to a drop-off area;- simulation (S50) of an effect on the state of the simulated system of the manipulation of the article to be grasped by the manipulator arm in response to the determination (S40);- establishment (S70) of a report (REP) indicative of (i) a representative duration for the manipulator arm to grasp the article to be grasped and bring it to the deposit zone and (ii) of problems possibly encountered during the simulated manipulation;- generation (S80) of a reward (REC) in the form of a number in response to the report (REP); and- adjustment (S90) of operational parameters of the system, comprising at least operational parameters (Param.PICK) of the handling arm of the step of determining in response the reward (REC).; The method according to claim 1, wherein the operational parameters (Param.PICK) of the manipulator arm (MA) are chosen from a choice of an article to be grasped among the visible simulated articles (IT) and a configuration of a gripping head (GH) of the manipulator arm (MA). The method according to claim 1 or 2, further comprising a step (S60) of determining a strategy for supplying articles to the picking area of the simulated system by means of the unit (CONV) for collective movement of the articles, the step (S90) of adjusting the operational parameters of the system further comprising an adjustment of operational parameters (Param.ALIM) of the unit (CONV) for collective movement of the articles. The method according to claim 3, in which the operational parameters (Param.ALIM) of the unit (CONV) for collective movement of the articles are chosen from parameters for controlling the starting and stopping of a conveyor, for modulating a conveying speed of the articles, and for sorting the articles. The method according to any one of claims 1 to 4, wherein the report (REP) is representative of a time required to carry out the actions of the manipulator arm, a length of paths of the manipulator arm, a success of the manipulation of an article and any problems encountered: multiple grip, collision, and / or fly-out.
6. The method according to any one of the preceding claims, in which the problems possibly encountered are chosen from a multiple grip of articles by the handling arm, a fly-out of an article being handled by the handling arm and a collision of an article being moved by the manipulator arm.
7. The method according to any one of claims 1 to 6, wherein the step (S40) of determining the trajectories of the simulated manipulator arm is carried out by means of a path search algorithm such as the Dijkstra algorithm. Method (200) for individual handling of articles by means of the system (SYS) having followed the learning method according to any one of claims 1 to 7, comprising the implementation of the following steps by means of a computer system (INF): - capturing (S220) a digital image (IMG.Real) and generating a topographic map of the picking area (PickZ) by means of a sensor (CAM1); - analyzing (S230) the digital image (IMG.Real) and the topographic map so as to detect and isolate articles visible in the digital image (IMG.Real), and calculating individual characteristics (CAR) of each of these articles; - defining (S240) a next article (IT) to be handled on the basis of the individual characteristics (CAR) of the articles and the operational parameters (Param.PICK); and- in response to the definition (S240) of the next item to be handled, manipulation (250) of the next item to be handled by means of the manipulator arm (MA). System (SYS) for implementing the method (200) for individual handling of articles according to claim 8, comprising:- a picking zone (PickZ);- a sensor (CAM1) configured to acquire images representative of the picking zone (PickZ);- a first unit (CONV) for collective movement of articles configured to bring articles to the picking zone (PickZ);- a manipulator arm (MA) configured to individually manipulate articles (IT) located in the picking zone; and- a computer system (INF) for implementing a logic application (APP) provided for controlling the manipulator arm and the first unit (CONV) for collective movement of articles.
Citation Information
Patent Citations
THERMAL DISTRIBUTION STATION
FR2301782A1
LUBRICATING COMPOUNDS THAT IMPROVE THE SWELLING OF SEALS
FR2305495A1
Lather tanning, soaking, liming,dyeing,processes - requiring shorter vat times by introducing gas at pressure above vapour pressure
FR2309639A1
Device for supplying flat items for a bucket conveyor
FR3019068A1
Installation for the Separation and Individualization of Heterogeneous Postal Items
FR3031046A1