Reinforced learning process for an individual item handling system
The reinforcement learning method optimizes individual article handling systems by simulating operations, analyzing images, and adjusting parameters to minimize handling errors and maximize throughput, effectively addressing the challenges of gripping quality and movement speed in sorting installations.
Patent Information
- Application Number
- FR2023014888
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-27
AI Technical Summary
Existing individual article handling systems in sorting installations face challenges in optimizing the compromise between gripping quality and movement speed, particularly due to variations in article packaging and the risk of handling errors such as fly-out, collision, and multiple gripping.
A reinforcement learning method is employed to optimize the operation of individual article handling systems. This method involves simulating the system's operation, analyzing digital images of the picking area, determining optimal grasping and movement strategies, and adjusting operational parameters to minimize handling errors and maximize throughput.
The reinforcement learning method enables the system to effectively manage risks associated with fly-out, collision, and multiple gripping, thereby optimizing the individual processing of articles and improving the overall throughput and quality of processing in sorting installations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Reinforced learning method for an individual article handling system TECHNICAL FIELD OF THE INVENTION
[0001] The invention relates to a learning method intended to optimize the operation of an individual article handling system, for example for the gripping and movement of these articles during a picking operation in a sorting installation. TECHNOLOGICAL BACKGROUND
[0002] The sorting of articles such as postal parcels is carried out within a sorting center comprising sorting lines intended to carry out automated sorting of articles, each processing a flow of articles. The articles are handled along the lines by means of collective conveyor systems but require, for example for individual picking or placement, to be picked individually during an operation called "pick and place", or "pick and place" in English terminology, implemented by means of a robotic manipulator arm, as described for example in document FR 3 019 068 Bl. An example of a manipulator arm using suction cups as a device for gripping articles is described in document FR 3 031 046 AL Of course, in order to make the sorting lines profitable, they must process as large a flow of articles as possible and the speed of execution of the "pick and place" operation is essential.
[0003] The function of the manipulator arm is to grasp an article arriving at a zone called a picking zone, forming for example one end of a conveyor belt of articles where they arrive in bulk, move this article to a predetermined location, then release it. This operation must be rapid while remaining reliable in order to limit handling errors.
[0004] A first example of handling error is the "fly-out" in English terminology, which consists of an unexpected release of the article while it is being moved at high speed by the arm and can result in an ejection of the article from the sorting line. The quality of gripping of an article can depend on the nature of its packaging: the geometry (flat or irregular) and the material (flexible or rigid) packaging the article have a great influence on the risk of fly-out, for example when the gripping is carried out by means of suction cups. More specifically, the quality of gripping of a rigid box having a flat and regular face by the suction cups of a handling arm is better than the quality of gripping flexible packaging, so that the acceleration imposed on the box by the handling arm without causing fly- out may be higher than that imposed on the flexible packaging by the same manipulator arm. Thus, it is understood that for a given level of reliability of the manipulation of an article by the manipulator arm, it will be able to impose on a rigid box a stronger acceleration and therefore a higher speed than that which it will be able to impose on a flexible packaging. Optimizing the overall throughput of processing of the articles therefore requires finding a compromise between the quality of the gripping of the articles and the speed of movement of the arm, this compromise having to take into account the nature of the packaging of each article
[0005] A second example of handling error is the collision between a picked item and an element of its environment, for example another item or a stack of items located on the path followed by the robotic arm to move it from its picking point to its destination point. Collisions can lead to items being dropped or damaged. Avoiding this type of error requires taking into account the positions of each of the items in the picking area as well as the relative positions of the items to each other, and possibly the characteristics of the handling device which could also collide with items.
[0006] A third example of error is that of the simultaneous gripping of more than one article by the manipulator arm, this is called "multiple gripping". It happens in fact that the gripping head of the manipulator arm accidentally grips more than one article. This is for example the case when a gripping head uses a plurality of suction cups to grip articles, and some of the suction cups attach to a first article and others attach to a second article. Such an event is accidental and preferably avoided because of its consequences, ranging from a sorting error to a jamming of a sorting machine receiving more than one article simultaneously, and sometimes even the breakage of an element of this sorting machine.
[0007] Each of the above-mentioned errors affects the quality of processing, potentially leading to
[0008] We see that the individual handling of articles arriving in bulk at a picking area requires anticipating one by one distinct problems relating to the consideration of equally distinct parameters. Conventionally, the control of the manipulator arm takes these problems into account, but each one is treated independently of the others with a view to obtaining a compromise between the quality and the speed of handling of the articles. Reference may be made to French patent applications FR2301782, FR2305495 and FR2309639. It remains difficult to optimize the independent treatment of each of these problems while optimizing the overall treatment of the flow of articles arriving to be treated. Statement of the invention
[0009] The independent handling of each of the problems that may arise during the individual handling of the articles can be described as open-loop operation, i.e. the effects of actions taken to address one of the problems are not taken into account in solving the other problems or for the overall objective of approaching an optimal compromise between the speed and quality of the individual handling of the articles.
[0010] An object of the invention is a reinforcement learning method for an individual article handling system, intended to teach an individual article handling system to pick an article arriving at a picking area, grasp it and move it to an arrival point.
[0011] In order to achieve this aim, one aspect of the invention is a method for reinforcement learning of a system for individual handling of articles, the system comprising a unit for collectively moving the articles to a picking area, a manipulator arm configured to individually grasp and move articles located in the picking area, a sensor configured to acquire images representative of the picking area, the method comprising the implementation of the following steps by means of a computer system: simulation of the system in state; generation of digital images representative of the picking area of the simulated system; analysis of at least one of the generated images and calculation of representative characteristics of simulated articles visible in the generated image;based on the characteristics calculated in the analysis step, determining an item to be grasped from among the simulated items and the trajectories of the simulated manipulator arm to grasp this item and bring it to a drop-off zone; simulating an effect on the state of the simulated system of the manipulation of the item to be grasped by the manipulator arm in response to the determination; establishing an indicative report (i) of a representative duration for the manipulator arm to grasp the item to be grasped and bring it to the drop-off zone and (ii) of problems possibly encountered during the simulated manipulation; generating a reward in the form of a number in response to the report; and adjusting operational parameters of the system, comprising at least operational parameters of the handling arm of the determination step in response to the reward.
[0012] At the end of its learning, the system is capable of managing the risks of fly-out, collision or even multiple picking, in order to optimize the individual processing of articles in a flow of articles arriving at the picking zone.
[0013] Furthermore, the pre-existing equipment of a sorting line may be sufficient to implement the method, thus allowing an improvement in the performance in quality and processing throughput of the system.
[0014] According to additional non-limiting characteristics of the first aspect of the invention, considered individually or in any technically feasible combination:
[0015] - the operational parameters of the manipulator arm can be chosen from a selecting an item to be grasped from among the visible simulated items and a configuration of a gripping head of the manipulator arm;
[0016] - the method may further comprise a step of determining a strategy supplying articles to the picking area of the simulated system by means of the collective article movement unit, the step of adjusting the operational parameters of the system further comprising an adjustment of operational parameters of the collective article movement unit (CONV);
[0017] - the operational parameters of the collective item movement unit are chosen from parameters for controlling the starting and stopping of a conveyor, modulating a conveying speed of the articles, and sorting the articles;
[0018] - the report may be representative of a time required to carry out the actions of the manipulator arm, a length of paths of the manipulator arm, a successful manipulation of an item and any problems encountered: multiple grip, collision, and / or fly-out;
[0019] - the problems possibly encountered can be chosen from a socket multiple items by the handling arm, a fly-out of an item being handled by the handling arm and a collision of an item being moved by the manipulator arm;
[0020] - the step of determining the trajectories of the simulated manipulator arm can be performed using a pathfinding algorithm such as Dijkstra's algorithm.
[0021] The invention extends to a method for individual handling of articles by means of the system having followed the learning method according to the invention, comprising the implementation of the following steps by means of a computer system: capturing a digital image and forming a topographical map of the sampling area by means of a sensor; capturing a digital image and generating a topographical map of the sampling area by means of a sensor; analyzing the digital image and the topographical map so as to detect and isolate articles visible in the digital image, and calculating individual characteristics of each of these articles; defining a next article to be handled on the basis of the individual characteristics of the articles and the operational parameters; and in response to the definition of the next article to be handled, manipulating the next article to be handled by means of the manipulator arm.
[0022] The invention also extends to a system for implementing the method of Individual handling of articles according to the invention, comprising a picking zone; a sensor configured to acquire images representative of the picking zone; a first unit for collective movement of articles configured to bring articles to the picking zone; a manipulator arm configured to individually manipulate articles located in the picking zone; and a computer system for implementing a logic application intended for controlling the manipulator arm and the first unit for collective movement of articles. BRIEF DESCRIPTION OF THE FIGURES
[0023] Other characteristics and advantages of the invention will emerge from the detailed description of the invention which follows with reference to the appended figures in which:
[0024] [Fig.l] [Fig.l] represents an overall view of an installation in which the method according to the invention can be applied;
[0025] [Fig.2] [Fig.2] illustrates a system to which the method can be applied learning according to the invention, integrated into the installation of [Fig.l];
[0026] [Fig.3] [Fig.3] illustrates logic modules implemented in a method of manipulation to which the learning according to the invention applies;
[0027] [Fig.4] [Fig.4] is a diagram illustrating an implementation of the method learning according to the invention;
[0028] [Fig.5] [Fig.5] schematically represents a gripping head provided with suction cups; and
[0029] [Fig.6] [Fig.6] is a diagram illustrating an implementation of the method of ma assembly of articles according to the invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] Embodiment
[0031] Figures 1 to 4 illustrate an embodiment of a reinforcement learning method for an individual item handling system, applied to the particular case taken as an example of a system integrated into a sorting installation for these items.
[0032] [Fig.l] illustrates an article sorting installation FAC, comprising a PickZ picking zone in which IT articles to be sorted brought by a truck Trck are brought in bulk by means of a first CONV unit for collective movement of articles. A robotic manipulator arm MA, constituting an individual article handling unit, is configured to pick up the IT articles one by one and place them on a second INS unit for collective movement of articles. The second unit for collective movement of articles is in the present embodiment an inserter INS configured to insert articles into a sorting machine SORT. This sorting machine may for example comprise a circular conveyor belt CCB receiving the articles from the INS inserter and the Out sorting outlets (three outlets shown here) to which articles are directed according to their final destination. The CONV and INS collective article movement units are systems for collective article movement and can independently consist of one or more conveyors such as conveyor belts or rollers. The installation can, for example, be designed to sort postal parcels and divide them into distribution rounds, each corresponding to one of the Out sorting outlets.
[0033] [Fig.2] illustrates more specifically a system SYS for individual article handling integrated into the FAC sorting installation, with the first unit CONV for collective movement of articles CONV with its picking zone PickZ, the second unit INS for collective movement of articles with its depositing zone DZ, and the manipulator arm MA for individual article handling with its gripping head GH.
[0034] The first CONV unit for collective movement of articles moves the IT articles from a bulk unloading area of the truck Trck to the PickZ picking area in a direction Dhy represented by an arrow in [Fig.2]. The gripping head GH can be provided with suction cups connected to a pumping system, each suction cup can be individually controlled so as to apply a vacuum to the surface of an article to be grasped in the PickZ picking area. [Fig.5] illustrates a gripping head GF provided with six suction cups SUC. The number and distribution of the suction cups to be used to grasp a particular article are part of the operational control parameters of the robotic arm, which must be controlled to adapt to the configuration (dimensions, orientation, etc.) specific to each article to be moved.
[0035] A CONT unit for controlling the sorting line, and therefore the SYS system for individual handling of articles, comprises a digital camera, such as a camera, as an image sensor CAM1, and a computer system INF comprising an electronic computer such as a processor and a computer memory, and being functionally connected to the sensor CAM1, to the manipulator arm MA and to the CONV unit for collective movement of articles.
[0036] The CAM1 sensor is configured to acquire images each representing a general view of the PickZ sampling area. The first CAM1 sensor can be adapted to the acquisition of a color image of the RGB type and a pair of digital images capable of forming a depth map of said scene according to known principles of photogrammetry. The depth map could be obtained by other techniques, such as laser triangulation or the time-of-flight method.
[0037] In the present embodiment, the articles IT brought into the picking zone PickZ by the first collective article movement unit, which are then manipulated by the manipulator arm MA so as to make them pass from a state disordered from the bulk unloading of the Trck truck on the first article movement unit CONV to an ordered state on the second INS unit for collective article movement. The manipulator arm MA has the function of individually manipulating these articles by grasping them one by one to deposit them one by one in a deposit zone DZ of the second INS movement unit.
[0038] The disordered state is characterized by an irregular distribution of the articles, possibly with articles piled up on top of each other or simply placed next to each other. The ordered state is characterized, for example, by placing the articles one by one on a continuously rotating conveyor belt, in such a way that each of the articles is placed in the middle of the conveyor belt, preferably at a constant pitch P and defined as the distance separating two of the articles placed consecutively on the conveyor belt by the manipulator arm.
[0039] The second unit INS for collectively moving articles is configured so as to move the articles placed therein towards an entrance of the sorting machine SORT, in a direction Dir2 which is defined as going from an area DZ for depositing the articles by the manipulator arm MA on the second unit for moving articles INS towards an entrance of the sorting machine SORT. The articles, entrusted in an orderly manner to the second unit for moving articles INS, can be correctly processed by the sorting machine SORT.
[0040] The ambition of the present invention is, in the present embodiment, a method for optimizing the processing of articles by the system illustrated by [Fig.2], in particular the control of the manipulator arm MA, but also the control of the first article movement unit CONV. This optimization method is based on reinforcement learning of manipulation of articles by the system SYS, applying firstly to the control of the manipulator arm MA.
[0041] [Fig.3] illustrates three logic blocks ALIM, VIS and PICK included in a logic application APP of the computer system INF, which can be used to describe the control of the system SYS by the control unit CONT. The application APP can be implemented by the computer system INF. The computer system INF also includes a simulation block SIM configured to simulate the system SYS and its operation.
[0042] The first supply logic block ALIM is designed to control the supply of articles to the picking zone. The ALIM block controls the first CONV unit for collective movement of articles: starting and stopping a conveyor belt and / or modulating the conveying speed of the articles, for example. This logic block is also responsible, where appropriate, for the sorting of the articles in order to aerate the bulk and facilitate the picking of the articles. The sorting can be implemented by any conventional means equipping the first unit of CONV displacement.
[0043] The second VIS visualization logic block is designed to analyze the images of the sampling area captured by the image sensor CAM1 and a topographic map derived from these images, according to conventional methods. These images include the articles brought from the sampling area in response to the control of the first ALIM logic block. The VIS block in particular performs a segmentation of the images so as to detect and isolate each of the visible articles, and calculates the individual characteristics of each of these isolated articles, such as their positions (heights, locations occupied in the plane formed by the sampling area), their orientations in an orthogonal reference associated with the sampling area (yaw, pitch, roll), and their colors and / or textures, the latter being indicative of the nature of the packaging of the articles (typically, rigid boxes such as cardboard boxes or flexible packaging such as plastic films).Patent applications FR2305495 and FR2301782 detail implementations of such functions.
[0044] The third logic block PICK is designed to determine a strategy for picking up the articles identified by the second logic block VIS.: choice of the next article to be picked up by the manipulator arm, approach trajectory of the latter and movement trajectory of the picked up article from its picking point to its destination point, here the drop-off zone DZ. The picking strategy is determined as a function of information on the location of the drop-off point DZ, the location of each of the articles, any overlaps between articles, the height of the articles, the configurations of the manipulator arm and the gripping head, possibly the choices of the gripping suction cups to be activated (implementation of the suction for each suction cup independently of the others) or, where appropriate, the choice of a drop-off zone, for example located on one of several inserters.
[0045] The SIM simulation block is based on a three-dimensional digital twin of the SYS system, which integrates a dynamic engine and an automation system that reproduce in particular realistic configurations of bulk items on the CONV conveyor unit and in the PickZ picking area. A texture extracted from an image of a real item is applied to each simulated item.
[0046] The three logic blocks ALIM, VIS and PICK can first be pre-trained by classical statistical approaches or heuristic methods determining operational parameters of these logic blocks optimizing the response of the system to a given situation, defined by the configuration of the articles to be processed and the configuration of the system itself, the manipulator arm in particular. Only subsequently, these blocks can benefit from reinforcement learning as described below. An advantage of this two-step learning is to reduce the convergence time of the learning algorithm. Alternatively, these logic blocks can be di properly trained using reinforcement learning to define their operational parameters.
[0047] In all cases, the operational parameters of the logic blocks must be determined in such a way that the system as a whole responds in the most appropriate way possible to a given situation, that is to say in such a way as to maximize the processing rate of the articles while minimizing the occurrence of problems such as collisions, fly-outs or even multiple takes. The occurrences of multiple takes and fly-outs can for example be determined by checking the number of articles deposited on the inserter INS. The occurrences of collisions can be determined by trajectory calculations, the topography of the environment and the characteristics of the articles and of the manipulator arm being known.
[0048] [Fig.4] illustrates a reinforcement learning method for logic blocks.
[0049] Principle
[0050] The principle of reinforcement learning methods is to train an intelligent agent to perform actions in a simulation so that it is able to perform them in real conditions once the learning is done. Reinforcement learning has two main elements: the agent, and the environment.
[0051] The agent is the algorithm that will seek to learn, and which is ultimately expected to be capable of controlling the manipulator arm and the first collective movement unit for supplying articles to the picking area. In this embodiment, it is made up of the logic application APP, which controls the manipulator arm MA and the first collective article movement unit CONV. It is this application whose operation is sought to be optimized by means of reinforcement learning.
[0052] The environment is defined by the individual characteristics of each of these articles (nature, positioning) of the PickZ picking zone, the characteristics of the handling arm, and those of the conveyor unit. It is in this environment that the agent is located and it is on this environment that he acts.
[0053] Three types of information are to be considered in the space defined by the agent and the environment: actions, states and rewards.
[0054] An action a is an action on the environment carried out by the agent following a decision on his part. This concerns in particular the choice of an article to be manipulated by the manipulator arm and the way in which the manipulator arm carries out this manipulation (path taken, etc.). It can also concern the control of the conveyor unit.
[0055] A state s is a situation at a time t of the agent in the environment. The state is updated after each action performed by the agent.
[0056] A reward r is a rating indicating the quality of an action applied to a given state of the environment, good (for example: article handled adequately, i.e. brought to the drop-off zone DZ) or bad (for example: occurrence of a collision, a multiple grip, a fly-out... A reward can be a positive number for good actions, and negative for bad actions.
[0057] Reinforcement learning consists of teaching the agent to act in such a way as to maximize the sum of the rewards generated. Thus, during learning, it will seek to identify the actions which will allow it to maximize the rewards generated in a given situation, i.e. a state of the environment.
[0058] Q-learning algorithm
[0059] A reinforcement learning algorithm is Q-learning, which uses the concept of quality (or value) of a state-action tuple, i.e. the total reward generated by an action a in a state s. An action is defined by the operational parameters of the element (manipulator arm and its gripping head, first unit of collective movement of articles) which implements it.
[0060] The formal definition of quality can be defined as follows for Q-learning:
[0061] Q ( s, a ) = r ( a; a ) + ym ax Q(s, a) Eq. 1
[0062] Equation Eq. 1 mathematically defines the Q-value of a state-action tuple, r(s,a) defining the reward of an action a of the agent applied to a state s of the environment, F representing a devaluation factor between 0 and 1, and max<2Q, ci) representing the maximum possible reward of the action a in a state s'. The states s and s' belong to a state space S. The action a belongs to an action space A.
[0063] These algorithms "learn" the quality Q(s, a) of each state-action tuple as the space A of actions a of each state s is explored. The Q-values are used to choose which action a is performed for each state s. At the beginning of the learning, the Q-values are randomly initialized. As the learning progresses, the learned Q-values will be increasingly exploited.
[0064] Implementation of Q-learning
[0065] An implementation 100 of the Q-learning algorithm using the computer system INF and the application APP is illustrated by [Fig.4]. This implementation is a learning phase of the logical application APP for controlling the system SYS.
[0066] The environment used for learning is, at least initially, a simulation of the SYS system, in order to avoid damaging the real system.
[0067] In a step S10, a SIMUL simulation of a state Ssys of the system SYS is generated by the SIM simulation module, as a function of operational parameters PARAM.ALIM for controlling the CONV article movement unit, and therefore the supply of articles to the picking area. These parameters control the starting and stopping of a conveyor belt, the modulation of the article conveying speed, or the separation of articles.
[0068] At a step S20, the application APP generates a digital image IMG.Virt in color and a depth map of the virtual sampling area, as they would be perceived by the sensor CAM1 in the real system. The rendering of the generated image is preferably photorealistic so as to bring the simulation closer to the operating conditions of the real system SYS.
[0069] In a step S30, the VIS logic block analyzes the digital image by segmenting the digital color image and by exploiting the depth map so as to detect and isolate each of the visible articles, and calculates the individual characteristics CAR of each of these articles: positions (heights, locations occupied in the plane formed by the sampling zone), orientations, colors and / or textures, as indicated above in the case of the analysis of real system images. The analysis is carried out according to operational parameters Param.VIS of the VIS block, predetermined by the practitioner for the implementation of the image analysis.
[0070] At a step S40, the logic block PICK performs a first task, which is to define a strategy STR for handling the articles by defining the next article IT to be handled on the basis of (i) the individual characteristics CAR of the articles calculated at step S30 and (ii) the characteristics of the manipulator arm and, where appropriate, its gripping head.
[0071] The operational parameters Param.PICK are control parameters of the manipulator arm MA, and are representative of the next article that it has been decided to grasp and the characteristics of this article: position, orientations, dimensions, or even nature of the packaging.
[0072] At this stage, the determination of the trajectories of the manipulator arm to pick up the next item to be handled and bring it to the drop-off area can be done by reinforcement learning, but, in this embodiment, they are determined by conventional pathfinding algorithms, such as the Dijkstra algorithm. This determination is based on the topological data of the scene obtained by means of the CAM1 sensor: RGB color image, depth map and the geometry and dimensions of the manipulator arm.
[0073] On the other hand, taking into account problem situations (multiple grips, nature of the article's packaging, collisions, fly-outs) is done by reinforcement learning.
[0074] Thus, among the operational parameters Param.PICK for controlling the manipulator arm, parameters of a first set of parameters which are dedicated to the de The completion of the manipulator arm trajectories can be determined by a path-finding algorithm, while parameters of a second set of parameters that are dedicated to addressing problem situations are determined by reinforcement learning.
[0075] At a step S50, the simulation block SIM simulates the application of the manipulation strategy defined at step S40, which corresponds, in terms of Q-learning, to an action a applied by the agent, which is here the manipulator arm, to a current state s of the environment to arrive at a new state s: movement of the manipulator arm to the article that it has been decided to grasp, grasping of this article, then transport and deposit of the article to its deposit point.
[0076] In a step S60, the supply logic block ALIM performs a second task ALIM?, which is to determine a strategy for managing the supply of articles to the picking zone. The supply is managed by means of the movement unit CONV. The supply management may consist of starting the movement unit CONV for a variable duration, determining a conveying speed of the articles by the conveying unit CONV, or applying a bulk article stripping operation by the conveying unit. The determination is made on the basis of the new configuration of the articles simulated by the simulation block SIM, possibly processed in the same manner as in steps S20 and S30.The inability of the manipulator arm to pick up an item may also represent a signal indicating the need for the CONV conveyor unit to empty the picking area of items, for example by tipping the items at the end of the conveyor belt into a collection bin.
[0077] At a step S70, the application establishes a REP report representative of the time required to carry out the actions of the manipulator arm, the length of the paths of the manipulator arm, the success of the manipulation of an article (actually deposited or not in the deposit zone) and any problems encountered: multiple gripping, collision, and / or fly-out.
[0078] At a step S80, the application APP calculates the reward REC for the actions performed at step S50 by the manipulator arm MA on the basis of the report established at step S70. The principle is to generate a reward that is all the higher as the manipulation of the article proceeds without incident, that is to say that potential problems such as multiple gripping and collision do not occur, and that the supply is fluid, that is to say allows a high flow rate of manipulated articles (information deriving from the time taken to perform the actions by the manipulator arm for moving each article).
[0079] The following table summarizes one way of generating rewards within the Q-learning algorithm. Rewards are generated in response to the REP report which indicates the overall performance of the system with respect to moving of the articles and the system throughput, without individually evaluating each of the actions (listed in Action Space A leading to these overall performances, and which correspond to the actions carried out by the system to fulfill the functions of steps S40 and S60 (first task and second task), implemented respectively by the PICK logic block and the ALIM logic block. The State Space S column lists the parameters taken into account in determining the actions to be performed (each defined by operational parameters of the mechanisms responsible for implementing them), which belong to the Action Space. Tasks State space S Action space A Rewards Choice of the item to be handled • Configuration of the manipulator arm at the start of the sequence • Grasp each item in the picking zone • Empty the scene (no item can be grabbed) • Configuration of the gripping head +1 for each item placed in the drop-off zone -1 for each occurrence of a problem (multiple grabs, collision, fly-out) Supply of items to the picking zone • Configuration of items in the picking zone • Configuration of items on the CONV unit • Height of the chosen item • Position of the drop-off zone • Start-up of the CONV movement unit, yes or no • Duration of start-up of the CONV movement unit • Movement speed • Require filling the scene
[0080] At a step S90, in response to the reward, the parameters Param.ALIM, and Param.PICK are adjusted, it is also possible to consider adjusting the parameters Param.VIS during this step. Steps S10 to S90 are repeated until the parameters converge sufficiently, according to a criterion established by the operator of the system SYS, or until an expected performance level of the system is reached.
[0081] Once the parameters are adjusted, the system is ready to be used for sorting items in the actual system. The handling method 200 is illustrated in [Fig.6].
[0082] At a step S210, in the SYS system, IT articles are brought to the picking area by the article moving unit CONV, on the basis of the adjusted Param.ALIM parameters.
[0083] In a step S220, the image sensor CAM1 captures a digital image IMG.Real of the picking area and a topographic map of this picking area is generated, including the articles which were brought to this area in step S210, possibly by means of a pair of digital images, this image and this map are stored in a memory of the computer system INF.
[0084] At a step S230, the logic block VIS analyzes the digital image IMG.Real and the topographic map so as to detect and isolate each of the visible articles, and calculates the individual characteristics CAR of each of these articles, in a similar manner to step S30 of the learning phase with the difference that real data and not virtual or simulation data are analyzed this time.
[0085] At a step S240, the logic block PICK defines a strategy STR for handling the articles by defining the next article IT to be handled on the basis of the individual characteristics CAR of the articles calculated at step S230 and the operational parameters Param.PICK.
[0086] At a step S250, the strategy defined at step S240 is implemented: the manipulator arm will grasp the article that it was decided to grasp and moves it to its destination on the inserter INS.
[0087] At a step S260, the logic block ALIM determines a strategy for managing the supply of articles to the picking zone, possibly on the basis of new data describing the picking zone obtained by means of the sensor.
[0088] The sequence of steps S210 to S260 is repeated cyclically as long as there are articles to be processed. Steps S230, S240, and S260 for implementing the real system are similar to steps S30, S40, and S60 of the learning phase, the only difference being the origin of the data used: simulations (image, topographic map) for learning 100 and data representative of a real situation for implementing the method 200 for handling articles.
[0089] In this document, the figures are not necessarily to scale. Some features and components may be shown exaggerated relative to other components or in a somewhat schematic form, and some details of conventional elements may not be shown in the interest of clarity and conciseness.
[0090] Of course, the invention is not limited to the embodiment described and variant embodiments can be made without departing from the scope of the invention as defined by the claims.
Claims
Claims
1. Method (100) for reinforcement learning of a system (SYS) for individual handling of articles, the system (SYS) comprising a unit (CONV) for collectively moving the articles (IT) to a picking zone (PickZ), a manipulator arm (MA) configured to individually grasp and move articles (IT) located in the picking zone (PickZ), a sensor (CAM1) configured to acquire images representative of the picking zone (PickZ), the method comprising the implementation of the following steps by means of a computer system (INF): - simulation (S 10) of the system (SYS) in state (SSYs); - generation (S20) of digital images (IMG.Virt) representative of the sampling area of the simulated system; - analysis (S30) of at least one of the generated images and calculation of characteristics (CAR) representative of simulated articles visible in the generated image; - on the basis of the characteristics calculated in the analysis step (S30), determination (S40) of an article to be grasped among the simulated articles and the trajectories of the simulated manipulator arm to grasp this article and bring it to a drop-off zone; - simulation (S50) of an effect on the state of the simulated system of the manipulation of the article to be grasped by the manipulator arm in response to the determination (S40); - establishment (S70) of an indicative report (REP) (i) of a representative duration for the manipulator arm to grasp the article to be grasped and bring it to the deposit zone and (ii) of any problems encountered during the simulated handling; - generation (S80) of a reward (REC) in the form of a number in response to the report (REP); and - adjustment (S90) of operational parameters of the system, comprising at least operational parameters (Param.PICK) of the handling arm of the step of determining the reward in response (REC).
2. The method according to claim 1, wherein the operational parameters (Param.PICK) of the manipulator arm (MA) are chosen from a choice of an article to be grasped among the visible simulated articles (IT) and a configuration of a gripping head (GH) of the manipulator arm (MA).
3. The method according to claim 1 or 2, further comprising a step (S60) of determining a strategy for supplying articles to the picking area of the simulated system by means of the unit (CONV) for collective movement of the articles, the step (S90) of adjusting the operational parameters of the system further comprising an adjustment of operational parameters (Param.ALIM) of the unit (CONV) for collective movement of the articles.
4. The method according to claim 3, in which the operational parameters (Param.ALIM) of the unit (CONV) for collective movement of the articles are chosen from parameters for controlling the starting and stopping of a conveyor, for modulating a conveying speed of the articles, and for separating the articles.
5. The method according to any one of claims 1 to 4, in which the report (REP) is representative of a time necessary to carry out the actions of the manipulator arm, a length of paths of the manipulator arm, a success of the manipulation of an article and the problems possibly encountered: multiple grip, collision, and / or fly-out.
6. 6. The method according to any one of the preceding claims, in which the problems possibly encountered are chosen from a multiple grip of articles by the handling arm, a fly-out of an article being handled by the handling arm and a collision of an article being moved by the manipulator arm.
7. 7. The method according to any one of claims 1 to 6, wherein the step (S40) of determining the trajectories of the simulated manipulator arm is carried out by means of a path finding algorithm such as the Dijkstra algorithm.
8. Method (200) for individual handling of articles by means of the system (SYS) having followed the learning method according to any one of claims 1 to 7, comprising the implementation of the following steps by means of a computer system (INF): - capturing (S220) a digital image (IMG.Real) and generating a topographic map of the picking area (PickZ) by means of a sensor (CAM1); - analyzing (S230) the digital image (IMG.Real) and the topographic map so as to detect and isolate articles visible in the digital image (IMG.Real), and calculating individual characteristics (CAR) of each of these articles; - defining (S240) a next article (IT) to be handled on the basis of the
9. individual characteristics (CAR) of items and operational parameters (Param.PICK); and - in response to the definition (S240) of the next item to be handled, manipulation (250) of the next item to be handled by means of the manipulator arm (MA). System (SYS) for implementing the method (200) for individual handling of articles according to claim 8, comprising: - a sampling zone (PickZ); - a sensor (CAM1) configured to acquire images representative of the sampling area (PickZ); - a first unit (CONV) for collective movement of articles configured to bring articles to the picking zone (PickZ); - a manipulator arm (MA) configured to individually manipulate items (IT) located in the picking area; and - a computer system (INF) for implementing a logic application (APP) intended for controlling the manipulator arm and the first unit (CONV) for collective movement of articles.
Citation Information
Patent Citations
THERMAL DISTRIBUTION STATION
FR2301782A1
LUBRICATING COMPOUNDS THAT IMPROVE THE SWELLING OF SEALS
FR2305495A1
Lather tanning, soaking, liming,dyeing,processes - requiring shorter vat times by introducing gas at pressure above vapour pressure
FR2309639A1
Device for supplying flat items for a bucket conveyor
FR3019068B1
Installation for the Separation and Individualization of Heterogeneous Postal Items
FR3031046A1