Method and apparatus for controlling palletizing robot by using reinforcement learning

A reinforcement learning-based learning model optimizes palletizing robot control for efficient item placement on pallets and buffers, enhancing loading efficiency and space utilization.

WO2025143672A1PCT designated stage expired Publication Date: 2025-07-03CMES INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/020524
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-17
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing palletizing robots lack optimal methods for efficiently loading items onto pallets, particularly in terms of space utilization and item placement, and there is a need for improved control strategies using reinforcement learning.

Method used

A learning model trained through reinforcement learning is used to determine item placement on pallets or buffers, generating action information based on state and reward information to optimize loading processes.

Benefits of technology

Improves space utilization and enhances the number of items loaded per pallet by optimizing item placement, and allows for temporary storage in buffers before final loading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024020524_03072025_PF_FP_ABST
    Figure KR2024020524_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A method and an apparatus for controlling a palletizing robot by using reinforcement learning are disclosed. According to one embodiment of the present disclosure, the method for controlling a palletizing robot may comprise the steps of: using a learning model to determine a first article from among a plurality of articles, and a first pallet or a first buffer from among one or more pallets and one or more buffers in which the first article is to be positioned; using the learning model so as to determine the space in which the first article is to be positioned and the direction in which the first article is to be positioned; and issuing a control command to the robot on the basis of the first pallet or the first buffer, the space in which the first article is to be positioned, and the direction in which the first article is to be positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for controlling a palletizing robot using reinforcement learning

[0001] The present invention relates to a method and device for controlling a palletizing robot using reinforcement learning. More specifically, the present invention relates to a method and device for training a learning model through reinforcement learning and controlling a palletizing robot using the learning model.

[0002] The content described below merely provides background information related to the present embodiment and does not constitute prior art.

[0003] In modern industrial society, the importance of industrial automation is growing across various sectors to increase productivity, improve product quality, and reduce production costs. Industrial robots are being used in various fields to achieve this automation. In particular, industrial robots are being used in areas such as parts assembly, part transport, product sorting, and product loading.

[0004] Recently, palletizing robots have been developed among industrial robots to distribute mass-produced goods. Palletizing is the process of stacking items onto pallets. Palletizing robots perform this palletizing process, handling items in pallet units.

[0005] Reinforcement learning is a method for training a learning model to interact with its environment to achieve a goal. The goal of reinforcement learning is to determine which actions the learning model should take to maximize rewards. Reinforcement learning is widely used in robotics and artificial intelligence. Therefore, reinforcement learning can also be applied to palletizing robots, and research is needed to utilize reinforcement learning to optimize palletizing robots' ability to load items onto pallets.

[0006] The purpose of this disclosure is to train a learning model using reinforcement learning and to control a palletizing robot using the learning model.

[0007] Additionally, according to one embodiment, the purpose is to temporarily position items in a buffer other than a pallet.

[0008] In addition, according to one embodiment, the purpose is to train a learning model using offline learning data and to control a palletizing robot using the learning model.

[0009] The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0010] According to the present disclosure, a method for controlling a palletizing robot may include a process of determining a first item among a plurality of items on a rail using a learning model and determining a first pallet or a first buffer among one or more pallets and one or more buffers on which the first item is to be positioned, a process of determining a space in which the first item is to be positioned and a direction in which the first item is to be positioned using the learning model, and a process of issuing a control command to the robot based on the space in which the first pallet or the first buffer and the first item are to be positioned and the direction in which the first item is to be positioned, wherein the learning model may be a model that generates action information using state information and reward information.

[0011] According to the present disclosure, a device for controlling a palletizing robot includes a memory and at least one processor, wherein the at least one processor determines a first item among a plurality of items on a rail using a learning model, determines a first pallet or a first buffer among one or more pallets and one or more buffers on which the first item is to be positioned, determines a space in which the first item is to be positioned and a direction in which the first item is to be positioned using the learning model, and issues a control command to the robot based on the space in which the first pallet or the first buffer and the first item are to be positioned and the direction in which the first item is to be positioned, and the learning model may be a model that generates action information using state information and reward information.

[0012] According to the present disclosure, a computer-readable recording medium is a computer-readable recording medium having stored thereon a command, which, when executed by the computer, causes the computer to perform a process of determining a first item among a plurality of items on a rail using a learning model and determining a first pallet or a first buffer among one or more pallets and one or more buffers on which the first item is to be positioned, a process of determining a space in which the first item is to be positioned and a direction in which the first item is to be positioned using the learning model, and a process of issuing a control command to a robot based on the space in which the first pallet or the first buffer and the first item are to be positioned and the direction in which the first item is to be positioned, wherein the learning model may be a model that generates action information using state information and reward information.

[0013] According to the present disclosure, a learning model is trained using reinforcement learning, and a palletizing robot is controlled using the learning model, thereby improving the utilization rate of loading space and increasing the number of items loaded on a pallet.

[0014] Additionally, according to one embodiment, there is an effect of temporarily placing goods in a buffer other than the pallet, thereby enabling goods to be loaded onto the pallet in an optimal condition.

[0015] In addition, according to one embodiment, there is an effect of training a learning model using offline learning data and controlling a palletizing robot using the learning model, thereby improving the qualitative evaluation index of the palletizing robot.

[0016] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0017] FIG. 1 is a block diagram illustrating a system for controlling a palletizing robot according to one embodiment of the present disclosure.

[0018] FIG. 2 is a drawing for explaining a method for a palletizing robot to perform palletizing according to one embodiment of the present disclosure.

[0019] FIG. 3 is a flowchart illustrating a process of training a learning model using reinforcement learning according to one embodiment of the present disclosure.

[0020] FIG. 4 is a drawing for explaining a process of controlling a palletizing robot according to one embodiment of the present disclosure.

[0021] Hereinafter, some embodiments of the present disclosure will be described in detail using exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, when describing the present disclosure, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present disclosure.

[0022] In describing components of embodiments according to the present disclosure, symbols such as first, second, i), ii), a), b) may be used. These symbols are only for distinguishing the components from other components, and the nature, order, or sequence of the components are not limited by the symbols. When a part in the specification is said to "include" or "have" a component, this does not mean that other components are excluded, but rather that other components may be included, unless explicitly stated otherwise.

[0023] The detailed description set forth below, together with the accompanying drawings, is intended to explain exemplary embodiments of the present disclosure and is not intended to represent the only embodiments in which the present disclosure may be practiced.

[0024] FIG. 1 is a block diagram illustrating a system for controlling a palletizing robot according to one embodiment of the present disclosure.

[0025] Referring to FIG. 1, a system for controlling a palletizing robot includes all or part of an interface unit (110), a detection unit (120), a control unit (130), and a robot (140). The system for controlling a palletizing robot and each component thereof may be implemented in hardware or software, or a combination of hardware and software. In addition, the functions of each component may be implemented in software, and one or more processors may be implemented to execute the functions of the software corresponding to each component.

[0026] The interface unit (110) provides an interface to the user. The interface unit (110) receives information from the user and transmits it to the control unit (130), or outputs results of control commands from the control unit (130) and results of the operation of the robot (140). For example, the control unit (130) can receive information about items, information about a conveyor belt, information about pallets, learning data, information about the robot (140), and information about buffers from the user using the interface unit (110).

[0027] The detection unit (120) detects items on the conveyor belt and items within buffers and pallets to obtain information on the items and transmits the obtained information on the items to the control unit (130). Here, the information on the items may include information on the size of the items, the number of items, the identifiers of the items, the shape of the items, etc. The detection unit (120) may include radar, a camera, a lidar, and other sensors. In addition to the items, the detection unit (120) may detect spaces within the conveyor belt, pallets, or buffers to obtain information on the pallets and buffers. The detection unit (120) transmits information on the pallets and buffers to the control unit (130). The information on the pallets may include information on the size of the pallets, the identifiers of the pallets, the space within the pallets, and the number of pallets. The information on the buffers may include information on the size of the buffers, the identifiers of the buffers, the space within the buffers, and the number of buffers.

[0028] The control unit (130) can issue a control command to the robot (140) to load items onto multiple pallets or multiple buffers using information on items, information on pallets, and information on buffers. The control unit (130) can issue a control command to the robot (140) using information received from the detection unit (120) and the interface unit (110). The control unit (130) can be a device (hereinafter, “robot control device”) that controls a palletizing robot using reinforcement learning of the present disclosure. The control unit (130) includes a learning model (131), a simulation unit (132), and a reward unit (133). The control unit (130) can issue a control command to the robot (140) using action information generated by the learning model (131).

[0029] The learning model (131) may correspond to a trained learning model. The learning model (131) may be a deep learning-based model. The control unit (130) may additionally include a learning unit (not shown) for pre-training the learning model (131). The learning unit may pre-train the learning model (131) using supervised learning, unsupervised learning, semi-supervised learning, and / or reinforcement learning.

[0030] The learning unit performs reinforcement learning to load items onto multiple pallets while satisfying various constraints by performing simulations to train the learning model (131). The learning model (131) receives state information from the simulation unit (132) and reward information from the reward unit (133). The learning model (131) generates action information using the received state information and reward information. For example, the action information may include information on whether to select an item on a conveyor belt, whether to select an item within a buffer, where to load an item within a pallet, what direction to load an item within a pallet, which buffer among multiple buffers to select, where to position an item within a buffer, what direction to position an item within a buffer, and which pallet among multiple pallets to select.

[0031] The learning unit can train the learning model (131) to generate action information that maximizes the reward value of the reward information generated by the reward unit (133). The learning unit can train the learning model (131) to generate action information that maximizes the cumulative reward value of the reward information generated by the reward unit (133). When a specific goal is set, the learning unit can train the learning model (131) by performing reinforcement learning to achieve the set goal.

[0032] The simulation unit (132) performs a simulation using the action information generated by the learning model (131). The simulation unit (132) performs the simulation to generate state information and transmits the state information to the learning model (131). The simulation unit (132) can generate state information using information on items, information on pallets, and information on buffers obtained from the detection unit (120). For example, the state information may include information on the size of items on the conveyor belt, the number of items on the conveyor belt, the identifiers of items on the conveyor belt, the number of items in a buffer, the size of items in a buffer, the identifiers of items in a buffer, the positions of items in a buffer, the identifiers of buffers, the number of items loaded on a pallet, the size of items loaded on a pallet, the positions of items loaded on a pallet, the identifiers of items loaded on a pallet, the size of an empty space in a pallet, the identifiers of pallets, the size of an empty space in a buffer, the size of buffers, the number of layers of items to be loaded, etc.

[0033] The compensation unit (133) generates compensation information using the results simulated by the simulation unit (132). For example, the compensation information may include information on a compensation value, the utilization rate of the loading space within the pallet, the stability of the loaded items, and the number of loaded items. Here, the compensation value may be a value that digitizes the information on the utilization rate of the loading space within the pallet, the stability of the loaded items, and the number of loaded items into a preset standardized value. The compensation unit (133) transmits the compensation information to the learning model (131). The compensation unit (133) may use a compensation function that generates compensation information as a compensation for each action information. The compensation unit (133) may be trained to generate compensation information to enable the learning model (131) to generate optimal action information.

[0034] The robot (140) may be a palletizing robot with multiple joints connected to perform precise movements as desired by the user. The robot (140) can accurately perform tasks such as transport or loading using the multiple joints. Instead of a single power source, multiple motors and reducers may be installed on each joint axis within the robot (140). The robot (140) may have a structure in which the movement of each joint axis is controlled by each motor. For example, the robot (140) may be a six-axis robot. The robot (140) picks up items on a conveyor belt under the control command of the control unit (130) and loads the items onto a pallet or places them in a buffer.

[0035] FIG. 2 is a drawing for explaining a method for a palletizing robot to perform palletizing according to one embodiment of the present disclosure.

[0036] Referring to FIG. 2, item 1 (210), item 2 (220), and item 3 (230) can move on a conveyor belt (240). Item 1 (210), item 2 (220), and item 3 (230) can be packed in boxes. More than three items can move on the conveyor belt (240). The detection unit (120) can detect the identifier, shape, weight, size, etc. of item 1 (210), item 2 (220), and item 3 (230). In addition to item 1 (210), item 2 (220), and item 3 (230), the detection unit (120) can detect the conveyor belt (240), the pallet (260), and the buffer (270). The robot (250) can select and pick up one of the items 1 (210), 2 (220), and 3 (230) by a control command from the control unit (130). That is, the robot (250) can select an item to be loaded from among the items 1 (210), 2 (220), and 3 (230) based on the control command from the control unit (130) rather than loading the items in the order of the items 1 (210), 2 (220), and 3 (230).

[0037] The robot (250) can select a pallet or buffer to be loaded by a control command from the control unit (130). The robot (250) can position the selected item in a specific direction by the control command from the control unit (130). Here, the specific direction may be a direction that takes into account the rotational direction and angle of the item, and the coordinates of each vertex of the item when the upper left vertex of the pallet is used as a reference. The robot (250) can load the selected item on a pallet (260) or position it on a buffer (270) by the control command from the control unit (130). There may be a plurality of pallets (260) and buffers (270).

[0038] For example, if the control unit (130) issues a control command to select item 3 (230) using the action information generated by the learning model (131), the robot (250) can pick up item 3 (230). If the control unit (130) issues a control command to position item 3 (230) in a specific space of a specific pallet among multiple pallets in a specific direction using the action information generated by the learning model (131), the robot (250) can position item 3 (230) in a specific space of a specific pallet in a specific direction.

[0039] For example, if the control unit (130) issues a control command to select item 2 (220) using the action information generated by the learning model (131), the robot (250) can pick up item 2 (220). If the control unit (130) issues a control command to position item 2 (220) in a specific space of a specific buffer among a plurality of buffers in a specific direction using the action information generated by the learning model (131), the robot (250) can position item 2 (220) in a specific space of a specific buffer in a specific direction. Thereafter, if the control unit (130) issues a control command to position item 2 (220) in a buffer in a specific space of a specific pallet among a plurality of pallets in a specific direction using the action information generated by the learning model (131), the robot (250) can position item 2 (220) in a specific space of a specific pallet in a specific direction.

[0040] FIG. 3 is a flowchart illustrating a process of training a learning model using reinforcement learning according to one embodiment of the present disclosure.

[0041] Referring to FIG. 3, the learning model (131) receives state information from the simulator unit (132) and reward information from the reward unit (133) (S310). The learning model (131) generates action information using the state information and reward information (S320). The learning model (131) can generate action information that maximizes the reward value of the reward information. The learning model (131) can transmit the action information to the simulation unit (132). The learning model (131) can be trained in advance through offline reinforcement learning.

[0042] Offline reinforcement learning is a method of training a learning model using sequence data obtained by a user directly loading items onto multiple pallets. The sequence data may include time-series data showing a user loading specific items onto a conveyor belt, onto a specific pallet, at a specific location, and in a specific direction; time-series data showing a user loading specific items onto a specific pallet, at a specific location, and in a specific direction within a buffer; and time-series data showing a user positioning specific items within a specific buffer, at a specific location, and in a specific direction. The control unit (130) may receive sequence data from the user using the interface (110). The learning unit may pre-train the learning model (131) using the sequence data.

[0043] If the learning model (131) is pre-trained by offline reinforcement learning, the stability of the learning can be increased, and the performance of the learning model (130) can be improved. If the learning model (131) is pre-trained by offline reinforcement learning, the learning unit can train the learning model (131) by reflecting qualitative factors that are difficult to quantitatively calculate using the reward value of the reward information. For example, the qualitative factors can be the stability of loaded items compared to the utilization rate of the loading space of the pallet, the loading pattern of the items by size or weight, and the loading method based on the user's know-how.

[0044] The simulation unit (132) performs a simulation using action information (S330). The simulation unit (132) performs the simulation and can generate status information using information on items, information on pallets, and information on buffers obtained from the detection unit (120). The simulation unit (132) can transmit the generated status information to the learning model (131). The compensation unit (133) generates compensation information using the simulated results (S340).

[0045] The simulation unit (132) can perform a simulation using action information that reflects the user's feedback. The action information that reflects the user's feedback may be action information determined by the user among the action information generated by the learning model (131) during the learning process. Among the action information generated by the learning model (131) during the learning process, the action information that generates the highest reward value of the reward information according to the user's judgment may be selected. The control unit (130) can provide the user with the action information generated by the learning model (131) during the learning process using the interface (110). The control unit (130) can receive the action information that reflects the user's feedback using the interface (110). The reward unit (133) can generate reward information using the simulation result using the action information that reflects the user's feedback.

[0046] The reward unit (133) can be trained using action information reflecting user feedback as learning data. Accordingly, the reward unit (133) can generate reward information using a reward function desired by the user. The reward function desired by the user may be a function that reflects qualitative factors that are difficult to quantitatively estimate.

[0047] The reward unit (133) determines whether the reward value of the reward information is greater than a threshold value (S350). Here, the threshold value may be a predefined arbitrary value. If it is determined that the reward value is not greater than the threshold value (S350-NO), the learning model (131) receives status information from the simulator unit (132) and receives reward information from the reward unit (133) (S310). Thereafter, the learning unit can perform reinforcement learning to continue training the learning model (131). If it is determined that the reward value is greater than the threshold value (S350-YES), the learning unit terminates reinforcement learning (S360).

[0048] FIG. 4 is a drawing for explaining a process of controlling a palletizing robot according to one embodiment of the present disclosure.

[0049] Referring to FIG. 4, the detection unit (120) obtains information on items moving on a conveyor belt (S410). In addition to information on the items, the detection unit (120) can obtain information on the conveyor belt, pallets, buffers, etc. The control unit (130) determines which items to load among the items on the conveyor belt (S420). The control unit (130) can determine which items to load using a learning model (131). The learning model (131) may be a model learned by performing reinforcement learning. The control unit (130) can determine which items to load using action information generated by the learning model (131). The robot (140) can pick up the determined items on the conveyor belt.

[0050] The control unit (130) determines one of one or more pallets and one or more buffers (S430). The control unit (130) can use the action information generated by the learning model (131) to determine one pallet or one buffer on which to load or place an item among the one or more pallets and one or more buffers. The control unit (130) determines a space in which to place an item within one or more pallets or one or more buffers (S440). The control unit (130) can use the action information generated by the learning model (131) to determine a space in which to place an item within one or more pallets or one or more buffers. The control unit (130) can use the action information generated by the learning model (131) to determine a direction in which to place an item. The control unit (130) issues a control command to the robot (140) to place an item at the determined location (S450). The robot (140) can place the item at the determined location in the determined direction based on the control command.

[0051] Each component of the device or method according to the present invention may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0052] Various implementations of the systems and techniques described herein may be implemented as digital electronic circuits, integrated circuits, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations of one or more computer programs executable on a programmable system. The programmable system includes at least one programmable processor (which may be a special purpose processor or a general purpose processor) coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device. Computer programs (also known as programs, software, software applications, or code) include instructions for the programmable processor and are stored on a "computer-readable recording medium."

[0053] A computer-readable recording medium includes any type of recording device that stores data that can be read by a computer system. Such a computer-readable recording medium may be a non-volatile or non-transitory medium such as a ROM, CD-ROM, magnetic tape, floppy disk, memory card, hard disk, magneto-optical disk, storage device, and may further include a transitory medium such as a data transmission medium. Furthermore, the computer-readable recording medium may be distributed across network-connected computer systems, so that computer-readable code can be stored and executed in a distributed manner.

[0054] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.

[0055] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of ​​the present embodiment, but rather to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.

[0056]

[0057] CROSS-REFERENCE TO RELATED APPLICATION

[0058] This patent application claims priority to Korean Patent Application No. 10-2023-0197541, filed on December 29, 2023, the entire contents of which are incorporated herein by reference.

Claims

1. In a method performed by a robot control device, A process of determining a first item among a plurality of items on a rail using a learning model and determining a first pallet or a first buffer among one or more pallets and one or more buffers on which to place the first item; A process of determining a space in which to position the first item and a direction in which to position the first item using the above learning model; Including a process of issuing a control command to a robot based on the space in which the first pallet or the first buffer and the first article are to be positioned and the direction in which the first article is to be positioned, The above learning model is a method for generating action information using state information and reward information.

2. In paragraph 1, The above learning model is a method in which the learning model is trained to generate action information that maximizes the reward value of the reward information.

3. In paragraph 1, A method wherein the above status information includes information on items on a conveyor belt, information on items in a buffer, information on items loaded on a pallet, information on pallets, and information on buffers.

4. In paragraph 1, A method wherein the above compensation information includes information on a compensation value, information on the utilization rate of a loading space within a pallet, information on the stability of loaded items, and information on the number of loaded items.

5. In paragraph 1, A method in which the above action information includes information on whether to select a certain item on a conveyor belt, information on whether to select a certain item within a buffer, information on where to load the item within a pallet, information on the direction in which to load the item within the pallet, information on which buffer among a plurality of buffers to select, information on whether to position the item within the buffer, information on the space in which to position the item within the buffer, information on the direction in which to position the item within the buffer, and information on which pallet among a plurality of pallets to select.

6. In paragraph 1, The above learning model is trained using sequence data, The above sequence data is a time-series data in which a user loads the above multiple items.

7. In paragraph 1, The above action information includes action information determined by the user, A method wherein the above reward information includes reward information generated using action information determined by the user.

8. In the robot control device, memory; and comprising at least one processor, said at least one processor comprising: Using a learning model, a first item among a plurality of items on a rail is determined, and a first pallet or a first buffer among one or more pallets and one or more buffers on which the first item is to be placed is determined. Using the above learning model, determine the space in which the first item will be positioned and the direction in which the first item will be positioned, Based on the space in which the first pallet or the first buffer and the first article are to be positioned and the direction in which the first article is to be positioned, a control command is issued to the robot, The above learning model is a device that generates action information using state information and reward information.

9. In paragraph 8, The above learning model is a device that learns to generate action information that maximizes the reward value of the reward information.

10. In paragraph 8, The above status information is a device including information on items on a conveyor belt, information on items in a buffer, information on items loaded on a pallet, information on pallets, and information on buffers.

11. In paragraph 8, The above compensation information is a device including information on a compensation value, information on the utilization rate of the loading space within the pallet, information on the stability of the loaded items, and information on the number of loaded items.

12. In paragraph 8, A device wherein the above action information includes information on whether to select a certain item on a conveyor belt, information on whether to select a certain item within a buffer, information on where to load the item within a pallet, information on the direction in which to load the item within the pallet, information on which buffer among a plurality of buffers to select, information on whether to position the item within the buffer, information on the space in which to position the item within the buffer, information on the direction in which to position the item within the buffer, and information on which pallet among a plurality of pallets to select.

13. In paragraph 8, The above learning model is trained using sequence data, The above sequence data is a device that is time-series data in which the user loads the above multiple items.

14. In paragraph 8, The above action information includes action information determined by the user, A device wherein the above reward information includes reward information generated using action information determined by the user.

15. A computer-readable recording medium having stored thereon a command, wherein the command, when executed by the computer, causes the computer to: A process of determining a first item among a plurality of items using a learning model and determining a first pallet or a first buffer among one or more pallets and one or more buffers on which to place the first item; A process of determining a space in which to position the first item and a direction in which to position the first item using the above learning model; A process of issuing a control command to a robot based on the space in which the first pallet or the first buffer and the first article are to be positioned and the direction in which the first article is to be positioned is executed, The above learning model is a computer-readable recording medium that generates action information using state information and reward information.

Citation Information

Patent Citations

  • Loading pattern calculation device for setting position where article is loaded

    JP2017094428A

  • Loading pattern calculation device to load multiple kinds of articles and loading device

    JP2017165565A

  • Palletizing reinforcement learning apparatus and method

    KR102551039B1

  • Apparatus and method for arranging thoughts information based on hash tag

    KR102697190B1

  • KR20220123975A