Machine learning device, machine learning method, and machine learning program

A machine learning device using reinforcement learning optimizes paper transport control in image forming devices by adapting to user-specific conditions, reducing jams and downtime through intelligent control adjustments.

JP7743171B2Active Publication Date: 2025-09-24KONICA MINOLTA INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2019134502
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2019-07-22
Publication Date
2025-09-24
Estimated Expiration
2039-07-22

AI Technical Summary

Technical Problem

Image forming devices face challenges in optimizing paper transport control due to varying user environments and conditions, leading to increased likelihood of jams and downtime, as conventional control methods are not adaptive enough to handle unanticipated conditions.

Method used

A machine learning device employing reinforcement learning to acquire positional information of conveyed objects, calculate rewards based on predetermined rules, and generate control information to optimize drive source behavior for continuous and simultaneous conveyance, considering factors like humidity, temperature, and paper type.

Benefits of technology

This approach enables adaptive control that reduces unnecessary downtime by learning optimal transport conditions for varying user environments and conditions, ensuring efficient paper conveyance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007743171000001
    Figure 0007743171000001
  • Figure 0007743171000002
    Figure 0007743171000002
  • Figure 0007743171000003
    Figure 0007743171000003
Patent Text Reader

Abstract

To create control information for a drive source in order to appropriately convey a to-be-conveyed object.SOLUTION: In a mechanical learning device that learns a behavior of a drive source, in a conveyance apparatus that continuously conveys at least two to-be-conveyed objects along a conveyance path, pieces of information about the positions of the at least two to-be-conveyed objects on the conveyance path are acquired based detection results from a detection unit provided on the conveyance path. Based on the acquired pieces of information about the positions, rewards are calculated according to a predetermined rule. Based on the acquired pieces of information about the positions and the calculated rewards, a behavior value in intensified learning is calculated to learn a behavior. Control information for causing the drive source to exhibit a behavior determined based on the result of the behavior is created and output.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a machine learning device, a machine learning method, and a machine learning program that learn the behavior of a drive source in a conveying device that controls the conveying of multiple moving objects, and in particular to a machine learning device, a machine learning method, and a machine learning program that learn the behavior of a drive source in an image forming device that controls the conveying of multiple sheets of paper. [Background technology]

[0002] Image forming devices such as MFPs (Multi-Functional Peripherals) are subject to different operating environments and conditions depending on the user, which changes the state of the machine, and therefore the likelihood of jams occurring due to bending or pulling of paper during transport.When a jam occurs, the machine must stop and maintenance must be performed, and the time until maintenance becomes downtime, so optimal control according to the machine state is required.

[0003] However, there are an enormous number of combinations of usage environments and usage situations, and designing control that takes into account all possible usage environments and situations requires a significant amount of development time. Conventionally, control has been designed to prevent jams from occurring under both the worst and most typical conditions, but this method may not provide optimal control under unanticipated conditions, resulting in a lack of customer satisfaction.

[0004] To address this issue, a method for determining control conditions for a device using machine learning has been proposed. For example, Patent Document 1 listed below discloses a machine learning device that learns conditions associated with adjustment of a current gain parameter in motor control, the machine learning device including: a state observation unit that acquires an integral gain function and a proportional gain function of a current control loop, acquires an actual current, and observes state variables constituted by the integral gain function and the proportional gain function, as well as at least one of an overshoot amount, an undershoot amount, and a rise time of the actual current in response to a step torque command; and a learning unit that learns conditions associated with adjustment of the current gain parameter according to a training data set constituted by the state variables. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-034844 Summary of the Invention [Problem to be solved by the invention]

[0006] However, in the case of an image forming device, the paper transport state changes depending on environmental conditions such as temperature and humidity, as well as printing conditions such as the lifespan of each part, slip rate, paper type, basis weight, size, print mode, and print rate, and the likelihood of jamming changes depending on the actual paper transport state. Therefore, even if the technology in Patent Document 1 is used, it is not possible to determine control conditions for transporting paper so that jams do not occur.

[0007] The present invention has been made in consideration of the above-mentioned problems, and its main purpose is to provide a machine learning device, a machine learning method, and a machine learning program that can generate control information for a drive source for appropriately transporting an object. [Means for solving the problem]

[0008] One aspect of the present invention is a machine learning device that learns the behavior of a driving source in a conveying device that continuously and simultaneously conveys at least two conveyed objects along a conveying path, the machine learning device including: a state information acquisition unit that acquires position information of the at least two conveyed objects on the conveying path based on a detection result of a detection unit provided on the conveying path; a reward calculation unit that calculates a reward in accordance with a predetermined rule based on the acquired position information; a learning unit that learns behavior by calculating an action value in reinforcement learning based on the acquired position information and the calculated reward; and a control information output unit that generates and outputs control information for causing the driving source to perform an action determined based on a learning result, wherein the reward calculation unit calculates a state information acquisition unit that acquires position information of the at least two conveyed objects on the conveying path based on a detection result of a detection unit provided on the conveying path; Continuously transported The reward is calculated by comparing the distance between the two conveyances with a predetermined distance.

[0009] One aspect of the present invention is a machine learning method for a machine learning device that learns the behavior of a driving source in a conveying device that continuously and simultaneously conveys at least two conveyed objects along a conveying path, the machine learning method including: a state information acquisition process that acquires position information of the at least two conveyed objects on the conveying path based on a detection result of a detection unit provided on the conveying path; a reward calculation process that calculates a reward in accordance with a predetermined rule based on the acquired position information; a learning process that learns an action by calculating an action value in reinforcement learning based on the acquired position information and the calculated reward; and a control information output process that generates and outputs control information for causing the driving source to perform an action determined based on a learning result, the reward calculation process Continuously transported The reward is calculated by comparing the distance between the two conveyances with a predetermined distance.

[0010] One aspect of the present invention is a machine learning program that runs on a machine learning device that learns the behavior of a driving source in a conveyance device that continuously and simultaneously conveys at least two conveyed objects along a conveyance path, the program causing a control unit of the machine learning device to execute: a state information acquisition process that acquires position information of the at least two conveyed objects on the conveyance path based on a detection result of a detection unit provided on the conveyance path; a reward calculation process that calculates a reward in accordance with a predetermined rule based on the acquired position information; a learning process that learns an action by calculating an action value in reinforcement learning based on the acquired position information and the calculated reward; and a control information output process that generates and outputs control information for causing the driving source to perform an action determined based on a learning result, the reward calculation process Continuously transported The reward is calculated by comparing the distance between the two conveyances with a predetermined distance. [Effects of the Invention]

[0011] According to the machine learning device, the machine learning method, and the machine learning program of the present invention, it is possible to generate control information for a drive source for appropriately transporting an object.

[0012] The reason is that in a machine learning device that learns the behavior of a driving source in a conveying device that continuously conveys at least two conveying objects along a conveying path, position information of at least two conveying objects on the conveying path is obtained based on the detection results of a detection unit installed on the conveying path, a reward is calculated based on the obtained position information in accordance with predetermined rules, and an action value in reinforcement learning is calculated based on the obtained position information and the calculated reward, thereby learning the behavior and generating and outputting control information for causing the driving source to perform the action determined based on the learning results. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a schematic diagram illustrating a configuration of a control system according to an embodiment of the present invention. [Figure 2]FIG. 10 is a schematic diagram showing another configuration of a control system according to an embodiment of the present invention. [Figure 3] 1 is a block diagram showing a configuration of a machine learning device according to an embodiment of the present invention. [Figure 4] 1 is a block diagram showing a configuration of an image forming apparatus according to an embodiment of the present invention; [Figure 5] 2 is a schematic diagram illustrating sensors and drive sources in a paper transport path of an image forming apparatus according to an embodiment of the present invention. FIG. [Figure 6] 2 is a schematic diagram illustrating input / output parameters in a paper transport path of an image forming apparatus according to an embodiment of the present invention. FIG. [Figure 7] 10 is a table showing the relationship between states and actions in a paper transport path of an image forming apparatus according to an embodiment of the present invention. [Figure 8] FIG. 2 is a block diagram illustrating the general operation of a control system according to an embodiment of the present invention. [Figure 9] 10A and 10B are schematic diagrams illustrating another configuration of sensors and drive sources in the paper transport path of the image forming apparatus according to the embodiment of the present invention. [Figure 10] FIG. 2 is a flowchart illustrating the operation of the machine learning device according to an embodiment of the present invention. [Figure 11] FIG. 10 is a flowchart illustrating the operation (pitch-based reward calculation process) of the machine learning device according to one embodiment of the present invention. [Figure 12] FIG. 4 is a flowchart illustrating the operation (target pitch condition setting process) of the machine learning device according to one embodiment of the present invention. [Figure 13] FIG. 10 is a flowchart illustrating the operation (reward calculation process based on operation time) of the machine learning device according to one embodiment of the present invention. [Figure 14] FIG. 10 is a flowchart illustrating the operation (target movement completion condition setting process) of the machine learning device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0014] As explained in the background art, image forming devices such as MFPs vary in their operating environment and conditions depending on the user, and as a result, the state of the machine changes, which in turn affects the likelihood of jams occurring due to factors such as bending and pulling of paper during transport. Given this background, optimal control tailored to the machine state is required, but the number of possible combinations of operating environments and conditions is enormous, and designing control that takes all possible operating environments and conditions into account requires a significant amount of development time. Therefore, conventionally, control designs are designed to prevent jams from occurring under both the worst-case and typical conditions, but this method may not result in optimal control under unanticipated conditions.

[0015] Therefore, in one embodiment of the present invention, machine learning (particularly reinforcement learning) of AI (artificial intelligence) is used to learn the behavior of the drive source based on the actual state of the transported object, which varies depending on the user's usage environment and usage conditions (humidity, temperature, lifespan, slip rate, paper type, basis weight, size, print mode, print rate, etc.), thereby achieving optimal control of the drive source.

[0016] Specifically, in a machine learning device that learns the behavior of a driving source in a conveying device that continuously conveys at least two conveying objects along a conveying path, positional information of at least two conveying objects on the conveying path is acquired based on the detection results of a detection unit installed on the conveying path, a reward is calculated based on the acquired positional information in accordance with predetermined rules, and an action value in reinforcement learning is calculated based on the acquired positional information and the calculated reward, thereby learning the behavior and generating and outputting control information for causing the driving source to perform the action determined based on the learning results.

[0017] For example, in a system including a machine learning device and an image forming device, when paper transport begins, the machine learning device acquires paper position information, calculates a reward according to preset rules, learns behavior by calculating an action value in reinforcement learning based on the position information and the reward, and generates and outputs control information for causing a drive source to perform an action determined based on the learning results. The image forming device acquires the control information and controls the drive source by updating firmware each time or all at once.

[0018] In this way, by applying reinforcement learning to the transport control of transported items such as paper, and calculating the action value by giving appropriate rewards for the target behavior, it becomes possible to learn for various conditions, and therefore it is possible to automatically construct transport control for transported items that is suitable for the user's usage environment and usage conditions, thereby reducing unnecessary downtime. [Example]

[0019] To further explain the above-described embodiment of the present invention in detail, a machine learning device, a machine learning method, and a machine learning program according to an embodiment of the present invention will be described with reference to FIGS. 1 to 14. FIGS. 1 and 2 are schematic diagrams illustrating the configuration of a control system according to this embodiment, and FIGS. 3 and 4 are block diagrams illustrating the configurations of a machine learning device and an image forming apparatus according to this embodiment, respectively. FIGS. 5 and 9 are schematic diagrams illustrating sensors and drive sources in a paper transport path of the image forming apparatus according to this embodiment, FIG. 6 is a schematic diagram illustrating input / output parameters in the paper transport path, and FIG. 7 is a table illustrating the relationship between states and actions in the paper transport path. FIG. 8 is a block diagram illustrating the general operation of the control system according to this embodiment, and FIGS. 10 to 14 are flowcharts illustrating the operation of the machine learning device according to this embodiment.

[0020] As shown in FIG. 1, the control system 10 of this embodiment is composed of a machine learning device 20 and a conveyance device (referred to as an image forming device 30 in this embodiment) that continuously conveys at least two conveyed objects along a conveyance path, which are connected via a communication network such as a LAN (Local Area Network) or WAN (Wide Area Network) defined by standards such as Ethernet (registered trademark), Token Ring, or FDDI (Fiber-Distributed Data Interface). As shown in FIG. 2, the machine learning device 20 may be included in the image forming device 30 (the control unit of the image forming device 30 functions as the machine learning device). Each device will be described in detail below, assuming the system configuration of FIG. 1.

[0021] [Machine learning device] The machine learning device 20 is a computer device that provides cloud services and learns the control conditions of the drive source of the image forming device 30. As shown in FIG. 3(a), the machine learning device 20 is composed of a control unit 21, a storage unit 25, a network I / F unit 26, a display unit 27, an operation unit 28, and the like.

[0022] The control unit 21 is composed of a CPU (Central Processing Unit) 22 and memories such as a ROM (Read Only Memory) 23 and a RAM (Random Access Memory) 24. The CPU 22 controls the overall operation of the machine learning device 20 by loading a control program stored in the ROM 23 or storage unit 25 into the RAM 24 and executing it. As shown in FIG. 3(b), the control unit 21 functions as a state information acquisition unit 21a, a reward calculation unit 21b, a learning unit 21c, a control information output unit 21d, and the like.

[0023] The status information acquisition unit 21a acquires status information (position information) of at least two transported objects on the transport path based on the detection results of a detection unit (sensor) installed on the transport path. This position information may be acquired from the detection results of the detection unit installed on the transport path, calculated from the detection results of the detection unit and the transport speed of the transported objects, calculated from the elapsed time since the output of control information and the transport speed of the transported objects, or calculated from the elapsed time since the output of control information and the number of pulses in the control information. When calculating this position information, any one of humidity, temperature, lifespan, slip rate, paper type, basis weight, size, print mode, and print rate can be taken into consideration. In other words, when the position information is acquired from the detection results of the detection unit, the position information inherently includes the user's usage environment and usage conditions, such as humidity, temperature, lifespan, slip rate, paper type, basis weight, size, print mode, and print rate. When the position information is calculated using the transport speed of the transported objects or the elapsed time since the output of control information, the position information can include the user's usage environment and usage conditions.

[0024] The reward calculation unit 21b calculates the reward according to predetermined rules based on the position information acquired by the status information acquisition unit 21a. The reward calculation unit 21b can calculate the reward by comparing the time it takes for one of the at least two transported objects to reach a second position on the transport route from a first position with a predetermined time, or by comparing the distance between two of the at least two transported objects with a predetermined distance. In the latter case, the reward corresponding to the transported object in the first region of the transport route is calculated by comparing it with the first predetermined distance, and the reward corresponding to the transported object in the second region of the transport route is calculated by comparing it with the second predetermined distance. If the distance between two of the at least two transported objects is shorter than the predetermined distance, the reward can be set to a negative value. The reward calculation unit 21b can also calculate the reward based on the acquired position information and the transport speed of adjacent driving sources. Furthermore, if the acquired position information does not change for a certain period of time, the reward calculation unit 21b can set the reward to a negative value. The reward calculation unit 21b can also calculate the reward according to the stopping positions of the at least two transported objects.

[0025] The learning unit 21c learns an action (a control condition of a driving source) by calculating an action value in reinforcement learning (Q-learning) based on the state information acquired by the state information acquisition unit 21a and the reward calculated by the reward calculation unit 21b. In this case, in addition to the acquired state information and the calculated reward, the learning can also take into consideration any one of humidity, temperature, lifespan, and slip rate, or any one of paper type, basis weight, size, print mode, and print rate.

[0026] Control information output unit 21d generates control information (a control signal, a control current, a frequency, etc.) for causing the drive source to perform an action (an action with the highest action value) determined based on the learning result of learning unit 21c, and outputs the control information to image forming device 30. Furthermore, when learning unit 21c has performed learning taking into consideration any one of humidity, temperature, lifespan, and slip rate, control information output unit 21d can generate control information taking into consideration any one of paper type, basis weight, size, print mode, and print rate, and when learning unit 21c has performed learning taking into consideration any one of paper type, basis weight, size, print mode, and print rate, control information output unit 21d can generate control information taking into consideration any one of humidity, temperature, lifespan, and slip rate.

[0027] The above-mentioned state information acquisition unit 21a, reward calculation unit 21b, learning unit 21c, and control information output unit 21d may be configured as hardware, or the control unit 21 may be configured as a machine learning program that causes the control unit 21 to function as the state information acquisition unit 21a, reward calculation unit 21b, learning unit 21c, and control information output unit 21d, and the machine learning program may be executed by the CPU 22.

[0028] The memory unit 25 is composed of a hard disk drive (HDD) or a solid state drive (SSD), and stores programs for the CPU 22 to control each part, status information (sensor detection information and drive source drive information) acquired from the image forming device 30, location information acquired by the status information acquisition unit 21a, rules for calculating rewards, action values ​​and learning results (Q table described later) calculated by the learning unit 21c, control information generated by the control information output unit 21d, etc.

[0029] The network I / F unit 26 is configured by a NIC (Network Interface Card), a modem, and the like, and connects the machine learning device 20 to a communication network and establishes a connection with the image forming device 30.

[0030] The display unit 27 is configured by a liquid crystal display (LCD) or an organic electroluminescence (EL) display, and displays various screens.

[0031] The operation unit 28 is composed of a mouse, keyboard, etc., and enables various operations.

[0032] [Image forming device] The image forming apparatus 30 is an apparatus that continuously conveys at least two conveyed objects (paper sheets) along a conveyance path. As shown in Fig. 4(a), this image forming apparatus is composed of a control unit 31, a storage unit 35, a network I / F unit 36, a display operation unit 37, an image processing unit 38, an image reading unit 39, a print processing unit 40, etc.

[0033] The control unit 31 is composed of a CPU 32 and memories such as a ROM 33 and a RAM 34, and the CPU 32 controls the overall operation of the image forming apparatus 30 by loading a control program stored in the ROM 33 or the storage unit 35 into the RAM 34 and executing it. As shown in FIG. 4(b), the control unit 31 functions as a transport control unit 31a that controls the transport of paper, and the transport control unit 31a functions as a status notification unit 31b, an update processing unit 31c, etc.

[0034] The status notification unit 31b monitors the detection unit (sensor) and drive source (motor, clutch, etc.) provided on the paper transport path of the print processing unit 40, and notifies the machine learning device 20 of status information such as detection information from the detection unit and drive information from the drive source (e.g., motor rotation frequency, paper transport distance per motor rotation, and paper transport speed corresponding to the motor rotation frequency).

[0035] The update processing unit 31c acquires control information from the machine learning device 20, and updates firmware that controls the operation of the drive source (motor, clutch, etc.) based on the control information. In this case, the firmware may be updated every time control information is acquired from the machine learning device 20, or the firmware may be updated collectively after multiple pieces of control information are acquired.

[0036] The storage unit 35 is configured by an HDD, an SSD, or the like, and stores programs for the CPU 32 to control each unit, information about the processing functions of the device itself, device information, image data generated by the image processing unit 38, and the like.

[0037] The network I / F unit 36 ​​is configured with a NIC, a modem, and the like, and connects the image forming device 30 to a communication network to establish communication with the machine learning device 20 and the like.

[0038] The display operation unit (operation panel) 37 is a touch panel or the like having a pressure-sensitive or capacitance-type operation unit (touch sensor) with transparent electrodes arranged in a grid pattern on the display unit, and displays various screens related to the printing process and enables various operations related to the printing process.

[0039] The image processing unit 38 functions as a RIP (Raster Image Processor), translating the print job to generate intermediate data, and performing rendering to generate image data in bitmap format. The image processing unit 38 also performs screen processing, tone correction, density balance adjustment, line thinning, halftone processing, and the like on the image data as needed. The image processing unit 38 then outputs the generated image data to the print processing unit 40.

[0040] The image reading unit (ADU) 39 optically reads image data from a document placed on a document table, and is composed of a light source that scans the document, an image sensor such as a CCD (Charge Coupled Device) that converts light reflected by the document into an electrical signal, an A / D converter that A / D converts the electrical signal, etc. The image reading unit 39 then outputs the read image data to the print processing unit 40.

[0041] Print processing unit 40 executes printing processing based on image data acquired from image processing unit 38 or image reading unit 39. This print processing unit 40 is composed of, for example, an exposure unit that irradiates and exposes with laser light based on image data, an image forming unit that has a photosensitive drum, a developing unit, a charging unit, a photosensitive cleaning unit, and a primary transfer roller and forms toner images of each color of CMYK, an intermediate belt that is rotated by a roller and functions as an intermediate transfer body that transports the toner image formed in the image forming unit to paper, a secondary transfer roller that transfers the toner image formed on the intermediate belt to paper, a fixing unit that fixes the toner image transferred to paper, a paper feed unit such as a tray that feeds paper, transport units such as paper feed rollers, registration rollers, loop rollers, reversing rollers, and paper discharge rollers (collectively referred to as transport rollers), sensors that detect the transport position of paper provided in the transport path of the transport unit, and a drive source that drives the transport unit (a motor and a clutch that switches the transmission of power from the motor). The sensor may be any sensor capable of detecting the position of the paper being conveyed, such as one that detects based on the ON / OFF state of light or the contact of an electrical contact, etc. The drive source may be any sensor capable of supplying power to drive the conveyance roller, and there are no particular restrictions on the type of motor or clutch, the motor power transmission structure, etc.

[0042] 1 to 4 are an example of the control system 10 of this embodiment, and the configuration and control of each device can be changed as appropriate. For example, while FIG. 4 illustrates the image forming device 30 as a device that continuously conveys at least two conveyed objects along a conveyance path, the image forming device 30 may also be a post-processing device that performs post-processing such as stapling and folding, a sorting device that classifies paper, or an inspection device that inspects images formed on paper. Also, while FIG. 1 illustrates the control system 10 as being composed of the machine learning device 20 and the image forming device 30, the control system 10 may also include a computer device in a development department or a sales company. In this case, the computer device may receive individual requests from users of the image forming device 30 and notify the machine learning device 20, and the machine learning device 20 may then change the product specifications in accordance with the individual requests.

[0043] Next, sensors and drive sources in the paper transport path of the image forming apparatus 30 will be described. FIG. 5 is a schematic diagram showing a paper transport path 41 in the print processing unit 40, with paper transported from left to right in the figure. This paper transport path 41, for example, is provided with a plurality of sensors 42 (20 sensors arranged at positions 1 to 20 in the figure). Also, the paper transport path 41 is provided with a plurality of rollers (black circles in the figure) that transport paper. The roller drive sources include, for example, a main motor 43, a fixing motor 44, and a paper discharge motor 45. The main motor 43 is provided with a paper feed clutch 43a and a timing clutch 43b that turn on / off the transmission of power from the motor, and the paper discharge motor 45 is provided with a paper discharge clutch 45a that turns on / off the transmission of power from the motor. Note that FIG. 5 is an example of the paper transport path 41, and the number and arrangement of sensors 42, rollers, motors, and clutches can be changed as appropriate.

[0044] In a paper transport path 41 configured as described above, as shown in FIG. 6, status information such as the detection results of the sensor 42 becomes input parameters, and control information such as the control signals, control currents, and frequencies of the main motor 43 (paper feed clutch 43a, timing clutch 43b), fixing motor 44, and paper discharge motor 45 (paper discharge clutch 45a) becomes output parameters, and the machine learning device 20 learns the relationship between these input parameters and output parameters.

[0045] Figure 7 shows tables that the machine learning device 20 uses when learning the relationship between input parameters and output parameters, where (a) shows details of states (ON / OFF combinations of each sensor 42), (b) shows details of actions (here, ON / OFF combinations of each clutch), and (c) is a Q table showing action values ​​(Q values) corresponding to combinations of states and actions.

[0046] In this table, the number of sensors 42 is 14, and the number of states Ns at this time is Ns = number of sensor states ^ number of sensors = 2^14 = 16384 In this table, the target of the action is the clutch, and the number of clutches is 3. The number of actions Na is: Na = number of clutch states ^ number of clutches = 2^3 = 8 Therefore, the size of the Q table is Q[Ns,Na]=Q[16384,8].

[0047] The machine learning device 20 calculates the reward when a certain action is taken in a certain state according to predetermined rules, calculates the action value (Q value) according to a predetermined formula and updates the Q table to learn the action so as to optimize the total reward, and determines the action based on the learning results (selects the action with the highest action value).

[0048] The learning coefficient is α, the discounted reward is γ, and the reward at time t is r t Then, the action value (Q(s t,a t )) is, for example, Q(s t ,a t )←(1-α)Q(s t ,a t )+α(r t+1 +γmaxQ(s t+1 ,a t+1 )) It can be calculated using the Q-learning formula such as:

[0049] 8 is a block diagram showing an outline of the sheet transport control of the control system 10 of this embodiment. The transport control unit 31a (status notification unit 31b) of the image forming device 30 acquires state information, such as sensor detection information for each step (predetermined time) and drive information of the drive source (e.g., motor rotation frequency), from output signals of sensors 42 and motors provided on the sheet transport path 41, and outputs the state information to the machine learning device 20. The state information acquisition unit 21a of the machine learning device 20 acquires position information of each sheet as a state variable based on the state information and notifies the reward calculation unit 21b. The reward calculation unit 21b calculates a reward based on the position information and notifies the learning unit 21c. The learning unit 21c learns the behavior by calculating an action value based on the position information of each sheet acquired from the state information acquisition unit 21a and the reward acquired from the reward calculation unit 21b, and notifies the control information output unit 21d of the learning result (state variables, each action, action value). The control information output unit 21d generates control information such as a control signal, a control current, and a frequency for causing the drive source to perform the action determined based on the learning result, and notifies the image forming device 30. The transport control unit 31a (update processing unit 31c) of the image forming device 30 updates firmware for driving the drive source such as a motor or a clutch in accordance with the control information acquired from the machine learning device 20, and controls the operation of the motor or clutch in accordance with the firmware.

[0050] In Figure 5, 20 sensors 42 are arranged on the paper transport path 41, and the position of the paper is determined from the output signal (ON / OFF) of each sensor 42. However, if drive information of the drive source is used, the number of actual sensors 42 (black triangles in the figure) may be reduced, and virtual sensors 42 (dotted hatched triangles in the figure) may be arranged based on the output signals of the sensors 42, the drive signal of the motor, and physical parameters (such as the paper transport distance per motor rotation and the paper transport speed according to the motor rotation frequency).

[0051] The following describes the machine learning method used in the machine learning device 20 of this embodiment. The CPU 22 of the control unit 21 of the machine learning device 20 loads a machine learning program stored in the ROM 23 or the storage unit 25 into the RAM 24 and executes it, thereby performing the processing of each step shown in the flowcharts of Figures 10 to 14.

[0052] First, when the print processing unit 40 of the image forming device 30 starts conveying paper, the control unit 21 (status information acquisition unit 21a) of the machine learning device 20 acquires status information such as detection information from the sensor 42 and drive information from the drive source from the control unit 31 (status notification unit 31b) of the image forming device 30, and acquires position information of the paper based on the status information (S101). This position information may be acquired from the detection information from the sensor 42, or may be calculated from the detection information from the sensor 42 and the drive information from the drive source. Furthermore, when calculating the position information, any one of humidity, temperature, lifespan, slip rate, paper type, basis weight, size, print mode, and print rate may be taken into consideration.

[0053] Next, the control unit 21 (reward calculation unit 21b) calculates a reward based on the position information of the paper (S102). The calculation of this reward will be described in detail later. Next, the control unit 21 (learning unit 21c) learns the behavior by calculating an action value (Q value) using the Q-learning calculation formula described above based on the position information of the paper acquired by the state information acquisition unit 21a and the reward calculated by the reward calculation unit 21b (S103), and updates the Q table (S104). At this time, the learning unit 21c can perform learning taking into consideration any one of humidity, temperature, lifespan, and slip rate in addition to the position information of the paper and the reward, or any one of paper type, basis weight, size, print mode, and print rate.

[0054] The control unit 21 (control information output unit 21d) then determines the next action based on the learning result (Q table) (S105), generates control information (such as a control signal, a control current, and a frequency) for causing the drive source to perform the determined action, and outputs the control information to the image forming device 30 (S106). At this time, if the learning unit 21c has performed learning taking into consideration any one of humidity, temperature, lifespan, and slip rate, the control information output unit 21d can generate control information taking into consideration any one of paper type, basis weight, size, print mode, and print rate. If the learning unit 21c has performed learning taking into consideration any one of paper type, basis weight, size, print mode, and print rate, the control information output unit 21d can generate control information taking into consideration any one of humidity, temperature, lifespan, and slip rate. Then, upon acquiring the control information from the machine learning device 20, the control unit 31 (update processing unit 31c) of the image forming device 30 updates firmware that controls the operation of the drive source based on the control signal, and drives the drive source in accordance with the updated firmware to transport paper. Then, the process returns to S101 and repeats the same process.

[0055] Next, the reward calculation in S102 will be described. There are several methods for calculating the reward, such as a method for calculating the reward based on the pitch of the paper sheets (the distance or time interval between the paper sheets) and a method for calculating the reward based on the operation time.

[0056] FIG. 11 shows an example of a method for calculating a reward based on the pitch (time interval) of paper. The reward calculation unit 21b sets a target pitch condition (S201). FIG. 12 shows the details of this step. First, a target section is set (S301). Next, a section determination is performed (S302). For section A, the target pitch is set to a first value (50 ms in this case) (S303). For section B, the target pitch is set to a second value (200 ms in this case) (S304). For section C, the target pitch is set to a third value (400 ms in this case) (S305). Returning to FIG. 11, the reward calculation unit 21b measures the sensor passing time (the time from when one paper sheet passes the sensor to when the next paper sheet passes the sensor, i.e., the transport time interval between two paper sheets) (S202) and determines whether the actual measurement in S202 is larger or smaller than the target in S201 (S203). If the actual result and the target are approximately equal, the paper is being transported as set, so the reward is set to a positive predetermined value (e.g., +1) (S204). If the actual measurement is greater than the target, the paper is not being transported as set, but there is no risk of the paper sheets colliding with each other, so the reward is set to 0 (S205). If the actual measurement is smaller than the target, there is a risk of the paper sheets colliding with each other, so the reward is set to a negative predetermined value (e.g., -1) (S206).

[0057] 11 and 12, the reward is calculated based on the time interval between two sheets of paper, but the reward may also be calculated based on the distance between two sheets of paper. In this case, if the distance between two sheets of paper is shorter than the target distance, the reward can be set to a negative predetermined value because there is a risk of the sheets colliding with each other.

[0058] FIG. 13 shows an example of a method for calculating a reward based on an operation time. The reward calculation unit 21b sets a target movement completion condition (S401). FIG. 14 shows the details of this step. First, the movement speed and total movement distance are obtained from changes in the position information of the paper, and the stop time during movement is calculated (S501). Then, from this information, the target movement completion condition is set to a predetermined value (600 ms in this case) (S502). Returning to FIG. 13, the reward calculation unit 21b measures the time from the start to the end of the operation (S402) and determines whether the actual measurement is larger or smaller than the target (S403). If the actual result and the target are approximately equal, the paper is being transported as set, so the reward is set to a positive predetermined value (e.g., 1) (S404). If the actual result is greater than the target, the paper is not being transported as set, but there is no risk of the paper sheets colliding with each other, so the reward is set to 0 (S405). If the actual result is less than the target, there is a risk of the paper sheets colliding with each other, so the reward is set to a negative predetermined value (e.g., -1) (S406).

[0059] Note that Figures 11 and 12 describe a method for calculating rewards based on paper pitch, and Figures 13 and 14 describe a method for calculating rewards based on operation time, but rewards may be calculated based on both paper pitch and operation time, or other parameters may be added to paper pitch and / or operation time to calculate rewards.

[0060] Furthermore, if the acquired position information does not change for a certain period of time, it is considered that a jam has occurred, so the reward may be set to a negative value, or the reward may be calculated according to the stop positions of at least two conveyed objects (according to whether or not they stopped at a predetermined stop position). Furthermore, depending on the drive state of the drive source (for example, if the conveying speeds of adjacent drive sources are different), bending or pulling of the paper may occur, so the reward may be calculated taking such defects into consideration.

[0061] As described above, by acquiring the position information of the paper, calculating the reward according to preset rules, learning the behavior by calculating the action value in reinforcement learning based on the position information and the reward, and outputting control information for causing the drive source to perform the action determined based on the learning results, it is possible to realize transport control of the transported item that is suitable for the user's usage environment and usage situation.

[0062] The present invention is not limited to the above-described embodiment, and the configuration and control can be modified as appropriate without departing from the spirit of the present invention.

[0063] For example, in the above embodiment, the machine learning method of the present invention was described as being applied to an image forming device that controls the transport of multiple sheets of paper and performs processing, but the machine learning method of the present invention can be similarly applied to any device that controls the transport of multiple moving objects and performs processing. [Industrial Applicability]

[0064] The present invention can be used in a machine learning device, a machine learning method, a machine learning program, and a recording medium on which the machine learning program is recorded, which learn the behavior of a drive source in a conveying device that controls the conveyance of multiple moving objects. [Explanation of symbols]

[0065] 10. Control System 20 Machine Learning Device 21 Control Unit 21a Status information acquisition unit 21b Remuneration Calculation Department 21c Learning Department 21d Control information output unit 22 CPU 23 ROM 24 RAM 25 Memory section 26 Network I / F section 27 Display section 28 Control section 30 Image forming device 31 Control Unit 31a Transport control section 31b Status notification section 31c Update processing section 32 CPU 33 ROM 34 RAM 35 Storage section 36 Network I / F section 37 Display operation section 38 Image processing section 39 Image reading unit 40 Print processing unit 41 Paper transport path 42 sensors 43 Main motor 43a Paper feed clutch 43b Timing clutch 44 Fuser motor 45 Paper ejection motor 45a Paper ejection clutch

Claims

1. A machine learning device that learns the behavior of a driving source in a conveying device that continuously and simultaneously conveys at least two conveyed objects along a conveying path, comprising: a status information acquiring unit that acquires position information of the at least two transported objects on the transport path based on a detection result of a detection unit provided on the transport path; a reward calculation unit that calculates a reward in accordance with a predetermined rule based on the acquired location information; a learning unit that learns an action by calculating an action value in reinforcement learning based on the acquired position information and the calculated reward; a control information output unit that generates and outputs control information for causing the driving source to perform an action determined based on a learning result, the remuneration calculation unit calculates the remuneration by comparing a distance between two consecutively transported objects among the at least two objects with a predetermined distance; A machine learning device characterized by:

2. the state information acquisition unit acquires the position information from the detection result of the detection unit, acquires the position information by calculating it from the detection result of the detection unit and the movement speed of the transported object, acquires the position information by calculating it from the elapsed time since the output of the control information and the movement speed of the transported object, or acquires the position information by calculating it from the elapsed time since the output of the control information and the number of pulses of the control information. The machine learning device according to claim 1 .

3. the remuneration calculation unit calculates a remuneration corresponding to the transported object in a first region of the transport route by comparing it with a first predetermined distance, and calculates a remuneration corresponding to the transported object in a second region of the transport route by comparing it with a second predetermined distance; The machine learning device according to claim 1 .

4. the remuneration calculation unit sets the remuneration to a negative value when a distance between two consecutively transported articles among the at least two articles is shorter than the predetermined distance; 4. The machine learning device according to claim 1, wherein the machine learning device is a computer.

5. The reward calculation unit further calculates a negative value for the reward when the transport speeds of the adjacent driving sources are different.

5. The machine learning device according to claim 1, wherein the machine learning device is a computer.

6. The reward calculation unit further sets the reward to a negative value when the acquired location information does not change for a certain period of time.

6. The machine learning device according to claim 1, wherein the machine learning device is a computer.

7. The reward calculation unit further calculates a negative value for the reward when the at least two conveyed objects stop at a predetermined stop position.

7. The machine learning device according to claim 1, wherein the machine learning device is a computer.

8. the drive source is a motor or a clutch that switches transmission of power from the motor, The control information is a control signal, a control current, and a frequency for operating the motor and / or the clutch. The machine learning device according to any one of claims 1 to 7.

9. the conveying device is an image forming device that conveys paper and prints on it; The machine learning device according to any one of claims 1 to 8.

10. the state information acquisition unit also takes into consideration any one of humidity, temperature, lifespan, slip rate, paper type, basis weight, size, print mode, and print rate when calculating the position information; The machine learning device according to claim 9 .

11. the learning unit performs learning taking into consideration any one of humidity, temperature, lifespan, and slip rate in addition to the acquired position information and the calculated reward; the control information output unit generates the control information taking into consideration any one of paper type, basis weight, size, print mode, and print rate. The machine learning device according to claim 9 or 10.

12. the learning unit performs learning taking into consideration any one of paper type, basis weight, size, print mode, and print rate in addition to the acquired position information and the calculated reward; the control information output unit generates the control information taking into consideration any one of humidity, temperature, lifespan, and slip rate. The machine learning device according to any one of claims 9 to 11.

13. A transport device comprising the machine learning device according to any one of claims 1 to 12.

14. A machine learning method for a machine learning device that learns behavior of a driving source in a conveying device that continuously and simultaneously conveys at least two conveyed objects along a conveying path, comprising: a state information acquisition process for acquiring position information of the at least two transported objects on the transport path based on a detection result of a detection unit provided on the transport path; a reward calculation process for calculating a reward in accordance with a predetermined rule based on the acquired location information; a learning process for learning an action by calculating an action value in reinforcement learning based on the acquired position information and the calculated reward; a control information output process for generating and outputting control information for causing the driving source to perform an action determined based on the learning result; the reward calculation process calculates the reward by comparing a distance between two consecutively transported objects among the at least two objects with a predetermined distance; A machine learning method characterized by:

15. A machine learning program that runs on a machine learning device that learns behavior of a drive source in a conveyance device that continuously and simultaneously conveys at least two conveyed objects along a conveyance path, the program comprising: a control unit of the machine learning device, a state information acquisition process for acquiring position information of the at least two transported objects on the transport path based on a detection result of a detection unit provided on the transport path; a reward calculation process for calculating a reward in accordance with predetermined rules based on the acquired location information; a learning process for learning an action by calculating an action value in reinforcement learning based on the acquired position information and the calculated reward; a control information output process for generating and outputting control information for causing the driving source to perform an action determined based on the learning result; the reward calculation process calculates the reward by comparing a distance between two consecutively transported objects among the at least two objects with a predetermined distance; A machine learning program characterized by:

Citation Information

Patent Citations

  • Paper feed device

    JP1993201564A

  • Image forming apparatus

    JP2006072101A

  • Image forming apparatus

    JP2007079011A

  • Image forming apparatus

    JP2013238682A

  • Machine learning device learning gain optimization, motor control device having the same, and the machine learning method

    JP2017034844A