POSITION CONTROL DEVICE AND POSITION CONTROL METHOD
The position control device uses a monocular camera and neural networks to correct positional errors in robot arms, improving manufacturing efficiency by enabling precise alignment and adaptability.
Patent Information
- Application Number
- DE112017007025
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2017-02-09
- Publication Date
- 2025-12-04
- Estimated Expiration
- 2037-02-09
AI Technical Summary
Existing position control systems for robot arms rely on manual teaching and struggle to correct positional errors using monocular cameras, leading to inefficiencies and limited adaptability in manufacturing environments.
A position control device utilizing a monocular camera and neural networks to calculate control amounts for robot arm movements, enabling precise alignment of objects by capturing images and adjusting positions based on learned relationships between images and control values.
Enables precise alignment of objects using a monocular camera, accommodating individual positional errors and deviations, enhancing adaptability and productivity in manufacturing processes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
field of technology
[0001] The present invention relates to a position control device and a position control method. State of the art
[0002] Generally, setting up a production system where a robot arm performs assembly work requires a manual instruction process called teaching, which is carried out by an operator. In this case, the robot only repeats movements for the position stored during the teaching process, meaning the robot may not be able to react correctly if manufacturing or assembly errors occur. Therefore, developing a position correction technique that allows for individual positional errors promises to enable robots to be used in even more areas and lead to productivity improvements.
[0003] A conventional technique allows for position corrections using a camera image during processes up to the insertion of the fastener (Patent 1). Furthermore, when multiple devices such as a force sensor and a stereo camera are used, it is possible to allow for positional errors that affect production assembly (insertion, clamping, etc.). However, to determine the amount of position correction, values such as the coordinates of the center of the fastener being gripped and the coordinates of the center of the fastener being inserted must be calculated from the image information, as disclosed in Patent 1. Since this calculation depends on the shapes of the fasteners, a designer must define the parameters for each fastener to be used.If the three-dimensional information is obtained from a device such as a distance camera, the calculation is relatively simple. However, if the three-dimensional information has to be extracted from two-dimensional image data, image processing algorithms must be developed for each connecting element, resulting in high development costs.
[0004] Further methods for controlling robots are described, for example, in the publications JP 2009 - 83 095 A and "View-based teaching / playback for manipulation by industrial robots," MORIYAMA, Yuki; MAEDA, Yusuke, Transactions of the Japan Society of Mechanical Engineers. C, vol. 79, 2013, no. 806, pp. 3597-3608. - ISSN 1884-8354. Neural networks are described, for example, in the publication "Intelligent control using neural networks," NARENDRA, Kumpati S.; MUKHOPADHYAY, Snehasis, IEEE Control Systems Magazine, vol. 12, 1992, no. 2, pp. 11-18. - ISSN 1066-033X. List of cited patent specifications
[0005] Patent specification 1: WO 98 / 17444A1 Summary of the invention; Technical task
[0006] It is difficult to control the mounting position using only information from a monocular camera.
[0007] The present revelation is intended to solve the problem and enable alignment using only a monocular camera. Technical solution
[0008] A position control device according to the disclosure serves to install a first object in a second object by causing a drive unit to move the first object. The device includes an imaging unit of a monocular camera for capturing an image containing two objects, the first object and the second object, and a control parameter generation unit. Among a variety of neural networks, each with an input layer into which information from a variety of images captured by the imaging unit is fed, and an output layer that outputs control amounts obtained based on a learning rule such as a stochastic gradient method, where each of the control amounts for moving the first object corresponds to each of the variety of images fed into the input layer, the control parameter generation unit is activated when the imaging unit captures an image for the first time.The control parameter generation unit selects a predetermined neural network and outputs the control amount from the output layer of the selected neural network as the control amount to cause the drive unit to move the first of the two objects contained in the image captured by the imaging unit. The control parameter generation unit captures an image captured by the imaging unit a second time or later, inputs information from the captured image into an input layer of a neural network selected from the multitude of neural networks based on a control amount output from an output layer of a neural network whose input layer received information from an image captured the previous time.and outputs the control amount output in an output layer of the selected neural network as the control amount to cause the drive unit to move the first of the two objects, which are the first object and the second object, contained in the image captured by the imaging unit. Advantageous effects of the invention
[0009] The position control device according to the disclosure serves to install a first object inside a second object by causing a drive unit to move the first object. The device includes an imaging unit of a monocular camera for capturing an image containing the two objects, namely the first object and the second object, and a control parameter generation unit. Among a plurality of neural networks, each of which has an input layer into which information from a plurality of images captured by the imaging unit is inputted and an output layer in which control amounts obtained based on a learning rule such as a stochastic gradient method are output, with each of the control amounts for moving the first object corresponding to each of the plurality of images input to the input layer, the control parameter generation unit...When the imaging unit first acquires an image, it selects a predetermined neural network and outputs the control amount from the output layer of the selected neural network as the control amount to cause the drive unit to move the first of the two objects contained in the image acquired by the imaging unit. The control parameter generation unit acquires an image acquired by the imaging unit a second time or later, inputs information from the acquired image into an input layer of a neural network selected from a plurality of neural networks based on a control amount output from an output layer of a neural network whose input layer received information from an image acquired the previous time.and outputs the control amount in an output layer of the selected neural network as the control amount to cause the drive unit to move the first of the two objects, which are the first and second objects contained in the image captured by the imaging unit. Consequently, the present disclosure enables alignment using only a monocular camera. Brief description of the drawings Fig. Figure 1 is an illustration showing a robot arm 100 according to embodiment 1, plugs 110 and sockets 120. Fig. Figure 2 is a functional configuration diagram showing the position control device according to embodiment 1. Fig. Figure 3 is a hardware structure diagram showing the position control device according to embodiment 1. Fig. Figure 4 is a flowchart showing the position control carried out by the position control device according to embodiment 1. Fig. Figure 5 is an exemplary scheme showing the camera images and the control amounts; the images are captured by the monocular camera 102 according to embodiment 1 at the initial insertion position and in its respective vicinity. Fig. Figure 6 is an example showing a neural network and a learning rule of the neural network according to embodiment 1. Fig. Figure 7 is a flowchart showing the neural network according to embodiment 1, in which a plurality of networks are used. Fig. Figure 8 is a functional configuration diagram of a position control device according to embodiment 2. Fig. Figure 9 is a hardware structure diagram of the position control device according to embodiment 2. Fig. Figure 10 shows a plug 110 and a socket 120 in an attempt to connect them according to embodiment 2. Fig. Figure 11 is a flowchart showing the path learning of the position control device according to embodiment 2. Fig. Figure 12 is a flowchart showing the path learning of a position control device according to embodiment 3. Fig. Figure 13 is an example showing a neural network and a learning rule of the neural network according to embodiment 3. Description of the embodiments: Embodiment 1
[0010] The following describes embodiments of the present invention.
[0011] In embodiment 1, a robot arm that learns the insertion positions of plugs and works in a production line, as well as a position control method for it, are described.
[0012] Configurations are described. Fig. Figure 1 shows a robot arm 100, connectors 110, and sockets 120 arranged according to embodiment 1. The robot arm 100 is equipped with a gripping unit 101 for gripping a connector 110, and a monocular camera 102 is attached to the robot arm 100 at a position where the gripping unit is within the monocular camera's field of view. While the gripping unit 101 grasps the connector 110 at the end of the robot arm 100, the monocular camera 102 is positioned such that the end of the gripped connector 110 and the socket 120 into which it is to be inserted are within its field of view.
[0013] Fig. Figure 2 is a functional configuration diagram showing the position control device according to embodiment 1.
[0014] As in Fig. As shown in Figure 2, the position control device includes the following: an imaging unit 201 for capturing images, wherein the imaging unit 201 performs a function of the in Fig. The monocular camera 102 shown in Figure 1 comprises: a control parameter generation unit 202 for generating a position control amount for the robot arm 100 using the captured images; a control unit 203 for controlling the current and voltage values to be provided to the drive unit 204 of the robot arm 100 using the position control amount; and a drive unit 204 for changing the position of the robot arm 100 based on the current and voltage values output by the control unit 203.
[0015] After receiving an image captured by the imaging unit 201, which is a function of the monocular camera 102, the control parameter generation unit 202 determines the control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) that corresponds to the position of the robot arm 100 (X, Y, Z, Ax, Ay, Az) and outputs the control amount to the control unit 203. Here, X, Y, and Z are the coordinates of the position of the robot arm 100, and Ax, Ay, and Az are the position angles of the robot arm 100.
[0016] The control unit 203 determines and controls the current or voltage values 204 for the devices that form the drive unit 204, based on the received control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) which corresponds to the position (X, Y, Z, Ax, Ay, Az) of the robot arm 100.
[0017] Each device that forms the drive unit 204 operates according to the current or voltage values received from the control unit 203 and causes the robot arm 100 to move into the position (X+ΔX, Y+ΔY, Z+ΔZ, Ax+ΔAx, Ay+ΔAy, Az+ΔAz).
[0018] Fig. Figure 3 is a hardware structure diagram of the position control device according to embodiment 1. The monocular camera 102 is connected to a processor 302 and a memory element 303 via an input / output interface 301, either wired or wirelessly. The input / output interface 301, the processor 302, and the memory element 303 implement the functions of the control parameter generation unit 202 in Fig. 2. The input / output interface 301 is communicatively connected, either wired or wirelessly, to a control circuit 304 corresponding to the control unit 203. The control circuit 304 is also electrically connected to a motor 305. The motor 305, which is connected to the drive unit 204 in Fig. 2 corresponds to a component for any device for performing position control. In the present embodiment, the motor 305 is used in an embodiment of the hardware corresponding to the drive unit 204; any hardware designed to provide the functionality of such position control would suffice. The monocular camera 102 and the input / output interface 301 can be separate bodies, and the input / output interface 301 and the control circuit 304 can be separate bodies.
[0019] The process will be described next.
[0020] Fig. Figure 4 is a flowchart showing the position control carried out by the position control device according to embodiment 1.
[0021] First, the gripping unit 101 of the robot arm 100 grasps a plug 110 during step S101. The position and orientation of the plug 110 are described in the Fig. The control unit 203 shown is pre-registered, and the process is carried out according to the control program also pre-registered in the control unit 203.
[0022] Next, in step S102, the robot arm 100 is moved closer to the insertion position of a socket 120. The approximate position and orientation of the socket 120 are shown in the Fig. The control unit 203 shown is pre-registered, and the position of the plug 110 is controlled according to the control program pre-registered in the control unit 203.
[0023] Next, at step S103, the control parameter generation unit 202 instructs the imaging unit 201 of the monocular camera 102 to take an image, and the monocular camera 102 takes an image that includes both the plug 110 gripped by the gripping unit 101 and the socket 120, which is the insertion target part.
[0024] Next, at step S104, the control parameter generation unit 202 receives the image from the imaging unit 201 and determines the control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz). In determining the control amount, the control parameter generation unit 202 uses the processor 302 and the memory element 303, which are located in Fig. Figure 3 shows the hardware and calculates the control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) using a neural network. The procedure for calculating the control amount using the neural network will be described later.
[0025] Next, at step S105, the control unit 203 receives the control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) output by the control parameter generation unit 202 and compares all components of the control amount with their respective previously determined limit values. If all components of the control amount are at most equal to their respective limit values, the process proceeds to step S107, and the control unit 203 controls the drive unit 204 such that the plug 110 is inserted into the socket 120.
[0026] If any of the components of the control amount is greater than its corresponding limit value, the control unit 203 controls the drive unit 204 at step S106 using the control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) output by the control parameter generation unit 202, and the process returns to step S103.
[0027] Next, this will be done at step S104 of Fig. Four methods for calculating the control amount using a neural network are described.
[0028] Before the calculation of the control amount using the neural network begins, a large number of image data sets and the required amount for a movement are acquired. This prepares the neural network to calculate the movement required until successful connection using the input image. For example, plug 110 and socket 120, whose positions are known, are plugged together, and plug 110 is grasped by the gripper 101 of the robot arm 100. The gripper 101 then moves in the known direction and pulls the plug out to the initial insertion position, and the monocular camera 102 captures a large number of images.Furthermore, if (0, 0, 0, 0, 0, 0) is specified as the control amount for the initial insertion position, not only the amount for a movement from the assembly state position to the initial insertion position with its image are captured, but also the amounts for a movement to positions near the initial insertion position or the control amounts (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) with their images.
[0029] Fig. Figure 5 is an exemplary scheme showing control amounts and their respective images, which are captured by the monocular camera 102 in embodiment 1 at the initial introduction position and in their respective vicinity.
[0030] The learning process then takes place based on a general learning rule of the neural network (such as a stochastic gradient method) using the multitude of data sets, each consisting of the amount for a movement from the assembly position to the initial insertion position, and the image taken by the monocular camera 102 at or near the initial insertion position.
[0031] There are different types of neural networks, such as a CNN or an RNN, and any of these types can be used for the present disclosure.
[0032] Fig. Figure 6 is a scheme showing an example of the neural network and the neural network learning rule according to embodiment 1.
[0033] The images received from the monocular camera 102 (such as the brightness and color difference of each pixel) are fed to the input layer, and the control values (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) are output in the output layer.
[0034] During the learning process of the neural network, the parameters of the hidden layer are optimized so that the output values of the output layer, obtained from the input images via the hidden layer, approximate the control values stored in connection with their respective images. One such approximation method is the stochastic gradient method.
[0035] Consequently, more accurate learning can be achieved, as in Fig. 5 shown, insofar as those images are obtained that correspond not only to the amount for a movement from the assembly position to the initial introduction position, but also to the amounts for movements to positions around the initial introduction position.
[0036] In Fig. Figure 5 shows a case in which the position of plug 110 relative to monocular camera 102 is fixed, and only the position of socket 120 is changed. However, the gripper unit 101 of the robot arm 100 does not always grasp plug 110 precisely at the intended position, and the position of plug 110 can deviate from its normal position due to individual discrepancies, etc. During the learning process, in a state where plug 110 has deviated from its exact position, data records of a control value and an image at the initial insertion position, as well as the positions in their respective vicinity, are acquired for the purpose of carrying out the learning process; thus, the learning process is carried out in such a way that individual discrepancies of both plug 110 and socket 120 can be taken into account.
[0037] It should be noted that the control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) is calculated at the time of recording in such a way that the amount for a movement from the assembly position to the initial insertion position is excluded. Therefore, the amount for a movement from the assembly position to the initial insertion position must be included for use in step S107 in Fig. 4. It should also be noted that the coordinates mentioned above are obtained with reference to the coordinate system of the monocular camera. If the coordinate system of the monocular camera does not match the coordinate system of the entire robot arm 100, the control unit 203 must therefore perform a coordinate conversion before controlling the robot arm 100.
[0038] In this embodiment, the coordinate system of the socket 120 is different from the coordinate system of the monocular camera 102 because the monocular camera is attached to the robot arm 100. If the monocular camera 102 and the socket 120 are in the same coordinate system, the conversion from the coordinate system for the monocular camera 102 to the coordinate system for the robot arm 100 is not necessary.
[0039] Next, the process and a sample process, which is shown in Fig. 4 is shown and described in detail.
[0040] In step S101, the robot arm 100 grasps a plug 110 in a pre-registered manner, and in step S102, the plug 110 is moved to a position located almost above the socket 120.
[0041] It should be noted that the position of the gripped plug 110 is not always the same immediately before it is gripped. If there is a slight deviation in the machine sequence that determines the position of plug 110, there is always a possibility that a minor positioning error may have occurred. For the same reason, socket 120 may also exhibit some defects.
[0042] Therefore, it is important that the images captured in step S103 show both plug 110 and socket 120, with the images as shown in Fig. 5 are captured by the imaging unit 201 of the monocular camera 102, which is attached to the robot arm 100. The position of the monocular camera 102 relative to the robot arm 100 is always fixed, so that the images contain information about the positional relationship between the plug 110 and the socket 120.
[0043] In step S104, the control amount (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) is generated by the control parameter generation unit 202 using the neural network, as shown in Fig. Figure 6 shows that the robot arm 100 is calculated after having previously learned the information about the position relationship. However, depending on the accuracy of the learning process, movement to the initial starting position using the control amount output by the control parameter generation unit 202 is impossible. In this case, the control parameter generation unit 202 can repeatedly perform calculations by executing the loop from step S103 to step S106 until all components of the control amount are at most equal to their respective limit values, as shown in step S105. The control unit 203 and the drive unit 204 then control the position of the robot arm 100.
[0044] The limit values shown at S 105 are determined based on the required accuracy of the plug 110 and the socket 120, which are to be mated together. If, for example, they are loosely mated, meaning that high accuracy is not initially required as a characteristic of the plug, large limit values can be set. Conversely, small limit values are set in the opposite case. In a normal manufacturing process, manufacturing tolerances, which are often predefined, can be used for the limit values.
[0045] If, depending on the accuracy of the learning process, movement to the insertion start position is impossible using the control amount output by the control parameter generation unit 202, a multiple insertion start positions can also be defined. If an insertion start position is defined without sufficient clearance between plug 110 and socket 120, they may collide before insertion and break, either individually or both. To avoid this risk, the insertion start position can be defined incrementally, depending on how often steps S103 to S106, described in Fig. The 4 steps shown are traversed. For example, a distance of 5 mm between plug 110 and socket 120 can be set for the first time, a distance of 20 mm for the second time, a distance of 10 mm for the third time, and so on.
[0046] Although the present embodiment is described using plug connections, the application of the technique is not limited to the assembly of connecting elements. This method is practically applicable, for example, in the mounting of IC chips on a substrate and, in particular, in the insertion of capacitors or the like with leads exhibiting a large dimensional error into holes in a substrate.
[0047] This method is applicable not only to the control of insertion into substrates, but also to position control in general, in order to derive a control value from a known relationship between images and control values. In the present disclosure, the neural network learns the relationships between images and control values; consequently, individual differences in the objects can be permitted during the execution of the object alignment.
[0048] In embodiment 1, an imaging unit 201, a control parameter generation unit 202, a control unit 203, and a drive unit 204 are provided. The imaging unit 201 captures an image containing two objects. The control parameter generation unit 202 feeds information from the captured image containing the two objects to an input layer of a neural network and outputs a position control value to the output layer of the neural network to control the position relationship between the two captured objects. The control unit 203 controls the current or voltage to control the position relationship between the two objects using the output position control value. The drive unit 204 changes the position of one of the two objects using the current or voltage.Even if there are differences between objects or errors in the positional relationship between the two objects, alignment can therefore only be performed with a monocular camera.
[0049] The above describes an embodiment in which only one neural network is used. In some cases, however, a multitude of neural networks must be used. This is because, as in the embodiment, if the inputs are images and the outputs are numerical values, the outputs can contain errors of approximately a few percent; this is due to the limited approximation accuracy among the numerical values. The determination result can depend on the distance from the initial input position to the position near it at step S102. Fig. 4 always answer "no", and therefore the process cannot occur at step S105. A variety of networks are used to handle such a case, as in Fig. 7 shown.
[0050] Fig. Figure 7 is a flowchart for using a multitude of neural networks, as opposed to using only the single neural network shown in embodiment 1 above. Step S104 is used in this process. Fig. 4 shown in detail. The multitude of parameters is shown in the Fig. The control parameter generation unit shown in section 2 is included.
[0051] In step S701, the control parameter generation unit 202 selects the network to be used based on the entered image.
[0052] If the number of loops is one or the previously received control amount is at least 25 mm, neural network 1 is selected, and the process proceeds to step S702. If the number of loops is at least two and the previously received control amount is at least 5 mm but less than 25 mm, neural network 2 is selected, and the process proceeds to step S703. If the number of loops is at least two and the previously received control amount is less than 5 mm, neural network 3 is selected, and the process proceeds to step S704. The neural network selected from S702 to S704 is used to calculate the control amount for each step.
[0053] Each neural network has learned based on the distance between plug 110 and socket 120 or the control amount. In the figure, for example, the training data ranges are changed incrementally; neural network 3 uses training data with errors in a range of ±1 mm and ±1 degree, and neural network 2 uses training data with errors in a range of ±1 mm to ±10 mm and ±1 degree to ±5 degrees. It is more efficient if the image areas used in each neural network do not overlap.
[0054] The in Fig. Example 7 shows three networks, but the number of networks is not limited. If such a method is used, a determination function must be prepared as a "network selection switch" in step S701 to determine which network should be used. This network selection switch can also be configured using a neural network. In this case, the input to the input layer is an image, and the output to the output layer is a network number. The image data set consists of an image used in all networks and a network number.
[0055] An example of using a variety of neural networks with connectors is described. However, the application of this technique is not limited to simply connecting components. This method is practically applicable, for example, in the assembly of IC chips onto a substrate, and especially in the insertion of capacitors or similar components with leads that have a large dimensional error into holes in a substrate.
[0056] This method, which utilizes a variety of neural networks, is applicable not only to the control of insertion into substrates but also to position control in general, in order to derive a control amount from a known relationship between images and control amounts. In the present disclosure, the neural network learns the relationships between images and control amounts; consequently, individual differences in the objects can be permitted during the execution of the object alignment, and the control amounts are precisely calculable.
[0057] In the example above, an imaging unit 201, a control parameter generation unit 202, a control unit 203, and a drive unit 204 are provided. The imaging unit 201 captures an image containing two objects. The control parameter generation unit 202 feeds information from the captured image containing the two objects to an input layer of a neural network and outputs a position control value to the neural network's output layer to control the position relationship between the two captured objects. The control unit 203 controls the current or voltage to control the position relationship between the two objects using the output position control value. The drive unit 204 changes the position of one of the two objects using the current or voltage. The control parameter generation unit 202 selects the neural network from a variety of neural networks.Even if there are differences between objects or errors in the positional relationship between the two objects, a more precise alignment can therefore be achieved. Design 2
[0058] In embodiment 1, the plug 110 and the socket 120, whose positions are known, are connected, and the plug 110 is grasped by the gripping unit 101 of the robot arm 100. The gripping unit 101 then moves in the known direction and pulls the plug out to the initial insertion position, and the monocular camera 102 captures a multitude of images. In embodiment 2, a case is described in which the connection position of the plug 110 and the socket 120 is unknown.
[0059] A method called reinforcement learning has already been investigated, in which a robot learns to independently acquire its own actions. In this method, a robot performs various actions using a trial-and-error approach, stores the action that leads to a better result, and ultimately obtains the optimized action. Unfortunately, optimizing the action requires a very large number of attempts.
[0060] Methods for reducing the number of attempts include a framework called "On Policy," which is commonly used in reinforcement learning. However, applying this framework to teaching a robot arm requires several improvements specifically related to the robot arm and its control signals, and it has not yet been implemented in practice.
[0061] In a configuration that will be described later for embodiment 2, the robot, as shown in embodiment 1, performs various actions using the trial-and-error method, and the action that leads to a good result is stored, thus reducing the number of trials currently required to optimize the action.
[0062] Next, the system configuration is described. What is not described here is the same as in Implementation 1. The entire hardware structure is exactly the same as in Implementation 1. Fig. 1 in embodiment 1, however, it differs in that the robot arm 100 is equipped with a force sensor 801 (which is in Fig. (not shown) is equipped to measure the load applied to the gripping unit 101.
[0063] Fig. Figure 8 is a functional configuration diagram showing a position control device according to embodiment 2. The difference to Fig. 2 consists in the additional provision of a force sensor 801 and a path determination unit 802, wherein the path determination unit 802 contains a critical unit 803, an actor unit 804, an evaluation unit 805 and a path determination unit 806.
[0064] Fig. Figure 9 is a hardware structure diagram showing the position control device according to embodiment 2. The only difference to Fig. 3 consists in the fact that the force sensor 801 is electrically and / or communicatively connected to an input / output interface 301. Furthermore, the input / output interface 301, the processor 302, and the memory element 303 perform the functions of the control parameter generation unit 202 and the path determination unit 802, which are described in Fig. Figure 8 shows the following: The force sensor 801, the monocular camera 102 and the input / output interface 301 can be separate bodies, and the input / output interface 301 and the control circuit 304 can be separate bodies.
[0065] Next, Fig. 8 described in detail.
[0066] The force sensor 801 measures the load applied to the gripping unit 101 of the robot arm 100; for example, it measures the value of the force exerted when the plug 110 and the socket 120, which are in Fig. 1 are shown, touching each other.
[0067] The Critic unit 803 and the Actor unit 80 correspond to the Critic and the Actor in traditional Reinforcement Learning.
[0068] Next, the traditional reinforcement learning method will be described.
[0069] In the present embodiment, a model called the Actor-Critic model is used in reinforcement learning (source: Reinforcement Learning: RS Sutton and AG Barto, published in Dec. 2000). The Actor Unit 804 and the Critic Unit 803 receive an environmental state via the Imaging Unit 201 and the Force Sensor 801. The Actor Unit 804 is responsible for receiving the environmental state I, obtained by the sensor, and outputting the control amount A to the robot controller. The Critic Unit 803 is configured to cause the Actor Unit 804 to learn an output A for the input I, thus ensuring correct mating with the Actor Unit 804.
[0070] Next, the traditional reinforcement learning method will be described.
[0071] In reinforcement learning, a quantity called the reward R is defined, and the Actor unit 804 learns the action A that maximizes the reward R. For example, if the operation to be learned is to connect a plug 110 and a socket 120, as shown in embodiment 1, then R = 1 if the connection is successful, and R = 0 otherwise. In this example, action A specifies a "motion correction amount" relative to the current position (X, Y, Z, Ax, Ay, Az), where A = (dX, dY, dZ, dAx, dAy, dAz). Here, X, Y, and Z are position coordinates with the center of the robot as the origin; Ax, Ay, and Az are amounts of rotations about the X-axis, the Y-axis, and the Z-axis, respectively. The "movement correction amount" is the amount for a movement from the current position to the initial insertion position during a first attempt at connecting with connector 110. The environmental condition orThe test result is monitored by the imaging unit 201 and the force sensor 801 and obtained as an image or a value thereof.
[0072] In reinforcement learning, the Critic Unit 803 learns a function called the State Value Function V(I). If time t = 1 (for example, when the first assembly attempt is made), action A(1) is performed in state I(1). However, if time t = 2 (for example, after the end of the first assembly attempt and before the start of the second), the environment changes to I(2), and the reward amount R(2) (the result of the first assembly attempt) is obtained. Next, several possible example update formulas are shown. The update formula for V(I) is defined as follows. δ=R(2)+γV(I(2))−V(I(1)) V(I(1))⇐V(I(1))+α δ
[0073] Here, δ is a prediction error, α is a learning rate, which is a positive real number from 0 to 1, and y is a subtraction factor, which is a positive real number from 0 to 1. The Actor Unit 804 updates A(I) as follows, with I as input and A(I) as output. If δ>0, A(I(1))⇐A(I(1))+α(A(1)−A(I(1))) If δ≤0, σ(I(1))⇐β σ(I(1))
[0074] Here, σ denotes the standard deviation of the outputs. In state I, the Actor unit adds random numbers with a distribution having a mean of 0 and dispersion σ. 2 to A(I). This randomly determines the motion correction amount for the second attempt, independent of the result of the first attempt.
[0075] It should be noted that of the various update formulas in the Actor-Critic model, any of the commonly used models can replace the update formula shown above as an example.
[0076] In such a configuration, the actuator unit 804 learns the appropriate action for each state. However, as shown in embodiment 1, the action is only performed once the learning process is complete. During the learning process, the path-setting unit 806 calculates the recommended action for the learning process and sends it to the control unit 203. In other words, during the learning process, the control unit 203 receives the motion signal as is from the path-setting unit 806 and controls the drive unit 204.
[0077] In this approach, learning according to the conventional actor-critic model only occurs after successful assembly. This is due to the definition that R = 1 for a successful assembly and R = 0 otherwise. Until a successful assembly, the motion correction amounts used in subsequent attempts are randomly generated. Therefore, until a successful assembly, the determination of the motion correction amount for the next attempt, which depends on the degree of failure of the previous attempt, is not performed. Even when other reinforcement learning models such as Q-learning are used, similar results are obtained as with conventional actor-critic models, because only the success or failure of the assembly is evaluated.In the present embodiment of the disclosure, a determination process is described in which the degree of failure is evaluated to calculate the amount of motion correction for the next attempt.
[0078] The evaluation unit 805 generates a function to perform the evaluation on each assembly attempt. Fig. Figure 10 shows the plug 110 and the socket 120 during a plug-in test according to embodiment 2.
[0079] Suppose it is an image, like in Fig. Figure 10(A) shows the result obtained from the attempt. This attempt failed because the connecting elements are in a position that deviates significantly from the intended assembly position. Here, the distance to success is measured and quantified to obtain an evaluation value indicating the degree of success. One example of a quantification method is to calculate the area (number of pixels) of the connecting element, which is the target part in the image, as shown in Figure 10(A). Fig. Figure 10(B) shows how the calculation is performed. If, in this procedure, the force sensor 801 of the robot arm 100 detects that the insertion of the plug 110 into the socket 120 has failed, if only the contact surface of the socket 120 has a coating or sticker of a color different from that of other background surfaces, it is easier to acquire data from the image and perform the calculation. In the procedure described above, only one camera was used; however, multiple aligned cameras can also capture images, and the results derived from each image can be integrated.
[0080] A similar evaluation can also be carried out by obtaining the number of pixels in the two-dimensional directions (in the X direction and the Y direction) instead of the area of the connecting element.
[0081] The processing in the path specification unit 806 is divided into two steps.
[0082] In the first step, the path-setting unit 806 learns the evaluation result processed in the evaluation unit 805 and the actual movement of the robot. Assume the motion correction amount for the robot is A and the evaluation value processed by the evaluation unit 805, which indicates the degree of success, is E. Then the path-setting unit 806 creates a function with A as input and E as output to perform the approximation. The function contains, for example, a radial basis function (RBF) mesh. The RBF is known as a function that allows for a simple approximation of various unknown functions.
[0083] For example, the kth input is assumed as follows. xk=(x_1k,⋯,x_ik,⋯x_Ik)
[0084] Then the output f(x) is defined as follows. f(x)=∑iJ wi φj(x) φj(x)=exp(−∑(xi−μi)2σ2)
[0085] Here, σ denotes the standard deviation; µ denotes the RBF midpoint.
[0086] The training data used by the RBF is not a single data point, but rather all data from the beginning of the experiment to the last. If the current experiment is, for example, the Nth experiment, N data sets are created. Through training, W = (w_1, w_J), as mentioned above, must be determined. Of the various determination methods, the RBF interpolation is illustrated below as an example.
[0087] Let's assume Formula 8 and Formula 9 are as follows. Φ=(φ1(x11)⋯φI(xI1)⋮⋱⋮φ1(x1N)⋯φI(xIN)) F=(f(x1),⋯,f(xN))
[0088] Then the learning takes place via Formula 10. W=Φ−1F
[0089] After the approximation via RBF interpolation, the minimum value is calculated using a general optimization procedure, such as gradient descent and particle swarm optimization (PSO), within the RBF network. This minimum value is then transferred to the Actor unit 804 as the recommended value for the next trial.
[0090] The above case can be explained more concretely by arranging the areas or the number of pixels in the two-dimensional direction for the respective motion correction amounts in a time series for each trial number as evaluation values, and using the arranged values to obtain the optimal solution. Even more simply, the motion correction amount can be obtained, whereupon a movement at a constant speed is initiated in the direction in which the number of pixels in the two-dimensional direction decreases.
[0091] Next, in Fig. The process is shown in section 11.
[0092] Fig. Figure 11 is a flowchart showing the path learning of the position control device according to embodiment 2.
[0093] In step S1101, the gripping unit 101 of the robot arm 100 first grasps the plug 110. The position and orientation of the plug 110 are controlled in the control unit 203. Fig. 8 pre-registered, and the process is carried out according to the tax program pre-registered in control unit 203.
[0094] Next, in step S1102, the robot arm 100 is moved closer to the insertion position of the socket 120. The approximate position and orientation of the socket 120 are stored in the control unit 203. Fig. 8 is pre-registered, and the position of connector 110 is controlled according to the control program pre-registered in control unit 203. The steps up to this point correspond to steps S101 to S102 in the Fig. 4 flowchart shown in embodiment 1.
[0095] In step S1103, the path determination unit 802 instructs the imaging unit 201 of the monocular camera 102 to take an image. The monocular camera 102 takes an image that includes both the plug 110, gripped by the gripping unit 101, and the socket 120, which is the insertion target part. The path determination unit 802 also instructs the control unit 203 and the monocular camera 102 to take images near the current position. The monocular camera is moved by the drive unit 204 to a variety of positions based on motion parameters as instructed by the control unit 203, and at each position, it takes an image that includes both the plug 110 and the socket 120, which is the insertion target part.
[0096] In step S1104, the actor unit 804 provides the path determination unit 802 with an amount for a movement to be attached to the control unit 203; the control unit 203 causes the drive unit 204 to move the robot arm 100 in such a way that an attempt can be made to connect the plug 110 and the socket 120, which is the insertion target part.
[0097] In step S1105, when the connecting elements touch each other while the robot arm 100 is moved by the drive unit 204, the evaluation unit 805 and the critical unit 803 of the path determination unit 802 store the values obtained from the force sensor 801 and the images obtained from the monocular camera 102 per unit of movement.
[0098] In step S1106, the evaluation unit 805 and the critical unit 803 check whether the assembly was successful.
[0099] In most cases, the assembly is unsuccessful at this point. Therefore, the evaluation unit 805 assesses the degree of success in step S1108 using the value in Fig. 10 described procedures and delivers the evaluation value, which indicates the degree of success in the alignment, to the path setting unit 806.
[0100] In step S1109, the path-setting unit 806 performs the learning process using the procedure described above and transmits a recommended value for the next trial to the actor unit 804. The critical unit 803 calculates a value according to the reward amount and outputs the value. The actor unit 804 receives the value. In step S1110, the actor unit 804 adds the value obtained according to the reward amount and output by the critical unit 803 to the recommended value for the next trial output by the path-setting unit 806 to obtain the movement correction amount. However, if the recommended value for the next trial output by the path-setting unit 806 alone can produce a sufficient effect, the value obtained according to the reward amount need not be added in this step.Furthermore, when calculating the movement correction amount, the Actor unit 804 can establish an addition ratio of the recommended value for the next attempt, output by the path setting unit 806, to the value obtained according to the reward amount and output by the Critic unit 803, thereby allowing the movement correction amount to be changed according to the reward amount.
[0101] Then, at step S1111, the actor unit 804 transmits the motion correction amount to the control unit 203, and the control unit 203 moves the gripping unit 101 of the robot arm 100.
[0102] The process then returns to step S1103, images are captured at the position to which the robot arm 100 is moved according to the motion correction amount, and then the assembly process is performed. These steps are repeated until the assembly is successful.
[0103] If the connection is successful, the Actor Unit 804 and the Critic Unit 803 learn at step S1107 in which environmental state I from step S1102 to step S1106 the connection is successful. Finally, the Path Determination Unit 802 transmits the learned data of the neural network to the Control Parameter Generation Unit 202 so that the process can be carried out according to embodiment 1.
[0104] It should be noted that the Actor Unit 804 and the Critic Unit 803 learn in step S1107 in which environmental state I the mating is successful; however, the Actor Unit 804 and the Critic Unit 803 may learn using the data obtained for all mating attempts from the beginning until the mating is successful. In embodiment 1, a case is described in which a plurality of neural networks are formed according to the control amount. If the position leading to a successful mating is known, it is possible in this context to simultaneously form a plurality of suitable neural networks according to the size of the control amount.
[0105] This description is based on the Actor-Critic model as a reinforcement learning module, however, another reinforcement learning model such as Q-Learning can also be used.
[0106] The RBF network is an example of the approximation function, but other approximation function methods (linear function, quadratic function, etc.) can also be used.
[0107] In the above exemplary evaluation procedure, a fastener with a surface whose color differed from that of other surfaces was used; however, the amount obtained using a different image processing technique for a difference between fasteners or the like can also be used for the evaluation.
[0108] As indicated for embodiment 1 and the present embodiment, the application of this technique is not limited to the assembly of connecting elements. This method is practically applicable, for example, in the mounting of IC chips on a substrate and, in particular, in the insertion of capacitors or the like with leads having a large dimensional error into holes in a substrate.
[0109] This method is applicable not only to the control of insertion into substrates, but also to position control in general, in order to derive a control amount from a known relationship between images and control amounts. In the present disclosure, the neural network learns the relationships between images and control amounts; consequently, individual differences in the objects can be permitted during the execution of the object alignment, and the control amounts are precisely calculable.
[0110] According to the present embodiment, when the Actor-Critic model is applied to learning the control amounts, the Actor unit 804 calculates the motion correction amounts for the trials by adding the value obtained by the Critic unit 803 according to the reward amount and the recommended value obtained by the Path-Determining unit 806 based on the evaluation value. Therefore, according to the present disclosure, the number of alignment trials can be significantly reduced, whereas conventional Actor-Critic models require a very large number of trials and errors before successful alignment.
[0111] It should be noted that in the present embodiment, it is described that the number of alignment attempts can be reduced by evaluating the misalignment images obtained by the imaging unit 201; however, the number of attempts can also be reduced using the values obtained from the force sensor 801 during the alignment attempts. For example, when connecting elements or aligning two objects that include an insertion, failure detection is typically carried out such that, if the value obtained from the force sensor 801 exceeds a limit value, the actuator unit 804 checks whether the two objects are in a position where the connection or insertion has been completed. In such a case, the following scenarios are conceivable. One of these scenarios is a case in which the connection or insertion...the insertion has not been fully completed (Case A), and another of these cases is a case in which the assembly or insertion has been fully completed, but a certain value has been reached during the assembly or insertion with the value obtained from the force sensor 801 (Case B).
[0112] In case A, a method can be used in which a learning process is carried out using both the value from force sensor 801 and the image. This is described in detail in embodiment 3.
[0113] In case B, a method can be used in which a learning process is performed using only the value from the force sensor 801, as described in embodiment 3. A similar effect can be achieved using a different method in which the reward R of the actor-critic model is defined such that at the time of success, R = (1 - A / F), and at the time of failure, R = 0. Here, F is the maximum load applied during assembly or insertion, and A is a positive constant. embodiment 3
[0114] In the present embodiment, a method for efficient data acquisition during the learning process is described, which is carried out after successful alignment in embodiment 2. It is assumed that anything not specifically mentioned here is the same as in embodiment 2. Consequently, with regard to the position control device according to embodiment 3, in Fig. 8 shows the function configuration diagram, while in Fig. 9 shows the hardware structure diagram.
[0115] Regarding the action, a procedure for more efficient learning data collection during the process at step S1107 is described below. Fig. 11 described in embodiment 2.
[0116] Fig. Figure 12 is a flowchart showing a path learning process of the position control device according to embodiment 3.
[0117] If the connection of plug 110 and socket 120 is unsuccessful during step S1107 of Fig. If step 11 is successful, the path definition unit 806 at step S1201 first initializes the variables i = 0, j = 1, and k = 1. Here, the variable i indicates the number of subsequent learning operations for the robot arm 100, the variable k indicates the number of learning operations after disconnecting the plug 110 and socket 120; and the variable j indicates the number of loops in the Fig. The flowchart shown in 12 is shown.
[0118] Next, at step S1202, the path-setting unit 806 transmits an amount for a movement via the actor unit 804 to the control unit 203 to retract by 1 mm from a state defined by the step S1104. Fig. The amount provided for assembly in 11 is used for a movement, and causes the drive unit 204 to move the robot arm 100 back accordingly. Then the variable i is incremented by one. Here, the amount provided for a movement is for a retraction of 1 mm, however, the distance is not limited to 1 mm, but can also be 0.5 mm, 2 mm, etc.
[0119] At step S1203, the path-setting unit 806 stores the current coordinates as O(i) (at this time i = 1).
[0120] In step S1204, the path-setting unit 806 generates a random amount for a movement (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) from point O(i) and transmits this random amount to the control unit 203 via the actor unit 804. The drive unit 204 then moves the robot arm 100 accordingly. Any amount within the range of motion of the robot arm 100 can be set as the maximum movement amount.
[0121] In step S1205, the Actor unit 804, at the position reached in step S1204, records the value obtained from the force sensor 801, which corresponds to the random amount for a movement (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz). In step S1206, the Critic unit 803 and the Actor unit 804 record the value (-ΔX, -ΔY, -ΔZ, -ΔAx, -ΔAy, -ΔAz), namely the movement value multiplied by -1, and the sensor value from the force sensor 801, which measures the force required to hold the plug 110, as a training data set.
[0122] In step S1207, the path-setting unit 806 assesses whether the number of captured records has reached the previously determined number J. If the number of records is lower, in step S1208 the variable j is incremented by one, and after returning to step S1204, the amount for a movement (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) is updated using an arbitrary random number to obtain another record. Steps S1204 through S1207 are then repeated until the number of records reaches the previously determined number J.
[0123] As soon as the number of data records reaches the previously determined number, the path definition unit 806 sets the variable j to one at step S1209 and checks at step S1210 whether the connection of plug 110 and socket 120 is disconnected.
[0124] If it is not resolved, the process returns to step S1202 via step S1211.
[0125] In step S1211, the path-setting unit 806 transmits the amount for a movement via the actor unit 804 to the control unit 203, so that the coordinates of the robot arm 100 are again available at O(i), i.e., the coordinates before the transmission of the random amounts for a movement. Then the drive unit 204 moves the robot arm 100 accordingly.
[0126] The loop then repeats from step S1202 to step S1210; that is, the following two processes are repeated until the connection between plug 110 and socket 120 is disconnected: the process in which a retraction of 1 mm or another distance occurs from a state caused by the amount of force provided for connection, and a corresponding retraction of the robot arm 100 takes place; and a process in which random amounts of force are provided for movement from the position after the retraction, and data from the force sensor 801 are acquired at this point. Once the connection between plug 110 and socket 120 has been disconnected, the process proceeds to step S1212.
[0127] In step S1212, the path setting unit 806 sets the value of variable i to I, where I is an integer greater than the value of variable i when it is determined that the connection between plug 110 and socket 120 is disconnected. The path setting unit 806 then transmits to the control unit 203, via the actuator unit 804, an amount for a retraction movement of 10 mm (with no limit to this value) from the state caused by the amount of movement provided for reconnection, and causes the drive unit 204 to retract the robot arm 100 accordingly.
[0128] In step S1213, the path definition unit 806 stores the coordinates of the position to which the robot arm 100 moved in step S1212 as the coordinates of the center position O (i + k).
[0129] In step S1214, the path-setting unit 806 generates a random amount for a movement (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) from the center position O(i+k) and transmits the generated random amount for a movement via the actor unit 804 to the control unit 203. Then the drive unit 204 moves the robot arm 100 accordingly.
[0130] In step S1215, the Critic Unit 803 and the Actor Unit 804 receive the image that was captured by the Imaging Unit 201 of the monocular camera 102 at the position of the robot arm 100, which has moved by the random amount for a movement (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz).
[0131] In step S1216, the Critic unit 803 and the Actor unit 804 record the amount (-ΔX, -ΔY, -ΔZ, -ΔAx, -ΔAy, -ΔAz) obtained by multiplying the random amount for a movement by -1 and the captured image as a training data set.
[0132] In step S1217, the path-setting unit 806 checks whether the number of captured records has reached the previously determined number J. If the number of records is lower, in step S1218 the variable j is incremented by one, and the process returns to step S1214. The path-setting unit 806 randomly changes the amount for a movement (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) and receives another record for a movement. Steps S1214 to S1217 are repeated until the number of records reaches the previously determined number J.
[0133] It should be noted that the random maximum amount for a move (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) at S1214 and the random maximum amount for a move at S1204 may be different.
[0134] The Actor Unit 804 and the Critic Unit 803 perform the learning process using the training data sets obtained through the procedure described above.
[0135] Fig. Figure 13 is a scheme showing an exemplary neural network and the learning rule of the neural network according to embodiment 3.
[0136] In embodiments 1 and 2, the learning method used by the force sensor 801 is not described. In embodiments 1 and 2, only the images are used for the input layer, whereas in embodiment 3, the values obtained by the force sensor 801 can be fed to the input layer instead of the images. The force sensor 801 may provide three values (a force and moments in two directions) or six values (forces in three directions and moments in three directions). The control value (ΔX, ΔY, ΔZ, ΔAx, ΔAy, ΔAz) is output in the output layer. It should be noted that as soon as the connection between plug 110 and socket 120 is disconnected, both the images and the values obtained by the force sensor 801 are simultaneously fed to the input layer.
[0137] During the learning process of the neural network, the parameters of the hidden layer are optimized so that the output values of the output layer approximate the control values. The output values are derived from the input images and the values from the force sensor 801 via the hidden layer, and the control values are stored with their respective images and their respective values from the force sensor 801. Finally, the path determination unit 802 transmits the learned data of the neural network to the control parameter generation unit 202 so that the process can be carried out according to embodiment 1.
[0138] The description of the present embodiment is based on the following assumptions. To perform the learning process, the robot arm 100 gradually moves backward from the connection position of the plug 110 and the socket 120 and then moves slightly toward the outer positions. Depending on the number of pixels in the image from the monocular camera 102, the learning process can only be carried out satisfactorily once the connection has been disconnected. However, if the monocular camera 102 produces images with a sufficiently high resolution, and the learning process can therefore be carried out satisfactorily, and if the images obtained when the robot arm 100 moves slightly toward the outer positions are used, the learning process can be carried out using only the images provided by the monocular camera 102.Furthermore, both the images from the monocular camera 102 and the values obtained from the force sensor 801 can be used even when the plug 110 and the socket 120 are plugged together.
[0139] In embodiments 1 and 2, a case is described in which a plurality of neural networks are used. Similarly, in the present embodiment, for example, different neural networks can be used when the plug 110 and the socket 120 are connected or disconnected. As described above, the learning process is more accurate when the values from the force sensor 801 are used to form the input layer when the plug 110 and the socket 120 are connected, and when the images are used to form the input layer when the connection is disconnected. Furthermore, even when only the images are used for learning, the learning process is more accurate if the processing for the connected state and the processing for the disconnected state are performed separately, because the configuration of the images corresponding to each state is different.
[0140] As described for embodiments 1 and 2, the application of this technique is not limited to the assembly of connecting elements. This method is practically applicable, for example, in the mounting of IC chips on a substrate and, in particular, in the insertion of capacitors or the like with leads having a large dimensional error into holes in a substrate.
[0141] This method is applicable not only to the control of insertion into substrates, but also to position control in general, in order to derive a control amount from a known relationship between images and control amounts. In the present disclosure, the neural network learns the relationships between images and control amounts; consequently, individual differences in the objects can be permitted during the execution of the object alignment, and the control amounts are precisely calculable.
[0142] In the present embodiment, for operations involving the alignment and insertion of two objects, a path-setting unit 806 and an actor unit 804 are provided for learning a control amount. The path-setting unit 806 provides the amount for a movement to remove an object from its insertion position and to locate it on and around the removal path. The actor unit 804 receives the object's position and the values from a force sensor 801 at that location to perform a learning operation by allowing the object's position to be the values for the output layer and the force sensor 801's values to be the values for the input layer. Consequently, training data can be acquired efficiently. Reference symbol list 100 robot arms 101 gripping unit 102 monocular camera 110 plugs 120 socket 201 imaging unit 202 Control parameter generation unit 203 Control unit 204 Drive unit 301 Input / Output Interface 302 processor 303 Storage element 304 Control circuit 305 engine 801 Force sensor 802 Path Determination Unit 803 Critic Unit 804 Actor Unit 805 rating unit 806 Path Definition Unit
Claims
[1] Position control device for installing a first object (110) in a second object (120) by causing a drive unit (204) to move the first object (110), the device comprising: an imaging unit (201) of a monocular camera (102) for capturing an image containing two objects, namely the first object (110) and the second object (120); and a control parameter generation unit (202), wherein among a multitude of neural networks, each of which an input layer into which information from a large number of images acquired by the imaging unit (201) is entered and an output layer in which control amounts obtained based on a learning rule such as a stochastic gradient method are output, wherein each of the control amounts for moving the first object (110) corresponds to one of the plurality of images input into the input layer, the control parameter generation unit (202) when the imaging unit (201) acquires an image for the first time, selects a predetermined neural network and outputs the control amount in the output layer of the selected neural network as a control amount to cause the drive unit (204) to move the first object (110) among the two objects, which are the first object (110) and the second object (120), contained in the image captured by the imaging unit (201), wherein the control parameter generation unit (202) an image captured a second time or later by the imaging unit (201), inputs information from the captured image into an input layer of a neural network, which is selected from the multitude of neural networks based on a control amount that is output to an output layer of a neural network whose input layer has received information from a previously captured image, and outputs the control amount in an output layer of the selected neural network as a control amount to cause the drive unit (204) to move the first object (110) of the two objects, which are the first object (110) and the second object (120) contained in the image captured by the imaging unit (201). [2] Position control device according to claim 1, further comprising a control unit (203) for controlling an current value or a voltage value for the drive unit (204) for moving the first object (110) using the control amount output by the control parameter generation unit (202). [3] Position control device according to claim 1 or 2, further comprising an actuator unit (804) and a path determination unit (806), wherein in an attempt the Actor unit (804) receives information from an image captured by the imaging unit (201), a motion correction amount that maximizes a defined reward, obtained based on a position of the first object (110) in an image first captured by the imaging unit (201) in the trial, is output as a motion control amount to cause the drive unit (204) to move the first object (110), and outputs a motion correction amount, obtained using a recommendation value based on the position of the first object (110) in an image acquired a second time or later by the imaging unit (201), as a motion control amount to cause the drive unit (204) to move the first object (110), wherein the path-setting unit (806), when the first object (110) is moved according to a motion control amount issued by the Actor unit (804), a recommendation value is received based on a rating value that indicates a degree of success, based on a positional relationship between the first object and the second object (120), and The received recommendation value is transmitted as a recommendation value to the Actor unit (804). [4] Position control method for installing a first object (110) in a second object (120) by causing a drive unit (204) to move the first object (110), the method comprising: Input of information from an image captured by an imaging unit (201) of a monocular camera (102) and containing two objects, a first object and a second object, into an input layer of a predetermined neural network among a plurality of neural networks that differ in the range of the control amount output in the output layer; Outputting a control amount, which is issued in the output layer and is an amount for a movement, which serves to move the first object (110) that corresponds to an image first captured by the imaging unit (201) and which is obtained based on a learning rule such as a stochastic gradient method, in an output layer of the predetermined neural network as a control amount to cause the drive unit (204) to move the first object (110) of the two objects, which are the first object (110) and the second object (120) contained in the image first captured by the imaging unit (201); Inputting information from an image containing two objects, namely the first object (110) and the second object (120), which were captured by the imaging unit (201) for the second or subsequent time, into an input layer of a neural network selected from the multitude of neural networks based on the size of the control amount, which is output in the output layer of the neural network whose input layer received the image captured by the imaging unit (201) the previous time; and Spending a control amount, which is an amount for a movement, which serves to move the first object (110) that corresponds to an image captured by the imaging unit (201), and is entered into the input layer of the selected neural network and which is obtained based on a learning rule such as a stochastic gradient method, in the output layer of the selected neural network as a control amount to cause the drive unit (204) to move the first object (110) of the two objects, which are the first object (110) and the second object (120) contained in an image captured by the imaging unit (201). [5] Position control method for installing a first object (110) in a second object (120) by causing a drive unit (204) to move the first object (110), the method comprising: Calculating a motion correction amount that maximizes a defined reward, based on a position of the first object (110) in information from an image containing two objects, the first object (110) and the second object (120), which is captured by an imaging unit (201) of a monocular camera (102) on a first attempt, as a motion control amount to cause the drive unit (204) to move the first object (110); Calculating a recommendation value based on a rating value that indicates a degree of success, based on a positional relationship between the first object and the second object (120), when the first object (110) is moved according to the calculated motion control amount; Calculating a motion correction amount based on the position of the first object (110) in an image captured by the imaging unit (201) on a second attempt, using the calculated recommended value as the motion control amount to cause the drive unit (204) to move the first object (110); when the imaging unit (201) captures an image on a third or subsequent attempt, calculating a recommendation value based on an evaluation value indicating a degree of success, based on a positional relationship between the first object and the second object (120), when the first object (110) is moved according to a motion control amount calculated the previous time; Calculating a motion correction amount from a position of the first object (110) in an image taken by the imaging unit (201) on a third or later attempt, using the calculated recommendation amount as a motion control amount to cause the drive unit (204) to move the first object (110); Inputting information from the image containing the two objects, namely the first object (110) and the second object (120), captured by the imaging unit (201), into an input layer of a neural network, which is selected from the multitude of neural networks, each of which has a different range of a control amount output in its output layer, based on the calculated motion control amount; and Spending a control amount, which is an amount for a movement, which serves to move the first object (110) that corresponds to an image captured by the imaging unit (201), and is entered into the input layer of the selected neural network and which is obtained based on a learning rule such as a stochastic gradient method, in the output layer of the selected neural network as a control amount to cause the drive unit (204) to move the first object (110) of the two objects, which are the first object (110) and the second object (120) contained in an image captured by the imaging unit (201).
Citation Information
Patent Citations
Control method of robot device, and the robot device
JP2009083095A
Force control robot system with visual sensor for inserting work
WO1998017444A1
JP002009083095A