Image processing device, image processing method, image processing program, and robot control system
By using an image processing device that masks grayscale images with three-dimensional information, the robot control system enhances object recognition accuracy during depalletizing tasks, addressing the limitations of existing systems.
Patent Information
- Application Number
- JP2022009614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing robot control systems face challenges in achieving high recognition accuracy during object recognition processing, particularly in depalletizing tasks where accurate identification of cargo is crucial.
The system employs an image processing device that acquires grayscale images and three-dimensional information of a predetermined space, masks portions of the image based on the three-dimensional information, and performs object recognition using the masked image.
This approach significantly improves recognition accuracy by reducing erroneous recognition and enhancing the system's ability to accurately identify objects in complex environments.
Smart Images

Figure 2025078887000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an image processing device, an image processing method, an image processing program, and a robot control system. [Background technology]
[0002] A robot control system that automates the task of unloading cargo from pallets (depalletizing) is currently being developed. In this robot control system, for example, an object recognition process is performed to identify cargo to be depalletized using RGB images captured by an RGB camera. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2020-075340 A [Patent Document 2] Patent No. 6211734 Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure improves the recognition accuracy in object recognition processing. [Means for solving the problem]
[0005] An image processing device according to an aspect of the present disclosure has, for example, the following configuration. one or more memories; one or more processors; The one or more processors: acquiring a grayscale image of a predetermined space including one or more objects and three-dimensional information of the predetermined space; masking a portion of the grayscale image based on the three-dimensional information; performing a predetermined process using the masked gradation image; Execute. [Brief description of the drawings]
[0006] [Figure 1] FIG. 1 is a diagram illustrating an example of a usage scene of a robot control system. [Diagram 2] FIG. 2 is a diagram illustrating an example of a system configuration of each phase of a robot control system. [Diagram 3] FIG. 2 illustrates an example of a hardware configuration of an image processing apparatus. [Figure 4] FIG. 1 is a first diagram illustrating an example of a functional configuration of an image processing device in a training phase. [Diagram 5] FIG. 11 is a first diagram showing an example of the functional configuration of a training unit. [Figure 6] FIG. 1 is a first diagram illustrating an example of a functional configuration of an image processing device in a depalletizing phase. [Figure 7] FIG. 11 is a diagram showing a specific example of object recognition processing. [Figure 8] 1 is a first flowchart showing the flow of processing in the entire robot control system. [Figure 9] 13 is a flowchart showing the flow of a training process. [Figure 10] 1 is a first flowchart showing the flow of a depalletizing process. [Figure 11] FIG. 11 is a diagram showing a specific example of a depalletizing process. [Figure 12] FIG. 2 is a second diagram illustrating an example of a functional configuration of an image processing device in a training phase. [Figure 13] FIG. 2 is a second diagram showing an example of the functional configuration of the training unit. [Figure 14] FIG. 2 is a second diagram illustrating an example of a functional configuration of the image processing device in the depalletizing phase. [Figure 15] 13 is a second flowchart showing the flow of the training process. [Figure 16] 13 is a second flowchart showing the flow of the depalletizing process. [Figure 17]FIG. 11 is a diagram illustrating an example of a functional configuration of an image processing device in a shooting condition adjustment phase. [Figure 18] 11 is a second flowchart showing the flow of processing in the entire robot control system. [Figure 19] 13 is a flowchart showing the flow of an imaging condition adjustment process. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0007] Hereinafter, each embodiment will be described with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same functional configurations are denoted by the same reference numerals, and redundant description will be omitted.
[0008] [First embodiment] <Robot control system usage scenarios> First, a usage scene of the robot control system according to the first embodiment will be described. Fig. 1 is a diagram showing an example of a usage scene of the robot control system. The example of Fig. 1 shows a scene in which the robot control system 100 is used for automating depalletizing.
[0009] 1 shows one or more rectangular parallelepiped-shaped cardboard boxes as luggage, which is an example of an object, stacked on a pallet 140. The example in Fig. 1 also shows the robot 110 picking up the cardboard boxes one by one in order from the top and lowering them onto a transport unit 160 such as a conveyor (depalletizing).
[0010] 1 is merely an example, and the work automated by the robot control system 100 is not limited to depalletizing, but may be other work. Furthermore, the type of cargo to be depalletized by the robot control system 100 is not limited to cardboard boxes, but may be other types of cargo. Furthermore, the shape of the cargo is not limited to a rectangular parallelepiped, but may be any other shape.
[0011] Also, the order in which the cardboard boxes are picked up does not have to be from the top down, and for example, cardboard boxes at a certain height may be picked up first. Alternatively, a specific area on the pallet 140 may be given priority, and cardboard boxes in the prioritized area may be picked up first from the top down.
[0012] Furthermore, the number of cardboard boxes picked up in one operation is not limited to one, and multiple cardboard boxes may be picked up in one operation. These work rules are predefined as work rule information, and the robot control system 100 performs the work in accordance with the work rule information.
[0013] In the present embodiment, the following description will be given on the assumption that a work rule that "picks up cardboard boxes one by one in order from the top" is prescribed as an example of work rule information.
[0014] Returning to the explanation of Fig. 1, as shown in Fig. 1, the robot control system 100 has a robot 110, an RGB camera 121, a depth camera 122, and an image processing device 130. Fig. 1(a) shows the robot control system 100 viewed in the positive direction of the y axis (direction from the front to the back of the paper) when the directions indicated by the reference numeral 171 are the x-axis direction and the z-axis direction, respectively. Meanwhile, Fig. 1(b) shows the robot control system 100 viewed in the negative direction of the x axis (direction from the front to the back of the paper) when the directions indicated by the reference numeral 172 are the y-axis direction and the z-axis direction, respectively.
[0015] The RGB camera 121 shown in FIGS. 1(a) and (b) is an example of an image generating device that generates an image. The RGB camera 121 is, for example, fixedly attached to the ceiling of the space in which the robot 110 is arranged, and captures cardboard boxes 150 stacked on a pallet 140 from directly above, thereby notifying the image processing device 130 of the RGB image. Note that in this disclosure, including the present embodiment, an example of handling an RGB image having color information for each pixel will be described, but instead of an RGB image, a grayscale image having brightness information for each pixel may also be handled. That is, in this disclosure, as an example of a gradation image having color or brightness gradation information (gradation value) for each pixel, a color image having color information, specifically, color information such as ·R(Red), ·G(Green), ·B (Blue), Here, we will explain an RGB image expressed using the values of the three primary colors of the three colors. Note that other examples of color images include a Lab image in which color information is expressed using values in the Lab color space, and an HSL image in which color information is expressed using values of hue, saturation, and brightness.
[0016] 1(a) and (b) is an example of a three-dimensional information generating device that generates three-dimensional information. The depth camera 122 is fixedly attached to the ceiling of the space in which the robot 110 is placed, for example, and captures an image of cardboard boxes 150 stacked on a pallet 140 from directly above, thereby notifying the image processing device 130 of the three-dimensional information.
[0017] The RGB image captured by the RGB camera 121 and the three-dimensional information captured by the depth camera 122 may be configured with the same resolution or different resolutions. However, when the RGB image and the three-dimensional information are processed in the image processing device 130, each pixel (two-dimensional coordinates (X coordinate, Y coordinate)) of the RGB image is associated with three-dimensional information (x coordinate, y coordinate, z coordinate) in addition to color information (R value, G value, B value). In the following description, unless otherwise specified, a pixel refers to a pixel in the RGB image.
[0018] Moreover, the image processing device 130 shown in FIGS. 1(a) and 1(b) outputs a control command for controlling the robot 110 based on the notified RGB image and three-dimensional information.
[0019] Specifically, the image processing device 130 identifies the area of the top surface of the cardboard box that has the highest height in the RGB image based on the three-dimensional information.
[0020] Furthermore, the image processing device 130 performs a mask process on color information associated with pixels other than the pixels included in the specified region in the RGB image. Note that, hereinafter, the masked RGB image generated by performing a mask process on color information associated with pixels other than the pixels included in the specified region in the RGB image is referred to as a "mask image."
[0021] In addition, the image processing device 130 performs object recognition processing using the mask image. Furthermore, based on the result of the object recognition processing, the image processing device 130 generates a control command for the robot 110 to depalletize the cardboard boxes identified as the depalletizing target, and transmits the control command to the robot 110.
[0022] FIG. 1(a) shows how the robot 110 picks up the tallest cardboard box and then rotates in the direction of the thick arrow 180 based on a control command sent from the image processing device 130.
[0023] FIG. 1(b) shows the state in which robot 110 has completed rotation in the direction of thick arrow 180 and has placed the picked-up cardboard box onto transport unit 160.
[0024] In this way, the robot control system 100 according to the first embodiment performs mask processing on the RGB image based on the three-dimensional information, and performs object recognition processing using the masked image. This makes it possible to reduce erroneous recognition and improve recognition accuracy, compared to, for example, performing object recognition processing using an RGB image that has not been subjected to mask processing.
[0025] <System configuration of robot control system> Next, a description will be given of the system configuration of the robot control system 100. Fig. 2 is a diagram showing an example of the system configuration of the robot control system.
[0026] As described above, the image processing device 130 performs object recognition processing using a mask image. Here, the object recognition unit used when the image processing device 130 performs object recognition processing (instance segmentation) is configured with, for example, a DNN (Deep Neural Network). The object recognition unit is trained by using the mask image as training data.
[0027] That is, as shown in FIG. 2, the system configuration of the robot control system 100 can be divided into a system configuration in a "training phase" (FIG. 2(a)) and a system configuration in a "depalletizing phase" (FIG. 2(b)).
[0028] As shown in FIG. 2(a), in the training phase, the robot control system 100 performs, for example, RGB camera 121 and A depth camera 122; CG (Computer Graphic) simulator 210, An acquired data storage unit 220; An image processing device 230; It is composed of:
[0029] In the case of the training phase, the RGB image captured by the RGB camera 121 and the three-dimensional information captured by the depth camera 122 may be stored in the acquired data storage unit 220. In addition, in the case of the training phase, the environment in which the robot 110 is placed is reproduced by the CG simulator 210, whereby a virtual RGB image and virtual three-dimensional information may be generated and stored in the acquired data storage unit 220.
[0030] In addition, in the case of the training phase, the RGB images and the three-dimensional images stored in the acquired data storage unit 220 are read out to the image processing device 230, where mask processing is performed to generate training data. Furthermore, in the case of the training phase, the image processing device 230 trains an object recognition unit (described in detail below) using the generated training data, and a trained object recognition unit is generated.
[0031] On the other hand, as shown in FIG. 2(b), in the depalletizing phase, the robot control system 100 RGB camera 121 and A depth camera 122; An image processing device 130 (including a trained object recognizer); Robot 110, It is composed of:
[0032] In the depalletizing phase, the RGB image captured by the RGB camera 121 and the three-dimensional information captured by the depth camera 122 are each notified to the image processing device 130, which performs mask processing and object recognition processing. Furthermore, in the depalletizing phase, a control command is generated based on the result of the object recognition processing, and the robot 110 is controlled. As a result, the robot 110 depalletizes the cardboard boxes to be depalletized.
[0033] In the depalletizing phase, the timing of photographing by the RGB camera 121 and the depth camera 122 is arbitrary, as long as the photographing is done before the cardboard boxes to be depalletized are picked up.
[0034] In addition, during the depalletizing phase, the frequency of photographing by the RGB camera 121 and the depth camera 122 is also arbitrary; for example, photographing may be performed every time one cardboard box is depalletized, or photographing may be performed once every time multiple cardboard boxes are depalletized.
[0035] 2 shows a case where different image processing devices are used in the training phase and the depalletizing phase. However, the image processing device 230 used in the training phase and the image processing device 130 used in the depalletizing phase may be the same image processing device.
[0036] <Hardware configuration of image processing device> Next, a description will be given of the hardware configuration of the image processing devices 130 and 230. Note that since the image processing devices 130 and 230 have similar hardware configurations, the hardware configuration of the image processing device 130 will be described here as a representative.
[0037] Fig. 3 is a diagram showing an example of the hardware configuration of an image processing device. As shown in Fig. 3, the image processing device 130 has, as components, a processor 301, a main storage device (memory) 302, an auxiliary storage device 303, a network interface 304, and a device interface 305. The image processing device 130 is realized as a computer in which these components are connected via a bus 306.
[0038] In the example of FIG. 3, the image processing device 130 is shown as having one of each component, but the image processing device 130 may have a plurality of the same components. In the example of FIG. 3, one image processing device 130 is shown, but the image processing program may be installed in a plurality of image processing devices, and each of the plurality of image processing devices may execute the same or different parts of the image processing program. In this case, the image processing device may take the form of distributed computing in which each image processing device communicates via a network interface 304 or the like to execute the entire process. In other words, the image processing device 130 may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize the function. In addition, the image processing device 130 may be configured to process various data transmitted from the RGB camera 121 and the depth camera 122 in one or more image processing devices provided on the cloud, and transmit the processing results to the image processing device of the customer.
[0039] Various calculations of the image processing device 130 may be executed in parallel using one or more processors, or using multiple image processing devices via the communication network 310. Moreover, various calculations may be distributed to multiple arithmetic cores in the processor 301 and executed in parallel. Moreover, a part or all of the processes, means, etc. of the present disclosure may be executed by an external device 320 (at least one of a processor and a storage device) provided on a cloud that can communicate with the image processing device 130 via the communication network 310. In this way, the image processing device 130 may take the form of parallel computing using one or more computers.
[0040] The processor 301 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) that at least controls a computer or performs calculations. The processor 301 may be a general-purpose processor, a dedicated processing circuit designed to perform a specific calculation, or a semiconductor device including both a general-purpose processor and a dedicated processing circuit. The processor 301 may include an optical circuit, or may include a calculation function based on quantum computing.
[0041] The processor 301 may perform various calculations based on various data and commands input from each device in the internal configuration of the image processing device 130, and may output calculation results and control signals to each device, etc. The processor 301 may control each component of the image processing device 130 by executing an OS (Operating System), applications, etc.
[0042] Furthermore, processor 301 may refer to one or more electronic circuits arranged on one chip, or to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other by wire or wirelessly.
[0043] The main memory 302 is a memory device that stores instructions executed by the processor 301 and various data, and the various data stored in the main memory 302 may be read by the processor 301. The auxiliary memory 303 is a memory device other than the main memory 302. Note that these memory devices refer to any electronic components capable of storing various data, and may be semiconductor memories. The semiconductor memories may be either volatile memories or non-volatile memories. The memory device for saving various data in the image processing device 130 may be realized by the main memory 302 or the auxiliary memory 303, or may be realized by an internal memory built into the processor 301.
[0044] Furthermore, a plurality of processors 301 may be connected (coupled) to one main storage device 302, or a single processor 301 may be connected. Alternatively, a plurality of main storage devices 302 may be connected (coupled) to one processor 301. When the image processing device 130 is configured with at least one main storage device 302 and a plurality of processors 301 connected (coupled) to the at least one main storage device 302, the image processing device 130 may include a configuration in which at least one of the plurality of processors 301 is connected (coupled) to at least one main storage device 302. This configuration may also be realized by the main storage devices 302 and the processors 301 included in a plurality of image processing devices 130. Furthermore, the image processing device 130 may include a configuration in which the main storage device 302 is integrated with the processor (for example, a cache memory including an L1 cache and an L2 cache).
[0045] The network interface 304 is an interface for connecting to the communication network 310 wirelessly or by wire. An appropriate interface, such as one conforming to an existing communication standard, may be used for the network interface 304. The network interface 304 may exchange various data with the robot 110 and other external devices 320 connected via the communication network 310. The communication network 310 may be any one of a wide area network (WAN), a local area network (LAN), a personal area network (PAN), etc., or a combination thereof, as long as information is exchanged between a computer and the robot 110 and other external devices 320. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication).
[0046] The device interface 305 is an interface such as a USB that directly connects to an external device 330 .
[0047] The external device 330 is a device connected to a computer. The external device 330 may be an input device, for example. The input device is, for example, a device such as a camera (including the RGB camera 121 and the depth camera 122 of this embodiment), a microphone, a motion capture device, various sensors, a keyboard, a mouse, a touch panel, etc., and provides acquired information to the computer. Alternatively, the external device 330 may be a device having an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0048] Moreover, the external device 330 may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound or the like. The output device may also be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0049] The external device 330 may be a storage device (memory). For example, the external device 330 may be a network storage or the like, or may be a storage such as a HDD.
[0050] Furthermore, the external device 330 may be a device having some of the functions of the constituent elements of the image processing device 130. In other words, the computer may transmit or receive some or all of the processing results of the external device 330.
[0051] <Details of the functional configuration of the image processing device in the training phase> Next, the functional configuration of the image processing device 230 in the training phase will be described in detail. Fig. 4 is a first diagram showing an example of the functional configuration of the image processing device in the training phase. As described above, an image processing program is installed in the image processing device 230, and in the training phase, the program is executed, and as shown in Fig. 4, ·3D information acquisition unit 410, A mask region identification unit 420, RGB image acquisition unit 430, Mask part 440, A training data generator 450; ·Training Department 460, It functions as:
[0052] The three-dimensional information acquisition unit 410 acquires the three-dimensional information read from the acquired data storage unit 220 and notifies the mask region identification unit 420 of the acquired three-dimensional information.
[0053] The mask area identifying unit 420 accepts "processing range information" and "work rule information" in advance. The processing range information is a range specified in an RGB image, and is, for example, a range on the pallet 140 that corresponds to a predetermined space in which cardboard boxes can be placed. The work rule information accepted by the mask area identifying unit 420 is information used when identifying pixels to be included in the mask area. In the case of this embodiment, the mask area identifying unit 420 accepts "Pick up cardboard boxes one by one in order from the top" as the work rule information.
[0054] Furthermore, the mask region identifying unit 420 specifies a range in the RGB image based on the processing range information. Furthermore, the mask region identifying unit 420 extracts three-dimensional information corresponding to the work rule information (in this embodiment, the three-dimensional information of the top surface of the tallest cardboard) from the three-dimensional information notified from the three-dimensional information acquisition unit 410. Furthermore, the mask region identifying unit 420 specifies pixels associated with the extracted three-dimensional information in the RGB image of the range specified based on the processing range information as pixels to be included in the "non-mask region" (pixels included in a region that satisfies a predetermined condition).
[0055] In another example of the process of identifying the non-mask region, the mask region identifying unit 420 identifies pixels associated with three-dimensional information that satisfies a condition determined based on the work rule information, among the pixels included in the RGB image of the range specified based on the processing range information, as pixels to be included in the "non-mask region." For example, in this embodiment, the mask region identifying unit 420 identifies pixels associated with three-dimensional information whose height (z coordinate) is equal to or greater than the top X percentile (e.g., X=99), among the pixels included in the RGB image of the range specified based on the processing range information.
[0056] That is, the mask region identifying section 420 identifies a non-mask region in the RGB image based on three-dimensional information corresponding to the processing range information and the work rule information.
[0057] Furthermore, the mask region specification section 420 specifies the region other than the non-mask region as a “mask region” and notifies the mask section 440 of this.
[0058] The RGB image acquisition unit 430 acquires the RGB image read from the acquired data storage unit 220 and notifies the mask region specification unit 420 and the mask unit 440 of the acquired RGB image.
[0059] The mask unit 440 performs mask processing on color information associated with pixels included in the “mask area” notified by the mask area identification unit 420, out of the RGB image notified by the RGB image acquisition unit 430, to generate a “mask image”.
[0060] The masking process performed by the mask unit 440 on the color information associated with the pixels included in the mask area includes the following steps: Converting color information (R values, G values, B values) associated with pixels other than those included in the non-mask area (pixels included in the mask area) in an RGB image into a single predetermined color information (for example, R values, G values, B values indicating black or gray, etc.), or Processing an RGB image other than the non-masked area (an image in the masked area) into an image containing multiple predetermined color information (for example, a gradation image), or - Applying filtering such as blurring to the image other than the non-masked area (the image in the masked area) of the RGB image, and processing the image in the masked area into an image that is blurred more than the image in the non-masked area. etc. are included.
[0061] Furthermore, the mask unit 440 notifies the training data generation unit 450 of the mask image.
[0062] The training data generating unit 450 associates the mask image notified by the mask unit 440 with the correct answer data and stores the mask image in the training data storage unit 470 as training data. Note that the correct answer data here refers to, for example, a recognition result in pixel units that is properly recognized when object recognition processing is performed using a mask image. Note that the recognition result in pixel units may be, for example, annotation data in which a different label is assigned to each pixel for each object (instance) to be recognized.
[0063] The training unit 460 trains the object recognition unit using the training data stored in the training data storage unit 470, and generates a trained object recognition unit. The training unit 460 also sets the generated trained object recognition unit in the image processing device 130 that operates in the depalletizing phase.
[0064] <Functional Structure of the Training Department> Next, a detailed description will be given of the functional configuration of the training unit 460. Fig. 5 is a first diagram showing an example of the functional configuration of the training unit.
[0065] As shown in FIG. 5, the training unit 460 includes an object recognition unit 510 and a comparison / modification unit 520, and trains the object recognition unit 510 using training data 500.
[0066] The training data 500 includes information items such as a "mask image" and "correct answer data." A mask image is stored in the "mask image." Furthermore, a pixel-by-pixel recognition result that should be properly recognized based on the corresponding "mask image" is stored in the "correct answer data."
[0067] The object recognition unit 510 is configured by a DNN, and outputs output data by inputting a "mask image" of the training data 500. This output data may be, for example, segmentation data in which the probability of each of different labels for each object is assigned to each pixel.
[0068] The comparison / modification unit 520 compares the output data output from the object recognition unit 510 with the "correct answer data" of the training data 500, and updates the model parameters of the object recognition unit 510 based on the comparison result. In this way, the training unit 460 generates a trained object recognition unit 510.
[0069] <Details of the functional configuration of the image processing device in the depalletizing phase> Next, the functional configuration of the image processing device 130 in the depalletizing phase will be described in detail. Fig. 6 is a first diagram showing an example of the functional configuration of the image processing device in the depalletizing phase. As described above, an image processing program is installed in the image processing device 130, and in the depalletizing phase, the program is executed, and as shown in Fig. 6, ·3D information acquisition unit 410, A mask region identification unit 420, RGB image acquisition unit 430, Mask part 440, A trained object recognizer 650; Robot control unit 660, It functions as:
[0070] Of these, the three-dimensional information acquisition section 410, the mask region identification section 420, the RGB image acquisition section 430, and the mask section 440 have already been described with reference to FIG. 4, and therefore description thereof will be omitted here.
[0071] The trained object recognition unit 650 is a trained object recognition unit that has been trained by the training unit 460 using the training data 500. The trained object recognition unit 650 performs object recognition processing by inputting the mask image notified by the mask unit 440, and outputs the recognition result.
[0072] The robot control unit 660 generates a control command for controlling the operation of the robot 110 based on the recognition result output by the trained object recognition unit 650, and transmits the control command to the robot 110. When generating the control command based on the recognition result, the robot control unit 660 refers to the work rule information and generates the control command in accordance with the work rule information.
[0073] <Example of object recognition processing> Next, a specific example of the object recognition process by the image processing device 130 in the depalletizing phase will be described. Fig. 7 is a diagram showing a specific example of the object recognition process.
[0074] In Fig. 7, reference numeral 710 indicates an RGB image acquired by the RGB image acquisition unit 430. Also in Fig. 7, reference numeral 711 indicates processing range information previously received by the mask region identification unit 420, which is displayed on the RGB image indicated by reference numeral 710. As indicated by reference numeral 711, the processing range information indicates an area on the pallet 140 where the cardboard is located.
[0075] 7, reference numeral 720 denotes a mask image generated by performing a mask process on the RGB image denoted by reference numeral 710. In the mask image denoted by reference numeral 720, reference numeral 721 denotes a non-masked region identified by the mask region identifying unit 420, and reference numeral 722 denotes a masked region identified by the mask region identifying unit 420.
[0076] The non-masked region indicated by reference numeral 721 is the region of the top surface of the cardboard box with the highest height. In the example indicated by reference numeral 721, the heights of the two cardboard boxes are approximately the same, so the regions of the top surfaces of the two cardboard boxes are identified as the non-masked region.
[0077] In addition, the masked area indicated by the symbol 722 indicates that the color information associated with each pixel of the RGB image indicated by the symbol 710 other than the pixels included in the non-masked area (pixels included in the masked area) has been converted to color information indicating black.
[0078] Also, in FIG. 7, reference numeral 730 indicates that object recognition processing is performed using the mask image shown in reference numeral 720, and the cardboard box shown in reference numeral 731 and the cardboard box shown in reference numeral 732 are each recognized as different objects (different instances).
[0079] <Overall processing flow of the robot control system> Next, a description will be given of the overall process flow of the robot control system 100. Fig. 8 is a first flowchart showing the overall process flow of the robot control system.
[0080] In step S801, the robot control system 100 transitions to a training phase and executes training processing. Note that details of a flowchart of the training processing will be described later.
[0081] In step S802, the robot control system 100 transitions to the depalletizing phase and executes the depalletizing process. Details of the flowchart of the depalletizing process will be described later.
[0082] <Training process flow> Next, the flowchart of the training process (step S801 in FIG. 8) will be described in detail. FIG 9 is a first flowchart showing the flow of the training process.
[0083] In step S901, the image processing device 230 acquires, from the acquired data storage unit 220, an RGB image captured by the RGB camera 121 or a virtual RGB image generated by the CG simulator 210. In addition, the image processing device 230 acquires, from the acquired data storage unit 220, three-dimensional information captured by the depth camera 122 or virtual three-dimensional information generated by the CG simulator 210.
[0084] In step S902, the image processing device 230 specifies a non-mask area in the RGB image based on the processing range information and the work rule information and on the three-dimensional information.
[0085] In step S903, the image processing apparatus 230 performs a masking process on the color information associated with pixels other than the pixels included in the non-masked area in the RGB image, and generates a masked image.
[0086] In step S904, the image processing apparatus 230 acquires correct answer data.
[0087] In step S905, the image processing apparatus 230 associates the masked image with the correct answer data, and generates training data.
[0088] In step S906, the image processing apparatus 230 trains the object recognition unit using the training data.
[0089] In step S907, the image processing apparatus 230 determines whether or not the training end condition is satisfied. If it is determined in step S907 that the training end condition is not satisfied (if NO in step S907), the process returns to step S901.
[0090] On the other hand, if it is determined in step S907 that the training end condition is satisfied (if YES in step S907), the process proceeds to step S908.
[0091] In step S908, the image processing apparatus 230 sets the trained object recognition unit, and ends the training process.
[0092] <Flow of the de-paretoizing process> Next, the details of the flowchart of the de-paretoizing process (step S802 in FIG. 8) will be described. FIG. 10 is a first flowchart showing the flow of the de-paretoizing process.
[0093] In step S1001, the image processing apparatus 130 acquires the RGB image captured by the RGB camera 121 and the three-dimensional information captured by the depth camera 122.
[0094] In step S1002, the image processing device 130 identifies a non-masked region in the RGB image within the range specified by the processing range information, based on three-dimensional information in accordance with the work rule information.
[0095] In step S1003, the image processing device 130 performs mask processing on color information associated with pixels other than those included in the non-mask region in the RGB image, to generate a mask image.
[0096] In step S1004, the image processing device 130 performs object recognition processing by inputting the generated mask image to a trained object recognition unit.
[0097] In step S1005, the image processing device 130 identifies cardboard boxes to be depalletized based on the result of the object recognition process. For example, when a plurality of cardboard boxes are recognized as a result of the object recognition process, the image processing device 130 identifies the plurality of cardboard boxes as the depalletizing targets.
[0098] In step S1006, the image processing device 130 generates a control command for depalletizing the cardboard boxes identified in the mask image as targets for depalletizing in accordance with the three-dimensional information and work rule information for the identified cardboard boxes, and outputs the control command to the robot 110. For example, when multiple cardboard boxes are identified as targets for depalletizing, the image processing device 130 determines the order of depalletizing in accordance with the work rule information, and generates a control command corresponding to the three-dimensional information of the cardboard boxes to be depalletized.
[0099] In step S1007, the image processing device 130 determines whether or not there is a next cardboard box to be depalletized. Specifically, the image processing device 130 determines whether or not there is a next cardboard box to be depalletized on the pallet 140 after all of the cardboard boxes to be depalletized identified in step S1005 have been depalletized.
[0100] In step S1007, if it is determined that there is a next cardboard box to be depalletized (YES in step S1007), the process proceeds to step S1008.
[0101] In step S1008, the image processing device 130 determines whether it is time to perform object recognition processing. If it is determined in step S1008 that it is not time to perform object recognition processing (NO in step S1008), the image processing device 130 waits until it is time to perform object recognition processing.
[0102] On the other hand, if it is determined in step S1008 that it is time to perform the object recognition process (YES in step S1008), the process returns to step S1001.
[0103] On the other hand, if it is determined in step S1007 that there is no next cardboard box to be depalletized (NO in step S1007), the depalletizing process ends.
[0104] <Examples of depalletizing processing> Next, a specific example of the depalletizing process by the robot control system 100 will be described. Fig. 11 is a diagram showing a specific example of the depalletizing process. As shown in Fig. 11(a), when four cardboard boxes are stacked on the pallet 140, the robot control system 100 recognizes the cardboard box 1101 as the cardboard box with the highest height. Fig. 11(b) shows the state in which the recognized cardboard box 1101 is picked up.
[0105] Next, as shown in Fig. 11(c), after picking up the cardboard box 1101, three cardboard boxes are stacked on the pallet 140, and the robot control system 100 recognizes the cardboard box 1102 as the tallest cardboard box. Fig. 11(d) shows the state in which the recognized cardboard box 1102 is picked up.
[0106] Next, as shown in Fig. 11(e), two cardboard boxes are placed on the pallet 140 after the cardboard box 1102 is picked up, and the robot control system 100 recognizes the cardboard box 1103 as the tallest cardboard box. Fig. 11(f) shows the state in which the recognized cardboard box 1103 is picked up.
[0107] Next, as shown in Fig. 11(g), one cardboard box is placed on the pallet 140 after picking up the cardboard box 1103, and the robot control system 100 recognizes the cardboard box 1104 as the tallest cardboard box. Fig. 11(h) shows the state in which the recognized cardboard box 1104 is picked up.
[0108] In this way, the robot control system 100 can realize depalletizing according to the work rule of "picking up cardboard boxes one by one in order from the top".
[0109] <Summary> As is clear from the above description, the robot control system 100 according to the first embodiment has the following features: For a given space in which one or more cardboard boxes are stacked on a pallet, an RGB image of the given space and 3D information of the given space are obtained by photographing the one or more cardboard boxes. Based on the acquired 3D information, masking is performed on some of the color information in the RGB image. - Object recognition is performed using the masked image after the masking process.
[0110] As a result, the robot control system 100 according to the first embodiment can improve the recognition accuracy in the object recognition process.
[0111] [Second embodiment] In the first embodiment, the mask unit 440 performs mask processing on color information associated with pixels included in the mask area. However, the target on which the mask unit 440 performs mask processing is not limited to color information, and the mask unit 440 may perform mask processing on color information and three-dimensional information. The second embodiment will be described below, focusing on the differences from the first embodiment.
[0112] <Details of the functional configuration of the image processing device in the training phase> First, the details of the functional configuration of the image processing device 230 in the training phase will be described. Fig. 12 is a second diagram showing an example of the functional configuration of the image processing device in the training phase. As in the first embodiment, an image processing program is installed in the image processing device 230, and by executing the program, each function similar to that of the first embodiment is realized in the training phase.
[0113] 12. Note that the difference from FIG. 4 described in the first embodiment is that in the case of FIG. 12, a three-dimensional information acquisition unit 410 notifies a mask region identification unit 420 and a mask unit 440 of three-dimensional information.
[0114] 4 described in the first embodiment is that the mask unit 440 performs mask processing on the color information and 3D information associated with the pixels included in the mask area to generate a mask image and mask 3D information. The mask processing performed on the 3D information includes deleting the 3D information (x coordinate, y coordinate, z coordinate) associated with the pixels included in the mask area, etc.
[0115] Also, the difference from FIG. 4 described in the first embodiment is that the training data generating unit 450 generates training data by associating the mask image and mask 3D information notified by the masking unit 440 with the ground truth data, and stores the training data in the training data storage unit 470.
[0116] <Functional Structure of the Training Department> Next, the details of the functional configuration of the training unit 460 in the second embodiment will be described. FIG. 13 is a second diagram showing an example of the functional configuration of the training unit. The difference from FIG. 5 described in the first embodiment is that the training data 1300 includes "3D mask information" as an information item, and the 3D mask information is stored. Also, the difference from FIG. 5 described in the first embodiment is that the "mask image" and "3D mask information" of the training data 1300 are input to the object recognition unit 510.
[0117] <Details of the functional configuration of the image processing device in the depalletizing phase> Next, the details of the functional configuration of the image processing device 130 in the depalletizing phase will be described. Fig. 14 is a second diagram showing an example of the functional configuration of the image processing device in the depalletizing phase. As in the first embodiment, an image processing program is installed in the image processing device 130, and by executing the program, each function similar to that of the first embodiment is realized in the depalletizing phase.
[0118] 14. Note that the difference from FIG. 6 described in the first embodiment is that in the case of FIG. 14, a three-dimensional information acquisition section 410 notifies a mask region identification section 420 and a mask section 440 of three-dimensional information.
[0119] Also, the difference from FIG. 6 described in the first embodiment is that the mask unit 440 performs mask processing on the color information and 3D information associated with the pixels included in the mask area to generate a mask image and mask 3D information.
[0120] Also, the difference from FIG. 6 described in the first embodiment is that the trained object recognition unit 650 performs object recognition processing based on the mask image and mask three-dimensional information notified by the mask unit 440, and outputs the recognition result.
[0121] <Training process flow> Next, the details of the flowchart of the training process in the second embodiment will be described. Fig. 15 is a second flowchart showing the flow of the training process. The difference from the first flowchart described using Fig. 9 is steps S1501 and S1502.
[0122] In step S1501, the image processing device 230 performs mask processing on color information and three-dimensional information associated with pixels other than those included in the non-mask region, and generates a mask image and mask three-dimensional information.
[0123] In step S1502, the image processing device 230 generates training data by associating the mask image and the mask three-dimensional information with the correct answer data.
[0124] <Depalletizing process flow> Next, the details of the flowchart of the depalletizing process in the second embodiment will be described. Fig. 16 is a second flowchart showing the flow of the depalletizing process. The differences from the first flowchart described using Fig. 10 are steps S1601 and S1602.
[0125] In step S1601, the image processing device 130 performs mask processing on color information and three-dimensional information associated with pixels other than those included in the non-mask region, and generates a mask image and mask three-dimensional information.
[0126] In step S1602, the image processing device 130 performs object recognition processing by inputting the generated mask image and the three-dimensional mask information to a trained object recognition unit.
[0127] <Summary> As is apparent from the above description, the robot control system 100 according to the second embodiment has the following features: For a given space in which one or more cardboard boxes are stacked on a pallet, an RGB image of the given space and 3D information of the given space are obtained by photographing the one or more cardboard boxes. Based on the acquired 3D information, masking is performed on some of the color information and 3D information of the RGB image. - Object recognition processing is performed using the mask image after mask processing and the mask 3D information.
[0128] As a result, the robot control system 100 according to the second embodiment can improve the recognition accuracy in the object recognition process.
[0129] [Third embodiment] In the above first embodiment, no mention is made of the shooting conditions (e.g., white balance, exposure, focus, etc.) of the RGB camera 121 when performing the training process (step S801 in FIG. 8), and it has been described that these are appropriately adjusted.
[0130] In contrast to this, in the third embodiment, a shooting condition adjustment phase is provided before the training process, and a case will be described in which the shooting conditions of the RGB camera 121 are adjusted. The description will be centered on the differences from the first embodiment.
[0131] <Functional configuration of the image processing device in the shooting condition adjustment phase> First, the details of the functional configuration of the image processing device 130 in the shooting condition adjustment phase will be described. Fig. 17 is a diagram showing an example of the functional configuration of the image processing device in the shooting condition adjustment phase. In the shooting condition adjustment phase, the image processing device 130 executes an image processing program, as shown in Fig. 17, ·3D information acquisition unit 410, A mask region identification unit 420, RGB image acquisition unit 430, Mask part 440, Shooting condition adjustment unit 1250, It functions as:
[0132] Among these, the functions of the three-dimensional information acquisition unit 410, the mask region specification unit 420, the RGB image acquisition unit 430, and the mask unit 440 have already been described, and therefore will not be described here. The mask unit 440 performs mask processing on color information associated with pixels included in the mask region of the RGB image, and notifies the shooting condition adjustment unit 1750 of the generated mask image.
[0133] The photographing condition adjustment unit 1750 adjusts the photographing conditions based on the mask image notified by the mask unit 440. The photographing conditions adjusted by the photographing condition adjustment unit 1750 include the white balance, exposure, focus, and the like of the RGB camera 121.
[0134] Moreover, the photographing condition adjustment unit 1750 transmits the adjusted photographing conditions to the RGB camera 121 and sets them.
[0135] In this way, in the robot control system 100 according to the third embodiment, the shooting conditions of the RGB camera 121 are adjusted based on the mask image. This makes it possible to set shooting conditions suitable for object recognition processing. As a result, the robot control system 100 according to the third embodiment can improve the recognition accuracy in the object recognition processing.
[0136] <Overall processing flow of the robot control system> Next, the overall process flow of the robot control system 100 according to the third embodiment will be described. Fig. 18 is a second flowchart showing the overall process flow of the robot control system. The difference from the first flowchart shown in Fig. 8 is that an imaging condition adjustment process (step S1801) is included before the training process (step S801).
[0137] In step S1801, the robot control system 100 transitions to the image capturing condition adjustment phase and executes the image capturing condition adjustment process. Details of the image capturing condition adjustment process will be described with reference to FIG.
[0138] <Flow of shooting condition adjustment process> FIG. 19 is a flowchart showing the flow of the shooting condition adjustment process.
[0139] In step S1901, the image processing device 130 acquires an RGB image captured by the RGB camera 121 and three-dimensional information captured by the depth camera 122.
[0140] In step S1902, the image processing device 130 specifies a non-mask area in the RGB image based on the processing range information and the work rule information and on the three-dimensional information.
[0141] In step S1903, the image processing device 130 performs mask processing on color information associated with pixels other than those included in the non-mask region in the RGB image, to generate a mask image.
[0142] In step S1904, the image processing device 130 adjusts the shooting conditions of the RGB camera 121 based on the mask image.
[0143] In step S1905, the image processing device 130 acquires an RGB image captured by the RGB camera 121 under the shooting conditions adjusted in step S1904.
[0144] In step S1906, the image processing device 130 performs mask processing on color information associated with pixels other than the pixels included in the non-mask region identified in step S1902 in the RGB image acquired in step S1905. In this way, the image processing device 130 generates a mask image.
[0145] In step S1907, the image processing device 130 evaluates the mask image generated in step S1906.
[0146] In step S1908, the image processing device 130 judges whether the photographing conditions have been optimized based on the evaluation result of the mask image, and if it is judged that the photographing conditions have not been optimized (NO in step S1908), the process returns to step S1904. At this time, if the arrangement of the cardboard boxes stacked on the pallet 140 is to be changed, the process returns to step S1901.
[0147] On the other hand, if it is determined in step S1908 that the optimization has been performed (YES in step S1908), the process proceeds to step S1909.
[0148] In step S1909, the image processing device 130 transmits and sets the optimized shooting conditions to the RGB camera 121, and ends the shooting condition adjustment process. As a result, in the training phase and the depalletizing phase, it is possible to acquire RGB images shot by the RGB camera 121 under the optimized shooting conditions.
[0149] <Summary> As is apparent from the above description, the robot control system 100 according to the third embodiment has the following features: For a given space in which one or more cardboard boxes are stacked on a pallet, an RGB image of the given space and 3D information of the given space are obtained by photographing the one or more cardboard boxes. Based on the acquired 3D information, masking is performed on some of the color information in the RGB image. · Adjust the shooting conditions of the RGB camera using the mask image after mask processing.
[0150] As a result, the robot control system 100 according to the third embodiment can achieve shooting conditions suitable for object recognition processing, thereby improving the recognition accuracy in the object recognition processing.
[0151] [Fourth embodiment] In the above embodiments, the RGB camera 121 and the depth camera 122 are described as being fixedly attached to the ceiling of the space in which the robot 110 is disposed, but the attachment positions of the RGB camera 121 and the depth camera 122 are not limited to the ceiling of the space. In addition, the attachment positions of the RGB camera 121 and the depth camera 122 are not limited to directly above the pallet 140. Furthermore, the attachment destination of the RGB camera 121 and the depth camera 122 is not limited to a non-movable object, and may be a movable object (for example, the robot 110).
[0152] In addition, in each of the above embodiments, the three-dimensional information acquisition unit 410 has been described as acquiring three-dimensional coordinates (x coordinate, y coordinate, z coordinate) captured by the depth camera 122 as the three-dimensional information. However, the three-dimensional information acquisition unit 410 may acquire three-dimensional information other than the three-dimensional coordinates (x coordinate, y coordinate, z coordinate). The three-dimensional information other than the three-dimensional coordinates (x coordinate, y coordinate, z coordinate) includes, for example, point cloud information, mesh information, and the like.
[0153] In the above first and second embodiments, in the depalletizing phase, the image processing device 130 has been described as having the robot control unit 660. However, in the depalletizing phase, the robot control unit 660 may be realized within the robot 110.
[0154] In the above first and second embodiments, in the depalletizing phase, the image processing device 130 has been described as having the trained object recognition unit 650. However, in the depalletizing phase, the trained object recognition unit 650 may be realized in a device other than the image processing device 130.
[0155] In the third embodiment, in the shooting condition adjustment phase, the image processing device 130 has been described as having the shooting condition adjustment unit 1750. However, in the shooting condition adjustment phase, the shooting condition adjustment unit 1750 may be realized in a device other than the image processing device 130.
[0156] In the above first and second embodiments, the object recognition unit is trained using masked data (masked image or masked three-dimensional information) as training data, and the masked data (masked image or masked three-dimensional information) is input to the trained object recognition unit. However, the object recognition unit may be trained using unmasked data (RGB image or masked three-dimensional information) as training data, and the masked data (masked image or masked three-dimensional information) may be input to the trained object recognition unit to perform object recognition processing.
[0157] [Other embodiments] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.
[0158] In addition, in this specification (including claims), when expressions such as "data as input / based on / according to / in response to" (including similar expressions) are used, unless otherwise specified, it includes cases where various data itself is used as input, and cases where various data that have been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) are used as input. In addition, when it is stated that a result is obtained "based on / according to / in response to data," it includes cases where the result is obtained based only on the data, and may also include cases where the result is obtained by being influenced by other data, causes, conditions, and / or states other than the data. In addition, when it is stated that "data is output," it includes cases where various data itself is used as output, and cases where various data that have been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) are output.
[0159] In addition, when the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms including any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, physically connection / coupling, etc. The terms should be interpreted appropriately depending on the context in which the terms are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in the terms without any restriction.
[0160] In addition, in this specification (including claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A is configured / set to actually perform operation B. For example, when element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). In addition, when element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.
[0161] In addition, when terms implying inclusion or possession (e.g., "comprising / including," "having," etc.) are used in this specification (including the claims), they are intended as open-ended terms that include cases in which something other than the object indicated by the object of the term is contained or possessed. When the object of such terms implying inclusion or possession is an expression that does not specify a quantity or suggests a singular number (an expression using the article "a" or "an"), the expression should be construed as not being limited to a specific number.
[0162] In addition, even if expressions such as "one or more" or "at least one" are used in some places in this specification (including the claims) and expressions that do not specify a quantity or suggest a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.
[0163] In addition, when a particular advantage / result is described in this specification for a particular configuration of an embodiment, it should be understood that the same effect can also be obtained for one or more other embodiments having the same configuration, unless there is a specific reason to the contrary. However, it should be understood that the presence or absence of the effect generally depends on various causes, conditions, and / or states, and that the effect is not necessarily obtained by the configuration. The effect is merely obtained by the configuration described in the embodiment when various causes, conditions, and / or states are satisfied, and the effect is not necessarily obtained in the invention according to the claim that specifies the configuration or a similar configuration.
[0164] In this specification (including the claims), when the term "optimize" or "optimization" is used, it includes determining a global optimum, determining an approximation of a global optimum, determining a local optimum, and determining an approximation of a local optimum, and should be interpreted appropriately according to the context in which the term is used. It also includes determining an approximation of these optimum values probabilistically or heuristically.
[0165] In addition, in this specification (including claims), when a plurality of pieces of hardware perform a predetermined process, each piece of hardware may cooperate to perform the predetermined process, or a portion of the hardware may perform all of the predetermined process. Also, a portion of the hardware may perform a portion of the predetermined process, and another piece of hardware may perform the remainder of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. The hardware may include an electronic circuit, a device including an electronic circuit, and the like.
[0166] Furthermore, in this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices (memories) may store only a portion of the data, or may store the entire data.
[0167] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, replacements, partial deletions, etc. are possible within the scope of the conceptual idea and intent of the present invention derived from the contents defined in the claims and their equivalents. For example, in all the above-mentioned embodiments, when numerical values or formulas are used in the explanation, they are shown as examples and are not limited to these. In addition, the order of each operation in the embodiments is shown as an example and is not limited to these.
Claims
1. one or more memories; one or more processors; The one or more processors: acquiring a gradation image of a predetermined space including one or more objects and three-dimensional information of the predetermined space; masking a portion of the grayscale image based on the three-dimensional information; performing a predetermined process using the masked gradation image; An image processing device that executes the above.
2. Masking the portion of the gradient image includes: identifying the portion of the grayscale image based on the three-dimensional information; a process of masking the portion of the gradation image identified by the identifying process; The image processing device according to claim 1 ,
3. The image processing apparatus according to claim 2 , wherein the specifying process specifies, as the part of the gradation image, two-dimensional coordinates of the gradation image that correspond to specific three-dimensional coordinates in the predetermined space.
4. The image processing device according to claim 2 , wherein the masking process masks gradation information of the partial pixels of the gradation image.
5. Masking the portion of the gradient image includes: identifying an area of an object that satisfies a predetermined condition from among the one or more objects based on the three-dimensional information; a process of masking gradation information associated with pixels other than the pixels included in the specified object area, out of the gradation information associated with each pixel of the gradation image; The image processing device according to claim 1 ,
6. The one or more processors further include: performing masking of a portion of the three-dimensional information based on the three-dimensional information; The image processing apparatus according to claim 5 , wherein performing the predetermined processing includes performing the predetermined processing using the masked grayscale image and the masked three-dimensional information.
7. Masking a portion of the three-dimensional information includes: a process of masking three-dimensional information associated with pixels other than pixels included in an area of the object that satisfies the predetermined condition, among three-dimensional information associated with each pixel of the gradation image; The image processing device according to claim 6 ,
8. The image processing apparatus according to claim 6 or 7, wherein the region of the object that satisfies the predetermined condition is a region of the object that has predetermined three-dimensional information.
9. The image processing device according to claim 8 , wherein the region of the object having the predetermined three-dimensional information is a region of an object of which a height of an upper surface has the predetermined three-dimensional information, among the one or more objects.
10. The image processing device according to claim 8 , wherein the predetermined processing includes an object recognition processing or an adjustment processing for adjusting a shooting condition when the gradation image is shot.
11. The image processing device according to claim 10 , wherein, when the predetermined process is an object recognition process, a region of an object that satisfies the predetermined condition is identified within a range in which the one or more objects can be placed.
12. The image processing device according to claim 10 , wherein the object recognition process is performed using a deep neural network.
13. The image processing device according to claim 12 , wherein the deep neural network is trained using a masked grayscale image as training data.
14. The image processing device according to claim 12 , wherein the deep neural network is trained using a masked grayscale image and a masked three-dimensional information as training data.
15. The image processing device according to claim 10 , wherein the photographing conditions adjusted in the adjustment process include a white balance, an exposure, and a focus of an image generating device that photographs the gradation image.
16. one or more processors, acquiring a gradation image of a predetermined space including one or more objects and three-dimensional information of the predetermined space; masking a portion of the grayscale image based on the three-dimensional information; performing a predetermined process using the masked gradation image; An image processing method that performs
17. one or more processors, acquiring a gradation image of a predetermined space including one or more objects and three-dimensional information of the predetermined space; masking a portion of the grayscale image based on the three-dimensional information; performing a predetermined process using the masked gradation image; An image processing program for executing the above.
18. An image processing device according to any one of claims 1 to 15, an image generating device for generating the gradation image; a three-dimensional information generating device for generating the three-dimensional information; a robot for depalletizing the one or more objects; and A robot control system having the following:
Citation Information
Patent Citations
Production of electrically conductive film
JP1987011734A
Operation system, control device and program
JP2020075340A