Method for generating training data used for learning about conveying devices.
The method generates training data using a 3D model and virtual images to improve forklift pallet detection, addressing the challenges of fork alignment and reducing training costs and effort, thereby enhancing operational efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SUMITOMO NACCO FORKLIFT CO LTD
- Filing Date
- 2022-07-01
- Publication Date
- 2026-05-13
Smart Images

Figure 0007857622000001 
Figure 0007857622000002 
Figure 0007857622000003
Abstract
Description
[Technical Field]
[0001] This invention relates to a conveying device for transporting pallets. [Background technology]
[0002] Forklifts and other conveying devices are designed to transport pallets that have a predetermined shape. Pallets are provided with holes (fork pockets), and the conveying device is equipped with forks that can be inserted into these fork pockets.
[0003] The operator of the conveying device needs to adjust the height of the forks, steer the device, and insert the forks into the fork pockets.
[0004] However, for an inexperienced operator, inserting the forks properly into the fork pockets is not easy. If the forks collide with a place other than the fork pocket, the load on the pallet may collapse. Even if the forks are inserted into the fork pockets, if the forks hit the inner walls of the pockets, the pallet may shift or undesirable stress may be applied to the forks. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2017-151650 [Overview of the project] [Problems that the invention aims to solve]
[0006] The inventors considered equipping a forklift with a detection device capable of detecting specific locations on a pallet. The detection device comprises a sensor that acquires images and a classifier (also called a discriminator) configured to detect specific locations based on the sensor's output image. The classifier is constructed using machine learning based on pre-acquired training data or instructional data. Training data is typically image data containing specific parts. In order for the classifier to acquire sufficient discriminative power, a vast amount of training data is required, obtained by taking images with varying shooting directions and distances.
[0007] If a forklift handles a wide variety of pallets, the cost of training increases dramatically because training data must be prepared for each type of pallet. In particular, if service personnel or users (hereinafter referred to as "humans") take numerous images of various types of pallets in the forklift's operating environment to generate training data, the amount of human effort required becomes substantial, meaning the cost of training becomes extremely high.
[0008] This invention has been made in view of the aforementioned problems, and one exemplary objective of a certain embodiment of it is to provide a technology that can reduce the effort and time required for learning. [Means for solving the problem]
[0009] One aspect of the present invention relates to a method for generating training data used for training a conveying device that transports pallets. The conveying device comprises an active distance image sensor that irradiates light and measures reflected light, and a processing unit that (i) can acquire a distance image in which pixel values represent distance based on the output of the distance image sensor, and (ii) has a classifier configured by machine learning using the training data so that it can identify a pallet or a specific part of the pallet that is a detection target based on the distance image. The method for generating training data includes the following steps. • Generate a 3D model of the palette. · Place a 3D model of a pallet in a 3D virtual space and generate a virtual distance image that would be obtained when the pallet is measured by a distance image sensor. · Generate a virtual luminance image showing the intensity distribution of the reflected light that the distance image sensor would receive. · Modify the virtual distance image based on the virtual luminance image. · Generate training data based on the modified virtual distance image. The virtual distance image is generated using parameters indicating the reflection characteristics of diffuse reflection and specular reflection in the measurement target. The values of the parameters indicating the reflection characteristics are determined using a measured luminance image obtained by actually measuring the measurement target placed in the real space with a distance image sensor.
[0010] This conveying device maps 3D point cloud data obtained by an active distance image sensor such as a TOF (Time Of Flight) sensor, LiDAR (Light Detection and Ranging), or a structured light camera to a 2D distance image, and by configuring a discriminator that takes the distance image as input, the specific accuracy of the pallet or its part can be improved.
[0011] Here, since the shapes and materials of pallets vary widely, it is not easy to prepare training data for use in learning the discriminator by photographing the actual pallets with a distance image sensor assuming all situations for all pallets. According to this aspect, by utilizing the 3D model, training data can be efficiently generated and learning can be made more efficient.
[0012] The inventors have found that the accuracy of distance images generated by actual distance image sensors is affected by the intensity of reflected light received by the image sensor / photodetector of the distance image sensor. Therefore, by generating a virtual luminance image and correcting the virtual distance image based on the virtual distance image, it is possible to generate an image that is close to the measured distance image generated by an actual distance image sensor. Furthermore, by using the measured luminance image obtained when measured with an actual distance image sensor to determine the values of parameters that indicate the reflection characteristics of the object being measured, and then generating a virtual luminance image using those parameter values, the calculation accuracy of the virtual luminance image can be further improved.
[0013] The parameters indicating the reflective properties may be determined for each material and / or color of the surface being measured. Since palettes come in a wide variety of materials and colors, determining the reflective property parameters for each palette material and color allows for the efficient generation of training data that can accommodate a wide variety of palettes.
[0014] The step of correcting the virtual distance image may include superimposing an error on the pixel values of the virtual distance image based on the pixel values of the virtual luminance image. For example, in areas with extremely high or extremely low luminance, an uncertainty corresponding to the pixel values of the luminance image may be superimposed on the pixel values of the virtual distance image.
[0015] The error characteristics may also be defined as the relationship between the pixel values and standard deviation of the luminance image. In this case, the original virtual distance image's pixel values, the mean value, and the standard deviation based on the luminance image's pixel values are given to a random number function, and the resulting random numbers can be used as the pixel values of the corrected virtual distance image.
[0016] Another aspect of this disclosure is also a method for generating training data. The method for generating training data includes the following steps: • Generate a 3D model of the palette. • A 3D model of the palette is placed in a 3D virtual space, and a virtual distance image is generated that would be obtained if the palette were measured by a distance image sensor. • Generate a virtual luminance image that shows the intensity distribution of reflected light that the distance image sensor will receive. • The virtual distance image is corrected based on missing data generated using the virtual distance image and virtual brightness image. • Generate training data based on the corrected virtual distance image.
[0017] Missing data tends to occur frequently at the edges of objects, and also frequently in areas where the reflected light incident on the depth image sensor is either too weak or too strong. Therefore, by generating missing data using both a virtual depth image correlated with the edges of objects and a virtual brightness image correlated with the intensity of reflected light, the accuracy of missing data estimation can be improved compared to using only one of them. By replacing areas in the virtual depth image where pixel loss is estimated with missing data, an image closer to the actual depth image can be generated.
[0018] Missing data may be generated by a missing data generator configured using machine learning, which takes as input measured distance images and measured brightness images obtained when a measurement target placed in real space is measured using a distance image sensor. By using machine learning with measured distance images and brightness images, the accuracy of estimating missing data can be improved.
[0019] The missing data generator may consist of a convolutional neural network and machine learning performed under the constraint that the kernel value at the center of the convolutional kernel is fixed. The fixed value may be 0. By fixing the kernel value at the center of the convolutional kernel, the influence of missing pixels in the measured depth and brightness images can be reduced, and the accuracy of missing data estimation can be improved.
[0020] The method for generating training data may further include the step of generating an error-superimposed distance image by superimposing an error on the pixel values of a virtual distance image based on the pixel values of a virtual brightness image. Missing data may be generated using the error-superimposed distance image, and the corrected virtual distance image may be generated by replacing the pixel values of the error-superimposed distance image with the missing data. By generating missing data using the error-superimposed distance image, the estimation accuracy of missing data can be improved. Furthermore, by applying missing data to the error-superimposed distance image, it is possible to generate a virtual distance image that is close to an actual measured distance image, taking into account both the effects of noise (error) and the effects of missing data.
[0021] Furthermore, any combination of the above components, or any substitution of components or expressions of the present invention between methods, apparatus, systems, etc., is also valid as an embodiment of the present invention. [Effects of the Invention]
[0022] According to this invention, the effort and time required for learning can be reduced. [Brief explanation of the drawing]
[0023] [Figure 1] This is a perspective view showing the external appearance of a forklift, which is one form of a conveying device. [Figure 2] This is a diagram showing an example of a forklift driver's seat. [Figure 3] Figures 3(a) to 3(h) show the pallets that the forklift is transporting. [Figure 4] This is a functional block diagram of a forklift according to an embodiment. [Figure 5] Figures 5(a) and 5(b) illustrate the installation locations of the distance image sensors. [Figure 6] This figure shows an example of a distance image. [Figure 7] Figures 7(a) and 7(b) schematically show the specific areas that are the target of detection. [Figure 8] This figure shows a distance image of a specific area in Figure 7(a). [Figure 9] This is a flowchart for generating training data used to train a classifier. [Figure 10] This figure shows a three-dimensional virtual space where the palette and distance image sensor models are placed. [Figure 11] Figures 11(a) to (c) illustrate the generation of training data based on background distance images. [Figure 12] This is a flowchart illustrating the method for generating training data when using an active-type distance image sensor. [Figure 13] This diagram illustrates the generation of an IR brightness image. [Figure 14] Figures 14(a) to 14(c) show the generation of training data based on the flowchart in Figure 12. [Figure 15] This figure shows an example of the relationship between the pixel values of the IR brightness image and the error superimposed on the virtual distance image. [Figure 16] Figures 16(a) and (b) show the correction of the virtual distance image based on the relationship in Figure 15. [Figure 17] This figure shows an example of input and output images for a missing data generator. [Figure 18] This diagram schematically shows an example of the network structure of a missing data generator. [Figure 19] Figure 19(a) shows an example of an actual distance image, and Figure 19(b) shows an example of a virtual distance image. [Figure 20] This diagram schematically shows the convolution kernel used in the convolution process of the missing data generator. [Figure 21] This is a flowchart of the method for generating training data according to Example 3. [Modes for carrying out the invention]
[0024] The present invention will be described below with reference to the drawings, based on preferred embodiments. The same or equivalent components, members, and processes shown in each drawing will be denoted by the same reference numerals, and redundant descriptions will be omitted as appropriate. Furthermore, the embodiments are illustrative and not limiting to the invention, and not all features or combinations thereof described in the embodiments are necessarily essential to the invention.
[0025] Figure 1 is a perspective view showing the external appearance of a forklift, which is one embodiment of a transport device. The forklift 600 comprises a chassis 602, forks 604, a lift 606, a mast 608, and wheels 610 and 612. The mast 608 is located in front of the chassis 602. The lift 606 is driven by a power source such as a hydraulic actuator (not shown in Figure 1) and moves up and down along the mast 608. Forks 604 for supporting loads are attached to the lift 606. The forklift 600 transports pallets with the two forks 604 inserted into holes (fork pockets) in the pallets (not shown).
[0026] Figure 2 shows an example of a forklift driver's seat 700. The driver's seat 700 includes an ignition switch 702, a steering wheel 704, a lift lever 706, an accelerator pedal 708, a brake pedal 710, a dashboard 714, and a forward / reverse lever 712.
[0027] The ignition switch 702 is a switch for starting the forklift 600. The steering wheel 704 is an operating means for steering the forklift 600. The lift lever 706 is an operating means for moving the lifting body 606 up and down. The accelerator pedal 708 is an operating means for controlling the rotation of the wheels for driving, and the operation of the forklift 600 is controlled by the operator adjusting the amount the pedal is pressed. When the operator presses the brake pedal 710, the brakes are applied. The forward / reverse lever 712 is a lever for switching the direction of travel of the forklift 600 between forward and reverse. In addition, an inching pedal (not shown) may be provided.
[0028] The operator must control the position of the vehicle by operating the steering wheel 704, accelerator pedal 708, and brake pedal 710, and control the position of the forks 604 by operating the lift lever 706.
[0029] Figures 3(a) to 3(h) show pallets that are transported by forklift 600. Pallet 800 has a top plate 806, holes 802, and beams 804. The shape of pallet 800 is defined by JIS standards, and there are many different shapes depending on the purpose and use. Figures 3(a) to 3(h) are denoted as S2, SU2, DP4, R4, RW2, D2, DU2, and D4 respectively in JIS standards. As shown in Figures 3(a) to 3(h), there are various shapes and structures of pallets, and it is not easy for an inexperienced operator to accurately operate the multiple input means (704, 706, 708, 710, 706) and accurately insert the forks into the fork pockets.
[0030] The following describes a configuration that can assist in the scooping operation of pallets. Figure 4 is a functional block diagram of a forklift 300 according to an embodiment. Figure 4 shows only the blocks related to assisting in the scooping operation, and other blocks are omitted.
[0031] The forklift 300 is equipped with a distance image sensor 302, a camera 303, a notification means 304, and a controller 310. The distance image sensor 302 is mounted on the front of the forklift 300 and acquires distance information in front of the vehicle. The distance image sensor 302 can be a ToF camera, LiDAR, a sensor using structured light, a stereo camera, etc., although this is not limited to the above. The distance image sensor 302 generates distance measurement data D1 which includes 3D point cloud data of structures and objects located in front of the vehicle.
[0032] Figures 5(a) and 5(b) illustrate the installation locations of the distance image sensor 302. Figure 5(a) is a top view. Since it is necessary to detect the position of the hole into which the fork should be inserted, it is preferable to place the distance image sensor 302 in the center of the two forks in the left-right direction. This allows the distance image sensor 302 to photograph the pallet from the front when the fork is inserted into the pallet. In particular, when detecting a specific part including a beam, as will be described later, placing the distance image sensor 302 in the center of the fork makes it possible to successfully capture a distance image of the specific part.
[0033] Figure 5(b) is a side view. In the vertical direction, it may be at the same height as the fork 604 as shown by 302c, but it is preferable to position it slightly above 302a (or below 302b). This makes it possible to reliably grip the upper surface S1 (or lower surface S2) of the top plate 806 of the pallet 800. This position also has the advantage of being less likely to be obstructed by the fork or other vehicle structures during the final adjustment phase of the fork height.
[0034] Returning to Figure 4, camera 303 captures the area in front of the forklift 300. The image data D3 captured by camera 303 is displayed on display 306. In Figure 4, camera 303 and distance image sensor 302 are shown as separate functional blocks, but this does not necessarily mean that they are hardware-independent of the distance image sensor 302; the distance image sensor 302 may also function as camera 303. For example, if the distance image sensor 302 is a TOF camera or a stereo camera, the monocular image obtained first from it (infrared intensity distribution, visible light monochrome image, or visible light color image) can be used as image data D3.
[0035] The controller 310 provides integrated control of the forklift 300. The controller 310 includes a distance image acquisition unit 330 and a classifier 340. The distance image acquisition unit 330 generates a distance image D2 based on distance measurement data D1, which includes point cloud data. The distance image D2 is image data having multiple pixels vertically and horizontally, where the pixel values represent distance. Figure 6 shows an example of a distance image. Note that if the output of the distance image sensor 302 is distance data, the processing of the distance image acquisition unit 330 is simplified or omitted.
[0036] Returning to Figure 4, the classifier 340 is configured using machine learning to identify the palette or a specific part of the palette based on the distance image D2.
[0037] The processing unit 350 identifies the location of the holes in the pallet based on the detection result of the classifier 340. Based on the detection result, it controls the notification means 304 and notifies the operator. The notification means 304 may include a display 306 and a speaker 308. The display 306 is used to visually notify the operator of the location of the holes in the pallet and other information. The speaker 308 can be used to audibly notify the operator of the location of the pallet and other information. Notifications by the notification means 304 will be described later.
[0038] The classifier 340 and the processing unit 350 can be composed of a microcontroller 320 having a processing unit (CPU) and memory, and a software program executed by the microcontroller 320. Furthermore, the distance image acquisition unit 330 can also be implemented as a function of the microcontroller 320. Therefore, the distance image acquisition unit 330, the classifier 340, and the processing unit 350 are not necessarily recognized as separate hardware, but rather represent functions realized by a combination of a processing unit and a software program.
[0039] Alternatively, the distance image acquisition unit 330 and the classifier 340 may be configured with separate hardware. For example, hardware such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit) may be mounted as the interface between the distance image sensor 302 and the microcontroller 320. In this case, the processing of the distance image acquisition unit 330 may be performed by this hardware.
[0040] The above describes the configuration of the forklift 300. With this forklift 300, by configuring a classifier that uses distance images obtained from a distance image sensor as input using machine learning, the accuracy of identifying pallets or parts thereof can be improved. The obtained detection results can then be reflected in the operation of the conveying device or notified to the operator, thereby assisting the operator's operation.
[0041] Another method for detecting the pallet or specific parts of it is to use 3D point cloud data and detect characteristic areas through pattern matching. However, this method requires measuring the object from the same direction every time, which makes it difficult to measure 3D point cloud data of characteristic areas with high density and accuracy.
[0042] Next, we will explain a specific example of the classifier 340 and its training. In one embodiment, the classifier 340 is composed of a combination of a linear SVM (Support Vector Machine) and HOG (Histograms of Oriented Gradients) features. Linear SVM is widely used as a classifier capable of determining whether the input belongs to one of two classes. In the field of general image processing, the input to a linear SVM is image data in which the pixel values are brightness, but in this embodiment, the input is a distance image, which is a difference.
[0043] Linear SVMs can also be trained to directly detect pairs of holes. However, as shown in Figures 3(a) to 3(h), the shape and size of the holes vary depending on the type of pallet. Specifically, the width of the holes differs depending on the pallet. Also, some pallets have a bottom plate, while others do not, and therefore the holes may be open on the bottom or closed on the bottom. For this reason, a classifier trained to directly identify holes may not be able to achieve a sufficiently high level of discrimination.
[0044] Therefore, in order to further improve the identification rate, the inventors focused on the holes 802 of the pallet 800 and the stiles 804 that separate the holes 802, and decided to configure the classifier 340 with a portion (specific part) including the stiles 804 as the detection target. The stiles 804 are an essential component of the pallet regardless of the specifications of the pallet 800, making them suitable as a detection target, and a high identification rate can be expected.
[0045] Figures 7(a) and 7(b) schematically show the specific part 801 that is the target of detection. The specific part 801 includes a part of the upper plate 806, a part of the lower plate 808, a beam 804, and two holes 802. This specific part 801 can be understood to include two U-shaped parts facing each other. Figure 8 shows a distance image of the specific part 801 in Figure 7(a).
[0046] As shown in Figure 3(b), there are also pallets that do not have a bottom plate. Therefore, as shown in Figure 7(b), the classifier 340 may be trained to recognize the T-shape excluding the bottom plate 808 as a specific part 801.
[0047] The processing unit 350 stores geometric information of the pallet. Based on the position of the specific part 801 identified by the classifier 340 and the geometric information, the processing unit 350 can determine the positions of the two holes into which the forks should be inserted through calculation. For example, the geometric information may include the position (coordinates) of the center of the pallet and the relative positional relationship between the holes. Based on the output of the classifier 340, the processing unit 350 can detect the center position of the pallet 800 and determine the positions of the holes into which the forks should be inserted as positions offset by a predetermined amount to the left and right from the center position.
[0048] Figure 7(c) shows a portion of another pallet. A pallet may have three (or more) holes, in which case one pallet may contain two (or more) girders, and two specific parts 801 may be detected adjacent to each other on one pallet. In this case, the centers of the two specific parts 801 (i.e., the centers of the two girders 804) may be determined as the center position of the entire pallet 800.
[0049] (Training of the classifier) Next, we will explain the training process for the classifier 340.
[0050] Figure 9 is a flowchart showing the generation of training data used for training the classifier 340. First, a 3D model Mp of the pallet 800 is generated (S100). The 3D model Mp of the pallet corresponds to the 3D CAD data of the pallet.
[0051] Furthermore, a model Ms for the distance image sensor 302 is generated (S102). The model for the distance image sensor 302 describes the field of view and input / output characteristics of the distance image sensor 302, and can be generated considering the field of view, resolution, measurement accuracy, temperature, frame rate, response characteristics, etc. The response characteristics may include wavelength sensitivity characteristics and transient response characteristics. The model for the distance image sensor 302 can take into account errors and defects based on the IR brightness image, object edges, distance to object points, observation direction, and measurement surface direction, as explained with reference to Figure 12 and later.
[0052] The 3D model Mp of the pallet 800 and the model Ms of the distance image sensor 302 are placed in the 3D virtual space of the computer simulator (S104). In this state, the output of the model Ms of the distance image sensor 302 is acquired as a virtual distance image I1 (S106), and training data I2 is generated based on the virtual distance image I1 (S108).
[0053] Processes S106 to S108 are repeated while changing the relative positional relationship between the palette and the distance image sensor models Mp and Ms (S110).
[0054] Figure 10 shows a three-dimensional virtual space 900 where the palette and the distance image sensor models Mp and Ms are placed. Here, the model Ms of the distance image sensor 302 is assumed to be placed at the origin. The dashed line represents the field of view of the distance image sensor 302. In step S110 of Figure 8, the position (x, y, z) of the palette model Mp is changed. In step S110, in addition to the position (x, y, z), the orientation (θp, θy, θr) may also be changed.
[0055] The above describes the method for generating training data according to the embodiment.
[0056] Given the wide variety of pallet shapes and materials, it is not easy to capture images of actual pallets using a distance image sensor, considering all possible situations, and prepare training data for classifiers. This approach allows for the efficient generation of training data and streamlined learning by utilizing a 3D model.
[0057] (Adding background information) In the process S106 shown in Figure 8, the background distance image I3 may be superimposed on the virtual distance image I1 to generate training data I2.
[0058] Figures 11(a) to 11(c) illustrate the generation of training data based on background distance images. Figure 11(a) is the virtual distance image I1 of the palette 800 obtained in Figure 10. Figure 11(b) represents the background distance image I3. For the background distance image I3, it is preferable to use the actual background distance image captured by the actual distance image sensor 302 in the environment in which the palette is actually used.
[0059] By combining the virtual distance image I1 of the palette in Figure 11(a) and the background distance image I3 in Figure 11(b), the combined distance image in Figure 11(c) is obtained. Part or all of this combined distance image can be used as training data.
[0060] By superimposing the virtual distance image I1 of the palette and the background distance image I3 to generate training data, the recognition rate in the environment in which the palette is actually used can be improved.
[0061] Furthermore, the background distance image I3 may be one generated by actual measurements of the distance image sensor 302. Since the sensing of the distance image sensor 302 is greatly affected by ambient light, using an actual captured image for the background allows for the generation of training data that takes ambient light into account, which is expected to improve the recognition rate.
[0062] Alternatively, a 3D model of a background object may be placed in a 3D virtual space, a virtual distance image of the background may be generated using the model Ms of the distance image sensor 302, and the virtual distance image of the palette and the virtual distance image of the background may be superimposed to use as training data.
[0063] Alternatively, a 3D model of a background object and a 3D model of a palette may be placed in a 3D virtual space, a virtual distance image including them may be generated, and part or all of this image may be used as training data.
[0064] (Error model) In the explanation so far, we have described a measurement model that considers only the model Ms of the distance image sensor 302, but in order to generate more accurate training data, we will also consider the error model M ERR It would be good to implement this.
[0065] Error Model M ERR This describes the influence of at least one of the following on the output of the distance image sensor: (i) the distance to the object being measured, (ii) the direction of the measurement surface, and (iii) the material of the object being measured. The influence on the output of the distance image sensor is greatest for the distance to the object being measured, the direction of the measurement surface, and the material of the object being measured, so the error model M ERR It is preferable to create the device considering at least the distance to the object to be measured, more preferably considering both the distance to the object and the direction of the measurement surface, and if you want to further improve accuracy, it is best to create it considering all three factors.
[0066] (i) Distance to the object to be measured Let's consider the input / output characteristics of a depth image sensor. In an ideal depth image sensor, the pixel values pv of the output depth image change linearly with respect to the input distance d. pv = a1 × d …(1)
[0067] However, in actual distance image sensors, the tilt a1 may differ depending on the sensor type and product. Alternatively, as shown in equation (2), a nonlinear term or an offset term a0 may be included. pv = a0 + a1 × d + a2 × d 2 +··· …(2) Therefore, error model M ERR Next, we will describe the relationship between the distance d between the depth image sensor and the object being measured, and the pixel value pv in the depth image.
[0068] (ii) Direction of the measurement surface The input / output characteristics represented by equation (2) may change depending on the direction of the measurement surface. For example, suppose that for an object in front of the distance image sensor 302, an input / output characteristic with a set of coefficients S1[a0, a1, ...] is obtained. In this case, the input / output characteristics for a different direction may not be described by the same set S1, and an input / output characteristic with a different set of coefficients S2[a0, a1, ...] may be obtained. Therefore, error model M ERR It is best to create it while considering the direction of the measurement surface.
[0069] (iii) Materials Depending on the sensing method and wavelength used by the distance image sensor 302, even when measuring objects at the same distance using the same distance image sensor 302, the pixel values of the distance image based on the output of the distance image sensor 302 may not be the same if the materials of the objects are different.
[0070] Here, the materials and surface materials of the pallet 800 vary widely, including resin, wood, metal, and painted versions thereof. Based on this understanding, the model Ms of the distance image sensor 302 is generated considering the effect that differences in the material being measured have on the output of the distance image sensor 302. In other words, the input / output characteristics of the distance image sensor 302 are measured in advance for each material, and the error model M takes this into account. ERR Prepare the following. The model Mp for Palette 800 allows you to specify the material for each part of Palette 800, or for the whole. This allows for the generation of more accurate training data.
[0071] (Environmental model) The measurement model includes the error model M. ERR In addition to, or instead of, environmental model MENV The environmental model M can be introduced. ENV describes the influence of at least one of environmental light and ambient temperature on the output of the distance image sensor.
[0072] Many distance image sensors 302 are active sensors that sense the distance to an object by irradiating the object with light and measuring the reflected light. Here, if there is environmental light of the same wavelength as the light in the environment where the distance image sensor 302 is used, noise will be superimposed on the distance image based on the output of the distance image sensor 302. The environmental model M ENV When introduced, noise dependent on environmental light can be superimposed on the virtual distance image I1, and the training data can be made closer to the distance image obtained in a more realistic environment.
[0073] Also, the measurement error changes due to the influence of the ambient temperature. Therefore, it is advisable to describe the influence of the ambient temperature in the environmental model.
[0074] (Error model of an active type distance image sensor) Regarding the generation of training data when using an active sensor that irradiates an object with infrared or near-infrared light (hereinafter referred to as IR light) as light and measures the IR reflected light as the distance image sensor 302, an explanation will be given.
[0075] FIG. 12 is a flowchart of a method for generating training data when using an active type distance image sensor. This flowchart includes processes S120 and S122 in addition to the flowchart of FIG. 9. Although not shown in FIG. 12, a background image may be added to the training data.
[0076] In process S120, an IR luminance image I4 showing the intensity distribution (IR luminance distribution) of the IR reflected light that the distance image sensor 302 will receive is generated. The resolution of the IR luminance image I4 may be the same as the resolution of the distance image sensor 302 (that is, the resolution of the virtual distance image I1), or may be different, but the pixels of the virtual distance image I1 and the pixels of the IR luminance image I4 can be associated.
[0077] In the subsequent process S122, the virtual distance image I1 is modified based on the IR brightness image I4. In the following process S108, training data I2 is generated based on the modified virtual distance image I1'.
[0078] Figure 13 illustrates the generation of the IR brightness image I4. The light source 400 corresponds to the light source that illuminates the IR light of the distance image sensor 302, and the camera 402 corresponds to the image sensor and light-receiving element of the distance image sensor 302.
[0079] Camera 402 measures the light reflected by object OBJ from the light emitted by light source 400. Figure 13 shows a case where light source 400 and camera 402 are located in different positions and the illumination direction (vector) Lv and observation direction Vv are different. However, in the actual distance image sensor 302, light source 400 and camera 402 are positioned in substantially the same location and facing substantially the same direction, and therefore the illumination direction Lv and observation direction Vv are approximately the same. Nv is the normal vector pointing to the normal at the reflection point.
[0080] The reflected light measured by camera 402 includes diffuse reflection from object OBJ and specular reflection. Diffuse reflection occurs isotropically, independent of the illumination and observation directions. The diffuse reflection intensity Id observed by camera 402 is given by the distance r between camera 402 and object OBJ, and the diffuse reflectance k of the object OBJ's surface. d It can be calculated based on the following: Diffuse reflectance k d This is determined by the material of the object's surface (OBJ) and is given as a parameter. The diffuse reflected light intensity Id is, for example, Id = (Nv·Lv)k d ·IR0 / r 2 This can be expressed as follows: Here, IR0 is the IR light intensity of light source 400 when light source 400 is treated as a point light source.
[0081] Specularly reflected light is strongest when the observation direction Vv coincides with the specular reflection direction (vector) Rv, and decreases as the observation direction Vv moves away from the specular reflection direction Rv. The directivity of the reflected light is determined by the surface properties of the reflective surface, such as surface roughness n, and its polarization characteristics. The specularly reflected light intensity Is observed by camera 402 is equal to the specular reflectance k s It can be calculated using parameters such as distance r and surface roughness n of the object OBJ. The specular reflection intensity Is is, for example, Is = (Rv·Nv)k s ·IR0 / r 2 It can be expressed as follows.
[0082] The generation of the IR luminance image I4 can be calculated as follows. For the pallet 800 and surrounding objects placed in a 3D virtual space, the parameters that define their respective reflection characteristics are, specifically, the diffuse reflectance k. d , specular reflectance k s The surface roughness n and other parameters are given. Then, using a known shading (rendering) method, the intensity distribution of IR reflected light received by the image sensor of the distance image sensor 302 is calculated, and an IR luminance image I4 is generated. For example, the IR luminance value I4i at each pixel i of the image sensor of the image sensor 302 is calculated by the sum of the diffuse reflected light intensity Isi and the specular reflected light intensity Idi at each pixel i, so I4i = Isi + Idi.
[0083] The diffuse reflectance k is a reflectance characteristic parameter used in the calculation of the IR brightness image I4. d , specular reflectance k s And the surface roughness n is measured using the I4 value of the IR brightness image. * This can be determined using the measured value I4 of the IR brightness image. * This can be obtained, for example, by using a flat plate as the object to be measured (OBJ), placing the plate in real space, and positioning an active distance sensor equipped with a light source and camera directly opposite the plate. Parameter k d , k s n is the calculated value I4 and the measured value I4 of the IR brightness image. *The difference between the calculated and measured values of the IR brightness image is determined by an optimization calculation that minimizes the error function f. Here, the error function fΔ is calculated between the calculated value I4 and the measured value I4 of the IR brightness image. * This is the sum of the differences Δi for each pixel i over all pixels. For example, difference Δi = |I4i - I4i * Let | be the error function fΔ = Σi|I4i - I4i * This can be expressed as |. For optimization calculations, known methods such as the Simulated Annealing method can be applied. The parameter constraints are 0≦k. d ≤1, 0 ≤ k s The following can be used: ≤1, 0≦n, 0≦IR0.
[0084] Measured value I4 * The reflection characteristic parameter k is calculated using an optimization calculation with an error function fΔ that represents the difference. d , k s When determining n, the reflection characteristic parameters are determined based on literature values, or measured values I4 * Compared to determining the reflection characteristic parameters without using this method, the calculation accuracy of the IR brightness image I4 can be improved. Reflection characteristic parameter k d , k s n can be determined for each object OBJ. If the materials and properties of the palette 800 and surrounding objects placed in the 3D virtual space are different, the reflection property parameters that have been determined in advance for each material can be applied.
[0085] Figures 14(a) to 14(c) illustrate the generation of training data based on the flowchart in Figure 12. Figure 14(a) shows how models of the pallet 800 and other objects are placed in a virtual three-dimensional space. The upper part of Figure 14(b) shows a virtual distance image I1 that does not contain errors and is generated based on the model in Figure 14(a). Figure 14(c) shows an IR luminance image I4 that is generated based on the model in Figure 14(a). The lower part of Figure 14(b) shows a corrected virtual distance image I1' based on the IR luminance image I4. The difference d'-d between the pixel value d of a certain pixel in the corrected virtual distance image I1' in Figure 14(b) and the corresponding pixel value d' in the uncorrected virtual distance image I1 in Figure 14(a) is correlated with the corresponding pixel value α in the IR luminance image I4. d'-d=f(α)
[0086] The above describes the method for generating training data when using active sensors. Next, we will explain its advantages.
[0087] The inventors have found that the accuracy of the distance image generated by an actual distance image sensor is affected by the intensity of the reflected light received by the image sensor / photodetector of the distance image sensor 302. Therefore, in processing S120, a virtual IR luminance image is generated, and in processing S122, noise based on the IR luminance image I4 is superimposed on the virtual distance image I1, or some pixels of the virtual distance image I1 are replaced with missing data, thereby generating an image that is close to the distance image generated by an actual distance image sensor.
[0088] Next, we will explain a specific example of the correction process for the virtual distance image I1 based on the IR luminance image I4.
[0089] (Example 1) In one embodiment, an error may be superimposed on the pixel values of the virtual distance image I1 based on the pixel values of the IR luminance image I4. Figure 15 shows an example of the relationship between the pixel values of the IR luminance image I4 and the distance error. The horizontal axis represents the pixel values α of the IR luminance image I4, and the vertical axis represents the distance error δ as a standard deviation SD.
[0090] The plot represents the measured values obtained for a certain distance image sensor 302, and the solid line shows the fitted function f(α). SD = f(α) = 1316α -0.98 +0.88
[0091] For example, if the pixel value α of IR luminance image I4 is 400, the standard deviation SD is 5 mm. This means that for a pixel in IR luminance image I4 with a pixel value α of 400, the pixel value (distance) in virtual distance image I1 has an error that follows a normal distribution with a statistically significant standard deviation of 5 mm.
[0092] In this example, the smaller the intensity of the IR reflected light, the larger the superimposed error. This means that the smaller the brightness of the IR reflected light, in other words, when the distance to the object point is large or the reflectivity of the object point is very low, the distance measurement by the distance image sensor 302 is considered inaccurate, and a large error is superimposed.
[0093] Let's focus on a certain pixel [x,y] in the virtual distance image I1, and assume its pixel value (i.e., distance) is d. When the pixel value (i.e., reflected light intensity) of the same pixel in the IR luminance image I4 is α, the pixel value d of the same pixel [x,y] in the corrected virtual distance image I1 is corrected as follows by superimposing a statistical error δ based on the standard deviation SD. d' = d + δd To generate δd, you can use a random number generator that allows you to specify the mean and standard deviation.
[0094] Figures 16(a) and (b) show the correction of the virtual distance image based on the relationship in Figure 15. Figures 16(a) and (b) show the central portion of the virtual distance image before and after correction, respectively, including the digit 804 of palette 800.
[0095] The holes in pallet 800 and their surrounding areas have small pixel values in the IR luminance image I4. In these areas with small pixel values in the IR luminance image I4, noise is superimposed on the pixel values (distance) of the virtual distance image.
[0096] Through the above processing, training data close to the distance image measured by the actual distance image sensor 302 can be obtained. Parameter k based on measured values d , k s By superimposing errors (noise) on the IR brightness image I4 generated using n, it is possible to obtain training data that approximates the actual distance image measured by the actual distance image sensor 302, thereby improving the recognition accuracy of the classifier. In particular, for a wide variety of palettes with different materials and colors, individual parameters k d , k s By applying n to obtain training data, the classification accuracy of the classifier can be improved.
[0097] Furthermore, depending on the distance measurement method of the distance image sensor 302, the distance may also become inaccurate if the intensity of the IR reflected light is too high. In this case, the error may be superimposed on the pixel values of the virtual distance image I1 not only in the parts of the IR luminance image I4 where the pixel values are small, but also in the parts where the pixel values of the IR luminance image I4 are large. In other words, the function f should be modified.
[0098] (Example 2) Depending on the type of distance image sensor 302, the distance image it generates may contain missing data (missing values) indicating that distance information could not be obtained. The inventors have recognized that missing data tends to occur frequently at the edges of objects. They have also recognized that missing data can occur in areas where the reflected light incident on the image sensor / photodetector of the distance image sensor 302 is too weak or too strong.
[0099] Therefore, in Example 2, based on the virtual distance image I1 and the IR luminance image I4 (also called the virtual luminance image I4), locations (pixels) where missing data may occur are estimated, and the pixel values of those pixels in the virtual distance image I1 are replaced with the missing data. For example, a missing data generator is constructed using machine learning, which takes the virtual distance image I1 and the virtual luminance image I4 as input and outputs the missing data. This makes it possible to generate missing data that takes both the distance image and the luminance image into consideration, thereby improving the accuracy of missing data estimation.
[0100] Figure 17 shows an example of the input and output images of the missing data generator 408. The input images of the missing data generator 408 are a virtual distance image I1 and a virtual luminance image I4. The output image of the missing data generator 408 is the missing image I5. The missing image I5 is a monochrome image in which the pixel value of pixels estimated to be missing is set to 0, and the pixel value of pixels estimated not to be missing is set to 1. In the missing image I5 of Figure 17, the image is inverted to monochrome for visibility, with pixels with a pixel value of 0 (missing) being shown as white, and pixels with a pixel value of 1 (not missing) being shown as black.
[0101] Figure 18 schematically shows an example of the network structure 410 of the missing data generator 408. The network structure 410 comprises an input layer 412, a feature extraction layer 414, a classification layer 416, and an output layer 418. The network structure 410 is a so-called deep neural network (DNN), and in particular a convolutional neural network (CNN) that utilizes convolutional processing.
[0102] The input layer 412 accepts a virtual distance image I1 and a virtual luminance image I4 as input data. The input layer 412 has a data size corresponding to the image sizes of the virtual distance image I1 and the virtual luminance image I4, and has 2 channels. The data size of the input layer 412 is not particularly limited, but for example, it is 424 x 512.
[0103] The feature extraction layer 414 extracts image features through convolution. The feature extraction layer 414 is configured to extract features through, for example, four convolution operations 420, 422, 424, and 426. In each of the four convolution operations 420 to 426, for example, the size of the convolution kernel can be set to 7x7, and ReLU (Rectified Linear Unit) can be used as the activation function. The data size handled by the feature extraction layer 414 may be the same as the data size of the input layer 412, and pooling may not be included. The number of channels in the feature extraction layer 414 is not particularly limited, but for example, it is 32.
[0104] The classification layer 416 classifies whether each pixel is missing or not based on the features extracted by the feature extraction layer 414. The classification layer 416 includes two convolution operations 428 and 430 for fully connecting multiple (e.g., 32) channels, a softmax operation 432 for calculating the probability of pixel loss, and a classification operation 434 that sets a threshold for the loss probability and outputs whether the pixel is missing or not.
[0105] The output layer 418 outputs the missing image I5 as output data. The output layer 418 has a data size corresponding to the image sizes of the virtual distance image I1 and the virtual luminance image I4, and has 1 channel.
[0106] Next, we will explain the machine learning of the missing data generator 408. When training the missing data generator 408, images of actual palettes captured by a depth image sensor are used as training data. As example data, we use the measured distance image I1r and measured brightness image I4r obtained from measurements taken by the depth image sensor. As ground truth data, we use the measured missing image I5r, in which the pixel values of pixels with missing pixel values (e.g., distance of 0) in the measured distance image I1r are set to 0, and the pixel values of pixels that do not have missing pixel values (e.g., distance is not 0) are set to 1.
[0107] The measured distance image I1r and measured luminance image I4r used for training the missing data generator 408 may contain missing data (missing values) because they are based on actual measurements. On the other hand, the virtual distance image I1 and virtual luminance image I4 used as input to the missing data generator 408 do not contain missing data because they are generated from a 3D model of the object. Therefore, there is a discrepancy between the training data and input data of the missing data generator 408 in terms of the presence or absence of missing data.
[0108] Figure 19(a) shows an example of a measured distance image I1r, and Figure 19(b) shows an example of a virtual distance image I1. In the measured distance image I1r in Figure 19(a), the pixel values of missing pixels become 0, and there are many pixels that appear black. On the other hand, in the virtual distance image I1 in Figure 19(b), there are no missing pixels, so there are no pixels with a pixel value of 0 that appear black.
[0109] If there is a discrepancy between the training data and input data of the missing data generator 408, the estimation accuracy of the missing data generator 408 after training will decrease. Therefore, during the training of the missing data generator 408, the weighting of the pixels of interest is arbitrarily reduced in the convolutional processing 420-426 of the feature extraction layer 414.
[0110] Figure 20 schematically shows the convolution kernel 440 used in the convolution processes 420-426 of the missing data generator 408. Figure 20 shows that the convolution kernel 440 is, for example, 7x7 in size and has 7x7 = 49 kernel values w11-w77. The convolution kernel 440 has a central kernel value w44 located at the center 442 of the kernel, which is a fixed value. The fixed value is, for example, 0. The central kernel value w44 at 442 is multiplied by the pixel of interest in the convolution process. Kernel values other than the central 442 are multiplied by the surrounding pixels of the pixel of interest in the convolution process. Kernel values other than the central 442 are adjusted during the machine learning process.
[0111] By using a constraint that fixes the kernel value of the central 442 of kernel 440 to 0, the pixel value of the pixel of interest can be ignored, and the convolution process can be performed using the pixel values of surrounding pixels. This reduces the influence of measured missing data contained in the measured distance image I1r and measured luminance image I4r during training on the estimation of missing pixels, and makes it possible to estimate missing pixels from the information of surrounding pixels. As a result, even when virtual distance image I1 and virtual luminance image I4 that do not contain missing data are used as input data, the accuracy of missing data estimation by the missing data generator 408 can be improved.
[0112] The convolution kernel 440, with its central kernel value fixed at 0, may be used in all four convolution operations 420-426 of the missing data generator 408, or in some of the four convolution operations 420-426. The central kernel value 442 may be a small value close to 0 (e.g., 0.1, 0.01, 0.001) rather than 0.
[0113] The missing data generator 408 generates a missing data image I5, which is used to correct the virtual distance image I1. Pixels that are missing in the missing data image I5 are corrected so that their pixel values in the virtual distance image I1 become 0 (e.g., distance is 0). Pixels that are not missing in the missing data image I5 retain their original pixel values in the virtual distance image I1 (e.g., finite distance). This allows for the generation of a corrected virtual distance image I1' that takes missing data into account.
[0114] According to Example 2, it is possible to emulate the occurrence of missing data in actual distance images and generate training data that closely resembles actual distance images.
[0115] (Example 3) Example 3 combines Examples 1 and 2, using an error-superimposed distance image, which is a virtual distance image with the error (noise) from Example 1 superimposed, as input to the missing data generator 408 of Example 2. The error-superimposed distance image is closer to the measured distance image I1r used for training the missing data generator 408 in that it contains error (noise). Therefore, by using the error-superimposed distance image as input, the accuracy of missing data estimation by the missing data generator 408 can be improved.
[0116] Figure 21 is a flowchart of the method for generating training data according to Example 3. First, the 3D model Mp of the pallet 800 and the model Ms of the distance image sensor 302 are placed in the 3D virtual space of the computer simulator (S200). Step S200 may be the same as the processing in S100 to S104 described above. Next, the output of the model Ms of the distance image sensor 302 is acquired as a virtual distance image I1 (S202). Step S202 may be the same as the processing in S106 described above.
[0117] Next, a virtual luminance image I4 is generated from the model of the distance image sensor using reflection characteristic parameters determined using measured values (S204). Step S204 may be the same as the process in S120 described above. Subsequently, an error-superimposed distance image I1d is generated by superimposing an error δ based on the pixel values of the virtual luminance image I4 onto the pixel values of the virtual distance image I1 (S206). The error-superimposed distance image I1d can be generated by the same method as in Example 1.
[0118] Next, missing data indicating missing pixels is generated based on the virtual luminance image I4 and the error superimposed distance image I1d (S208). In step S208, the missing data generator 408 generates a missing image I5 by inputting the virtual luminance image I4 and the error superimposed distance image I1d into the missing data generator 408. Subsequently, the pixel values of the error superimposed distance image I1d are replaced with missing data to generate a corrected virtual distance image I1' (S210). Finally, training data I2 is generated based on the corrected virtual distance image I1' (S212). Step S212 may be the same process as S108 described above.
[0119] In the flow shown in Figure 21, similar to S110 described above, diverse training data can be prepared by repeating steps S202 to S212 while changing the relative position of the palette and the distance image center. Furthermore, by changing the reflection characteristic parameters according to the material and texture of the palette and repeating steps S202 to S212, training data corresponding to a wide variety of palettes can be prepared.
[0120] The present invention has been described above based on examples. Those skilled in the art will understand that the present invention is not limited to the above embodiments, that various design changes are possible, and that various modifications are possible, and that such modifications also fall within the scope of the present invention. Such modifications will be described below.
[0121] Although the embodiments described used a forklift as an example of a conveying device, the present invention is not limited to this and can be applied to various conveying devices that handle pallets, such as hand lifts (hand pallets), AGVs (Automatic Guided Vehicles), and fork loaders (wheel loaders with forks). [Explanation of Symbols]
[0122] 600 forklifts 602 Car body 604 Fork 606 Lifting mechanism 608 Mast 610 Front Wheel 612 Rear wheel 700 cockpits 702 Ignition Switch 704 Steering Wheel 706 Lift Lever 708 Accelerator pedal 710 Brake pedal 712 Forward / Reverse Lever 714 Dashboard 300 forklifts 302 Distance Image Sensor 303 Camera 304 Notification means 306 displays 308 speakers 310 Controller 320 Microcontrollers 330 Distance image acquisition unit 340 classifiers 350 Processing Unit 360 Pattern Matching Section 362 Template Image Generator D1 Distance measurement data D2 distance image D3 image data D4 partial distance image D5 Location information D6 template image 800 pallets 801 Specific part 802 holes 804 digits
Claims
1. A method for generating training data used for learning a conveying device that transports pallets, The aforementioned transport device is An active distance image sensor that irradiates light from a light source and measures the reflected light, (i) A calculation unit having a classifier configured by machine learning using the training data so as to be able to acquire a distance image in which the pixel values represent distance based on the output of the distance image sensor, and (ii) a detection target which is a palette or a specific part of the palette based on the distance image, It is equipped with, The aforementioned generation method is The steps include generating a three-dimensional model of the aforementioned palette, The steps include: placing three-dimensional models of the palette and the distance image sensor in a three-dimensional virtual space, and generating a virtual distance image that would be obtained when the palette is measured by the distance image sensor; The steps include: arranging three-dimensional models of the palette, the light source, and the distance image sensor in a three-dimensional virtual space, and generating a virtual brightness image showing the intensity distribution of the reflected light that the distance image sensor will receive; The steps include modifying the virtual distance image based on the virtual luminance image, The process includes the step of generating the training data based on the corrected virtual distance image, The virtual luminance image is generated using parameters that show the reflection characteristics of diffuse and specular reflection in the object being measured. The generation method is characterized in that the values of the parameters indicating the reflection characteristics are determined using a measured brightness image obtained when the object to be measured, which is placed in real space, is measured by the distance image sensor.
2. The generation method according to claim 1, characterized in that the parameters indicating the reflection characteristics are determined for each material and / or color of the surface of the object to be measured.
3. The generation method according to claim 1 or 2, characterized in that the step of correcting the virtual distance image includes superimposing an error on the pixel values of the virtual distance image based on the pixel values of the virtual luminance image.
4. A method for generating training data used for learning a conveying device that transports pallets, The aforementioned transport device is An active distance image sensor that irradiates light and measures the reflected light, (i) A calculation unit having a classifier configured by machine learning using the training data so as to be able to acquire a distance image in which the pixel values represent distance based on the output of the distance image sensor, and (ii) a detection target which is a palette or a specific part of the palette based on the distance image, It is equipped with, The aforementioned generation method is The steps include generating a three-dimensional model of the aforementioned palette, The steps include: placing a three-dimensional model of the palette in a three-dimensional virtual space and generating a virtual distance image that would be obtained when the palette is measured by the distance image sensor; The steps include generating a virtual brightness image showing the intensity distribution of the reflected light that the distance image sensor will receive, The steps include correcting the virtual distance image based on the missing data generated using the virtual distance image and the virtual brightness image, A generation method characterized by comprising the step of generating training data based on the corrected virtual distance image.
5. The generation method according to claim 4, characterized in that the missing data is generated by a missing data generator configured by machine learning, which takes as input a measured distance image and a measured brightness image obtained when a measurement target placed in real space is measured by the distance image sensor.
6. The generation method according to claim 5, characterized in that the missing data generator comprises a convolutional neural network and is performed by machine learning under the constraint that the kernel value at the center of the convolutional kernel is fixed.
7. The generation method according to claim 6, characterized in that the fixed value is 0.
8. The generation method further comprises the step of generating an error-superimposed distance image by superimposing an error on the pixel values of the virtual distance image based on the pixel values of the virtual brightness image, The missing data is generated using the error superimposed distance image, The generation method according to any one of claims 4 to 7, characterized in that the corrected virtual distance image is generated by replacing the pixel values of the error superimposed distance image with the missing data.