Learning method
Through the weakly supervised learning method, the monocular camera model is trained using image areas with known distance size relationships, which solves the problem of data set preparation in monocular camera ranging and improves the learning efficiency and accuracy of the model.
Patent Information
- Application Number
- CN202111042419.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-15
- Filing Date
- 2021-09-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-09-07
AI Technical Summary
In the prior art, when using images from a monocular camera to measure the distance of a subject, a large number of image data sets with correct distance labels are required for model training, but these data sets are difficult to prepare.
Through a weakly supervised learning method, multiple images with known distance size relationships are used to capture and cut out local areas using the imaging device, and the distance relationship between these areas is learned through statistical models to achieve model training.
Even without the correct label image data, the learning efficiency and accuracy of the model can be improved, and the dataset preparation process can be simplified.
Smart Images

Figure CN114638354B_ABST
Abstract
Description
[0001] This application is based on Japanese Patent Application No. 2020-207634 (filing date: December 15, 2020) and claims the benefit of priority therefrom. This application incorporates the entire contents of that application by reference thereto. Technical Field
[0002] Embodiments of the present invention relate to a learning method. Background Art
[0003] In order to obtain the distance to a subject, a technique using images captured by two imaging devices (cameras), a stereo camera (multi-camera) is known. In recent years, however, a technique has been developed to obtain the distance to a subject using an image captured by one imaging device (monocular camera).
[0004] Here, in order to obtain the distance to a subject using an image as described above, it is considered to use a statistical model generated by applying a machine learning algorithm such as a neural network.
[0005] However, in order to generate a highly accurate statistical model, it is necessary to make the statistical model learn a huge learning dataset (a set of learning images and correct values related to the distance to the subject in the learning images), but it is not easy to prepare this dataset. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a learning method that can improve the ease of learning in a statistical model.
[0007] According to an embodiment, there is provided a learning method for causing a statistical model to learn, the statistical model being configured to take an image including a subject as an input and output the distance to the subject. The learning method includes the steps of: obtaining a first image and a second image including a subject captured by an imaging device; and causing the statistical model to learn based on a first distance output from the statistical model with a first region, which is at least a part of the first image, as an input and a second distance output from the statistical model with a second region, which is at least a part of the second image, as an input. The magnitude relationship between a third distance to the subject included in the first image and a fourth distance to the subject included in the second image is known, and the learning includes causing the statistical model to learn such that the magnitude relationship between the first distance and the second distance is equal to the magnitude relationship between the third distance and the fourth distance. Brief Description of the Drawings
[0008] Figure 1 FIG. is an example showing the configuration of a distance measurement system in a first embodiment.
[0009] Figure 2 This is a diagram showing an example of the system configuration of an image processing apparatus.
[0010] Figure 3 This is a diagram for explaining an overview of the operation of a distance measurement system.
[0011] Figure 4 This is a diagram for explaining the principle of predicting the distance to a subject.
[0012] Figure 5 This is a diagram for explaining the patch method of predicting distance based on a captured image.
[0013] Figure 6 This is a diagram showing an example of information related to image patching.
[0014] Figure 7 This is a diagram for explaining an overview of the learning method of a general statistical model.
[0015] Figure 8 This is a diagram for explaining a learning dataset.
[0016] Figure 9 This is a diagram for explaining an overview of the learning method of the statistical model of the present embodiment.
[0017] Figure 10 This is a diagram for explaining a learning image for statistical model learning.
[0018] Figure 11 This is a block diagram showing an example of the functional configuration of a learning processing unit.
[0019] Figure 12 This is a flowchart showing an example of the processing sequence of an image processing apparatus when learning a statistical model.
[0020] Figure 13 This is a flowchart showing an example of the processing sequence of an image processing apparatus when obtaining distance information from a captured image.
[0021] Figure 14 This is a flowchart showing an example of the processing sequence of an image processing apparatus when learning a statistical model in the second embodiment.
[0022] Explanation of reference numerals
[0023] 1. Distance measurement system, 2. Imaging device, 3. Image processing device, 21. Lens, 22. Image sensor, 31. Statistical model storage unit, 32. Image acquisition unit, 33. Distance acquisition unit, 34. Output unit, 35. Learning processing unit, 35a. Discrimination unit, 35b. Calculation unit, 35c. Learning unit, 221. First sensor, 222. Second sensor, 223. Third sensor, 301. CPU, 302. Non-volatile memory, 303. RAM, 303A. Image processing program, 304. Communication device, 305. Bus. Detailed implementation mode
[0024] Hereinafter, the implementation mode will be described with reference to the drawings.
[0025] (First implementation mode)
[0026] First, the first implementation mode will be described. Figure 1 An example of the structure of the distance measurement system in this implementation mode is shown. Figure 1 The shown distance measurement system 1 is used to capture an image and use the captured image to obtain (measure) the distance from the imaging location to the subject. In addition, the distance described in this implementation mode can be either an absolute distance or a relative distance.
[0027] As Figure 1 shown, the distance measurement system 1 includes an imaging device 2 and an image processing device 3. In this implementation mode, it is assumed that the distance measurement system 1 includes the imaging device 2 and the image processing device 3 as independent devices for description, but the distance measurement system 1 can also be implemented as one device (distance measurement device) in which the imaging device 2 functions as an imaging unit and the image processing device 3 functions as an image processing unit. In addition, the image processing device 3 can also operate as a server that executes various cloud computing services, for example.
[0028] The imaging device 2 is used to capture various images. The imaging device 2 includes a lens 21 and an image sensor 22. The lens 21 and the image sensor 22 correspond to the optical system (monocular camera) of the imaging device 2.
[0029] The light reflected by the subject enters the lens 21. The light incident on the lens 21 passes through the lens 21. The light passing through the lens 21 reaches the image sensor 22 and is received (detected) by the image sensor 22. The image sensor 22 converts the received light (photoelectric conversion) into an electrical signal, thereby generating an image composed of a plurality of pixels.
[0030] In addition, the image sensor 22 is implemented by, for example, a CCD (Charge Coupled Device) image sensor and a CMOS (Complementary Metal Oxide Semiconductor) image sensor. The image sensor 22 includes, for example, a first sensor (R sensor) 221 that detects light in the red (R) band, a second sensor (G sensor) 222 that detects light in the green (G) band, and a third sensor (B sensor) 223 that detects light in the blue (B) band. The image sensor 22 can receive light in the corresponding band through the first to third sensors 221 to 223, and generate sensor images (R image, G image, and B image) corresponding to each band (color component). That is, the image captured by the imaging device 2 is a color image (RGB image), and the R image, G image, and B image are included in this image.
[0031] In addition, in the present embodiment, it is described that the image sensor 22 includes the first to third sensors 221 to 223, but the image sensor 22 only needs to be configured to include at least one of the first to third sensors 221 to 223. In addition, the image sensor 22 may be configured to include a sensor for generating, for example, a monochrome image instead of the first to third sensors 221 to 223.
[0032] In the present embodiment, the image generated based on the light transmitted through the lens 21 is an image affected by the aberration of the optical system (lens 21), and includes blur generated by the aberration.
[0033] Figure 1 The illustrated image processing device 3 includes a statistical model storage unit 31, an image acquisition unit 32, a distance acquisition unit 33, an output unit 34, and a learning processing unit 35 as functional structures.
[0034] A statistical model is stored in the statistical model storage unit 31, and this statistical model is used to obtain the distance of the subject from the image captured by the imaging device 2. The statistical model stored in the statistical model storage unit 31 is generated by learning the blur that is generated in the image affected by the aberration of the optical system and that changes non-linearly according to the distance to the subject in the image. According to such a statistical model, by inputting an image into this statistical model, the distance to the subject in the image can be predicted (output) as a predicted value corresponding to this image.
[0035] In addition, the statistical model is set to be generated by applying various known machine learning algorithms such as neural networks, linear recognizers, or random forests. Additionally, neural networks that can be applied in this embodiment can include, for example, convolutional neural networks (CNN: Convolutional Neural Network), fully coupled neural networks, and recurrent neural networks.
[0036] The image acquisition unit 32 acquires the image captured by the imaging device 2 (image sensor 22) from the imaging device 2.
[0037] The distance acquisition unit 33 uses the image acquired by the image acquisition unit 32 to acquire distance information indicating the distance to the subject in the image. In this case, the distance acquisition unit 33 inputs the image into the statistical model stored in the statistical model storage unit 31 to acquire distance information indicating the distance to the subject in the image.
[0038] The output unit 34 outputs the distance information acquired by the distance acquisition unit 33 in a mapping form arranged in correspondence with the image in terms of position, for example. In this case, the output unit 34 can output image data composed of pixels with the distance represented by the distance information as pixel values (i.e., output the distance information as image data). When the distance information is output as image data in this way, this image data can be displayed as a distance image representing the distance with, for example, colors. The distance information output by the output unit 34 can also be used, for example, to calculate the size of the subject in the image captured by the imaging device 2.
[0039] The learning processing unit 35 performs a process of learning the statistical model stored in the statistical model storage unit 31 using, for example, the image acquired by the image acquisition unit 32. Details of the process performed by the learning processing unit 35 will be described later.
[0040] In addition, in Figure 1 the example shown, the image processing device 3 is described as including each unit 31 to 35, but the image processing device 3 can also be composed of, for example, a distance measurement device including the image acquisition unit 32, the distance acquisition unit 33, and the output unit 34, and a learning device including the statistical model storage unit 31, the image acquisition unit 32, and the learning processing unit 35.
[0041] Figure 2 Represents Figure 1 An example of the system configuration of the image processing device 3 shown. The image processing device 3 includes a CPU 301, a non-volatile memory 302, a RAM 303, and a communication device 304. Additionally, the image processing device 3 has a bus 305 that connects the CPU 301, the non-volatile memory 302, the RAM 303, and the communication device 304 to each other.
[0042] The CPU 301 is a processor for controlling the operations of various components within the image processing apparatus 3. The CPU 301 can be either a single processor or composed of multiple processors. The CPU 301 executes various programs loaded from the non-volatile memory 302 into the RAM 303. These programs include an operating system (OS) and various application programs. The application programs include an image processing program 303A.
[0043] The non-volatile memory 302 is a storage medium used as an auxiliary storage device. The RAM 303 is a storage medium used as a main storage device. In Figure 2 only the non-volatile memory 302 and the RAM 303 are shown, but the image processing apparatus 3 may also include other storage devices such as an HDD (Hard Disk Drive) and an SSD (Solid State Drive).
[0044] In addition, in the present embodiment, Figure 1 the statistical model storage unit 31 shown is implemented, for example, by the non-volatile memory 302 or other storage devices.
[0045] In addition, in the present embodiment, it is assumed that Figure 1 a part or all of the image acquisition unit 32, the distance acquisition unit 33, the output unit 34, and the learning processing unit 35 shown are implemented by causing the CPU 301 (i.e., the computer of the image processing apparatus 3) to execute the image processing program 303A, that is, by software. The image processing program 303A can be distributed by being stored in a computer-readable storage medium or downloaded to the image processing apparatus 3 via a network.
[0046] Here, it is described that the CPU 301 executes the image processing program 303A, but a part or all of the units 32 to 35 can also be implemented using, for example, a GPU (not shown) instead of the CPU 301. In addition, a part or all of the units 32 to 35 can be implemented either by hardware such as an IC (Integrated Circuit) or by a combination of software and hardware.
[0047] The communication device 304 is a device configured to perform wired communication or wireless communication. The communication device 304 includes a transmission unit for transmitting signals and a reception unit for receiving signals. The communication device 304 performs communication with external devices via a network, communication with external devices existing in the vicinity, etc. The external devices include the imaging device 2. In this case, the image processing apparatus 3 can receive images from the imaging device 2 via the communication device 304.
[0048] Although in Figure 2Although it is omitted in the figure, the image processing device 3 may further include an input device such as a mouse or a keyboard and a display device such as a display.
[0049] Next, with reference to Figure 3 , a summary of the operation of the distance measurement system 1 in the present embodiment will be described.
[0050] In the distance measurement system 1, the imaging device 2 (image sensor 22) generates an image affected by the aberration of the optical system (lens 21) as described above.
[0051] The image processing device 3 (image acquisition unit 32) acquires the image generated by the imaging device 2 and inputs the image to the statistical model stored in the statistical model storage unit 31.
[0052] Here, according to the statistical model in the present embodiment, the distance (predicted value) of the subject in the image input as described above is output. Thus, the image processing device 3 (distance acquisition unit 33) can acquire distance information indicating the distance (distance to the subject in the image) output from the statistical model.
[0053] In this way, in the present embodiment, distance information can be obtained from the image captured by the imaging device 2 using the statistical model.
[0054] Here, with reference to Figure 4 , the principle of predicting the distance to the subject in the present embodiment will be briefly described.
[0055] In the image captured by the imaging device 2 (hereinafter referred to as the captured image), as described above, blurring occurs due to the aberration (lens aberration) of the optical system of the imaging device 2. Specifically, since the refractive index of light when passing through the lens 21 having aberration varies for each wavelength band, for example, when the position of the subject deviates from the focal position (the position focused in the imaging device 2), the light of each wavelength band does not converge at one point but reaches different points. This appears as blurring (chromatic aberration) in the image.
[0056] In addition, in the captured image, blurring (color, size, and shape) that varies non-linearly according to the distance to the subject in the captured image (i.e., the position of the subject relative to the imaging device 2) is observed.
[0057] Therefore, in the present embodiment, as Figure 4 shown, the blurring (blurring information) 402 generated in the captured image 401 is used as a physical clue related to the distance to the subject 403 and analyzed by the statistical model, thereby predicting the distance 404 to the subject 403.
[0058] Hereinafter, with reference to Figure 5, an example of the method for predicting distance based on a captured image in a statistical model will be described. Here, the patch method will be described.
[0059] As Figure 5 shown, in the patch method, a local region (hereinafter referred to as an image patch) 401a is cut out (extracted) from the captured image 401.
[0060] In this case, for example, the entire region of the captured image 401 can be divided into a matrix shape, and the divided partial regions can be sequentially cut out as the image patches 401a, or the captured image 401 can be recognized and the image patches 401a can be cut out in such a way as to include the region where the subject (image) is detected. In addition, the image patch 401a can partially overlap with other image patches 401a.
[0061] In the patch method, the output distance is the predicted value corresponding to the image patch 401a cut out as described above. That is, in the patch method, information related to each of the image patches 401a cut out from the captured image 401 is used as input, and the distance 404 to the subject included in each of the image patches 401a is predicted.
[0062] Figure 6 An example of the information related to the image patch 401a input to the statistical model in the above patch method is shown.
[0063] In the patch method, for the R image, G image, and B image included in the captured image 401, gradient data (gradient data of the R image, gradient data of the G image, and gradient data of the B image) of the image patch 401a cut out from the captured image 401 are respectively generated. Such generated gradient data is input to the statistical model.
[0064] In addition, the gradient data corresponds to the difference (difference value) of the pixel values between each pixel and the pixel adjacent to it. For example, when the image patch 401a is extracted as a rectangular region of n pixels (X-axis direction) × m pixels (Y-axis direction), gradient data (that is, gradient data of each pixel) is generated by arranging the difference values calculated for each pixel in the image patch 401a with respect to the pixel adjacent to the right in a matrix shape of n rows × m columns.
[0065] The statistical model uses the gradient data of the R image, the gradient data of the G image, and the gradient data of the B image to predict the distance based on the blur generated in each of these images. In Figure 6 , the case where the gradient data of the R image, the G image, and the B image are respectively input to the statistical model is shown, but it can also be a structure in which the gradient data of the RGB image is input to the statistical model.
[0066] Here, in the present embodiment, by using the statistical model as described above, it is possible to obtain the distance (distance information indicating the distance) of the subject included in the image from the image. However, in order to improve the accuracy of the distance output from the statistical model, it is necessary to make the statistical model learn.
[0067] Hereinafter, with reference to Figure 7 , a summary of the learning method of a general statistical model will be described. The learning of the statistical model is performed by inputting information related to an image (hereinafter, referred to as a learning image) 501 prepared for this learning into the statistical model, and feeding back the error (loss) between the distance 502 output (predicted) from the statistical model and the correct value 503 to the statistical model. In addition, the correct value 503 refers to the actual distance (measured value) from the shooting location of the learning image 501 to the subject included in the learning image 501, and is also referred to as a correct label, etc. In addition, feedback means updating the parameters (for example, weight coefficients) of the statistical model in such a way as to reduce the error.
[0068] Specifically, in the case where the above-described patch method is applied as a method of predicting the distance from the captured image in the statistical model, for each image patch (local region) cut out from the learning image 501, information (gradient data) related to the image patch is input to the statistical model, and the statistical model outputs the distance 502 as a predicted value corresponding to each image patch. The error obtained by comparing the distance 502 output in this way with the correct value 503 is fed back to the statistical model.
[0069] In the above general learning method of the statistical model, it is necessary to prepare a learning image (that is, a learning data set including the learning image and the correct label which is the distance that should be obtained from the learning image) given the correct label as shown in Figure 8 . In order to obtain the correct label, it is necessary to measure the actual distance of the subject included in the learning image every time the learning image is shot. In order to improve the accuracy of the statistical model, it is necessary to make the statistical model learn a lot of learning data sets, so it is not easy to prepare such a lot of learning data sets.
[0070] Here, in order to make the statistical model learn, it is necessary to evaluate (feedback) the loss (error), which is calculated based on the distance output from the statistical model by inputting the learning image (image patch). However, in the present embodiment, it is assumed that the measured value of the distance to the subject included in the learning image is unknown, but weak supervised learning of the rank loss (order loss) calculated based on a plurality of learning images whose size relationship of the distance is known is performed.
[0071] In addition, weakly-supervised learning based on rank loss is a method of learning based on the relative order relationship (rank) between data. In the present embodiment, a statistical model is learned according to the ranks of two images based on the distance from the imaging device 2 to the subject.
[0072] Here, it is assumed that, as Figure 9 shown, there are five subjects S1 to S5 whose actual distances from the imaging device 2 are unknown, but the magnitude relationship (rank) of the distances is known. In addition, among the subjects S1 to S5, the subject S1 is located at the position closest to the imaging device 2, and the subject S5 is located at the position farthest from the imaging device 2. When the imaging device 2 images such subjects S1 to S5 and the images including each of the subjects S1 to S5 are set as images x1 to x5, the ranks (orders) of the respective images corresponding to the distances to the subjects S1 to S5 included in the images x1 to x5 are as follows: the rank of the image x1 is "1", the rank of the image x2 is "2", the rank of the image x3 is "3", the rank of the image x4 is "4", and the rank of the image x5 is "5".
[0073] It is assumed that, for such images x1 to x5, a statistical model is used to predict, for example, the distance to the subject S2 included in the image x2 and the distance to the subject S5 included in the image x5.
[0074] In this case, if a statistical model that has been sufficiently learned and has high accuracy is used, the distance output from the statistical model by inputting the image x2 should be smaller than the distance output from the statistical model by inputting the image x5.
[0075] That is, in the present embodiment, it is assumed that, for example, in the case where the magnitude relationship between two images x i and the image x k is known, based on the premise that the relationship "if rank(x i ) > rank(x k ) then f θ (x i ) > f θ (x k )" holds, a loss (rank loss) that maintains such a relationship is used to learn the statistical model.
[0076] In this case, rank(x i ) represents the rank (order) assigned to the image x i , and rank(x k ) represents the rank (order) assigned to the image x k . In addition, f θ (x i ) represents the output obtained by inputting the image x iThe distance output from the statistical model f θ (i.e., the predicted value corresponding to the image x i ) is represented by f θ (x k ). That is, the distance output from the statistical model f k through the input image x θ (i.e., the predicted value corresponding to the image x k ). In addition, θ in f θ is a parameter of the statistical model.
[0077] In addition, an image in which the magnitude relationship of the distance from the above-described imaging device 2 to the subject is known can be easily obtained, for example, by imaging sequentially while moving the imaging device 2 in a direction away from the subject S fixed at a predetermined position as shown Figure 10 .
[0078] Generally, in the image captured by the imaging device 2, an identification number (for example, consecutive numbers) is attached in the order in which it is captured. Therefore, in the present embodiment, the identification number attached to the image is used as the rank of the image. That is, when the identification number is small, it can be determined that the distance to the subject included in the image to which the identification number is attached is small (close), and when the identification number is large, it can be determined that the distance to the subject included in the image to which the identification number is attached is large (far).
[0079] In addition, in the image captured by the imaging device 2, in addition to the above-described identification number, the date and time when the image was captured is also attached. Therefore, in the case of imaging images sequentially while moving the imaging device 2 in a direction away from the subject as described above, the magnitude relationship of the distances to the subjects included in the respective images (i.e., the front-back relationship of the ranks of the images) can also be determined based on the date and time attached to the image.
[0080] Here, it has been described that the imaging device 2 is moved in a direction away from the subject while imaging the image, but it may also be configured to image the images sequentially while moving the imaging device 2 in a direction approaching the subject. In this case, when the identification number is small, it can be determined that the distance to the subject included in the image to which the identification number is attached is large (far), and when the identification number is large, it can be determined that the distance to the subject included in the image to which the identification number is attached is small (close).
[0081] In addition, Figure 10 shows a subject having a planar shape, but as such a subject, for example, a television monitor or the like can be used. Here, a subject having a planar shape has been described, but the subject may also be another object having a different shape or the like.
[0082] Hereinafter, a detailed description will be given of the learning processing unit 35 included in the image processing apparatus 3 shown below. Figure 1 Figure 1 Figure 11 FIG. 5 is a block diagram showing an example of the functional configuration of the learning processing unit 35.
[0083] As shown in FIG. Figure 11 5, the learning processing unit 35 includes a discrimination unit 35a, a calculation unit 35b, and a learning unit 35c.
[0084] Here, in the case of learning a statistical model in the present embodiment, the image acquisition unit 32 acquires a plurality of learning images not given the above correct label. In addition, it is assumed that the above identification number is attached to the learning images.
[0085] Based on the identification numbers (orders) attached to each of two learning images out of the plurality of learning images acquired by the image acquisition unit 32, the discrimination unit 35a discriminates the magnitude relationship of the distances of the subjects included in each of the learning images (hereinafter, simply referred to as the magnitude relationship between images).
[0086] Based on the distances output by inputting each of the two learning images whose magnitude relationship has been discriminated by the discrimination unit 35a into the statistical model, and the magnitude relationship between the learning images discriminated by the discrimination unit 35a, the calculation unit 35b calculates the order loss.
[0087] Based on the order loss calculated by the calculation unit 35b, the learning unit 35c causes the statistical model stored in the statistical model storage unit 31 to be learned. The statistical model after the learning by the learning unit 35c is stored in the statistical model storage unit 31 (that is, overwrites the statistical model stored in the statistical model storage unit 31).
[0088] Next, with reference to the flowchart of FIG. Figure 12 6, an example of the processing sequence of the image processing apparatus 3 when learning the statistical model will be described.
[0089] Here, it is assumed that a learned statistical model (previously learned model) is stored in advance in the statistical model storage unit 31 for description, but this statistical model can be generated, for example, by learning an image captured by the imaging device 2, or can be generated by learning an image captured by an imaging device (or lens) different from the imaging device 2. That is, in the present embodiment, at least a statistical model for outputting the distance of the subject included in the image by taking the image as an input needs to be prepared in advance. In addition, the statistical model prepared in advance in the present embodiment can be, for example, a statistical model in a randomly initialized state (unlearned statistical model), etc.
[0090] First, the image acquisition unit 32 acquires a plurality of learning images (hereinafter, referred to as a set of learning images) (step S1). The set of learning images acquired in step S1 is, for example, a set of images captured by the imaging device 2.
[0091] When performing the process of step S1, the learning processing unit 35 selects (acquires) any two learning images, for example, from the set of learning images acquired in step S1 (step S2). In the following description, the two learning images selected in step S2 are set as image x i and image x k .
[0092] When performing the process of step S2, the learning processing unit 35 cuts out an arbitrary region from each of image x i and image x k (step S3). Specifically, the learning processing unit 35 cuts out a region that is at least a part of this image x i from image x i . Similarly, the learning processing unit 35 cuts out a region that is at least a part of this image x k from image x k . In addition, the regions respectively cut out from image x i and image x k in step S3 correspond to the above-mentioned image patches, and are, for example, rectangular regions of n pixels × m pixels.
[0093] Here, it is described that a specified region (image patch) is cut out from image x i and image x k respectively, but this specified region may also be a region that occupies the entirety of image x i and image x k .
[0094] In addition, in the following description, for the sake of convenience, the region cut out from image x i in step S3 is simply set as image x i , and the region cut out from image x k in this step S3 is simply set as image x k .
[0095] Here, in the present embodiment, since the magnitude relationship of the distances to the subjects included in the learning images is known, the discrimination unit 35a included in the learning processing unit 35 discriminates the magnitude relationship between the image x i and the image x k selected in step S2 (the distances to image x i and image x k(The magnitude relationship of the distances of the subjects included in each) (Step S4). The image x i and the image x k The magnitude relationship between them can be determined based on the identification numbers respectively attached to the image x i and the image x k .
[0096] When performing the process of Step S4, the calculation unit 35b included in the learning processing unit 35 uses the statistical model stored in the statistical model storage unit 31 to obtain the distance (predicted value) of the subject included in the image x i and the distance (predicted value) of the subject included in the image x k (Step S5).
[0097] In Step S5, by inputting the image x i (that is, an n-pixel × m-pixel image patch cut out from the image x i ), the distance f θ (x i ) output from the statistical model is obtained, and the image x k (that is, an n-pixel × m-pixel image patch cut out from the image x k ), thereby obtaining the distance f θ (x k ).
[0098] Next, the calculation unit 35b calculates the rank loss (a loss considering the magnitude relationship between the image x i and the image x k based on the distances obtained in Step S5 (hereinafter, referred to as the predicted values corresponding to the image x i and the image x k respectively) (Step S6).
[0099] In Step S6, a loss (rank loss) reflecting whether the magnitude relationship of the predicted values corresponding to the image x i and the image x k respectively is equal to the magnitude relationship between the image x i and the image x k is calculated.
[0100] Here, for example, according to "Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. Learning to rank using gradient descent. In Proceedings of the 22nd international conference on Machine learning, pages 89-96, 2005.", the function representing the ranking loss (ranking loss function) is defined by the following equation (1).
[0101] L rank (x i , x k ) = -y ik (f θ (x k ) - f θ (x i )) + softplus(f θ (x k ) - f θ (x i )) Equation (1)
[0102]
[0103] softplus(x) = log(1 + e x ) Equation (3)
[0104] In this equation (1), L rank (x i , x k ) represents the ranking loss, and y ik corresponds to a label indicating whether the magnitude relationship of the predicted values corresponding to each of the above-mentioned images x i and the image x k is equal to the magnitude relationship between the image x i and the image x k (i.e., the predicted values of the statistical model satisfy the known magnitude relationship). In addition, as shown in Equation (2), when rank(x i ) > rank(x k ), y ik is 1, and when rank(x i ) < rank(x k ), y ik is 0. When k(x i ) > and k(x k ) and k(xi ) k (x k ) in which rank(x i ) > rank(x k ) and rank(x i ) < rank(x k ) are equivalent to the discrimination results of the size relationship between the image x i and the image x k in the above step S4.
[0105] Additionally, the softplus of Equation (1) is a function called soft positive that is used as an activation function and is defined as in Equation (3).
[0106] According to such a rank loss function, when the size relationship of the predicted values corresponding to the image x i and the image x k respectively is equal to the size relationship between the image x i and the image x k , the calculated rank loss (value) becomes smaller, and when the size relationship of the predicted values corresponding to the image x i and the image x k respectively is not equal to the size relationship between the image x i and the image x k , the calculated rank loss (value) becomes larger.
[0107] Next, the learning unit 35c included in the learning processing unit 35 uses the rank loss calculated in step S6 to make the statistical model learn (step S7). The learning of the statistical model is performed by updating the parameters θ of the statistical model, and the update of the parameters θ is performed according to the following optimization problem as in Equation (4).
[0108]
[0109] Here, N in Equation (4) represents the above-mentioned learning image set. Although it is omitted in Figure 12 , the processing of steps S2 to S6 is performed for each group of any image x i and the image x k (regions cut out from each).
[0110] In this case, through Equation (4), the parameter θ' (i.e., the updated parameter) that minimizes the sum of the rank losses L i and the image x k calculated for each group of the image x rank (x i , x k ) can be obtained.
[0111] In addition, when a neural network or a convolutional neural network or the like is applied to the statistical model in the present embodiment (that is, the statistical model is composed of a neural network or a convolutional neural network or the like), the error backpropagation method that calculates the above formula (4) in the opposite direction is used in the learning (update of the parameter θ) of the statistical model. According to this error backpropagation method, the gradient of the rank loss is calculated, and the parameter θ is updated according to this gradient.
[0112] In step S7, by updating the parameter θ of the statistical model to the parameter θ' obtained using the above formula (4), the statistical model can be made to learn the learning image set obtained in step S1.
[0113] In addition, in the present embodiment, for example, a predetermined number of images x i and the image x k are used as an object to perform the Figure 12 processing shown, but the statistical model can also be further learned by repeatedly performing the Figure 12 processing shown.
[0114] In addition, the learning method using the rank loss function such as the above formula (1) is called RankNet, but in the present embodiment, the statistical model can also be learned by other learning methods. Specifically, as the learning method of the statistical model in the present embodiment, for example, FRank, RankBoot, Ranking SVM, or IR SVM can also be used. That is, in the present embodiment, if the learning model is learned in such a manner that the magnitude relationship between the predicted values corresponding to the images x i and the image x k is equal to the magnitude relationship between the images x i and the image x k respectively (that is, learning is performed under the constraints related to the respective ranks of the learning images), various loss functions can be used.
[0115] Next, with reference to the Figure 13 flowchart, an example of the processing order of the image processing device 3 when obtaining distance information from a captured image using the statistical model that has learned the learning image set by performing the Figure 11 processing shown above will be described.
[0116] First, the imaging device 2 (image sensor 22) generates a captured image including the subject by imaging the subject whose distance from the imaging device 2 is measured. This captured image is an image affected by the aberration of the optical system (lens 21) of the imaging device 2 as described above.
[0117] The image acquisition unit 32 included in the image processing apparatus 3 acquires a captured image from the imaging device 2 (step S11).
[0118] Next, the distance acquisition unit 33 inputs information related to the captured image (each of the image patches) acquired in step S11 into the statistical model stored in the statistical model storage unit 31 (step S12). In addition, the information related to the captured image input into the statistical model in step S12 includes the gradient data of each pixel constituting the captured image.
[0119] When the process of step S12 is executed, the distance to the subject is predicted in the statistical model, and the statistical model outputs the predicted distance. Thereby, the distance acquisition unit 33 acquires distance information indicating the distance output from the statistical model (step S13). In addition, the distance information acquired in step S13 includes, for example, the distance of each image patch constituting the captured image acquired in step S11.
[0120] When the process of step S13 is executed, the output unit 34 outputs the distance information acquired in this step S13, for example, in the form of a map arranged corresponding to the captured image in terms of position (step S14). In addition, in the present embodiment, it has been described that the distance information is output in the form of a map, but the distance information may also be output in other forms.
[0121] As described above, in the present embodiment, an image x including a subject captured by the imaging device 2 i and the image x k (the first and second images) are obtained, and the statistical model is made to learn based on the distance (the first distance) output from the statistical model with the image x i (at least a part of the image x i i.e., the first region) as the input and the distance (the second distance) output from the statistical model with the image x k (at least a part of the image x k i.e., the second region) as the input. In the present embodiment, the magnitude relationship between the distance (the third distance) to the subject included in the image x i and the distance (the fourth distance) to the subject included in the image x k (i.e., the magnitude relationship between the image x i and the image x k ) is known, and the statistical model is made to learn in such a way that the magnitude relationship between the predicted value (the first distance) corresponding to the image x i and the predicted value (the second distance) corresponding to the image x k is equal to the front-back relationship between the image x i and the image x k .
[0122] In this embodiment, with such a structure, even for learning images that are not given correct labels (teaching labels), the statistical model can be made to learn, thus improving the ease of learning in this model.
[0123] In addition, in this embodiment, it is assumed that while moving the imaging device 2 away from the subject fixed at a prescribed position, for example, imaging is performed for a plurality of learning images including image x i and image x k Thereby, based on the identification numbers (for example, consecutive numbers) attached to each of these learning images in the order of imaging, the magnitude relationship of the distances to the subjects included in each learning image can be easily determined.
[0124] In addition, a plurality of learning images including image x i and image x k can also be imaged while moving the imaging device 2 closer to the subject, for example.
[0125] In addition, in this embodiment, it has been described that the magnitude relationship of the distances to the subjects included in each of the plurality of learning images is determined based on the identification numbers attached to the learning images, but this magnitude relationship can also be determined based on the position of the imaging device 2 when imaging the learning images with the subject fixed as described above. Such a position of the imaging device 2 only needs to be attached to the learning images.
[0126] Here, for example, there is a case where an internal sensor (such as a gyro sensor or an acceleration sensor) is mounted on the imaging device 2, and based on the signal detected by this internal sensor, the movement (trajectory) of the imaging device 2 can be calculated. In this case, the position of the imaging device 2 when imaging the above-described learning images can be obtained based on the operation of the imaging device 2 calculated from the signal obtained by the internal sensor.
[0127] In addition, for example, when imaging learning images using a workbench having a moving mechanism for moving the imaging device 2, the position of the imaging device 2 when imaging the learning images can also be obtained based on the position of this workbench.
[0128] In addition, as the subject included in the learning images in this embodiment, for example, a television monitor having a planar shape can be used. When using such a television monitor as the subject, since various images can be switched and displayed on the television monitor, the statistical model can learn various color patterns (of learning images).
[0129] Furthermore, in the present embodiment, it has been described that when the statistical model is learned, any two learning images are selected from the set of learning images (that is, the learning images are randomly selected). However, as the two learning images, for example, learning images in which the difference in the distance to the subject becomes equal to or greater than a predetermined value may be preferentially selected. In addition, although the distances (measured values) to the subjects included in the respective learning images are unknown, since the order in which each of these learning images is captured (that is, the magnitude relationship of the distances to the subjects) is known based on the identification numbers, for example, by selecting two learning images in which the difference in the identification numbers attached to the learning images is equal to or greater than a predetermined value, it is possible to select images that are presumed to have a difference in the distance to the subject equal to or greater than a predetermined value. Thereby, misrecognition (confusion) of the magnitude relationship between the learning images can be eliminated.
[0130] In addition, when the learning images are captured, due to the operation of the imaging device 2, there may be a situation where images are continuously captured although the subject has not moved. Therefore, two learning images in which the difference in the captured times (date and time) is equal to or greater than a predetermined value may also be preferentially selected.
[0131] In addition, when the statistical model is learned, an arbitrary region is cut out from each of the two learning images selected from the set of learning images (that is, the region is randomly cut out). However, the region may be cut out, for example, based on a predetermined regularity corresponding to the position, pixel value, etc. in each learning image.
[0132] In addition, in the present embodiment, the patch method has been described as an example of the method for predicting the distance based on the image in the statistical model. However, as the method for predicting the distance based on the image, for example, a screen unification method in which the entire region of the image is input to the statistical model and a predicted value (distance) corresponding to the entire region is output may also be used.
[0133] In addition, in the present embodiment, it has been described that the statistical model is generated by learning the learning images affected by the aberration of the optical system (blur that changes non-linearly according to the distance to the subject included in the learning images). However, the statistical model may be generated, for example, by learning the learning images based on the light that has passed through a filter (color filter, etc.) provided at the opening of the imaging device 2 (that is, the blur that is intentionally generated in the image by the filter and changes non-linearly according to the distance to the subject).
[0134] (Second Embodiment)
[0135] Next, a second embodiment will be described. Regarding the structure of the distance measurement system (imaging device and image processing device) in this embodiment, since it is the same as that of the foregoing first embodiment, when describing the structure of the distance measurement system in this embodiment, Figure 1 and the like are appropriately used. Here, the points different from the foregoing first embodiment will be mainly described. Figure 1 And so on. Here, the points different from the foregoing first embodiment will be mainly described.
[0136] In the foregoing first embodiment, it has been described that the statistical model outputs the distance of the subject included in the image. However, in this embodiment, the statistical model is configured to output the degree of unreliability (hereinafter referred to as unreliability) with respect to this distance (that is, the predicted value) together with the distance. The difference between this embodiment and the foregoing first embodiment is that the rank loss (rank loss function) reflecting the unreliability output from the statistical model is used to make the statistical model learn. In addition, the unreliability is represented by a real number of 0 or more, for example, and the larger the value, the higher the reliability. The calculation method of the unreliability is not limited to a specific method, and various known methods can be applied.
[0137] Hereinafter, with reference to Figure 14 the flowchart of, an example of the processing sequence of the image processing device 3 when making the statistical model learn in this embodiment will be described.
[0138] First, the processing of steps S21 to S24 corresponding to the processing of steps S1 to S4 shown in the foregoing Figure 12 is executed.
[0139] When the processing of step S24 is executed, the calculation unit 35b included in the learning processing unit 35 uses the statistical model stored in the statistical model storage unit 31 to obtain the distance of the subject included in the obtained image x i and the unreliability with respect to this distance (the predicted value and unreliability corresponding to the image x i ), and the distance of the subject included in the image x k and the unreliability with respect to this distance (the predicted value and unreliability corresponding to the image x k ) (step S25).
[0140] Here, if the above unreliability is represented by σ, in step S5, the distance f i output from the statistical model f i by inputting the image x θ (that is, an n-pixel × m-pixel image patch cut out from the image x θ ) into the statistical model, and the unreliability σ i are obtained, and the image x i (that is, from the image x k ) is obtained.k The distance f output from the statistical model f for the input of the cut-out n-pixel × m-pixel image patch θ and the unreliability σ θ (x k ) k .
[0141] Next, the calculation unit 35b calculates a rank loss based on the distance and unreliability obtained in step S25 (step S26).
[0142] In the above-described first embodiment, it was described that the rank loss is calculated using Equation (1), but the function (rank loss function) representing the rank loss in the present embodiment is defined as Equation (5) below.
[0143]
[0144] σ = max(σ i , σ k ) Equation (6)
[0145] In this Equation (5), L uncrt (x i , x k ) represents the rank loss calculated in the present embodiment, and L rank (x i , x k ) is the same as L rank (x i , x k ) in Equation (1) in the above-described first embodiment.
[0146] Here, for example, when a region without texture or a light-saturated (i.e., washed-out) region is cut out in step S23, it is difficult to output a high-precision distance (i.e., a correctly predicted distance) from the statistical model. However, in the above-described first embodiment, even in a region where such clues for predicting the distance are absent or scarce (hereinafter, referred to as a difficult-to-predict region), learning is performed in such a way as to satisfy the size relationship between the image x i and the image x k , so overfitting may occur. In this case, the statistical model is optimized for the difficult-to-predict region, and the generality of the statistical model is reduced.
[0147] Therefore, in the present embodiment, as shown in Equation (5) above, the unreliability σ is added to the loss function, and thus the rank loss considering the predictability (unpredictability) in the above-described difficult-to-predict region is calculated. In addition, σ in Equation (5) is, as defined in Equation (6), the unreliability with a larger value among the unreliability σ i and the unreliability σ k .
[0148] According to the ranking loss function (unreliability ranking loss function) as in Equation (5), L cannot be reduced (decreased) in the difficult-to-predict region. rank (x i ,x k ) When this is the case, by increasing at least one of the unreliability degrees σ i and the unreliability degree σ k (that is, the unreliability degree σ), it is possible to adjust to reduce the ranking loss L in the present embodiment, that is, L uncrt (x i ,x k ). However, in order to prevent L uncrt (x i ,x k ) from decreasing excessively due to an excessive increase in the unreliability degree σ, a second term is added to the right side of Equation (5) as a penalty.
[0149] In addition, the ranking loss function shown in Equation (5) can be obtained, for example, by expanding the definition formula of the uneven dispersion.
[0150] When the process of step S26 is executed, the process of step S27 corresponding to the process of step S7 shown in Figure 12 is executed. In addition, in this step S27, as long as L of Equation (4) described in the first embodiment above rank (x i ,x k ) is used as L uncrt (x i ,x k ), the statistical model can be made to learn.
[0151] As described above, in the present embodiment, when the statistical model is made to learn in such a way that the ranking loss calculated based on the predicted values (first distance and second distance) corresponding to the image x i and the image x k is minimized, the ranking loss is adjusted based on at least one of the unreliability degrees (first and second unreliability degrees) corresponding to the image x i and the image x k output from the statistical model.
[0152] In the present embodiment, with such a structure, it is possible to mitigate the influence of the above-mentioned difficult-to-predict region on the learning of the statistical model, and thus it is possible to achieve learning of a highly accurate statistical model.
[0153] (Third Embodiment)
[0154] Next, a third embodiment will be described. Since the structure of the distance measurement system (imaging device and image processing device) in this embodiment is the same as that of the foregoing first embodiment, when the structure of the distance measurement system is described in this embodiment, Figure 1 etc. are appropriately used. Here, the points different from the foregoing first embodiment will be mainly described. Figure 1 Here, the points different from the foregoing first embodiment will be mainly described.
[0155] The difference between this embodiment and the foregoing first embodiment is that the statistical model is learned in such a way that the size relationship between two learning images is satisfied and the deviation of the distances (predicted values) corresponding to two different regions within the same learning image is minimized. In addition, in this embodiment, it is assumed that a television monitor or the like having a planar shape is used as the subject included in the learning image.
[0156] Hereinafter, an example of the processing sequence of the image processing device 3 when the statistical model is learned in this embodiment will be described. Here, for convenience, Figure 12 is used to describe with a flowchart. Figure 12 is used to describe with a flowchart.
[0157] First, the processes of steps S1 and S2 described in the foregoing first embodiment are executed. In the following description, the two learning images selected in step S2 are set as image x i and image x k .
[0158] When the process of step S2 is executed, the learning processing unit 35 cuts out an arbitrary region from each of image x i and image x k (step S3).
[0159] Here, in the foregoing first embodiment, it was described that one region was cut out from image x i and image x k respectively, but in this embodiment, for example, two regions are cut out from image x i and one region is cut out from image x k .
[0160] In addition, in the foregoing first embodiment, it was described that a region occupying the whole of image x i and image x k could be cut out, but in this embodiment, it is set that regions (image patches) of a part of image x i and image x k are cut out.
[0161] In the following description, for convenience, the two regions cut out from image x i in step S3 are set as image x i1 and image xi2 that is cut out from the image x in the step S3 k is simply set as the image x k .
[0162] When the process of step S3 is executed, the processes of steps S4 and S5 described in the foregoing first embodiment are executed. In addition, in step S5, the distances f i1 output from the statistical model f θ by inputting the image x θ (x i1 ), the distances f i2 output from the statistical model f θ by inputting the image x θ (x i2 ), and the distances f k output from the statistical model f θ by inputting the image x θ (x k ) are obtained.
[0163] Next, the calculation unit 35b calculates a rank loss (step S6) based on the distances (predicted values corresponding to the image x i1 , the image x i2 , and the image x k ) obtained in step S5.
[0164] Here, since the subject included in the learning image in the present embodiment has a planar shape, the distances to the subjects included in the same learning image are the same. In the present embodiment, focusing on this point, the statistical model is learned in such a way that the deviation between the predicted values corresponding to the image x i1 and the image x i2 (that is, two regions cut out from the same image x i ) is minimized.
[0165] In this case, the function (rank loss function) representing the rank loss in the present embodiment is defined as the following formula (7).
[0166] L intra (x i1 , x k , x i2 ) = L rank (x i1 , x k ) + λ|f θ (x i1 ) - f θ (x i 2)| Formula (7)
[0167] rank(x i1) ≠ rank(x k ),rank(x i1 ) = rank(x i2 ) Equation (8)
[0168] In this Equation (7), L intra (x i1 , x i2 , x k ) represents the rank loss calculated in this embodiment. L rank (x i1 , x k ) is equivalent to L rank (x i , x k ) in Equation (1) of the foregoing first embodiment. That is, L rank (x i1 , x k ) calculates the image x i in Equation (1) as the image x i1 .
[0169] In addition, the second term on the right side of Equation (7) represents the deviation (difference) between the distance (predicted value) corresponding to the image x i1 and the distance (predicted value) corresponding to the image x i2 . In this second term, λ is an arbitrary coefficient (λ > 0) used to achieve balance with the first term on the right side.
[0170] In addition, in this embodiment, the image x i1 and the image x i2 are respectively regions cut from the same image x i . Therefore, the size relationship among the image x i1 , the image x i2 and the image x k (that is, the front - back relationship of the ranks of the image x i1 , the image x i2 and the image x k ) satisfies Equation (8).
[0171] When the process of step S6 is executed, the process of step S7 described in the foregoing first embodiment is executed. In this step S7, as long as L kran (x i , x k ) in Equation (4) described in the foregoing first embodiment is used as L intra (x i1 , x i2 , x k ), the statistical model can be made to learn.
[0172] As described above, in the present embodiment, by minimizing the difference between the distances (the first distance and the fifth distance) output from the statistical model for each of the two regions (the first and third regions) cut out from the image x i The configuration in which the statistical model is learned in such a manner enables learning of a statistical model with higher accuracy that takes into account the deviation of the distances corresponding to the respective regions within the same learning image, as compared with the aforementioned first embodiment.
[0173] In the present embodiment, it is assumed that the rank loss is calculated by taking into account the deviation of the distances corresponding to the respective regions within the image x i and the image x k Among them, for the image x i It has been described that the rank loss is calculated while taking into account the deviation of the distances corresponding to the respective regions within the image. However, for example, as in the following formula (9), a rank loss function that calculates the rank loss while further taking into account the deviation of the distances corresponding to the respective regions within the image x k can be used.
[0174] L intra (x i1 , x k1 , x i2 , x k2 ) = L rank (x i1 , x k1 ) + λ|f θ (x i1 ) - f θ (x i2 )| + λ|f θ (x k1 ) - f θ (x k2 )| Formula (9)
[0175] In addition, in formula (9), the two regions cut out from the image x k are respectively represented as the image x k1 and the image x k2 .
[0176] In addition, the present embodiment can also be configured to be combined with the aforementioned second embodiment. In this case, a rank loss function such as the following formula (10) can be used.
[0177]
[0178] According to at least one of the above-described embodiments, it is possible to provide a learning method, a program, and an image processing apparatus that can improve the ease of learning in a statistical model.
[0179] Several embodiments of the present invention have been described, but these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other ways, and various omissions, substitutions, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope or gist of the invention, and are also included in the scope of the invention described in the claims and its equivalents.
[0180] In addition, the above embodiments can be summarized into the following technical solutions.
[0181] [Technical Solution 1]
[0182] A learning method for causing a statistical model to learn, the statistical model being used to output the distance to a subject when an image including the subject is input.
[0183] The learning method includes the following steps:
[0184] Obtaining a first image and a second image including a subject captured by an imaging device; and
[0185] Based on a first distance output from the statistical model when a first region, which is at least a part of the first image, is input and a second distance output from the statistical model when a second region, which is at least a part of the second image, is input, causing the statistical model to learn.
[0186] The magnitude relationship between a third distance to the subject included in the first image and a fourth distance to the subject included in the second image is known.
[0187] The learning includes: causing the statistical model to learn in such a way that the magnitude relationship between the first distance and the second distance is equal to the magnitude relationship between the third distance and the fourth distance.
[0188] [Technical Solution 2]
[0189] According to the above Technical Solution 1,
[0190] The statistical model outputs the first distance and a first degree of unreliability of the first distance when the first region is input, and outputs the second distance and a second degree of unreliability of the second distance when the second region is input.
[0191] The learning includes: causing the statistical model to learn in such a way that a rank loss calculated based on the first distance and the second distance output from the statistical model is minimized.
[0192] The rank loss is adjusted based on at least one of the first degree of unreliability and the second degree of unreliability.
[0193] [Technical Solution 3]
[0194] According to the above Technical Solution 1 or 2,
[0195] The statistical model takes, as an input, a third region that is at least a part of the first image and different from the first region, and outputs a fifth distance.
[0196] The learning includes: making the statistical model learn in such a way that the difference between the first distance and the fifth distance is minimized.
[0197] [Technical Solution 4]
[0198] According to the above Technical Solutions 1 to 3,
[0199] The first image and the second image are captured by the imaging device while the imaging device is moving in a direction away from the subject.
[0200] An identification number is assigned to the first image and the second image, and this identification number represents the order in which the images are captured by the imaging device.
[0201] The magnitude relationship between the third distance and the fourth distance is determined based on the identification numbers assigned to the first image and the second image.
[0202] [Technical Solution 5]
[0203] According to the above Technical Solutions 1 to 3,
[0204] The first image and the second image are captured by the imaging device while the imaging device is moving in a direction approaching the subject.
[0205] An identification number is assigned to the first image and the second image, and this identification number represents the order in which the images are captured by the imaging device.
[0206] The magnitude relationship between the third distance and the fourth distance is determined based on the identification numbers assigned to the first image and the second image.
[0207] [Technical Solution 6]
[0208] According to the above Technical Solutions 1 to 3,
[0209] The magnitude relationship between the third distance and the fourth distance is determined based on the position of the imaging device when the first image and the second image are captured by the imaging device.
[0210] [Technical Solution 7]
[0211] According to the technical solution 6,
[0212] When the imaging device captures the first image and the second image, the position of the imaging device is obtained by a sensor mounted on the imaging device.
[0213] [Technical solution 8]
[0214] According to the technical solution 6,
[0215] When the imaging device captures the first image and the second image, the position of the imaging device is obtained based on the position of a moving mechanism that moves the imaging device.
[0216] [Technical solution 9]
[0217] According to the technical solutions 1 to 8,
[0218] The shape of the subject is a planar shape.
[0219] [Technical solution 10]
[0220] According to the technical solutions 1 to 9,
[0221] The difference between the third distance and the fourth distance is equal to or greater than a predetermined value.
[0222] [Technical solution 11]
[0223] According to the technical solutions 1 to 10,
[0224] The difference between the first time when the first image is captured and the second time when the second image is captured is equal to or greater than a predetermined value.
[0225] [Technical solution 12]
[0226] According to the technical solutions 1 to 11,
[0227] The statistical model is generated by learning a blur that is generated in an image affected by the aberration of an optical system and that varies non-linearly according to the distance to the subject included in the image.
[0228] [Technical solution 13]
[0229] According to the technical solutions 1 to 11,
[0230] The statistical model is generated by learning a blur that is generated in an image generated based on light that has passed through a filter and that varies non-linearly according to the distance to the subject included in the image.
[0231] [Technical solution 14]
[0232] A program that causes a statistical model to learn, the statistical model being for taking an image including a subject as an input and outputting a distance to the subject
[0233] The program causes a computer to perform the following processing:
[0234] Acquire a first image and a second image including a subject captured by an imaging device; and
[0235] Cause the statistical model to learn based on a first distance output from the statistical model with a first region that is at least a part of the first image as an input and a second distance output from the statistical model with a second region that is at least a part of the second image as an input
[0236] The magnitude relationship between a third distance to the subject included in the first image and a fourth distance to the subject included in the second image is known
[0237] The learning includes: causing the statistical model to learn in such a manner that the magnitude relationship between the first distance and the second distance is equal to the magnitude relationship between the third distance and the fourth distance
[0238] [Technical solution 15]
[0239] An image processing device that causes a statistical model to learn, the statistical model being for taking an image including a subject as an input and outputting a distance to the subject
[0240] The image processing device includes:
[0241] An acquisition unit that acquires a first image and a second image including a subject captured by an imaging device; and
[0242] A learning unit that causes the statistical model to learn based on a first distance output from the statistical model with a first region that is at least a part of the first image as an input and a second distance output from the statistical model with a second region that is at least a part of the second image as an input
[0243] The magnitude relationship between a third distance to the subject included in the first image and a fourth distance to the subject included in the second image is known
[0244] The learning includes: causing the statistical model to learn in such a manner that the magnitude relationship between the first distance and the second distance is equal to the magnitude relationship between the third distance and the fourth distance
Claims
1. A learning method for causing a statistical model to learn, the statistical model being configured to take an image including a subject as an input and output a distance to the subject. The learning method includes the following steps: Obtaining a first image and a second image including the subject captured by an imaging device; and Causing the statistical model to learn based on a first distance output from the statistical model with a first region, which is at least a part of the first image, as an input and a second distance output from the statistical model with a second region, which is at least a part of the second image, as an input. The magnitude relationship between a third distance to the subject included in the first image and a fourth distance to the subject included in the second image is known. The learning includes: Causing the statistical model to learn such that the magnitude relationship between the first distance and the second distance is equal to the magnitude relationship between the third distance and the fourth distance.
2. The learning method according to claim 1, wherein the statistical model takes the first region as an input and outputs the first distance and a first uncertainty degree of the first distance, takes the second region as an input and outputs the second distance and a second uncertainty degree of the second distance, The learning includes: causing the statistical model to learn such that a rank loss calculated based on the first distance and the second distance output from the statistical model is minimized, wherein the rank loss is adjusted based on at least one of the first uncertainty degree and the second uncertainty degree.
3. The learning method according to claim 1 or 2, wherein the statistical model takes a third region, which is at least a part of the first image and different from the first region, as an input and outputs a fifth distance, The learning includes: causing the statistical model to learn such that a difference between the first distance and the fifth distance is minimized.
4. The learning method according to claim 1, wherein the first image and the second image are captured by the imaging device while the imaging device is moving in a direction away from the subject, assigning identification numbers to the first image and the second image, the identification numbers indicating the order in which the imaging device captures images, wherein the magnitude relationship between the third distance and the fourth distance is determined based on the identification numbers attached to the first image and the second image.
5. The learning method according to claim 1, wherein the first image and the second image are captured by the imaging device while the imaging device is moving in a direction approaching the subject, assigning identification numbers to the first image and the second image, the identification numbers indicating the order in which the imaging device captures images, wherein the magnitude relationship between the third distance and the fourth distance is determined based on the identification numbers attached to the first image and the second image.
6. The learning method according to claim 1, wherein the magnitude relationship between the third distance and the fourth distance is determined based on the position of the imaging device when the imaging device captures the first image and the second image.
7. The learning method according to claim 6, The position of the imaging device when the first image and the second image are captured by the imaging device is obtained by a sensor mounted on the imaging device.
8. The learning method according to claim 6, The position of the imaging device when the first image and the second image are captured by the imaging device is obtained based on the position of a moving mechanism that moves the imaging device.
9. The learning method according to claim 1, The shape of the subject is a planar shape.
10. The learning method according to claim 1, The difference between the third distance and the fourth distance is equal to or greater than a predetermined value.
Citation Information
Patent Citations
Vehicle ranging method based on deep neural network
CN110068302A
Image processing device
CN111683193A