Processing apparatus, processing method, and program
The processing device corrects distance values using neighborhood pixel information to generate accurate high-density depth maps, addressing inaccuracies from differing viewpoints and enhancing 3D distance information precision.
Patent Information
- Application Number
- JP2023191223
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-21
AI Technical Summary
Existing methods for generating dense depth maps from 3D distance information fail to accurately reproduce distance values near object boundaries due to differing viewpoints between 3D distance measuring devices and cameras, leading to inaccurate high-density three-dimensional distance information.
A processing device that generates a distance image from a second viewpoint using distance information from a first viewpoint and modifies distance values based on neighborhood pixel values to correct inaccuracies, employing a depth correction unit and interpolation techniques.
Enables the acquisition of highly accurate and high-density three-dimensional distance information by correcting distance values near object boundaries, resulting in improved depth map accuracy.
Smart Images

Figure 2025078918000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a processing device, a processing method, and a program. [Background technology]
[0002] In recent years, 3D distance information, which is used for various purposes, is required to be highly accurate and high density. 3D distance measuring devices such as LiDAR sensors are known as a means of acquiring highly accurate 3D distance information.
[0003] On the other hand, high-density three-dimensional distance information is difficult to apply to moving objects because it takes time to acquire, and the amount of information is large, which increases the processing load. Non-Patent Document 1 discloses a method of converting three-dimensional distance information into a depth map format. A depth map is image data in which a value corresponding to the distance to a subject (for example, a value proportional to the distance) is stored as a pixel value for each pixel. In general, a three-dimensional distance measuring device acquires data at a constant sampling interval, so that when projected onto a depth map, pixels that do not store pixel values are generated. By inputting such a sparse depth map together with an image captured by a camera into a CNN (Convolutional Neural Network), a dense depth map in which pixel values are stored in all pixels can be obtained. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Xinjing Cheng, Peng Wang and Ruigang Yang,”Depth Estimation via Affinity Learned with Convolutional Spatial Propagation Network” Summary of the Invention [Problem to be solved by the invention]
[0005] When the viewpoint positions of the three-dimensional distance measuring device and the camera are different, the occlusion relationship between the objects may be different between the viewpoints of the three-dimensional distance measuring device and the camera. That is, a situation may occur in which a distant object that is not occluded when viewed from the three-dimensional distance measuring device is occluded by a nearby object when viewed from the camera. If a depth map viewed from the camera is generated under such a situation, the depth map is sparse, and a value corresponding to the distance of the distant object is stored in the depth map. Therefore, if a dense depth map is generated using the method of Non-Patent Document 1, the distance value near the object boundary becomes inaccurate, and accurate three-dimensional distance information cannot be reproduced.
[0006] An object of the present invention is to provide a processing device capable of acquiring highly accurate and high density three-dimensional distance information. [Means for solving the problem]
[0007] A processing device as one aspect of the present invention is characterized in having a generation means for generating a distance image viewed from a second viewpoint using distance information to a subject viewed from a first viewpoint, and a modification means for deciding whether to modify the distance value depending on the distance value of a pixel included in a neighborhood area of a reference pixel in the distance image. Effect of the Invention
[0008] According to the present invention, it is possible to provide a processing device capable of acquiring highly accurate and high density three-dimensional distance information. [Brief description of the drawings]
[0009] [Figure 1] 1 is a block diagram of a three-dimensional distance information processing apparatus according to a first embodiment. [Diagram 2] 13 is a flowchart showing three-dimensional distance information processing. [Diagram 3] FIG. 1 is a conceptual diagram of acquiring three-dimensional distance information and an image. [Figure 4] 11 is a flowchart showing a depth correction process according to the first embodiment. [Diagram 5]FIG. 4 is an explanatory diagram of the effect of the configuration of the first embodiment. [Figure 6] FIG. 11 is a conceptual diagram illustrating the operation of a depth correction unit according to the second embodiment. [Figure 7] 13 is a flowchart showing a depth correction process according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, the same reference numerals are used to refer to the same components, and duplicated descriptions will be omitted. EXAMPLES
[0011] 1 is a block diagram of a three-dimensional distance information processing device (processing device) 100 according to this embodiment. The three-dimensional distance information processing device 100 includes a three-dimensional distance measuring unit 101, an imaging unit 102, a system memory 109, a non-volatile memory 110, and a control unit 120.
[0012] The three-dimensional distance measuring unit 101 is, for example, a LiDAR sensor. The LiDAR sensor includes a light output unit that outputs laser light to irradiate a light beam (irradiation light beam) on the surface of the object, and a receiving unit that receives a reflected light beam from the surface of the object. The distance data to the surface of the structure in the irradiation direction is obtained using the time until the reflected light beam returns, or the phase difference between the irradiation light beam and the reflected light beam, and three-dimensional distance information (distance information to the subject as viewed from the first viewpoint) is obtained by combining the data with information on the irradiation direction. Note that the three-dimensional distance measuring unit 101 is not limited to a LiDAR sensor. For example, it may be a device that measures distance using electromagnetic waves or sound waves other than laser light. In addition, the three-dimensional distance information is assumed to be composed of at least one of a point cloud, a voxel, a polygon, a mesh, an implicit function display, a depth map, a parallax map, and a depth map.
[0013] The imaging unit 102 is, for example, a camera. The camera includes a lens unit, an imaging element that converts an optical image into an electrical signal, and an A / D converter that converts an analog signal into a digital signal. An optical image is input to the imaging element via the lens unit, and the imaging element converts the converted electrical signal into a digital signal to obtain an image (captured image).
[0014] The system memory 109 is a rewritable volatile memory such as a DRAM, and stores constants and variables during operation of the control unit 120, data read from the non-volatile memory 110, and the like.
[0015] The non-volatile memory 110 includes a three-dimensional distance information storage unit 103, a system storage unit 104, and an image storage unit 105. The non-volatile memory 110 is an electrically erasable and recordable memory, and for example, an EEPROM or the like is used. The three-dimensional distance information storage unit 103 stores the three-dimensional distance information acquired from the three-dimensional distance measuring unit 101 based on a specified format. For example, the information is saved in a point cloud format. The system storage unit 104 stores operation programs and operation constants of each block of the control unit 120. The programs referred to here are programs for executing various flowcharts described later in this embodiment. The image storage unit 105 stores images acquired from the imaging unit 102.
[0016] The control unit 120 includes at least one processor, and controls the entire three-dimensional distance information processing device 100. The control unit 120 includes a projection unit (generation means) 106, a depth correction unit (correction means, setting means) 107, and a depth interpolation unit (interpolation means) 108. The control unit 120 realizes each process of this embodiment, which will be described later, by executing a program recorded in the system storage unit 104. Note that the various controls performed by the control unit 120 may be performed by one piece of hardware, or the processes may be shared and performed by multiple pieces of hardware (e.g., multiple processors or circuits).
[0017] The projection unit 106 generates a depth map (distance image seen from a second viewpoint) seen from the imaging unit 102 from the three-dimensional distance information recorded in the three-dimensional distance information storage unit 103 and the relative position information of the three-dimensional distance measuring unit 101 and the imaging unit 102 recorded in the system storage unit 104. Here, the depth map is image data in which a value corresponding to the distance to the subject (for example, a value proportional to the distance) is stored as a pixel value in each pixel. In this embodiment, a value obtained by projecting the position vector of the subject seen from the imaging unit 102 in the optical axis direction is used as the pixel value, but the length of the position vector may also be used. Also, in this embodiment, a case where the distance image is a depth map will be described, but a parallax map or a depth map may also be used.
[0018] In this embodiment, a pixel value of 0 is assigned to a pixel corresponding to a distance to the subject of 0 [mm], and a pixel value of 255 is assigned to a pixel corresponding to a distance to the subject of equal to or greater than a threshold (for example, 100000 [mm] = 100 [m]). A pixel value of 0 is stored for pixels that do not have a corresponding distance to the subject. In the following description, pixel values on the depth map are referred to as distance values to distinguish them from pixel values in an image acquired by the imaging unit 102. This value does not necessarily have to match the physical distance.
[0019] Since the LiDAR sensor scans the laser light based on a certain angular resolution, the obtained depth map generally has a certain number of pixels with missing distance values. Although it is possible to increase the number of scans and perform finer sampling, this method cannot be used when a moving object is present. In addition, the amount of data for the three-dimensional distance information is also huge. In the following description, a depth map with pixels with missing distance values is referred to as a sparse depth map. Note that the number of pixels in the depth map may be different from the number of pixels in the image acquired by the imaging unit 102. The sparse depth map is stored in the system storage unit 104 and the system memory 109.
[0020] The depth modification unit 107 determines whether to modify the distance value according to the distance value of a pixel included in a neighborhood area of a reference pixel of the sparse depth map output by the projection unit 106. For example, the depth modification unit 107 may determine whether to modify the distance value according to a difference between the distance value of a reference pixel included in the neighborhood area of the reference pixel (reference distance value) as a reference and the distance value of each pixel included in the neighborhood area, or the magnitude of the distance value of each pixel included in the neighborhood area.
[0021] The depth interpolation unit 108 generates a depth map (hereinafter, dense depth map) in which distance values are stored in all pixels, based on the sparse depth map corrected by the depth correction unit 107 and the image acquired by the imaging unit 102. For example, the dense depth map can be generated by using a convolutional neural network (CNN) from a combination of an image and a corresponding sparse depth map. In addition, the interpolation may be performed using at least one of a filter, machine learning, super-resolution, and a conditional random field.
[0022] FIG. 2 is a flowchart showing three-dimensional distance information processing. In step S201, the image storage unit 105 acquires and records an image from the imaging unit 102. In step S202, the three-dimensional distance information storage unit 103 acquires and stores three-dimensional distance information from the three-dimensional distance measurement unit 101. In step S203, the projection unit 106 generates a depth map viewed from the imaging unit 102 from the three-dimensional distance information recorded in the three-dimensional distance information storage unit 103 and the relative position information of the three-dimensional distance measurement unit 101 and the imaging unit 102 recorded in the system storage unit 104. Note that in this embodiment, the relative position information of the three-dimensional distance measurement unit 101 and the imaging unit 102 is known, but in reality, this is not limited to this. For example, matching of the three-dimensional point group and feature points in the image may be performed, and calculation may be performed using a known calibration method. In step S204, the depth correction unit 107 executes a depth correction process for correcting the distance value of the depth map. In step S205, the depth interpolator 108 generates a dense depth map based on the modified depth map and the image.
[0023] The process (depth correction process) of step S204 in FIG. 2 will be described below. FIG. 3 is a conceptual diagram of acquiring three-dimensional distance information and an image using the three-dimensional distance measuring unit 101 and the imaging unit 102. In FIG. 3(a), an area 303 is an area of the wall 301 that is blocked by the object 302 when viewed from the imaging unit 102. In FIG. 3(b), an area 320 of the depth map 310 generated by the projection unit 106 and viewed from the imaging unit 102 is an area corresponding to the object 302. Here, for convenience of illustration, the distance values outside the area 320 are expressed as 0, but it is assumed that other values are actually stored. Within the area 320, a distance value corresponding to the distance to the wall 301 is stored in pixel 311, and a distance value corresponding to the distance to the object 302 is stored in pixel 312.
[0024] Since the three-dimensional distance measuring unit 101 and the imaging unit 102 are at different positions in space, the wall 301 and the subject 302 have different parallax. That is, when viewed from the imaging unit 102, the area 303 is occluded by the subject 302 and does not appear on the depth map. However, since the data acquired by the three-dimensional distance measuring unit 101 is sparse, the area 320 contains a mixture of distance values originating from the wall 301 and distance values originating from the subject 302. Therefore, when an interpolation process is performed, the distance values in the area 320 are averaged, and a distance value different from the original value is stored. In FIG. 3, the parallax between the three-dimensional distance measuring unit 101 and the imaging unit 102 is highlighted, but the actual parallax is not as large as shown in the figure, so a phenomenon occurs in which the distance value becomes inaccurate near the boundary as the subject is closer to the imaging unit 102.
[0025] In order to solve the above problem, in this embodiment, the distance value derived from the wall 301 in the area 320 is corrected according to the flow of Fig. 4. Fig. 4 is a flowchart showing the depth correction process of this embodiment.
[0026] In step S401, the depth modification unit 107 obtains a distance value (reference distance value) D of a pixel of interest (reference pixel) of the depth map.
[0027] In step S402, the depth modification unit 107 determines whether the distance value D is greater than 0. If the depth modification unit 107 determines that the distance value D is greater than 0, it executes the process of step S404, and if it determines that the distance value D is not greater than 0, i.e., if the distance value D is equal to 0, it executes the process of step S403.
[0028] In step S403, the depth modification unit 107 updates the pixel of interest.
[0029] In step S404, the depth correction unit 107 determines whether the distance value d of the peripheral pixel included in the neighborhood of the pixel of interest satisfies a predetermined condition. In this embodiment, the depth correction unit 107 determines whether the difference value dD between the distance value d of the peripheral pixel and the distance value D is greater than a threshold value (first predetermined value) Th1. Here, the neighborhood is set by the depth correction unit 107, is a square region of (2b+1)×(2b+1) pixels centered on the pixel of interest, and may include the pixel of interest. For example, b is 1 [pix]. The shape of the neighborhood is not limited to a square, and may be other shapes such as a rectangle or a circle. The size of the neighborhood may be changed according to the distance value D. For example, when the three-dimensional distance measuring unit 101 and the imaging unit 102 are on approximately the same plane, the parallax between the two is proportional to the value 1 / D, so the size of the neighborhood may be proportional to the value 1 / D. This utilizes the characteristic that almost no parallax occurs between distant subjects. When the depth correction unit 107 determines that the difference value dD is greater than the threshold value Th1, it executes the process of step S405, and when it determines that the difference value dD is smaller than the threshold value Th1, it executes the process of step S406. When the difference value dD is equal to the threshold value Th1, it is possible to arbitrarily set which step the depth correction unit 107 executes. Also, in this embodiment, when the distance value d of the surrounding pixel is greater than the threshold value with respect to the distance value D of the pixel of interest, the distance value d of the surrounding pixel is corrected, but the distance value d of the surrounding pixel may be corrected when it is greater than the threshold value (second predetermined value). This corresponds to correcting the distance value corresponding to the distance to the subject that is located farther than a certain distance.
[0030] In step S405, the depth modification unit 107 modifies the distance value d of the surrounding pixels. In this embodiment, the distance value d is modified to 0. This corresponds to deleting the distance value corresponding to the distance to a distant subject present in the nearby area. The reason why the distance value d may be set to 0 is that the distance value is appropriately filled in from the surrounding pixels by the process (interpolation process) of step S205. Note that the distance value d may be modified to a value close to the distance value D.
[0031] In step S406, the depth correction unit 107 determines whether the above process has been performed on all pixels as the pixel of interest. If the depth correction unit 107 determines that the above process has been performed on all pixels as the pixel of interest, it ends this flow, and if it determines that the above process has not been performed, it executes the process of step S403.
[0032] As described above, in this embodiment, by correcting the sparse depth map seen by the imaging unit 102, it is possible to obtain a dense depth map that reproduces the boundaries of the subject with high accuracy by interpolation processing.
[0033] Note that when the distance value D of the pixel of interest is equal to or greater than a predetermined value (equal to or greater than a third predetermined value), the occurrence of parallax is suppressed, so that the depth modification process may not be executed.
[0034] In addition, in this embodiment, an example of correcting the distance value of the neighboring region has been described, but the present invention is not limited to this. For example, when the difference D-d_min between the minimum value d_0 (which may be the median or average value) of the distance value of the neighboring region and the distance value D of the reference pixel is greater than a first predetermined threshold, the distance value D of the reference pixel may be corrected to the minimum value d_0.
[0035] 5 is an explanatory diagram of the effect of the configuration of this embodiment. An image 501 is an image captured by the imaging unit 102 of a subject 510. By performing the depth correction process, it is possible to obtain a highly accurate and high-density depth map 503 with less loss of the subject compared to a depth map 502 obtained when the depth correction process was not performed.
[0036] As described above, according to the configuration of this embodiment, it is possible to acquire highly accurate and high-density three-dimensional distance information. EXAMPLES
[0037] In this embodiment, the depth modification unit 107 determines whether to modify the depth map by also using information on the image acquired by the imaging unit 102. Note that in this embodiment, only configurations different from the first embodiment will be described, and descriptions of the same configurations as the first embodiment will be omitted.
[0038] FIG. 6 is a conceptual diagram showing the operation of the depth correction unit 107 of this embodiment, showing a state in which the region 320 is superimposed on the image 501 and the depth map 503. The depth correction unit 107 corrects the distance value in the region 320 to the distance value of the object 510. As a result of the interpolation process, as shown in FIG. 6, there is a possibility that the region having the distance value of the object 510 may be expanded outside the original object boundary. In particular, when the object boundary is complicated, the adverse effect becomes large. Therefore, in this embodiment, in the image 501, the pixel value of the corresponding pixel of interest and the pixel value of the surrounding pixels included in the neighboring region are compared to switch whether or not to perform correction. In the following description, it is assumed that the number of pixels in the image 501 and the depth map 503 are the same, but it may be different. In that case, the pixel value on the image 501 corresponding to the pixel on the depth map 503 is acquired.
[0039] FIG. 7 is a flowchart showing the depth correction process of this embodiment.
[0040] The processes of steps S701 to S703 and steps S705 to S707 are similar to the processes of steps S401 to S406 in FIG. 4, respectively, and therefore will not be described.
[0041] In step S704, the depth correction unit 107 judges whether the absolute difference |Ii| between the pixel value 1 of the pixel of interest in the image 501 acquired by the imaging unit 102 and the pixel value i of the peripheral pixel included in the neighboring region is greater than a threshold value (fourth predetermined value) Th2. If the depth correction unit 107 judges that the absolute difference |Ii| is greater than the threshold value Th2, it executes the process of step S705, and if it judges that the absolute difference |Ii| is smaller than the threshold value Th2, it executes the process of step S703. Note that when the absolute difference |Ii| is equal to the threshold value Th2, it is possible to arbitrarily set which step the depth correction unit 107 executes. In addition, when comparing pixel values, a value other than the absolute difference value may be used. Here, since the purpose of the process of this step is to identify the object boundary in the image 501, the magnitude of the difference is important, and the sign may be either positive or negative. On the other hand, in the process of step S705, since a relatively distant object is identified, the sign of the difference in the distance value is important.
[0042] In this embodiment, luminance is used as the pixel value, but RGB information may be used. In that case, it is possible to use the average value of the results of each RGB channel. Also, the method is not limited to the above as long as it is possible to compare the difference in pixel values.
[0043] In this embodiment, an example has been described in which it is determined whether or not to modify the depth map 503 based on the pixel values of the image 501, but the shape of the neighboring region may be changed. In that case, in step S704, pixels whose absolute difference value is greater than the threshold value Th2 may be excluded from the neighboring region.
[0044] As described above, according to the configuration of this embodiment, in addition to the effects of the first embodiment, it is possible to reproduce the subject boundary with high accuracy. [Other Examples] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-mentioned embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0045] The disclosure of this embodiment includes the following configurations and methods. (Configuration 1) a generating means for generating a distance image viewed from a second viewpoint using distance information to a subject viewed from a first viewpoint; and modifying means for determining whether or not to modify the distance value in accordance with the distance value of a pixel included in a neighborhood area of a reference pixel in the distance image. (Configuration 2) The processing device according to configuration 1, wherein the modification means modifies the distance value when a difference between the distance value and a reference distance value of the reference pixel is greater than a first predetermined value, and does not modify the distance value when the difference is smaller than the first predetermined value. (Configuration 3) 2. The processing device according to claim 1, wherein the correction means corrects the distance value when the distance value is greater than a second predetermined value, and does not correct the distance value when the distance value is smaller than the second predetermined value. (Configuration 4) 4. The processing device according to any one of configurations 1 to 3, wherein the distance information is acquired using at least one of a laser beam, an electromagnetic wave, and a sound wave. (Configuration 5) The distance information is composed of at least one of a point cloud, a voxel, a polygon, a mesh, an implicit function representation, a depth map, a disparity map, and a depth map; The processing device according to any one of configurations 1 to 4, wherein the distance image is any one of a depth map, a disparity map, and a depth map. (Configuration 6) the neighboring region is configured by a plurality of pixels including the reference pixel, 6. The processing device according to any one of configurations 1 to 5, wherein the modifying means modifies a distance value of at least one of the plurality of pixels. (Configuration 7) 7. The processing device according to any one of configurations 1 to 6, wherein the modification means does not execute a process of modifying the distance value when the reference distance value of the reference pixel is equal to or greater than a third predetermined value. (Configuration 8) 8. The processing device according to any one of configurations 1 to 7, further comprising setting means for setting the neighboring region. (Configuration 9) 9. The processing device according to configuration 8, wherein the setting means sets the size of the neighboring region in accordance with a reference distance value of the reference pixel. (Configuration 10) 10. The processing device according to any one of configurations 1 to 9, wherein the second viewpoint is a viewpoint of an imaging means that acquires a captured image. (Configuration 11) 11. The processing device according to configuration 10, wherein the correction means determines whether or not to perform processing for correcting the distance value depending on a pixel value of the captured image. (Configuration 12) The processing device according to configuration 11, characterized in that the correction means determines whether or not to execute the processing based on the magnitude of the difference between the pixel value of a pixel corresponding to the reference pixel of the captured image and the pixel value of a pixel corresponding to a pixel included in the neighboring region. (Configuration 13) The processing device described in configuration 12, characterized in that the correction means does not perform the processing when an absolute value of a difference between a pixel value of a pixel corresponding to the reference pixel of the captured image and a pixel value of a pixel corresponding to a pixel included in the neighboring area is greater than a fourth predetermined value, and performs the processing when the absolute value of the difference is smaller than the fourth predetermined value. (Configuration 14) The method further includes a setting unit for setting the neighboring region, 14. The processing device according to any one of configurations 10 to 13, wherein the setting means sets the neighboring area in accordance with pixel values of the captured image. (Configuration 15) The processing device according to any one of configurations 10 to 14, further comprising an interpolation means for interpolating distance values of pixels having missing distance values, using the distance image corrected by the correction means and a captured image. (Configuration 16) 16. The processing device according to configuration 15, wherein the interpolation means performs the interpolation using at least one of a filter, machine learning, super-resolution, and a conditional random field. (Method 1) converting distance information to a subject viewed from a first viewpoint into a distance image viewed from a second viewpoint; and correcting the distance value in accordance with distance values of pixels included in a neighborhood area of a reference pixel in the distance image. (Configuration 17) A program for causing a computer to execute the processing method according to Method 1.
[0046] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]
[0047] 100 Three-dimensional distance information processing device (processing device) 106 Projection unit (generation means) 107 Depth correction unit (correction means)
Claims
1. a generating means for generating a distance image viewed from a second viewpoint using distance information to a subject viewed from a first viewpoint; and modifying means for determining whether or not to modify the distance value in accordance with the distance value of a pixel included in a neighborhood area of a reference pixel in the distance image.
2. 2. The processing device according to claim 1, wherein the modification means modifies the distance value when a difference between the distance value and a reference distance value of the reference pixel is greater than a first predetermined value, and does not modify the distance value when the difference is smaller than the first predetermined value.
3. 2. The processing device according to claim 1, wherein said correcting means corrects said distance value when said distance value is greater than a second predetermined value, and does not correct said distance value when said distance value is smaller than said second predetermined value.
4. 4. The processing device according to claim 1, wherein the distance information is obtained using at least one of a laser beam, an electromagnetic wave, and a sound wave.
5. The distance information is composed of at least one of a point cloud, a voxel, a polygon, a mesh, an implicit function representation, a depth map, a disparity map, and a depth map; The processing device according to claim 1 , wherein the distance image is any one of a depth map, a disparity map, and a depth map.
6. the neighboring region is configured by a plurality of pixels including the reference pixel, 4. The processing device according to claim 1, wherein the modifying means modifies the distance value of at least one of the plurality of pixels.
7. 4. The processing device according to claim 1, wherein the modifying means does not execute the process of modifying the distance value when the reference distance value of the reference pixel is equal to or greater than a third predetermined value.
8. 4. The processing apparatus according to claim 1, further comprising a setting unit for setting the neighboring region.
9. 9. The processing device according to claim 8, wherein said setting means sets the size of said neighboring region in accordance with a reference distance value of said reference pixel.
10. 4. The processing device according to claim 1, wherein the second viewpoint is a viewpoint of an image capturing device that captures a captured image.
11. 11. The processing device according to claim 10, wherein the correction means determines whether or not to perform processing for correcting the distance value depending on a pixel value of the captured image.
12. The processing device according to claim 11, characterized in that the correction means determines whether or not to execute the processing based on the magnitude of the difference between the pixel value of a pixel corresponding to the reference pixel of the captured image and the pixel value of a pixel corresponding to a pixel included in the neighboring area.
13. The processing device according to claim 12, characterized in that the correction means does not perform the processing when an absolute value of a difference between a pixel value of a pixel corresponding to the reference pixel of the captured image and a pixel value of a pixel corresponding to a pixel included in the neighboring area is greater than a fourth predetermined value, and performs the processing when the absolute value of the difference is smaller than the fourth predetermined value.
14. The method further includes a setting unit for setting the neighboring region, The processing device according to claim 10 , wherein the setting means sets the neighboring area in accordance with pixel values of the captured image.
15. 11. The processing device according to claim 10, further comprising an interpolation unit that uses the distance image corrected by the correction unit and the captured image to interpolate distance values of pixels having missing distance values.
16. The processing device according to claim 15 , wherein the interpolation means performs the interpolation using at least one of a filter, machine learning, super-resolution, and a conditional random field.
17. converting distance information to a subject viewed from a first viewpoint into a distance image viewed from a second viewpoint; and correcting the distance value in accordance with distance values of pixels included in a neighborhood area of a reference pixel in the distance image.
18. A program causing a computer to execute the processing method according to claim 17.