Processing apparatus, processing method, and processing program
The processing apparatus reconstructs three-dimensional structures by identifying and estimating voxel values, interpolating non-visible regions, and assigning confidence levels, addressing incomplete reconstructions and outdoor data challenges.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2023-03-08
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods struggle to accurately reconstruct three-dimensional structures from images due to factors like camera position, device accuracy, and textureless regions, leading to incomplete reconstructions and difficulties in creating synthetic data for outdoor environments.
A processing apparatus and method that identifies and estimates voxel values for visible regions, interpolates values for non-visible regions, and assigns confidence levels to these voxels, allowing for the reconstruction of three-dimensional structures without relying on pre-labeled synthetic data.
Enables the reconstruction of three-dimensional structures, particularly in outdoor environments, with improved accuracy and reduced reliance on synthetic data, by complementing unrecovered portions and enabling tasks like semantic segmentation.
Smart Images

Figure 0007845570000001 
Figure 0007845570000002 
Figure 0007845570000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a processing device, a processing method, and a processing program.
Background Art
[0002] When restoring a three-dimensional region as a three-dimensional structure from a color image and / or a depth image obtained by imaging the three-dimensional region, portions that are not restored as the three-dimensional structure may occur due to factors such as the camera position, the accuracy of the device, and textureless regions. Under such circumstances, Patent Document 1, Non-Patent Document 1, and Non-Patent Document 2 disclose techniques for complementing portions that were not restored as the three-dimensional structure.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
[0005] The technology described in Patent Document 1 performs interpolation of areas that could not be restored as a three-dimensional structure by referring to the labels of adjacent distance measurement points. For this reason, it was necessary to assign a label to each distance measurement point according to the type of subject. Furthermore, the completion of three-dimensional structures described in Non-Patent Documents 2 and 3 is achieved by learning using complete three-dimensional structure information, such as synthetic data. Indoor three-dimensional structures can be accurately reconstructed by using appropriate synthetic data (e.g., the SUNCG dataset). However, since there are a wide variety of objects outdoors, and the terrain included as the background is also highly varied, it is difficult to create synthetic data for learning outdoor three-dimensional structures.
[0006] One aspect of the present invention has been made in view of the above-mentioned problems, and one example of its purpose is to provide a new technique for supplementing areas that were not restored when reconstructing a three-dimensional structure from an image of a three-dimensional region. [Means for solving the problem]
[0007] A processing apparatus according to one aspect of the present invention includes: identification means for identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging a three-dimensional region in a voxel space representing the three-dimensional region as a group of voxels not to be estimated; estimation and interpolation means for calculating the voxel values of estimable voxels, which are included in the group of voxels to be estimated, obtained by excluding the group of voxels not to be estimated from the voxel space, by referring to the one or more images, and for interpolating the voxel values of unestimable voxels, which are not estimated from the one or more images, by referring to the voxel values of the estimable voxels; and assigning means for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, wherein the assigning means assigns a confidence level to the unestimable voxels that is lower than the confidence level assigned to the estimable voxels.
[0008] A processing method relating to one aspect of the present invention includes: at least one processor identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging a three-dimensional region in a voxel space representing the three-dimensional region as a group of voxels to be not estimated; calculating the voxel values of estimable voxels, which are among the voxels included in the group of voxels to be estimated, obtained by excluding the group of voxels to be not estimated from the voxel space, by referring to the one or more images, and supplementing the voxel values of unestimable voxels, which are not among the voxels to be estimated from the one or more images, by referring to the voxel values of the estimable voxels; and assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, and assigning a lower confidence level to the unestimable voxels than the confidence level assigned to the estimable voxels.
[0009] A program for causing a computer to function as a processing unit according to one aspect of the present invention includes: an identification means for identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging a three-dimensional region in a voxel space representing a three-dimensional region as a group of voxels not to be estimated; an estimation and interpolation means for calculating the voxel values of estimable voxels, which are included in the group of voxels to be estimated, obtained by excluding the group of voxels not to be estimated from the voxel space, by referring to the one or more images, and for interpolating the voxel values of unestimable voxels, which are not estimated from the one or more images, by referring to the voxel values of the estimable voxels; and an assignment means for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, wherein the assignment means assigns a confidence level to the unestimable voxels that is lower than the confidence level assigned to the estimable voxels. [Effects of the Invention]
[0010] According to one aspect of the present invention, a novel technique can be provided for supplementing areas that were not restored when reconstructing a three-dimensional structure from an image of a three-dimensional region. [Brief explanation of the drawing]
[0011] [Figure 1] This is a block diagram showing the configuration of the processing apparatus 1 according to exemplary embodiment 1 of the present invention. [Figure 2] This is a flowchart showing the flow of processing method S1 according to exemplary embodiment 1 of the present invention. [Figure 3] This figure shows an example of processing performed by the application unit 13 of the processing apparatus 1 according to exemplary embodiment 1 of the present invention. [Figure 4] This figure shows an example of processing performed by the application unit 13 of the processing apparatus 1 according to exemplary embodiment 1 of the present invention. [Figure 5] This is a block diagram showing the configuration of the processing apparatus 1A according to exemplary embodiment 2 of the present invention. [Figure 6] It is a flowchart showing the flow of the processing method S1A according to the exemplary embodiment 2 of the present invention. [Figure 7] It is a block diagram showing the configuration of the processing device 1B according to the exemplary embodiment 3 of the present invention. [Figure 8] It is a flowchart showing the flow of the processing method S1B according to the exemplary embodiment 3 of the present invention. [Figure 9] It is a block diagram showing the configuration of the processing program according to the exemplary embodiment 4 of the present invention.
Mode for Carrying Out the Invention
[0012] 〔Exemplary Embodiment 1〕 The first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for the exemplary embodiments described later.
[0013] (Outline of the processing device 1) When creating a voxel group representing the three-dimensional region from a color image and / or a depth image obtained by imaging the three-dimensional region, areas that are blind spots due to obstacles and areas without texture, etc., are not restored as voxels and may become voxels without feature amounts. For a voxel group including voxels without feature amounts, even when tasks such as three-dimensional semantic segmentation and object detection tasks are executed, there is a risk that the accuracy will be low. The processing device 1 according to this exemplary embodiment is a device for improving the accuracy of the tasks to be executed by complementing the unrecovered portions of the three-dimensional structure restored for executing the tasks.
[0014] (Configuration of the processing device 1) The configuration of the processing device 1 according to this exemplary embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the processing device 1.
[0015] As shown in FIG. 1, the processing device 1 includes a specifying unit 11, an estimating / complementing unit 12, and an assigning unit 13. In this exemplary embodiment, the specifying unit 11, the estimating / complementing unit 12, and the assigning unit 13 are configurations that respectively implement specifying means, estimating / complementing means, and assigning means.
[0016] The specifying unit 11 is a configuration for specifying a non-estimation target voxel group in a voxel space representing a three-dimensional region. As the non-estimation target voxel group, a voxel group that is not included in any viewing angle of one or more images obtained by imaging the three-dimensional region is specified. The specifying unit 11 takes the viewing angle information of the image as an input.
[0017] The estimating / complementing unit 12 is a configuration for calculating the voxel values of estimable voxels among the voxels included in the estimation target voxel group and complementing the voxel values of non-estimable voxels. The estimation target voxel group is a voxel group obtained by excluding the non-estimation target voxel group from the voxel space. An estimable voxel refers to a voxel among the voxels included in the estimation target voxel group for which the voxel value can be estimated from one or more images. The voxel values of estimable voxels are calculated by referring to one or more images. A non-estimable voxel refers to a voxel for which the voxel value cannot be estimated from one or more images. The voxel values of non-estimable voxels are complemented by referring to the voxel values of estimable voxels. The estimating / complementing unit 12 takes one or more images as an input.
[0018] The assigning unit 13 is a configuration for assigning a reliability to the voxel value of each voxel included in the estimation target voxel group. As the reliability assigned to a non-estimable voxel, a lower reliability than the reliability assigned to an estimable voxel is assigned.
[0019] (Flow of the processing method S1) The flow of the processing method S1 according to this exemplary embodiment will be described with reference to FIG. 2. FIG. 2 is a flowchart showing the flow of the processing method S1.
[0020] As shown in Figure 2, processing method S1 includes a specific processing S11, an estimation / complementary processing S12, and an assignment processing S13.
[0021] The identification process S11 is a process for identifying a group of non-estimated voxels in the voxel space representing a three-dimensional region. The group of non-estimated voxels identified is a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging the three-dimensional region.
[0022] The estimation and imputation process S12 calculates the voxel values of estimable voxels and imputes the voxel values of unestimable voxels among the voxels included in the group of voxels to be estimated, which is obtained by excluding the group of voxels not to be estimated from the voxel space. An estimable voxel is a voxel included in the group of voxels to be estimated, which is obtained by excluding the group of voxels not to be estimated from the voxel space, and whose voxel value can be estimated from one or more images. The voxel value of an estimable voxel is calculated by referring to one or more images. An unestimable voxel is a voxel whose voxel value cannot be estimated from one or more images. The voxel value of an unestimable voxel is imputed by referring to the voxel value of an estimable voxel.
[0023] The assignment process S13 is a process for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated. The confidence level assigned to voxels that cannot be estimated is lower than the confidence level assigned to voxels that can be estimated.
[0024] (Effects of Processing Unit 1) As described above, the apparatus 1 according to this exemplary embodiment employs a configuration that complements the voxel values of unestimable voxels in the group of voxels to be estimated. According to the processing apparatus 1 according to this exemplary embodiment, the voxel values of unestimable voxels can be complemented without labeling according to the type of subject or learning based on pre-created synthetic data before complementing the voxel values. Therefore, it is possible to provide a new technique for complementing areas that could not be restored when reconstructing a three-dimensional structure from an image of a three-dimensional region. In particular, this disclosure can also restore three-dimensional structures for which labeling according to the type of subject and creating synthetic data with complete information on the three-dimensional structure in advance are difficult. For this reason, it can be applied to processing outdoor three-dimensional regions and three-dimensional regions with many unknown subjects, for which the application of conventional techniques was difficult. (Example image) The image can be a color image and / or a depth image. Alternatively, for example, it can be an image in which color information (RGB values) and depth information are assigned to each pixel.
[0025] (Specific example of section 11) The identification unit 11 can identify a group of non-estimated target voxels by referring to the camera position and orientation of each camera that has acquired one or more images, and the field of view of the camera.
[0026] (Specific example of estimation / complementary section 12) The voxel value of each voxel is a value that includes the three-dimensional coordinates of the voxel and the feature quantities of the voxel determined based on the image. The feature quantities of each voxel determined based on the image can be, for example, RGB values and normal values. Examples of unpredictable voxels include occluded voxels and missing voxels. Occluded voxels are voxels corresponding to points that were not included as subjects in any of the one or more images. Missing voxels are voxels corresponding to points for which no depth value was given in the image, or points for which the subject was included in multiple images but the depth values were inconsistent between the images.
[0027] Occlusion voxels refer to voxels that correspond to points that were not captured by any camera, such as points that were not restored due to the camera's movement and position when acquiring each image, or points that were not restored due to blind spots caused by objects in the foreground. Occlusion voxels may also include voxels that correspond to points that were included as subjects in one of the multiple color images. Points that are included in only one color image cannot have their depth value estimated, so their voxel values are not calculated as estimable voxels, but they can be interpolated as occlusion voxels included in the group of voxels to be estimated.
[0028] A missing voxel refers to a voxel that corresponds to a point in an image for which no depth value was assigned, or a point that is included as a subject in multiple images but whose depth value is inconsistent across those images. Such points in a three-dimensional region are, for example, points that are included as subjects in multiple two-dimensional images but for which no texture is assigned, even through edge processing.
[0029] For example, the voxel value of the nearest estimable voxel can be used to interpolate the voxel value of an unestimable voxel.
[0030] (Specific example of the granting unit 13) The assignment unit 13 can, for example, assign a confidence level to a shielding voxel, such that the confidence level decreases as the distance from the shielding voxel to the nearest estimable voxel increases.
[0031] An example of assigning confidence to occluded voxels will be explained with reference to Figure 3. Figure 3 shows a cross-sectional view of the group of voxels to be estimated, generated from multiple two-dimensional images of a step as a three-dimensional region. Each square represents one voxel. In the upper part of Figure 3, voxels shown in black are estimated voxels, and voxels shown with diagonal lines are occluded voxels. Voxels shown in white are voxels that do not correspond to the subject.
[0032] Referring to the lower part of Figure 3, an example of the confidence level assignment process in the assignment unit 13 is shown. The assignment unit 13 assigns a confidence level of 1.0 to an estimable voxel. The assignment unit 13 assigns a confidence level of 0.9 to an obscured voxel facing an estimable voxel, 0.8 to an obscured voxel one square away, and 0.7 to an obscured voxel two squares away, and does not assign a confidence level to voxels further away. In this way, the assignment unit 13 assigns a confidence level to obscured voxels that decreases as the distance to the nearest estimable voxel increases.
[0033] The assignment unit 13 can assign a confidence level to each missing voxel corresponding to a point where the depth values were inconsistent across multiple images. This confidence level decreases as the difference between the depth value set for the pixel corresponding to the missing voxel among the pixels constituting the image and the depth value that should be set for that pixel, estimated from the position of the missing voxel in voxel space, increases. The depth value of each pixel in the image can be, for example, a depth value determined based on multiple color images, a depth value acquired along with the image by an RGB-D camera that images a three-dimensional region, a depth value acquired by a Lidar used in the same position and orientation as the color image, and a depth value of a depth image acquired by a Lidar. The depth value that should be set for the pixel, estimated from the position of each missing voxel in voxel space, can be determined by a projection method.
[0034] An example of assigning confidence to missing voxels will be explained with reference to Figure 4. Each figure in Figure 4 shows the projection plane when a group of estimated voxels generated from multiple two-dimensional images captured as a three-dimensional region is projected onto one of the multiple two-dimensional images. Here, the estimated voxels include missing voxels.
[0035] In the upper panel of Figure 4, gray voxels represent pixels corresponding to estimated voxels, while diagonal lines represent pixels corresponding to missing voxels. Among the missing voxels, those with thick diagonal lines represent voxels corresponding to subject X. White voxels are voxels that do not correspond to the subject.
[0036] The middle section of Figure 4 schematically shows, using numbers from 1 to 6, the difference between the depth value of each pixel in the projected image (projected onto a two-dimensional image) and the depth value of that pixel in the corresponding depth image for each missing voxel. A larger value indicates a larger difference in depth values.
[0037] The lower panel of Figure 4 shows the results of assigning confidence levels to each missing voxel based on the magnitude of the difference. Voxels with a difference of 1 were assigned a confidence level of 0.9, voxels with a difference of 2 were assigned a confidence level of 0.8, and voxels with a difference of 3 were assigned a confidence level of 0.7. Voxels with a difference of 4 or more were not assigned a confidence level.
[0038] (Modified version of processing device 1) The processing unit 1 described above is configured such that the estimation / completion unit 12 and the assignment unit 13 perform sequential and independent processing. However, it is not limited to this configuration, and the estimation / completion unit 12 and the assignment unit 13 can perform processing in parallel, such as in processing using a TSDF (truncated signed distance function). For example, in the process of calculating the voxel values of estimable voxels among the voxels included in the group of voxels to be estimated, which is obtained by excluding the group of voxels not to be estimated from the voxel space, and interpolating the voxel values of unestimable voxels, the estimation / completion unit 12 and the assignment unit 13 can assign a confidence level to the unestimable voxel by referring to a score based on the distance from nearby estimable voxels.
[0039] [Exemplary Embodiment 2] A second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Components having the same function as those described in Exemplary Embodiment 1 will be denoted by the same reference numerals, and their descriptions will be omitted as appropriate.
[0040] As shown in Figure 5, the processing unit 1A comprises a identification unit 11, an estimation / completion unit 12, an assignment unit 13, a calculation unit 14, and a segmentation unit 15. In this exemplary embodiment, the identification unit 11, the estimation / completion unit 12, the assignment unit 13, the calculation unit 14, and the segmentation unit 15 are configured to realize the identification means, estimation / completion means, assignment means, calculation means, and segmentation means, respectively.
[0041] (Configuration of Processing Unit 1A) The identification unit 11, estimation / completion unit 12, and assignment unit 13 of the processing device 1A are components that have the same functions as the identification unit 11, estimation / completion unit 12, and assignment unit 13 of the processing device 1.
[0042] The calculation unit 14 is configured to calculate the feature quantities of each voxel included in the group of voxels to be estimated. The feature quantities of each voxel are calculated by referring to the voxel value of the voxel and the confidence level assigned to the voxel.
[0043] The segmentation unit 15 is configured to perform semantic segmentation of the target voxel group and determine the labels of the voxels. Semantic segmentation is performed by referring to the feature quantities of each voxel included in the target voxel group.
[0044] (Processing method S1A flow) The flow of processing method S1A according to this exemplary embodiment will be explained with reference to Figure 6. Figure 6 is a flowchart showing the flow of processing method S1A.
[0045] As shown in Figure 6, the processing method S1A includes a specific processing S11, an estimation / completion processing S12, an assignment processing S13, a calculation processing S14, and a segmentation processing S15.
[0046] Calculation process S14 is a process for calculating the feature quantities of each voxel included in the group of voxels to be estimated. The feature quantities of each voxel are calculated by referring to the voxel value of the voxel and the confidence level assigned to the voxel.
[0047] Segmentation process S15 is a process for performing semantic segmentation on the group of voxels to be estimated and determining the labels of those voxels. Semantic segmentation is performed by referring to the feature quantities of each voxel included in the group of voxels to be estimated.
[0048] (Effects of processing device 1A) Furthermore, in the processing apparatus 1A of this exemplary embodiment, a configuration is employed to perform semantic segmentation on the group of voxels to be estimated, which are obtained by supplementing the voxels that cannot be estimated. Therefore, in addition to the effects achieved by the processing apparatus 1 of the exemplary embodiment 1, the processing apparatus 1A of this exemplary embodiment can be used to perform highly accurate semantic segmentation.
[0049] (Modified version of segmentation unit 15) The processing unit 1A may be configured to perform tasks other than semantic segmentation, either in place of or in addition to the segmentation unit 15. For example, it may be equipped with a task execution unit for performing an object detection task, such as detecting an object by enclosing a three-dimensional region around it with a bounding box. Alternatively, it may be equipped with a task execution unit for performing a completion task.
[0050] [Exemplary Embodiment 3] A third exemplary embodiment of the present invention will be described in detail with reference to the drawings. Components having the same function as those described in Exemplary Embodiment 1 and Exemplary Embodiment 2 will be denoted by the same reference numerals, and their descriptions will not be repeated.
[0051] As shown in Figure 7, the processing unit 1B comprises a identification unit 11, an estimation / completion unit 12, an assignment unit 13, a calculation unit 14, a segmentation unit 15, and a learning unit 16. In this exemplary embodiment, the identification unit 11, estimation / completion unit 12, assignment unit 13, calculation unit 14, segmentation unit 15, and learning unit 16 are configured to implement the identification means, estimation / completion means, assignment means, calculation means, segmentation means, and learning means, respectively.
[0052] (Configuration of Processing Unit 1B) The identification unit 11, estimation / completion unit 12, assignment unit 13, calculation unit 14, and segmentation unit 15 of the processing device 1B are components that have the same functions as the identification unit 11, estimation / completion unit 12, and assignment unit 13 of the processing device 1, and the calculation unit 14 and segmentation unit 15 of the processing device 1A.
[0053] The learning unit 16 is configured to generate a model using machine learning, which the calculation unit 14 uses to calculate features. The learning unit 16 uses training data for machine learning that includes, as features of unestimable voxels, estimated features estimated from the features of estimable voxels surrounding the unestimable voxel, as the ground truth labels. (Processing method S1B flow) The flow of processing method S1A according to this exemplary embodiment will be explained with reference to Figure 8. Figure 8 is a flowchart showing the flow of processing method S1B.
[0054] As shown in Figure 8, the processing method S1B includes a specific processing S11, an estimation / completion processing S12, an assignment processing S13, a calculation processing S14, a segmentation processing S15, and a learning processing S16.
[0055] The learning process S16 is configured to generate a model used to calculate features using machine learning. In the learning process S16, training data is used for machine learning that includes estimated features, which are estimated from the features of the estimable voxels surrounding the unestimable voxel, as the ground truth labels for the unestimable voxel.
[0056] (Effects of processing device 1B) As described above, the processing apparatus 1B according to this exemplary embodiment employs a configuration in which the model used to calculate features is generated by machine learning. Therefore, the processing apparatus 1B according to this exemplary embodiment has the effect of being able to obtain a model that can perform highly accurate semantic segmentation.
[0057] (Specific example from Learning Section 16) The learning unit 16 uses data to which the correct labels have been assigned as features of unestimable voxels as training data. The training data is created from a group of voxels to be estimated, generated by the estimation / interpolation unit 12. For the training data of the estimable voxels included in the group of voxels to be estimated, the correct label features are assigned by referring to the voxel values. Furthermore, for the training data of the unestimable voxels included in the group of voxels to be estimated, the values can be a copy of the correct label of the nearest voxel to the voxel, or a value in which the features of each neighboring voxel are reflected in the features of the unestimable voxel according to the distance between multiple neighboring voxels.
[0058] The learning unit 16 can update the model used by the calculation unit 14 to calculate features based on a first deviation, which is the degree of difference between the generated training data and the feature quantities of each voxel calculated by the calculation unit 14. The calculation of the first deviation uses the training data and the feature quantities of each voxel calculated based on the estimable voxels and the non-estimable voxels, respectively. For example, the learning unit 16 can update the model used by the calculation unit 14 to calculate features so that the sum of the first deviations is minimized.
[0059] The learning unit 16 can further perform learning using training data of images in which each pixel in each of the one or more images used to generate the estimated voxel group has been assigned a correct label. For example, the learning unit 16 can update the model used by the calculation unit 14 to calculate features based on a second deviation, which is the degree of difference between the two-dimensional image training data and the voxel features calculated by the calculation unit 14 corresponding to the pixels. For example, the learning unit 16 can calculate the sum of the second deviations and use this to update the model used by the calculation unit 14 to calculate features.
[0060] Here, for each of the one or more images, the position of the two-dimensional image can be determined within the estimated voxel group by referring to the estimated camera position and orientation at the time the two-dimensional image was captured. This allows the voxel within the estimated voxel group corresponding to each pixel in the two-dimensional image to be determined.
[0061] (Modified version of processing unit 1B) The calculation unit 14 can further calculate the feature quantities of each pixel in each of the two-dimensional images for each of the one or more images.
[0062] The segmentation unit 15 may be configured to perform semantic segmentation based on the feature quantities of each voxel within the estimated voxel group and the feature quantities of each two-dimensional image.
[0063] The learning unit 16 can perform learning using training data of one or more images, in which each pixel in one or more images is assigned a correct label. The learning unit 16 can update the model used by the calculation unit 14 to calculate features based on a third deviation, which is the degree of difference between the training data of one or more images and the feature quantities of each pixel in one or more images calculated by the calculation unit 14. For example, the learning unit 16 can calculate the sum of the second deviations and use this to update the model used by the calculation unit 14 to calculate features. With this configuration, semantic segmentation can be performed by referencing both the feature quantities of each voxel and the feature quantities of each pixel, thereby achieving more robust semantic segmentation.
[0064] (A modified version of Learning Section 16) When the processing unit 1B includes a task execution unit instead of a segmentation unit 15, the learning unit 16 can learn using training data created according to the tasks executed by the task execution unit.
[0065] [Examples of implementation using software] Some or all of the functions of processing units 1, 1A, and 1B may be implemented by hardware such as integrated circuits (IC chips) or by software.
[0066] In the latter case, the processing units 1, 1A, and 1B are implemented, for example, by a computer that executes instructions for a program, which is software that implements each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 9. Computer C comprises at least one processor C1 and at least one memory C2. The memory C2 stores a program P that causes computer C to operate as processing units 1, 1A, and 1B. In computer C, the processor C1 reads the program P from the memory C2 and executes it, thereby implementing each function of processing units 1, 1A, and 1B.
[0067] For processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof can be used. For memory C2, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof can be used.
[0068] Computer C may also be equipped with RAM (Random Access Memory) for loading program P at runtime and for temporarily storing various data. Furthermore, computer C may be equipped with communication interfaces for sending and receiving data with other devices. Additionally, computer C may be equipped with input / output interfaces for connecting input / output devices such as keyboards, mice, displays, and printers.
[0069] Furthermore, program P can be recorded on a non-temporary, tangible recording medium M that is readable by computer C. Such a recording medium M could be, for example, tape, disk, card, semiconductor memory, or programmable logic circuitry. Computer C can acquire program P via such a recording medium M. Program P can also be transmitted via a transmission medium. Such a transmission medium could be, for example, a communication network or broadcast waves. Computer C can also acquire program P via such a transmission medium.
[0070] [Additional Note 1] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the embodiments described above are also included in the technical scope of the present invention.
[0071] [Additional Note 2] Some or all of the embodiments described above may also be described as follows. However, the present invention is not limited to the embodiments described below.
[0072] (Note 1) A processing device comprising: identification means for identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging a three-dimensional region in a voxel space representing the three-dimensional region as a group of voxels to be not estimated; estimation and interpolation means for calculating the voxel values of estimable voxels, which are included in the group of voxels to be estimated, obtained by excluding the group of voxels to be not estimated from the one or more images, by referring to the one or more images, and interpolating the voxel values of unestimable voxels, which are not estimated from the one or more images, by referring to the voxel values of the estimable voxels; and assigning means for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, wherein the assigning means assigns a confidence level to unestimable voxels that is lower than the confidence level assigned to estimable voxels.
[0073] According to the above configuration, it is possible to provide a new technique for supplementing areas that were not reconstructed when reconstructing a three-dimensional structure from an image captured in a three-dimensional region.
[0074] (Note 2) The apparatus according to Appendix 1, wherein the unestimable voxels include occluded voxels corresponding to points not included as subjects in any of the one or more images, and missing voxels corresponding to points for which no depth value was assigned in any of the one or more images, or points for which, even if included as subjects in the multiple images, the depth values were inconsistent between the images.
[0075] According to the above configuration, unpredictable voxels can be identified in more detail based on their causal factors.
[0076] (Note 3) The processing apparatus according to Appendix 2, wherein the assigning means assigns a confidence level to each shielding voxel, the confidence level decreases as the distance from the shielding voxel to the nearest estimable voxel increases.
[0077] With the above configuration, pseudo-labels can be assigned to occluding voxels.
[0078] (Note 4) The processing apparatus according to Appendix 2 or 3, wherein the assigning means assigns a confidence level to each missing voxel corresponding to a point where the depth values were inconsistent between the images, the confidence level decreases as the difference between the depth value set for the pixel corresponding to the missing voxel among the pixels constituting the image and the depth value that should be set for the pixel estimated from the position of the missing voxel in the voxel space increases.
[0079] With the above configuration, it is possible to assign pseudo-labels to missing voxels.
[0080] (Note 5) A calculation means for calculating the feature quantity of each voxel included in the aforementioned group of voxels to be estimated by referring to the voxel value of the voxel and the confidence level assigned to the voxel, The processing apparatus according to any one of the appendices 1 to 4, further comprising: a segmentation means for performing semantic segmentation of the estimated target voxel group by referring to the feature quantities of each voxel included in the estimated target voxel group, and determining the label of the voxel.
[0081] The above configuration allows for highly accurate semantic segmentation.
[0082] (Note 6) The processing apparatus according to Appendix 4, further comprising a learning means for generating a model used by the calculation means to calculate the features, wherein the learning means uses training data in the machine learning that includes, as features of the unestimable voxels, estimated features estimated from the features of the estimable voxels surrounding the unestimable voxels as ground truth labels.
[0083] With the above configuration, it is possible to obtain a model that can perform highly accurate semantic segmentation.
[0084] (Note 7) A processing method comprising: at least one processor identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging a three-dimensional region in a voxel space representing the three-dimensional region as a group of voxels to be not estimated; calculating the voxel values of estimable voxels, which are among the voxels included in the group of voxels to be estimated, obtained by excluding the group of voxels to be not estimated from the voxel space, by referring to the one or more images, and supplementing the voxel values of unestimable voxels, which are not among the voxels that can be estimated from the one or more images, by referring to the voxel values of the estimable voxels; and assigning a confidence level indicating the confidence level of the voxel value of each voxel included in the group of voxels to be estimated, and assigning a confidence level lower to the unestimable voxels than the confidence level assigned to the estimable voxels.
[0085] According to the above configuration, it is possible to provide a new technique for supplementing areas that were not reconstructed when reconstructing a three-dimensional structure from an image captured in a three-dimensional region.
[0086] (Note 8) A program for causing a computer to function as a processing unit, wherein at least one processor of the computer is configured to function as: identification means for identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging a three-dimensional region in a voxel space representing the three-dimensional region as a group of voxels not to be estimated; estimation and interpolation means for calculating the voxel values of estimable voxels, which are included in the group of voxels to be estimated, obtained by excluding the group of voxels not to be estimated from the voxel space, by referring to the one or more images, and for interpolating the voxel values of unestimable voxels, which are not estimable from the one or more images, by referring to the voxel values of the estimable voxels; and assigning means for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, wherein the assigning means assigns a confidence level to the unestimable voxels that is lower than the confidence level assigned to the estimable voxels.
[0087] According to the above configuration, it is possible to provide a new technique for supplementing areas that were not reconstructed when reconstructing a three-dimensional structure from an image captured in a three-dimensional region.
[0088] [Additional Note 3] Some or all of the embodiments described above can also be expressed as follows:
[0089] A processing device comprising at least one processor, the processor performing the following steps: identification process for identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging a three-dimensional region in a voxel space representing the three-dimensional region as a group of voxels to be not estimated; estimation and interpolation process for calculating the voxel values of estimable voxels, which are included in the group of voxels to be estimated, obtained by excluding the group of voxels to be not estimated from the voxel space, by referring to the one or more images, and interpolating the voxel values of unestimable voxels, which are not estimated from the one or more images, by referring to the voxel values of the estimable voxels; and assigning means for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, wherein the assigning means performs an assignment process for assigning a confidence level to unestimable voxels that is lower than the confidence level assigned to estimable voxels.
[0090] Furthermore, this processing unit may also be equipped with memory, and this memory may store a program that causes the processor to execute the specific processing, the estimation / completion processing, and the assignment processing. This program may also be recorded on a computer-readable, non-temporary, tangible recording medium. [Explanation of symbols]
[0091] 1, 1A, 1B ... Processing Unit 11...Specific section 12 ···Estimated / Supplementary Section 13 ··· Assignment Section 14 ···Calculation Section 15 ···Segmentation Department 16 ···Learning Department
Claims
1. In a voxel space representing a three-dimensional region, a means for identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging the three-dimensional region is identified as a group of non-estimated voxels. An estimation and interpolation means that calculates the voxel values of estimable voxels, which are included in the group of voxels to be estimated, obtained by excluding the group of non-estimated voxels from the voxel space, by referring to the one or more images, and interpolates the voxel values of unestimateable voxels, which are not estimable from the one or more images, by referring to the voxel values of the estimable voxels, A processing device comprising: an assigning means for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, wherein the assigning means assigns a confidence level to an unestimable voxel that is lower than the confidence level assigned to an estimable voxel.
2. The apparatus according to claim 1, wherein the unestimable voxels include occluded voxels corresponding to points that were not included as subjects in any of the one or more images, and missing voxels corresponding to points for which no depth value was assigned in any of the one or more images, or points for which the depth values were inconsistent between the multiple images even if the points were included as subjects in the multiple images.
3. The processing apparatus according to claim 2, wherein the assigning means assigns a confidence level to each shielding voxel, the confidence level decreases as the distance from the shielding voxel to the nearest estimable voxel increases.
4. The processing apparatus according to claim 2 or 3, wherein the assigning means assigns a confidence level to each missing voxel corresponding to a point where the depth values were inconsistent among the plurality of images, the confidence level decreases as the difference between the depth value set for the pixel corresponding to the missing voxel among the pixels constituting the image and the depth value that should be set for the pixel estimated from the position of the missing voxel in the voxel space increases.
5. A calculation means for calculating the feature quantity of each voxel included in the aforementioned group of voxels to be estimated by referring to the voxel value of the voxel and the confidence level assigned to the voxel, The apparatus according to claim 1 or 2, further comprising: a segmentation means for performing semantic segmentation of the estimated target voxel group by referring to the feature quantities of each voxel included in the estimated target voxel group, and determining the label of the voxel.
6. The calculation means further comprises a learning means for generating a model used by machine learning to calculate the feature quantities, The processing apparatus according to claim 5, wherein the learning means uses training data in the machine learning process that includes, as features of the unestimable voxels, estimated features estimated from the features of the estimable voxels surrounding the unestimable voxels as the correct labels.
7. At least one processor identifies a group of voxels in a voxel space representing a three-dimensional region that are not included in any of the fields of view of one or more images obtained by imaging the three-dimensional region as a group of voxels to be not estimated, Among the voxels included in the group of voxels to be estimated, obtained by excluding the group of non-estimated voxels from the voxel space, the voxel values of the estimable voxels, for which the voxel values can be estimated from the one or more images, are calculated by referring to the one or more images, and the voxel values of the unestimateable voxels, for which the voxel values cannot be estimated from the one or more images, are supplemented by referring to the voxel values of the estimable voxels. A processing method that includes assigning a confidence level indicating the confidence level of the voxel value of each voxel included in the group of voxels to be estimated, and assigning a confidence level lower to the voxels that cannot be estimated than the confidence level assigned to the voxels that can be estimated.
8. A program for causing a computer to function as a processing unit, wherein at least one processor provided by the computer, In a voxel space representing a three-dimensional region, a means for identifying a group of voxels that are not included in any of the fields of view of one or more images obtained by imaging the three-dimensional region is identified as a group of non-estimated voxels. An estimation and interpolation means that calculates the voxel values of estimable voxels, which are included in the group of voxels to be estimated, obtained by excluding the group of non-estimated voxels from the voxel space, by referring to the one or more images, and interpolates the voxel values of unestimateable voxels, which are not estimable from the one or more images, by referring to the voxel values of the estimable voxels, A processing program that functions as an assigning means for assigning a confidence level to the voxel value of each voxel included in the group of voxels to be estimated, wherein the assigning means assigns a lower confidence level to the voxels that cannot be estimated than the confidence level assigned to the voxels that can be estimated.
Citation Information
Patent Citations
Method and device for preparing parallactic picture
JP1995254074A
Silhouette extraction apparatus, and method and program
JP2018205788A
Signal processor, signal processing method, program, and moving object
JP2019028861A
Computer program, image processing device, image processing method and voxel data
JP2019114035A
Image processing apparatus, image processing system, image processing method and program
JP2021018570A