Image processing apparatus, image processing method, and computer program
The image processing device addresses the issue of unnatural rotation of three-dimensional images by detecting and correcting errors and missing parts, ensuring natural appearance through 3D data correction techniques.
Patent Information
- Application Number
- JP2024106475
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2026-01-20
AI Technical Summary
Existing image processing devices generate three-dimensional images that appear unnatural when rotated, with issues such as stretched or missing facial features.
An image processing device that includes an information acquisition unit, a 3D data generation processing unit, an error region detection mechanism, and an image adjustment unit to correct errors and missing parts in 3D data, using techniques like upsampling, image suppression, and interpolation.
Generates three-dimensional images that appear natural and comfortable to view even when rotated, by detecting and correcting errors and missing parts using machine learning and image processing techniques.
Smart Images

Figure 2026009418000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, a computer program, and the like. [Background technology]
[0002] For example, Patent Document 1 describes an image processing device that can generate a three-dimensional image by simultaneously acquiring distance information when capturing a single still image and processing the image based on the distance information. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2024-8596 Summary of the Invention [Problem to be solved by the invention]
[0004] However, with the configuration of Patent Document 1, for example, when a face is photographed and a three-dimensional image generated by image processing is rotated, when the face is viewed from the side, the sides of the face (ears, etc.) may be stretched, or, for example, parts of the sides of the neck (cheeks, neck, etc.) may be missing from the three-dimensional image.
[0005] An object of the present invention is to provide an image processing device that can generate an image that does not look unnatural when a stereoscopic image generated by image processing is rotated. [Means for solving the problem]
[0006] An image processing device according to an embodiment of the present invention includes: an information acquisition unit that acquires an image of a subject and distance information; a 3D data generation processing unit that generates 3D data of the subject based on the image and the distance information; an error region detection means for detecting an error region of the 3D data; and an image adjustment unit for adjusting the image information of the error area. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide an image processing device and the like that can generate an image that does not look unnatural when a three-dimensional image generated by image processing is rotated. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a functional block diagram showing an example of the configuration of an imaging device 100 according to a first embodiment of the present invention. [Figure 2] 4 is a flowchart showing an example of image processing according to the first embodiment. [Figure 3] 10 is a flowchart illustrating a detailed example of the 3D data generation process in step S21. [Figure 4] (A) is a diagram showing an example of a three-dimensional image obtained in step S20, (B) is a diagram showing an example of a distance map, (C) is a diagram showing an example of a mesh image generated based on point cloud data, and (D) is a diagram showing an example of a texture image of a three-dimensional shape generated based on the mesh image of Figure 4(C). [Figure 5] 10 is a flowchart showing a detailed example of the adjustment process in step S22. [Figure 6] 6A is a flowchart showing a detailed example of the processing for detecting errors in 3D data in step S50, and FIG. 6B is a diagram showing an example of semantic labeling in the flowchart of FIG. 6A applied to a three-dimensional image. [Figure 7] 7A is a flowchart showing another example of the process of detecting an error in 3D data in step S50, and FIG. 7B is a diagram showing an example of a distance map in the process flow of FIG. 7A. [Figure 8]8A is a flowchart showing yet another example of the process of detecting an error in 3D data in step S50, and FIG. 8B is a diagram showing an example of a large polygon area of a stereoscopic image in the process flow of FIG. 8A. [Figure 9] 10 is a flowchart showing an example of image adjustment processing in step S51. [Figure 10] 10 is a flowchart showing another example of the image adjustment process in step S51. [Figure 11] 10 is a flowchart showing an example of image processing according to the second embodiment. [Figure 12] 11A is a flowchart showing an example of the loss detection process in step S1103, FIG. 11B is a diagram showing an example of a 3D image when there is no loss, and FIG. 11C is a diagram showing an example of a 3D image when there is a loss. [Figure 13] 10A is a flowchart showing another example of the missing part detection process in step S1103, FIG. 10B is a diagram showing an example of a three-dimensional image when a missing part occurs, and FIG. 10C is a diagram showing an example of the shape of the three-dimensional image of the model. [Figure 14] 11A is a flowchart showing yet another example of the loss portion detection process in step S1103, and FIG. 11B is a diagram showing an example of the edge of a 3D image when a loss occurs. [Figure 15] 11 is a flowchart showing an example of image interpolation processing in step S1104. [Figure 16] 11 is a flowchart showing another example of the image interpolation process in step S1104. [Figure 17] 11 is a flowchart showing yet another example of the image interpolation process in step S1104. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. In each drawing, the same members or elements are designated by the same reference numerals, and duplicate descriptions will be omitted or simplified.
[0010] <Embodiment 1> Fig. 1 is a diagram showing an example of the configuration of an image capture device 100 according to a first embodiment of the present invention. Note that some of the functional blocks shown in Fig. 1 are realized by causing a CPU or the like serving as a computer (not shown) included in the image capture device 100 to execute a computer program stored in a memory serving as a storage medium (not shown).
[0011] However, some or all of these functions may be implemented by hardware. Examples of such hardware include dedicated circuits (ASICs) and processors (reconfigurable processors, DSPs). Furthermore, the functional blocks shown in Figure 1 do not have to be housed in the same housing, and may be configured as separate devices connected to each other via signal paths.
[0012] The imaging device 100 is applicable to digital still cameras, digital video cameras, in-vehicle cameras, surveillance cameras, smartphones, etc. The imaging device 100 includes an optical system 1, an imaging element 2, an image processing unit 3, a compression / decompression unit 4, a control unit 5, an operation unit 6, an image display unit 7, and an image recording unit 8. The imaging device 100 in this embodiment functions as an image signal processing device.
[0013] The optical system 1 includes a lens, a lens driving mechanism, a mechanical shutter mechanism, an aperture mechanism, etc. Among these, the movable parts are driven based on control signals from the control unit 5.
[0014] The imaging element 2 is, for example, an XY address type CMOS (Complementary Metal Oxide Semiconductor) image sensor, and performs imaging operations in response to control signals from the control unit 5. Furthermore, the imaging signal is digitized by an AD conversion circuit included in the imaging element 2 and output to the image processing unit 3 as an image signal.
[0015] In the image sensor 2 of this embodiment, for example, a first photoelectric conversion element and a second photoelectric conversion element are arranged side by side in each pixel. A common microlens is disposed on the light incident surface of the first photoelectric conversion element and the second photoelectric conversion element. As a result, light from different exit pupils of the photographing lens included in the optical system 1 is incident on the first photoelectric conversion element and the second photoelectric conversion element, respectively.
[0016] Therefore, there is parallax between the first image signal obtained from the group of first photoelectric conversion elements of the plurality of pixels and the second image signal obtained from the group of second photoelectric conversion elements of the plurality of pixels. Note that the image sensor 2 can read out a signal obtained by adding the signals of the first photoelectric conversion elements and the second photoelectric conversion elements for each pixel as image data for display.
[0017] Alternatively, the image sensor 2 may be configured to separately output the first image signal and the second image signal, or to separately read out the image data added for each pixel and the first image signal, thereby allowing the downstream image processor 3 to calculate the second image signal by subtracting the first image signal from the added image data.
[0018] The image processing unit 3 generates a distance image (distance map) by calculating distance information to the subject based on the correlation distance between the first image signal and the second image signal obtained from the image sensor 2. Furthermore, a three-dimensional image is generated based on the image signal and the distance image (distance map), as will be described later.
[0019] Under the control of the control unit 5, the image processing unit 3 also performs image processing such as noise correction and white balance processing on the digitized image signal input from the image sensor 2. The image processing unit 3 also generates a control signal for controlling the focus lens of the optical system 1 based on the distance information described above, and generates control signals for controlling the accumulation time and aperture of the image sensor based on the luminance information of the image signal.
[0020] The image signals and control information that have been subjected to image processing in the image processing unit 3 are output to the control unit 5. Note that at least a part of the image processing for generating a stereoscopic image may be performed in an external image processing device separate from the imaging device 100.
[0021] The compression / decompression unit 4 operates under the control of the control unit 5, and performs compression / encoding processing of image signals and decompression / decoding processing of encoded data of still images. It may also perform compression / encoding / decompression / decoding processing of moving images.
[0022] The control unit 5 is, for example, a microcontroller including a central processing unit (CPU), a read-only memory (ROM), a random access memory (RAM), and the like.
[0023] The CPU serving as a computer of the control unit 5 executes a computer program stored in a storage medium such as a ROM, thereby providing overall control of each unit of the entire imaging device 100. The operation unit 6 is made up of various operation members such as a shutter release button, and outputs control signals to the control unit 5 in response to input operations by the user. Examples of input operations by the user include setting the recording mode for still images, video, etc., and exposure control (aperture, accumulation time of the image sensor, ISO sensitivity).
[0024] The image display unit 7 supplies an image signal to a display device such as an LCD (Liquid Crystal Display) to display an image. The image recording unit 8 is connected to, for example, a portable recording medium and stores compressed and encoded image data files.
[0025] Note that a distance image (distance map) may be further linked to the image data file and recorded in the image recording unit 8. Alternatively, the first image signal and the second image signal may be recorded as an image data file in the image recording unit 8. Alternatively, the image recording unit 8 may record the display image data and the first image signal that are added for each pixel, so that the second image signal can be calculated later.
[0026] By doing as described above, it is possible to generate a three-dimensional image by reading the image data file, distance image (distance map), etc. from the image recording unit 8 at any timing after shooting. Note that the imaging device 100 may have a communication unit, and for example, it can transmit the image data and distance image (distance map), etc. recorded in the image recording unit 8 to an external image processing device. Therefore, 3D data (three-dimensional image) can be generated in the external image processing device.
[0027] Fig. 2 is a flowchart showing an example of image processing according to embodiment 1. Note that the operations of the steps in the flowchart of Fig. 2 and other flowcharts in the following description are sequentially performed by a CPU or the like serving as a computer in the control unit 5 executing a computer program stored in memory.
[0028] 2 starts when an instruction to generate a three-dimensional image is input, for example, via the operation unit 6. In step S20, an image is captured using, for example, the image sensor 2. Alternatively, image data already stored in the image recording unit 8 is acquired.
[0029] In step S21, a 3D data generation process is performed. That is, a stereoscopic image is generated based on image data, a distance image (distance map), etc. Note that 3D data in this embodiment refers to a stereoscopic image from the front up to a predetermined rotation angle range. Note that a detailed example of step S21 will be described with reference to FIG. 3 and FIGS. 4(A) to (D).
[0030] In step S22, adjustment processing is performed on the 3D data, and in step S23, the adjusted 3D data is output. Note that detailed examples of step S22 will be described with reference to FIGS.
[0031] 3 is a flowchart illustrating a detailed example of the 3D data generation process in step S21. In step S30, a new distance map is created. Alternatively, when image data is acquired from the image recording unit 8 in step S20, if a distance map associated with the image data is read, that distance map is acquired.
[0032] Here, step S30 and step S20 function as an information acquisition step (information acquisition unit) for acquiring an image of the subject and distance information.
[0033] Fig. 4(A) is a diagram showing an example of a stereoscopic image obtained in step S20, and Fig. 4(B) is a diagram showing an example of a distance map. In Fig. 4(B), higher density indicates greater distance. In step S31, point cloud conversion is performed based on the distance map data to obtain point cloud data.
[0034] In step S32, a mesh image is generated based on the point cloud data. Fig. 4(C) is a diagram showing an example of a mesh image generated based on the point cloud data. In step S33, a texture image is generated based on the mesh image.
[0035] Fig. 4(D) is a diagram showing an example of a texture image of a three-dimensional shape generated based on the mesh image of Fig. 4(C). This texture image is output as 3D data. After the processing of step S33 is completed, the process proceeds to the adjustment processing of step S22 in Fig. 2. Here, steps S31 to S33 function as a 3D data generation processing step (3D data generation processing unit) that generates 3D data of the subject based on the image and distance information.
[0036] 5 is a flowchart showing a detailed example of the adjustment process in step S22. In step S50, error detection of 3D data is performed. Here, step S50 functions as an error area detection step (error area detection means) that detects an error area in the 3D data. A detailed example of error detection of 3D data in step S50 will be described with reference to FIGS. 6 to 8.
[0037] If an error is detected in the 3D data in step S50, an image adjustment process is performed in step S51. Here, step S51 functions as an image adjustment step (image adjustment unit) that adjusts image information in the error area.
[0038] If no error is detected in the 3D data in step S50, the processing flow in Fig. 5 ends and the process proceeds to step S23 in Fig. 2. A detailed example of the image adjustment process in step S51 will be described with reference to Figs.
[0039] Fig. 6(A) is a flowchart showing a detailed example of the process of detecting errors in 3D data in step S50, and Fig. 6(B) is a diagram showing an example of semantic labeling in the flowchart of Fig. 6(A) applied to a 3D image. In the example of Fig. 6(A), error regions in the 3D data are detected by determining the semantic regions of the 3D data based on machine learning.
[0040] In step S60 of Fig. 6(A), the 3D data generated in step S21 is acquired. In step S61, semantic labeling processing is performed on the 3D data. That is, by image recognition, each partial region of the image is classified into, for example, face, ear, neck, etc., and each partial region is labeled as face, ear, neck, etc., as shown in Fig. 6(B).
[0041] In step S62, the layout of the model labels is compared and collated. That is, the layout of the model labels is compared with the layout of the model labels that have been machine-learned in advance.
[0042] In step S63, it is determined whether or not there is an error label. Specifically, as a result of comparing with the arrangement of the model label, it is checked whether or not there is an unnatural arrangement in the arrangement of the semantic label (for example, the ears are positioned above the face), whether or not the balance of the sizes of the face, ears, neck, etc. is within the normal range, etc.
[0043] If the result of step S63 is No, the flow of Fig. 6A ends and the process proceeds to step S23. On the other hand, if the result of step S63 is Yes, the process determines the error area in step S64. That is, the label containing the error is determined as the error area.
[0044] Then, in step S65, it is determined whether the area of the error region is equal to or greater than a predetermined value. If the determination in step S65 is No, the processing flow of Fig. 6A is terminated and the process proceeds to step S23. On the other hand, if the determination in step S65 is Yes, the process proceeds to the image adjustment process of step S51.
[0045] Fig. 7(A) is a flowchart showing another example of the process of detecting an error in 3D data in step S50, and (B) is a diagram showing an example of a distance map in the process flow of Fig. 7(A). In the process flow of Fig. 7(A), regions in the 3D data at distances equal to or greater than a predetermined value are detected as error regions based on distance information.
[0046] In step S70 of Fig. 7(A), a distance map such as that shown in Fig. 7(B) is acquired. Meanwhile, image data is acquired in step S71. It is assumed that the image data acquired in step S71 includes pupil position information in, for example, EXIF format. That is, information on the positions (coordinates) of the pupils of the face that is the main subject is included in the image data as metadata.
[0047] In step S72, the distance value of the pupil position is calculated based on the information on the pupil position (coordinates) and the distance map. In step S73, a threshold value indicating an allowable range from the distance value of the pupil position is determined. Note that the threshold value may be determined based on, for example, the ratio between the resolution of the polygon formed by the point cloud data in step S31 of FIG. 3 and the resolution of the subject.
[0048] Alternatively, the threshold value may be a distance value at which the gradient of the distance change is equal to or greater than a predetermined value, or a preset difference value of the distance may be used as the threshold value.
[0049] In step S74, it is determined whether there is an area in the distance map that is equal to or greater than the threshold value. That is, it is determined whether there is an area that is at a distance equal to or greater than the threshold value from the pupil position distance. If the determination in step S74 is No, the processing flow in FIG. 7A is terminated and the process proceeds to step S23.
[0050] On the other hand, if the answer in step S74 is Yes, in step S75, the area that is a distance greater than or equal to the threshold from the pupil position distance is determined to be an error area, and in step S76, it is determined whether the area of the error area is greater than or equal to a predetermined value.
[0051] If the determination in step S76 is No, the processing flow of Fig. 7A ends and the process proceeds to step S23. On the other hand, if the determination in step S76 is Yes, the process proceeds to the image adjustment process of step S51.
[0052] FIG. 8A is a flowchart showing yet another example of the processing for detecting errors in 3D data in step S50, and FIG. 8B is a diagram showing an example of a large polygonal region of a stereoscopic image in the processing flow of FIG. 8A.
[0053] In step S80, polygon data (mesh data) is acquired based on the point cloud data in step S31. Next, in step S81, the area of each polygon (each mesh) is calculated based on the polygon data (mesh data) to create an area map. Then, in step S82, it is determined whether the area of each polygon (each mesh) is equal to or less than a threshold value.
[0054] If the determination in step S82 is No, the processing flow of Fig. 8(A) is terminated, and the process proceeds to step S23. On the other hand, if the determination in step S82 is Yes, in step S83, a region where the area of each polygon (each mesh) is larger than the threshold is determined as an error region. Fig. 8(B) shows an example of a region where the area of each polygon (each mesh) is larger than the threshold.
[0055] Then, in step S84, it is determined whether the area of the error region is equal to or larger than a predetermined value, and if the determination is No, the processing flow of Fig. 8(A) is terminated and the process proceeds to step S23. On the other hand, if the determination is Yes in step S84, the process proceeds to the image adjustment process of step S51.
[0056] As described above, error detection is performed in step S50 by executing any one of the processes in Figures 6 to 8. Note that a combination of the processes in Figures 6 to 8 may be executed in step S50, and in this embodiment, it is sufficient to execute at least one of the error detection processes in Figures 6 to 8 in step S50.
[0057] Fig. 9 is a flowchart showing an example of the image adjustment process in step S51. In the example shown in Fig. 9, the image adjustment means performs at least one of upsampling and image suppression on the error region.
[0058] In step S90, 3D data is acquired, and in step S91, information on the error region detected in step S50 is acquired. Next, in step S92, upsampling processing of the error region is performed.
[0059] That is, if the error region is shrinking, for example, the sampling of that region is increased to widen the region. For example, in the process of Figure 6(A), if the width of the ear region is narrow compared to the model label, upsampling is performed to widen the width.
[0060] In this case, the upsampling resolution is set to a resolution that makes the size of each polygon the same in the depth direction and horizontal direction, or a resolution that makes the area the same as the polygon size of the pupil of the face.
[0061] Next, in step S93, image suppression processing (for example, filtering and smoothing processing) is performed to make the boundary between the area upsampled in step S92 and other areas less noticeable. Then, in step S94, UV coordinate values are updated. After step S93, the process proceeds to step S23. Here, step S93 performs at least one image suppression processing.
[0062] 10 is a flowchart showing another example of the image adjustment process in step S51. In step S101, 3D data is acquired, and in step S102, information on the error region detected in step S50 is acquired. Meanwhile, in step S103, a distance map is acquired.
[0063] Next, in step S104, a distance map of the error region is extracted, and in step S105, image suppression processing is performed to make the image less noticeable. The suppression processing in step S105 includes processing to reduce at least one of the brightness, saturation, contrast, and transparency of the error region, for example.
[0064] Furthermore, the suppression process to make the image less noticeable is gradually strengthened according to the amount of error (for example, the difference in distance from the pupil position). That is, the suppression process is strengthened according to the degree of error in the error area (error amount, etc.). The size of the polygon may also be used as the amount of error.
[0065] That is, the larger the area of each polygon (mesh) shown in Figure 8(B), the stronger the suppression process may be. Also, since the amount of error often increases toward the periphery of the subject, the suppression process may be gradually strengthened toward the periphery of the subject.
[0066] As described above, according to the first embodiment, an error area in the 3D data is detected, and an image suppression process is performed to make the error area less noticeable, thereby obtaining a 3D image that creates a less uncomfortable feeling. Note that the image adjustment process in step S51 may be a combination of the processes shown in Figures 9 and 10, and it is sufficient to execute at least one of the processes shown in Figures 9 and 10.
[0067] <Embodiment 2> In the second embodiment, an error area in the 3D data is detected and the error area is interpolated to obtain a 3D image that gives a less uncomfortable feeling. Note that the first and second embodiments may be combined.
[0068] 11 is a flowchart showing an example of image processing according to embodiment 2. Steps S1101 and S1102 are the same as steps S20 and S21 in Fig. 2, respectively, and therefore will not be described here. In step S1103, a missing portion in the image is detected, and if detected, the process proceeds to step S1104, and if not, the process proceeds to step S1105.
[0069] Note that step S1103 functions as an error region detection unit that detects error regions in the 3D data. That is, in the second embodiment, missing portions are detected as error regions in the 3D data. In this way, the error region includes a missing region in the 3D data.
[0070] Figure 12(A) is a flowchart showing an example of the missing portion detection process in step S1103, (B) is a diagram showing an example of a 3D image when there is no missing portion, and (C) is a diagram showing an example of a 3D image when there is a missing portion. Note that Figure 12(B) is the same as Figure 6(B), and shows a state in which the image is distorted but has no missing portion. Note that in Figure 12(A), missing areas of the 3D data are detected based on the positional relationship of the semantic areas of the 3D data.
[0071] On the other hand, in Fig. 12(C), a defect occurs in the neck area. In Fig. 12(A), steps S1201 to S1203 are the same as steps S60 to S62 in Fig. 6, and therefore a description thereof will be omitted.
[0072] In step S1204, it is determined whether there is a missing label. That is, the result of comparing with the arrangement of model labels previously learned by machine learning is checked to see whether there is a missing label in the arrangement of the semantic labels. For example, as shown in FIG. 12(C), if the neck is not covered by hair and the width of the neck is narrow, it is determined that there is a missing label in the neck label area. For example, if part of the neck is covered by hair, it may be determined that there is no missing label.
[0073] If the determination in step S1204 is No, that is, if it is determined that there is no missing portion, the processing flow in FIG. 12(A) ends, and the process proceeds to the output processing of step S1105 in FIG. 11. On the other hand, if the determination in step S1204 is Yes, the missing portion is determined in step S1205. That is, the partial region in which the missing portion occurs is determined to be the missing portion. After the processing of step S1205, the process proceeds to the interpolation processing of step S1104 in FIG. 11.
[0074] 13A is a flowchart showing another example of the missing portion detection process in step S1103, (B) is a diagram showing an example of a 3D image when a missing portion occurs, and (C) is a diagram showing an example of the shape of the 3D image of the model. In the example shown in Fig. 13, the error region detection means detects an error (missing portion) in the shape of the 3D data, and detects the missing region based on the result of comparing the 3D data with the 3D shape of a model stored in advance.
[0075] In step S1301, 3D data is acquired, and in step S1302, the three-dimensional shape of the model as shown in Fig. 13(C) is acquired. Note that the model shape is a pre-created shape of the model as viewed from the front, for example. Note that the model shape as viewed from diagonal left and right fronts, for example, may also be used.
[0076] In step S1303, the position and size of the 3D data acquired in step S1301 are matched to the position and size of the shape of the model acquired in step S1302. At this time, the position and size of the shape of the model may be matched to the 3D data as described above based on the shooting conditions (distance to the subject, lens zoom information, etc.) and organ information.
[0077] Then, in step S1304, it is determined whether the non-corresponding region, where the position and size of the 3D data and the shape of the model do not match, is equal to or smaller than a predetermined area. If the determination in step S1304 is Yes, the processing flow of Fig. 13(A) ends and the process proceeds to the output processing in step S1105. On the other hand, if the determination in step S1304 is No, in step S1305 the non-corresponding region is determined to be a missing portion, and the process proceeds to the interpolation processing in step S1104.
[0078] Figure 14(A) is a flowchart showing yet another example of the missing portion detection process in step S1103, and (B) is a diagram showing an example of the edge of a three-dimensional image when a missing portion has occurred. In the example of Figure 14(A), a portion below the average value based on the distance distribution of the edge is detected as a missing portion. That is, in the example of Figure 14(A), the error region detection means detects errors in the 3D data image, and in particular detects a missing region based on the distance distribution of the edges of the 3D data.
[0079] 14A, 3D data is acquired in step S1401, edges are detected from the 3D data in step S1402, and then distance distribution of the edges detected in step S1402 is acquired in step S1403.
[0080] Then, in step S1404, it is determined whether the range of the distance distribution acquired in step S1403 is equal to or smaller than a predetermined range. If the determination in step S1404 is Yes, the processing flow of Fig. 14(A) is terminated, and the process proceeds to the output processing of step S1105.
[0081] On the other hand, if the determination in step S1404 is No, the process proceeds to step S1405, where the average value of the edges is calculated, and in step S1406, the average value calculated in step S1405 is compared with the edge distance. Then, in step S1407, a missing portion is determined based on the comparison result in step S1406. After the processing in step S1407, the interpolation processing in step S1104 is performed.
[0082] 14, the edge shape of the 3D data may be calculated for each distance, and if the edge shape at each distance deviates from a predetermined pattern stored in advance by more than a predetermined percentage, the deviated portion may be determined to be a missing portion. In other words, the portion where the cross-sectional shape of the 3D data for each distance deviates from the predetermined pattern by more than a predetermined amount may be determined to be a missing portion.
[0083] 15 is a flowchart showing an example of the image interpolation process in step S1104. Step S1104 functions as an image adjustment unit that adjusts image information in an error area. That is, in the second embodiment, the image adjustment unit interpolates the error area (missing area).
[0084] 15 shows an example of interpolating the shape and image of a missing part using machine learning. In step S1501, 3D data is acquired, and in step S1502, information such as the position and distance of the missing part is acquired.
[0085] Then, in step S1503, the shape after repair when the defect portion acquired in step S1502 is repaired is estimated based on the machine-learned model. Furthermore, in step S1504, filtering processing is performed so as to smooth out the step in the shape of the boundary between the defect portion after repair and the other portions.
[0086] The filtering process in step S1504 includes, for example, a process of calculating a moving average of boundary edges, a process of calculating a weighted average of overlapping regions, etc. In this way, in steps S1503 and S1504, the shape of the error region (missing region) of the 3D data is interpolated using machine learning.
[0087] Meanwhile, in step S1505, a post-repair image is estimated based on the information about the missing portion acquired in step S1502 and the machine-learned model. Furthermore, in step S1506, a filtering process is performed to smooth out any step in the image at the boundary between the repaired missing portion and the remaining portion. In this way, in steps S1505 and S1506, the image of the error region (missing region) in the 3D data is interpolated using machine learning.
[0088] The filtering process in step S1506 also includes, for example, a process of calculating a moving average of the brightness, saturation, etc. of the boundary, and a process of calculating a weighted average of the brightness, saturation, etc. of the overlapping region.
[0089] Furthermore, in step S1507, the shape after restoration obtained in step S1504 and the image after restoration obtained in step S1506 are synthesized, and in step S1508, the interpolated portion is suppressed.
[0090] The suppression processing in step S1508 may be the same as the suppression processing in step S105 in Fig. 10. That is, suppression processing such as reducing the brightness, saturation, or contrast of the interpolated portion is performed.
[0091] Next, Fig. 16 is a flowchart showing another example of the image interpolation process in step S1104. Fig. 16 shows an example in which the shape of the missing portion is extrapolated, and the image of the missing portion is interpolated using machine learning.
[0092] In step S1601, 3D data is acquired, and in step S1602, information such as the position and distance of the missing part is acquired. In addition, in step S1603, the edge area of the missing part is acquired, and in step S1604, the edge is extrapolated.
[0093] Furthermore, in step S1605, filtering is performed to smooth out any step differences in the shape of the boundary between the repaired defective portion and the remaining portion. That is, processing is performed to reduce the step differences in the boundary between the interpolated region and the remaining region.
[0094] It should be noted that step S1605 may be the same as step S1504. In this way, in steps S1604 and S1605, the shape of the error region of the 3D data is interpolated by extrapolation.
[0095] Note that steps S1606 to S1609 are the same as steps S1505 to S1508 in Fig. 15, and therefore description thereof will be omitted. As in the processing flow shown in Fig. 16, the shape of the missing portion may be determined by extrapolation or interpolation.
[0096] Fig. 17 is a flowchart showing yet another example of the image interpolation process in step S1104. Fig. 17 shows an example in which the shape of the missing part is interpolated using a model shape, and the image is interpolated using machine learning. Steps S1701 and S1702 are the same processes as steps S1502 and S1503 in Fig. 15, and therefore their explanation will be omitted.
[0097] In step S1703, pre-stored model shape data is acquired. Then, in step S1704, the 3D data is fitted to the model shape to interpolate the missing portion. That is, the shape of the error region (missing portion) in the 3D data is interpolated using the pre-stored model shape data.
[0098] Steps S1705 to S1709 are not described here because they are the same as steps S1504 to S1508 in Fig. 15. As shown in Fig. 17, missing portions may be interpolated using pre-stored model shape data.
[0099] In the first embodiment, an error area in the 3D data is detected and an image adjustment process is performed to obtain a 3D image with reduced discomfort, while in the second embodiment, a missing portion of the 3D data is detected and interpolated to obtain a 3D image with reduced discomfort. However, the processes in the first and second embodiments may be combined as appropriate. This allows a 3D image with no discomfort to be obtained even if, for example, part of the 3D data is distorted or missing.
[0100] Note that error information regarding the amount of error (degree of error), error area, etc. may be stored as metadata of the image file of the 3D data. Then, image adjustment such as suppression processing or interpolation processing may be performed on the error area based on the error information stored as metadata of the image file.
[0101] The present invention has been described above in detail based on its preferred embodiments, but the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible based on the spirit of the present invention, and these are not excluded from the scope of the present invention.
[0102] The present invention also includes those that realize the functions of the above embodiments using, for example, at least one processor such as a CPU, memory, or circuit (for example, ASIC). Also, multiple processors may be used to perform distributed processing.
[0103] In order to realize part or all of the control in the above-described embodiments, a computer program that realizes the functions of the above-described embodiments may be supplied to an image processing device or the like via a network or various storage media. Then, a computer (or a CPU, MPU, or the like) in the image processing device or the like may read and execute the program. In this case, the program and the storage medium storing the program constitute the present invention. The present invention also includes the following combinations.
[0104] (Configuration 1) An image processing device comprising: an information acquisition unit that acquires an image and distance information of a subject; a 3D data generation processing unit that generates 3D data of the subject based on the image and the distance information; an error area detection means that detects an error area in the 3D data; and an image adjustment means that adjusts the image information of the error area.
[0105] (Configuration 2) The image processing device according to configuration 1, wherein the image adjustment means interpolates the error region.
[0106] (Configuration 3) The image processing device according to configuration 1 or 2, wherein the error region detection means detects an error in the shape of the 3D data or the image.
[0107] (Configuration 4) An image processing device described in any one of configurations 1 to 3, characterized in that the error area detection means detects the error area in the 3D data by determining the semantic area of the 3D data based on machine learning.
[0108] (Configuration 5) The image processing device described in any one of configurations 1 to 4, characterized in that the error area detection means detects an area in the 3D data at a distance equal to or greater than a predetermined value as the error area based on the distance information.
[0109] (Configuration 6) The image processing device according to any one of configurations 1 to 5, wherein the image adjustment means performs at least one of upsampling processing and suppression processing of the image on the error region.
[0110] (Configuration 7) The image processing device according to configuration 6, wherein the suppression processing includes processing for reducing at least one of brightness, saturation, contrast, and transparency of the image in the error region.
[0111] (Configuration 8) The image processing device according to configuration 7, wherein the suppression processing is performed in a stronger manner depending on the degree of error in the error region.
[0112] (Configuration 9) The image processing device according to any one of configurations 1 to 8, characterized in that the suppression processing is performed more strongly in the peripheral areas of the subject.
[0113] (Configuration 10) The image processing device according to any one of configurations 1 to 9, wherein the error region includes a missing region in the 3D data.
[0114] (Configuration 11) The image processing device according to Configuration 10, wherein the error area detection means detects the missing area of the 3D data based on the arrangement relationship of semantic areas of the 3D data.
[0115] (Configuration 12) The image processing device according to configuration 19 or 11, wherein the error region detection means detects the defective region based on a distance distribution of the edges of the 3D data.
[0116] (Configuration 13) The image processing device according to any one of configurations 10 to 12, wherein the error area detection means detects the missing area based on a comparison result between the 3D data and a model shape stored in advance.
[0117] (Configuration 14) The image processing device according to any one of configurations 1 to 13, wherein the image adjustment means interpolates the image of the error region of the 3D data using machine learning.
[0118] (Configuration 15) The image processing device according to any one of configurations 1 to 14, wherein the image adjustment means interpolates the shape of the error region of the 3D data using machine learning.
[0119] (Configuration 16) The image processing device according to any one of configurations 1 to 15, wherein the image adjustment means interpolates the shape of the error region of the 3D data by extrapolation.
[0120] (Configuration 17) The image processing device according to any one of configurations 1 to 16, wherein the image adjustment means interpolates the shape of the error region of the 3D data using pre-stored model shape data.
[0121] (Configuration 18) An image processing device according to any one of configurations 1 to 17, characterized in that the image adjustment means performs processing to reduce the step at the boundary between the area where the error area of the 3D data has been interpolated and the other area.
[0122] (Configuration 19) An image processing device according to any one of configurations 1 to 18, characterized in that error information regarding the error area is stored as metadata of an image file, and the image adjustment means adjusts the image information of the error area based on the error information stored as the metadata.
[0123] (Method) An image processing method comprising an information acquisition step of acquiring an image and distance information of a subject, a 3D data generation processing step of generating 3D data of the subject based on the image and the distance information, an error area detection step of detecting an error area in the 3D data, and an image adjustment step of adjusting image information of the error area.
[0124] (Program) A computer program for controlling each unit of the image processing device according to any one of configurations 1 to 19 by a computer. [Explanation of symbols]
[0125] 100: Imaging device 1:Optical system 2: Image sensor 3: Image processing section 5: Control unit 6:Operation unit 7: Image display section 8: Image recording unit
Claims
1. an information acquisition unit that acquires an image of a subject and distance information; a 3D data generation processing unit that generates 3D data of the subject based on the image and the distance information; an error region detection means for detecting an error region of the 3D data; and an image adjustment unit for adjusting image information of the error area.
2. 2. The image processing apparatus according to claim 1, wherein said image adjustment means interpolates said error region.
3. 2. The image processing apparatus according to claim 1, wherein the error area detection means detects an error in the shape of the 3D data or in the image.
4. The image processing device according to claim 1 , wherein the error region detection means detects the error region in the 3D data by determining a semantic region of the 3D data based on machine learning.
5. The image processing apparatus according to claim 1 , wherein the error region detection means detects, based on the distance information, a region in the 3D data whose distance is equal to or greater than a predetermined value as the error region.
6. 2. The image processing apparatus according to claim 1, wherein the image adjustment means performs at least one of upsampling processing and image suppression processing on the error region.
7. 7. The image processing apparatus according to claim 6, wherein the suppression processing includes processing for reducing at least one of brightness, saturation, contrast, and transparency of the image in the error region.
8. 8. The image processing device according to claim 7, wherein the suppression processing is performed in a stronger manner depending on the degree of error in the error region.
9. 8. The image processing device according to claim 7, wherein the suppression processing is performed more strongly at the periphery of the subject.
10. The image processing device according to claim 1 , wherein the error region includes a missing region in the 3D data.
11. 11. The image processing apparatus according to claim 10, wherein the error area detection means detects the missing area of the 3D data based on a layout relationship of semantic areas of the 3D data.
12. 11. The image processing apparatus according to claim 10, wherein the error region detection means detects the defective region based on a distance distribution of an edge of the 3D data.
13. 11. The image processing apparatus according to claim 10, wherein the error area detection means detects the defective area based on a result of comparison between the 3D data and a pre-stored model shape.
14. The image processing device according to claim 1 , wherein the image adjustment means interpolates the image of the error region of the 3D data using machine learning.
15. The image processing device according to claim 1 , wherein the image adjustment means interpolates the shape of the error region of the 3D data using machine learning.
16. 2. The image processing apparatus according to claim 1, wherein the image adjustment means interpolates the shape of the error region of the 3D data by extrapolation.
17. 2. The image processing apparatus according to claim 1, wherein the image adjustment means interpolates the shape of the error region of the 3D data using pre-stored model shape data.
18. 2. The image processing device according to claim 1, wherein the image adjustment means performs processing to reduce a step at a boundary between an area where the error area of the 3D data has been interpolated and the other area.
19. 2. The image processing device according to claim 1, wherein error information regarding the error area is stored as metadata of the image file, and the image adjustment means adjusts the image information of the error area based on the error information stored as the metadata.
20. an information acquisition step of acquiring an image of a subject and distance information; a 3D data generation processing step for generating 3D data of the subject based on the image and the distance information; an error region detection step of detecting an error region of the 3D data; an image adjustment step of adjusting image information of the error region.
21. A computer program for controlling each unit of the image processing device according to any one of claims 1 to 19 by a computer.
Citation Information
Patent Citations
Image processing apparatus, control method, program, and image processing system
JP2024008596A