Information processing device, information processing method, and recording medium
The dual-camera system effectively detects and corrects partial depth estimation errors in three-dimensional shape data by generating and displaying errors for user input correction, improving the accuracy of three-dimensional models.
Patent Information
- Application Number
- PCT/JP2024/025823
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-22
AI Technical Summary
Existing technologies for generating three-dimensional shape data from images captured by cameras often suffer from partial depth estimation errors, particularly at boundaries like the face-neck junction, which are difficult to detect and correct.
The system employs dual camera and illumination units to capture images from different directions, generates three-dimensional shape data, detects partial depth estimation errors, and allows user-driven correction through error display and input, using phase unwrapping and integer multiple adjustments.
Accurately identifies and corrects depth estimation errors in three-dimensional shape data, enhancing the precision and usability of the generated models.
Smart Images

Figure JP2024025823_22012026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] The present disclosure relates to the technical fields of an information processing device, an information processing method, and a recording medium.
[0002] Known examples of this type of device include one that generates data representing a three-dimensional shape from images captured by a camera. For example, Patent Document 1 discloses a technology in which a person's head is captured by a right and left camera unit, and the three-dimensional shape is calculated from the images captured by each camera unit and output to a display or the like.
[0003] International Publication No. 2022 / 034694
[0004] An object of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve upon the techniques disclosed in prior art documents.
[0005] One aspect of the information processing device disclosed herein includes a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object with light from the first direction; a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object with light from the second direction; a generation means that generates three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; a detection means that detects a partial depth estimation error in the three-dimensional shape data based on an image photographed by at least one of the first camera and the second camera; an output means that outputs error display information that displays the three-dimensional shape data so that a portion in which the depth estimation error has been detected is displayed differently from other portions in which the depth estimation error has not been detected; and a correction means that corrects the portion in which the depth estimation error has been detected using an image photographed by at least one of the first camera and the second camera, based on instructions input by a user after outputting the error display information.
[0006] One aspect of the information processing method disclosed herein is an information processing method for controlling an information processing device including: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object from the first direction; and a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object from the second direction, wherein the information processing method generates three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera, detects partial depth estimation errors in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera, outputs error display information that displays the three-dimensional shape data so that a portion in which the depth estimation error is detected is displayed differently from other portions in which the depth estimation error is not detected, and corrects the portion in which the depth estimation error is detected using an image photographed by at least one of the first camera and the second camera based on instructions input by a user after outputting the error display information.
[0007] One aspect of the recording medium disclosed herein is an information processing method for controlling an information processing device including: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object from the first direction; and a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object from the second direction, wherein the information processing method generates three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; detects partial depth estimation errors in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera; outputs error display information that displays the three-dimensional shape data so that a portion where the depth estimation error is detected is displayed differently from other portions where the depth estimation error is not detected; and corrects the portion where the depth estimation error is detected using an image photographed by at least one of the first camera and the second camera based on instructions input by a user after the error display information has been output.
[0008] 1 is a block diagram showing a hardware configuration of a first information processing device. 2 is a perspective view showing the configuration of the first information processing device. 3 is a top view showing the configuration of the first information processing device. 4 is a block diagram showing a functional configuration of the first information processing device. 5 is a flowchart showing the flow of operations of the first information processing device. 6 is a plan view showing an example of display by the first information processing device. 7 is a flowchart showing the flow of operations of an error detection operation by a second information processing device. 8 is a schematic diagram showing a specific example of operations of an error detection operation by the second information processing device. 9 is a flowchart showing the flow of operations of a third information processing device. 10 is a schematic diagram showing an example of display by the third information processing device. 11 is a block diagram showing the functional configuration of a fourth information processing device. 12 is a flowchart showing the flow of operations of the fourth information processing device.
[0009] Hereinafter, embodiments of an information processing device, an information processing method, and a recording medium will be described with reference to the drawings.
[0010] First Embodiment A first information processing apparatus will be described with reference to FIGS. 1 to 6. FIG.
[0011] (Hardware Configuration) First, the hardware configuration of the first information processing apparatus will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the hardware configuration of the first information processing apparatus.
[0012] 1, the first information processing device 1 includes a processor 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, a storage device 14, an input device 15, an output device 16, a first unit 21, and a second unit 22. The processor 11, RAM 12, ROM 13, storage device 14, input device 15, output device 16, first unit 21, and second unit 22 are connected to each other via a data bus 17. Note that the data bus 17 may be an interface other than a data bus (e.g., a LAN, a USB, etc.).
[0013] The processor 11 loads a computer program. For example, the processor 11 is configured to load a computer program stored in at least one of the RAM 12, the ROM 13, and the storage device 14. Alternatively, the processor 11 may load a computer program stored in a computer-readable storage medium using a storage medium reading device (not shown). The processor 11 may acquire (i.e., load) the computer program from a device (not shown) located outside the first information processing device 1 via a network interface. The processor 11 executes the loaded computer program to perform various processes. When the processor 11 executes the loaded computer program, functional blocks related to the processes performed by the first information processing device 1 are realized within the processor 11. In other words, the processor 11 may function as a controller that executes each control in the first information processing device 1.
[0014] The processor 11 may be configured as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a quantum processor. The processor 11 may be configured as one of these, or may be configured to use multiple processors in parallel.
[0015] The RAM 12 temporarily stores computer programs executed by the processor 11. The RAM 12 temporarily stores data that the processor 11 temporarily uses while it is executing the computer programs. The RAM 12 may be, for example, a dynamic random access memory (D-RAM) or a static random access memory (SRAM). Alternatively, other types of volatile memory may be used instead of the RAM 12.
[0016] The ROM 13 stores computer programs executed by the processor 11. The ROM 13 may also store fixed data. The ROM 13 may be, for example, a programmable read-only memory (PROM) or an erasable read-only memory (EPROM). Alternatively, other types of non-volatile memory may be used instead of the ROM 13.
[0017] The storage device 14 stores data that is to be saved long-term by the first information processing device 1. The storage device 14 may operate as a temporary storage device for the processor 11. The storage device 14 may store computer programs executed by the processor 11. The storage device 14 may include, for example, at least one of a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device.
[0018] The input device 15 is a device that receives input instructions from a user of the first information processing device 1. The input device 15 may include, for example, at least one of a keyboard, a mouse, and a touch panel. The input device 15 may also be, for example, a device that includes a microphone and is capable of voice input.
[0019] The output device 16 is a device that outputs information related to the first information processing device 1 to the outside. For example, the output device 16 may be a display device (e.g., a display or a monitor) that can display information related to the first information processing device 1. The output device 16 may also be a speaker or the like that can output information related to the information processing device 1 as audio.
[0020] The first unit 21 is configured to include a first camera 211 and a first irradiating unit 212. The first camera 211 is arranged to capture an image of the target from a first direction. The first irradiating unit 212 is arranged to irradiate light onto the target from the first direction (i.e., the shooting direction of the first camera 211). Similarly, the second unit 22 is configured to include a second camera 221 and a second irradiating unit 222. The second camera 221 is arranged to capture an image of the target from a second direction. The second irradiating unit 222 is arranged to irradiate light onto the target from the second direction (i.e., the shooting direction of the second camera 221). The specific arrangement of the first unit 21 and the second unit 22 will be described in detail later.
[0021] The first camera 211 and the second camera 221 may include a solid-state imaging element such as a CCD (Charge Coupled Device) image sensor, a CMOS (Complementary Metal Oxide Semiconductor) image sensor, etc. The first camera 211 and the second camera 221 may also include an optical system that forms an image of the subject on the imaging surface of the solid-state imaging element, a signal processing circuit that processes the output of the solid-state imaging element to obtain a luminance value for each pixel, etc.
[0022] The first irradiating unit 212 and the second irradiating unit 222 may be configured to project a periodic light pattern. For example, the first irradiating unit 212 and the second irradiating unit 222 may be configured to project sinusoidal wave patterns with different periods. The first irradiating unit 212 and the second irradiating unit 222 may be, for example, a DLP (Digital Light Processing) projector or a liquid crystal projector.
[0023] 1. For example, the first information processing device 1 may be configured to include only the processor 11, RAM 12, and ROM 13 among the above-mentioned components. In this case, the storage device 14, input device 15, output device 16, first unit 21, and second unit 11 may be provided as devices external to the first information processing device 1. Furthermore, some of the calculation functions of the first information processing device 1 may be realized by an external server, a cloud, or the like.
[0024] (Specific Device Configuration) Next, a specific configuration of the first information processing device 1 (particularly, an example of the arrangement of the first unit 21 and the second unit 22) will be described with reference to Figures 2 and 3. Figure 2 is a perspective view showing the configuration of the first information processing device. Figure 3 is a top view showing the configuration of the first information processing device.
[0025] 2 and 3 , the first information processing device 1 is configured as a device that photographs the head of the subject 50 (specifically, a facial region including the face of the subject 50 and a neck region including the neck of the subject 50). The subject 50 may be photographed, for example, with the subject sitting in front of the first information processing device 1. Alternatively, the subject 50 may be photographed with the subject standing in front of the first information processing device 1.
[0026] The first unit 21 in the first information processing device 1 is disposed on the left front side as viewed from the target 50. Therefore, the first camera 211 photographs the head of the target 50 from the left front side. Furthermore, the first irradiating section 212 irradiates light from the left front side toward the head of the target 50. The second unit 22 in the first information processing device 1 is disposed on the right front side as viewed from the target 50. Therefore, the second camera 221 photographs the head of the target 50 from the right front side. Furthermore, the second irradiating section 222 irradiates light from the right front side toward the head of the target 50.
[0027] The first information processing device 1 includes a third camera 300 that captures an image of the head of the target 50 from the front of the target 50. The third camera 300 may be a camera similar to the first camera 211 and the second camera 221 described above. However, the third camera 300 is not an essential component of the first information processing device 1.
[0028] (Functional Configuration) Next, the functional configuration of the first information processing device 1 will be described with reference to Fig. 4. Fig. 4 is a block diagram showing the functional configuration of the first information processing device.
[0029] 4, the first information processing device 1 is configured to include, as components for realizing its functions, a generation unit 110, a detection unit 120, an output unit 130, and a correction unit 140. Each of the generation unit 110, the detection unit 120, the output unit 130, and the correction unit 140 may be a processing block realized by the above-mentioned processor 11 (see FIG. 1).
[0030] The generation unit 110 is configured to acquire an image of the target 50 captured by the first camera 211 (hereinafter referred to as the "first image" as appropriate) and an image of the target 50 captured by the second camera 221 (hereinafter referred to as the "second image" as appropriate). The generation unit 110 is then configured to generate three-dimensional shape data of the target 50 using the acquired first and second images. That is, the generation unit 110 is configured to generate data indicating the three-dimensional shape of the target's head using two images captured from different directions. The generation unit 110 first calculates phase values of each part of the target's head from the first and second images captured by projecting light of a predetermined pattern (e.g., a striped pattern) onto the target's head. Specifically, the generation unit 110 calculates relative phase values based on the luminance values of the captured images, performs phase unwrapping, and calculates absolute phase values. The generation unit 110 then estimates the depth of the target's head based on the calculated absolute phase values. Examples of such three-dimensional measurement methods include a sinusoidal grating phase shift method, a method that combines a sinusoidal wave pattern with a different period or a brightness gradient pattern in addition to the sinusoidal grating phase shift method, etc. These methods can achieve both high measurement accuracy and high measurement speed, and are suitable for measuring the face of a person, for example, who has difficulty remaining still for a long period of time.
[0031] The detection unit 120 is configured to detect partial depth estimation errors occurring in the 3D shape data generated by the generation unit 110. Here, a depth estimation error refers to an error occurring in the depth estimation result when generating 3D shape data. For example, when using the sinusoidal grating phase shift method, if an incorrect connection occurs during the phase unwrapping process, 3D shape data that differs from the original shape is generated. Note that incorrect connection during the phase unwrapping process is likely to occur in areas where the depth position changes discontinuously. Therefore, for example, around the boundary between the face region and the neck region of a human head, an incorrect 3D shape is likely to be generated. More specifically, 3D shape data may be generated in which the neck region protrudes forward relative to the face region. The detection unit 120 detects such depth estimation errors based on at least one of the first image and the second image. Note that a more specific method for detecting depth estimation errors will be described in detail in another embodiment described later.
[0032] The output unit 130 is configured to output information for displaying 3D shape data regarding depth estimation errors detected by the detection unit 120 so that the portion where the depth estimation error occurs can be identified. More specifically, the output unit 130 outputs information for displaying 3D shape data such that the portion where the depth estimation error is detected and the portion where the depth estimation error is not detected are displayed in different ways (hereinafter, appropriately referred to as "error display information"). For example, the output unit 130 may display the 3D data such that the portion where the depth estimation error is detected and the portion where the depth estimation error is not detected are displayed in different colors. In this case, the output unit 130 may display the portion where the depth estimation error is detected in a conspicuous color (e.g., red or yellow). Alternatively, the output unit 130 may display the portion where the depth estimation error is not detected in the normal manner, while adding movement (e.g., adding animation such as blinking) to the portion where the depth estimation error is detected. The error display information output by the output unit 130 may be output to a display unit included in the output device 16 (see FIG. 1 ), for example. The display unit may be one that displays at least to the user of the device (for example, the person photographing the object 50).
[0033] After the output unit 130 outputs the error display information (i.e., after a display corresponding to the error display information is performed), the correction unit 140 corrects the portion of the 3D shape data where a depth estimation error has been detected, based on an instruction (hereinafter referred to as a "correction instruction") input by the user of the device. Specifically, when the correction instruction is input, the correction unit 140 corrects the depth estimation error in the 3D shape data using at least one of the first image and the second image. That is, the correction unit 140 corrects the 3D shape data so that no depth estimation error has occurred. For example, if an incorrect connection occurs during phase unwrapping, the phase value will, in principle, be shifted by an integer multiple of 2π. Therefore, the correction unit 140 may correct the phase value of the portion where a depth estimation error has been detected by adding or subtracting an integer multiple of 2π to or subtracting from the phase value, and then perform a process to correct the depth based on the correspondence between the phase value and the depth. The correction instruction may be input by a user operation. For example, the user may input the correction instruction by touching the screen of the display unit configured as a touch panel. Alternatively, the user may input correction instructions using an input device such as a mouse or keyboard.
[0034] (Flow of Operation) Next, the flow of operation by the first information processing device 1 will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the flow of operation of the first information processing device.
[0035] 5, when the operation of the first information processing device 1 is started, the first camera 211 and the second camera 221 first capture a first image and a second image of the target 50 (step S101). Note that multiple first images and multiple second images may be captured.
[0036] Next, the generation unit 110 generates three-dimensional shape data of the object 50 using the first image and the second image (step S102). After that, the detection unit 120 detects depth estimation errors occurring in the three-dimensional shape data generated by the generation unit 110 (step S103).
[0037] Next, the output unit 130 outputs error display information regarding the depth estimation error detected by the detection unit 120 (step S104). After the output unit 130 outputs the error display information, the first information processing device 1 enters a standby state for accepting a correction instruction from the user. Note that if the detection unit 120 does not detect a depth estimation error (i.e., if no depth estimation error has occurred), the output unit 130 does not need to output the error display information. In this case, the series of operations may end at this stage.
[0038] Subsequently, when a correction instruction is input from the user (step S105: YES), the correction unit 140 corrects the portion of the 3D shape data where a depth estimation error was detected using at least one of the first image and the second image (step S106). On the other hand, when a correction instruction is not input from the user (step S105: NO), the above-mentioned step S106 may be omitted. That is, the correction unit 140 may end the series of operations without correcting the depth estimation error in the 3D shape data.
[0039] The correction unit 140 may output the corrected three-dimensional shape data. For example, the correction unit 140 may display the corrected three-dimensional shape data on a display unit or the like. The correction unit 140 may also output the corrected three-dimensional shape data to an external database or the like for storage.
[0040] (Display Example) Next, a display example when the first information processing device 1 operates will be described with reference to Fig. 6. Fig. 6 is a plan view showing a display example by the first information processing device.
[0041] In Fig. 6 , when the first information processing device 1 operates, the output unit 130 outputs error display information in accordance with a depth estimation error detected from the three-dimensional shape data. As a result, a display in accordance with the error display information is performed on the display unit. Specifically, the three-dimensional shape data generated by the generation unit 110 is displayed so that the portion where a depth estimation error is detected and the portion where a depth estimation error is not detected are displayed in different ways. In the example shown in Fig. 6 , three-dimensional shape data is generated in which the neck region of the object 50 protrudes forward relative to the face region due to a depth estimation error. Therefore, the neck region where the depth estimation error occurs is displayed in a different color from the other portions.
[0042] In addition, a recalculation button is also displayed together with the three-dimensional shape data corresponding to the error display information. This recalculation button is used by the user to input a correction instruction. When the user presses the recalculation button, the first information processing device 1 accepts the input of the correction instruction.
[0043] When a correction instruction is input, the correction unit 140 corrects the three-dimensional shape data. As a result, the corrected three-dimensional shape data is displayed on the display unit. In the example shown in FIG. 6, the neck region of the object 50, which had protruded due to a depth estimation error, is corrected. Therefore, in the corrected three-dimensional shape data, the face region and neck region of the object 50 are displayed correctly.
[0044] (Technical Effects) Next, technical effects obtained by the first information processing device 1 will be described.
[0045] 1 to 6 , the first information processing device 1 detects a partial depth estimation error in the three-dimensional shape data, and displays the three-dimensional shape data so that the portion where the depth estimation error was detected is displayed differently from the other portions. Then, when a correction instruction is input by a user viewing the display, the depth estimation error in the three-dimensional shape data is corrected. In this way, it is possible to present the portion where the depth estimation error occurred in an easy-to-understand manner to the user, and to perform appropriate correction of the three-dimensional shape data in response to a user request.
[0046] Second Embodiment A second information processing device 1 will be described with reference to Figures 7 and 8. The second information processing device 1 differs in some configurations and operations from the first information processing device 1 described above, but other parts may be similar to the first information processing device 1. Therefore, the following will describe in detail the parts that differ from the first embodiment, and will omit explanations of other overlapping parts as appropriate.
[0047] (Error Detection Operation) First, the error detection operation by the second information processing device 1 (i.e., the operation when the detection unit 120 detects a depth estimation error) will be described with reference to Fig. 7. Fig. 7 is a flowchart showing the flow of the error detection operation by the second information processing device.
[0048] As shown in FIG. 7 , when the error detection operation by the second information processing device 1 begins, the detection unit 120 first extracts a target region containing a target from at least one of the first image and the second image (hereinafter referred to as the “captured image”) (step S201). The target region is an area excluding background areas where no target is present, and may be extracted as, for example, a region combining the target's face region and neck region. Note that, when extracting the target's face region or neck region, the results of projecting and capturing a predetermined pattern may be used. For example, the results of projecting multiple sine wave patterns (e.g., four images with phases shifted by 1 / 4π) are combined to generate an amplitude image (i.e., a luminance image). Then, threshold processing is performed on the generated amplitude image to extract the target's face region or neck region. This enables robust processing of extracting the face region or neck region.
[0049] Next, the detection unit 120 perspectively projects the three-dimensional shape data generated by the generation unit 110 onto the extracted target region (step S202). The perspective projection can be realized by performing calibration in advance using a calibration object with known dimensions (for example, a checkered pattern plate, etc.) and determining the internal parameters and external parameters of the first camera 211 and the second camera 221.
[0050] If a target region is extracted from each of the first and second images, the detection unit 120 may perspectively project three-dimensional shape data onto the target region extracted from the first image and the target region extracted from the second image. On the other hand, if a target region is extracted from either the first or second image, the detection unit 120 may perspectively project three-dimensional shape data onto one of the extracted target regions. The detection unit 120 may perspectively project three-dimensional shape data of the right half of the target onto the target region extracted from the first image (i.e., the target region extracted from the image capturing the left side of the target's face). Similarly, the detection unit 120 may perspectively project three-dimensional shape data of the left half of the target onto the target region extracted from the second image (i.e., the target region extracted from the image capturing the right side of the target's face).
[0051] Next, the detection unit 120 determines whether or not the perspectively projected 3D shape data contains any portion that does not fit within the target area (step S203). If any portion does not fit within the target area (step S203: YES), the detection unit 120 detects the portion that does not fit within the target area as a portion where a depth estimation error has occurred (step S204). That is, the detection unit 120 detects the portion that protrudes from the target area when the 3D shape data is perspectively projected as a portion where a depth estimation error has occurred.
[0052] If there is no part that does not fit within the target area (step S203: NO), the processing of step S204 described above may be omitted. That is, if the perspective projection of the 3D shape data fits entirely within the target area, the detection unit 120 does not detect a depth estimation error.
[0053] (Example of Operation) Next, a specific example of the above-mentioned error detection operation will be described with reference to Fig. 8. Fig. 8 is a schematic diagram showing a specific example of the error detection operation by the second information processing device.
[0054] 8, when the error detection operation is started, the detection unit 120 extracts an object region from the captured image of the object 50. The object region may be extracted as an area showing the silhouette of the object, as shown in FIG.
[0055] Next, the detection unit 120 perspectively projects the three-dimensional shape data generated by the generation unit 110 onto the extracted target region. Note that the three-dimensional shape data that is perspectively projected here has a depth estimation error in the neck region of the target 50.
[0056] As a result of perspectively projecting the three-dimensional shape data onto the target area, part of the neck area does not fit within the target area and protrudes outside the target area. The detection unit 120 detects this protruding area as a part where a depth estimation error has occurred.
[0057] In the above-described error detection operation, clustering of the three-dimensional data (i.e., the three-dimensional point cloud) may be performed in advance. In this case, after detecting the protruding region, all of the clusters (i.e., partial regions) to which the protruding region corresponds may be detected as regions where a depth estimation error has occurred. Alternatively, in the error detection operation, regions may be divided into regions for each facial feature in advance on an image captured by a camera. In this case, after detecting the protruding region, all of the regions corresponding to the part to which the protruding region corresponds may be detected as regions where a depth estimation error has occurred. In this way, by performing a process of dividing the regions in advance, the error detection operation can be performed robustly even in cases where detection omissions occur due to noise during measurement or calibration errors.
[0058] (Technical Effects) Next, technical effects obtained by the second information processing device 1 will be described.
[0059] As described with reference to Figures 7 and 8, in the second information processing device 1, three-dimensional shape data is perspectively projected onto a target area extracted from a captured image. Then, a portion of the perspectively projected three-dimensional shape data that does not fit within the target area is detected as a portion where a depth estimation error has occurred. In this manner, it is possible to easily and accurately detect a portion of the three-dimensional shape data where a depth estimation error has occurred. As a result, it is possible to appropriately correct the depth estimation error and generate accurate three-dimensional shape data.
[0060] <Third embodiment> A third information processing device 1 will be described with reference to Figures 9 and 10. The third information processing device 1 differs in some configurations and operations from the first and second information processing devices 1 described above, but other parts may be similar to the first and second information processing devices 1. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate.
[0061] (Operation Flow) First, the operation flow of the third information processing device 1 will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the operation flow of the third information processing device. Note that in Fig. 9, the same processes as those shown in Fig. 5 are denoted by the same reference numerals.
[0062] 9 , when the operation of the third information processing device 1 starts, the first camera 211 and the second camera 221 first capture a first image and a second image of the target 50 (step S101). Then, the generation unit 110 generates three-dimensional shape data of the target 50 using the first image and the second image (step S102).
[0063] Next, the output unit 130 outputs information (hereinafter, appropriately referred to as "initial display information") for displaying the three-dimensional shape data generated by the generation unit 110 (step S301). After the output unit 130 outputs the initial display information, the first information processing device 1 enters a standby state for receiving an instruction from the user requesting detection of a depth estimation error (hereinafter, appropriately referred to as "detection instruction").
[0064] Subsequently, when a detection instruction is input from the user (step S302: YES), the detection unit 120 detects a depth estimation error occurring in the 3D shape data generated by the generation unit 110 (step S103). On the other hand, when a detection instruction is not input from the user (step S302: NO), the subsequent processing may be omitted. That is, the detection unit 120 may end the series of operations without estimating a depth estimation error in the 3D shape data.
[0065] Next, the output unit 130 outputs error display information regarding the depth estimation error detected by the detection unit 120 (step S104). After the output unit 130 outputs the error display information, the first information processing device 1 enters a standby state for receiving a correction instruction from the user.
[0066] Subsequently, when a correction instruction is input from the user (step S105: YES), the correction unit 140 corrects the portion of the 3D shape data where a depth estimation error was detected using at least one of the first image and the second image (step S106). On the other hand, when a correction instruction is not input from the user (step S105: NO), the above-mentioned step S106 may be omitted. That is, the correction unit 140 may end the series of operations without correcting the depth estimation error in the 3D shape data.
[0067] (Display Example) Next, a display example when the third information processing device 1 operates will be described with reference to Fig. 10. Fig. 10 is a plan view showing a display example by the third information processing device.
[0068] As shown in Fig. 10 , when the third information processing device 1 operates, the output unit 130 outputs initial display information according to the three-dimensional shape data generated by the generation unit 110. As a result, a display according to the initial display information is performed on the display unit. Specifically, the three-dimensional shape data generated by the generation unit 110 is displayed. In the example shown in Fig. 10 , due to a depth estimation error, three-dimensional shape data is generated in which the neck region of the target 50 protrudes forward relative to the face region. However, at this stage, the depth estimation error has not been detected.
[0069] A measurement result check button is also displayed together with the three-dimensional shape data corresponding to the initial display information. This measurement result check button is used by the user to input a detection instruction. When the user presses the measurement result check button, the third information processing device 1 accepts the input of the detection instruction.
[0070] When a detection instruction is input, the detection unit 120 performs an error detection operation and detects a portion where a depth estimation error has occurred. Then, the output unit 130 outputs error display information in accordance with the detected depth estimation error. As a result, a display in accordance with the error display information is performed on the display unit. Specifically, the 3D shape data generated by the generation unit 110 is displayed in such a way that a portion where a depth estimation error has been detected and a portion where a depth estimation error has not been detected are displayed in different manners.
[0071] Thereafter, when the user presses the recalculation button displayed together with the three-dimensional shape data, the first information processing device 1 accepts input of a correction instruction. When the correction instruction is input, the correction unit 140 corrects the three-dimensional shape data. As a result, the corrected three-dimensional shape data is displayed on the display unit.
[0072] (Technical Effects) Next, technical effects obtained by the third information processing device 1 will be described.
[0073] As described with reference to Figures 9 and 10, in the third information processing device 1, when three-dimensional shape data is generated, initial display information is output and the three-dimensional shape data is displayed. Then, in response to a detection instruction input by the user after the three-dimensional shape data is displayed, an error detection operation is performed. In this way, the generated three-dimensional shape data can be presented to the user, allowing the user to decide whether or not to detect a depth estimation error. Therefore, unnecessary error detection operations can be prevented from being performed.
[0074] <Fourth embodiment> A fourth information processing device 1 will be described with reference to Figures 11 and 12. The fourth information processing device 1 differs in some configurations and operations from the first to third information processing devices 1 described above, but other parts may be similar to the first to third information processing devices 1. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate.
[0075] (Functional Configuration) First, the functional configuration of the fourth information processing device 1 will be described with reference to Fig. 11. Fig. 11 is a block diagram showing the functional configuration of the fourth information processing device. Note that in Fig. 11, the same elements as those shown in Fig. 4 are denoted by the same reference numerals.
[0076] 11, the fourth information processing device 1 is configured to include, as components for realizing its functions, a generation unit 110, a detection unit 120, an output unit 130, a correction unit 140, and a guide unit 150. That is, the fourth information processing device 1 further includes a guide unit 150 in addition to the configuration of the first embodiment (see FIG. 4). Note that the guide unit 150 may be a processing block realized by the above-described processor 11 (see FIG. 1).
[0077] The guide unit 150 is configured to be able to determine whether or not both a face region and a neck region are included in the captured first and second images. If the first and second images do not include both a face region and a neck region, the guide unit 150 outputs guide information for capturing the first and second images that include both a face region and a neck region. Note that if the first and second images include both a face region and a neck region, the guide unit 150 does not need to output guide information.
[0078] The guide unit 150 can determine whether both a face region and a neck region are included by, for example, performing semantic segmentation on the first image and the second image and determining whether both a face region and a neck region can be extracted. Alternatively, the guide unit 150 can determine whether both a face region and a neck region are included by using a feature point detector that detects the positions of the chin and Adam's apple and determining whether both the chin and Adam's apple are detected. Such a determination may be made using a learning model constructed by machine learning. The learning model may include, for example, a neural network trained by deep learning.
[0079] The guide information may be guidance information that prompts the user to move to an appropriate position (i.e., a position where both the face region and the neck region are within the imaging range of the first camera 211 and the second camera 221). For example, the guide unit 150 may output guide information that outputs a message such as "Please move a little further away from the camera" or "Please move your face up a little more." Such a message may be displayed as text on a display or the like, or as an illustration. It may also be output as audio by a speaker or the like.
[0080] Furthermore, the guide information may be control information that changes the imaging range by controlling the first camera 211 and the second camera 221. For example, the guide unit 150 may output guide information that controls the zoom or pan of the first camera 211 and the second camera 221. Alternatively, the guide unit 150 may output guide information that changes the angle or position of the first camera 211 and the second camera 221.
[0081] (Operation Flow) Next, the operation flow of the fourth information processing device 1 will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the operation flow of the fourth information processing device. Note that in Fig. 12, the same processes as those shown in Fig. 5 are denoted by the same reference numerals.
[0082] As shown in FIG. 12, when the operation of the fourth information processing device 1 is started, first, the first camera 211 and the second camera 221 capture a first image and a second image of the target 50 (step S101).
[0083] Next, the guide unit 150 determines whether or not both the face region and the neck region are included in the first image and the second image (step S401). If both the face region and the neck region are not included in the first image and the second image (step S401: NO), the guide unit 150 outputs guide information (step S402). After the guide information is output, the process of step S101 may be executed again. That is, after the guide information has been used to place both the face region and the neck region within the imaging ranges of the first camera 211 and the second camera 221, the first image and the second image of the target 50 may be captured again.
[0084] On the other hand, if both the face region and the neck region are included in the first image and the second image (step S401: YES), the generation unit 110 generates three-dimensional shape data of the object 50 using the first image and the second image (step S102). Thereafter, the detection unit 120 detects a depth estimation error occurring in the three-dimensional shape data generated by the generation unit 110 (step S103).
[0085] Next, the output unit 130 outputs error display information regarding the depth estimation error detected by the detection unit 120 (step S104). After the output unit 130 outputs the error display information, the first information processing device 1 enters a standby state for receiving a correction instruction from the user.
[0086] Subsequently, when a correction instruction is input from the user (step S105: YES), the correction unit 140 corrects the portion of the 3D shape data where a depth estimation error was detected using at least one of the first image and the second image (step S106). On the other hand, when a correction instruction is not input from the user (step S105: NO), the above-mentioned step S106 may be omitted. That is, the correction unit 140 may end the series of operations without correcting the depth estimation error in the 3D shape data.
[0087] (Technical Effects) Next, technical effects obtained by the fourth information processing device 1 will be described.
[0088] As described with reference to Figures 11 and 12, the fourth information processing device 1 outputs guide information when the captured image does not include both the face region and the neck region of the subject 50. In this way, an image including the face region and the neck region can be captured, making it possible to appropriately generate 3D shape data of a human head including the face region and the neck region. As already described, depth estimation errors may occur in the 3D shape data of a human head around the boundary between the face region and the neck region. However, the information processing device 1 according to this embodiment can detect areas where depth estimation errors occur and appropriately correct them.
[0089] The scope of each embodiment also includes a processing method in which a program that operates the configuration of each embodiment to realize the functions of the above-described embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and the program is executed on a computer. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, each embodiment includes not only a recording medium on which the above-described program is recorded, but also the program itself.
[0090] Examples of recording media that can be used include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, magnetic tapes, non-volatile memory cards, and ROMs. Furthermore, the scope of each embodiment is not limited to programs that execute processes by themselves, but also includes programs that execute processes by operating on an OS in conjunction with other software or expansion board functions. Furthermore, the program itself may be stored on a server, and part or all of the program may be downloadable from the server to a user terminal. The program may be provided to the user in, for example, a SaaS (Software as a Service) format.
[0091] <Supplementary Notes> The above-described embodiment may be further described as in the following supplementary notes, but is not limited to the following.
[0092] (Supplementary Note 1) The information processing device described in Supplementary Note 1 is an information processing device including: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object with light from the first direction; a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object with light from the second direction; a generation means that generates three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; a detection means that detects a partial depth estimation error in the three-dimensional shape data based on an image photographed by at least one of the first camera and the second camera; an output means that outputs error display information that displays the three-dimensional shape data such that a portion in which the depth estimation error has been detected and other portions in which the depth estimation error has not been detected are displayed in different ways; and a correction means that, based on an instruction input by a user after outputting the error display information, corrects the portion in which the depth estimation error has been detected using an image photographed by at least one of the first camera and the second camera.
[0093] (Supplementary Note 2) The information processing device described in Supplementary Note 2 is the information processing device described in Supplementary Note 1, wherein the detection means perspectively projects the three-dimensional shape data onto a coordinate system of an image captured by at least one of the first camera and the second camera, determines whether the perspectively projected three-dimensional shape data falls within the area of the target in the image captured by at least one of the first camera and the second camera, and detects a portion of the perspectively projected three-dimensional shape data that does not fall within the area of the target as the depth estimation error.
[0094] (Appendix 3) The information processing device described in Appendix 3 is the information processing device described in Appendix 1 or 2, wherein the output means outputs initial display information that displays the three-dimensional shape data generated by the generation means before the detection means detects the depth estimation error, and outputs the error display information based on instructions input by a user after the initial display information is output.
[0095] (Appendix 4) The information processing device described in Appendix 4 is the information processing device described in any one of Appendices 1 to 3, wherein the generation means generates three-dimensional shape data including the face region and neck region of the target, and further includes a guide means that outputs guide information for capturing an image including both the face region and the neck region when an image captured by at least one of the first camera and the second camera does not include both the face region and the neck region.
[0096] (Supplementary Note 5) The information processing method described in Supplementary Note 5 is an information processing method for controlling an information processing device including: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object with light from the first direction; and a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object with light from the second direction, the information processing method comprising: generating three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; detecting a partial depth estimation error in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera; outputting error display information for displaying the three-dimensional shape data such that a portion in which the depth estimation error has been detected and another portion in which the depth estimation error has not been detected are displayed in different ways; and correcting the portion in which the depth estimation error has been detected using an image photographed by at least one of the first camera and the second camera based on an instruction input by a user after the error display information has been output.
[0097] (Supplementary Note 6) The recording medium described in Supplementary Note 6 is a recording medium on which a computer program is recorded to execute an information processing method for controlling an information processing device including: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object from the first direction with light; and a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object from the second direction with light, the information processing method comprising: generating three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; detecting a partial depth estimation error in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera; outputting error display information for displaying the three-dimensional shape data such that a portion in which the depth estimation error has been detected and another portion in which the depth estimation error has not been detected are displayed in different ways; and correcting the portion in which the depth estimation error has been detected using an image photographed by at least one of the first camera and the second camera based on an instruction input by a user after the error display information has been output.
[0098] (Supplementary Note 7) The computer program described in Supplementary Note 7 is an information processing method for controlling an information processing device including: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object with light from the first direction; and a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object with light from the second direction, the computer program causing the information processing method to execute the following steps: generate three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; detect partial depth estimation errors in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera; output error display information for displaying the three-dimensional shape data such that a portion in which the depth estimation error has been detected and other portions in which the depth estimation error has not been detected are displayed in different ways; and correct the portion in which the depth estimation error has been detected using an image photographed by at least one of the first camera and the second camera based on an instruction input by a user after the error display information has been output.
[0099] This disclosure may be modified as appropriate within the scope that does not contradict the gist or idea of the invention that can be read from the claims and the entire specification, and information processing devices, information processing methods, and recording media that involve such modifications are also included in the technical idea of this disclosure.
[0100] REFERENCE SIGNS LIST 1 Information processing device 11 Processor 12 RAM 13 ROM 14 Storage device 15 Input device 16 Output device 17 Data bus 21 First unit 211 First camera 212 First irradiation unit 22 Second unit 221 Second camera 222 Second irradiation unit 300 Third camera 50 Object 110 Generation unit 120 Detection unit 130 Correction unit 140 Output unit 150 Guide unit
Claims
1. An information processing device comprising: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object with light from the first direction; a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object with light from the second direction; a generation means that generates three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; a detection means that detects partial depth estimation errors in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera; an output means that outputs error display information that displays the three-dimensional shape data so that a portion where the depth estimation error is detected is displayed differently from other portions where the depth estimation error is not detected; and a correction means that corrects the portion where the depth estimation error is detected using an image photographed by at least one of the first camera and the second camera based on instructions input by a user after outputting the error display information.
2. The information processing device of claim 1, wherein the detection means perspectively projects the three-dimensional shape data onto a coordinate system of an image captured by at least one of the first camera and the second camera, determines whether the perspectively projected three-dimensional shape data fits within the area of the target in the image captured by at least one of the first camera and the second camera, and detects any portion of the perspectively projected three-dimensional shape data that does not fit within the area of the target as the depth estimation error.
3. An information processing device as described in claim 1 or 2, wherein the output means outputs initial display information for displaying the three-dimensional shape data generated by the generation means before the depth estimation error is detected by the detection means, and outputs the error display information based on instructions input by a user after the initial display information is output.
4. An information processing device as described in claim 1 or 2, wherein the generating means generates three-dimensional shape data including the face area and neck area of the subject, and further comprises a guide means that outputs guide information for capturing an image including both the face area and the neck area when an image captured by at least one of the first camera and the second camera does not include both the face area and the neck area.
5. An information processing method for controlling an information processing device comprising: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that irradiates the object with light from the first direction; and a second unit having a second camera that photographs the object from a second direction and a second illumination unit that irradiates the object with light from the second direction, the information processing method comprising: generating three-dimensional shape data of the object using an image photographed by the first camera and an image photographed by the second camera; detecting partial depth estimation errors in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera; outputting error display information that displays the three-dimensional shape data so that the portion where the depth estimation error is detected is displayed differently from other portions where the depth estimation error is not detected; and correcting the portion where the depth estimation error is detected using an image photographed by at least one of the first camera and the second camera based on instructions input by a user after outputting the error display information.
6. A recording medium having recorded thereon a computer program for controlling an information processing device comprising: a first unit having a first camera that photographs an object from a first direction and a first illumination unit that illuminates the object from the first direction; and a second unit having a second camera that photographs the object from a second direction and a second illumination unit that illuminates the object from the second direction, the information processing method comprising: generating three-dimensional shape data of the object using images photographed by the first camera and images photographed by the second camera; detecting partial depth estimation errors in the three-dimensional shape data based on images photographed by at least one of the first camera and the second camera; outputting error display information that displays the three-dimensional shape data so that the portions where the depth estimation errors are detected are displayed differently from other portions where the depth estimation errors are not detected; and correcting the portions where the depth estimation errors are detected using images photographed by at least one of the first camera and the second camera based on instructions input by a user after the error display information has been output.
Citation Information
Patent Citations
Depth estimating device, depth estimating method, and depth estimate program
JP2010251878A
Image processor, image processing method, and program
JP2012138787A
Stereoscopic discrimination image generating apparatus, stereoscopic discrimination image generating method, and stereoscopic discrimination image generating program
JP2014017684A
3D shape measurement device, 3D shape measurement method, program, and recording medium
JP2022022326A
Method and control unit for operating an autonomous vehicle
US20200156633A1