Information processing device, information processing method, and program
The information processing device addresses inefficiencies in three-dimensional model generation by estimating virtual point speed and providing guidance for optimal camera movement, enhancing image quality and reducing capture time.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
- Filing Date
- 2024-11-07
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for generating three-dimensional models from images are inefficient in terms of time and prone to capturing low-quality images due to unclear image quality deterioration during user movement.
An information processing device that estimates the speed of a virtual point between two camera captures and outputs auxiliary information to guide optimal camera movement for high-quality image capture, using processor-acquired images and memory to determine appropriate imaging speeds.
Reduces the time required for imaging and prevents the capture of low-quality images by providing users with intuitive guidance on camera movement speed.
Smart Images

Figure 2026082413000001_ABST
Abstract
Description
Technical Field
[0006] ,
[0007] , ,
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Conventionally, a method of detecting the position and orientation of an imaging device from a plurality of images (frames) and generating a three-dimensional model is known.
[0003] For example, Patent Document 1 discloses a model forming apparatus that forms a three-dimensional model of an object starting from three-dimensional model data of the object that has already been acquired. The model forming apparatus disclosed in Patent Document 1 includes a recognition unit that recognizes an unformed portion of the model of the object, and an imaging instruction information unit that obtains imaging instruction information related to imaging of the unformed portion of the model. Imaging of the object by the imaging unit is performed in accordance with the imaging instruction information obtained by the imaging instruction information unit.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] It is desired that the time required for imaging an image used for generating such a three-dimensional model can be shortened, and that a high-quality image that can generate a three-dimensional model with high accuracy can be captured.
[0006] The present disclosure provides an information processing apparatus and the like that can easily suppress the time related to imaging and can easily suppress the capture of low-quality images.
Means for Solving the Problems
[0007] An information processing device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor uses the memory to acquire a first image captured by a camera at a first time and a second image captured by the camera at a second time, estimates the speed of a virtual point assuming that the virtual point moved between the first and second time points from a first position corresponding to a first coordinate in a first camera coordinate system based on the position and orientation of the camera when the first image was captured, to a second position corresponding to a second coordinate having the same value as the first coordinate in a second camera coordinate system based on the position and orientation of the camera when the second image was captured, and outputs auxiliary information relating to the speed of the virtual point.
[0008] An information processing method according to one aspect of the present disclosure acquires a first image captured by a camera at a first time and a second image captured by the camera at a second time, estimates the speed of a virtual point assuming that the virtual point moved between the first and second time points from a first position corresponding to a first coordinate in a first camera coordinate system based on the position and orientation of the camera when the first image was captured, to a second position corresponding to a second coordinate having the same value as the first coordinate in a second camera coordinate system based on the position and orientation of the camera when the second image was captured, and outputs auxiliary information regarding the speed of the virtual point.
[0009] A program relating to one aspect of this disclosure is a program for a computer to execute the information processing method described above. [Effects of the Invention]
[0010] According to this disclosure, it is possible to provide an information processing device that can easily reduce the time required for imaging and also easily prevent the acquisition of low-quality images. [Brief explanation of the drawing]
[0011] [Figure 1] Figure 1 is a block diagram showing the configuration of an information processing system according to an embodiment. [Figure 2] Figure 2 is a diagram illustrating a first example of a virtual point according to the embodiment. [Figure 3] Figure 3 is a diagram illustrating a second example of a virtual point according to the embodiment. [Figure 4] Figure 4 is a diagram illustrating a third example of a virtual point according to the embodiment. [Figure 5] Figure 5 is a diagram illustrating a fourth example of a virtual point according to the embodiment. [Figure 6] Figure 6 is a diagram illustrating a fifth example of a virtual point according to the embodiment. [Figure 7] Figure 7 is a flowchart showing the processing procedure of the estimation device according to the embodiment. [Figure 8] Figure 8 shows a display image presented by the display device according to the embodiment. [Figure 9] Figure 9 is a block diagram showing the configuration of an information processing device according to an embodiment. [Figure 10] Figure 10 is a flowchart showing the information processing method according to the embodiment. [Modes for carrying out the invention]
[0012] (Summary of this disclosure) Traditionally, there is a technique called Visual-SLAM (also simply called VSLAM) that performs SLAM (Simultaneous Localization and Mapping) using an image (frame) as input. For example, in a process using VSLAM, tracking, local mapping, loop closing, and mapping (specifically, real-time point cloud mapping) are performed in a multi-threaded manner simultaneously with image acquisition by a depth sensor.
[0013] In tracking, the position and orientation of the imaging device are estimated based on each input image. This process involves matching feature points detected in each image (feature point matching).
[0014] In local mapping, among the input multiple images, the positions of the main images (key frames) and the map (3D map) are optimized (also called Bundle Adjustment). For example, if all the input images are used for this optimization, the processing amount will increase. Therefore, for example, only a plurality of key frames among the input multiple images are optimized.
[0015] In loop closing, when the imaging device repeatedly captures images while moving, a process for eliminating errors when returning to the same place is performed. In this process, for example, it is determined for each key frame whether it is the same place or not. For example, in this process, feature point matching between images is performed. For example, when it is determined that two images are images captured at the same place, the whole is optimized from the map (3D map) and the positions of the key frames using the errors of the two images.
[0016] In mapping, for each key frame, the overlap of depth values is checked with the previous and next key frames. In this process, depth values with no overlap are excluded as noise. Also, in this process, the information (point cloud information) of the 3D points for all key frames is integrated. As a result, for example, the imaged location is represented by a 3D map, or in other words, a 3D model (for example, a 3D mesh).
[0017] In this way, for example, VSLAM technology is used to generate (create) a 3D model.
[0018] Here, when the user (imager) uses an application that utilizes VSLAM, for example, no matter what the state is, such as the speed at which the user moves while imaging, it is not clear when the quality of the image, that is, the quality of the 3D model generated from the image, will deteriorate. Therefore, there are problems such as blurred images (or equivalently, blurry images), that is, low-quality images, or the time taken for imaging becomes long because the user excessively carefully captures images.
[0019] Therefore, the purpose of this disclosure is to provide an information processing device, etc., that can easily reduce the time required for imaging and also easily prevent the acquisition of low-quality images.
[0020] Example 1 is an information processing device comprising a processor and memory, wherein the processor uses the memory to acquire a first image captured by a camera at a first time and a second image captured by the camera at a second time, estimates the speed of a virtual point assuming that the virtual point moved between the first and second time points from a first position corresponding to a first coordinate in a first camera coordinate system based on the position and orientation of the camera when the first image was captured, to a second position corresponding to a second coordinate having the same value as the first coordinate in a second camera coordinate system based on the position and orientation of the camera when the second image was captured, and outputs auxiliary information regarding the speed of the virtual point.
[0021] When images are repeatedly and continuously captured, it takes considerable effort for the user to check whether each image is blurry or not. Therefore, the information processing device estimates the speed of a set virtual point. If the estimated speed of the virtual point is too fast, there is a high probability that blurry images will be captured instead of sharp ones. On the other hand, if the user moves excessively slowly to avoid the virtual point moving too fast, the time required for capturing images will be unnecessarily long if many images need to be taken. Therefore, according to the information processing device, by checking auxiliary information, the user capturing images using the camera can move at an appropriate speed and more easily capture images that are less prone to blurring. Thus, according to the information processing device, it is possible to reduce the time required for capturing images and to reduce the capture of low-quality images.
[0022] Example 2 is an information processing device as described in Example 1, wherein distance information indicating the distance from the camera to a predetermined object is acquired, and the first coordinates are determined based on the distance information.
[0023] According to this, for example, when an image containing a specific object is captured, setting a virtual point at the position of the specific object makes it easier to suppress blurring of the specific object in the image.
[0024] Example 3 is an information processing device as described in Example 1, wherein the first coordinates may be determined based on depth information indicating the depth corresponding to a pixel included in the first image.
[0025] This makes it easier to pinpoint the specific location of an object in an image. For example, when an image containing such an object is captured, setting a virtual point at the object's location makes it easier to suppress blurring of the object in the image.
[0026] Example 4 is an information processing device as described in Example 1, wherein the first coordinates may be determined based on depth information indicating the depth corresponding to a pixel included in the first image, which is calculated using segmentation.
[0027] According to this, the first coordinate can be determined based on the depth corresponding to a specific pixel contained in the first image.
[0028] Example 5 is an information processing device as described in Example 1, in which the first coordinates may be determined based on feature point information indicating the position of feature points obtained by VSLAM (Visual Simultaneous Localization and Mapping) using the first image.
[0029] According to this, for example, by setting virtual points at the locations of feature points, it becomes easier to suppress the blurring of objects containing feature points that appear in the image.
[0030] Example 6 is an information processing device as described in Example 1, wherein the first coordinates may be determined based on matching information relating to the matching of feature points obtained by VSLAM using the first image.
[0031] According to this, for example, by setting virtual points at the locations of common feature points present in both the first and second images, it becomes easier to suppress blurring of objects containing feature points that appear in both images. Therefore, for example, it becomes easier to suppress the decrease in accuracy of the three-dimensional model generated by VSLAM using the first and second images.
[0032] Example 7 is an information processing device described in any of Examples 1 to 6, wherein a first determination is made as to whether the speed of the virtual point is equal to or greater than a predetermined first threshold, and in the output of the auxiliary information, the auxiliary information based on the determination result of the first determination may be output.
[0033] According to this, users can more easily decide whether or not to move the camera slowly to take an image by checking the auxiliary information.
[0034] Example 8 is an information processing device as described in Example 7, further comprising a determination of whether the speed of the virtual point is less than or equal to a predetermined second threshold that is lower than the predetermined first threshold, and in the output of the auxiliary information, the auxiliary information based on the determination results of the first and second determinations may be output.
[0035] According to this, users can more easily decide whether or not to move the camera quickly to take an image by checking the auxiliary information.
[0036] Example 9 is an information processing device described in Example 7 or Example 8, wherein the predetermined first threshold value may be calculated based on the shutter speed of the camera.
[0037] The degree to which an image blurs varies depending on the shutter speed. Based on this, an appropriate predetermined first threshold can be set to make it easier for the user to capture an image that is not blurred.
[0038] Example 10 is an information processing device as described in Example 9, wherein the predetermined first threshold may be calculated such that the faster the shutter speed of the camera, the higher the predetermined first threshold.
[0039] The faster the shutter speed, the less likely the image is to blur, even if the virtual point is moving quickly. This allows for setting an appropriate predetermined first threshold to make it easier for the user to capture a sharp image without unnecessarily slowing down the user's movements.
[0040] Example 11 is an information processing device described in any of Examples 7 to 10, wherein in the first determination, if it is determined that the speed of the virtual point is equal to or greater than the predetermined first threshold, the output of the auxiliary information may include a first warning information to reduce the speed at which the user performing the shooting using the camera moves.
[0041] This makes it easy for users to intuitively understand how fast they should move to take images.
[0042] Example 12 is an information processing device as described in Example 8, wherein in the first determination, if it is determined that the speed of the virtual point is equal to or greater than the predetermined first threshold, the output of the auxiliary information may include auxiliary information containing first warning information to reduce the speed at which the user taking pictures with the camera moves, and in the second determination, if it is determined that the speed of the virtual point is equal to or less than the predetermined second threshold, the output of the auxiliary information may include auxiliary information containing second warning information to increase the speed at which the user taking pictures with the camera moves.
[0043] This makes it easy for users to intuitively understand how fast they should move to take images.
[0044] Example 13 is an information processing device described in any of Examples 1 to 12, wherein the auxiliary information may include speed information indicating the speed of the virtual point.
[0045] This makes it easy for users to intuitively understand how fast they should move to take images.
[0046] Example 14 is an information processing device according to any of Examples 1 to 13, wherein the auxiliary information is saturation information indicating the saturation corresponding to the speed of the virtual point, and may also include saturation information indicating the saturation of the image when the display device displays an image containing the auxiliary information.
[0047] This makes it easy for users to intuitively understand how fast they should move to take images.
[0048] Example 15 is an information processing method that acquires a first image captured by a camera at a first time step and a second image captured by the camera at a second time step, estimates the speed of a virtual point assuming that the virtual point moved between the first and second time steps from a first position corresponding to a first coordinate in a first camera coordinate system based on the position and orientation of the camera when the first image was captured, to a second position corresponding to a second coordinate having the same value as the first coordinate in a second camera coordinate system based on the position and orientation of the camera when the second image was captured, and outputs auxiliary information regarding the speed of the virtual point.
[0049] According to this, the information processing device described in Example 1 will have the same effect.
[0050] Example 16 is a program for a computer to execute the information processing method described in Example 15.
[0051] According to this, the information processing device described in Example 1 will have the same effect.
[0052] The embodiments of this disclosure will be described in detail below with reference to the drawings. The embodiments described below are all specific examples of this disclosure. Therefore, the numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Accordingly, any components in the following embodiments that are not described in the independent claims of this disclosure will be described as optional components.
[0053] Furthermore, each figure is a schematic diagram and not necessarily a strictly accurate representation. Therefore, for example, the scale may not necessarily match in each figure. Also, in each figure, substantially identical components are given the same reference numerals, and redundant explanations are omitted or simplified.
[0054] Furthermore, in this specification, ordinal numbers such as "first," "second," etc., do not indicate the number or order of components unless otherwise specified, but are used to avoid confusion and to distinguish similar components.
[0055] Furthermore, in this specification, when we describe, for example, a value greater than or equal to a threshold and a value less than a threshold in comparison, it means that the value is distinguished by that threshold, and may mean that the value is greater than the threshold and less than or equal to the threshold, respectively.
[0056] (Embodiment) [composition] First, the configuration of the information processing system according to the embodiment will be described.
[0057] Figure 1 is a block diagram showing the configuration of the information processing system 400 according to the embodiment.
[0058] The information processing system 400 is a system that generates a three-dimensional model from multiple images (frames). The information processing system 400 comprises an estimation device 100, a presentation device 200, and an imaging device 300.
[0059] The estimation device 100 is a computer that estimates the position and orientation of the imaging device 300 (hereinafter also simply referred to as the position and orientation of the imaging device 300) using multiple images acquired from the imaging device 300, such as a camera. The estimation device 100 also generates a three-dimensional model using the multiple images and the position and orientation of the imaging device 300. For example, the estimation device 100 generates a three-dimensional map (hereinafter also simply referred to as a map) as a three-dimensional model. The estimation device 100 outputs map information indicating the generated map to a display device 200, such as a display, thereby displaying the map indicated by the map information on the display device 200.
[0060] The estimation device 100 is implemented, for example, by a communication interface for communicating with the presentation device 200 and the imaging device 300, a non-volatile memory where the program is stored, a volatile memory which is a temporary storage area for executing the program, input / output ports for sending and receiving signals, and a processor for executing the program. The communication interface may be implemented, for example, by an antenna and a wireless communication circuit to enable wireless communication, or by a connector to which a communication line is connected to enable wired communication.
[0061] The estimation device 100 includes a position and orientation estimation unit 110, a virtual point determination unit 120, an image quality calculation unit 130, and a storage unit 140.
[0062] The position and orientation estimation unit 110 is a processing unit that estimates (calculates) the position and orientation of the imaging device 300. For example, the position and orientation estimation unit 110 uses a plurality of images captured by the imaging device 300 to generate position and orientation information that indicates the position and orientation of the imaging device 300 at the time each image was captured. Here, the orientation of the imaging device 300 refers to at least one of the imaging direction of the imaging device 300 and the tilt of the imaging device 300. The imaging direction of the imaging device 300 is the direction of the optical axis of the imaging device 300. The tilt of the imaging device 300 is the rotation angle of the imaging device 300 around the optical axis from the reference orientation.
[0063] Furthermore, the position and orientation estimation unit 110 generates map information that shows a three-dimensional model (three-dimensional map) composed of a three-dimensional point cloud indicating the three-dimensional position, along with the position and orientation information.
[0064] First, the position and orientation estimation unit 110 acquires multiple images from the imaging device 300. Specifically, the position and orientation estimation unit 110 acquires multiple images captured by the imaging device 300 at different times. For example, the position and orientation estimation unit 110 acquires a first image captured by the imaging device 300 at a first time, and a second image captured by the imaging device 300 at a second time different from the first time. The position and orientation estimation unit 110 then performs feature point matching on the multiple images captured by the imaging device 300. That is, the position and orientation estimation unit 110 extracts feature points from each of the multiple images and extracts sets of similar points (corresponding points) that are similar across the multiple images from the extracted set of feature points. Next, the position and orientation estimation unit 110 uses the extracted sets of similar points to estimate the position and orientation of the imaging device 300 for each image.
[0065] Furthermore, the map information indicates multiple map points, each of which is a three-dimensional point (specifically, a point indicating the position of a three-dimensional point). The position and orientation estimation unit 110 generates the map information by performing triangulation using the results of the feature point matching process and the position and orientation of the imaging device 300.
[0066] The generated map information may include, in addition to information indicating the position of each map point, information indicating the visibility of each map point, information indicating the color of each map point, and information indicating the surface shape around each map point (for example, information indicating the normal). The map information may also be corrected to high-precision information by undergoing optimization processing together with the position and orientation of the imaging device 300.
[0067] The position and orientation estimation unit 110 generates, for example, a mesh model represented by a plane with map points as vertices, as a three-dimensional model. The three-dimensional model is not limited to a mesh model; it may be any three-dimensional model. In addition, for example, point cloud data obtained from LiDAR or depth images obtained from a distance sensor may also be used to generate the three-dimensional model.
[0068] The position and orientation estimation method used by the position and orientation estimation unit 110 to estimate the position and orientation of the imaging device 300 is, for example, VSLAM, but is not particularly limited. Position information may be obtained from an external device, such as a server device (not shown), from the imaging device 300, or from the user via an operating device such as a touch panel, mouse, or keyboard. The operating device may be a microphone that accepts voice commands.
[0069] The virtual point determination unit 120 is a processing unit that determines (sets) a virtual point. Specifically, the virtual point determination unit 120 determines the coordinates (also called virtual point coordinates) in the camera coordinate system based on the position and orientation of the imaging device 300. For example, the virtual point determination unit 120 determines the virtual point coordinates in the camera coordinate system when the position of the imaging device 300 is taken as the origin. It is assumed that a virtual point is located at the determined coordinates.
[0070] The virtual point coordinates determined by the virtual point determination unit 120 can be arbitrarily determined and are not particularly limited. For example, the virtual point coordinates may be determined from feature points in the image, or they may be set by the user. For example, the virtual point coordinates may be determined based on information indicating the distance between the user (imaging device 300) and the object to be imaged, which is set in advance. For example, the distance may be set to sensor effective range = 1m. In this case, the virtual point coordinates may be set to, for example, (0, 0, 1). Alternatively, if the image captured by the imaging device 300 contains depth information, the depth (depth position) of all pixels in the image may be determined as the virtual point coordinates. Alternatively, if the object to be imaged is predetermined, the pixels corresponding to the object to be imaged in the image may be identified by segmentation, and the depth of the identified pixels may be determined as the virtual point coordinates. Furthermore, the virtual point coordinates may be determined using the position and orientation of the imaging device 300 without using depth information. In addition, the position of the map point corresponding to the matched feature point pixel may be calculated as the virtual point coordinates. A specific example of how to determine the coordinates of a virtual point will be discussed later.
[0071] The image quality calculation unit 130 is a processing unit that calculates (estimates) the quality of the image captured by the imaging device 300. Specifically, the image quality calculation unit 130 determines whether the image captured by the imaging device 300 is suitable for generating a three-dimensional map, that is, whether it is a high-quality image.
[0072] The image quality calculation unit 130 estimates the speed of a virtual point, assuming that the virtual point moved between the first and second time points, for example, from a first position corresponding to a first coordinate in a first camera coordinate system based on the position and orientation of the imaging device 300 when the first image captured by the imaging device 300 at the first time point was taken, to a second position in a second camera coordinate system based on the position and orientation of the imaging device 300 when the second image captured by the imaging device 300 at the second time point was taken. The image quality calculation unit 130 then determines the quality of the image (for example, the second image) based on the estimated speed of the virtual point.
[0073] The second time point is, for example, a time point later than the first time point. If the imaging device 300 performs continuous imaging, the second time point is, for example, the time when an image is captured immediately after the first time point.
[0074] Furthermore, the second time point does not have to be the time when the image was captured immediately after the first time point. The first and second time points may be arbitrarily determined in advance. In other words, two images captured at arbitrary timings may be used to calculate the speed of the virtual point.
[0075] The first camera coordinate system and the second camera coordinate system are, for example, both three-axis Cartesian coordinate systems. In both the first and second camera coordinate systems, the position of the imaging device 300 is set as the origin. In other words, if the position of the imaging device 300 changes, the coordinates (values) of the origins transformed from the first camera coordinate system and the second camera coordinate system will be different, even if the same coordinate system (for example, the world coordinate system) is transformed.
[0076] The first and second coordinates are both specific examples of the virtual point coordinates described above. Here, the first and second coordinates are coordinates with the same value. Specifically, if the first coordinate is determined to be (A, B, C) (for example, A, B, and C are all real numbers), then the second coordinate is also (A, B, C). In other words, if the position and orientation of the imaging device 300 change, the first and second coordinates may become different coordinates (values) when they are transformed into the same coordinate system. Therefore, the image quality calculation unit 130 transforms the first and second coordinates into the same coordinate system and calculates the speed of the virtual point assuming that the virtual point moved from the transformed first coordinate to the transformed second coordinate.
[0077] Note that the same coordinate system mentioned above may be any coordinate system. For example, the same coordinate system mentioned above may be the world coordinate system, the first camera coordinate system, or the second camera coordinate system. For example, the image quality calculation unit 130 may convert the second coordinate to the world coordinate system and then to the first coordinate system. In this case, for example, the image quality calculation unit 130 calculates the speed of the virtual point assuming that the virtual point moved from the first coordinate to the second coordinate which has been converted to the first camera coordinate system.
[0078] The speed of a virtual point can be calculated, for example, by the following equation (1).
[0079]
number
[0080] Note that vp is the coordinate of a virtual point in the second camera coordinate system. Also, Tcwt-1 is the transformation formula (determinant) for converting from the world coordinate system to the first camera coordinate system. Also, Tcwt is the transformation formula (determinant) for converting from the second camera coordinate system to the world coordinate system. Δt is the elapsed time from the first time point to the second time point. Also, length is a function that calculates the distance between two points.
[0081] The vector quantity (vs) calculated by equation (1) above is the velocity of the virtual point. For example, the speed of the virtual point can be calculated by converting the vs calculated by equation (1) above into a scalar quantity.
[0082] The number of virtual points (specifically, virtual point coordinates) determined may be one or multiple. For example, if multiple virtual points (specifically, virtual point coordinates) are determined, the average value (average coordinate) of the multiple virtual point coordinates may be adopted as one virtual point coordinate (average virtual point coordinate). Alternatively, the speed of all virtual points may be calculated, and the maximum value of the calculated virtual point speeds may be adopted as the speed of one virtual point. Alternatively, the speed of all virtual points may be calculated, and the average value of the calculated virtual point speeds may be adopted as the speed of one virtual point.
[0083] Furthermore, the image quality calculation unit 130 outputs auxiliary information regarding the calculated speed of the virtual points. For example, the image quality calculation unit 130 outputs information regarding the image quality based on the speed of the virtual points as auxiliary information. For example, the image quality calculation unit 130 makes a determination (first determination) as to whether the speed of the virtual points is above a predetermined threshold (first threshold), and outputs auxiliary information based on the determination result of this determination. For example, the image quality calculation unit 130 calculates this determination result as the image quality. The predetermined first threshold indicates, for example, the limit of the speed at which the imaging device 300 can stably capture images without blurring due to changes in the position and orientation of the imaging device 300 (the speed of changes in the position and orientation of the imaging device 300, in other words, the speed of the user's actions). For example, the image quality calculation unit 130 determines that the image quality is poor if the speed of the virtual points is above the predetermined first threshold, and determines that the image quality is good if the speed of the virtual points is below the predetermined first threshold. The image quality calculation unit 130 outputs, for example, auxiliary information including the judgment result to the presentation device 200. The image quality calculation unit 130 then uses the presentation device 200 to present this judgment result (i.e., auxiliary information) to the user.
[0084] The image quality may be calculated using either the speed of the virtual point or the velocity of the virtual point. In this embodiment, the explanation uses the speed of the virtual point, but the speed of the virtual point may be read as the velocity of the virtual point.
[0085] Furthermore, the auxiliary information may include information based on the above determination result. For example, the image quality calculation unit 130 outputs auxiliary information that includes information about the user's actions during image acquisition, based on the determination result. For example, if the speed of the virtual point is greater than or equal to a predetermined first threshold, the image quality calculation unit 130 outputs auxiliary information that includes information indicating that the user should reduce their movement speed (also called first warning information), and if the speed of the virtual point is less than the predetermined first threshold, it outputs auxiliary information that includes information indicating that there is no problem with the user's movement speed.
[0086] Furthermore, a predetermined second threshold, which is lower than the predetermined first threshold and different from the predetermined first threshold, may be set separately. The information processing device 10 (specifically, the image quality calculation unit 130) may perform a determination (second determination) as to whether the speed of the virtual point is less than or equal to a predetermined threshold (second threshold), and in outputting auxiliary information, it may output auxiliary information based on the determination result of this determination. For example, if the speed of the virtual point is greater than the predetermined second threshold, the image quality calculation unit 130 may output auxiliary information including information indicating that there is no problem with the speed at which the user is moving, and if it is less than or equal to the predetermined second threshold, it may output auxiliary information including information indicating an instruction to the user to increase their speed (also called second warning information).
[0087] In determining whether a value is above or below a threshold, it is also acceptable to determine whether the value is greater than or below a certain threshold, or whether it is less than or equal to a different threshold.
[0088] Furthermore, the predetermined thresholds (specifically, a predetermined first threshold and a predetermined second threshold) may be arbitrarily determined in advance and are not particularly limited. Also, for example, the image quality calculation unit 130 may calculate the predetermined thresholds based on predetermined information.
[0089] The image quality calculation unit 130 may calculate a predetermined first threshold based, for example, on the shutter speed of the imaging device 300 (specifically, the shutter speed at which the imaging device 300 captures an image). For example, the image quality calculation unit 130 calculates a predetermined first threshold such that the faster the shutter speed of the imaging device 300, the higher the predetermined first threshold. For example, the image quality calculation unit 130 sets the predetermined first threshold to the third threshold when the shutter speed of the imaging device 300 is the first speed, and sets the predetermined first threshold to the fourth threshold, which is higher than the third threshold, when the shutter speed of the imaging device 300 is the second speed, which is faster than the first speed. The predetermined second threshold may also be calculated based on the shutter speed of the imaging device 300 in the same way as the predetermined first threshold. For example, the image quality calculation unit 130 calculates a predetermined second threshold such that the faster the shutter speed of the imaging device 300, the higher the predetermined second threshold. For example, the image quality calculation unit 130 sets a predetermined second threshold to the fifth threshold when the shutter speed of the imaging device 300 is the third speed, and sets a predetermined second threshold to the sixth threshold, which is higher than the fifth threshold, when the shutter speed of the imaging device 300 is the fourth speed, which is faster than the third speed.
[0090] The shutter speed information may be obtained by the estimation device 100 in any way. For example, the shutter speed information may be obtained from the imaging device 300, or from the user via the operating device.
[0091] Alternatively, for example, a predetermined value may be set in advance, and this predetermined value may be weighted using the shutter speed value, with the weighted predetermined value being adopted as a predetermined threshold.
[0092] Furthermore, for example, if the image quality calculation unit 130 determines that the speed of the virtual point is greater than or equal to a predetermined first threshold, it outputs auxiliary information including first warning information to reduce the speed at which the user performing the imaging using the imaging device 300 moves. In other words, the auxiliary information may include first warning information if it is determined that the speed of the virtual point is greater than or equal to a predetermined first threshold. On the other hand, for example, if the image quality calculation unit 130 determines that the speed of the virtual point is less than a predetermined first threshold, it may output auxiliary information that does not include first warning information. Also, for example, if the image quality calculation unit 130 determines that the speed of the virtual point is less than a predetermined first threshold, it may output auxiliary information that indicates that the speed at which the user performing the imaging using the imaging device 300 moves is appropriate.
[0093] Furthermore, for example, if the image quality calculation unit 130 determines that the speed of the virtual point is less than or equal to a predetermined second threshold, it outputs auxiliary information including a second warning to encourage the user performing the imaging using the imaging device 300 to increase the speed at which they move. In other words, the auxiliary information may include the second warning if it is determined that the speed of the virtual point is less than or equal to a predetermined second threshold. On the other hand, for example, if the image quality calculation unit 130 determines that the speed of the virtual point is greater than a predetermined second threshold, it may output auxiliary information that does not include the second warning. Also, for example, if the image quality calculation unit 130 determines that the speed of the virtual point is greater than a predetermined second threshold, it may output auxiliary information that indicates that the speed at which the user performing the imaging using the imaging device 300 moves is appropriate.
[0094] As described above, for example, if the image quality calculation unit 130 determines that the speed of the virtual point is greater than or equal to a predetermined first threshold, it outputs auxiliary information including first warning information; if it determines that the speed of the virtual point is less than a predetermined first threshold and greater than a predetermined second threshold, it outputs auxiliary information including information indicating that the user's movement speed is appropriate; and if it determines that the speed of the virtual point is less than or equal to a predetermined second threshold, it outputs auxiliary information including second warning information.
[0095] Furthermore, for example, the image quality calculation unit 130 outputs auxiliary information that includes speed information indicating the speed of the virtual point. In other words, the auxiliary information may include speed information.
[0096] Warning information (first warning information and second warning information) and speed information are, for example, information to be presented to the user, such as by a presentation device 200.
[0097] Furthermore, for example, the image quality calculation unit 130 may output auxiliary information that includes saturation information indicating saturation according to the speed of the virtual point. For example, the auxiliary information is saturation information indicating saturation according to the speed of the virtual point, and may include saturation information indicating the saturation of the image when the presentation device 200 displays an image (for example, the presentation image 600 described later) that includes the auxiliary information (specifically, an image based on the auxiliary information).
[0098] Saturation information is information that indicates, for example, an instruction to increase the saturation of the normal image display (for example, making the display redder) the faster the virtual point moves. The saturation indicated by the saturation information may be arbitrarily determined. Saturation information may also be information that indicates, for example, an instruction to decrease the saturation of the normal image display (for example, making the display darker (blacker)) the faster the virtual point moves. Alternatively, for example, saturation information may be information that indicates, for example, a first saturation when the speed of the virtual point is less than a predetermined first threshold, and a second saturation with higher saturation than the first saturation when the speed of the virtual point is equal to or greater than the predetermined first threshold.
[0099] The auxiliary information may include color information indicating brightness and / or hue corresponding to the speed of the virtual point, and may also include saturation information indicating the saturation of the image when the display device 200 displays an image containing the auxiliary information.
[0100] Each processing unit, such as the position and orientation estimation unit 110, the virtual point determination unit 120, and the image quality calculation unit 130, is implemented by, for example, a processor such as a CPU (Central Processing Unit) and a memory in which a control program executed by the processor is stored.
[0101] The memory unit 140 is a storage device that stores various types of information. For example, the memory unit 140 stores multiple images acquired from the imaging device 300, as well as sensor information, such as shutter speed information, which will be described later. In addition, for example, the memory unit 140 stores map information generated by the estimation device 100, and information indicating virtual point coordinates. The memory unit 140 is implemented by, for example, an SSD (Solid State Drive) or an HDD (Hard Disk Drive).
[0102] The display device 200 is a display device that presents various types of information. For example, the display device 200 displays images generated by the imaging device 300. Also, for example, the display device 200 presents auxiliary information.
[0103] This is implemented by a display device 200, for example, a display device such as a screen. The display device 200 may also include a speaker for audio output and a vibrator for transmitting vibrations to the user as means of presenting information to the user. For example, if the image quality calculation unit 130 determines that the speed of the virtual point is greater than or equal to a predetermined first threshold, it may vibrate the display device 200 and / or the imaging device 300, but if it determines that the speed of the virtual point is less than the predetermined first threshold, it may not vibrate the display device 200 and / or the imaging device 300.
[0104] The imaging device 300 is a camera that generates multiple images by imaging a subject and outputs the generated multiple images to the estimation device 100. For example, the imaging device 300 outputs multiple images obtained by imaging a subject (object) from different viewpoints (imaging position and imaging direction) to the estimation device 100. For example, multiple images are captured when a user using the imaging device 300 moves while taking images.
[0105] The imaging device 300 includes, for example, a storage device such as an HDD or SSD, a control device which consists of a memory and a processor such as a CPU that executes a control program stored in the memory, an image sensor, and an optical system which consists of a lens for controlling the light input to the image sensor. The image sensor is, for example, a sensor capable of detecting R (Red), G (Green), and B (Blue). The image sensor is implemented by, for example, an image sensor such as a CCD (Charge Coupled Device) or CMOS (Complementary Metal Oxide Semiconductor). In other words, in this embodiment, the imaging device 300 captures (generates) an RGB image.
[0106] The imaging device 300 may also include various sensors, such as a distance sensor (depth sensor), a ToF (Time of Flight) sensor or a laser sensor such as LiDAR (Light Detection and Ranging), or an IMU (Inertial Measurement Unit). The IMU is a sensor that includes, for example, at least one of an acceleration sensor, a rotational angular acceleration sensor, and a gyroscope sensor.
[0107] The imaging device 300 outputs sensor information, including multiple images, to the estimation device 100, for example, via a communication interface (not shown) provided within the device.
[0108] The sensor information may include, for example, multiple images captured by the imaging device 300, shutter speed information indicating the shutter speed of the imaging device 300, depth information (e.g., depth image obtained by a distance sensor), point cloud data obtained by a laser sensor, and inertial information (e.g., acceleration information obtained by an IMU).
[0109] The estimation device 100, the presentation device 200, and the imaging device 300 may be implemented as a single device. Furthermore, the multiple processing units included in each device may be distributed across multiple devices.
[0110] For example, in a use case where a user captures an image of a subject while it is moving using a device they own (such as a smartphone, tablet, or personal computer), the estimation device 100, the presentation device 200, and the imaging device 300 may be included in that device. In another use case, the imaging device 300 may be included in a mobile device operated by the user (e.g., a vehicle, robot, or drone), and the estimation device 100 and presentation device 200 may be included in the user's device. Furthermore, in either case, some of the processing units included in each device may be included in other devices connected to the user's device via a communication network, such as a server. For example, some or all of the processing units included in the estimation device 100 may be included in the server.
[0111] Furthermore, the transmission and reception of information (data) between the devices shown in Figure 1 may be carried out by any method, such as wired communication or wireless communication. These communications may also be carried out directly between devices or indirectly via other communication devices or servers. Moreover, these transmissions and receptions may also be the transfer of data within a single device.
[0112] Furthermore, for example, the information processing system 400 may include an operating device for user operation. The estimation device 100 may acquire the user's operation results (input results) from the operating device. Also, the presentation device 200 and the operating device may be implemented as an integrated unit, such as a touch panel display.
[0113] [Specific example] Next, we will explain virtual points in detail. Figures 2 to 6 show an example in which the imaging device 300 takes an image at time t-1, moves to position t, and takes another image at time t. Furthermore, in the following specific examples, the camera coordinate system based on the position and orientation of the imaging device 300 when the image was taken at time t-1 will be referred to as the first camera coordinate system, and the camera coordinate system based on the position and orientation of the imaging device 300 when the image was taken at time t will be referred to as the second camera coordinate system. In addition, it will be explained that the coordinates of the virtual points in both the first and second camera coordinate systems are (A, B, C). Furthermore, the coordinates of the virtual points in each figure are shown in the world coordinate system.
[0114] Figure 2 is a diagram illustrating a first example of a virtual point according to the embodiment. Figure 2(a) schematically shows the position and orientation of the imaging device 300 at time t-1 in the world coordinate system, and Figure 2(b) schematically shows the position and orientation of the imaging device 300 at time t in the world coordinate system. In this example, after the imaging device 300 captures image 500 at time t-1, both the position and orientation of the imaging device 300 are changed, and the imaging device 300 captures image 501 at time t.
[0115] Furthermore, in this example, the coordinates (A1, B1, C1) in the world coordinate system corresponding to the coordinates (A, B, C) in the first camera coordinate system, based on the position and orientation of the imaging device 300 when image 500 was captured, coincide with the coordinates (A1, B1, C1) in the world coordinate system corresponding to the coordinates (A, B, C) in the second camera coordinate system, based on the position and orientation of the imaging device 300 when image 501 was captured. In other words, in this example, it is estimated that the virtual point did not move from time t-1 to time t. To put it another way, in this example, the speed of the virtual point is 0.
[0116] In this way, the speed of the virtual point is calculated from the position and orientation of the imaging device 300 before and after the time series and the virtual point coordinates. For example, even if the imaging device 300 moves between time t-1 and time t, if the virtual point does not change from its reference position, the speed of the virtual point will be 0, and the image corresponding to the virtual point coordinates will not be blurred. In other words, if the virtual point is not moving, it is thought that there is no blurring (or significantly less blurring) in the image of the virtual point due to the movement of the imaging device 300, and therefore image 501 is estimated to be a high-quality image.
[0117] Figure 3 is a diagram illustrating a second example of a virtual point according to the embodiment. Figure 3(a) schematically shows the position and orientation of the imaging device 300 at time t-1 in the world coordinate system, and Figure 3(b) schematically shows the position and orientation of the imaging device 300 at time t in the world coordinate system. In this example, after the imaging device 300 captures image 502 at time t-1, only the position of the imaging device 300 is changed, and the imaging device 300 captures image 503 at time t.
[0118] Furthermore, in this example, the coordinates (A1, B1, C1) in the world coordinate system corresponding to the coordinates (A, B, C) in the first camera coordinate system, which is based on the position and orientation of the imaging device 300 when image 502 was captured, do not match the coordinates (A2, B2, C2) in the world coordinate system corresponding to the coordinates (A, B, C) in the second camera coordinate system, which is based on the position and orientation of the imaging device 300 when image 503 was captured. In other words, in this example, it is estimated that the virtual point moved from time t-1 to time t. For example, if the coordinates of (A2, B2, C2) in the first coordinate system are (A3, B3, C3), then the velocity of the virtual point can be calculated, for example, as (A3-A1, B3-B1, C3-C1). For example, the speed of the virtual point can be calculated from the calculated velocity of the virtual point.
[0119] In this way, when a virtual point is moving, its speed is calculated, and the image quality is calculated based on the calculated speed of the virtual point.
[0120] Figure 4 is a diagram illustrating a third example of a virtual point according to the embodiment. Figure 4(a) schematically shows the position and orientation of the imaging device 300 at time t-1 in the world coordinate system, and Figure 4(b) schematically shows the position and orientation of the imaging device 300 at time t in the world coordinate system. In this example, after the imaging device 300 captures image 504 at time t-1, both the position and orientation of the imaging device 300 are changed, and the imaging device 300 captures image 505 at time t.
[0121] Furthermore, in this example, the coordinates (A1, B1, C1) in the world coordinate system corresponding to the coordinates (A, B, C) in the first camera coordinate system, which is based on the position and orientation of the imaging device 300 when image 504 was captured, do not match the coordinates (A2, B2, C2) in the world coordinate system corresponding to the coordinates (A, B, C) in the second camera coordinate system, which is based on the position and orientation of the imaging device 300 when image 505 was captured. In other words, in this example, it is estimated that the virtual point moved from time t-1 to time t. For example, if the coordinates of (A2, B2, C2) in the first coordinate system are (A3, B3, C3), then the velocity of the virtual point can be calculated, for example, as (A3-A1, B3-B1, C3-C1). For example, the speed of the virtual point can be calculated from the calculated velocity of the virtual point.
[0122] Thus, the virtual point may also move when both the position and orientation of the imaging device 300 are changed. For example, if the rotation element of the imaging device 300 increases, the speed of the virtual point may also increase. In such cases, the speed of the virtual point is calculated, and the image quality is calculated based on the calculated speed of the virtual point.
[0123] Figure 5 is a diagram illustrating a fourth example of a virtual point according to the embodiment. Figure 5(a) schematically shows the position and orientation of the imaging device 300 at time t-2 in the world coordinate system, Figure 5(b) schematically shows the position and orientation of the imaging device 300 at time t-1 in the world coordinate system, and Figure 5(c) schematically shows the position and orientation of the imaging device 300 at time t in the world coordinate system. In this example, the imaging device 300 captures image 506 at time t-2, then moves to the position at time t-1, captures image 507 at time t-1, then moves to the position at time t, and captures image 508 at time t.
[0124] The estimation device 100 (for example, the position and orientation estimation unit 110) generates a three-dimensional map using, for example, multiple images captured before time t-1. In generating the three-dimensional map, for example, the estimation device 100 extracts feature points from each of the multiple images and extracts sets of similar points (corresponding points) that are similar between the multiple images from among the extracted feature points. Next, the position and orientation estimation unit 110 generates map points by performing triangulation using the extracted sets of similar points and the position and orientation of the imaging device 300. The virtual point determination unit 120 determines the coordinates of the generated map points as virtual point coordinates. For example, the virtual point determination unit 120 determines that a target feature point in image 506 and a target feature point in image 507 are similar points, generates a target map point from these target feature points, and determines the coordinates of the generated target map point as virtual point coordinates.
[0125] Thus, in determining the virtual point coordinates, map information generated using VSLAM may be used. For example, if the imaging device 300 moves between time t-2 and time t-1, map information may be generated from triangulation. For example, the position of the map point indicated by the generated map information may be adopted as the virtual point coordinates. For example, the virtual point determination unit 120 determines the virtual point coordinates based on feature point information indicating the position of feature points obtained by VSLAM using image 507 (more specifically, at least image 507). Alternatively, for example, the virtual point determination unit 120 determines the virtual point coordinates based on matching information regarding the matching of feature points obtained by VSLAM using image 507 (more specifically, at least image 507). The matching information is, for example, map information. Specifically, the matching information is information indicating the position of map points (target map points) included in the map information, calculated from the feature points (target feature points) included in image 507.
[0126] Furthermore, in this example, the coordinates (A1, B1, C1) in the world coordinate system corresponding to the coordinates (A, B, C) in the first camera coordinate system, which is based on the position and orientation of the imaging device 300 when image 507 was captured, do not match the coordinates (A2, B2, C2) in the world coordinate system corresponding to the coordinates (A, B, C) in the second camera coordinate system, which is based on the position and orientation of the imaging device 300 when image 508 was captured. In other words, in this example, it is estimated that the virtual point moved from time t-1 to time t. For example, if the coordinates of (A2, B2, C2) in the first coordinate system are (A3, B3, C3), then the velocity of the virtual point can be calculated, for example, as (A3-A1, B3-B1, C3-C1). For example, the speed of the virtual point can be calculated from the calculated velocity of the virtual point.
[0127] Thus, the virtual point coordinates may be determined from the feature points of the image. For example, the coordinates of the feature points included in the image may be used as the virtual point coordinates.
[0128] Alternatively, for example, coordinates corresponding to at least a portion of the line segments contained in the image may be used as virtual point coordinates.
[0129] Furthermore, coordinates corresponding to at least some of the edges included in the image may be used as virtual point coordinates.
[0130] Furthermore, the virtual point coordinates can be arbitrarily determined in advance and are not particularly limited.
[0131] Alternatively, multiple virtual point coordinates may be set, a single virtual point coordinate which is the average of the multiple virtual point coordinates may be calculated, and the speed of the virtual point may be calculated based on the calculated single virtual point coordinate.
[0132] Alternatively, multiple virtual point coordinates may be set, and the speed of each virtual point may be calculated. The maximum value among these speeds may then be ultimately adopted as the virtual point's speed.
[0133] Figure 6 is a diagram illustrating a fifth example of a virtual point according to the embodiment. Figure 6(a) schematically shows the position and orientation of the imaging device 300 at time t-1 in the world coordinate system, and Figure 6(b) schematically shows the position and orientation of the imaging device 300 at time t in the world coordinate system. In this example, after the imaging device 300 captures image 509 at time t-1, only the position of the imaging device 300 is changed, and at time t, the imaging device 300 captures image 510.
[0134] The estimation device 100 acquires distance information indicating the distance between the object (object to be imaged) and the imaging device 300 from a distance measuring sensor such as LiDAR. If such distance information is acquired, virtual point coordinates may be determined based on the distance information. In other words, for example, the virtual point determination unit 120 may acquire distance information indicating the distance from the imaging device 300 to a predetermined object, and determine virtual point coordinates based on the distance information.
[0135] Thus, distance information may be used to determine the virtual point coordinates. For example, distance information corresponding to the image at time t-1 may be obtained from a distance measuring sensor or the like, and the virtual point coordinates may be determined based on the obtained distance information. For example, the coordinates of feature points in the image that are included in the imaged object, whose distances indicated by the distance information are known, are determined as the virtual point coordinates. In this example, the imaged object is a house, and the positions corresponding to the feature points of the object shown in Figure 6 within that house are determined as the virtual point coordinates.
[0136] Furthermore, map information (specifically, map points) may be generated based on distance information indicating the distance between an object and the imaging device 300, obtained from, for example, a distance measuring sensor such as LiDAR. For example, the coordinates of map points in the map information that are included in an imaging target whose distance indicated by the distance information is known may be determined as virtual point coordinates.
[0137] The predetermined object can be arbitrarily defined and is not particularly limited. Information indicating the predetermined object may be stored in the storage unit 140 in advance, or it may be obtained from the user via an operating device or the like. For example, the coordinates of any map point among a plurality of map points indicating the predetermined object included in the map information are determined as virtual point coordinates.
[0138] Furthermore, distance information may be acquired by the estimation device 100 in any way. Distance information may be acquired from the distance measuring sensor, such as a ToF sensor or LiDAR, when the distance between the imaging device 300 and a predetermined object is measured and the measurement result is obtained as distance information from the distance measuring sensor, or distance information may be acquired from the user via an operating device.
[0139] Furthermore, the distance information may also be depth information indicating the depth corresponding to a pixel included in the image. In other words, for example, the virtual point determination unit 120 may determine the virtual point coordinates based on depth information indicating the depth corresponding to a pixel included in the image 509.
[0140] The depth information may be acquired by the estimation device 100 by any method. The depth information may be acquired by the imaging device 300 which is equipped with a depth sensor, or by the user via an operating device. The imaging device 300 may also be a depth camera, and the multiple images acquired from the imaging device 300 may each be depth images.
[0141] Alternatively, for example, the virtual point determination unit 120 may determine the virtual point coordinates based on depth information indicating the depth corresponding to pixels included in the image (e.g., image 509), which is calculated using segmentation. The virtual point determination unit 120 labels each pixel included in image 509 using segmentation such as semantic segmentation, and uses this result to calculate the depth of a specific pixel (e.g., one or more pixels in which a predetermined imaging target is captured).
[0142] Thus, for example, by performing segmentation on the image, an image corresponding to a predetermined target being captured within that image may be adopted as a virtual point coordinate.
[0143] [Processing Procedure] Next, the processing procedure of the estimation device 100 will be described. For example, based on the image captured by the imaging device 300, the estimation device 100 generates a map, which is an example of a three-dimensional model, in real time, determines the position of virtual points, and evaluates the image quality (specifically, calculates the speed of the virtual points).
[0144] Figure 7 is a flowchart showing the processing procedure of the estimation device 100 according to the embodiment.
[0145] First, the imaging device 300 generates an image by capturing the subject. Next, the imaging device 300 transmits sensor information, including the captured image (specifically, the image generated by capturing the subject), to the estimation device 100. The imaging device 300 also captures the subject and transmits sensor information, including the captured image, to the estimation device 100.
[0146] In this way, the imaging device 300 repeatedly captures images of the subject while moving, for example, and transmits a series of images to the estimation device 100.
[0147] The estimation device 100 acquires sensor information from the imaging device 300 (S110).
[0148] Next, the estimation device 100 estimates the position and orientation of the imaging device 300 based on the sensor information acquired from the imaging device 300 (S120). The estimation device 100 also generates a map based on the estimated position and orientation of the imaging device 300 and the sensor information.
[0149] Next, the estimation device 100 determines the position of the virtual point (virtual point coordinates) (S130). The estimation device 100 may determine the position of the virtual point using the method described above, or it may determine the virtual point coordinates to predetermined coordinates.
[0150] Next, the estimation device 100 calculates the image quality (S150). For example, the estimation device 100 calculates the speed of a virtual point based on the most recently captured image (e.g., the second image above) and the image captured immediately before it (e.g., the first image above), and generates auxiliary information related to the calculated speed.
[0151] Next, the estimation device 100 outputs the presentation information to the presentation device 200, causing the presentation device to display the presentation information (S150). The presentation information includes, for example, auxiliary information and images captured by the imaging device 300. For example, the estimation device 100 causes the presentation device 200 to display the auxiliary information along with the image with the most recent capture time.
[0152] Figure 8 shows the presentation image 600 presented by the presentation device 200 according to the embodiment.
[0153] The presented image 600 is an example of presented information, and includes, for example, an image 630 captured by the imaging device 300 and an image based on auxiliary information. In this example, the image based on auxiliary information is superimposed on the image captured by the imaging device 300. The presented image 600 is updated, for example, every frame. In other words, for example, the presented image 600 is updated every time image 630 changes to the latest image.
[0154] Furthermore, in this embodiment, the image based on the auxiliary information includes a speed image 610 and a warning image 620.
[0155] The speed image 610 is an image based on speed information. For example, the speed image 610 shows the speed of a virtual point and a threshold image 611 that indicates a predetermined first threshold.
[0156] Warning image 620 is an image based on the first warning information. For example, warning image 620 is displayed on the display device 200 as "Warning: Move slowly" when the speed of the virtual point is greater than or equal to a predetermined first threshold.
[0157] Speed information may be presented using scales or other visual aids, as shown in speed image 610, or it may be presented using numerical values (characters), or it may be presented in any manner.
[0158] Furthermore, the presented image 600 may include an image that shows a predetermined second threshold.
[0159] Furthermore, if the speed of the virtual point is below a predetermined second threshold, the display device 200 may display an image based on the second warning information, such as "Warning: Move quickly."
[0160] Furthermore, the warning information (first warning information and second warning information) may be information that shows text (images), as in the warning image 620; information that shows a display mode such as flashing the display or light source used by the user; or information that shows sound.
[0161] Furthermore, the presented information only needs to include supplementary information, and may or may not include images captured by the imaging device 300.
[0162] The accuracy of the position and orientation (pose) of the imaging device 300 used in VSLAM is strongly influenced by the movement (blur) of the image captured by the imaging device 300 when the imaging device 300 moves or rotates. For example, if the imaging device 300 is positioned to surround an object from the periphery, the blur of the image relative to the object will be reduced. By setting a virtual point at the position of an object in front of the imaging device 300 and prompting the user to operate in a way that does not exceed a predetermined speed, it is possible to capture images with less blur. By presenting the user with the speed of movement of the virtual point, it is possible to clearly demonstrate how to take high-quality images, enabling appropriate and efficient imaging.
[0163] As described above, the estimation device 100 can present the user with limits on their actions (e.g., limits on the user's movement speed) to prevent blurring before it occurs in the image. It can also present limits on the user's movement to maintain image quality and the quality of the three-dimensional model generated using the image, allowing the user to capture images efficiently. This disclosure is applicable, for example, to autonomous mobile robots. Furthermore, it can suppress the capture of low-quality images. In other words, it can capture images without blurring. For example, this disclosure is applicable from applications that generate three-dimensional models to applications that generate images and videos. Since this disclosure can be implemented by determining (calculating) the position of virtual points, it can suppress the capture of low-quality images with a small amount of processing power.
[0164] [summary] Figure 9 is a block diagram showing the configuration of the information processing device 10 according to the embodiment. Figure 10 is a flowchart showing the information processing method according to the embodiment.
[0165] The estimation device 100 described above is a specific example of the information processing device 10. For example, the information processing device 10 comprises a processor 11 and a memory 12. The processor 11 is connected to the memory 12 and, in operation, uses the memory 12 to execute the information processing method shown in Figure 10. In other words, for example, the estimation device 100 executes the information processing method shown in Figure 10.
[0166] First, the information processing device 10 acquires a first image captured by the camera at a first time point, and a second image captured by the camera at a second time point (S10).
[0167] The camera is, for example, the imaging device 300 described above. The first time is, for example, the time t-1 described above. The second time is, for example, the time t described above. The first image is, for example, the images 500, 502, 504, 507, or 509 described above. The second image is, for example, the images 501, 503, 505, 508, or 510 described above.
[0168] The camera generates multiple images by continuously capturing images at predetermined time intervals when operated by the user. The first and second images are, for example, two consecutive images from the multiple images generated in this way. The predetermined time interval is 10 fps (frames per second), but can be set arbitrarily.
[0169] Next, the information processing device 10 estimates the speed of a virtual point, assuming that the virtual point moved between the first and second time points, from a first position corresponding to the first coordinate in the first camera coordinate system, which is based on the position and orientation of the camera when the first image was captured, to a second position corresponding to the second coordinate, which has the same value as the first coordinate, in the second camera coordinate system, which is based on the position and orientation of the camera when the second image was captured (S20).
[0170] The first camera coordinate system is a three-axis Cartesian coordinate system that uses the position of the imaging device 300 at time t-1 as a reference (for example, the origin), and has three axes: a first axis parallel to the imaging direction of the imaging device 300 at time t, a second axis perpendicular to the first axis, and a third axis perpendicular to both the first and second axes. The second camera coordinate system is a three-axis Cartesian coordinate system that uses the position of the imaging device 300 at time t as a reference (for example, the origin), and has three axes: a first axis parallel to the imaging direction of the imaging device 300 at time t, a second axis perpendicular to the first axis, and a third axis perpendicular to both the first and second axes.
[0171] The first and second coordinates are, for example, coordinates in the world coordinate system. For example, if the first coordinate is (A, B, C) (where A, B, and C are arbitrary numbers), then the second coordinate is also (A, B, C). For example, the information processing device 10 converts the coordinates (A, B, C) in the first camera coordinate system and the coordinates (A, B, C) in the second camera coordinate system to coordinates in the same coordinate system, and uses the conversion result to calculate the speed of the virtual point. This same coordinate system may be the first camera coordinate system, the second camera coordinate system, the world coordinate system, or any other coordinate system.
[0172] Next, the information processing device 10 outputs auxiliary information regarding the speed of the virtual point (S30).
[0173] Auxiliary information includes, for example, the speed information (e.g., information showing the speed image 610) and / or warning information (e.g., information showing the warning image 620).
[0174] When images are repeatedly and continuously captured, it takes considerable effort for the user to check whether each image is blurry or not. Therefore, the information processing device estimates the speed of a set virtual point. If the estimated speed of the virtual point is too fast, there is a high probability that blurry images will be captured instead of sharp images. As a result, by checking the auxiliary information, the user taking images with the camera can capture images that are less likely to be blurry. Therefore, the information processing device can easily suppress the capture of blurry images, that is, low-quality images.
[0175] Furthermore, for example, the information processing device 10 acquires distance information indicating the distance from the camera to a predetermined object, and determines a first coordinate based on the distance information.
[0176] According to this, for example, when an image containing a specific object is captured, setting a virtual point at the position of the specific object makes it easier to suppress blurring of the specific object in the image.
[0177] The first coordinate may be arbitrarily determined in advance.
[0178] Furthermore, for example, the information processing device 10 determines the first coordinates based on depth information indicating the depth corresponding to the pixels included in the first image.
[0179] This makes it easier to pinpoint the specific location of an object in an image. For example, when an image containing such an object is captured, setting a virtual point at the object's location makes it easier to suppress blurring of the object in the image.
[0180] Furthermore, for example, the information processing device 10 determines the first coordinates based on depth information that indicates the depth corresponding to the pixels included in the first image, which is calculated using segmentation.
[0181] The information processing device 10 labels each pixel in the first image using segmentation, such as semantic segmentation, and uses this result to calculate the depth of a specific pixel in which a predetermined image target or the like is captured.
[0182] According to this, the first coordinate can be determined using the depth corresponding to a specific pixel in the first image.
[0183] Furthermore, for example, the information processing device 10 determines the first coordinates based on feature point information indicating the position of feature points obtained by VSLAM using the first image.
[0184] According to this, for example, by setting virtual points at the locations of feature points, it becomes easier to suppress the blurring of objects containing feature points that appear in the image.
[0185] Furthermore, for example, the information processing device 10 determines the first coordinates based on matching information regarding the matching of feature points obtained by VSLAM using the first image.
[0186] Matching information is, for example, map information. Specifically, matching information is information that indicates the location of a map point, which is included in the map information.
[0187] According to this, for example, by setting virtual points at the locations of common feature points present in both the first and second images, it becomes easier to suppress blurring of objects containing feature points that appear in both images. Therefore, for example, it becomes easier to suppress the decrease in accuracy of the three-dimensional model generated by VSLAM using the first and second images.
[0188] Furthermore, for example, the information processing device 10 performs a first determination to determine whether the speed of the virtual point is equal to or greater than a predetermined first threshold, and outputs auxiliary information based on the determination result of the first determination.
[0189] According to this, users can more easily decide whether or not to move the camera slowly to take an image by checking the auxiliary information.
[0190] Furthermore, for example, the information processing device 10 performs a second determination to determine whether the speed of the virtual point is less than or equal to a predetermined second threshold, which is lower than a predetermined first threshold, and outputs auxiliary information based on the determination results of the first and second determinations.
[0191] According to this, users can more easily decide whether or not to move the camera quickly to take an image by checking the auxiliary information.
[0192] Furthermore, for example, the information processing device 10 calculates a predetermined first threshold based on the camera's shutter speed.
[0193] The degree to which an image blurs varies depending on the shutter speed. Based on this, an appropriate predetermined first threshold can be set to make it easier for the user to capture an image that is not blurred.
[0194] Furthermore, for example, in calculating a predetermined first threshold, the information processing device 10 calculates the predetermined first threshold such that the faster the camera's shutter speed, the higher the predetermined first threshold becomes.
[0195] The faster the shutter speed, the less likely the image is to blur, even if the virtual point is moving quickly. This allows for setting an appropriate predetermined first threshold to make it easier for the user to capture a sharp image without unnecessarily slowing down the user's movements.
[0196] Furthermore, for example, if the information processing device 10 determines in the first determination that the speed of the virtual point is equal to or greater than a predetermined first threshold, it outputs auxiliary information that includes a first warning information to reduce the speed at which the user performing the photography using the camera moves.
[0197] The first warning information is, for example, information used to display a display such as the warning image 620 shown above on a display device 200 or the like.
[0198] This makes it easy for users to intuitively understand how fast they should move to take images.
[0199] Furthermore, for example, if the information processing device 10 determines in the first determination that the speed of the virtual point is greater than or equal to a predetermined first threshold, it outputs auxiliary information that includes a first warning to reduce the speed at which the user taking pictures with the camera moves. If the information processing device 10 determines in the second determination that the speed of the virtual point is less than or equal to a predetermined second threshold, it outputs auxiliary information that includes a second warning to increase the speed at which the user taking pictures with the camera moves.
[0200] This makes it easy for users to intuitively understand how fast they should move to take images.
[0201] Furthermore, for example, the information processing device 10 includes speed information indicating the speed of a virtual point as auxiliary information.
[0202] Speed information is information used to display a display such as the speed image 610 shown above on a display device such as the presentation device 200.
[0203] This makes it easy for users to intuitively understand how fast they should move to take images.
[0204] Furthermore, for example, the information processing device 10 includes saturation information that indicates the saturation of an image when the display device 200 displays an image containing the auxiliary information, with the auxiliary information being saturation information that indicates the saturation of the image when the display device 200 displays an image containing the auxiliary information.
[0205] This makes it easy for users to intuitively understand how fast they should move to take images.
[0206] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium. For example, this disclosure may be implemented as the above-mentioned information processing method executed by a computer, as a program for a computer to execute the above-mentioned information processing method, or as a computer-readable non-temporary recording medium on which the program is recorded.
[0207] (Other embodiments) The information processing device 10 and other components according to the embodiments of this disclosure have been described above, but this disclosure is not limited to these embodiments.
[0208] For example, the estimation device 100, the display device 200, and the imaging device 300 may be separate or integrated. For example, the estimation device 100, the display device 200, and the imaging device 300 may be implemented by a display 0, a camera, and a computer capable of communicating with them. Alternatively, the estimation device 100, the display device 200, and the imaging device 300 may be implemented by a single computer, such as a tablet terminal having a camera and a display.
[0209] Furthermore, for example, each processing unit included in the information processing device 10 according to the above embodiment is typically implemented as an LSI, which is an integrated circuit. These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.
[0210] Furthermore, integrated circuit implementation is not limited to LSIs; it can also be achieved using dedicated circuits or general-purpose processors. Field-Programmable Gate Arrays (FPGAs), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow for the reconfiguration of the connections and settings of circuit cells within the LSI, may also be used.
[0211] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0212] Furthermore, for example, this disclosure may be implemented as an information processing method executed by an information processing device 10 or the like.
[0213] Furthermore, the division of functional blocks in a block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. Additionally, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.
[0214] Furthermore, for example, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.
[0215] The above describes information processing systems 400 and information processing devices 10, etc., according to one or more embodiments, based on embodiments. However, this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications to the embodiments that a person skilled in the art could conceive, or forms constructed by combining components from different embodiments, may also be included within the scope of one or more embodiments. [Industrial applicability]
[0216] This disclosure is applicable to systems that generate three-dimensional models. [Explanation of symbols]
[0217] 10 Information Processing Devices 11 processors 12 memory 100 Estimator 110 Position and orientation estimation unit 120 Virtual Point Determination Unit 130 Image Quality Calculation Unit 140 Storage section 200 Presentation device 300 Imaging devices 400 Information Processing Systems Images 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 630 600 images presented 610 Speed Images 611 Threshold Image 620 Warning Images
Claims
1. Processor and Equipped with memory, The processor uses the memory to: A first image captured by the camera at a first time point, and a second image captured by the camera at a second time point are acquired. Assuming that the virtual point moved between the first and second time points, from a first position corresponding to a first coordinate in a first camera coordinate system based on the position and orientation of the camera when the first image was captured, to a second position corresponding to a second coordinate having the same value as the first coordinate in a second camera coordinate system based on the position and orientation of the camera when the second image was captured, the speed of the virtual point is estimated. Output auxiliary information regarding the speed of the aforementioned virtual point. Information processing device.
2. The camera acquires distance information indicating the distance to a predetermined object. Based on the distance information, the first coordinate is determined. The information processing apparatus according to claim 1.
3. Based on depth information indicating the depth corresponding to the pixels included in the first image, the first coordinates are determined. The information processing apparatus according to claim 1.
4. The first coordinates are determined based on depth information indicating the depth corresponding to a pixel included in the first image, which is calculated using segmentation. The information processing apparatus according to claim 1.
5. Based on the feature point information indicating the position of feature points obtained by VSLAM (Visual Simultanate Localization and Mapping) using the first image, the first coordinates are determined. The information processing apparatus according to claim 1.
6. Based on matching information regarding the matching of feature points obtained by VSLAM using the first image, the first coordinates are determined. The information processing apparatus according to claim 1.
7. A first determination is made as to whether the speed of the virtual point is equal to or greater than a predetermined first threshold. In the output of the auxiliary information, the auxiliary information based on the determination result of the first determination is output. The information processing apparatus according to claim 1.
8. Furthermore, a second determination is made as to whether the speed of the virtual point is less than or equal to a predetermined second threshold, which is lower than the predetermined first threshold. In the output of the auxiliary information, the auxiliary information is output based on the determination results of the first determination and the second determination. The information processing apparatus according to claim 7.
9. Based on the shutter speed of the camera, the predetermined first threshold is calculated. The information processing apparatus according to claim 7 or 8.
10. In calculating the predetermined first threshold, the predetermined first threshold is calculated such that the faster the shutter speed of the camera, the higher the predetermined first threshold. The information processing apparatus according to claim 9.
11. In the first determination, if it is determined that the speed of the virtual point is equal to or greater than the predetermined first threshold, the output of the auxiliary information includes a first warning information to reduce the speed at which the user performing the shooting using the camera moves. The information processing apparatus according to claim 7.
12. In the first determination, if it is determined that the speed of the virtual point is equal to or greater than the predetermined first threshold, the output of the auxiliary information includes a first warning information to reduce the speed at which the user performing the shooting using the camera moves. In the second determination, if it is determined that the speed of the virtual point is less than or equal to the predetermined second threshold, the output of the auxiliary information includes a second warning information to increase the speed at which the user performing the shooting using the camera moves. The information processing apparatus according to claim 8.
13. The aforementioned auxiliary information includes speed information indicating the speed of the virtual point. The information processing apparatus according to claim 1.
14. The auxiliary information is saturation information indicating the saturation corresponding to the speed of the virtual point, and includes saturation information indicating the saturation of the image when the display device displays the image containing the auxiliary information. The information processing apparatus according to claim 1.
15. A first image captured by the camera at a first time point, and a second image captured by the camera at a second time point are acquired. Assuming that the virtual point moved between the first and second time points, from a first position corresponding to a first coordinate in a first camera coordinate system based on the position and orientation of the camera when the first image was captured, to a second position corresponding to a second coordinate having the same value as the first coordinate in a second camera coordinate system based on the position and orientation of the camera when the second image was captured, the speed of the virtual point is estimated. Output auxiliary information regarding the speed of the aforementioned virtual point. Information processing methods.
16. For a computer to execute the information processing method described in claim 15, program.