Information processing device, information processing method, and program

JP2026131269APending Publication Date: 2026-08-14CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-03
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

【0010】 本発明によれば、撮像画像における現実物体に当該現実物体の3次元モデルを射影した2次元モデルを重畳する際に、現実物体と2次元モデルとの誤差を低減することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026131269000001_ABST
    Figure 2026131269000001_ABST
Patent Text Reader

Abstract

When superimposing a 2D model, which is obtained by projecting a 3D model of a real object onto a captured image, the error between the real object and the 2D model is reduced. [Solution] The information processing device includes: a first acquisition means for acquiring a first feature by projecting the features of a 3D model corresponding to a real object in an image of real space onto the image plane of the image; a second acquisition means for acquiring a second feature corresponding to the first feature from the image; and a correction means for correcting the 2D model obtained by projecting the 3D model onto the image plane of the image, based on the difference between the first feature and the second feature, so that it is displayed superimposed on the real object in the image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Extended Reality (XR) technology is becoming more widespread. XR includes, for example, virtual reality (VR), mixed reality (MR), and augmented reality. This includes Reality (AR), etc. In XR, a computer graphics (CG) model is presented to the user via a head-mounted device or a handheld device, which is rendered in accordance with the movement of the device. A head-mounted device is, for example, a head-mounted display (HMD), and a handheld device is, for example, a smartphone.

[0003] Within XR, specifically MR, there are use cases where a nearly identical 3D model is superimposed onto an object existing in the real world. One such use case is when you want to make a real object, such as a car, appear as if its color and texture have been changed. In this use case, a CG model with the same shape as the real car but with altered color and texture is superimposed on the real object. Another use case is achieving occlusion between a real object and a CG model. The depth information of the real object used to achieve occlusion between a real object and a CG model is obtained by superimposing a 3D reconstruction model of the real world, and a pre-created 3D model, onto the real world.

[0004] Most MR devices present, for the captured images, a 2D image obtained by synthesizing a 3D model projected onto the image plane of the captured image with a 2D model after presenting the 2D image through a panel such as a liquid crystal or an organic EL. When superimposing a 2D model obtained by projecting the 3D model of a real object onto the image plane on the real object in the captured image, an error (misalignment) between the real object and the 2D model occurs due to various factors. Representative causes of the error include the accuracy of calibration by camera parameters, the accuracy of the 3D model, the accuracy of 3D alignment / fitting between the 3D model and the real space, and the accuracy of the position and orientation of the imaging camera. However, it is difficult to perfectly align the real object and the 2D model by improving these accuracies.

[0005] Patent Document 1 discloses a technique for synthesizing a real space image with corrected image quality such as color tone and a virtual object and displaying a synthesized image with less discomfort. Patent Document 2 discloses a technique for accurately creating a three-dimensional portrait of a person by modifying a 3D model so as to fit a two-dimensional face photograph.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0007] Even if the image quality of the virtual object is adjusted to match the real space image or the 3D model is modified to fit the 2D image, the error when superimposing the 2D model obtained by projecting the 3D model on the real object in the captured image cannot be eliminated.

[0008] Therefore, the present invention aims to provide an information processing device that reduces the error between a real object and a two-dimensional model when superimposing a two-dimensional model, which is obtained by projecting a three-dimensional model of a real object onto a real object in an captured image. [Means for solving the problem]

[0009] The information processing device according to the present invention is characterized by comprising: a first acquisition means for acquiring a first feature by projecting the features of a three-dimensional model corresponding to a real object in an image of real space onto the image plane of the image; a second acquisition means for acquiring a second feature corresponding to the first feature from the image; and a correction means for correcting a two-dimensional model obtained by projecting the three-dimensional model onto the image plane of the image, based on the difference between the first feature and the second feature, so that it is displayed superimposed on the real object in the image. [Effects of the Invention]

[0010] According to the present invention, when superimposing a two-dimensional model obtained by projecting a three-dimensional model of a real object onto a real object in an captured image, it is possible to reduce the error between the real object and the two-dimensional model. [Brief explanation of the drawing]

[0011] [Figure 1] This figure illustrates the configuration of an information processing device according to Embodiment 1. [Figure 2] This is a flowchart illustrating the processing of the information processing device according to Embodiment 1. [Figure 3] This is a diagram illustrating the processing of the information processing device according to Embodiment 1. [Figure 4] This diagram illustrates the configuration of an information processing device according to Modification Example 1. [Figure 5] This is a flowchart illustrating the processing of the information processing device according to Modification Example 1. [Figure 6] This diagram illustrates the processing of an information processing device related to Modification Example 1. [Figure 7]This diagram illustrates the configuration of an information processing device according to Modification Example 2. [Figure 8] This is a flowchart illustrating the processing of the information processing device according to Modification Example 2. [Figure 9] This figure illustrates the configuration of an information processing device according to Embodiment 2. [Figure 10] This flowchart illustrates the process of generating a depth image database. [Figure 11] This diagram illustrates the process of generating a depth image database. [Figure 12] This is a flowchart for the process of reprojecting the depth image at the current position and orientation. [Figure 13] This diagram illustrates the process of reprojecting the depth image at the current position and orientation. [Figure 14] This diagram illustrates the hardware configuration of an information processing device. [Modes for carrying out the invention]

[0012] <Embodiment 1> Embodiment 1 is an embodiment in which a 3D model is superimposed as computer graphics on a corresponding real object. The configuration of the information processing device according to Embodiment 1 will be described with reference to Figures 1 and 14. Figure 1 is a diagram illustrating the configuration of the information processing device 101 according to Embodiment 1. Although the explanation will be given using an example where the information processing device 101 is an HMD, the information processing device 101 may be a handheld device such as a smartphone. Furthermore, the information processing device 101 may be placed inside the HMD as a device that realizes the functions according to the present invention. The information processing device 101 may be realized as a single piece of hardware, or as a group of hardware composed of multiple housings.

[0013] The imaging unit 102 captures a real-world scene (real space) for displaying superimposed computer graphics to the user wearing the HMD. The imaging unit 102 may be a two-lens stereo camera if it is presenting stereo images to the HMD. The imaging unit 102 may also be a monocular camera.

[0014] The position and orientation acquisition unit 103 is the position and orientation of the imaging unit 102, as seen from the user's perspective. The position and orientation used to generate computer graphics (CG) are acquired. If the imaging unit 102 is a stereo camera, the position and orientation acquisition unit 103 acquires the position and orientation of each of the left and right cameras. The method of acquiring the position and orientation by the imaging unit 102 is not particularly limited. The position and orientation acquisition unit 103 can acquire the position and orientation of the imaging unit 102 using the cameras of the imaging unit 102, a camera provided separately from the imaging unit 102, or an IMU (Inertial Measurement Unit). The position and orientation acquisition unit 103 may use an inside-out method using technologies such as Simultaneous Localization and Mapping (SLAM) or Odometry as a method for acquiring the position and orientation of the imaging unit 102. Alternatively, the position and orientation acquisition unit 103 may acquire the position and orientation of the imaging unit 102 using an outside-in method using a motion capture system or the like.

[0015] The 3D model alignment unit 104 superimposes a 3D model corresponding to a target real object onto the real object in a predetermined 3D coordinate system (e.g., the world coordinate system). The 3D model alignment unit 104 acquires the position and orientation of the 3D model and can align the 3D model based on the position and orientation of the real object. Specifically, the 3D model alignment unit 104 may perform alignment by placing markers on the real object and detecting the placed markers. The 3D model alignment unit 104 may also perform alignment based on input from the user via a controller or GUI. Furthermore, the 3D model alignment unit 104 may automatically acquire the position and orientation of the 3D model using an ICP (Iterative Closest Point) algorithm utilizing edge information and vertices, as well as machine learning techniques.

[0016] The 3D model data storage unit 105 stores 3D model (3D shape) data corresponding to real objects. When the 3D model data is used for CG rendering, the 3D model data includes not only shape information of the 3D model but also material information that expresses color and texture. The 3D model data may also include information associated with the 3D model, such as the 3D position of markers placed on the 3D model and information on feature points on the 3D model.

[0017] The CG rendering unit 106 renders the CG using the 3D model data held by the 3D model data holding unit 105, the position and orientation of the 3D model, the position and orientation of the imaging unit 102, and the camera parameters of the imaging unit 102. CG rendering can be performed by a GPU, but is not limited to this. When displaying a portion of the image captured by the imaging unit 102 after cropping, the CG rendering unit 106 performs rendering using the camera parameters after cropping.

[0018] The feature acquisition unit 107 acquires feature points from the rendered CG image and feature points from the image captured by the imaging unit 102, and obtains information on the correspondence between these feature points. The error acquisition unit 108 acquires a correction vector to correct the error between the feature points acquired by the feature acquisition unit 107. The correction unit 109 corrects the CG using the correction vector information acquired by the error acquisition unit 108.

[0019] The synthesis unit 110 generates a composite image by combining the CG corrected by the correction unit 109 with the real image captured by the imaging unit 102. If the 3D shape and position and orientation of other real objects other than the real object on which the CG is superimposed are known, the correction unit 109 also corrects the depth image of the CG. By correcting the depth image of the CG, the synthesis unit 110 can generate a composite image that represents the occlusion between the CG and the real image.

[0020] The display unit 111 displays the composite image generated by the synthesis unit 110. When displaying in stereo, the display unit 111 arranges two display panels and displays images corresponding to the left eye and the right eye respectively. Display it on the display panel.

[0021] The hardware configuration of the information processing device 101 will be described with reference to Figure 14. Figure 14 shows an example in which the information processing device 101 is configured as an HMD. In the example in Figure 14, the information processing device 101 is equipped with the display unit 111 described in Figure 1 and is connected to the imaging unit 102 via the interface (I / F) 1412. The information processing device 101 is a CPU (Central The information processing unit 101 includes a Processing Unit 1401 and a ROM (Read Only Memory) 1402. The information processing unit 101 also includes a storage medium drive 1405, an external storage device 1406, and a RAM (Random Access Memory) 1407. Each device constituting the information processing unit 101 is interconnected via a bus 1410.

[0022] If the information processing device 101 is configured as a separate device from the HMD, the information processing device 101 may include input devices such as a mouse 1408 and a keyboard 1409, and may also include a monitor 1411 as an external display device to the HMD. In this case, the display unit 111 is included in the HMD.

[0023] The CPU 1401 controls the entire information processing device 101 using the programs and data stored in the ROM 1402 and RAM 1407. The CPU 1401 executes various processes by the information processing device 101 by operating as the functional unit described in Figure 1.

[0024] ROM 1402 stores computer configuration data and boot programs. RAM 1407 temporarily stores programs and data loaded from external storage devices 1406 and storage media drives 1405. RAM 1407 also temporarily stores data received from external devices such as the imaging unit 102 via I / F 1412. RAM 1407 also functions as a work area used by the CPU 1401 when executing various processes.

[0025] I / F1412 has an analog video port or a digital input / output port such as IEEE1394 for connecting the imaging unit 102. Data received via I / F1412 is input to RAM1407 or external storage device 1406.

[0026] The storage medium drive 1405 reads programs and data recorded on storage media such as CD-ROMs and DVD-ROMs, and writes programs and data to storage media. Note that some or all of the programs and data described as being stored in the external storage device 1406 may also be recorded on the storage media. The programs and data read by the storage medium drive 1405 from the storage media are output to the external storage device 1406 or RAM 1407.

[0027] The external storage device 1406 is a storage device capable of storing large amounts of information, such as a hard disk drive. The external storage device 1406 stores the OS (operating system), programs for the CPU 1401 to execute the processing of each functional unit of the information processing device 101, and various types of data. The various types of data include data used in the processing of each functional unit of the information processing device 101, such as information on the 3D model and information on the position and orientation of the imaging unit 102. The programs and various types of data stored in the external storage device 1406 are loaded into the RAM 1407 as appropriate, according to the control of the CPU 1401. The CPU 1401 can realize the processing of the information processing device 101 by executing the programs loaded into RAM 1407 using the various types of data loaded into RAM 1407.

[0028] The mouse 1408 and keyboard 1409 are examples of input devices, and the user can input various commands to the CPU 1401 by operating them. The monitor 1411 includes a CRT (Cathode Ray Tube) or an LCD panel, etc. The monitor 1411 can display the same image as the HMD display unit 111.

[0029] Referring to Figures 2, 3(A) to 3(E), a process for reducing errors that occur when the CG rendering unit 106 renders a CG image is superimposed onto a corresponding real object will be described. Figure 2 is a flowchart illustrating the processing of the information processing device 101 according to Embodiment 1. Figures 3(A) to 3(E) are diagrams for explaining the processing of the information processing device according to Embodiment 1. Figure 3(A) is an image of the real object 301 to which the CG is to be superimposed.

[0030] In step S201, the CG rendering unit 106 projects a 3D model corresponding to a real object in an image of real space onto the image plane of the captured image and renders it as a CG (2D model). The CG rendering unit 106 simply needs to project the shape data (vertices, edges, faces) of the 3D model onto the image plane of the captured image.

[0031] Figure 3(B) shows the rendered CG302. The CG rendering unit 106 may generate a depth image of CG302, not limited to a color image of CG302, in order to achieve occlusion with other real objects.

[0032] The processing in steps S202 to S203 is performed by the feature acquisition unit 107. In step S202, the feature acquisition unit 107 acquires (extracts) first feature points from the CG. The feature acquisition unit 107 can acquire first feature points from, for example, the geometric features (e.g., edges and vertices) of the shape of the 3D model itself. Figure 3(C) shows how multiple first feature points 304 have been acquired from the rendered CG 302. The method of acquiring feature points is not particularly limited. In order to facilitate matching with feature points of the real image in step S203, the feature acquisition unit 107 may acquire first feature points from an image to which at least one of the following image processing methods has been applied: grayscale conversion, contrast adjustment, and edge detection. For example, the feature acquisition unit 107 may convert the CG image to grayscale, adjust the contrast, convert it to an edge image, and then acquire feature points.

[0033] In the following explanation, an example of acquiring feature points will be used, but the feature acquisition unit 107 is not limited to acquiring feature points; it may also acquire edges or faces as features. The first feature point corresponds to the first feature. The algorithm for acquiring features is not particularly limited. The feature acquisition algorithm may be SIFT (Scale Invariant Feature Transform) and FAST (Features from Accelerated Segment Test). Alternatively, the feature acquisition algorithm may be ORB (Oriented FAST and Rotated BRIEF) and AKAZE (Accelerated KAZE). The feature acquisition unit 107 may acquire feature points using machine learning methods such as SUPERPOINT.

[0034] If the CG has textures that do not exist on the real object, the feature acquisition unit 107 may mask the area where the texture is applied, excluding it from the area for acquiring feature points. Through this process, the feature acquisition unit 107 can acquire multiple first feature points from the CG.

[0035] In step S203, the feature acquisition unit 107 acquires a second feature point corresponding to the first feature point 304 from the real object in the image captured by the imaging unit 102. If there are multiple first feature points 304, the feature acquisition unit 107 acquires a second feature point corresponding to each of them. Figure 3(D) shows that multiple second feature points 305 have been acquired from the real object in the image. This shows the state in which it was observed. The second characteristic point corresponds to the second characteristic.

[0036] Since the image captured by the imaging unit 102 has different image quality from the CG image, the feature acquisition unit 107 may acquire second feature points from an image obtained by applying at least one of the following image processing methods to the CG image: grayscale conversion, contrast adjustment, or edge detection.

[0037] The method for associating the first feature point 304 with the second feature point 305 is not particularly limited. The feature acquisition unit 107 may acquire the second feature point using the same method as the first feature point and associate them using the feature quantities of each feature point, or it may use a tracking method such as the Lucas-Kanade method to associate them. The feature acquisition unit 107 may also use a machine learning method such as LightGlue to associate the first feature point 304 with the second feature point 305. Since the second feature point corresponding to the first feature point is thought to exist in the vicinity of the first feature point, the feature acquisition unit 107 only needs to search for the second feature point in the vicinity of the first feature point. The feature acquisition unit 107 may associate the first feature point acquired from a CG image with the second feature point searched in its vicinity, or it may associate the second feature point acquired from a real image with the first feature point searched in its vicinity.

[0038] In step S204, the error acquisition unit 108 acquires a correction vector from a first feature point obtained from the computer graphics (CG) and a second feature point corresponding to the first feature point and obtained from the captured image. The correction vector is used to correct the error between the first feature point and the second feature point.

[0039] Specifically, the error acquisition unit 108 can obtain a correction vector 306 that corrects the coordinates with errors in the CG (2D model) to the coordinates of the corresponding real object by subtracting the image coordinates of the first feature point 304 from the image coordinates of the second feature point 305. In other words, the error acquisition unit 108 can correct the CG using the correction vector obtained by subtracting the coordinates indicating the position of the first feature corresponding to the second feature from the coordinates indicating the position of the second feature.

[0040] Since there are multiple first feature points and multiple second feature points corresponding to each first feature point, the error acquisition unit 108 acquires multiple correction vectors 306 (hereinafter also referred to as a group of correction vectors). To suppress errors in the correction vectors 306, the error acquisition unit 108 may remove correction vectors 306 that are longer than a threshold. Alternatively, the error acquisition unit 108 may suppress errors in the correction vectors 306 by removing outliers using a method such as RANSAC (Random Sample Consensus).

[0041] In step S205, the correction unit 109 corrects the CG (2D model) using the correction vectors obtained in step S204 so that it appears superimposed on the real object in the captured image. The correction unit 109 may, for example, perform Delaunay triangulation using a plurality of first feature points to generate a group of polygons, and then perform texture mapping on the generated polygons using the CG image. The correction unit 109 may also deform and correct the CG by moving the first feature points, which are the vertices of the texture-mapped polygons, according to the correction vector group.

[0042] The correction unit 109 may generate a homography transformation matrix for the entire image based on a group of correction vectors. The homography transformation matrix is ​​a matrix that transforms the coordinates indicating the position of a first feature to the coordinates indicating the position of a second feature corresponding to the first feature. The correction unit 109 can correct the CG by performing a homography transformation on the CG image using the generated homography transformation matrix.

[0043] Alternatively, the correction unit 109 may divide the CG image into grid-like regions and perform a homography transformation for each region. In this case, the correction unit 109 uses the correction vectors acquired for each region to generate a homography transformation matrix for each region and performs a homography transformation for each region.

[0044] Figure 3(E) shows the corrected CG307. The composite unit 110 combines the CG307 corrected in step S205 with the captured image and displays the combined image on the display unit 111.

[0045] When occlusion is achieved between CG307 and a real object other than the real object corresponding to CG307, the synthesis unit 110 uses the CG depth information. The correction unit 109 applies the same transformation as the CG correction in step S205 to the depth image generated by the CG rendering unit 106 in step S201, thereby obtaining a depth image that matches the contour of the corrected CG.

[0046] In the above embodiment 1, the information processing device 101 acquires a first feature point and a second feature point from a two-dimensional model obtained by projecting a three-dimensional model corresponding to a real object onto an captured image, and from the real object in the captured image, respectively. The information processing device 101 corrects the two-dimensional model (CG) based on the difference of the acquired feature points. As a result, the information processing device 101 can accurately superimpose the two-dimensional model onto the real object.

[0047] [Example 1] In Embodiment 1, the information processing device 101 obtains a correction vector for correcting the CG by associating a first feature point in the rendered CG with a second feature point in the real object in the image captured by the imaging unit 102. However, depending on the pattern on the real object, the shape of the real object, the shape of the CG, or the texture of the CG, it may be difficult to associate the feature points of the CG with the feature points of the real object. In Modification 1, by attaching multiple markers to the real object on which the CG is superimposed, the information processing device 101 can robustly estimate the correction vector using the feature points detected from the markers. In the description of Modification 1, explanations that overlap with Embodiment 1 are omitted.

[0048] Figure 4 is a diagram illustrating the configuration of the information processing device 101 according to Modification 1. The information processing device 101 according to Modification 1 includes a marker coordinate holding unit 401 and a marker projection unit 402 in addition to the configuration of the information processing device 101 according to Embodiment 1. Furthermore, the information processing device 101 according to Modification 1 includes a marker detection unit 403 instead of a feature acquisition unit 107.

[0049] The marker coordinate holding unit 401 holds the relative position and orientation of the marker in the coordinate system of the 3D model, corresponding to the marker attached to the real object. The relative position and orientation of the marker with respect to the 3D model may be determined by pre-determining a common origin position with the 3D model and measuring it in advance using a ruler or measuring instrument. The relative position and orientation between markers may be determined by image-based distance measurement using images captured from multiple viewpoints.

[0050] The markers attached to real objects may have shapes that can be recognized by their IDs. The feature acquisition unit 107 can easily associate the first feature point with the second feature point if the shape of the marker can be determined based on the marker's ID. When an ID is assigned to a marker, the marker coordinate holding unit 401 holds each marker in association with its respective ID and its relative position and orientation to the 3D model.

[0051] The marker projection unit 402 projects the marker position on the 3D model held by the marker coordinate holding unit 401 onto the image plane of the captured image. The marker projection unit 402 performs the process in step S502 of Figure 5. The marker detection unit 403 detects that the marker projection unit 402 has detected the captured image The marker detection unit 403 detects a marker (2D model) projected onto the image plane and acquires feature points. The marker detection unit 403 performs the process shown in step S503 of Figure 5.

[0052] Referring to Figures 5 and 6(A) to 6(E), the process by which the information processing device 101 according to Modification 1 corrects the computer graphics (CG) will be explained. Figure 5 is a flowchart illustrating the process of the information processing device 101 according to Modification 1. For processes that are the same as those in the flowchart of Figure 2, the same reference numerals are used and detailed explanations are omitted. Figures 6(A) to 6(E) are diagrams for explaining the process of the information processing device 101 according to Modification 1. Figure 6(A) is an image of the real object 601 on which the CG is superimposed. A marker is attached to the real object 601. The second feature point 602 is a feature point obtained from the marker attached to the real object 601.

[0053] In step S201, the CG rendering unit 106 renders the CG. Figure 6(B) shows the rendered CG302.

[0054] In step S502, the marker projection unit 402 projects markers on a 3D model corresponding to a real object onto the image plane of the captured image using the same parameters as those used for rendering the CG image. That is, markers on the 3D model are projected onto the image plane of the captured image using the position and orientation of the imaging unit 102 and the camera parameters used when projecting the 3D model onto the image plane of the captured image. Note that the marker projection unit 402 is not limited to markers; it may also project predetermined feature points specified by the user onto the image plane of the captured image using the same parameters as those used for rendering the CG image.

[0055] The marker projection unit 402 can project a marker on a 3D model onto the image plane of the captured image using the information held in the marker coordinate holding unit 401. Figure 6(C) shows how the marker is projected onto the image plane of the captured image. The marker projection unit 402 acquires a feature (for example, the position of the marker) obtained from the information projected onto the image plane of the captured image as a first feature point. Note that the marker projection unit 402 is not limited to markers; if a predetermined feature point specified by the user is projected onto the image plane of the captured image, the projected predetermined feature point is acquired as the first feature point.

[0056] In step S503, the marker detection unit 403 detects markers attached to real objects in the captured image taken by the imaging unit 102. The markers can be detected using known techniques, and the specific method for detecting the markers is not particularly limited. The marker detection unit 403 acquires features obtained from the detected markers (for example, the position of the markers) as second feature points 602.

[0057] In step S504, the error acquisition unit 108 acquires a correction vector 604 based on the difference between the first feature point 603 and the second feature point 602. The error acquisition unit 108 acquires the correction vector 604 by subtracting the image coordinates of the second feature point acquired in step S503, which corresponds to the first feature point, from the image coordinates of the first feature point acquired in step S502.

[0058] In step S205, the correction unit 109 uses the correction vector 604 acquired in step S504 to correct the CG (2D model) so that it is displayed superimposed on the real object in the captured image. The details of the correction process are the same as in Embodiment 1 and are therefore omitted.

[0059] In Modification 1, the information processing device 101 acquires a second feature point and a first feature point from a marker attached to a real object in the captured image, and from a two-dimensional marker obtained by projecting a marker on a three-dimensional model onto the captured image. Based on the difference between the acquired first feature point and the second feature point, the information processing device 101 acquires a second feature point and a first feature point from a two-dimensional marker corresponding to the real object to which the marker is attached. The 2D model, which is obtained by projecting the 3D model onto the captured image, is corrected. The information processing device 101 acquires feature points from markers of predetermined shapes, so it can stably and accurately superimpose the 2D model onto the real object.

[0060] In the above modified example 1, an example was described in which a 2D model is corrected using feature points obtained from a marker or predetermined feature points specified by the user. However, the information processing device 101 may also correct the 2D model by combining these feature points with those described in Embodiment 1. That is, the information processing device 101 may correct the 2D model by arbitrarily combining feature points obtained from a marker, predetermined feature points specified by the user, and feature points obtained from a 2D model obtained by projecting a real object and a 3D model onto an image. By combining various feature points, the information processing device 101 can superimpose the 2D model onto a real object with greater accuracy.

[0061] [Differentiation 2] In Embodiment 1, the information processing device 101 obtains a correction vector by associating a first feature point obtained (extracted) from a rendered CG image with a second feature point obtained from an image captured by the imaging unit 102 and taking the difference. Since the process of obtaining the correction vector as correction information for correcting a 2D model to be composited onto a real object takes time, the delay time until the final composite image is displayed may increase. Therefore, in Modification 2, the information processing device 101 reduces the delay time until the composite image is displayed by obtaining the correction information by prediction.

[0062] The information processing device 101 takes time to acquire the first feature point and the second feature point corresponding to the first feature point in the process of acquiring correction vectors for correcting a 2D model. The information processing device 101 reduces the time required to acquire correction information for the current frame by predicting the correction information for the current frame based on the correction information up to the previous frame. When predicting information is used, there is a concern that the accuracy of the correction will decrease. However, if the main cause of the superposition error of the CG image is the calibration error of the device, and if there is little jitter in the time-series change of the position and orientation of the imaging unit 102, it is assumed that there is time-series continuity in the change of error, so the decrease in the accuracy of the correction is limited.

[0063] In the following description of Modification 2, content that overlaps with Embodiment 1 will be omitted. Figure 7 is a diagram illustrating the configuration of the information processing device 101 according to Modification 2. The information processing device 101 according to Modification 2 has the same functional parts as the information processing device 101 according to Embodiment 1. In order to reduce the delay in the correction information acquisition process, the correction unit 109 acquires position and orientation information of the imaging unit 102 from the position and orientation acquisition unit 103 for use in predicting the correction information. Therefore, in Figure 7, a path is added for transmitting the position and orientation information of the imaging unit 102 from the position and orientation acquisition unit 103 to the correction unit 109.

[0064] Figure 8 is a flowchart illustrating the processing of the information processing device 101 according to Modification 2. The processing shown in Figure 8 is a process that reduces the superposition error of CG onto real objects while reducing the delay time. For processes that are the same as those in the flowchart of Figure 2, the same reference numerals are used and detailed explanations are omitted.

[0065] Systems that reduce CG rendering latency generally utilize image processing techniques known as warping and time warp. Figure 8 illustrates an example of combining the latency reduction process related to Modification Example 2 with warping and time warp.

[0066] Note that the warping and time warp processes in step S803 do not need to be executed. If the process in step S803 is not executed, the processes in steps S804 to S808 can also perform correction by warping and time warp. On the other hand, if the process in step S803 is executed, the amount of correction using errors on the image plane becomes smaller, thus reducing the impact of prediction errors.

[0067] Steps S201 to S202 are the same as in Embodiment 1, so their explanation will be omitted. In step S803, the correction unit 109 performs image processing such as warping and time warp on the CG image to compensate for the delay of the CG image that occurs during rendering. By performing image processing such as warping and time warp, the correction unit 109 can deform the CG image so that it becomes a CG image viewed from the latest position and orientation of the imaging unit 102. The method of warping is not particularly limited. For example, the correction unit 109 may use homography transformation, or it may use Positional Timewarp, which performs geometrically more accurate deformation using a depth image.

[0068] In step S804, the feature acquisition unit 107 acquires (extracts) second feature points from the real object in the captured image that correspond to the first feature points after the warping process in step S803. If there are multiple first feature points, the feature acquisition unit 107 acquires a second feature point corresponding to each of the first feature points.

[0069] The position of the first feature point acquired in step S202 is moved by the warping process performed in step S803. The feature acquisition unit 107 updates the position of the first feature point using a tracking method such as the Lucas-Kanade method and feature point matching.

[0070] The feature acquisition unit 107 acquires a second feature point based on the updated position of the first feature point. The method for matching the updated first feature point with the second feature point is the same as the process in step S203.

[0071] In step S805, the error acquisition unit 108 acquires a correction vector for correcting the error between a first feature point and a second feature point corresponding to that first feature point. If there are multiple first feature points, the error acquisition unit 108 acquires a group of correction vectors based on each first feature point and its corresponding second feature point. The method for acquiring the correction vectors is the same as in step S204 of Embodiment 1.

[0072] In step S806, the correction unit 109 obtains a homography transformation matrix for correcting the CG image based on the correction vector (group of correction vectors). The correction unit 109 stores the homography transformation matrix in the memory unit, linking it to the current frame (the image captured in step S201). The homography transformation matrix obtained in step S806 is not used to correct the CG image of the current frame, but is used to predict the homography transformation matrix for frames after the current frame.

[0073] In step S807, the correction unit 109 predicts the homography transformation matrix to be used to correct the CG image of the current frame based on the time-series changes of the homography transformation matrix associated with past frames stored in the memory unit. The method of prediction is not particularly limited. The correction unit 109 may make a linear prediction based on the changes in each element of the homography transformation matrix up to the previous frame, or it may make a prediction using a Kalman filter that treats each element of the homography matrix as a state variable. The correction unit 109 may also make a prediction using machine learning such as a recurrent neural network.

[0074] In step S808, the correction unit 109 corrects the CG using the homography transformation matrix predicted in step S807. Note that the process shown in Figure 8 is performed frame by frame (captured image). This is the process performed in step S803. After the process in step S803, the processes in steps S804 to S806 and the processes in steps S807 to S808 may be executed in parallel.

[0075] In the modified example 2, the information processing device 101 predicts and generates a homography transformation matrix for correcting the CG in the current frame (currently captured image) based on correction vectors acquired in past frames (captured images). Since the information processing device 101 predicts the homography transformation matrix to be used in the current frame from the homography transformation matrix of past frames without performing time-consuming processes such as acquiring feature points, it is possible to reduce the delay time until the CG image is displayed.

[0076] <Embodiment 2> Recent advancements in deep learning have made it possible to generate depth images with high contour accuracy and minimal loss of detail by inputting stereo images. Depth images allow for the acquisition of 3D points corresponding to each pixel, based on the depth information of each pixel and camera parameters. By utilizing depth images estimated from various viewpoints, it becomes possible to generate 3D models to achieve occlusion between real objects and CG objects. However, if the depth of the inferred depth image is inaccurate, or if there are errors in the position and orientation information of the imaging unit 901, the superposition accuracy of the corresponding real object on the image may decrease when viewed from a different viewpoint.

[0077] Embodiment 2 is an embodiment that corrects the depth image by obtaining the error between the result (position) of projecting the 3D points obtained from the corresponding feature points on the depth image onto the image plane of the captured image, and feature points obtained from the feature points obtained from the real object in the captured image. By correcting the depth image to match the real object, it is possible to accurately achieve occlusion between the real object and the CG object (3D model) even when viewed from various viewpoints.

[0078] Figure 9 is a diagram illustrating the configuration of the information processing device 900 according to Embodiment 2. The hardware configuration of the information processing device 900 according to Embodiment 2 is the same as the hardware configuration of the information processing device 101 according to Embodiment 1 described in Figure 14.

[0079] The imaging unit 901 is generally a stereo camera. In recent years, there has also been research on depth estimation using monocular cameras, so the number of cameras is not particularly limited. The depth image estimation unit 902 takes time t n In this system, a depth image is estimated or inferred using the image obtained from the imaging unit 901.

[0080] The position and orientation acquisition unit 903 acquires the position and orientation of the imaging unit 901. If the imaging unit 901 is a stereo camera, the position and orientation acquisition unit 903 acquires the position and orientation of each of the left and right cameras. The method for acquiring the position and orientation of the imaging unit 901 is not particularly limited. The position and orientation acquisition unit 903 may acquire the position and orientation of the imaging unit 901 using the cameras of the imaging unit 901, a camera provided separately from the imaging unit 901, or an inside-out method using an IMU (Inertial Measurement Unit). The inside-out method can be realized, for example, by technologies such as Simultaneous Localization and Mapping (SLAM) and Odometry. The position and orientation acquisition unit 903 may also acquire the position and orientation of the imaging unit 901 using an outside-in method using a motion capture system or the like.

[0081] The feature acquisition unit 904 acquires feature points from the image captured by the imaging unit 901. The depth image database 905 stores the depth image estimated or inferred by the depth image estimation unit 902 and metadata related to the depth image as a single depth image keyframe. The depth image database 905 stores multiple depth image keyframes.

[0082] The depth image selection unit 906 currently (time t c By reprojecting a depth image that matches the viewpoint of ) Select a suitable depth image keyframe for generation. The depth image reprojection unit 907 generates the depth image at the current time (time t c ) in line with the perspective of time t cReproject (project) it onto the captured image. Also, the depth image reprojection unit 907 acquires the coordinates of the 3D points corresponding to the feature points in the captured image based on the depth information at the coordinate values of the feature points of the real object. The depth image reprojection unit 907 projects the acquired coordinates of the 3D points onto the captured image at the current time (time t c ). Aligning with the viewpoint at the current time (time t c ), it reprojects (projects) onto the captured image at time t

[0083] The feature point matching unit 908 matches the feature points at time t n included in the depth image keyframe information with the feature points at time t c . The feature point matching unit 908 acquires the feature points at time t n corresponding to the feature points at time t c as the second feature points. The details of the matching process between the feature points at time t n and the feature points at time t c will be described in step S1203 of FIG. 12.

[0084] The error acquisition unit 909 acquires a correction vector from the coordinates of the first and second feature points in the captured image. The details of the acquisition process of the correction vector will be described in step S1204 of FIG. 12. The correction unit 910 corrects the depth image reprojected by the depth image reprojection unit 907 using the correction vector acquired by the error acquisition unit 909.

[0085] Referring to FIGS. 10 and 11(A) to 11(D), a method for generating depth image keyframe information at time t n will be described. FIG. 10 is a flowchart illustrating a process for generating a depth image database 905 that stores depth image keyframe information. The depth image keyframe information is generated based on information from the imaging unit 901, the depth image estimation unit 902, the feature acquisition unit 904, and the position and orientation acquisition unit 903. FIGS. 11(A) to 11(D) are diagrams for explaining the process of generating the depth image database 905.

[0086] In step S1001, the depth image estimation unit 902 determines the time t n The depth image 1101 (Figure 11(B)) is estimated from the image 1100 (Figure 11(A)) of the real object captured by the imaging unit 901. The algorithm for estimating the depth image 1101 is not particularly limited. The depth image estimation unit 902 may estimate the depth image using machine learning techniques such as HITNET, RAFT-Stereo, and CREStereo, or it may estimate the depth image using methods such as semi-global matching.

[0087] In step S1002, the feature acquisition unit 904 performs the time t n Next, feature points 1102 are obtained from the image 1100 acquired by the imaging unit 901. Figure 11(C) shows feature points 1102 of image 1100. The coordinates of the acquired feature points 1102 are (u ni ,v ni )

[0088] The algorithm for acquiring feature points is not particularly limited. The feature acquisition unit 904 may use methods such as SIFT, FAST, ORB, AKAZE, or machine learning methods such as SuperPoint to acquire feature points. In order to improve the robustness of the feature points, the feature acquisition unit 904 may perform contrast adjustment and various filtering processes as preprocessing on the image from which feature points are acquired.

[0089] In step S1003, the position and orientation acquisition unit 903 acquires the time t n The position and orientation of the imaging unit 901 are obtained. In step S1004, the depth image reprojection unit 907 compares the depth image estimated in step S1001 with the image coordinates of the feature points obtained in step S1002 and obtains the 3D position corresponding to the pixel of feature point 1102. As shown in Figure 11(D), the depth image reprojection unit 907 obtains the coordinates of the corresponding 3D position based on the coordinates of the pixel of feature point 1102 and the depth information (depth value) of the depth image 1101 at the coordinates of the pixel of feature point 1102. The obtained 3D position coordinates are (x ni ,y ni ,z niThe depth image reprojection unit 907 uses the depth information of the pixels of the feature point 1102 and the camera of the imaging unit 901. By using the focal length and principal point position information, a projection matrix can be calculated, and the inverse of the calculated projection matrix can be used to obtain the 3D position.

[0090] In step S1005, the depth image database 905 is stored at time t obtained in steps S1001 to S1004. n Depth image, time t n The feature points, the three-dimensional position of the feature points, and the position and orientation of the imaging unit 901 are stored as depth image keyframe information. The feature point information may include, in addition to the coordinate information on the captured image, the feature quantities of the feature points, the captured image itself, and at least one of the image patches around the coordinates of the feature points on the captured image.

[0091] Refer to Figures 12 and 13(A) to 13(E), time t c The depth image is reprojected to match the viewpoint at time t, and the reprojected depth image is displayed at time t c This section describes a correction method for aligning the captured image with the contours of real objects. Figure 12 shows the current time (time t c This is a flowchart for the process of reprojecting the depth image at the specified position and orientation.

[0092] Figures 13(A) to 13(E) show the current time (time t c This figure illustrates the process of reprojecting the depth image at the position and orientation of the imaging unit 901 of the ). Figure 13(A) shows the time t n Figure 13(B) shows the image 1100 of a real object and the feature points 1102 obtained (extracted) from the image 1100. Figure 13(B) shows the depth image shown in Figure 11(B) at time t c Based on the position and orientation of the imaging unit 901 at time t c The depth image 1301 is shown, which is a reprojection of the image captured at time t. The first feature point 1302 in the depth image 1301 represents the three-dimensional position of feature point 1102 at time t c Based on the position and orientation of the imaging unit 901 at time t cThis is the position reprojected onto the captured image. The first feature point 1302 corresponds to feature point 1102 obtained from the image 1100 of the real object.

[0093] Figure 13(C) shows time t c Figure 13(D) shows the captured image 1303 of the real object and the second feature point 1304 obtained from the captured image 1303. Figure 13(D) shows the first feature point 1302 projected onto the captured image 1303, the second feature point 1304 corresponding to the first feature point 1302 in the captured image 1303, and the correction vector 1305 obtained from the first feature point 1302 and the second feature point 1304. Figure 13(E) shows the corrected depth image 1306.

[0094] In step S1201, the depth image selection unit 906 selects the time t c Using information such as the position and orientation of the imaging unit 901 at time t c To generate depth images, one or more depth image keyframes are selected from the depth image database 905. The depth image selection unit 906 selects t c Select at least one depth image keyframe that was acquired at the position and orientation closest to the current position and orientation.

[0095] The depth image selection unit 906 selects the time t c Multiple depth image keyframes located within a predetermined distance and angle range from the position and orientation of the imaging unit 901 may be selected. The depth image reprojection unit 907 can compensate for any loss or defects that occur when reprojecting the depth image by combining the multiple depth image keyframes.

[0096] In step S1202, the depth image reprojection unit 907 operates at time t n Depth image and time t n The 3D position of feature points on a real object in the captured image and time t c Based on the position and orientation of the imaging unit 901 when the image was captured, at time t cThe depth image reprojection unit 907 reprojects (projects) the three-dimensional positions of feature points on real objects in the captured image and acquires these feature points as the first feature points 1302.

[0097] The depth image reprojection unit 907 at time t n By using the depth image and information on camera parameters such as the position and orientation of the imaging unit 901, focal length, and principal point position, at time t n Each of the depth images It is possible to convert pixels into three-dimensional points. The depth image reprojection unit 907, at time t c Using the position and orientation of the imaging unit 901 at time t, and a projection matrix corresponding to camera parameters such as focal length and principal point position, n It is possible to generate a depth image (distance image) by reprojecting the 3D points of each pixel in the depth image.

[0098] time t n and time t c Due to differences in the position and distance of real objects, missing pixels may occur in the reprojected depth image. The depth image reprojection unit 907 may appropriately compensate for the missing pixels by using a median filter or by drawing 3D points where the distance difference is less than or equal to a predetermined distance as larger points.

[0099] Figure 13(B) shows the depth image 1301 reprojected in step S1202. Time t n The three-dimensional position of feature point 1102 on the real object in the captured image is determined in the same way as in the depth image, at time t c It can be reprojected onto the captured image at time t. n The feature point 1102 on the real object in the captured image is at time t c The feature point reprojected using the position and orientation of the imaging unit 901 is the first feature point 1302.

[0100] In step S1203, the feature acquisition unit 904 and the feature point matching unit 908 perform the following at time t cA second feature point 1304 on the real object in the captured image is obtained. The second feature point 1304 corresponds to the time t of the depth image keyframe information selected in step S1201. n This is the feature point corresponding to the first feature point 1302.

[0101] First, the feature acquisition unit 904 detects time t n The feature point 1102 on the real object in the captured image (Figure 13(A)) corresponds to time t c The second feature point 1304 on the real object in the captured image (Figure 13(C)) is obtained.

[0102] The feature acquisition unit 904 uses the depth image keyframe information to determine time t c A second feature point 1304 can be obtained from the captured image. The feature acquisition unit 904, at time t n The second feature point 1304 can be obtained using the same method as when the first feature point 1102 was obtained. The feature point matching unit 908 may match the first feature point 1302 and the second feature point 1304 using their respective feature quantities.

[0103] The feature point matching unit 908 may associate the first feature point 1302 with the second feature point 1304 using a tracking method such as the Lucas-Kanade method. Alternatively, the feature point matching unit 908 may associate the first feature point 1302 with the second feature point 1304 using a machine learning method such as LightGlue. The search range for the second feature point 1304 can be narrowed down based on the location information of the first feature point 1302. By narrowing down the search range for the second feature point 1304, the information processing device 900 can reduce its processing load and increase its robustness.

[0104] time t n Feature point 1102 on the real object in the captured image is at time t c The second feature point 1304 corresponds to the first feature point 1302. Also, feature point 1102 corresponds to the first feature point 1302. Therefore, the feature point matching unit 908 determines the time t cThe second feature point 1304 and the first feature point 1302 can be associated with each other.

[0105] In step S1204, the error acquisition unit 909 acquires a correction vector from a first feature point and a second feature point corresponding to that first feature point. If there are multiple first feature points, the error acquisition unit 909 acquires a group of correction vectors from each first feature point and the second feature point corresponding to each first feature point.

[0106] Specifically, the error acquisition unit 909 uses the image coordinates of the second feature point 1304 to determine the first feature The correction vector 1305 can be obtained by subtracting the image coordinates of point 1302. The error acquisition unit 909 may remove correction vectors 1305 that are longer than a threshold in order to suppress errors in the correction vector 1305. Alternatively, the error acquisition unit 909 may suppress errors in the correction vector 1305 by removing outliers using a method such as RANSAC.

[0107] In step S1205, the correction unit 910 corrects the depth image using the correction vector 1305. The correction unit 910 may, for example, perform Delaunay triangulation using a plurality of first feature points to generate a group of polygons, and then perform texture mapping on the generated polygons using a CG image. The correction unit 910 may also deform and correct the depth image by moving the first feature points, which are the vertices of the texture-mapped polygons, according to the correction vector group.

[0108] The correction unit 910 may generate a homography transformation matrix for the entire image based on the correction vector group. The correction unit 910 can correct the depth image by performing a homography transformation on the depth image using the generated homography transformation matrix. Alternatively, the correction unit 910 may divide the depth image into grid-like regions and perform a homography transformation on each region. In this case, the correction unit 910 generates a homography transformation matrix for each region using the correction vectors obtained for each region and performs a homography transformation on each region. Figure 13(E) shows the corrected depth image 1306.

[0109] In Embodiment 2, time t n The camera of the imaging unit 901 captures an image of the real space for generating a depth image, and at time t c In the above explanation, it was assumed that the camera capturing the real-world image onto which the depth image is superimposed is the same camera. On the other hand, when using depth image information in MR to represent occlusion between CG and real objects, the camera capturing the real-world image for display in MR may be different from the camera capturing the real-world image for generating the depth image. When capturing the image for generating the depth image and the real-world image for display (hereinafter also referred to as the display image) with different cameras, it is desirable to superimpose the depth image so as to match the real object in the display image.

[0110] When using images captured by different cameras as display images, in step S1202, the depth image reprojection unit 907 can perform a reprojection process using the position and orientation of the camera capturing the display image and its camera parameters.

[0111] Furthermore, in step S1203, the feature acquisition unit 904 records time t cFor the image capture, the image captured by the camera that captures the display image can be used. When the feature acquisition unit 904 acquires the second feature point 1304 in Figure 13(C) which corresponds to the feature point 1102 in Figure 13(A), it may perform various filtering such as contrast adjustment and edge extraction, resizing, and cropping on each captured image. By performing these image processing steps, the feature acquisition unit 904 can easily associate the feature points even if the cameras that capture the images for which the feature point 1102 and the second feature point 1304 are acquired are different.

[0112] In the above embodiment 2, the information processing device 900 corrects the depth image based on the difference between a first feature point obtained by projecting a 3D point obtained from the feature points of the depth image onto the image plane of the captured image, and a second feature point in the captured image that corresponds to the first feature point. As a result, the information processing device 900 can accurately superimpose the CG object (3D model) onto the real object even when viewed from various viewpoints. Furthermore, the information processing device 900 can accurately achieve occlusion between the CG object and the real object.

[0113] Furthermore, the object of the present invention can also be achieved by the following method, namely, the above embodiment. A recording medium (or storage medium) containing program code for software that implements the function is supplied to the system or device. The storage medium supplied to the system or device is a computer-readable storage medium. The computer (or CPU, MPU) of the system or device reads and executes the program code stored on the recording medium. The program code read from the recording medium itself implements the function of the embodiment, and the recording medium containing the program code constitutes the present invention.

[0114] An operating system (OS) running on a computer executes program code read by the computer and performs some or all of the actual processing based on the instructions in the program code. The present invention also includes cases in which the functions of the above embodiments are realized by the processing of the OS or the like.

[0115] The program code read from the recording medium may be written to the memory of a function expansion card inserted into a computer or a function expansion unit connected to a computer. The present invention also includes cases in which a CPU 1201 or the like provided in the function expansion card or function expansion unit performs some or all of the actual processing based on the instructions of the written program code, thereby realizing the functions of the above embodiment. The recording medium to which the present invention is applied stores the program code corresponding to the processing described in each flowchart.

[0116] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). Multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) may share the processing to control the entire device.

[0117] Furthermore, the above-mentioned processors are processors in a broad sense, including general-purpose processors and specialized processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Specialized processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).

[0118] Furthermore, the embodiments described above (including modified examples) are merely examples, and configurations obtained by appropriately modifying or changing the above-described configurations within the scope of the gist of the present invention are also included in the present invention. Configurations obtained by appropriately combining the above-described configurations are also included in the present invention.

[0119] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit that implements one or more functions.

[0120] This embodiment includes the following configurations, methods, and programs. (Composition 1) A first acquisition means that acquires a first feature by projecting the features of a 3D model corresponding to a real object in an image captured in real space onto the image plane of the captured image, A second acquisition means for acquiring a second feature corresponding to the first feature from the captured image, Correction means for correcting a two-dimensional model obtained by projecting the three-dimensional model onto the image plane of the captured image, based on the difference between the first and second features, so that it is displayed superimposed on the real object in the captured image. An information processing device characterized by having the following features. (Configuration 2) The first feature described above is obtained from an image obtained by applying at least one of the following image processing techniques to the image of the two-dimensional model: grayscale conversion, contrast adjustment, and edge detection. The information processing device according to configuration 1, characterized by the above. (Composition 3) The second feature described above is obtained from an image obtained by applying at least one of the following image processing methods to the captured image of the real object: grayscale conversion, contrast adjustment, and edge detection. An information processing device according to configuration 1 or 2, characterized by the above. (Composition 4) The second feature includes features obtained from markers attached to the real object in the captured image. An information processing device according to any one of configurations 1 to 3, characterized by the above. (Composition 5) The first feature described above includes the characteristic that the marker on the three-dimensional model is obtained from information projected onto the image plane of the captured image. The information processing apparatus according to configuration 4, characterized by the features described above. (Composition 6) The second feature includes a predetermined feature point specified by the user for the real object in the captured image. An information processing device according to any one of configurations 1 to 5, characterized by the above. (Composition 7) The first feature is that the predetermined feature point on the three-dimensional model includes a position projected onto the image plane of the captured image. The information processing device according to configuration 6, characterized by the features described therein. (Composition 8) The correction means corrects the two-dimensional model using a correction vector obtained by subtracting the coordinates indicating the position of the first feature corresponding to the second feature from the coordinates indicating the position of the second feature. An information processing device according to any one of configurations 1 to 7, characterized by the above. (Composition 9) The correction means generates a homography transformation matrix based on the correction vector that transforms the coordinates indicating the position of the first feature into coordinates indicating the position of the second feature corresponding to the first feature, and corrects the two-dimensional model using the homography transformation matrix. The information processing apparatus according to configuration 8, characterized by the above. (Composition 10) The correction means predicts and generates the homography transformation matrix for correcting the two-dimensional model in the current captured image based on the correction vector obtained in the past captured image. The information processing apparatus according to configuration 9, characterized by the features described therein. (Composition 11) The correction means generates polygons from a plurality of first features, performs texture mapping on the generated polygons using an image of the 2D model, and corrects the 2D model by moving the vertices of the texture-mapped polygons using the correction vector. The information processing apparatus according to configuration 8, characterized by the above. (Composition 12) The first acquisition means acquires the first features from a second depth image obtained by reprojecting a first depth image of the real object, estimated from the captured image captured at a first time, onto the captured image captured at a second time, based on the position and orientation of the imaging unit that captured the image at a second time, which is later than the first time. An information processing device according to any one of configurations 1 to 11, characterized by the above. (Composition 13) The first acquisition means acquires the three-dimensional position of a feature point acquired from the captured image captured at the first time based on the depth information of the first depth image, and acquires the point obtained by projecting the three-dimensional position of the feature point onto the captured image captured at the second time based on the position and orientation of the imaging unit at the second time, as the first feature. The second acquisition means acquires, in the captured image captured at the second time, a point corresponding to the feature point as the second feature. The information processing device according to configuration 12, characterized by the features described herein. (Composition 14) The imaging unit that captures the image at the second time is different from the imaging unit that captures the image for generating the first depth image at the first time. An information processing device according to configuration 12 or 13, characterized by the above. (method) The first step is to obtain a feature by projecting the features of a 3D model corresponding to a real object in an image of real space onto the image plane of the image, A step of obtaining a second feature corresponding to the first feature from the captured image, The steps include correcting the two-dimensional model obtained by projecting the three-dimensional model onto the image plane of the captured image, based on the difference between the first and second features, so that it appears superimposed on the real object in the captured image. An information processing method characterized by having the following features. (program) A program for causing a computer to function as one of the means of the information processing device described in any of configurations 1 to 14. [Explanation of symbols]

[0121] 101: Information processing unit, 107: Feature acquisition unit, 108: Error acquisition unit, 109: Correction unit, CPU: 1401

Claims

1. A first acquisition means that acquires a first feature by projecting the features of a three-dimensional model corresponding to a real object in an image captured in real space onto the image plane of the captured image, A second acquisition means for acquiring a second feature corresponding to the first feature from the captured image, Correction means for correcting a two-dimensional model obtained by projecting the three-dimensional model onto the image plane of the captured image, based on the difference between the first and second features, so that it is displayed superimposed on the real object in the captured image. An information processing device characterized by having the following features.

2. The first feature described above is obtained from an image obtained by applying at least one of the following image processing techniques to the image of the two-dimensional model: grayscale conversion, contrast adjustment, and edge detection. The information processing apparatus according to feature 1.

3. The second feature described above is obtained from an image obtained by applying at least one of the following image processing methods to the captured image of the real object: grayscale conversion, contrast adjustment, and edge detection. The information processing apparatus according to feature 1.

4. The second feature described above includes features obtained from markers attached to the real object in the captured image. The information processing apparatus according to feature 1.

5. The first feature described above includes the characteristic that the marker on the three-dimensional model is obtained from information projected onto the image plane of the captured image. The information processing apparatus according to feature 4.

6. The second feature described above includes a predetermined feature point specified by the user for the real object in the captured image. The information processing apparatus according to feature 1.

7. The first feature is that the predetermined feature point on the three-dimensional model includes a position projected onto the image plane of the captured image. The information processing apparatus according to feature 6.

8. The correction means corrects the two-dimensional model using a correction vector obtained by subtracting the coordinates indicating the position of the first feature corresponding to the second feature from the coordinates indicating the position of the second feature. The information processing apparatus according to feature 1.

9. The correction means generates a homography transformation matrix based on the correction vector that transforms the coordinates indicating the position of the first feature into coordinates indicating the position of the second feature corresponding to the first feature, and corrects the two-dimensional model using the homography transformation matrix. The information processing apparatus according to feature 8.

10. The correction means predicts and generates the homography transformation matrix for correcting the two-dimensional model in the current captured image based on the correction vector obtained in the past captured image. The information processing apparatus according to feature 9.

11. The correction means generates polygons from a plurality of first features, performs texture mapping on the generated polygons using the image of the two-dimensional model, and corrects the two-dimensional model by moving the vertices of the texture-mapped polygons using the correction vector. The information processing apparatus according to feature 8.

12. The first acquisition means acquires the first features from a second depth image obtained by reprojecting a first depth image of the real object, estimated from the captured image captured at a first time, onto the captured image captured at a second time, based on the position and orientation of the imaging unit that captured the image at a second time, which is later than the first time. The information processing apparatus according to feature 1.

13. The first acquisition means acquires the three-dimensional position of a feature point acquired from the captured image captured at the first time based on the depth information of the first depth image, and acquires the point obtained by projecting the three-dimensional position of the feature point onto the captured image captured at the second time based on the position and orientation of the imaging unit at the second time, as the first feature. The second acquisition means acquires, in the captured image captured at the second time, a point corresponding to the feature point as the second feature. The information processing apparatus according to feature 12.

14. The imaging unit that captures the image at the second time is different from the imaging unit that captures the image for generating the first depth image at the first time. The information processing apparatus according to feature 12.

15. The first step is to obtain a feature by projecting the features of a three-dimensional model corresponding to a real object in an image of real space onto the image plane of the image, A step of obtaining a second feature corresponding to the first feature from the captured image, A step of correcting a two-dimensional model obtained by projecting the three-dimensional model onto the image plane of the captured image, based on the difference between the first and second features, so that it is displayed superimposed on the real object in the captured image. An information processing method characterized by having the following features.

16. A program for causing a computer to function as one of the means of an information processing apparatus described in any one of claims 1 to 14.

Citation Information

Patent Citations

  • Three-dimensional portrait creation device

    JP2013097588A

  • Image processing apparatus, image processing method, and program

    JP2024047799A