Apparatus and method
The apparatus and method align real and virtual space images using transformation matrices and feature point matching to overcome the challenge of creating accurate 3D spatial models from 2D images, achieving precise 3D spatial estimation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NTT DOCOMO INC
- Filing Date
- 2024-11-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing methods struggle to accurately create a 3D spatial model that perfectly matches real space from a 2D image, as determining the correspondence between 2D and 3D coordinates is challenging.
An apparatus and method that includes a real-space image acquisition unit, a virtual-space image acquisition unit, a generation unit for a position transformation matrix, and a position adjustment unit to align objects in real and virtual space images, using feature point extraction and matching, and transformation parameters to accurately convert real-space images into virtual-space images.
Enables accurate conversion of real-space images into virtual-space images, allowing precise alignment and estimation of 3D spatial models from 2D images.
Smart Images

Figure 2026088671000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus and method for handling three-dimensional spatial models. [Background technology]
[0002] There is a need for technology to estimate 3D space based on 2D images. To do this, it is necessary to determine the positional relationship between an image of real space captured by a camera and a corresponding 3D virtual space. There is a technology that determines the positional relationship between the camera image and 3D space by placing objects with known positions, such as markers, in real space and capturing an image that includes the markers. Furthermore, for example, Patent Document 1 below describes a technology related to learning a model for estimating the position and orientation of an object from an image of real space. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-81081 [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] As mentioned above, in order to estimate a 3D space from a 2D image, it is necessary to determine various parameters related to the projection of the 3D space onto the 2D image, as well as the correspondence between coordinates in the 2D image and coordinates in the 3D space. However, even if this correspondence can be determined, it is difficult to create a 3D spatial model that perfectly matches real space.
[0005] Therefore, the present invention has been made in view of the above problems, and aims to provide an apparatus and method that can match a three-dimensional spatial model with real space. [Means for solving the problem]
[0006] To solve the above problems, the system includes: a real-space image acquisition unit that acquires real-space images captured by a real-space camera; a virtual-space image acquisition unit that acquires virtual-space images captured by a virtual camera installed at a position in the virtual space corresponding to the position of the real-space camera, representing a virtual space represented by a virtual-space model corresponding to the real space; a generation unit that generates a position transformation matrix for alignment from the reference image coordinates of the real-space image and the virtual reference image coordinates; and a position adjustment unit that uses the position transformation matrix to align objects included in the real-space image with the virtual-space image. [Effects of the Invention]
[0007] According to this disclosure, it is possible to accurately convert real-space images captured by a camera that images real space into virtual-space images of a virtual space. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram showing the functional configuration of the information processing system and information processing apparatus of this embodiment. [Figure 2] This figure shows an example of generating a virtual space model from a real-world image. [Figure 3] Figure 3(a) shows a first example of setting up a virtual camera in a virtual space. Figure 3(b) shows a second example of setting up a virtual camera in a virtual space. [Figure 4] This figure shows examples of virtual space images and real space images, as well as examples of feature point extraction and matching processes. [Figure 5] This figure shows an example of the process for calculating reference space coordinates, which are the coordinates of reference feature points in a virtual space. [Figure 6] This figure shows an example of reinstalling a virtual camera. [Figure 7] This figure shows an example of how to superimpose the imaging areas between the virtual space image from a specific virtual camera and the virtual space image from a re-installed virtual camera. [Figure 8] It is a flowchart showing the processing content of an information processing method in an information processing apparatus. [Figure 9] It is a diagram showing the configuration of an information processing program. [Figure 10] It is a diagram showing the real camera 100 arranged in the real space of the present disclosure, and the obtained real space image rp and virtual space vs. [Figure 11] It is a diagram showing a virtual space image vp. [Figure 12] It is a diagram showing an overview of the operation of the imaging position determination process of the virtual camera of the present disclosure and the subsequent affine transformation process. [Figure 13] It is a block diagram showing the functional configuration of the information processing apparatus 10a. [Figure 14] It is a diagram showing the functional configuration of a conversion matrix generation unit 101 that obtains an affine conversion matrix for each division unit. [Figure 15] It is a diagram showing a specific example of an affine conversion matrix. [Figure 16] It is a flowchart showing the operation of the conversion matrix generation unit 101 in the information processing apparatus 10a of the present disclosure. [Figure 17] It is a diagram showing divided images of different division units. [Figure 18] It is a diagram showing an enlarged divided image obtained by enlarging a divided image. [Figure 19] It is a diagram showing that an averaging process is performed on the output value from an affine conversion matrix. [Figure 20] It is a hardware block diagram of an information processing apparatus.
Mode for Carrying Out the Invention
[0009] An embodiment of an information processing system according to the present invention will be described with reference to the drawings. In addition, when possible, the same parts are denoted by the same reference numerals, and duplicate explanations are omitted.
[0010] Figure 1 shows the functional configuration of the information processing system and information processing device according to this embodiment. The information processing system 1 of this embodiment is a system that obtains information representing the position in virtual space of a camera that images real space in order to estimate a three-dimensional space from a two-dimensional image without using installed objects such as markers, and is configured as an information processing device 10 as an example.
[0011] As shown in Figure 1, the information processing device 10 functionally comprises a real-space image acquisition unit 11, a setting unit 12, a virtual-space image acquisition unit 13, a feature point extraction unit 14, a feature point matching unit 15, a parameter calculation unit 16, an output unit 17, a reset unit 18, and a position calculation unit 19. These functional units 11 to 19 may be configured in a single device as illustrated in Figure 1, or they may be distributed across multiple devices.
[0012] Each of the functional units 11 to 19 of the information processing device 10 is configured to access storage means such as the real-space image storage unit 21 and the virtual-space model storage unit 22. Each of the storage units 21 to 22 may be provided in the information processing device 10, or it may be configured in another device that is accessible from the information processing device 10, as illustrated in Figure 1.
[0013] Next, the various functions of the information processing device 10 will be described. The real-space image acquisition unit 11 acquires real-space images captured in real space. Specifically, the real-space image acquisition unit 11 may acquire real-space images from a camera that captures real space, or it may acquire real-space images from the real-space image storage unit 21, which is a storage means that stores captured real-space images.
[0014] Figure 2 shows an example of generating a virtual space model from a real-world image. As shown in Figure 2, the real-world image acquisition unit 11 acquires a real-world image from the real-world space rs. Also, as shown in Figure 2, a virtual space model vm representing the virtual space corresponding to the real-world space rs is generated based on the real-world space rs.
[0015] The virtual space model (vm) may be generated based on the real-world image (rp) by any known method. For example, in real-world image acquisition, the virtual space model (vm) may be generated based on the distance to an object measured by LiDAR (Light Detection and Ranging). Alternatively, the virtual space model (vm) may be generated by estimating the depth of what is represented in each pixel of the two-dimensional real-world image (rp) using a known technique, converting the depth and color value of each pixel into point cloud data, and generating a three-dimensional virtual space by a three-dimensional display based on the point cloud data.
[0016] The generated virtual space model (vm) may be stored in the virtual space model storage unit 22. The virtual space model storage unit 22 is a storage means for storing pre-generated virtual space models (vm).
[0017] The setting unit 12 places a virtual camera in the virtual space represented by the virtual space model vm. Specifically, the setting unit 12 sets the position of the virtual camera in the virtual space represented by the virtual space model vm. The position of the virtual camera is the viewpoint position for capturing a virtual space image, which is an image projected from the virtual space model vm.
[0018] Figure 3 shows an example of the placement of virtual cameras in a virtual space, where Figures 3(a) and 3(b) show first and second examples of the placement positions of virtual cameras in the virtual space, as viewed from above the virtual space vs., respectively. As shown in Figures 3(a) and 3(b), the setting unit 12 may set up multiple virtual cameras vc in the virtual space vs.
[0019] In the first example shown in Figure 3(a), the setting unit 12 may install multiple virtual cameras vc along the outer perimeter of the virtual space vs, with the normal direction of the outer perimeter as the imaging direction. Alternatively, as shown in Figure 3(b), the setting unit 12 may install multiple virtual cameras vc within the virtual space vs (for example, around the central part of the virtual space vs), with the direction of the outer perimeter of the virtual space vs as the imaging direction. By installing multiple virtual cameras vc in this way, the virtual space corresponding to the space represented in the real-world image rp can be comprehensively imaged.
[0020] The virtual space image acquisition unit 13 acquires at least one virtual space image of the virtual space vs using the virtual camera vc. The virtual space image acquisition unit 13 may acquire a virtual space image using a known method based on the virtual space vs represented by the virtual space model vm.
[0021] Figure 4 shows examples of virtual space images and real space images, as well as examples of feature point extraction and matching processes. Specifically, the virtual space image acquisition unit 13 acquires a virtual space image vp by projecting the virtual space vs onto a virtual screen, using the position of the virtual camera vc set in the virtual space vs, which is represented based on the virtual space model vm, as the viewpoint position.
[0022] The feature point extraction unit 14 extracts feature points from the real-space image rp, which is an image of the real-space rs, and at least one virtual-space image vp. The feature point matching unit 15 matches the feature points of the real-space image rp with the feature points of the virtual-space image vp. The feature point extraction unit 14 and the feature point matching unit 15 may use known methods for feature point extraction and feature point matching, for example, they may use known methods such as SIFT and AKAZE.
[0023] In the example shown in Figure 4, the feature point extraction unit 14 extracts feature points vfp1 and vfp2 from the virtual space image vp. The feature point extraction unit 14 also extracts feature points rfp1 and rfp2 from the real space image rp. Feature points are extracted, for example, based on the difference in pixel values of adjacent pixels in the image. For example, endpoints and corners of reference objects represented in the image, such as a door or a desk, are extracted as feature points.
[0024] Furthermore, the feature point matching unit 15 matches each of the feature points vfp1 and vfp2 of the virtual space image vp with each of the feature points rfp1 and rfp2 of the real space image rp, based on the feature quantities of each feature point, as shown by the code fm.
[0025] The parameter calculation unit 16 calculates transformation parameters related to imaging in the real space by substituting the reference space coordinates, which are the three-dimensional coordinates of the reference feature point in the virtual space vs, and the reference image coordinates, which are the coordinates of the reference feature point in the real space image rp, into a predetermined transformation formula.
[0026] A transformation formula is an expression that describes the relationship between the three-dimensional coordinates of a specific point in three-dimensional space and the two-dimensional coordinates of that specific point in an image captured from that three-dimensional space, using predetermined transformation parameters. If the three-dimensional coordinates in three-dimensional space are (X, Y, Z) and the two-dimensional coordinates in the two-dimensional image are (u, v), then an example of a predetermined transformation formula is expressed by the following transformation formula (1).
number
[0027] Then, if we let t be the 3D reference space coordinates and d' be the 2D virtual reference image coordinates, which are the reference feature points in the virtual space image vp, and if A is the transformation matrix that transforms the reference space coordinates t to the virtual reference image coordinates d', then the relationship between these coordinates is expressed by the following equation (2). d'=tA ···(2) Furthermore, if we denote the two-dimensional reference image coordinates as d, and the transformation matrix B is expressed by the transformation parameters of transformation equation (1) and transforms the reference spatial coordinates t to the reference image coordinates d, then the relationship between these coordinates can be expressed as shown in equation (3) below. d = tB ... (3) As shown in equation (2), the transformation matrix A is known because it is a transformation matrix (transformation parameter) that shows the relationship between the 3D coordinates indicating the position in the virtual space vs and the 2D coordinates of that position in the image captured by the virtual camera vc placed in the virtual space vs. Therefore, the parameter calculation unit 16 can calculate the reference space coordinates t based on the transformation matrix A and the virtual reference image coordinates d'.
[0028] Specifically, the parameter calculation unit 16 may calculate the reference spatial coordinates t using a technique known as collision detection. Figure 5 schematically shows an example of the process of calculating reference spatial coordinates using collision detection.
[0029] As shown in Figure 5, since the transformation matrix A is known, the position of the virtual camera vc is also known, so the parameter calculation unit 16 can calculate a three-dimensional straight line ln that passes through the virtual reference image coordinates d' of the reference feature point fp in the virtual space image vp. The parameter calculation unit 16 then calculates the point where the plane (mesh) constituting the virtual space model vm intersects with the straight line ln as the reference space coordinates t of the reference feature point fp.
[0030] The parameter calculation unit 16 calculates the transformation matrix B in equation (3) as a transformation parameter for imaging in real space, based on the reference spatial coordinate t and reference image coordinate d of the reference feature point fp.
[0031] The setting unit 12 may install multiple virtual cameras vc in the virtual space vs, as shown in the example described with reference to Figures 3(a) and 3(b). The virtual space image acquisition unit 13 then acquires multiple virtual space images vp acquired by each of the multiple virtual cameras vc.
[0032] The parameter calculation unit 16 may calculate the transformation parameters for imaging in real space based on at least one virtual space image vp selected from among a plurality of virtual space images vp based on the number of reference feature points fp.
[0033] Since the number of transformation parameters in the transformation equation (1) that constitutes the transformation matrix B is 15, the parameter calculation unit 16 constructs the transformation equation (1) using the reference spatial coordinates t and reference image coordinates d of at least 8 reference feature points fp, and calculates the transformation parameters by solving the multiple transformation equations that have been constructed as equations. Accordingly, the parameter calculation unit 16 may calculate the transformation parameters related to imaging in real space based on a virtual space image vp having a number of reference feature points fp equal to or greater than a predetermined threshold.
[0034] Specifically, the parameter calculation unit 16 may calculate transformation parameters related to imaging in the real space using a virtual space image vp having 8 or more reference feature points fp. Alternatively, the parameter calculation unit 16 may calculate transformation parameters related to imaging in the real space using multiple virtual space images vp such that the total number of reference feature points fp is equal to or greater than a predetermined threshold.
[0035] Furthermore, the parameter calculation unit 16 may calculate transformation parameters related to imaging in the real space based on each of the multiple virtual space images vp. This calculates transformation parameters associated with each virtual space image vp.
[0036] The output unit 17 outputs the transformation parameters (corresponding to the transformation matrix B) related to imaging in real space, calculated by the parameter calculation unit 16, as camera information that substantially represents the position in virtual space of the camera that captured the real-space image.
[0037] Furthermore, the output unit 17 may output as camera information a transformation parameter obtained by processing the calculated transformation parameters using a predetermined statistical method. Specifically, the output unit 17 may output as camera information a transformation parameter obtained by processing the calculated transformation parameters using the least squares method or the like. In this way, by obtaining a transformation parameter that is finally output by statistically processing each transformation parameter calculated based on the virtual space image vp acquired by each of the multiple virtual cameras vc, it becomes possible to improve the accuracy of the transformation parameter output as camera information.
[0038] The resetting unit 18 applies the conversion parameters for imaging in the real space, calculated by the parameter calculation unit 16, to the virtual camera vc, which has been applied as conversion parameters for imaging in the virtual space vs, and then reinstalls the virtual camera vc in the virtual space vs.
[0039] When initially setting up the virtual camera vc and acquiring the virtual space image vp, the transformation parameters applied to the virtual camera vc may be set arbitrarily. After the transformation parameters for imaging the real space are calculated based on the initial virtual space image vp, the calculated transformation parameters can be applied to the virtual camera vc. Once the calculated transformation parameters are applied to the virtual camera vc, it is highly likely that it will have a positional relationship and characteristics similar to the camera that captured the real space image rp.
[0040] The virtual space image acquisition unit 13 acquires a virtual space image vp captured by a virtual camera vc that has been reinstalled in the virtual space vs. The feature point extraction unit 14 extracts feature points from the virtual space image vp captured by the reinstalled virtual camera vc, and the feature point matching unit 15 performs feature point matching. Then, the parameter calculation unit 16 calculates conversion parameters based on the virtual space image vp captured by the reinstalled virtual camera vc. In this way, the accuracy can be improved by recalculating the conversion parameters based on the virtual space image vp captured by the virtual camera vc to which the calculated conversion parameters have been applied.
[0041] The resetting unit 18 may reinstall at least one virtual camera vc within a predetermined range around a specific virtual camera vc, which is one of the virtual cameras vc set by the setting unit 12 that captured the virtual space image vp having the most reference feature points fp.
[0042] Figure 6 shows an example of the re-installation of a virtual camera. In the example shown in Figure 6, the re-installation unit 18 extracts a specific virtual camera vc1, which is the virtual camera that captured the virtual space image vp having the most reference feature points fp among the virtual cameras vc installed in the virtual space vs during the initial or previous calculation of the transformation parameters. Since the specific virtual camera vc1 is the virtual camera that captured the virtual space image vp having many feature points that match the feature points in the real space image rp, it is highly probable that it is a camera installed at a position in the virtual space vs that corresponds to the position of camera rc that captured the real space image rp.
[0043] The resetting unit 18 reinstalls the virtual camera vc within a predetermined range around the specific virtual camera vc1. The resetting unit 18 may, for example, reinstall the virtual camera at a position within a predetermined distance from the position of the specific virtual camera vc1. As illustrated in Figure 6, the resetting unit 18 may reinstall virtual cameras vc2 and vc3 at positions adjacent to the specific virtual camera vc1 along the outer perimeter of the virtual space vs.
[0044] In this way, by recalculating the transformation parameters based on the virtual space image vp derived from a virtual camera vc that has been reinstalled within a predetermined range around a specific virtual camera vc1, it becomes possible to further improve the accuracy.
[0045] Furthermore, the resetting unit 18 may re-install at least one virtual camera vc so that a virtual space image vp is captured in which a portion of the virtual space image vp captured by the specific virtual camera vc1 overlaps with a portion of the virtual space image vp that includes the reference feature point fp. Figure 7 shows an example of how the imaging areas between the virtual space image captured by the specific virtual camera and the virtual space image captured by the re-installed virtual camera are superimposed.
[0046] As illustrated in Figure 7, when a virtual space image vp having an imaging region ts0 is captured by a specific virtual camera vc1, the resetting unit 18 repositions the virtual camera vc to capture an imaging region that is superimposed on the imaging region ts0 and a portion of the region including the reference feature point fp. Specifically, the resetting unit 18 repositions the virtual camera vc to a position where imaging regions ts1, ts2, ts3, and ts4, which include the region containing the reference feature point fp located at the edge of the imaging region ts0, are captured as the virtual space image vp. By acquiring the virtual space image vp with the virtual camera vc repositioned in this way, it becomes possible to improve the accuracy of the conversion parameters related to lens distortion.
[0047] Referring again to Figure 1, the position calculation unit 19 calculates the position of an object in three-dimensional space using transformation parameters included in the camera information, based on the two-dimensional coordinates that indicate the position of an object such as a person, represented in the real-space image rp captured from the real-space rs.
[0048] In this way, by using the two-dimensional coordinates of an object in a real-space image rp, which is an image of the real space, and a virtual space model vm corresponding to that real space, it becomes possible to estimate the position of the object represented in the real-space image rp in three-dimensional space.
[0049] Figure 8 is a flowchart showing the processing details of the information processing method in the information processing system 1. In step S1, the real-space image acquisition unit 11 acquires a real-space image rp captured of the real space.
[0050] In step S2, a virtual space model vm representing the virtual space vs corresponding to the real space rs is generated based on the real space image rp. The virtual space model vm may be generated by the processor of the information processing device 10 or by another device. The generated virtual space model vm may be stored in the virtual space model storage unit 22.
[0051] In step S3, the setting unit 12 installs a virtual camera vc in the virtual space vs represented by the virtual space model vm. In step S4, the virtual space image acquisition unit 13 acquires at least one virtual space image vp of the virtual space vs captured by the virtual camera vc.
[0052] In step S5, the feature point extraction unit 14 extracts feature points from both the real-space image rp and the virtual-space image vp. In step S6, the feature point matching unit 15 matches the feature points of the real-space image rp with the feature points of the virtual-space image vp.
[0053] In step S7, the parameter calculation unit 16 determines whether the number of reference feature points fp, which are matched feature points, is equal to or greater than a threshold. If it is determined that the number of reference feature points fp is equal to or greater than the threshold, the process proceeds to step S8. On the other hand, if it is not determined that the number of reference feature points fp is equal to or greater than the threshold, the process returns to step S3.
[0054] In step S8, the parameter calculation unit 16 obtains the reference spatial coordinates t of the reference feature point fp based on the three-dimensional coordinates indicating the position in the virtual space vs, the transformation parameters (transformation matrix A) indicating the relationship between the three-dimensional coordinates indicating the position in the virtual space vs and the two-dimensional coordinates of that position in the image captured by the virtual camera vc placed in the virtual space vs, and the virtual reference image coordinates d'.
[0055] In step S9, the parameter calculation unit 16 calculates transformation parameters (transformation matrix B) related to imaging in real space based on the reference spatial coordinates t and the reference image coordinates d.
[0056] In step S10, the parameter calculation unit 16 determines whether or not to terminate the process. For example, if the camera parameter fluctuation falls below a threshold, it is determined that the process will terminate. If it is determined not to terminate the process, that is, to reinstall the virtual camera vc in order to improve the accuracy of the calculated conversion parameters, the process proceeds to step S11. If it is determined to terminate the process, the process proceeds to step S12.
[0057] In step S11, the resetting unit 18 applies the conversion parameters calculated in step S9 to the virtual camera vc. Then, the process returns to step S3 in order to repeat the calculation of conversion parameters based on the reinstallation of the virtual camera vc. If the process returns to step S3 after going through step S11, the resetting unit 18 reinstalls the virtual camera vc.
[0058] In step S12, the output unit 17 outputs the transformation parameters (transformation matrix B) related to imaging in real space, calculated by the parameter calculation unit 16, as camera information that substantially represents the position in virtual space of the camera that captured the real-space image rp.
[0059] Next, with reference to Figure 9, an information processing program for causing a computer to function as the information processing device 10 of this embodiment will be described. Figure 9 is a diagram showing the configuration of the information processing program. The information processing program P1 is composed of a main module m10 that comprehensively controls information processing in the information processing device 10, a real-space image acquisition module m11, a setting module m12, a virtual-space image acquisition module m13, a feature point extraction module m14, a feature point matching module m15, a parameter calculation module m16, an output module m17, a reset module m18, and a position calculation module m19. Each of the modules m11 to m19 realizes the respective functions for each of the functional units 11 to 19.
[0060] The information processing program P1 may be transmitted via a transmission medium such as a communication line, or it may be stored in a recording medium M1, as shown in Figure 9.
[0061] According to the information processing system, information processing device 10, information processing method, and information processing program P1 of this embodiment described above, feature points extracted from the real-space image and the virtual-space image are matched as reference feature points, and by associating the two-dimensional reference image coordinates of the reference feature points in the real-space image with the three-dimensional reference space coordinates of the reference feature points in the virtual space, it becomes possible to calculate the transformation parameters in the transformation formula. Since the calculated transformation parameters constitute information that represents the relative relationship between the position in the real-space image and the position in the virtual space, it becomes possible to obtain information that represents the position in the virtual space of a camera that is essentially imaging the real space using camera information consisting of transformation parameters.
[0062] The information processing apparatus and information processing method relating to this disclosure may have the following configurations. The operation and effects of each configuration are described below.
[0063] An information processing device relating to one aspect of this disclosure is a feature point extraction unit that extracts feature points from a real-space image captured of real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to real space and represented by a virtual-space model, and the position of the virtual camera is a viewpoint position for capturing an image projected from the virtual-space model, the feature point extraction unit, the feature point matching unit that matches feature points of the real-space image with feature points of the virtual-space image, the reference space coordinates in the virtual space of the reference feature points which are the matched feature points, and the reference image coordinates which are the coordinates of the reference feature points in the real-space image A parameter calculation unit that calculates transformation parameters related to imaging in real space by substituting into a predetermined transformation formula that expresses the relationship between the 3D coordinates of a specific point in 3D space and the 2D coordinates of the specific point in an image captured in 3D space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on the transformation parameters related to imaging in virtual space by a virtual camera and the virtual reference image coordinates which are the coordinates of a reference feature point in the virtual space image; and an output unit that outputs the transformation parameters related to imaging in real space as camera information representing the position in virtual space of the camera that captured the real space image.
[0064] An information processing method relating to one aspect of this disclosure is a feature point extraction step performed by a processor, which extracts feature points from a real-space image captured of real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to real space and represented by a virtual-space model, and the position of the virtual camera is a viewpoint position for capturing an image projected from the virtual-space model; a feature point matching step that matches feature points of the real-space image with feature points of the virtual-space image; and the reference space coordinates in the virtual space of the reference feature points which are the matched feature points and the coordinates of the reference feature points in the real-space image. A parameter calculation step for calculating transformation parameters related to imaging in real space by substituting quasi-image coordinates into a predetermined transformation formula that expresses the relationship between the 3D coordinates of a specific point, which is a specific point in 3D space, and the 2D coordinates of the specific point in an image captured in 3D space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on transformation parameters related to imaging in virtual space by a virtual camera and virtual reference image coordinates, which are the coordinates of a reference feature point in the virtual space image; and an output step for outputting the transformation parameters related to imaging in real space as camera information representing the position in virtual space of the camera that captured the real space image.
[0065] Based on the above aspects, feature points extracted from both the real-space image and the virtual-space image are matched as reference feature points, and by associating the two-dimensional reference image coordinates of the reference feature points in the real-space image with the three-dimensional reference space coordinates of the reference feature points in the virtual space, it becomes possible to calculate the transformation parameters in the transformation formula. Since the calculated transformation parameters constitute information that represents the relative relationship between the position in the real-space image and the position in the virtual space, it becomes possible to obtain information that represents the virtual space position of the camera that is essentially capturing the real space using camera information consisting of transformation parameters.
[0066] Furthermore, an information processing device relating to other aspects may further include a setting unit for installing multiple virtual cameras, and a parameter calculation unit may calculate conversion parameters for imaging the real space based on at least one virtual space image selected from among multiple virtual space images acquired by each of the multiple virtual cameras based on the abundance of reference feature points.
[0067] Based on the above aspects, the virtual space image selected based on the abundance of reference feature points among multiple virtual space images is likely to be an image captured by a virtual camera positioned close to the location of the camera that captured the real-world image. Therefore, the accuracy can be improved by calculating transformation parameters based on such a virtual space image.
[0068] Furthermore, in information processing devices relating to other aspects, the setting unit may install at least multiple virtual cameras along the outer perimeter of the virtual space with the normal direction of the outer perimeter as the imaging direction, or install multiple virtual cameras within the virtual space with the direction of the outer perimeter of the virtual space as the imaging direction.
[0069] Based on the above aspects, multiple virtual cameras can comprehensively capture images of the virtual space corresponding to the space represented in the real-world image.
[0070] Furthermore, in information processing devices relating to other aspects, the parameter calculation unit may calculate transformation parameters for imaging in real space based on a virtual space image having a number of reference specified points equal to or greater than a threshold for reference feature points.
[0071] Based on the above aspects, by setting a threshold corresponding to the number of unknown transformation parameters in the transformation formula, the transformation parameters can be reliably determined.
[0072] Furthermore, an information processing device relating to other aspects may further include a resetting unit that resets a virtual camera in the virtual space to which the transformation parameters for imaging the real space calculated by the parameter calculation unit have been applied as transformation parameters for imaging the virtual space, and the feature point extraction unit may extract feature points from a virtual space image captured by the resetting virtual camera.
[0073] Based on the above aspects, a virtual camera to which the calculated transformation parameters have been applied is highly likely to have positional relationships and characteristics similar to those of the camera that captured the real-world image. By recalculating the transformation parameters based on the virtual-world image captured by the virtual camera to which the calculated transformation parameters have been applied, it is possible to improve their accuracy.
[0074] Furthermore, in information processing devices relating to other aspects, the resetting unit may reinstall at least one virtual camera within a predetermined range around a specific virtual camera, which is a virtual camera that captured the virtual space image having the most reference feature points.
[0075] Based on the above aspects, it is highly probable that the specific virtual camera is a camera placed in the virtual space close to the camera that captured the real-world image. Therefore, by recalculating the transformation parameters based on the virtual space image obtained from a virtual camera reinstalled within a predetermined range around the specific virtual camera, it is possible to further improve its accuracy.
[0076] Furthermore, in information processing devices relating to other aspects, the resetting unit may reconfigure at least one virtual camera so that a virtual space image is captured in which a portion of the virtual space image captured by a specific virtual camera overlaps with a portion of the region containing reference feature points.
[0077] According to the above aspects, a virtual space image is acquired by a reinstalled virtual camera that includes reference feature points located at the edges of the imaging area of the virtual space image captured by a specific virtual camera. This makes it possible to improve the accuracy of the transformation parameters related to lens distortion.
[0078] Also, in the information processing apparatus according to another aspect, the parameter calculation unit calculates conversion parameters related to imaging in the real space based on each of a plurality of virtual space images, and the output unit outputs, as camera information, the conversion parameters obtained by processing the calculated plurality of conversion parameters by a predetermined statistical method.
[0079] According to the above aspect, conversion parameters that are calculated based on virtual space images obtained by each of a plurality of virtual cameras and are finally output after being statistically processed are obtained. Thereby, it becomes possible to improve the accuracy of the conversion parameters output as camera information.
[0080] Also, in the information processing apparatus according to another aspect, the conversion parameters include six predetermined external camera parameters related to the projection of a three-dimensional space onto a two-dimensional image, four internal camera parameters, and five parameters related to the lens distortion of the camera. The parameter calculation unit uses the following conversion formula (1) that represents the relationship between the three-dimensional coordinates (X, Y, Z) in the three-dimensional space and the two-dimensional coordinates (u, v) in the two-dimensional image [Number] wherein, by substituting the three-dimensional reference space coordinates into the three-dimensional coordinates (X, Y, Z) and substituting the two-dimensional reference image coordinates into the two-dimensional coordinates (u, v), the external camera parameters R, t having three degrees of freedom, the internal camera parameters f x , f y , c x , c y , and the parameters k1, k2, k3, p1, p2 related to the lens distortion of the camera may be calculated.
[0081] According to the above aspect, conversion parameters consisting of 15 variables can be calculated as camera information.
[0082] Next, we will describe a process that accurately projects objects such as people included in real-space images captured by a real-space camera onto a virtual-space image vp of a virtual space composed of a virtual-space model, using information (conversion parameters) that represent the position of the camera capturing the real-space image in the virtual space. The purpose of this disclosure is to accurately project objects (such as people) in the real space captured by a real-space camera (hereinafter referred to as the real-space camera) onto a virtual-space model (virtual-space image).
[0083] Figure 10 shows a real-world camera 100 placed in the real-world space of this disclosure, and the real-world image rp and virtual space vs obtained therefrom. The real-world camera 100 corresponds to camera rc in the above. In this disclosure, it is necessary that the reference space coordinates (or virtual reference image coordinates of the virtual-world image vp), which are the three-dimensional coordinates of the reference feature points in the virtual-world image vp, and the reference image coordinates, which are the coordinates of the said reference feature points in the real-world image rp, correspond precisely. However, slight discrepancies may occur in the virtual-world model that provides the virtual space.
[0084] Figure 11 shows a virtual space image vp. In the figure, the door S, indicated by the thick border in the real world, does not coincide with the edge of the door in the virtual space image vp. That is, there is a slight discrepancy between the edge of the door (the thick bordered area) represented in the real space image rp and the edge of the door in the virtual space image vp. The virtual space image vp is an image captured by a virtual camera positioned according to the above process, but due to the characteristics of the virtual space model, there may be a slight discrepancy with the real space image rp. Therefore, if the real space image rp includes objects such as people, those objects will also be represented with a discrepancy in the virtual space image vp.
[0085] When such a discrepancy occurs, if an object such as a person is imaged in real space and then projected into virtual space, a discrepancy will occur. In this disclosure, an affine transformation process is performed on an object in real space to absorb the slight discrepancy that occurs when an object such as a person in real space is imaged by a real camera 100 in real space and then projected into virtual space.
[0086] Figure 12 is a diagram illustrating the operation overview of the virtual camera imaging position determination process and subsequent affine transformation process of the present disclosure. The real space image acquisition unit 11 acquires a real space image rp (S101), and the virtual space image acquisition unit 13 acquires a virtual space image vp (S102). The feature point extraction unit 14 extracts feature points from the real space image rp and the virtual space image vp, respectively, and the feature point matching unit 15 performs feature point matching (S103). Then, the parameter calculation unit 16 performs the transformation parameter calculation process (S104).
[0087] Then, if the parameter calculation unit 16 determines that the variation in the transformation parameters is below a certain level, the process proceeds to S106, where the virtual camera position and focal length are changed according to the transformation parameters, and a virtual space image vp is obtained (S106, S102). This process is repeated until the variation in the transformation parameters falls below a certain level (threshold) (S105). Once the variation in the transformation parameters falls below a certain level, the transformation matrix generation unit 101 and the position adjustment unit 102 perform affine transformation processing (S107). Processes S101 to S106 described above are schematic representations of the processes shown in Figure 8, and in substance, the processes are carried out according to Figure 8.
[0088] In this way, by calculating appropriate transformation parameters for the virtual camera, the position of the real camera 100 in the virtual space can be determined, and the virtual space image vp can be obtained using this position as the virtual camera's position. Then, by compositing the real space image rp (containing the objects) of the real camera 100, which has undergone an affine transformation, onto the virtual space image vp, accurate alignment between the virtual space image vp and the real space image rp (containing the objects) can be achieved.
[0089] The functional configuration of the information processing device 10a for performing the above operations will now be described. Figure 13 is a block diagram showing the functional configuration of the information processing device 10a. As shown in the figure, the information processing device 10a includes a real-space image acquisition unit 11, a setting unit 12, a virtual-space image acquisition unit 13, a feature point extraction unit 14, a feature point matching unit 15, a parameter calculation unit 16, an output unit 17, a reset unit 18, a position calculation unit 19, a transformation matrix generation unit 101, and a position adjustment unit 102. In addition to the functions of the information processing device 10 for obtaining the transformation parameters described above, this information processing device 10a includes a transformation matrix generation unit 101 for generating an affine transformation matrix and a position adjustment unit 102 for performing the affine transformation.
[0090] The transformation matrix generation unit 101 is responsible for generating affine transformation matrices. This transformation matrix generation unit 101 generates affine transformation matrices based on the feature points (reference image coordinates) of the real-space image rp and the corresponding feature points (virtual reference image coordinates) of the virtual-space image vp. Further details will be described later.
[0091] The position adjustment unit 102 synthesizes objects in the real space captured by the real camera 100 into a virtual space image, and fine-tunes the position of the objects using an affine transformation matrix. This virtual space image vp is an image captured from the position of the virtual camera calculated by the position calculation unit 19. The objects are people, etc., obtained from the real space image rp captured by the real camera 100, for example, the skeleton of the object extracted by OpenPose. OpenPose is executed by the real space image acquisition unit 11.
[0092] OpenPose is a known technique that can detect the positions of joints or skeletons (hereinafter referred to as skeletons or bones) of a human body from an image or video and estimate a person's pose based on this. This technique is used in various fields such as sports motion analysis, dance choreography, and interactive applications. In this disclosure, assuming that the real-world camera 100 is a surveillance camera, the aim is to reflect the skeletons of people and other objects contained in the real-world image onto a virtual-world image. In this disclosure, OpenPose is used as the technique for extracting the skeletons of objects, but it is not limited to this. For example, MediaPipe, PoseNet, HRNet (High-Resolution Network), or AlphaPose may be used.
[0093] This information processing device 10a can determine the position of a virtual camera in the virtual space based on the real-space image captured by the real-space camera 100, and generate an affine transformation matrix based on the difference between the virtual space image (feature points) obtained by capturing based on that position and the real-space image (feature points). Then, by inputting the skeleton of the object captured by the real-space camera 100 into the generated affine transformation matrix, the precise coordinates on the virtual space image vp can be determined. This affine transformation matrix can absorb the difference between the real-space image and the virtual space image.
[0094] Furthermore, in this disclosure, the transformation matrix generation unit 101 divides the real space image and the virtual space image into predetermined division units and obtains an affine transformation matrix for each division unit. Figure 14 shows the functional configuration of the transformation matrix generation unit 101 that obtains an affine transformation matrix for each division unit.
[0095] This transformation matrix generation unit 101 is composed of an image segmentation unit 101a, a feature point verification unit 101b, an affine transformation matrix generation unit 101c, and an affine transformation matrix storage unit 101d.
[0096] The image division unit 101a is the part that divides the real-space image rp and the virtual-space image vp into predetermined division units. In this disclosure, the image division unit 101a divides the real-space image and the virtual-space image into a 3x3 division unit as the first division unit and a 4x4 division unit as the second division unit to obtain multiple divided images. Of course, the image division unit 101a may also divide into division units other than those specified above.
[0097] The feature point verification unit 101b checks whether each divided image contains a predetermined number of feature points and enlarges the divided image until the predetermined number of feature points is reached. In this disclosure, the feature point verification unit 101b enlarges the divided image until there are three or more feature points. As described above, feature points are reference feature points, and the feature point verification unit 101b enlarges the image so that there are three or more virtual reference image coordinates d' and reference image coordinates d, respectively. Since an affine transformation matrix generally contains six variables, at least three feature points are required to solve it.
[0098] The affine transformation matrix generation unit 101c is responsible for generating an affine transformation matrix for each segmented image. The affine transformation matrix generation unit 101c calculates the affine transformation matrix based on the virtual reference image coordinates and reference image coordinates of each segmented image. Since the shift from the virtual space image vp differs depending on the position in the real space image rp, it is preferable to calculate an affine transformation matrix that corresponds to that position.
[0099] The affine transformation matrix storage unit 101d is the part that stores the generated affine transformation matrix in association with the divided image. For example, the affine transformation matrix generation unit 101c assigns an identifier to the divided image and generates an affine transformation matrix in association with each identifier.
[0100] The position adjustment unit 102 (see Figure 13) adjusts the position of the composite object in the virtual space image by placing the object in the real space image captured by the real camera 100 into an affine transformation matrix. In this disclosure, the output value of the affine transformation matrix is a value that has been adjusted to match the shift in the virtual space model, specifically the position of the object in the real space image.
[0101] In this disclosure, as described above, the position of the virtual camera in the virtual space is determined by the transformation parameters. This position is the same as the position of the real camera 100 in the real space. Therefore, the real space image rp and the virtual space image vp are roughly identical. However, due to the characteristics of the virtual space model, the real space image rp and the virtual space image vp are not exactly the same, so an affine transformation matrix is used to fine-tune the discrepancy. In this disclosure, the skeleton of an object in the real space captured by the real camera 100 is extracted, and by performing an affine transformation on that skeleton, the object (skeleton) can be synthesized at an appropriate position in the virtual space image.
[0102] Figure 15 shows a concrete example of an affine transformation matrix. This affine transformation matrix can output the virtual reference image coordinates (u',v') of a virtual space image vp by taking the reference image coordinates (u,v) of a real space image rp as input. Here, a, b, c, and d are parameters that indicate the rotation of the image, and t x t y This parameter indicates translation.
[0103] The transformation matrix generation unit 101 can obtain the affine transformation matrix by solving the affine transformation matrix shown in Figure 15 using at least three reference image coordinates (u,v) of the real-space image rp and the corresponding virtual reference image coordinates (u',v') of the virtual-space image vp. That is, the above parameters a, b, c, d, t x t y It is possible to find this.
[0104] In this disclosure, affine transformation is a process for correcting a virtual space model that contains errors. Generally, since affine transformation uses six variables, it is possible to solve it as a general polynomial if there are at least three matching feature points.
[0105] Next, the affine transformation process of the information processing device 10a of this disclosure (processing S107 in Figure 12) will be described. Figure 16 is a flowchart showing the operation of the transformation matrix generation unit 101 in the information processing device 10a of this disclosure.
[0106] In the process S105 shown in Figure 12, once the transformation parameters for the transformation matrix B are obtained and the position of the virtual camera is identified, the image division unit 101a divides the virtual space image vp obtained by the virtual camera at that position and the real space image rp captured by the real camera 100 into a first division unit and a second division unit, respectively (S201). The feature point count verification unit 101b checks the number of feature points contained in each of the divided images of the real space image rp and the virtual space image vp (for example, the divided images rdv311, vdv311, rdv411, vdv411 in Figure 17) (S202). These feature points are those obtained in the process S5 shown in Figure 8. The feature point count verification unit 101b then determines whether each divided image contains a predetermined number (for example, 3) or more feature points (S203).
[0107] The feature point count verification unit 101b determines that the divided image does not have a predetermined number of feature points or more, and then processes the divided image to enlarge it and obtain the enlarged divided image rdv311e (S204). For example, in Figure 18, the divided image rdv311, which is a 3x3 division unit, is enlarged on the left and bottom sides. This enlargement process is carried out until the number of feature points exceeds a predetermined number. The same process is performed on other divided images. The direction and width of the enlargement can be arbitrary, but for example, the divided image rdv311 may be enlarged by 10% vertically and horizontally in the downward and right directions.
[0108] The affine transformation matrix generation unit 101c generates an affine transformation matrix for each segmented image using the feature points (reference image coordinates d(u,v) and virtual reference image coordinates d'(u',v')) in the segmented image (or the enlarged segmented image) and stores it in the affine transformation matrix storage unit 101d (S205). The affine transformation matrix generated here is associated with each segmented image before enlargement, not with the enlarged segmented image. It is generated using feature points contained in the enlarged segmented image, but in that case, the segmented image before enlargement does not have feature points (for example, a white wall), so it is based on the idea that it is better to use feature points in the vicinity of that area.
[0109] The feature point verification unit 101b returns to process S202 if there are other segmented images for which an affine transformation matrix needs to be generated, and performs processing on those other segmented images (S202). The affine transformation matrix generation unit 101c generates an affine transformation matrix for all segmented images (S206:NO), and the position adjustment unit 102 performs the generated affine transformation processing on the images of objects (such as people) included in the real-space image rp captured by the real-world camera 100 (S207).
[0110] The position adjustment unit 102 performs position adjustment processing using a different affine transformation matrix for each divided image. Furthermore, it processes the output values by performing an averaging process on the output values of the affine transformation matrices of the divided images obtained from different division units. Other processing methods besides averaging may also be used, such as weighted averaging.
[0111] Figure 19 shows the process of performing an averaging operation on the output values from the affine transformation matrix. In Figure 19(a), the real-space image rp, which includes an object such as a person, is divided into a first division unit. Here, for example, points P1 and P2 are included. Points P1 and P2 are pixels that constitute the object. In this disclosure, if the object is a person, the person's movements are represented by the person's skeleton (bones). In this disclosure, if points P1 and P2 of the real-space image rp are pixels that constitute the person's skeleton, then in the virtual-space image vp, points P1 and P2 are transformed into points P100 and P200 by the affine transformation matrix of the division image G11. This affine transformation matrix is an affine transformation matrix calculated in correspondence with the position of the division image G11 when the real-space image rp is divided into 3x3 division units. In Figure 19(b), points P100 and P200 are transformed to different positions from points P1 and P2.
[0112] Similarly, in Figures 19(c) and (d), the real-space image rp is divided into 4x4 units, and positions P1 and P2 are transformed into positions P210 and P220 using an affine transformation matrix. Here, since position P1 is included in the divided image G21 and position P2 is included in the divided image G22, the transformation process is performed using the affine transformation matrix calculated for each respective image.
[0113] Then, as shown in Figure 19(e), the positions P100 and P210, which are transformed from position P1 respectively, are used to calculate the average of these positions, thereby obtaining position P100a. Since the distortion of the virtual space model differs depending on the location, multiple affine transformation matrices are required. In this disclosure, the distortion can be corrected or mitigated by obtaining and integrating multiple affine transformation matrices using the procedure described above. Position P100a is the position in the virtual space image corresponding to point P1, and is the position where the distortion of the virtual space model is reduced.
[0114] Next, the effects of the information processing device 10a of this disclosure will be explained. In this disclosure, the real space image acquisition unit 11 acquires a real space image rp captured by the real camera 100. The virtual space image acquisition unit 13 acquires a virtual space image vp captured by a virtual camera vc installed at a position in the virtual space vs corresponding to the position of the real camera 100, representing the virtual space vs which is represented by a virtual space model vm corresponding to the real space.
[0115] The transformation matrix generation unit 101 generates a positional transformation matrix for alignment from the reference image coordinates d of the real-space image rp and the virtual reference image coordinates d'. More specifically, the transformation matrix generation unit 101 generates a positional transformation matrix that performs a transformation process to displace the reference image coordinates d of the real-space image rp to the virtual reference image coordinates d' of the virtual-space image vp corresponding to the reference image coordinates d. This positional transformation matrix is, for example, an affine transformation matrix. The position adjustment unit 102 uses the positional transformation matrix to align objects (such as skeletons of people) included in the real-space image with the virtual-space image. The positional transformation matrix is generated in advance, and then imaging is performed by the real-space camera 100, followed by positional adjustment using the generated positional transformation matrix.
[0116] This configuration allows for the high-precision projection of objects contained in real-world images into virtual space (virtual space images). Generally, there is a slight discrepancy between the virtual space represented by a virtual space model attempting to represent real space and the real space itself. In this disclosure, this discrepancy can be absorbed by using positional transformation matrices such as affine transformations, enabling accurate projection. As a result of this accurate projection, the three-dimensional position of objects such as people in virtual space can be precisely determined.
[0117] Furthermore, objects such as people exist in the real space. The real-world camera 100 captures images of the real space including the objects, extracts only the skeletons of those objects, and projects them onto the virtual space, thereby compositing the objects into the virtual space. As described in this disclosure, by performing positional adjustments such as affine transformations, the projection of objects onto the virtual space can be performed with high precision. Then, using the above transformation formulas (1) to (3), the 3D coordinates of the virtual space can be determined from the 2D coordinates of the objects in the projected virtual space image, and the positional relationship of the objects in the 3D model can be accurately grasped.
[0118] In this disclosure, the position calculation unit 19 functions as an estimation unit that estimates the position of a virtual camera and calculates a transformation parameter to be applied to the virtual camera vc. This transformation parameter is obtained by feature point matching between the real-space image rp and the virtual-space image vp. The virtual camera vc to which this transformation parameter is applied can then be estimated to correspond to the shooting position of the virtual camera in the virtual space (i.e., the shooting position of the real camera 100 in the real space).
[0119] The virtual space image acquisition unit 13 then acquires a virtual space image vp taken from a virtual camera vc (i.e., the real camera shooting position) to which the conversion parameters calculated by the position calculation unit 19 have been applied.
[0120] Thus, when there are many common reference feature points between the real-space image rp and the virtual-space image vp, the position of the real-space camera 100 and the position of the virtual camera in the virtual space can be considered to be corresponding positions. As described above, the virtual-space model vm may have some discrepancies with the real space, but for affine transformation processing, it is necessary for the virtual-space image vp and the real-space image rp to coincide to some extent. In this disclosure, the above transformation parameters (transformation matrix B) are used to estimate the position of the virtual camera in order to make the virtual-space image vp and the real-space image rp coincide to some extent, but this is not the only method. Any method can be used to make the position of the virtual camera vc in the virtual space coincide to some extent with the position of the real camera 100 in the real space. Alternatively, the position of the virtual camera vc in the virtual space and the position of the real camera 100 in the real space may be fixed positions so that they can be associated.
[0121] Furthermore, the transformation matrix generation unit 101 may generate a position transformation matrix (affine transformation matrix) for each of the divided images obtained by dividing the real-space image rp into several parts. The position adjustment unit 102 uses the position transformation matrix to adjust the position of objects included in the real-space image to the virtual space image.
[0122] This allows for the generation of positional transformation matrices appropriate to each segmented image, enabling correction tailored to each segmented image.
[0123] More specifically, the process is as follows: The image division unit 101a divides the real-space image rp and the virtual-space image vp to obtain divided images. The affine transformation matrix generation unit 101c then generates a first positional transformation matrix for each of the multiple first divided image regions (for example, divided image G11 in Figure 19) obtained by dividing the real-space image rp into first division units (for example, 3x3 division units). The affine transformation matrix generation unit 101c also generates a second positional transformation matrix for each of the multiple second divided images (for example, divided image G21 in Figure 19) obtained by dividing the real-space image rp into second division units (for example, 4x4 division units) that are different from the first division units.
[0124] The position adjustment unit 102 then calculates a first target coordinate (for example, position P100 in Figure 19) for each first segmented image using a first position transformation matrix, and calculates a second target coordinate (for example, position P210 in Figure 19) for each second segmented image using a second position transformation matrix. The position adjustment unit 102 calculates a target coordinate (for example, position P100a in Figure 19) based on the first and second target coordinates, and performs position adjustment using this target coordinate. This target coordinate is based, for example, on the average value of the first and second target coordinates.
[0125] This configuration allows for smooth correction of the seams between segmented images by using values based on multiple affine transformation results, such as the average value.
[0126] Furthermore, the transformation matrix generation unit 101 (feature point verification unit 101b) enlarges the segmented image if there are no reference image coordinates, or if the number of reference image coordinates is less than the threshold, until the number of reference image coordinates is equal to or greater than the threshold. Then, the feature point verification unit 101b uses the reference image coordinates in the enlarged segmented image to determine the position transformation matrix of the segmented image before enlargement.
[0127] This configuration allows for accurate correction by securing the necessary feature points even in segmented images with few feature points.
[0128] Furthermore, the transformation matrix generation unit 101 (feature point count verification unit 101b) does not perform the transformation matrix generation process if, in the real-space image rp or virtual-space image vp, there are no reference image coordinates d or virtual reference image coordinates d' that serve as feature points, or if the number of reference image coordinates d or virtual reference image coordinates d' is less than or equal to a threshold.
[0129] In real-space image rp and virtual-space image vp, if no feature points exist, transformation matrices cannot be generated even with segmentation processing, so this processing is omitted. This reduces the processing load.
[0130] In this disclosure, the affine transformation process is important among the above processes, but it is not essential to prepare an affine transformation matrix for each segmented image or to enlarge the segmented images.
[0131] The computer may be an information processing program that causes it to function as the information processing device 10a of this embodiment. The information processing program P1 shown in Figure 9 is composed of a main module m10 that comprehensively controls information processing in the information processing device 10, a real-space image acquisition module m11, a setting module m12, a virtual-space image acquisition module m13, a feature point extraction module m14, a feature point matching module m15, a parameter calculation module m16, an output module m17, a reset module m18, and a position calculation module m19. Each of the modules m11 to m19 realizes the respective functions for each of the functional units 11 to 19.
[0132] In addition, the information processing program P1 may further include a transformation matrix generation module that functions as a transformation matrix generation unit 101, and a position adjustment module that functions as a position adjustment unit 102.
[0133] The block diagram shown in Figure 13 represents functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may also be realized by combining the above one device or the above multiple devices with software.
[0134] Functions include, but are not limited to, judgment, decision, judgment, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, assumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, a functional block (configuration part) that enables transmission is called a transmitting unit or transmitter. As mentioned above, the method of implementation is not particularly limited.
[0135] For example, the information processing device 10 (including the information processing device 10a) in one embodiment of the present invention may function as a computer. Figure 20 is a diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. Physically, the information processing device 10 may be configured as a computer device including a processor 1001, memory 1002, storage 1003, communication device 1004, input device 1005, output device 1006, bus 1007, etc.
[0136] In the following explanation, the term "device" can be replaced with "circuit," "device," "unit," etc. The hardware configuration of the information processing device 10 may include one or more of the devices shown in Figure 20, or it may be configured to omit some of the devices.
[0137] Each function in the information processing device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, allowing the processor 1001 to perform calculations and control communication by the communication device 1004, as well as the reading and / or writing of data to the memory 1002 and storage 1003.
[0138] The processor 1001 controls the entire computer, for example, by running the operating system. The processor 1001 may consist of a central processing unit (CPU) that includes interfaces with peripheral devices, control devices, arithmetic units, registers, etc. For example, the various functional units 11 to 19 shown in Figure 1, the transformation matrix generation unit 101, or the position adjustment unit 102 may be implemented in the processor 1001.
[0139] Furthermore, the processor 1001 reads programs (program code), software modules, and data from the storage 1003 and / or communication device 1004 into the memory 1002 and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described in the above embodiment. For example, each of the functional units 11 to 19 of the information processing device 10, the transformation matrix generation unit 101, or the position adjustment unit 102 may be implemented by a control program stored in the memory 1002 and operated by the processor 1001. Although the above-described processes have been explained as being executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented on one or more chips. The program may also be transmitted from a network via a telecommunications line.
[0140] Memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 1002 may also be called a register, cache, main memory, etc. Memory 1002 can store executable programs (program code), software modules, etc., for carrying out an information processing method according to one embodiment of the present invention.
[0141] The storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. The storage 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including memory 1002 and / or storage 1003.
[0142] The communication device 1004 is hardware (transceiver / receiver device) for communicating between computers via a wired and / or wireless network, and is also referred to as a network device, network controller, network card, communication module, etc.
[0143] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).
[0144] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may consist of a single bus or different buses may be used for communication between devices.
[0145] Furthermore, the information processing device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0146] Information notification is not limited to the embodiments described herein and may be carried out by other means. For example, information notification may be carried out by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or combinations thereof. RRC signaling may also be called RRC messages, and may be, for example, RRC Connection Setup messages, RRC Connection Reconfiguration messages, etc.
[0147] Each aspect / embodiment described in this disclosure may be applied to at least one of the following systems: LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth (registered trademark), and other appropriate systems, as well as next-generation systems extended based thereon. Furthermore, multiple systems may be applied in combination (for example, a combination of at least one of LTE and LTE-A with 5G).
[0148] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described herein may be reordered, provided they are consistent with each other. For example, the methods described herein present various step elements in an exemplary order and are not limited to that specific order.
[0149] The specific operations described in this disclosure as being performed by a base station may, in some cases, be performed by its upper node. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal can be performed by the base station and at least one other network node (for example, an MME or S-GW, but not limited to these). Although the above example illustrates a case where there is one other network node besides the base station, it may also be a combination of multiple other network nodes (for example, an MME and an S-GW).
[0150] Information can be output from a higher layer (or lower layer) to a lower layer (or higher layer). Input and output may also occur via multiple network nodes.
[0151] Input and output information may be stored in a specific location (e.g., memory) or managed in a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be sent to other devices.
[0152] The determination may be made by a value represented by 1 bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).
[0153] Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).
[0154] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.
[0155] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name.
[0156] Furthermore, software, instructions, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cable, fiber optic cable, twisted pair, and digital subscriber lines (DSL) and / or wireless technologies such as infrared, radio, and microwave, these wired and / or wireless technologies are included in the definition of a transmission medium.
[0157] The information, signals, etc. described in this disclosure may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0158] In addition, terms described in this disclosure and / or terms necessary for understanding this specification may be replaced with terms having the same or similar meaning.
[0159] The terms “system” and “network” as used in this disclosure are interchangeable.
[0160] Furthermore, the information, parameters, etc., described in this disclosure may be expressed as absolute values, relative values from a given value, or by corresponding other information. For example, wireless resources may be indicated by an index.
[0161] The names used for the parameters described above are not restrictive in any way. Furthermore, the formulas and other expressions using these parameters may differ from those expressly disclosed in this disclosure. Various channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable name, and therefore, the various names assigned to these various channels and information elements are not restrictive in any way.
[0162] As used in this disclosure, the terms “determining” and “determining” may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiry (e.g., searching in a table, database, or other data structure), and ascertaining. “Determining” may also include, for example, receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, and accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."
[0163] As used in this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."
[0164] Where the terms “first,” “second,” etc., are used in this disclosure, no reference to those elements shall generally limit the quantity or order of those elements. These terms may be used herein as a convenient way to distinguish between two or more elements. Accordingly, references to the first and second elements shall not imply that only two elements may be employed therein, or that the first element must precede the second element in any way.
[0165] In the configuration of each of the above devices, "means" may be replaced with "part," "circuit," "device," etc.
[0166] To the extent that “include,” “including,” and their variations are used herein or in the claims, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used herein or in the claims is not intended to be exclusive OR.
[0167] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.
[0168] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different."
[0169] The information processing system 1 disclosed herein may have the following configuration.
[0170] [1] A feature point extraction unit that extracts feature points from a real-space image captured of a real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to the real space and represented by a virtual-space model, and the position of the virtual camera is the viewpoint position for capturing an image projected from the virtual-space model, the feature point extraction unit, A feature point matching unit that matches the feature points of the real-world image with the feature points of the virtual-world image, A parameter calculation unit calculates the transformation parameters for imaging the real space by substituting the reference space coordinates in the virtual space of the reference feature points, which are the matched feature points, and the reference image coordinates, which are the coordinates of the reference feature points in the real space image, into a predetermined transformation formula that expresses the relationship between the 3D coordinates of a specific point, which is a specific point in 3D space, and the 2D coordinates of the specific point in an image captured of the 3D space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on the transformation parameters for imaging the virtual space by the virtual camera and the virtual reference image coordinates, which are the coordinates of the reference feature points in the virtual space image. An output unit that outputs the conversion parameters related to imaging the real space as camera information representing the position of the camera that captured the real space image in the virtual space, An information processing device equipped with the following features.
[0171] [2] The system further comprises a setting unit for installing multiple virtual cameras, The parameter calculation unit calculates the transformation parameters for imaging the real space based on at least one of the multiple virtual space images acquired by each of the multiple virtual cameras, selected based on the abundance of the reference feature points. [1] The information processing device described above.
[0172] [3] The setting unit installs at least a plurality of virtual cameras along the outer perimeter of the virtual space, with the normal direction of the outer perimeter as the imaging direction, or Multiple virtual cameras are installed within the virtual space with the outer periphery of the virtual space as the imaging direction. [2] The information processing device described above.
[0173] [4] The parameter calculation unit calculates the conversion parameters for imaging the real space based on the virtual space image having a number of reference specified points equal to or greater than a threshold for the reference feature points. An information processing device as described in any one of the items [1] to [3].
[0174] [5] The system further includes a resetting unit that re-installs the virtual camera in the virtual space to which the conversion parameters for imaging the real space calculated by the parameter calculation unit have been applied as conversion parameters for imaging the virtual space. The feature point extraction unit extracts feature points from the virtual space image captured by the reset virtual camera. [4] The information processing device described above.
[0175] [6] The resetting unit reinstalls at least one virtual camera within a predetermined range around a specific virtual camera, which is the virtual camera that captured the virtual space image having the most reference feature points. [5] The information processing device described above.
[0176] [7] The resetting unit repositions at least one of the virtual cameras so that a virtual space image is captured in which a portion of the virtual space image captured by the specific virtual camera overlaps with a portion of the region including the reference feature points. The information processing device described in [6].
[0177] [8] The parameter calculation unit calculates the conversion parameters related to imaging the real space based on each of the plurality of virtual space images, The output unit outputs the converted parameters obtained by processing the calculated plurality of converted parameters using a predetermined statistical method as camera information. An information processing device as described in any one of items [1] to [7].
[0178] [9] The transformation parameters include six predetermined camera external parameters, four camera internal parameters, and five parameters related to the projection of a three-dimensional space onto a two-dimensional image, respectively. The parameter calculation unit uses the following transformation formula (1) to represent the relationship between three-dimensional coordinates (X, Y, Z) in three-dimensional space and two-dimensional coordinates (u, v) in a two-dimensional image.
number
[0179]
[10] A feature point extraction step for extracting feature points from a real-space image captured of a real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to the real space and represented by a virtual-space model, and the position of the virtual camera is a viewpoint position for capturing an image projected from the virtual-space model, A feature point matching step that matches the feature points of the real-world image with the feature points of the virtual-world image, A parameter calculation step to calculate the transformation parameters for imaging the real space by substituting the reference space coordinates in the virtual space of the reference feature points, which are the matched feature points, and the reference image coordinates, which are the coordinates of the reference feature points in the real space image, into a predetermined transformation formula that expresses the relationship between the 3D coordinates of a specific point, which is a specific point in 3D space, and the 2D coordinates of the specific point in an image captured of the 3D space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on the transformation parameters for imaging the virtual space by the virtual camera and the virtual reference image coordinates, which are the coordinates of the reference feature points in the virtual space image. Output step of outputting the conversion parameters related to imaging of the real space as camera information representing the position of the camera that captured the real space image in the virtual space, An information processing method performed by a processor, comprising:
[0180] Furthermore, the apparatus of this disclosure may have the following configuration.
[0181] [1A] A real-space image acquisition unit that acquires real-space images captured by a real-space camera, A virtual space image acquisition unit acquires virtual space images captured by a virtual camera installed at a location in the virtual space corresponding to the location of the real camera, representing a virtual space model that corresponds to the real space. A generation unit that generates a positional transformation matrix for alignment from the reference image coordinates of the real space image and the virtual reference image coordinates, A position adjustment unit that uses the position transformation matrix to align objects included in the real-space image with the virtual-space image, A device equipped with the following features.
[0182] [2A] The system includes an estimation unit that estimates the shooting position of a virtual camera in the virtual space by using feature point matching between the real-world image and the virtual-world image, The virtual space image acquisition unit acquires a virtual space image taken from the virtual camera shooting position estimated by the estimation unit. The apparatus described in [1A].
[0183] [3A] The generating unit is A position transformation matrix is generated that performs a transformation process to displace the reference image coordinates of the real space image to the virtual reference image coordinates of the virtual space image corresponding to the reference image coordinates. The apparatus described in [1A] or [2A].
[0184] [4A] The aforementioned positional transformation matrix is an affine transformation matrix. The apparatus described in [3A].
[0185] [5A] The generation unit generates a position transformation matrix for each of the divided images obtained by dividing the real space image into several parts. The position adjustment unit performs position adjustment using the position transformation matrix. The apparatus described in [3A] or [4A].
[0186] [6A] The generating unit is For each of the multiple first division image regions obtained by dividing the real-space image into first division units, a first positional transformation matrix is generated for each division image. For a plurality of second division images obtained by dividing the real-space image using a second division unit different from the first division unit, a second positional transformation matrix is generated for each of the division images. The position adjustment unit is, For each of the first segmented images, the first transformed coordinates are calculated using the first positional transformation matrix. For each of the two divided images, the second transformed coordinates are calculated using the second positional transformation matrix. Based on the first and second target coordinates, the target coordinates are calculated, and the position is adjusted using these target coordinates. The apparatus described in [5A].
[0187] [7A] The generating unit is For segmented images where no reference image coordinates exist, or where the number of such reference image coordinates is less than the threshold, the segmented image is enlarged until the number of reference image coordinates is equal to or greater than the threshold. Using the reference image coordinates in the enlarged segmented image, the positional transformation matrix of the segmented image before enlargement is obtained. The apparatus described in [5A] or [6A].
[0188] [8A] The generating unit is If, in the aforementioned real-space image or virtual-space image, there are no reference image coordinates or virtual reference image coordinates that serve as feature points, or if the number of reference image coordinates or virtual reference image coordinates is less than or equal to a threshold, the generation process is not performed. The apparatus described in any one of [5A] to [7A].
[0189] [9A] In the method of the apparatus, A real-space image acquisition step involves acquiring a real-space image captured by a real-world camera, A virtual space image acquisition step involves acquiring a virtual space image captured by a virtual camera installed at a location in the virtual space corresponding to the location of the real camera, in a virtual space represented by a virtual space model corresponding to the real space. A generation step of generating a positional transformation matrix for alignment from the reference image coordinates of the real space image and the virtual reference image coordinates, A position adjustment step in which the position of an object included in the real space image is aligned with the virtual space image using the position transformation matrix, A method for providing this. [Explanation of symbols]
[0190] 1... Information processing system, 10... Information processing device, 11... Real-world image acquisition unit, 12... Setting unit, 13... Virtual-world image acquisition unit, 14... Feature point extraction unit, 15... Feature point matching unit, 16... Parameter calculation unit, 17... Output unit, 18... Reset unit, 19... Position calculation unit, 21... Real-world image storage unit, 22... Virtual-world model storage unit, 101... Transformation matrix generation unit, 101a... Image segmentation unit, 101b... Feature point count confirmation unit, 101c... Affine transformation matrix Generation unit, 101d... Affine transformation matrix storage unit, 102... Position adjustment unit, M1... Recording medium, m11... Real space image acquisition module, m12... Setting module, m13... Virtual space image acquisition module, m14... Feature point extraction module, m15... Feature point matching module, m16... Parameter calculation module, m17... Output module, m18... Reset module, m19... Position calculation module, P1... Information processing program.
Claims
1. A real-space image acquisition unit that acquires real-space images captured by a real-space camera, A virtual space image acquisition unit acquires virtual space images captured by a virtual camera installed at a location in the virtual space corresponding to the location of the real camera, representing a virtual space model that corresponds to the real space. A generation unit generates a positional transformation matrix for alignment from the reference image coordinates of the real space image and the virtual reference image coordinates of the virtual space image. A position adjustment unit that uses the position transformation matrix to align objects included in the real-space image with the virtual-space image, A device equipped with the following features.
2. The system includes an estimation unit that estimates the shooting position of a virtual camera in the virtual space by using feature point matching between the real-world image and the virtual-world image, The virtual space image acquisition unit acquires a virtual space image taken from the shooting position of the virtual camera estimated by the estimation unit. The apparatus according to claim 1.
3. The generating unit is A position transformation matrix is generated that performs a transformation process to displace the reference image coordinates of the real space image to the virtual reference image coordinates of the virtual space image corresponding to the reference image coordinates. The apparatus according to claim 1.
4. The aforementioned positional transformation matrix is an affine transformation matrix. The apparatus according to claim 3.
5. The generation unit generates a position transformation matrix for each of the divided images obtained by dividing the real space image into several parts. The position adjustment unit performs position adjustment using the position transformation matrix. The apparatus according to claim 3.
6. The generating unit is For each of the multiple first divided image regions obtained by dividing the real-space image into first division units, a first positional transformation matrix is generated for each of the divided images. For a plurality of second division images obtained by dividing the real-space image using a second division unit different from the first division unit, a second positional transformation matrix is generated for each of the division images. The position adjustment unit is, For each of the first segmented images, the first transformed coordinates are calculated using the first positional transformation matrix. For each of the two divided images, the second transformed coordinates are calculated using the second position transformation matrix. Based on the first and second target coordinates, the target coordinates are calculated, and the position is adjusted using these target coordinates. The apparatus according to claim 5.
7. The generating unit is For segmented images where no reference image coordinates exist, or where the number of such reference image coordinates is less than the threshold, the segmented image is enlarged until the number of reference image coordinates is equal to or greater than the threshold. Using the reference image coordinates in the enlarged segmented image, the positional transformation matrix of the segmented image before enlargement is obtained. The apparatus according to claim 5.
8. The generating unit is If, in the aforementioned real-space image or virtual-space image, there are no reference image coordinates or virtual reference image coordinates that serve as feature points, or if the number of reference image coordinates or virtual reference image coordinates is less than or equal to a threshold, the generation process is not performed. The apparatus according to claim 5.
9. In the method of the apparatus, A real-space image acquisition step that acquires a real-space image captured by a real-world camera, A virtual space image acquisition step involves acquiring a virtual space image captured by a virtual camera installed at a location in the virtual space corresponding to the location of the real camera, in a virtual space represented by a virtual space model corresponding to the real space. A generation step of generating a positional transformation matrix for alignment from the reference image coordinates of the real space image and the virtual reference image coordinates of the virtual space image, A position adjustment step is performed using the position transformation matrix to align the objects included in the real space image with the virtual space image, A method for providing this.