Device and method
The apparatus and method align real-space and virtual-space images using a system of units to generate and adjust position transformation matrices, addressing the challenge of accurately creating three-dimensional spatial models from two-dimensional images.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NTT DOCOMO INC
- Filing Date
- 2025-10-07
- Publication Date
- 2026-05-28
AI Technical Summary
Existing technologies face challenges in accurately creating a three-dimensional spatial model that perfectly matches the real world, even when the correspondence between two-dimensional and three-dimensional coordinates is determined.
An apparatus and method that includes a real-space image acquisition unit, a virtual-space image acquisition unit, a generation unit for generating a position transformation matrix, and a position adjustment unit to align objects in real-space and virtual-space images, using a system with functional units like real-space image acquisition, virtual-space image acquisition, feature point extraction and matching, parameter calculation, and position adjustment to achieve accurate alignment.
Enables the conversion of real-space images into virtual-space images with high accuracy, allowing for precise estimation of three-dimensional spatial models from two-dimensional images.
Smart Images

Figure JP2025035604_28052026_PF_FP_ABST
Abstract
Description
Apparatus and method
[0001] The present invention relates to an apparatus and method for handling three-dimensional spatial models.
[0002] There is a need for technology to estimate three-dimensional space based on two-dimensional images. To do this, it is necessary to determine the positional relationship between an image of the real world captured by a camera and a corresponding three-dimensional virtual space. There is a technology that determines the positional relationship between the camera image and the three-dimensional space by placing objects with known positions, such as markers, in the real world and capturing an image that includes the markers. Furthermore, for example, Patent Document 1 below describes a technology related to learning a model for estimating the position and orientation of an object from an image of the real world.
[0003] Japanese Patent Publication No. 2022-81081
[0004] As mentioned above, in order to estimate a three-dimensional space from a two-dimensional image, it is necessary to determine various parameters related to the projection of the three-dimensional space onto the two-dimensional image, as well as the correspondence between coordinates in the two-dimensional image and coordinates in the three-dimensional space. However, even if this correspondence can be determined, it is difficult to create a three-dimensional spatial model that perfectly matches the real world.
[0005] Therefore, the present invention has been made in view of the above problems, and aims to provide an apparatus and method that can match a three-dimensional spatial model with real space.
[0006] To solve the above problems, the system includes: a real-space image acquisition unit that acquires real-space images captured by a real-space camera; a virtual-space image acquisition unit that acquires virtual-space images captured by a virtual camera installed at a position in the virtual space corresponding to the position of the real-space camera, representing a virtual space represented by a virtual-space model corresponding to the real space; a generation unit that generates a position transformation matrix for alignment from the reference image coordinates of the real-space image and the virtual reference image coordinates; and a position adjustment unit that uses the position transformation matrix to align objects included in the real-space image with the virtual-space image.
[0007] According to this disclosure, it is possible to accurately convert real-space images captured by a camera that images real space into virtual-space images of a virtual space.
[0008] This is a block diagram showing the functional configuration of the information processing system and information processing device of this embodiment. This is a diagram showing an example of generating a virtual space model from a real space image. Figure 3(a) shows a first example of setting up a virtual camera in a virtual space. Figure 3(b) shows a second example of setting up a virtual camera in a virtual space. This is a diagram showing an example of a virtual space image and an example of a real space image, as well as an example of feature point extraction and matching processing. This is a diagram showing an example of the calculation process for reference space coordinates, which are the coordinates of reference feature points in a virtual space. This is a diagram showing an example of reinstalling a virtual camera. This is a diagram showing an example of how to superimpose the imaging areas between a virtual space image from a specific virtual camera and a virtual space image from a reinstalled virtual camera. This is a flowchart showing the processing content of the information processing method in the information processing device. This is a diagram showing the configuration of an information processing program. This is a diagram showing a real camera 100 placed in the real space of this disclosure, and the real space image rp and virtual space vs obtained thereby. This is a diagram showing a virtual space image vp. This is a diagram showing an overview of the operation of the process for determining the imaging position of the virtual camera of this disclosure and the subsequent affine transformation process. This is a block diagram showing the functional configuration of the information processing device 10a. This figure shows the functional configuration of the transformation matrix generation unit 101, which calculates an affine transformation matrix for each division unit. This figure shows a specific example of an affine transformation matrix. This flowchart shows the operation of the transformation matrix generation unit 101 in the information processing device 10a of this disclosure. This figure shows divided images of different division units. This figure shows enlarged divided images obtained by enlarging the divided images. This figure shows that an averaging process is performed on the output values from the affine transformation matrix. This is a hard block diagram of the information processing device.
[0009] Embodiments of the information processing system according to the present invention will be described with reference to the drawings. Where possible, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.
[0010] Figure 1 shows the functional configuration of the information processing system and information processing device according to this embodiment. The information processing system 1 of this embodiment is a system that obtains information representing the position in virtual space of a camera that captures real space in order to estimate a three-dimensional space from a two-dimensional image without using installed objects such as markers, and is configured as an information processing device 10 as an example.
[0011] As shown in Figure 1, the information processing device 10 functionally comprises a real-space image acquisition unit 11, a setting unit 12, a virtual-space image acquisition unit 13, a feature point extraction unit 14, a feature point matching unit 15, a parameter calculation unit 16, an output unit 17, a reset unit 18, and a position calculation unit 19. These functional units 11 to 19 may be configured in a single device as illustrated in Figure 1, or they may be distributed across multiple devices.
[0012] Each of the functional units 11 to 19 of the information processing device 10 is configured to access storage means such as the real-space image storage unit 21 and the virtual-space model storage unit 22. Each of the storage units 21 to 22 may be provided in the information processing device 10, or, as illustrated in Figure 1, may be configured in other devices that are accessible from the information processing device 10.
[0013] Next, the various functions of the information processing device 10 will be described. The real-space image acquisition unit 11 acquires real-space images captured in real space. Specifically, the real-space image acquisition unit 11 may acquire real-space images from a camera that captures real space, or it may acquire real-space images from a real-space image storage unit 21, which is a storage means that stores captured real-space images.
[0014] Figure 2 shows an example of generating a virtual space model from a real-world image. As shown in Figure 2, the real-world image acquisition unit 11 acquires a real-world image from the real-world space rs. Also, as shown in Figure 2, a virtual space model vm representing the virtual space corresponding to the real-world space rs is generated based on the real-world space rs.
[0015] The virtual space model vm may be generated based on the real-space image rp by any known method. For example, in the acquisition of a real-space image, the virtual space model vm may be generated based on the distance to an object measured by LiDAR (Light Detection and Ranging). Alternatively, the virtual space model vm may be generated by estimating the depth of what is represented in each pixel of a two-dimensional real-space image rp using a known technique, converting it into point cloud data based on the depth and color value of each pixel, and generating a three-dimensional virtual space by a three-dimensional display based on the point cloud data.
[0016] The generated virtual space model vm may be stored in the virtual space model storage unit 22. The virtual space model storage unit 22 is a storage means for storing pre-generated virtual space models vm.
[0017] The setting unit 12 places a virtual camera in the virtual space represented by the virtual space model vm. Specifically, the setting unit 12 sets the position of the virtual camera in the virtual space represented by the virtual space model vm. The position of the virtual camera is the viewpoint position for capturing a virtual space image, which is an image projected from the virtual space model vm.
[0018] Figure 3 shows an example of the placement of virtual cameras in a virtual space, where Figures 3(a) and 3(b) show first and second examples of the placement positions of virtual cameras in the virtual space as viewed from above the virtual space vs., respectively. As shown in Figures 3(a) and 3(b), the setting unit 12 may set up multiple virtual cameras vc in the virtual space vs.
[0019] In the first example shown in Figure 3(a), the setting unit 12 may install multiple virtual cameras vc along the outer perimeter of the virtual space vs, with the normal direction of the outer perimeter as the imaging direction. Alternatively, as shown in Figure 3(b), the setting unit 12 may install multiple virtual cameras vc within the virtual space vs (for example, around the central part of the virtual space vs), with the direction of the outer perimeter of the virtual space vs as the imaging direction. By installing multiple virtual cameras vc in this way, the virtual space corresponding to the space represented in the real-world image rp can be comprehensively imaged.
[0020] The virtual space image acquisition unit 13 acquires at least one virtual space image obtained by imaging the virtual space vs using the virtual camera vc. The virtual space image acquisition unit 13 may acquire a virtual space image by a known method based on the virtual space vs represented by the virtual space model vm.
[0021] FIG. 4 is a diagram showing an example of a virtual space image, an example of a real space image, and an example of feature point extraction and matching processing. Specifically, the virtual space image acquisition unit 13 projects the virtual space vs onto a virtual screen with the position of the virtual camera vc set in the virtual space vs represented based on the virtual space model vm as the viewpoint position, thereby acquiring the virtual space image vp.
[0022] The feature point extraction unit 14 extracts feature points from the real space image rp obtained by imaging the real space rs and at least one virtual space image vp. The feature point matching unit 15 matches the feature points of the real space image rp with the feature points of the virtual space image vp. The feature point extraction unit 14 and the feature point matching unit 15 may perform feature point extraction and feature point matching by a known method. For example, they may perform feature point extraction and matching by known methods such as SIFT and AKAZE.
[0023] In the example shown in FIG. 4, the feature point extraction unit 14 extracts feature points vfp1 and vfp2 from the virtual space image vp. Also, the feature point extraction unit 14 extracts feature points rfp1 and rfp2 from the real space image rp. Feature points are extracted based on, for example, the difference in pixel values of adjacent pixels in the image. For example, the endpoints and corner portions of reference objects such as doors or desks represented in the image are extracted as feature points.
[0024] Further, the feature point matching unit 15 matches each of the feature points vfp1 and vfp2 of the virtual space image vp with each of the feature points rfp1 and rfp2 of the real space image rp as indicated by the symbol fm based on the feature amounts of the respective feature points.
[0025] The parameter calculation unit 16 calculates conversion parameters related to imaging in the real space by substituting the reference space coordinates, which are the three-dimensional coordinates of the reference feature points, which are the matched feature points, in the virtual space vs, and the reference image coordinates, which are the coordinates of the reference feature points in the real space image rp, into a predetermined conversion formula.
[0026] The conversion formula is a formula that represents the relationship between the three-dimensional coordinates of a specific point, which is a specific point in the three-dimensional space, and the two-dimensional coordinates of the specific point in the captured image that captured the three-dimensional space, using predetermined conversion parameters. Assuming the three-dimensional coordinates in the three-dimensional space are (X, Y, Z) and the two-dimensional coordinates in the two-dimensional image are (u, v), an example of the predetermined conversion formula is represented by the following conversion formula (1). The above conversion formula (1) includes the camera external parameters R (degree of freedom 3), t (degree of freedom 3), the camera internal parameters f x , f y , c x , c y , and the parameters k 1 , k 2 , k 3 , p 1 , p 2 as conversion parameters. The camera external parameter R represents a 3×3 matrix indicating the orientation of the camera. t represents a translation vector and is a three-dimensional vector indicating the position of the camera. The camera internal parameters f x , f y indicate the focal length in the horizontal direction and the focal length in the vertical direction. The camera internal parameters c x , c y indicate the x coordinate and the y coordinate of the center of the image. The parameters k 1 , k 2 , k 3 , p 1 , p 2 indicate radial distortion and tangential distortion.
[0027] Then, taking the three-dimensional reference space coordinates as t and the virtual reference image coordinates, which are the two-dimensional coordinates of the reference feature points in the virtual space image vp, as d', and assuming that the transformation matrix for transforming the reference space coordinates t into the virtual reference image coordinates d' is represented by the transformation parameters of the transformation formula (1) as A, the relationship between these coordinates is represented by the following formula (2). d' = tA... (2) Further, taking the two-dimensional reference image coordinates as d and assuming that the transformation matrix for transforming the reference space coordinates t into the reference image coordinates d is represented by the transformation parameters of the transformation formula (1) as B, the relationship between these coordinates is represented as in the following formula (3). d = tB... (3) As shown in formula (2), since the transformation matrix A is a transformation matrix (transformation parameters) indicating the relationship between the three-dimensional coordinates indicating the position in the virtual space vs and the two-dimensional coordinates of that position in the image captured by the virtual camera vc arranged in the virtual space vs, it is known. Therefore, the parameter calculation unit 16 can calculate the reference space coordinates t based on the transformation matrix A and the virtual reference image coordinates d'.
[0028] Specifically, the parameter calculation unit 16 may calculate the reference space coordinates t using a technique called hit determination. FIG. 5 is a diagram schematically showing an example of the calculation process of the reference space coordinates by hit determination.
[0029] As shown in FIG. 5, since the position of the virtual camera vc is also known because the transformation matrix A is known, the parameter calculation unit 16 can calculate a three-dimensional straight line ln passing through the virtual reference image coordinates d' of the reference feature point fp in the virtual space image vp. Then, the parameter calculation unit 16 calculates the point where the plane (mesh) constituting the virtual space model vm and the straight line ln intersect as the reference space coordinates t of the reference feature point fp.
[0030] The parameter calculation unit 16 calculates the transformation matrix B in formula (3) as the transformation parameters related to the imaging in the real space based on the reference space coordinates t and the reference image coordinates d of the reference feature point fp.
[0031] The setting unit 12 may install multiple virtual cameras vc in the virtual space vs, as shown in the example described with reference to Figures 3(a) and 3(b). The virtual space image acquisition unit 13 then acquires multiple virtual space images vp acquired by each of the multiple virtual cameras vc.
[0032] The parameter calculation unit 16 may calculate the transformation parameters for imaging in real space based on at least one virtual space image vp selected from among a plurality of virtual space images vp based on the abundance of reference feature points fp.
[0033] Since the number of transformation parameters in the transformation equation (1) that constitutes the transformation matrix B is 15, the parameter calculation unit 16 constructs the transformation equation (1) using the reference spatial coordinates t and reference image coordinates d of at least 8 reference feature points fp, and calculates the transformation parameters by solving the multiple transformation equations that have been constructed as equations. Accordingly, the parameter calculation unit 16 may calculate the transformation parameters related to imaging in real space based on a virtual space image vp having a number of reference feature points fp equal to or greater than a predetermined threshold.
[0034] Specifically, the parameter calculation unit 16 may calculate transformation parameters related to imaging in the real space using a virtual space image vp having eight or more reference feature points fp. Alternatively, the parameter calculation unit 16 may calculate transformation parameters related to imaging in the real space using multiple virtual space images vp such that the total number of reference feature points fp is equal to or greater than a predetermined threshold.
[0035] Furthermore, the parameter calculation unit 16 may calculate conversion parameters related to imaging in the real space based on each of the multiple virtual space images vp. This calculates conversion parameters associated with each virtual space image vp.
[0036] The output unit 17 outputs the transformation parameters (corresponding to the transformation matrix B) related to imaging in real space, calculated by the parameter calculation unit 16, as camera information that substantially represents the position in virtual space of the camera that captured the real-space image.
[0037] Furthermore, the output unit 17 may output a conversion parameter obtained by processing the calculated conversion parameters using a predetermined statistical method as camera information. Specifically, the output unit 17 may output a conversion parameter obtained by processing the calculated conversion parameters using the least squares method or the like as camera information. In this way, by statistically processing each conversion parameter calculated based on the virtual space image vp acquired by each of the multiple virtual cameras vc, it is possible to improve the accuracy of the conversion parameters output as camera information.
[0038] The resetting unit 18 applies the conversion parameters for imaging in the real space, calculated by the parameter calculation unit 16, to the virtual camera vc, which has been applied as conversion parameters for imaging in the virtual space vs, and then reinstalls the virtual camera vc in the virtual space vs.
[0039] When initially setting up the virtual camera vc and acquiring the virtual space image vp, the transformation parameters applied to the virtual camera vc may be set arbitrarily. After the transformation parameters for imaging the real space are calculated based on the initial virtual space image vp, the calculated transformation parameters can be applied to the virtual camera vc. A virtual camera vc to which the calculated transformation parameters have been applied is highly likely to have a positional relationship and characteristics similar to the camera that captured the real space image rp.
[0040] The virtual space image acquisition unit 13 acquires a virtual space image vp captured by a virtual camera vc that has been reinstalled in the virtual space vs. The feature point extraction unit 14 extracts feature points from the virtual space image vp captured by the reinstalled virtual camera vc, and the feature point matching unit 15 performs feature point matching. Then, the parameter calculation unit 16 calculates conversion parameters based on the virtual space image vp captured by the reinstalled virtual camera vc. In this way, the accuracy can be improved by recalculating the conversion parameters based on the virtual space image vp captured by the virtual camera vc to which the calculated conversion parameters have been applied.
[0041] The resetting unit 18 may reinstall at least one virtual camera vc within a predetermined range around a specific virtual camera vc, which is one of the virtual cameras vc set by the setting unit 12 that captured the virtual space image vp having the most reference feature points fp.
[0042] Figure 6 shows an example of the re-installation of a virtual camera. In the example shown in Figure 6, the re-installation unit 18 extracts a specific virtual camera vc1, which is the virtual camera that captured the virtual space image vp having the most reference feature points fp among the virtual cameras vc installed in the virtual space vs during the calculation of the initial or previous conversion parameters. Since the specific virtual camera vc1 is a virtual camera that captured the virtual space image vp having many feature points that match the feature points in the real space image rp, there is a high probability that it is a camera installed at a position in the virtual space vs that corresponds to the position of the camera rc that captured the real space image rp.
[0043] The resetting unit 18 reinstalls the virtual camera vc within a predetermined range around the specific virtual camera vc1. The resetting unit 18 may, for example, reinstall the virtual camera at a position within a predetermined distance from the position of the specific virtual camera vc1. As illustrated in Figure 6, the resetting unit 18 may reinstall virtual cameras vc2 and vc3 at positions adjacent to the specific virtual camera vc1 along the outer perimeter of the virtual space vs.
[0044] In this way, the accuracy can be further improved by recalculating the conversion parameters based on the virtual space image vp derived from a virtual camera vc that has been reinstalled within a predetermined range around a specific virtual camera vc1.
[0045] Furthermore, the resetting unit 18 may re-install at least one virtual camera vc so that a virtual space image vp is captured in which a portion of the virtual space image vp captured by the specific virtual camera vc1 overlaps with a portion of the virtual space image vp that includes the reference feature point fp. Figure 7 shows an example of how the imaging areas between the virtual space image captured by the specific virtual camera and the virtual space image captured by the re-installed virtual camera are superimposed.
[0046] As illustrated in Figure 7, when a virtual space image vp having an imaging region ts0 is captured by a specific virtual camera vc1, the resetting unit 18 repositions the virtual camera vc to capture an imaging region that is superimposed on the imaging region ts0 in a portion of the region including the reference feature point fp. Specifically, the resetting unit 18 repositions the virtual camera vc to a position where imaging regions ts1, ts2, ts3, and ts4, which include the region containing the reference feature point fp located at the edge of the imaging region ts0, are captured as the virtual space image vp. By acquiring the virtual space image vp with the virtual camera vc in this way, it becomes possible to improve the accuracy of the conversion parameters related to lens distortion.
[0047] Referring again to Figure 1, the position calculation unit 19 calculates the position of an object in three-dimensional space based on the two-dimensional coordinates that indicate the position of an object such as a person, represented in the real-space image rp captured from the real-space rs, using transformation parameters included in the camera information.
[0048] In this way, by using the two-dimensional coordinates of an object in a real-space image rp captured from real space, and a virtual space model vm corresponding to that real space, it becomes possible to estimate the position of the object represented in the real-space image rp in three-dimensional space.
[0049] Figure 8 is a flowchart showing the processing details of the information processing method in the information processing system 1. In step S1, the real-space image acquisition unit 11 acquires a real-space image rp captured from the real space.
[0050] In step S2, a virtual space model vm representing the virtual space vs corresponding to the real space rs is generated based on the real space image rp. The virtual space model vm may be generated by the processor of the information processing device 10 or by another device. The generated virtual space model vm may be stored in the virtual space model storage unit 22.
[0051] In step S3, the setting unit 12 installs a virtual camera vc in the virtual space vs represented by the virtual space model vm. In step S4, the virtual space image acquisition unit 13 acquires at least one virtual space image vp of the virtual space vs captured by the virtual camera vc.
[0052] In step S5, the feature point extraction unit 14 extracts feature points from both the real-space image rp and the virtual-space image vp. In step S6, the feature point matching unit 15 matches the feature points of the real-space image rp with the feature points of the virtual-space image vp.
[0053] In step S7, the parameter calculation unit 16 determines whether the number of reference feature points fp, which are matched feature points, is equal to or greater than a threshold. If it is determined that the number of reference feature points fp is equal to or greater than a threshold, the process proceeds to step S8. On the other hand, if it is not determined that the number of reference feature points fp is equal to or greater than a threshold, the process returns to step S3.
[0054] In step S8, the parameter calculation unit 16 obtains the reference spatial coordinates t of the reference feature point fp based on the transformation parameters (transformation matrix A) that show the relationship between the three-dimensional coordinates indicating the position in the virtual space vs and the two-dimensional coordinates of that position in the image captured by the virtual camera vc placed in the virtual space vs, and the virtual reference image coordinates d'.
[0055] In step S9, the parameter calculation unit 16 calculates transformation parameters (transformation matrix B) related to imaging in real space based on the reference spatial coordinates t and the reference image coordinates d.
[0056] In step S10, the parameter calculation unit 16 determines whether or not to terminate the process. For example, if the camera parameter fluctuation falls below a threshold, it is determined that the process will terminate. If it is determined not to terminate the process, that is, to reinstall the virtual camera vc in order to improve the accuracy of the calculated conversion parameters, the process proceeds to step S11. If it is determined to terminate the process, the process proceeds to step S12.
[0057] In step S11, the resetting unit 18 applies the conversion parameters calculated in step S9 to the virtual camera vc. Then, the process returns to step S3 in order to repeat the calculation of the conversion parameters based on the reinstallation of the virtual camera vc. When the process returns to step S3 via step S11, the resetting unit 18 reinstalls the virtual camera vc.
[0058] In step S12, the output unit 17 outputs the transformation parameters (transformation matrix B) related to imaging in real space, calculated by the parameter calculation unit 16, as camera information that substantially represents the position in virtual space of the camera that captured the real-space image rp.
[0059] Next, with reference to Figure 9, an information processing program for causing a computer to function as the information processing device 10 of this embodiment will be described. Figure 9 is a diagram showing the configuration of the information processing program. The information processing program P1 is composed of a main module m10 that comprehensively controls information processing in the information processing device 10, a real-space image acquisition module m11, a setting module m12, a virtual-space image acquisition module m13, a feature point extraction module m14, a feature point matching module m15, a parameter calculation module m16, an output module m17, a reset module m18, and a position calculation module m19. Each of the modules m11 to m19 realizes the respective functions for each of the functional units 11 to 19.
[0060] The information processing program P1 may be transmitted via a transmission medium such as a communication line, or it may be stored in a recording medium M1, as shown in Figure 9.
[0061] According to the information processing system, information processing device 10, information processing method, and information processing program P1 of this embodiment described above, feature points extracted from the real-space image and the virtual-space image are matched as reference feature points, and by associating the two-dimensional reference image coordinates of the reference feature points in the real-space image with the three-dimensional reference space coordinates of the reference feature points in the virtual space, it becomes possible to calculate the transformation parameters in the transformation formula. Since the calculated transformation parameters constitute information that represents the relative relationship between the position in the real-space image and the position in the virtual space, it becomes possible to obtain information that represents the position in the virtual space of a camera that is essentially imaging the real space using camera information consisting of transformation parameters.
[0062] The information processing apparatus and information processing method relating to this disclosure may have the following configurations. The operation and effects of each configuration are described below.
[0063] An information processing device relating to one aspect of this disclosure is a feature point extraction unit that extracts feature points from a real-space image captured of real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to real space and represented by a virtual-space model, and the position of the virtual camera is a viewpoint position for capturing an image projected from the virtual-space model, the feature point extraction unit, the feature point matching unit that matches feature points of the real-space image with feature points of the virtual-space image, the reference space coordinates in the virtual space of the reference feature points which are the matched feature points, and the reference image coordinates which are the coordinates of the reference feature points in the real-space image A parameter calculation unit that calculates transformation parameters related to imaging in real space by substituting into a predetermined transformation formula that expresses the relationship between the three-dimensional coordinates of a specific point in three-dimensional space and the two-dimensional coordinates of the specific point in an image captured in three-dimensional space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on transformation parameters related to imaging in virtual space by a virtual camera and virtual reference image coordinates which are the coordinates of a reference feature point in the virtual space image; and an output unit that outputs the transformation parameters related to imaging in real space as camera information representing the position in virtual space of the camera that captured the real space image.
[0064] An information processing method relating to one aspect of this disclosure is a feature point extraction step performed by a processor, which extracts feature points from a real-space image captured of real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to real space and represented by a virtual-space model, and the position of the virtual camera is a viewpoint position for capturing an image projected from the virtual-space model; a feature point matching step that matches feature points of the real-space image with feature points of the virtual-space image; and the reference space coordinates in the virtual space of the reference feature points which are the matched feature points and the coordinates of the reference feature points in the real-space image. A parameter calculation step for calculating transformation parameters related to imaging in real space by substituting quasi-image coordinates into a predetermined transformation formula that expresses the relationship between the three-dimensional coordinates of a specific point, which is a specific point in three-dimensional space, and the two-dimensional coordinates of the specific point in an image captured in three-dimensional space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on transformation parameters related to imaging in virtual space by a virtual camera and virtual reference image coordinates, which are the coordinates of a reference feature point in the virtual space image; and an output step for outputting the transformation parameters related to imaging in real space as camera information representing the position in virtual space of the camera that captured the real space image.
[0065] Based on the above aspects, feature points extracted from both the real-space image and the virtual-space image are matched as reference feature points, and by associating the two-dimensional reference image coordinates of the reference feature points in the real-space image with the three-dimensional reference space coordinates of the reference feature points in the virtual space, it becomes possible to calculate the transformation parameters in the transformation formula. Since the calculated transformation parameters constitute information that represents the relative relationship between the position in the real-space image and the position in the virtual space, it becomes possible to obtain information that represents the virtual space position of the camera that is essentially capturing the real space using camera information consisting of transformation parameters.
[0066] Furthermore, an information processing device relating to other aspects may further include a setting unit for installing multiple virtual cameras, and a parameter calculation unit may calculate conversion parameters for imaging the real space based on at least one virtual space image selected from among multiple virtual space images acquired by each of the multiple virtual cameras based on the abundance of reference feature points.
[0067] Based on the above aspects, the virtual space image selected based on the abundance of reference feature points among multiple virtual space images is likely to be an image captured by a virtual camera positioned close to the location of the camera that captured the real-world image. Therefore, the accuracy can be improved by calculating transformation parameters based on such a virtual space image.
[0068] Furthermore, in information processing devices relating to other aspects, the setting unit may install at least multiple virtual cameras along the outer perimeter of the virtual space with the normal direction of the outer perimeter as the imaging direction, or install multiple virtual cameras within the virtual space with the direction of the outer perimeter of the virtual space as the imaging direction.
[0069] Based on the above aspects, multiple virtual cameras can comprehensively capture images of the virtual space corresponding to the space represented in the real-world image.
[0070] Furthermore, in information processing devices relating to other aspects, the parameter calculation unit may calculate transformation parameters for imaging in real space based on a virtual space image having a number of reference specified points equal to or greater than a threshold for reference feature points.
[0071] Based on the above aspects, by setting a threshold corresponding to the number of unknown transformation parameters in the transformation formula, the transformation parameters can be reliably determined.
[0072] Furthermore, an information processing device relating to other aspects may further include a resetting unit that resets a virtual camera in the virtual space to which the transformation parameters for imaging the real space calculated by the parameter calculation unit have been applied as transformation parameters for imaging the virtual space, and the feature point extraction unit may extract feature points from a virtual space image captured by the resetting virtual camera.
[0073] Based on the above aspects, a virtual camera to which the calculated transformation parameters have been applied is highly likely to have positional relationships and characteristics similar to those of the camera that captured the real-world image. By recalculating the transformation parameters based on the virtual-world image captured by the virtual camera to which the calculated transformation parameters have been applied, it is possible to improve their accuracy.
[0074] Furthermore, in information processing devices relating to other aspects, the resetting unit may reinstall at least one virtual camera within a predetermined range around a specific virtual camera, which is a virtual camera that captured the virtual space image having the most reference feature points.
[0075] Based on the above aspects, it is highly probable that the specific virtual camera is a camera placed in the virtual space close to the camera that captured the real-world image. Therefore, by recalculating the transformation parameters based on the virtual space image obtained from a virtual camera reinstalled within a predetermined range around the specific virtual camera, it is possible to further improve its accuracy.
[0076] Furthermore, in information processing devices relating to other aspects, the resetting unit may reconfigure at least one virtual camera so that a virtual space image is captured in which a portion of the virtual space image captured by a specific virtual camera overlaps with a portion of the region containing reference feature points.
[0077] According to the above aspects, a virtual space image is acquired by a reinstalled virtual camera that includes reference feature points located at the edges of the imaging area of the virtual space image captured by a specific virtual camera. This makes it possible to improve the accuracy of the transformation parameters related to lens distortion.
[0078] Furthermore, in information processing devices relating to other aspects, the parameter calculation unit may calculate transformation parameters related to imaging in the real space based on each of a plurality of virtual space images, and the output unit may output the transformation parameters obtained by processing the calculated plurality of transformation parameters using a predetermined statistical method as camera information.
[0079] Based on the aspects described above, each transformation parameter calculated based on the virtual space image acquired by each of the multiple virtual cameras is statistically processed to obtain the final output transformation parameter. This makes it possible to improve the accuracy of the transformation parameters output as camera information.
[0080] Furthermore, in the information processing device relating to other aspects, the conversion parameters include six predetermined camera external parameters, four camera internal parameters, and five parameters relating to the lens distortion of the camera, and the parameter calculation unit uses the following conversion formula (1) to represent the relationship between the three-dimensional coordinates (X, Y, Z) in the three-dimensional space and the two-dimensional coordinates (u, v) in the two-dimensional image. In this case, by substituting the 3D reference space coordinates into 3D coordinates (X, Y, Z) and the 2D reference image coordinates into 2D coordinates (u, v), we obtain camera external parameters R and t, and camera internal parameters f, each with 3 degrees of freedom. x , f y , c x , c y , and the parameter k related to camera lens distortion. 1 ,k 2 ,k 3 , p 1 , p 2 Alternatively, you may calculate the following:
[0081] Based on the above aspects, a transformation parameter consisting of 15 variables can be calculated as camera information.
[0082] Next, we will describe a process that accurately projects objects such as people included in real-space images captured by a real-space camera onto a virtual-space image vp of a virtual space constructed by a virtual-space model, using information (conversion parameters) that represent the position of the camera capturing the real-space image in the virtual space. The purpose of this disclosure is to accurately project objects (such as people) in the real space captured by a real-space camera (hereinafter referred to as the real-space camera) onto a virtual-space model (virtual-space image).
[0083] Figure 10 shows a real-world camera 100 placed in the real-world space of this disclosure, and the real-world image rp and virtual space vs obtained therefrom. The real-world camera 100 corresponds to camera rc in the above. In this disclosure, it is necessary that the reference space coordinates (or virtual reference image coordinates of the virtual-world image vp), which are the three-dimensional coordinates of the reference feature points in the virtual-world image vp, and the reference image coordinates, which are the coordinates of the said reference feature points in the real-world image rp, correspond precisely. However, slight discrepancies may occur in the virtual-world model that provides the virtual space.
[0084] Figure 11 shows a virtual space image vp. In the figure, the door S, indicated by the thick border in the real space, does not coincide with the edge portion of the door in the virtual space image vp. That is, there is a slight discrepancy between the edge portion of the door (thick border portion) represented in the real space image rp and the edge portion of the door in the virtual space image vp. The virtual space image vp is an image captured by a virtual camera positioned according to the above process, but due to the characteristics of the virtual space model, there may be a slight discrepancy between it and the real space image rp. Therefore, if the real space image rp includes objects such as people, those objects will also be represented with a discrepancy in the virtual space image vp.
[0085] When such a discrepancy occurs, if an object such as a person is imaged in real space and then projected into virtual space, a discrepancy will occur. In this disclosure, an affine transformation process is performed on an object in real space to absorb the slight discrepancy that occurs when an object such as a person in real space is imaged by a real camera 100 in real space and then projected into virtual space.
[0086] Figure 12 is a diagram illustrating the operation overview of the virtual camera imaging position determination process and the subsequent affine transformation process of the present disclosure. The real space image acquisition unit 11 acquires a real space image rp (S101), and the virtual space image acquisition unit 13 acquires a virtual space image vp (S102). The feature point extraction unit 14 extracts feature points from the real space image rp and the virtual space image vp, respectively, and the feature point matching unit 15 performs feature point matching (S103). Then, the parameter calculation unit 16 performs the transformation parameter calculation process (S104).
[0087] Then, if the parameter calculation unit 16 determines that the variation in the transformation parameters is below a certain level, the process proceeds to S106, where the virtual camera position and focal length are changed according to the transformation parameters, and a virtual space image vp is obtained (S106, S102). This process is repeated until the variation in the transformation parameters falls below a certain level (threshold) (S105). When the variation in the transformation parameters falls below a certain level, the transformation matrix generation unit 101 and the position adjustment unit 102 perform affine transformation processing (S107). Processes S101 to S106 described above are schematic representations of the processes shown in Figure 8, and in substance, the processes are carried out according to Figure 8.
[0088] In this way, by calculating appropriate transformation parameters for the virtual camera, the position of the real camera 100 in the virtual space can be determined, and a virtual space image vp can be obtained using that position as the position of the virtual camera. Then, by compositing the real space image rp (containing the objects) of the real camera 100, which has undergone an affine transformation, onto the virtual space image vp, accurate alignment between the virtual space image vp and the real space image rp (containing the objects) can be achieved.
[0089] The functional configuration of the information processing device 10a for performing the above operations will now be described. Figure 13 is a block diagram showing the functional configuration of the information processing device 10a. As shown in the figure, the information processing device 10a includes a real-space image acquisition unit 11, a setting unit 12, a virtual-space image acquisition unit 13, a feature point extraction unit 14, a feature point matching unit 15, a parameter calculation unit 16, an output unit 17, a reset unit 18, a position calculation unit 19, a transformation matrix generation unit 101, and a position adjustment unit 102. In addition to the functions of the information processing device 10 for obtaining the transformation parameters described above, this information processing device 10a includes a transformation matrix generation unit 101 for generating an affine transformation matrix and a position adjustment unit 102 for performing the affine transformation.
[0090] The transformation matrix generation unit 101 is the part that generates the affine transformation matrix. This transformation matrix generation unit 101 generates the affine transformation matrix based on the feature points (reference image coordinates) of the real-space image rp and the corresponding feature points (virtual reference image coordinates) of the virtual-space image vp. Details will be described later.
[0091] The position adjustment unit 102 synthesizes objects in the real space captured by the real camera 100 into a virtual space image, and at that time, fine-tunes the position of the object's synthesis using an affine transformation matrix. This virtual space image vp is an image captured from the position of the virtual camera calculated by the position calculation unit 19. The object is a person or the like obtained from the real space image rp captured by the real camera 100, for example, the skeleton of the object extracted by OpenPose. OpenPose is executed by the real space image acquisition unit 11.
[0092] OpenPose is a known technique that can detect the positions of joints or skeletons (hereinafter referred to as skeletons or bones) of the human body from images or videos and estimate a person's pose based on this. This technique is used in various fields such as sports motion analysis, dance choreography, and interactive applications. In this disclosure, assuming that the real-world camera 100 is a surveillance camera, the aim is to reflect the skeletons of people and other objects contained in the real-world image onto a virtual-world image. In this disclosure, OpenPose is used as the technique for extracting the skeletons of objects, but it is not limited to this. For example, MediaPipe, PoseNet, HRNet (High-Resolution Network), or AlphaPose may also be used.
[0093] This information processing device 10a can determine the position of a virtual camera in the virtual space based on the real-world image captured by the real-world camera 100, and generate an affine transformation matrix based on the difference between the virtual space image (feature points) obtained by capturing based on that position and the real-world image (feature points). Then, by inputting the skeleton of the object captured by the real-world camera 100 into the generated affine transformation matrix, the precise coordinates on the virtual space image vp can be determined. This affine transformation matrix can absorb the difference between the real-world image and the virtual space image.
[0094] Furthermore, in this disclosure, the transformation matrix generation unit 101 divides the real space image and the virtual space image into predetermined division units and obtains an affine transformation matrix for each division unit. Figure 14 shows the functional configuration of the transformation matrix generation unit 101 that obtains an affine transformation matrix for each division unit.
[0095] This transformation matrix generation unit 101 is configured to include an image segmentation unit 101a, a feature point verification unit 101b, an affine transformation matrix generation unit 101c, and an affine transformation matrix storage unit 101d.
[0096] The image division unit 101a is the part that divides the real-space image rp and the virtual-space image vp into predetermined division units. In this disclosure, the image division unit 101a divides the real-space image and the virtual-space image into a 3x3 division unit as the first division unit and a 4x4 division unit as the second division unit to obtain multiple divided images. Of course, the image division unit 101a may also divide into division units other than those specified above.
[0097] The feature point count verification unit 101b checks whether each divided image contains a predetermined number of feature points or more, and enlarges the divided image until the predetermined number of feature points or more is reached. In this disclosure, the feature point count verification unit 101b enlarges the divided image until there are three or more feature points. As described above, feature points are reference feature points, and the feature point count verification unit 101b enlarges the image so that there are three or more virtual reference image coordinates d' and reference image coordinates d, respectively. Since an affine transformation matrix generally contains six variables, at least three feature points are required to solve it.
[0098] The affine transformation matrix generation unit 101c is the part that generates an affine transformation matrix for each segmented image. The affine transformation matrix generation unit 101c determines the affine transformation matrix based on the virtual reference image coordinates and reference image coordinates of each segmented image. Since the shift from the virtual space image vp differs depending on the position in the real space image rp, it is preferable to determine an affine transformation matrix that corresponds to that position.
[0099] The affine transformation matrix storage unit 101d is the part that stores the generated affine transformation matrix in association with the divided image. For example, the affine transformation matrix generation unit 101c assigns an identifier to the divided image and generates an affine transformation matrix in association with each identifier.
[0100] The position adjustment unit 102 (see Figure 13) adjusts the position of the composite object in the virtual space image by placing the object in the real space image captured by the real camera 100 into an affine transformation matrix. In this disclosure, the output value of the affine transformation matrix is a value obtained by adjusting the position of the object in the real space image to match the shift in the virtual space model.
[0101] In this disclosure, as described above, the position of the virtual camera in the virtual space is determined by the transformation parameters. This position is the same as the position of the real camera 100 in the real space. Therefore, the real space image rp and the virtual space image vp are roughly identical. However, due to the characteristics of the virtual space model, the real space image rp and the virtual space image vp are not exactly the same, so an affine transformation matrix is used to fine-tune the discrepancy. In this disclosure, the skeleton of an object in the real space captured by the real camera 100 is extracted, and by performing an affine transformation on that skeleton, the object (skeleton) can be synthesized at an appropriate position in the virtual space image.
[0102] Figure 15 shows a specific example of an affine transformation matrix. This affine transformation matrix can output the virtual reference image coordinates (u', v') of a virtual space image vp by inputting the reference image coordinates (u, v) of a real space image rp. Here, a, b, c, and d are parameters that indicate the rotation of the image, and t x ,t y This parameter indicates translation.
[0103] The transformation matrix generation unit 101 can obtain the affine transformation matrix by solving the affine transformation matrix shown in Figure 15 using at least three reference image coordinates (u, v) of the real space image rp and the corresponding virtual reference image coordinates (u', v') of the virtual space image vp. That is, the above parameters a, b, c, d, t x ,t y It is possible to find this.
[0104] In this disclosure, affine transformation is a process for correcting a virtual space model that contains errors. Generally, since affine transformation uses six variables, it is possible to solve it as a general polynomial if there are at least three matching feature points.
[0105] Next, the affine transformation process of the information processing apparatus 10a of the present disclosure (processing S107 in Figure 12) will be described. Figure 16 is a flowchart showing the operation of the transformation matrix generation unit 101 in the information processing apparatus 10a of the present disclosure.
[0106] In the process S105 of Figure 12, once the transformation parameters of the transformation matrix B are obtained and the position of the virtual camera is identified, the image division unit 101a divides the virtual space image vp obtained by the virtual camera at that position and the real space image rp captured by the real camera 100 into a first division unit and a second division unit, respectively (S201). The feature point count confirmation unit 101b confirms the number of feature points contained in one of the divided images of the real space image rp and the virtual space image vp (for example, the divided images rdv311, vdv311, rdv411, vdv411 in Figure 17) (S202). These feature points are those obtained in the process S5 of Figure 8. The feature point count confirmation unit 101b then determines whether each divided image contains a predetermined number (for example, three) or more feature points (S203).
[0107] If the feature point count verification unit 101b determines that the divided image does not have a predetermined number of feature points or more, it processes the divided image to enlarge it and obtains an enlarged divided image rdv311e (S204). For example, in Figure 18, the divided image rdv311 of the 3x3 division unit is enlarged to the left and bottom. This enlargement process is carried out until the number of feature points exceeds a predetermined number. The same process is performed on other divided images. The direction and width of the enlargement can be arbitrary, but for example, the divided image rdv311 may be enlarged by 10% vertically and horizontally in the downward and rightward directions.
[0108] The affine transformation matrix generation unit 101c generates an affine transformation matrix for each segmented image using the feature points (reference image coordinates d(u,v) and virtual reference image coordinates d'(u',v')) in the segmented image (or the enlarged segmented image) and stores it in the affine transformation matrix storage unit 101d (S205). The affine transformation matrix generated here is associated with each segmented image before enlargement, not with the enlarged segmented image. It is generated using feature points contained in the enlarged segmented image, but in that case, the segmented image before enlargement does not have feature points (for example, a white wall), so it is based on the idea that it is better to use feature points in the vicinity of that (for example, a white wall).
[0109] If the feature point verification unit 101b has other segmented images for which it is necessary to generate an affine transformation matrix, it returns to processing S202 and processes those other segmented images (S202). Once the affine transformation matrix generation unit 101c has generated an affine transformation matrix for all segmented images (S206: NO), the position adjustment unit 102 performs the generated affine transformation processing on the images of objects (such as people) included in the real-space image rp captured by the real-world camera 100 (S207).
[0110] The position adjustment unit 102 performs position adjustment processing using a different affine transformation matrix for each divided image. Furthermore, it processes the output values by performing an averaging process on the output values of the affine transformation matrices of the divided images obtained from different division units. Other processing methods besides averaging may also be used. For example, a weighted average may be used.
[0111] Figure 19 shows the process of performing an averaging operation on the output values from the affine transformation matrix. In Figure 19(a), the real-space image rp, which includes an object such as a person, is divided into a first division unit. Here, for example, points P1 and P2 are included. Points P1 and P2 are pixels that constitute the object. In this disclosure, if the object is a person, the person's movements are represented by the person's skeleton (bones). In this disclosure, if points P1 and P2 of the real-space image rp are pixels that constitute the person's skeleton, then in the virtual-space image vp, points P1 and P2 are transformed into points P100 and P200 by the affine transformation matrix of the division image G11. This affine transformation matrix is an affine transformation matrix calculated in correspondence with the position of the division image G11 when the real-space image rp is divided into 3x3 division units. In Figure 19(b), points P100 and P200 are transformed to different positions from points P1 and P2.
[0112] Similarly, in Figures 19(c) and (d), the real-space image rp is divided into 4x4 units, and positions P1 and P2 are transformed into positions P210 and P220 using an affine transformation matrix. Here, since position P1 is included in the divided image G21 and position P2 is included in the divided image G22, the transformation process is performed using the affine transformation matrix calculated for each.
[0113] Then, as shown in Figure 19(e), the positions P100 and P210, which are transformed from position P1 respectively, are used to calculate the average of these positions, thereby obtaining position P100a. Since the distortion of the virtual space model differs depending on the location, multiple affine transformation matrices are required. In this disclosure, the distortion can be corrected or mitigated by obtaining and integrating multiple affine transformation matrices using the procedure described above. Position P100a is the position in the virtual space image corresponding to point P1, and is the position where the distortion of the virtual space model is reduced.
[0114] Next, the effects of the information processing device 10a of this disclosure will be explained. In this disclosure, the real space image acquisition unit 11 acquires a real space image rp captured by the real camera 100. The virtual space image acquisition unit 13 acquires a virtual space image vp captured by a virtual camera vc installed at a position in the virtual space vs corresponding to the position of the real camera 100, in the virtual space vs represented by a virtual space model vm corresponding to the real space.
[0115] The transformation matrix generation unit 101 generates a positional transformation matrix for alignment from the reference image coordinates d and virtual reference image coordinates d' of the real-space image rp. More specifically, the transformation matrix generation unit 101 generates a positional transformation matrix that performs a transformation process to displace the reference image coordinates d of the real-space image rp to the virtual reference image coordinates d' of the virtual-space image vp corresponding to the reference image coordinates d. This positional transformation matrix is, for example, an affine transformation matrix. The position adjustment unit 102 uses the positional transformation matrix to align objects (such as skeletons of people) included in the real-space image with the virtual-space image. The positional transformation matrix is generated in advance, and then imaging is performed by the real-space camera 100, followed by positional adjustment using the generated positional transformation matrix.
[0116] This configuration allows for the high-precision projection of objects contained in real-world images into virtual space (virtual space images). Generally, there is a slight discrepancy between the virtual space represented by a virtual space model attempting to represent real space and the real space itself. In this disclosure, this discrepancy can be absorbed by using positional transformation matrices such as affine transformations, enabling accurate projection. As a result of this accurate projection, the three-dimensional position of objects such as people in virtual space can be precisely determined.
[0117] Furthermore, objects such as people exist in the real space. The real-world camera 100 captures images of the real space including the objects, extracts only the skeletons of those objects, and projects them onto the virtual space, thereby compositing the objects into the virtual space. As described in this disclosure, by performing positional adjustments such as affine transformations, the projection of objects into the virtual space can be performed with high precision. Then, using the above transformation formulas (1) to (3), the 3D coordinates of the virtual space can be determined from the 2D coordinates of the objects in the projected virtual space image, and the positional relationship of the objects in the 3D model can be accurately grasped.
[0118] In this disclosure, the position calculation unit 19 functions as an estimation unit that estimates the position of a virtual camera and calculates a transformation parameter to be applied to the virtual camera vc. This transformation parameter is obtained by feature point matching between the real-space image rp and the virtual-space image vp. The virtual camera vc to which this transformation parameter is applied can then be estimated to correspond to the shooting position of the virtual camera in the virtual space (i.e., the shooting position of the real camera 100 in the real space).
[0119] The virtual space image acquisition unit 13 then acquires a virtual space image vp taken from a virtual camera vc (i.e., the real camera shooting position) to which the conversion parameters calculated by the position calculation unit 19 have been applied.
[0120] Thus, when there are many common reference feature points between the real-space image rp and the virtual-space image vp, the position of the real-space camera 100 in that real space and the position of the virtual camera in the virtual space can be considered to be corresponding positions. As described above, the virtual-space model vm may have some discrepancies with the real space, but for affine transformation processing, it is necessary for the virtual-space image vp and the real-space image rp to coincide to some extent. In this disclosure, the above transformation parameters (transformation matrix B) are used to estimate the position of the virtual camera in order to make the virtual-space image vp and the real-space image rp coincide to some extent, but this is not the only method. Any method can be used to make the position of the virtual camera vc in the virtual space coincide to some extent with the position of the real camera 100 in the real space. Alternatively, the position of the virtual camera vc in the virtual space and the position of the real camera 100 in the real space may be fixed positions so that they can be associated.
[0121] Furthermore, the transformation matrix generation unit 101 may generate a position transformation matrix (affine transformation matrix) for each of the divided images obtained by dividing the real-space image rp into several parts. The position adjustment unit 102 uses the position transformation matrix to adjust the position of objects included in the real-space image to the virtual space image.
[0122] This allows for the generation of positional transformation matrices appropriate to each segmented image, enabling correction tailored to each segmented image.
[0123] More specifically, the process is as follows: The image division unit 101a divides the real-space image rp and the virtual-space image vp to obtain divided images. The affine transformation matrix generation unit 101c then generates a first positional transformation matrix for each of the multiple first divided image regions (for example, the divided image G11 in Figure 19) obtained by dividing the real-space image rp into first division units (for example, 3x3 division units). The affine transformation matrix generation unit 101c also generates a second positional transformation matrix for each of the multiple second divided images (for example, the divided image G21 in Figure 19) obtained by dividing the real-space image rp into second division units (for example, 4x4 division units) that are different from the first division units.
[0124] The position adjustment unit 102 then calculates a first target coordinate (for example, position P100 in Figure 19) for each first divided image using a first position transformation matrix, and calculates a second target coordinate (for example, position P210 in Figure 19) for each second divided image using a second position transformation matrix. The position adjustment unit 102 calculates a target coordinate (for example, position P100a in Figure 19) based on the first and second target coordinates, and performs position adjustment using this target coordinate. This target coordinate is based, for example, on the average value of the first and second target coordinates.
[0125] This configuration allows for smooth correction of the seams between segmented images by using values based on multiple affine transformation results, such as the average value.
[0126] Furthermore, the transformation matrix generation unit 101 (feature point verification unit 101b) enlarges the segmented image if there are no reference image coordinates, or if the number of reference image coordinates is less than the threshold, until the number of reference image coordinates is equal to or greater than the threshold. Then, the feature point verification unit 101b uses the reference image coordinates in the enlarged segmented image to determine the position transformation matrix of the segmented image before enlargement.
[0127] This configuration allows for accurate correction by securing the necessary feature points even in segmented images with few feature points.
[0128] Furthermore, the transformation matrix generation unit 101 (feature point count verification unit 101b) does not perform the transformation matrix generation process if, in the real-space image rp or virtual-space image vp, there are no reference image coordinates d or virtual reference image coordinates d' that serve as feature points, or if the number of reference image coordinates d or virtual reference image coordinates d' is less than or equal to a threshold.
[0129] In real-space images (rp) and virtual-space images (vp), if no feature points exist, transformation matrices cannot be generated even with segmentation processing, so this processing is omitted. This reduces the processing load.
[0130] In this disclosure, the affine transformation process is important among the above processes, but it is not essential to prepare an affine transformation matrix for each segmented image or to enlarge the segmented images.
[0131] The computer may be used as an information processing program to make it function as the information processing device 10a of this embodiment. The information processing program P1 shown in Figure 9 is composed of a main module m10 that comprehensively controls information processing in the information processing device 10, a real-space image acquisition module m11, a setting module m12, a virtual-space image acquisition module m13, a feature point extraction module m14, a feature point matching module m15, a parameter calculation module m16, an output module m17, a reset module m18, and a position calculation module m19. Each of the modules m11 to m19 realizes the respective functions for each of the functional units 11 to 19.
[0132] In addition, the information processing program P1 may further include a transformation matrix generation module that functions as a transformation matrix generation unit 101, and a position adjustment module that functions as a position adjustment unit 102.
[0133] The block diagram shown in Figure 13 represents functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may also be realized by combining software with the one or more devices described above.
[0134] Functions include, but are not limited to, judgment, decision, determination, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, assumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, a functional block (configuration part) that enables transmission is called a transmitting unit or transmitter. In all cases, as mentioned above, the method of implementation is not particularly limited.
[0135] For example, the information processing device 10 (including the information processing device 10a) in one embodiment of the present invention may function as a computer. Figure 20 is a diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. Physically, the information processing device 10 may be configured as a computer device including a processor 1001, memory 1002, storage 1003, communication device 1004, input device 1005, output device 1006, bus 1007, etc.
[0136] In the following explanation, the term "device" can be replaced with "circuit," "device," "unit," etc. The hardware configuration of the information processing device 10 may include one or more of the devices shown in Figure 20, or it may be configured to omit some of the devices.
[0137] Each function in the information processing device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, allowing the processor 1001 to perform calculations and control communication by the communication device 1004, as well as the reading and / or writing of data in the memory 1002 and storage 1003.
[0138] The processor 1001 controls the entire computer, for example, by running an operating system. The processor 1001 may consist of a central processing unit (CPU) that includes interfaces with peripheral devices, control devices, arithmetic units, registers, etc. For example, the various functional units 11 to 19, the transformation matrix generation unit 101, or the position adjustment unit 102 shown in Figure 1 may be implemented in the processor 1001.
[0139] Furthermore, the processor 1001 reads programs (program code), software modules, and data from the storage 1003 and / or communication device 1004 into the memory 1002, and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described in the above embodiment. For example, each of the functional units 11 to 19 of the information processing device 10, the transformation matrix generation unit 101, or the position adjustment unit 102 may be implemented by a control program stored in the memory 1002 and operated by the processor 1001. Although the above-described processes have been explained as being executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented on one or more chips. The program may also be transmitted from a network via a telecommunications line.
[0140] The memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The memory 1002 may also be called a register, cache, main memory, etc. The memory 1002 can store executable programs (program code), software modules, etc., for carrying out an information processing method according to one embodiment of the present invention.
[0141] The storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. The storage 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including memory 1002 and / or storage 1003.
[0142] The communication device 1004 is hardware (transceiver / receiver device) for communicating between computers via a wired and / or wireless network, and is also referred to as a network device, network controller, network card, communication module, etc.
[0143] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).
[0144] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may consist of a single bus, or different buses may be used for communication between devices.
[0145] Furthermore, the information processing device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.
[0146] The notification of information is not limited to the embodiments described herein and may be carried out by other means. For example, the notification of information may be carried out by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or combinations thereof. RRC signaling may also be called RRC messages, and may be, for example, RRC Connection Setup messages, RRC Connection Reconfiguration messages, etc.
[0147] Each aspect / embodiment described in this disclosure may be applied to at least one of the following: LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA®, GSM®, CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi®), IEEE 802.16 (WiMAX®), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth®, and other appropriate systems, as well as next-generation systems extended based thereon. Furthermore, multiple systems may be applied in combination (for example, a combination of at least one of LTE and LTE-A with 5G).
[0148] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described in this disclosure may be reordered, provided they do not contradict each other. For example, the methods described in this disclosure present various step elements using exemplary order and are not limited to the specific order presented.
[0149] The specific operations described in this disclosure as being performed by a base station may, in some cases, be performed by its upper node. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal can be performed by the base station and at least one other network node (for example, an MME or S-GW, but not limited to these). Although the above example illustrates the case where there is one other network node besides the base station, it may also be a combination of multiple other network nodes (for example, an MME and an S-GW).
[0150] Information can be output from a higher layer (or lower layer) to a lower layer (or higher layer). Input and output may also occur via multiple network nodes.
[0151] Input and output information may be stored in a specific location (e.g., memory) or managed in a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be sent to other devices.
[0152] The determination may be made by a value represented by one bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).
[0153] Each aspect / embodiment described in this disclosure may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).
[0154] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.
[0155] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name.
[0156] Furthermore, software, instructions, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cable, fiber optic cable, twisted pair, and digital subscriber lines (DSL) and / or wireless technologies such as infrared, radio, and microwave, these wired and / or wireless technologies are included in the definition of a transmission medium.
[0157] The information, signals, etc. described in this disclosure may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0158] In addition, terms described in this disclosure and / or terms necessary for understanding this specification may be replaced with terms having the same or similar meaning.
[0159] The terms “system” and “network” as used in this disclosure are interchangeable.
[0160] Furthermore, the information, parameters, etc., described in this disclosure may be expressed as absolute values, relative values from a given value, or by corresponding other information. For example, wireless resources may be indicated by an index.
[0161] The names used for the parameters described above are not restrictive in any way. Furthermore, the formulas and other expressions using these parameters may differ from those expressly disclosed in this disclosure. Various channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable name, and therefore, the various names assigned to these various channels and information elements are not restrictive in any way.
[0162] As used in this disclosure, the terms “determining” and “determining” may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, or inquiring (e.g., searching in a table, database, or other data structure), or ascertaining. “Determining” may also include receiving (e.g., receiving information), transmitting (e.g., sending information), inputting, outputting, or accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."
[0163] As used in this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."
[0164] Where the terms “first,” “second,” etc., are used in this disclosure, no reference to those elements shall generally limit the quantity or order of those elements. These terms may be used herein as a convenient way to distinguish between two or more elements. Accordingly, references to the first and second elements shall not imply that only two elements may be employed therein, or that the first element must precede the second element in any way.
[0165] In the configuration of each of the above devices, "means" may be replaced with "part," "circuit," "device," etc.
[0166] To the extent that “include,” “including,” and their variations are used herein or in the claims, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used herein or in the claims is not intended to be exclusive OR.
[0167] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.
[0168] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different."
[0169] The information processing system 1 disclosed herein may have the following configuration.
[0170] [1] A feature point extraction unit that extracts feature points from a real-space image captured of a real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to the real space and represented by a virtual-space model, and the position of the virtual camera is a viewpoint position for capturing an image projected from the virtual-space model; and a feature point matching unit that matches the feature points of the real-space image with the feature points of the virtual-space image. An information processing device comprising: a parameter calculation unit that calculates the transformation parameters for imaging the real space by substituting the reference space coordinates in the virtual space of the reference feature points, which are the matched feature points, and the reference image coordinates, which are the coordinates of the reference feature points in the real space image, into a predetermined transformation formula that expresses the relationship between the three-dimensional coordinates of a specific point, which is a specific point in three-dimensional space, and the two-dimensional coordinates of the specific point in an image captured of the three-dimensional space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on the transformation parameters for imaging the virtual space by the virtual camera and the virtual reference image coordinates, which are the coordinates of the reference feature points in the virtual space image; and an output unit that outputs the transformation parameters for imaging the real space as camera information representing the position in the virtual space of the camera that captured the real space image.
[0171] [2] The information processing apparatus according to [1], further comprising a setting unit for setting up a plurality of virtual cameras, wherein the parameter calculation unit calculates the conversion parameters for imaging the real space based on at least one of the plurality of virtual space images selected from among a plurality of virtual space images acquired by each of the plurality of virtual cameras based on the abundance of the reference feature points.
[0172] [3] The information processing apparatus according to [2], wherein the setting unit installs at least a plurality of virtual cameras along the outer perimeter of the virtual space with the normal direction of the outer perimeter as the imaging direction, or installs a plurality of virtual cameras within the virtual space with the direction of the outer perimeter of the virtual space as the imaging direction.
[0173] [4] The information processing apparatus according to any one of [1] to [3], wherein the parameter calculation unit calculates the conversion parameters for imaging the real space based on the virtual space image having a number of reference specific points equal to or greater than a threshold for the reference feature points.
[0174] [5] The information processing apparatus according to [4], further comprising: a resetting unit which resets the virtual camera to which the conversion parameters for imaging the real space calculated by the parameter calculation unit have been applied as conversion parameters for imaging the virtual space, the feature point extraction unit which extracts feature points from the virtual space image captured by the reset virtual camera.
[0175] [6] The information processing apparatus according to [5], wherein the resetting unit reinstalls at least one virtual camera within a predetermined range around a specific virtual camera which is the virtual camera that captured the virtual space image having the most reference feature points.
[0176] [7] The information processing apparatus according to [6], wherein the resetting unit resets at least one of the virtual cameras so that a virtual space image is captured in which a part of the virtual space image captured by the specific virtual camera overlaps with a part of the region including the reference feature point.
[0177] [8] The information processing apparatus according to any one of [1] to [7], wherein the parameter calculation unit calculates the conversion parameters relating to imaging the real space based on each of the plurality of virtual space images, and the output unit outputs the conversion parameters obtained by processing the plurality of calculated conversion parameters by a predetermined statistical method as camera information.
[0178] [9] The conversion parameters include six predetermined external camera parameters, four internal camera parameters, and five parameters relating to the projection of a three-dimensional space onto a two-dimensional image, and the parameter calculation unit uses the following conversion formula (1) to represent the relationship between three-dimensional coordinates (X, Y, Z) in three-dimensional space and two-dimensional coordinates (u, v) in a two-dimensional image. In this case, by substituting the three-dimensional reference space coordinates into the three-dimensional coordinates (X, Y, Z) and the two-dimensional reference image coordinates into the two-dimensional coordinates (u, v), the camera external parameters R and t and the camera internal parameter f, each having 3 degrees of freedom, are obtained. x , f y , c x , c y , and the parameter k relating to the lens distortion of the camera. 1 ,k 2 ,k 3 , p 1 , p 2 An information processing apparatus according to any one of items [1] to [8], which calculates .
[0179]
[10] A feature point extraction step for extracting feature points from a real-space image captured of a real space and at least one virtual-space image, wherein the virtual-space image is an image captured by a virtual camera installed in the virtual space of a virtual space corresponding to the real space and represented by a virtual-space model, and the position of the virtual camera is a viewpoint position for capturing an image projected from the virtual-space model; and a feature point matching step for matching the feature points of the real-space image with the feature points of the virtual-space image. An information processing method executed by a processor, comprising: a parameter calculation step of calculating the transformation parameters for imaging the real space by substituting the reference space coordinates in the virtual space of the reference feature points, which are the matched feature points, and the reference image coordinates, which are the coordinates of the reference feature points in the real space image, into a predetermined transformation formula that expresses the relationship between the three-dimensional coordinates of a specific point, which is a specific point in three-dimensional space, and the two-dimensional coordinates of the specific point in an image captured of the three-dimensional space using predetermined transformation parameters, wherein the reference space coordinates are calculated based on the transformation parameters for imaging the virtual space by the virtual camera and the virtual reference image coordinates, which are the coordinates of the reference feature points in the virtual space image; and an output step of outputting the transformation parameters for imaging the real space as camera information representing the position in the virtual space of the camera that captured the real space image.
[0180] Furthermore, the apparatus of this disclosure may have the following configuration.
[0181] [1A] A device comprising: a real-space image acquisition unit that acquires a real-space image captured by a real-space camera; a virtual-space image acquisition unit that acquires a virtual-space image captured by a virtual camera installed at a position in the virtual space corresponding to the position of the real-space camera, of a virtual space represented by a virtual-space model corresponding to the real space; a generation unit that generates a position transformation matrix for alignment from the reference image coordinates of the real-space image and the virtual reference image coordinates; and a position adjustment unit that uses the position transformation matrix to align objects included in the real-space image with the virtual-space image.
[0182] [2A] The apparatus according to [1A], comprising: an estimation unit that estimates the shooting position of a virtual camera in a virtual space by using feature point matching between the real space image and the virtual space image, wherein the virtual space image acquisition unit acquires a virtual space image taken from the virtual camera shooting position estimated by the estimation unit.
[0183] [3A] The apparatus according to [1A] or [2A], wherein the generation unit generates a position transformation matrix that performs a transformation process to displace the reference image coordinates of the real space image to the virtual reference image coordinates of the virtual space image corresponding to the reference image coordinates.
[0184] [4A] The apparatus according to [3A], wherein the position transformation matrix is an affine transformation matrix.
[0185] [5A] The apparatus according to [3A] or [4A], wherein the generation unit generates a position transformation matrix for each divided image obtained by dividing the real space image into several parts, and the position adjustment unit performs position adjustment using the position transformation matrix.
[0186] [6A] The apparatus according to [5A], wherein the generation unit generates a first position transformation matrix for each of the multiple first divided image regions obtained by dividing the real space image in a first division unit, generates a second position transformation matrix for each of the multiple second divided images obtained by dividing the real space image in a second division unit different from the first division unit, and the position adjustment unit calculates a first destination coordinate for each of the first divided images using the first position transformation matrix, calculates a second destination coordinate for each of the second divided images using the second position transformation matrix, calculates a destination coordinate based on the first destination coordinate and the second destination coordinate, and performs position adjustment using the destination coordinate.
[0187] [7A] The apparatus according to [5A] or [6A], wherein the generating unit enlarges the divided image until the number of reference image coordinates becomes equal to or greater than the threshold, for divided images in which no reference image coordinates exist or the number of reference image coordinates is less than the threshold, and uses the reference image coordinates in the enlarged divided image to determine the positional transformation matrix of the divided image before enlargement.
[0188] [8A] The apparatus according to any one of [5A] to [7A], wherein the generation unit does not perform generation processing if there are no reference image coordinates or virtual reference image coordinates that serve as feature points in the real space image or the virtual space image, or if the number of reference image coordinates or virtual reference image coordinates is less than or equal to a threshold.
[0189] [9A] A method of the apparatus comprising: a real space image acquisition step of acquiring a real space image captured by a real camera; a virtual space image acquisition step of acquiring a virtual space image captured by a virtual camera installed at a position in the virtual space corresponding to the position of the real camera, of a virtual space represented by a virtual space model corresponding to the real space; a generation step of generating a position transformation matrix for alignment from the reference image coordinates of the real space image and the virtual reference image coordinates; and a position adjustment step of aligning an object included in the real space image to the virtual space image using the position transformation matrix.
[0190] 1... Information processing system, 10... Information processing device, 11... Real-world image acquisition unit, 12... Setting unit, 13... Virtual-world image acquisition unit, 14... Feature point extraction unit, 15... Feature point matching unit, 16... Parameter calculation unit, 17... Output unit, 18... Reset unit, 19... Position calculation unit, 21... Real-world image storage unit, 22... Virtual-world model storage unit, 101... Transformation matrix generation unit, 101a... Image segmentation unit, 101b... Feature point count confirmation unit, 101c... Affine transformation matrix Generation unit, 101d... Affine transformation matrix storage unit, 102... Position adjustment unit, M1... Recording medium, m11... Real space image acquisition module, m12... Setting module, m13... Virtual space image acquisition module, m14... Feature point extraction module, m15... Feature point matching module, m16... Parameter calculation module, m17... Output module, m18... Reset module, m19... Position calculation module, P1... Information processing program.
Claims
1. A device comprising: a real-space image acquisition unit that acquires a real-space image captured by a real-space camera; a virtual-space image acquisition unit that acquires a virtual-space image captured by a virtual camera installed at a position in the virtual-space corresponding to the position of the real-space camera, representing a virtual-space model corresponding to the real-space; a generation unit that generates a position transformation matrix for alignment from the reference image coordinates of the real-space image and the virtual reference image coordinates of the virtual-space image; and a position adjustment unit that uses the position transformation matrix to align objects included in the real-space image with the virtual-space image.
2. The apparatus according to claim 1, comprising: an estimation unit that estimates the shooting position of a virtual camera in a virtual space by using feature point matching between the real space image and the virtual space image, wherein the virtual space image acquisition unit acquires a virtual space image taken from the shooting position of the virtual camera estimated by the estimation unit.
3. The apparatus according to claim 1, wherein the generation unit generates a position transformation matrix that performs a transformation process to displace the reference image coordinates of the real space image to the virtual reference image coordinates of the virtual space image corresponding to the reference image coordinates.
4. The apparatus according to claim 3, wherein the positional transformation matrix is an affine transformation matrix.
5. The apparatus according to claim 3, wherein the generation unit generates a position transformation matrix for each of the divided images obtained by dividing the real space image into several parts, and the position adjustment unit performs position adjustment using the position transformation matrix.
6. The apparatus according to claim 5, wherein the generation unit generates a first position transformation matrix for each of the multiple first division image regions obtained by dividing the real space image in a first division unit, generates a second position transformation matrix for each of the multiple second division images obtained by dividing the real space image in a second division unit different from the first division unit, and the position adjustment unit calculates a first destination coordinate for each of the first division images using the first position transformation matrix, calculates a second destination coordinate for each of the second division images using the second position transformation matrix, calculates a destination coordinate based on the first destination coordinate and the second destination coordinate, and performs position adjustment using the destination coordinate.
7. The apparatus according to claim 5, wherein the generation unit, for divided images in which no reference image coordinates exist or in which the number of such coordinates is less than a threshold, enlarges the divided images until the number of reference image coordinates is equal to or greater than a threshold, and uses the reference image coordinates in the enlarged divided images to determine the positional transformation matrix of the divided images before enlargement.
8. The apparatus according to claim 5, wherein the generation unit does not perform generation processing if there are no reference image coordinates or virtual reference image coordinates that serve as feature points in the real space image or the virtual space image, or if the number of reference image coordinates or virtual reference image coordinates is less than or equal to a threshold.
9. A method of the apparatus comprising: a real-space image acquisition step of acquiring a real-space image captured by a real-space camera; a virtual-space image acquisition step of acquiring a virtual-space image captured by a virtual camera installed at a position in the virtual-space corresponding to the position of the real-space camera, of a virtual-space represented by a virtual-space model corresponding to the real-space; a generation step of generating a position transformation matrix for alignment from the reference image coordinates of the real-space image and the virtual reference image coordinates of the virtual-space image; and a position adjustment step of aligning an object included in the real-space image with the virtual-space image using the position transformation matrix.
Citation Information
Patent Citations
Image alignment method, image alignment device and terminal equipment
CN112348863A
Method and device for constructing three-dimensional map, electronic equipment and storage medium
CN117011481A
Registration method, device and system and storage medium
CN118115546A