Information processing system, information processing apparatus, program, and information processing method
The information processing system addresses misalignment issues between three-dimensional models and real-space images by using a coordinate transformation matrix for self-position correction, enabling high-precision virtual space generation and accurate virtual information overlay.
Patent Information
- Application Number
- JP2023197082
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-06-02
AI Technical Summary
Three-dimensional models used as map data may have dimensional and orientational differences with real-space images, leading to misalignment when virtual information is superimposed on real-space images.
An information processing system that includes a terminal for acquiring real-time images of the real space and an information processing apparatus that stores a coordinate transformation matrix. This system performs self-position estimation and correction using the matrix to align the real-space images with three-dimensional models, enabling high-precision virtual space generation.
The system effectively generates a high-precision virtual space by correcting self-position information based on the coordinate transformation matrix, ensuring accurate overlay of virtual information on real-space images.
Smart Images

Figure 2025083618000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system, an information processing apparatus, a program, and an information processing method, and more particularly to an information processing system, an information processing apparatus, a program, and an information processing method using a three-dimensional model.
Background Art
[0002] In recent years, for example, in a video of the real space displayed on a terminal such as a smartphone or a glasses-type wearable device, a video by computer graphics is superimposed as virtual information, and it is possible to enjoy the harmony between the real space and the virtual information. An information processing technology for providing a virtual space as an extended reality space in which the real space is extended has been proposed.
[0003] Patent Document 1 proposes a technique for displaying virtual information according to a user's operation in an information processing technique for superimposing an image of virtual information on an image of the real space and displaying it on a user's display device.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] By the way, this type of information processing technology is used to generate a virtual space by using a three-dimensional model, also called a BIM (Building Information Modeling) model, as map data and superimposing an image of the real space on this map data, and the generated virtual space is used for various purposes.
[0006] Such a three-dimensional model may have differences in the dimensions and orientations between the objects of the three-dimensional model and the images of the real space, or the modeling accuracy may be lower than that of the images of the real space. Therefore, when the three-dimensional model is used as map data and the images of the real space are overlaid, the map data and the images of the real space may be misaligned.
[0007] The present invention has been made in view of the above circumstances, and an object thereof is to provide an information processing system, an information processing apparatus, a program, and an information processing method capable of easily generating a high-precision virtual space.
Means for Solving the Problems
[0008] The information processing system according to the present invention for achieving the above object includes a terminal that acquires an image of the real space in real time as a first image, and an information processing apparatus that is accessed by the terminal and stores a coordinate transformation matrix indicating the correspondence between the image of the real space acquired in advance as a second image and the second image and the three-dimensional model in coordinates. The system performs a self-position estimation process for estimating the self-position of the terminal by collating the first image and the second image, and a self-position information correction process for correcting the self-position information regarding the self-position estimated in the self-position estimation process based on the coordinate transformation matrix.
[0009] According to this, based on the coordinate transformation matrix indicating the correspondence between the second image and the three-dimensional model in coordinates, the self-position information of the terminal estimated by collating the first image and the second image is corrected. Therefore, when virtual information is displayed on the first image of the terminal, a high-precision virtual space can be easily generated by utilizing the three-dimensional model.
[0010] Here, the first image and the second image include both moving images and still images.
[0011] This information processing system performs a virtual information display process for overlaying virtual information on the first image based on the self-position information corrected in the self-position information correction process.
[0012] The coordinate transformation matrix processed by this information processing system is generated by calculating the similarity of feature vectors based on the feature quantities extracted from the point cloud output from the second image and the feature quantities extracted from the point cloud output from the three-dimensional model, and collating the point cloud output from the second image and the point cloud output from the three-dimensional model based on the calculated similarity.
[0013] Similarly, the coordinate transformation matrix is generated by dividing the point cloud output from the second image into the shape of the point cloud and partial features of the point cloud, replacing the shape of the point cloud and the partial features of the point cloud with graph structures respectively, dividing the point cloud output from the three-dimensional model into the shape of the point cloud and partial features of the point cloud, replacing the shape of the point cloud and the partial features of the point cloud with graph structures respectively, and collating the graph structure based on the second image and the graph structure based on the three-dimensional model.
[0014] Similarly, the coordinate transformation matrix is generated by dividing the point cloud output from the second image into a plurality of regions to extract feature quantities from the point cloud, dividing the point cloud output from the three-dimensional model into a plurality of regions to extract feature quantities from the point cloud, calculating the similarity of feature vectors based on the feature quantities extracted from the point cloud of the second image and the feature quantities extracted from the point cloud of the three-dimensional model, and collating the point cloud output from the second image and the point cloud output from the three-dimensional model based on the calculated similarity.
[0015] Furthermore, similarly, the coordinate transformation matrix is generated by partially collating the point cloud output from the second image and the point cloud output from the three-dimensional model, using the collation result as teacher data, and generating it by a trained model obtained by performing machine learning on the learning data associated with the point cloud output from the second image and the point cloud output from the three-dimensional model using the teacher data.
[0016] An information processing apparatus according to the present invention for achieving the above object is an information processing apparatus including a processor, a memory storing a program, and a storage unit, wherein a coordinate transformation matrix indicating, in coordinates, the correspondence between an image of the real space acquired in advance and an image of the real space and a three-dimensional model generated by collating them is stored in the storage unit, and the apparatus executes a self-position estimation process for estimating the self-position of the terminal by collating an image of the real space acquired by the terminal in real time with the image of the real space acquired in advance, and a self-position information correction process for correcting self-position information regarding the self-position estimated in the self-position estimation process based on the coordinate transformation matrix.
[0017] A program according to the present invention for achieving the above object causes an information processing apparatus implemented by a computer in which a coordinate transformation matrix indicating, in coordinates, the correspondence between an image of the real space acquired in advance and an image of the real space and a three-dimensional model generated by collating them is stored in a storage unit to execute a self-position estimation process for estimating the self-position of the terminal by collating an image of the real space acquired by the terminal in real time with the image of the real space acquired in advance, and a self-position information correction process for correcting self-position information regarding the self-position estimated in the self-position estimation process based on the coordinate transformation matrix.
[0018] An information processing method according to the present invention for achieving the above object causes an information processing apparatus implemented by a computer in which a coordinate transformation matrix indicating, in coordinates, the correspondence between an image of the real space acquired in advance and an image of the real space and a three-dimensional model generated by collating them is stored in a storage unit to execute a self-position estimation process for estimating the self-position of the terminal by collating an image of the real space acquired by the terminal in real time with the image of the real space acquired in advance, and a self-position information correction process for correcting self-position information regarding the self-position estimated in the self-position estimation process based on the coordinate transformation matrix.
Advantages of the Invention
[0019] According to this invention, a high-precision virtual space can be easily generated.
Brief Description of the Drawings
[0020]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0021] Next, based on FIGS. 1 to 13, an information processing system according to an embodiment of the present invention will be described.
[0022] FIG. 1 is a block diagram for explaining an outline of the configuration of the information processing system according to the present embodiment. As shown in the figure, the information processing system 10 mainly includes a plurality of user terminals 20 which are terminals and an information processing apparatus 30, and these are connected to each other via a network N such as the Internet.
[0023] In the present embodiment, the user terminal 20 is held by a plurality of users 1 who use a service using the information processing system 10 provided by the operator 2 described below, and the information processing apparatus 30 is managed by the operator 2 who provides a service using the information processing system 10.
[0024] In the present embodiment, the service using the information processing system 10 is a service that provides a virtual space, in which an object which is virtual information is superimposed on an image in the real space, to the user terminal 20.
[0025] Next, the specific configuration of each part of the information processing system 10 will be described.
[0026] In the present embodiment, the user terminal 20 is implemented by a smartphone which is a portable information terminal. However, for example, it may be implemented by a tablet-type computer, a desktop or notebook computer, or any information processing terminal (information processing technology) such as a glasses-type wearable device that executes other projection processing.
[0027] FIG. 2 is a block diagram for explaining an outline of the configuration of the user terminal 20. As shown in the figure, the user terminal 20 mainly includes a control unit 21, a camera 22, a display 23, and sensors 24.
[0028] In this embodiment, the control unit 21 controls each part of the user terminal 20 such as the camera 22, the display 23, and the sensors 24, and is composed of, for example, a processor, a memory, a storage, a transmission / reception unit, and the like.
[0029] In this embodiment, the control unit 21 stores an application for executing a service provided by the information processing system 10 or a browser capable of browsing a website, and based on the processing of the program in the information processing apparatus 30, the service is executed in the user terminal 20 via the application or the browser.
[0030] The camera 22 captures an image with any arbitrary object as a subject and acquires the image. In this embodiment, an image of the real space is acquired in real time as the first image.
[0031] In this embodiment, various screen interfaces and the like of the service executed in the user terminal 20 are displayed on the display 23.
[0032] This display 23 is a so-called touch panel that receives input of information by contact with the display surface, and is implemented by various technologies such as a resistive film method and a capacitance method.
[0033] In this embodiment, information for operating an application that executes a service is input via this display 23. This information is input based on an arbitrary operation of the user 1 on the display 23 (for example, an operation of tapping or swiping the screen, or an operation of dragging and dropping an icon or the like displayed on the screen).
[0034] In this embodiment, the sensors 24 are composed of a gyro sensor, an acceleration sensor, and the like, and detect the position and orientation of the user terminal 20 in the real space.
[0035] In this embodiment, the position of the user terminal 20 on a second image, which will be described later, is estimated as the self-position based on the position and orientation of the user terminal 20 detected by these sensors 24 and the first image acquired by the camera 22.
[0036] In this embodiment, the information processing apparatus 30 shown in FIG. 1 is implemented by a computer, for example, a desktop or laptop computer.
[0037] FIG. 3 is a block diagram for explaining the outline of the configuration of the information processing apparatus 30. As shown in the figure, the information processing apparatus 30 mainly includes a processor 31, a memory 32, a storage 33, a transmission / reception unit 34, and an input / output unit 35, and these are electrically connected to each other via a bus 36.
[0038] The processor 31 is an arithmetic unit that controls the operation of the information processing apparatus 30, controls the transmission and reception of data between elements, and performs processes necessary for executing an application program.
[0039] In this embodiment, the processor 31 is, for example, a CPU (Central Processing Unit), and executes an application program developed in the memory 32 described below to perform each process.
[0040] The memory 32 is implemented by a main memory device composed of a volatile storage device such as a DRAM (Dynamic Random Access Memory).
[0041] This memory 32 is used as a working area for the processor 31, and at the same time, stores a BIOS (Basic Input / Output System) executed when the information processing apparatus 30 is started up, and various setting information and the like.
[0042] The storage 33 stores data and the like used for various processes by an application program or the like.
[0043] The transmission / reception unit 34 connects the information processing apparatus 30 to the network N. This transmission / reception unit 34 may conform to a wireless communication standard such as Wi-Fi, or may be equipped with a short-range communication interface such as Bluetooth (registered trademark) or BLE (Bluetooth Low Energy).
[0044] To the input / output unit 35, information input devices such as a keyboard and a mouse, and output devices such as a display are connected as necessary. In the present embodiment, a keyboard, a mouse, and a display are respectively connected.
[0045] The bus 36 transmits, for example, address signals, data signals, and various control signals among the connected processor 31, memory 32, storage 33, transmission / reception unit 34, and input / output unit 35.
[0046] FIG. 4 is a block diagram for explaining an outline of functions of the information processing apparatus 30 of the information processing system 10. As shown in the drawing, the information processing apparatus 30 includes a three-dimensional city model storage unit 30a, a three-dimensional city model point cloud output unit 30b, a three-dimensional city model point cloud storage unit 30c, a second image point cloud storage unit 30d, a collation processing unit 30e, a coordinate transformation matrix generation unit 30f, a coordinate transformation matrix storage unit 30g, a first image point cloud output unit 30h, a self-position estimation processing unit 30i, a self-position information correction processing unit 30j, a virtual information storage unit 30k, and a virtual information display processing unit 30l.
[0047] The three-dimensional city model storage unit 30a, the three-dimensional city model point cloud storage unit 30c, the second image point cloud storage unit 30d, the coordinate transformation matrix storage unit 30g, and the virtual information storage unit 30k are realized by partitioning the storage area of the storage 33.
[0048] On the other hand, the three-dimensional city model point cloud output unit 30b, the collation processing unit 30e, the coordinate transformation matrix generation unit 30f, the first image point cloud output unit 30h, the self-position estimation processing unit 30i, the self-position information correction processing unit 30j, and the virtual information display processing unit 30l are realized by the processor 31 executing a program stored in the memory 32.
[0049] In the three-dimensional urban model storage unit 30a, in the present embodiment, a three-dimensional urban model, which is a three-dimensional model, is stored.
[0050] FIG. 5 is a diagram for explaining the outline of the three-dimensional urban model. As shown in the figure, in the present embodiment, the three-dimensional urban model M is information for realizing, in the information processing apparatus 30, a three-dimensional model identical to buildings and the like existing in the real space, and is composed of various information such as appearance information of buildings and the like grasped from the three-dimensional model, design information, equipment information, etc. of buildings and the like corresponding to the three-dimensional model, and is also information called a BIM (Building Information Modeling) model.
[0051] The three-dimensional urban model point cloud output unit 30b shown in FIG. 4 outputs, in the present embodiment, a point cloud based on the pixel information of the three-dimensional urban model M as data, and the output point cloud of the three-dimensional urban model M is stored in the three-dimensional urban model point cloud storage unit 30c. An example of the point cloud of the three-dimensional urban model M is shown in FIG. 6.
[0052] In the second image point cloud storage unit 30d shown in FIG. 4, in the present embodiment, the point cloud output based on the pixel information of the image of the real space previously acquired as the second image by the business operator 2 is stored as data. An example of the point cloud of the second image is substantially the same as the example of the point cloud of the three-dimensional urban model M shown in FIG. 6.
[0053] The collation processing unit 30e executes a process of collating the second image and the three-dimensional urban model M, and in the present embodiment, as shown in FIG. 7, executes a first process S1, a second process S2, a third process S3, a fourth process S4, a fifth process S5, and a sixth process S6.
[0054] For these first process S1 to sixth process S6, from the viewpoint of optimally collating the second image and the three-dimensional urban model M, one arbitrarily selected process may be executed, or a plurality of arbitrarily selected processes may be combined and executed, or all the processes may be executed.
[0055] Figure 8 is a flowchart for explaining the outline of the process of the first process S1 in the collation processing unit 30e. As shown in the figure, in the first process S1, first, in step S10, feature amounts are extracted from the point group of the second image stored in the second image point group storage unit 30d.
[0056] On the other hand, in step S11, which is executed before or after step S10 or simultaneously with step S10, feature amounts are extracted from the point group of the three-dimensional city model M stored in the three-dimensional city model point group storage unit 30c.
[0057] In the subsequent step S12, for the feature amounts extracted from the point group of the second image and the feature amounts extracted from the point group of the three-dimensional city model M, the similarity of the feature vectors is calculated based on the cosine similarity.
[0058] Next, in step S13, based on the calculated similarity of the feature vectors, the point group of the second image and the point group of the three-dimensional city model M are collated to associate (correspond) the point group of the second image with the point group of the three-dimensional city model M.
[0059] Figure 9 is a flowchart for explaining the outline of the process of the second process S2 in the collation processing unit 30e. As shown in the figure, in the second process S2, first, in step S20, the scales of the point group of the second image and the point group of the three-dimensional city model M are each normalized.
[0060] Next, in step S21, a process is executed to divide the point group output from the second image normalized in step S20 into the shape of the macro-structural point group and the partial features of the micro-structural point group.
[0061] In the subsequent step S22, the shape of the point group and the partial features of the point group divided in step S21 are each replaced with a graph structure that connects between the respective point groups.
[0062] On the one hand, in step S23, a process is executed to classify the point cloud output from the three-dimensional city model M normalized in step S20 into the shape of the macro-structural point cloud and the partial features of the micro-structural point cloud.
[0063] Furthermore, in step S24, the shape of the point cloud and the partial features of the point cloud classified in step S23 are respectively replaced with a graph structure that connects between each point cloud.
[0064] Subsequently, in step S25, the graph structure of the shape of the point cloud of the second image is compared with the graph structure of the shape of the point cloud of the three-dimensional city model M, and the graph structure of the partial features of the point cloud of the second image is compared with the graph structure of the partial features of the point cloud of the three-dimensional city model M, and the feature quantities of the point cloud of the second image and the feature quantities of the point cloud of the three-dimensional city model M are associated (corresponded).
[0065] FIG. 10 is a diagram for explaining the outline of the third process S3 in the collation processing unit 30e. As shown in the figure, in the third process S3, the point cloud output from the second image and the point cloud output from the three-dimensional city model M are each divided into a plurality of regions in a grid pattern.
[0066] In the present embodiment, after the point cloud output from the second image and the point cloud output from the three-dimensional city model M are divided into a grid pattern, feature quantities are extracted from each of the point cloud of the second image and the point cloud of the three-dimensional city model M.
[0067] After that, in the same manner as in the first process S1, for the feature quantities extracted from the point cloud of the second image and the feature quantities extracted from the point cloud of the three-dimensional city model M, the similarity of the feature vectors is calculated based on the cosine similarity, and based on the calculated similarity of the feature vectors, the point cloud of the second image and the point cloud of the three-dimensional city model M are collated, and the point cloud of the second image and the point cloud of the three-dimensional city model M are associated (corresponded).
[0068] In this way, after dividing the point cloud output from the second image and the point cloud output from the three-dimensional city model M into a grid pattern, by comparing the point cloud of the second image with the point cloud of the three-dimensional city model M, the association (correspondence) between the point cloud of the second image and the point cloud of the three-dimensional city model M can be executed quickly and efficiently.
[0069] In the fourth process S4 shown in FIG. 7, in this embodiment, similar to the third process S3, first, after dividing the point cloud output from the second image and the point cloud output from the three-dimensional city model M into a plurality of regions in a grid pattern, the point cloud of the second image and the point cloud of the three-dimensional city model M are each converted into pixel information (pixels).
[0070] After that, the similarity is calculated for the pixel information obtained by converting the point cloud of the second image and the pixel information obtained by converting the point cloud of the three-dimensional city model M, and based on the calculated similarity, the point cloud of the second image is compared with the point cloud of the three-dimensional city model M to associate (correspond) the point cloud of the second image with the point cloud of the three-dimensional city model M.
[0071] In the fifth process S5, in this embodiment, first, only the characteristic parts in each of the point clouds of the second image and the three-dimensional city model M are partially compared to associate the point cloud of the second image with the point cloud of the three-dimensional city model M, and the resulting data is used as teacher data to generate learning data in association with the point cloud of the second image and the point cloud of the three-dimensional city model M.
[0072] Subsequently, in the fifth process S5, a learned model is generated by performing machine learning using the generated learning data. As methods for performing machine learning, various algorithms such as neural networks, random forests, and SVM (Support Vector Machine) are appropriately used.
[0073] Using this learned model, the association (correspondence) between the point cloud of the second image and the point cloud of the three-dimensional city model M is performed.
[0074] In the present embodiment, the sixth process S6 executes a process of adjusting the resolution, scale, noise level, etc. of the point cloud of the second image and the point cloud of the three-dimensional city model M between the point cloud of the second image and the three-dimensional city model M.
[0075] Since this sixth process S6 is a preprocess for performing the matching between the point cloud of the second image and the point cloud of the three-dimensional city model M, only the sixth process S6 is not executed alone.
[0076] In the present embodiment, the coordinate transformation matrix generation unit 30f shown in FIG. 4 executes a process of generating a coordinate transformation matrix based on the association (correspondence) between the point cloud of the second image and the point cloud of the three-dimensional city model M executed in the matching by the matching processing unit 30e.
[0077] In the present embodiment, this coordinate transformation matrix is metadata for correcting the self-position information (first self-position information), which is information regarding the position (self-position) of the user terminal 20 on the second image estimated by the self-position estimation processing unit 30i described later, to the self-position information (second self-position information) on the three-dimensional city model M so as to be an optimal position on the three-dimensional city model M.
[0078] In the present embodiment, the coordinate transformation matrix is stored in the coordinate transformation matrix storage unit 30g.
[0079] In the present embodiment, the first image point cloud output unit 30h executes a process of outputting, as data, a point cloud based on pixel information from an image of the real space acquired in real time as the first image by the user terminal 20 via the camera 22.
[0080] In the present embodiment, the self-position estimation processing unit 30i collates the point cloud of the first image acquired in real time by the user terminal 20 with the point cloud of the second image stored in the second image point cloud storage unit 30d based on the position and orientation of the user terminal 20 detected by the sensors 24 of the user terminal 20, and generates self-position information for estimating the self-position of the user terminal 20 on the second image (self-position estimation processing).
[0081] In this embodiment, the self-position information correction processing unit 30j executes processing (self-position information correction processing) for correcting the self-position information estimated by the self-position estimation processing to the self-position information on the three-dimensional city model M that is the optimal position on the three-dimensional city model M based on the coordinate transformation matrix.
[0082] FIG. 11(a) is a diagram for explaining the outline of the self-position information correction processing. As shown in the figure, the self-position information correction processing can be expressed using a matrix (determinant).
[0083] This self-position information correction processing executes processing for converting the coordinates (x1, y1, z1) based on the point cloud of the first image grasped as self-position information into the coordinates (x2, y2, z2) based on the point cloud of the three-dimensional city model M grasped as self-position information that is the optimal position on the three-dimensional city model M.
[0084] Here, (tx, ty, tz) represents the translation amount with respect to the coordinate axes (x, y, z), (a11, a12, a13, a21, a22, a23, a31, a32, a33) represents the linear transformation coefficients with respect to the coordinate axes (x, y, z), and (0, 0, 0, 1) represents the homogeneous coordinate system.
[0085] The result of the processing by this self-position information correction processing is shown in FIG. 11(b).
[0086] In this embodiment, virtual information (objects) to be overlaid on the first image is stored in the virtual information storage unit 30k shown in FIG. 4.
[0087] In this embodiment, the virtual information display processing unit 30l executes processing for overlaying virtual information on the first image based on the self-position information corrected by the self-position information correction processing (virtual information display processing).
[0088] Next, the outline of the processing for generating the coordinate transformation matrix according to this embodiment will be described.
[0089] FIG. 12 is a flowchart for explaining an outline of a process for generating a coordinate transformation matrix. As shown in the figure, first, in step S30, a point cloud based on the pixel information of the three-dimensional city model M stored in the three-dimensional city model storage unit 30a is output as data.
[0090] In this embodiment, the output point cloud of the three-dimensional city model M is stored in the three-dimensional city model point cloud storage unit 30c.
[0091] In the subsequent step S31, the point cloud of the three-dimensional city model M is compared with the point cloud of the second image stored in the second image point cloud storage unit 30d. In this embodiment, among the first process S1 to the sixth process S6 related to the comparison, any one selected process may be executed, a combination of any selected multiple processes may be executed, or all processes may be executed.
[0092] After performing the association (correspondence) between the point cloud of the second image and the point cloud of the three-dimensional city model M executed in the comparison between the point cloud of the three-dimensional city model M and the point cloud of the second image, in step S32, a process for generating a coordinate transformation matrix is executed.
[0093] In this embodiment, the generated coordinate transformation matrix is stored in the coordinate transformation matrix storage unit 30g.
[0094] Next, an outline of the process of the information processing system according to this embodiment will be described.
[0095] FIG. 13 is a flowchart for explaining an outline of the process of the information processing system. As shown in the figure, when user 1 acquires a first image in real time via user terminal 20 in step S40, information processing apparatus 30 acquires the first image.
[0096] When the information processing apparatus 30 acquires the first image, in step S41, a point cloud based on pixel information is output as data from the first image, and in step S42, the point cloud of the first image is collated with the point cloud of the second image stored in the second image point cloud storage unit 30d to estimate the self-position information of the user terminal 20 (self-position estimation process).
[0097] Subsequently, in step S43, based on the coordinate transformation matrix stored in the coordinate transformation matrix storage unit 30g, the estimated self-position information is corrected (self-position information correction process).
[0098] In the present embodiment, by the self-position information correction process, the self-position information estimated by the self-position estimation process is corrected to self-position information that is in the optimal position on the three-dimensional city model M.
[0099] Thereafter, in step S44, based on the self-position information corrected by the self-position information correction process, virtual information is superimposed on the first image (virtual information display process). In this case, various types of information such as design information and equipment information of buildings and the like corresponding to the three-dimensional models included in the three-dimensional city model M are also superimposed on the first image.
[0100] In this way, since the self-position information of the user terminal 20 estimated by collating the first image and the second image is corrected based on the coordinate transformation matrix indicating the correspondence between the second image and the three-dimensional city model M in coordinates, when displaying virtual information on the first image of the user terminal 20, an existing three-dimensional city model M can be utilized to generate a high-precision virtual space.
[0101] This virtual space may be any virtual space realized on the user terminal 20, such as, for example, an AR (Augumented Reality) space (augmented reality space), a VR (Virtual Reality) space (virtual reality space), or a so-called metaverse space where multiple users 1 can simultaneously access via avatars and trade items and the like.
[0102] Note that the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of the invention.
[0103] In the above embodiment, the case where the point group output from the second image and the point group output from the three-dimensional city model M are collated to generate the coordinate transformation matrix has been described. However, for example, the point group output from the second image and the three-dimensional city model M itself may be collated to generate the coordinate transformation matrix.
[0104] In the above embodiment, the case where the three-dimensional model is the three-dimensional city model M has been described. However, for example, various three-dimensional models such as CAD drawings and arbitrary three-dimensional images may be included.
[0105] In the above embodiment, the case where the information processing apparatus 30 is implemented by a computer managed by the business operator 2 has been described. However, for example, the information processing apparatus 30 may be a computer implemented in a cloud environment.
Explanation of Reference Numerals
[0106] 1 User 2 Business Operator 10 Information Processing System 20 User Terminal 30 Information Processing Apparatus
Claims
1. A terminal that acquires in real time an image of the real space as a first image, An information processing apparatus that is accessed by the terminal and stores a coordinate transformation matrix indicating, in coordinates, the correspondence between the image of the real space acquired in advance as a second image and the second image and the three-dimensional model generated by collating the second image and the three-dimensional model, Self-position estimation processing for estimating the self-position of the terminal by collating the first image and the second image, Self-position information correction processing for correcting self-position information regarding the self-position estimated by the self-position estimation processing based on the coordinate transformation matrix, An information processing system that executes the above processes.
2. Executes virtual information display processing for superimposing virtual information on the first image based on the self-position information corrected by the self-position information correction processing, The information processing system according to Claim 1.
3. The coordinate transformation matrix Calculates the similarity of the feature vectors based on the feature amounts extracted from the point group output from the second image and the feature amounts extracted from the point group output from the three-dimensional model, Is generated by collating the point group output from the second image and the point group output from the three-dimensional model based on the calculated similarity, The information processing system according to Claim 1 or 2.
4. The coordinate transformation matrix Divides the point group output from the second image into the shape of the point group and partial features of the point group, and replaces the shape of the point group and the partial features of the point group with graph structures, respectively, Divides the point group output from the three-dimensional model into the shape of the point group and partial features of the point group, and replaces the shape of the point group and the partial features of the point group with graph structures, respectively, Is generated by collating the graph structure based on the second image and the graph structure based on the three-dimensional model, The information processing system according to Claim 1 or 2.
5. The coordinate transformation matrix Divides the point group output from the second image into a plurality of regions and extracts feature amounts from the point group, and divides the point group output from the three-dimensional model into a plurality of regions and extracts feature amounts from the point group, Calculates the similarity of the feature vectors based on the feature amounts extracted from the point group of the second image and the feature amounts extracted from the point group of the three-dimensional model, Generated by collating the point cloud output from the second image and the point cloud output from the three-dimensional model based on the calculated similarity. The information processing system according to claim 1 or 2. **Claim 6** The coordinate transformation matrix is Partially collate the point cloud output from the second image and the point cloud output from the three-dimensional model, and use the collation result as teacher data. Machine learning is performed on the learning data associated with the point cloud output from the second image and the point cloud output from the three-dimensional model by a learned model generated thereby. The information processing system according to claim 1 or 2. **Claim 7** An information processing apparatus including a processor, a memory storing a program, and a storage unit, A coordinate transformation matrix indicating, in coordinates, the correspondence between a pre-acquired image of the real space and the three-dimensional model generated by collating the image of the real space and the three-dimensional model is stored in the storage unit. An ego-position estimation process for estimating the ego-position of the terminal by collating an image of the real space acquired in real time by the terminal with the pre-acquired image of the real space; An ego-position information correction process for correcting the ego-position information regarding the ego-position estimated in the ego-position estimation process based on the coordinate transformation matrix; An information processing apparatus that executes the above. **Claim 8** In an information processing apparatus implemented by a computer in which a coordinate transformation matrix indicating, in coordinates, the correspondence between a pre-acquired image of the real space and the three-dimensional model generated by collating the image of the real space and the three-dimensional model is stored in a storage unit, An ego-position estimation process for estimating the ego-position of the terminal by collating an image of the real space acquired in real time by the terminal with the pre-acquired image of the real space; An ego-position information correction process for correcting the ego-position information regarding the ego-position estimated in the ego-position estimation process based on the coordinate transformation matrix; A program for causing the above to be executed. **Claim 9** An information processing apparatus implemented by a computer in which a coordinate transformation matrix indicating, in coordinates, the correspondence between a pre-acquired image of the real space and the three-dimensional model generated by collating the image of the real space and the three-dimensional model is stored in a storage unit, An ego-position estimation process for estimating the ego-position of the terminal by collating an image of the real space acquired in real time by the terminal with the pre-acquired image of the real space; An own-position information correction process that corrects own-position information regarding the own position estimated by the own-position estimation process based on the coordinate conversion matrix, An information processing method for executing the above.
Citation Information
Patent Citations
Game processing program, game processing method, and game processing device
JP2019201942A