Information processing apparatus, and program
The information processing device enhances realism in three-dimensional models by aligning virtual camera positions and orientations with wide-angle images, enabling realistic renderings of buildings and objects.
Patent Information
- Application Number
- JP2025027386
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-22
- Filing Date
- 2025-02-21
- Publication Date
- 2025-09-03
AI Technical Summary
Three-dimensional area model information lacks realism due to its geometric nature, making it impractical to apply texture images to numerous buildings effectively.
An information processing device that acquires three-dimensional area model information, wide-angle image data, estimates camera positions, and aligns virtual camera positions and orientations to create realistic renderings by comparing virtual space images with wide-angle images.
Provides realistic information using three-dimensional model data through simple processing, ensuring accurate alignment and rendering of buildings and objects in a virtual space.
Smart Images

Figure 2025129064000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device and a program that utilizes three-dimensional area model information. [Background technology]
[0002] In recent years, three-dimensional area model information, including information on three-dimensional models of buildings, has been widely released, and various examples of its use are being considered (Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Ministry of Land, Infrastructure, Transport and Tourism, "Project PLATEAU: Open Data Completion of 3D City Models of 56 Cities Nationwide," [online], August 6, 2021, [Retrieved February 8, 2024], Internet<URL:https: / / www.mlit.go.jp / report / press / toshi03_hh_000078.html> Summary of the Invention [Problem to be solved by the invention]
[0004] However, because the information on 3D models is basically geometric information, it is currently not possible to provide information with a sense of reality. Therefore, although it is conceivable to use photographs of real spaces and apply them as textures to the surfaces of the 3D models, it is not realistic to prepare texture images for each of the shapes of the 3D models of numerous buildings.
[0005] The present invention has been made in consideration of the above-mentioned situation, and one of its objects is to provide an information processing device and program that can provide realistic information using three-dimensional model information in a relatively simple manner. [Means for solving the problem]
[0006] One aspect of the present invention that solves the problems of the above-mentioned conventional examples is an information processing device that includes a model acquisition means for acquiring three-dimensional area model information including information on a three-dimensional model of a building; an image acquisition means for acquiring wide-angle image data captured while moving through the area in which the building is located; an estimated position acquisition means for acquiring information representing the trajectory of movement of a camera that captured the acquired wide-angle image data and estimating an estimated capturing position of the wide-angle image data on the trajectory; and a determination means for placing the three-dimensional model of the building in a specified three-dimensional virtual space, and determining position and orientation information of a virtual camera in the virtual space that corresponds to the trajectory of movement of the camera that captured the wide-angle image data based on a comparison between a rendering image of the virtual space from the viewpoint of a virtual camera set in the virtual space and the wide-angle image data, wherein the wide-angle image data, the coordinate information of the virtual space determined corresponding to the wide-angle image data, and the acquired three-dimensional area model information are subjected to specified processing. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide realistic information using three-dimensional model information in a relatively simple manner by utilizing wide-angle image data captured while moving. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram illustrating an example of the configuration of an information processing device according to an embodiment of the present invention. [Figure 2] 1 is a functional block diagram illustrating an example of an information processing device according to an embodiment of the present invention. [Figure 3] 10A and 10B are explanatory diagrams illustrating an example of alignment processing by an information processing device according to an embodiment of the present invention. [Figure 4] FIG. 2 is a flowchart illustrating an example of the operation of the information processing device according to the embodiment of the present invention. [Figure 5] FIG. 10 is a flowchart illustrating another example of the operation of the information processing device according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0009] An embodiment of the present invention will be described with reference to the drawings. As shown in Fig. 1, an information processing device 1 according to the embodiment of the present invention includes a control unit 11, a storage unit 12, an operation unit 13, a display unit 14, and an interface unit 15.
[0010] The control unit 11 is a program-controlled device such as a CPU (processor), and operates according to a program stored in the storage unit 12. In one example of the present embodiment, the control unit 11 operates according to the program stored in the storage unit 12 to acquire three-dimensional area model information such as a city model, including information on three-dimensional models of buildings. The control unit 11 also acquires wide-angle image data captured while moving through the area where the buildings are located.
[0011] The control unit 11 acquires information representing the trajectory of the movement of the camera that captured the acquired wide-angle image data, estimates the estimated capturing position of the wide-angle image data on the trajectory, and places a three-dimensional model of the building in a predetermined three-dimensional virtual space, and determines coordinate information in the virtual space corresponding to the trajectory of the movement of the camera that captured the wide-angle image data based on a comparison between a rendering image of the virtual space from a viewpoint set in the virtual space and the wide-angle image data.
[0012] The control unit 11 then subjects the wide-angle image data, the virtual space coordinate information determined in accordance with the wide-angle image data, and the three-dimensional area model information to predetermined processing. The detailed operation of the control unit 11 will be described later.
[0013] The storage unit 12 is a memory device, a disk device, or the like, and stores programs executed by the control unit 11. The storage unit 12 also operates as a work memory for the control unit 11.
[0014] The operation unit 13 is a keyboard, a mouse, etc., which accepts user operations and outputs information indicating the content of the operations to the control unit 11. The display unit 14 is a display, etc., which displays and outputs images, etc., in accordance with instructions input from the control unit 11.
[0015] The interface unit 15 includes a USB (Universal Serial Bus) interface and a network interface, and acquires various information from disk devices connected via this interface unit 15 and server devices on a network that are communicatively connected via the interface unit 15, and outputs the information to the control unit 11. In addition, in accordance with instructions from the control unit 11, the interface unit 15 stores or transmits various information to the disk devices, server devices, etc.
[0016] Next, the operation of the control unit 11 will be described. By executing a program stored in the storage unit 12, the control unit 11 is configured to functionally include a registration processing unit 20 and a browsing processing unit 30, as illustrated in Fig. 2. The registration processing unit 20 functionally includes a model acquisition unit 21, an image acquisition unit 22, an estimated position acquisition unit 23, a rendering processing unit 24, and a determination processing unit 25, and the browsing processing unit 30 functionally includes a position and orientation information acquisition unit 31 and a rendering processing unit 32.
[0017] [Registration processing behavior] When the control unit 11 operates as the alignment processing unit 20, the model acquisition unit 21 acquires three-dimensional area model information including information on three-dimensional models of buildings. In one example of the present embodiment, this three-dimensional area model information may be three-dimensional city model data (such as PLATEAU LoD1 data provided by the Ministry of Land, Infrastructure, Transport and Tourism of Japan). Such data generally includes the shapes of buildings in a given area as three-dimensional solid model data (i.e., as three-dimensional geometric data).
[0018] In this embodiment, the three-dimensional area model information does not have to be data including the three-dimensional model of the building, but may include at least information representing the perimeter of a structure such as a building or road in plan view (such as latitude and longitude information of points on the perimeter) and altitude information representing the height of the building from the ground surface. The altitude information representing the height of the building may be the median or average value of the height of each part of the building, or may be a predetermined constant value.
[0019] As an example, the 3D regional model information may be a two-dimensional image of buildings and other structures viewed from above (obtained from a plan view), with the base being a two-dimensional image of the buildings and other structures, with columnar shapes of a certain height used as a substitute for the buildings (pseudo LoD1 data). Here, the two-dimensional image may be generated based on satellite or aerial photographs. As a specific example, by using the basic map information provided by the Ministry of Land, Infrastructure, Transport and Tourism of Japan (https: / / www.gsi.go.jp / kiban / ), it is possible to obtain information on the perimeter of buildings and their elevation (height above the ground surface) (referred to as two-dimensional information), and these can be used to create pseudo LoD1 data.
[0020] In this example, the 3D solid model data of each building does not reflect the height of the actual building, but this does not cause any problems in subsequent processing. Also, when obtaining a rendering image, if the height of the virtual camera is set to, for example, about 2 m (when obtaining a rendering image from near the ground), it is preferable to set the fixed height of the buildings in the above pseudo LoD1 data to a value sufficiently larger than the height value of the virtual camera position, such as 20 m.
[0021] The image acquisition unit 22 moves through an area represented by three-dimensional model area information that the model acquisition unit 21 can acquire (an area where a building represented by information included in the three-dimensional model area information is located), and acquires wide-angle image data R1, R2, ..., RN captured by a camera tilted by an angle λ around a line segment (xN-x1) connecting the start point and end point of the movement trajectory at multiple points x1, x2, ..., xN (where x is a vector quantity representing three-dimensional coordinates) on the movement trajectory. Here, the wide-angle image data Ri (i = 1, 2, ..., N) may be, for example, 360-degree image data (omnidirectional image data) or image data captured over a relatively wide angle range centered on the line of sight of the camera that captured the image. In other words, it is assumed here that the wide-angle image data is captured while the camera moves along the movement trajectory while keeping the camera angle constant. In the following example, it is assumed that this wide-angle image data is a captured image projected onto a sphere centered on the shooting position, converted into an equirectangular coordinate image (ERP image). In this case, pixels of a predetermined color, such as solid black, are arranged in the non-imaged area.
[0022] The estimated position acquisition unit 23 estimates the position xi and angle λ (hereinafter referred to as the "estimated image capture position") of the camera that captured each of the acquired wide-angle image data Ri, and estimates and acquires information representing the trajectory of movement of the camera. Here, the angle of the camera may be expressed as an inclination with respect to the ground (horizontal plane) or an angle around the trajectory of movement.
[0023] As an example, the estimated position acquisition unit 23 estimates the position and orientation x1, x2, ..., xN and angle λ (here, the angle around the line segment connecting the initial position x1 and the end position xN) of a camera that captured wide-angle image data in a global coordinate system (information on latitude, longitude, altitude, and rotation angle) or a local coordinate system for each region (which may be expressed as relative coordinates from a known point, such as latitude, longitude, and altitude, within or outside the region) using a method such as the widely known Visual SLAM (S. Sumikura, M. Shibuya, and K. Sakurada. Openvslam: A versatile visual slam framework. In ACMMM, 2019), and estimates the trajectory of its movement. Since processing such as Visual SLAM is widely known, detailed description thereof will be omitted here. In the following example, the estimated position acquisition unit 23 receives information settings from the user in a global coordinate system (latitude, longitude, altitude, and rotation angle information) for at least one of the starting point and ending point of the movement trajectory (the settings here do not need to be precise as they will be corrected by later processing), and converts the coordinates of each point on the trajectory representing the estimated movement trajectory of the camera that captured the wide-angle image data (the position where each wide-angle image data was captured) into values in this global coordinate system and outputs them.
[0024] The rendering processing unit 24 places the three-dimensional models of buildings included in the three-dimensional regional model information acquired by the model acquisition unit 21 in a specified three-dimensional virtual space, and generates a rendering image of the virtual space from a viewpoint set within the virtual space.
[0025] Specifically, the rendering processing unit 24 sets up a virtual three-dimensional space and determines a conversion relationship between the coordinates of that coordinate system (which may be an XYZ Cartesian coordinate system) and the coordinates in the real global coordinate system (which includes the range of latitude, longitude, and altitude represented by the three-dimensional area model information acquired by the model acquisition unit 21).The rendering processing unit 24 then uses the determined conversion relationship to convert the coordinates (coordinates in the global coordinate system) of the three-dimensional model of the building included in the three-dimensional area model information into the coordinate system of the virtual space, and places the solid model represented by the three-dimensional model in the virtual space.This processing is similar to the general processing of placing three-dimensional model data in a virtual space, so a detailed description will be omitted here.
[0026] Furthermore, the rendering processing unit 24 converts the camera position and orientation information x1, x2, ..., xN (global coordinates are used here) at each point where the wide-angle image data was captured, estimated by the estimated position acquisition unit 23, and the trajectory of the camera movement represented by λ, into trajectory information in virtual space coordinates using the above conversion relationship, and obtains the camera position and orientation information x'1, x'2, ..., x'N in the virtual space that corresponds to the position and orientation information of the camera that captured the wide-angle image data.
[0027] The rendering processing unit 24 refers to the data of the three-dimensional model of the building placed in the virtual space, places a virtual camera (virtual camera) at a position represented by each of the coordinates x′i (i = 1, 2, ...) in the virtual space obtained above, sets the tilt of the virtual camera to an angle λ that represents the attitude of the virtual camera, and renders an image of the virtual space viewed from the virtual camera. The range of the image to be rendered here is set to an angle of view that at least partially includes the angle of view of the wide-angle image data acquired by the image acquisition unit 22. In this example of the present embodiment, the rendering processing unit 24 sets a virtual camera based on the position and orientation information x′i (i = 1, 2, ...), and renders a 360-degree celestial sphere image R′i (i = 1, 2, ...) corresponding to each piece of position and orientation information as an image in the same coordinate system as the wide-angle image data described above (hence, in this example, as an ERP image in equirectangular coordinates). In this case, too, pixels of a predetermined color, such as solid black, are arranged in the area not captured by the virtual camera. The rendering processing unit 24 does not need to obtain rendering images corresponding to all position and orientation information, but may select a portion of it (for example, i=1, p+1, 2p+1... (where p is an integer greater than or equal to 2)) to obtain rendering images.
[0028] The determination processing unit 25 compares the rendered image R'i obtained by the rendering processing unit 24 rendering the virtual space from the viewpoint of the virtual camera set based on the position and orientation information x'i with the wide-angle image data Ri estimated to have been captured at the global coordinate xi corresponding to the coordinate x'i, and based on the result of the comparison, determines the coordinate information in the virtual space corresponding to the movement trajectory of the camera that captured each wide-angle image data (at least the starting coordinates xs, xe of the estimated camera movement trajectory) and the tilt angle λ of the camera.
[0029] As an example, the determination processing unit 25 compares the shape of the building depicted in the rendering image with the shape of the building captured in the wide-angle image data, and determines coordinate information in the virtual space corresponding to the trajectory of movement of the camera that captured the wide-angle image data. Specifically, as shown in Figure 3, the determination processing unit 25 in this example extracts pixel areas representing the building from both the rendering image R'i generated by the rendering processing unit 24 using coordinate x'i as the viewpoint, and wide-angle image data Ri estimated to have been captured at global coordinate x'i corresponding to the coordinate x'i (S1), (S2).
[0030] The method for extracting the pixel area representing the building from the rendering image R'i here involves extracting the area in which the three-dimensional model of the building is drawn, and any of a variety of well-known methods can be used, so detailed explanation will be omitted here.
[0031] In addition, methods such as semantic segmentation (S. Orhan and Y. Bastanlar. Semantic segmentation of outdoor panoramic images. Springer SIVP, Vol. 16, No. 3, pp. 643-650, 2022) can be used to extract pixel areas representing buildings from wide-angle image data Ri.
[0032] The determination processing unit 25 calculates the number of pixels di that is the difference between the pixel area representing the building extracted from the rendering image R'i and the pixel area representing the building extracted from the wide-angle image data Ri, and calculates the sum Σdi (S3). The determination processing unit 25 updates at least a part of the estimated position and orientation information x'i, for example, the start point and end point x'1, x'N, and λ (S4), and causes the rendering processing unit 24 to generate a rendering image R'i rendered using the updated position and orientation information x'i.
[0033] The determination processing unit 25 then extracts the pixel area representing the building from the rendering image R'i (S1), returns to step S3, and calculates the number of pixels di that are the difference between the pixel area representing the building extracted from the rendering image R'i and the pixel area representing the building extracted from the wide-angle image data Ri, and calculates the sum Σdi. The determination processing unit 25 repeats this process thereafter to find x'1, x'N, and λ that minimize the sum Σdi (alignment process).
[0034] The method of iteratively determining x'1, x'N, and λ that minimize this sum Σdi while updating it can be a widely known method such as a random search. When Σdi falls below a predetermined threshold or the number of iterations exceeds a predetermined threshold, the determination processing unit 25 terminates the above process and determines x'i (i=1, 2, ..., N) and λ.
[0035] The determination processing unit 25 then stores the processing results, including the three-dimensional area model information, the determined x′i (i=1, 2, ... N) and λ (i.e., the position and orientation information of the virtual camera), and the corresponding wide-angle image data Ri, in the memory unit 12.
[0036] The results of this processing are used for predetermined processing such as browsing processing, which will be described next.
[0037] [Browsing process behavior] In addition, the control unit 11 of this embodiment acquires the processing results generated in the above-mentioned alignment processing operation (three-dimensional area model information, virtual camera position and orientation information x'i (i = 1, 2, ... N) and λ for rendering processing, and wide-angle image data Ri), and performs the browsing processing exemplified below.
[0038] The control unit 11 that performs this processing executes a program stored in the storage unit 12 to functionally realize a configuration as a browsing processing unit 30. As already described, this browsing processing unit 30 includes a position and orientation information acquisition unit 31 and a rendering processing unit 32.
[0039] The position and orientation information acquisition unit 31 acquires information (hereinafter referred to as target information, since this is the information to be processed) including the three-dimensional area model information generated in the alignment process, the x′i (i=1, 2, ... N) and λ (i.e., the position and orientation information of the virtual camera) determined here, and the corresponding wide-angle image data Ri.
[0040] The rendering processing unit 32 places a three-dimensional model of the building included in the three-dimensional regional model information of the target information within a three-dimensional virtual space using the information acquired from the target information, selects one of the position and orientation information within the virtual space, and sets a virtual camera using the selected position and orientation information.
[0041] The position and orientation information to be selected here may be determined by a position designation operation by the user. Note that, initially, predetermined position and orientation information such as x'1,λ may be selected.
[0042] The rendering processing unit 32 generates, by rendering, a 360-degree celestial sphere image R′i of the virtual space, which is the viewpoint from the virtual camera set using the selected position and orientation information. As already exemplified, the celestial sphere image R′i here may also be an ERP image in an equirectangular coordinate system.
[0043] The rendering processing unit 32 further sets each pixel of the image R′i obtained by the rendering to the value of the corresponding pixel in the wide-angle image data Ri that corresponds to the position and orientation information xi,λ of the virtual camera that is included in the target information and that was set to generate the rendered image R′i (pixel value setting), and outputs the set image. This output is performed, for example, by displaying the result of the pixel value setting on the display unit 14.
[0044] At this time, the rendering processing unit 32 may receive an instruction from the user (an instruction to adjust the angle of view) to set which part (angle of view) of the obtained spherical image should be displayed as a result of the pixel value setting, and determine the range to be displayed in accordance with the instruction.
[0045] The position and orientation information of the virtual camera used in the rendering process is determined so that the buildings included in the rendering result R'i match the buildings captured in the wide-angle image data Ri, so the pixels in the rendering result R'i and the wide-angle image data Ri correspond, and areas in the wide-angle image data Ri where there are no buildings (such as roads or sky) match the corresponding areas in the rendering result R'i.
[0046] The rendering processing unit 32 of this embodiment may also place at least one type of three-dimensional object in a virtual space and perform rendering processing on the three-dimensional object using the position and orientation information xi,λ of the virtual camera. Here, the three-dimensional object may include a type of three-dimensional object (e.g., an avatar) that can be moved in the virtual space by a user's operation.
[0047] In this example, the rendering processing unit 32 places three-dimensional objects in the virtual space together with three-dimensional models of buildings included in the three-dimensional area model information of the target information. As already mentioned, the placement positions of at least some of the three-dimensional objects to be placed (hereinafter referred to as movable objects when a distinction needs to be made) in the virtual space may be controlled by the user.
[0048] Furthermore, when the height position (position in the Z-axis direction) of the bottom side (i.e., the ground) of the three-dimensional model of the building is Z=0, the rendering processing unit 32 may arrange at least a part of this three-dimensional object (which may be a movable object) so that its bottom surface is always in contact with Z=0. In this case, too, the positions of the three-dimensional object, which is a movable object, in the X-axis direction and the Y-axis direction may be changed by the user's object movement operation. A three-dimensional object arranged in this way will be rendered in a state of contact with the ground.
[0049] The rendering processing unit 32 renders a 360-degree omnidirectional image of this virtual space using a virtual camera set using one of the position and orientation information xi,λ in the target information selected by the user's position designation operation, and obtains the resulting rendered image R′i.
[0050] In this case, the 3D object may be occluded by the 3D model of the building, i.e., the area of the 3D object occluded by the 3D model of the building (or the area occluded by another 3D object) will not be drawn in the rendered image.
[0051] Furthermore, in this embodiment, the rendering processing unit 32 may perform rendering of a three-dimensional object by applying a texture map, a bump map, or the like that has been previously associated with the three-dimensional object. As a result, for example, a three-dimensional object of water is rendered with its texture expressed using a texture map of the water surface, or the like.
[0052] At this time, the rendering processing unit 32 sets the pixels of the image R′i of the rendering result, excluding the pixels where the three-dimensional object is drawn, to the pixel values of the corresponding pixels of the wide-angle image data Ri that correspond to the position and orientation information xi,λ of the virtual camera used to obtain this rendering result (pixel value setting).
[0053] The rendering processing unit 32 then outputs the image after setting these pixel values.
[0054] Furthermore, in this embodiment, when an object movement operation of a three-dimensional object is received from the user, the position of the three-dimensional object in the virtual space is moved in accordance with the instruction of the object movement operation, and the process is repeated from rendering.
[0055] When placing a three-dimensional object in this manner, the range of movement of the three-dimensional object within the virtual space may be restricted by the three-dimensional model of the building. Specifically, if the position of the three-dimensional object after movement is included in the area represented by the three-dimensional model, the rendering processing unit 32 returns the position of the three-dimensional object to the position before movement. This prevents the three-dimensional object from entering the interior of the three-dimensional model of the building (collision detection is performed).
[0056] [Operation] This embodiment has the above configuration and operates, for example, as follows: While moving along a road in Akihabara, which is located in Chiyoda Ward, Tokyo, Japan, a user captures spherical images at various locations along the path of movement using a camera capable of capturing spherical images (for example, THETA (registered trademark) manufactured by Ricoh Company, Ltd.) (moving images can be captured while moving).
[0057] The user also inputs the captured spherical image Ri (i=1, 2...) (corresponding to wide-angle image data of the present invention) to the information processing device 1 of this embodiment, and operates the information processing device 1 to acquire three-dimensional area model information (as already mentioned, PLATEAU LoD1 data provided by the Ministry of Land, Infrastructure, Transport and Tourism of Japan may be used) including information on three-dimensional models of buildings in the same area as the captured area, and then executes the following process.
[0058] 4, in a process started by a user instruction, the information processing device 1 estimates the positions xi and angles λ of the cameras that captured each of the input spherical images Ri, and obtains information on the estimated image capturing positions (S11). As already described, this process can be performed by employing a widely known method such as Visual SLAM.
[0059] The information processing device 1 places the three-dimensional models of the buildings included in the acquired three-dimensional area model information in a three-dimensional virtual space (S12). Furthermore, the information processing device 1 uses the information included in the three-dimensional area model information to determine the conversion relationship between the coordinates of the virtual space and the coordinates (information of latitude, longitude, and altitude) in the real global coordinate system.
[0060] Then, the information processing device 1 converts the camera position and orientation information x1, x2, ..., xN at each point where the wide-angle image data was captured, estimated in step S11, and the trajectory of camera movement represented by λ, into trajectory information in the coordinates of the virtual space set in step S12 using the conversion relationship determined above, and obtains the camera position and orientation information x'1, x'2, ..., x'N in the virtual space that corresponds to the position and orientation information of the camera that captured the wide-angle image data (setting of initial position and orientation information: S13).
[0061] The information processing device 1 refers to data of the three-dimensional model of the building arranged in the virtual space, sets the virtual camera using each piece of position and orientation information obtained in step S13, and renders a celestial sphere image R'i (i=1, 2, ...) of the virtual space (S14). As already described, it is not necessary to obtain a celestial sphere image R'i corresponding to all pieces of position and orientation information here.
[0062] Then, the information processing device 1 obtains a difference (a portion where pixel areas differ) between the pixel area (shape of the building) of the building drawn in the rendering image R′i obtained in step S14 and the pixel area (shape of the building) of the building captured in the corresponding celestial sphere image Ri, extracted from the celestial sphere image Ri by a semantic segmentation method or the like (S15).
[0063] The information processing device 1 obtains the number of pixels di of the difference, and calculates the sum Σdi of the difference for each piece of position and orientation information (S16).
[0064] Furthermore, the information processing device 1 updates at least a part of the estimated position and orientation information x'i, for example, x'1, x'N, and λ (S17), and returns to step S14 to repeat the process. Note that in step S17, if the sum Σdi of differences based on a plurality of mutually different sets of position and orientation information has been obtained in the past, at least a part of the position and orientation information x'i, for example, x'1, x'N, and λ, is updated using a method such as a gradient method based on the difference between the sums and the difference between the corresponding sets of position and orientation information so as to reduce Σdi.
[0065] When Σdi falls below a predetermined threshold value or the number of repetitions exceeds a predetermined threshold value, the information processing device 1 terminates the above-mentioned repetitive process and determines the position and orientation information x′i (i=1, 2, . . . N) and λ to be the last updated values (S18).
[0066] Following the above process, the information processing device 1 performs browsing processing illustrated in Fig. 5. In this processing, the information processing device 1 acquires the position and orientation information x'i (i = 1, 2, ..., N) and λ determined in the processing of step S18 in Fig. 4, and also acquires the acquired three-dimensional area model information and the omnidirectional image Ri (information acquisition: S21).
[0067] The information processing device 1 places a plurality of types of pre-specified three-dimensional objects A, B, etc. in a predetermined virtual space together with the three-dimensional models of buildings included in the three-dimensional area model information obtained in step S21 (S22). Initially, the positions of these three-dimensional objects are set in advance for each type.
[0068] In the following example, the information processing device 1 arranges a three-dimensional object A that resembles the shape of a human body (and to which a surface texture of a human body has been set) and a three-dimensional object B that represents a water mass that is substantially rectangular (with height Z=H, where H may initially be "0") (an opaque three-dimensional object that represents an area filled with water and has a water surface texture set).
[0069] In this case, the information processing device 1 arranges the three-dimensional model of the building so that the height position (position in the Z-axis direction) of the bottom side (i.e., the ground) of the three-dimensional model is Z = 0, and also arranges the three-dimensional objects A and B so that their bottom surfaces are in contact with Z = 0 (i.e., they exist on the ground surface). Furthermore, it is assumed that the three-dimensional object A is a movable object.
[0070] The information processing device 1 selects one of the acquired position and orientation information, xi,λ. For example, initially, the information processing device 1 selects a predetermined one of the position and orientation information, xi,λ (S23). Note that if a user operation is performed in a later process, one of the position and orientation information, xi,λ, may be selected based on the operation.
[0071] The information processing device 1 uses the virtual camera set using the selected position and orientation information to render a 360-degree spherical image (ERP image) of the virtual space set in step S22, and obtains the resulting rendered image R′i (S24).
[0072] Furthermore, the information processing device 1 sets the pixels of the image R′i of the rendering result, excluding the pixels where the three-dimensional object is drawn, to the pixel value of the corresponding pixel in the wide-angle image data Ri that corresponds to the position and orientation information xi,λ of the virtual camera used to obtain this rendering result (pixel value setting: S25).
[0073] Then, the information processing device 1 outputs the image in which the pixel values have been set (S26).
[0074] Next, the information processing device 1 accepts a user instruction operation (S27). The instruction operation may include at least an object movement operation of a three-dimensional object and a position specification operation. If the information processing device 1 accepts a position specification operation, the information processing device 1 returns to step S23 and repeats the process from selecting one piece of position and orientation information xi,λ based on the accepted position specification operation.
[0075] On the other hand, in step S27, when an object movement operation is performed, the information processing device 1 follows the instruction and checks whether or not the position of at least one of the three-dimensional objects (for example, three-dimensional object A in this case) is moved in the virtual space, overlapping with the area in the virtual space in which the three-dimensional model of the building is located (determining whether it overlaps with the building: S28), and if it is determined that it overlaps (S28: Yes), returns to step S27 and continues processing (i.e., does not move the three-dimensional object).
[0076] Furthermore, in step S28, if it is determined that the position of the three-dimensional object after movement does not overlap with the area in the virtual space where the three-dimensional model of the building is located (S28: No), the position of three-dimensional object A is set to the position after movement, and the process returns to step S22 to update the position of the three-dimensional object and repeat the process.
[0077] This process allows the user to move a human-shaped three-dimensional object A against the background of the projection surface of an actually captured spherical image, and since the movement of the three-dimensional object A is restricted where it comes into contact with a building, the three-dimensional object A does not overlap with the building, creating an unnatural situation. Furthermore, when the three-dimensional object A is hidden by a building, such as when it enters a road behind a building from the position of the virtual camera serving as the viewpoint, the three-dimensional object A is hidden in accordance with the shape of the building, providing realistic information.
[0078] Furthermore, in the example of the present embodiment, a substantially omnidirectional image is displayed as the background. Therefore, even if the three-dimensional model of a building does not include vegetation (such as trees planted on a street) or structures such as utility poles, these structures are displayed naturally, thereby improving the sense of realism.
[0079] Furthermore, in step S27, the information processing device 1 may set the height direction size H of the three-dimensional object B in response to an instruction from the user. This setting may be arbitrarily performed by the user, or may be performed based on the predicted water level in the event of a flood in the corresponding area.
[0080] In this example, the water level h in the global coordinates specified by the user is converted using the conversion relationship between the virtual space coordinate system and the global coordinate system to the height H of the corresponding three-dimensional object B in the virtual space, and a rectangular three-dimensional object B set to the converted height is placed.
[0081] In this way, a realistic situation when a flood occurs in a city can be simulated and displayed.
[0082] [Variations] In the example of this embodiment, when performing alignment processing by comparing the rendering result of a virtual space in which a three-dimensional model is placed with wide-angle image data, the image areas of buildings contained in both the rendering result and the wide-angle image data are compared, but this embodiment is not limited to this, and other methods may be adopted as long as the position and orientation information of the virtual camera can be set so as to match the rendering result with the wide-angle image data.
[0083] For example, the information processing device 1 may execute the above process by using common feature points (such as sides and corners of structures such as buildings and roads) from both the rendering result and the wide-angle image data.
[0084] Specifically, the information processing device 1 aligns feature points (which may be a point cloud obtained by SLAM processing) obtained from the wide-angle image data with corresponding points in the rendering result of the virtual space in which the three-dimensional model is arranged. Since the process of aligning such a point cloud with an image can be performed using a widely known method, a detailed description thereof will be omitted here.
[0085] Furthermore, in the description of the operation of this embodiment, the information processing device 1 that performs the alignment processing and the information processing device 1 that performs the browsing processing are the same, but this embodiment is not limited to this, and the information processing device 1 that performs the alignment processing and the information processing device 1 that performs the browsing processing may be separate devices. [Explanation of symbols]
[0086] 1 Information processing device, 11 Control unit, 12 Memory unit, 13 Operation unit, 14 Display unit, 15 Interface unit, 20 Alignment processing unit, 21 Model acquisition unit, 22 Image acquisition unit, 23 Estimated position acquisition unit, 24 Rendering processing unit, 25 Decision processing unit, 30 Browsing processing unit, 31 Position and orientation information acquisition unit, 32 Rendering processing unit.
Claims
1. a model acquisition means for acquiring three-dimensional area model information including information on three-dimensional models of buildings; an image acquisition means for acquiring wide-angle image data captured while moving around the area where the building is located; an estimated position acquisition means for acquiring information representing the trajectory of movement of a camera that captured the acquired wide-angle image data, and estimating an estimated capturing position of the wide-angle image data on the trajectory; a determination means for arranging a three-dimensional model of the building in a predetermined three-dimensional virtual space, and determining position and orientation information of a virtual camera in the virtual space corresponding to a trajectory of movement of the camera that captured the wide-angle image data based on a comparison between a rendering image of the virtual space from a viewpoint of a virtual camera set in the virtual space and the wide-angle image data; and an information processing device comprising: the wide-angle image data, coordinate information of the virtual space determined in accordance with the wide-angle image data, and the acquired three-dimensional area model information, and the information processing device is subjected to predetermined processing.
2. 2. The information processing device according to claim 1, The determination means is an information processing device that compares the shape of the building depicted in the rendering image with the shape of the building captured in the wide-angle image data, and determines the position and orientation information of the virtual camera in the virtual space corresponding to the trajectory of movement of the camera that captured the wide-angle image data.
3. 2. The information processing device according to claim 1, The predetermined processing includes processing for arranging a three-dimensional object in the virtual space that can be moved within the virtual space by user operation, setting a virtual camera that renders the three-dimensional object based on position and orientation information of the virtual camera in the virtual space that corresponds to the trajectory of movement of the camera that captured the wide-angle image data determined by the determination means, rendering the three-dimensional object, and setting and outputting at least a portion of the pixel values of the image resulting from the rendering based on the wide-angle image data that corresponds to the position and orientation of the virtual camera.
4. 4. The information processing device according to claim 3, An information processing device in which the movement range of the three-dimensional object within the virtual space is restricted by a three-dimensional model of the building.
5. 2. The information processing device according to claim 1, The predetermined processing further includes processing of placing at least one type of three-dimensional object in the virtual space, setting a position and angle of a virtual camera that renders the three-dimensional object based on an estimated imaging position of the wide-angle image data, rendering the virtual space including the three-dimensional object, and setting and outputting at least a portion of the pixel values of the image resulting from the rendering based on the wide-angle image data that corresponds to the position and orientation of the virtual camera.
6. 6. The information processing device according to claim 5, The information processing device, in the predetermined processing, further includes processing for arranging in the virtual space a three-dimensional object of a type that can be moved within the virtual space by a user's operation.
7. 7. The information processing device according to claim 6, An information processing device in which a range of movement within the virtual space of a three-dimensional object of a type that can be moved within the virtual space by an operation of the user is restricted by a three-dimensional model of the building.
8. an accepting means for accepting target information including three-dimensional area model information including information on a three-dimensional model of a building, wide-angle image data captured while moving through an area where the building is located, and position and orientation information of a virtual camera in a predetermined three-dimensional virtual space determined using information representing the trajectory of movement of the camera that captured the wide-angle image data, wherein the three-dimensional model of the building is placed in the virtual space, and the position and orientation information is determined so that when a rendering image of the virtual space from the viewpoint of the camera set based on the position and orientation information is superimposed on target wide-angle image data that is one of the wide-angle image data and displayed, the rendering result of the building based on the three-dimensional model in the rendering image matches the building captured in the target wide-angle image data; a rendering means for setting the virtual space, arranging a three-dimensional model of the building included in the target information in the virtual space, arranging a predetermined three-dimensional object, and outputting an image obtained by rendering the virtual space from a viewpoint of a virtual camera set based on the position and orientation information of the virtual camera included in the target information; An information processing device comprising:
9. Computer, a model acquisition means for acquiring three-dimensional area model information including information on three-dimensional models of buildings; an image acquisition means for acquiring wide-angle image data captured while moving around the area where the building is located; an estimated position acquisition means for acquiring information representing the trajectory of movement of a camera that captured the acquired wide-angle image data, and estimating an estimated capturing position of the wide-angle image data on the trajectory; a determination means for arranging a three-dimensional model of the building in a predetermined three-dimensional virtual space, and determining coordinate information in the virtual space corresponding to the trajectory of movement of a camera that captured the wide-angle image data based on a comparison between a rendering image of the virtual space from a viewpoint set in the virtual space and the wide-angle image data; It functions as A program for subjecting the wide-angle image data, the coordinate information of the virtual space determined in accordance with the wide-angle image data, and the acquired three-dimensional area model information to predetermined processing.