Imaging device, imaging method, and program

The imaging device and method generate virtual images from desired positions using distance and model information, addressing installation and positioning challenges, and enabling accurate image representation.

JP7868634B2Active Publication Date: 2026-06-02SONY GROUP CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2024-06-06
Publication Date
2026-06-02

Smart Images

  • Figure 0007868634000001
    Figure 0007868634000001
  • Figure 0007868634000002
    Figure 0007868634000002
  • Figure 0007868634000003
    Figure 0007868634000003
Patent Text Reader

Abstract

To make it possible to easily obtain an image captured from a desired position.SOLUTION: By using distance information from an imaging position to a subject and model information, a virtual image obtained by imaging the subject from a virtual imaging position different from the imaging position is generated from a captured image obtained by imaging the subject from the imaging position. The present technology can be applied to, for example, an imaging apparatus that images a subject.SELECTED DRAWING: Figure 25
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This technology relates to an imaging device, an imaging method, and a program, and more particularly to an imaging device, an imaging method, and a program that enable easy acquisition of images captured from a desired position. [Background technology]

[0002] For example, Patent Document 1 describes a technique for obtaining a virtual image captured from a virtual imaging position different from the actual imaging position. This technique involves using multiple imaging devices to image a subject from various imaging positions, and then generating high-precision 3D data from the resulting images. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2019-103126 [Overview of the project] [Problems that the invention aims to solve]

[0004] The technology described in Patent Document 1 requires numerous imaging devices to be placed in various locations. Therefore, due to the cost of the imaging devices and the effort required for installation, it is often not easily implemented.

[0005] Furthermore, when arranging multiple imaging devices, care must be taken to prevent one device from appearing in the image of another, and to ensure that the subject does not collide with the imaging device when the subject is moving. Therefore, it is not always possible to place imaging devices in any arbitrary position.

[0006] This technology was developed in light of these circumstances, and aims to make it possible to easily obtain images captured from a desired position. [Means for solving the problem]

[0007] The imaging device of this technology uses distance information from the imaging position to the subject and model information to capture a virtual image of the subject from a virtual imaging position different from the actual imaging position. This is a virtual subject after interpolation, in which the occluded portions of the virtual subject, as viewed from the virtual imaging position, have been interpolated. The imaging device comprises a generation unit that generates a corrected model and generates using the corrected model, and a UI (User Interface) for specifying the virtual imaging position, wherein the model information is a virtual subject of 3D data generated from an image obtained by auxiliary imaging, and the UI has a first operation unit that is operated when determining the center of a spherical coordinate system that represents the virtual imaging position, a second operation unit that is operated when changing the azimuth angle of the virtual imaging position in the spherical coordinate system, a third operation unit that is operated when changing the elevation angle of the virtual imaging position in the spherical coordinate system, and a fourth operation unit that is operated when changing the distance between the center of the spherical coordinate system and the virtual imaging position.

[0008] The imaging method or program of this technology uses distance information from the imaging position to the subject and model information to capture a virtual image of the subject from a virtual imaging position different from the actual imaging position. This is a virtual subject after interpolation, in which the occluded portions of the virtual subject, as viewed from the virtual imaging position, have been interpolated. An imaging method, or a program for causing a computer to function as such an imaging device, includes generating a corrected model, controlling the display of a UI (User Interface) that generates and specifies the virtual imaging position using the corrected model, the model information being a virtual subject of 3D data generated from an image obtained by auxiliary imaging, and the UI having a first operation unit operated when determining the center of a spherical coordinate system representing the virtual imaging position, a second operation unit operated when changing the azimuth angle of the virtual imaging position in the spherical coordinate system, a third operation unit operated when changing the elevation angle of the virtual imaging position in the spherical coordinate system, and a fourth operation unit operated when changing the distance between the center of the spherical coordinate system and the virtual imaging position.

[0009] In the imaging device, imaging method, and program of this technology, distance information from the imaging position to the subject and model information are used to capture a virtual image of the subject from a virtual imaging position different from the actual imaging position. This is a virtual subject after interpolation, in which the occluded portions of the virtual subject, as viewed from the virtual imaging position, have been interpolated. The system generates a corrected model, uses the corrected model to generate a user interface (UI) for specifying the virtual imaging position, and controls the display of the UI. The model information is a virtual subject of 3D data generated from an image obtained by auxiliary imaging. The UI includes a first operation unit operated when determining the center of a spherical coordinate system representing the virtual imaging position, a second operation unit operated when changing the azimuth angle of the virtual imaging position in the spherical coordinate system, a third operation unit operated when changing the elevation angle of the virtual imaging position in the spherical coordinate system, and a fourth operation unit operated when changing the distance between the center of the spherical coordinate system and the virtual imaging position.

[0010] The imaging device may be a standalone device or an internal block that makes up a single device.

[0011] Furthermore, the program can be provided by transmitting it via a transmission medium or by recording it on a recording medium. [Brief explanation of the drawing]

[0012] [Figure 1] This figure shows an example of the imaging setup. [Figure 2] This diagram shows examples of imaging conditions and the resulting images. [Figure 3] This figure shows other examples of imaging conditions and the images captured under those conditions. [Figure 4] This figure shows another example of the imaging setup. [Figure 5] This figure shows an image obtained by photographing a person and a building from above and in front of the person. [Figure 6] This is a top view illustrating an example of an imaging situation where it is not possible to capture images from a distance away from a person. [Figure 7] This diagram illustrates the perspective projection transformation that occurs when imaging is performed by an imaging device. [Figure 8] This figure shows an example of an imaging scenario in which a subject located on a single object surface is imaged. [Figure 9] This is a top view showing wide-angle imaging, where a wide-angle lens is used to capture an image of the subject from a close position. [Figure 10] This is a top view showing telephoto imaging, where a subject is captured from a distant imaging position using a telephoto lens. [Figure 11] This diagram illustrates an example of the process of obtaining a virtual image. [Figure 12] This figure shows an example of imaging conditions when the subject is located on multiple `dist` surfaces. [Figure 13] This is a top view showing wide-angle imaging, where a wide-angle lens is used to capture an image of the subject from a close position. [Figure 14] This is a top view showing telephoto imaging, where a subject is captured from a distant imaging position using a telephoto lens. [Figure 15] This figure shows the process of close-range imaging and the image obtained from that imaging. [Figure 16] This figure shows the process of long-distance imaging and the image obtained from that imaging. [Figure 17] This is a top view showing the imaging process. [Figure 18] This diagram illustrates the pixel value mapping when generating a virtual image obtained from a virtual long-distance image, based on an image obtained from an actual close-range image. [Figure 19] This is another diagram illustrating the pixel value mapping when generating a virtual image obtained from a long-distance image, which is a virtual imaging technique, based on an image obtained from an actual close-range image, which is a virtual imaging technique. [Figure 20] This diagram illustrates an example of a method for interpolating occlusion areas, which involves filling in the pixels in the occluded portion of an area. [Figure 21]This diagram illustrates another example of the process of obtaining a virtual image obtained through virtual imaging, based on information obtained from actual imaging. [Figure 22] This is a plan view showing examples of captured images and virtual images. [Figure 23] This diagram illustrates how to represent a virtual imaging position when performing virtual imaging. [Figure 24] This is a plan view showing an example of the UI used when a user specifies a virtual imaging location. [Figure 25] This block diagram shows an example configuration of one embodiment of an imaging device to which this technology is applied. [Figure 26] This is a flowchart illustrating an example of the processing in the generation section. [Figure 27] This is a block diagram showing an example configuration of one embodiment of a computer to which this technology is applied. [Modes for carrying out the invention]

[0013] <Relationship between imaging distance and captured image>

[0014] Figure 1 shows an example of the imaging situation using an imaging device.

[0015] Figure 1 shows the imaging conditions using the third-angle projection method.

[0016] In Figure 1, as viewed from the perspective of the imaging device, a person is standing in front of the building, and the imaging device is capturing images of both the person and the building from the front of the person.

[0017] The following describes the image actually captured by the imaging device (captured image) in the imaging situation shown in Figure 1.

[0018] Figure 2 shows an example of the imaging conditions by the imaging device and the resulting image.

[0019] Figure 2A is a top view showing the imaging setup, and Figure 2B shows the image captured under the imaging setup shown in Figure 2A.

[0020] In Figure 2A, the dashed line represents the field of view of the imaging device, and the space within the field of view is captured by the imaging device. In Figure 2A, the dashed line represents the field of view occupied by the main subject (the person).

[0021] In the imaging scenario shown in Figure 2A, the distance between the person and the building is relatively far compared to the distance between the imaging device and the person when imaging is performed (imaging distance). Therefore, although the width of the building behind the person is actually wider than the width of the person, in the captured image, the width of the building appears narrower than the width of the person. This is because distant objects appear smaller due to perspective.

[0022] Figure 3 shows another example of the imaging conditions by the imaging device and the captured image obtained under those conditions.

[0023] Figure 3A is a top view showing the imaging setup, and Figure 3B shows the image captured under the imaging setup shown in Figure 3A.

[0024] In Figure 3A, similar to Figure 2A, the dashed line represents the field of view of the imaging device, and the dashed line represents the field of view occupied by the person.

[0025] In Figure 3A, the same subjects—people and buildings—are captured from a position further away from the subject than in Figure 2A, using a telephoto lens with a narrow field of view (or a zoom lens with a longer focal length).

[0026] In the imaging scenario shown in Figure 3A, the captured image, as shown in Figure 3B, shows the building wider than the person, just as in reality. This is because, in the imaging scenario shown in Figure 3A, the imaging distance between the imaging device and the person is greater than in the case of Figure 2A, making the distance between the person and the building relatively smaller and reducing the sense of perspective.

[0027] As described above, even when imaging the same subject (people and buildings), the content (composition) of the resulting image will differ depending on the imaging distance between the subject and the imaging device.

[0028] The fact that the content of an image differs depending on the imaging distance is of great importance in visual expression. To give a simple example, if you want to obtain an image with a vast mountain range in the background, you need to use a wider-angle lens and get closer to the subject when taking the image. Conversely, if you want to obtain an image that minimizes the appearance of clutter in the background, you need to use a longer-telephoto lens and move further away from the subject when taking the image.

[0029] In principle, if imaging is performed from infinity, the size ratio of people and buildings in the captured image will be equal to the actual ratio. Therefore, in order to obtain images that accurately reflect the actual size ratio for architectural or academic purposes, it is necessary to perform imaging from a greater distance.

[0030] Figure 4 shows another example of the imaging situation using the imaging device.

[0031] In Figure 4, as in Figure 1, the imaging conditions are shown using the third-angle projection method.

[0032] In the imaging setup shown in Figure 1, the optical axis of the imaging device roughly coincides with the direction from the person to the building. In this case, the image captured by the imaging device will be an image that does not easily represent the sense of distance between the person and the building, as shown in Figure 2B and Figure 3B.

[0033] In Figure 4, the person and the building are imaged from above and in front of the person, with the optical axis of the imaging device pointed towards the person. In this case, the direction of the optical axis of the imaging device is different from the direction from the person to the building, and an image can be obtained that expresses the sense of distance between the person and the building.

[0034] Figure 5 shows an image obtained by photographing a person and a building from above and in front of the person.

[0035] By positioning the optical axis of the imaging device towards the person from above and in front of the person, and imaging both the person and the building, it is possible to obtain an overhead image that expresses the sense of distance between the person and the building, as if viewed from above.

[0036] To achieve visual expression that meets the intended purpose, it is necessary to capture images of the subject from various positions.

[0037] However, in reality, it is not always possible to take images from any position. For example, even if you want to take an image from a distance away from a person, as shown in Figure 3A, in reality, it may not be possible to take an image from a distance away from the person.

[0038] Figure 6 is a top view illustrating an example of an imaging situation where it is not possible to capture images from a distance away from a person.

[0039] In Figure 6, a wall is present in front of the person. Therefore, when imaging a person from the front, it is physically impossible to move the imaging device behind the wall in front of the person, making it impossible to image the person from a long distance.

[0040] Furthermore, as shown in Figure 4, when photographing a person and a building from above and in front of the person, it is possible to take images from a certain height by using a tripod or step ladder. However, when using a tripod or step ladder, the limit is only a few meters above the subject. Moreover, using a tripod or step ladder reduces mobility at the shooting site.

[0041] In recent years, drones have been used to capture images from almost directly above the subject, but the flight time of the drone, and consequently the image capture time, is limited by the capacity of the battery installed in the drone.

[0042] Furthermore, operating a drone is not always easy, and outdoors, it is affected by weather conditions such as rain and wind. In addition, drones cannot be used in areas where drone flights are restricted or prohibited due to the high density of people.

[0043] This technology makes it possible to easily obtain images of a subject from a desired position, even when it is not possible to freely choose an imaging position. For example, this technology can generate images of a subject from imaging positions such as those shown in Figure 3A or Figure 5, from the image shown in Figure 2B, which was taken under imaging conditions shown in Figure 2A.

[0044] Furthermore, Patent Document 1 describes a technique for generating a virtual image (as if it were taken from an arbitrary virtual virtual imaging position) from 3D data generated using 3D data produced from 3D data generated from images obtained from a subject using a large number of imaging devices.

[0045] However, the technology described in Patent Document 1 requires numerous imaging devices to be placed in various locations, and due to the cost of the imaging devices and the effort required for installation, it is often not easy to realize the imaging conditions described in Patent Document 1.

[0046] Furthermore, when arranging multiple imaging devices, it becomes necessary to prevent one device from appearing in the image of another, and to ensure that the subject does not collide with the imaging device when the subject is moving. Therefore, it is not always possible to place imaging devices in any arbitrary position.

[0047] This technology uses distance information from the imaging position to the subject and corresponding model information to generate a virtual image of the subject from an image captured from the imaging position, by creating a virtual image of the subject from a virtual imaging position different from the actual imaging position. As a result, this technology makes it possible to easily obtain a virtual image of the subject captured from a desired virtual imaging position without having to install a large number of imaging devices.

[0048] The following describes a method for generating a virtual image of a subject from an image captured from a certain imaging position. This method involves generating a virtual image of a subject from an image captured from a certain imaging position, for example, from an image captured from a close imaging position using a wide-angle lens (or a zoom lens with a shortened focal length), as shown in Figure 2A, to an image captured from a distant imaging position using a telephoto lens, as shown in Figure 3A.

[0049] <Perspective projection transformation>

[0050] Figure 7 illustrates the perspective projection transformation that occurs when imaging is performed by the imaging device.

[0051] Figure 7 shows the relationship between the actual object on the object surface where the object exists and the image on the imaging surface of the image sensor that performs photoelectric conversion in the imaging device.

[0052] Figure 7 is a top view of an object standing vertically on the ground, as seen from above. The horizontal direction represents the horizontal position relative to the ground. The following explanation also applies to the vertical direction, which is perpendicular to the ground, as shown in a side view of an object standing vertically on the ground.

[0053] The distance from the object surface to the lens of the imaging device (the imaging distance between the subject on the object surface and the imaging device) is called the object distance, L obj It is expressed as follows: The distance from the lens to the image sensor is called the image distance, L img This is expressed as follows: The position on the object surface, that is, the distance from the optical axis of the imaging device on the object surface, is X obj This is expressed as follows: The position on the imaging plane, that is, the distance from the optical axis of the imaging device on the imaging plane, is X img It is represented as follows.

[0054] Object distance L obj , image distance L img , distance (position) obj、and the distance (position) X img satisfies Equation (1).

[0055] X img / X obj =L img / L obj ···(1)

[0056] From Equation (1), the position X obj on the imaging surface corresponding to the position of the subject on the object surface img can be expressed by Equation (2).

[0057] X img =L img / L obj ×X obj ···(2)

[0058] Equation (2) represents a transformation called so-called perspective projection transformation.

[0059] The perspective projection transformation of Equation (2) is physically (optically) performed when actually imaging the subject with the imaging device.

[0060] Also, from Equation (1), the position X img on the object surface corresponding to the position on the imaging surface obj can be expressed by Equation (3).

[0061] X obj =L obj / L img ×X img ···(3)

[0062] Equation (3) represents the inverse transformation of the perspective projection transformation (inverse perspective projection transformation).

[0063] To perform the inverse perspective projection transformation of Equation (3), the object distance L obj , the image distance L img , and the position X img of the subject on the imaging surface are required.

[0064] In an imaging device that images a subject, the image distance L img and the position X of the subject on the imaging plane img It can recognize (acquire) it.

[0065] Therefore, to perform the inverse perspective projection transformation in equation (3), the object distance L obj (Distance information) needs to be recognized in some way.

[0066] Position X of the subject on the object surface relative to each pixel of the imaging surface obj To obtain this, the object distance L has a resolution of pixel units or close to it. obj This will be necessary.

[0067] Object distance L obj Any method can be used to obtain the distance to the subject. For example, the so-called stereo method can be used, which calculates the distance to the subject from the parallax obtained using multiple image sensors that perform photoelectric conversion. Alternatively, a method can be used in which a predetermined optical pattern is projected onto the subject, and the distance to the subject is calculated from the shape of the optical pattern projected onto the subject. Another method is called ToF (Time of Flight), which calculates the distance to the subject from the time it takes for the reflected light to return from the subject after the laser light is projected onto it. Furthermore, a method can be used to calculate the distance to the subject using the image plane phase difference method, which is one of the so-called autofocus methods. In addition, a combination of several of the above methods can be used to calculate the distance to the subject.

[0068] Below, the object distance L obj Assuming that it can be recognized in some way, the distance L from the subject to the actual object is calculated using perspective projection transformation and inverse perspective projection transformation. obj This section describes a method for generating a virtual image of a subject that would be captured if it were captured from a virtual imaging position located at a different distance from the actual subject.

[0069] <How to generate virtual images>

[0070] Figure 8 shows an example of an imaging scenario in which a subject located on a single object surface is imaged.

[0071] In Figure 8, as in Figure 1, the imaging conditions are shown using the third-angle projection method.

[0072] In the imaging setup shown in Figure 8, the subject lies on a single object plane, and this object plane is parallel to the imaging plane of the imaging device. Therefore, the object plane is perpendicular to the optical axis of the imaging device.

[0073] Figure 9 is a top view showing wide-angle imaging, where a wide-angle lens is used to capture the subject from a position close to the subject, under the imaging conditions shown in Figure 8.

[0074] In the wide-angle image in Figure 9, the position on the object plane is X obj The subject is at an object distance L obj_W The image is taken from an imaging position that is a distance away. The image distance during wide-angle imaging is L img_W The position of the subject on the imaging plane is X img_W It has become that way.

[0075] Figure 10 is a top view showing telephoto imaging, where a telephoto lens is used to capture an image of the subject from an imaging position far from the subject, in the imaging situation shown in Figure 8.

[0076] In the telephoto image in Figure 10, the position on the object plane is X obj Then, the same subject as in the wide-angle imaging case in Figure 9, at object distance L obj_T The image is taken from an imaging position that is a distance away. The image distance during telephoto imaging is L img_T The position of the subject on the imaging plane is X img_T It has become that way.

[0077] Applying equation (3) for the inverse transform of perspective projection to the wide-angle image in Figure 9 yields equation (4) for the inverse transform of perspective projection.

[0078] X obj =L obj_W / Limg_W ×X img_W ...(4)

[0079] Applying the perspective projection transformation equation (2) to the telephoto image in Figure 10 yields the perspective projection transformation equation (5).

[0080] X img_T =L img_T / L obj_T ×X obj ...(5)

[0081] X on the left side of equation (4) obj The X on the right side of equation (5) obj By substituting these values, we can obtain equation (6).

[0082] X img_T =(L img_T / L obj_T )×(L obj_W / L img_W )×X img_W ...(6)

[0083] Here, the coefficient k is defined by equation (7).

[0084] k=(L img_T / L obj_T )×(L obj_W / L img_W ) ...(7)

[0085] Using equation (7), equation (6) can be reduced to the simple proportion equation (8).

[0086] X img_T =k × X img_W ...(8)

[0087] By using equations (8) and (6), we can obtain wide-angle imaging using a wide-angle lens, in this case, the position X on the imaging plane in close-range imaging from a close distance. img_WTherefore, telephoto imaging using a telephoto lens, here, the position X on the imaging plane in long-distance imaging from a long distance. img_T This allows us to obtain information such as the image data obtained from actual imaging at close range, and then, assuming that imaging was performed at long range, we can obtain information about the virtual image data that would be obtained from that long-range imaging.

[0088] The above explanation described imaging from different distances from the subject, using close-up imaging with a wide-angle lens and long-distance imaging with a telephoto lens as examples. However, the above explanation can be applied when imaging from any distance from the subject using a lens of any focal length.

[0089] In other words, according to equations (8) and (6), based on information such as the captured image obtained by imaging from a certain imaging position using a lens of a certain focal length, it is possible to obtain information about the captured image (virtual image) obtained when imaging from a different imaging position (virtual imaging position) using a lens of a different focal length.

[0090] Here, imaging using a lens with a certain focal length from a certain imaging position is an actual imaging operation, and is therefore also called actual imaging. On the other hand, imaging using a lens with a different focal length from a different imaging position (a virtual imaging position) is not an actual imaging operation, and is therefore also called virtual imaging.

[0091] Figure 11 illustrates an example of the process of obtaining a virtual image obtained through virtual imaging, based on information obtained from actual imaging, using equation (8).

[0092] The conceptual meaning of obtaining equation (6) from equations (4) and (5) is as follows:

[0093] Position X of the subject on the imaging plane img_W X is the position of a point on a 3D object in three-dimensional space projected onto the imaging plane of an image sensor, which is a 2D plane. img_WBy performing the inverse perspective transformation of equation (4), the position X of a point on the object in 3D space (object plane) can be obtained. obj You can obtain this.

[0094] The position X on the subject in the 3D space obtained in this way obj For this, by performing the perspective projection transformation of equation (5), the object distance L from the subject is obtained. obj_W A virtual imaging position that is different from the actual imaging position, i.e., the object distance L from the subject. obj_T It is possible to obtain information about a virtual image obtained when imaging is performed from a virtual imaging position that is only a short distance away.

[0095] Equation (6) shows the apparent position (of a variable X) of a point on an object in three-dimensional space. obj The position X of the subject on the imaging plane during wide-angle imaging is erased and is a two-dimensional plane. img_W Therefore, the position X of the subject on the imaging plane during telephoto imaging as another two-dimensional plane. img_T This is a conversion to . However, in the process of deriving equation (6) from equations (4) and (5), the position X on the subject in 3D space is used. obj However, it has already been decided.

[0096] The process of obtaining a virtual image through virtual imaging, based on information obtained from actual imaging, consists of actual imaging, generation of a virtual subject (model), and virtual imaging, as shown in Figure 11.

[0097] In actual imaging, an object in physical space (3D space) is transformed into a perspective projection onto an image sensor by an optical system (physical lens optical system) such as a physical lens in the imaging device, generating a 2D image (actual image). The perspective projection transformation in actual imaging is performed optically, using the physical imaging position of the imaging device (physical imaging position) as a parameter.

[0098] In the generation of virtual subjects, the inverse perspective projection transform of equation (4) is calculated using distance information from the imaging position to the subject, obtained separately through measurements, etc., on the image captured by actual imaging, and a subject model of the subject is virtually reproduced (generated) in three-dimensional space. This virtually reproduced subject is also called a virtual subject (model).

[0099] In virtual imaging, the virtual subject is (virtually) imaged by performing a perspective projection transformation according to equation (5) through calculation, and a virtual image (virtual imaged image) is generated. In virtual imaging, the virtual imaging position when the virtual subject is imaged is specified as a parameter, and the virtual subject is imaged from that virtual imaging position.

[0100] <Position of the subject on the imaging plane when the subject is located on multiple object surfaces>

[0101] Figure 12 shows an example of imaging conditions when the subject is located on multiple object surfaces.

[0102] In Figure 12, the imaging conditions are shown using the third-angle projection method, similar to Figure 1.

[0103] Figure 8 assumes that the subject has a single object surface, but in actual imaging, the subject often exists on multiple object surfaces. Figure 12 shows the imaging situation when the subject exists on multiple object surfaces.

[0104] In the imaging setup shown in Figure 12, from the perspective of the imaging device, a second subject exists behind the first subject, which corresponds to the subject in Figure 8.

[0105] For the first subject, equations (6) and (8) are used to determine the position X of the subject on the imaging plane during actual imaging, for example, in close-range imaging. img_W This represents, for example, the position X of the subject on the imaging plane in long-distance imaging, as a virtual image. img_T It can be converted to [this].

[0106] The same transformation can be performed on the second subject as well.

[0107] Figure 13 is a top view showing wide-angle imaging, where a wide-angle lens is used to capture the subject from a position close to the subject, under the imaging conditions shown in Figure 12.

[0108] Figure 14 is a top view showing telephoto imaging, where a telephoto lens is used to capture an image of the subject from a position far from the subject, in the imaging situation shown in Figure 12.

[0109] Figures 13 and 14 are diagrams that add an object plane and an imaging plane to the second subject, compared to Figures 9 and 10.

[0110] In Figures 13 and 14, the first object plane is the object plane of the first subject, and the second object plane is the object plane of the second subject. Since the second subject is imaged simultaneously with the first subject as the background of the first subject, the imaging planes are the same for the first subject and the second subject in both Figures 13 and 14.

[0111] In the wide-angle image in Figure 13, the position on the second object plane is X obj2 The second subject is at object distance L obj_W2 The image is taken from an imaging position that is a distance away. The image distance during wide-angle imaging is L img_W The position of the second subject on the imaging plane is X img_W2 This is the case. Since the first and second subjects are imaged simultaneously, the image distance during wide-angle imaging is the same as in Figure 9. img_W This is the result. Furthermore, if we let d be the distance between the first object plane and the second object plane, then d is given by the equation d = L obj_W2 -L obj_W It is represented as follows.

[0112] In the telephoto image in Figure 14, the position on the second object plane is X obj2 The second subject is at object distance L obj_T2 The image is taken from an imaging position that is a distance away. The image distance during telephoto imaging is L img_T The position of the second subject on the imaging plane is Ximg_T2 It is. Since the first subject and the second subject are imaged simultaneously, the image distance during telephoto imaging is the same L as in the case of FIG. 10 img_T is obtained. If the distance between the first object plane and the second object plane is d, then d is given by the formula d = L obj_T2 - L obj_T is represented by.

[0113] When the formula (3) of perspective projection inverse transformation is applied to the wide - angle imaging of FIG. 13, the formula (9) of perspective projection inverse transformation can be obtained.

[0114] X obj2 = L obj_W2 / L img_W × X img_W2 ···(9)

[0115] When the formula (2) of perspective projection transformation is applied to the telephoto imaging of FIG. 14, the formula (10) of perspective projection transformation can be obtained.

[0116] X img_T2 = L img_T / L obj_T2 × X obj2 ···(10)

[0117] By substituting X obj2 on the left - hand side of formula (9) into X obj2 on the right - hand side of formula (10), formula (11) can be obtained.

[0118] X img_T2 =(L img_T / L obj_T2 )×(L obj_W2 / L img_W )× X img_W2 ···(11)

[0119] Here, the coefficient k2 is defined by formula (12).

[0120] k2=(L img_T / L obj_T2 )×(L obj_W2 / L img_W ) ...(12)

[0121] Using equation (12), equation (11) can be reduced to the simple proportion equation (13).

[0122] X img_T2 =k²×X img_W2 ...(13)

[0123] By using equations (13) and (11), we can obtain the position X on the imaging plane in wide-angle imaging using a wide-angle lens, in this case, close-range imaging from a close distance. img_W2 Therefore, telephoto imaging using a telephoto lens, here, the position X on the imaging plane in long-distance imaging from a long distance. img_T2 You can obtain this.

[0124] Therefore, by applying equation (8) to the pixels of the image obtained by actual imaging, such as close-range imaging, that capture the first subject on the first object plane, and by applying equation (13) to the pixels of the image obtained by the second subject on the second object plane, it is possible to map the pixels of the image obtained by close-range imaging to the pixels of a virtual image obtained by virtual imaging, such as long-range imaging.

[0125] <Occlusion>

[0126] Figure 15 shows the close-range imaging process under the imaging conditions shown in Figure 12, and the image obtained from that close-range imaging.

[0127] Specifically, Figure 15A is a top view showing close-range imaging of the subject from a close position using a wide-angle lens, in the imaging situation shown in Figure 12. Figure 15B is a plan view showing the image obtained by the close-range imaging in Figure 15A, and is equivalent to a front view of the imaging surface viewed from the front.

[0128] Figure 15A is a diagram that, compared to Figure 13, has dotted lines added as auxiliary lines passing through the center of the lens from the endpoints of the first and second subjects, respectively.

[0129] Figure 16 shows the long-distance imaging process under the imaging conditions shown in Figure 12, and the image obtained from that long-distance imaging.

[0130] In other words, Figure 16A is a top view showing long-distance imaging, where a telephoto lens is used to image the subject from a position far from the subject, in the imaging situation shown in Figure 12. Figure 16B is a plan view showing the image obtained by the long-distance imaging in Figure 16A, and, as in the case of Figure 15, is equivalent to a front view of the imaging surface seen from the front.

[0131] Figure 16A is a diagram that, compared to Figure 14, has dotted lines added as auxiliary lines passing through the center of the lens from the endpoints of the first and second subjects, respectively.

[0132] For the sake of simplicity, let's assume that in both near-range and far-range imaging, the size of the first subject on the imaging surface (captured image) is the same.

[0133] The size of the second subject on the imaging plane (captured image) is larger in the far-distance imaging shown in Figure 16 than in the near-distance imaging shown in Figure 15. This phenomenon, where the size of the second subject on the imaging plane is larger in far-distance imaging than in near-distance imaging, is due to perspective, as explained in Figures 2 and 3.

[0134] Figure 17 is a top view showing the imaging process, created by superimposing the top view of A in Figure 15 and the top view of A in Figure 16, with some parts omitted.

[0135] In Figure 17, part M of the second subject is captured in long-distance imaging, but is not captured in close-distance imaging because it is obscured by the first subject.

[0136] When a subject occupies multiple surfaces, occlusion can occur, where a first subject in the foreground obscures a second subject in the background, making it invisible.

[0137] The portion M of the second subject is visible in long-distance imaging, but becomes occluded and invisible in close-up imaging because it is hidden by the first subject. This occluded portion of the second subject M is also called the occluded area (missing part).

[0138] In actual close-range imaging, the occluded portion M of the second subject is not captured. Therefore, when generating a virtual image obtained from a long-range image using equations (8) and (13) based on the image obtained from close-range imaging, the pixel values ​​for the occluded portion M of the second subject are not obtained in the virtual image, and thus it is missing.

[0139] Figure 18 illustrates the mapping of pixel values ​​when generating a virtual image obtained from a long-distance imaging technique, based on an image obtained from an actual close-range imaging technique.

[0140] In the image obtained by close-range imaging as an actual imaging technique shown in the upper part of Figure 18 (close-range image), part M of the second subject is in shadow of the first subject and is an occluded area.

[0141] When generating a virtual image obtained by long-distance imaging as a virtual imaging method (Figure 18 lower) based on the captured image (Figure 18 upper), the position X of the first subject in the captured image (close-range captured image) img_W and position X of the second subject img_W2 The pixel value of the pixel is, using equations (8) and (13), the position X of the first subject in the virtual image (far-distance captured image). img_T and position X of the second subject img_T2 These are mapped as the pixel values ​​of each pixel.

[0142] In the lower virtual image of Figure 18, the pixel values ​​of the pixels that capture part M of the second subject should be mapped to the shaded area. However, in the upper captured image of Figure 18, part M of the second subject is not captured, and therefore the pixel values ​​of part M of the second subject cannot be obtained. Consequently, in the lower virtual image of Figure 18, the pixel values ​​of part M of the second subject cannot be mapped to the shaded area, resulting in missing pixel values.

[0143] As described above, when a subject exists on multiple object surfaces, occlusion occurs in the occluded portion, such as part M of the second subject, resulting in a loss of pixel values.

[0144] Figure 19 is another diagram illustrating the mapping of pixel values ​​when generating a virtual image obtained from a long-distance imaging technique, based on an image obtained from an actual close-range imaging technique.

[0145] In Figure 19, image picW is an image obtained by actual close-range imaging, while image picT is a virtual image obtained by virtual long-range imaging.

[0146] Furthermore, in Figure 19, the horizontal axis of the 2D coordinate system represents the horizontal position X of the captured image picW. img_W and X img_W2 This represents the horizontal position X of the virtual image picT. img_T and X img_T2 It represents.

[0147] Furthermore, in Figure 19, line L1 represents equation (8), and line L2 represents equation (13).

[0148] Position X of the first subject in the captured image picW img_W The pixel (pixel value) is at its position X img_W The position X of the first subject in the virtual image picT, obtained by equation (8) using the input X, is given by the input X. img_T It is mapped to the pixels (pixel values) of the element.

[0149] Position X of the second subject in the captured image picW img_W2 The pixels are at their position X img_W2 The position X of the second subject in the virtual image picT, obtained by equation (13) using the input X, is determined by the input X. img_T2 It is mapped to the pixels.

[0150] In the virtual image picT, the areas marked with diagonal lines are occlusion areas where the corresponding parts are not captured in the image picW, and therefore pixels (pixel values) are missing.

[0151] <Completion of occlusion areas>

[0152] Figure 20 illustrates an example of an occlusion interpolation method that replaces pixels in the occlusion area.

[0153] Various methods can be used to interpolate occlusion areas.

[0154] One method for interpolating occlusion areas is to use neighboring pixels to interpolate the pixels (or their pixel values) in the occlusion area. Any method can be used to interpolate these pixels, such as the nearest neighbor method, bilinear interpolation, or bicubic interpolation.

[0155] In the nearest neighbor method, the pixel values ​​of neighboring pixels are used directly as the pixel values ​​of the occluded pixels. In the bilinear method, the average of the pixel values ​​of the surrounding pixels of the occluded pixels is used as the pixel value of the occluded pixels. In the bicubic method, the interpolated value obtained by performing 3D interpolation using the pixel values ​​of the surrounding pixels of the occluded pixels is used as the pixel value of the occluded pixels.

[0156] If the occlusion area is, for example, a monotonous wall image, interpolation using pixels in the vicinity of the occlusion area can be used to interpolate the occlusion area, thereby generating a virtual image that is (almost) identical to the image obtained when the image is captured from the virtual imaging position where the virtual image is captured. A virtual image that is similar to the image obtained when the image is captured from the virtual imaging position is also called a highly reproducible virtual image.

[0157] Furthermore, as a method of interpolating pixels in an occluded area using pixels in its vicinity, for example, if the occluded area is an image with a texture such as a rough wall surface, a method can be employed to interpolate the occluded area by duplicating a certain area surrounding the occluded area.

[0158] The method of interpolating pixels in an occluded area using pixels in its vicinity assumes that the occluded area will have an image similar to that of its neighbors.

[0159] Therefore, if the occlusion area does not resemble the surrounding area (i.e., if the occlusion area is unique compared to its surroundings), the method of interpolating the pixels of the occlusion area using pixels from its surroundings may not yield a highly reproducible virtual image.

[0160] For example, if a wall has graffiti and the graffiti area is an occluded area, a method of interpolating the pixels of the occluded area using neighboring pixels will not be able to reproduce the graffiti, and a highly accurate virtual image cannot be obtained.

[0161] If the occlusion area does not resemble the surrounding area, in order to obtain a highly reproducible virtual image, in addition to the primary image (original image), auxiliary imaging can be performed from a different imaging position than the primary image to capture the occlusion area that occurs in the primary image. Then, the image obtained from this auxiliary imaging can be used to complement the occlusion area that occurs in the primary image.

[0162] Figure 20 is a top view illustrating the primary and auxiliary imaging performed as actual imaging of the first and second subjects.

[0163] In Figure 20, the actual imaging with position p201 as the imaging location is performed as the primary imaging, while the actual imaging with positions p202 and p203, which are shifted to the left and right of position p201, is performed as auxiliary imaging.

[0164] In this case, the primary imaging from imaging position p201 cannot capture the portion M of the second subject that is the occluded area. However, the auxiliary imaging from imaging positions p202 and p203 can capture the portion M of the second subject that is the occluded area in the primary imaging.

[0165] Therefore, by generating a virtual image obtained from a virtual imaging process based on the primary image obtained from imaging position p201, and then interpolating the pixel values ​​of the second subject (which constitutes the occlusion portion) in that virtual image using images obtained from auxiliary imaging positions p202 and p203, a highly reproducible virtual image can be obtained.

[0166] Primary and auxiliary imaging can be performed simultaneously or at different times using multiple imaging devices.

[0167] Furthermore, both primary and auxiliary imaging can be performed using a single imaging device, such as a multi-camera system with multiple imaging systems.

[0168] Furthermore, primary and auxiliary imaging can be performed at different times using a single imaging device with a single imaging system. For example, for stationary subjects, auxiliary imaging can be performed before or after primary imaging.

[0169] Occlusion interpolation can be performed using only some information, such as color and texture, from the image obtained through auxiliary imaging. Furthermore, occlusion interpolation can also be performed in combination with other methods.

[0170] As described above, occlusion can be compensated for using images obtained from auxiliary imaging, as well as images obtained from other primary imaging, such as images obtained from primary imaging performed in the past.

[0171] For example, if the second subject, which serves as the background for the first subject, is a well-known building such as Tokyo Tower, then images of such a well-known building taken from various positions in the past may be stored in image libraries such as stock photo services.

[0172] If an image captured by actual imaging contains a well-known (or famous) building, and the portion of that building is an occluded area, the occluded area can be filled in using previously captured images of the same well-known building stored in the image library. In addition, the occluded area can be filled in using images published on the internet or other networks, such as photographs published on websites that provide map search services.

[0173] Occlusion can be interpolated using images, or using data (information) other than images.

[0174] For example, if the second subject, which serves as the background for the first subject, is a building, and architectural data related to that building, such as its shape, surface finish, and paint color, is publicly available on a web server or the like, then the occlusion can be interpolated by estimating the pixel values ​​of the occlusion area using such architectural data.

[0175] When an occlusion occurs in an image containing a building, and this occlusion is to be filled in using previously captured images stored in an image library, or using architectural data, it is necessary to identify the building, or in this case, the second subject. The second subject can be identified by image recognition targeting the captured image containing the second subject, or by identifying the location where the image was actually captured. The location where the image was actually captured can be identified by referring to metadata of the captured image, such as EXIF ​​(Exchangeable image file format) information.

[0176] Actual imaging is performed under conditions where the subject is illuminated, for example, by a light source such as sunlight.

[0177] On the other hand, when interpolating occlusion areas using past images (images taken in the past) or architectural data, the actual light source (illumination) at the time of imaging is not reflected in the occlusion areas.

[0178] Therefore, for example, if occlusion areas in images captured under actual sunlight are compensated for using past images or architectural data, the color of the occlusion area (or former occlusion area) may appear unnatural compared to the color of other areas.

[0179] Therefore, when interpolating occlusion areas in images captured using actual imaging under sunlight, using past images or architectural data, if meteorological data related to weather conditions is available, the interpolation of occlusion areas can be performed using past images or architectural data, and then the color tone correction of the occlusion areas can be performed using the meteorological data.

[0180] In actual imaging conducted under sunlight, the intensity and color temperature of the light illuminating the subject are affected by the weather. If meteorological data is available, it is possible to identify the weather conditions at the time of actual imaging from that data, and then estimate the illumination light information, such as the intensity and color temperature of the light illuminating the subject during actual imaging conducted under sunlight, from that weather condition.

[0181] Furthermore, the occlusion areas can be interpolated using past captured images and architectural data, and the color tone of the occlusion areas can be corrected so that the color of the occlusion areas matches the color of the subject when it is illuminated by the light represented by the illumination information.

[0182] By performing the color correction described above, the color of the occlusion area can be made to look more natural compared to the color of the other areas, thereby obtaining a highly reproducible virtual image.

[0183] In addition, occlusion can be compensated for, for example, using a machine learning model.

[0184] For example, if it is possible to actually perform both near-field and far-field imaging, the images obtained from actual near-field and far-field imaging can be used as training data to train a model. For instance, the model can be trained to take the image obtained from near-field imaging (which is performed as actual imaging) as input and output an image of the occlusion portion of a virtual image obtained from far-field imaging (which is performed as virtual imaging).

[0185] In this case, by inputting images obtained from actual close-range imaging into the trained model, it is possible to obtain images of the occlusion portion of virtual images obtained from virtual long-range imaging, and then use these images to interpolate the occlusion portion.

[0186] The interpolation method for resolving occlusion is not particularly limited. However, by employing an interpolation method that can be performed with one imaging device, or a small number of imaging devices, it is possible to suppress a decrease in mobility at the imaging site and easily obtain an image (virtual image) captured from a desired position (virtual imaging position). In particular, by employing an interpolation method that can be performed with one imaging device, mobility at the imaging site can be maximized.

[0187] Figure 21 illustrates another example of the process of obtaining a virtual image obtained through virtual imaging, based on information obtained from actual imaging.

[0188] In Figure 21, the process of obtaining a virtual image through virtual imaging based on information obtained from actual imaging is the same as in Figure 11, consisting of actual imaging, generation of a virtual subject (model), and virtual imaging. However, in Figure 21, occlusion interpolation is added compared to the case in Figure 11.

[0189] In actual imaging, a two-dimensional image (actual image) is generated (captured), similar to the case in Figure 11.

[0190] In the generation of virtual subjects, distance information from the actual imaging position to the subject and corresponding model information are used to reproduce (generate) a corrected virtual subject from the image obtained during actual imaging.

[0191] Countermeasure model information refers to knowledge information for dealing with occlusion, and includes, for example, one or more of the following: previously captured images (past captured images), images obtained through auxiliary imaging (auxiliary captured images), building data, and meteorological data.

[0192] In generating a virtual subject, as in the case of Figure 11, the virtual subject is first generated by performing an inverse perspective projection transform of the image obtained from actual imaging using distance information.

[0193] Furthermore, in the generation of a virtual subject, a virtual imaging position is given as a parameter, and in the subsequent virtual imaging, the portion of the virtual subject to be imaged from the virtual imaging position is identified.

[0194] Then, within the imaged portion of the virtual subject, the missing pixels (or their pixel values) in the captured image—that is, the occlusion areas that appear as occlusion when the virtual subject is viewed from the virtual imaging position—are interpolated using the corresponding model information. This interpolated virtual subject is then generated as a corrected model of the virtual subject.

[0195] In virtual imaging, a virtual image is generated by perspective projection transformation, similar to the case in Figure 11.

[0196] However, in the virtual imaging shown in Figure 21, the object of the perspective projection transformation is not the virtual subject itself, which is generated by performing the inverse perspective projection transformation of the image obtained from the actual imaging. Instead, it is a corrected model of the virtual subject in which the occlusion portion of the virtual subject has been interpolated. This differs from the case in Figure 11.

[0197] In the virtual imaging shown in Figure 21, a perspective projection transformation is performed on the corrected model, so that the corrected model is (virtually) imaged from the virtual imaging position, and a virtual image (virtual captured image) is generated.

[0198] Regarding the interpolation of occlusion areas, by performing interpolation only on the occlusion areas that are occluded when viewed from the virtual imaging position to the virtual subject, the scope of interpolation can be kept to the absolute minimum. This allows auxiliary imaging to be performed in addition to the primary imaging, with the auxiliary imaging limited to the minimum necessary area, thereby suppressing a decrease in mobility during imaging.

[0199] Furthermore, in auxiliary imaging, capturing a slightly wider area than the minimum necessary allows for minor adjustments to the virtual imaging position after the image is acquired.

[0200] <Other methods for generating virtual images>

[0201] Figure 22 is a plan view showing examples of captured images and virtual images.

[0202] In the above case, using near-range and far-range imaging as examples, a generation method was described in which a virtual image would be captured if imaging were performed in the other position, based on an image actually captured in one of the near-range or far-range imaging positions. Specifically, a virtual imaging position was defined as a position that differs only by the imaging distance from the actual imaging position, by moving along the optical axis of the imaging device at the time of imaging. A generation method was described in which a virtual image would be captured from this virtual imaging position, based on the actual image.

[0203] The virtual image generation method described above can also be applied when the virtual imaging position is a position moved in a direction not aligned with the optical axis of the imaging device from the imaging position of the actual image. In other words, the virtual image generation described above can be applied not only when generating a virtual image of a subject taken from a position moved along the optical axis of the imaging device from the imaging position of the actual image, but also when generating a virtual image (another virtual image) of a subject taken from a position moved in a direction not aligned with the optical axis of the imaging device.

[0204] When the optical axis of the imaging device is pointed towards the subject and the subject is imaged, if the virtual imaging position is defined as the position moved along the optical axis of the imaging device at the time of actual imaging from the imaging position of the captured image, then the optical axis of the imaging device will coincide in the actual image and the virtual image.

[0205] On the other hand, if the virtual imaging position is defined as a position moved from the actual imaging position in a direction not aligned with the optical axis of the imaging device during actual imaging, then the optical axis of the imaging device will differ between the actual imaging and the virtual imaging.

[0206] The case where a virtual imaging position is defined as a position moved in a direction not aligned with the optical axis of the imaging device from the actual imaging position of the captured image is, for example, when actual imaging is performed under the imaging conditions shown in Figure 1, and virtual imaging is performed under the imaging conditions shown in Figure 4.

[0207] Figure 22A shows the imaging setup as in Figure 1, and displays the image obtained from the actual imaging.

[0208] Figure 22B shows the imaging setup in Figure 4, and displays an overhead view image obtained from the actual imaging process.

[0209] Figure 22C shows a virtual image obtained (generated) by performing a virtual imaging on a virtual subject, which is generated using distance information based on the captured image in Figure 22A, with the imaging position in Figure 4 as the virtual imaging position.

[0210] In the virtual image obtained by performing virtual imaging on the virtual subject itself, which is generated using distance information based on the captured image in Figure 22A, the parts of the person and the upper part of the building that are not visible in the captured image in the imaging situation in Figure 1 become occlusion areas, as shown by the diagonal lines in Figure 22C.

[0211] By performing interpolation on the occlusion areas, a virtual image similar to the captured image in Figure 22B can be obtained.

[0212] That is, in the imaging part of the virtual subject, the occlusion part that is occluded when viewing the virtual subject from the virtual imaging position is complemented using the countermeasure model information, and by performing perspective projection transformation on the corrected model that is the virtual subject after the complementation, a virtual image close to the captured image of B in FIG. 22 can be obtained.

[0213] As a method for complementing the occlusion part, methods such as interpolating the occlusion part using pixels in the vicinity of the occlusion part described above, using a captured image obtained by auxiliary imaging, using a captured image captured in the past, using a learning model learned by machine learning, using building data, etc. can be adopted.

[0214] <Virtual UI for imaging>

[0215] FIG. 23 is a diagram for explaining a method of expressing the virtual imaging position in the case of performing virtual imaging.

[0216] In FIG. 23, the imaging situation is shown by the third angle method.

[0217] For the imaging device, by physically (actually) installing the imaging device, the imaging position of actual imaging is determined. In this technology, in addition to the imaging position of actual imaging, a virtual imaging position is required, and it is necessary to specify the virtual imaging position.

[0218] As a method for specifying the virtual imaging position, for example, a method of automatically specifying, as the virtual imaging position, a position that has moved a predetermined distance in a predetermined direction with respect to the imaging position can be adopted. Also, as a method for specifying the virtual imaging position, for example, a method of having the user specify it can be adopted.

[0219] Hereinafter, the UI when the user specifies the virtual imaging position will be described. Before that, the method of expressing the virtual imaging position will be described.

[0220] In this embodiment, the virtual imaging position is expressed in a spherical coordinate system (coordinates) with the position of the subject (intersecting the optical axis of the imaging device) as the center (origin), as shown in FIG. 23.

[0221] Here, the intersection of the optical axis of the imaging device (physically existing imaging device), that is, the optical axis of the optical system of the imaging device, and the subject is defined as the center of the subject. The optical axis of the imaging device is assumed to pass through the center of the imaging element of the imaging device and coincide with a straight line perpendicular to the imaging element.

[0222] The optical axis connecting the center of the imaging element of the imaging device and the center of the subject (the optical axis of the imaging device) is referred to as the physical optical axis, and the optical axis connecting the virtual imaging position and the center of the subject is referred to as the virtual optical axis.

[0223] When the center of the subject is the center of the spherical coordinate system, in the spherical coordinate system, the virtual imaging position is the rotation amount (azimuth angle) φ in the azimuth direction with respect to the physical optical axis of the virtual optical axis v the rotation amount (elevation angle) θ in the elevation direction v and the distance r between the subject on the virtual optical axis and the virtual imaging position. v It can be expressed by.

[0224] In FIG. 23, the distance r r represents the distance between the subject on the physical optical axis and the imaging position.

[0225] Also, in FIG. 23, the distance r shown in the top view v represents the distance along the virtual optical axis, not the distance component on the plane.

[0226] FIG. 24 is a plan view showing an example of a UI that is operated when the user designates a virtual imaging position.

[0227] In FIG. 24, the UI has operation buttons such as a C button, a TOP button, a BTM button, a LEFT button, a RIGHT button, a SHORT button, a LONG button, a TELE button, and a WIDE button.

[0228] The UI can be configured using control buttons, rotary dials, joysticks, touch panels, and other control elements. Furthermore, when configuring the UI using control buttons, the button arrangement is not limited to that shown in Figure 24.

[0229] An imaging device applying this technology can generate a virtual image in real time that is similar to an image obtained by capturing a subject from a virtual imaging position, and output it in real time to a display unit such as a viewfinder. In this case, the display unit can display the virtual image in real time as a so-called through image. By viewing the virtual image displayed as a through image on the display unit, the user of the imaging device can enjoy the feeling of capturing a subject from a virtual imaging position.

[0230] In a spherical coordinate system, it is necessary to determine the center of the spherical coordinate system in order to represent the virtual imaging position.

[0231] In the imaging device, for example, when the UI's C button is pressed, the position of the point where the optical axis of the imaging device intersects with the subject is determined to be the center of the spherical coordinate system. Then, the virtual imaging position is the actual imaging position, i.e., the azimuth angle φ. v =0, elevation angle θ v = 0, distance r v =r r It will be set to this.

[0232] Azimuth angle φ v This can be specified by operating the LEFT or RIGHT button. When the LEFT button is pressed on the imaging device, the azimuth angle φ is adjusted by a predetermined amount. v The azimuth angle changes in the negative direction. Also, when the RIGHT button is pressed, the azimuth angle φ changes by a predetermined amount. v It changes in the positive direction.

[0233] Elevation angle θ vThis can be specified by operating the TOP button or BTM button. When the TOP button is pressed, the elevation angle θ is raised by a predetermined amount. v The angle changes in the positive direction. Also, when the BTM button is pressed, the elevation angle θ changes by a predetermined amount. v It changes in the negative direction.

[0234] distance r v This can be specified by operating the SHORT or LONG button. When the SHORT button is pressed, the distance r is set by a predetermined fixed amount or by a fixed magnification. v The value changes in the negative direction. Also, when the LONG button is pressed, the distance r changes by a predetermined fixed amount or by a fixed magnification. v It changes in the positive direction.

[0235] In Figure 24, the UI includes, as described above, the C button, TOP button, BTM button, LEFT button, RIGHT button, SHORT button, and LONG button, which are related to specifying the virtual imaging position, as well as the TELE button and WIDE button, which are used to specify the focal length of the virtual imaging device (hereinafter also referred to as the virtual focal length) when performing virtual imaging from the virtual imaging position.

[0236] When the TELE button is pressed, the virtual focal length changes in the direction of increasing by a predetermined amount or by a certain magnification. Conversely, when the WIDE button is pressed, the virtual focal length changes in the direction of decreasing by a predetermined amount or by a certain magnification.

[0237] Depending on the virtual focal length, for example, the image distance L in equations (4) and (5) img_W and L img_T This will be decided.

[0238] Note that the azimuth angle φ is relative to the operation of the UI control buttons. v The methods for causing such changes are not limited to those described above. For example, if an operation button is pressed and held, the azimuth angle φ will change as long as the press and hold continues.v continuously changing the virtual imaging position and virtual focal length, etc., or changing the azimuth angle φ according to the time during which the long press continues v By making changes such as increasing the amount of change in the virtual imaging position and virtual focal length, etc., convenience can be improved.

[0239] Also, the method of specifying the virtual imaging position is not limited to the method of operating the UI. For example, a method of detecting the user's line of sight, detecting the fixation point on which the user is fixating from the detection result of the line of sight, and specifying the virtual imaging position can be adopted. In this case, the position of the fixation point is specified (set) as the virtual imaging position.

[0240] Furthermore, in the imaging device, when displaying a virtual image obtained by virtual imaging from the virtual imaging position specified by operating the UI, etc., in real time, the occlusion portion can be displayed so that the user can recognize it.

[0241] Here, in the virtual image in which the occlusion portion is complemented and generated, the accuracy of the information of the complemented portion where the occlusion portion is complemented may be inferior to the image obtained by actually imaging from the virtual imaging position.

[0242] Therefore, in the imaging device, after the virtual imaging position is specified, the virtual image can be displayed on the display unit so that the user can recognize the occlusion portion that becomes an occlusion in the virtual image obtained by virtual imaging from that virtual imaging position.

[0243] In this case, the user of the imaging device can recognize which part of the subject will be occluded by looking at the virtual image displayed on the display unit. By recognizing which part of the subject will be occluded, the user can then consider the actual imaging position so that important parts of the subject do not become occluded. In other words, the user can consider the imaging position so that important parts of the subject are captured in the image obtained from the actual imaging.

[0244] Furthermore, by allowing the user of the imaging device to recognize which parts of the subject will be occluded, the auxiliary imaging can be performed in a way that captures the occluded parts of the subject, as explained in Figure 20, when performing primary and auxiliary imaging.

[0245] To enable the user to recognize the occlusion areas in the virtual image, the display method for showing the virtual image on the display unit can include, for example, displaying the occlusion areas in the virtual image in a specific color, or inverting the gradation of the occlusion areas at a predetermined interval such as 1 second.

[0246] <One embodiment of an imaging device to which this technology is applied>

[0247] Figure 25 is a block diagram showing an example configuration of one embodiment of an imaging device such as a digital camera to which this technology is applied.

[0248] In Figure 25, the imaging device 100 includes an imaging optical system 2, an image sensor 3, a distance sensor 5, an inverse conversion unit 7, a correction unit 9, a conversion unit 11, a display unit 13, a UI 15, a storage unit 17, recording units 21 to 23, and an output unit 24. The imaging device 100 can be used for capturing both video and still images.

[0249] The imaging optical system 2 focuses light from the subject onto the image sensor 3 to form an image. As a result, the subject in three-dimensional space is transformed into a perspective projection onto the image sensor 3.

[0250] The image sensor 3 receives light from the imaging optical system 2 and performs photoelectric conversion to generate an image 4, which is a two-dimensional image having pixel values ​​corresponding to the amount of light received, and supplies it to the inverse conversion unit 7.

[0251] The distance sensor 5 measures and outputs distance information 6 to each point of the subject. The distance information 6 output by the distance sensor 5 is supplied to the inverse conversion unit 7.

[0252] Furthermore, the distance information 6 of the subject can be measured by an external device and supplied to the inverse conversion unit 7. In this case, the imaging device 100 can be configured without providing the distance sensor 5.

[0253] The inverse transformation unit 7 uses distance information 6 from the distance sensor 5 to perform an inverse perspective projection transformation of the image 4 captured from the image sensor 3, generating and outputting a virtual subject as 3D data 8.

[0254] The correction unit 9 interpolates the occlusion portion of the virtual subject as 3D data 8 output by the inverse transform unit 7, and outputs the interpolated virtual subject as the corrected model 10.

[0255] The conversion unit 11 performs a perspective projection transformation on the corrected model 10 output by the correction unit 9, and outputs a virtual image 12, which is a two-dimensional image obtained as a result.

[0256] The display unit 13 displays the virtual image 12 output by the conversion unit 11. When the conversion unit 11 outputs the virtual image 12 in real time, the display unit 13 can display the virtual image 12 in real time.

[0257] The UI15 is configured as shown in Figure 24, for example, and is operated by the imager, for example, the user of the imaging device 100. The user can perform operations on the UI15 to specify the virtual imaging position 16 while viewing the virtual image displayed on the display unit 13.

[0258] UI15 sets and outputs the virtual imaging position 16 in response to user operations.

[0259] The correction unit 9 performs interpolation of the occlusion portion that occurs when viewing the virtual subject from the virtual imaging position 16 output by the UI 15.

[0260] In other words, the correction unit 9 identifies the occlusion areas that occur when the virtual subject is viewed from the virtual imaging position 16. Subsequently, the correction unit 9 interpolates the occlusion areas, and the interpolated virtual subject is output as the corrected model 10.

[0261] In the conversion unit 11, a virtual image 12, which is a two-dimensional image obtained by capturing the corrected model 10 output by the correction unit 9 from the virtual imaging position 16 output by the UI 15, is generated by perspective projection transformation of the corrected model 10.

[0262] Therefore, the display unit 13 displays in real time the virtual image 12 obtained by capturing the corrected model 10 from the virtual imaging position 16 set according to the user's operation of the UI 15. This allows the user to specify the virtual imaging position 16 from which the desired virtual image 12 can be obtained by operating the UI 15 while viewing the virtual image 12 displayed on the display unit 13.

[0263] Furthermore, in the correction unit 9, occlusion can be interpolated by using pixels in the vicinity of the occlusion. In addition, occlusion can be interpolated using externally obtained countermeasure model information, such as past captured images 18, building data 19, or weather data 20.

[0264] Furthermore, occlusion can be compensated for using other countermeasure model information, such as images obtained through auxiliary imaging or trained models that have been trained using machine learning.

[0265] When auxiliary imaging is performed, and the auxiliary imaging is performed prior to the main imaging, the virtual subject as 3D data 8 generated from the captured image 4 obtained by the auxiliary imaging is stored in the storage unit 17 by the inverse transformation unit 7.

[0266] In other words, the memory unit 17 stores a virtual subject as 3D data 8 generated from the captured image 4 obtained by auxiliary imaging in the inverse transformation unit 7.

[0267] The virtual subject as 3D data 8 generated from the captured image 4 obtained by auxiliary imaging, which is stored in the memory unit 17, can be used in the correction unit 9 to compensate for occlusion portions of the virtual subject as 3D data 8 generated from the captured image 4 obtained by the main imaging performed after the auxiliary imaging, when viewed from the virtual imaging position 16.

[0268] When auxiliary imaging is performed, and the auxiliary imaging is performed after the main imaging, the virtual subject as 3D data 8 generated from the captured image 4 obtained by the main imaging is recorded in the recording unit 23 by the inverse conversion unit 7.

[0269] In other words, the recording unit 23 stores a virtual subject as 3D data 8 generated from the captured image 4 obtained by the main imaging in the inverse conversion unit 7.

[0270] The correction unit 9 can perform the interpolation of occlusion portions of the virtual subject as 3D data 8 generated from the captured image 4 obtained by the main imaging, which is recorded in the recording unit 23, using the virtual subject as 3D data 8 generated from the captured image 4 obtained by auxiliary imaging performed after the main imaging.

[0271] Therefore, when auxiliary imaging is performed, and the auxiliary imaging is performed after the primary imaging, the occlusion portion of the virtual subject as 3D data 8 generated from the image obtained by the primary imaging is interpolated after the auxiliary imaging has been performed.

[0272] When such occlusion areas are interpolated, it becomes difficult to generate a virtual image 12 in real time from the captured image 4 obtained by the primary imaging. Therefore, if real-time generation of a virtual image is required, the auxiliary imaging must be performed prior to the primary imaging, not after.

[0273] In Figure 25, the recording unit 21 records the virtual image 12 output by the conversion unit 11. The virtual image 12 recorded in the recording unit 21 can be output to the display unit 13 and the output unit 24.

[0274] The recording unit 22 records the corrected model 10 output by the correction unit 9.

[0275] For example, the correction unit 9 can interpolate a wider area than the occlusion portion, including the occlusion portion when the virtual subject is viewed from the virtual imaging position 16 from the UI 15 (the area that includes the portion of the virtual subject that becomes a new occlusion portion when the virtual imaging position 16 changes slightly and the virtual subject is viewed from the changed virtual imaging position 16). The recording unit 22 can record the corrected model 10 in which such a wide area interpolation has been performed.

[0276] In this case, the corrected model 10, which has been interpolated over a wide area and recorded in the recording unit 22, can be used as the target of the perspective projection transformation in the transformation unit 11 to generate a virtual image 12 with the virtual imaging position 16 slightly modified (fine-tuned). Therefore, after capturing the image 4 that served as the basis for the corrected model 10 with interpolated over a wide area, a virtual image 12 can be generated with the virtual imaging position 16 slightly modified using the corrected model 10 with interpolated over a wide area recorded in the recording unit 22.

[0277] The recording unit 23 records the virtual subject as 3D data 8 output by the inverse transformation unit 7, that is, the virtual subject before the correction unit 9 interpolates the occlusion portion. The virtual subject recorded in the recording unit 23 can be referenced as unprocessed, so-called true data to verify its authenticity, for example, if a virtual image 12 is used in news or the like, and the authenticity of a part of that virtual image 12 is questioned.

[0278] Furthermore, the recording unit 23 can record the captured image 4 that served as the basis for generating the virtual subject as 3D data 8, either along with the virtual subject as 3D data 8, or in place of the virtual subject as 3D data 8.

[0279] The output unit 24 is an interface that outputs data to the outside of the imaging device 100, and outputs the virtual image 12 output by the conversion unit 11 to the outside in real time.

[0280] When an external device (not shown) is connected to the output unit 24, the conversion unit 11 can output the virtual image 12 in real time, and the output unit 24 can then distribute the virtual image 12 to the external device in real time.

[0281] For example, when an external display unit (not shown) is connected to the output unit 24, and the conversion unit 11 outputs a virtual image 12 in real time, the virtual image 12 is output from the output unit 24 to the external display unit in real time, and the virtual image 12 is displayed on the external display unit in real time.

[0282] In the imaging device 100 configured as described above, the inverse transformation unit 7 generates a virtual subject as 3D data 8 by performing an inverse perspective projection transformation of the captured image 4 from the image sensor 3 using distance information 6 from the distance sensor 5.

[0283] The correction unit 9 uses the corresponding model information such as past captured images 18 to interpolate the occlusion portions that appear when viewing the virtual subject as 3D data 8 generated by the inverse transform unit 7 from the virtual imaging position 16 from the UI 15, thereby obtaining a corrected model 10 of the interpolated virtual subject.

[0284] The conversion unit 11 uses the corrected model 10 obtained by the correction unit 9 to generate a virtual image 12 by perspective projection conversion, which is an image of the corrected model 10 taken from a virtual imaging position 16.

[0285] Therefore, the inverse conversion unit 7, the correction unit 9, and the conversion unit 11 constitute a generation unit that uses distance information 6 from the imaging position to the subject and corresponding model information to generate a virtual image 12 from an image 4 taken of the subject from the imaging position, which is obtained by imaging the subject from a virtual imaging position 16 different from the actual imaging position.

[0286] Figure 26 is a flowchart illustrating an example of the processing in the generation unit.

[0287] In step S1, the generation unit uses distance information 6 and countermeasure model information (knowledge information) to deal with occlusion in past captured images 18, etc., to generate a virtual image 12 from the captured image 4, which is captured from a virtual imaging position 16 different from the captured image 4.

[0288] Specifically, in step S11, the inverse transformation unit 7 of the generation unit generates a virtual subject as 3D data 8 by performing an inverse perspective projection transformation of the captured image 4 using the distance information 6, and the process proceeds to step S12.

[0289] In step S12, the correction unit 9 uses the corresponding model information, such as past captured images 18, to interpolate the occlusion portions that appear when viewing the virtual subject as 3D data 8 generated by the inverse transformation unit 7 from the virtual imaging position 16. This corrects the virtual subject as 3D data 8, generating a corrected model 10 (3D data 8 with the occlusion portions interpolated), and the process proceeds to step S13.

[0290] In step S13, the conversion unit 11 uses the corrected model 10 generated by the correction unit 9 to generate a virtual image of the corrected model 10 captured from the virtual imaging position 16 by perspective projection conversion.

[0291] According to the imaging device 100, even in situations where it is difficult to image a subject from a desired imaging position (viewpoint), a virtual image can be generated by using an image captured from a certain imaging position (viewpoint) that can be imaged, distance information from that imaging position to the subject, and handling model information as auxiliary information other than distance information obtained separately. This allows the virtual image to be generated from a virtual imaging position, which is a desired imaging position different from the actual imaging position. Therefore, an image (virtual image) captured from a desired position can be easily obtained.

[0292] According to the imaging device 100, for example, as shown in Figure 6, in an imaging situation where a wall is present in front of a person, it is possible to generate a virtual image that appears as if it were taken from a position behind the wall in front of the person.

[0293] Furthermore, the imaging device 100 can generate virtual images that appear as if the user had approached the subject and taken the image, even in imaging situations where the user cannot get close to the subject, such as when the user is indoors or in a vehicle and taking an image of the outside through a window.

[0294] Furthermore, the imaging device 100 can generate virtual images such as the overhead view shown in Figure 5 without using, for example, a stepladder or a drone.

[0295] Furthermore, according to the imaging device 100, for example, if the subject is a person and the person's gaze is not directed towards the imaging device, the position at the end of the gaze can be set as the virtual imaging position, thereby generating a so-called camera-facing virtual image.

[0296] Furthermore, the imaging device 100 can generate a virtual image that captures the view from the user's perspective by setting the virtual imaging position to the position of the user's eyeballs on their head. By displaying this virtual image on a glasses-type display, a parallax-free electronic pair of glasses can be constructed.

[0297] <Description of a computer using this technology>

[0298] Next, the series of processes performed by the inverse conversion unit 7, correction unit 9, and conversion unit 11 that constitute the generation unit described above can be carried out by hardware or by software. When the series of processes are carried out by software, the program that constitutes the software is installed on a general-purpose computer or the like.

[0299] Figure 27 is a block diagram showing an example configuration of one embodiment of a computer on which the program that performs the series of processes described above is installed.

[0300] The program can be pre-recorded on the hard disk 905 or ROM 903, which are recording media built into the computer.

[0301] Alternatively, the program can be stored (recorded) on a removable recording medium 911 driven by the drive 909. Such a removable recording medium 911 can be provided as so-called packaged software. Examples of removable recording media 911 include flexible disks, CD-ROMs (Compact Disc Read Only Memory), MO (Magneto Optical) disks, DVDs (Digital Versatile Discs), magnetic disks, semiconductor memory, etc.

[0302] In addition to installing the program from the removable storage medium 911 as described above, the program can also be downloaded to the computer via a communication network or broadcasting network and installed on the built-in hard disk 905. That is, the program can be transferred wirelessly to the computer from a download site via a satellite for digital satellite broadcasting, or transferred via a wired connection to the computer via a network such as a LAN (Local Area Network) or the Internet.

[0303] The computer has a built-in CPU (Central Processing Unit) 902, and an input / output interface 910 is connected to the CPU 902 via a bus 901.

[0304] When the CPU 902 receives a command from the user via the input / output interface 910, such as by operating the input unit 907, it executes a program stored in the ROM (Read Only Memory) 903 accordingly. Alternatively, the CPU 902 loads a program stored in the hard disk 905 into the RAM (Random Access Memory) 904 and executes it.

[0305] As a result, the CPU 902 performs processing according to the flowchart described above, or processing according to the configuration of the block diagram described above. The CPU 902 then outputs the processing results as needed, for example, via the input / output interface 910 from the output unit 906, or transmits them from the communication unit 908, or records them on the hard disk 905.

[0306] The input section 907 consists of a keyboard, mouse, microphone, etc. The output section 906 consists of an LCD (Liquid Crystal Display), speakers, etc.

[0307] In this specification, the processes performed by a computer according to a program do not necessarily have to be performed chronologically in the order described in the flowchart. That is, the processes performed by a computer according to a program include processes that are executed in parallel or individually (e.g., parallel processing or object-based processing).

[0308] Furthermore, the program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Moreover, the program may be transferred to a remote computer for execution.

[0309] Furthermore, in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device in which multiple modules are housed in one enclosure, are both considered systems.

[0310] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.

[0311] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.

[0312] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.

[0313] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.

[0314] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0315] Furthermore, this technology can take the following configuration.

[0316] <1> The system includes a generation unit that uses distance information from the imaging position to the subject and model information to generate a virtual image of the subject from an image captured from the imaging position, by capturing the subject from a virtual imaging position different from the imaging position. Imaging device. <2> The generation unit generates a corrected model from the captured image using the distance information and the model information, and generates the virtual image using the corrected model. <1> The imaging device described above. <3> The aforementioned model information is knowledge information for dealing with occlusion. <1> or <2> The imaging device described above. <4> The generating unit is By performing an inverse perspective projection transform on the captured image using the aforementioned distance information, a virtual subject is generated. Using the aforementioned model information, a corrected model is generated in which the virtual subject is corrected by interpolating the occlusion portion that is occluded when the virtual subject is viewed from the virtual imaging position. Using the corrected model, a virtual image is generated by perspective projection transformation, capturing the corrected model from the virtual imaging position. <3> The imaging device described above. <5> The system further comprises a recording unit for recording the virtual subject or the corrected model. <4> The imaging device described above. <6> The aforementioned model information includes one or more of the previously captured images, building data related to buildings, and weather data related to weather. <3> or <5> An imaging device as described in any of the following. <7> The system further includes a UI (User Interface) for specifying the virtual imaging position. <1> The imaging device described above. <8> The virtual image is output to the display unit in real time. <1> or <7> An imaging device as described in any of the following. <9> The aforementioned UI is A first operating unit that is operated when determining the center of the spherical coordinate system representing the virtual imaging position, A second operating unit that is operated when changing the azimuth angle of the virtual imaging position in the spherical coordinate system, A third operating unit that is operated when changing the elevation angle of the virtual imaging position in the spherical coordinate system, A fourth operating unit that is operated when changing the distance between the center of the spherical coordinate system and the virtual imaging position. has <7> The imaging device described above. <10> The UI further includes a fifth operating section which is operated when changing the focal length of a virtual imaging device when performing virtual imaging from the virtual imaging position. <9> The imaging device described above. <11> The UI continuously changes the virtual imaging position or the focal length while the operation of any of the first to fifth operation units is being continued. <10> The imaging device described above. <12> The UI changes the amount of change in the virtual imaging position or the focal length according to the duration of the operation of any of the first to fifth operation units. <10> The imaging device described above. <13> The UI specifies the point of focus that the user is looking at as the virtual imaging position. <1> or <12> An imaging device as described in any of the following. <14> Using distance information from the imaging position to the subject and model information, a virtual image is generated from an image of the subject taken from the imaging position, representing the subject taken from a virtual imaging position different from the actual imaging position. An imaging method that includes the following. <15> A generation unit generates a virtual image of the subject taken from a virtual imaging position different from the actual imaging position, using distance information from the imaging position to the subject and model information. A program that makes a computer function. [Explanation of symbols]

[0317] 2 Imaging optical system, 3 Image sensor, 5 Distance sensor, 7 Inverse transformer, 9 Correction unit, 11 Transformer, 13 Display unit, 15 UI, 17 Memory unit, 21 to 23 Recording unit, 24 Output unit, 901 Bus, 902 CPU, 903 ROM, 904 RAM, 905 Hard disk, 906 Output unit, 907 Input unit, 908 Communication unit, 909 Drive, 910 Input / Output interface, 911 Removable recording medium

Claims

1. A generation unit generates a corrected model of a virtual subject by using distance information from the imaging position to the subject and model information to capture a virtual image of the subject from a virtual imaging position different from the imaging position, and by interpolating the occlusion portions that are occluded when the virtual subject is viewed from the virtual imaging position, and then generating using the corrected model. The UI (User Interface) for specifying the virtual imaging position and Equipped with, The aforementioned model information is a virtual subject of 3D data generated from captured images obtained through auxiliary imaging. The aforementioned UI is A first operating unit that is operated when determining the center of the spherical coordinate system representing the virtual imaging position, A second operating unit that is operated when changing the azimuth angle of the virtual imaging position in the spherical coordinate system, A third operating unit that is operated when changing the elevation angle of the virtual imaging position in the spherical coordinate system, A fourth operating unit that is operated when changing the distance between the center of the spherical coordinate system and the virtual imaging position. has Imaging device.

2. The UI further includes a fifth operating section which is operated when changing the focal length of a virtual imaging device when performing virtual imaging from the virtual imaging position. The imaging apparatus according to claim 1.

3. The UI continuously changes the virtual imaging position or the focal length while any of the first to fifth operating units is being operated. The imaging apparatus according to claim 2.

4. The UI changes the amount of change in the virtual imaging position or the focal length according to the duration of time that operation of any of the first to fifth operation units is being continued. The imaging apparatus according to claim 2.

5. Using distance information from the imaging position to the subject and model information, a corrected model of the virtual subject is generated by interpolating the occlusion portions that appear as the virtual subject is viewed from the virtual imaging position, based on a virtual image of the subject taken from a virtual imaging position different from the said imaging position, and then generating using the corrected model. Control the display of the UI (User Interface) for specifying the virtual imaging position, The aforementioned model information is a virtual subject of 3D data generated from captured images obtained through auxiliary imaging. The aforementioned UI is A first operating unit that is operated when determining the center of the spherical coordinate system representing the virtual imaging position, A second operating unit that is operated when changing the azimuth angle of the virtual imaging position in the spherical coordinate system, A third operating unit that is operated when changing the elevation angle of the virtual imaging position in the spherical coordinate system, A fourth operating unit that is operated when changing the distance between the center of the spherical coordinate system and the virtual imaging position. Control the display to have An imaging method that includes the following.

6. Using distance information from the imaging position to the subject and model information, a corrected model of the virtual subject is generated by interpolating the occlusion portions that appear as the virtual subject is viewed from the virtual imaging position, based on a virtual image of the subject taken from a virtual imaging position different from the said imaging position, and then generating using the corrected model. Control the display of the UI (User Interface) for specifying the virtual imaging position, The aforementioned model information is a virtual subject of 3D data generated from captured images obtained through auxiliary imaging. The aforementioned UI A first operating unit that is operated when determining the center of the spherical coordinate system representing the virtual imaging position, A second operating unit that is operated when changing the azimuth angle of the virtual imaging position in the spherical coordinate system, A third operating unit that is operated when changing the elevation angle of the virtual imaging position in the spherical coordinate system, A fourth operating unit that is operated when changing the distance between the center of the spherical coordinate system and the virtual imaging position. A program to control the display so that it has [a certain feature].