Input device, input method of input device, output device and output method of output device
The input device simplifies the correlation of inspection results with spatial location data by displaying spherical images and associating diagnostic information, enhancing report creation efficiency and accuracy.
Patent Information
- Application Number
- JP2025099668
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-15
AI Technical Summary
The manual correlation of location information on a 3D spatial model with inspection results during on-site inspections is cumbersome and time-consuming, particularly when compiling reports on defects like cracks.
An input device that displays a spherical image of a structure and receives inputs of diagnostic information along with position information, storing them in association for easier report creation.
Facilitates easier aggregation of information with position data, reducing human errors and report creation burden by intuitively linking annotations to inspection targets.
Smart Images

Figure 2025120451000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an input device and an input method for an input device, as well as an output device and an output method for an output device. [Background technology]
[0002] In recent years, on-site inspections of buildings have become increasingly important. In particular, many things can only be understood on-site, such as inspections for repairs, construction work planning, equipment renewal, and design. There is a known technology for capturing images of the inspection target during on-site inspections using an imaging device and storing the captured image data, thereby reducing the labor required for inspection work and enabling information sharing.
[0003] Furthermore, Patent Document 1 discloses a technology in which location information of a predetermined location on a three-dimensional space model is included in search conditions, and this information is configured to be extractable, thereby enabling the sharing of surrounding conditions based on the location information. Summary of the Invention [Problem to be solved by the invention]
[0004] When inspecting cracks and other defects, the inspection results must be compiled and submitted in a report. In this case, inspectors can search for associated information based on location information indicating the location on the 3D spatial model. However, in the past, when compiling the inspection results in a report, the task of correlating the location information on the 3D spatial model with the defect was performed manually, which made the report creation process cumbersome and time-consuming.
[0005] The present invention has been made in view of the above, and has an object to make it easier to aggregate information associated with position information on an image. [Means for solving the problem]
[0006] In order to solve the above-described problems and achieve the object, the present invention provides an input device for inputting a diagnosis result of a diagnosis target in a structure, the input device including: a display means for displaying a spherical image of the structure; and a receiving means for receiving an input of a position indicating the diagnosis target relative to the spherical image and storing position information indicating the received position in a storage means, wherein the display means displays the spherical image and a diagnostic information input screen for receiving an input of diagnostic information including a diagnosis result of the diagnosis target on a single screen, and the receiving means receives an input of diagnostic information according to the diagnostic information input screen and stores the diagnostic information and the position information in association with each other in the storage means. [Effects of the Invention]
[0007] According to the present invention, it is possible to more easily aggregate information associated with position information on an image. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram schematically illustrating the appearance of an example of an imaging device applicable to the first embodiment. [Figure 2] FIG. 2 is a diagram showing the structure of an example of an imaging body applicable to the first embodiment. [Figure 3] FIG. 3 is a three-sided view schematically showing the appearance of the imaging device according to the first embodiment. [Figure 4] FIG. 4 is a block diagram showing an example of the configuration of the imaging device according to the first embodiment. [Figure 5] FIG. 5 is a diagram showing an example of the arrangement of the battery and the circuit unit in the imaging device according to the first embodiment. [Figure 6] FIG. 6 is a block diagram showing an example of the configuration of an information processing device as an input device for inputting annotations, which is applicable to the first embodiment. [Figure 7] FIG. 7 is a functional block diagram of an example illustrating functions of the information processing device according to the first embodiment. [Figure 8]FIG. 8 is a diagram for explaining how the imaging lens applicable to the first embodiment projects three-dimensional incident light into two dimensions. [Figure 9] FIG. 9 is a diagram for explaining the tilt of the imaging device according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating a format of a spherical image applicable to the first embodiment. [Figure 11] FIG. 11 is a diagram for explaining correspondence between each pixel position of a hemispherical image and each pixel position of a celestial sphere image according to a conversion table, which is applicable to the first embodiment. [Figure 12] FIG. 12 is a diagram for explaining vertical correction applicable to the first embodiment. [Figure 13] FIG. 13 is a diagram schematically illustrating an example of the configuration of an information processing system applicable to the first embodiment. [Figure 14] FIG. 14 is a diagram for explaining an example of imaging using the imaging device according to the first embodiment. [Figure 15-1] FIG. 15A is a diagram for schematically explaining the annotation input process in the information processing system according to the first embodiment. [Figure 15-2] FIG. 15-2 is a diagram for schematically explaining the annotation input process in the information processing system according to the first embodiment. [Figure 15-3] FIG. 15-3 is a diagram for schematically explaining the annotation input process in the information processing system according to the first embodiment. [Figure 16] FIG. 16 is a flowchart illustrating an example of an annotation input process according to the first embodiment. [Figure 17] FIG. 17 is a diagram for explaining image cropping processing applicable to the first embodiment. [Figure 18] FIG. 18 is a diagram showing an example of a screen displayed on a display device by a UI unit according to the first embodiment. [Figure 19]FIG. 19 is a diagram showing an example of a floor plan image applicable to the first embodiment. [Figure 20] FIG. 20 is a diagram showing an example of a displayed annotation input screen according to the first embodiment. [Figure 21] FIG. 21 is a diagram showing an example of an annotation input screen that has been switched to an area designation screen, which is applicable to the first embodiment. [Figure 22] FIG. 22 is a diagram showing an example of a display of a screen for specifying a cutout region according to the first embodiment. [Figure 23] FIG. 23 is a diagram showing an example in which a marker is displayed on a cut-out image displayed in a cut-out image display area according to the first embodiment. [Figure 24] FIG. 24 is a diagram showing an example of a display screen related to report data creation according to the first embodiment. [Figure 25] FIG. 25 is a diagram illustrating an example of report data applicable to the first embodiment. [Figure 26] FIG. 26 is a diagram showing an example in which cracks are observed on a wall surface, which is applicable to the modified example of the first embodiment. [Figure 27] FIG. 27 is a diagram showing an example of an annotation input screen for inputting an annotation for a specified crack according to a modified example of the first embodiment. [Figure 28] FIG. 28 is a diagram showing an example of an annotation input screen that has been switched to an area designation screen, which is applicable to the modified example of the first embodiment. [Figure 29] FIG. 29 is a diagram showing an example of a display of a screen for specifying a cutout region according to a modification of the first embodiment. [Figure 30] FIG. 30 is a diagram schematically showing an imaging device according to the second embodiment. [Figure 31] FIG. 31 is a diagram showing an example of an imaging range that can be imaged by each imaging body applicable to the second embodiment. [Figure 32]FIG. 32 is a functional block diagram illustrating an example of a function of an information processing device serving as an input device for inputting annotations according to the second embodiment. [Figure 33] FIG. 33 is a diagram showing an example of images captured from five different viewpoints by an imaging device applicable to the second embodiment and synthesized for each imaging body. [Figure 34] FIG. 34 is a flowchart illustrating an example of a process for creating a three-dimensional reconstruction model applicable to the second embodiment. [Figure 35] FIG. 35 is a diagram for explaining triangulation applicable to the second embodiment. [Figure 36] FIG. 36 is a diagram for explaining the principle of EPI applicable to the second embodiment. [Figure 37] FIG. 37 is a diagram for explaining the principle of EPI applicable to the second embodiment. [Figure 38] FIG. 38 is a diagram for explaining that the slope m is a value based on a curve when the EPI is configured using a panoramic image, which is applicable to the second embodiment. [Figure 39] FIG. 39 is a diagram for explaining that the slope m is a value based on a curve when the EPI is configured using a panoramic image, which is applicable to the second embodiment. [Figure 40] FIG. 40 is a diagram showing an example of a plane for which parallax is preferentially calculated based on an image captured by an imaging device according to the second embodiment. [Figure 41] FIG. 41 is a diagram showing an example of a large space including a large building as a target for creating a 3D reconstruction model according to the second embodiment. [Figure 42] FIG. 42 is a block diagram illustrating an example of a configuration of an imaging device according to an embodiment. [Figure 43] FIG. 43 is a block diagram showing an example of a configuration of a control unit and a memory in the imaging device according to the second embodiment. [Figure 44] FIG. 44 is a diagram for explaining position designation according to the second embodiment. [Figure 45]FIG. 45 is a diagram showing an example in which a position is specified by a line according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of an input device, an input method for an input device, and an output device and an output method for an output device will be described in detail with reference to the accompanying drawings.
[0010] (First embodiment) In the first embodiment, an annotation is input by specifying a position on a spherical image of a diagnosis target captured using an imaging device capable of capturing an image of the celestial sphere. The input annotation is stored in association with the position.
[0011] By using a spherical image to capture an image of a diagnostic target, it is possible to capture the diagnostic target without omission with fewer captures. Furthermore, since annotations are input by specifying a position on the spherical image, it is possible to intuitively grasp the target to which the annotation is input, and it is possible to prevent input omissions. Furthermore, since the input annotations are saved in association with the position on the spherical image, it is easy to grasp the correspondence between the annotations and the target to which the annotation is input when creating a report based on the annotations, and it is possible to reduce the burden associated with the report creation process. This also makes it possible to prevent human errors from occurring when creating reports.
[0012] A celestial sphere image is an image captured with a 360° angle of view in each of two orthogonal planes (for example, a horizontal plane and a vertical plane) relative to the imaging position, i.e., a 4π steradian angle of view relative to the imaging position. An annotation is information based on the results of observing and diagnosing a diagnostic object. An annotation can include one or more items predetermined for each type of diagnostic object and any comment. An annotation can further include an image.
[0013] [Image capture device applicable to the first embodiment] Fig. 1 is a diagram schematically illustrating the appearance of an example of an imaging device applicable to the first embodiment. In Fig. 1, imaging device 1a has imaging lens 20a provided on a first surface of a substantially rectangular parallelepiped housing 10a, and imaging lens 20b provided on a second surface opposite to the first surface at a position corresponding to imaging lens 20a. Imaging elements are provided within housing 10a corresponding to each of imaging lenses 20a and 20b.
[0014] Light incident on each of the imaging lenses 20a and 20b is irradiated onto the corresponding imaging element via an imaging optical system including the imaging lenses 20a and 20b, which is provided in the housing 10a. The imaging element is, for example, a CCD (Charge Coupled Device), which is a light receiving element that converts the irradiated light into electric charges. However, the imaging element is not limited to this, and may also be a CMOS (Complementary Metal Oxide Semiconductor) image sensor.
[0015] As will be described in detail later, each driver that drives each image sensor controls the shutter of the image sensor in response to a trigger signal and reads out the electric charges converted from light from each image sensor. Each driver converts the electric charges read out from each image sensor into electrical signals, and converts these electrical signals into captured images as digital data and outputs them. Each captured image output from each driver is stored in a memory, for example.
[0016] In the following description, a configuration including the pair of imaging lenses 20a and 20b, and the imaging optical systems and imaging elements corresponding to the pair of imaging lenses will be referred to as imaging body 21. For convenience, the operation of outputting a captured image based on light incident on imaging lenses 20a and 20b in response to a trigger signal will be referred to as "imaging by imaging body 21." When it is necessary to distinguish between imaging lenses 20a and 20b, it will be referred to as "imaging by imaging lens 20a," for example.
[0017] The shutter button 30 is a button that, when operated, commands imaging by the imaging body 21. When the shutter button 30 is operated, imaging is synchronously performed by the imaging lenses 20a and 20b in the imaging body 21.
[0018] As shown in Fig. 1, the housing 10a of the imaging device 1a includes an imaging section area 2a where the imaging body 21 is arranged, and an operation section area 3a where the shutter button 30 is arranged. The operation section area 3a includes a grip section 31 with which the user holds the imaging device 1a. The grip section 31 has a non-slip surface to make it easy for the user to hold and operate. In addition, a fixing section 32 is provided on the bottom of the imaging device 1a, i.e., the bottom of the grip section 31, for fixing the imaging device 1a to a tripod or the like.
[0019] Next, the structure of the imaging body 21 will be described in more detail. FIG. 2 is a diagram showing the structure of an example of the imaging body 21 applicable to the first embodiment. In FIG. 2, the imaging body 21 includes imaging optical systems 201a and 201b, each including imaging lenses 20a and 20b, and imaging elements 200a and 200b, each of which is a CCD or CMOS sensor. Each of the imaging optical systems 201a and 201b is configured as, for example, a fisheye lens with 6 groups and 7 lenses. This fisheye lens has a total angle of view of 180° (=360° / n, where n is the number of optical systems=2) or more, preferably greater than 180°, more preferably 185° or more, and more preferably 190° or more.
[0020] Each of the imaging optical systems 201a and 201b includes a prism 202a or 202b that changes the optical path by 90°. The six-group, seven-element fisheye lenses included in each of the imaging optical systems 201a and 201b can be divided into a group on the incident side of the prisms 202a or 202b and a group on the exit side (the image sensor 200a or 200b side). For example, in the imaging optical system 201a, light incident on the imaging lens 20a is incident on the prism 202a via the lenses belonging to the group on the incident side of the prism 202a. The light incident on the prism 202a has its optical path changed by 90° and is irradiated onto the image sensor 200a via the lenses, aperture stop, and filter belonging to the group on the exit side of the prism 202a.
[0021] The optical elements (lenses, prisms 202a and 202b, filters, and aperture stops) of the two imaging optical systems 201a and 201b are positioned relative to the image sensors 200a and 200b. More specifically, the optical elements of the imaging optical systems 201a and 201b are positioned so that the optical axes of the optical elements are perpendicular to the centers of the light-receiving areas of the corresponding image sensors 200a and 200b, and so that the light-receiving areas become the image planes of the corresponding fisheye lenses. In the image sensor 21, the imaging optical systems 201a and 201b have the same specifications and are combined in opposite directions so that their optical axes coincide.
[0022] Fig. 3 is a three-dimensional view schematically illustrating the appearance of an imaging device 1a according to an embodiment. Fig. 3(a), Fig. 3(b), and Fig. 3(c) are examples of a top view, a front view, and a side view of the imaging device 1a, respectively. As shown in Fig. 3(c), an imaging body 21 is configured by an imaging lens 20a and an imaging lens 20b that is arranged on the rear side of the imaging lens 20a at a position corresponding to the imaging lens 20a.
[0023] 3(a) and 3(c), angle α indicates an example of the angle of view (imaging range) of imaging lenses 20a and 20b. As described above, each of imaging lenses 20a and 20b included in imaging body 21 captures images with angle α greater than 180°. Therefore, in order to prevent housing 10a from appearing in each captured image captured by each imaging lens 20a and 20b, the first and second surfaces of housing 10a are chamfered on both sides of center line C of each imaging lens 20a and 20b in accordance with the angle of view of each imaging lens 20a and 20b, as shown as surfaces 231, 232, 233, and 234 in FIGS. 3(a), 3(b), and 3(c).
[0024] The center line C is a line that passes vertically through the center of each imaging lens 20a and 20b when the orientation of the imaging device 1a is set so that the direction of the vertex (pole) of each hemispherical image captured by the imaging device 1a is parallel to the vertical direction.
[0025] The imaging body 21, by combining the imaging lenses 20a and 20b, has an imaging range that covers the entire celestial sphere centered on the center of the imaging body 21. That is, as described above, the imaging lenses 20a and 20b each have an angle of view of 180° or more, preferably greater than 180°, and more preferably greater than 185°. Therefore, by combining the imaging lenses 20a and 20b, for example, the imaging range of the plane perpendicular to the first plane and the imaging range of the plane parallel to the first plane can each be 360°, and by combining these, an imaging range of the entire celestial sphere is realized.
[0026] [Configuration of signal processing in the imaging device according to the first embodiment] Fig. 4 is a block diagram showing an example of the configuration of the imaging device 1a according to the first embodiment. In Fig. 4, parts corresponding to those in Figs. 1 and 2 described above are given the same reference numerals, and detailed description thereof will be omitted.
[0027] 4, imaging device 1a includes, as an imaging system configuration, imaging body 21 including imaging elements 200a and 200b and driving units 210a and 210b, and signal processing units 211a and 211b. Also, imaging device 1a includes, as an imaging control and signal processing system configuration, a CPU (Central Processing Unit) 2000, a ROM (Read Only Memory) 2001, a memory controller 2002, a RAM (Random Access Memory) 2003, a trigger I / F 2004, a SW (Switch) circuit 2005, a data I / F 2006, a communication I / F 2007, and an acceleration sensor 2008, all of which are connected to a bus 2010. Furthermore, imaging device 1a includes a battery 2020 for supplying power to each of these units.
[0028] First, the configuration of the imaging system will be described. In the configuration of the imaging system, imaging element 200a and driver 210a correspond to imaging lens 20a, respectively. Similarly, imaging element 200b and driver 210b correspond to imaging lens 20b, respectively. These imaging element 200a and driver 210a, as well as imaging element 200b and driver 210b, are each included in imaging body 21.
[0029] In the imaging body 21, the driver 210a drives the imaging element 200a and reads out charges from the imaging element 200a in response to a trigger signal supplied from the trigger I / F 2004. The driver 210a outputs one frame of a captured image based on the charges read out from the imaging element 200a in response to one trigger signal. The driver 210a converts the charges read out from the imaging element 200a into electrical signals and outputs the signals. The signals output from the driver 210a are supplied to the signal processor 211a. The signal processor 211a performs predetermined signal processing such as noise removal and gain adjustment on the signals supplied from the driver 210a, converts the analog signals into digital signals, and outputs the resulting captured image as digital data. This captured image is a hemispherical image (fisheye image) obtained by capturing an area of the entire celestial sphere that is a hemisphere corresponding to the angle of view of the imaging lens 20a.
[0030] In the imaging body 21, the imaging element 200b and the driving section 210b have the same functions as the imaging element 200a and the driving section 210a described above, and therefore their description will be omitted here. Also, the signal processing section 211b has the same functions as the signal processing section 211a described above, and therefore their description will be omitted here.
[0031] The hemispherical images output from the signal processing units 211 a and 211 b are stored in the RAM 2003 via the memory controller 2002 .
[0032] Next, the configuration of the imaging control and signal processing system will be described. The CPU 2000 operates in accordance with a program pre-stored in, for example, the ROM 2001, using a portion of the storage area of the RAM 2003 as a work memory, and controls the overall operation of this imaging device 1a. The memory controller 2002 controls the storage and reading of data in the RAM 2003 in accordance with instructions from the CPU 2000.
[0033] The SW circuit 2005 detects an operation on the shutter button 30 and passes the detection result to the CPU 2000. When the CPU 2000 receives a detection result from the SW circuit 2005 indicating that an operation on the shutter button 30 has been detected, it outputs a trigger signal. The trigger signal is output via a trigger I / F 2004, branches, and is supplied to each of the drive units 210a and 210b.
[0034] The data I / F 2006 is an interface for performing data communication with an external device. For example, a USB (Universal Serial Bus) or Bluetooth (registered trademark) can be used as the data I / F 2006. The communication I / F 2007 is connected to a network and controls communication with the network. The network connected to the communication I / F 2007 may be either wired or wireless, or may be connectable to both wired and wireless networks.
[0035] In the above description, the CPU 2000 outputs a trigger signal in response to the detection result by the SW circuit 2005, but this is not limited to this example. For example, the CPU 2000 may output a trigger signal in response to a signal supplied via the data I / F 2006 or the communication I / F 2007. Furthermore, the trigger I / F 2004 may generate a trigger signal in response to the detection result by the SW circuit 2005 and supply it to each of the driving units 210a and 210b.
[0036] The acceleration sensor 2008 detects three-axis acceleration components and passes the detection results to the CPU 2000. The CPU 2000 detects the vertical direction based on the detection results of the acceleration sensor 2008 and calculates the tilt of the imaging device 1a with respect to the vertical direction. The CPU 2000 adds tilt information indicating the tilt of the imaging device 1a to each hemispherical image captured by the imaging lenses 20a and 20b in the imaging body 21 and stored in the RAM 2003.
[0037] The battery 2020 is a secondary battery such as a lithium ion secondary battery, and serves as a power supply unit that supplies power to each unit within the imaging device 1a that requires power supply. The battery 2020 includes a charge / discharge control circuit that controls charging and discharging of the secondary battery.
[0038] FIG. 5 is a diagram showing an example of the arrangement of the battery 2020 and the circuit unit 2030 in the imaging device 1a according to the embodiment. As described above, the battery 2020 and the circuit unit 2030 are provided inside the housing 10a. Furthermore, of the battery 2020 and the circuit unit 2030, at least the battery 2020 is fixed inside the housing 10a by fixing means such as adhesive or screws. Note that FIG. 5 corresponds to the above-mentioned FIG. 3(b) and shows an example of the imaging device 1a as viewed from the front. The circuit unit 2030 includes, for example, each element of the above-mentioned imaging control and signal processing systems, and is configured, for example, on one or more circuit boards.
[0039] Here, it has been described that the battery 2020 and the circuit unit 2030 are placed at the position, but this is not limited to this example. For example, if the circuit unit 2030 is sufficiently small, it is sufficient that at least the battery 2020 is placed at the position.
[0040] In this configuration, a trigger signal is output from the trigger I / F 2004 in response to an operation of the shutter button 30. The trigger signal is supplied to the driving units 210a and 210b at the same time. The driving units 210a and 210b read out electric charges from the imaging elements 200a and 200b in synchronization with the supplied trigger signal.
[0041] The driving units 210a and 210b convert the electric charges read from the imaging elements 200a and 200b into electrical signals and supply them to the signal processing units 211a and 211b, respectively. The signal processing units 211a and 211b perform the predetermined processing described above on the signals supplied from the driving units 210a and 210b, convert these signals into respective hemispherical images, and output the respective hemispherical images. Each hemispherical image is stored in the RAM 2003 via the memory controller 2002.
[0042] Each hemispherical image stored in the RAM 2003 is sent to an external information processing device via a data I / F 2006 or a communication I / F 2007 .
[0043] [Outline of image processing according to the first embodiment] Next, image processing for a hemispherical image according to the first embodiment will be described. Fig. 6 is a block diagram showing the configuration of an example of an information processing device as an input device for inputting annotations, which is applicable to the first embodiment. In Fig. 6, the information processing device 100a includes a CPU 1000, a ROM 1001, a RAM 1002, a graphic I / F 1003, a storage 1004, a data I / F 1005, a communication I / F 1006, and an input device 1011, which are all connected by a bus 1030, and further includes a display device 1010 connected to the graphic I / F 1003.
[0044] The storage 1004 is a nonvolatile memory such as a flash memory, and stores programs and various data for the operation of the CPU 1000. A hard disk drive may be used for the storage 1004. The CPU 1000 operates in accordance with the programs stored in the ROM 1001 and the storage 1004, using the RAM 1002 as a work memory, and controls the overall operation of the information processing device 100a.
[0045] The graphic I / F 1003 generates a display signal that can be displayed by the display device 1010 based on a display control signal generated by the CPU 1000 according to a program, and sends the generated display signal to the display device 1010. The display device 1010 includes, for example, an LCD (Liquid Crystal Display) and a drive circuit that drives the LCD, and displays a screen according to the display signal sent from the graphic I / F 1003.
[0046] The input device 1011 outputs a signal according to a user operation and accepts a user input. In this example, the information processing device 100a is a tablet personal computer, and is configured as a touch panel 1020 in which the display device 1010 and the input device 1011 are integrally formed. The input device 1011 allows the display of the display device 1010 to be seen through, and outputs position information according to a contact position. The information processing device 100a is not limited to a tablet personal computer, and may be, for example, a desktop personal computer.
[0047] The data I / F 1005 is an interface for performing data communication with an external device. For example, USB or Bluetooth (registered trademark) can be used as the data I / F 1005. The communication I / F 1006 is connected to a network via wireless communication and controls communication with the network. The communication I / F 1006 may be connected to the network via wired communication. Here, it is assumed that the imaging device 1a and the information processing device 100a are connected via a wired connection via the data I / F 1005.
[0048] 7 is a functional block diagram illustrating an example of functions of the information processing device 100a according to the first embodiment. In FIG. 7, the information processing device 100a includes an image acquisition unit 110, an image processing unit 111, an auxiliary information generation unit 112, a UI (User Interface) unit 113, a communication unit 114, and an output unit 115.
[0049] The image acquisition unit 110, image processing unit 111, incidental information generation unit 112, UI unit 113, communication unit 114, and output unit 115 are realized by the input program according to the first embodiment running on the CPU 1000. However, without being limited to this, some or all of the image acquisition unit 110, image processing unit 111, incidental information generation unit 112, UI unit 113, communication unit 114, and output unit 115 may be configured by hardware circuits that operate in cooperation with each other.
[0050] The UI unit 113 is a receiving unit that receives user input in response to a user operation on the input device 1022 and executes processing in response to the received user input. The UI unit 113 can store information received in response to the user input in the RAM 1002 or the storage 1004. The UI unit 113 also serves as a display unit that generates a screen to be displayed on the display device 1010. The communication unit 114 controls communication via the data I / F 1005 and the communication I / F 1006.
[0051] The image acquisition unit 110 acquires each hemispherical image captured by the imaging lenses 20a and 20b from the imaging device 1a and tilt information added to each hemispherical image via the data I / F 1005. The image processing unit 111 executes image conversion processing to generate one omnidirectional image by joining the hemispherical images acquired by the image acquisition unit 110. The incidental information generation unit 112 generates incidental information (annotations) at designated positions on the omnidirectional image generated by the image processing unit 111 in accordance with a user input received by the UI unit 113. The output unit 115 creates and outputs a report in a predetermined format based on the annotations generated by the incidental information generation unit 112.
[0052] An input program for realizing each function of the information processing device 100a according to the first embodiment is provided by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD (Compact Disk), a flexible disk (FD), or a DVD (Digital Versatile Disk). Alternatively, the input program may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. The input program may also be configured to be provided or distributed via a network such as the Internet.
[0053] The input program has a modular configuration including an image acquisition unit 110, an image processing unit 111, an incidental information generation unit 112, a UI unit 113, a communication unit 114, and an output unit 115. In terms of actual hardware, when the CPU 1000 reads out the input program from a storage medium such as the storage 1004 and executes it, the above-mentioned units are loaded onto a main storage device such as the RAM 1002, and the image acquisition unit 110, the image processing unit 111, the incidental information generation unit 112, the UI unit 113, the communication unit 114, and the output unit 115 are generated on the main storage device.
[0054] [Image transformation applicable to the first embodiment] Next, image conversion processing applicable to the first embodiment will be described with reference to Figures 8 to 12. The image conversion processing described below is processing executed by, for example, the image processing unit 111 in the information processing device 100a.
[0055] 8 is a diagram for explaining how imaging lenses 20a and 20b applicable to the first embodiment project three-dimensional incident light into two dimensions. Note that, in the following explanation, unless otherwise specified, imaging lens 20a will be used as a representative of imaging lenses 20a and 20b.
[0056] In Fig. 8(a), the imaging lens 20a includes a fisheye lens 24 (imaging optical system 201a) and an imaging element 200a. The axis perpendicular to the light-receiving surface of the imaging element 200a is defined as the optical axis. In the example of Fig. 6(a), the incident angle φ is expressed as the angle with respect to the optical axis, with the intersection of the plane tangent to the edge of the fisheye lens 24 and the optical axis defined as the vertex.
[0057] A fisheye image (hemispherical image) captured by a fisheye lens 24 with an angle of view exceeding 180° is an image of a scene equivalent to a hemisphere from the imaging position. However, as shown in FIGS. 8(a) and 8(b), a hemispherical image 22 is generated at an image height h corresponding to an incident angle φ, the relationship of which is determined by a projection function f(φ). Note that in FIG. 8(b), areas on the imaging surface of the image sensor 200a that are not illuminated by light from the fisheye lens 24 are indicated as invalid areas by being filled in black. The projection function f(φ) varies depending on the properties of the fisheye lens 24. For example, there is a fisheye lens 24 that uses a projection method known as the equidistant projection method, which is expressed by the following equation (1) where φ is the image height h, the focal length f, and the angle between the incident direction and the optical axis (incident angle). We will use this fisheye lens 24 in this example.
[0058]
number
[0059] 9 is a diagram for explaining the tilt of the imaging device 1a according to the first embodiment. In FIG. 9, the vertical direction coincides with the z-axis in the Cartesian coordinate system of the x, y, and z three-dimensional directions of the global coordinate system. When this direction is parallel to the center line C of the imaging device 1a shown in FIG. 3(b), the camera is not tilted. When these are not parallel, the imaging device 1a is in a tilted state.
[0060] The imaging device 1a associates each hemispherical image captured by the imaging lenses 20a and 20b with the output value output from the acceleration sensor 2008 at the time of capturing the image, and stores the associated images in, for example, the RAM 2003. The information processing device 100a acquires each hemispherical image and the output value of the acceleration sensor 2008 stored in the RAM 2003 from the imaging device 1a.
[0061] In the information processing device 100a, the image processing unit 111 calculates the tilt angle from the gravity vector (hereinafter referred to as tilt angle α) and the tilt angle β in the xy plane (hereinafter referred to as tilt angle β) using the output value of the acceleration sensor 2008 acquired from the image capturing device 1a, using the following equations (2) and (3). Here, value Ax represents the value of the x0-axis component of the output value of the acceleration sensor 2008 in the camera coordinate system, value Ay represents the value of the y0-axis component of the output value of the acceleration sensor 2008 in the camera coordinate system, and value Az represents the value of the z0-axis component of the output value of the acceleration sensor 2008 in the camera coordinate system. The image processing unit 111 calculates the tilt angles α and β from the values of the axial components of the acceleration sensor 2008 in response to a trigger signal. The image processing unit 111 associates the calculated tilt angles α and β with each hemispherical image acquired from the image capturing device 1a and stores them in, for example, the storage 1004.
[0062]
number
[0063]
number
[0064] The image processing unit 111 generates a spherical image based on each hemispherical image acquired from the image capturing device 1a and the tilt angles α and β associated with each hemispherical image.
[0065] Fig. 10 is a diagram illustrating a format of a celestial sphere image applicable to the first embodiment. Fig. 10(a) shows an example of a format when a celestial sphere image is represented on a plane, and Fig. 10(b) shows an example of a format when a celestial sphere image is represented on a spherical surface. When represented on a plane, the format of the celestial sphere image is an image having pixel values corresponding to angle coordinates (φ, θ) with horizontal angles ranging from 0° to 360° and vertical angles ranging from 0° to 180°, as shown in Fig. 10(a). The angle coordinates (φ, θ) correspond to each coordinate point on the spherical surface shown in Fig. 10(b), and are similar to latitude and longitude coordinates on a globe.
[0066] The relationship between the planar coordinate values of an image captured with a fisheye lens and the spherical coordinate values of a celestial sphere image can be established by using the projection function f (h=f(θ)) as described in Fig. 8. As a result, by converting and combining (combining) two partial images (hemispherical images) captured with a fisheye lens, a planar celestial sphere image as shown in Fig. 10(a) can be created.
[0067] In the first embodiment, a conversion table that associates each pixel position of a hemispherical image with each pixel position in a celestial sphere image based on the plane shown in FIG. 10(a) is created in advance and stored in, for example, the storage 1004 of the information processing device 100a. Table 1 shows an example of this conversion table. FIG. 11 is a diagram for explaining the association between each pixel position of a hemispherical image and each pixel position of a celestial sphere image based on the conversion table, which is applicable to the first embodiment.
[0068] [Table 1]
[0069] As shown in Table 1, the conversion table has a data set of coordinate values (θ, φ) [pix: pixel] of the converted image and corresponding coordinate values (x, y) [pix] of the pre-conversion image for all coordinate values of the converted image. A converted image can be generated from a captured hemispherical image (pre-conversion image) according to the conversion table shown in Table 1. Specifically, as shown in FIG. 11, based on the correspondence between pre-conversion coordinates and post-conversion coordinates shown in the conversion table of Table 1, each pixel of the converted image can be generated by referring to the pixel value of the coordinate value (x, y) [pix] of the pre-conversion image corresponding to the coordinate value (θ, φ) [pix].
[0070] The conversion table shown in Table 1 reflects distortion correction assuming that the direction of the center line C of the imaging device 1a is parallel to the vertical direction. By applying a correction process to this conversion table according to the tilt angles α and β, it is possible to perform a correction (called vertical correction) that makes the center line C of the imaging device 1a parallel to the vertical direction.
[0071] Fig. 12 is a diagram for explaining vertical correction applicable to the first embodiment. Fig. 12(a) shows a camera coordinate system, and Fig. 12(b) shows a global coordinate system. In Fig. 12(b), the three-dimensional Cartesian coordinates of the global coordinate system are expressed as (x1, y1, z1), and the spherical coordinates are expressed as (θ1, φ1). In Fig. 12(a), the three-dimensional Cartesian coordinates of the camera coordinate system are expressed as (x0, y0, z0), and the spherical coordinates are expressed as (θ0, φ0).
[0072] The image processing unit 111 performs a vertical correction calculation to convert spherical coordinates (θ1, φ1) to spherical coordinates (θ0, φ0) using the following equations (4) to (9). First, to correct the tilt, a rotational transformation using three-dimensional orthogonal coordinates is required, so the image processing unit 111 converts spherical coordinates (θ1, φ1) to three-dimensional orthogonal coordinates (x1, y1, z1) using equations (4) to (6).
[0073]
number
[0074]
number
[0075]
number
[0076] Next, the image processing unit 111 performs rotational coordinate transformation shown in equation (7) using the tilt angle (α, β) to transform the global coordinate system (x1, y1, z1) into the camera coordinate system (x0, y0, z0). In other words, equation (6) defines the tilt angle (α, β).
[0077]
number
[0078] This means that the global coordinate system is first rotated by α around the z-axis and then rotated by β around the x-axis to become the camera coordinate system. Finally, the image processing unit 111 converts the three-dimensional Cartesian coordinates (x0, y0, z0) of the camera coordinate system back into spherical coordinates (θ0, φ0) using equations (8) and (9).
[0079]
number
[0080]
number
[0081] In the above description, coordinate conversion is performed by executing a vertical correction calculation, but this is not limited to this example. For example, multiple conversion tables corresponding to tilt angles (α, β) may be created and stored in advance. This allows the vertical correction calculation process to be omitted, thereby speeding up processing.
[0082] [Outline of input processing according to the first embodiment] Next, the annotation input process according to the first embodiment will be described in more detail. Fig. 13 is a diagram schematically illustrating an example of the configuration of an information processing system applicable to the first embodiment. In Fig. 13(a), the information processing system includes an imaging device 1a, an information processing device 100a connected to the imaging device 1a via wired or wireless communication, and a server 6 connected to the information processing device 100a via a network 5.
[0083] Each hemispherical image captured by each imaging lens 20a and 20b in the imaging device 1a is transmitted to the information processing device 100a. The information processing device 100a converts each hemispherical image transmitted from the imaging device 1a into a celestial sphere image as described above. Furthermore, the information processing device 100a adds annotations to the converted celestial sphere image in accordance with a user input. The information processing device 100a transmits, for example, the celestial sphere image and the annotations to the server 6 via the network 5. The server 6 stores and manages the celestial sphere image and the annotations transmitted from the information processing device 100a.
[0084] Furthermore, the information processing device 100a acquires a spherical image and annotations from the server 6, and creates report data based on the acquired spherical image and annotations. Without being limited to this, the information processing device 100a can also create report data based on a spherical image and annotations stored in the information processing device 100a. The information processing device 100a transmits the created report data to the server 6.
[0085] 13(a) and 13(b), the server 6 is shown as being configured by a single computer, but this is not limited to this example. That is, the server 6 may be configured so that its functions are distributed among a plurality of computers. The server 6 may also be configured as a cloud system on the network 5.
[0086] 13(b), a configuration in which the imaging device 1a and the information processing device 100a are each connected to a network 5 is also possible. In this case, each hemispherical image captured by the imaging device 1a is transmitted to the server 6 via the network 5. The server 6 performs the conversion process on each hemispherical image as described above, converts each hemispherical image into a celestial sphere image, and stores the converted image. The information processing device 100a adds annotations to the celestial sphere image stored in the server 6.
[0087] However, in the first embodiment, the information processing device 100a may not be connected to the server 6, and the processing may be performed locally in the information processing device 100a.
[0088] FIG. 14 is a diagram illustrating an example of imaging using the imaging device 1a according to the first embodiment. For example, the imaging device 1a is used to capture an image of the entire celestial sphere inside a building 4 that is a diagnosis target. Imaging may be performed at multiple positions inside the building 4. Because the entire imaging range of the celestial sphere can be captured with a single imaging operation, it is possible to prevent omissions in imaging, including the ceiling inside the building 4. Also, in FIG. 14, imaging is performed inside the building 4 using the imaging device 1a, but this is not limited to this example, and the exterior surface of the building 4 can also be captured using the imaging device 1a. In this case, for example, by capturing images multiple times while moving around the entire perimeter of the building 4, it is possible to capture images of the exterior surface (for example, wall surfaces) of the building 4 without omissions.
[0089] Annotation input processing in the information processing system according to the first embodiment will be outlined with reference to Figs. 15-1 to 15-3. The information processing system according to the first embodiment provides several functions for assisting annotation input to a spherical image. Among the provided auxiliary functions, a function for automatically estimating an imaging position, a 3D panorama automatic tour function for realizing movement of the viewpoint among a plurality of spherical images, and an annotation input and confirmation function will be outlined here.
[0090] Fig. 15-1 is a diagram for explaining the automatic estimation function of the imaging position according to the first embodiment. In Fig. 15-1, an image 5010 of a predetermined area of a celestial sphere image based on each hemispherical image obtained by capturing, for example, the interior of a building 4 to be diagnosed in a celestial sphere is displayed on a screen 500 displayed on a display device 1010. Also displayed on the screen 500 is a floor plan image 5011 of the building 4 that has been acquired in advance.
[0091] In the example of FIG. 15-1, pin markers 50401, 50402, 50403, . . . indicating the image capturing positions where images were captured are displayed on a floor plan image 5011.
[0092] For example, using the imaging device 1a, an image of a diagnosis target is captured multiple times while moving the imaging position, and multiple pairs of hemispherical images are acquired by the imaging lenses 20a and 20b. In the information processing device 100a, the image processing unit 111 generates multiple omnidirectional images based on each of the multiple pairs of hemispherical images, and performs a matching process between the generated omnidirectional images using a known method. The image processing unit 111 estimates the relative imaging positions of each omnidirectional image based on the results of the matching process and each hemispherical image before conversion of each omnidirectional image. While maintaining the relative positional relationship between the estimated imaging positions, the image processing unit 111 rotates and scales each pair of imaging positions to associate them with the floor plan image 5011, thereby estimating each imaging position in the building 4.
[0093] Furthermore, in the information processing device 100a, the UI unit 113 displays a cursor 5041 that can be moved to any position in response to a user operation on the screen 500. The UI unit 113 displays, on the screen 500, an image 5010 based on the spherical image that corresponds to the pin marker designated by the cursor 5041, among the pin markers 50401, 50402, 50403, .
[0094] In this way, by automatically estimating the imaging positions at which a plurality of omnidirectional images are captured, the user can easily understand the relationship between the plurality of omnidirectional images, thereby improving work efficiency.
[0095] Note that the UI unit 113 can display tags at positions designated by the user on the image 5010. In the example of Fig. 15-1, tags 5030a and 5030b are displayed for the image 5010. Although details will be described later, the tags 5030a and 5030b indicate positions of targets to which annotations have been input, and are associated with coordinates of the omnidirectional image including the image 5010.
[0096] 15-2 is a diagram for explaining the 3D panorama automatic tour function according to the first embodiment. For example, referring to FIG. 15-2(a), the pin marker 5040 in the floor plan image 5011 10 is specified, and pin marker 5040 10 It is assumed that an image 5010a (see FIG. 15-2(b)) corresponding to the position of the pin marker 5040 is displayed on the screen 500. In this state, when, for example, a cursor 541 is instructed to move diagonally forward to the left (as indicated by the arrow in the drawing), the UI unit 133 instructs the pin marker 5040 to move diagonally forward to the left. 10 From the pin marker 5040 10 The pin marker 5040 displayed diagonally in front of the right of 11 and move the pin marker 5040 11 Similarly, the UI unit 113 displays an image 5010b based on the spherical image corresponding to the pin marker 5040 on the screen 500. 10 From the pin marker 5040 10 The pin marker 5040 displayed diagonally in front of the right of 12 When the pin marker 5040 is moved to the 12 An image 5010c based on the spherical image corresponding to the image 5010 is displayed on the screen 500.
[0097] In this way, by specifying movement within screen 500, the specified pin marker moves to the adjacent pin marker in the direction of movement, for example, pin markers 50401, 50402, 50403, ..., and accordingly, image 5010 displayed on screen 500 switches. This allows the user to operate the device as if he or she were observing the diagnostic target in sequence, thereby making it possible to improve the efficiency of the diagnostic work.
[0098] FIG. 15-3 is a diagram for explaining an annotation input and confirmation function according to the first embodiment. The UI unit 113 allows a user to specify a desired position on an image 5010 on a screen 500 using a cursor 5041, whereby a tag 5030a indicating the position is displayed superimposed on the image 5010. Also, an annotation input screen 600 for inputting an annotation corresponding to the position of the tag 5030a is displayed on the screen 500 together with the image 5010. That is, the annotation input screen 600 is a diagnostic information input screen for inputting diagnostic information related to a diagnostic target. Note that the annotation input screen 600 is highlighted in FIG. 15-3 for ease of explanation.
[0099] As will be described in detail later, the annotation input screen 600 allows input to preset items and input of arbitrary comments. The annotation input screen 600 also allows editing of information that has already been input. Furthermore, the annotation input screen 600 also allows input of an image within a specified range including the position indicated by the tag 5030a in the image 5010, or an image acquired from an external source, as an annotation.
[0100] In this way, a position is specified on the spherical image and an annotation is input for the specified position, which makes it easy to associate the annotation with the image, thereby improving work efficiency.
[0101] [Details of input processing according to the first embodiment] Next, the annotation input process according to the first embodiment will be described in more detail. Fig. 16 is a flowchart showing an example of the annotation input process according to the first embodiment.
[0102] 16 , the information processing device 100a acquires each hemispherical image of the diagnosis target captured by the imaging lenses 20a and 20b in the imaging device 1a and the output value of the acceleration sensor 2008, generates a celestial sphere image based on the hemispherical images, and stores the image in, for example, the storage 1004. In the following, to avoid complexity, the description will be given assuming one celestial sphere image.
[0103] In step S100, in the information processing device 100a, the image acquisition unit 110 acquires the spherical image stored in the storage 1004. In the next step S101, the UI unit 113 cuts out an image of a predetermined region from the spherical image acquired by the image acquisition unit 110 in step S100, and displays the image on the screen 500.
[0104] Fig. 17 is a diagram for explaining image cropping processing by the UI unit 113, which is applicable to the first embodiment. Note that Fig. 17 corresponds to Fig. 10(a) described above. The UI unit 113 crops an image of a partial area 511 of a celestial sphere image 510 generated from two hemispherical images, and displays the image 5010 on the screen 500. In step S101, the image of the area 511 that is predetermined as an initial value is cropped. Here, the UI unit 113 replaces, for example, angular coordinates (φ, θ) of the celestial sphere image 510 with coordinates (x, y) suitable for display on the screen 500, and displays the image 5010.
[0105] In the next step S102, in response to a user operation, the UI unit 113 moves the region 511 within the spherical image 510. Furthermore, the UI unit 113 can change the size of the region 511 relative to the spherical image 510 in response to a user operation.
[0106] 18 is a diagram showing an example of a screen 500 displayed on the display device 1010 by the UI unit 113 according to the first embodiment. In this example, the UI unit 113 displays an image 5010 of the area 511 shown in Fig. 17 over the entire surface of the screen 500. On the screen 500, buttons 5020, 5021, 5022, and 5023, and menu buttons 5050 and 5051 are further displayed.
[0107] Button 5020 is a button for loading the spherical image 510. Button 5021 is a button for saving annotation data, which is set on the screen 500 and will be described later. The current screen 500 itself may also be saved in response to an operation of button 5021. Button 5022 is a button for changing the scale, and an instruction screen for instructing enlargement and reduction of the image 5010 on the screen 500 is displayed in response to the operation. Button 523 is a button for displaying an annotation input screen 600, which will be described later, on the screen 500.
[0108] The menu button 5050 is a button for displaying a menu for switching the display mode of the screen 500. The menu button 5051 is a button for switching the viewpoint for the image 5010.
[0109] Also, on the screen 500, an area 5024 displays each imaging position estimated using the function of automatically estimating the imaging position described with reference to FIG. 15-1, in relative position using pin markers 50411, 50412, . . .
[0110] Here, the information processing device 100a acquires in advance a floor plan image 5011 including a diagnosis target as exemplified in FIG. 19 and stores it in, for example, the storage 1004. The UI unit 113 can display this floor plan image 5011 on the screen 500. For example, the UI unit 113 can display the floor plan image 5011 by superimposing it on an area 5024 on the screen 500. The UI unit 113 can also display pin markers 50411, 50412, ... indicating imaging positions on the floor plan image 5011 displayed on the screen 500. By performing a predetermined input operation, the user can adjust the positions of the pin markers 50411, 50412, ... on the floor plan image 5011 to match the positions where imaging was actually performed while maintaining the relative positional relationship between the pin markers 50411, 50412, ....
[0111] In the next step S103, the UI unit 113 waits for an operation on the button 5023, and when the button 5023 is operated, the UI unit 113 displays the annotation input screen 600 on the screen 500. Then, in response to a predetermined operation on the annotation input screen 600, a position on the image 5010 is specified.
[0112] FIG. 20 is a diagram showing an example in which an annotation input screen 600 according to the first embodiment is displayed on the screen 500. Note that in FIG. 20, parts common to those in FIG. 19 described above are assigned the same reference numerals, and detailed description thereof will be omitted. In FIG. 20, a marker 602 is displayed superimposed on the position specified in step S103. Also in FIG. 20, the annotation input screen 600 is displayed superimposed on the image 5010. The UI unit 113 associates position information indicating the specified position with identification information (referred to as an annotation ID) that identifies the annotation, and stores the associated information in, for example, the RAM 1002. Here, the position information is, for example, coordinates in the omnidirectional image 510 including the image 5010.
[0113] The annotation input screen 600 may be displayed on the same screen 500 as the image 5010, and is not limited to being displayed superimposed on the image 5010. For example, the annotation input screen 600 and the image 5010 may be displayed in different areas within the screen 500, or the annotation input screen 600 may be displayed in a window separate from the image 5010. Furthermore, the display position of the annotation input screen 600 on the screen 500 can be changed in response to a user operation.
[0114] 20, the annotation input screen 600 is provided with tabs 601a, 601b, and 601c for switching functions on the annotation input screen 600. The tab 601a is a tab for editing annotations. The tab 601b is a tab for inserting a detailed image, and in response to an operation, a file selection screen provided by, for example, the OS (Operating System) of the information processing device 100a is displayed, allowing a desired image to be selected. The tab 601c is a tab for selecting a cutout area, and in response to an operation, the display on the screen 500 is switched to a display for specifying an area on the image 5010.
[0115] 20, a button 6011 is a button for adding a tag to an image 5010. When the button 6011 is operated, for example, a position designation operation is performed on the image 5010 using a cursor 5041, the UI unit 113 displays a marker 602 at the designated position.
[0116] In the next step S104, the UI unit 113 waits for an operation on each of the tabs 601a to 601c on the annotation input screen 600 and the button 5021 on the screen 500, and when an operation is performed on any of them, the UI unit 113 executes the corresponding process.
[0117] When the annotation input screen 600 is displayed (initial screen) in response to the operation of the button 5023, or when the tab 601a is specified, the UI unit 113 determines that an annotation is to be input, and proceeds to step S110. In the next step S111, the UI unit 113 displays the annotation input screen 600 including an annotation input area 6010 for inputting an annotation, as shown in FIG.
[0118] In the example of Fig. 20, the annotation input area 6010 includes input areas 6010a to 6010d. In this example, the input area 6010a is an input area for selecting a predetermined item using a pull-down menu, and each type included in the predetermined item "equipment type" can be selected from the item. The input area 6010b is an input area for inputting text information. This input area 6010b is used to input text information with a certain number of characters, such as a product model number.
[0119] The input area 6010c is an input area in which predetermined items are listed and in which a check mark is entered in the check box of an appropriate item. A plurality of items can be specified in the input area 6010c. Furthermore, the input area 6010c displays items corresponding to the content selected in the input area 6010a, for example. In this example, the input area 6010c displays items corresponding to "Guide Light" selected in "Facility Type" in the input area 6010a.
[0120] The input field 6010d is an input field for inputting, for example, remarks, and allows input of text information in a free format. The number of characters is unlimited, or a sentence of a certain length can be input. The UI unit 113 stores the content input in the annotation input field 6010 in the RAM 1002 in association with position information indicating the position specified in step S103.
[0121] Note that even while annotation input screen 600 is being displayed in step S111, it is possible to operate buttons 5020 to 5023, tabs 601a to 601c, and menu buttons 5050 and 5051 arranged on screen 500. In addition, it is also possible to change the position within spherical image 510 of region 511 into which image 5010 is cut out.
[0122] Furthermore, while the annotation input screen 600 is being displayed in step S111, an operation for specifying the next position on the image 5010 is also possible. In step S150, the UI unit 113 determines whether or not the next position on the image 5010 has been specified. If the UI unit 113 determines that the next position has been specified (step S150, "Yes"), the process returns to step S102.
[0123] On the other hand, if the UI unit 113 determines that the next position has not been specified (step S150, "No"), the process returns to step S104. In this case, the state immediately before the determination in step S150 is maintained.
[0124] If tab 601b is selected in step S104, UI unit 113 determines that detailed image insertion is to be performed, and proceeds to step S120. In the next step S121, UI unit 113 displays an image file selection screen on screen 500. For example, UI unit 113 can use a file selection screen provided as a standard function by the OS (Operating System) installed in information processing device 100a as the image file selection screen. At this time, it is preferable to filter file names based on, for example, file extensions so that only the file names of corresponding image files are displayed.
[0125] In the next step S122, the UI unit 113 reads the image file selected on the image file selection screen. In the next step S123, the UI unit 113 displays the image of the image file read in step S122 in a predetermined area on the annotation input screen 600. The UI unit 113 associates the file name, including the path, of the read image file with location information indicating the location specified in step S103, and stores the file name in, for example, RAM 1002. After the processing of step S123, the processing proceeds to the above-mentioned step S150.
[0126] When tab 601c is selected in step S104, UI unit 113 determines that a cutout region is to be selected for image 5010, and proceeds to step S130. In the next step S131, UI unit 113 switches the display of annotation input screen 600 to a region designation screen for designating a cutout region as a diagnostic region including the image to be diagnosed. Fig. 21 is a diagram showing an example of annotation input screen 600 switched to the region designation screen, which is applicable to the first embodiment.
[0127] 21 , the annotation input screen 600 switched to the area designation screen includes a cut image display area 6020, a check box 6021, and a selection button 6022. The cut image display area 6020 displays a cut image cut out from the area designated for the image 5010. In the initial state where no area is designated for the image 5010, the cut image display area 6020 displays, for example, a blank space. However, the present invention is not limited to this, and an image of an area preset for the position designated in step S103 may be cut out from the image 5010 and displayed in the cut image display area 6020.
[0128] The check box 6021 is a button for displaying a marker image as an icon in the cut image display area 6020. The select button 6022 is a button for starting to specify a cut area for the image 5010.
[0129] Fig. 22 is a diagram showing an example of the display of a screen 500 for specifying a cutout area in response to an operation on a selection button 6022 according to the first embodiment. As shown in Fig. 22, the UI unit 113 displays a frame 603 indicating the cutout area on the screen 500 in response to an operation on the selection button 6022, and also erases the display of the annotation input screen 600 from the screen 500. The frame 603 is displayed, for example, corresponding to an area set in advance for the position specified in step S103.
[0130] The UI unit 113 can change the size, shape, and position of the frame 603 in response to a user operation. The shape is limited to, for example, a rectangle, but the ratio of the lengths of the long and short sides can be changed. When the size, shape, and position of the frame 603 is changed, if the changed frame 603 does not include the position specified in step S103, the UI unit 113 can display a warning screen indicating this.
[0131] 16, the UI unit 113 determines whether the area specification is complete in step S132. For example, when the display device 1010 and the input device 1011 are configured as a touch panel 1020 and the area specification using the frame 603 is performed by a drag operation in which the user moves a finger while the finger is in contact with the touch panel 1020, it can be determined that the drag operation has ended and the area specification using the frame 603 has been completed when the finger is released from the touch panel 1020.
[0132] If the UI unit 113 determines that the area designation is not complete (step S132, "No"), the process returns to step S132. On the other hand, if the UI unit 113 determines that the area designation is complete (step S132, "Yes"), the process proceeds to step S133.
[0133] In step S133, the UI unit 113 acquires, as a cut-out image, an image of the area specified by the frame 603 from the image 5010. At this time, the UI unit 113 acquires coordinates in the spherical image 510 including the image 5010, for example, of each vertex of the frame 603, and stores the acquired coordinates in, for example, the RAM 1002 in association with position information indicating the position specified in step S103.
[0134] In the next step S134, the UI unit 113 displays the cropped image acquired in step S133 in the cropped image display area 6020 on the annotation input screen 600. The UI unit 113 also stores the cropped image acquired in step S133 in a file with a predetermined file name to create a cropped image file. The UI unit 113 stores the cropped image file in, for example, the RAM 1002.
[0135] In the next step S135, the UI unit 113 determines whether or not to display an icon for the cut-out image displayed in the cut-out image display area 6020. If the check box 6021 on the annotation input screen 600 is checked, the UI unit 113 determines that an icon is to be displayed (step S135, "Yes"), and proceeds to step S136.
[0136] In step S136, UI unit 113 displays marker 602' at a position corresponding to the position specified in step S103 on the cropped image displayed in cropped image display area 6020, as exemplified in Fig. 23. In this example, marker 602' has the same shape as marker 602 displayed on image 5010, but this is not limited to this example, and marker 602' and marker 602 may have different shapes.
[0137] After the process of step S136, the process proceeds to the above-mentioned step S150. Furthermore, if it is determined in step S135 that the icon should not be displayed (step S135, "No"), the process also proceeds to step S150.
[0138] If, for example, button 5021 is operated in step S104, the UI unit 113 determines that the input contents on the annotation input screen 600 are to be saved, and causes the process to proceed to step S140. In step S140, the UI unit 113 instructs the incidental information generation unit 112 to save each piece of data input on the annotation input screen 600.
[0139] In the next step S141, the incidental information generation unit 112, in accordance with an instruction from the UI unit 113, creates annotation data based on the data stored in the RAM 1002 by the UI unit 113 in each of the above-mentioned processes (e.g., step S103, step S111, step S123, step S133, step S134, etc.). Annotation data is created for each position specified in step S103. The incidental information generation unit 112 stores the created annotation data in, for example, the storage 1004 of the information processing device 100a. Alternatively, the incidental information generation unit 112 may transmit the created annotation data to the server 6 via the network 5 and have the server 6 store the data.
[0140] Tables 2 to 7 show examples of the configuration of annotation data according to the first embodiment.
[0141] [Table 2]
[0142] [Table 3]
[0143] [Table 4]
[0144] [Table 5]
[0145] [Table 6]
[0146] [Table 7]
[0147] Table 2 shows the overall structure of annotation data. The annotation data includes shape data, attribute data, image region data, and an annotation data ID. Examples of the structure of this shape data, attribute data, and image region data are shown in Tables 3, 4, and 5, respectively.
[0148] An example of shape data is shown in Table 3. The shape data indicates the shape formed by the positions specified by the user in step S103 of Figure 16, and includes the items "primitive type," "number of vertices," and "list of vertex coordinates."
[0149] Table 6 shows example values defined for the "Primitive Type" item. In this example, the values "Point," "Line," and "Polygon" are defined for the "Primitive Type" item. If the value of the "Primitive Type" item is "Point," the value of the "Number of Vertices" item is set to "1." If the value of the "Primitive Type" item is "Line," the value of the "Number of Vertices" item is set to "2 or more." Also, if the value of the "Primitive Type" item is "Polygon," the value of the "Number of Vertices" item is set to "3 or more." The "Vertex Coordinate List" item defines each vertex shown in the "Number of Vertices" item using the (φ,θ) coordinate system. Here, the closed area enclosed by the straight lines connecting the coordinates listed in the "Vertex Coordinate List" item in order from the top becomes the range of the "Polygon" defined by the primitive type.
[0150] In the above example, one point is specified in step S103 of Fig. 16, so the item "primitive type" is set to the value "point" and the number of vertices is set to "1." Also, the item "vertex coordinate list" describes only the coordinates of the one specified point.
[0151] Table 4 shows an example of attribute data. The attribute data includes the following items: "creator," "creation date and time," "location," "subject," "source image," "investigation date and time," "response method," "attribute information list," and "tag."
[0152] The value of the item "Creator" is obtained, for example, from the login information for the information processing device 100a or an input program for realizing the functions of the information processing device 100a according to the first embodiment. The value of the item "Creation Date and Time" is obtained from the system time based on the clock of the information processing device 100a. The value of the item "Location" is obtained from the imaging position on the floor plan image 5011 if the floor plan image 5011 can be obtained. If the floor plan image 5011 cannot be obtained, the value of the item "Location" is obtained from user input.
[0153] The value of the "target" item is obtained from user input. The value of the "source image" item is the image name (file name) referenced when the annotation was created on the annotation input screen 600. For example, the file name of the cropped image file displayed in the cropped image display area 6020 described above is used as the value of the "source image" item. The value of the "investigation date and time" item is obtained from the timestamp of the image file described in the "source image" item. The value of the "response method" item is obtained from user input.
[0154] The value of the item "Attribute Information List" is described as a list of attribute information exemplified in Table 7. In Table 7, the attribute information includes the items "Type," "Name," and "Attribute Value List." The value of the item "Type" is the type of diagnostic object, such as equipment or abnormality, and is, for example, the value of input field 6010a on annotation input screen 600 shown in FIG. 20, where the equipment type is entered. The value of the item "Name" is the specific name of the diagnostic object, and is, for example, the value of input field 6010b on annotation input screen 600 shown in FIG. 20, where the product model number is entered. The value of the item "Attribute Value List" is, for example, a list of pairs of attribute names and attribute values, where the names corresponding to each checkbox in input field 6010c on annotation input screen 600 shown in FIG. 20 are the attribute names and the values of the checkboxes are the attribute values. The attribute names included in the item "Attribute Value List" vary depending on the value of the item "Type."
[0155] Returning to the explanation of Table 4, the value of the item "tag" is acquired from a user input. For example, the value (for example, a note) entered in the input field 6010d on the annotation input screen 600 is used as the value of the item "tag."
[0156] Table 5 shows an example of image region data. The image region data indicates, for example, the coordinates of the top left, bottom left, top right, and bottom right vertices of the cropped image specified in step S132 of FIG. 16 by angular coordinates (φ, θ) in spherical image 510.
[0157] 16 , when the auxiliary information generation unit 112 completes saving the annotation data in step S141, the auxiliary information generation unit 112 causes the process to proceed to step S142. Here, the auxiliary information generation unit 112 associates, with the annotation data, information indicating each of the omnidirectional images 510 acquired when the annotation was input (for example, a file name including a path). Taking FIG. 15-1 as an example, the auxiliary information generation unit 112 associates, with the annotation data, information indicating each of the omnidirectional images 510, 510, ... corresponding to each of the pin markers 50401, 50402, 50403, ....
[0158] In step S142, the UI unit 113 determines whether an instruction to end the annotation input process by the input program has been issued. For example, after the save process in step S141, the UI unit 113 displays an end instruction screen for instructing whether to end or continue the annotation input process. If the UI unit 113 determines that an instruction to end the annotation input process has been issued (step S142, "Yes"), it terminates the series of processes according to the flowchart in Fig. 16. On the other hand, if the UI unit 113 determines that an instruction to continue the annotation input process has been issued (step S142, "No"), it transitions the process to the above-mentioned step S150.
[0159] [Output processing according to the first embodiment] Next, a description will be given of the output process according to the first embodiment. In the information processing device 100a, the output unit 115 creates report data that summarizes the diagnostic results of the diagnostic target, based on the annotation data created as described above.
[0160] As an example, in the information processing device 100a, the UI unit 113 reads each piece of annotation data and each of the spherical images 510, 510, ... associated with the annotation data, for example, in response to an operation of a button 5020 on the screen 500. The UI unit 113 cuts out a predetermined area 511 of one of the read spherical images 510, 510, ..., and displays it as an image 5010 on the screen 500. At this time, the UI unit 113 displays, for example, a tag corresponding to each piece of annotation data at a position on the image 5010 that corresponds to the value of the item "list of vertex coordinates" in the shape data of each piece of annotation data.
[0161] Fig. 24 is a diagram showing an example of the display of a screen 500 related to report data creation according to the first embodiment. In Fig. 24, tags 5030c and 5030d are displayed on the screen 500 according to the item "tag" included in the attribute data of each annotation data, for example. Corresponding to these tags 5030c and 5030d, values (e.g., notes) entered in the item "tag" are displayed as comments 604a and 604b. In response to the specifications for these tags 5030c and 5030d, the output unit 115 creates report data based on the annotation data corresponding to the specified tags.
[0162] Without being limited to this, the output unit 115 can also perform a narrowed search for each item of attribute data of each annotation data using conditions specified by the user, and create report data based on each annotation data obtained as a result of the search.
[0163] 25 is a diagram illustrating an example of report data applicable to the first embodiment. In this example, the report data includes, for each piece of annotation data, a record including the items of "inspection date" and "inspector name," and the items of "room name," "target," "site photo," "content," "necessity of action," and "remarks."
[0164] In the report data record, the items "Inspection Date" and "Inspector Name" can be obtained from the items "Creation Date and Time" and "Creator" in the attribute data of the annotation data, for example, by referring to Table 4 above. The item "Object" in the report data can be obtained from the item "Name" in the item "Attribute Information List" in the attribute data of the annotation data, for example, by referring to Tables 4 and 7. Furthermore, the items "Room Name," "Content," and "Response Necessity" in the report data can be obtained from the items "Location," "Tag," and "Response Method" in the attribute data of the annotation data, respectively. Furthermore, the item "Current Status Photo" in the report data embeds an image obtained based on the image name described in the item "Source Image" in the attribute data of the annotation data, for example, by referring to Table 4.
[0165] In the report data, the item "Notes" can be obtained, for example, according to a user input when the report data is created.
[0166] 25 is an example, and the items included in the record are not limited to the above example. The items included in the record can be changed and set according to user input.
[0167] The output unit 115 outputs the report data created in this manner in a predetermined data format. For example, the output unit 115 can output the report data in a data format such as a commercially available document creation application program, a table creation application program, or a presentation material creation application program. The output unit 115 can also output the report data in a data format specialized for printing and display, such as PDF (Portable Document Format).
[0168] Note that the report data shown in FIG. 25 is not limited to being output by the output unit 115. For example, an image embedded in the "site photo" item can be associated with link information that calls up a panoramic automatic tour function (see FIG. 15-2) that starts from the image capture position corresponding to the image. The UI unit 113 displays an image representing the report data shown in FIG. 25 on the screen 500. When an image to be displayed in the "site photo" item is specified by a user operation, the UI unit 113 switches the display on the screen 500 to, for example, the display shown in FIG. 15-2(a) in accordance with the link information associated with the image.
[0169] Alternatively, it is also possible to associate link information to a specific website with the image embedded in the "site photo" item.
[0170] Furthermore, in the report data, an item (referred to as "Time Series Photo" item) for embedding time-series images can be provided instead of the "Current Status Photo" item. Time-series images, for example, images in which multiple images taken at different times in a certain location and imaging range are arranged in chronological order, can be captured in advance and stored, for example, in the information processing device 100a in association with location information indicating the location. When displaying a screen based on the report data, the output unit 115 displays time-series images containing multiple images in the "Time Series Photo" item. This makes it possible to easily grasp the progress of work, such as construction work. Furthermore, by comparing images taken at different times, changes over time can be easily confirmed.
[0171] 24, annotation data used to create report data is specified by specifying each tag 5030c and 5030d displayed on screen 500, but this is not limited to this example. For example, it is also possible to create and display a list including thumbnail images, comments, attribute information, etc. for each piece of annotation data. Thumbnail images can be used by reducing the size of the cut image.
[0172] Furthermore, comments by people other than the report creator (inspector) can be associated with the report data. This allows for a thread function, such as a designer replying to the inspector's report data, and the inspector then replying with a comment to that reply.
[0173] [Modification of the first embodiment] Next, a modified example of the first embodiment will be described. In the first embodiment described above, as explained using Figures 20 to 23, etc., the position was specified by a point in step S103 of Figure 16. This is not limited to this example, and for example, as explained in the item "Primitive type" in Tables 3 and 6, the position can also be specified by a line or a polygon.
[0174] An example of specifying a position using a line in step S103 in Fig. 16 will be described below. Fig. 26 is a diagram showing an example in which a crack 610 is observed on a wall surface, which is applicable to a modified example of the first embodiment. In Fig. 26 and Figs. 27 to 29 described below, parts corresponding to Figs. 20 to 23 above are given the same reference numerals, and detailed description thereof will be omitted.
[0175] 26, a linear crack 610 is observed in image 5010 displayed on screen 500. In step S103 of FIG. 16, the user specifies the crack 610 by, for example, tracing (dragging) the image of the crack 610 on image 5010 on touch panel 1020. The position information representing the crack 610 is, for example, a collection of position information for multiple points where the positions of successively adjacent points are smaller than a threshold value.
[0176] Fig. 27 corresponds to Fig. 20 described above and is a diagram showing an example of an annotation input screen 600 for inputting an annotation for a specified crack 610 according to a modified example of the first embodiment. In Fig. 27, the annotation input screen 600 includes an annotation input area 6010' for inputting an annotation. In this example, the annotation input area 6010' includes input areas 6010d and 6010e.
[0177] The input area 6010e is an area for inputting information about the state of deformation such as crack 610, and includes the input items "Location," "Type of Deformation," "Width," "Length," and "Status." Of these, the items "Type of Deformation" and "Status" are input areas for selecting predetermined items using pull-down menus. Furthermore, the user inputs names and values for the items "Location" and the items "Width" and "Length." The items "Width" and "Length" are not fixed items, but can be items that correspond to the content selected in the item "Type of Deformation."
[0178] 27, a button 6012 is a button for adding a deformation type. In response to an operation on the button 6012, the UI unit 113 adds and displays, for example, a set of the above-mentioned item "deformation type," the items "width" and "length," and the item "status" to the annotation input screen 600.
[0179] The designation of the cutout area for this linearly designated position can be performed in the same manner as described in the first embodiment using Figures 21 and 22. Figure 28 is a diagram showing an example of an annotation input screen 600 that has been switched to an area designation screen, which is applicable to a modified example of the first embodiment. In this example, a crack 610' corresponding to the crack 610 on the image 5010 is displayed in the cutout image display area 6020. Other than that, there are no differences from Figure 21 described above.
[0180] 29 is a diagram showing an example of the display of a screen 500 for specifying a cutout area in response to an operation on a selection button 6022, according to a modification of the first embodiment. In this case as well, a frame 603 is specified so as to include the crack 610. If part or all of the crack 610 is not included in the frame 603, the UI unit 113 can display a warning screen indicating this.
[0181] Similarly, it is also possible to specify the position and cutout area for polygons with a size greater than a triangle (including concave polygons).
[0182] [Second embodiment] Next, a second embodiment will be described. In the first embodiment described above, annotation data is created based on a spherical image 510 having two-dimensional coordinate information. In contrast, in the second embodiment, a three-dimensional image having three-dimensional coordinate information is further used when creating annotation data.
[0183] [Imaging device according to the second embodiment] Fig. 30 is a diagram schematically showing an imaging device according to a second embodiment. In Fig. 30, parts common to those in Fig. 1 described above are assigned the same reference numerals, and detailed description thereof will be omitted. In Fig. 30, imaging device 1b has a first surface of a substantially rectangular parallelepiped housing 10b, on which are provided a plurality of (five in this example) imaging lenses 20a1, 20a2, 20a3, 20a4, and 20a5, and a shutter button 30. Within housing 10b, imaging elements are provided corresponding to the imaging lenses 20a1, 20a2, ..., 20a5, respectively.
[0184] Furthermore, a plurality of imaging lenses 20b1, 20b2, 20b3, 20b4, and 20b5 are provided on a second surface of housing 10b behind the first surface. Similar to the above-described imaging lenses 20a1, 20a2, ..., 20a5, imaging elements corresponding to these imaging lenses 20b1, 20b2, ..., 20b4, and 20b5 are provided within housing 10b.
[0185] , 20a5 and 20b1, 20b2, . . . , 20b5, the pairs of which have the same height from the bottom surface of the housing 10b constitute imaging bodies 211, 212, 213, 214, and 215, respectively.
[0186] Note that the same configurations as the imaging lenses 20a and 20b described in the first embodiment can be applied to the imaging lenses 20a1, 20a2, ..., 20a5 and the imaging lenses 20b1, 20b2, ..., 20b5, and therefore detailed description thereof will be omitted here. Also, the imaging bodies 211, 212, 213, 214, and 215 correspond to the imaging body 21 described above.
[0187] In the second embodiment, imaging lenses 20a1, 20a2, ..., 20a5 are arranged at equal intervals, with the distance between each adjacent imaging lens being distance d. Furthermore, imaging lenses 20a1 and 20b1, imaging lenses 20a2 and 20b2, imaging lenses 20a3 and 20b3, imaging lenses 20a4 and 20b4, and imaging lenses 20a5 and 20b5 are arranged in housing 10b so that their heights from, for example, the bottom surface of housing 10b are the same.
[0188] Furthermore, imaging lenses 20a5 and 20b5 of imaging body 215, which is arranged at the bottom among imaging bodies 211, 212, 213, 214, and 215, are each arranged at a height h from the bottom surface. Also, for example, imaging lenses 20a1 to 20a5 are arranged such that, starting from the lowest imaging lens 20a5, imaging lenses 20a4, 20a3, 20a2, and 20a1 are spaced apart by a distance d and extend from the bottom surface side to the top surface side of housing 10a, with the lens centers aligned with the center line of housing 10a in the long side direction.
[0189] Shutter button 30 is a button for instructing imaging by each of imaging lenses 20a1, 20a2, ..., 20a5 and each of imaging lenses 20b1, 20b2, ..., 20b5 in response to operation. When shutter button 30 is operated, imaging is synchronously performed by each of imaging lenses 20a1, 20a2, ..., 20a5 and each of imaging lenses 20b1, 20b2, ..., 20b5.
[0190] 30, the housing 10b of the imaging device 1b is configured to include an imaging section area 2b in which the imaging bodies 211 to 215 are arranged, and an operation section area 3b in which the shutter button 30 is arranged. The operation section area 3b is provided with a grip section 31 that allows the user to hold the imaging device 1b, and a fixing section 32 is provided on the bottom surface of the grip section 31 for fixing the imaging device 1b to a tripod or the like.
[0191] Although the imaging device 1b has been described here as including five imaging bodies 211 to 215, this is not limited to this example. That is, the imaging device 1b may include six or more imaging bodies 21, or two to four imaging bodies 21, as long as it includes a plurality of imaging bodies 21.
[0192] 31 is a diagram showing an example of an imaging range that can be imaged by each of the imaging bodies 211, 212, ..., 215 that can be applied to the second embodiment. Each of the imaging bodies 211, 212, ..., 215 has a similar imaging range. In Fig. 31, the imaging ranges of the imaging bodies 211, 212, ..., 215 are shown as a representative by the imaging range of the imaging body 211.
[0193] 31, the Z axis is defined as the direction in which imaging lenses 20a1, 20a2, ..., 20a5 are aligned, and the X axis is defined as the direction of the optical axes of imaging lenses 20a1 and 20b1. The Y axis is defined as the direction that is included in a plane that intersects the Z axis at a right angle and also intersects the X axis at a right angle.
[0194] The imaging body 211, by combining the imaging lenses 20a1 and 20b1, has an imaging range that covers the entire celestial sphere with the center of the imaging body 211 as the center. That is, as described above, the imaging lenses 20a1 and 20b1 each have an angle of view of 180° or more, preferably greater than 180°, and more preferably greater than 185°. Therefore, by combining the imaging lenses 20a1 and 20b1, the imaging range A on the XY plane and the imaging range B on the XZ plane can each be 360°, and this combination realizes an imaging range of the entire celestial sphere.
[0195] Furthermore, the imaging bodies 211, 212, ..., 215 are arranged at intervals of a distance d in the Z-axis direction. Therefore, each set of hemispherical images captured by the imaging bodies 211, 212, ..., 215 with the entire celestial sphere as their imaging range is an image with a different viewpoint in the Z-axis direction by the distance d.
[0196] At this time, in the embodiment, imaging at each of the imaging lenses 20a1 to 20a5 and each of the imaging lenses 20b1 to 20b5 is performed synchronously in response to the operation of the shutter button 30. Therefore, by using the imaging device 1b according to the second embodiment, it is possible to obtain five sets of hemispherical images of the first and second surfaces of the casing 10b, each of which is captured at the same timing and from viewpoints that differ by a distance d in the Z-axis direction.
[0197] The five celestial sphere images generated from the five sets of hemispherical images captured at the same timing with different viewpoints at a distance d in the Z-axis direction are images aligned along the same epipolar line extending in the Z-axis direction.
[0198] [Outline of image processing according to the second embodiment] Next, image processing according to the second embodiment will be described. Fig. 32 is a functional block diagram of an example for describing the function of an information processing device 100b according to the second embodiment as an input device for inputting annotations. Note that the hardware configuration of the information processing device 100b can be directly applied to the configuration of the information processing device 100a described using Fig. 6, so detailed description here will be omitted. Also, in Fig. 32, parts common to those in Fig. 7 above will be assigned the same reference numerals, and detailed description will be omitted.
[0199] 32, like the information processing device 100a described above, the information processing device 100b includes an image acquisition unit 110, an image processing unit 111, an auxiliary information generation unit 112, a UI unit 113, a communication unit 114, and an output unit 115. Furthermore, the information processing device 100b includes a 3D information generation unit 120.
[0200] Of these, image acquisition unit 110 acquires each hemispherical image captured by each of imaging lenses 20a1-20a5 and each of imaging lenses 20b1-20b5 of imaging device 1b. Image processing unit 111 generates five omnidirectional images corresponding to imaging bodies 211-215, each with a different viewpoint in the Z-axis direction by a distance d, from each hemispherical image acquired by image acquisition unit 110, by the processing described using FIGS. 8-12 and equations (1)-(9).
[0201] Note that, since the image capture device 1b according to the second embodiment is assumed to be installed so that the center line connecting the centers of the imaging lenses 20a1-20a5 is parallel to the vertical direction when capturing images, tilt correction processing for each hemispherical image can be omitted. Alternatively, the image capture device 1b may be provided with the acceleration sensor 2008 as described above, and the vertical direction may be detected based on the detection result of the acceleration sensor 2008, and the tilt of the image capture device 1b with respect to the vertical direction may be obtained and tilt correction performed.
[0202] 18, 20 to 23, and 26 to 29 using one of the five celestial sphere images corresponding to the imaging bodies 211 to 215, for example, a celestial sphere image generated from a set of hemispherical images captured by the imaging body 211. The supplementary information generation unit 112 also generates annotation data based on the celestial sphere image generated from the set of hemispherical images captured by the imaging body 211.
[0203] The 3D information generator 120 generates three-dimensional information using five spherical images generated by the image processor 111, each having a different viewpoint in the Z-axis direction at a distance d.
[0204] [3D information generation process applicable to the second embodiment] Next, a three-dimensional information generation process that is executed by the 3D information generation unit 120 and that can be applied to the second embodiment will be described.
[0205] Fig. 33 is a diagram showing examples of images captured from five different viewpoints by imaging device 1b, which is applicable to the second embodiment, and synthesized for each of imaging bodies 211, 212, ..., 215. Fig. 33(a) shows an example of subject 60, and Figs. 33(b), 33(c), 33(d), 33(e), and 33(f) show examples of omnidirectional images 3001, 3002, 3003, 3004, and 3005, respectively, in which captured images of the same subject 60 captured from five different viewpoints are synthesized. As shown in Figs. 33(b) to 33(f), each of omnidirectional images 3001 to 3005 includes an image of subject 60 that is shifted slightly depending on the distance d between each of imaging bodies 211 to 215.
[0206] 33 shows that imaging device 1b captures an image of subject 60 present on the first surface (front surface) side of imaging device 1b, but this is for the sake of explanation, and in reality, imaging device 1b can capture images of subject 60 that surround imaging device 1b. In this case, spherical images 3001 to 3005 are, for example, images based on equirectangular projection, and each image has its left and right sides representing the same position, and its top and bottom sides each representing a single point. That is, spherical images 3001 to 3005 shown in FIG. 33 are images that have been converted and partially cropped from images based on equirectangular projection for the sake of explanation.
[0207] The projection method of the spherical images 3001 to 3005 is not limited to equirectangular projection. For example, the spherical images 3001 to 3005 may be images using cylindrical projection if there is no need to have a large angle of view in the Z-axis direction.
[0208] 34 is a flowchart illustrating an example of a process for creating a 3D reconstruction model applicable to the second embodiment. Each process in this flowchart is executed by the information processing device 100b. It is also assumed that the imaging device 1b already has 10 hemispherical images captured by the imaging bodies 211 to 215 stored in its built-in memory.
[0209] In step S10, the image acquisition unit 110 acquires from the imaging device 1b each of the hemispherical images captured by the imaging bodies 211 to 215. The image processing unit 111 combines the acquired hemispherical images for each of the imaging bodies 211 to 215 to create five omnidirectional images 3001 to 3005 captured from multiple viewpoints, as shown in FIG.
[0210] In the next step S11, the 3D information generation unit 120 selects one of the omnidirectional images 3001 to 3005 created in step S10 as a reference omnidirectional image (hereinafter referred to as omnidirectional image 3001). The 3D information generation unit 120 calculates the parallax of the other omnidirectional images 3002 to 3005 with respect to the selected reference omnidirectional image for all pixels of the images.
[0211] Here, the principle of a parallax calculation method applicable to the second embodiment will be described. The basic principle of performing parallax calculation using captured images captured by an image sensor such as the image sensor 200a is a method using triangulation. Triangulation will be described with reference to FIG. 35. In FIG. 35, cameras 400a and 400b include lenses 401a and 401b and image sensors 402a and 402b, respectively. Using triangulation, a distance D from a line connecting the cameras 400a and 400b to an object 403 is calculated from imaging position information in the captured images captured by each of the image sensors 402a and 402b.
[0212] In FIG. 35, the value f indicates the focal length of each of lenses 401a and 401b. The length of the line connecting the optical axis centers of lenses 401a and 401b is defined as base line length B. In the example of FIG. 30, the distance d between each of imaging bodies 211 to 215 corresponds to base line length B. The difference between imaging positions i1 and i2 of object 403 on imaging elements 402a and 402b is parallax q. Because the similarity relationship of triangles holds, D:f=B:q, and therefore distance D can be calculated using equation (10).
[0213]
number
[0214] In equation (10), since the focal length f and the base length B are known, the task of the process is to calculate the parallax q. Since the parallax q is the difference between the imaging positions i1 and i2, the fundamental task of parallax calculation is to detect the correspondence between the imaging positions in the images captured by the imaging elements 402a and 402b. Generally, this matching process of finding corresponding positions between multiple images is realized by searching for each parallax on the epipolar line based on the epipolar constraint.
[0215] The disparity search process can be realized using various calculation methods. For example, a block matching process using a normalized cross-correlation coefficient (NCC) shown in equation (11) can be applied. Alternatively, a high-density disparity calculation process using semi-global matching (SGM) can also be applied. The method to be used for calculating the disparity can be selected appropriately depending on the application of the 3D reconstruction model to be finally generated. In equation (11), the value p represents the pixel position, and the value q represents the disparity.
[0216]
number
[0217] Based on these cost functions, the correspondence between each pixel on the epipolar line is calculated, and the calculation result that is considered to be the most similar is selected. In NCC, i.e., equation (11), the numerical value C(p,q) NCC The pixel position with the maximum cost can be regarded as a corresponding point. In SGM, the pixel position with the minimum cost is regarded as a corresponding point.
[0218] As an example, a case where disparity is calculated using the NCC of equation (11) will be outlined below. The block matching method acquires pixel values of an area cut out as an M-pixel by N-pixel block centered on an arbitrary reference pixel in a reference image, and pixel values of an area cut out as an M-pixel by N-pixel block centered on an arbitrary target pixel in a target image. Based on the acquired pixel values, the similarity between the area including the reference pixel and the area including the target pixel is calculated. The similarity is compared while moving the M-pixel by N-pixel block within the image to be searched, and the target pixel in the block at the position with the highest similarity is set as the corresponding pixel for the reference pixel.
[0219] In equation (11), the value I(i,j) represents the pixel value of a pixel in a pixel block in the reference image, and the value T(i,j) represents the pixel value of a pixel in a pixel block in the target image. The calculation of equation (11) is performed while moving the pixel block in the target image, which corresponds to the pixel block of M pixels by N pixels in the reference image, pixel by pixel, to obtain the numerical value C(p,q). NCC The pixel position where is the maximum value is searched for.
[0220] Even when the imaging device 1b according to the second embodiment is used, the parallax is basically calculated using the above-described principle of triangulation. Here, the imaging device 1b includes five imaging bodies 211 to 215 and is capable of capturing five omnidirectional images 3001 to 3005 at a time. That is, the imaging device 1b according to the second embodiment can capture three or more captured images simultaneously. Therefore, the above-described principle of triangulation is applied in an expanded form to the second embodiment.
[0221] For example, as shown in equation (12), the total sum of disparities q of the costs for each camera separated by a base line length B can be used as the cost to detect corresponding points in each image captured by each camera.
[0222]
number
[0223] As an example, let us assume that the first, second, and third cameras are arranged on the epipolar line in the order of the first camera, the second camera, and the third camera. In this case, cost calculations are performed using the above-mentioned NCC or SGM for each of the pair of the first and second cameras, the pair of the first and third cameras, and the pair of the second and third cameras. The final distance D to the target can be calculated by summing up the costs calculated for each of these camera pairs and finding the minimum value of the sum.
[0224] The matching process described here can also be applied to the matching process in the automatic imaging position estimation function described with reference to FIG. 15-1.
[0225] However, as a method for calculating disparity in the embodiment, a stereo image measurement method using EPI (Epipolar Plane Image) can be applied. For example, EPI can be created by regarding spherical images 3001 to 3005 based on captured images captured by imaging bodies 211 to 215 in imaging device 1b as the same images captured by a camera moving at a constant speed. By using EPI, it is possible to more easily search for corresponding points between spherical images 3001 to 3005 than, for example, the method using triangulation described above.
[0226] As an example, in the imaging device 1b, for example, the imaging body 211 is used as a reference, and the distances d between the imaging body 211 and each of the imaging bodies 212 to 215 are 1- 2, d 1-3 , d 1-4 and d 1-5 Based on the calculation results, the horizontal and vertical axes (x, y) of the spherical images 3001 to 3005 and the distance D (= 0, distance d 1-2 ,…,d 1-5 A three-dimensional spatial image is created. A cross-sectional image of this spatial image in the yD plane is created as an EPI.
[0227] In the EPI created as described above, points on an object present in each of the original spherical images 3001 to 3005 are represented as a single straight line. The slope of this line changes depending on the distance from the image capture device 1b to the point on the object. Therefore, by detecting the line included in the EPI, it is possible to determine corresponding points between each of the spherical images 3001 to 3005. Furthermore, the slope of the line can be used to calculate the distance from the image capture device 1b to the object corresponding to the line.
[0228] The principle of EPI will be explained using Figures 36 and 37. Figure 36(a) shows a set of multiple images 4201, 4202, ..., each of which is a cylindrical image. Figure 36(b) schematically shows an EPI 422 cut out from the set of images 4201, 4202, ..., along a plane 421. In the example of Figure 36(a), the shooting position axis of each image 4201, 4202, ... is taken in the depth direction, and the set of images 4201, 4202, ... are superimposed and converted into 3D data as shown in Figure 36(a). The set of images 4201, 4202, ..., converted into 3D data, is cut out along a plane 421 parallel to the depth direction to produce the EPI 422 shown in Figure 36(b).
[0229] In other words, the EPI 422 is an image obtained by extracting lines with the same X coordinate from each of the images 4201, 4202, . . . and arranging the extracted lines side by side with the X coordinates of the images 4201, 4202, . . . in which they are included.
[0230] Fig. 37 is a diagram showing the principle of EPI applicable to the second embodiment. Fig. 37(a) is a schematic diagram of Fig. 36(b) described above. In Fig. 37(a), the horizontal axis u is the depth direction in which the images 4201, 4202, ... are superimposed and represents parallax, and the vertical axis v corresponds to the vertical axis of each of the images 4201, 4202, .... EPI 422 refers to an image in which captured images are superimposed in the direction of the base length B.
[0231] This change in base length B is represented by distance ΔX in Figure 37(b). Note that in Figure 37(b), positions C1 and C2 correspond to the optical centers of lenses 401a and 401b, respectively, in Figure 8. Furthermore, positions u1 and u2 are positions relative to positions C1 and C2, respectively, and correspond to imaging positions i1 and i2, respectively, in Figure 8.
[0232] When the images 4201, 4202, ... are arranged in the direction of the base line length B in this way, the positions of corresponding points on the images 4201, 4202, ... are expressed as a straight line or curve with a slope m on the EPI 422. This slope m is the parallax q used when calculating the distance D. The slope m becomes smaller as the distance D becomes closer, and becomes larger as the distance D becomes farther. This straight line or curve with a slope m that varies depending on the distance D is called a feature point locus.
[0233] The gradient m is expressed by the following formula (13). In formula (13), the value Δu is the difference between positions u1 and u2, which are the imaging points in FIG. 37(b), and can be calculated by formula (14). From the gradient m, the distance D is calculated by formula (15). Here, in formulas (13) to (15), the value v represents the moving speed of the camera, and the value f represents the frame rate of the camera. In other words, formulas (13) to (15) are calculation formulas when each panoramic image is captured at a frame rate f while the camera is moving at a constant speed v.
[0234]
number
[0235]
number
[0236]
number
[0237] Note that when a panoramic image is used as the image constituting the EPI, the slope m is a value based on a curve. This will be described with reference to FIGS. 38 and 39. In FIG. 38, spheres 4111, 4112, and 4113 each have the structure of imaging body 21 and represent the panoramic image captured by camera #0, camera #ref, and camera #(n-1) arranged on a straight line. The distance (baseline length) between camera #0 and camera #ref is distance d2, and the distance between camera #ref and camera #(n-1) is distance d1. Hereinafter, spheres 4111, 4112, and 4113 will be referred to as panoramic images 4111, 4112, and 4113, respectively.
[0238] The imaging position of the target point P on the omnidirectional image 4111 is at an angle φ n-1 Similarly, the imaging positions of the target point P on the spherical images 4112 and 4113 are at angles φ ref and a position having an angle of φ0.
[0239] Figure 39 shows these angles φ0 and φ ref and φ n-1 39 is plotted on the vertical axis and the positions of each camera #0, #ref, and #(n-1) on the horizontal axis. As can be seen from Fig. 39, the feature point loci indicated as the imaging positions in each of the spherical images 4111, 4112, and 4113 and the positions of each camera #0, #ref, and #(n-1) are not straight lines, but are approximated by curve 413 based on straight lines 4121 and 4122 connecting the points.
[0240] When calculating the parallax q of the entire circumference using the omnidirectional images 3001 to 3005, a method of directly searching for corresponding points on the curve 413 from the omnidirectional images 3001 to 3005 as described above may be used, or a method of converting the omnidirectional images 3001 to 3005 into a projection system such as a pinhole and searching for corresponding points based on each converted image may be used.
[0241] 38, of the celestial sphere images 4111, 4112, and 4113, the celestial sphere image 4112 is set as a reference image (ref), the celestial sphere image 4111 is set as the (n-1)th target image, and the celestial sphere image 4113 is set as the 0th target image. Based on the celestial sphere image 4112 that is the reference image, the corresponding points between the celestial sphere images 4111 and 4113 are determined by a parallax q n-1 and q0. This disparity q n-1 and q0 can be obtained using various known methods, such as the above-mentioned equation (11).
[0242] Using EPI to create a 3D reconstruction model makes it possible to process a large number of panoramic images in a unified manner. In addition, using the gradient m makes the calculation more robust, as it does not just involve point-by-point correspondence.
[0243] Returning to the description of the flowchart in FIG. 34, after the disparity calculation in step S11, the 3D information generation unit 120 proceeds to step S12. In step S12, the 3D information generation unit 120 performs a correction process on the disparity information indicating the disparity calculated in step S11. As the correction process on the disparity information, correction based on the Manhattan-world hypothesis, line segment correction, or the like can be applied. In the next step S13, the 3D information generation unit 120 converts the disparity information corrected in step S12 into 3D point cloud information. In the next step S14, the 3D information generation unit 120 performs a smoothing process, a meshing process, or the like as necessary on the 3D point cloud information into which the disparity information has been converted in step S13. By the process up to step S14, a 3D reconstruction model based on each of the omnidirectional images 3001 to 3005 is generated.
[0244] The processes of steps S11 to S14 described above can be performed using open-source distributed SfM (Structure-from-Motion) software, MVS (Multi-View Stereo) software, etc. The input program operated on the information processing device 100b includes, for example, the SfM software, the MVS software, etc.
[0245] As explained with reference to Fig. 31, the imaging device 1b according to the second embodiment has the imaging bodies 211-215 arranged on the Z axis. Therefore, the distance from the imaging device 1b to each subject is calculated preferentially in the radial direction on a plane 40 shown in Fig. 40, which is perpendicular to the direction in which the imaging lenses 20a1-20a5 are aligned. The term "preferential" here refers to the ability to generate a 3D reconstruction model for the angle of view.
[0246] That is, in each direction on plane 40 (radial direction in FIG. 40), the angle of view of each of imaging bodies 211-215 can include the entire circumference of 360°, making it possible to calculate distances for the entire circumference. On the other hand, in the Z-axis direction, the overlapping portion of the angle of view (imaging range) of each of imaging bodies 211-215 becomes large. Therefore, in the Z-axis direction, parallax becomes small in the omnidirectional images 3001-3005 based on the captured images captured by each of imaging bodies 211-215 around the angle of view of 180°. Therefore, it is difficult to calculate distances for the entire circumference of 360° in the direction of a plane including the Z-axis.
[0247] Here, as illustrated in FIG. 41, a large space including large buildings 50, 50, ... is considered as a target for creating a 3D reconstruction model. In FIG. 41, the X-axis, Y-axis, and Z-axis coincide with the X-axis, Y-axis, and Z-axis shown in FIG. 31. In this case, modeling at the full angle of view (360°) is in the direction of a plane 40 represented by the XY axes. Therefore, it is preferable to arrange a plurality of imaging bodies 211 to 215 in an aligned manner in the Z-axis direction perpendicular to this plane 40.
[0248] [Configuration of signal processing in the imaging device according to the second embodiment] Next, a configuration related to signal processing of the imaging device 1b according to the second embodiment will be described. Fig. 42 is a block diagram showing an example of the configuration of the imaging device 1b according to the embodiment. In Fig. 42, parts corresponding to those in Fig. 30 described above are given the same reference numerals, and detailed description thereof will be omitted.
[0249] In FIG. 42, imaging device 1b includes imaging elements 200a1, 200a2, ..., 200a5, imaging elements 200b1, 200b2, ..., 200b5, driving units 210a1, 210a2, ..., 210a5, driving units 210b1, 210b2, ..., 210b5, buffer memories 211a1, 211a2, ..., 211a5, and buffer memories 211b1, 211b2, ..., 211b5.
[0250] Of these, imaging elements 200a1, 200a2, ..., 200a5, driving units 210a1, 210a2, ..., 210a5, and buffer memories 211a1, 211a2, ..., 211a5 are components corresponding to imaging lenses 20a1, 20a2, ..., 20a5, respectively, and are included in imaging bodies 211, 212, ..., 215. Note that, to avoid complexity, only imaging body 211 of imaging bodies 211 to 215 is shown in Fig. 42.
[0251] Similarly, the imaging elements 200b1, 200b2, ..., 200b5, the driving units 210b1, 210b2, ..., 210b5, and the buffer memories 211b1, 211b2, ..., 211b5 are components corresponding to the imaging lenses 20b1, 20b2, ..., 20b5, respectively, and are included in the imaging bodies 211, 212, ..., 215, respectively.
[0252] The imaging device 1b further includes a control unit 220, a memory 221, and a switch (SW) 222. The switch 222 corresponds to the shutter button 30 shown in Fig. 30. For example, the state in which the switch 222 is closed corresponds to the state in which the shutter button 30 is operated.
[0253] A description will now be given of the imaging body 211. The imaging body 211 includes an imaging element 200a1, a driving section 210a1, and a buffer memory 211a1, and an imaging element 200b1, a driving section 210b1, and a buffer memory 211b1.
[0254] The driving unit 210a1 and the imaging element 200a1, as well as the driving unit 210b1 and the imaging element 200b1, are equivalent to the driving unit 210a and the imaging element 200a, as well as the driving unit 210b and the imaging element 200b described using Figure 4, so detailed description here will be omitted.
[0255] The buffer memory 211a1 is a memory capable of storing at least one frame of captured image, and the captured image output from the driving section 210a1 is temporarily stored in the buffer memory 211a1.
[0256] It should be noted that the other imaging bodies 212 to 215 have the same functions as the imaging body 211, and therefore their explanations will be omitted here.
[0257] The control unit 220 controls the overall operation of the imaging device 1b. When the control unit 220 detects a transition of the switch 222 from an open state to a closed state, the control unit 220 outputs a trigger signal. The trigger signal is simultaneously supplied to each of the drive units 210a1, 210a2, ..., 210a5 and each of the drive units 210b1, 210b2, ..., 210b5.
[0258] The memory 221 reads out each captured image from each buffer memory 211a1, 211a2, ..., 211a5 and each buffer memory 211b1, 211b2, ..., 211b5 under the control of the control unit 220 in response to the output of the trigger signal, and stores each read captured image. Each captured image stored in the memory 221 can be read out by the information processing device 100b connected to the imaging device 1b.
[0259] The battery 2020 is a secondary battery such as a lithium ion secondary battery, and serves as a power supply unit that supplies power to each unit within the imaging device 1b that requires power supply. The battery 2020 includes a charge / discharge control circuit that controls charging and discharging of the secondary battery.
[0260] Fig. 43 is a block diagram showing an example of the configuration of a control unit 220 and a memory 221 in an imaging device 1b according to the second embodiment. In Fig. 43, parts that are common to those in Fig. 4 above are given the same reference numerals, and detailed description thereof will be omitted.
[0261] 43, the control unit 220 includes a CPU 2000, a ROM 2001, a trigger I / F 2004, a switch (SW) circuit 2005, a data I / F 2006, and a communication I / F 2007, and these units are communicatively connected to a bus 2010. The memory 221 includes a RAM 2003 and a memory controller 2002, and the memory controller 2002 is further connected to the bus 2010. Power is supplied from a battery 2020 to each of the CPU 2000, the ROM 2001, the memory controller 2002, the RAM 2003, the trigger I / F 2004, the switch circuit 2005, the data I / F 2006, the communication I / F 2007, and the bus 2010.
[0262] The memory controller 2002, in accordance with instructions from the CPU 2000, controls the storage and reading of data into and from the RAM 2003. The memory controller 2002, in accordance with instructions from the CPU 2000, also controls the reading of captured images (hemispherical images) from each of the buffer memories 211a1, 211a2, ..., 211a5 and each of the buffer memories 211b1, 211b2, ..., 211b5.
[0263] Switch circuit 2005 detects the transition of switch 222 between the closed state and the open state, and passes the detection result to CPU 2000. When CPU 2000 receives the detection result from switch circuit 2005 that switch 222 has transitioned from the open state to the closed state, it outputs a trigger signal. The trigger signal is output via trigger I / F 2004, branches, and is supplied to each of drive units 210a1, 210a2, ..., 210a5 and each of drive units 210b1, 210b2, ..., 210b5, respectively.
[0264] Although the CPU 2000 has been described as outputting a trigger signal in response to the detection result by the switch circuit 2005, this is not limited to this example. For example, the CPU 2000 may output a trigger signal in response to a signal supplied via the data I / F 2006 or the communication I / F 2007. Furthermore, the trigger I / F 2004 may generate a trigger signal in response to the detection result by the switch circuit 2005 and supply it to each of the drive units 210a1, 210a2, ..., 210a5 and each of the drive units 210b1, 210b2, ..., 210b5.
[0265] In this configuration, when control unit 220 detects a transition of switch 222 from an open state to a closed state, it generates and outputs a trigger signal. The trigger signal is supplied simultaneously to each of drive units 210a1, 210a2, ..., 210a5 and each of drive units 210b1, 210b2, ..., 210b5. Each of drive units 210a1, 210a2, ..., 210a5 and each of drive units 210b1, 210b2, ..., 210b5 reads out electric charges from each of image sensors 200a1, 200a2, ..., 200a5 and each of image sensors 200b1, 200b2, ..., 200b5 in synchronization with the supplied trigger signal.
[0266] Each driving unit 210a1, 210a2, ..., 210a5 and each driving unit 210b1, 210b2, ..., 210b5 converts the charges read out from each imaging element 200a1, 200a2, ..., 200a5 and each imaging element 200b1, 200b2, ..., 200b5 into a captured image, and stores the converted captured images in each buffer memory 211a1, 211a2, ..., 211a5 and each buffer memory 211b1, 211b2, ..., 211b5.
[0267] At a predetermined timing after output of the trigger signal, the control unit 220 instructs the memory 221 to read out the captured images from the buffer memories 211a1, 211a2, ..., 211a5 and the buffer memories 211b1, 211b2, ..., 211b5. In response to this instruction, the memory controller 2002 in the memory 221 reads out the captured images from the buffer memories 211a1, 211a2, ..., 211a5 and the buffer memories 211b1, 211b2, ..., 211b5, and stores the read captured images in a predetermined area of the RAM 2003.
[0268] When the information processing device 100b is connected to the imaging device 1b via, for example, the data I / F 2006, the information processing device 100b requests, via the data I / F 2006, reading of each captured image (hemispherical image) stored in the RAM 2003. In response to this request, the CPU 2000 instructs the memory controller 2002 to read each captured image from the RAM 2003. In response to this instruction, the memory controller 2002 reads each captured image from the RAM 2003 and transmits the read captured images to the information processing device 100b via the data I / F 2006. The information processing device 100b executes the process according to the flowchart of FIG. 34 based on each captured image transmitted from the imaging device 1b.
[0269] Fig. 43 shows an example of the configuration of a control unit 220 and a memory 221 in an imaging device 1b according to the second embodiment. In Fig. 43, parts that are common to those in Fig. 4 above are given the same reference numerals, and detailed description thereof will be omitted.
[0270] 43, the control unit 220 includes a CPU 2000, a ROM 2001, a trigger I / F 2004, a switch (SW) circuit 2005, a data I / F 2006, and a communication I / F 2007, and these units are communicatively connected to a bus 2010. The memory 221 includes a RAM 2003 and a memory controller 2002, and the memory controller 2002 is further connected to the bus 2010. Power is supplied from a battery 2020 to each of the CPU 2000, the ROM 2001, the memory controller 2002, the RAM 2003, the trigger I / F 2004, the switch circuit 2005, the data I / F 2006, the communication I / F 2007, and the bus 2010.
[0271] The memory controller 2002, in accordance with instructions from the CPU 2000, controls the storage and reading of data into and from the RAM 2003. The memory controller 2002, in accordance with instructions from the CPU 2000, also controls the reading of captured images (hemispherical images) from each of the buffer memories 211a1, 211a2, ..., 211a5 and each of the buffer memories 211b1, 211b2, ..., 211b5.
[0272] The switch circuit 2005 detects the transition of the switch 222 between the closed state and the open state, and passes the detection result to the CPU 2000. When the CPU 2000 receives the detection result from the switch circuit 2005 that the switch 222 has transitioned from the open state to the closed state, it outputs a trigger signal. The trigger signal is output via the trigger I / F 2004, branches, and is sent to each of the drive units 210a1, 210a2, ..., 210a5 and each of the drive units 210b 1 , 210b2, ..., 210b5, respectively.
[0273] Although the CPU 2000 has been described as outputting a trigger signal in response to the detection result by the switch circuit 2005, this is not limited to this example. For example, the CPU 2000 may output a trigger signal in response to a signal supplied via the data I / F 2006 or the communication I / F 2007. Furthermore, the trigger I / F 2004 may generate a trigger signal in response to the detection result by the switch circuit 2005 and supply it to each of the drive units 210a1, 210a2, ..., 210a5 and each of the drive units 210b1, 210b2, ..., 210b5.
[0274] In this configuration, when the control unit 220 detects a transition of the switch 222 from an open state to a closed state, the control unit 220 generates and outputs a trigger signal. The trigger signal is supplied simultaneously to the drive units 210a1, 210a2, ..., 210a5 and the drive units 210b1, 210b2, ..., 210b5. The drive units 210a1, 210a2, ..., 210a5 and the drive units 210b1, 210b2, ..., 210b5 synchronize with the supplied trigger signal to drive the image pickup elements 200a1, 200a2, ..., 200a5 and the image pickup elements 200b1, 200b 2 , ..., read out the charge from 200b5.
[0275] Each driving unit 210a1, 210a2, ..., 210a5 and each driving unit 210b1, 210b2, ..., 210b5 converts the charges read out from each imaging element 200a1, 200a2, ..., 200a5 and each imaging element 200b1, 200b2, ..., 200b5 into a captured image, and stores the converted captured images in each buffer memory 211a1, 211a2, ..., 211a5 and each buffer memory 211b1, 211b2, ..., 211b5.
[0276] At a predetermined timing after output of the trigger signal, the control unit 220 instructs the memory 221 to read out the captured images from the buffer memories 211a1, 211a2, ..., 211a5 and the buffer memories 211b1, 211b2, ..., 211b5. In response to this instruction, the memory controller 2002 in the memory 221 reads out the captured images from the buffer memories 211a1, 211a2, ..., 211a5 and the buffer memories 211b1, 211b2, ..., 211b5, and stores the read captured images in a predetermined area of the RAM 2003.
[0277] When the information processing device 100b is connected to the imaging device 1b via, for example, the data I / F 2006, the information processing device 100b requests, via the data I / F 2006, reading of each captured image (hemispherical image) stored in the RAM 2003. In response to this request, the CPU 2000 instructs the memory controller 2002 to read each captured image from the RAM 2003. In response to this instruction, the memory controller 2002 reads each captured image from the RAM 2003 and transmits the read captured images to the information processing device 100b via the data I / F 2006. The information processing device 100b executes the process according to the flowchart of FIG. 34 based on each captured image transmitted from the imaging device 1b.
[0278] In the imaging device 1b according to the second embodiment, the battery 2020 and the circuit section 2030 are provided inside the housing 10b, similarly to the imaging device 1a according to the first embodiment described with reference to FIG. 5. Of the battery 2020 and the circuit section 2030, at least the battery 2020 is fixed inside the housing 10b by fixing means such as adhesive or screws. The circuit section 2030 includes at least the elements of the control section 220 and the memory 221 described above. The control section 220 and the memory 221 are configured on, for example, one or more circuit boards. In the imaging device 1b, the battery 2020 and the circuit section 2030 are arranged on an extension of the row in which the imaging lenses 20a1, 20a2, ..., 20a5 are aligned.
[0279] Here, it has been described that the battery 2020 and the circuit unit 2030 are placed at the position, but this is not limited to this example. For example, if the circuit unit 2030 is sufficiently small, it is sufficient that at least the battery 2020 is placed at the position.
[0280] By arranging the battery 2020 and the circuit unit 2030 in this manner, the imaging lenses 20a1, 20a2, ..., 20a5 (and the imaging lenses 20b1, 20b2, ..., 20b 5 ) can be arranged, the width of the surfaces (front and rear) can be narrowed. This prevents a part of the housing 10b of the imaging device 1b from being included in each of the captured images captured by the imaging bodies 211, 212, ..., 215, and makes it possible to calculate the parallax with higher accuracy. For the same reason, it is preferable to make the width of the housing 10b of the imaging device 1b as narrow as possible. This also applies to the imaging device 1a according to the first embodiment.
[0281] [Annotation input method according to the second embodiment] Next, an annotation input method according to the second embodiment will be described. As described above, in the information processing device 100b, the UI unit 113 displays a screen for inputting an annotation using a celestial sphere image based on each hemispherical image captured by one of the five imaging bodies 211 to 215 of the imaging device 1b (for example, the imaging body 211). That is, the UI unit 113 cuts out an image of a partial area of the celestial sphere image and displays the cut-out image on the screen 500 as, for example, an image 5010 shown in FIG. 20 .
[0282] At this time, the UI unit 113 acquires the three-dimensional point cloud information generated according to the flowchart of Fig. 34 from the 3D information generation unit 120. The UI unit 113 can switch the display on the screen 500 between the image 5010 and the three-dimensional point cloud information in response to, for example, an operation on the menu button 5050. The UI unit 113 displays each point included in the three-dimensional point cloud information in a different color according to, for example, distance.
[0283] Also in the second embodiment, the annotation input process is performed according to the flowchart of Fig. 16. Here, the position designation in step S103 in the flowchart of Fig. 16 is performed using three-dimensional coordinates based on three-dimensional point group information.
[0284] Fig. 44 is a diagram illustrating position specification according to the second embodiment. As shown in Fig. 44(a), consider a case where objects 7001, 7002, and 7003 having a three-dimensional structure are arranged in a three-dimensional space represented by mutually orthogonal X-, Y-, and Z-axes, and the objects are captured by imaging device 1b. Three-dimensional point cloud information generated based on each celestial sphere image acquired by imaging device 1b includes, for example, three-dimensional information at each point on the surface of each of objects 7001 to 7003 facing the imaging position from the imaging position.
[0285] Fig. 44(b) shows an example of an image 5010 obtained by cutting out a part of the celestial sphere image captured by the imaging body 211 of the imaging device 1b in the state of Fig. 44(a). The UI unit 13 displays the image 5010 shown in Fig. 44(b) on the screen 500. It is assumed that the user has specified positions of points on objects 7001 and 7002, as shown by markers 602a and 602b, for example, in the image 5010 shown in Fig. 44(b).
[0286] Here, in the second embodiment, it is possible to acquire three-dimensional point cloud information of the space captured by the imaging device 1b, including the objects 7001 to 7003. Therefore, position specification is performed on positions indicated by the three-dimensional coordinates of the objects 7001 and 7002, as shown as markers 602a' and 602b' in FIG.
[0287] Furthermore, by specifying a position using three-dimensional coordinates based on three-dimensional point cloud information, it is possible to calculate the length of the line or the area of the polygon when specifying a position using a line or polygon, as described in the modified example of the first embodiment above.
[0288] Fig. 45 is a diagram showing an example in which a position is specified using a line according to the second embodiment. Similar to Fig. 44(a), Fig. 45 shows objects 7004 and 7005 having three-dimensional structures arranged in a three-dimensional space represented by mutually orthogonal X-, Y-, and Z-axes. Furthermore, in the cylindrical object 7004, a crack 610a is observed on the cylindrical surface in a direction circumferential around the cylindrical surface. In the rectangular parallelepiped (quadrature) object 7005, a crack 610b is observed across an edge.
[0289] As such, it is difficult to calculate the length of cracks 610a and 610b that have depth from a two-dimensional captured image. For example, the length of crack 610 included in image 5010 in FIG. 26 described above can be measured only if the wall surface on which crack 610 exists is parallel to the surface of image 5010. In the second embodiment, three-dimensional point cloud information of these objects 7004 and 7005 is acquired, so the lengths of cracks 610a and 610b that have depth can be easily calculated. The same applies to calculating the area and perimeter of a polygon.
[0290] Although the above-described embodiments are preferred examples of the present invention, the present invention is not limited to these, and various modifications can be made without departing from the spirit of the present invention. [Explanation of symbols]
[0291] 1a, 1b Imaging device 4 Buildings 6 Server 20a, 20a1, 20a2, 20a3, 20a4, 20a5, 20b, 20b1, 20b2, 20b3, 20b4, 20b5 imaging lenses 21,211,212,213,214,215 imaging body 30 Shutter button 100a, 100b Information processing device 110 Image acquisition unit 111 Image processing unit 112 Additional information generation unit 113 UI section 114 Communications Department 115 Output section 120 3D information generation section 200a, 200b, 200a1, 200a2, 200a5, 200b1, 200b2, 200b5 image sensor 210a, 210b, 210a1, 210a2, 210a5, 210b1, 210b2, 210b5 drive units 211a1, 211a2, 211a5, 211b1, 211b2, 211b5 buffer memory 3001,3002,3003,3004,3005,510 Spherical images 500 screens 511 area 600 Annotation input screen 601a, 601b, 601c tabs 602,602' marker 610,610a,610b Cracks 2002 Memory Controller 2003 RAM 2004 Trigger I / F 2005 Switch Circuit 5010, 5010a, 5010b, 5010c Images 5011 Floor plan image 5030a, 5030b, 5030c, 5030d Tags 50401, 50402, 50403, 5040 10 ,5040 11 ,5040 12 Pin Marker 6010 Annotation input area 6010a, 6010b, 6010c, 6010d, 6010e Input area [Prior art documents] [Patent documents]
[0292] [Patent Document 1] Japanese Patent Application Publication No. 2018-013878
Claims
1. An input device for inputting a diagnosis result of a diagnosis target in a structure, a display means for displaying a spherical image of the structure; a receiving unit that receives an input of a position indicating the diagnostic target on the spherical image and stores the received position information indicating the position in a storage unit; and The display means displaying the spherical image and a diagnostic information input screen that accepts input of diagnostic information including a diagnostic result of the diagnostic object on a single screen; The receiving means accepting input of the diagnostic information in accordance with the diagnostic information input screen, and storing the diagnostic information and the position information in the storage means in association with each other; 1. An input device comprising:
2. The display means An image showing the diagnostic target is superimposed on the spherical image in accordance with the position information and displayed.
2. The input device according to claim 1.
3. The receiving means Accepting an input of a diagnostic area including the diagnostic object in the spherical image, and storing the position information corresponding to the accepted diagnostic area and the diagnostic information for the diagnostic object in association with each other in the storage means.
3. The input device according to claim 1 or 2.
4. The display means An image of the diagnostic area in the spherical image is displayed on the diagnostic information input screen.
4. The input device according to claim 3.
5. The display means An image different from the spherical image is displayed on the diagnostic information input screen.
5. The input device according to claim 1, wherein the input device is a touch panel.
6. The display means The diagnostic information input screen is displayed superimposed on the spherical image.
2. The input device according to claim 1.
7. The receiving means receiving an input of a set of positions indicating the diagnostic target using a plurality of positions on the spherical image, and storing position information indicating the received set of positions in the storage means; 7. The input device according to claim 1, wherein the input device is a touch panel.
8. The display means A drawing showing the structure of the structure and the spherical image are displayed on one screen.
8. The input device according to claim 1, wherein the input device is a touch panel.
9. the spherical image is a three-dimensional image having three-dimensional information, The receiving means receiving an input of the position having three-dimensional information indicating the diagnostic target; 9. The input device according to claim 1, wherein the input device comprises: a first input unit;
10. An input method for an input device for inputting a diagnosis result of a diagnosis target in a structure, comprising: a display step of displaying a spherical image of the structure; a receiving step of receiving an input of a position indicating the diagnostic target on the spherical image and storing the received position information indicating the position in a storage unit; and The display step includes: displaying the spherical image and a diagnostic information input screen that accepts input of diagnostic information including a diagnostic result of the diagnostic object on a single screen; The receiving step includes: accepting input of the diagnostic information in accordance with the diagnostic information input screen, and storing the diagnostic information and the position information in the storage means in association with each other; 10. An input method for an input device comprising:
11. An input device for inputting a diagnosis result of a structure, a display control means for displaying a spherical image of the structure; a designation unit for designating a diagnostic target on the displayed spherical image; an input means for inputting diagnostic information for the diagnostic object; a storage means for storing the position of the diagnostic target designated on the spherical image and the diagnostic information in association with each other; have An input device characterized by:
12. An output device that outputs a diagnosis result of a diagnosis target in a structure, a diagnostic information acquiring means for acquiring, from a storage means in which position information indicating a position of the diagnostic target with respect to the celestial sphere image and diagnostic information including a diagnosis result of the diagnostic target are stored in association with each other, the position information and the diagnostic information associated with the position information; an output unit that outputs the diagnostic information acquired by the diagnostic information acquisition unit based on the location information associated with the diagnostic information; have An output device characterized by:
13. The diagnostic information acquisition means further acquiring, from the storage means, an image based on the omnidirectional image, which is associated with the position information and stored in the storage means; The output means The image acquired by the diagnostic information acquisition means is included in the diagnostic information and output based on the position information. The output device of claim 12.
14. a celestial sphere image acquisition means for acquiring the celestial sphere image; a display means for displaying the celestial sphere image acquired by the celestial sphere image acquisition means, and for superimposing information indicating the diagnostic information associated with the position information on the celestial sphere image at a position indicated by the position information on the celestial sphere image, and Further having 14. The output device according to claim 12 or 13.
15. a receiving unit configured to receive a designation input for designating information indicating the diagnostic information to be displayed in a superimposed manner on the spherical image; The output means The diagnostic information designated by the designated input received by the receiving means is output.
15. The output device according to claim 12.
16. An output method for an output device that outputs a diagnosis result of a diagnosis target in a structure, comprising: a diagnostic information acquiring step of acquiring, from a storage means in which position information indicating a position of the diagnostic target with respect to the omnidirectional image and diagnostic information including a diagnosis result of the diagnostic target are stored in association with each other, the position information and the diagnostic information associated with the position information; an output step of outputting the diagnostic information acquired by the diagnostic information acquisition step based on the position information associated with the diagnostic information; have 10. An output method for an output device comprising:
Citation Information
Patent Citations
Test object management support device, test object management support method and program
JP2012252499A
Inspection result output device, inspection result output method, and program
JP2016126769A
System, method, and program for supporting information sharing
JP2018013877A
Data acquisition and encoding process linking physical objects with virtual data for manufacturing, inspection, maintenance and repair
US20170199645A1
Fluid leakage data management apparatus and management system
WO2016056297A1