Three-dimensional data processing system and three-dimensional data processing method

By using an omnidirectional camera and correcting the 3D data with feature points and markers, the system addresses the challenges of absolute scale and coordinate system alignment in 3D data processing, achieving simplified and accurate corrections.

JP2025080543APending Publication Date: 2025-05-26KK TOSHIBA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023193761
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-26

AI Technical Summary

Technical Problem

Conventional 3D data processing using monocular cameras faces challenges in obtaining absolute scale and correcting the coordinate system to match the actual physical space.

Method used

The system employs an omnidirectional camera to capture images from multiple positions, extracts feature points, estimates camera position and orientation, and corrects the 3D data's coordinate system and scale to align with the actual physical space using markers or fixed objects.

Benefits of technology

This approach simplifies the correction process for the coordinate system and scale in 3D data processing, enabling more accurate and efficient conversion of 3D data to match the actual physical space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080543000001_ABST
    Figure 2025080543000001_ABST
Patent Text Reader

Abstract

To simplify the correction of a coordinate system and scale in three-dimensional data processing using an omnidirectional image captured by an omnidirectional camera.SOLUTION: A three-dimensional data processing system 1 is configured to: estimate the position and orientation of an omnidirectional camera 5 during capture from the correspondence between feature points and omnidirectional images; restore first-density point cloud data indicative of three-dimensional distribution of first-density feature points on the basis of the feature points and the position and orientation of the omnidirectional camera 5 during capture; perform a correction to align the coordinate system and scale of the first-density point cloud data with the absolute coordinate system and absolute scale of the real physical space 40 in which the images were captured by the omnidirectional camera 5 on the basis of at least one of the position and orientation of the omnidirectional camera 5 during capture and the information related to the feature points; and restore second-density point cloud data indicative of the three-dimensional distribution of feature points at a second density higher than the first density from the correspondence between the corrected first-density point cloud data and the omnidirectional images.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to 3D data processing technology.

Background Art

[0002] In plants such as power plants or factories, regular patrol inspections are carried out for construction progress management and abnormality confirmation. Conventional maintenance inspections are performed by manual recording work, and technologies aimed at labor saving and mechanization of a series of inspection work are required. Among these technologies, there is a 3D reconstruction technology that converts the on-site situation into 3D data using images taken by a camera during patrol inspections. With this technology, it is possible to convert the 3D space of the inspection location into 3D data from a video taken while moving during patrol inspections, and it becomes possible to grasp the situation such as the latest equipment layout of plants such as power plants or factories three-dimensionally. For example, if this reconstructed 3D data is stored in a server, the on-site situation at the time of inspection can be confirmed from any location. In addition, it can be utilized for grasping changes from the time of construction or daily changes in renovation work. By converting this into 3D data, it becomes possible to grasp the 3D dimensional information of the space, so it can be used for grasping the dimensions of the location to be confirmed and interference confirmation when arranging equipment. In addition, if it is managed centrally together with equipment information of the equipment and inspection data (records such as images, meter values, inspection results, etc.) and linked to the 3D data, intuitive confirmation becomes possible.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When using a stereo camera that combines two cameras, when performing three-dimensional reconstruction from an image, it is possible to obtain three-dimensional data with an absolute scale (actual scale) that matches the actual space based on the dimensions between the cameras. However, when using a monocular camera composed of a single camera, since the three-dimensional data does not have an absolute scale, correction to the absolute scale is required. In addition, since the coordinate system of the three-dimensional data is based on the position and orientation of the camera at the start of shooting, it is necessary to correct it to the reference and coordinate system of the actual space.

[0005] Embodiments of the present invention have been made in consideration of such circumstances, and an object thereof is to simplify the correction work of the coordinate system and scale in three-dimensional data processing using the omnidirectional image acquired by the omnidirectional camera.

Means for Solving the Problems

[0006] The three-dimensional data processing system according to an embodiment of the present invention receives a plurality of omnidirectional images captured by an omnidirectional camera at at least a first position and a second position different from the first position, extracts a plurality of feature points indicating characteristic parts of at least one object shown in the omnidirectional images, estimates the position and orientation of the omnidirectional camera at the time of shooting from the correspondence between the feature points and the omnidirectional images, restores first density point cloud data indicating the three-dimensional distribution of the feature points of the first density based on the feature points and the position and orientation of the omnidirectional camera at the time of shooting, performs correction to match the coordinate system and scale of the first density point cloud data to the absolute coordinate system and absolute scale of the actual physical space where shooting was performed by the omnidirectional camera based on at least one of the position and orientation of the omnidirectional camera at the time of shooting and information regarding the feature points, and restores second density point cloud data indicating the three-dimensional distribution of the feature points of a second density denser than the first density from the correspondence between the corrected first density point cloud data and the omnidirectional images, and is configured to include one or more computers.

Effects of the Invention

[0007] According to an embodiment of the present invention, in 3D data processing using an omnidirectional image acquired by an omnidirectional camera, the coordinate system and scale correction operations can be simplified.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Embodiments for Carrying Out the Invention

[0009] (First Embodiment) Hereinafter, embodiments of a three-dimensional data processing system and a three-dimensional data processing method will be described in detail with reference to the drawings. First, the first embodiment will be described with reference to FIGS. 1 to 10.

[0010] Reference numeral 1 in FIG. 1 is a three-dimensional data processing system according to the first embodiment. A three-dimensional data processing method is implemented using this three-dimensional data processing system 1. The three-dimensional data processing system 1 is mainly used when restoring (generating) three-dimensional information such as the position of an object at an inspection location. In order to accurately restore the three-dimensional information, an accurate position of the acquired three-dimensional data is required. Therefore, the three-dimensional data processing system 1 can also estimate an accurate position.

[0011] The three-dimensional data processing system 1 includes a mobile body 2 and a control computer 3. The mobile body 2 is, for example, an inspection robot that travels on the floor surface. The control computer 3 controls the mobile body 2 and processes the images captured by the mobile body 2.

[0012] The mobile body 2 includes a plurality of wheels 4 provided at the lower part of its housing. The mobile body 2 travels by rotationally driving the wheels 4. In addition, an omnidirectional camera 5 (first imaging unit) and a measurement camera 6 (second imaging unit) as imaging devices are mounted on the upper part of the housing of the mobile body 2. The omnidirectional camera 5 and the measurement camera 6 capture images of the periphery of the mobile body 2.

[0013] Note that although an inspection robot that travels on the floor surface is exemplified as the moving body 2, other embodiments may also be used. For example, the moving body 2 may be an inspection drone that flies in the air. Further, the moving body 2 may be a portable terminal that can be carried by an inspector. The inspector may carry the omnidirectional camera 5 (moving body 2) and estimate the shooting position at a later date from the image of the omnidirectional camera 5. Further, the moving body 2 may be a helmet worn by the inspector, and the omnidirectional camera 5 may be attached to this helmet. Further, the omnidirectional camera 5 may be attached to somewhere on the inspector's body. That is, the inspector himself / herself may be configured as the moving body 2.

[0014] The omnidirectional camera 5 is a 360° camera that can simultaneously shoot in all directions. This omnidirectional camera 5 generates an omnidirectional image by simultaneously shooting images of the periphery of the moving body 2 with a plurality of image sensors with fish-eye lenses, for example. Further, optical components such as a convex mirror may be used, and the surrounding scenery may be guided to one image sensor to shoot an omnidirectional image. With this omnidirectional camera 5, it is possible to simultaneously shoot an all-sphere image that shows the entire omnidirectional view of the top, bottom, left, and right around the moving body 2. Note that the captured image does not have to be an all-sphere image, and may be a panoramic image that shows a 360-degree range in the horizontal direction (omnidirectional left and right).

[0015] Here, a form in which one omnidirectional camera 5 shoots an omnidirectional image is exemplified, but other forms may also be used. For example, an omnidirectional image may be generated by synthesizing images captured by a plurality of omnidirectional cameras 5.

[0016] Note that the omnidirectional camera 5 may be a consumer product. The output resolution of the omnidirectional camera 5 may be arbitrarily determined, and may be, for example, 4K or FullHD. Further, the image data may be image data composed of three RGB channels, or may be image data of one grayscale channel.

[0017] The omnidirectional camera 5 can continuously generate an omnidirectional image as a spherical image, that is, generate a spherical image as a moving picture. The omnidirectional image is an image obtained by converting an image captured by the omnidirectional camera 5 into an image that can be viewed as a panoramic view on a plane, and the conversion method is, for example, the orthographic cylindrical projection method.

[0018] In the first embodiment, the omnidirectional camera 5 performs shooting while the moving body 2 is moving. That is, the omnidirectional camera 5 performs shooting of omnidirectional images at least at a first position and a second position different from the first position. A plurality of omnidirectional images captured by the omnidirectional camera 5 are input to the control computer 3. Then, the control computer 3 extracts a plurality of feature points indicating characteristic parts of at least one object shown in the omnidirectional image, and estimates the position and orientation of the omnidirectional camera 5 at the time of shooting from the correspondence relationship between the feature points and the omnidirectional image.

[0019] Furthermore, the control computer 3 restores (generates) first density point cloud data indicating the three-dimensional distribution of feature points of the first density based on the feature points and the position and orientation of the omnidirectional camera 5 at the time of shooting. Then, the control computer 3 corrects the coordinate system and scale of the first density point cloud data to match the absolute coordinate system and absolute scale of the actual physical space 40 (FIG. 5) where shooting is performed by the omnidirectional camera 5. This correction is performed based on, for example, at least one of the position and orientation of the omnidirectional camera 5 at the time of shooting and the information regarding the feature points. Here, the control computer 3 restores (generates) second density point cloud data indicating the three-dimensional distribution of feature points of a second density denser than the first density from the correspondence relationship between the corrected first density point cloud data and the omnidirectional image.

[0020] The coordinate system and scale of the point cloud data are values recorded in an environment map 50 (Fig. 5), which is a virtual space generated by the control computer 3, and are values that accumulate errors or change according to the acquisition situation. For example, even if the point cloud data is acquired by the same device, the same value may not be obtained each time. On the other hand, the absolute coordinate system and absolute scale refer to the coordinate system (actual coordinate system) and scale (actual scale, size) in the physical space 40 (Fig. 5), which is the actual real space, and are values uniquely determined according to the size of the actual physical space 40 and are fixed values.

[0021] The measurement camera 6 captures a detailed image of an object such as a device or structure to be inspected. The measurement camera 6 is, for example, an infrared camera that captures an infrared image of the inspection location and measures the temperature distribution in the image. Here, an example where the measurement camera 6 is mounted on the same moving body 2 as the omnidirectional camera 5 will be described, but it may be provided on a moving body 2 different from the omnidirectional camera 5. Note that the measurement camera 6 is not limited to an infrared camera and may be a camera that captures wavelengths other than infrared. In the following description, the image captured by the measurement camera 6 is referred to as a measurement image.

[0022] Note that the measurement camera 6 may be, for example, a pan-tilt-zoom camera. A pan-tilt-zoom camera has a pan function that enables panning in the horizontal direction, a tilt function that enables tilting in the vertical direction, and a zoom function that enables zooming in (telephoto) and zooming out (wide angle).

[0023] The images captured by the omnidirectional camera 5 and the measurement camera 6 exemplify videos captured at an arbitrary frame rate. Note that the images captured by the omnidirectional camera 5 and the measurement camera 6 may also be still images captured at regular intervals.

[0024] The three-dimensional data processing system 1 restores the three-dimensional information of the inspection location from the omnidirectional image captured by the omnidirectional camera 5. Further, the three-dimensional data processing system 1 estimates the position of the moving body 2 based on the image captured by the omnidirectional camera 5. Since the omnidirectional camera 5 is provided at the upper center of the moving body 2, the description will be made assuming that the shooting position of the omnidirectional camera 5 is the same as the position of the moving body 2.

[0025] As shown in FIG. 4, the inspection location is, for example, indoors in a building 30 in a predetermined plant such as a power plant, a chemical plant, or a factory. A large number of inspection objects 31 are arranged in the plant. Note that the inspection location may be a predetermined commercial facility or public facility where a large building 30 exists.

[0026] The moving body 2 travels indoors in the building 30. A large number of devices or structures serving as inspection objects 31 are provided indoors, and images of these are captured. Then, an inspector located at a remote location checks the inspection objects 31 by viewing the images. In addition, a plurality of markers 32 as specific subjects are attached to the indoor wall surfaces. The three-dimensional data processing system 1 refers to the positions of these markers 32 and corrects the position of the moving body 2.

[0027] Next, the system configuration of the three-dimensional data processing system 1 will be described with reference to the block diagrams shown in FIGS. 2 to 3.

[0028] The moving body 2 includes an omnidirectional camera 5 and a measurement camera 6. Although not particularly shown, in addition to this, the moving body 2 includes predetermined devices such as a traveling motor, a communication device, and a computer for the moving body.

[0029] For example, the traveling motor is mounted inside the housing of the moving body 2 and drives the wheels 4 to rotate (FIG. 1). By controlling the rotation of each wheel 4 by the traveling motor, the moving body 2 can move forward, backward, and turn.

[0030] The control computer 3 includes a processing circuit 7, a storage unit 8, an input unit 9, an output unit 10, and a communication unit 11. The control computer 3 may be provided at a remote location away from the moving body 2. This control computer 3 can communicate with the moving body 2 via a predetermined network.

[0031] The communication unit 11 is a wireless communication device that performs wireless communication. Through this communication unit 11, the control computer 3 can communicate with the moving body 2 and other computers. For example, the control computer 3, the moving body 2, and other computers are connected to each other via a predetermined communication line such as the Internet, a LAN (Local Area Network), a WAN (Wide Area Network), or a mobile communication network.

[0032] The control computer 3 controls the movement of the moving body 2. Although an example of a mode in which the control computer 3 automatically controls the movement of the moving body 2 is illustrated, other modes may also be possible. For example, the control computer 3 may receive an input operation of a user located at a remote location and control the movement of the moving body 2. That is, the control computer 3 may also be a remote operation unit for controlling the moving body 2 by a manual operation of the user.

[0033] The control computer 3 has hardware resources such as a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), and an SSD (Solid State Drive). By the CPU executing various programs, information processing by software is realized using the hardware resources. Furthermore, the three-dimensional data processing method of the present embodiment is realized by causing the control computer 3 to execute various programs.

[0034] Note that each component of the three-dimensional data processing system 1 does not necessarily have to be provided in one computer. For example, one three-dimensional data processing system 1 may be realized by a plurality of computers connected to each other via a network. Also, the three-dimensional data processing system 1 may be mounted on individual computers respectively.

[0035] The processing circuit 7 is, for example, a circuit including a CPU, a GPU (Graphics Processing Unit), a dedicated or general-purpose processor. This processor realizes various functions by executing various programs stored in the storage unit 8. Also, the processing circuit 7 may be configured by hardware such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Various functions can also be realized by these hardware. Also, the processing circuit 7 can also realize various functions by combining software processing by a processor and a program and hardware processing.

[0036] The storage unit 8 stores various information necessary for estimating the position of the moving body 2 based on information stored in a predetermined database. For example, the storage unit 8 records a marker management table (FIG. 7) and a position recording table (FIG. 8). Note that the database is a collection of information stored in a memory, an HDD, an SSD, or the cloud and organized so that it can be searched or accumulated. Also, the storage unit 8 stores various information such as images taken by the omnidirectional camera 5 and feature points.

[0037] Predetermined information is input to the input unit 9 according to the operation of a user who uses the control computer 3. The input unit 9 includes input devices such as a mouse, a keyboard, and a touch panel. That is, predetermined information is input to the control computer 3 according to the operation of these input devices.

[0038] The output unit 10 outputs predetermined information. The control computer 3 includes a device for displaying images such as a display that outputs the analysis results. The output unit 10 is, for example, a display. Note that the display may be separate from or integrated with the computer main body.

[0039] Although a display is exemplified as the device for displaying images, other modes may also be used. For example, image display may be performed using a head-mounted display or a projector. Further, a printer that prints information on a paper medium may be used instead of the display. That is, a head-mounted display, a projector, or a printer may be included as the object controlled by the output unit 10.

[0040] The configuration of the processing circuit 7 is shown in FIG. 3. This processing circuit 7 includes an image input unit 12, a feature point extraction unit 13, a marker detection unit 14, a position and orientation estimation unit 15, a first restoration unit 16, a coordinate scale correction unit 17, a range setting unit 18, a second restoration unit 19, a parallax correction unit 20, an information mapping unit 21, and a display control unit 22. These functions of the processing circuit 7 are realized by a program stored in a memory, HDD, or SSD being executed by a CPU. Note that the processing circuit 7 may include components other than those shown in FIG. 3, or some of the components shown in FIG. 3 may be omitted.

[0041] The image input unit 12 receives a plurality of omnidirectional images generated by the omnidirectional camera 5 and a plurality of measurement images generated by the measurement camera 6. These images are input to the image input unit 12 by wireless communication, but these images may also be input by a recording medium such as a flash memory.

[0042] The feature point extraction unit 13 extracts the first feature points existing in the omnidirectional image. The first feature points are the pixels that become corners extracted from each pixel on the image. A corner is a point where the pixel value changes greatly in all directions. For example, the image captured by the omnidirectional camera 5 is analyzed, and feature points with prominent contrast to the surroundings, such as parts of an object like a corner, are extracted, and point cloud data is generated with a plurality of feature points.

[0043] For the feature quantity calculation for extracting the first feature points, for example, ORB (Oriented FAST and Rotated BRIEF) is used. Then, the feature point extraction unit 13 detects the same first feature points from the similarity of the first feature points among a plurality of omnidirectional images that are consecutive frames, and associates the first feature points with each other. As the method for calculating the similarity of the first feature points, methods such as Brute Force, Brute Force-L1, Brute Force-Hamming, or Flann Based are used.

[0044] The marker detection unit 14 detects a plurality of markers 32 (specific subjects) that exist in a state of being fixed in advance at predetermined positions. Markers 32 are provided in advance at the inspection locations (Figure 4) where the moving body 2 travels. These markers 32 are, for example, attached to several locations on the indoor wall surface and are used for correction when estimating the self-position of the moving body 2. Here, the self-position of the moving body 2 is the shooting positions of the omnidirectional camera 5 and the measurement camera 6.

[0045] The control computer 3 stores the three-dimensional coordinates (absolute coordinate system) and sizes (absolute scale) of these markers 32 in the actual physical space 40 (Figure 5). Also, the control computer 3 generates an environmental map 50 (Figure 5) including information on the surrounding environment according to the movement of the moving body 2. For example, the control computer 3 estimates the self-position of the moving body 2 and records the position and trajectory of the moving body 2 in the environmental map 50.

[0046] That is, the control computer 3 stores the three-dimensional position in the physical space 40 of at least one marker 32 that exists in a state fixed in advance at the location photographed by the omnidirectional camera 5. Then, the control computer 3 performs correction to match the coordinate system and scale of the first density point cloud data to the absolute coordinate system and absolute scale of the physical space 40 based on the three-dimensional position in the physical space 40 of the marker 32 shown in the omnidirectional image.

[0047] In the first embodiment, the information regarding the feature points includes the information regarding the specific subject. Here, it is preferable that there are at least two specific subjects shown in the omnidirectional image. In addition, when the distance and direction from the omnidirectional camera 5 to one specific subject can be accurately acquired by a predetermined sensor (not shown), the coordinate system and scale of the first density point cloud data can be corrected based on one specific subject.

[0048] The position and orientation estimation unit 15 estimates the external parameters, which are the position and orientation of the omnidirectional camera 5, based on the correspondence relationship between the first feature points in a plurality of consecutive omnidirectional images that are consecutive frames.

[0049] The position and orientation estimation unit 15 simultaneously estimates the external parameters (R1, T1) of the omnidirectional camera 5 and the coordinates of the first feature points in the omnidirectional image such that the reprojection error between the first feature points projected by, for example, the non-linear least squares method and their back-projected points is minimized. Here, R1 indicates, for example, the rotational movement with respect to a reference point such as the shooting start point of the omnidirectional camera 5. T1 indicates the translational movement with respect to the reference point. In addition, there may be a case where the point cloud data of the feature points having an absolute scale in advance is stored in the storage unit 8. In this case, the estimated external parameters of the omnidirectional camera 5 and the first feature points may be converted from local coordinates to global coordinates. Here, the local coordinates are the coordinates of the virtual space used by the processing circuit 7, and the global coordinates are the coordinates of the physical space 40. That is, the external parameters of the omnidirectional camera 5 and the first feature points may be corrected according to the known scale at the site. In the following description, the coordinates of the first feature points may sometimes be simply referred to as the first feature points.

[0050] Note that the position and orientation estimation unit 15 only needs to be able to estimate the external parameters (R1, T1) of the omnidirectional camera 5. For example, the external parameters (R1, T1) of the omnidirectional camera 5 may be calculated based on the changes in the external parameters (R1, T1) measured by a motion sensor (not shown) provided on the omnidirectional camera 5 (mobile body 2), or the external parameters (R1, T1) may be measured by a satellite positioning system such as GPS (Global Positioning System).

[0051] The first restoration unit 16 estimates the three-dimensional distribution of the feature points. This estimation is performed based on the first feature points and the external parameters (R1, T1) of the omnidirectional camera 5. Here, the first feature points are associated among a plurality of omnidirectional images, which are a plurality of entire-sphere panorama images G1 that are frames taken at different times (at different positions). The external parameters (R1, T1) of the omnidirectional camera 5 are the parameters in the entire-sphere panorama image G1 where the first feature points are acquired.

[0052] Note that although the entire-sphere panorama image G1, which is a spherical image, is exemplified as the omnidirectional image, the omnidirectional image may also be a cylindrical image.

[0053] FIG. 9 is a diagram schematically showing an example of the processing of the first restoration unit 16. Assume that there is a first feature point P1a in the omnidirectional panorama image G1a corresponding to the external parameters (R1a, T1a) and a first feature point P1b in the omnidirectional panorama image G1b corresponding to the external parameters (R1b, T1b). Then, a point P0 in the physical space 40 of the feature points indicated by these first feature points P1a and P1b is shown. As shown in this FIG. 9, the storage unit 8 stores, as the three-dimensional coordinates of the point P0, the intersection of a projection line L1a connecting the external parameters (R1a, T1a) and the first feature point P1a and a projection line L1b connecting the external parameters (R1b, T1b) and the first feature point P1b. By performing this process for a plurality of first feature points, the three-dimensional distribution of the feature points is estimated, and first density point group data, which is sparse three-dimensional point group data, is generated. Note that the first density point group data may be three-dimensional point group data that is at least sparser than the second density point group data.

[0054] The coordinate scale correction unit 17 corrects the coordinate system and scale of the first density point group data to match the absolute coordinate system and absolute scale of the actual physical space 40 (FIG. 5) in which shooting is performed by the omnidirectional camera 5.

[0055] For example, the coordinate scale correction unit 17 designates a portion that serves as a reference for the coordinates and scale in the image. The coordinate scale correction unit 17 performs this process for a plurality of angles and performs correction by inputting the absolute scale of the designated portion. Note that it is also possible to automatically extract based on a reference image registered in advance. When the reference portions repeatedly exist at a certain interval, by detecting each of them, it is possible to improve the accuracy of the correction.

[0056] The range setting unit 18 sets a range for restoring the second density point cloud data based on the position and orientation at the time of shooting the omnidirectional image in accordance with the absolute coordinate system and absolute scale of the physical space 40. This range setting may be automatically performed by the range setting unit 18, or an arbitrary range may be set according to the user's input operation. Then, the second restoration unit 19 trims the feature points outside the range and restores the second density point cloud data within the range. In this way, noise that appears as a feature point outside the range can be removed.

[0057] For example, when shooting is performed in the indoor space of the building 30 (Fig. 4), the feature points that appear in the space outside the indoor space of the building 30 can be regarded as noise. By setting the indoor area of the building 30 as the range for restoring the second density point cloud data, the range setting unit 18 can remove the feature points that are noise outside the indoor area of the building 30.

[0058] Since the first density point cloud data includes the feature points in the entire area captured by the omnidirectional camera 5, as a result of restoration as the second density point cloud data, it includes noise parts, that is, data outside the area to be used as the second density point cloud data. Therefore, not only does the restoration result become difficult to understand, but the capacity of the second density point cloud data also increases. The range setting unit 18 limits the range output as the restoration result, and sets the range for outputting the restoration result for the first density point cloud data corrected to the absolute coordinates and absolute scale by the coordinate scale correction unit 17.

[0059] Only within the set range is output as the restoration result. By this process, it becomes possible to delete unnecessary noise, that is, data outside the usage range, and the utilization efficiency of the data can be improved. Also, for unnecessary noise existing within the range, by setting a threshold for the density of the feature points, feature points with low density can be deleted as noise, and it is also possible to improve the accuracy of the second density point cloud data.

[0060] The second restoration unit 19 estimates the three-dimensional shape of an object in the omnidirectional panoramic image G1 based on the first feature points associated between the first density point group data and the plurality of omnidirectional panoramic images G1, and generates second density point group data which is dense three-dimensional point group data. The processing of the second restoration unit 19 can be realized by the multi-view stereo method. Note that multi-view stereo is a method that extends the restoration of three-dimensional shapes based on stereo matching, that is, matching between stereo images, to also use three or more images simultaneously.

[0061] As the multi-view stereo method, for example, a local method may be used in which a matching cost is calculated using the shading values in the peripheral region of a pixel, and one corresponding point with the lowest cost (highest similarity) is determined. Further, after calculating the matching cost, a global method may be used in which it is formulated as an optimization problem for the entire image to obtain the overall disparity (disparity image). The optimization problem here includes, for example, an energy minimization problem. Alternatively, semi-global matching, which is an intermediate method between the local and global methods, may be used. Note that semi-global matching is a method of obtaining a disparity image by using the matching cost for each pixel, which is a feature of the local method, aggregating the cost by approximation for the energy function, which is a feature of the global method, and solving the optimization problem.

[0062] Note that the first restoration unit 16 and the second restoration unit 19 may generate colored three-dimensional point group data based on the color information of the omnidirectional panoramic image G1.

[0063] The disparity correction unit 20 corrects the disparity between the omnidirectional image captured by the omnidirectional camera 5 and the measurement image captured by the measurement camera 6 based on the positional relationship between the omnidirectional camera 5 and the measurement camera 6.

[0064] When the positional relationship between the omnidirectional camera 5 and the measurement camera 6 is fixed regardless of time and is known, the external parameters between the omnidirectional camera 5 and the measurement camera 6 may be calculated in advance. In this case, as the initial value of the external parameters, the radius of the spherical surface attached to the omnidirectional image is input. Further, the offsets in the (u, v) directions of the omnidirectional image, the displacements between the omnidirectional camera 5 and the measurement camera 6 in the (x, y, z) directions, and the displacements between the omnidirectional camera 5 and the measurement camera 6 in terms of rotation about the (x, y, z) axes are input.

[0065] Then, the parallax correction unit 20 corrects the parallax between the omnidirectional image and the measurement image based on these input initial values. When the positional relationship between the omnidirectional camera 5 and the measurement camera 6 varies with time, after calculating the relative positional relationship based on the respective external parameters of the omnidirectional camera 5 and the measurement camera 6, the above-described processing is executed.

[0066] The information mapping unit 21 associates the pixels of the omnidirectional image and the measurement image with each other and superimposes the measurement image on the second density point cloud data. First, the images (frames) captured by the omnidirectional camera 5 and the measurement camera 6 according to their respective shooting times are acquired. The frame rates of the omnidirectional camera 5 and the measurement camera 6 may be different, and in some cases, the frame rate of the omnidirectional camera 5 may be smaller. In this case, frames are acquired at specified intervals based on the frame rate of the omnidirectional camera 5. The measurement camera 6 acquires a frame at a time close to the frame acquisition time of the omnidirectional camera 5. Then, the frames acquired at these close times are associated with each other. Further, based on these associated frames and the above-described parallax correction, the pixels of the omnidirectional image and the measurement image are associated with each other.

[0067] Since the positions of the feature points of the second density point cloud data corresponding to the pixels of the omnidirectional image are calculated as described above, by associating the pixels of the omnidirectional image and the measurement image with each other, it becomes possible to associate the second density point cloud data with the pixels of the measurement image. Then, the information mapping unit 21 superimposes the measurement image on the second density point cloud data corresponding to the omnidirectional image.

[0068] Note that the information mapping unit 21 may extract the color information of the second density point cloud data from the measurement image based on the positional relationship between the omnidirectional camera 5 and the measurement camera 6. That is, the parallax between the omnidirectional image and the measurement image does not necessarily have to be corrected by the parallax correction unit 20. Then, the information mapping unit 21 may directly extract the color information of the measurement image corresponding to the second density point cloud data from the positional relationship between the omnidirectional camera 5 and the measurement camera 6, and superimpose the second density point cloud data and the measurement image.

[0069] The display control unit 22 controls, for example, the display of the display. As described above, the display control unit 22 displays an image in which the second density point cloud data and the measurement image are superimposed on the display. For example, the image may be displayed on the display of a terminal such as a tablet or a smartphone owned by the user.

[0070] Further, the display control unit 22 may have a function of switching the presence or absence of superimposing the measurement image on the second density point cloud data. For example, a switching button is provided on the display, and when the user presses the switching button, it is possible to switch between an image in which the measurement image is superimposed on the second density point cloud data and an image of only the second density point cloud data.

[0071] Known techniques are used for the self-position estimation of the first embodiment. For this self-position estimation, for example, SLAM (Simultaneous Localization and Mapping) is used. By this technique, the self-position of the moving body 2 can be obtained, and the amount of movement can be obtained. In particular, VSLAM (Visual Simultaneous Localization and Mapping) is used. By using this VSLAM, it is possible to record images and positions only with the omnidirectional camera 5.

[0072] The control computer 3 calculates the changes in the position and orientation of the moving body 2, which is the shooting position of the omnidirectional camera 5, based on the images captured by the omnidirectional camera 5. Therefore, self-position estimation is possible even in places such as indoors where the satellite positioning system cannot be used. Note that the orientation of the moving body 2 includes information on the posture such as the inclination of the moving body 2.

[0073] VSLAM is a technology that extracts feature points of surrounding objects using the information obtained by the omnidirectional camera 5. When the omnidirectional camera 5 moves, the image changes accordingly, and the feature points in the image also move. By tracking the movement of these feature points three-dimensionally, it is possible to simultaneously obtain the three-dimensional point cloud data of the inspection location and the position of the omnidirectional camera 5. Also, the posture (orientation) of the omnidirectional camera 5 can be obtained from the positional relationship between the image of the omnidirectional camera 5 and the positions of the respective feature points.

[0074] Here, the feature points of the objects shown in the omnidirectional image are extracted. For example, by calculating the trajectory of the movement of the omnidirectional camera 5 (moving body 2) from a predetermined location with a known position as the starting point, the current position and orientation can be calculated. This VSLAM is a technology that can estimate the self-position without previously obtaining the three-dimensional point cloud data of the inspection location. Then, at the inspection location, by calculating the respective feature points of the objects shown in the multiple frames constituting the omnidirectional image, the self-position can be estimated from the amount of movement of the feature points.

[0075] The three-dimensional point cloud data obtained by VSLAM is recorded in the environmental map 50 (Fig. 5). That is, the three-dimensional coordinates of the multiple feature points of the object are recorded in the environmental map 50. The control computer 3 estimates the three-dimensional coordinates indicating the position of the moving body 2 in the environmental map 50 based on the feature points recorded in the environmental map 50.

[0076] The position and orientation obtained by this VSLAM are the accumulation of changes in relative position and orientation from the start time of the VSLAM process. Therefore, if there is an error (positional deviation) between the environmental map 50 and the physical space 40 (Fig. 5) at the start time of the VSLAM process, the exact position of the moving body 2 cannot be estimated. Thus, by correcting the deviation of the environmental map 50 to align it with the physical space 40, the position and orientation of the moving body 2 corresponding to the inspection location can be estimated.

[0077] In the first embodiment, the marker 32 is used to align the environmental map 50 with the physical space 40. As shown in Fig. 6, the marker 32 is a figure that can be recognized by image recognition by the control computer 3, and a readable code 33 is printed on a predetermined mount 34. For example, a matrix-type two-dimensional code, a so-called QR code (registered trademark), is printed on the mount 34. Note that the marker 32 may be a known AR marker. This marker 32 is attached to the wall surface of the inspection location.

[0078] Furthermore, the colors of the mounts 34 of the markers 32 are different from each other. For example, the respective mounts 34 are color-coded in red, blue, yellow, and green. The colors of the mounts 34 of the plurality of markers 32 provided in at least the same room should be arranged so as to be different from each other. Note that there may be a plurality of markers 32 of the same color in the same room, but even in that case, the colors of the plurality of markers 32 provided on the same wall surface should be arranged so as to be different from each other.

[0079] As shown in Fig. 7, in the marker management table, the color of the marker 32, the location ID indicating the building 30 where the marker 32 is provided, and the coordinates of the marker 32 are registered in association with the marker ID, which is identification information that can individually identify the marker 32.

[0080] Note that the location ID is identification information that can individually identify a plurality of buildings 30 or a plurality of rooms. For example, an absolute coordinate system of the physical space 40 is set for each location ID.

[0081] In addition, the coordinates of the marker 32 registered in the marker management table are the coordinates of a preset physical space 40 (FIG. 5). In this way, the coordinates in the physical space 40 can be specified from the marker 32 shown in the image captured by the omnidirectional camera 5, and the position of the moving body 2 can be estimated from the marker 32.

[0082] Next, the flow of the correction process for the coordinate system and scale of the environmental map 50 (FIG. 5) including the first density point cloud data will be described in detail. Note that the above-mentioned drawings may be referred to.

[0083] Note that the flow of the process described below is an example, and there may be other process flows. Also, the order of each process is not necessarily fixed, and the order of some processes may be reversed. Also, some processes may be executed in parallel with other processes.

[0084] First, the omnidirectional camera 5 captures an omnidirectional image around the moving body 2 (FIG. 2), and this omnidirectional image is sent to the control computer 3 and input to the image input unit 12 (FIG. 3) of the processing circuit 7. Here, the feature point extraction unit 13 (FIG. 3) extracts a plurality of feature points of the object shown in the omnidirectional image from the omnidirectional image obtained from the omnidirectional camera 5 based on the luminance gradient and similarity on the image.

[0085] The feature point extraction unit 13 records the three-dimensional coordinates of the extracted feature points in the environmental map 50 (FIG. 5). Also, the feature point extraction unit 13 compares the feature points already recorded in the storage unit 8 with the newly extracted feature points based on the vectors possessed by each feature point. That is, the feature point extraction unit 13 compares whether the feature points are of the same object. For example, when the feature points extracted from the first frame of the image are already recorded in the environmental map 50, it is compared whether the newly acquired feature points extracted from the second frame are related to the already recorded feature points of the first frame.

[0086] That is, the feature point extraction unit 13 compares the feature points newly recorded in the environmental map 50 with the known feature points. When the feature point extraction unit 13 determines that the new feature points are related to the known feature points, the feature point extraction unit 13 associates the new feature points with the object related to the known feature points and records them in the environmental map 50.

[0087] The position and orientation estimation unit 15 (Fig. 3) estimates the position where the image was taken from the three-dimensional coordinates of the respective feature points of the objects recorded in the environmental map 50. That is, the position and orientation estimation unit 15 estimates the coordinates of the moving body 2 at the time when the image of the new feature points was taken from the coordinates of the new feature points. The position and orientation estimation unit 15 records the estimated coordinates of the moving body 2 in the environmental map 50.

[0088] The position and orientation estimation unit 15 measures the distance from the moving body 2 to the new feature points based on the coordinates of the moving body 2 and the position of the new feature points in the image. The position in the image is two-dimensional coordinates when the vertical and horizontal dimensions of one image are used as the dimensional axes. For example, when there are feature points not recorded in the environmental map 50, the position and orientation estimation unit 15 calculates the distance from the feature points to the omnidirectional camera 5 from the position of the feature points in the image based on the principle of triangulation. The position and orientation estimation unit 15 records the measured distance in the environmental map 50 together with the coordinates. In this way, the accurate position of the object can be recorded in the environmental map 50.

[0089] Here, the marker detection unit 14 (Fig. 3) detects a marker 32 capable of obtaining the relative coordinates with respect to the omnidirectional camera 5 from the image of the omnidirectional camera 5. The marker 32 to be detected here is registered in advance in a marker management table (Fig. 7).

[0090] The marker detection unit 14 tracks the marker 32 shown in each frame constituting the image. When the marker 32 cannot be detected from the image, this marker detection unit 14 estimates the coordinates of the marker 32 from the movement amount of the moving body 2. Then, based on the estimated coordinates of the marker 32 and the environmental map 50, the marker detection unit 14 specifies the range in which the marker 32 may be shown in the image, and detects the marker 32 from the specified range. In this way, the accuracy of detecting the marker 32 from the image can be improved.

[0091] For example, when the distance from the omnidirectional camera 5 to the marker 32 becomes long, the detection accuracy of the marker 32 decreases. Therefore, the marker detection unit 14 reads the shooting position, which is the position of the moving body 2, from the environmental map 50, extracts the range (area) of the image where the marker 32 is estimated to exist, and performs the process of detecting the marker 32 with parameters that match the range. In this way, by adjusting the parameters for image processing, the detection accuracy of the marker 32 can be improved.

[0092] Here, the coordinates of the marker 32 and the shooting position, which are the detection results of the marker 32, are recorded in the environmental map 50. Furthermore, the coordinates of the location where each marker 32 is fixed are recorded in advance in the coordinate system that the user wants to manage. This coordinate system that the user wants to manage is, for example, the coordinate system of the physical space 40.

[0093] The marker detection unit 14 associates the relative coordinates of the shooting position obtained from the marker 32 with the coordinates of the location where the marker 32 is fixed. The marker detection unit 14 records the shooting position obtained from the marker 32 in the environmental map 50.

[0094] The coordinate scale correction unit 17 (Fig. 3) obtains an arbitrary transformation matrix so that the same image frames match for the shooting position obtained from the feature points of the object and the shooting position obtained from the marker 32. Then, the coordinate scale correction unit 17 corrects the shooting position obtained from the feature points by applying the shooting position obtained from the marker 32.

[0095] Here, the coordinate scale correction unit 17 corrects the coordinate system and scale of the environment map 50 (Fig. 5) to match the absolute coordinate system and absolute scale of the physical space 40. This correction is performed, for example, based on the coordinates in the physical space 40 (Fig. 5) of the marker 32 shown in the image captured by the omnidirectional camera 5 and the position in the image. Note that this correction includes a mode of aligning the origin and the direction of the coordinate axes of the environment map 50 with those of the physical space 40.

[0096] As shown in Fig. 5, the physical space 40 is, for example, the indoor space of the building 30. Taking a predetermined position in this physical space 40 as the origin of the coordinate system, the coordinates 41 of each marker 32 in the physical space 40 are stored in the control computer 3. The coordinates 41 of each marker 32 are acquired in advance at the inspection location by the user and input into the control computer 3.

[0097] Also, the environment map 50 is a virtual space generated and updated by the control computer 3 as the mobile body 2 moves. Taking a predetermined position in this environment map 50 as the origin of the coordinate system, based on the coordinates 41 of the marker 32 in the physical space 40, the coordinates 51 of the marker 32 estimated in advance are also recorded in the environment map 50.

[0098] Here, when the mobile body 2 moves in the actual physical space 40, an error occurs between the actual movement path 42 and the movement path 52 recorded in the environment map 50. Therefore, the coordinate system and scale of the environment map 50 are corrected so that the coordinates 51 of the marker 32 recorded in the environment map 50 match the coordinates 41 of the marker 32 in the actual physical space 40.

[0099] The coordinate scale correction unit 17 (Fig. 3) first obtains the coordinates 43 of the shooting position of the omnidirectional camera 5 in the physical space 40 based on the coordinates 41 of the marker 32 in the physical space 40 and the position of the marker 32 in the image using a projection matrix. Then, the coordinate scale correction unit 17 corrects the coordinate system and scale of the environmental map 50 so that the coordinates 53 of the moving body 2 in the environmental map 50 match the coordinates 43 of the actual shooting position. In this way, the coordinate system and scale of the environmental map 50 can be corrected from the coordinates 43 of the shooting position of the omnidirectional camera 5.

[0100] Note that the scale of the environmental map 50 refers to the magnification or reduction rate of the environmental map 50. This scale also depends on the scale at the start of the VSLAM process. Therefore, by correcting the scale, the accurate shooting position of the omnidirectional camera 5, that is, the accurate position of the moving body 2 in the physical space 40, can be estimated.

[0101] In this way, when the moving body 2 performs inspection, the self-position (shooting position) is measured from the image of the omnidirectional camera 5, and at the same time, the correction of the coordinate system and the scale of the coordinate system are performed.

[0102] Conventional self-position estimation (VSLAM) based on feature points can be performed without the marker 32, but the coordinate system and scale of the environmental map 50 become indeterminate. On the other hand, self-position estimation using the marker 32 can obtain the self-position in the coordinate system and scale of the environmental map 50 only within the range where the marker 32 can be seen. In this embodiment, by estimating the self-position by VSLAM and correcting the self-position with the marker 32, the accuracy of estimating the self-position can be improved.

[0103] As shown in Fig. 3, the shooting position, which is the coordinates of the corrected environmental map 50 and the moving body 2, is recorded. Here, the coordinates of the moving body 2 are registered in the position recording table (Fig. 8).

[0104] As shown in FIG. 8, in the position recording table, the coordinates of the moving body 2 at the time of acquisition of each frame and the marker ID of the marker 32 shown in each frame are registered in association with the frame number of the image. Note that the number of marker IDs associated with one frame number is synonymous with the number of markers 32 detected when the coordinates of the moving body 2 are estimated.

[0105] In this way, the estimated coordinates of the moving body 2 are associated with the frame at the time when the image capturing the object is taken. By doing so, it becomes easier for the marker detection unit 14 to obtain the coordinates of the feature points that overlap with each other in a plurality of frames, and it also becomes easier to estimate the coordinates of the moving body 2.

[0106] Also, the estimated coordinates of the moving body 2 are associated with the frame at the time when the image capturing the marker 32 is taken. By doing so, it becomes easier for the marker detection unit 14 (FIG. 3) to track the marker 32 shown in the image of each frame.

[0107] The feature point extraction unit 13 (FIG. 3) obtains, for example, the transformed coordinates such that the respective feature points overlap, using the frame number of the image in the position recording table (FIG. 8) as a key. By doing so, the position and orientation estimation unit 15 (FIG. 3) can estimate the imaging position based on each feature point. Based on the result, the coordinate scale correction unit 17 (FIG. 3) can correct the coordinate system and scale of the environmental map 50. Therefore, the inspector can easily collate and confirm the inspection target 31 in the layout diagram of the inspection location.

[0108] Further, the control computer 3 preferentially uses the coordinates of the moving body 2 estimated when detecting the second number of markers 32, which is larger than the first number, rather than the coordinates of the moving body 2 estimated when detecting the first number of markers 32, for processing to estimate the position of the moving body 2. That is, the control computer 3 preferentially uses the coordinates of the moving body 2 estimated when there are many markers 32 shown in one frame for processing to estimate the position of the moving body 2. When the detected number of markers 32 is larger, the position (imaging position) of the moving body 2 can be estimated more accurately.

[0109] For example, the feature points to be recorded are recorded in association with the number of markers 32 shown in the image at the time of distance measurement. Then, the position and orientation estimation unit 15 can derive a more reliable estimation result of the positional relationship by preferentially using the one with a larger number of markers 32 shown in the image for processing.

[0110] The coordinate scale correction unit 17 (Fig. 3) performs correction to match the coordinate system and scale of the environment map 50 (first density point cloud data) with the absolute coordinate system and absolute scale of the physical space 40 based on the coordinates and positions in the image of at least three markers 32. In this way, the correction of the coordinate system and scale of the three-dimensional environment map 50 can be performed more accurately.

[0111] Also, conventionally, when the distance from the omnidirectional camera 5 to the marker 32 becomes long, the size of the marker 32 shown in the image becomes small, and there are cases where it cannot be detected. In this case, the estimation accuracy of the shooting position by the marker 32 decreases. Here, although it is possible to increase the range in which the marker 32 is detected by attaching a large number of markers 32 to the wall surface, the work of attaching the markers 32 is troublesome.

[0112] The colors of the mounts 34 of the markers 32 in this embodiment are different from each other. Therefore, even for the marker 32 that is far from the omnidirectional camera 5, its position can be extracted based on the color of the marker 32. For example, as shown in the marker management table (Fig. 7), the control computer 3 stores the color information of the marker 32 in association with the identification information of the marker 32. Then, the control computer 3 detects the position of the marker 32 in the image based on these colors.

[0113] The marker detection unit 14 (FIG. 3) extracts a specified color range (multiple pixels) from a range (area) in an image where the marker 32 may exist, and identifies the position of the marker 32 in the image by assuming that the marker 32 exists at the center of gravity of the range. The coordinates of the marker 32 are recorded in the environment map 50. In this way, even if the marker 32 is far away, it is possible to identify the marker 32 from its color.

[0114] Furthermore, when the marker 32 is recorded as a colored point instead of the marker 32 itself, the marker detection unit 14 may process the positions of a plurality of points as a PnP problem. For example, the relationship between the coordinates of the marker 32 in the physical space 40 and the position of the marker 32 in the image (two-dimensional coordinates in the image) may be analyzed as a PnP problem, and the shooting position may be estimated by the marker 32.

[0115] In the first embodiment, the user attaches a small number of markers 32 to the wall surface of the building 30. By simply performing this task, the user can correct the scale and coordinate system of the environment map 50. This reduces the amount of manual correction work that has been required in the past, and makes it easier to operate the mobile object 2 that performs automatic inspection.

[0116] Next, the three-dimensional data processing executed by the three-dimensional data processing system 1 will be described with reference to the flowchart of Fig. 10. The above-mentioned drawings may be referred to. The following steps are at least a part of the processing included in the three-dimensional data processing, and other steps may be included in the three-dimensional data processing.

[0117] First, in step S1, the omnidirectional camera 5 (FIG. 2) generates (takes) an omnidirectional image (spherical panoramic image G1) of the inspection location, and the measurement camera 6 (FIG. 2) generates (takes) a measurement image of the inspection location. These images are input to the control computer 3 (FIG. 2). Then, the image input unit 12 (FIG. 3) of the processing circuit 7 acquires these images.

[0118] In the next step S2, the feature point extraction unit 13 (Fig. 3) extracts the first feature points existing in the omnidirectional image and associates the same first feature points among a plurality of omnidirectional images.

[0119] In the next step S3, the position and orientation estimation unit 15 (Fig. 3) estimates the external parameters (R1, T1) of the omnidirectional camera 5 and the coordinates of the first feature point P1 in the omnidirectional camera 5 based on the first feature points associated by the feature point extraction unit 13.

[0120] In the next step S4, the first restoration unit 16 (Fig. 3) estimates the three-dimensional distribution of the feature points based on the first feature points associated among a plurality of omnidirectional images and the external parameters (R1, T1) of the omnidirectional camera 5 in the omnidirectional image. Then, the first restoration unit 16 generates first density point cloud data of the first density.

[0121] In the next step S5, the coordinate scale correction unit 17 corrects the coordinate system and scale of the first density point cloud data to match the absolute coordinate system and absolute scale of the actual physical space 40 (Fig. 5) where shooting is performed by the omnidirectional camera 5. Here, the coordinate scale correction unit 17 identifies the absolute coordinate system and absolute scale based on the marker 32 detected by the marker detection unit 14 and performs correction to match them.

[0122] In the next step S6, the range setting unit 18 sets a range for restoring the second density point cloud data based on the position and orientation at the time of shooting the omnidirectional image in accordance with the absolute coordinate system and absolute scale of the physical space 40.

[0123] In the next step S7, the second restoration unit 19 (Fig. 3) estimates the three-dimensional shape of the object in the omnidirectional image based on the first feature points associated among a plurality of omnidirectional images and the first density point cloud data, and generates second density point cloud data of a second density denser than the first density.

[0124] In the next step S8, the parallax correction unit 20 (Fig. 3) corrects the parallax between the omnidirectional image and the measurement image based on the positional relationship between the omnidirectional camera 5 and the measurement camera 6.

[0125] In the next step S9, the information mapping unit 21 (Fig. 3) associates the pixels of the omnidirectional image and the measurement image with each other, and based on this association, superimposes the measurement image on the second density point cloud data.

[0126] In the next step S10, the display control unit 22 (Fig. 3) displays the image obtained by superimposing the measurement image and the second density point cloud data on the output unit 10 such as a display. Then, the three-dimensional data processing is completed.

[0127] Note that the display control unit 22 stores or outputs the position and orientation of the omnidirectional camera 5, the coordinates of the feature points, and the result of correcting the angle. Also, the display control unit 22 may not store or output the result. For example, it may store or output in the form of the initially estimated numerical values and the transformation matrix. If the initially estimated numerical values and the transformation matrix can be obtained, the result can be obtained in subsequent processing.

[0128] Note that when the three-dimensional data processing system 1 repeats this process, it may utilize various information such as the omnidirectional image, measurement image, or feature points stored previously.

[0129] Note that the processing circuit 7 as a modification example has an update unit. When the omnidirectional camera 5 newly generates an omnidirectional image, the update unit calculates the parallax between the external parameters (R1, T1) at that time and the external parameters of the omnidirectional camera 5 already recorded. Then, based on this parallax, the update unit updates the pixels of the omnidirectional image associated with the pixels of the measurement image as the pixels of the newly generated omnidirectional image. In this way, it becomes possible to superimpose the second density point cloud data corresponding to the latest inspection location and the measurement image.

[0130] In the first embodiment, first, sparse first density point cloud data is restored, and coordinate system and scale corrections are performed. Then, by restoring dense second density point cloud data, the processing load related to coordinate system and scale corrections can be reduced. Note that "restoring" means reproducing the state of the physical space 40 and includes newly generating three-dimensional point cloud data. That is, the term "restoring" includes the meaning of "generating".

[0131] Also, since a measurement image is superimposed on the second density point cloud data, which is dense three-dimensional point cloud data, the user can three-dimensionally check the situation at any point at the inspection location even from a remote location. Furthermore, the user can easily detect the presence or absence of abnormalities at the inspection location. In addition, since the measurement image is superimposed on the restored dense point cloud data without superimposing it on the distorted omnidirectional image, the accuracy of this superimposition can be improved, and the inspection results can be grasped more accurately.

[0132] (Second Embodiment) Next, the second embodiment will be described with reference to FIGS. 11 to 13. Note that the same reference numerals are given to the same components as those shown in the above-described embodiments, and redundant descriptions are omitted.

[0133] As shown in FIG. 11, the processing circuit 7 of the second embodiment has a conversion unit 60 in addition to the configuration of the first embodiment (FIG. 3). These functions of the processing circuit 7 are realized by a program stored in a memory, HDD, or SSD being executed by a CPU.

[0134] The conversion unit 60 maps the omnidirectional image, i.e., the equirectangular format full-sphere panoramic image G1 (Fig. 12), onto each face of a virtual cube set to an arbitrary size in advance, and generates a converted image G2 (Fig. 12) in cube map format with distortion corrected. Then, the position and orientation estimation unit 15 calculates the virtual camera positions and orientations corresponding to each face of the converted image G2. This calculation is performed based on the position and orientation at the time of shooting by the omnidirectional camera 5 and the correspondence relationship between the relative orientations of each face of the full-sphere panoramic image G1 (omnidirectional image) and the converted image G2 (virtual cube). Further, the conversion unit 60 projects the feature points onto the converted image G2.

[0135] The conversion unit 60 projects the first feature point in the full-sphere panoramic image G1 onto a predetermined face of the converted image G2 as a second feature point. That is, based on the coordinates of the first feature point P1 in the full-sphere panoramic image G1 and the association of the first feature points among a plurality of full-sphere panoramic images G1, the coordinates of the second feature point P2 on a predetermined face in the converted image G2 can be calculated.

[0136] The conversion unit 60, for example, converts (Toei) the equirectangular format full-sphere panoramic image G1 onto a predetermined face of the converted image G2. The converted image G2 in cube map format is generated by each of these converted faces.

[0137] Fig. 12 shows an example of the processing when converting the equirectangular format full-sphere panoramic image G1 taken from the omnidirectional camera 5 at a predetermined position (R1, T1) into a converted image G2 in cube map format.

[0138] The full-sphere panoramic image G1 is, for example, a spherical image captured on the surface of a sphere with a radius r. Therefore, each pixel (first feature point) of the full-sphere panoramic image G1 can be represented in polar coordinates. As a result, since the arrangement relationship of each face of the converted image G2 in the front, rear, left, right, up, and down directions and the polar coordinates of the first feature point P1 are known, the first feature point P1 can be projected onto each face of the converted image G2 as the second feature point P2. Note that the point P0 shows an example of the shooting target point of the second feature point P2 in the physical space 40.

[0139] Through the processing of the conversion unit 60, a converted image G2 having front, rear, left, right, top, and bottom surfaces with less distortion than the omnidirectional panoramic image G1 is generated. Here, the case where the omnidirectional panoramic image G1 is projected onto each plane of a hexahedron is described, but the plane to be projected can also be a polyhedron having six or more planes. Also, it is possible to use a hemispherical panoramic image instead of the omnidirectional panoramic image G1. In this case, the same processing can be performed by using a converted image G2 having front, left, right, top, and bottom surfaces instead of projecting onto a hexahedron.

[0140] Note that any surface of the converted image G2 may be used. For example, four converted images G2 corresponding to the side surfaces may be used. Also, each surface of the converted image G2 is, for example, a square image, and its size is, for example, 642 pixels, but other shapes and sizes may also be used.

[0141] The conversion unit 60 calculates the conversion external parameters (R2, T2) of the omnidirectional camera 5 corresponding to the converted image G2 by using the external parameters (R1, T1) of the omnidirectional camera 5 and the relative postures between each surface of the virtual cube represented by the converted image G2 and the omnidirectional panoramic image G1. That is, the conversion unit 60 calculates the external parameters of the omnidirectional camera 5 that captured each converted image G2 as the conversion external parameters (R2, T2) assuming that each converted image G2 was captured with a finite field of view. When the external parameters of the omnidirectional camera 5 are (R1, T1) and the conversion external parameters corresponding to the converted image G2 of the omnidirectional camera 5 are (R2, T2), the conversion external parameters (R2, T2) can be calculated by the following formula using the inverse matrix Inv(R).

[0142] (R2,T2)={Inv(R2)}R1(R1,T1)

[0143] The first restoration unit 16 estimates the three-dimensional distribution of the feature points based on the second feature points associated between a plurality of converted images G2, which are frames captured at different times (at different positions), and the conversion external parameters (R2, T2) of the omnidirectional camera 5 in the converted image G2.

[0144] Here, it is assumed that there exist a second feature point P2a in the converted image G2a corresponding to the conversion external parameters (R2a, T2a) and a second feature point P2b in the converted image G2b corresponding to the conversion external parameters (R2b, T2b). Further, it is assumed that there exists a point P0 in the real space of the feature points indicated by these second feature points P2a and P2b.

[0145] In this case, the intersection point of the projection line L2a connecting the conversion external parameters (R2a, T2a) and the second feature point P2a and the projection line L2b connecting the conversion external parameters (R2b, T2b) and the second feature point P2b is recorded as the three-dimensional coordinates of the point P0. By executing this process for a plurality of second feature points, the three-dimensional distribution of the feature points is estimated, and first density point group data, which is sparse three-dimensional point group data, is generated.

[0146] The second restoration unit 19 estimates the three-dimensional shape of the object in the converted image G2 based on the first density point group data and the second feature points associated with each other among the plurality of converted images G2, and generates second density point group data, which is dense three-dimensional point group data.

[0147] Note that the first restoration unit 16 and the second restoration unit 19 may generate colored three-dimensional point group data based on the color information of the converted image G2.

[0148] The information mapping unit 21 associates the pixels of the measurement image and the converted image G2 with each other, and superimposes the measurement image on the second density point group data.

[0149] Next, the three-dimensional data processing executed by the three-dimensional data processing system 1 will be described using the flowchart of FIG. 13. In the three-dimensional data processing of the second embodiment, steps S3A to S3B are added to the three-dimensional data processing (FIG. 10) of the first embodiment, and the other steps are substantially the same as those of the first embodiment.

[0150] As shown in FIG. 13, in step S3A that proceeds to the next step after step S3, the conversion unit 60 (FIG. 11) generates a converted image G2 (FIG. 12) obtained by converting the omnidirectional image, i.e., the full-sphere panoramic image G1 (FIG. 12), into a cube map format. At this time, since the first feature points of the full-sphere panoramic image G1 are also projected onto the converted image G2, the coordinates of the second feature point P2 in the converted image G2 and the association of the second feature points among the plurality of converted images G2 are also known.

[0151] In the next step S3B, the conversion unit 60 calculates the conversion external parameters (R2, T2) of the omnidirectional camera 5 corresponding to each surface of the converted image G2. This calculation is performed based on the external parameters (R1, T1) of the omnidirectional camera 5 and the relative postures of the full-sphere panoramic image G1 and the converted image G2.

[0152] In the subsequent steps S4 to S10, the "omnidirectional image" in the 3D data processing (FIG. 10) of the first embodiment is simply replaced with the "converted image G2", and the "external parameters (R1, T1)" are replaced with the "conversion external parameters (R2, T2)", and the processing is substantially the same.

[0153] In the second embodiment, in addition to the same effects as those of the first embodiment, the three-dimensional point cloud data is generated using the converted image G2 without distortion. Therefore, the accuracy of the three-dimensional point cloud data can be further improved. Furthermore, since both the measurement image and the converted image G2 are flat images, the accuracy of associating these pixels with each other can be improved compared to the case of using the spherical full-sphere panoramic image G1.

[0154] (Third Embodiment) Next, the third embodiment will be described with reference to FIG. 14. Note that the same components as those shown in the above-described embodiments are denoted by the same reference numerals, and redundant descriptions are omitted.

[0155] The processing circuit 7 of the third embodiment is the same as the configuration (FIG. 11) of the second embodiment.

[0156] The conversion unit 60 (FIG. 11) first calculates the external conversion parameters (R2, T2). Then, the feature point extraction unit 13 newly extracts the second feature points in the converted image G2 and associates the same second feature points among the plurality of converted images G2. Next, the feature point extraction unit 13 calculates the coordinates of the second feature point P2 in the converted image G2 based on the external conversion parameters (R2, T2) and the second feature points associated among the plurality of converted images G2.

[0157] Next, the 3D data processing executed by the 3D data processing system 1 will be described with reference to the flowchart of FIG. 14. In the 3D data processing of the third embodiment, steps S3C to S3D are added to the 3D data processing (FIG. 13) of the second embodiment, and the other steps are substantially the same as those of the second embodiment.

[0158] As shown in FIG. 14, in step S3C that follows step S3B, the feature point extraction unit 13 extracts the second feature points in the converted image G2 and associates the same second feature points among the plurality of converted images G2.

[0159] In the next step S3D, the feature point extraction unit 13 calculates the coordinates of the second feature point P2 in the converted image G2 based on the external conversion parameters (R2, T2) and the second feature points associated among the plurality of converted images G2.

[0160] The subsequent steps S4 to S10 are the same as the 3D data processing (FIG. 13) of the second embodiment.

[0161] In the third embodiment, when the conversion unit 60 converts the omnidirectional panoramic image G1 into the converted image G2, it is not necessary to project the first feature points in the omnidirectional panoramic image G1 onto the converted image G2. Therefore, the processing load can be reduced.

[0162] (Fourth Embodiment) Next, a fourth embodiment will be described with reference to FIGS. 15 to 17. Note that the same components as those shown in the above-described embodiments are denoted by the same reference numerals, and redundant descriptions thereof are omitted.

[0163] As shown in FIG. 15, the processing circuit 7 of the fourth embodiment has an object detection unit 61 in addition to the configuration of the first embodiment (FIG. 3). These functions of the processing circuit 7 are realized by a program stored in a memory, HDD, or SSD being executed by a CPU.

[0164] In the fourth embodiment, a predetermined object present at the inspection location is used instead of the marker 32 (FIG. 4). For example, an arbitrary type of object is provided in a state fixed in advance at the location where the moving body 2 (FIG. 4) travels. When the coordinates of this arbitrary type of object are newly recorded, the object becomes a field marker as a specific subject. Further, such an object may be stored in advance as a reference object.

[0165] In the fourth embodiment, the information regarding the feature points includes the information regarding the reference object. For example, the control computer 3 stores the three-dimensional position in the physical space 40 of at least one reference object that exists in a state fixed in advance at the location photographed by the omnidirectional camera 5. Then, the control computer 3 performs correction to match the coordinate system and scale of the first density point group data with the absolute coordinate system and absolute scale of the physical space 40 based on the three-dimensional position and scale of the reference object shown in the omnidirectional image in the physical space 40.

[0166] As shown in Fig. 17, various types of objects can be used as on-site markers. For example, the objects that can be on-site markers include devices such as motors or pumps 70, instruments 71 or valves 72 provided on the devices, nameplates 73 attached to the devices to display the brand names of the devices, control panels or power distribution panels, and other members such as support structures. That is, structures that do not move in position at the inspection location and whose shapes and dimensions are known become on-site markers. From the position, the orientation information of each structure, and the dimension information, the absolute coordinates and the absolute scale can be defined. These objects appearing in the image of the inspection location captured by the omnidirectional camera 5 are processed by image recognition technology, and the positions of the objects in the image are specified. The types of on-site markers are set in advance by the user.

[0167] In the fourth embodiment, the on-site marker management table is stored in the storage unit 8 (Fig. 2). As shown in Fig. 16, in the on-site marker management table, the type of the object, the location ID indicating the building 30 where the object is installed, and the coordinates of the object are registered in association with the marker ID, which is identification information that can individually identify the object that is the on-site marker. These pieces of information are registered each time an arbitrarily determined object is detected from the image of the omnidirectional camera 5. That is, the storage unit 8 stores the positions and types of arbitrary objects.

[0168] Also, the coordinates of the objects registered in the on-site marker management table are the coordinates of the preset physical space 40 (Fig. 5). That is, the types and coordinates of the objects are stored in the storage unit 8 (Fig. 2) in association with the identification information that can individually identify the objects. In this way, the coordinates in the physical space 40 can be specified from the objects appearing in the image captured by the omnidirectional camera 5, and the position of the moving body 2 can be estimated from those objects.

[0169] For example, the on-site marker first registered when the moving body 2 travels through the inspection location can be used to estimate its own position when the moving body 2 travels through the inspection location next.

[0170] Next, the process flow of correcting the coordinate system and scale of the environment map 50 (Fig. 5) including the first density point group data will be described in detail. Note that the above-mentioned drawings may be referred to.

[0171] In the fourth embodiment, in addition to the processing of the first embodiment, processing for detecting an object of the type that becomes a field marker is executed. Here, the object detection unit 61 detects a rectangular region (Fig. 17) of an image that surrounds the object in the image captured by the omnidirectional camera 5. Further, the object detection unit 61 tracks the object appearing in each frame constituting the image.

[0172] The object detection unit 61 estimates the position of the object. Here, the position of the object estimated by the object detection unit 61 is first recorded in the environment map 50. Since this environment map 50 is corrected to match the physical space 40, the coordinates of the object after the correction will match the physical space 40.

[0173] That is, when the object detection unit 61 detects an arbitrary type of object from the image captured by the omnidirectional camera 5, three-dimensional coordinates indicating the position of the object in the physical space 40 are calculated based on the coordinates of the known marker 32 or the known field marker. The coordinates of the object after this correction are registered in the field marker management table (Fig. 16) as a new field marker (specific subject). In this way, the newly detected object can be used as a field marker.

[0174] Thereafter, the coordinates of the field marker (object) are used for estimating the self-position of the mobile body 2 in the same manner as the coordinates of the marker 32. In this way, even when the marker 32 cannot be detected, the self-position of the mobile body 2 can be estimated using the object instead of the marker 32.

[0175] In the fourth embodiment, the user attaches a small number of markers 32 to the wall surface of the building 30, or acquires the types of objects fixed at the inspection locations and inputs them to the control computer 3. By simply performing these operations, the user can correct the scale and coordinate system of the environmental map 50. Therefore, the work that has conventionally been corrected manually is reduced, and the operation of the moving body 2 that performs automatic inspection becomes easier.

[0176] (Fifth Embodiment) Next, the fifth embodiment will be described with reference to FIGS. 18 to 19. Note that the same components as those shown in the above-described embodiments are denoted by the same reference numerals, and redundant descriptions are omitted.

[0177] As shown in FIG. 18, the processing circuit 7 of the fifth embodiment does not have a marker detection unit 14 (FIG. 3). Instead, it has an object detection unit 61 and an arrangement diagram adjustment unit 62. These functions of the processing circuit 7 are realized by a program stored in a memory, HDD, or SSD being executed by a CPU.

[0178] In the fifth embodiment, the marker 32 (FIG. 4) of the first embodiment is not used, and an object (on-site marker: specific subject) present at the inspection location is used instead. Also, an object management table (FIG. 19) is stored in the storage unit 8 (FIG. 2). What is registered in this object management table is information regarding any type of object in a fixed state in advance.

[0179] As shown in FIG. 19, the object management table associates an object ID, which is identification information that can individually identify an object (on-site marker), with the type of the object, a location ID indicating the building 30 where the object is provided, and the coordinates of the object.

[0180] Also, the coordinates of the objects registered in the object management table are coordinates in a preset physical space 40 (FIG. 5). In this way, the coordinates in the physical space 40 can be specified from the objects shown in the image captured by the omnidirectional camera 5, and the position of the moving body 2 can be estimated from the objects.

[0181] In the fifth embodiment, the information regarding the feature points includes the layout diagram. For example, the storage unit 8 (Fig. 2) of the control computer 3 stores the layout diagram of the objects arranged at the locations photographed by the omnidirectional camera 5.

[0182] Here, the layout diagram adjustment unit 62 fits at least one of the position and orientation of the omnidirectional camera 5 during photographing and the feature points to the layout diagram. By this fitting, the layout diagram adjustment unit 62 corrects the coordinate system and scale of the first density point cloud data to match the absolute coordinate system and absolute scale of the physical space 40.

[0183] For example, the time-series data of the photographing positions of the omnidirectional camera 5 corresponds to the movement route when photographing with the omnidirectional camera 5. Therefore, by superimposing and correcting the coordinates and scale on the layout diagram of the inspection location, it becomes possible to correct the first density point cloud data with respect to the absolute coordinate system and absolute scale. At this time, the data of the feature points may also be superimposed on the layout diagram to correct the coordinate system and scale. In addition, since the positions of the walls and the surface portions of the devices tend to have dense feature points, it is possible to accurately align them when superimposing them on the layout diagram.

[0184] Next, the flow of the correction process for the coordinate system and scale of the environment map 50 (Fig. 5) including the first density point cloud data will be described in detail. Note that the above-mentioned drawings may be referred to.

[0185] In the fifth embodiment, in addition to the process of extracting the feature points of the first embodiment, a process of detecting a preset object as a field marker is executed.

[0186] The object detection unit 61 detects an object (field marker) capable of obtaining the relative coordinates with respect to the omnidirectional camera 5 from the omnidirectional image photographed by the omnidirectional camera 5. The object to be detected is registered in an object management table (Fig. 19) in advance. Here, the object detection unit 61 detects a rectangular region of the image (Fig. 17) surrounding the object in the image.

[0187] The object detection unit 61 tracks an object captured in each frame constituting an image. When the object cannot be detected from the image, the object detection unit 61 estimates the coordinates of the object from the amount of movement of the moving body 2. Then, based on the estimated coordinates of the object and the environmental map 50, the object detection unit 61 specifies a range where the object may be captured in the image, and detects the object from the specified range. In this way, the accuracy of detecting the object from the image can be improved.

[0188] Here, the coordinates of the object and the shooting position, which are the object detection results, are recorded in the environmental map 50. Further, the coordinates of the location where each object is fixed are pre-recorded in a coordinate system that the user wants to manage. This coordinate system that the user wants to manage is, for example, the coordinate system of the physical space 40.

[0189] The relative coordinates of the shooting position obtained from the object are associated with the coordinates of the location where the object is fixed. Here, the shooting position obtained from the object is recorded in the environmental map 50.

[0190] Also, the type and coordinates of the object are recorded on the layout diagram, 3D point cloud data, or 3D CAD. Further, the object detection unit 61 calculates the position (two-dimensional coordinates) and type of the object in the image. The results of these recordings and calculations may be compared and processed as a PnP problem. For example, the relationship between the three-dimensional coordinates of the same type of object and the position in the image may be analyzed, and the shooting position of the omnidirectional camera 5 when the object is photographed may be estimated.

[0191] Also, when an object not described in the layout diagram is detected and the same object is detected from a plurality of frames, the three-dimensional coordinates of the object may be estimated from the shooting position of the omnidirectional camera 5, and the estimated result may be recorded.

[0192] In the fifth embodiment, the user acquires the type and coordinates of the object and inputs them to the control computer 3. By simply performing this operation by the user, the scale and coordinate system of the environmental map 50 can be corrected. Therefore, the work that has been conventionally corrected manually is reduced, and the operation of the moving body 2 that performs automatic inspection becomes easier.

[0193] As described above, the present invention has been described based on the first to fifth embodiments. However, the configuration applied in any one of the embodiments may be applied to other embodiments, or the configurations applied in each embodiment may be combined.

[0194] In the above example, the control computer 3 constituting the three-dimensional data processing system 1 is illustrated as performing various processes (including various processes such as setting). However, other modes may also be possible. For example, among the above-described various processes, a user may perform some of the processes, and the computer may receive the input of the processing result and use it for its own processing.

[0195] In addition, an image recognition technique based on machine learning may be used for detecting an object in an image by the object detection unit 61. For example, the image of the control panel is machine-learned, and when the user designates the width (scale) of the control panel, the width can be automatically recognized from the control panel in the image. Furthermore, a spatial recognition technique based on machine learning may be used for setting the range by the range setting unit 18.

[0196] For example, the three-dimensional data processing system 1 may include a computer equipped with artificial intelligence (AI: Artificial Intelligence) that performs machine learning. In addition, the three-dimensional data processing system 1 may include a deep learning unit that extracts a specific pattern from a plurality of patterns based on deep learning.

[0197] For the analysis using this computer, an analysis technique based on the learning of artificial intelligence can be used. For example, a learned model generated by machine learning using a neural network, other learned models generated by machine learning, a deep learning algorithm, a mathematical algorithm such as regression analysis can be used. In addition, the forms of machine learning include forms such as clustering and deep learning.

[0198] The three-dimensional data processing system 1 may be composed of one computer equipped with a neural network, or may be composed of a plurality of computers equipped with neural networks.

[0199] Here, a neural network is a mathematical model that expresses the characteristics of brain functions through computer simulation. For example, it shows a model in which artificial neurons (nodes) forming a network through synaptic connections change the synaptic connection strength through learning and acquire problem-solving ability. Furthermore, the neural network acquires problem-solving ability through deep learning.

[0200] An intermediate layer having a plurality of layers is provided in the neural network. Each layer of this intermediate layer is composed of a plurality of units. Also, by preliminarily training a multi-layer neural network using learning data (teacher data), it is possible to automatically extract feature amounts existing in the pattern of changes in the state of a circuit or system. Note that in a multi-layer neural network, on the user interface, an arbitrary number of intermediate layers, an arbitrary number of units, an arbitrary learning rate, an arbitrary number of learning times, and an arbitrary activation function can be set.

[0201] Note that a reward function may be set for each type of information item to be learned, and deep reinforcement learning in which the information item with the highest value is extracted based on the reward function may be used for the neural network.

[0202] For example, a CNN (Convolution Neural Network) with proven performance in image recognition is used. In this CNN, the intermediate layer is composed of a convolutional layer and a pooling layer. The convolutional layer obtains a feature map by performing filter processing on nodes close to it in the previous layer. The pooling layer further shrinks the feature map output from the convolutional layer to obtain a new feature map. At this time, by obtaining the maximum value of the pixels included in the region of interest in the feature map, it is possible to absorb any slight deviation in the position of the feature amount.

[0203] The convolutional layer extracts local features of an image, and the pooling layer performs a process of summarizing local features. In these processes, the image is downsized while maintaining the features of the input image. That is, in a convolutional neural network (CNN), the amount of information contained in the image can be greatly compressed (abstracted). Then, using the abstracted image image stored in the neural network, the input image can be recognized and the image can be classified.

[0204] Note that there are various methods in machine learning, such as autoencoders, LSTM (Long Short-Term Memory), SDF (Signed Distance Function), GAN (Generative Adversarial Network), and RNN (Recurrent Neural Network). These methods may be applied to the above-described embodiments.

[0205] The learned model includes an input layer, an intermediate layer, and an output layer. Input data is input to the input layer. The parameters of the intermediate layer are pre-trained by machine learning using training data. The output layer outputs output data indicating the result processed by the intermediate layer in response to the input of the input data to the input layer. Also, the learned model is machine-learned using training data that takes at least one of the input data or data simulating it as input and at least one of the output data or data simulating it as output.

[0206] Note that in the above flowchart, an example form in which each step is executed serially is illustrated, but the order of each step is not necessarily fixed, and the order of some steps may be reversed. Also, some steps may be executed in parallel with other steps.

[0207] The aforementioned control computer 3 includes a control device, a storage device, an output device, an input device, and a communication interface. Here, the control device includes highly integrated processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), and a dedicated chip. The storage device includes a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc. The output device includes a display panel, a head-mounted display, a projector, a printer, etc. The input device includes a mouse, a keyboard, a touch panel, etc. This control computer 3 can be realized with a hardware configuration using an ordinary computer.

[0208] Note that the program or the learned model executed by the aforementioned control computer 3 is provided by being pre-embedded in a ROM or the like. Additionally or alternatively, this program or learned model is provided by being stored in a computer-readable non-transitory storage medium as an installable or executable file. This storage medium includes a CD-ROM, a CD-R, a memory card, a DVD, a flexible disk (FD), etc.

[0209] Also, the program or the learned model executed by this control computer 3 may be stored in a computer connected to a network such as the Internet and provided by being downloaded via the network. That is, the program or the learned model may be provided from cloud computing resources. Also, a server on the cloud may execute the program or the learned model and only the processing result thereof may be provided via the cloud. Also, this control computer 3 can be configured by connecting separate modules that independently perform each function of the components to each other with a network or a dedicated line and combining them.

[0210] According to at least one embodiment described above, by providing the coordinate scale correction unit 17, in the three-dimensional data processing using the omnidirectional image acquired by the omnidirectional camera 5, the correction work of the coordinate system and the scale can be simplified.

[0211] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, changes, and combinations can be made without departing from the gist of the invention. These embodiments or their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and its equivalent scope.

Explanation of Reference Numerals

[0212] 1... Three-dimensional data processing system, 2... Moving body, 3... Control computer, 4... Wheels, 5... Omnidirectional camera, 6... Measuring camera, 7... Processing circuit, 8... Storage unit, 9... Input unit, 10... Output unit, 11... Communication unit, 12... Image input unit, 13... Feature point extraction unit, 14... Marker detection unit, 15... Position and orientation estimation unit, 16... First restoration unit, 17... Coordinate scale correction unit, 18... Range setting unit, 19... Second restoration unit, 20... Parallax correction unit, 21... Information mapping unit, 22... Display control unit, 30... Building, 31... Inspection target object, 32... Marker, 33... Code, 34... Mounting board, 40... Physical space, 41... Coordinates, 42... Movement route, 43... Coordinates, 50... Environment map, 51... Coordinates, 52... Movement route, 53... Coordinates, 60... Conversion unit, 61... Object detection unit, 62... Layout adjustment unit, 70... Pump, 71... Instrument, 72... Valve, 73... Nameplate.

Claims

1. A plurality of omnidirectional images captured by an omnidirectional camera at at least a first position and a second position different from the first position are input, and a plurality of feature points indicating characteristic parts of at least one object shown in the omnidirectional images are extracted. Based on the correspondence between the feature points and the omnidirectional images, the position and orientation of the omnidirectional camera at the time of shooting are estimated. Based on the feature points and the position and orientation of the omnidirectional camera at the time of shooting, first density point cloud data indicating the three-dimensional distribution of the feature points of the first density is restored. Based on at least one of the position and orientation of the omnidirectional camera at the time of shooting and the information regarding the feature points, correction is performed to match the coordinate system and scale of the first density point cloud data with the absolute coordinate system and absolute scale of the actual physical space in which shooting was performed by the omnidirectional camera. Based on the correspondence between the corrected first density point cloud data and the omnidirectional images, second density point cloud data indicating the three-dimensional distribution of the feature points of a second density denser than the first density is restored. A three-dimensional data processing system comprising one or more computers, configured as follows.

2. The computer maps the omnidirectional image onto each surface of a virtual cube set in advance to an arbitrary size and generates a transformed image with distortion corrected. Based on the position and orientation of the omnidirectional camera at the time of shooting and the correspondence between the relative orientations of the omnidirectional image and each surface of the virtual cube, virtual camera positions and orientations corresponding to each surface of the virtual cube are calculated. projects the feature points onto the transformed image. configured as follows. The three-dimensional data processing system according to Claim 1.

3. The computer stores a layout diagram of the object arranged at the location where shooting is performed by the omnidirectional camera. By fitting at least one of the position and orientation of the omnidirectional camera at the time of shooting and the feature points to the layout diagram, correction is performed to match the coordinate system and scale of the first density point cloud data with the absolute coordinate system and absolute scale of the physical space. configured as follows. The three-dimensional data processing system according to Claim 1 or Claim 2.

4. The computer stores the three-dimensional position in the physical space of at least one specific subject that exists in a fixed state in advance at the location where shooting is performed by the omnidirectional camera. Based on the three-dimensional position of the specific subject in the physical space captured in the omnidirectional image, perform correction to align the coordinate system and scale of the first density point cloud data with the absolute coordinate system and absolute scale of the physical space. Is configured as follows. The three-dimensional data processing system according to claim 1 or claim 2.

5. The computer Stores the three-dimensional position of at least one reference object existing in a state fixed in advance at the location photographed by the omnidirectional camera in the physical space, Based on the three-dimensional position and scale of the reference object captured in the omnidirectional image in the physical space, perform correction to align the coordinate system and scale of the first density point cloud data with the absolute coordinate system and absolute scale of the physical space. Is configured as follows. The three-dimensional data processing system according to claim 1 or claim 2.

6. The computer Based on the position and orientation at the time of shooting the omnidirectional image adjusted to the absolute coordinate system and absolute scale of the physical space, set the range for restoring the second density point cloud data, Trim the feature points outside the range and restore the second density point cloud data within the range. Is configured as follows. The three-dimensional data processing system according to claim 1 or claim 2.

7. A plurality of omnidirectional images captured by an omnidirectional camera at at least a first position and a second position different from the first position are input, and a plurality of feature points indicating characteristic parts of at least one object captured in the omnidirectional image are extracted. Estimate the position and orientation of the omnidirectional camera at the time of shooting from the correspondence relationship between the feature points and the omnidirectional image. Based on the feature points and the position and orientation of the omnidirectional camera at the time of shooting, restore first density point cloud data indicating the three-dimensional distribution of the feature points. Based on at least one of the position and orientation of the omnidirectional camera at the time of shooting and the information regarding the feature points, perform correction to align the coordinate system and scale of the first density point cloud data with the absolute coordinate system and absolute scale of the actual physical space where shooting is performed by the omnidirectional camera. Restore second density point cloud data indicating the three-dimensional distribution of the feature points with a second density denser than the first density from the correspondence relationship between the corrected first density point cloud data and the omnidirectional image. One or more computers execute the process. Three-dimensional data processing method.

Citation Information

Patent Citations

  • Controlling of copying welding

    JP1984004976A