Video management system, method, and program

The video management system addresses the challenge of reconstructing 3D point clouds in large spaces by employing omnidirectional cameras to capture and link feature points, facilitating quick and accurate 3D point cloud creation and shooting position estimation.

JP2026021906APending Publication Date: 2026-02-12KK TOSHIBA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123139
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately reconstructing 3D point clouds for wide areas like power plants or factories due to the need for labor-intensive manual linking of video footage and shooting positions, and existing methods like laser scanning and 2D image reconstruction are time-consuming or inaccurate, especially when indoor movements cause significant field-of-view changes.

Method used

A video management system using an omnidirectional camera to capture celestial sphere frames, extract feature points, define a 3D point cloud, and associate shooting positions and directions, enabling quick and accurate reconstruction of 3D point clouds by linking feature points across frames and converting orientations.

Benefits of technology

Enables rapid creation of accurate 3D point clouds of wide spaces with structures, allowing precise estimation of shooting positions and directions, even in indoor environments, by using omnidirectional video and feature point extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021906000001_ABST
    Figure 2026021906000001_ABST
Patent Text Reader

Abstract

To provide a video management technology for creating a 3D point group of a space where structures are arranged in a wide range in a short time, and for highly accurately estimating the position of a photographing object.SOLUTION: The video management system 10 includes a first acquiring unit 21 that acquires an omnidirectional video image 15 captured by a first image-capturing part 11, a clipping unit 18 that clips a reference planar image 17 from a celestial sphere frame 16 in a designated azimuth 19 defined in a first coordinates system of the first image-capturing part 11, an extracting unit 25 that extracts a feature point 26 defined in the first coordinates system from the celestial sphere frame 16, a defining unit 27 that defines the same feature point 26 in a plurality of celestial sphere frames 16 as a 3D point cloud 23 in a second coordinates system and defines an image-capturing position 29 and posture information 37 of the first image-capturing part 11 that captures each celestial sphere frame 16 in the second coordinates system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to a technology for managing video captured by a camera. [Background technology]

[0002] At power plants and factories, patrol inspections are conducted periodically to manage construction progress and check for abnormalities. Traditionally, patrol inspections involve manual work records, as well as video recordings taken with digital cameras and smartphones. These video records are managed as inspection history by recording information such as the location of the footage along with the footage of the subject.

[0003] When patrol inspections are carried out outdoors, it is possible to record the shooting position and direction using a positioning system such as a GPS (Global Positioning System). On the other hand, when patrol inspections are carried out indoors, the shooting position and direction must be recorded by another means. In this case, a large amount of video of the subject and the shooting position and direction will be recorded separately, and the task of linking the two for management purposes becomes enormous and labor-intensive.

[0004] A known technique for estimating the shooting position and direction from video of a target object is as follows. This technique involves building a database that stores reference video footage shot in advance and links it to the shooting position and direction. This known technique searches the database for video footage similar to that shot during periodic inspection patrols. The shooting position and direction linked to the searched reference video are then estimated as the shooting position and direction of the corresponding video footage. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Publication No. 2022-105442 Summary of the Invention [Problem to be solved by the invention]

[0006] According to the above-mentioned conventional technology, the constructed database defines the reference image and the shooting position and direction as a 3D point cloud. To create a 3D point cloud of the site space that is the subject of the patrol inspection, 3D measurement by laser scanning and technology for reconstructing from multiple planar images are required.

[0007] Of these, laser scanning technology has the ability to acquire high-precision, high-density 3D point clouds that contain scale information. However, when the target area is wide, such as a power plant or factory, laser scanning requires 3D measurements to be taken at multiple locations, and creating a 3D point cloud for the entire site can be time-consuming.

[0008] Furthermore, technology for reconstructing 3D point clouds from multiple 2D images cannot accurately reconstruct 3D point clouds unless the positional relationships between a large number of 2D images can be accurately determined. For this reason, when shooting indoors or at close range, even a small movement can result in a large amount of movement in the image, the field of view can change significantly when the subject turns around, the positional relationship cannot be determined when the continuity of the field of view between images is broken, and the accuracy of the positional relationship deteriorates where the area of ​​continuity is narrow, posing challenges in accurately reconstructing 3D point clouds.

[0009] The embodiments of the present invention have been developed with these circumstances in mind, and aim to provide a video management technology that can quickly create a 3D point cloud of a space where structures are located over a wide area, and estimate the position of the subject being photographed with high accuracy. [Means for solving the problem]

[0010] a first acquisition unit that acquires omnidirectional video captured by moving a first image capture unit in a space in which a structure is located while the first image capture unit is moved to cover the entire space; a cropping unit that crops out a reference planar image from each of a plurality of celestial sphere frames that constitute the omnidirectional video, based on a specified orientation defined in a first coordinate system of the first image capture unit; an extraction unit that extracts feature points of the structure defined in the first coordinate system from each of the celestial sphere frames; a definition unit that defines the same feature points across the plurality of celestial sphere frames in a second coordinate system as a 3D point cloud of the space, and defines in the second coordinate system the shooting position and attitude information of the first image capture unit that captured each of the celestial sphere frames; a conversion unit that converts the specified orientation into a shooting direction defined in the second coordinate system based on the attitude information; and a database that registers the reference planar images in association with the corresponding shooting positions and shooting directions. [Effects of the Invention]

[0011] An embodiment of the present invention provides a video management technology that quickly creates a 3D point cloud of a space where structures are arranged over a wide area, and estimates the position of the subject being photographed with high accuracy. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a block diagram of a video management system according to a first embodiment of the present invention. [Figure 2] FIG. 10 is a block diagram of a video management system according to a second embodiment of the present invention. [Figure 3] The first flat image was recorded by the second camera unit, showing the space at the inspection site. [Figure 4] The second flat image was recorded by the second camera unit, showing the space at the inspection site. [Figure 5] 1 is an area map of the inspection site showing the shooting positions and directions of the first and second planar images. [Figure 6] 1 is a flowchart illustrating steps of a video management method and an algorithm of a video management program according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] (First embodiment) Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. Fig. 1 is a block diagram of a video management system 10A(10) according to a first embodiment of the present invention. The video management system 10A(10) includes a first acquisition unit 21 that acquires omnidirectional video 15 captured by moving a first image capture unit 11 in a space where a structure is located, so as to cover the entire space; a cropping unit 18 that crops out a reference planar image 17 from each of a plurality of celestial sphere frames 16 based on a specified orientation 19 defined in a first coordinate system of the first image capture unit 11; and an extraction unit 25 that extracts feature points 26 of the structure defined in the first coordinate system from each of the celestial sphere frames 16.

[0014] The video management system 10A(10) further includes a definition unit 27 that defines identical feature points 26 in multiple celestial sphere frames 16 as a spatial 3D point cloud 23 in a second coordinate system and defines, in the second coordinate system, a shooting position 29 and attitude information 37 of the first image capture unit 11 that captured each of the celestial sphere frames 16; a conversion unit 38 that converts the specified orientation 19 into a shooting direction 39 defined in the second coordinate system based on the attitude information 37; and a database 36 that registers reference planar images 17 in association with the corresponding shooting positions 29 and shooting directions 39.

[0015] The first imaging unit 11 is generally called an omnidirectional camera (360° camera), and any commercially available product may be used as long as it is capable of capturing omnidirectional video 15. The omnidirectional video 15 is a representation of the entire video scene as a celestial sphere panoramic video. This celestial sphere panoramic video can be displayed on a plane by projecting it using equirectangular projection, for example. The omnidirectional video 15 is not limited to a video, and may be a still image captured intermittently while moving.

[0016] The space photographed by the first photographing unit 11 is assumed to be an indoor space such as a power plant or factory, where structures are widely arranged (see FIG. 5). Patrol inspections are periodically carried out in the space to manage construction progress and check for abnormalities. Such patrols are not limited to being carried out by personnel, but may also be carried out by inspection robots. The inspector or inspection robot carries the first photographing unit 11 and photographs the surroundings and the equipment to be inspected while moving around.

[0017] In this way, the first image capturer 11 captures omnidirectional images 15 that cover the entire space while moving through the space in which the structure is located. The first acquisition unit 21 then acquires the omnidirectional images 15 temporarily stored in a storage medium (not shown) all at once, or continuously acquires the omnidirectional images 15 from the first image capturer 11 while it is capturing images. The acquired omnidirectional images 15 are composed of a time series of multiple celestial sphere frames 16 along the movement route of the first image capturer 11.

[0018] The extraction unit 25 extracts feature points 26 of structures from each of the multiple celestial sphere frames 16. The extraction unit 25 extracts omnidirectional feature points 26 from each of the celestial sphere frames 16. Because the multiple celestial sphere frames 16 are omnidirectional, feature points 26 are extracted from the multiple celestial sphere frames 16 without changing the field of view of the omnidirectional image 15. This makes it easy for the determination unit 27, described below, to compare feature points 26 in areas with relatively little movement or when turning around. This maintains continuity of the field of view between the multiple celestial sphere frames 16, allowing a 3D point cloud 23 to be reconstructed with high accuracy, even for indoor objects.

[0019] Here, the feature points 26 correspond to pixels that correspond to the corners of the structure among the pixels that make up the celestial sphere frame 16. These corners are characterized by pixel brightness being high in all directions. To extract the feature points, for example, ORB (Oriented Fast and Rotated Brief) is used. Furthermore, in each of the celestial sphere frames 16, the feature points 26 are defined in the first coordinate system of the first image capture unit 11.

[0020] The cropping unit 18 crops out a partial reference planar image 17, which has a flat surface, from each of the celestial sphere frames 16, which have a spherical surface. The portion of the cropped reference planar image 17 that has been converted from a spherical surface to a flat surface is corrected using a known method. The position from which the reference planar image 17 is cropped from the celestial sphere frame 16 is specified by a specified orientation 19 defined in the first coordinate system of the first imaging unit 11. This specified orientation 19 is set in steps in all directions based on the center point of the celestial sphere frame 16. Therefore, multiple reference planar images 17, each specified for one of the multiple specified orientations 19, are cropped from one celestial sphere frame 16.

[0021] The determination unit 27 compares the feature points 26 between consecutive celestial sphere frames 16 to detect identical feature points. The correspondence between the feature points 26 can be determined based on a similarity calculation. Examples of similarity calculation methods include Brute Force, Brute Force-L1, Brute Force-Hamming, and Flann-Based.

[0022] Furthermore, the definition unit 27 generates a 3D point cloud 23 that defines a space from the feature points 26 that make up each of the multiple celestial sphere frames 16. The second coordinate system that defines this 3D point cloud 23 does not originally have an absolute scale, but may be aligned with a global coordinate system that represents the space in which the structure is located.

[0023] Each of the multiple celestial sphere frames 16 is associated with an imaging position 29 in space of the first image capture unit 11. The definition unit 27 defines the imaging position 29 of the first image capture unit 11 in the second coordinate system together with the 3D point cloud 23 based on the relative positional relationship of feature points 26 in the first coordinate system in the multiple celestial sphere frames 16.

[0024] Similarly, the definition unit 27 also defines the attitude information 37 of the first image capture unit 11 in the second coordinate system together with the 3D point cloud 23 based on the relative positional relationships of the feature points 26 in the first coordinate system in multiple celestial spherical frames 16.

[0025] Note that a known technique such as SLAM (Simultaneous Localization and Mapping) is used to define the image capture position 29 and the orientation information 37. In particular, by using VSLAM (Visual Simultaneous Localization and Mapping), not only the image capture position 29 and the orientation information 37 but also the 3D point cloud 23 can be defined.

[0026] The conversion unit 38 converts the first coordinate system of the first image capturing unit 11 into the second coordinate system of the 3D point cloud 23. As a result, the specified orientation 19 defined in the first coordinate system is converted into an image capturing direction 39 defined in the second coordinate system based on the attitude information 37 defined in the second coordinate system.

[0027] Each of the multiple celestial sphere frames 16 is associated with an imaging position 29 of the first imaging unit 11 in space, and therefore the reference planar image 17 extracted from the common celestial sphere frame 16 is linked to the common imaging position 29. The specified orientation 19 defined in the first coordinate system of the first imaging unit 11 is also converted to an imaging direction 39 defined in the second coordinate system and linked to the reference planar image 17. As a result, the reference planar image 17 registered in the database 36 is linked to the corresponding imaging position 29 and imaging direction 39, and is therefore defined in the second coordinate system.

[0028] (Second embodiment) Next, a second embodiment of the present invention will be described with reference to FIGS. 2 to 5. FIG. 2 is a block diagram of a video management system 10B (10) according to the second embodiment of the present invention. The video management system 10B of the second embodiment is configured by adding the configuration described below to the configuration of the video management system 10A of the first embodiment described above. Note that in FIG. 2, the configuration common to the first and second embodiments is omitted by referencing the description in FIG. 1. The second embodiment is implemented after all of the reference planar images 17 linked to the shooting positions 29 and shooting directions 39 have been registered in the database 36 in the first embodiment.

[0029] The video management system 10B (10) further includes a configuration added to the video management system 10A, that is, a projection unit 43 that defines an area map 44 of the space onto which the 3D point cloud 23 is projected in a second coordinate system.

[0030] The 3D point cloud 23 generated in the first embodiment has a relative positional relationship and does not have a real scale or absolute coordinates. Therefore, in order to project the 3D point cloud 23 onto the area map 44, the projection unit 43 needs to align the coordinate system of the area map 44 with the second coordinate system, or align the second coordinate system with the coordinate system of the area map 44. Note that if the second coordinate system employs a global coordinate system that represents the space in which the structure is located, the coordinate system of the area map 44 may also employ the global coordinate system.

[0031] The video management system 10B (10) further adds components to the video management system 10A, and is equipped with a second acquisition unit 22 that acquires planar images 46 (46a, 46b) recorded by the second imaging unit 12 of structures arranged in space, a matching unit 48 that compares the planar images 46 with the database 36 and selects a reference planar image 17 that has a high similarity, a linking unit 47 that links the shooting position 29 and shooting direction 39 linked to the selected reference planar image 17 with the corresponding planar image 46, and a display unit 45 that displays the shooting position 29 and shooting direction 39 linked to the planar image 46 on the area map 44.

[0032] The second imaging unit 12 may be any device capable of capturing a still planar image 46, and may be a commercially available digital camera, smartphone, or the like. The planar image 46 may be a still image extracted from a video. The second acquisition unit 22 either acquires a plurality of planar images 46 temporarily stored in a separate storage medium all at once, or acquires the planar images 46 sequentially from the second imaging unit 12 currently capturing images.

[0033] The matching unit 48 matches the large number of reference planar images 17 stored in the database 36 with the planar video 46. Then, the matching unit 48 selects the reference planar image 17 that is most similar to the reference planar image 17 from the large number of reference planar images 17, and extracts the shooting position 29 and shooting direction 39 associated with this selected reference planar image 17.

[0034] By linking with a highly similar reference planar image 17 by the linking unit 47, the planar video 46 is considered to have been captured at the imaging position 29 and imaging direction 39 in the 3D point cloud 23 (second coordinate system).

[0035] However, each of the multiple reference planar images 17 specified by the specified orientations 19, which are set in steps in all directions, has a discrete and non-continuous representation. For this reason, even if the comparison unit 48 selects a reference planar image 17 as having a high similarity, the actual similarity with the planar image 46 may be insufficient. In this case, the correction unit 49 corrects the shooting position 29 and shooting direction 39 so as to improve the similarity of the reference planar image 17 to the planar image 46. This correction unit 49 can derive the optimal corrected shooting position 29 and shooting direction 39 of the planar image 46 using a deep neural network.

[0036] Fig. 3 shows a first planar image 46a (46) of the space of the inspection site recorded by the second photographing unit 12 (Fig. 2). Fig. 4 shows a second planar image 46b (46) of the space of the inspection site recorded by the second photographing unit 12 (Fig. 2). Fig. 5 shows an area map 44 of the inspection site, showing the photographing position 29a and photographing direction 39a of the first planar image 46a (Fig. 3) and the photographing position 29b and photographing direction 39b of the second planar image 46b (Fig. 4).

[0037] 5, the display unit 45 (FIG. 2) displays the shooting positions 29 (29a, 29b) and shooting directions 39 (39a, 39b) associated with the planar image 46 on the area map 44. Note that, when correction is performed by the correction unit 49, the display unit 45 displays the corrected shooting positions 29 (29a, 29b) and shooting directions 39 (39a, 39b) on the area map 44.

[0038] This makes it possible to easily create a database in which a space is captured in omnidirectional video 15 by a first capturing unit 11 such as an omnidirectional camera, a large number of reference planar images 17 are extracted from the video, and the reference planar images 17 are linked to the capturing positions 29 and capturing directions 39 defined in the coordinate system of the space. This makes it possible to recognize with high accuracy the capturing positions 29 and capturing directions 39 of planar images 46 captured by a second capturing unit 12 such as a digital camera or smartphone.

[0039] The steps of the video management method and the algorithm of the video management program according to this embodiment will be described with reference to the flowchart in Figure 6. First, omnidirectional video 15 is acquired by moving the first image capture unit 11 to capture video covering the entire space (S11). Next, a reference planar image 17 is extracted from each of a plurality of celestial sphere frames 16 based on a specified orientation 19 defined in a first coordinate system (S12). Furthermore, feature points 26 of structures defined in the first coordinate system are extracted from each of the celestial sphere frames 16 (S13).

[0040] Next, the same feature points 26 in multiple celestial sphere frames 16 are defined in a second coordinate system as a spatial 3D point cloud 23 (S14). Then, the shooting position 29 and attitude information 37 of the first image capturing unit 11 that captured each of the celestial sphere frames 16 are defined in the second coordinate system (S15).

[0041] Next, the specified orientation 19 defined in the first coordinate system is converted into an imaging direction 39 defined in the second coordinate system based on the attitude information 37 (S16). Then, the reference planar image 17 is linked to the corresponding imaging position 29 and imaging direction 39 and registered in the database 36 (S17).

[0042] Next, a planar image 46 (46a, 46b) of a structure arranged in space is acquired by recording the image using the second imaging unit 12 (S18). The planar image 46 is then compared with the database 36 to select a reference planar image 17 with a high similarity (S19).

[0043] Next, the photographing position 29 and the photographing direction 39 associated with the selected reference planar image 17 are linked to the corresponding planar video 46 (S20). Then, the photographing position 29 and the photographing direction 39 associated with the planar video 46 are displayed on the display unit 45 in the area map 44 (S21; END). Note that the area map 44 is defined in advance in the second coordinate system, with the 3D point cloud 23 projected thereon.

[0044] According to at least one of the embodiments of the video management system described above, by extracting feature points from a celestial spherical frame of omnidirectional video to generate a 3D point cloud and then cutting out a reference plane image in any specified direction from the celestial spherical frame, it is possible to quickly create a 3D point cloud of a space in which structures are arranged over a wide range, and to estimate the position of the subject being photographed with high accuracy.

[0045] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, modifications, and combinations can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as the inventions described in the claims and their equivalents.

[0046] The video management system described above includes a control device with a highly integrated processor such as a dedicated chip, FPGA (Field Programmable Gate Array), GPU (Graphics Processing Unit), or CPU (Central Processing Unit), a storage device such as ROM (Read Only Memory) or RAM (Random Access Memory), an external storage device such as HDD (Hard Disk Drive) or SSD (Solid State Drive), a display device such as a monitor, input devices such as a mouse and keyboard, and a communication I / F, and can be realized with a hardware configuration using a normal computer. Therefore, the components of the video management system can also be realized by a computer processor and can be operated by a video management program.

[0047] The video management program may be provided in advance by being embedded in a ROM, etc. Alternatively, the program may be provided by being stored in an installable or executable file format on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD, or flexible disk (FD).

[0048] The video management program according to this embodiment may be stored on a computer connected to a network such as the Internet and provided by downloading it via the network. The video management system may also be configured by connecting and combining separate modules that independently perform the functions of the components via a network or dedicated lines. [Explanation of symbols]

[0049] 10 (10A, 10B)...image management system, 11...capturing unit, 15...omnidirectional image, 16...celestial sphere frame, 17...reference planar image, 18...cutting unit, 19...specified orientation, 21...first acquisition unit, 22...second acquisition unit, 23...3D point cloud, 25...extraction unit, 26...feature points, 27...determination unit, 29 (29a, 29b)...capturing position, 36...database, 37...posture information, 38...conversion unit, 39 (39a, 39b)...capturing direction, 43...projection unit, 44...area map, 45...display unit, 46 (46a, 46b)...planar image, 47...linking unit, 48...matching unit, 49...correction unit.

Claims

1. a first acquisition unit that acquires omnidirectional video images that cover the entire space while moving a first image capture unit in the space where the structure is arranged; a cropping unit that crops out a reference plane image from each of a plurality of celestial sphere frames that constitute the omnidirectional video image based on a designated orientation defined in a first coordinate system of the first image capturing unit; an extractor that extracts feature points of the structure defined in the first coordinate system from each of the celestial sphere frames; a definition unit that defines the same feature points in a plurality of the celestial sphere frames as a 3D point cloud in a second coordinate system, and defines, in the second coordinate system, imaging position and attitude information of the first imaging unit that captured each of the celestial sphere frames; a conversion unit that converts the specified orientation into an imaging direction defined in the second coordinate system based on the attitude information; a database that registers the reference planar image in association with the corresponding shooting position and shooting direction.

2. 2. The video management system according to claim 1, A video management system comprising a projection unit that projects the 3D point cloud to define an area map of the space in the second coordinate system.

3. 3. The video management system according to claim 2, a second acquisition unit that acquires a planar image of the structure arranged in the space captured by a second imaging unit; a comparison unit that compares the planar image with the database and selects the reference planar image with high similarity; a linking unit that links the imaging position and the imaging direction associated with the selected reference planar image to the corresponding planar video; a display unit that displays the shooting position and the shooting direction linked to the planar image on the area map.

4. 4. The video management system according to claim 3, a correction unit that corrects the shooting position and the shooting direction so that the similarity of the reference planar image to the planar video image is improved; The display unit displays the corrected shooting position and shooting direction on the area map.

5. acquiring an omnidirectional image capturing the entire space while moving a first image capturing unit in the space where the structure is arranged; extracting a reference plane image from each of a plurality of celestial sphere frames constituting the omnidirectional video image based on a designated orientation defined in a first coordinate system of the first image capture unit; extracting feature points of the structure defined in the first coordinate system from each of the celestial sphere frames; defining the same feature points in a plurality of the celestial sphere frames as a 3D point cloud in a second coordinate system, and defining, in the second coordinate system, imaging position and attitude information of the first imaging unit that captured each of the celestial sphere frames; converting the specified orientation into an imaging direction defined in the second coordinate system based on the attitude information; and registering the reference planar image in a database by linking it to the corresponding shooting position and shooting direction.

6. On the computer, acquiring an omnidirectional image capturing the entire space while moving the first image capturing unit in the space where the structure is arranged; extracting a reference plane image from each of a plurality of celestial sphere frames constituting the omnidirectional video image based on a designated orientation defined in a first coordinate system of the first image capture unit; extracting feature points of the structure defined in the first coordinate system from each of the celestial frames; defining the same feature points in a plurality of the celestial sphere frames as a 3D point cloud in a second coordinate system, and defining, in the second coordinate system, imaging position and attitude information of the first imaging unit that captured each of the celestial sphere frames; converting the specified orientation into an imaging direction defined in the second coordinate system based on the attitude information; and registering the reference planar image in a database in association with the corresponding imaging position and imaging direction.

Citation Information

Patent Citations

  • Information processing device, information processing method and program

    JP2022105442A