Image analysis device, image analysis method, and image analysis program
The video analysis device aligns two-dimensional feature quantities from three-dimensional and video data to generate associated three-dimensional data, facilitating real-time monitoring of construction progress and water levels without continuous measurement or equipment maintenance.
Patent Information
- Application Number
- JP2024153677
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2044-03-25
AI Technical Summary
Existing methods for real-time three-dimensional measurement in surveillance environments require significant effort for data acquisition and alignment, making it difficult to accurately monitor construction progress without conversion distortion.
A video analysis device that extracts and aligns two-dimensional feature quantities from three-dimensional data and video data using perspective projection, allowing for the generation of associated three-dimensional data without continuous measurement.
Enables real-time monitoring of construction progress and water levels without conversion distortion, reducing the need for continuous measurement and equipment maintenance.
Smart Images

Figure 0007698779000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video analysis device, a video analysis method, and a video analysis program.
Background Art
[0002] In recent years, with the spread of surveillance cameras, it has become possible to monitor the images of the environment and objects. In addition, there is a technique for measuring a quantity such as the volume or weight of an object located in a predetermined space using three-dimensional data (see, for example, Patent Document 1).
[0003] Patent Document 1 describes an invention of a material quantity calculation device. The material quantity calculation device described in Patent Document 1 includes an acquisition unit that acquires a first three-dimensional model indicating a predetermined space and a second three-dimensional model indicating the predetermined space and different from the first three-dimensional model, an alignment unit that aligns the first three-dimensional model and the second three-dimensional model based on the attribute information each of the first three-dimensional model and the second three-dimensional model has, and a calculation unit (difference calculation unit) that calculates the amount of difference between the first three-dimensional model and the second three-dimensional model and outputs the attribute information the difference has and difference information indicating the amount of the difference.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] When it is desired to perform measurement in real time, for the acquisition of three-dimensional information, compared with a surveillance camera, it takes more effort to acquire, and it is necessary to acquire attribute information for alignment for comparison, which requires a lot of effort. An object of the present invention is to enable the progress of construction work and the like at a construction site to be easily grasped without conversion distortion.
Means for Solving the Problem
[0006] To solve the above-described problems, the video analysis apparatus of the present invention includes a three-dimensional data feature quantity conversion unit that extracts feature quantities by assuming two dimensions by perspective projection at a plurality of viewpoints with respect to three-dimensional data obtained by measuring an object with a three-dimensional measuring device, a three-dimensional data feature quantity holding unit that holds the feature quantities extracted by the three-dimensional data feature quantity conversion unit, a video feature quantity conversion unit that extracts feature quantities of video data obtained by photographing the object with a camera, a video feature quantity holding unit that holds the feature quantities extracted by the video feature quantity conversion unit, and the feature quantities of the three-dimensional data stored in the three-dimensional data feature quantity holding unit and the feature quantities of the video data stored in the video feature quantity holding unit Only the feature amount part of the three-dimensional data A spatial alignment unit that compares and determines two-dimensional projection data having an angular field of view closest to the video data, and a three-dimensional data generation unit that associates the video data with the three-dimensional data based on the two-dimensional projection data determined by the spatial alignment unit and generates associated three-dimensional data.
[0007] The video analysis method of the present invention includes steps in which a three-dimensional data feature quantity conversion unit extracts feature quantities by assuming two dimensions by perspective projection at a plurality of viewpoints with respect to three-dimensional data obtained by measuring an object with a three-dimensional measuring device, the three-dimensional data feature quantity holding unit stores the feature quantities extracted by the three-dimensional data feature quantity conversion unit, a video feature quantity conversion unit extracts feature quantities of video data obtained by photographing the object with a camera, the video feature quantity holding unit stores the feature quantities extracted by the video feature quantity conversion unit, and a spatial alignment unit compares the feature quantities of the three-dimensional data stored in the three-dimensional data feature quantity holding unit and the feature quantities of the video data stored in the video feature quantity holding unit Only the feature amount part of the three-dimensional data to determine two-dimensional projection data having an angular field of view closest to the video data, and a three-dimensional data generation unit associates the video data with the three-dimensional data based on the two-dimensional projection data determined by the spatial alignment unit and generates associated three-dimensional data.
[0008] The video analysis program of the present invention causes a computer to perform procedures for extracting feature amounts by assuming two dimensions by perspective projection from a plurality of viewpoints on three-dimensional data obtained by measuring an object with a three-dimensional measuring device, a procedure for storing the feature amounts extracted from the three-dimensional data in a three-dimensional data feature amount holding unit, a procedure for extracting the feature amounts of video data obtained by photographing the object with a camera, a procedure for storing the feature amounts extracted from the video data in a video feature amount holding unit, and comparing the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount holding unit with the feature amounts of the video data stored in the video feature amount holding unit, and Only the feature amount part of the three-dimensional data determining two-dimensional projection data having the closest viewing angle to the video data, and based on the determined two-dimensional projection data, associating the video data with the three-dimensional data and generating associated three-dimensional data. Other means will be described in the mode for carrying out the invention.
Advantages of the Invention
[0009] According to the present invention, it becomes possible to easily grasp the progress of construction work etc. at a construction site etc. without distortion of conversion.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments for carrying out the present invention will be described in detail with reference to each figure. 《First Embodiment》 The video analysis device of the first embodiment measures the water level in real time only by analyzing the video of a surveillance camera. As a result, it is not necessary to install meters for measuring the water level, maintenance of water level measuring instruments and equipment is not required, and secular errors in water level measurement can be eliminated by auto-calibration. Furthermore, the water level can be measured in real time, and it is superior in terms of cost and aesthetics.
[0012] FIG. 1 is a schematic configuration diagram showing a video analysis device 1 according to the first embodiment. The video analysis device 1 includes a video feature quantity conversion unit 12, a video feature quantity holding unit 13, a three-dimensional data feature quantity conversion unit 15, a three-dimensional data feature quantity holding unit 16, a spatial alignment unit 17, a measurement unit 18, and a display control unit 100.
[0013] The three-dimensional data 14 is obtained by measuring a certain object in advance with a three-dimensional measurement device such as LIDAR (Laser Imaging Detection and Ranging). The video data 11 is video data of the same object as the object photographed by LIDAR or the like, and is photographed by a two-dimensional camera. The video data 11 is, for example, a video stream in which a plurality of video frames are arranged in time series, but may also be a still image.
[0014] The three-dimensional data feature quantity conversion unit 15 calculates three-dimensional feature quantities obtained by converting feature quantities assuming two-dimensional projection data through perspective projection from a plurality of viewpoints of the input three-dimensional data 14. That is, the three-dimensional data feature quantity conversion unit 15 extracts feature quantities assuming two dimensions through perspective projection from a plurality of viewpoints with respect to the three-dimensional data obtained by measuring an object with a three-dimensional measuring device. The three-dimensional data feature quantity conversion unit 15 extracts, as the feature quantity, boundary information of a surface detected from the normal lines included in the three-dimensional data of the object. Here, the object is, for example, a bridge pier, and boundary information of a surface detected from the normal lines included in the three-dimensional data of the bridge pier is extracted as a feature quantity and utilized for water level measurement. Note that the boundary information of the surface detected from the normal lines included in the three-dimensional data of the structure is edge information of the shape of the structure.
[0015] The three-dimensional data feature quantities calculated by the three-dimensional data feature quantity conversion unit 15 are held in the three-dimensional data feature quantity holding unit 16. The three-dimensional data 14 may be measured only once at first, and it is not necessary to continuously measure the object. Also, if three-dimensional information already exists, that may be used.
[0016] The video feature quantity conversion unit 12 calculates video feature quantities obtained by converting the feature quantities of the input video data 11. That is, the video feature quantity conversion unit 12 extracts the feature quantities of the video data obtained by photographing an object with a camera. The video feature quantity conversion unit 12 extracts, as the feature quantity, edge information of the video data 11 obtained by photographing the object. Here, the object is, for example, a bridge pier. In water level measurement, the video feature quantity conversion unit 12 extracts, as the feature quantity, edge information in the video obtained by photographing the bridge pier. Also, the video feature quantity conversion unit 12 may use the three-dimensional data 14 when extracting the feature quantity from the video data 11 obtained by photographing the object. Thereby, the video feature quantity conversion unit 12 can suitably extract the feature quantity of the bridge pier. The video feature quantities calculated by the video feature quantity conversion unit 12 are held in the video feature quantity holding unit 13. Further, the video feature quantity conversion unit 12 may extract feature quantities using a plurality of video frames with different time series. Thereby, the video feature quantity conversion unit 12 can extract feature quantities regardless of the noise of individual frames. Further, the video feature quantity conversion unit 12 may extract feature quantities based on the position information of the three-dimensional measuring device and the position information of the camera. Thereby, the video feature quantity conversion unit 12 can suitably extract feature quantities while considering how the surface that can be suitably measured by the three-dimensional measuring device is photographed by the camera.
[0017] The spatial alignment unit 17 compares the data of the video feature quantity holding unit 13 and the three-dimensional data feature quantity holding unit 16, and determines the two-dimensional projection data closest to the video data 11. That is, the spatial alignment unit 17 compares the feature quantity of the three-dimensional data stored in the three-dimensional data feature quantity holding unit 16 with the feature quantity of the video data stored in the video feature quantity holding unit, and determines two-dimensional projection data with an angle of view closest to the video data. In this comparison, since the data of the video feature quantity holding unit 13 may detect not only the edges of the structure but also the edges generated on the video due to shadows or the like, it is preferable to perform the comparison only on the feature quantity part of the three-dimensional data feature quantity holding unit 16. The spatial alignment unit 17 may perform comparison at a predetermined height or more based on the three-dimensional data 14. The predetermined height is, for example, the water surface. Also, for this comparison, it is good to first compare two-dimensional projection data with a rough viewpoint and angle of view, narrow down the general viewpoint and angle of view, and then hierarchically calculate the closest angle of view by changing the fine viewpoint and angle of view.
[0018] The measurement unit 18 measures the size of the object in the video by comparing the feature quantity in the two-dimensional projection data determined by the spatial alignment unit 17 with the feature quantity of the video data stored in the video feature quantity holding unit. The display control unit 100 superimposes and displays the video data and three-dimensional data after alignment, and the warning and danger water levels on the three-dimensional data. Thereby, it is possible to show the user that the current water level should be noted or is dangerous.
[0019] As a specific measurement method, the measurement unit 18 converts the draft line information in the video into the height information of the three-dimensional data by comparing the feature amount of the draft line in the video data with the feature amount of the three-dimensional data stored in the three-dimensional data feature amount holding unit 16.
[0020] In addition, the measurement unit 18 measures the water level by comparing the feature amounts of video data at different times based on the feature amount of the video data at a certain time, and converting it based on the difference. For example, when the water level of a river increases due to a rise in water level, an occlusion area where structures such as bridges do not appear in the video due to water will occur and will not be detected as a feature amount of the video. When this is expressed as a histogram of the feature amounts for each height of the water level, there will be a large difference in the histogram at the increased part of the water level between the reference feature amount and the feature amount when the water level has increased. The measurement unit 18 measures the increase in the water level from the part where this difference occurs, and converts it to the water level by taking into account the height relative to the reference value. At this time, when determining the water level, the measurement unit 18 may compare with the reference of the histogram from the higher water level side, and select a location where the value has changed more than a certain threshold, or a location where the tendency is significantly different from the surrounding height. Thereby, based on the video data captured by the surveillance camera, the water level information of the river can be calculated in real time. In addition, the measurement unit 18 performs measurement in short time units to measure the water level, but it is not necessary to perform the spatial alignment process every time.
[0021] FIG. 2 is an example of video data of a bridge pier. The bridge pier is installed in the river, and the shape of the structure and the water surface appear on the surface of the bridge pier. By calculating the height information of the three-dimensional data of this shape and the water surface, and comparing it with the feature amount of the reference video data, the water level information of the river can be obtained. In order to obtain the shape information of this structure, it is preferable to detect the boundary line (edge) information between different surfaces according to the situation of the normal line, which is a vector indicating the surface of the surface, from the three-dimensional data, and use this as a feature amount.
[0022] FIG. 3 is a flowchart of the measurement visualization process. First, the three-dimensional data feature quantity conversion unit 15 acquires high-precision three-dimensional data 14 (step S10). Then, the three-dimensional data feature quantity conversion unit 15 extracts boundary information of the surface detected from the normal line included in the three-dimensional data from this three-dimensional data 14 (step S11). Next, the spatial alignment unit 17 associates the feature quantity of the image with the boundary information of the surface included in the three-dimensional data (step S12). The measurement unit 18 measures the difference in occlusion where the feature points appearing in the two-dimensional projection data are hidden in the video data among the associated feature quantities of the image, creates a height-specific match quantity histogram (step S13), and identifies the water level by comparing it with the histogram at the reference time (step S14), and then ends the process of FIG. 3.
[0023] FIG. 4 is a data flow diagram of the video analysis device 1 of the first embodiment. At the time of data collection, the LIDAR 31 performs three-dimensional measurement of the object to generate the point cloud 321. Then, the video analysis device 1 generates the mesh 322 and delivers it to the Raycast unit 33. The Raycast unit 33 corresponds to, for example, the Raycast of the game engine Unity (registered trademark), and executes a process of emitting a transparent light ray from a certain specific object and acquiring the coordinates of another object that the light ray hits.
[0024] At the time of camera installation, based on the camera parameter 331 indicating the position and direction of the camera that captures the video, the Raycast unit 33 determines the two-dimensional projection data closest to the video data when the object is captured by the camera from the object indicated by the mesh 322. This two-dimensional projection data is stored in the virtual legal image feature database 34. The virtual legal image feature database 34 further stores the LIDAR / camera conversion matrix 341 and the Raycast resolution data 342.
[0025] During operation, the monitoring camera 35 captures an object to generate an RGB image 351. The video analysis device 1 generates feature data 352 from the RGB image 351 and delivers it to the feature matching unit 36 and the feature comparison unit 37. The feature matching unit 36 compares the feature points of the two-dimensional projection data stored in the virtual law image feature database 34 with the feature data 352. Then, the feature matching unit 36 outputs the most matching normal image feature 361. The feature comparison unit 37 compares the feature data 352, the most matching normal image feature 361, and the feature amount at the reference time to generate a height-specific match amount histogram 371. The water level estimation unit 38 estimates the current water level by determining this height-specific match amount histogram 371.
[0026] FIG. 5 is a configuration diagram of the specific hardware of the video analysis device 1 according to the first embodiment. The video analysis device 1 is, for example, a computer, and a feature processing plugin 43 and a water level measurement plugin 44 are incorporated therein. The video analysis device 1 further stores a shooting information database 481, three-dimensional data 482, an image feature database 483, a feature matching processing result 484, and an analysis result 485. Connected to the video analysis device 1 are a monitoring camera 41, a video management device 42, a three-dimensional measurement device 45 such as a LIDAR, and an initial data generation device 46.
[0027] The monitoring camera 41 captures video and delivers it to the video management device 42. The video management device 42 stores the video captured by the monitoring camera 41 and outputs the image data to the feature processing plugin 43 and the water level measurement plugin 44.
[0028] The three-dimensional measurement device 45 is, for example, a LIDAR. It performs three-dimensional measurement of the object, writes the shooting location ID into the shooting information database 481, and writes the mesh data and the shooting location ID into the three-dimensional data 482. The initial data generation device 46 is configured to include a virtual normal image generation unit 461 and a feature amount extraction unit 462. The virtual normal image generation unit 461 generates a normal image from the shooting location ID and the mesh data. The feature amount extraction unit 462 generates a feature image and a depth image from the normal image. The initial data generation device 46 stores the normal image, the feature image, and the depth image in the image feature database 483.
[0029] The feature processing plugin 43 includes a feature extraction unit 431 and a feature matching unit 432. The feature extraction unit 431 extracts features from the image data and outputs them to the feature matching unit 432. The feature matching unit 432 receives the features of the image data and the shooting location ID and the camera parameters from the shooting information database 481.
[0030] The feature matching unit 432 passes the shooting location ID to the image feature database 483 and reads the feature image corresponding to the shooting location ID. The feature matching unit 432 matches the features of the image data with the feature image and selects a normal image ID. The feature matching unit 432 stores the normal image ID and the shooting location ID in the feature matching processing result 484.
[0031] The water level measurement plugin 44 includes a feature extraction unit 441 and a water level measurement unit 442. The feature extraction unit 441 receives the image data from the video management device 42. The feature extraction unit 441 extracts feature points from the normal image ID and the image data and outputs them to the water level measurement unit 442. The water level measurement unit 442 measures the water level from the feature points, the normal image ID, the feature image, and the depth image. The water level measurement unit 442 stores the shooting location ID, the shooting time, the image data, and the water level in the analysis result database 485. The analysis result display device 47 acquires and displays the water level and the like from the analysis result database 485. Thus, it is possible to measure the water level from the video of the monitoring camera 41. In this embodiment, an example is given in which three-dimensional data is acquired in advance by LiDAR, but a method of utilizing the already acquired three-dimensional information may also be used.
[0032] "Second Embodiment" The video analysis device of the second embodiment converts a construction site or the like into an aerial view video with a free viewpoint without distortion, enabling easy understanding of the progress of the construction work and the like.
[0033] FIG. 6 is a schematic configuration diagram showing a video analysis device 1A according to the second embodiment. The video analysis device 1A includes a video feature amount conversion unit 12, a video feature amount holding unit 13, a three-dimensional data feature amount conversion unit 15, a three-dimensional data feature amount holding unit 16, a spatial alignment unit 17, and a three-dimensional data generation unit 19.
[0034] The three-dimensional data 14 is, for example, the measurement of an object such as a construction site with a three-dimensional measurement device such as LiDAR (Laser Imaging Detection and Ranging) each time. The video data 11 is a video of the same object as the object such as a construction site photographed by LiDAR or the like, and is photographed by a two-dimensional camera.
[0035] The three-dimensional data feature amount conversion unit 15 calculates a three-dimensional data feature amount obtained by converting the feature amount assuming two-dimensional projection data by perspective projection at a plurality of viewpoints of the input three-dimensional data 14. The three-dimensional data feature amount calculated by the three-dimensional data feature amount conversion unit 15 is held in the three-dimensional data feature amount holding unit 16. The three-dimensional data 14 is acquired each time the construction progresses.
[0036] The video feature amount conversion unit 12 calculates a video feature amount obtained by converting the feature amount of the input video data 11. The video feature amount calculated by the video feature amount conversion unit 12 is held in the video feature amount holding unit 13.
[0037] The spatial alignment unit 17 compares the data in the video feature amount holding unit 13 and the three-dimensional data feature amount holding unit 16, and determines the two-dimensional projection data closest to the video data 11. Based on the two-dimensional projection data determined by the spatial alignment unit 17, the three-dimensional data generation unit 19 generates texture-mapped three-dimensional data by mapping the texture of the video data onto the three-dimensional data. That is, the three-dimensional data generation unit 19 associates the video data with the mesh information generated from the three-dimensional data corresponding to the two-dimensional projection data determined by the spatial alignment unit 17. By rendering this three-dimensional data using, for example, the game engine Unity (registered trademark), the object can be viewed from a free viewpoint.
[0038] FIG. 7 is a data flow diagram of the video analysis device 2 of the modified example. The video analysis device 2 includes a texture extraction unit 22, a mesh data conversion unit 25, and a viewpoint conversion rendering unit 26, and a plurality of cameras 21a, 21b, a long-range LiDAR 23, and a portable LiDAR 24 are connected thereto.
[0039] The plurality of cameras 21a, 21b are installed at different viewpoints and capture videos of the object from the plurality of viewpoints. The long-range LiDAR 23 measures the three-dimensional data of this object from a point where the distance to the object is long. The portable LiDAR 24 measures the three-dimensional data of this object from a point close to the object.
[0040] The mesh data conversion unit 25 integrates the three-dimensional data of the object and converts it into mesh data. The mesh data generated by the mesh data conversion unit 25 is output to the texture extraction unit 22.
[0041] The texture extraction unit 22 receives video input from the plurality of cameras 21a, 21b. Then, the texture extraction unit 22 converts the three-dimensional mesh data into two-dimensional projection data as seen from each viewpoint of the cameras 21a, 21b, associates each two-dimensional projection data with each video, and extracts the texture from the video for pasting on the mesh data. Thereby, the texture extraction unit 22 creates mesh data with the texture pasted thereon.
[0042] The viewpoint conversion rendering unit 26 is, for example, the game engine Unity (registered trademark), and renders mesh data with textures attached from an arbitrary viewpoint. Thereby, it is possible to observe an object, for example, a construction site, from a desired viewpoint. In addition, in this embodiment, three-dimensional information is sequentially captured by LiDAR, but a method of utilizing already acquired three-dimensional information may also be used. The already acquired three-dimensional information is, for example, BIM (Building Information Modeling), CIM (Construction Information Modeling), etc.
[0043] 《Third Embodiment》 The video analysis device of the third embodiment measures the length and volume of temporary materials on-site simply by taking a picture with a mobile terminal such as a smartphone based on the three-dimensional data acquired by LiDAR. The video analysis device grasps the situation of the material bundle from the video taken by the mobile terminal and performs matching processing with the length analysis information of the cross-section. Thereby, for example, a material with a length of 20 m can be measured from a place 20 m away, and the length of the material can be classified in real time in 10-cm increments.
[0044] The video analysis device of the third embodiment realizes the measurement of the length and volume of the load on the truck at once. Thereby, the work that was conventionally carried out manually one by one by climbing onto the loading platform can be carried out safely and efficiently.
[0045] FIG. 8 is a schematic configuration diagram showing the video analysis device 1B according to the third embodiment. The video analysis device 1 includes a video feature quantity conversion unit 12, a video feature quantity holding unit 13, a three-dimensional data feature quantity conversion unit 15, a three-dimensional data feature quantity holding unit 16, a video recognition unit 101, and a measurement unit 102.
[0046] The three-dimensional data 14 is the measurement of a certain object by a three-dimensional measuring device such as LIDAR (Laser Imaging Detection and Ranging) in advance. The video data 11 is a video of the same object as the object photographed by LIDAR or the like, and is photographed by a two-dimensional camera.
[0047] The three-dimensional data feature quantity conversion unit 15 calculates a three-dimensional data feature quantity obtained by converting the feature quantity assuming two-dimensional projection data by perspective projection at a plurality of viewpoints of the input three-dimensional data 14. The three-dimensional data feature quantity calculated by the three-dimensional data feature quantity conversion unit 15 is held in the three-dimensional data feature quantity holding unit 16. The three-dimensional data 14 may be measured only once at first, and it is not necessary to continuously measure the object.
[0048] The video feature quantity conversion unit 12 calculates a video feature quantity obtained by converting the feature quantity of the input video data 11. The video feature quantity calculated by the video feature quantity conversion unit 12 is held in the video feature quantity holding unit 13.
[0049] The video recognition unit 101 performs image recognition of the object from the video data 11, recognizes whether it is a bundle of target materials, and when the objects overlap, recognizes the number or quantity of the individuals. Then, the video recognition unit 101 compares the feature quantity data in the video feature quantity holding unit 13 with the feature quantity data in the three-dimensional data feature quantity holding unit 16, determines the two-dimensional projection data closest to the video data 11, and recognizes the three-dimensional data corresponding to the recognized object.
[0050] The measurement unit 102 measures the size of the object according to the range of the three-dimensional data of the object recognized by the video recognition unit 101.
[0051] Specifically, this size measurement means measuring the length, width, and height based on the three-dimensional data of the recognized object (for example, a bundle), calculating the volume from that, detecting the number and shape of the materials included in the bundle from the video, and performing individual length measurements respectively.
[0052] Figure 9 is a flowchart of the measurement visualization process. First, the user measures three-dimensional data with LiDAR from the side of the material (step S20). Next, the user takes a picture of the cross-section and outline of the material with a camera (step S21). The measured three-dimensional data and the captured video data are input into the video analysis device 1B. The video recognition unit 101 of the video analysis device 1B recognizes the material bundle from the video data 11 (step S22).
[0053] The video recognition unit 101 of the video analysis device 1B performs a matching process on the video data 11 and the three-dimensional data 14 to grasp the situation of the material bundle (step S23). Thereby, the part of the material bundle in the three-dimensional data 14 can be recognized. Furthermore, the video recognition unit 101 of the video analysis device 1B analyzes the number of target materials from the cross-sectional shape. Then, when the measurement unit 102 of the video analysis device 1B analyzes the length in the longitudinal direction of the material (step S25), the process of FIG. 9 ends.
[0054] Hereinafter, the configuration and effects of the present invention will be described.
[0055] [1] A three-dimensional data feature amount conversion unit (15) that extracts feature amounts by assuming two dimensions through perspective projection at a plurality of viewpoints with respect to three-dimensional data (14) measured by a three-dimensional measuring device for an object, A three-dimensional data feature amount holding unit (16) that stores the feature amounts extracted by the three-dimensional data feature amount conversion unit, A video feature amount conversion unit (12) that extracts the feature amounts of video data (11) obtained by photographing the object with a camera, A video feature amount holding unit (13) that stores the feature amounts extracted by the video feature amount conversion unit, A spatial alignment unit (17) that compares the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount holding unit (16) with the feature amounts of the video data stored in the video feature amount holding unit (13) and determines two-dimensional projection data with the viewing angle closest to the video data (11), Based on the two-dimensional projection data determined by the spatial alignment unit (17), a three-dimensional data generation unit (19) that associates the video data with the three-dimensional data and generates the associated three-dimensional data. A video analysis device characterized by having the above.
[0056] As a result, it becomes possible to easily grasp the progress of construction etc. at a construction site etc. without distortion of conversion.
[0057] [2] A viewpoint conversion rendering unit (26) that renders the associated three-dimensional data generated by the three-dimensional data generation unit (19) from an arbitrary viewpoint is provided. The video analysis device according to [1], characterized by the above.
[0058] As a result, it becomes possible to overlook the object from a free viewpoint.
[0059] [3] The three-dimensional data feature quantity conversion unit (15) extracts, as the feature quantity, boundary information of a surface detected from the normal line included in the three-dimensional data obtained by measuring the object. The video analysis device according to [1], characterized by the above.
[0060] As a result, boundary information of a surface detected from the normal line can be obtained from the three-dimensional information, and the water level can be measured from the relationship between the boundary surface information and the waterline or the edge information of the video.
[0061] [4] The video feature quantity conversion unit (12) extracts, as the feature quantity, edge information of the video data (11) obtained by photographing the object. The video analysis device according to [1], characterized by the above.
[0062] As a result, it becomes possible to easily associate the video data with the three-dimensional data.
[0063] [5] When extracting the feature amount from the video data (11) obtained by photographing the object, the video feature amount conversion unit (12) uses the three-dimensional data. The video analysis apparatus according to [1], characterized in that.
[0064] As a result, the video data and the three-dimensional data can be accurately associated with each other.
[0065] [6] The video feature amount conversion unit (12) extracts the feature amount using a plurality of video frames with different time series. The video analysis apparatus according to [1], characterized in that.
[0066] As a result, noise in the video data can be removed, and the height, length, or size of an object can be preferably measured.
[0067] [7] The spatial alignment unit (17) performs comparison at a predetermined height or more based on the three-dimensional data (14). The video analysis apparatus according to [1], characterized in that.
[0068] As a result, it is possible to prevent erroneously measuring an impossible water level.
[0069] [8] The video feature amount conversion unit (12) extracts a feature amount based on the position information of the three-dimensional measuring device and the position information of the camera. The video analysis apparatus according to [1], characterized in that.
[0070] As a result, it is possible to reduce the cost of calculating two-dimensional projection data when viewed from the camera from the three-dimensional data measured by the three-dimensional measuring device.
[0071] [9] The video feature amount conversion unit (12) extracts a feature amount from the frame information of a plurality of frames of the video data (11). The video analysis apparatus according to [1], characterized in that.
[0072] Thereby, noise in the video data can be removed, and the height, length, or size of an object can be suitably measured.
[0073]
[10] A display control unit (100) that superimposes and displays the video data (11) and three-dimensional data after alignment, and the water levels of caution and danger on the three-dimensional data, The video analysis apparatus according to [1], further comprising:
[0074] Thereby, a user can be notified of a water level warning together with the basis data.
[0075]
[11] For the three-dimensional data (14) obtained by measuring an object with a three-dimensional measuring device, a step in which a three-dimensional data feature amount conversion unit (15) extracts a feature amount assuming two dimensions by perspective projection at a plurality of viewpoints; A step of storing the feature amount extracted by the three-dimensional data feature amount conversion unit (15) in a three-dimensional data feature amount holding unit (16); A step in which a video feature amount conversion unit (12) extracts a feature amount of video data obtained by photographing the object with a camera; A step of storing the feature amount extracted by the video feature amount conversion unit (12) in a video feature amount holding unit (13); A step in which a spatial alignment unit (17) compares the feature amount of the three-dimensional data stored in the three-dimensional data feature amount holding unit (16) with the feature amount of the video data stored in the video feature amount holding unit (13), and determines two-dimensional projection data having an angle of view closest to the video data; A step in which a three-dimensional data generation unit (19) associates the video data with the three-dimensional data based on the two-dimensional projection data determined by the spatial alignment unit (17), and generates associated three-dimensional data; A video analysis method, characterized by comprising:
[0076] This enables the progress of construction work, etc. at a construction site, etc. to be easily grasped without distortion due to conversion.
[0077]
[12] On a computer, For three-dimensional data (14) obtained by measuring an object with a three-dimensional measuring device, a procedure for extracting feature amounts by assuming two dimensions through perspective projection from a plurality of viewpoints, A procedure for storing the feature amounts extracted from the three-dimensional data in a three-dimensional data feature amount holding unit (16), A procedure for extracting the feature amounts of video data (11) obtained by photographing the object with a camera, A procedure for storing the feature amounts extracted from the video data in a video feature amount holding unit (13), A procedure for comparing the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount holding unit (16) with the feature amounts of the video data stored in the video feature amount holding unit (13) and determining two-dimensional projection data having the viewing angle closest to the video data, A procedure for associating the video data with the three-dimensional data based on the determined two-dimensional projection data and generating associated three-dimensional data, A video analysis program for executing the above.
[0078] This enables the progress of construction work, etc. at a construction site, etc. to be easily grasped without distortion due to conversion.
[0079] (Modification example) The present invention is not limited to the above-described embodiments and includes various modification examples. For example, the above-described embodiments have been described in detail for easy understanding of the present invention and are not necessarily limited to those having all the configurations described. It is possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Further, it is possible to add, delete, or replace a part of the configuration of each embodiment with other configurations.
[0080] Each of the above-described configurations, functions, processing units, processing means, etc. may be realized in part or in whole by hardware such as an integrated circuit. Each of the above-described configurations, functions, etc. may also be realized by software by a processor interpreting and executing a program for realizing each function. Information such as a program, table, file, etc. for realizing each function can be placed in a recording device such as a memory, hard disk, SSD (Solid State Drive), or a recording medium such as a flash memory card, DVD (Digital Versatile Disk).
[0081] In each embodiment, the control lines and information lines show those considered necessary for explanation, and not necessarily all the control lines and information lines on the product. In reality, it may be considered that almost all the components are interconnected.
Explanation of Reference Numerals
[0082] 1, 1A, 1B Video Analysis Device 11 Video Data 12 Video Feature Quantity Conversion Unit 13 Video Feature Quantity Holding Unit 14 Three-Dimensional Data 15 Three-Dimensional Data Feature Quantity Conversion Unit 16 Three-Dimensional Data Feature Quantity Holding Unit 17 Spatial Alignment Unit 18 Measurement Unit 2 Video Analysis Device 22 Texture Extraction Unit 25 Mesh Data Conversion Unit 26 Viewpoint Conversion Rendering Unit 21a Camera 21b Camera 23 Long-Range LiDAR 24 Portable LiDAR 31 LIDAR 321 Point Cloud 322 Mesh 33 Raycast Unit 331 Camera Parameters 34 Hypothetical Legal Image Feature Database 341 LIDAR / Camera Transformation Matrix 342 Raycast Resolution Data 35 Surveillance Camera 351 RGB Image 352 Feature Data 36 Feature Matching Unit 37 Feature Comparison Unit 361 Normal Image Feature 371 Matching Quantity Histogram by Height 38 Water Level Estimation Unit 43 Feature Processing Plug-in 44 Water Level Measurement Plug-in 481 Shooting Information Database 482 3D Data 483 Image Feature Database 41 Surveillance Camera 42 Video Management Device 45 3D Measurement Device 46 Initial Data Generation Device 461 Hypothetical Normal Image Generation Unit 462 Feature Quantity Extraction Unit 431 Feature Extraction Unit 432 Feature Matching Unit 441 Feature Extraction Unit 442 Water Level Measurement Unit 485 Analysis Result Database 47 Analysis Result Display Device 19 3D Data Generation Unit
Claims
1. a three-dimensional data feature quantity conversion unit that extracts feature quantities by assuming two dimensions through perspective projection from a plurality of viewpoints from three-dimensional data obtained by measuring an object using a three-dimensional measuring device; a three-dimensional data feature quantity storage unit for storing the feature quantity extracted by the three-dimensional data feature quantity conversion unit; an image feature conversion unit that extracts features of image data obtained by photographing the object with a camera; a video feature storage unit for storing the feature extracted by the video feature conversion unit; a spatial registration unit that compares the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit with the feature amounts of the video data stored in the video feature amount storage unit, using only the feature amounts of the three-dimensional data, to determine two-dimensional projection data having an angle of view closest to that of the video data; a three-dimensional data generating unit that links the video data to the three-dimensional data based on the two-dimensional projection data determined by the spatial positioning unit, and generates linked three-dimensional data; A video analysis device comprising:
2. a viewpoint conversion rendering unit that renders the linked three-dimensional data generated by the three-dimensional data generation unit from an arbitrary viewpoint; 2. The video analysis device according to claim 1.
3. the three-dimensional data feature quantity conversion unit extracts, as the feature quantity, boundary information of a surface detected from a normal line included in the three-dimensional data obtained by measuring the object; 2. The video analysis device according to claim 1.
4. the image feature conversion unit extracts edge information of the image data capturing the object as the feature.
2. The video analysis device according to claim 1.
5. the image feature conversion unit uses the three-dimensional data when extracting features from the image data of the object; 2. The video analysis device according to claim 1.
6. The video feature conversion unit extracts the feature using a plurality of video frames having different time series.
2. The video analysis device according to claim 1.
7. The image feature conversion unit extracts features based on position information of the three-dimensional measuring device and position information of the camera.
2. The video analysis device according to claim 1.
8. The video feature conversion unit extracts features from a plurality of frames of the video data.
2. The video analysis device according to claim 1.
9. A step in which a three-dimensional data feature quantity conversion unit extracts feature quantities from three-dimensional data obtained by measuring an object using a three-dimensional measuring device, assuming two dimensions by perspective projection from a plurality of viewpoints; storing the feature quantity extracted by the three-dimensional data feature quantity conversion unit in a three-dimensional data feature quantity storage unit; A video feature conversion unit extracts features of video data captured by a camera of the object; storing the feature extracted by the video feature conversion unit in a video feature storage unit; a spatial positioning unit comparing the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit with the feature amounts of the video data stored in the video feature amount storage unit using only the feature amounts of the three-dimensional data to determine two-dimensional projection data having an angle of view closest to that of the video data; a three-dimensional data generating unit linking the video data to the three-dimensional data based on the two-dimensional projection data determined by the spatial registration unit, and generating linked three-dimensional data; A video analysis method comprising:
10. On the computer, A procedure for extracting feature quantities from three-dimensional data obtained by measuring an object using a three-dimensional measuring device, assuming two-dimensionality by perspective projection from multiple viewpoints; storing the feature amount extracted from the three-dimensional data in a three-dimensional data feature amount storage unit; A step of extracting features of video data of the object captured by a camera; storing the feature amount extracted from the video data in a video feature amount storage unit; a step of comparing the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit with the feature amounts of the video data stored in the video feature amount storage unit, using only the feature amounts of the three-dimensional data, to determine two-dimensional projection data having an angle of view closest to that of the video data; linking the video data to the three-dimensional data based on the determined two-dimensional projection data to generate linked three-dimensional data; A video analysis program for executing the above.
Citation Information
Patent Citations
Enhanced virtual environment
JP2006503379A
Three-dimensional model processing device and camera calibration system
JP2016170610A
Object amount calculation device and object amount calculation method
WO2020179438A1
Cited By
Image synthesis device, image synthesis method, and program
JP7775526B1