Video analysis device, video analysis method, and video analysis program
The video analysis device facilitates real-time measurement of object dimensions by aligning three-dimensional and video data features, addressing the effort required for attribute information comparison and reducing measurement errors.
Patent Information
- Application Number
- JP2024048203
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2044-03-25
AI Technical Summary
Obtaining three-dimensional information for real-time measurements and alignment requires significant effort and work, particularly for attribute information comparison using surveillance cameras.
A video analysis device and method that includes a three-dimensional data feature conversion unit to extract features from three-dimensional data, a video feature conversion unit to extract features from video data, and a spatial alignment unit to compare these features, determining two-dimensional projection data with the closest angle of view for accurate measurements.
Enables real-time measurement of object height, length, or size using image data, reducing the need for continuous measurement and minimizing errors in water level estimation.
Smart Images

Figure 2025147785000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video analysis device, a video analysis method, and a video analysis program. [Background technology]
[0002] In recent years, the widespread use of surveillance cameras has made it possible to monitor the environment and objects through video. There is also a technology that uses three-dimensional data to measure quantities such as the volume or weight of an object located in a specified space (see, for example, Patent Document 1).
[0003] Patent Document 1 describes an invention of a quantity calculation device. The quantity calculation device described in Patent Document 1 includes an acquisition unit that acquires a first three-dimensional model that represents a predetermined space and a second three-dimensional model that represents the predetermined space and is different from the first three-dimensional model, an alignment unit that aligns the first three-dimensional model with the second three-dimensional model based on attribute information that each of the first and second three-dimensional models has, and a calculation unit (difference calculation unit) that calculates the amount of difference between the first and second three-dimensional models and outputs attribute information that the difference has and difference information that indicates the amount of the difference. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] International Publication No. 2020 / 179438 Summary of the Invention [Problem to be solved by the invention]
[0005] When performing measurements in real time, obtaining three-dimensional information requires more effort than using a surveillance camera, and attribute information must be obtained for alignment purposes for comparison, which requires a lot of work. Therefore, an object of the present invention is to measure the height, length, or size of an object in real time using an image. [Means for solving the problem]
[0006] In order to solve the above-mentioned problems, the video analysis device of the present invention is characterized by having a three-dimensional data feature conversion unit that extracts features by assuming two dimensions through perspective projection from multiple viewpoints from three-dimensional data measured by a three-dimensional measurement device; a three-dimensional data feature storage unit that stores the features extracted by the three-dimensional data feature conversion unit; a video feature conversion unit that extracts features from video data captured by a camera of the object; a video feature storage unit that stores the features extracted by the video feature conversion unit; and a spatial alignment unit that compares the features of the three-dimensional data stored in the three-dimensional data feature storage unit with the features of the video data stored in the video feature storage unit, and determines two-dimensional projection data with an angle of view closest to the video data.
[0007] The video analysis method of the present invention is characterized by comprising the steps of: a three-dimensional data feature conversion unit extracting features from three-dimensional data measured by a three-dimensional measuring device by assuming two dimensions through perspective projection from multiple viewpoints; storing the features extracted by the three-dimensional data feature conversion unit in a three-dimensional data feature storage unit; a video feature conversion unit extracting features from video data obtained by photographing the object with a camera and storing the features extracted by the video feature conversion unit in the video feature storage unit; and a spatial alignment unit comparing the features of the three-dimensional data stored in the three-dimensional data feature storage unit with the features of the video data stored in the video feature storage unit to determine two-dimensional projection data with an angle of view closest to the video data.
[0008] The video analysis program of the present invention causes a computer to execute the following steps: extracting features from three-dimensional data measured by a three-dimensional measuring device, assuming two dimensions by perspective projection from multiple viewpoints; storing the features extracted from the three-dimensional data in a three-dimensional data feature storage unit; extracting features from video data captured by a camera of the object; storing the features extracted from the video data in a video feature storage unit; and comparing the features of the three-dimensional data stored in the three-dimensional data feature storage unit with the features of the video data stored in the video feature storage unit to determine two-dimensional projection data with an angle of view closest to the video data. Other means will be described in the detailed description of the invention. [Effects of the Invention]
[0009] According to the present invention, it is possible to measure the height, length, or size of an object in real time using an image. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a schematic configuration diagram showing a video analysis device according to a first embodiment. [Figure 2] This is an example of video data of a bridge pier. [Figure 3] 10 is a flowchart of a measurement visualization process. [Figure 4] FIG. 2 is a data flow diagram of the video analysis device of the first embodiment. [Figure 5] 1 is a diagram illustrating a hardware configuration of a video analysis device according to a first embodiment. [Figure 6] FIG. 10 is a schematic configuration diagram showing a video analysis device according to a second embodiment. [Figure 7] FIG. 10 is a data flow diagram of the video analysis device of the second embodiment. [Figure 8] FIG. 10 is a schematic configuration diagram showing a video analysis device according to a third embodiment. [Figure 9] 10 is a flowchart of a measurement visualization process. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. First Embodiment The video analysis device of the first embodiment measures water levels in real time by analyzing only video footage from surveillance cameras. This eliminates the need to install meters to measure water levels, eliminates the need for maintenance of water level measurement instruments and equipment, and makes it possible to cancel out errors in water level measurement over time through auto-calibration. Furthermore, while being able to measure water levels in real time, it is advantageous in terms of cost and aesthetics.
[0012] FIG. 1 is a schematic diagram showing the configuration of a video analysis device 1 according to the first embodiment. The video analysis device 1 is configured to include a video feature conversion unit 12, a video feature storage unit 13, a three-dimensional data feature conversion unit 15, a three-dimensional data feature storage unit 16, a spatial alignment unit 17, a measurement unit 18, and a display control unit 100.
[0013] The three-dimensional data 14 is obtained by measuring a certain object in advance using a three-dimensional measuring device such as LIDAR (Laser Imaging Detection and Ranging). The video data 11 is video data of the same object as the object photographed by the LIDAR or the like, and is photographed by a two-dimensional camera. The video data 11 is, for example, a video stream in which multiple video frames are arranged in time series, but may also be still images.
[0014] The three-dimensional data feature quantity conversion unit 15 converts the input three-dimensional data 14 into three-dimensional feature quantities by assuming two-dimensional projection data through perspective projection from multiple viewpoints. In other words, the three-dimensional data feature quantity conversion unit 15 extracts feature quantities from three-dimensional data measured with a three-dimensional measurement device, assuming two dimensions through perspective projection from multiple viewpoints. The three-dimensional data feature quantity conversion unit 15 extracts, as the feature quantities, surface boundary information detected from normals included in the three-dimensional data measured with the object. Here, the object is, for example, a bridge pier, and the surface boundary information detected from normals included in the three-dimensional data of the pier is extracted as a feature quantity and used for water level measurement. The surface boundary information detected from normals included in the three-dimensional data of the structure is edge information of the structure's shape.
[0015] The three-dimensional data feature values calculated by the three-dimensional data feature value conversion unit 15 are stored in the three-dimensional data feature value storage unit 16. The three-dimensional data 14 may be obtained by measuring the object only once at the beginning, and it is not necessary to measure the object continuously. Furthermore, if three-dimensional information already exists, that information may be used.
[0016] The image feature conversion unit 12 calculates image features by converting the features of the input image data 11. In other words, the image feature conversion unit 12 extracts features from image data of an object captured by a camera. The image feature conversion unit 12 extracts edge information from the image data 11 of the object as features. Here, the object is, for example, a bridge pier. In water level measurement, the image feature conversion unit 12 extracts edge information from the image of the bridge pier as features. Furthermore, the image feature conversion unit 12 preferably uses three-dimensional data 14 when extracting features from the image data 11 of the object. This enables the image feature conversion unit 12 to suitably extract the features of the bridge pier. The image features calculated by the image feature conversion unit 12 are stored in the image feature storage unit 13. Furthermore, it is preferable that the image feature conversion unit 12 extracts features using multiple image frames with different time series. This allows the image feature conversion unit 12 to extract features regardless of noise in individual frames. Furthermore, it is preferable that the image feature conversion unit 12 extracts features based on position information of the 3D measurement device and position information of the camera. This allows the image feature conversion unit 12 to extract features appropriately while taking into consideration how a surface that can be suitably measured by the 3D measurement device is photographed by the camera.
[0017] The spatial positioning unit 17 compares the data in the image feature amount storage unit 13 and the data in the three-dimensional data feature amount storage unit 16, and determines the two-dimensional projection data that is closest to the image data 11. In other words, the spatial positioning unit 17 compares the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit 16 with the feature amounts of the image data stored in the image feature amount storage unit, and determines the two-dimensional projection data that has the angle of view closest to the image data. When making this comparison, it is advisable to make the comparison using only the feature portion of the three-dimensional data feature storage unit 16, since the data in the image feature storage unit 13 may detect not only the edges of structures but also edges that appear on the image due to shadows, etc. The spatial positioning unit 17 may perform comparison at a predetermined height or higher based on the three-dimensional data 14. The predetermined height may be, for example, the water surface. Furthermore, for this comparison, it is advisable to first compare two-dimensional projection data at a rough viewpoint and angle of view, narrow down the general viewpoint and angle of view, and then calculate the hierarchically closest angle of view by making finer changes to the viewpoint and angle of view.
[0018] The measurement unit 18 measures the size of the object in the video by comparing the features in the two-dimensional projection data determined by the spatial positioning unit 17 with the features of the video data stored by the video feature storage unit. The display control unit 100 displays the image data and three-dimensional data after alignment, and the water levels indicating caution and danger superimposed on the three-dimensional data, thereby indicating to the user that the current water level requires caution or is dangerous.
[0019] As a specific measurement method, the measurement unit 18 converts the waterline information in the video into height information of the three-dimensional data by comparing the waterline feature values in the video data with the feature values of the three-dimensional data stored in the three-dimensional data feature value storage unit 16.
[0020] The measurement unit 18 also uses the feature values of the video data at a certain time as a reference, compares the feature values of the video data at a different time, and measures the water level based on the difference. This is because, for example, when the water level of a river rises due to flooding, occlusion areas occur in which structures such as bridges are hidden by the water, preventing them from being detected as feature values of the image. If this is expressed as a histogram of feature values for each water level, a large difference will occur in the histogram between the reference feature value and the feature value when the water level increases. The measurement unit 18 measures the increase in water level from the area where this difference occurs and converts it into a water level by taking the height into account relative to the reference value. In determining the water level, the measurement unit 18 may compare the histogram with the reference value starting from the highest water level and select points where the value has changed more than a certain threshold, or points where the trend differs significantly from the surrounding heights. This allows river water level information to be calculated in real time based on video data captured by surveillance cameras. Furthermore, the measurement unit 18 performs measurements in short time units to measure the water level, but in this case, it is not necessary to perform the spatial positioning process every time.
[0021] Figure 2 is an example of video data of a bridge pier. Bridge piers are installed in rivers, and the shape and water surface of the structure are visible on the surface of the pier. By calculating height information from the 3D data of this shape and water surface and comparing it with the feature quantities of the reference image data, it is possible to obtain river water level information. In order to obtain this shape information of the structure, it is preferable to detect boundary (edge) information between different faces from the 3D data based on the state of the normals, which are vectors that indicate the surface of the faces, and use this as the feature quantity.
[0022] FIG. 3 is a flowchart of the measurement visualization process. First, the three-dimensional data feature quantity conversion unit 15 acquires high-precision three-dimensional data 14 (step S10). Then, the three-dimensional data feature quantity conversion unit 15 extracts, from this three-dimensional data 14, boundary information of surfaces detected from normals included in the three-dimensional data (step S11). Next, the spatial positioning unit 17 associates the image feature amounts with the boundary information of the surfaces included in the three-dimensional data (step S12). The measurement unit 18 measures the difference in occlusion, where feature points that appear in the two-dimensional projection data are hidden in the video data, among the associated image feature amounts, to create a height-specific match amount histogram (step S13), and identifies the water level by comparing it with the histogram for the reference time (step S14), thereby terminating the processing of FIG.
[0023] FIG. 4 is a data flow diagram of the video analysis device 1 of the first embodiment. When collecting data, the LIDAR 31 performs three-dimensional measurement of the object and generates a point cloud 321. Then, the video analysis device 1 generates a mesh 322 and passes it to the raycast unit 33. The raycast unit 33 corresponds to, for example, the raycast of the Unity (registered trademark) game engine, and executes a process of emitting a transparent ray of light from a specific object and acquiring the coordinates of another object that the ray hits.
[0024] When the camera is installed, the Raycast unit 33 determines, from the object represented by the mesh 322, two-dimensional projection data that is closest to the image data captured by the camera, based on camera parameters 331 that indicate the position and direction of the camera capturing the image. This two-dimensional projection data is stored in a virtual legal image feature database 34. The virtual legal image feature database 34 further stores a LIDAR / camera transformation matrix 341 and Raycast resolution data 342.
[0025] During operation, the surveillance camera 35 captures an object and generates an RGB image 351. The video analysis device 1 generates feature data 352 from the RGB image 351 and passes it to the feature matching unit 36 and the feature comparison unit 37. The feature matching unit 36 compares the feature points of the two-dimensional projection data stored in the virtual normal image feature database 34 with the feature data 352. The feature matching unit 36 then outputs the best-matching normal image feature 361. The feature comparison unit 37 compares the feature data 352, the best-matching normal image feature 361, and the feature amount at the reference time to generate a height-specific match amount histogram 371. The water level estimation unit 38 estimates the current water level by evaluating this height-specific match amount histogram 371.
[0026] FIG. 5 is a diagram showing a specific hardware configuration of the video analysis device 1 of the first embodiment. The video analysis device 1 is, for example, a computer, and incorporates a feature processing plug-in 43 and a water level measurement plug-in 44. The video analysis device 1 further stores a shooting information database 481, three-dimensional data 482, an image feature database 483, feature matching processing results 484, and analysis results 485. The video analysis device 1 is connected to a surveillance camera 41, a video management device 42, a three-dimensional measurement device 45 such as a LIDAR, and an initial data generation device 46.
[0027] The surveillance camera 41 captures video and passes it on to the video management device 42. The video management device 42 stores the video captured by the surveillance camera 41 and outputs the image data to a feature processing plug-in 43 and a water level measurement plug-in 44.
[0028] The three-dimensional measuring device 45 is, for example, a LIDAR, which measures the object in three dimensions, writes the photography location ID to the photography information database 481, and writes the mesh data and the photography location ID to the three-dimensional data 482. The initial data generation device 46 includes a virtual normal image generation unit 461 and a feature extraction unit 462. The virtual normal image generation unit 461 generates a normal image from the shooting location ID and mesh data. The feature extraction unit 462 generates a feature image and a depth image from the normal image. The initial data generation device 46 stores the normal image, feature image, and depth image in an image feature database 483.
[0029] The feature processing plug-in 43 includes a feature extraction unit 431 and a feature matching unit 432. The feature extraction unit 431 extracts features from image data and outputs them to the feature matching unit 432. The feature matching unit 432 receives the features of the image data as input, and also receives the shooting location ID and camera parameters from the shooting information database 481.
[0030] The feature matching unit 432 transfers the shooting location ID to the image feature database 483 and reads the feature image corresponding to the shooting location ID. The feature matching unit 432 matches the features of the image data with the feature image and selects a normal image ID. The feature matching unit 432 stores the normal image ID and the shooting location ID in the feature matching processing result 484.
[0031] The water level measurement plug-in 44 includes a feature extraction unit 441 and a water level measurement unit 442. Image data is input to the feature extraction unit 441 from the video management device 42. The feature extraction unit 441 extracts feature points from the normal image ID and image data and outputs them to the water level measurement unit 442. The water level measurement unit 442 measures the water level from the feature points, normal image ID, feature image, and depth image. The water level measurement unit 442 stores the shooting location ID, shooting time, image data, and water level in the analysis result database 485. The analysis result display device 47 acquires and displays the water level and the like from the analysis result database 485. This makes it possible to measure the water level from the image captured by the monitoring camera 41. Note that although the present embodiment has been described as an example in which three-dimensional data is acquired in advance using LiDAR, a method of utilizing already acquired three-dimensional information may also be used.
[0032] Second Embodiment The video analysis device of the second embodiment creates a bird's-eye view of a construction site or the like from any viewpoint without any distortion due to conversion, making it easy to grasp the progress of the construction work.
[0033] FIG. 6 is a schematic diagram showing the configuration of a video analysis device 1A according to the second embodiment. The video analysis device 1A is configured to include a video feature conversion unit 12, a video feature storage unit 13, a three-dimensional data feature conversion unit 15, a three-dimensional data feature storage unit 16, a spatial alignment unit 17, and a three-dimensional data generation unit 19.
[0034] The three-dimensional data 14 is obtained by measuring an object such as a construction site each time using a three-dimensional measuring device such as LiDAR (Laser Imaging Detection and Ranging). The video data 11 is an image of the same object as the object at the construction site photographed by LiDAR or the like, but photographed by a two-dimensional camera.
[0035] The three-dimensional data feature quantity conversion unit 15 calculates three-dimensional data feature quantities by converting the feature quantities by assuming two-dimensional projection data through perspective projection of the input three-dimensional data 14 from multiple viewpoints. The three-dimensional data feature quantities calculated by the three-dimensional data feature quantity conversion unit 15 are stored in the three-dimensional data feature quantity storage unit 16. The three-dimensional data 14 is acquired each time the construction work progresses.
[0036] The video feature quantity transforming unit 12 calculates video feature quantities by transforming the features of the input video data 11. The video feature quantity transformed by the video feature quantity transforming unit 12 is stored in the video feature quantity storing unit 13.
[0037] The spatial positioning unit 17 compares the data in the video feature amount storage unit 13 and the data in the three-dimensional data feature amount storage unit 16 to determine the two-dimensional projection data that is closest to the video data 11 . The three-dimensional data generation unit 19 generates texture-mapped three-dimensional data by mapping the texture of the video data onto the three-dimensional data based on the two-dimensional projection data determined by the spatial registration unit 17. In other words, the three-dimensional data generation unit 19 links the video data to mesh information generated from the three-dimensional data corresponding to the two-dimensional projection data determined by the spatial registration unit 17. By rendering this three-dimensional data using, for example, a game engine such as Unity (registered trademark), it is possible to view the object from a free viewpoint.
[0038] FIG. 7 is a data flow diagram of the video analysis device 2 of the modified example. The video analysis device 2 includes a texture extraction unit 22, a mesh data generation unit 25, and a viewpoint conversion rendering unit 26, and is connected to a plurality of cameras 21a and 21b, a long-distance LiDAR 23, and a portable LiDAR 24.
[0039] The multiple cameras 21a, 21b are installed at different viewpoints and capture images of the object from multiple viewpoints. The long-range LiDAR 23 measures three-dimensional data of the object from a point far away from the object. The portable LiDAR 24 measures three-dimensional data of the object from a point close to the object.
[0040] The mesh data generating unit 25 integrates the three-dimensional data of the object to generate mesh data. The mesh data generated by the mesh data generating unit 25 is output to the texture extracting unit 22.
[0041] The texture extraction unit 22 receives input of images from multiple cameras 21 a and 21 b. The texture extraction unit 22 then converts the three-dimensional mesh data into two-dimensional projection data as seen from each viewpoint of the cameras 21 a and 21 b, associates each piece of two-dimensional projection data with each image, and extracts textures from the images to be applied to the mesh data. In this way, the texture extraction unit 22 creates mesh data to which the textures have been applied.
[0042] The viewpoint conversion rendering unit 26 is, for example, a game engine Unity (registered trademark), and renders the textured mesh data from an arbitrary viewpoint, thereby making it possible to observe an object, for example, a construction site, from a desired viewpoint. In this embodiment, three-dimensional information is sequentially captured by LiDAR, but a method of utilizing already acquired three-dimensional information, such as BIM (Building Information Modeling) or CIM (Construction Information Modeling), may also be used.
[0043] Third Embodiment The video analysis device of the third embodiment measures the length and volume of heavy temporary construction materials on the spot, simply by taking a picture with a mobile device such as a smartphone, based on three-dimensional data acquired by LiDAR. The video analysis device grasps the condition of the material bundle from the video captured by the mobile device and performs a matching process with the length analysis information of the cross section. This makes it possible to measure, for example, a 20-meter-long material from a distance of 20 meters and classify the length of the material in 10-cm increments in real time.
[0044] The video analysis device of the third embodiment can measure the length and volume of all the cargo on a truck at once, allowing safe and efficient measurement of the cargo, which previously required workers to climb onto the loading platform one by one.
[0045] FIG. 8 is a schematic diagram showing the configuration of a video analysis device 1B according to the third embodiment. The video analysis device 1 is configured to include a video feature conversion unit 12, a video feature storage unit 13, a three-dimensional data feature conversion unit 15, a three-dimensional data feature storage unit 16, a video recognition unit 101, and a measurement unit 102.
[0046] The three-dimensional data 14 is obtained by measuring an object in advance using a three-dimensional measuring device such as a LIDAR (Laser Imaging Detection and Ranging). The video data 11 is an image of the same object photographed by the LIDAR or the like, and is photographed by a two-dimensional camera.
[0047] The three-dimensional data feature quantity conversion unit 15 calculates three-dimensional data feature quantities by converting the feature quantities by assuming two-dimensional projection data through perspective projection of the input three-dimensional data 14 from multiple viewpoints. The three-dimensional data feature quantity calculated by the three-dimensional data feature quantity conversion unit 15 is stored in the three-dimensional data feature quantity storage unit 16. The three-dimensional data 14 may be obtained by measuring only once initially, and it is not necessary to measure the object continuously.
[0048] The video feature quantity transforming unit 12 calculates video feature quantities by transforming the features of the input video data 11. The video feature quantity transformed by the video feature quantity transforming unit 12 is stored in the video feature quantity storing unit 13.
[0049] The video recognition unit 101 performs image recognition of the object from the video data 11, recognizes whether it is a bundle of target materials, and if the objects are overlapping, recognizes the number of individual pieces. The video recognition unit 101 then compares the feature data in the video feature storage unit 13 with the feature data in the three-dimensional data feature storage unit 16 to determine the two-dimensional projection data that is closest to the video data 11, and recognizes the three-dimensional data that corresponds to the recognized object.
[0050] The measurement unit 102 measures the size of the object based on the range of the three-dimensional data of the object recognized by the image recognition unit 101.
[0051] Specifically, this size measurement involves measuring the length, width, and length of the recognized object (for example, a bundle of materials) based on 3D data, and then calculating the volume from that data, or detecting the number and shape of the materials contained in the bundle of materials from the image and measuring the length of each individual item.
[0052] FIG. 9 is a flowchart of the measurement visualization process. First, the user measures three-dimensional data from the side of the material using LiDAR (step S20). Next, the user photographs the cross section and outline of the material using a camera (step S21). The measured three-dimensional data and the photographed video data are input to the video analysis device 1B. The video recognition unit 101 of the video analysis device 1B performs image recognition of the material bundle from the video data 11 (step S22).
[0053] The video recognition unit 101 of the video analysis device 1B performs a matching process between the video data 11 and the three-dimensional data 14 to grasp the state of the material bundle (step S23). This makes it possible to recognize the material bundle portion of the three-dimensional data 14. Furthermore, the video recognition unit 101 of the video analysis device 1B analyzes the number of target materials from the cross-sectional shape (step S24). After that, the measurement unit 102 of the video analysis device 1B analyzes the longitudinal length of the materials (step S25), and the processing of FIG. 9 ends.
[0054] The configuration and effects of the present invention will be described below.
[0055] [1] a three-dimensional data feature quantity conversion unit (15) that extracts feature quantities from three-dimensional data (14) measured by a three-dimensional measurement device by assuming two dimensions through perspective projection from multiple viewpoints; a three-dimensional data feature quantity storage unit (16) for storing the feature quantities extracted by the three-dimensional data feature quantity conversion unit; an image feature quantity conversion unit (12) that extracts features from image data (11) of the object captured by a camera; a video feature storage unit (13) that stores the feature extracted by the video feature conversion unit (12); a spatial positioning unit (17) that compares the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit (16) with the feature amounts of the video data stored in the video feature amount storage unit (13) to determine two-dimensional projection data that has the closest angle of view to the video data (11); A video analysis device comprising:
[0056] This makes it possible to measure the height, length, or size of an object from the image.
[0057] [2] The three-dimensional data feature quantity conversion unit (15) extracts, as the feature quantity, boundary information of a surface detected from a normal line included in the three-dimensional data obtained by measuring the object. 2. The video analysis device according to claim 1.
[0058] This allows us to obtain boundary information for surfaces detected from normals from three-dimensional information, and measure the water level from the relationship between this boundary surface information and the waterline and edge information of the image.
[0059] [3] The video feature conversion unit (12) extracts edge information from the video data (11) of the object as the feature. 2. The video analysis device according to claim 1.
[0060] This makes it possible to easily associate the video data with the three-dimensional data.
[0061] [4] The image feature conversion unit (12) uses the three-dimensional data when extracting features from the image data (11) of the object. 2. The video analysis device according to claim 1.
[0062] This allows the video data and the three-dimensional data to be associated with each other without error.
[0063] [5] The video feature conversion unit (12) extracts the feature using a plurality of video frames in different time series. 2. The video analysis device according to claim 1.
[0064] This makes it possible to remove noise from the video data and suitably measure the height, length, or size of an object.
[0065] [6] The spatial positioning unit (17) performs a comparison at a predetermined height or higher based on the three-dimensional data (14). 2. The video analysis device according to claim 1.
[0066] This prevents the incorrect measurement of impossible water levels.
[0067] [7] The image feature conversion unit (12) extracts features based on the position information of the three-dimensional measuring device and the position information of the camera. 2. The video analysis device according to claim 1.
[0068] This makes it possible to reduce the cost of calculating two-dimensional projection data as seen from a camera from three-dimensional data measured by a three-dimensional measurement device.
[0069] [8] a measurement unit (102) that measures the size of the object in the video by comparing the feature amount of the two-dimensional projection data determined by the spatial positioning unit (17) with the feature amount of the video data stored in the video feature amount storage unit (13); 2. The video analysis device according to claim 1, further comprising:
[0070] This makes it possible to easily calculate the size of an object captured in the video data.
[0071] [9] The video feature conversion unit (12) extracts features from a plurality of frames of the video data (11). 9. The video analysis device according to claim 8.
[0072] This makes it possible to remove noise from the video data and suitably measure the height, length, or size of an object.
[0073]
[10] a display control unit (100) that displays the image data (11) and three-dimensional data after alignment, and water levels of caution and danger superimposed on the three-dimensional data; 9. The video analysis device according to claim 8, further comprising:
[0074] This allows the user to be notified of a water level warning along with the underlying data.
[0075]
[11] a three-dimensional data generation unit (19) that links the video data to mesh information generated from the three-dimensional data corresponding to the two-dimensional projection data determined by the spatial positioning unit (17); 2. The video analysis device according to claim 1, further comprising:
[0076] This allows the object measured by the three-dimensional measuring device to be viewed from multiple viewpoints.
[0077]
[12] an image recognition unit (101) that performs image recognition of an object to be detected from the image data; a measurement unit (102) that measures the size of an object in the image portion recognized by the image recognition unit (101) based on mesh information linked by the image recognition unit (101); 2. The video analysis device according to claim 1, further comprising:
[0078] This makes it easy to measure the size of the object.
[0079]
[13] The image recognition unit (101) recognizes the number of pieces or the number of individuals, The measuring unit (102) measures the length of each of the individuals. 13. The video analysis device according to claim 12.
[0080] This makes it easy to measure the number and length of objects.
[0081]
[14] A step in which a three-dimensional data feature quantity conversion unit (15) extracts feature quantities from three-dimensional data (14) obtained by measuring an object using a three-dimensional measurement device, assuming two dimensions by perspective projection from multiple viewpoints; a step of storing the feature quantity extracted by the three-dimensional data feature quantity conversion unit (15) in a three-dimensional data feature quantity storage unit (16); A video feature conversion unit (12) extracts features from video data of the object captured by a camera; storing the feature extracted by the video feature conversion unit (12) in a video feature storage unit (13); a step in which a spatial positioning unit (17) compares the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit (16) with the feature amounts of the video data stored in the video feature amount storage unit (13) to determine two-dimensional projection data having an angle of view closest to the video data; A video analysis method comprising:
[0082] This makes it possible to measure the height, length, or size of an object from the image.
[0083]
[15] On the computer, A procedure for extracting feature quantities from three-dimensional data (14) obtained by measuring an object using a three-dimensional measuring device, assuming two dimensions by perspective projection from multiple viewpoints; a step of storing the feature amount extracted from the three-dimensional data in a three-dimensional data feature amount storage unit (16); A step of extracting feature amounts from video data (11) of the object captured by a camera; a step of storing the feature amount extracted from the video data in a video feature amount storage unit (13); a step of comparing the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit (16) with the feature amounts of the video data stored in the video feature amount storage unit (13) to determine two-dimensional projection data having an angle of view closest to the video data; A video analysis program for executing the above.
[0084] This makes it possible to measure the height, length, or size of an object from the image.
[0085] (Variation) The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. It is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is also possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0086] The above-described configurations, functions, processing units, processing means, etc. may be realized in part or in whole by hardware such as an integrated circuit. The above-described configurations, functions, etc. may be realized by software by a processor interpreting and executing a program that realizes each function. Information such as the programs, tables, and files that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or on a storage medium such as a flash memory card or a DVD (Digital Versatile Disk).
[0087] In each embodiment, the control lines and information lines shown are those that are considered necessary for the explanation, and not all control lines and information lines in the product are necessarily shown. In reality, it can be considered that almost all components are interconnected. [Explanation of symbols]
[0088] 1, 1A, 1B Video analysis equipment 11 Video data 12 Video feature conversion unit 13 Video feature storage unit 14 Three-dimensional data 15 3D data feature conversion unit 16 Three-dimensional data feature storage unit 17 Spatial alignment section 18 Measurement section 2. Video analysis equipment 22 Texture Extraction Unit 25 Mesh Data Generation Department 26 Viewpoint conversion rendering section 21a Camera 21b Camera 23 Long Range LiDAR 24 Portable LiDAR 31 LIDAR 321 point cloud 322 mesh 33 Raycast section 331 Camera Parameters 34 Virtual Legal Image Feature Database 341 LIDAR / Camera Transformation Matrix 342 Raycast resolution data 35 Surveillance Camera 351 RGB images 352 feature data 36 Feature Matching Unit 37 Feature Comparison Section 361 Normal Image Features 371 Height-specific match quantity histogram 38 Water level estimation section 43 Feature Processing Plug-ins 44 Water Level Measurement Plugin 481 Photography Information Database 482 3D data 483 Image Feature Database 41 Surveillance Camera 42 Video Management Device 45 Three-dimensional measuring device 46 Initial Data Generator 461 Virtual normal image generation unit 462 Feature Extraction Unit 431 Feature Extraction Unit 432 Feature Matching Unit 441 Feature Extraction Unit 442 Water Level Measurement Unit 485 Analysis Results Database 47 Analysis result display device 19 Three-dimensional data generation unit
Claims
1. a three-dimensional data feature quantity conversion unit that extracts feature quantities by assuming two dimensions through perspective projection from multiple viewpoints from three-dimensional data obtained by measuring an object using a three-dimensional measurement device; a three-dimensional data feature quantity storage unit for storing the feature quantities extracted by the three-dimensional data feature quantity conversion unit; an image feature conversion unit that extracts features of image data obtained by capturing the object with a camera; a video feature storage unit for storing the feature extracted by the video feature conversion unit; a spatial registration unit that compares the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit with the feature amounts of the video data stored in the video feature amount storage unit, and determines two-dimensional projection data that has an angle of view closest to the video data; A video analysis device comprising:
2. the three-dimensional data feature quantity conversion unit extracts, as the feature quantity, boundary information of a surface detected from a normal line included in the three-dimensional data obtained by measuring the object; 2. The video analysis device according to claim 1.
3. the video feature conversion unit extracts edge information from the video data capturing the object as the feature.
2. The video analysis device according to claim 1.
4. the video feature quantity conversion unit uses the three-dimensional data when extracting features from the video data of the object; 2. The video analysis device according to claim 1.
5. the video feature conversion unit extracts the feature using a plurality of video frames in different time series; 2. The video analysis device according to claim 1.
6. the spatial positioning unit performs a comparison at a predetermined height or higher based on the three-dimensional data; 2. The video analysis device according to claim 1.
7. the image feature conversion unit extracts features based on position information of the three-dimensional measuring device and position information of the camera; 2. The video analysis device according to claim 1.
8. a measurement unit that measures a size of the object in the video by comparing a feature of the two-dimensional projection data determined by the spatial positioning unit with a feature of the video data stored by the video feature storage unit; 2. The video analysis device according to claim 1, further comprising:
9. the video feature conversion unit extracts features from information on a plurality of frames of the video data; 9. The video analysis device according to claim 8.
10. a display control unit that displays the image data and three-dimensional data after alignment, and water levels of caution and danger superimposed on the three-dimensional data; The video analysis device according to claim 8, further comprising:
11. a three-dimensional data generation unit that associates the video data with mesh information generated from the three-dimensional data corresponding to the two-dimensional projection data determined by the spatial positioning unit; 2. The video analysis device according to claim 1, further comprising:
12. an image recognition unit that performs image recognition of an object to be detected from the image data; a measuring unit that measures the size of an object in the image portion recognized by the image recognition unit based on mesh information linked by the image recognition unit; 2. The video analysis device according to claim 1, further comprising:
13. The image recognition unit recognizes the number or quantity of the individual pieces, The measurement unit measures the length of each of the individuals.
13. The video analysis device according to claim 12.
14. a step in which a three-dimensional data feature quantity conversion unit extracts feature quantities from three-dimensional data obtained by measuring an object using a three-dimensional measurement device, assuming two dimensions by perspective projection from multiple viewpoints; storing the feature quantity extracted by the three-dimensional data feature quantity conversion unit in a three-dimensional data feature quantity storage unit; a video feature conversion unit extracting features from video data of the object captured by a camera; storing the feature extracted by the video feature conversion unit in a video feature storage unit; a spatial positioning unit comparing the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit with the feature amounts of the video data stored in the video feature amount storage unit, and determining two-dimensional projection data having an angle of view closest to that of the video data; A video analysis method comprising:
15. On the computer, A procedure for extracting feature quantities from three-dimensional data measured by a three-dimensional measurement device by assuming two dimensions through perspective projection from multiple viewpoints; a step of storing the feature amounts extracted from the three-dimensional data in a three-dimensional data feature amount storage unit; A step of extracting feature amounts from video data of the object captured by a camera; a step of storing the feature amount extracted from the video data in a video feature amount storage unit; a step of comparing the feature amounts of the three-dimensional data stored in the three-dimensional data feature amount storage unit with the feature amounts of the video data stored in the video feature amount storage unit, and determining two-dimensional projection data having an angle of view closest to that of the video data; A video analysis program for executing the above.
Citation Information
Patent Citations
Image processing apparatus and environmental information observation system
JP2007018347A
Water level observation system by image processing
JP2008057994A
Water level measurement device and water level measurement method
WO2020121376A1
Object amount calculation device and object amount calculation method
WO2020179438A1