Stockpile volume measurement system based on surveillance camera
Through a material stack volume measurement system based on a monitoring camera, combined with deep learning technology for image processing and three-dimensional modeling, the existing lidar measurement methods are solved, and fully automatic, accurate and efficient material stack volume measurement is achieved.
Patent Information
- Application Number
- CN202411765863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-04
AI Technical Summary
The existing lidar-based material stack measurement methods cannot achieve fully automatic, accurate and efficient measurements, and are cumbersome to operate and have long data acquisition time.
The stack volume measurement system based on the monitoring camera is adopted, and the combination of the control center unit, network switch unit and monitoring camera module is used to extract image features, feature matching, motion structure modeling and dense reconstruction using deep learning technology to realize grid processing and volume calculation of the three-dimensional model.
It realizes the measurement of large material piles automatically, accurately and efficiently, with simple operation and short data acquisition time, reducing the cost of lidar use, accurate point cloud splicing and high measurement efficiency.
Smart Images

Figure CN119251282B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of material pile volume measurement, and in particular to a material pile volume measurement system based on a monitoring camera. Background Art
[0002] As a major resource consumer, my country requires a large amount of coal, ore and other mineral resources as raw materials in production and life. Most of these mineral resources are stored in ports, mines or related enterprises in the form of large bulk stockpiles. The price of mineral resources is high, and the stock changes dynamically, so it is necessary to regularly check the volume or weight of the stockpiles.
[0003] At present, most of the large stockpile measurements are based on LiDAR, which is mainly divided into two modes: mobile and fixed. Among them, the mobile mode requires the radar to be installed on the top of the vehicle or on the cantilever, which depends on the movement of the vehicle or the cantilever. Not only is the operation cumbersome, but the data collection time is also long; while the fixed mode requires the combination of the pan-tilt head and the radar to scan the stockpile, and each measurement point is measured separately. Not only is the point cloud stitching inaccurate, but the background elements need to be interactively eliminated, and a complete measurement takes a long time. Therefore, the LiDAR-based measurement method cannot achieve fully automatic, accurate, and efficient measurement of large stockpiles.
[0004] In view of this, the present application proposes a stockpile volume measurement system based on a surveillance camera. Summary of the invention
[0005] The object of the present invention is to provide a stockpile volume measurement system based on a monitoring camera to solve the above problems.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A stockpile volume measurement system based on a monitoring camera comprises a control center unit, a network switch unit and a monitoring camera module, wherein the control center unit and the network switch unit are connected via a local area network, and the monitoring camera module is connected to the network switch unit via a network cable to form a local area network;
[0008] The control center unit is used for data processing and storage, receiving data from the monitoring camera module and performing analysis and archiving;
[0009] The network switch unit is used to connect all monitoring camera modules with the control center unit to achieve stable data transmission;
[0010] The monitoring camera module is used to set a plurality of preset shooting points, shoot at different positions to obtain two-dimensional image data, and calibrate the data to ensure shooting errors.
[0011] Furthermore, before the monitoring camera module sets the preset shooting point, the camera is calibrated to ensure that the intrinsic distortion coefficient of each monitoring camera is preserved, the calibration error is confirmed, and then the multiple monitoring cameras in the monitoring camera module are installed in the factory and a bridge is built.
[0012] Furthermore, when the control center unit processes and stores data, the following steps are also included:
[0013] S1. Image feature extraction and image feature matching based on deep learning: The image acquired from the surveillance camera is subjected to distortion correction to obtain an image without distortion, and then the image is sent to a deep neural network for feature extraction, and the extracted feature points are subjected to feature matching to obtain the spatial position relationship between images;
[0014] S2, Structure from Motion based on deep learning: After deep learning-based image feature extraction and image feature matching, triangulation is performed to calculate the position of feature points in three-dimensional space based on the known camera position and feature points in the image, thereby generating a new SfM model;
[0015] S3, dense reconstruction: After obtaining the SFM model, the model is meshed using dense multi-view stereo pipeline technology;
[0016] S4, texture mapping: mapping the densely reconstructed mesh;
[0017] S5. Volume calculation: Through the above steps, a three-dimensional model of the field can be obtained, and the volume of the materials in the field can be measured based on the three-dimensional model.
[0018] Furthermore, in the step S1, when extracting features from the image, two sets of local features are matched by jointly searching for correspondences and rejecting unmatched points, and the transportation cost is estimated by solving a differentiable optimal transportation problem, and the cost is predicted by a graph neural network.
[0019] Furthermore, in step S1, when calculating the spatial position relationship between the images, the position coordinates of the feature points in the two images are input. and And the feature description vector corresponding to the feature point and The position coordinates p i Contains x, y coordinate values and detection confidence c, i.e. p i =(x, y, c) i , position coordinate p i After being processed by the encoder, it is combined with the feature description vector d i Add together to get the local features. The relevant calculation formula is as follows:
[0020] (0) x i =d i +MLP enc (p i ).
[0021] Furthermore, after obtaining the local features, the calculation process of the feature point aggregation information to be calculated through the attention mechanism is as follows:
[0022] m ε→i =∑ j:(i,j)∈ε α ij v j ;
[0023] in This process is similar to retrieving data from a database. i represents the query vector, k j represents a key, and v j Indicates the value corresponding to each key.
[0024] Furthermore, during the gridding process in step S3, a quasi-dense point set is extracted from the image, and multiple points are matched in pairs between different views. From the above matching, a quasi-dense 3D point cloud is generated by reconstructing and optionally merging the triangulated 3D points. Given a reference image X ref And the original image X src ={X m |m=1...M}, for the I-th pixel on the reference To estimate the depth θ l and comparison and The color similarity of is as follows:
[0025]
[0026] in is the occlusion label, is the normalized shadow,
[0027] is uniformly distributed, is the color similarity between two patches, σ ρ is the smoothing coefficient.
[0028] Further, the processed point cloud is fed to the second stage, which constructs a Delaunay triangulation from it, then robustly extracts the initial surface from the facets of this triangulation, filters out most of the outliers, improves the quality of the recovered surface by using a hybrid photoconsistency and smoothness criterion, and finally obtains the mesh.
[0029] Furthermore, in step S5, the volume measurement includes the following steps:
[0030] S51, determining the material pile ground: performing plane fitting on the mesh grid to obtain the normal vector of the material pile ground, so that the direction of the normal vector is perpendicular to the ground and upward;
[0031] S52, grid coordinate conversion: converting the grid to the ground coordinate system;
[0032] S53, extracting wall point cloud and determining wall sequence number: extracting point cloud from the grid, estimating the normal vector of each point cloud, performing threshold processing through the z value of the normal vector, obtaining candidate wall point cloud, fitting the point cloud to a plane, obtaining all wall plane coefficients, and determining the order and distance of the walls according to the color of the marking points on the wall;
[0033] S54, constructing a material pile coordinate system: the coordinate system z-axis is the ground normal vector, the wall normal vector is the x-axis, and the y-axis is the coordinate system of the model according to the cross product of the x-axis and the z-axis;
[0034] S55, splitting the pile grid: constructing a space body according to the marking points of the map, the marking points on the ground and the wall, and then manually giving the x, y coordinates of each pile, and finally cutting out each pile;
[0035] S56. Calculate the volume enclosed by the segmented material pile grid and the material pile ground. By traversing each triangle in the triangular grid, calculate the volume contribution composed of each triangle and the xoy plane, and accumulate it into the total volume. Then, obtain the volume scaling factor through the distance between the two reconstructed walls and the actual wall distance or the distance between the reconstructed marking points and the actual marking point distance, and obtain the actual material pile volume.
[0036] Furthermore, in step S56, the volume calculation formula of the triangle and the xoy plane is as follows:
[0037] z_avg=(v0.z()+v1.z()+v2.z()) / 3.0;
[0038] vol = (area * u_z * z_avg);
[0039] Where z_avg is the average height of the three vertices of the triangle, area is the area of the triangle, and u_z is the component in the z direction.
[0040] Beneficial effects of the present invention:
[0041] In the present invention, a new SfM model can be generated through image feature extraction and image feature matching based on deep learning and motion structure based on deep learning, and then the meshing processing of the model can be realized through dense reconstruction, and the openMVS technology is used to texture the three-dimensional model, and the texture is projected onto the mesh model by plane projection, and texture coordinates are assigned to each mesh vertex so that the texture is mapped onto the mesh surface, and then the texture image is mapped to the mesh model, and corresponding texture pixels are assigned to each mesh face or vertex according to the texture coordinates; finally, the texture mapping result is used for rendering to display the mesh model with texture on the screen, and then volume measurement is performed based on the The actual volume of materials in the field is calculated with the help of the three-dimensional model, and the measurement of large stockpiles can be completed automatically, accurately and efficiently. By introducing an attention mechanism network based on a flexible context aggregation mechanism, the neural network can jointly reason about the underlying 3D scene and function allocation. Compared with traditional hand-designed geometry, neural network technology learns geometric transformations and prior knowledge of the 3D world through end-to-end training of image pairs, which can improve the accuracy of feature matching. The operation is relatively simple, the data collection time is short, and the cost of using lidar is reduced when scanning stockpiles. There is no need to measure the measurement points separately, the point cloud stitching is more accurate, the measurement takes less time to complete, and the measurement efficiency is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a system block diagram of the stockpile volume measurement system based on the monitoring camera of the present invention;
[0043] Figure 2 The figure is a flow chart of the stockpile volume measurement system based on monitoring camera of the present invention.
[0044] In the figure: 1. Control center unit; 2. Network switch unit; 3. Monitoring camera module. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] Example 1: Please refer to Figure 1-Figure 2 ,This design proposes an implementation method, a stockpile volume measurement system based on monitoring cameras, including a control center unit 1, a network switch unit 2 and a monitoring camera module 3, the control center unit 1 and the network switch unit 2 are connected through a local area network, and the monitoring camera module 3 is connected to the network switch unit 2 through a network cable to form a local area network;
[0047] The control center unit 1 is used for data processing and storage, receiving data from the monitoring camera module 3 and analyzing and archiving it;
[0048] The network switch unit 2 is used to connect all the monitoring camera modules 3 with the control center unit 1 to achieve stable data transmission;
[0049] The monitoring camera module 3 is used to set multiple preset shooting points, shoot at different positions to obtain two-dimensional image data, and calibrate the data to ensure shooting errors; before setting the preset shooting points in the monitoring camera module 3, the camera is calibrated to ensure that the intrinsic distortion coefficient of each monitoring camera is preserved, the calibration error is confirmed, and then the multiple monitoring cameras in the monitoring camera module 3 are installed in the factory and a bridge is built.
[0050] In this embodiment, the camera is first calibrated to ensure that the intrinsic distortion coefficient of each monitoring camera is preserved and the calibration error is less than 0.5 pixels. By controlling the calibration error, the accuracy of the monitoring camera's image data acquisition can be guaranteed, the number of invalid images can be reduced, and the workload of subsequent image data processing can be reduced. Then, the monitoring cameras are installed in the factory and a bridge is built to achieve the installation and use of multiple monitoring cameras; then, preset shooting points are set for each camera and the preset points are calibrated to reduce position deviations. Finally, photos are taken at the preset points to obtain the required two-dimensional image data, thereby achieving preliminary acquisition of image data. Then, the data is stably transmitted to the control center unit 1 through the network switch unit 2, and the control center unit 1 is used to quickly process and store data to avoid data loss.
[0051] Embodiment 2: When the control center unit 1 processes and stores data, the following steps are also included:
[0052] Step 1: Image feature extraction and image feature matching based on deep learning: The image acquired from the surveillance camera is subjected to distortion correction to obtain an image without distortion, and then the image is sent to a deep neural network for feature extraction, and the extracted feature points are subjected to feature matching to obtain the spatial position relationship between images;
[0053] When extracting features from an image, two sets of local features are matched by jointly finding correspondences and rejecting mismatched points. The transportation cost is estimated by solving a differentiable optimal transportation problem, and the cost is predicted by a graph neural network.
[0054] When calculating the spatial position relationship between images, the position coordinates of the feature points in the two images are input and And the feature description vector corresponding to the feature point and The position coordinates p i Contains x, y coordinate values and detection confidence c, i.e. p i =(x, y, c) i , position coordinate p i After being processed by the encoder, it is combined with the feature description vector d i Add together to get the local features. The relevant calculation formula is as follows:
[0055] (0) x i =d i +MLP enc (p i ).
[0056] After obtaining the local features, the calculation process of the feature point aggregation information to be calculated through the attention mechanism is as follows:
[0057] m ε→i =∑ j:(i,j)∈ε a ij v j ;
[0058] in This process is similar to retrieving data from a database. i represents the query vector, k j represents a key, and v j Indicates the value corresponding to each key;
[0059] By introducing an attention mechanism network based on a flexible context aggregation mechanism, the neural network can jointly reason about the underlying 3D scene and function allocation. Compared with traditional hand-designed geometry, the neural network technology learns geometric transformations and prior knowledge of the 3D world through end-to-end training of image pairs, which can improve the accuracy of feature matching.
[0060] Step 2: Structure from motion based on deep learning: After deep learning-based image feature extraction and image feature matching, triangulation is performed to calculate the position of feature points in three-dimensional space based on the known camera position and feature points in the image, thereby generating a new SfM model.
[0061] Step 3: Dense reconstruction: After obtaining the SFM model, the model is meshed using dense multi-view stereo pipeline technology;
[0062] During the gridding process, a quasi-dense point set is extracted from the image, multiple points are matched pairwise between different views, and from the above matches, a quasi-dense 3D point cloud is generated by reconstructing and optionally merging the triangulated 3D points. Given a reference image X ref And the original image X src ={X m|m=1...M}, for the I-th pixel on the reference To estimate the depth θ l and comparison and The color similarity of is as follows:
[0063]
[0064] in is the occlusion label, is the normalized shadow,
[0065] is uniformly distributed, is the color similarity between two patches, σ ρ is the smoothing coefficient;
[0066] The processed point cloud is fed to the second stage which builds a Delaunay triangulation from it, then robustly extracts the initial surface from the facets of this triangulation, filters out most outliers, improves the quality of the recovered surface by using a hybrid photoconsistency and smoothness criterion, and finally obtains the mesh.
[0067] Step 4: Texture mapping: Map the densely reconstructed mesh, use openMVS technology to texture the 3D model, project the texture onto the mesh model using a plane projection method, assign texture coordinates to each mesh vertex so that the texture can be mapped onto the mesh surface, and then map the texture image onto the mesh model, assigning corresponding texture pixels to each mesh face or vertex according to the texture coordinates;
[0068] Step 5: Volume calculation: Through the above steps, a three-dimensional model of the field can be obtained, and the volume of the materials in the field can be measured according to the three-dimensional model;
[0069] Volume measurement involves the following steps:
[0070] 51. Determine the material pile ground: perform plane fitting on the mesh grid to obtain the normal vector of the material pile ground, making the direction of the normal vector perpendicular to the ground and upward;
[0071] 52. Grid coordinate conversion: convert the grid to the ground coordinate system;
[0072] 53. Extract wall point cloud and determine wall sequence number: extract point cloud from the grid, estimate the normal vector of each point cloud, perform threshold processing through the z value of the normal vector, obtain candidate wall point cloud, fit the point cloud to the plane, obtain all wall plane coefficients and determine the order and distance of the wall according to the color of the marked points on the wall;
[0073] 54. Construct the material pile coordinate system: the coordinate system z-axis is the ground normal vector, the wall normal vector is the x-axis, and the y-axis is the coordinate system of the model based on the cross product of the x-axis and the z-axis;
[0074] 55. Split the pile grid: construct a space body based on the marked points of the map, the marked points on the ground and the wall, then manually give the x, y coordinates of each pile, and finally cut out each pile;
[0075] 56. Calculate the volume enclosed by the segmented pile grid and the pile ground. By traversing each triangle in the triangular grid, calculate the volume contribution composed of each triangle and the xoy plane, and accumulate it into the total volume. Then, obtain the volume scaling factor through the distance between the two reconstructed walls and the actual wall distance or the distance between the reconstructed marking points and the actual marking point distance, and obtain the actual pile volume. The volume calculation formula composed of the triangle and the xoy plane is as follows:
[0076] z_avg=(v0.z()+v1.z()+v2.z()) / 3.0;
[0077] vol = (area * u_z * z_avg);
[0078] Where z_avg is the average height of the three vertices of the triangle, area is the area of the triangle, and u_z is the component in the z direction.
[0079] In the present invention, a new SfM model can be generated through image feature extraction and image feature matching based on deep learning and motion structure based on deep learning, and then the meshing processing of the model can be realized through dense reconstruction, and the openMVS technology is used to texture the three-dimensional model, and the texture is projected onto the mesh model by plane projection, and texture coordinates are assigned to each mesh vertex so that the texture is mapped onto the mesh surface, and then the texture image is mapped to the mesh model, and corresponding texture pixels are assigned to each mesh face or vertex according to the texture coordinates; finally, the texture mapping result is used for rendering to display the mesh model with texture on the screen, and then volume measurement is performed based on the The actual volume of materials in the field is calculated with the help of the three-dimensional model, and the measurement of large stockpiles can be completed automatically, accurately and efficiently. By introducing an attention mechanism network based on a flexible context aggregation mechanism, the neural network can jointly reason about the underlying 3D scene and function allocation. Compared with traditional hand-designed geometry, neural network technology learns geometric transformations and prior knowledge of the 3D world through end-to-end training of image pairs, which can improve the accuracy of feature matching. The operation is relatively simple, the data collection time is short, and the cost of using lidar is reduced when scanning stockpiles. There is no need to measure the measurement points separately, the point cloud stitching is more accurate, the measurement takes less time to complete, and the measurement efficiency is high.
[0080] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A stockpile volume measurement system based on a surveillance camera, characterized in that: The system comprises a control center unit (1), a network switch unit (2) and a monitoring camera module (3), wherein the control center unit (1) and the network switch unit (2) are connected via a local area network, and the monitoring camera module (3) is connected to the network switch unit (2) via a network cable, thereby forming a local area network; The control center unit (1) is used for data processing and storage, receiving data from the monitoring camera module (3) and performing analysis and archiving; The network switch unit (2) is used to connect all monitoring camera modules (3) and the control center unit (1) to achieve stable data transmission; The monitoring camera module (3) is used to set a plurality of preset shooting points, shoot at different positions to obtain two-dimensional image data, and calibrate the data to ensure shooting errors. Before setting the preset shooting points, the camera is calibrated to ensure that the internal distortion coefficient of each monitoring camera is preserved, the calibration error is confirmed, and then the plurality of monitoring cameras in the monitoring camera module (3) are installed in the factory building, and a bridge frame is built; When the control center unit (1) processes and stores data, the following steps are also included: S1. Image feature extraction and image feature matching based on deep learning: The image acquired from the surveillance camera is subjected to distortion correction to obtain a distortion-free image, and then the image is sent to a deep neural network for feature extraction, and the extracted feature points are subjected to feature matching to obtain the spatial position relationship between images; S2, Structure from Motion based on deep learning: After deep learning-based image feature extraction and image feature matching, triangulation is performed to calculate the position of feature points in three-dimensional space based on the known camera position and feature points in the image, thereby generating a new SfM model; S3, dense reconstruction: After obtaining the SFM model, the model is gridded using the dense multi-view stereo pipeline technology. During the gridding process, a quasi-dense point set is extracted from the image, and multiple points are matched in pairs between different views. From the above matching, a quasi-dense 3D point cloud is generated by reconstructing and optionally merging the triangulated 3D points. Given a reference image X ref And the original image X src ={X m |m=1...M}, for the lth pixel on the reference To estimate the depth θ l and comparison and The color similarity of is as follows: in is the occlusion label, is the normalized shadow, is uniformly distributed, is the color similarity between two patches, σ ρ is the smoothing coefficient; The processed point cloud is fed to the second stage, which constructs a Delaunay triangulation from it, then robustly extracts the initial surface from the facets of this triangulation, filters out most outliers, and improves the quality of the recovered surface by using a hybrid photoconsistency and smoothness criterion to finally obtain a mesh. S4, texture mapping: mapping the densely reconstructed mesh; S5. Volume calculation: Through the above steps, a three-dimensional model of the field can be obtained. The volume of the materials in the field is measured according to the three-dimensional model. The volume measurement includes the following steps: S51, determining the material pile ground: performing plane fitting on the mesh grid to obtain the normal vector of the material pile ground, so that the direction of the normal vector is perpendicular to the ground and upward; S52, grid coordinate conversion: converting the grid to the ground coordinate system; S53, extracting wall point cloud and determining wall sequence number: extracting point cloud from the grid, estimating the normal vector of each point cloud, performing threshold processing through the z value of the normal vector, obtaining candidate wall point cloud, fitting the point cloud to a plane, obtaining all wall plane coefficients, and determining the order and distance of the walls according to the color of the marking points on the wall; S54, constructing a material pile coordinate system: the coordinate system z-axis is the ground normal vector, the wall normal vector is the x-axis, and the y-axis is the coordinate system of the model according to the cross product of the x-axis and the z-axis; S55, splitting the pile grid: constructing a space body according to the marking points of the map, the marking points on the ground and the wall, and then manually giving the x, y coordinates of each pile, and finally cutting out each pile; S56, calculating the volume enclosed by the segmented material pile grid and the material pile ground, traversing each triangle in the triangular grid, calculating the volume contribution composed of each triangle and the xoy plane, and accumulating it into the total volume, and then obtaining the volume scaling factor through the distance between the two reconstructed walls and the actual wall distance or the distance between the reconstructed marking points and the actual marking point distance, and obtaining the actual material pile volume; The volume calculation formula of the triangle and the xoy plane is as follows: z_avg=(v0.z()+v1.z()+v2.z()) / 3.0; vol = (area * u_z * z_avg); Where z_avg is the average height of the three vertices of the triangle, area is the area of the triangle, and u_z is the component in the z direction.
2. The system for measuring the volume of a stockpile based on a monitoring camera according to claim 1, characterized in that: In step S1, when extracting features from an image, two sets of local features are matched by jointly searching for correspondences and rejecting unmatched points, and the transportation cost is estimated by solving a differentiable optimal transportation problem, and the cost is predicted by a graph neural network.
3. The system for measuring the volume of a stockpile based on a monitoring camera according to claim 1, characterized in that: In step S1, when calculating the spatial position relationship between the images, the position coordinates of the feature points in the two images are input. and And the feature description vector corresponding to the feature point and The position coordinates p i Contains x, y coordinate values and detection confidence c, that is, pi = (x, y, c) i , position coordinate p i After being processed by the encoder, it is combined with the feature description vector d i Add together to get the local features. The relevant calculation formula is as follows: (0) x i =d i +MLP enc (p i )。 4. The system for measuring the volume of a stockpile based on a monitoring camera according to claim 3, characterized in that: After obtaining the local features, the calculation process of the feature point aggregation information to be calculated through the attention mechanism is as follows: in This process is similar to retrieving data from a database. i represents the query vector, k j represents a key, and v j Indicates the value corresponding to each key.
Citation Information
Patent Citations
Coal inventory system and method based on vision
CN112150629A
Cited By
Material coiling method and system based on single-line laser radar
CN122085297A