Steel coil storage yard checking method and device based on three-dimensional reconstruction
Through a three-dimensional reconstruction method, using historical multi-view video data and preset rules to train the network model, the problems of low efficiency and large error in the inventory of steel coil yards are solved, and efficient target positioning and automated management are achieved.
Patent Information
- Application Number
- CN202510415617.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has problems such as low efficiency, large errors and difficulty in obtaining positioning information in the inventory of steel coil yards. Especially when the yard objects are simple in texture, similar shapes and complex distribution, the feature matching method is prone to failure and it is difficult to achieve efficient automated management and logistics scheduling.
Using a three-dimensional reconstruction method, keyframes are determined through historical multi-view video data, sparse three-dimensional point clouds and dense three-dimensional reconstruction data are generated, combined with preset enhanced processing rules and sliding window method, the target yard inventory network model is trained, and the multi-view target result fusion center is used for detection to generate yard inventory results.
It improves the efficiency of steel coil yard target inventory, reduces errors, achieves more accurate positioning and automated management, and improves user experience.
Smart Images

Figure CN120278639A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method and device for inventorying a steel coil yard based on three-dimensional reconstruction. Background Art
[0002] Currently, in the process of industrial production and quality control, inventorying the steel coil yard is an important part of realizing automated management and efficient logistics scheduling. However, the methods for inventorying the targets in the steel coil yard mainly rely on manual operations, which will bring problems such as increased operation and management costs and large inventory errors.
[0003] Due to the problem of perspective limitation in the yard, it is difficult to use a single-angle image to collect complete yard information. Therefore, in the prior art, object detection methods are used to inventory the objects in the steel coil yard, and feature matching methods are used to find the objects that are repeatedly counted in multi-angle images. However, due to the simple texture and similar shapes of the yard objects, the feature matching method is prone to failure in the yard environment, resulting in inaccurate inventory; the objects in the yard are arranged closely and distributed complexly, and the occlusion problem cannot be solved in the real collected picture data; it is difficult to correspond the object detection results to the real world, and it is difficult to obtain additional information such as positioning. Therefore, the above methods are not easy to be popularized in fields such as yard logistics scheduling and cannot meet the actual use needs.
[0004] In recent years, three-dimensional reconstruction methods have shown obvious advantages in solving problems related to scene understanding. Such methods can fuse the data in large-scale pictures to obtain a fused representation of the complete scene. Combining the three-dimensional reconstruction method and the object detection method has the potential to overcome the problem of perspective limitation. However, the information redundancy of the three-dimensional reconstruction results is high, and directly conducting an inventory will lead to difficulties in collecting labeled data and excessive inventory calculation overhead, which limits the practical application of related methods in large-scale yards.
[0005] As can be seen from the above, how to improve the efficiency of inventorying the targets in the steel coil yard during the inventory process of the steel coil yard based on three-dimensional reconstruction is an urgent problem to be solved at present. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a method and device for inventorying a steel coil yard based on three-dimensional reconstruction, which can improve the efficiency of inventorying the targets in the steel coil yard during the inventory process of the steel coil yard based on three-dimensional reconstruction. The specific scheme is as follows:
[0007] In the first aspect, the present application provides a method for inventorying a steel coil yard based on three-dimensional reconstruction, including:
[0008] Determine key frames based on the feature matching relationships among the video frames in the historical multi-view video data corresponding to the steel coil yard, and determine the three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames; the three-dimensional reconstruction data includes imaging information, sparse three-dimensional point cloud data, and dense three-dimensional reconstruction data;
[0009] Determine a three-dimensional cuboid bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, determine the perspective imaging information corresponding to the three-dimensional cuboid bounding box based on a number of viewing angles, and then determine the historical multi-modal rendering map based on each of the perspective imaging information and the dense three-dimensional reconstruction data;
[0010] Process the historical multi-modal rendering map using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and use the training set to train an initial yard inventory network model to obtain a target yard inventory network model;
[0011] Detect the current multi-modal rendering map corresponding to the steel coil yard using the target yard inventory network model and a preset multi-view target result fusion center to obtain a yard inventory result corresponding to the steel coil yard.
[0012] Optionally, the determining key frames based on the feature matching relationships among the video frames in the historical multi-view video data corresponding to the steel coil yard, and determining the three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames, includes:
[0013] Determine corresponding multi-view picture data based on the historical multi-view video data corresponding to the steel coil yard, determine the feature matching relationships among the video frames in each of the multi-view picture data, and then determine key frames based on the feature matching relationships;
[0014] Determine the sparse three-dimensional point cloud data corresponding to the steel coil yard and the imaging information corresponding to each of the multi-view picture data based on each of the key frames using a preset sparse reconstruction process; the imaging information includes imaging points, imaging fields of view, and imaging angles corresponding to the multi-view picture data;
[0015] Determine the dense three-dimensional reconstruction data corresponding to the steel coil yard based on the multi-view picture data, the imaging information, and the sparse three-dimensional point cloud data.
[0016] Optionally, the determining a three-dimensional cuboid bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, and determining the perspective imaging information corresponding to the three-dimensional cuboid bounding box based on a number of viewing angles, includes:
[0017] Processing the imaging information and the sparse three-dimensional point cloud data based on a preset imaging point determination rule to obtain a plurality of imaging points;
[0018] Using a preset three-dimensional cuboid bounding box generation rule and each of the imaging points, and determining the vertex coordinates corresponding to the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, and determining a three-dimensional cuboid bounding box corresponding to the steel coil yard based on each of the vertex coordinates;
[0019] Using a preset perspective imaging information generation rule and performing a generation operation on the three-dimensional cuboid bounding box based on a plurality of viewing angles to obtain corresponding perspective imaging information.
[0020] Optionally, the determining a historical multi-modal rendering map based on each of the perspective imaging information and the dense three-dimensional reconstruction data includes:
[0021] Using a preset block rendering rule and determining a historical multi-modal rendering map of a plurality of initial local RGB images and a plurality of initial local depth maps based on each of the perspective imaging information and the dense three-dimensional reconstruction data; the initial local RGB images and the initial local depth maps do not overlap with each other;
[0022] Performing an alignment and stitching operation on each of the initial local RGB images and each of the initial local depth maps using a preset alignment and stitching rule to obtain a historical multi-modal rendering map including a target local RGB image and a target local depth map; wherein, the viewing angles corresponding to each of the historical multi-modal rendering maps are all different.
[0023] Optionally, the determining a historical multi-modal rendering map based on each of the perspective imaging information and the dense three-dimensional reconstruction data includes:
[0024] Generating a plurality of initial local RGB images using an initial target projection matrix and based on each of the perspective imaging information and the dense three-dimensional reconstruction data, then generating a target projection matrix using a preset matrix generation rule and based on each of the perspective imaging information and the dense three-dimensional reconstruction data, and processing each of the initial local RGB images using the target projection matrix to obtain a target local RGB image;
[0025] Generating a plurality of initial local depth maps based on a preset imaging field of view and based on each of the perspective imaging information and the dense three-dimensional reconstruction data, and then processing the initial local depth maps based on the angle between the viewing angle corresponding to each imaging point and the vertical direction to obtain a target local depth map;
[0026] Wherein, the historical multi-modal rendering map includes the target local RGB image and the target local depth map, and the target local RGB image and the target local depth map have a one-to-one corresponding relationship in pixel positions.
[0027] Optionally, the historical multi-modal rendering map is processed using preset enhancement processing rules, preset annotation rules, and a preset sliding window method to obtain a training set, including:
[0028] Perform spatial enhancement processing and Gaussian filtering on each of the target local RGB images to obtain images to be cut, and perform normalization processing, Gaussian noise addition processing, and Gaussian filtering on each of the target local depth maps to obtain depth maps to be cut;
[0029] Cut the image to be cut of the modal rendering map according to the preset cutting rules to obtain a number of sub-rendered images, determine the corresponding sliding window value based on the picture width of the historical multi-modal rendering map, and use the sliding window value to determine the detection frames to be screened; the image to be cut of the modal rendering map includes the image to be cut and the depth map to be cut;
[0030] Determine the first picture size corresponding to each sub-rendered image in each of the detection frames to be screened and the second picture size corresponding to each historical multi-modal rendering map in each of the detection frames to be screened, and determine the corresponding picture ratio based on each of the first picture sizes and the corresponding second picture sizes;
[0031] Set the detection frames to be screened corresponding to the picture ratios with values not less than the preset ratio threshold among each of the picture ratios as target detection frames, and convert the relative position parameters corresponding to the target detection frames from the coordinate system corresponding to the historical multi-modal rendering map to the coordinate system corresponding to the sub-rendered image to obtain target position parameters;
[0032] Use the preset annotation rules to annotate each of the sub-rendered images to obtain corresponding rendering annotation results, and then use the target position parameters and based on the target detection frames to process each of the rendering annotation results to obtain a training set.
[0033] Optionally, the detection of the current multi-modal rendering map corresponding to the steel coil yard using the target yard inventory network model and the preset multi-view target result fusion center to obtain the yard inventory result corresponding to the steel coil yard includes:
[0034] Use the target yard inventory network model to perform target detection on each of the sub-rendered images to obtain corresponding sub-detection results, and use the preset multi-view target result fusion center and based on each of the sub-detection results to determine the projection of the top-down view in the current multi-modal rendering map to obtain a projection result; the target yard inventory network model is an end-to-end one-stage target detection network model based on a convolutional neural network;
[0035] Using the preset result to fuse the detection frames corresponding to the projection results with overlapping degrees meeting the preset fusion conditions in each of the projection results to obtain a fusion result, and determining the detection frame confidence corresponding to the fusion result based on the area and confidence corresponding to the fusion result;
[0036] Using the non-maximum suppression method to process each of the fusion results to obtain a corresponding suppression result, and deleting the detection frames corresponding to the results with overlapping degrees less than the preset overlapping degree threshold and the confidence less than the preset confidence threshold in the suppression result to obtain the detection results to be processed;
[0037] Using the preset multi-view target result fusion center and based on the preset imaging parameters and preset perspectives to process each of the detection results to be processed to obtain the detection results to be screened, and setting the detection results with confidence greater than the preset confidence threshold in each of the detection results to be screened as the target detection results;
[0038] Counting the number of targets in the target detection results to obtain the yard inventory result corresponding to the steel coil yard.
[0039] In a second aspect, the present application provides a steel coil yard inventory device based on three-dimensional reconstruction, including:
[0040] A three-dimensional reconstruction data determination module, configured to determine key frames based on the feature matching relationship between each video frame in the historical multi-view video data corresponding to the steel coil yard, and determine the three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames; the three-dimensional reconstruction data includes imaging information, sparse three-dimensional point cloud data, and dense three-dimensional reconstruction data;
[0041] A rendering diagram generation module, configured to determine a three-dimensional rectangular parallelepiped bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, determine the perspective imaging information corresponding to the three-dimensional rectangular parallelepiped bounding box based on a plurality of viewing angles, and then determine the historical multi-modal rendering diagram based on each of the perspective imaging information and the dense three-dimensional reconstruction data;
[0042] A yard inventory model determination module, configured to process the historical multi-modal rendering diagram using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and use the training set to train an initial yard inventory network model to obtain a target yard inventory network model;
[0043] A yard inventory result determination module, configured to use the target yard inventory network model and a preset multi-view target result fusion center to detect the current multi-modal rendering diagram corresponding to the steel coil yard to obtain the yard inventory result corresponding to the steel coil yard.
[0044] Optionally, the yard inventory model determination module includes:
[0045] A depth map normalization unit, configured to perform spatial enhancement processing and Gaussian filtering processing on each of the target local RGB images to obtain an image to be cut, and perform normalization processing, Gaussian noise addition processing, and Gaussian filtering processing on each of the target local depth maps to obtain a depth map to be cut;
[0046] A detection box determination unit, configured to cut the image to be cut modal rendering map according to a preset cutting rule to obtain a plurality of sub-rendering maps, determine a corresponding sliding window value based on the picture width of the historical multi-modal rendering map, and use the sliding window value to determine a detection box to be screened; the image to be cut modal rendering map includes the image to be cut and the depth map to be cut;
[0047] A picture ratio determination unit, configured to determine a first picture size corresponding to each sub-rendering map in each of the detection boxes to be screened and a second picture size corresponding to each historical multi-modal rendering map in each of the detection boxes to be screened, and determine a corresponding picture ratio based on each of the first picture sizes and the corresponding second picture sizes;
[0048] A position parameter generation unit, configured to set the detection box to be screened corresponding to the picture ratio with a value not less than a preset ratio threshold in each of the picture ratios as a target detection box, and convert the relative position parameter corresponding to the target detection box from the coordinate system corresponding to the historical multi-modal rendering map to the coordinate system corresponding to the sub-rendering map to obtain a target position parameter;
[0049] A training set generation unit, configured to label each of the sub-rendering maps using a preset annotation rule to obtain a corresponding rendering map annotation result, and then process each of the rendering map annotation results using the target position parameter and based on the target detection box to obtain a training set.
[0050] Optionally, the yard inventory result determination module includes:
[0051] A projection result generation unit, configured to perform target detection on each of the sub-rendering maps using the target yard inventory network model to obtain a corresponding sub-detection result, and use a preset multi-view target result fusion center and determine a projection of the top view in the current multi-modal rendering map based on each of the sub-detection results to obtain a projection result; the target yard inventory network model is an end-to-end one-stage target detection network model based on a convolutional neural network;
[0052] A confidence determination unit, configured to use a preset result fusion result to perform a fusion process on the detection frames corresponding to the projection results with an overlap degree meeting a preset fusion condition in each of the projection results, to obtain a fusion result, and determine a detection frame confidence corresponding to the fusion result based on the area and confidence corresponding to the fusion result;
[0053] A first detection result generation unit, configured to use a non-maximum suppression method to process each of the fusion results, to obtain a corresponding suppression result, and delete the detection frames corresponding to the results with an overlap degree less than a preset overlap degree threshold and a confidence less than a preset confidence threshold in the suppression result, to obtain a detection result to be processed;
[0054] A second detection result generation unit, configured to use a preset multi-view target result fusion center and based on preset imaging parameters and a preset view angle to process each of the detection results to be processed, to obtain a detection result to be screened, and set the detection results with a confidence greater than the preset confidence threshold in each of the detection results to be screened as target detection results;
[0055] A quantity statistics unit, configured to count the quantity of targets in the target detection results, to obtain a yard inventory result corresponding to the steel coil yard.
[0056] As can be seen from the above, before performing the inventory of the steel coil yard based on 3D reconstruction in this application, it is necessary to determine key frames based on the feature matching relationship between each video frame in the historical multi-view video data corresponding to the steel coil yard, and determine 3D reconstruction data corresponding to the steel coil yard based on each key frame; the 3D reconstruction data includes imaging information, sparse 3D point cloud data, and dense 3D reconstruction data; determine a 3D cuboid boundary box enclosing the steel coil yard based on the imaging information and the sparse 3D point cloud data, and determine the view imaging information corresponding to the 3D cuboid boundary box based on a plurality of observation angles, and then determine a historical multi-modal rendering map based on each view imaging information and the dense 3D reconstruction data; process the historical multi-modal rendering map using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method, to obtain a training set, and use the training set to train an initial yard inventory network model, to obtain a target yard inventory network model; use the target yard inventory network model and a preset multi-view target result fusion center to detect the current multi-modal rendering map corresponding to the steel coil yard, to obtain a yard inventory result corresponding to the steel coil yard.
[0057] It can be seen that, in this application, first, key frames need to be determined based on the feature matching relationships between video frames in the historical multi-perspective video data corresponding to the steel coil yard, and three-dimensional reconstruction data corresponding to the steel coil yard is determined based on the key frames; subsequently, a three-dimensional cuboid bounding box enclosing the steel coil yard is determined based on the imaging information and sparse three-dimensional point cloud data, and perspective imaging information corresponding to the three-dimensional cuboid bounding box is determined based on a number of viewing angles, and then historical multi-modal rendering maps are determined based on the perspective imaging information and dense three-dimensional reconstruction data; furthermore, the historical multi-modal rendering maps are processed using preset enhancement processing rules, preset annotation rules, and a preset sliding window method to obtain a training set, and the initial yard inventory network model is trained using the training set to obtain a target yard inventory network model; finally, the target yard inventory network model and a preset multi-perspective target result fusion center are used to detect the current multi-modal rendering map corresponding to the steel coil yard to obtain a yard inventory result corresponding to the steel coil yard. In this way, the efficiency of inventorying targets in the steel coil yard is improved, thereby enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0059] Figure 1 It is a flowchart of a method for inventorying a steel coil yard based on three-dimensional reconstruction disclosed in this application;
[0060] Figure 2 It is a schematic diagram of the execution logic for inventorying a steel coil yard based on three-dimensional reconstruction disclosed in this application;
[0061] Figure 3 It is a schematic diagram of the structure of a device for inventorying a steel coil yard based on three-dimensional reconstruction disclosed in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0063] Currently, in the process of industrial production and quality control, the inventory count of the steel coil yard is an important part of realizing automated management and efficient logistics scheduling. However, the methods for inventory counting of the targets in the steel coil yard mainly rely on manual operations, which brings problems such as increased operation and management costs and large inventory counting errors. Due to problems such as the limitation of the yard perspective, the relatively simple texture of the yard objects, and the similar shapes, the feature matching method is prone to failure in the yard environment, and it is difficult to use a single-angle image to collect complete yard information. In addition, the objects in the yard are arranged closely and distributed complexly, and the occlusion problem cannot be solved in the actually collected picture data; it is difficult to correspond the object detection results to the real world, and it is difficult to obtain additional information such as positioning. For this reason, the present application provides a method for inventory counting of the steel coil yard based on three-dimensional reconstruction, which can improve the efficiency of inventory counting of the targets in the steel coil yard.
[0064] See Figure 1 As shown, an embodiment of the present invention discloses a method for inventory counting of the steel coil yard based on three-dimensional reconstruction, including:
[0065] Step S11: Determine key frames based on the feature matching relationships between the video frames in the historical multi-perspective video data corresponding to the steel coil yard, and determine three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames; the three-dimensional reconstruction data includes imaging information, sparse three-dimensional point cloud data, and dense three-dimensional reconstruction data.
[0066] In this embodiment, industrial cameras are used to collect multi-perspective video data of the steel coil yard, and key frames are intercepted according to the feature matching relationships between the video frames in the video data. Then, the intercepted key frames are used for sparse reconstruction and dense three-dimensional reconstruction. That is, the sparse three-dimensional point cloud, imaging information in the sparse reconstruction result, and the dense three-dimensional reconstruction result together constitute the three-dimensional reconstruction data of the steel coil yard. In a specific implementation manner, as Figure 2 shown, the embodiment of the present application first shoots through a number of industrial cameras arranged on the gantry crane in the steel coil yard to collect multi-perspective video data, and obtains three-dimensional reconstruction data after operations such as key frame interception, sparse reconstruction, and dense reconstruction.
[0067] Specifically, determining key frames based on the feature matching relationships between the video frames in the historical multi-view video data corresponding to the steel coil yard, and determining the three-dimensional reconstruction data corresponding to the steel coil yard based on each key frame may include: determining corresponding multi-view picture data based on the historical multi-view video data corresponding to the steel coil yard, determining the feature matching relationships between the video frames in each multi-view picture data, and then determining key frames based on the feature matching relationships; using a preset sparse reconstruction process and determining sparse three-dimensional point cloud data corresponding to the steel coil yard and imaging information corresponding to each multi-view picture data based on each key frame; the imaging information includes imaging points, imaging fields of view, and imaging angles corresponding to the multi-view picture data; determining dense three-dimensional reconstruction data corresponding to the steel coil yard based on the multi-view picture data, the imaging information, and the sparse three-dimensional point cloud data.
[0068] Step S12: Determine a three-dimensional rectangular boundary box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, determine perspective imaging information corresponding to the three-dimensional rectangular boundary box based on a number of viewing angles, and then determine historical multi-modal rendering maps based on each piece of the perspective imaging information and the dense three-dimensional reconstruction data.
[0069] In this embodiment, after obtaining the three-dimensional reconstruction data corresponding to the steel coil yard, the embodiments of the present application need to select a number of angles, perform block rendering operations approximating parallel light imaging respectively, and match and splice the obtained block rendering results to obtain a high-precision multi-modal rendering map of the complete yard. Among them, the process of calculating multi-modal rendering maps at multiple viewing angles using the three-dimensional reconstruction data is to estimate the vertex coordinates of the three-dimensional rectangular boundary box enclosing the steel coil yard in the sparse three-dimensional point cloud using the sparse three-dimensional points in the three-dimensional reconstruction data and the imaging points in the imaging information, and calculate the new perspective imaging information required for complete rendering of the blocks of the three-dimensional rectangular boundary box enclosing the steel coil yard from multiple viewing angles. Specifically, determining a three-dimensional rectangular boundary box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, and determining perspective imaging information corresponding to the three-dimensional rectangular boundary box based on a number of viewing angles may include: processing the imaging information and the sparse three-dimensional point cloud data based on a preset imaging point determination rule to obtain a number of imaging points; using a preset three-dimensional rectangular boundary box generation rule and each imaging point, and determining the vertex coordinates corresponding to the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, and determining the three-dimensional rectangular boundary box corresponding to the steel coil yard based on each vertex coordinate; using a preset perspective imaging information generation rule and performing a generation operation on the three-dimensional rectangular boundary box based on a number of viewing angles to obtain corresponding perspective imaging information.
[0070] Further, after obtaining the new perspective imaging information and the dense three-dimensional reconstruction data, the embodiments of the present application need to render non-overlapping local RGB (Red Green Blue) images and local depth maps in blocks, and perform alignment and stitching operations on the local RGB images and the local depth maps according to the new perspective imaging information, so as to obtain multi-modal rendered images corresponding to multiple viewing angles respectively. Specifically, determining the historical multi-modal rendered images based on the imaging information of each perspective and the dense three-dimensional reconstruction data may include: determining the historical multi-modal rendered images of a plurality of initial local RGB images and a plurality of initial local depth maps by using a preset block rendering rule and based on the imaging information of each perspective and the dense three-dimensional reconstruction data; the initial local RGB images and the initial local depth maps do not overlap with each other; using a preset alignment and stitching rule to perform alignment and stitching operations on each initial local RGB image and each initial local depth map to obtain a historical multi-modal rendered image including a target local RGB image and a target local depth map; wherein, the viewing angles corresponding to each historical multi-modal rendered image are all different.
[0071] It is worth mentioning that during the rendering process, in order to reduce the edge errors caused by stitching, in a specific embodiment, for the RGB image, the embodiments of the present application need to adjust the projection matrix used for imaging, so that the imaging process is equivalent to using parallel light. For the depth map, since the imaging focus required for calculating the depth is needed, the embodiments of the present application need to use an imaging field smaller than a preset imaging field threshold for approximate parallel light imaging, so as to reduce the incoordination of the edges of the stitched image. In addition, when the viewing angle used has an angle with the vertical direction, the embodiments of the present application need to adjust the depth map according to the imaging point position and the imaging angle during stitching, so as to reduce the stepped noise in the stitching result. In addition, since the RGB image and the depth map have different corresponding precisions, the embodiments of the present application need to save them separately, and the RGB image and the depth map have a one-to-one correspondence in pixel positions.
[0072] Specifically, determining the historical multi-modal rendering map based on the imaging information of each perspective and the dense three-dimensional reconstruction data may include: using the initial target projection matrix and generating several initial local RGB images based on the imaging information of each perspective and the dense three-dimensional reconstruction data, then generating a target projection matrix based on the preset matrix generation rule and the imaging information of each perspective and the dense three-dimensional reconstruction data, and using the target projection matrix to process each initial local RGB image to obtain a target local RGB image; generating several initial local depth maps based on the preset imaging field of view and the imaging information of each perspective and the dense three-dimensional reconstruction data, and then processing the initial local depth maps based on the angle between the viewing angle corresponding to each imaging point and the vertical direction to obtain target local depth maps; wherein, the historical multi-modal rendering map includes the target local RGB image and the target local depth map, and the target local RGB image and the target local depth map are in a one-to-one correspondence relationship in terms of pixel positions.
[0073] Step S13, process the historical multi-modal rendering map by using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and use the training set to train an initial yard inventory network model to obtain a target yard inventory network model.
[0074] In this embodiment, after obtaining the historical multi-modal rendering map, the embodiments of the present application need to perform data enhancement operations on the RGB image and the depth map included in the historical multi-modal rendering map respectively. In a specific implementation manner, the embodiments of the present application need to use the sliding window method to process the multi-modal rendering map to obtain overlapping sub-block multi-modal rendering maps, where the sub-block multi-modal rendering maps include sub-block RGB images and sub-block rendering maps. Subsequently, perform HSV (Hue Saturation Value) space enhancement processing operations and Gaussian filtering processing operations on the sub-block RGB images, perform normalization processing operations, Gaussian noise addition processing operations, and Gaussian filtering processing operations on the sub-block depth maps. Subsequently, perform manual annotation on the multi-modal rendering map obtained after data enhancement to obtain a yard inventory training data set of a labeled steel coil yard.
[0075] Specifically, by using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to process historical multi-modal rendering images, a training set can be obtained, which may include: performing spatial enhancement processing and Gaussian filtering on each target local RGB image to obtain an image to be cut, and performing normalization processing, Gaussian noise addition processing, and Gaussian filtering on each target local depth image to obtain a depth image to be cut; performing cutting on the image to be cut of the multi-modal rendering image according to a preset cutting rule to obtain a number of segmented rendering images, determining a corresponding sliding window value based on the picture width of the historical multi-modal rendering image, and using the sliding window value to determine a detection box to be screened; the image to be cut of the multi-modal rendering image includes the image to be cut and the depth image to be cut; determining the first picture size corresponding to each segmented rendering image in each detection box to be screened and the second picture size corresponding to each historical multi-modal rendering image in each detection box to be screened, and determining a corresponding picture ratio based on each first picture size and the corresponding second picture size; setting the detection box to be screened corresponding to the picture ratio with a value not less than a preset ratio threshold in each picture ratio as a target detection box, and converting the relative position parameter corresponding to the target detection box from the coordinate system corresponding to the historical multi-modal rendering image to the coordinate system corresponding to the segmented rendering image to obtain a target position parameter; using a preset annotation rule to annotate each segmented rendering image to obtain a corresponding rendering image annotation result, and then using the target position parameter and based on the target detection box to process each rendering image annotation result to obtain a training set.
[0076] It is worth mentioning that during the process of cutting the image to be cut of the multi-modal rendering image, in the embodiment of the present application, it is necessary to determine a ratio according to the width of the picture to select sliding windows of different sizes based on the ratio. Subsequently, calculate the ratio of the area of the detection box in the picture to the area of the detection box in the complete picture, and retain the detection box with a ratio reaching the threshold, and adjust the relative position parameter to be used as the true value of the detection box of the cut image. In a specific implementation manner, since the local difference of the depth image is not significant compared to the global difference, after the picture is segmented, in the embodiment of the present application, it is necessary to perform a normalization operation on the depth image so that its value range is between 0 and 1. Subsequently, after performing an enhancement operation on the data, add the depth data to the HSV space of the RGB image to train the model using the obtained training set.
[0077] Step S14: Use the target yard inventory network model and a preset multi-view target result fusion center to detect the current multi-modal rendering image corresponding to the steel coil yard, and obtain a yard inventory result corresponding to the steel coil yard.
[0078] Further, in a specific embodiment, the yard inventory network model is an end-to-end one-stage object detection network based on a convolutional neural network. The yard inventory network model is used to perform object detection on multi-view segmented multi-modal rendering images, thereby obtaining corresponding segmented detection results. The preset multi-view object result fusion center is used to process the segmented detection results to obtain projections within the complete rendering image from the top-down view. Finally, weight distribution is performed on the confidence levels of the detection results of several detection frames and the area overlap degree of the detection frames, and the non-maximum suppression method is used to determine the object detection results corresponding to the steel coil yard based on the segmented detection results, and the object results are statistically analyzed to obtain the inventory results of the steel coil yard.
[0079] Specifically, using the target yard inventory network model and the preset multi-view object result fusion center to detect the current multi-modal rendering image corresponding to the steel coil yard to obtain the yard inventory results corresponding to the steel coil yard may include: using the target yard inventory network model to perform object detection on each segmented rendering image to obtain corresponding segmented detection results, and using the preset multi-view object result fusion center and based on the segmented detection results to determine the projection within the current multi-modal rendering image from the top-down view to obtain a projection result; the target yard inventory network model is an end-to-end one-stage object detection network model based on a convolutional neural network; using the preset result fusion result to fuse the detection frames corresponding to the projection results with overlapping degrees satisfying the preset fusion conditions among the projection results to obtain a fusion result, and determining the detection frame confidence corresponding to the fusion result based on the area and confidence level corresponding to the fusion result; using the non-maximum suppression method to process each fusion result to obtain a corresponding suppression result, and deleting the detection frames corresponding to the results with overlapping degrees less than the preset overlap degree threshold and confidence levels less than the preset confidence level threshold in the suppression result to obtain the detection results to be processed; using the preset multi-view object result fusion center and based on the preset imaging parameters and preset viewing angles to process each detection result to be processed to obtain the detection results to be screened, and setting the detection results with confidence levels greater than the preset confidence level threshold among the detection results to be screened as the object detection results; statistically analyzing the number of objects in the object detection results to obtain the yard inventory results corresponding to the steel coil yard.
[0080] It is worth mentioning that in the embodiment of the present application, the target detection model is only used to detect the local rendering map in the training stage. And in the inference stage, first, the detection frames with an overlap degree less than the preset overlap degree threshold in the block detection results of the rendering pictures from the same perspective are fused, and the confidence of the fused detection frame is calculated according to their respective areas and confidences. Subsequently, using the non-maximum suppression method, the detection frames with an overlap degree less than the preset overlap degree threshold and a confidence less than the preset confidence threshold are deleted. Then, according to the imaging parameters, the detection results of the rendering pictures from different perspectives are projected onto the same perspective. In a specific implementation, the vertical top-down perspective is selected as the target perspective for projection, and the detection results from different angles are weighted and fused to obtain the fused target detection result, and the number of targets in the steel coil yard is counted as the inventory result of the steel coil yard.
[0081] It can be seen that in the embodiment of the present application, first, key frames need to be determined based on the feature matching relationship between each video frame in the historical multi-perspective video data corresponding to the steel coil yard, and the three-dimensional reconstruction data corresponding to the steel coil yard is determined based on each key frame; subsequently, a three-dimensional cuboid bounding box enclosing the steel coil yard is determined based on the imaging information and the sparse three-dimensional point cloud data, and the perspective imaging information corresponding to the three-dimensional cuboid bounding box is determined based on several viewing angles. Then, the historical multi-modal rendering map is determined based on each perspective imaging information and the dense three-dimensional reconstruction data; furthermore, the historical multi-modal rendering map is processed using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and the initial yard inventory network model is trained using the training set to obtain the target yard inventory network model; finally, the current multi-modal rendering map corresponding to the steel coil yard is detected using the target yard inventory network model and a preset multi-perspective target result fusion center to obtain the yard inventory result corresponding to the steel coil yard. In this way, the efficiency of inventorying the targets in the steel coil yard is improved.
[0082] Correspondingly, as shown in Figure 3 the present application also provides a steel coil yard inventory device based on three-dimensional reconstruction, including:
[0083] A three-dimensional reconstruction data determination module 11, configured to determine key frames based on the feature matching relationship between each video frame in the historical multi-perspective video data corresponding to the steel coil yard, and determine the three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames; the three-dimensional reconstruction data includes imaging information, sparse three-dimensional point cloud data, and dense three-dimensional reconstruction data;
[0084] A rendering image generation module 12, configured to determine a three-dimensional cuboid bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, determine perspective imaging information corresponding to the three-dimensional cuboid bounding box based on a plurality of viewing angles, and then determine historical multi-modal rendering images based on each of the perspective imaging information and the dense three-dimensional reconstruction data;
[0085] A yard inventory model determination module 13, configured to process the historical multi-modal rendering images by using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and use the training set to train an initial yard inventory network model to obtain a target yard inventory network model;
[0086] A yard inventory result determination module 14, configured to use the target yard inventory network model and a preset multi-perspective target result fusion center to detect the current multi-modal rendering image corresponding to the steel coil yard, and obtain a yard inventory result corresponding to the steel coil yard.
[0087] As can be seen from the above, before performing the inventory of the steel coil yard based on three-dimensional reconstruction in the embodiment of the present application, it is necessary to determine key frames based on the feature matching relationship between each video frame in the historical multi-perspective video data corresponding to the steel coil yard, and determine the three-dimensional reconstruction data corresponding to the steel coil yard based on each key frame; the three-dimensional reconstruction data includes imaging information, sparse three-dimensional point cloud data, and dense three-dimensional reconstruction data; determine a three-dimensional cuboid bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, determine perspective imaging information corresponding to the three-dimensional cuboid bounding box based on a plurality of viewing angles, and then determine historical multi-modal rendering images based on each of the perspective imaging information and the dense three-dimensional reconstruction data; process the historical multi-modal rendering images by using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and use the training set to train an initial yard inventory network model to obtain a target yard inventory network model; use the target yard inventory network model and a preset multi-perspective target result fusion center to detect the current multi-modal rendering image corresponding to the steel coil yard, and obtain a yard inventory result corresponding to the steel coil yard. In this way, the efficiency of inventorying the targets in the steel coil yard is improved.
[0088] In some specific embodiments, the three-dimensional reconstruction data determination module 11 may specifically include:
[0089] A key frame determination unit, configured to determine corresponding multi-perspective picture data based on the historical multi-perspective video data corresponding to the steel coil yard, determine the feature matching relationship between each video frame in each of the multi-perspective picture data, and then determine key frames based on the feature matching relationship;
[0090] An imaging information determining unit, configured to use a preset sparse reconstruction process and determine sparse three-dimensional point cloud data corresponding to the steel coil yard and imaging information corresponding to each of the multi-view picture data based on each of the key frames; the imaging information includes imaging points, an imaging field of view, and an imaging angle corresponding to the multi-view picture data;
[0091] A three-dimensional reconstruction data determining subunit, configured to determine dense three-dimensional reconstruction data corresponding to the steel coil yard based on the multi-view picture data, the imaging information, and the sparse three-dimensional point cloud data.
[0092] In some specific embodiments, the rendering graph generation module 12 may specifically include:
[0093] An imaging point determining unit, configured to process the imaging information and the sparse three-dimensional point cloud data based on a preset imaging point determination rule to obtain a plurality of imaging points;
[0094] A three-dimensional cuboid bounding box determining unit, configured to use a preset three-dimensional cuboid bounding box generation rule and each of the imaging points, and determine vertex coordinates corresponding to the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, and determine a three-dimensional cuboid bounding box corresponding to the steel coil yard based on each of the vertex coordinates;
[0095] A perspective imaging information determining unit, configured to use a preset perspective imaging information generation rule and perform a generation operation on the three-dimensional cuboid bounding box based on a plurality of observation angles to obtain corresponding perspective imaging information.
[0096] In some specific embodiments, the rendering graph generation module 12 may specifically include:
[0097] A first rendering graph generation unit, configured to use a preset block rendering rule and determine a historical multi-modal rendering graph of a plurality of initial local RGB images and a plurality of initial local depth maps based on each of the perspective imaging information and the dense three-dimensional reconstruction data; the initial local RGB images and the initial local depth maps do not overlap with each other;
[0098] A second rendering graph generation unit, configured to perform an alignment and splicing operation on each of the initial local RGB images and each of the initial local depth maps by using a preset alignment and splicing rule to obtain a historical multi-modal rendering graph including a target local RGB image and a target local depth map; wherein, the observation angles corresponding to each of the historical multi-modal rendering graphs are all different.
[0099] In some specific embodiments, the rendering graph generation module 12 may specifically include:
[0100] A local RGB image determination unit, configured to use an initial target projection matrix and generate a plurality of initial local RGB images based on each of the perspective imaging information and the dense three-dimensional reconstruction data, then generate a target projection matrix based on each of the perspective imaging information and the dense three-dimensional reconstruction data using a preset matrix generation rule, and process each of the initial local RGB images using the target projection matrix to obtain target local RGB images;
[0101] A local depth map determination unit, configured to generate a plurality of initial local depth maps based on a preset imaging field of view and based on each of the perspective imaging information and the dense three-dimensional reconstruction data, and then process the initial local depth maps based on the angle between the observation angle corresponding to each imaging point and the vertical direction to obtain target local depth maps; wherein, the historical multi-modal rendering maps include the target local RGB images and the target local depth maps, and the target local RGB images and the target local depth maps are in a one-to-one correspondence relationship in terms of pixel positions.
[0102] In some specific embodiments, the yard inventory model determination module 13 may specifically include:
[0103] A depth map normalization unit, configured to perform spatial enhancement processing and Gaussian filtering processing on each of the target local RGB images to obtain images to be cut, and perform normalization processing, Gaussian noise addition processing and Gaussian filtering processing on each of the target local depth maps to obtain depth maps to be cut;
[0104] A detection box determination unit, configured to cut the images to be cut of the multi-modal rendering maps according to a preset cutting rule to obtain a plurality of segmented rendering maps, determine a corresponding sliding window value based on the picture width of the historical multi-modal rendering maps, and use the sliding window value to determine detection boxes to be screened; the images to be cut of the multi-modal rendering maps include the images to be cut and the depth maps to be cut;
[0105] A picture ratio determination unit, configured to determine a first picture size corresponding to each of the segmented rendering maps in each of the detection boxes to be screened and a second picture size corresponding to each of the historical multi-modal rendering maps in each of the detection boxes to be screened, and determine a corresponding picture ratio based on each of the first picture sizes and the corresponding second picture sizes;
[0106] A position parameter generation unit, configured to set the detection boxes to be screened corresponding to the picture ratios with values not less than a preset ratio threshold in each of the picture ratios as target detection boxes, and convert the relative position parameters corresponding to the target detection boxes from the coordinate system corresponding to the historical multi-modal rendering maps to the coordinate system corresponding to the segmented rendering maps to obtain target position parameters;
[0107] A training set generation unit is configured to label each of the segmented rendering images using a preset annotation rule to obtain corresponding rendering image annotation results, and then process each of the rendering image annotation results using the target position parameter and based on the target detection box to obtain a training set.
[0108] In some specific embodiments, the yard inventory result determination module 14 may specifically include:
[0109] A projection result generation unit is configured to perform target detection on each of the segmented rendering images using the target yard inventory network model to obtain corresponding segmented detection results, and use a preset multi-view target result fusion center and based on each of the segmented detection results to determine the projection of the top-down view within the current multi-modal rendering image to obtain a projection result; the target yard inventory network model is an end-to-end one-stage target detection network model based on a convolutional neural network.
[0110] A confidence determination unit is configured to fuse the detection boxes corresponding to the projection results with an overlap degree satisfying a preset fusion condition in each of the projection results using a preset result fusion result to obtain a fusion result, and determine the detection box confidence corresponding to the fusion result based on the area and confidence corresponding to the fusion result.
[0111] A first detection result generation unit is configured to process each of the fusion results using a non-maximum suppression method to obtain corresponding suppression results, and delete the detection boxes corresponding to the results with an overlap degree less than a preset overlap degree threshold and a confidence less than a preset confidence threshold in the suppression results to obtain detection results to be processed.
[0112] A second detection result generation unit is configured to process each of the detection results to be processed using a preset multi-view target result fusion center and based on preset imaging parameters and a preset view to obtain detection results to be screened, and set the detection results with a confidence greater than the preset confidence threshold in each of the detection results to be screened as target detection results.
[0113] A quantity statistics unit is configured to count the number of targets in the target detection results to obtain a yard inventory result corresponding to the steel coil yard.
[0114] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts between each embodiment can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0115] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0116] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0117] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0118] The technical solutions provided in this application have been introduced in detail above. Specific examples have been used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for inventory taking of steel coil storage yards based on three-dimensional reconstruction, characterized in that Including: Determining key frames based on the feature matching relationships between video frames in the historical multi-view video data corresponding to the steel coil yard, and determining three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames; the three-dimensional reconstruction data includes imaging information, sparse three-dimensional point cloud data, and dense three-dimensional reconstruction data; Determining a three-dimensional cuboid bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, determining perspective imaging information corresponding to the three-dimensional cuboid bounding box based on a number of viewing angles, and then determining a historical multi-modal rendering map based on each of the perspective imaging information and the dense three-dimensional reconstruction data; Processing the historical multi-modal rendering map using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and training an initial yard inventory network model using the training set to obtain a target yard inventory network model; Detecting a current multi-modal rendering map corresponding to the steel coil yard using the target yard inventory network model and a preset multi-view target result fusion center to obtain a yard inventory result corresponding to the steel coil yard.
2. The method for inventorying a steel coil yard based on 3D reconstruction according to claim 1, wherein The determining key frames based on the feature matching relationships between video frames in the historical multi-view video data corresponding to the steel coil yard, and determining three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames, includes: Determining corresponding multi-view picture data based on the historical multi-view video data corresponding to the steel coil yard, determining the feature matching relationships between video frames in each of the multi-view picture data, and then determining key frames based on the feature matching relationships; Determining sparse three-dimensional point cloud data corresponding to the steel coil yard and imaging information corresponding to each of the multi-view picture data based on a preset sparse reconstruction process and each of the key frames; the imaging information includes imaging points, imaging fields of view, and imaging angles corresponding to the multi-view picture data; Determining dense three-dimensional reconstruction data corresponding to the steel coil yard based on the multi-view picture data, the imaging information, and the sparse three-dimensional point cloud data.
3. The method for inventory taking of steel coil storage yard based on 3D reconstruction according to claim 1, wherein The determining a three-dimensional cuboid bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, and determining perspective imaging information corresponding to the three-dimensional cuboid bounding box based on a number of viewing angles, includes: Processing the imaging information and the sparse three-dimensional point cloud data based on a preset imaging point determination rule to obtain a number of imaging points; Determining vertex coordinates corresponding to the steel coil yard using a preset three-dimensional cuboid bounding box generation rule and each of the imaging points, and based on the imaging information and the sparse three-dimensional point cloud data, and determining a three-dimensional cuboid bounding box corresponding to the steel coil yard based on each of the vertex coordinates; Generating corresponding perspective imaging information by performing a generation operation on the three-dimensional cuboid bounding box based on a preset perspective imaging information generation rule and a number of viewing angles.
4. The method for inventory taking of steel coil yards based on 3D reconstruction according to any one of claims 1 to 3, characterized in that, The determining a historical multi-modal rendering map based on each of the perspective imaging information and the dense three-dimensional reconstruction data, includes: Using the preset chunk rendering rules and based on the imaging information of each of the perspectives and the dense three-dimensional reconstruction data, determine a historical multi-modal rendering map of a number of initial local RGB images and a number of initial local depth maps; the initial local RGB images and the initial local depth maps do not overlap with each other; Using the preset alignment and stitching rules, perform alignment and stitching operations on each of the initial local RGB images and each of the initial local depth maps to obtain a historical multi-modal rendering map including a target local RGB image and a target local depth map; wherein, the viewing angles corresponding to each of the historical multi-modal rendering maps are all different.
5. The method for inventorying steel coil yards based on 3D reconstruction according to claim 4, wherein The determining of the historical multi-modal rendering map based on the imaging information of each of the perspectives and the dense three-dimensional reconstruction data includes: Using the initial target projection matrix and based on the imaging information of each of the perspectives and the dense three-dimensional reconstruction data, generate a number of initial local RGB images, then use the preset matrix generation rules and based on the imaging information of each of the perspectives and the dense three-dimensional reconstruction data, generate a target projection matrix, and use the target projection matrix to process each of the initial local RGB images to obtain a target local RGB image; Based on the preset imaging field of view and based on the imaging information of each of the perspectives and the dense three-dimensional reconstruction data, generate a number of initial local depth maps, and then process the initial local depth maps based on the angle between the viewing angle corresponding to each imaging point and the vertical direction to obtain a target local depth map; Wherein, the historical multi-modal rendering map includes the target local RGB image and the target local depth map, and the target local RGB image and the target local depth map are in a one-to-one correspondence relationship in terms of pixel positions.
6. The method for inventory taking of steel coil yard based on three-dimensional reconstruction according to claim 5, wherein, Using the preset enhancement processing rules, preset annotation rules and preset sliding window method to process the historical multi-modal rendering map to obtain a training set, including: Perform spatial enhancement processing and Gaussian filtering processing on each of the target local RGB images to obtain images to be sliced, and perform normalization processing, Gaussian noise addition processing and Gaussian filtering processing on each of the target local depth maps to obtain depth maps to be sliced; Slice the images to be sliced modal rendering map according to the preset slicing rules to obtain a number of sliced rendering maps, and determine the corresponding sliding window value based on the image width of the historical multi-modal rendering map, and use the sliding window value to determine the detection frames to be screened; the images to be sliced modal rendering map includes the images to be sliced and the depth maps to be sliced; Determine the first image size corresponding to each of the sliced rendering maps in each of the detection frames to be screened and the second image size corresponding to each of the historical multi-modal rendering maps in each of the detection frames to be screened, and determine the corresponding image ratio based on each of the first image sizes and the corresponding second image sizes; Set the detection frames to be screened corresponding to the image ratios whose values are not less than the preset ratio threshold in each of the image ratios as target detection frames, and convert the relative position parameters corresponding to the target detection frames from the coordinate system corresponding to the historical multi-modal rendering map to the coordinate system corresponding to the sliced rendering map to obtain target position parameters; Label each of the segmented rendering images using preset annotation rules to obtain corresponding rendering image annotation results, and then use the target position parameters and based on the target detection boxes to process each of the rendering image annotation results to obtain a training set.
7. The method for inventory taking of steel coil yard based on three-dimensional reconstruction according to claim 6, wherein The detecting the current multi-modal rendering image corresponding to the steel coil yard using the target yard inventory network model and a preset multi-view target result fusion center to obtain a yard inventory result corresponding to the steel coil yard includes: Using the target yard inventory network model to perform target detection on each of the segmented rendering images to obtain corresponding segmented detection results, and using a preset multi-view target result fusion center and based on each of the segmented detection results to determine the projection of the top-down view within the current multi-modal rendering image to obtain a projection result; the target yard inventory network model is an end-to-end one-stage target detection network model based on a convolutional neural network; Using a preset result fusion result to perform fusion processing on the detection boxes corresponding to the projection results with an overlap degree satisfying a preset fusion condition among each of the projection results to obtain a fusion result, and determining a detection box confidence corresponding to the fusion result based on the area and confidence corresponding to the fusion result; Using a non-maximum suppression method to process each of the fusion results to obtain corresponding suppression results, and deleting the detection boxes corresponding to the results with an overlap degree less than a preset overlap degree threshold and a confidence less than a preset confidence threshold among the suppression results to obtain detection results to be processed; Using a preset multi-view target result fusion center and based on preset imaging parameters and a preset view to process each of the detection results to be processed to obtain detection results to be screened, and setting the detection results with a confidence greater than the preset confidence threshold among each of the detection results to be screened as target detection results; Counting the number of targets in the target detection results to obtain a yard inventory result corresponding to the steel coil yard.
8. A steel coil yard inventory device based on three-dimensional reconstruction, characterized in that, Including: A three-dimensional reconstruction data determination module, configured to determine key frames based on the feature matching relationship between each video frame in the historical multi-view video data corresponding to the steel coil yard, and determine three-dimensional reconstruction data corresponding to the steel coil yard based on each of the key frames; the three-dimensional reconstruction data includes imaging information, sparse three-dimensional point cloud data, and dense three-dimensional reconstruction data; A rendering image generation module, configured to determine a three-dimensional cuboid bounding box enclosing the steel coil yard based on the imaging information and the sparse three-dimensional point cloud data, and determine view imaging information corresponding to the three-dimensional cuboid bounding box based on a plurality of viewing angles, and then determine historical multi-modal rendering images based on each of the view imaging information and the dense three-dimensional reconstruction data; A yard inventory model determination module, configured to process the historical multi-modal rendering images using a preset enhancement processing rule, a preset annotation rule, and a preset sliding window method to obtain a training set, and use the training set to train an initial yard inventory network model to obtain a target yard inventory network model; The yard inventory result determination module is used to detect the current multi-modal rendering diagram corresponding to the steel coil yard by using the target yard inventory network model and the preset multi-perspective target result fusion center, and obtain the yard inventory result corresponding to the steel coil yard.
9. The steel coil yard inventory device based on 3D reconstruction according to claim 8, wherein The yard inventory model determination module includes: The depth map normalization unit is used to perform spatial enhancement processing and Gaussian filtering processing on each of the target local RGB images to obtain an image to be cut, and perform normalization processing, Gaussian noise addition processing and Gaussian filtering processing on each of the target local depth maps to obtain a depth map to be cut; The detection box determination unit is used to cut the image to be cut modal rendering diagram according to a preset cutting rule to obtain a plurality of sub-rendered diagrams, determine a corresponding sliding window value based on the picture width of the historical multi-modal rendering diagram, and use the sliding window value to determine the detection box to be screened; the image to be cut modal rendering diagram includes the image to be cut and the depth map to be cut; The picture ratio determination unit is used to determine the first picture size corresponding to each sub-rendered diagram in each of the detection boxes to be screened and the second picture size corresponding to each of the historical multi-modal rendering diagrams in each of the detection boxes to be screened, and determine the corresponding picture ratio based on each of the first picture sizes and the corresponding second picture sizes; The position parameter generation unit is used to set the detection box to be screened corresponding to the picture ratio with a value not less than a preset ratio threshold in each of the picture ratios as the target detection box, and convert the relative position parameter corresponding to the target detection box from the coordinate system corresponding to the historical multi-modal rendering diagram to the coordinate system corresponding to the sub-rendered diagram to obtain the target position parameter; The training set generation unit is used to label each of the sub-rendered diagrams by using a preset annotation rule to obtain the corresponding rendered diagram annotation result, and then process each of the rendered diagram annotation results by using the target position parameter and based on the target detection box to obtain the training set.
10. The steel coil yard inventory device based on 3D reconstruction according to claim 8, characterized in that, The yard inventory result determination module includes: The projection result generation unit is used to perform target detection on each of the sub-rendered diagrams by using the target yard inventory network model to obtain the corresponding sub-detection result, and use the preset multi-perspective target result fusion center and based on each of the sub-detection results to determine the projection of the top view in the current multi-modal rendering diagram to obtain the projection result; the target yard inventory network model is an end-to-end one-stage target detection network model based on a convolutional neural network; The confidence determination unit is used to fuse the detection boxes corresponding to the projection results with an overlap degree meeting the preset fusion condition in each of the projection results by using the preset result fusion result to obtain the fusion result, and determine the detection box confidence corresponding to the fusion result based on the area and confidence corresponding to the fusion result; The first detection result generation unit is used to process each of the fusion results by using the non-maximum suppression method to obtain the corresponding suppression result, and delete the detection boxes corresponding to the results with an overlap degree less than the preset overlap degree threshold and a confidence less than the preset confidence threshold in the suppression result to obtain the detection result to be processed; The second detection result generation unit is configured to process each of the to-be-processed detection results by using a preset multi-view target result fusion center and based on preset imaging parameters and a preset viewing angle, obtain to-be-screened detection results, and set the detection results with a confidence level greater than the preset confidence threshold in each of the to-be-screened detection results as target detection results; The quantity statistics unit is configured to count the number of targets in the target detection results to obtain a yard inventory result corresponding to the steel coil yard.
Citation Information
Cited By
Steel coil checking method, device and equipment based on three-dimensional reconstruction and medium
CN121120577A
A steel reel point method, device, equipment and medium based on three-dimensional reconstruction
CN121120577B