A steel reel point method, device, equipment and medium based on three-dimensional reconstruction

By processing the top-view depth map based on the stockyard inventory network model and steel coil stacking rules, the problem of low inventory reliability caused by occlusion between steel coils was solved, and efficient and accurate steel coil quantity statistics were achieved.

CN121120577BActive Publication Date: 2026-08-04INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD
Filing Date
2025-09-08
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

During warehouse inventory checks, existing 3D reconstruction methods suffer from problems such as complex target feature representation, high noise levels in the site, and large amounts of information redundancy due to mutual obstruction between steel coils, resulting in low inventory reliability.

Method used

By utilizing a stockyard inventory network model and steel coil stacking rules to process the top-view depth map in the target multimodal data, including image acquisition, 3D reconstruction, point cloud data processing, and simulation data construction, 3D Gaussian reconstruction data and top-view depth map of the steel coil stockyard are generated, and steel coil inventory is performed using stacking rules.

Benefits of technology

It improves the reliability of steel coil inventory, avoids the problem of being unable to inventory lower-layer steel coils due to mutual obstruction in steel coil stacking, and achieves efficient and accurate steel coil quantity statistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120577B_ABST
    Figure CN121120577B_ABST
Patent Text Reader

Abstract

The application discloses a steel coil stack point method and device based on three-dimensional reconstruction, equipment and medium, relates to the technical field of computer vision, and comprises the following steps: collecting images of a steel coil yard from different angles to obtain target image data, and generating three-dimensional reconstruction data corresponding to the steel coil yard based on the target image data; obtaining target multi-modal data corresponding to the steel coil yard based on the three-dimensional reconstruction data; processing the overhead depth map in the target multi-modal data by using a yard point network model, so as to point the steel coils in the steel coil yard; wherein the yard point network model can point the steel coils according to the stacking rules of the steel coils. By processing the overhead depth map in the target multi-modal data by using the yard point network model and the steel coil stacking rules, the problem that the lower steel coils cannot be pointed due to the mutual shielding of the steel coils in the steel coil stack is avoided, and the reliability of the steel coil pointing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method, apparatus, equipment and medium for counting steel coils based on three-dimensional reconstruction. Background Technology

[0002] With the development of the global economy, the logistics industry has experienced rapid growth in recent years. As a crucial link in the bulk commodity logistics system, the accurate, efficient, and safe operation of storage yards is key to the smooth functioning of the logistics system. Inventory counting is one of the core operations in storage yards, with the primary goal of determining the stock of goods in the warehouse. Efficient and accurate inventory counting helps storage yards quickly complete the receipt and delivery of goods, freeing up more time for tasks such as cargo hoisting and transshipment, thereby increasing the throughput of the storage yard.

[0003] Currently, in the process of warehouse inventory, point cloud data is usually obtained by 3D reconstruction, and steel coils are inventoried on the point cloud data. However, due to the mutual occlusion between steel coils, this method has problems such as complex target feature representation, high noise in the site, and large amount of information redundancy, which reduces the reliability of the inventory process. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for steel coil inventory based on 3D reconstruction. This method utilizes a stockpile inventory network model and steel coil stacking rules to process the top-view depth map in the target multimodal data, avoiding the problem of being unable to inventory lower-layer coils due to mutual occlusion in the steel coil stack, thus improving the reliability of steel coil inventory. The specific solution is as follows:

[0005] Firstly, this application provides a method for steel coil inventory based on three-dimensional reconstruction, including:

[0006] Images of the steel coil stockpile are acquired from different perspectives to obtain target image data corresponding to the stockpile, and three-dimensional reconstruction data corresponding to the stockpile is generated based on the target image data. The three-dimensional reconstruction data includes calibration information corresponding to the target image data, sparse three-dimensional point cloud data corresponding to the stockpile, and three-dimensional Gaussian reconstruction data. The three-dimensional Gaussian reconstruction data is a three-dimensional scene representation generated based on the target image data and the sparse three-dimensional point cloud data.

[0007] Based on the three-dimensional reconstruction data, target multimodal data corresponding to the steel coil stockpile is obtained; wherein, the target multimodal data includes a top-view depth map of the steel coil stockpile;

[0008] The overhead depth map in the target multimodal data is processed using a stockyard inventory network model to inventory the steel coils in the stockyard based on the stacking rules of the steel coils; wherein the steel coils in the stockyard are stacked in a stacked structure and different steel coils block each other.

[0009] The process by which the stockpile inventory network model processes the target multimodal data includes: identifying targets in the top-view depth map of the target multimodal data to obtain the position markings of the steel coils in the stockpile from a top-view perspective, and determining the stacking rules of the steel coils based on the position markings and the top-view depth map; and taking inventory of the steel coils in the stockpile based on the stacking rules to obtain corresponding inventory results; wherein the inventory results include the quantity information of the steel coils.

[0010] The step of determining the stacking rules of steel coils based on the location markings and the top-view depth map includes: dividing the steel coils in the steel coil yard into different stacks according to the location markings corresponding to the steel coils in the steel coil yard and the spatial relationship between each steel coil; sampling the steel coils in the same stack in the top-view depth map to obtain the average depth of each target steel coil in the same stack, wherein the target steel coil is the steel coil visible in the top-view depth map; determining the average depth as a target clustering feature; obtaining a steel coil hierarchical division scheme based on the target clustering feature and different clustering evaluation indicators; determining the hierarchical distribution of each steel coil in the same stack based on the steel coil hierarchical division scheme; and determining the stacking rules based on each hierarchical distribution.

[0011] Optionally, generating the three-dimensional reconstruction data corresponding to the steel coil yard based on the target image data includes:

[0012] Using sparse reconstruction technology and the target image data, calibration information corresponding to the target image data and sparse 3D point cloud data are generated; wherein, the calibration information corresponds one-to-one with the target image data, and the calibration information includes the optical center position of the camera in 3D space and the line of sight of the camera;

[0013] The target image data, the calibration information, and the sparse three-dimensional point cloud data are used to generate the three-dimensional Gaussian reconstruction data corresponding to the steel coil stack.

[0014] Optionally, obtaining the target multimodal data corresponding to the steel coil stockpile based on the three-dimensional reconstruction data includes:

[0015] The initial point cloud model of the steel coil stockpile is obtained using the three-dimensional Gaussian reconstruction data in the three-dimensional reconstruction data;

[0016] The initial point cloud model is aligned using the optical center position in the calibration information to obtain the target point cloud model;

[0017] The grid size corresponding to the target point cloud model is determined using the elbow rule, and the target point cloud model is rasterized based on the grid size to obtain the top-view depth map.

[0018] Optionally, before processing the top-view depth map in the target multimodal data using the stockyard inventory network model, the method further includes:

[0019] A three-dimensional simulation space is constructed, and an initial virtual steel coil stack is randomly generated in the three-dimensional simulation space according to the stacking rules of the steel coils in the steel coil yard; wherein, the initial virtual steel coil stack exists in the three-dimensional simulation space in the form of a point cloud;

[0020] Noise data is added to the initial virtual steel coil stack to obtain the corresponding target virtual steel coil stack; wherein, the noise data includes tilt noise and filling noise, the tilt noise is the noise generated by rotating the initial virtual steel coil along the X-axis, Y-axis and Z-axis in the three-dimensional simulation space by a target angle respectively, and the filling noise is the noise generated in the gaps of the initial virtual steel coil stack using a target noise generation algorithm;

[0021] An initial simulated top-down depth map is generated using the target virtual steel coil stack, and Burmester noise and highly obfuscated noise are added to the initial simulated top-down depth map to obtain the target simulated top-down depth map. The initial simulated top-down depth map exists in the form of an unsigned binary number, and the highly obfuscated noise is the noise generated by performing a linear operation on the depth value of each pixel in the initial simulated top-down depth map to map the original depth value to a predetermined numerical range.

[0022] The target simulation top-down depth map is automatically annotated to obtain the corresponding annotated simulation top-down depth map, and the annotated simulation top-down depth map is used to generate a target training set.

[0023] Optionally, before processing the top-view depth map in the target multimodal data using the stockyard inventory network model, the method further includes:

[0024] An initial network model is obtained, and the initial network model is trained offline using the target training set to obtain the yard inventory network model; wherein the yard inventory network model consists of several layers of convolutional network structure.

[0025] Secondly, this application provides a steel coil inventory device based on three-dimensional reconstruction, comprising:

[0026] An image acquisition module is used to acquire images of the steel coil stack from different perspectives to obtain target image data corresponding to the steel coil stack, and to generate three-dimensional reconstruction data corresponding to the steel coil stack based on the target image data; wherein, the three-dimensional reconstruction data includes calibration information corresponding to the target image data, sparse three-dimensional point cloud data corresponding to the steel coil stack, and three-dimensional Gaussian reconstruction data, and the three-dimensional Gaussian reconstruction data is a three-dimensional scene representation generated based on the target image data and the sparse three-dimensional point cloud data;

[0027] The data acquisition module is used to acquire target multimodal data corresponding to the steel coil stockpile based on the three-dimensional reconstruction data; wherein, the target multimodal data includes a top-view depth map of the steel coil stockpile;

[0028] The steel coil inventory module is used to process the top-view depth map in the target multimodal data using a stockyard inventory network model, so as to inventory the steel coils in the steel coil stockyard based on the stacking rules of the steel coils; wherein the steel coils in the steel coil stockyard are stacked in a stacked structure, and different steel coils block each other.

[0029] The stockpile inventory network model includes: a stacking rule determination submodule, used to perform target identification on the top-view depth map in the target multimodal data to obtain the position markings of the steel coils in the steel coil stockpile from a top-view perspective, and to determine the stacking rules of the steel coils based on the position markings and the top-view depth map; and a steel coil inventory unit, used to perform inventory of the steel coils in the steel coil stockpile based on the stacking rules to obtain corresponding inventory results; wherein the inventory results include the quantity information of the steel coils.

[0030] The stacking rule determination submodule includes: a stacking division unit, used to divide the steel coils in the steel coil yard into different steel coil stacks according to the location markings corresponding to the steel coils in the steel coil yard and the spatial positional relationship between each steel coil; an average depth acquisition unit, used to sample the steel coils in the same steel coil stack in the top-view depth map to obtain the average depth of each target steel coil in the same steel coil stack, wherein the target steel coil is the steel coil visible in the top-view depth map; a hierarchical division scheme acquisition unit, used to determine the average depth as the target clustering feature, and obtain a steel coil hierarchical division scheme based on the target clustering feature and different clustering evaluation indicators; and a stacking rule determination unit, used to determine the hierarchical distribution of each steel coil in the same steel coil stack based on the steel coil hierarchical division scheme, and determine the stacking rule according to each hierarchical distribution.

[0031] Thirdly, this application provides an electronic device, comprising:

[0032] Memory, used to store computer programs;

[0033] A processor is used to execute the computer program to implement the aforementioned steel coil inventory method based on three-dimensional reconstruction.

[0034] Fourthly, this application provides a computer-readable storage medium for storing a computer program that, when executed by a processor, implements the aforementioned steel coil inventory method based on three-dimensional reconstruction.

[0035] This application first acquires images of the steel coil stacking yard from different perspectives to obtain target image data corresponding to the steel coil stacking yard, and then generates 3D reconstruction data corresponding to the steel coil stacking yard based on the target image data. The 3D reconstruction data includes calibration information corresponding to the target image data, sparse 3D point cloud data corresponding to the steel coil stacking yard, and 3D Gaussian reconstruction data. The 3D Gaussian reconstruction data is a 3D scene representation generated based on the target image data and the sparse 3D point cloud data. Then, target multimodal data corresponding to the steel coil stacking yard is obtained based on the 3D reconstruction data. The target multimodal data includes a top-view depth map of the steel coil stacking yard. Finally, a stacking yard inventory network model is used to process the top-view depth map in the target multimodal data to inventory the steel coils in the steel coil stacking yard based on the stacking rules of the steel coils. The steel coils in the steel coil stacking yard have a stacked structure, with different steel coils occluding each other. The process of the stacking yard inventory network model processing the target multimodal data includes: processing the top-view depth map in the target multimodal data... The process involves: identifying targets to obtain position markers of steel coils in the steel coil yard from a top-down perspective; determining stacking rules for the steel coils based on the position markers and the top-down depth map; and conducting an inventory of the steel coils in the steel coil yard based on the stacking rules to obtain corresponding inventory results. The inventory results include the quantity information of the steel coils. Determining the stacking rules based on the position markers and the top-down depth map includes: stacking the steel coils in the steel coil yard according to the position markers corresponding to the steel coils and the spatial relationships between the steel coils. The steel coils are divided into different stacks. Samples are taken from the same stack in the top-view depth map to obtain the average depth of each target coil in the same stack, wherein the target coils are those visible in the top-view depth map. The average depth is determined as a target clustering feature, and a coil hierarchy division scheme is obtained based on the target clustering feature and different clustering evaluation indicators. The hierarchy distribution of each coil in the same stack is determined based on the coil hierarchy division scheme, and the stacking rules are determined according to the hierarchy distribution. Therefore, this application uses a stockyard inventory network model to process the top-view depth map in the target multimodal data, avoiding direct identification of point cloud data. This avoids the problems of complex target feature representation, high noise in the site, and large information redundancy, thus improving the reliability of steel coil inventory. By using a stockyard inventory network model and the coil stacking rules to inventory the steel coils, the problem of being unable to inventory lower layers of coils due to mutual occlusion in the stack is avoided. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0037] Figure 1 This is a flowchart of a steel coil inventory method based on three-dimensional reconstruction disclosed in this application;

[0038] Figure 2 This is a schematic diagram of the process of a steel coil inventory method based on three-dimensional reconstruction disclosed in this application;

[0039] Figure 3 This is a schematic diagram of a steel coil inventory device based on three-dimensional reconstruction disclosed in this application;

[0040] Figure 4 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] Current methods for inventorying steel coils suffer from problems such as complex target feature representations, high noise levels in the field, and significant information redundancy, reducing the reliability of the inventory process. To address this, this application provides a steel coil inventory method based on 3D reconstruction. By utilizing a stockpile inventory network model and steel coil stacking rules to process the top-down depth map in the target multimodal data, this method avoids the problem of lower-layer steel coils being unable to be inventoried due to mutual occlusion within the steel coil stack, thus improving the reliability of the steel coil inventory.

[0043] See Figure 1 As shown, this embodiment of the invention discloses a method for inventorying steel coils based on three-dimensional reconstruction, including:

[0044] Step S11: Acquire images of the steel coil stockpile from different perspectives to obtain target image data corresponding to the steel coil stockpile, and generate three-dimensional reconstruction data corresponding to the steel coil stockpile based on the target image data; wherein, the three-dimensional reconstruction data includes calibration information corresponding to the target image data, sparse three-dimensional point cloud data corresponding to the steel coil stockpile, and three-dimensional Gaussian reconstruction data, and the three-dimensional Gaussian reconstruction data is a three-dimensional scene representation generated based on the target image data and the sparse three-dimensional point cloud data.

[0045] The steel coil inventory method in this embodiment mainly includes: collecting multi-view image data (i.e., target image data) of the steel coil yard; using the multi-view image data to calculate three-dimensional reconstruction data; using the three-dimensional reconstruction data to render multi-modal data of the steel coil yard, wherein the multi-modal data includes a top-down depth map and point cloud; constructing a large amount of simulation data offline; using the simulation dataset to train a convolutional neural network to obtain a yard inventory network model for steel coil yard inventory; collecting multi-modal data of the steel coil yard online; processing the multi-modal data through the yard inventory network model to obtain inventory results, wherein the inventory results include counting information of the steel coils.

[0046] like Figure 2 As shown, in this embodiment, image data is first obtained by taking pictures using several industrial anti-vibration cameras arranged on the steel coil yard gantry. Then, the obtained multi-view image data is transmitted to the cloud server using a fiber optic network, and sparse reconstruction and Gaussian reconstruction steps are performed to obtain three-dimensional reconstruction data.

[0047] This embodiment can use multiple industrial anti-vibration cameras mounted on a gantry crane to collect multi-view image data, and use the multi-view image data to perform sparse reconstruction and Gaussian 3D reconstruction.

[0048] The 3D reconstruction data in this embodiment includes calibration information, sparse 3D point cloud data, and 3D Gaussian reconstruction data. Correspondingly, the process of generating the 3D reconstruction data corresponding to the steel coil stack based on the target image data can specifically include: using sparse reconstruction technology and the target image data to generate calibration information and sparse 3D point cloud data corresponding to the target image data; wherein, the calibration information corresponds one-to-one with the target image data, and the calibration information includes the position of the optical center of the camera in 3D space and the direction of the camera's line of sight; and using the target image data, calibration information, and sparse 3D point cloud data to generate the 3D Gaussian reconstruction data corresponding to the steel coil stack.

[0049] The above process involves generating calibration information for multi-view images and sparse 3D point cloud data corresponding to the steel coil stockpile through a sparse reconstruction process based on multi-view image data; and then generating 3D Gaussian reconstruction data of the steel coil stockpile using the multi-view image data, calibration information, and sparse 3D point cloud data.

[0050] Step S12: Obtain target multimodal data corresponding to the steel coil stockpile based on the three-dimensional reconstruction data; wherein, the target multimodal data includes a top-view depth map of the steel coil stockpile.

[0051] In this embodiment, the process of obtaining target multimodal data corresponding to the steel coil stockpile based on 3D reconstruction data may specifically include: obtaining an initial point cloud model of the steel coil stockpile using 3D Gaussian reconstruction data from the 3D reconstruction data; aligning the initial point cloud model using the optical center position in the calibration information to obtain a target point cloud model; determining the grid size corresponding to the target point cloud model using the elbow rule, and rasterizing the target point cloud model based on the grid size to obtain a top-view depth map.

[0052] The above process involves using the 3D Gaussian reconstruction data from the 3D reconstruction data to calculate the point cloud model of the steel coil yard (i.e., the initial point cloud model); aligning the steel coil point cloud according to the point cloud model and the optical center position in the calibration information; and using the aligned point cloud (i.e., the target point cloud model) to export the top-view depth map, thus obtaining the top-view depth map of the steel coil yard.

[0053] Specifically, the 3D Gaussian reconstruction data is used to obtain the positions of object surfaces in the scene using surface estimation methods. Then, point clouds are generated using appropriate density to obtain a point cloud model. The point cloud model and the calibration information are used to orthogonalize the point cloud, and then the point cloud is projected using a rasterization method to obtain a top-view depth map.

[0054] It should be noted that when aligning the point cloud using the point cloud model and calibration information, it involves aligning the three coordinate axes of the point cloud coordinate system with the length, width, and height of the actual storage yard, as well as determining the positive direction of the point cloud's vertical line. This alignment helps in finding the orientation for viewing the entire site from above, which is crucial for deriving the top-view depth map. The algorithm attempts to perform rasterized projection of the point cloud at different granularities, calculates the fill rate of the depth map at different granularities, and uses this as a feature to determine the appropriate granularity parameters using the elbow rule, ultimately deriving the final top-view depth map.

[0055] Step S13: Process the top-view depth map in the target multimodal data using the stockyard inventory network model to inventory the steel coils in the steel coil stockyard based on the stacking rules of the steel coils; wherein the steel coils in the steel coil stockyard are stacked structures, and different steel coils block each other.

[0056] In this embodiment, before processing the top-view depth map in the target multimodal data using the stockpile inventory network model, the method further includes: constructing a three-dimensional simulation space; randomly generating an initial virtual steel coil stack in the three-dimensional simulation space according to the stacking rules of steel coils in the steel coil stockpile; wherein, the initial virtual steel coil stack exists in the form of a point cloud in the three-dimensional simulation space; adding noise data to the initial virtual steel coil stack to obtain the corresponding target virtual steel coil stack; wherein, the noise data includes tilt noise and filling noise, the tilt noise is the noise generated by rotating the initial virtual steel coil along the X-axis, Y-axis and Z-axis in the three-dimensional simulation space by the target angle, and the filling noise is the noise generated by using the target noise. The sound generation algorithm generates noise in the gaps of the initial virtual steel coil stack; an initial simulated top-down depth map is generated using the target virtual steel coil stack, and Burmester noise and highly obfuscated noise are added to the initial simulated top-down depth map to obtain the target simulated top-down depth map. The initial simulated top-down depth map exists in the form of an unsigned binary number, and the highly obfuscated noise is the noise generated by performing linear operations on the depth values ​​of each pixel in the initial simulated top-down depth map to map the original depth values ​​to a predetermined numerical range; the target simulated top-down depth map is automatically labeled to obtain the corresponding labeled simulated top-down depth map, and the labeled simulated top-down depth map is used to generate the target training set.

[0057] Specifically, based on the stacking rules of multi-layer steel coil yards, multi-layer steel coil stacks conforming to physical rules (i.e., initial virtual steel coil stacks) are randomly generated on a simulated area of ​​a certain size, and the results are represented as a colored simulated point cloud. To highlight data features, fill noise and tilt noise are added to the simulated point cloud to generate an enhanced simulated point cloud (i.e., the target virtual steel coil stack). The enhanced simulated point cloud is used to generate a simulated top-down depth map, and berthing noise and highly obfuscated noise are added around it to generate an enhanced simulated top-down depth map (i.e., the target simulated top-down depth map). The enhanced simulated top-down depth map is automatically labeled to obtain a labeled simulated steel coil yard inventory training dataset (i.e., the target training set).

[0058] It should be noted that the aforementioned filling noise refers to the generation of noise structures in the gaps of the simulated steel coil stacks using a certain algorithm. The height of the noise structures is lower than that of the steel coil stacks on both sides, which has the effect of blurring the boundaries on the top-view depth map. The high-level confusion noise refers to the random discarding of some values ​​at the first and last positions of the depth map when generating the top-view depth map from the simulated point cloud, and mapping the depth information of the point cloud into the remaining usable range, which prevents steel coils with fixed layer heights from only appearing at fixed depths. The tilt noise refers to the rotation of the simulated point cloud (i.e., the virtual steel coil stack) along the x, y, and z coordinate axes by a small angle after it is generated, so that the simulated steel coils are not always parallel to the edges on the simulated coordinate system and the simulated top-view depth map, thus increasing data diversity.

[0059] Furthermore, in the top-down depth map, each pixel value represents the vertical distance from the top of the steel coil to the camera at that location. The depth map is typically stored using 8-bit unsigned integers (0-255), representing the height level of the pixel when the height of the corresponding scene is divided into 256 equal parts. The stored value has a scaling factor k with the actual height, which varies with the height of the stockpile.

[0060] The following example illustrates the process of adding obfuscated noise:

[0061] Step 1: Randomly discard some values ​​from the beginning and end of the depth map:

[0062] Randomly select the retention thresholds for the high and low bits, for example:

[0063] preserved_high=249;

[0064] preserved_low=10;

[0065] The original 256-bit depth value is linearly mapped into the range [preserved_low,preserved_high].

[0066] The specific algorithm is as follows:

[0067] confused_d=ROUND((origin_d / 255)*(preserved_high-preserved_low)+preserved_low);

[0068] The ROUND function is a rounding function used to map a calculated decimal number to a depth value of integer type.

[0069] Technical effect: It is equivalent to performing a linear transformation on the depth values, moving the depth map to a smaller range while preserving the original order relationship and corresponding proportion of the depth information.

[0070] Physical meaning: This is equivalent to linearly compressing the actual depth range (e.g., 0-20 meters) into a small numerical range (0-255), and then breaking the fixed relationship between layer height and depth through another linear mapping. For example, before the treatment, the top depth of the first layer of steel coils in a double-layer steel coil yard is often 128, and the top depth of the second layer is 255. After the treatment, these depth values ​​will change randomly.

[0071] Technical significance: By remapping depth numerical values, a key contradiction in simulation training is resolved: while retaining the relative depth relationship (such as flatness) between steel coils in the same layer, the model avoids memorizing the fixed correspondence between "layer height and depth", thereby preventing model overfitting.

[0072] Understandably, the generation of a simulation site (i.e., a 3D simulation space) requires certain hyperparameter specifications, such as the size of the site, the size of the steel coils, and the maximum number of stacked layers. The simulation data generation method can utilize these hyperparameters to randomly generate a static point cloud of a steel coil stacking area that conforms to physical rules. Random factors include the location of the steel coil stacks, the specific shape of each stack, and the total number of steel coils.

[0073] It should be noted that since the generation of the simulated site is done by an algorithm, the position of the simulated steel coil in three-dimensional space can be obtained. After projection, the position of the steel coil visible on the depth map from a top view can also be obtained.

[0074] In this embodiment, before processing the top-down depth map in the target multimodal data using the yard inventory network model, the method further includes: obtaining an initial network model and training the initial network model offline using the target training set to obtain the yard inventory network model; wherein, the yard inventory network model consists of several layers of convolutional network structure; that is, the above-mentioned yard inventory network model is a one-stage target detection neural network composed of multiple layers of convolutional network structure, and the above-mentioned network model can be additionally trained offline using the yard inventory training dataset.

[0075] The above process specifically involves using simulated data for data augmentation followed by training a convolutional neural network to enhance the model's ability to identify steel coil targets on the top-down depth map. Specifically, this includes adding padding noise, highly obfuscated noise, and tilt noise to the simulated point cloud to generate an enhanced simulated point cloud. This enhanced simulated point cloud is then used to generate a simulated top-down depth map, with berm noise added around it to create another enhanced simulated top-down depth map. Finally, the enhanced simulated top-down depth map and its annotations are used to further train the pre-trained model, adjusting the weights to obtain the steel coil detection model (i.e., the stockpile inventory network model).

[0076] In this embodiment, the process of the stockyard inventory network model processing the target multimodal data may specifically include: identifying targets in the top-view depth map of the target multimodal data to obtain the position markings of the steel coils in the stockyard from a top-view perspective, and determining the stacking rules of the steel coils based on the position markings and the top-view depth map; and taking inventory of the steel coils in the stockyard based on the stacking rules to obtain the corresponding inventory results; wherein, the inventory results include the quantity information of the steel coils.

[0077] The above process involves using a stockpile inventory network model to identify targets on a top-down depth map to obtain the location markings of steel coils from a top-down perspective. The stacking morphology of the steel coil stockpile is then analyzed using the location markings and depth map. Based on the analysis of the stacking morphology, the inventory count of multi-layer steel coil stockpiles is completed.

[0078] In this embodiment, a simulated top-down depth map is obtained by using a simulation data generation method. Gaussian noise and noise filling techniques are added to highlight the projected features of the steel coils, enabling the trained neural network to accurately identify the steel coils under the real depth map. Using a simulation data generation method reduces the difficulty of data collection and computational overhead. Because it does not rely on real labeled data, the recognition algorithm still exhibits robustness when facing steel coil storage yards with different characteristics. This solves the problem of difficult inventory management in multi-layered steel coil storage yards.

[0079] The process of determining the stacking rules of steel coils based on location markings and top-view depth maps can specifically include: dividing the steel coils in the steel coil yard into different stacks according to the location markings of the steel coils and the spatial relationship between each steel coil; sampling the steel coils in the same stack in the top-view depth map to obtain the average depth of each target steel coil in the same stack, wherein the target steel coil is the steel coil visible in the top-view depth map; determining the average depth as the target clustering feature, obtaining a steel coil hierarchical division scheme based on the target clustering feature and different clustering evaluation indicators; determining the hierarchical distribution of each steel coil in the same stack based on the steel coil hierarchical division scheme, and determining the stacking rules according to the hierarchical distribution.

[0080] The above process involves online acquisition of image data from the steel coil yard, calculation of 3D reconstruction results to obtain a top-down depth map, and then target detection using the aforementioned yard inventory network model. The height of the steel coils is then analyzed to determine their layer, and combined with the location and layer of visible steel coils, the stacking structure is inferred, ultimately completing the inventory of multiple layers of steel coils.

[0081] It should be noted that when performing stacking morphology analysis, the layer order of visible steel coils within the same stack is required. This embodiment can classify the steel coils based on their stacks by marking their positions from a top-down perspective, and then perform cluster analysis using the average height of the steel coils as a feature to obtain the stacking layer of each steel coil. Finally, this information is combined to complete the inventory.

[0082] Therefore, this application utilizes a stockyard inventory network model to process the top-view depth map in the target multimodal data, avoiding direct identification of point cloud data. This avoids the problems of complex target feature representation, high noise in the site, and large information redundancy, thus improving the reliability of steel coil inventory. By using the stockyard inventory network model and the stacking rules of steel coils to inventory the steel coils, the problem of being unable to inventory the lower layer of steel coils due to mutual occlusion in the steel coil stack is avoided.

[0083] See Figure 3 As shown, this embodiment of the invention discloses a steel coil inventory device based on three-dimensional reconstruction, comprising:

[0084] The image acquisition module 11 is used to acquire images of the steel coil stack from different perspectives to obtain target image data corresponding to the steel coil stack, and generate three-dimensional reconstruction data corresponding to the steel coil stack based on the target image data; wherein, the three-dimensional reconstruction data includes calibration information corresponding to the target image data, sparse three-dimensional point cloud data corresponding to the steel coil stack, and three-dimensional Gaussian reconstruction data, and the three-dimensional Gaussian reconstruction data is a three-dimensional scene representation generated based on the target image data and the sparse three-dimensional point cloud data;

[0085] The data acquisition module 12 is used to acquire target multimodal data corresponding to the steel coil stockpile based on the three-dimensional reconstruction data; wherein, the target multimodal data includes a top-view depth map of the steel coil stockpile;

[0086] The steel coil inventory module 13 is used to process the top-view depth map in the target multimodal data using a stockyard inventory network model, so as to inventory the steel coils in the steel coil stockyard based on the stacking rules of the steel coils; wherein the steel coils in the steel coil stockyard are stacked structures, and different steel coils block each other.

[0087] The stockpile inventory network model includes: a stacking rule determination submodule, used to perform target identification on the top-view depth map in the target multimodal data to obtain the position markings of the steel coils in the steel coil stockpile from a top-view perspective, and to determine the stacking rules of the steel coils based on the position markings and the top-view depth map; and a steel coil inventory unit, used to perform inventory of the steel coils in the steel coil stockpile based on the stacking rules to obtain corresponding inventory results; wherein the inventory results include the quantity information of the steel coils.

[0088] The stacking rule determination submodule includes: a stacking division unit, used to divide the steel coils in the steel coil yard into different steel coil stacks according to the location markings corresponding to the steel coils in the steel coil yard and the spatial positional relationship between each steel coil; an average depth acquisition unit, used to sample the steel coils in the same steel coil stack in the top-view depth map to obtain the average depth of each target steel coil in the same steel coil stack, wherein the target steel coil is the steel coil visible in the top-view depth map; a hierarchical division scheme acquisition unit, used to determine the average depth as the target clustering feature, and obtain a steel coil hierarchical division scheme based on the target clustering feature and different clustering evaluation indicators; and a stacking rule determination unit, used to determine the hierarchical distribution of each steel coil in the same steel coil stack based on the steel coil hierarchical division scheme, and determine the stacking rule according to each hierarchical distribution.

[0089] In some specific embodiments, the image acquisition module 11 may specifically include:

[0090] A calibration information generation unit is used to generate calibration information corresponding to the target image data and sparse three-dimensional point cloud data using sparse reconstruction technology and the target image data; wherein, the calibration information corresponds one-to-one with the target image data, and the calibration information includes the optical center position of the camera in three-dimensional space and the line of sight of the camera;

[0091] The Gaussian reconstruction data generation unit is used to generate the three-dimensional Gaussian reconstruction data corresponding to the steel coil stack using the target image data, the calibration information, and the sparse three-dimensional point cloud data.

[0092] In some specific embodiments, the data acquisition module 12 may specifically include:

[0093] The point cloud model acquisition unit is used to acquire the initial point cloud model of the steel coil stockpile using the three-dimensional Gaussian reconstruction data in the three-dimensional reconstruction data.

[0094] A point cloud model alignment unit is used to align the initial point cloud model using the optical center position in the calibration information to obtain a target point cloud model.

[0095] The top-view depth map acquisition unit is used to determine the grid size corresponding to the target point cloud model using the elbow rule, and to rasterize the target point cloud model based on the grid size to obtain the top-view depth map.

[0096] In some specific embodiments, the steel coil inventory module 13 further includes:

[0097] A virtual steel coil stacking generation unit is used to construct a three-dimensional simulation space. According to the stacking rules of steel coils in the steel coil yard, an initial virtual steel coil stack is randomly generated in the three-dimensional simulation space; wherein, the initial virtual steel coil stack exists in the form of a point cloud in the three-dimensional simulation space.

[0098] A noise addition unit is used to add noise data to the initial virtual steel coil stack to obtain a corresponding target virtual steel coil stack; wherein, the noise data includes tilt noise and filling noise, the tilt noise is the noise generated by rotating the initial virtual steel coil along the X-axis, Y-axis and Z-axis in the three-dimensional simulation space by a target angle, and the filling noise is the noise generated in the gaps of the initial virtual steel coil stack using a target noise generation algorithm;

[0099] The simulated top-down depth map generation unit is used to generate an initial simulated top-down depth map using the target virtual steel coil stack, and to add Burmester noise and highly obfuscated noise to the initial simulated top-down depth map to obtain a target simulated top-down depth map. The initial simulated top-down depth map exists in the form of an unsigned binary number, and the highly obfuscated noise is noise generated by performing linear operations on the depth values ​​of each pixel in the initial simulated top-down depth map to map the original depth values ​​to a predetermined numerical range.

[0100] The training set generation unit is used to automatically annotate the target simulation top-down depth map to obtain the corresponding annotated simulation top-down depth map, and use the annotated simulation top-down depth map to generate the target training set.

[0101] In some specific embodiments, the steel coil inventory module 13 further includes:

[0102] The model training unit is used to obtain an initial network model and train the initial network model offline using the target training set to obtain the yard inventory network model; wherein the yard inventory network model consists of several layers of convolutional network structure.

[0103] Furthermore, embodiments of this application also disclose an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0104] Figure 4 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the steel coil inventory method based on three-dimensional reconstruction disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0105] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0106] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0107] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the three-dimensional reconstruction-based steel coil inventory method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0108] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method for inventorying steel coils based on three-dimensional reconstruction. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0109] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0110] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0111] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0112] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0113] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for inventorying steel coils based on three-dimensional reconstruction, characterized in that, include: Images of the steel coil stockpile are acquired from different perspectives to obtain target image data corresponding to the stockpile, and three-dimensional reconstruction data corresponding to the stockpile is generated based on the target image data. The three-dimensional reconstruction data includes calibration information corresponding to the target image data, sparse three-dimensional point cloud data corresponding to the stockpile, and three-dimensional Gaussian reconstruction data. The three-dimensional Gaussian reconstruction data is a three-dimensional scene representation generated based on the target image data and the sparse three-dimensional point cloud data. Based on the three-dimensional reconstruction data, target multimodal data corresponding to the steel coil stockpile is obtained; wherein, the target multimodal data includes a top-view depth map of the steel coil stockpile; The overhead depth map in the target multimodal data is processed using a stockyard inventory network model to inventory the steel coils in the stockyard based on the stacking rules of the steel coils; wherein the steel coils in the stockyard are stacked in a stacked structure and different steel coils block each other. The process by which the stockpile inventory network model processes the target multimodal data includes: identifying targets in the top-view depth map of the target multimodal data to obtain the position markings of the steel coils in the stockpile from a top-view perspective, and determining the stacking rules of the steel coils based on the position markings and the top-view depth map; and taking inventory of the steel coils in the stockpile based on the stacking rules to obtain corresponding inventory results; wherein the inventory results include the quantity information of the steel coils. The step of determining the stacking rules of steel coils based on the location markings and the top-view depth map includes: dividing the steel coils in the steel coil yard into different stacks according to the location markings corresponding to the steel coils in the steel coil yard and the spatial relationship between each steel coil; sampling the steel coils in the same stack in the top-view depth map to obtain the average depth of each target steel coil in the same stack, wherein the target steel coil is the steel coil visible in the top-view depth map; determining the average depth as a target clustering feature; obtaining a steel coil hierarchical division scheme based on the target clustering feature and different clustering evaluation indicators; determining the hierarchical distribution of each steel coil in the same stack based on the steel coil hierarchical division scheme; and determining the stacking rules based on each hierarchical distribution.

2. The steel coil inventory method based on three-dimensional reconstruction according to claim 1, characterized in that, The process of generating three-dimensional reconstruction data corresponding to the steel coil yard based on the target image data includes: Using sparse reconstruction technology and the target image data, calibration information corresponding to the target image data and sparse 3D point cloud data are generated; wherein, the calibration information corresponds one-to-one with the target image data, and the calibration information includes the optical center position of the camera in 3D space and the line of sight of the camera; The target image data, the calibration information, and the sparse three-dimensional point cloud data are used to generate the three-dimensional Gaussian reconstruction data corresponding to the steel coil stack.

3. The steel coil inventory method based on three-dimensional reconstruction according to claim 2, characterized in that, The process of obtaining the target multimodal data corresponding to the steel coil stockpile based on the three-dimensional reconstruction data includes: The initial point cloud model of the steel coil stockpile is obtained using the three-dimensional Gaussian reconstruction data in the three-dimensional reconstruction data; The initial point cloud model is aligned using the optical center position in the calibration information to obtain the target point cloud model; The grid size corresponding to the target point cloud model is determined using the elbow rule, and the target point cloud model is rasterized based on the grid size to obtain the top-view depth map.

4. The steel coil inventory method based on three-dimensional reconstruction according to claim 1, characterized in that, Before processing the top-view depth map in the target multimodal data using the stockyard inventory network model, the process further includes: A three-dimensional simulation space is constructed, and an initial virtual steel coil stack is randomly generated in the three-dimensional simulation space according to the stacking rules of the steel coils in the steel coil yard; wherein, the initial virtual steel coil stack exists in the three-dimensional simulation space in the form of a point cloud; Noise data is added to the initial virtual steel coil stack to obtain the corresponding target virtual steel coil stack; wherein, the noise data includes tilt noise and filling noise, the tilt noise is the noise generated by rotating the initial virtual steel coil along the X-axis, Y-axis and Z-axis in the three-dimensional simulation space by a target angle respectively, and the filling noise is the noise generated in the gaps of the initial virtual steel coil stack using a target noise generation algorithm; An initial simulated top-down depth map is generated using the target virtual steel coil stack, and Burmester noise and highly obfuscated noise are added to the initial simulated top-down depth map to obtain the target simulated top-down depth map. The initial simulated top-down depth map exists in the form of an unsigned binary number, and the highly obfuscated noise is the noise generated by performing a linear operation on the depth value of each pixel in the initial simulated top-down depth map to map the original depth value to a predetermined numerical range. The target simulation top-down depth map is automatically annotated to obtain the corresponding annotated simulation top-down depth map, and the annotated simulation top-down depth map is used to generate a target training set.

5. The steel coil inventory method based on three-dimensional reconstruction according to claim 4, characterized in that, Before processing the top-view depth map in the target multimodal data using the stockyard inventory network model, the process further includes: An initial network model is obtained, and the initial network model is trained offline using the target training set to obtain the yard inventory network model; wherein the yard inventory network model consists of several layers of convolutional network structure.

6. A steel coil inventory device based on three-dimensional reconstruction, characterized in that, include: An image acquisition module is used to acquire images of the steel coil stack from different perspectives to obtain target image data corresponding to the steel coil stack, and to generate three-dimensional reconstruction data corresponding to the steel coil stack based on the target image data; wherein, the three-dimensional reconstruction data includes calibration information corresponding to the target image data, sparse three-dimensional point cloud data corresponding to the steel coil stack, and three-dimensional Gaussian reconstruction data, and the three-dimensional Gaussian reconstruction data is a three-dimensional scene representation generated based on the target image data and the sparse three-dimensional point cloud data; The data acquisition module is used to acquire target multimodal data corresponding to the steel coil stockpile based on the three-dimensional reconstruction data; wherein, the target multimodal data includes a top-view depth map of the steel coil stockpile; The steel coil inventory module is used to process the top-view depth map in the target multimodal data using a stockyard inventory network model, so as to inventory the steel coils in the steel coil stockyard based on the stacking rules of the steel coils; wherein the steel coils in the steel coil stockyard are stacked in a stacked structure, and different steel coils block each other. The stockpile inventory network model includes: a stacking rule determination submodule, used to perform target identification on the top-view depth map in the target multimodal data to obtain the position markings of the steel coils in the steel coil stockpile from a top-view perspective, and to determine the stacking rules of the steel coils based on the position markings and the top-view depth map; and a steel coil inventory unit, used to perform inventory of the steel coils in the steel coil stockpile based on the stacking rules to obtain corresponding inventory results; wherein the inventory results include the quantity information of the steel coils. The stacking rule determination submodule includes: a stacking division unit, used to divide the steel coils in the steel coil yard into different steel coil stacks according to the location markings corresponding to the steel coils in the steel coil yard and the spatial positional relationship between each steel coil; an average depth acquisition unit, used to sample the steel coils in the same steel coil stack in the top-view depth map to obtain the average depth of each target steel coil in the same steel coil stack, wherein the target steel coil is the steel coil visible in the top-view depth map; a hierarchical division scheme acquisition unit, used to determine the average depth as the target clustering feature, and obtain a steel coil hierarchical division scheme based on the target clustering feature and different clustering evaluation indicators; and a stacking rule determination unit, used to determine the hierarchical distribution of each steel coil in the same steel coil stack based on the steel coil hierarchical division scheme, and determine the stacking rule according to each hierarchical distribution.

7. The steel coil inventory device based on three-dimensional reconstruction according to claim 6, characterized in that, The image acquisition module includes: A calibration information generation unit is used to generate calibration information corresponding to the target image data and sparse three-dimensional point cloud data using sparse reconstruction technology and the target image data; wherein, the calibration information corresponds one-to-one with the target image data, and the calibration information includes the optical center position of the camera in three-dimensional space and the line of sight of the camera; The Gaussian reconstruction data generation unit is used to generate the three-dimensional Gaussian reconstruction data corresponding to the steel coil stack using the target image data, the calibration information, and the sparse three-dimensional point cloud data.

8. The steel coil inventory device based on three-dimensional reconstruction according to claim 7, characterized in that, The data acquisition module includes: The point cloud model acquisition unit is used to acquire the initial point cloud model of the steel coil stockpile using the three-dimensional Gaussian reconstruction data in the three-dimensional reconstruction data. A point cloud model alignment unit is used to align the initial point cloud model using the optical center position in the calibration information to obtain a target point cloud model. The top-view depth map acquisition unit is used to determine the grid size corresponding to the target point cloud model using the elbow rule, and to rasterize the target point cloud model based on the grid size to obtain the top-view depth map.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steel coil inventory method based on three-dimensional reconstruction as described in any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the steel coil inventory method based on three-dimensional reconstruction as described in any one of claims 1 to 5.