A Goods Management Method, Device and Related Components Based on Depth Estimation Algorithm

Through the depth estimation calculation method, the depth information of the cargo image is obtained, combined with area and occlusion analysis, the problem of insufficient identification accuracy in the elevated library is solved, and the high-accuracy cargo quantity estimation is achieved in complex scenarios.

CN119741699BActive Publication Date: 2025-07-11NEW TREND INT LOGIS TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510258304.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

Traditional 2D image processing technology is difficult to accurately perceive the spatial position relationship of goods in elevated warehouses, especially in the case of non-full stacking, and is prone to missed detection and missed detection in occlusion scenarios, which cannot effectively verify physical constraints, resulting in insufficient accuracy of automatic inventory.

Method used

The depth estimation calculation method is used to obtain the depth information of the cargo image, and the occlusion analysis and consistency analysis are performed in combination with the depth information and area information. The depth information of the cargo is obtained through the depth estimation calculation method, and the area analysis and occlusion detection are carried out to ensure the accuracy of the estimation of the cargo quantity.

Benefits of technology

It improves the accuracy of cargo quantity identification, can accurately identify the occlusion relationship and consistency of cargo in complex scenarios, ensures the reliability of identification results, and provides high-accurate cargo quantity estimates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741699B_ABST
    Figure CN119741699B_ABST
Patent Text Reader

Abstract

The present invention discloses a goods management method, device and related components based on a depth estimation algorithm. The method includes: collecting a goods image of goods to be estimated; obtaining depth information of the goods image based on the depth estimation algorithm; performing goods area analysis on the goods according to the depth information to obtain area information; performing occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information, and performing consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information; estimating the quantity of goods based on the goods occlusion information and the goods consistency information, and outputting an estimation result. In the embodiments of the present invention, depth information is extracted from the goods image through the depth estimation algorithm, which can accurately analyze the goods area, thereby effectively estimating the quantity of goods and outputting an estimation result with high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of warehouse management, and particularly to a goods management method, device and related components based on a depth estimation algorithm. Background Art

[0002] In the automated management of high-bay warehouses, the rapid and accurate inventory of goods is a key link in improving logistics efficiency. Traditional inventory methods rely on manual operations, which are time-consuming and error-prone, especially in an environment like a high-bay warehouse where space is limited and there are numerous storage locations. With the development of technology, taking side photos of storage locations during the movement of stackers and then using image processing technology to calculate the stack type and quantity of goods has become an innovative automatic inventory method. This method can significantly reduce the need for manual inventory and improve the efficiency and accuracy of inventory.

[0003] However, traditional 2D image processing technology has limitations in dealing with such problems. First, 2D images lack depth information, which makes it difficult for algorithms to accurately perceive the spatial position relationships of goods. Especially when dealing with non-full pallets, the recognition error rate is relatively high. Second, 2D object detection is weak in dealing with occlusion scenarios and is prone to missed detections and false detections. Especially in high-bay warehouses, there may be complex occlusion relationships between goods. In addition, the spatial reasoning ability of 2D image processing technology is limited, unable to verify physical constraints and lacking three-dimensional consistency checks, which further restricts its application in the automatic inventory of high-bay warehouses.

[0004] To solve these problems, researchers have tried various methods. For example, some studies have tried to improve the recognition accuracy by improving image classification algorithms, but the effect of this method on the recognition of non-full pallets is still limited. In addition, some studies have also tried to combine deep learning technology and use deep learning models such as convolutional neural networks (CNNs) to improve the recognition ability of goods, but these methods still face challenges in dealing with occlusion and spatial position relationships.

[0005] In summary, although the application prospect of automatic inventory technology in high-bay warehouses is broad, the current technology still faces many challenges, especially in dealing with complex scenarios and improving recognition accuracy. Therefore, how to ensure the accuracy of goods recognition in different scenarios is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0006] Embodiments of the present invention provide a goods management method, device, computer device and storage medium based on a depth estimation algorithm, aiming to improve the accuracy of goods quantity recognition.

[0007] In a first aspect, embodiments of the present invention provide a goods management method based on a depth estimation algorithm, including:

[0008] Collect a goods image of the goods to be estimated;

[0009] Obtain depth information of the goods image based on a depth estimation algorithm;

[0010] Perform goods area analysis on the goods according to the depth information to obtain area information;

[0011] Perform occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information, and perform consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information;

[0012] Estimate the quantity of the goods based on the goods occlusion information and the goods consistency information, and output the estimation result.

[0013] In a second aspect, an embodiment of the present invention provides a goods management device based on a depth estimation algorithm, including:

[0014] An image acquisition unit for collecting a goods image of the goods to be estimated;

[0015] A depth estimation unit for obtaining depth information of the goods image based on a depth estimation algorithm;

[0016] A first analysis unit for performing goods area analysis on the goods according to the depth information to obtain area information;

[0017] A second analysis unit for performing occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information, and performing consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information;

[0018] A quantity estimation unit for estimating the quantity of the goods based on the goods occlusion information and the goods consistency information, and outputting the estimation result.

[0019] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the goods management method based on the depth estimation algorithm as described in the first aspect is implemented.

[0020] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the goods management method based on the depth estimation algorithm as described in the first aspect is implemented.

[0021] An embodiment of the present invention provides a goods management method, device, computer device and storage medium based on a depth estimation algorithm. The method includes: collecting a goods image of goods to be estimated; obtaining depth information of the goods image based on the depth estimation algorithm; performing goods area analysis on the goods according to the depth information to obtain area information; performing occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information, and performing consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information; estimating the quantity of goods based on the goods occlusion information and the goods consistency information, and outputting an estimation result. In the embodiment of the present invention, depth information is extracted from the collected goods image through the depth estimation algorithm, and the goods area can be accurately analyzed. Even when there is occlusion between goods, the occlusion relationship of the goods can be accurately identified through the combined analysis of the depth information and the area information. In addition, the embodiment of the present invention further ensures the reliability of the recognition result through consistency analysis. By comprehensively considering the goods occlusion information and the consistency information, the quantity of goods can be effectively estimated, and an estimation result with high accuracy can be output. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic flowchart of a goods management method based on a depth estimation algorithm provided by an embodiment of the present invention;

[0024] Figure 2 It is a schematic sub - flowchart of a goods management method based on a depth estimation algorithm provided by an embodiment of the present invention;

[0025] Figure 3 It is a schematic block diagram of a goods management device based on a depth estimation algorithm provided by an embodiment of the present invention;

[0026] Figure 4 It is a schematic sub - block diagram of a goods management device based on a depth estimation algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0029] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] Please refer to the following Figure 1 , the embodiments of the present invention provide a goods management method based on a depth estimation algorithm, specifically including: steps S101 to S105.

[0032] Step S101, collect a goods image of the goods to be estimated;

[0033] Step S102, obtain the depth information of the goods image based on the depth estimation algorithm;

[0034] Step S103, perform goods area analysis on the goods according to the depth information to obtain area information;

[0035] Step S104, perform occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information, and perform consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information;

[0036] Step S105, perform goods quantity estimation based on the goods occlusion information and the goods consistency information, and output the estimation result.

[0037] In this embodiment, first, an image containing the goods to be estimated is collected. Then, the depth information in the image is extracted using a depth estimation algorithm. Next, the goods area is identified by analyzing the depth information to obtain area information. Further, occlusion analysis and consistency analysis are respectively performed in combination with the depth information and the area information to determine the occlusion situation and consistency characteristics of the goods. Finally, the quantity of the goods is estimated based on the occlusion information and the consistency information, and the final estimation result is output.

[0038] In this embodiment, the depth information is extracted from the collected goods image through a depth estimation algorithm, which can accurately analyze the goods area. Even when there is occlusion between goods, the occlusion relationship of the goods can be accurately identified through the combined analysis of the depth information and the area information. In addition, this embodiment further ensures the reliability of the recognition result through consistency analysis. By comprehensively considering the goods occlusion information and the consistency information, the quantity of the goods can be effectively estimated, and a highly accurate estimation result can be output. Therefore, the goods management method provided in this embodiment has important application value in the field of goods management, can provide more accurate goods quantity information for industries such as warehousing and logistics, and thus improve the overall operation efficiency and accuracy.

[0039] In practical application scenarios, configure the hardware facilities, such as setting up cameras for collecting images of goods. The installation location of the cameras is crucial. Two (or more) cameras can be installed on the top of the stacker's load platform, and these two cameras will be installed facing each other, respectively responsible for photographing the opposite storage locations. This installation method has several obvious advantages: the cameras will not interfere with the loading and unloading operations of the goods, and at the same time, since the cameras are far from the positions to be photographed, it is beneficial to keep the photographed pictures from being distorted. For the parameter requirements of the cameras, there are the following basic requirements: the resolution should be at least 2 million pixels (1920×1080), but it is recommended to use 4 million pixels (2560×1440) to obtain clearer images. The frame rate should be at least 30fps to adapt to the moving speed of the stacker. The shutter type should be a global shutter to avoid distortion of the picture during movement. The dynamic range should be at least 70dB to cope with different lighting environments. In addition, the cameras should support external triggering to cooperate with the motion control of the stacker. In terms of lens parameters, the selection of the focal length should be based on the distance from the camera installation position to the goods (usually 3 - 5 meters), and it is recommended to use an adjustable focal length lens of 12 - 25mm to ensure that the proportion of the goods in the picture is not less than 60%. Regarding the aperture requirements, the adjustable range of the F value should be F1.4 - F16, and the working aperture is recommended to be set between F2.8 - F4.0. The depth of field requirements should ensure the clarity of the entire goods area. To control distortion, a low-distortion industrial lens should be used to ensure that the radial distortion is less than 1% and the tangential distortion is less than 0.5%. Finally, the resolution requirement of the lens is that the central resolution should reach or exceed 100lp / mm, and the edge resolution should reach or exceed 80lp / mm.

[0040] The camera calibration process includes several key steps. First, read the factory calibration parameters from the camera SDK. These parameters include the internal parameter matrix (focal length, principal point), distortion coefficients (radial, tangential), and resolution parameters. After obtaining these data, it is necessary to verify the validity of the calibration parameters and record the calibration certificate number and date. Next is the installation location calibration. This step involves using the pose estimation tool provided by the camera SDK and performing quick calibration based on the reference points of the shelf. During the calibration process, the shelf columns can be used as the vertical reference, and the edge of the pallet can be used as the horizontal reference. Finally, record the relative position relationship between the camera and the shelf.

[0041] Simplified calibration of depth mapping is another important link in the process. Utilize the built-in depth mapping function of the SDK and verify it with standard goods. The specific operation is to place standard goods at a typical working distance, collect a single image for depth verification, and ensure that the depth error is within the allowable range (less than 2 cm). After completing the verification, record the depth mapping parameters.

[0042] Image quality control is an important part of ensuring high-quality images output by the camera. First, the SDK is used to automatically adjust parameters, including automatic exposure control (the target gray value is set between 160 - 180), automatic gain control, automatic white balance, and enabling the anti-shake function. In addition, it is necessary to monitor the image quality in real time through the SDK, including image sharpness, contrast, and noise level, and set up an automatic alarm and adjustment mechanism to ensure that the image quality always meets the standards.

[0043] In subsequent maintenance management, daily inspections and management of recalibration trigger conditions are required. Daily inspections include using the self-diagnostic function of the SDK to check the camera status, monitor image quality indicators, and record operation logs. When the recalibration conditions are triggered, camera replacement is required, respond to abnormal calibration parameters reported by the SDK, handle situations where the image quality deteriorates significantly, and perform recalibration when the measurement accuracy exceeds the tolerance (verified by standard goods).

[0044] When collecting the images of goods, when the stacker passes by the side of the goods location, it will trigger the camera to take pictures of the goods location through the notification of the upper-level software. After the shooting is completed, the subsequent algorithm system will process the information obtained from the shooting and identify the quantity of goods. Then, this information will be compared with the data in the warehouse management system to achieve the inventory count.

[0045] In one embodiment, as Figure 2 shown, step S102 includes: steps S201 - S205.

[0046] Step S201, preprocess the goods image;

[0047] Step S202, input the preprocessed goods image into the monocular depth estimation model, and the monocular depth estimation model outputs the corresponding depth image;

[0048] Step S203, perform relative depth to absolute distance conversion on the depth image, and after the processing is completed, perform bilateral filtering denoising and outlier detection processing on the depth image;

[0049] Step S204, generate a three-dimensional point cloud image from the depth image, and perform point cloud denoising processing on the three-dimensional point cloud image;

[0050] Step S205, based on the three-dimensional point cloud image after point cloud denoising processing, extract the depth information of the goods.

[0051] In this embodiment, before estimating the depth of the goods image, the input image needs to be preprocessed. For example, the goods image can be first scaled to 384×384 pixels, which is the balance point between computational efficiency and accuracy. Then, the brightness of the goods image is normalized to the range of [0,1] to meet the requirements of the model. In addition, Gaussian filtering is used for denoising, and the kernel size of the filter is 3×3, and the standard deviation σ is set to 0.5.

[0052] The execution of depth estimation depends on the Depth Anything model (i.e., the aforementioned monocular depth estimation model). The input of this model is the preprocessed RGB image, and the output is a relative depth map of 384×384. The range of depth is from 0 (nearest) to 1 (farthest). To improve the inference speed, FP16 half-precision calculation can be adopted, which can double the speed. The inference speed can also be increased by 1.5 times through TensorRT optimization.

[0053] Furthermore, for the depth image output by the Depth Anything model, post-processing of the depth map is performed, and the post-processing of the depth map includes converting depth to absolute distance, bilateral filtering for denoising, and outlier processing. First, the conversion of depth to absolute distance is based on camera parameters and calibration data, and its actual effective range can be 2-6 meters, which can match the actual size of the shelf. Second, bilateral filtering for denoising can use parameters of spatial σ = 3 and range σ = 0.1, which can remove noise while maintaining edges. Finally, outlier processing can set values outside the range (such as less than 2 meters or greater than 6 meters) to invalid, and isolated points can also be removed through 8-neighborhood checking.

[0054] In a specific embodiment, the specific process of converting depth to absolute distance is as follows:

[0055] Perform mapping conversion based on camera calibration parameters:

[0056] 1. First, obtain the internal parameter matrix and distortion coefficients of the camera;

[0057] 2. Establish a geometric transformation relationship according to the installation height and pitch angle of the camera.

[0058] Let the installation height of the camera be H (usually the shelf height + 0.5 meters), and the downward tilt angle of the camera be θ (usually 15°). For any point P(x,y) on the image:

[0059] (1) First, establish an image coordinate system: The origin is located at the center of the image, the x-axis is positive to the right, and the y-axis is positive downward;

[0060] (2) Calculate the line-of-sight vector of this point:

[0061] Horizontal deflection angle α = arctan((x - c x ) / f x );

[0062] Vertical deflection angle β = arctan((y - c y ) / f y );

[0063] Among them, (c x , c y ) is the principal point coordinate, and f x , f y is the focal length parameter;

[0064] (3)Spatial geometric relationship:

[0065] The included angle φ between the line of sight vector and the vertical direction is φ = θ + β;

[0066] The horizontal projection angle remains α;

[0067] (4)Actual distance calculation:

[0068] Vertical distance Dv = H / cos(φ);

[0069] Horizontal distance Dh = Dv * tan(φ);

[0070] Spatial straight-line distance D = sqrt(Dv² + Dh²).

[0071] 3. Using the principle of similar triangles, map the relative depth value to the actual distance:

[0072] Assume that the relative depth value range is [0, 1]. According to the known nearest distance (for example, d min = 2m) and the farthest distance (for example, d max = 6m). Note: The nearest and farthest distances are determined by the camera performance. For any relative depth value d, its corresponding actual distance D is calculated as:

[0073] D = d min + d * (d max - d min ).

[0074] 4. Consider the correction of the viewing angle factor:

[0075] Calculate the included angle between the line of sight and the vertical direction according to the pixel position, and correct the depth value using the cosine theorem

[0076] (1)Viewing angle distortion correction:

[0077] Due to the non-verticality of the camera's perspective, the same depth plane appears trapezoidally distributed on the image. Therefore, the depth values in the edge region need to be compensated according to the perspective:

[0078] a. Calculate the distance r from the pixel point to the center of the image;

[0079] b. Obtain the angle ω between the point and the optical axis, where ω = arctan(r / f), and f is the equivalent focal length;

[0080] c. Apply cosine correction: D 实际 = D 测量 / cos(ω)

[0081] (2)Perspective effect correction:

[0082] For the case where the surface of the goods is inclined or stacked in a stepped manner:

[0083] a. Detect the normal vector of the local plane;

[0084] b. Calculate the angle γ between the line-of-sight vector and the normal vector;

[0085] c. Correct the depth value: D 修正 = D 实际 * cos(γ);

[0086] (3)Edge transition processing:

[0087] In the edge region of the goods, set the width of the transition zone (usually 2 - 3 pixels), and perform weighted averaging of the depth values within the transition zone to avoid measurement errors caused by sudden depth changes;

[0088] (4)Calibration of correction coefficients:

[0089] Collect data using a standard test object at different perspectives, thereby establishing a perspective - correction coefficient correspondence table, and calculating the correction coefficients for intermediate perspectives based on interpolation.

[0090] 5. Fine-tune using the calibration coefficients:

[0091] Obtain the correction coefficients using a standard test object, and then compensate for the systematic errors according to the correction coefficients.

[0092] 6. Installation, verify the conversion accuracy after calibration (optional):

[0093] In the calibration process, a standard test object can be used to verify the conversion results to ensure that the measurement error is within the range of ±1 cm.

[0094] In another specific embodiment, when generating a three-dimensional point cloud image from the depth image, it can be specifically implemented based on the back-projection of the camera internal parameters, that is, converting the pixel coordinates into camera coordinates and converting the depth value into a spatial distance. Specifically, the mapping transformation from the image plane to the camera coordinate system involves processing each pixel point (u, v) on the depth map. Among them, u represents the pixel position in the horizontal direction (from 0 to the image width - 1), v represents the pixel position in the vertical direction (from 0 to the image height - 1), and d represents the absolute depth value corresponding to this pixel. Next, the coordinate transformation is performed using the camera internal parameter matrix to normalize the pixel coordinates to the camera principal point. The specific calculation method is as follows:

[0095] x' = (u - c x ) / f x , y' = (v - c y ) / f y ;

[0096] Among them, (c x , c y ) is the principal point coordinate, and f x , f y are the focal length parameters.

[0097] Finally, calculate the three-dimensional space coordinates:

[0098] X = x' * d, Y = y' * d, Z = d;

[0099] Through the above steps, the spatial points (X, Y, Z) in the camera coordinate system can be obtained.

[0100] Further, when performing point cloud noise reduction processing on the three-dimensional point cloud image, it mainly includes statistical outlier removal and voxel downsampling. Specifically, when processing point cloud data, the first step is to perform statistical outlier removal, which includes local density analysis. For each point P, its local neighborhood characteristics are calculated, and the search radius R is set to 2 cm, which is related to the roughness of the cargo surface. Within this radius range, the number of neighboring points N is counted, and the average distance D_avg to the neighboring points is calculated. The outlier determination rules include setting the minimum number of neighboring points threshold N_min to 10, and calculating the mean μ and standard deviation σ of the average distance in the local area. If the number of neighboring points N is less than N_min, or the absolute value of the difference between the average distance D_avg and the mean μ is greater than 2σ, then this point is marked as an outlier. After the obvious outliers are removed in the first round, the local statistical characteristics are updated, and secondary optimization is performed to remove edge noise points. Next is voxel downsampling. First, the space is divided to establish a three-dimensional grid of 1 cm × 1 cm × 1 cm, and the point cloud space is evenly divided into voxel units. For each voxel unit, the spatial coordinates of all points inside are counted, and the centroid point is calculated as the representative point. At the same time, attributes such as the depth value and reflection intensity of the points are saved. In the sampling optimization stage, the empty voxels are processed, the positions of the empty voxels are marked, and records are retained for future space occupancy analysis.

[0101] In one embodiment, step S103 includes:

[0102] Perform spatial alignment processing on the cargo in combination with the depth information, and construct a shelf coordinate system after the spatial alignment processing is completed;

[0103] Based on the shelf coordinate system, use the depth edge detection algorithm to locate the pallet for carrying the cargo, and extract the pallet features according to the positioning result;

[0104] Perform hierarchical analysis on the cargo based on the shelf coordinate system;

[0105] Analyze the space occupancy of the cargo in combination with the pallet features and the results of the hierarchical analysis to determine whether there is cargo on each pallet layer, and set the judgment result as the area information.

[0106] When obtaining regional information through depth information in this embodiment, first, the depth information is used to perform spatial alignment processing on the goods, and a shelf coordinate system is established. Then, the depth edge detection algorithm is used to locate the pallet and extract the features of the pallet. Next, the goods are analyzed layer by layer according to the shelf coordinate system. Finally, combining the pallet features and the results of the layer-by-layer analysis, the spatial occupancy of the goods is analyzed, whether there are goods on each layer of the pallet is judged, and this information is recorded as regional information. It can be understood that in actual warehouse management, multiple shelves are set, each shelf has multiple layers, and multiple pallets are set on each layer to store goods. Therefore, this embodiment obtains the regional information of the goods through information such as shelves and pallets.

[0107] In a specific embodiment, the combining the depth information to perform spatial alignment processing on the goods and constructing a shelf coordinate system after the spatial alignment processing is completed includes:

[0108] Based on the depth information, obtain the depth mutation band in the depth image, and extract the target depth mutation band that meets the preset depth value from the depth mutation band;

[0109] Extract the linear structure from the three-dimensional point cloud image, and determine the parallel line segment group according to the extracted linear structure;

[0110] Select the first candidate coordinate system horizontal axis and the first candidate coordinate system vertical axis that are perpendicular to each other according to the target depth mutation band; and select the second candidate coordinate system horizontal axis and the second candidate coordinate system vertical axis that are perpendicular to each other according to the parallel line segment group;

[0111] Combine the first candidate coordinate system horizontal axis and the second candidate coordinate system horizontal axis to select the coordinate system horizontal axis, and combine the first candidate coordinate system vertical axis and the second candidate coordinate system vertical axis to select the coordinate system vertical axis, so as to construct the shelf coordinate system.

[0112] When performing spatial alignment and establishing the shelf coordinate system, the coordinate system can be established based on the structural characteristics of the shelf. Specifically, the main reference feature is the horizontal shelf bracket. First, the determination of the horizontal bracket feature is carried out, including the detection of depth map features, such as the detection of depth mutation bands. The width of the mutation band should conform to the bracket specifications, usually in the range of 8 - 12 cm, which depends on the bracket specifications and the actual project scenario. The depth value should protrude 2 - 5 cm relative to the background, also depending on the bracket specifications and the actual project scenario. Then, the structural feature verification is carried out. On the one hand, this involves detecting the linear structure in the point cloud. The RANSAC method can be used to fit the main straight line segments. By randomly sampling, two points are selected each time to form a straight line hypothesis, and the distance from other points to the hypothesis straight line is calculated (the threshold is 1 cm). The number of support points (points with a distance less than the threshold) is counted, and iterative optimization is carried out to find the straight line with the most support points. After removing the points on the fitted straight line, continue to find the next main straight line. The straight line segments whose lengths meet the standard span of the shelf are screened, and the parallelism and perpendicularity between the straight line segments are calculated. On the other hand, it involves multi-layer structure analysis, including detecting a group of parallel line segments with equally spaced distribution (the standard spacing is 1.5 m ± 5 cm), determining the direction of the main parallel line family, and finding the line segments perpendicular to the main parallel line family as column candidates. Subsequently, taking the connection point of the horizontal bracket and the column as the reference point, the trend of the horizontal bracket determines the horizontal reference, and the vertical relationship of multiple horizontal brackets determines the Y-axis direction.

[0113] The auxiliary reference features include the vertical edge of the column as the auxiliary confirmation of the Y-axis, the horizontal plane of the horizontal bracket as the reference for the XZ plane, and the bracket spacing as the scale reference (standard spacing).

[0114] The criteria for determining the horizontal bracket include that any one of the following conditions can be used to determine the horizontal bracket candidate: Condition 1 is that the width of the depth mutation band is in the range of 8 - 12 cm and protrudes 2 - 5 cm relative to the background; Condition 2 is that the length of the straight line segment obtained by RANSAC fitting is greater than 2 m and the number of support points exceeds 80% of the total number of points; Condition 3 is that more than two parallel and equally spaced (1.5 m ± 5 cm) line segments are detected. The final confirmation condition is to meet at least one of the following: Confirmation Condition 1 is to meet any two of the above three conditions simultaneously; Confirmation Condition 2 is to meet Condition 1 and detect a parallel line segment 1.5 m ± 5 cm above and below it; Confirmation Condition 3 is to meet Condition 2 and detect a vertical line segment (column) at each of its two ends.

[0115] The determination criteria for column features include depth features, structural features, and column determination criteria. Among them, for depth features, the width of the depth mutation zone is required to be within the range of 15 - 20 cm, protruding 8 - 12 cm relative to the background, and the depth value remains stable along the vertical direction (the change is less than 1 cm). For structural features, the length of the straight line segment obtained by RANSAC fitting is required to be greater than 3 m (the standard floor height of the shelf), the number of support points exceeds 85% of the total number of points, and the included angle between the inclination angle of the straight line segment and the gravity direction is less than 5 degrees. The column determination criteria include preliminary determination (meeting all of the following conditions): the width and protrusion of the depth mutation zone meet the requirements, and the length and number of support points of the RANSAC fitting straight line segment meet the thresholds; it also includes confirmation determination, meeting at least one of the following: Confirmation condition 1 is that transverse brackets are detected within 1 m on both sides of the straight line segment; Confirmation condition 2 is that another column meeting the preliminary determination is detected at a parallel distance of 2.7 m ± 10 cm.

[0116] Preferably, the coordinate system is precisely optimized. First, through the least - squares fitting method based on multiple reference points, the accuracy of the data is ensured. Then, through the transverse bracket edge extraction and straight line fitting technology, the structure of the coordinate system is further refined. To ensure the accuracy of the vertical direction reference, a vertical direction reference is established, and the origin positioning and the determination of the coordinate axis direction are carried out. In addition, by calculating the coordinate transformation matrix, a rigid - body transformation from the camera coordinate system to the shelf coordinate system can be achieved. To optimize the calibration process and consider measurement errors, corresponding measures can be taken, and finally, the transformation accuracy is controlled within the range of ±1 cm.

[0117] In another actual application process, when analyzing the goods area, it is first necessary to locate the pallet, which involves depth edge detection and pallet feature extraction. Depth edge detection uses the Sobel operator to extract depth discontinuities, and the depth gradient threshold is set to be greater than 0.5 m per pixel. Pallet feature extraction focuses on its standard size of 1.2 m × 1.0 m, with a height of about 15 cm, and allows an error within the range of ±2 cm.

[0118] Next, goods layer analysis is carried out, including analysis in the vertical and horizontal directions. In the vertical direction, through height histogram calculation, the bin size is set to 5 cm, with a range from 0 to 2 m (relative to the pallet surface). The standard spacing for layer - spacing detection is 30 cm ± 2 cm, the minimum spacing is 25 cm, and the maximum spacing is 35 cm. The confirmation of the number of layers is based on the standard stack pattern (28 pieces or 30 pieces), and the rationality of the total height is verified. In the horizontal direction, the standard spacing for column - spacing detection is 25 cm ± 2 cm, evenly divided based on the pallet width. The determination of the column position is numbered from left to right, considering the standard layout of the stack pattern.

[0119] Finally, spatial occupancy determination is performed, which includes local feature analysis and status determination rules. In local feature analysis, the voxel size for point cloud density calculation is 2 cm × 2 cm × 2 cm, and the density threshold is more than 50 points per voxel. Depth continuity is checked through local plane fitting, and the residual threshold is set at ±2 cm. In the status determination rules, the conditions for having goods include that the point cloud density meets the threshold, the depth continuity is good, and the height meets the hierarchical requirements; the conditions for having no goods include low point cloud density, discontinuous depth, or directly observing the background.

[0120] Furthermore, for depth continuity checking, it can specifically include local plane fitting and continuity evaluation. Among them, when performing local plane fitting, first, a checking window with a size of 10 cm × 10 cm needs to be selected and slid with a step size of 2 cm (the step size is the same as the voxel size). Next, the RANSAC method is used for plane fitting. This method randomly samples and selects 3 non - collinear points each time to calculate the plane equation ax + by + cz + d = 0. Then, calculate the distances from other points to this plane and count the number of support points (i.e., the number of points with a distance less than 2 cm). Through iterative optimization, the optimal plane can be found. For the judgment of plane quality, there are two criteria: the proportion of support points needs to be greater than 85% to be considered an effective plane; the angle between the plane normal vector and the line of sight should be less than 60°. In terms of continuity evaluation, it is necessary to compare the planes of adjacent windows to ensure that the angle between the plane normal vectors is less than 15° and the plane distance difference is less than 2 cm. The determination criterion for regional continuity is that the proportion of continuous planes needs to be greater than 80%, and the maximum discontinuous interval does not exceed 4 cm. Finally, the residual threshold is set at ±2 cm, which is consistent with the accuracy of the system.

[0121] In one embodiment, the step S104 includes:

[0122] The ray - tracing technology based on line - of - sight projection is used to detect the occlusion of the goods, and the occlusion degree is determined for the occluded goods; wherein, the occlusion degree includes complete occlusion, partial folding, and no occlusion;

[0123] Combining the depth information, regional information, and occlusion degree, the status of the completely occluded goods is judged to determine whether the goods actually exist;

[0124] The results of the occlusion degree and status judgment are summarized as the occlusion information.

[0125] In this embodiment, when performing occlusion analysis, it is first necessary to detect the occlusion area and infer the state. Occlusion detection can be achieved through the line-of-sight projection method, that is, project the line of sight from the camera position to the target position and check the depth values of the passing points. According to the degree of occlusion, occlusion can be divided into three cases: complete occlusion, partial occlusion, and no occlusion. For example, complete occlusion means that the line of sight is completely blocked, partial occlusion means that more than 50% of the area is still visible, and no occlusion means that more than 90% of the area is visible.

[0126] For the completely occluded positions, state inference is required. First, perform a support check to verify the state of the underlying goods and ensure that the support area is greater than 80%. Then, conduct an adjacent state analysis to check the adjacent positions on the same layer and consider the standard stack pattern rules. Finally, apply physical constraints, including gravity support rules, stability requirements, and standard spacing limitations.

[0127] In the result verification stage, it is necessary to verify the physical rules. The gravity support verification requires that the upper-layer goods must be supported by the lower layer and the support area should be sufficient. The stability check needs to ensure that the stacking height is reasonable and the interlayer misalignment is within the allowable range. The spacing verification includes horizontal spacing and vertical spacing, for example, which are required to be 25 cm ± 2 cm and 30 cm ± 2 cm respectively. Through these steps, the accuracy of occlusion processing and the stability of the goods stacking can be ensured.

[0128] In a specific embodiment, when using the ray tracing technology based on line-of-sight projection to detect the occlusion of the goods, first, by constructing a ray equation, from the camera position (C x ,C y ,C z ) to the target position (T x ,T y ,T z ), generate the ray parameter equation P(t) = C + t(T - C), where the value range of t is [0,1]. Then, conduct sampling detection, set the sampling interval Δt to 0.01, which corresponds to an actual distance of about 1 cm. For each sampling point P(t), calculate the theoretical depth dt = ||P(t) - C|| and the actual depth dr (obtained from the depth map), and thus obtain the depth difference Δd = dr - dt. According to the depth difference Δd, the occlusion situation can be determined: if Δd < -2 cm, it is confirmed that there is occlusion, and the position and depth difference are recorded; if |Δd| ≤ 2 cm, it is considered that there is no occlusion.

[0129] To further analyze the occlusion degree of the target area, grid division is first performed. For example, the cell size is set to 2 cm × 2 cm, and the total number of grids is 625 (25×25 grids). By performing ray tracing and counting the number of occlusion points Nb, the occlusion ratio R = Nb / N can be calculated. According to the occlusion ratio, the occlusion level can be divided into five levels: complete occlusion (R≥90%), severe occlusion (70%≤R<90%), partial occlusion (30%≤R<70%), slight occlusion (10%≤R<30%), and basically no occlusion (R<10%).

[0130] Finally, convert the occlusion degree to confidence (visibility). For the case of basically no occlusion, the confidence 1 - R will be greater than or equal to 90%.

[0131] In one embodiment, the step S104 further includes:

[0132] Performing in-layer consistency detection on the goods based on the depth information and region information; wherein, the steps of the in-layer consistency detection include: controlling the depth change of the central region of each layer of the goods area according to a preset control strategy, and performing edge processing on each layer of the goods area according to the structure of the central region depth change control; the edge processing includes verifying the rationality of depth mutation and correcting the edge depth;

[0133] Performing inter-layer relationship detection on the goods based on the depth information and region information; wherein, the steps of the inter-layer relationship detection include: calculating the perspective deviation of the adjacent layer of the goods area, and performing depth correction according to the result of the perspective deviation calculation;

[0134] Summarize the results of the in-layer consistency detection and the inter-layer relationship detection, and use them as the goods consistency information.

[0135] When performing depth consistency verification in this embodiment, it is necessary to ensure the accuracy of in-layer consistency and inter-layer relationships. For example, the in-layer consistency requires that the depth change within the same layer does not exceed 3 cm, and the influence of edge effects needs to be excluded. Another example is that the consideration of inter-layer relationships includes ensuring that the depth difference between adjacent layers is 30 cm ± 2 cm, and the influence of perspective needs to be taken into account. These standards jointly ensure the accuracy and reliability of depth data. It can be understood that at the edge of the goods, due to depth mutation, measurement errors are likely to occur, so some verification and correction are required. The floor heights involved in this embodiment are all set based on the height of one layer of goods in a preset business scenario, and are not fixed values.

[0136] Specifically, when performing in-layer consistency checks, it is first necessary to focus on the control of depth changes in the central region. The core region is defined as the area more than 5 cm away from the edge. The depth change threshold should be less than 2 cm, and the depths of adjacent measurement points should show a gradual change trend. For the edge region, special treatment is required. The edge is defined as the range of ±5 cm from the cargo boundary. The rationality of the depth mutation needs to be verified, and the depth difference from the background should be between 25 and 35 cm (standard cargo thickness). At the same time, the width of the mutation band should consider the viewing angle and resolution and be controlled within 2 to 4 cm. The edge depth correction should be based on the local plane fitting result, and the edge depth value is extrapolated through the plane equation, but the correction is only applied when the fitting credibility exceeds 85%.

[0137] When verifying the inter-layer relationship, it is first necessary to ensure that the depth difference between adjacent layers meets the predetermined standard, such as 30 cm ± 2 cm. Next, compensation for the viewing angle effect needs to be carried out, which mainly includes two aspects: the camera tilt effect and the horizontal viewing angle effect. For the camera tilt effect, it is necessary to calculate the downward tilt angle θ of the camera (usually 15°) and the angle α between the line of sight and the normal vector of the cargo surface, and apply the depth correction coefficient cos(α). For the horizontal viewing angle effect, it is necessary to perform a transformation from the image plane to the world coordinate system, calculate the horizontal offset angle β, and then apply the lateral depth correction cos(β).

[0138] The depth value correction is divided into the vertical direction and the horizontal direction. In the vertical direction, the actual storey height can be obtained by dividing the measured storey height by cos(α), but it is necessary to ensure that the α angle is less than 60° to ensure the reliability of the measurement. In the horizontal direction, the actual spacing can be obtained by dividing the measured spacing by cos(β), and it is necessary to ensure that the β angle is less than 45° to ensure the measurement accuracy.

[0139] Then, the correction value verification is carried out, including physical constraint checks and continuity verification. The physical constraint check ensures that the corrected storey height is within the range of 30 cm ± 1 cm, and the corrected spacing is within the range of 25 cm ± 1 cm. The continuity verification involves checking the depth gradual change of adjacent measurement points and performing outlier detection (values exceeding 3 standard deviations σ). Through these steps, the accuracy of the inter-layer relationship and the reliability of the measurement data can be ensured.

[0140] After that, the verification result processing is carried out. First, the conditions for passing the verification need to be clarified. For in-layer consistency, three sub-conditions need to be met: the depth change in the central region is less than 2 cm, the depth mutation in the edge region is reasonable, and the continuity meets the requirements. For the inter-layer relationship, it is required that the storey height meets the standard (such as 30 cm ± 2 cm), and the data should be reasonable after viewing angle compensation.

[0141] If the verification fails, corresponding handling strategies need to be adopted. For the case of intra-layer consistency failure, for example, if the deviation is less than 4 cm, the local plane fitting result can be used for correction, and the correction situation is recorded while reducing the confidence level (e.g., -0.2). If the deviation is greater than or equal to 4 cm, it is marked as an abnormal area and the re-detection process is triggered. If the re-detection still fails, the status can be downgraded to "awaiting manual confirmation". For the case of inter-layer relationship failure, if the layer height is abnormal, first check whether there are problems with the goods dumping or support status. If the physical state is normal but the depth is abnormal, the perspective compensation is re-executed. If the compensation fails, it is marked as "awaiting manual confirmation". If the data is still abnormal after perspective compensation, it is necessary to check whether the camera parameters are accurate and re-execute the calibration process. If the problem persists, the manual processing process is entered.

[0142] In one embodiment, the step S105 includes:

[0143] Calculating the confidence level of the goods based on the goods occlusion information and goods consistency information to obtain the confidence score of the goods;

[0144] Comparing the confidence score with a preset score to obtain a score comparison result, and constructing a stack pattern matrix according to the confidence score;

[0145] Combining the score comparison result and the stack pattern matrix to determine whether the goods actually exist, and estimating the quantity of the goods accordingly.

[0146] When outputting the estimated result of the goods quantity in this embodiment, the position status of the goods is first summarized to obtain the confidence score. The status collection includes the classification of directly observed positions, inferred positions, and abnormal positions. The visibility of the directly observed positions needs to be greater than 90%, the depth features are clear, and a weight coefficient of 1.0 is assigned. The inferred positions are inferred based on physical rules, and the confidence score is between 0.8 and 0.95. The abnormal positions have unclear features and rule conflicts and require manual confirmation. The confidence calculation involves basic scores. For example, the directly observed position is set to 1.0, partial occlusion is 0.9, and full inference is 0.8. In terms of quantity statistics, the basic statistics include counting the positions with goods, only counting the positions with a confidence score greater than 0.8, and counting by layer. The distribution analysis generates a stack pattern matrix and calculates the filling rate.

[0147] Furthermore, verify the output results, including pallet type rule verification and quantity rationality verification. Pallet type rule verification requires comparison with the standard pallet type to check for abnormal distributions. Quantity rationality is ensured through comparison with historical data and business rule verification. Furthermore, output processing includes the generation of a result report, covering the total quantity statistics, hierarchical quantities, and confidence distribution, as well as anomaly markings, including location numbers, anomaly types, and handling suggestions. In terms of data storage, it is necessary to save image data (original images, depth maps, visualization results), processing parameters (algorithm configurations, threshold settings), and statistical results (in JSON format, including timestamps).

[0148] In actual application scenarios, adopt a method of real-time quality monitoring and feedback for goods management. Specifically, first, quality monitoring includes real-time indicator monitoring, where the depth quality requires the proportion of valid depth points to exceed 95%, and the noise level is controlled below 1 centimeter. In terms of recognition quality, mainly monitor the confidence distribution and anomaly ratio to ensure that the latter is below 1%. In terms of performance indicators, the processing delay must be less than 1 second, and at the same time, the GPU utilization rate is maintained below 80%. To adapt to the changing environment, automatic adjustment strategies can be implemented, including depth threshold adjustment based on environmental light and distance changes, as well as algorithm parameter adjustment of adaptive noise thresholds and dynamic confidence requirements.

[0149] In terms of the feedback mechanism, on the one hand, it is achieved through immediate feedback, which includes visual display of depth map superposition, status markings, and anomaly prompts, as well as status updates after each processing, and these status updates contain confidence information. On the other hand, it is achieved through historical data, which includes logging of processing parameters, intermediate results, and anomaly situations, as well as performance tracking, such as accuracy statistics, processing efficiency, and resource usage.

[0150] In a practical application scenario, taking a non-full pallet location with a 30-piece stack type (6 layers and 5 columns) as an example, the goods management method provided by this embodiment is illustrated. First, during the data input process, an original RGB image with a resolution of 2560×1440 is used. This image is captured by a camera from a distance of about 4 meters, and the camera lens is tilted downward by 15 degrees to capture the target scene. The focal length of the camera is set to 16 mm, providing a wide field of view of 60 degrees horizontally and 40 degrees vertically. To ensure the accuracy of image processing, the internal parameter matrix and distortion coefficients of the camera are obtained from the software development kit (SDK). Subsequently, depth estimation is performed. First, image preprocessing is carried out, including scaling the image to 384×384 pixels, normalizing the brightness to the range of [0,1], and applying a 3×3 Gaussian filter to remove noise. Next, the processed image is input into the DepthAnything model to obtain a relative depth map of 384×384 pixels, where the depth range is from 0 (nearest) to 1 (farthest). Then, depth post-processing is performed by converting the relative depth to the actual distance based on the camera parameters, applying bilateral filtering to remove depth noise, and detecting and correcting abnormal depth values (greater than 6 meters or less than 2 meters), thus completing the depth post-processing step.

[0151] Then, during the process of region analysis, first, pallet positioning is carried out. Through detection, the position of the pallet edge is determined as (x = 320, y = 980), and the depth of the pallet plane is measured to be 3.8 meters. Based on this information, a local coordinate system with the pallet center as the origin is established for subsequent analysis. Next, layer-by-layer analysis of the pallet is performed. Vertically, 5 significant height intervals can be detected, with a layer spacing of approximately 30 cm, and the positions of 6 different layers including the pallet layer are identified. Horizontally, 4 column intervals are detected, with a column spacing of about 25 cm, and the positions of 5 columns are determined accordingly. When performing spatial occupancy analysis, a depth threshold of ±5 cm is set to ensure the accuracy of the analysis. The point cloud density at each position is analyzed to accurately determine the presence or absence of goods. The following determination results are obtained, where 1 indicates there is goods and 0 indicates there is no goods:

[0152] Layer 6: 1 1 1 0 0 / / The first 3 positions have goods;

[0153] Layer 5: 1 1 1 1 0 / / The first 4 positions have goods;

[0154] Layer 4: 1 1 1 1 1 / / All positions have goods;

[0155] Layer 3: 1 1 1 1 1 / / All positions have goods;

[0156] Layer 2: 1 1 1 1 1 / / All positions have goods;

[0157] Layer 1: 1 1 1 1 1 / / All in stock (including pallets);

[0158] When performing occlusion processing, occlusion detection is first carried out. The detection results show that at the last two positions of the 6th layer and the last position of the 5th layer, the goods are directly visible, and it is confirmed that there are no goods at these positions. For other positions, the visibility of the goods is good. During the verification process, gravity support inspection is carried out to ensure that all goods have lower layer support. In addition, the rationality of the spacing is also checked to confirm that all spacings are within the standard range. Finally, the depth consistency is checked to ensure that the depth changes of each layer are continuous.

[0159] In this quantity statistics, 30 positions were observed in detail. The results show that there are 30 directly observable positions, 0 occluded positions, 27 positions with goods, and 3 empty positions. In terms of result verification, it is confirmed that the layout conforms to the 30-piece stack standard, the quantity is within a reasonable range, and no physical rules are violated. The output results show that the total quantity is 23 pieces, the confidence level is as high as 98%, and no abnormal marks are found. In addition, the distribution characteristics show that there is a shortage in the top layer part. In the feedback and recording link, the system can generate depth visualization images to clearly display the layout and status of the goods. The specific position of each good will be clearly marked to ensure the accuracy of the information. In addition, the system will also display the quantity statistics results, providing intuitive data support for inventory management and decision-making.

[0160] In another actual application scenario, the effect of the goods management method provided by this embodiment is evaluated. First, in terms of depth estimation technology, this embodiment successfully applies the Depth Anything model to the depth estimation of goods, achieving accurate acquisition of depth information in a single view, significantly reducing the system complexity, and eliminating the need for multi-camera calibration. Second, in terms of spatial analysis technology, this embodiment develops a goods spatial positioning method based on depth information, which can accurately identify non-full stack situations and establish a complete occlusion processing mechanism. Finally, in terms of system integration and innovation, this embodiment simplifies the camera deployment requirements, only requires a single camera configuration, optimizes the image acquisition process, improves the acquisition efficiency, and maximizes the coordinated control of the stacker movement.

[0161] Specifically, when performing technical performance evaluation, depth estimation performance evaluation is one of the key links. First, for depth estimation accuracy evaluation, standard test objects are required to be measured in the range of 2 - 6 meters, record the absolute error and relative error, and statistically analyze the error distribution characteristics. Second, for effective depth point evaluation, it is required to calculate the proportion of effective points, analyze the spatial distribution characteristics, evaluate the noise level, and test the edge effect. Through these methods, the performance of the depth estimation technology can be comprehensively understood.

[0162] When evaluating the recognition performance, the whole stack recognition evaluation is first focused on. This includes testing under standard lighting conditions to ensure the performance of the system in an ideal environment. In addition, different stack types are also tested to verify the adaptability of the system to various stacking methods. Through these tests, the recognition accuracy rate can be statistically calculated, and the error types can be analyzed to further optimize the recognition algorithm.

[0163] For the non-full-stack recognition evaluation, a variety of typical missing scenarios are designed to simulate the incomplete stacking situations that may occur in actual applications. By testing different missing ratios, the position accuracy of the system when recognizing incomplete stacks is evaluated, and the difficulties in the recognition process are analyzed, so as to propose improvement measures.

[0164] For the occlusion scenario evaluation, a variety of occlusion test scenarios are designed to simulate the occlusion situations that may occur in the real world. By testing different occlusion degrees, the inference accuracy of the system under occlusion conditions is evaluated, and the failure cases are analyzed to identify and solve potential problems. Through these comprehensive evaluations, the recognition performance of the system can be comprehensively understood, and a basis can be provided for subsequent improvement work.

[0165] When conducting the adaptability evaluation, the environmental adaptability evaluation is first focused on, which includes the test of light adaptability. Different light intensities can be tested to evaluate the impact of backlight on the system and analyze the shadow effect. In addition, the impact of light changes on the system performance can also be tested.

[0166] Secondly, the business scenario adaptability evaluation is also an indispensable part. In terms of stack type adaptability, we will test the performance of the standard stack type, evaluate the adaptability of the system to stack type variants, and analyze the effect of stack type switching. At the same time, the processing ability of abnormal stack types can also be tested.

[0167] In terms of motion adaptability, the performance of the system at different speeds can be tested, the impact of acceleration on the system can be evaluated, and the image quality can be analyzed. In addition, the synchronization performance of the system can also be tested to ensure stable operation in various motion scenarios. Through these comprehensive tests, the adaptability of the system can be comprehensively evaluated to ensure that it can meet the expected performance standards in various environments and business scenarios.

[0168] When conducting an economic evaluation, it is first necessary to meticulously account for costs. Hardware costs include itemizing the required equipment, estimating installation costs, forecasting maintenance costs, and evaluating future upgrade costs. Operating costs involve calculating energy consumption, labor costs, maintenance expenses, and training costs. The benefit evaluation is divided into two parts: direct benefits and indirect benefits. Direct benefits include improved inventory-taking efficiency, labor savings, increased accuracy, and time savings. Indirect benefits are reflected in enhanced management efficiency, improved inventory accuracy, better decision-making support, and increased customer satisfaction. By comprehensively evaluating these costs and benefits, a scientific and reasonable judgment can be made about the economy of the project.

[0169] Figure 3 FIG. is a schematic block diagram of a goods management device 300 provided by an embodiment of the present invention. The device 300 includes:

[0170] An image acquisition unit 301 for acquiring a goods image of the goods to be estimated;

[0171] A depth estimation unit 302 for obtaining depth information of the goods image based on a depth estimation algorithm;

[0172] A first analysis unit 303 for analyzing the goods area of the goods according to the depth information to obtain area information;

[0173] A second analysis unit 304 for performing occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information, and performing consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information;

[0174] A quantity estimation unit 305 for estimating the quantity of the goods based on the goods occlusion information and the goods consistency information and outputting an estimation result.

[0175] In one embodiment, as Figure 4 shown, the depth estimation unit 302 includes:

[0176] A preprocessing unit 401 for preprocessing the goods image;

[0177] A model processing unit 402 for inputting the preprocessed goods image into a monocular depth estimation model, and outputting a corresponding depth image by the monocular depth estimation model;

[0178] A post-processing unit 403 for performing relative depth to absolute distance processing on the depth image, and performing bilateral filtering denoising and outlier detection processing on the depth image after the processing is completed;

[0179] A point cloud generation unit 404 is configured to generate a three-dimensional point cloud image from the depth image and perform point cloud noise reduction processing on the three-dimensional point cloud image;

[0180] A point cloud noise reduction unit 405 is configured to extract the depth information of the goods based on the three-dimensional point cloud image after point cloud noise reduction processing.

[0181] In one embodiment, the first analysis unit 303 includes:

[0182] A spatial alignment unit is configured to perform spatial alignment processing on the goods in combination with the depth information and construct a shelf coordinate system after the spatial alignment processing is completed;

[0183] A feature extraction unit is configured to locate the pallet for carrying the goods based on the shelf coordinate system by using a depth edge detection algorithm and extract pallet features according to the positioning result;

[0184] A hierarchical analysis unit is configured to perform hierarchical analysis on the goods based on the shelf coordinate system;

[0185] An information setting unit is configured to analyze the space occupancy of the goods in combination with the pallet features and the results of hierarchical analysis to determine whether there are goods on each layer of the pallet and set the determination result as the area information.

[0186] In one embodiment, the second analysis unit 304 includes:

[0187] An occlusion detection unit is configured to perform occlusion detection on the goods by using a ray tracing technique based on line-of-sight projection and determine the occlusion degree for the occluded goods; wherein, the occlusion degree includes complete occlusion, partial folding, and no occlusion;

[0188] A status judgment unit is configured to combine the depth information, area information, and occlusion degree to perform status judgment on the completely occluded goods to determine whether the goods actually exist;

[0189] A first summary unit is configured to summarize the results of the occlusion degree and status judgment as the occlusion information.

[0190] In one embodiment, the second analysis unit 304 further includes:

[0191] An intra-layer detection unit is configured to perform intra-layer consistency detection on the goods based on the depth information and area information; wherein, the steps of the intra-layer consistency detection include: controlling the depth change of the central area of each layer of the goods area according to a preset control strategy and performing edge processing on each layer of the goods area according to the structure of the central area depth change control; the edge processing includes verification of the rationality of depth mutation and correction of the edge depth.

[0192] An interlayer detection unit for detecting the interlayer relationship of the goods based on the depth information and the area information; wherein, the steps of the interlayer relationship detection include: calculating the perspective deviation of the adjacent-layer goods areas, and performing depth correction according to the result of the perspective deviation calculation;

[0193] A second summarizing unit for summarizing the results of the in-layer consistency detection and the interlayer relationship detection, and using them as the goods consistency information.

[0194] In one embodiment, the quantity estimation unit 305 includes:

[0195] A confidence calculation unit for calculating the confidence of the goods based on the goods occlusion information and the goods consistency information to obtain the confidence score of the goods;

[0196] A matrix construction unit for comparing the confidence score with a preset score to obtain a score comparison result, and constructing a stack type matrix according to the confidence score;

[0197] A goods judgment unit for judging whether the goods actually exist by combining the score comparison result and the stack type matrix, so as to estimate the quantity of the goods.

[0198] In one embodiment, the space alignment unit includes:

[0199] A mutation band extraction unit for obtaining the depth mutation band in the depth image based on the depth information, and extracting the target depth mutation band that meets the preset depth value from the depth mutation band;

[0200] A linear structure extraction unit for extracting the linear structure from the three-dimensional point cloud image and determining the parallel line segment group according to the extracted linear structure;

[0201] A candidate selection unit for selecting the first candidate coordinate system horizontal axis and the first candidate coordinate system vertical axis that are perpendicular to each other according to the target depth mutation band; and selecting the second candidate coordinate system horizontal axis and the second candidate coordinate system vertical axis that are perpendicular to each other according to the parallel line segment group;

[0202] A coordinate system construction unit for combining the first candidate coordinate system horizontal axis and the second candidate coordinate system horizontal axis to select the coordinate system horizontal axis, and combining the first candidate coordinate system vertical axis and the second candidate coordinate system vertical axis to select the coordinate system vertical axis, so as to construct the shelf coordinate system.

[0203] Since the embodiments of the device part correspond to the embodiments of the method part, for the embodiments of the device part, please refer to the description of the embodiments of the method part, which will not be elaborated here.

[0204] An embodiment of the present invention also provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include various media capable of storing program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0205] An embodiment of the present invention also provides a computer device, which may include a memory and a processor. When the processor calls the computer program stored in the memory, the steps provided in the above embodiments can be implemented. Of course, the computer device may also include various network interfaces, power supplies and other components.

[0206] The embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description of the method part. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0207] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

Claims

1. A goods management method based on a depth estimation algorithm, characterized in that Including: Collecting a goods image of the goods to be estimated; Obtaining depth information of the goods image based on a depth estimation algorithm; Performing goods area analysis on the goods according to the depth information to obtain area information; Performing occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information, and performing consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information; Estimating the quantity of the goods based on the goods occlusion information and the goods consistency information, and outputting an estimation result; The goods area includes multiple layers, and performing consistency analysis on the goods by combining the depth information and the area information to obtain goods consistency information, including: Performing intra-layer consistency detection on the goods based on the depth information and the area information; wherein, the steps of the intra-layer consistency detection include: controlling the depth change of the central area of each layer of the goods area according to a preset control strategy, and performing edge processing on each layer of the goods area according to the structure of the central area depth change control; the edge processing includes verifying the rationality of depth mutation and correcting the edge depth; Performing inter-layer relationship detection on the goods based on the depth information and the area information; wherein, the steps of the inter-layer relationship detection include: calculating the perspective deviation of adjacent layers of the goods area, and performing depth correction according to the result of the perspective deviation calculation; Summarizing the results of the intra-layer consistency detection and the inter-layer relationship detection, and using them as the goods consistency information.

2. The goods management method based on the depth estimation algorithm according to claim 1, wherein The obtaining depth information of the goods image based on the depth estimation algorithm includes: Preprocessing the goods image; Inputting the preprocessed goods image into a monocular depth estimation model, and outputting a corresponding depth image by the monocular depth estimation model; Performing relative depth to absolute distance processing on the depth image, and performing bilateral filtering denoising and outlier detection processing on the depth image after the processing is completed; Generating a three-dimensional point cloud image of the depth image, and performing point cloud denoising processing on the three-dimensional point cloud image; Extracting the depth information of the goods based on the three-dimensional point cloud image after the point cloud denoising processing.

3. The goods management method based on the depth estimation algorithm according to claim 2, wherein The performing goods area analysis on the goods according to the depth information to obtain area information includes: Performing spatial alignment processing on the goods by combining the depth information, and constructing a shelf coordinate system after the spatial alignment processing is completed; Based on the shelf coordinate system, positioning the pallet for carrying the goods by using a depth edge detection algorithm, and extracting pallet features according to the positioning result; Performing layer analysis on the goods based on the shelf coordinate system; Analyzing the space occupancy of the goods by combining the pallet features and the results of the layer analysis to determine whether there are goods on each layer of the pallet, and setting the determination result as the area information.

4. The goods management method based on the depth estimation algorithm according to claim 1, wherein The performing occlusion analysis on the goods by combining the depth information and the area information to obtain goods occlusion information includes: Performing occlusion detection on the goods by using a ray tracing technology based on line-of-sight projection, and determining the occlusion degree for the occluded goods; wherein, the occlusion degree includes complete occlusion, partial folding, and no occlusion; Combined with the depth information, region information, and occlusion degree, the state of the completely occluded goods is judged to determine whether the goods actually exist; The occlusion degree and the result of the state judgment are summarized as the occlusion information.

5. The goods management method based on the depth estimation algorithm according to claim 1, wherein Based on the goods occlusion information and goods consistency information, the quantity of the goods is estimated and the estimation result is output, including: Based on the goods occlusion information and goods consistency information, the confidence level of the goods is calculated to obtain the confidence score of the goods; The confidence score is compared with a preset score to obtain the score comparison result, and a stack type matrix is constructed according to the confidence score; Combined with the score comparison result and the stack type matrix, it is judged whether the goods actually exist, and the quantity of the goods is estimated accordingly.

6. The goods management method based on the depth estimation algorithm according to claim 3, characterized in that Combined with the depth information, spatial alignment processing is performed on the goods, and a shelf coordinate system is constructed after the spatial alignment processing is completed, including: Based on the depth information, the depth mutation band in the depth image is obtained, and the target depth mutation band that meets the preset depth value is extracted from the depth mutation band; Linear structures are extracted from the three-dimensional point cloud image, and a group of parallel line segments is determined according to the extracted linear structures; According to the target depth mutation band, a first candidate coordinate system horizontal axis and a first candidate coordinate system vertical axis that are perpendicular to each other are selected; and according to the group of parallel line segments, a second candidate coordinate system horizontal axis and a second candidate coordinate system vertical axis that are perpendicular to each other are selected; Combined with the first candidate coordinate system horizontal axis and the second candidate coordinate system horizontal axis, the coordinate system horizontal axis is selected, and combined with the first candidate coordinate system vertical axis and the second candidate coordinate system vertical axis, the coordinate system vertical axis is selected, so as to construct the shelf coordinate system.

7. A goods management device based on a depth estimation algorithm, characterized in that, Including: An image acquisition unit for acquiring a goods image of the goods to be estimated; A depth estimation unit for obtaining the depth information of the goods image based on a depth estimation algorithm; A first analysis unit for performing goods region analysis on the goods according to the depth information to obtain region information; A second analysis unit for performing occlusion analysis on the goods by combining the depth information and region information to obtain goods occlusion information, and performing consistency analysis on the goods by combining the depth information and region information to obtain goods consistency information; A quantity estimation unit for estimating the quantity of the goods based on the goods occlusion information and goods consistency information and outputting an estimation result; The goods region includes multiple layers, and the second analysis unit includes: An intra-layer detection unit for performing intra-layer consistency detection on the goods based on the depth information and region information; wherein, the steps of the intra-layer consistency detection include: controlling the depth change of the central region of each layer of the goods region according to a preset control strategy, and performing edge processing on each layer of the goods region according to the structure of the central region depth change control; the edge processing includes depth mutation rationality verification and edge depth correction; An inter-layer detection unit for performing inter-layer relationship detection on the goods based on the depth information and region information; wherein, the steps of the inter-layer relationship detection include: calculating the perspective deviation of adjacent layers of the goods region, and performing depth correction according to the result of the perspective deviation calculation; A second summarization unit, configured to summarize the results of in-layer consistency detection and the results of inter-layer relationship detection, and use the same as the goods consistency information.

8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the goods management method based on the depth estimation algorithm according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the goods management method based on the depth estimation algorithm according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video image depth estimation method based on image segmentation

    CN105069808A

  • Object quantity estimation method and device

    CN105096292A

  • Object number detection method, device and system for stacked objects

    CN112258452A