Compliance detection method based on depth stereoscopic vision
The three-dimensional structure information of the packaging is obtained through a deep stereo camera and combined with image processing technology, the existing detection methods are solved, and efficient automated detection is achieved.
Patent Information
- Application Number
- CN202510459127.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
AI Technical Summary
The existing packaging compliance detection methods rely on manual measurement, are time-consuming and labor-intensive and susceptible to human factors. The existing image-based detection methods can only obtain two-dimensional information, limiting the detection accuracy and reliability.
The three-dimensional structure information of the packaging is obtained by using a depth stereo camera, and through image processing technology, the maximum value algorithm of the sliding window is used to detect the size and shape of the packaging structure, combined with depth image preprocessing and extreme value mapping, automatic detection is achieved.
It improves detection accuracy and reliability, realizes automated detection, reduces labor costs, and improves detection efficiency.
Smart Images

Figure CN120368845A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automated detection, and specifically to a compliance detection method based on depth stereo vision. This method obtains the depth image of the packaging structure through a depth stereo camera, and uses image processing technology to detect the compliance of the size and shape of the packaging. Background Art
[0002] In the packaging industry, ensuring the compliance of packaging is crucial for the transportation, storage, and sales of products. Most traditional packaging compliance detection methods rely on manual measurement and inspection. This method is not only time-consuming and laborious, but also easily affected by human factors, resulting in low accuracy of detection results. With the development of computer vision technology, image-based automated detection methods have gradually been applied to packaging compliance detection. However, most existing image-based detection methods can only obtain two-dimensional information of the packaging and cannot accurately obtain the three-dimensional structure of the packaging, thus limiting the detection accuracy and reliability. Summary of the Invention
[0003] To solve the above problems, the present invention proposes a compliance detection method based on depth stereo vision. This method obtains the three-dimensional structure information of the packaging through a depth stereo camera, and uses image processing technology to accurately detect the size and shape of the packaging.
[0004] To solve the above technical problems, the technical solution provided by the present invention is: A compliance detection method based on depth stereo vision, characterized by including the following steps:
[0005] Step 1: First, install a group of depth stereo cameras at the top position of the source of the packaging conveyor belt with the shooting direction perpendicular to the downward direction, ensuring that the field of view can cover all the space of the object to be detected and both sides of the conveyor belt;
[0006] Step 2: After the depth stereo camera obtains the depth image of the packaging structure to be detected, perform preprocessing operations on this depth information. The specific operations are in sequence: truncation, bidirectional two-dimensionalization, filtering, binarization, and contour detection, and then the contour model of the packaging can be obtained;
[0007] Step 3: Perform a maximum value algorithm based on a sliding window on the above contour model to obtain the contour extreme values within the threshold, and then map the two-dimensional coordinates of the extreme values back to the world coordinates in the depth map, and then the possible positions of the unqualified packaging structure size and the specific overlimit values can be given.
[0008] Step 4: If an overlimit value is detected in the packaging appearance, slide the camera to obtain the overlimit values detected at different positions.
[0009] As an improvement, the contour model of the package in step two includes length L, height H, width W, depth camera installation height Hc, ultra-high setting value Hm, and ultra-wide setting value Wm; since the parameters of the package shell are fixed, they are directly input during system deployment. When it is detected that the length of the package shell exceeds the set value, for the three-dimensional point cloud set obtained from the depth map:
[0010] {P1, P2, P3,...Pn};
[0011] For any point Pk in
[0012] if it satisfies the condition: x Pk.x < Hm - t x or Pk.x > Hc - H + t or
[0013] or Pk.z < 0 or Pk.z > L; x and t y are the truncation thresholds in the X and Y directions respectively, both of which are non-negative values and are set according to the environment.
[0014] As an improvement, the two-way two-dimensionalization means that for each pixel point (u, v) in the depth image and the corresponding point P in the camera coordinate system, (u, v, P.y) and (u, v, P.z) are taken respectively to obtain two two-dimensional images with only height information and only distance information;
[0015] The binarization can adopt the local threshold method. Assuming that the current pixel point coordinate is (u, v), the neighborhood centered on this point is r*r, and g(u, v) represents the gray value at (u, v). Then the threshold calculation formula for the pixel point (u, v) is:
[0016]
[0017] where:
[0018]
[0019] R is the dynamic range of the standard deviation. For 8-bit grayscale images, R = 128; for 10-bit grayscale images, R = 512;
[0020] For the contour detection, a strategy of only detecting the outermost contour is adopted, and only the corner coordinates of the circumscribed rectangle of the contour are stored.
[0021] As an improvement, the sliding window in step three has two parameters: window size and sliding step; among them, the sliding step is smaller than the window size.
[0022] As an improvement, the process of finding the extreme value in step three is as follows: The size of the two-dimensional depth image is W×H, and the sliding window is (i*s, 0), ((i + 1)*s, H)
[0023] , where the initial value of i is 1, s is the sliding step length, and the minimum value point (u min , v min ) of v within this window is obtained. If (i + 1)*s < W, the window is slid to the right, that is, i is incremented by 1, and the minimum value point (u′ min , v′ min ) of v within the window is calculated again. If u′ min - u′ min < s, it is considered that the extreme values of the two windows come from the same problem-wrapped object. Therefore, only the smaller value and window are retained; continue to slide until (i + 1)*s ≥ W. At this time, the minimum and maximum values in the two-dimensional depth image can be calculated. By subtracting these extreme values from the minimum or maximum value parameter in the set packaging shell parameters, it can be determined whether these points are qualified, and the specific abnormal numerical range can be given.
[0024] As an improvement, the sliding window described in step three has two parameters: window size and sliding step length; among them, the sliding step length is less than the window size.
[0025] As an improvement, in step three, in order to gradually find the optimal solution, the sliding window size can be set to 20, and the sliding step length that can be set can be set to 10.
[0026] As an improvement, in step one, the camera pre-calibrates the camera initialization parameters using a calibration board, specifically including the forward distortion, tangential distortion of each camera, and the rotation matrix and translation matrix between the two cameras when a binocular camera is required; after the camera is installed, the external parameters of the camera are calibrated using position punctuation or conveyor belt corner points, that is, the translation matrix and rotation matrix between the required binocular camera and the conveyor belt, obtaining the reference distance and vector matrix parameters in the spatial dimension, and presetting the data of length L, height H, and width W according to different detection object commodity packaging categories.
[0027] As an improvement, the depth stereo camera is composed of a group of binocular cameras that are two cameras on the same straight line. One group has a larger field of view and is used to cover the detection range of the near distance, and the other group has a smaller field of view and is used for the detection of the far distance. The bottom of the package at a long distance and the top of the package at a short distance are detected to judge the overall compliance of the package.
[0028] As an improvement, the binocular cameras use hardware synchronization.
[0029] As an improvement, the depth stereo camera can be detachably connected to the package through a rigid connecting piece and shoots in the direction of the conveyor belt.
[0030] The advantages of the present invention compared with the prior art are as follows: The present invention obtains the three-dimensional structure information of the package through a depth stereo camera, improving the accuracy and reliability of detection. The present invention adopts a maximum value algorithm based on a sliding window, which can accurately detect the positions where the package structure size is unqualified and the specific overlimit values. The depth stereo camera of the present invention has flexible configuration and can be adjusted and optimized according to actual needs. The present invention realizes automatic detection, greatly improving the detection efficiency and reducing the labor cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flowchart for processing depth images and calculating overlimit values in the implementation of the present invention;
[0032] Figure 2 It is a schematic diagram of the system in the embodiment of the present invention;
[0033] Figure 3 It is a shooting schematic diagram of the depth stereo camera in the embodiment of the present invention;
[0034] Figure 4 It is a schematic diagram of system assembly in the embodiment of the present invention
[0035] Figure 5 It is a structural depth image collected by the present invention for edge-arc-shaped articles;
[0036] Figure 6 It is a structural depth image collected by the present invention for cubic articles;
[0037] Figure 7 It is a structural depth image collected by the present invention for stacked square articles. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following further describes the present invention in detail with reference to the drawings.
[0039] When the present invention is specifically implemented,
[0040] 1. Detection object: Packaging box.
[0041] 2. Specification size: 280mm * 150mm.
[0042] 3. Defect types: Material clamping, large wrinkles, cross wrinkles, misalignment, false sealing;
[0043] 4. Detection results
[0044] 4.1. Classify alarms and display defect pictures in real time;
[0045] 4.2. Classify and independently reject defective products;
[0046] 4.3. Save defect data for 24 months.
[0047] 5. The following methods are adopted in the detection scheme:
[0048] 5.1. Adopt the combined light source imaging mode to basically eliminate the shadow interference caused by the reflection of the thin film and improve the image quality.
[0049] 5.2. Adopt the polarized light method to eliminate the interference caused by three-dimensionality to the greatest extent.
[0050] 5.3. Adopt the reflection method and use a white conveyor belt. There are obvious color difference defects in the product.
[0051] 5.4. Adopt the feature analysis + deep learning method to realize defect detection in various scenarios.
[0052] The working principle of the present invention: I. Installation and calibration of the depth stereo camera Installation: The depth stereo camera is fixedly installed at the top position of the source of the packaging conveyor belt, and the shooting direction is vertically downward to ensure that the field of view of the camera can cover the entire space of the object to be detected and both sides of the conveyor belt.
[0053] The depth stereo camera is composed of a group of binocular cameras, two binocular cameras in a straight line. One group has a larger field of view and is used to cover the detection range at close range, and the other group has a smaller field of view and is used for distant detection; detect the bottom of the packaging at a long distance and the top of the packaging at a short distance to judge the overall compliance of the packaging.
[0054] Calibration: Use a calibration board to calibrate the initial parameters of the camera. These parameters include the forward distortion and tangential distortion of each camera, as well as the rotation matrix and translation matrix between the two cameras when using binocular cameras.
[0055] After the camera is installed, use position punctuation or conveyor belt corner points to calibrate the external parameters of the camera to obtain the translation matrix and rotation matrix between the binocular camera and the conveyor belt, as well as the reference distance and vector matrix parameters in the spatial dimension; and preset the data of length L, height H, and width W according to different types of detected object commodity packages. These parameters provide a basis for subsequent three-dimensional reconstruction and dimension measurement.
[0056] II. Acquisition and preprocessing of depth images Acquisition of depth images: The depth stereo camera emits laser light through a laser emitter and uses the camera to collect the reflected light. After algorithm analysis, the 3D contour data of the target object is obtained, so as to obtain the depth image of the packaging structure to be detected.
[0057] Preprocessing: Truncation: Remove the noise and outliers in the depth image.
[0058] The contour model of the package includes length L, height H, width W, depth camera mounting height Hc, ultra-high setting value Hm, and ultra-wide setting value Wm. Since the parameters of the package shell are fixed and directly input during system deployment, when it is detected that the length of the package shell exceeds the set value, for the three-dimensional point cloud set obtained from the depth map:
[0059] {P1, P2, P3,...Pn};
[0060] For any point Pk in
[0061] if it satisfies the condition: x Pk.x < Hm - t x or Pk.x > Hc - H + t or
[0062] or Pk.z < 0 or Pk.z > L; x and t y are the truncation thresholds in the X and Y directions respectively, both non-negative values, set according to the environment.
[0063] Two-way two-dimensionalization: That is, process the three-dimensional feature information in two dimensional directions (height and distance information). For each pixel point (u, v) in the depth image and the corresponding camera coordinate point P, take (u, v, P.y) and (u, v, P.z) respectively, and convert them into two-dimensional images of height information and distance information, obtaining two two-dimensional images respectively reflecting height and distance information.
[0064] Filtering: Smooth the image to reduce the interference of noise.
[0065] Binarization: Use the local threshold method to convert the grayscale image into a binary image for subsequent contour detection. This processing is mainly for the depth image in grayscale form.
[0066] Assume that the current pixel point coordinate is (u, v), the neighborhood centered on this point is r*r, and g(u, v) represents the grayscale value at (u, v). Then the threshold calculation formula for the pixel point (u, v) is:
[0067]
[0068] Where:
[0069]
[0070] R is the dynamic range of the standard deviation. For 8-bit grayscale images, R = 128; for 10-bit grayscale images, R = 512. The r parameter in the r*r parameter of the detection area is a preset detection parameter according to the appearance characteristics of different detected items.
[0071] Contour detection: For the edge points of the packaging image, adopt the strategy of only detecting the outermost contour, and store the corner coordinates of the circumscribed rectangle of the contour to provide a basis for subsequent dimension measurement and compliance judgment. Traverse the coordinates of the edge contour, and use the bubble algorithm to select the maximum value as the corner coordinates. The specified size is a preset parameter
[0072] III. Maximum value algorithm based on sliding window
[0073] Sliding window setting: Set a sliding window that slides on the depth image to find the extreme points in the image. The size and step of the sliding window can be adjusted according to the actual situation.
[0074] Finding the maximum value: Find the minimum value point within the sliding window, and then as the window slides, compare the minimum value points in adjacent windows. If it is considered that the extreme value sources of the two windows are the same problem packaging object, only retain the smaller value and window. Continue to slide until the entire depth image is traversed.
[0075] The size of the two-dimensional depth image is W×H, and the sliding window is (i*s, 0), ((i + 1)*s, H), where the initial value of i is 1, s is the sliding step, and find the minimum value point (u min , v min ) of v within this window. If (i + 1)*s < W, slide the window to the right, that is, increment i by 1, and calculate the minimum value point (u′ min , v′ min ) of w within the window again. If u′ min -u min < s, that is, it is considered that the extreme value sources of the two windows are the same problem packaging object. Therefore, only retain the smaller value and window; continue to slide until (i + 1)*s ≥ W. At this time, the minimum and maximum values in the two-dimensional depth image can be calculated. Subtract the minimum or maximum value parameter in the set packaging shell parameters from these extreme values to obtain whether these points are qualified and give the specific abnormal value range.
[0076] To gradually find the optimal solution, the sliding window size can be set to 20, and the adjustable sliding step can be set to 10.
[0077] Mapping back to world coordinates: Map the two-dimensional coordinates of the found extreme values back to the world coordinates within the depth map to obtain the possible positions where the packaging structure size is unqualified and the specific overlimit values.
[0078] IV. Compliance judgment
[0079] Compare the measured packaging structure size with the preset compliance standard to judge whether the packaging is compliant. If the size exceeds the preset range, mark it as non-compliant and give the specific abnormal value range.
[0080] In summary, the compliance detection method based on deep stereo vision obtains the three-dimensional structure information of the package through a deep stereo camera, and uses image processing technology to accurately measure the size and shape of the package and judge its compliance. This method realizes automatic detection, greatly improves the detection efficiency, reduces the labor cost, and at the same time improves the accuracy and reliability of the detection.
[0081] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0082] In the present invention, unless otherwise clearly defined and limited, terms such as "mounted", "connected", "connected to", "fixed" and the like should be construed in a broad sense. For example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0083] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through other features between them. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature.
[0084] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0085] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.
Claims
1. A compliance detection method based on depth stereo vision, characterized in that, The method includes the following steps: Step 1: First, install a set of depth stereo cameras at the top of the source of the packaging conveyor belt with the shooting direction vertically downward, ensuring that the field of view can cover the entire space of the object to be detected and both sides of the conveyor belt; Step 2: After the depth stereo camera obtains the depth image of the packaging structure to be detected, perform preprocessing operations on the depth information. The specific operations are as follows: truncation, bidirectional two-dimensionalization, filtering, binarization, and contour detection, and then the contour model of the packaging can be obtained; Step 3: Perform a maximum value algorithm based on a sliding window on the above contour model to obtain the contour extreme values within the threshold, and then map the two-dimensional coordinates of the extreme values back to the world coordinates in the depth map, and then the possible positions and specific overlimit values of the unqualified packaging structure size can be given; Step 4: If an overlimit value is detected in the packaging appearance, slide the camera to obtain the overlimit values detected at different positions.
2. The compliance detection method based on depth stereo vision according to claim 1, wherein: In Step 2, the contour model of the packaging includes length L, height H, width W, the installation height Hc of the depth camera, the overheight setting value Hm, and the overwidth setting value Wm; since the parameters of the packaging shell are fixed, they are directly input during system deployment. When it is detected that the length of the packaging shell exceeds the set value, for the three-dimensional point cloud set obtained from the depth map: {P1, P2, P3,...Pn}; For any point Pk, if it satisfies the condition: Pk.x < Hm - t x or Pk.x > Hc - H + t x or or Pk.z < 0 or Pk.z > L; Then set Pk.x = Pk.y = Pk.z = 0, where t x and t y are the truncation thresholds in the X and Y directions respectively, both of which are non-negative values and are set according to the environment.
3. The compliance detection method based on depth stereo vision according to claim 2, wherein: The bidirectional two-dimensionalization means that for each pixel point (u, v) in the depth image and its corresponding point P in the camera coordinate system, (u, v, P.y) and (u, v, P.z) are taken respectively to obtain two two-dimensional images with only height information and only distance information; The binarization uses the local threshold method. Assume that the current pixel point coordinate is (u, v), the neighborhood centered on this point is r*r, and g(u, v) represents the gray value at (u, v). Then the threshold calculation formula for the pixel point (u, v) is: Where: R is the dynamic range of the standard deviation. For 8-bit grayscale images, R = 128; for 10-bit grayscale images, R = 512; For the contour detection, a strategy of only detecting the outermost contour is adopted, and only the corner coordinates of the circumscribed rectangle of the contour are stored.
4. The compliance detection method based on depth stereo vision according to claim 1, wherein: The process of finding the extreme value in Step 3 is as follows: The size of the two-dimensional depth image is W×H, and the sliding window is (i*s, 0), ((i + 1)*s, H), where the initial value of i is 1 and s is the sliding step. Find the minimum value point (u min , v min ) of v within this window. If (i + 1)*s < W, slide the window to the right, that is, increment i by 1, and calculate the minimum value point (u′ min , v′ min ) of v within the window again. If u′ min -u min < s, it is considered that the extreme value sources of the two windows are the same problem-wrapped object. Therefore, only keep the smaller value and window; continue sliding until (i + 1)*s ≥ W. At this time, the minimum and maximum values in the two-dimensional depth image can be calculated. Subtract these extreme values from the minimum or maximum value parameter in the set packaging shell parameters to obtain whether these points are qualified and give the specific abnormal numerical range.
5. The compliance detection method based on depth stereo vision according to claim 4, wherein: In the sliding window in Step 3, there are two parameters: window size and sliding step; among them, the sliding step is smaller than the window size.
6. The compliance detection method based on depth stereo vision according to claim 5, wherein: In Step 3, in order to gradually find the optimal solution, the sliding window size can be set to 20, and the sliding step that can be set can be set to 10.
7. The compliance detection method based on depth stereo vision according to claim 1, wherein: In Step 1, the camera is pre-calibrated with a calibration board to initialize the camera parameters. After the camera is installed, the external parameters of the camera are calibrated using position punctuation or conveyor belt corner points to obtain the reference distance and vector matrix parameters in the spatial dimension; and according to different types of commodity packaging of the detection object, the data of length L, height H, and width W are preset.
8. The compliance detection method based on depth stereo vision according to claim 1, wherein: The depth stereo camera is composed of a set of binocular cameras that are two cameras on the same straight line. One set has a larger field of view and is used to cover the detection range at close range, and the other set has a smaller field of view and is used for detection at a distance; the bottom of the packaging at a long distance and the top of the packaging at a short distance are detected to judge the overall compliance of the packaging.
9. The compliance detection method based on depth stereo vision according to claim 8, wherein: The binocular cameras use hardware synchronization.
10. The compliance detection method based on depth stereo vision according to claim 7, wherein: The depth stereo camera can be detachably connected to the package through a rigid connecting member and shoots in the direction of the conveyor belt.