Lightweight dynamic volume measurement method and system

Through the end-to-end processing method of RGB cameras and deep cameras combined with deep learning, the problem of insufficient measurement accuracy under heavy hardware and complex environments in the prior art is solved, and lightweight and high-precision package volume measurement is achieved, which is suitable for modern logistics automation.

CN120580447AActive Publication Date: 2025-09-02GUANGZHOU GENYE INFORMATION TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511025230.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-09-02
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

The existing dynamic volume measurement technology cannot meet industrial needs in complex environments such as high speed, reflective, and low light. The hardware is heavy and the deployment cost is high. The instance segmentation capability is insufficient in multi-target adhesion and stacking.

Method used

The package image is obtained simultaneously by using RGB cameras and depth cameras. Through the end-to-end processing method of deep learning, the same package is fragmented into one class with a clustering algorithm, and the package segmentation is performed and the volume is calculated, reducing the number of sensors and optimizing the algorithm structure.

Benefits of technology

It realizes high-precision, low-cost and easy-to-maintenance package volume measurement in complex environments, reduces system costs and deployment difficulties, and meets the lightweight and precise needs of modern logistics automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580447A_ABST
    Figure CN120580447A_ABST
Patent Text Reader

Abstract

The invention aims to provide a lightweight dynamic volume measurement method and system. The method comprises the following steps: acquiring an RGB image and a depth image of a package; the depth map is preprocessed; inputting the RGB image and the processed depth image into a segmentation model; the segmentation model outputs parcel information; and calculating the package volume according to the package information. By reducing the number of sensors and optimizing the algorithm structure, real-time and high-precision volume measurement of a high-speed moving parcel is realized; meanwhile, weak light noise suppression and segmentation robustness in a multi-target adhesion and stacking scene are considered, so that the overall system cost and deployment difficulty are remarkably reduced, and the technical requirements of modern logistics automation for portability, precision, low cost and easy maintenance are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a lightweight dynamic volume measurement method and system. Background Art

[0002] Existing dynamic volume measurement technologies primarily rely on the following solutions, but all have significant shortcomings, making it difficult to achieve a balanced balance of real-time performance, accuracy, and cost control: Line laser scanning. A single laser beam projects a contour line onto the surface of an object. Height information is calculated based on the position offset of the laser line on the imaging plane. This is then combined with an encoder or multi-frame stitching to form a point cloud and calculate volume. This solution is easy to deploy in small boxes or embedded environments and offers a certain balance between speed and accuracy. However, single-line scanning inherently only produces sparse, strip-like point clouds. It struggles to capture surface irregularities, small protrusions, or large, highly reflective areas, leading to data discontinuities or blind spots. In strong or highly reflective environments, laser signal attenuation leads to increased noise, compromising depth calculations. Multi-camera arrays and structured light / ToF systems achieve higher measurement accuracy by using multiple cameras or structured light devices to scan from different angles. However, these systems require large hardware footprints and are complex to calibrate and maintain. In high-speed conveyor belt environments, synchronizing multiple cameras is challenging, and alignment failures due to motion blur or posture changes can easily cause measurement errors. Single depth cameras (ToF / binocular) offer small size and manageable costs, but depth data can be prone to jitter and holes in parallax blind spots, reflective / low-light conditions, or motion blur, making them difficult to meet the reliability requirements of industrial scenarios. In particular, insufficient RGB information in low-texture or repetitive-texture environments can lead to segmentation and matching failures, resulting in increased volume calculation errors. Multimodal sensor fusion, which integrates multiple data sources—infrared, laser, RGB, and depth—can improve measurement robustness and has been applied in intelligent sorting and refined warehousing. However, the error sources of each sensor vary significantly, the fusion algorithm is complex, and real-time performance is difficult to guarantee. Hardware combinations are still heavily weighted, requiring a balance between cost and deployment difficulty. In summary, existing technologies either suffer from heavy hardware and high deployment costs, or their measurement accuracy and stability fail to meet industrial requirements in complex environments such as high speed, reflective, and low-light conditions. Instance segmentation capabilities are also significantly insufficient in situations where multiple objects are clumped or stacked. To achieve a truly balanced balance of lightweight and high precision, breakthroughs in algorithm and hardware design are essential, avoiding simply stacking multiple sensors or blindly relying on a single device. Summary of the Invention

[0003] The purpose of the present invention is to provide a lightweight dynamic volume measurement method and system. This method achieves real-time, high-precision volume measurement of high-speed moving packages by reducing the number of sensors and optimizing the algorithm structure. At the same time, it takes into account the suppression of weak light noise and multi-target adhesion and segmentation robustness in stacking scenarios, so as to significantly reduce the overall system cost and deployment difficulty, and meet the technical requirements of modern logistics automation for "lightweight, precise, low-cost, and easy maintenance."

[0004] A lightweight dynamic volume measurement method, comprising: Get the RGB image and depth image of the package; Preprocessing the depth map; Input the RGB image and the processed depth image into a segmentation model; The segmentation model outputs package information; The package volume is calculated according to the package information.

[0005] Preferably, the preprocessing of the depth map includes: Placing a checkerboard calibration plate on an unloaded synchronous belt plane, and calculating a rotation matrix between the synchronous belt plane and the camera lens plane based on the RGB image; Correcting the entire synchronous belt XY plane to be parallel to the camera lens XY plane according to the rotation matrix between the synchronous belt plane and the camera lens plane; Select the region of interest in the depth map; performing noise filtering and image enhancement on the region of interest; The average depth of the synchronous belt plane is calculated and used as the reference plane.

[0006] Preferably, inputting the RGB image and the processed depth image into a segmentation model comprises: Get the feature information of RGB image and depth image; The characteristic information is classified to obtain package fragments: Merging the package fragments to obtain single package information; Verify the accuracy of the stated package information.

[0007] Preferably, calculating the package volume according to the package information includes: The package information includes: the xy position, width, height and segmentation mask of the detection box of each package; Get the height information of each package under the corresponding area on the image, and integrate it to obtain the actual package volume data.

[0008] Preferably, acquiring feature information of the RGB image and the depth image includes: An adaptive feature extraction network is constructed to extract feature information of RGB images and depth images.

[0009] Preferably, classifying the characteristic information to obtain package fragments includes: Grayscale classification is performed based on the grayscale value of the image, and areas with similar grayscale values ​​are classified as one category; Feature classification is performed based on the feature information of the RGB image and the depth image. On the premise of satisfying the grayscale classification, areas with similar features are classified as one category.

[0010] Preferably, the step of fusing the package fragments to obtain single package information comprises: performing basic fusion on the package fragments; performing detail fusion on the package fragments; The detail fusion result and the basic fusion result are combined to obtain a single package.

[0011] Preferably, verifying the accuracy of the package information includes: Construct model loss function; Calculating the difference between the package and the true model; If the difference between the wrapped and true models is greater than the preset value, the adjustable weights, regularization coefficients, and mixing ratios of the loss function are adjusted; Repeat the verification until the difference between the package and the real model meets the preset value.

[0012] A lightweight dynamic volume measurement system, comprising: Get the RGB image and depth image of the package; Preprocessing the depth map; Input the RGB image and the processed depth image into a segmentation model; The segmentation model outputs package information; The package volume is calculated according to the package information.

[0013] An electronic device includes: a chip, a processor and a memory, wherein the memory is used to store computer program code, and the computer program code includes computer instructions. When the chip executes the computer instructions, the electronic device performs a lightweight dynamic volume measurement method.

[0014] The beneficial effects of the present invention are as follows: the present invention simultaneously acquires package images through an RGB camera and a depth camera, thereby sampling more comprehensive package information. Then, during package segmentation of the overall image, a clustering algorithm is used to group fragments of packages belonging to the same entity into one category. Multiple package fragments are then pieced together to restore the package. This overall end-to-end process uses the images acquired by the camera directly as input to the model, which ultimately directly outputs the segmented package images. Volume calculation can then be performed directly on the images. The present invention utilizes an end-to-end processing method based on deep learning to more efficiently and conveniently acquire package information and calculate package volume. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0017] Figure 1 This is a flow chart of a lightweight dynamic volume measurement method of the present invention; Figure 2 A flowchart for preprocessing the depth map of the present invention; Figure 3 FIG. 1 is a flowchart of the package segmentation process of the present invention; Figure 4 The figure is a schematic diagram of the hardware structure of an electronic device of the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0019] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0020] In addition, the descriptions of "first", "second", etc. in the present invention are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0021] Existing technologies for image processing of logistics packages face numerous challenges. These include cumbersome hardware and high deployment costs, as well as measurement accuracy and stability that fail to meet industrial requirements in complex environments such as high speeds, reflective surfaces, and low light conditions. Instance segmentation capabilities are also significantly insufficient in situations where multiple objects are adhered to or stacked. To achieve a true balance between lightweight and high precision, breakthroughs in algorithm and hardware design are essential, avoiding the simplistic stacking of multiple sensors or blind reliance on a single device.

[0022] The present invention simultaneously acquires package images through an RGB camera and a depth camera, sampling more comprehensive package information. Then, during package segmentation of the overall image, a clustering algorithm is used to group fragments of the same package into a single category. Multiple package fragments are then pieced together to restore the package. This overall end-to-end process uses the camera images directly as input to the model, which ultimately outputs the segmented package images. Volume calculation is then performed directly on the images. The present invention utilizes an end-to-end processing method based on deep learning to more efficiently and conveniently acquire package information and calculate package volume.

[0023] Example 1 A lightweight dynamic volume measurement method, reference Figure 1 ,include: S100, obtaining the RGB image and depth image of the package; In this embodiment of the present invention, before taking a photo, the RGB image and depth map need to be aligned and calibrated. Methods that can be used include camera calibration. The RGB image captures the visible light information of a scene through the three color channels of red (R), green (G), and blue (B), presenting the appearance characteristics of an object, such as color, texture, and illumination, reflecting the two-dimensional visual information seen by the human eye. The depth map, on the other hand, records the absolute or relative distance of an object from the camera through the grayscale value or distance value of each pixel, expressing the three-dimensional geometry of the scene. While lacking color details, it contains spatial positional relationships. The combination of the two provides both appearance and geometry information, and is widely used in fields such as 3D reconstruction and augmented reality.

[0024] S200, preprocessing the depth map; Depth map preprocessing uses a series of algorithms (such as filtering, hole filling, and normalization) to remove noise, outliers, and missing data from the original depth map while enhancing valid information, thereby improving the accuracy and consistency of the depth data. This helps improve the robustness and accuracy of subsequent tasks (such as 3D reconstruction, object detection, or pose estimation) and ensures better alignment and fusion of depth information with other modal data (such as RGB images).

[0025] S300, inputting the RGB image and the processed depth image into the segmentation model; Segmentation models divide processed images into regions or objects with specific semantic or structural significance, thereby identifying and locating key targets in the image (such as objects, human bodies, or parts of scenes). Through pixel-level or instance-level classification, they provide more refined semantic understanding for computer vision tasks (such as target recognition, medical image analysis, and autonomous driving), and support subsequent processing such as object measurement, behavior analysis, and 3D reconstruction.

[0026] S400, the segmentation model outputs package information; In the embodiment of the present invention, the package information includes package logistics information such as package type and package size.

[0027] S500: Calculate the package volume according to the package information.

[0028] This method calculates the volume of a package based on segmented images. Using a depth camera or RGB-D sensor to obtain a depth map of the package, the volume is calculated using integration or voxelization. This method is widely applicable to automated dimensional measurement in logistics, warehousing, and other fields.

[0029] Preferably, reference Figure 2 , S200, pre-processing the depth map includes: S210, placing a checkerboard calibration plate on an unloaded synchronous belt plane, and calculating a rotation matrix between the synchronous belt plane and the camera lens plane based on the RGB image; Preprocess the depth map captured by the stereo camera. Due to the actual installation of the stereo camera, there will be a certain degree of deviation in the horizontal tilt angle. The camera lens plane and the timing belt plane cannot be completely parallel. Therefore, when there is no package, the depth map captured by the stereo camera will show the unloaded timing belt plane at a certain angle (i.e., it is not horizontal).

[0030] S220, correcting the entire synchronous belt XY plane to be parallel to the camera lens XY plane according to the rotation matrix between the synchronous belt plane and the camera lens plane; Use a checkerboard calibration plate and place it on the unloaded synchronous belt plane. Use the camera to obtain RGB images and calculate the rotation matrix between the synchronous belt plane and the camera lens plane. Using the calculated relationship matrix, calibrate the entire synchronous belt XY plane to be parallel to the camera lens XY plane.

[0031] S230, selecting a region of interest in the depth map; Select ROI (Region of Interest). A fixed-size rectangular area is selected at the center of the depth map (for example, the middle 30% × 30% or 50% × 50% of the image resolution, depending on the actual camera resolution and plane ratio). The center is selected because the surrounding area may have edge clipping, large distortion, or physical barriers that affect the depth value. These are easy to cause misjudgment.

[0032] S240, performing noise filtering and image enhancement on the region of interest; Noise filtering is performed on this ROI, and median filtering (3×3 or 5×5) is performed on the depth value to remove occasional scattered points or depth jumps.

[0033] S250: Calculate the average depth of the synchronous belt plane and use the average depth as the reference plane.

[0034] Then calculate the average depth of this synchronous belt plane and use this average depth as the reference plane. Simply subtract this reference value from the depth value of the surface above the object to get the height of any point on the object surface from this plane.

[0035] Preferably, reference Figure 3 , S300, inputting the RGB image and the processed depth image into the segmentation model includes: S310, acquiring feature information of the RGB image and the depth image; RGB images primarily contain appearance features such as color, texture, lighting, and edges. Combining pixels from the red, green, and blue channels, they present an object's surface details and color distribution, making them suitable for tasks such as target recognition and scene understanding. Depth maps, on the other hand, record the distance from the object to the camera and represent the scene's 3D geometry using grayscale or depth values. These images contain depth features such as the object's spatial position, shape, and outline, and can be used for applications such as 3D reconstruction, obstacle avoidance, and dimensional measurement. The combination of these two provides rich visual semantics and geometric structure information, enhancing environmental perception.

[0036] S320, classify the feature information to obtain package fragments: In this embodiment of the present invention, information about logistics packages is primarily extracted through a feature extraction network. This information primarily includes appearance features (such as color, texture, label text, barcode / QR code), geometric features (such as size, shape, volume, and edge contours), material features (such as hardness and reflective properties), and spatial position (such as placement angle and position). This information can be obtained through RGB imagery to obtain surface visual data, and through depth maps or 3D point cloud computing to calculate size and volume. Combining this multimodal data enables automated sorting, volume measurement, damage detection, and classification optimization of packages, thereby improving logistics efficiency and intelligence.

[0037] S330, merging the package fragments to obtain single package information; In an embodiment of the present invention, feature information of packages is extracted, the feature information of multiple packages is classified, fragments belonging to the same package are grouped together, and then the complete package information is obtained by fusing the multiple package fragments. Compared with the method of restoring the entire three-dimensional image and then performing package segmentation, the processing steps of the present invention are simpler, the amount of data processed is smaller, and it is more effective in the case of stacked packages. In the case of stacked packages, the clustering algorithm of the present invention has the advantage of being faster. It only needs to determine which fragments belong to which package, and then restore the information of the entire package based on the information of some fragments.

[0038] S340, verify the accuracy of the package information.

[0039] The loss function is closely related to prediction accuracy. As the optimization objective during model training, it directly guides how the model adjusts parameters to reduce prediction error. Smaller loss function values ​​generally indicate smaller deviations between the model's predictions and the real data, resulting in higher prediction accuracy. Conversely, larger loss values ​​indicate poor model performance. However, the choice of loss function must be tailored to the task (for example, cross entropy is commonly used for classification tasks, while mean squared error is commonly used for regression tasks). An inappropriate loss function can cause the model optimization direction to deviate from actual requirements. Even if the loss value is reduced, overfitting or underfitting may occur, which in turn affects generalization accuracy.

[0040] In image processing, loss functions quantify the difference between a model's predicted output (such as classification labels, segmentation masks, generated images, etc.) and the true value, providing a clear adjustment direction for optimization algorithms (such as gradient descent). By minimizing the loss function, the model can gradually adjust its parameters to improve task performance (such as enhancing image quality, improving segmentation accuracy, or improving reconstruction). Common loss functions include mean squared error, cross entropy, and perceptual loss, and their design directly affects the model's convergence speed and learning effectiveness.

[0041] Preferably, in S500, calculating the package volume according to the package information includes: Package information includes: the xy position, width, height, and segmentation mask of each package's detection box; Get the height information of each package under the corresponding area on the image, and integrate it to obtain the actual package volume data.

[0042] The package information detected by the model includes the xy position, width, height, and segmentation mask of each package's detection box. The model then combines this with the depth map to obtain the height of each package within its corresponding area on the image, which is then integrated to obtain the actual volume data.

[0043] Preferably, S320, obtaining feature information of the RGB image and the depth image includes: Construct an adaptive feature extraction network to extract the feature information of RGB image and depth image, which is expressed as: ; Among them, Q is the query vector of the input sequence transformation mapping, K is the key vector of the input sequence transformation mapping, and V is the value vector of the input sequence transformation mapping. is the softmax function, is the similarity between vector Q and vector K.

[0044] In feature extraction, the Softmax function serves as the output layer activation function for classification tasks, converting the extracted feature vectors (such as the output of a fully connected layer) into a probability distribution, thereby clarifying the prediction confidence for each category. However, in models based on attention mechanisms or feature weights (such as the visual Transformer), Softmax can highlight important features and suppress irrelevant information by normalizing the correlation scores between features (such as the query-key similarity in self-attention), indirectly optimizing the spatial or channel weight distribution of features and enhancing the model's focus on key features. Its essence is to improve the discriminability or interpretability of features through probabilistic mapping.

[0045] In an embodiment of the present invention, feature information about the package is extracted from the RGB image and the depth image through the feature extraction network of the segmentation model, wherein the extracted feature information includes: grayscale value, texture, text, and other features that can identify the specificity of the package. Each individual package can be identified by these specific features. In a logistics and transportation scenario, there are many packages stacked together. During the image acquisition process, it is difficult to obtain a complete image of each package. Therefore, it is necessary to collect these specific features to prepare for specific identification of each package.

[0046] Preferably, S320, classifying the characteristic information to obtain package fragments includes: Grayscale classification is performed based on the grayscale value of the image, and areas with similar grayscale values ​​are classified as one category. The grayscale classification process is expressed as: ; in, is the maximum grayscale value within the segmentation area, is the minimum grayscale value within the segmentation area, is the average grayscale value within the segmented area, is the splitting threshold, is the merging threshold; According to the grayscale value classification, the brightness difference of pixels or regions is directly used for simple threshold segmentation or clustering. It is suitable for distinguishing targets with obvious light and dark contrast (such as separating objects in black and white images) and is computationally efficient.

[0047] Feature classification is performed based on the feature information of the RGB image and the depth image. On the premise of satisfying the grayscale classification, regions with similar features are classified as one category. The feature classification process is expressed as follows: ; in, is the feature center corresponding to the kth class, is the i-th eigenvector of the image feature map, is a non-normalized relational function, For the conversion network, the network structure is 1×1 convolution layer-BN normalization layer-relu activation layer. To convert the network.

[0048] Feature-based classification relies on higher-dimensional semantic or structural features (such as texture, shape, depth, or abstract features from deep learning). It can handle complex scenarios (such as medical image analysis and object recognition) and is more robust. The former focuses on low-level physical attributes, while the latter relies on high-level semantic understanding. Combining the two offers a balanced approach to efficiency and accuracy.

[0049] In this embodiment of the present invention, packages that look similar but are incompletely photographed can be classified using specific features. For complete package images, their volume information can be directly calculated. For obstructed package images, it is necessary to classify the exposed fragments that can be obtained, and then restore the entire package using the fragments before calculating its volume information.

[0050] Preferably, S340, merging the package fragments to obtain single package information includes: Perform basic fusion on the package fragments, expressed as: ; in, is the initial fusion result, and is an adjustable weight, Represents the coordinate value of the pixel; The basic information of the wrapped fragments mainly represents the contour information of the fused image, so in order to effectively retain the main energy information of the source image.

[0051] Perform detail fusion on the package fragments to obtain the detail fusion result, which is expressed as: ; Among them, R() represents the feature extraction network.

[0052] The detail layer primarily contains the details and texture information of the source image. Processing of the detail layer directly affects the clarity of the fused image and the severity of edge distortion. In the above formula, there are four feature extraction layers. Different fusion details are obtained in different feature extraction layers, and ultimately the fusion details of each layer are combined to form the detail fusion result.

[0053] The detail fusion results and the basic fusion results are combined to obtain a single package.

[0054] Optionally, the fragment fusion method can also use a feature matching-based method (such as SIFT and SURF to extract key points and eliminate mismatches through RANSAC to achieve geometric alignment) or an optimization algorithm-based method (such as minimizing color differences or structural differences at the splicing boundaries and eliminating seams through Poisson fusion or Laplace pyramid blending), and finally output a coherent complete image.

[0055] In this embodiment of the present invention, the fragments of a package can be of any shape. The segmentation model can restore and reshape the entire package based on the fragments' spatial location, geometric characteristics, and degree of stacking, thereby accurately calculating the volume of each package. The present method requires minimal data processing and computational effort. Accurate package volume information can be obtained through a simple three-step process of feature extraction, feature classification, and feature fusion, significantly improving detection efficiency and production efficiency.

[0056] Preferably, S350, verifying the accuracy of the package information includes: Construct the model loss function, expressed as: ; in, is the predicted value of the i-th package, is the true value of the i-th package, which is valid only when the predicted value is in the target area, k is the total number of packages, is the regularization coefficient, is the regularized mixing ratio, is the Lasso regularization term, is the ridge regularization term, is an adjustable weight; Calculate the difference between the package and the true model; If the difference between the wrapped and true models is greater than the preset value, the adjustable weights, regularization coefficients, and mixing ratios of the loss function are adjusted; Adjusting loss function parameters (such as weight coefficients, margins, or temperature coefficients) is primarily used to balance the contributions of different tasks or classes, thereby optimizing the model's training dynamics. For example, when there is class imbalance, the loss weight of the minority class can be increased to prevent the model from favoring the majority class; in multi-task learning, the ratio of the losses of each task can be adjusted to coordinate the convergence rate; or the margin value can be adjusted to control the strictness of feature separation. Parameter adjustments directly affect the direction of gradient updates and can alleviate overfitting, improve generalization, or focus the model on a specific objective (such as precision rather than recall), ultimately improving the model's performance on test data.

[0057] Repeat the verification until the difference between the package and the real model meets the preset value.

[0058] Example 2 A lightweight dynamic volume measurement system, comprising: Get the RGB image and depth image of the package; Preprocess the depth map; Input the RGB image and the processed depth map into the segmentation model; The segmentation model outputs package information; Calculate the package volume based on the package information.

[0059] Example 3 An electronic device includes: a chip, a processor and a memory, the memory is used to store computer program code, the computer program code includes computer instructions, and when the chip executes the computer instructions, the electronic device performs a lightweight dynamic volume measurement method.

[0060] refer to Figure 4 The electronic device 2 includes a processor 21, a memory 22, an input device 23, and an output device 24. The processor 21, the memory 22, the input device 23, and the output device 24 are coupled via a connector, which may include various interfaces, transmission lines, or buses, etc., but this is not limited in the present embodiment. It should be understood that in various embodiments of the present invention, coupling refers to mutual connection in a specific manner, including direct connection or indirect connection through other devices, such as connection via various interfaces, transmission lines, buses, etc.

[0061] The processor 21 may be one or more graphics processing units (GPUs). If the processor 21 is a GPU, the GPU may be a single-core GPU or a multi-core GPU. Alternatively, the processor 21 may be a processor group consisting of multiple GPUs, with the multiple processors coupled to each other via one or more buses. Alternatively, the processor may be another type of processor, and this is not limited in this embodiment of the present invention.

[0062] The memory 22 can be used to store computer program instructions and various computer program codes, including program codes for executing the embodiments of the present invention. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), and is used for related instructions and data.

[0063] The input device 23 is used to input data and / or signals, and the output device 24 is used to output data and / or signals. The output device 24 and the input device 23 can be independent devices or an integrated device.

[0064] The present invention simultaneously acquires package images through an RGB camera and a depth camera, sampling more comprehensive package information. Then, during package segmentation of the overall image, a clustering algorithm is used to group fragments of the same package into a single category. Multiple package fragments are then pieced together to restore the package. This overall end-to-end process uses the camera images directly as input to the model, which ultimately outputs the segmented package images. Volume calculation is then performed directly on the images. The present invention utilizes an end-to-end processing method based on deep learning to more efficiently and conveniently acquire package information and calculate package volume.

[0065] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A lightweight dynamic volume measurement method, characterized in that: include: Get the RGB image and depth image of the package; Preprocessing the depth map; Input the RGB image and the processed depth image into a segmentation model; The segmentation model outputs package information; The package volume is calculated according to the package information.

2. A lightweight dynamic volume measurement method according to claim 1, characterized in that: The preprocessing of the depth map comprises: Placing a checkerboard calibration plate on an unloaded synchronous belt plane, and calculating a rotation matrix between the synchronous belt plane and the camera lens plane based on the RGB image; Correcting the entire synchronous belt XY plane to be parallel to the camera lens XY plane according to the rotation matrix between the synchronous belt plane and the camera lens plane; Select the region of interest in the depth map; performing noise filtering and image enhancement on the region of interest; The average depth of the synchronous belt plane is calculated and used as the reference plane.

3. A lightweight dynamic volume measurement method according to claim 1, characterized in that: Inputting the RGB image and the processed depth image into the segmentation model comprises: Get the feature information of RGB image and depth image; The characteristic information is classified to obtain package fragments: Merging the package fragments to obtain single package information; Verify the accuracy of the stated package information.

4. A lightweight dynamic volume measurement method according to claim 1, characterized in that: Calculating the package volume according to the package information includes: The package information includes: the xy position, width, height and segmentation mask of the detection box of each package; Get the height information of each package under the corresponding area on the image, and integrate it to obtain the actual package volume data.

5. A lightweight dynamic volume measurement method according to claim 3, characterized in that: Acquiring feature information of the RGB image and the depth image includes: An adaptive feature extraction network is constructed to extract feature information of RGB images and depth images.

6. A lightweight dynamic volume measurement method according to claim 3, characterized in that: The classifying the characteristic information to obtain package fragments includes: Grayscale classification is performed based on the grayscale value of the image, and areas with similar grayscale values ​​are classified as one category; Feature classification is performed based on the feature information of the RGB image and the depth image. On the premise of satisfying the grayscale classification, areas with similar features are classified as one category.

7. A lightweight dynamic volume measurement method according to claim 4, characterized in that: The merging of the package fragments to obtain single package information includes: performing basic fusion on the package fragments; performing detail fusion on the package fragments; The detail fusion result and the basic fusion result are combined to obtain a single package.

8. A lightweight dynamic volume measurement method according to claim 3, characterized in that: Verifying the accuracy of the package information includes: Construct model loss function; Calculating the difference between the package and the true model; If the difference between the wrapped and true models is greater than the preset value, the adjustable weights, regularization coefficients, and mixing ratios of the loss function are adjusted; Repeat the verification until the difference between the package and the real model meets the preset value.

9. A lightweight dynamic volume measurement system, characterized in that: include: Get the RGB image and depth image of the package; Preprocessing the depth map; Input the RGB image and the processed depth image into a segmentation model; The segmentation model outputs package information; The package volume is calculated according to the package information.

10. An electronic device, characterized in that: include: A chip, a processor and a memory, wherein the memory is used to store computer program code, wherein the computer program code includes computer instructions. When the chip executes the computer instructions, the electronic device executes a lightweight dynamic volume measurement method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Logistics package volume measurement method, storage medium and terminal

    CN111383258A

  • Logistics part volume measurement method and device

    CN113379826A

  • Weak texture object pose estimation method and system

    CN113538569A

  • Parcel static volume measurement method and system

    CN116255912A

  • Mobile parcel volume measurement method and system based on RGB-D camera

    CN116993806A