Box body identification method and system

By using a segmented network model based on Transformer encoder in box recognition, combined with color diagrams and depth diagrams, the problem of low box recognition accuracy in the prior art is solved, and box recognition and automatic loading and unloading with higher accuracy are achieved.

CN119919927APending Publication Date: 2025-05-02POTEVIO LOGISTICS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411906116.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing box recognition methods have insufficient detection accuracy, which is easy to mistakenly divide the box, making it difficult to achieve high-accurate automatic loading and unloading in modern logistics systems.

Method used

A segmented network model based on Transformer encoder is used to combine color diagrams and depth diagrams for box recognition. By establishing a feature extractor and position encoder, using Transformer combined encoder and decoder, the detection head is set to output the boundary point coordinates and types of the box.

Benefits of technology

It improves the accuracy of box recognition, can more accurately identify the boundaries of the carton, provides more accurate data support for the robotic arm to grab the carton, and improves the degree of unloading automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919927A_ABST
    Figure CN119919927A_ABST
Patent Text Reader

Abstract

The invention provides a box body identification method. The method comprises the following steps: acquiring a color map and a depth map of a training box body in an identification scene of a target box body; making a data set according to the color map and the depth map of the training box body; a segmentation network model based on a Transform encoder is established; training a segmentation network model according to the data set; testing the segmentation network model; the box type and boundary point coordinates of the target box are recognized through the tested segmentation network model, a depth map is added to input data of the network model on the basis that a data source is a color map, the recognition effect can be improved by combining the distance between the carton and a camera, network design is improved on the basis of a Transform architecture, and the recognition efficiency is improved. The boundary of the carton can be identified more accurately, more accurate data support is provided for segmenting the point cloud according to the carton segmentation result when the mechanical arm grabs the carton, and the method has important practical significance for improving the automation degree of unloading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of box recognition, and in particular to a box recognition method and system. Background Art

[0002] With the rapid development of science and technology, the degree of automation in modern logistics systems is getting higher and higher, and visual recognition technology is developing rapidly. Automatic loading and unloading requires vision to identify the position of cartons. Currently, it mainly relies on 2D recognition and 3D point clouds to identify the spatial coordinates and posture of cartons. Among them, the 2D recognition effect plays a very important role in the accuracy of the overall results.

[0003] The existing carton detection methods are mainly based on target detection or instance segmentation. For example, algorithms such as YOLO are used to output the center point and rectangular bounding box of the carton, and algorithms such as Vit are used to output multiple boundary points of the carton. The above deep learning-based box detection methods have certain limitations. The detection accuracy is not high enough and it is easy to segment the box incorrectly. Summary of the invention

[0004] The object of the present invention is to provide a box recognition method to improve the accuracy of box recognition.

[0005] In order to achieve the above-mentioned purpose, the present invention provides the following technical solutions: a box recognition method, comprising: obtaining a color image and a depth image of a training box in the recognition scene of the target box; preparing a data set according to the color image and the depth image of the training box; establishing a segmentation network model based on a Transformer encoder; training the segmentation network model according to the data set; testing the segmentation network model; and identifying the box type and boundary point coordinates of the target box through the segmentation network model that has completed the test.

[0006] Furthermore, the image size of the color image is the same as the image size of the depth image, and the pixels of the color image are aligned with the pixels of the depth image.

[0007] Furthermore, the establishment of the data set of the color image and the depth image of the training box includes: collecting the data of the training box in the recognition scene of the target box; annotating the color image and the depth image according to the data of the training box in the recognition scene of the target box; using the color image and the depth image with annotated information as the data set; dividing the data set into a training set, a validation set and a test set; wherein the annotation information includes the position information of the four boundary points of the training box in the image coordinate system and the box category.

[0008] Furthermore, the establishment of a segmentation network model based on a Transformer encoder includes: establishing a feature extractor and a position encoder, so that the position encoder and the feature extractor receive the color image and the depth image; establishing a Transformer joint encoder and a Transformer decoder, so that the Transformer joint encoder receives the output data of the feature extractor and the position encoder, and the output data of the Transformer joint encoder is decoded by the Transformer decoder; setting a detection head, so that the Transformer decoder outputs the boundary point coordinates and the box type of the training box or the target box through the detection head.

[0009] Furthermore, the detection head is a feed-forward neural network.

[0010] Furthermore, the training of the segmentation network model according to the data set of color images and depth images includes: confirming the training parameters of the segmentation network model; confirming the loss function used to train the segmentation network model; loading the training set through the segmentation network model; training the segmentation network model; evaluating the performance of the segmentation network model through the validation set; determining whether the evaluation performance of the segmentation network model meets the training indicators; when the evaluation performance of the segmentation network model meets the training indicators, the segmentation network model after the evaluation performance meets the training indicators is used as the trained segmentation network model.

[0011] Furthermore, the testing of the segmentation network model includes: loading the test set through the segmentation network model; clustering the output results of the segmentation network model; judging whether the accuracy of the output results of the clustered segmentation network model reaches a preset accuracy; when the accuracy of the output results of the segmentation network model reaches a preset accuracy, the segmentation network model is considered to have completed the test.

[0012] Furthermore, the loss function includes an MSE loss function, a GIoU loss function and a Dice loss function.

[0013] Further, the identifying the box type and boundary point coordinates of the target box by the segmentation network model that has completed the test includes: identifying the boundary point coordinates of the target box by the segmentation network model that has completed the test; clustering all the target boxes according to the boundary point coordinates of the target box and the known types of the target boxes, so that the type of each target box matches a known type of the target box.

[0014] On the other hand, a box recognition system is provided, including: an image acquisition module, which acquires a color image and a depth image of a training box in the recognition scene of the target box; a data set establishment module, which establishes a data set of the color image and the depth image of the training box; a model establishment module, which establishes a segmentation network model based on a Transformer encoder; a model training module, which trains the segmentation network model according to the data set of the color image and the depth image; and a recognition module, which recognizes the boundary of the target box by using the trained segmentation network model.

[0015] From the analysis, it can be known that the present invention discloses a box recognition method. The input data of the network model of the present invention adds a depth map on the basis of the data source of the color map, which can improve the recognition effect in combination with the distance of the carton from the camera, and improves the network design based on the Transformer architecture, which can more accurately identify the boundary of the carton, and provide more accurate data support for segmenting the point cloud according to the carton segmentation result when the robot arm grabs the carton, which has important practical significance for improving the degree of unloading automation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings constituting a part of the present application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. Among them:

[0017] Figure 1 Flowchart of an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. Each example is provided by way of explanation of the present invention and does not limit the present invention. In fact, it will be clear to those skilled in the art that modifications and variations may be made in the present invention without departing from the scope or spirit of the present invention. For example, a feature shown or described as a part of one embodiment may be used in another embodiment to produce yet another embodiment. Therefore, it is desired that the present invention encompasses such modifications and variations within the scope of the appended claims and their equivalents.

[0019] One or more examples of the present invention are shown in the accompanying drawings. The detailed description uses numbers and letter labels to refer to features in the drawings. Like or similar labels in the drawings and description have been used to refer to like or similar parts of the present invention. As used herein, the terms "first", "second", "third", and "fourth", etc. are used interchangeably to distinguish one component from another and are not intended to indicate the location or importance of individual components.

[0020] like Figure 1As shown, according to an embodiment of the present invention, a box identification method is provided, comprising:

[0021] Step S101, obtaining a color image and a depth image of a training box in a recognition scene of a target box;

[0022] The image size of the color image is the same as that of the depth image, and the pixels of the color image are aligned with the pixels of the depth image.

[0023] Usually, a color camera and a depth camera are used to take pictures of the training box to obtain the source data. The camera here can be an RGBD camera or a set of calibrated color cameras and depth cameras. The color image output by the camera is consistent in size with the depth image and the pixels are aligned.

[0024] Step S102, creating a data set based on the color image and depth image of the training box.

[0025] Establishing a data set of color images and depth images of a training box includes: collecting data of the training box in a recognition scenario of a target box; annotating the color image and the depth image according to the data of the training box in the recognition scenario of the target box; using the color image and the depth image after annotating information as the data set; dividing the data set into a training set, a validation set and a test set; wherein the annotated information includes position information of four boundary points of the training box in an image coordinate system and a box category.

[0026] Step S103, establishing a segmentation network model based on the Transformer encoder.

[0027] The above-mentioned establishment of a segmentation network model based on the Transformer encoder specifically includes: establishing a feature extractor and a position encoder, so that the position encoder and the feature extractor receive the color image and the depth image; establishing a Transformer joint encoder and a Transformer decoder, so that the Transformer joint encoder receives the output data of the feature extractor and the position encoder, and the output data of the Transformer joint encoder is decoded by the Transformer decoder; setting a detection head so that the Transformer decoder and the above-mentioned Transformer joint encoder merge data for the existing two Transformer encoders.

[0028] The detection head outputs the boundary point coordinates of the training box or the target box, the box type and the output data confidence. The above-mentioned output data confidence is the data credibility of the output box boundary point coordinates and the box type, which indicates the accuracy of the model for the currently recognized box category. Generally, the closer it is to 1, the more certain it is that the current recognition result is a box.

[0029] The types of boxes are classified according to the size of the boxes, and the boxes are divided into several categories according to the size of the boxes.

[0030] The main body of the decoder is the decoder part of the DETR network.

[0031] The feature extractor mentioned above is ResNet-50.

[0032] The detection head mentioned above is a feed-forward neural network.

[0033] The specific operation steps of the segmentation network model are as follows: confirm that the input data is a color image and a depth image, copy the input color image and depth image data into two groups, input them into the feature extractor and position encoder respectively, then load the output data of the feature extractor and position encoder into their respective Transformer encoders, and merge the output results of the two. After merging, input them into the Transformer decoder for decoding, and after decoding, the detection head of the fully connected layer outputs the boundary point coordinates of the target box, the box type, and the output data confidence.

[0034] Step S104: training a segmentation network model according to the data set.

[0035] The segmentation network model trained based on the color and depth map datasets includes:

[0036] Confirm the training parameters of the segmentation network model. Usually, the batch size can be set to 1, the epoch (training round) can be set to 300, the learning rate can be set to 0.00001, and the Adam optimizer can be used. Confirm the loss function used to train the segmentation network model; load the training set through the segmentation network model; train the segmentation network model; evaluate the performance of the segmentation network model through the validation set; determine whether the evaluation performance of the segmentation network model meets the training indicators; when the evaluation performance of the segmentation network model meets the training indicators, the segmentation network model whose evaluation performance meets the training indicators is regarded as the segmentation network model that has completed training.

[0037] The loss functions include MSE loss function, GIoU loss function and Dice loss function. The bitchsize is set to 1, the epoch is set to 300, the learning rate is 0.00001, and the Adam optimizer is used.

[0038] The performance of the segmentation network model is evaluated through the validation set. Specifically, the validation set is used to evaluate the evaluation function MIOU, Precision, and Recall size to estimate the network performance.

[0039] Step S105, testing the segmentation network model.

[0040] Testing the segmentation network model includes: loading the test set through the segmentation network model; clustering the output results of the segmentation network model; judging whether the accuracy of the output results of the clustered segmentation network model reaches a preset accuracy; when the accuracy of the output results of the segmentation network model reaches the preset accuracy, the segmentation network model is considered to have completed the test.

[0041] The above-mentioned identifying the box type and boundary point coordinates of the target box by completing the test of the segmentation network model specifically includes: identifying the boundary point coordinates of the target box by completing the test of the segmentation network model; clustering all target boxes according to the boundary point coordinates of the target box and the known types of the target box, so that the type of each target box matches the known type of a target box.

[0042] Because there will be errors in the size of the box recognized in the actual industrial scene, clustering can be used to determine which boxes have the same size. The size error problem of individual boxes can be reduced by processing the mean of the data after clustering or matching the size of the boxes in the known database. Box clustering processing is to process the results after the segmentation network model outputs the results. Usually, the output results of the segmentation network model need to be processed for box size adaptation so that the size of each box will change to the size after the carton clustering processing. Specifically, it can be understood that, for example, there are n boxes in total, and the center point of each box is used as the coordinate origin, the side of the box on the y-axis of the camera coordinate system is used as a point on the y-axis, and the side of the box on the x-axis is used as a point on the x-axis, recorded as (x_n, y_n). In this way, the clustering of box sizes is converted into clustering of points, and then the DBSCAN clustering algorithm is used to cluster these points, thereby realizing the clustering of box sizes.

[0043] The size of each box can then be changed to the data mean of the class to which it belongs after clustering or to match the size of the box in a known database.

[0044] Step S106, identifying the box type and boundary point coordinates of the target box through the segmentation network model that has completed the test.

[0045] The present invention also discloses a box recognition system, comprising: an image acquisition module, which acquires a color image and a depth image of a training box in a recognition scene of a target box; a data set establishment module, which establishes a data set of the color image and the depth image of the training box; a model establishment module, which establishes a segmentation network model based on a Transformer encoder; a model training module, which trains the segmentation network model according to the data set of the color image and the depth image; and a recognition module, which recognizes the boundary of the target box by using the trained segmentation network model.

[0046] From the above description, it can be seen that the above-mentioned embodiments of the present invention achieve the following technical effects: the input data of the network model designed by the present invention adds a depth map on the basis of the data source of the color map, which can improve the recognition effect by combining the distance of the carton from the camera, and improves the network design based on the Transformer architecture, which can more accurately identify the boundaries of the carton, and provide more accurate data support for segmenting the point cloud according to the carton segmentation results when the robot arm grabs the carton, which has important practical significance for improving the degree of unloading automation.

[0047] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A box identification method, characterized in that: include: Obtain the color image and depth image of the training box in the recognition scene of the target box; Creating a data set based on the color image and depth image of the training box; Establish a segmentation network model based on Transformer encoder; Training the segmentation network model according to the data set; Testing the segmentation network model; The segmentation network model that completes the test identifies the box type and boundary point coordinates of the target box.

2. A box identification method according to claim 1, characterized in that: The image size of the color image is the same as that of the depth image, and the pixels of the color image are aligned with the pixels of the depth image.

3. A box identification method according to claim 1, characterized in that: The data set for establishing the color image and depth image of the training box includes: Collecting data of the training box in the recognition scene of the target box; Annotating the color image and the depth image according to the data of the training box in the recognition scene of the target box; The color image and the depth image with the annotated information are used as the data set; Dividing the data set into a training set, a validation set, and a test set; The annotation information includes the position information of four boundary points of the training box in the image coordinate system and the box category.

4. A box identification method according to claim 3, characterized in that: The establishment of a segmentation network model based on a Transformer encoder includes: Establishing a feature extractor and a position encoder, so that the position encoder and the feature extractor receive the color map and the depth map; Establishing a Transformer joint encoder and a Transformer decoder, so that the Transformer joint encoder receives output data of the feature extractor and the position encoder, and the output data of the Transformer joint encoder is decoded by the Transformer decoder; A detection head is set so that the Transformer decoder outputs the boundary point coordinates and the box type of the training box or the target box through the detection head.

5. A box identification method according to claim 4, characterized in that: The detection head is a feed-forward neural network.

6. A box identification method according to claim 3, characterized in that: The step of training the segmentation network model according to the data set of the color image and the depth image comprises: Confirming the training parameters of the segmentation network model; Determine the loss function used to train the segmentation network model; Loading the training set through the segmentation network model; Training the segmentation network model; Evaluating the performance of the segmentation network model through the validation set; Determining whether the evaluation performance of the segmentation network model meets the training index; When the evaluation performance of the segmentation network model meets the training index, the segmentation network model after the evaluation performance meets the training index is used as the segmentation network model after training.

7. A box identification method according to claim 3, characterized in that: The testing of the segmentation network model comprises: Loading the test set through the segmentation network model; Clustering the output results of the segmentation network model; Determining whether the accuracy of the output result of the segmentation network model after clustering reaches a preset accuracy; When the accuracy of the output result of the segmentation network model reaches a preset accuracy, the segmentation network model is considered to have completed the test.

8. A box identification method according to claim 6, characterized in that: The loss functions include MSE loss function, GIoU loss function and Dice loss function.

9. A box identification method according to claim 1, characterized in that: The segmentation network model that completes the test to identify the box type and boundary point coordinates of the target box includes: Identify the boundary point coordinates of the target box by using the segmentation network model that completes the test; All the target boxes are clustered according to the boundary point coordinates of the target boxes and the known types of the target boxes, so that the type of each target box matches a known type of the target box.

10. A box identification system, characterized in that: include: An image acquisition module obtains a color image and a depth image of the training box in the recognition scene of the target box; A data set establishment module, establishing a data set of color images and depth images of the training box; Model building module, building a segmentation network model based on Transformer encoder; A model training module, for training the segmentation network model according to the data sets of the color image and the depth image; The recognition module recognizes the boundary of the target box by using the trained segmentation network model.