Method and device for counting warehouse stacked items based on machine vision
By applying machine vision-based object detection model and density clustering algorithm in the warehouse, automatic disk counting of items stacked in warehouses is realized, solving the problems of high cost and low intelligence in existing methods, and achieving efficient and accurate counting tasks.
Patent Information
- Application Number
- CN202210156816.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The existing warehouse storage counting method has high equipment cost or storage costs, low intelligence, and is not easy to promote.
Using a machine vision-based method, the front and top surfaces of stacked items are classified and marked by constructing an object detection model, and combined with the density clustering algorithm to convert the counting statistics algorithm, and finally the online disk library counting is realized.
It realizes automatic disk storage counting of items stacked in warehouses, has high accuracy and robustness, reduces equipment and storage costs, and improves intelligence.
Smart Images

Figure CN114548868B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a method and device for counting warehouse stacked items based on machine vision. Background Art
[0002] Warehousing is the core link of modern logistics. With the rapid development of technologies such as artificial intelligence and computer vision, the development of warehousing technology has reached the stage of intelligence based on the informatization and automation of warehousing. Among the various functions of warehousing solutions, the counting of warehouse items is a crucial link. Traditional counting tasks are mostly completed manually by warehouse management workers. This work is generally completed in a concentrated period of time. It is very labor-intensive for warehouse management workers and is prone to errors.
[0003] In the related technologies, some intelligent inventory methods based on RFID (Radio Frequency Identification) require electronic tags to be added to each piece of goods in the warehouse, which is often difficult to achieve in warehouses that store general goods. In addition, there are inventory systems that use the "visual comparison" method, but this method simply compares the pictures of the goods when they enter the warehouse with the pictures when they leave the warehouse. If the algorithm thinks the difference is too large, it will be handed over to manual identification. The above methods all require large equipment costs or storage costs, and the degree of intelligence is low, which is not easy to promote.
[0004] In recent years, in the field of computer vision, CNN (Convolutional Neural Network) based object detection models have emerged one after another. They have been proven to be far superior to traditional methods in many fields such as autonomous driving, face detection, and pedestrian detection. However, existing visual object detection methods have not been fully applied in the field of warehouse inventory counting. Therefore, the warehouse stacking object counting method based on machine vision needs further research.
[0005] Application Contents
[0006] The present application provides a method and device for counting warehouse stacked items based on machine vision to solve the problems of high equipment cost or storage cost, low intelligence level, and difficulty in promotion.
[0007] In a first aspect, an embodiment of the present application provides a method for counting warehouse stacked objects based on machine vision, comprising the following steps: constructing a target detection model for classifying and labeling the front and top surfaces of the stack, the target detection model comprising a feature extraction network and a detection / classification network; dividing a training set and a validation set of the stacking image into batches of a predetermined size and performing preprocessing; selecting any batch in the preprocessed training set to input the target detection model for forward propagation, calculating the multi-task loss of the output value of the target detection model and the classification label, updating the weight of the target detection model based on the loss value and a preset optimizer by back propagation, and obtaining a stacking target detection model through multiple updates until an update end condition is met; for the detection box result obtained by the stacking target detection model, converting the detection box result into a counting result using a density-based clustering counting statistical algorithm; and performing online stacking object counting on warehouse stacking data using the stacking target detection model and the counting statistical algorithm.
[0008] Optionally, in one embodiment of the present application, the target detection model is a model structure based on Faster R-CNN, and the feature extraction network based on the model structure based on Faster R-CNN is a VGG16 network, a ResNet network or a ResNeXt network.
[0009] Optionally, in one embodiment of the present application, dividing the training set and the validation set of the stacked images into batches of a predetermined size and performing preprocessing includes:
[0010] Scaling the stacked image to the predetermined size according to the same aspect ratio by image scaling;
[0011] Use image horizontal flipping to flip the image horizontally with a probability of 0.5;
[0012] The histogram equalization algorithm is used to perform histogram equalization on the brightness V component in the HSV space of the whole image.
[0013] Optionally, in one embodiment of the present application, the multi-task loss includes a cross entropy classification loss and a smoothL1 loss for bounding box regression, wherein the aspect ratio of the Anchor in the region proposal network layer is {1:2, 1:1, 2:1} and its size is {8, 16, 32}.
[0014] Optionally, in one embodiment of the present application, the update end condition includes: the loss value is less than a preset threshold or the number of updates reaches a preset number of updates.
[0015] Optionally, in one embodiment of the present application, the density clustering algorithm is a clustering algorithm based on DBSACN, wherein the distance between detection box samples is expressed as follows:
[0016] Distance1(bbox 1 , bbox 2 )=|y 1min -y 2min |+|y 1max -y 2max |,
[0017] Distance2(bbox 1 , bbox 2 )=1 / |y 1min -y 2max |+1 / |y 1max -y 2min |,
[0018] Distonce(bbox 1 , bbox 2 )=Distance1(bbox 1 , bbox 2 )+λDistance2(bbox 1 , bbox 2 ),
[0019] Among them, Distance1(bbox 1 , bbox 2 ) is the sum of the distances between the upper and lower edges of the two boxes, Distance2(bbox 1 , bbox 2 )The second distance is the penalty term for the distance between the upper and lower boxes, Distance(bbox 1 , bbox 2 ) is Distance1(bbox 1 , bbox 2 ) and Distance2(bbox 1 , bbox 2 ) is the weighted sum of these two distances;
[0020] And, the counting statistics algorithm is:
[0021] N=(N layer -1)*N cargo-perlayer +N top ,
[0022] Among them, N cargo-perlayer is the number of boxes stacked in each layer, N layer is the total number of positive layers obtained by the clustering algorithm, N top is the top-level box obtained by the object detection model.
[0023] The second aspect of the present application provides a warehouse stacked object counting device based on machine vision, including: a model building module, used to build a target detection model for classifying and labeling the front and top surfaces of the stack, the target detection model including a feature extraction network and a detection / classification network; a data preprocessing module, used to divide the training set and the verification set of the stacking image into batches of a predetermined size and perform preprocessing; a model training module, used to select any batch in the preprocessed training set to input the target detection model for forward propagation, calculate the multi-task loss of the output value of the target detection model and the classification label, update the weight of the target detection model based on the loss value and a preset optimizer back propagation, and obtain the stacking target detection model through multiple updates until the update end condition is met; a conversion module, used to convert the detection box result obtained by the stacking target detection model into a counting result using a density-based clustering counting statistical algorithm; and a counting module, used to use the stacking target detection model and the counting statistical algorithm to perform online stacking object counting on the warehouse stacking data.
[0024] Optionally, in one embodiment of the present application, the target detection model is a model structure based on Faster R-CNN, and the feature extraction network based on the model structure based on Faster R-CNN is a VGG16 network, a ResNet network or a ResNeXt network.
[0025] Optionally, in one embodiment of the present application, the data preprocessing module is specifically used to:
[0026] Scaling the stacked image to the predetermined size according to the same aspect ratio by image scaling;
[0027] Use image horizontal flipping to flip the image horizontally with a probability of 0.5;
[0028] The histogram equalization algorithm is used to perform histogram equalization on the brightness V component in the HSV space of the whole image.
[0029] Optionally, in one embodiment of the present application, the update end condition includes: the loss value is less than a preset threshold or the number of updates reaches a preset number of updates.
[0030] Therefore, this application has at least the following beneficial effects:
[0031] By collecting and detecting stacking image data and dividing it into training set and verification set; preprocessing and data expansion of image data; using deep neural network detection model to locate and classify the front and top surfaces of the stacking, training on the training set until the iteration reaches the preset conditions; using the trained network to detect other stacking image data online; using the proposed three-dimensional counting algorithm to convert the detection results obtained by the deep neural network into counting results. Thus, the automatic inventory counting task of warehouse stacking items is realized, with strong robustness and high accuracy. Therefore, the problems of high equipment cost or storage cost, low degree of intelligence, and difficulty in promotion are solved.
[0032] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0034] Figure 1 A flowchart of a method for counting warehouse stacked items based on machine vision according to an embodiment of the present application;
[0035] Figure 2 This is a diagram of the overall network structure of target detection in a vision-based warehouse stacked object counting method provided according to an embodiment of the present application;
[0036] Figure 3 A partial structural diagram of a target detection feature extraction network of a vision-based warehouse stacked object counting method provided according to an embodiment of the present application;
[0037] Figure 4 A schematic diagram of the execution logic of a method for counting warehouse stacked items based on vision according to one embodiment of the present application;
[0038] Figure 5 This is an example diagram of a warehouse stacked item counting device based on machine vision according to an embodiment of the present application.
[0039] Explanation of the reference numerals: model building module-100, data preprocessing module-200, model training module-300, conversion module-400 and counting module-500. DETAILED DESCRIPTION
[0040] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0041] The following describes a method, device, electronic device and storage medium for counting warehouse stacked items based on machine vision according to an embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a method for counting warehouse stacked items based on machine vision. In this method, each surface of warehouse stacked items can be identified and located using only one stacking photo, and the detection results can be counted and statistically processed by an algorithm, ultimately achieving accurate counting of the number of stacked items, and proposing a more efficient and energy-saving implementation plan for the warehouse item counting task. As a result, the problems of high equipment cost or storage cost, low degree of intelligence, and difficulty in promotion are solved.
[0042] Specifically, Figure 1 A flowchart of a method for counting warehouse stacked items based on machine vision provided in an embodiment of the present application.
[0043] like Figure 1 As shown, the warehouse stacked object counting method based on machine vision includes the following steps:
[0044] In step S101, a target detection model is constructed to classify and label the front and top surfaces of the stack. The target detection model includes a feature extraction network and a detection / classification network.
[0045] Optionally, in one embodiment of the present application, the target detection model is a model structure based on Faster R-CNN, and the feature extraction network based on the model structure of Faster R-CNN is a VGG16 network, a ResNet network or a ResNeXt network.
[0046] It should be noted that in the embodiments of the present application, the above target detection model is based on the two-stage target detection model Faster R-CNN (Faster Regions with Convolutional Neural Network), and its network structure is as follows: Figure 2 Specifically, after a color stacking image is input into the target detection model, it will first pass through a convolutional feature extraction network, and convert the input image into a feature map with smaller size and higher channel dimension through convolutional layers and pooling layers.
[0047] After extracting features using the feature extraction network, you can use RPN (Region Proposal Network), that is, the region proposal network, to perform target detection and region positioning on the obtained feature map, use a rectangular anchor box to mark the position of the target on the feature map, and compare it with the IoU distance calculated by the annotation of the training sample to find the initial proposed region ROI (Region of Proposal) that is close to the real target.
[0048] After extracting the proposed region from the feature map, the proposed region and the original feature map are input into the candidate region pooling layer together. The ROIPooling operation is used to normalize the feature regions of different sizes to the same size, which facilitates the subsequent classification head and bounding box regression head to perform further classification and bounding box regression operations.
[0049] In the embodiment of the present application, the feature extraction network part may be VGG16, and the specific internal structure of the network is as follows: Figure 3 As shown. Specifically, the Conv is a 3x3 convolutional layer, and a zero padding operation is performed on the feature map; the Pooling layer is a 2x2 pooling layer. As a preference, the feature network may also select a network structure such as ResNet, ResNext, etc. in addition to VGG16.
[0050] Among them, the classification head is two cascaded fully connected layers. Its first layer reduces the dimension of the fixed-size ROI feature map extracted by the ROIPooling layer to 4096 dimensions, and the second layer reduces the dimension of the feature vector after dimensionality reduction to a preset number of categories (the number of categories in the embodiment of the present invention is 3 dimensions, and the categories include background, front and top surface) to obtain the final classification result.
[0051] The bounding box regression head is also composed of two cascaded fully connected layers. The first layer reduces the fixed-size ROI feature map extracted by the ROIPooling layer to 4096 dimensions and shares parameters with the classification layer. The second layer further reduces the dimension of the feature vector after dimension reduction to 4 times the preset number of classifications, where 4 represents the (y min ,x min ,y max ,x max ).
[0052] In step S102, the training set and the validation set of the stacked images are divided into batches of a predetermined size and preprocessed.
[0053] Optionally, in one embodiment of the present application, the training set and the validation set of the stacked images are divided into batches of a predetermined size and preprocessed, including: scaling the stacked images to a predetermined size with an equal aspect ratio using image scaling; flipping the images horizontally with a probability of 0.5 using image horizontal flipping; and performing histogram equalization on the brightness V component in the HSV space of the entire image using a histogram equalization algorithm.
[0054] It can be understood that the above-mentioned image scaling and image horizontal flipping preprocessing are aimed at expanding the data set and increasing the data volume, and the histogram equalization is aimed at balancing the illumination of the input image and, to a certain extent, performing high-quality optimization processing on the stacking images in complex and low-quality industrial scenes.
[0055] In step S103, any batch in the preprocessed training set is selected to input the target detection model for forward propagation, the multi-task loss of the output value of the target detection model and the classification label is calculated, and the weight of the target detection model is updated based on the loss value and the preset optimizer back propagation. The stacked target detection model is obtained through multiple updates until the update end condition is met.
[0056] Optionally, in one embodiment of the present application, the multi-task loss includes a cross entropy classification loss and a smoothL1 loss for bounding box regression, wherein the aspect ratio of the Anchor in the region proposal network layer is {1:2, 1:1, 2:1}, and its size is {8, 16, 32}. At the same time, in an embodiment of the present application, the update end condition is that the loss value is less than a preset threshold or the number of updates reaches a preset number of updates.
[0057] Specifically, in the embodiment of the present application, the VGG16 network weights pre-trained on ImageNet are used as the initial weights of the feature extraction network, the learning rate is set to 0.001, and SGD is used as the optimizer to train the network parameters, where the multi-task loss function is:
[0058]
[0059] Where i is the number of the training image; p i is the probability that the image belongs to a certain category, is the label of the category to which the image belongs, t i is the border coordinate of the image (y min , x min ,y max , x max ), A label for its coordinates.
[0060] Among them, L cls Using cross entropy loss, L regThe smooth_L1 loss is adopted. In the embodiment of the present application, the value of λ is 1.
[0061] In step S104, the detection frame results obtained by the stacked object detection model are converted into counting results using a density-based clustering counting statistics algorithm.
[0062] Optionally, in one embodiment of the present application, a learning rate decay strategy is adopted to reduce the learning rate by half every 10 epochs, for a total of 20 epochs.
[0063] In step S105, the warehouse stacking data is subjected to online stacking object count using a stacking target detection model and a counting statistics algorithm.
[0064] Optionally, in one embodiment of the present application, the density clustering algorithm is a clustering algorithm modified based on DBSACN, wherein the distance between detection box samples is expressed as follows:
[0065] Distance1(bbox 1 , bbox 2 )=|y 1min -y 2min |+|y 1max -y 2max |,
[0066] Distance2(bbox 1 , bbox 2 )=1 / |y 1min -y 2max |+1 / |y 1max -y 2min |,
[0067] Distance(bbox 1 , bbox 2 )=Distonce1(bbox 1 , bbox 2 )+λDistance2(bbox 1 , bbox 2 ),
[0068] The first distance is the sum of the distances between the upper and lower edges of the two boxes, and the second distance is the penalty term for the distance between the upper and lower boxes, the purpose of which is to make the distance between the upper and lower layers as far as possible. The final distance is the weighted sum of these two distances, where λ is 1.
[0069] And, the counting algorithm is as follows:
[0070] N=(N layer -1)*N cargo-perlayer +Ntop ,
[0071] Among them, N cargo-perlayer is the number of boxes in each stack, which is the stacking information that can be obtained in advance. N layer is the total number of front surface layers obtained by the clustering algorithm, N top is the total number of top surface detections obtained by the detection network.
[0072] Specifically, after obtaining the detection results, that is, the classification and positioning results of the top frame and the front frame, the detection results need to be converted into counting results. In the embodiment of the present application, the formula for the counting result is modeled as:
[0073] N=(N layer -1)*N cargo-perlayer +N top
[0074] That is, the number of objects in a stack can be expressed as the total number of layers minus one multiplied by the number of stacks on each layer, plus the number of stacks on the top layer. This is based on the a priori that objects in a stack must be placed slowly before the next layer can be placed. cargo-perlayer is the stacking information that can be obtained in advance. Therefore, the key to this algorithm is to obtain the remaining two parameters N layer 、N top .
[0075] The counting result of the top box is the sum of the top surface detection results obtained in the detection model. The total number of stacking layers requires a stratification algorithm for the front detection results obtained by the detection model. The embodiment of this application uses a density-based clustering algorithm to illustrate it. The specific algorithm is shown in Table 1.
[0076] Table 1 Density-based clustering algorithms
[0077]
[0078] Among them, for the two detection boxes bbox(y min ,x min ,y max ,x max ) The distance of the sample is defined as follows:
[0079] Distance1(bbox 1 , bbox 2 )=|y 1min -y 2min |+|y 1max -y 2max |
[0080] Distance2(bbox 1 , bbox 2)=1 / |y 1min -y 2max |+1 / |y 1max -y 2min |
[0081] Distance(bbox 1 , bbox 2 )=Distance1(bbox 1 , bbox 2 )+λDistance2(bbox 1 , bbox 2 )
[0082] The first distance is the sum of the distances between the upper and lower edges of the two boxes, and the second distance is the penalty term for the distance between the upper and lower boxes, the purpose of which is to make the distance between the upper and lower layers as far as possible. The final distance is the weighted sum of these two distances, where λ is 1.
[0083] A method for counting warehouse stacked items based on machine vision of the present application is described in detail below through a specific embodiment.
[0084] Figure 4 The execution logic of the warehouse stacked object counting method based on machine vision in the embodiment of the present application is shown, such as Figure 4 As shown, the warehouse stacked object counting method of the embodiment of the present application specifically includes the following steps:
[0085] Step 1: Build two types of target detection models for the front and top surfaces of the stack based on the deep neural network target detection model. The target detection model includes a feature extraction network and a detection / classification network.
[0086] Step 2: Divide the training set and validation set into batches of set size, and perform image scaling, image horizontal flipping, and image histogram equalization preprocessing.
[0087] Step 3: Select any batch in the training set, forward propagate the input data through the target detection network, calculate the multi-task loss of the output value and the label, and update the model weights based on the loss value and the preset optimizer through backpropagation.
[0088] Step 4: Repeat step 3 until the loss is lower than the set threshold or the set number of training times is reached to obtain the final stacking object detection model.
[0089] Step 5: For the detection box results obtained by the target detection model, use the density-based clustering counting statistics algorithm to convert the detection results into counting results.
[0090] Step 6: Use the trained deep neural network model and counting algorithm to perform online inventory counting of warehouse stacking data.
[0091] Offline stage: collect the stacked images required for training and divide them into training samples and verification samples; construct Figure 2 A deep target detection neural network model is developed, and the training samples and validation samples are preprocessed separately; the neural network model is forward-propagated using the training set, and the training error is back-propagated, and the neural network model is calculated after each round of iteration and the accuracy of target detection is predicted on the validation set until the preset training step is reached; the target detection results given by the trained deep neural network model are used to density cluster the front surface boxes, and the density clustering threshold is selected according to the actual clustering effect; the final stacking counting result is given by combining the front clustering results and the top surface detection results.
[0092] Online stage: The stacking images acquired by the camera on the warehouse stacker in real time are set in the same way as in the training stage to obtain the detection frame results. The front detection results are layered using the density clustering algorithm. This result is robust to the situation where the bottom layer is missed. The number of layers obtained by layering is calculated with the number of top surfaces detected on the top surface to obtain the final count result.
[0093] According to the machine vision-based warehouse stacked item inventory counting method proposed in the embodiment of the present application, the problem of efficient counting of stacked items in warehouse inventory tasks can be effectively solved. The front and top surfaces of the stack are detected and located through a deep neural network target detection model, and the detection result frames are counted using a density clustering-based hierarchical algorithm to obtain the final result. Real-time online counting of warehouse stacked items can be fully achieved with only one camera installed on the stacker. At the same time, compared with other inventory counting methods, the embodiments of the present application do not require additional electronic tags to the stored items, nor do they require any human participation, and require lower computing and storage costs, and have strong scalability.
[0094] Next, a device for counting stacked items in a warehouse based on machine vision according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0095] Figure 5 This is an example diagram of a warehouse stacked item counting device based on machine vision according to an embodiment of the present application.
[0096] like Figure 5 As shown, the machine vision-based warehouse stacked object counting device 10 includes: a model building module 100, a data preprocessing module 200, a model training module 300, a conversion module 400 and a counting module 500.
[0097] Among them, the model construction module 100 is used to construct a target detection model for classifying and labeling the front and top surfaces of the stack, and the target detection model includes a feature extraction network and a detection / classification network; the data preprocessing module 200 is used to divide the training set and the verification set of the stacking image into batches of a predetermined size and perform preprocessing; the model training module 300 is used to select any batch in the preprocessed training set to input the target detection model for forward propagation, calculate the multi-task loss of the output value of the target detection model and the classification label, and update the weight of the target detection model based on the loss value and the preset optimizer back propagation, and obtain the stacking target detection model through multiple updates until the update end condition is met; the conversion module 400 is used to convert the detection box result obtained by the stacking target detection model into a counting result using a density-based clustering counting statistical algorithm; and the counting module 500 is used to use the stacking target detection model and the counting statistical algorithm to perform online stacking item inventory counting on the warehouse stacking data.
[0098] Optionally, in one embodiment of the present application, the target detection model is a model structure based on Faster R-CNN, and the feature extraction network based on the model structure of Faster R-CNN is a VGG16 network, a ResNet network or a ResNeXt network.
[0099] Optionally, in one embodiment of the present application, the data preprocessing module 200 is specifically configured to:
[0100] scaling the stacked image to a predetermined size with equal aspect ratio using image scaling;
[0101] Use image horizontal flipping to flip the image horizontally with a probability of 0.5;
[0102] The histogram equalization algorithm is used to perform histogram equalization on the brightness V component in the HSV space of the whole image.
[0103] Optionally, in one embodiment of the present application, the update end condition includes: the loss value is less than a preset threshold or the number of updates reaches a preset number of updates.
[0104] It should be noted that the aforementioned explanation of the embodiment of the warehouse stacked object inventory counting method based on machine vision is also applicable to the warehouse stacked object inventory counting device based on machine vision of this embodiment, and will not be repeated here.
[0105] According to the machine vision-based warehouse stacked item inventory counting device proposed in the embodiment of the present application, through a camera and sufficient data support (which can be easily obtained in a warehouse with a lot of stacked items), efficient real-time counting of warehouse stacked items is fully realized, without the need for additional human assistance, which can effectively save labor and reduce the workload of warehouse managers. The embodiment of the present application also does not require excessive consumption of hardware resources. There is no need to add additional electronic tags to the stored items, nor is there a need to use RFID scanning equipment, which requires lower computing costs and storage costs, and has strong scalability. It has strong robustness. At the same time, the embodiment of the present application first pre-processes the input image data to enhance the network's output robustness to the input data; and then the counting algorithm also has good filtering results for underlying missed detections and false detections, which can avoid counting errors to a certain extent and obtain results with higher accuracy.
[0106] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0107] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0108] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
Claims
1. A method for counting warehouse stacked items based on machine vision. It is characterized in that The following steps are involved: Constructing a target detection model for classifying and labeling the front and top surfaces of the stack, wherein the target detection model is a model structure based on Faster R-CNN, and the target detection model includes a feature extraction network and a detection / classification network, and the feature extraction network based on the model structure of Faster R-CNN is a VGG16 network, a ResNet network, or a ResNeXt network; The collected stacking images are divided into training set and validation set; Divide the training set and validation set of the stacked images into batches of a predetermined size and perform preprocessing; Select any Batch in the preprocessed training set and input it into the target detection model for forward propagation, calculate the multi-task loss of the output value of the target detection model and the classification label, update the weight of the target detection model based on the loss value and the preset optimizer back propagation, and obtain the stacked target detection model through multiple updates until the update end condition is met, and calculate the target detection model after each round of iteration and predict the accuracy of target detection on the verification set until the preset training step is reached; specifically, after a stacked image is input into the target detection model, it first passes through a convolutional feature extraction network, and converts the input image into a feature map with a smaller size and higher channel dimension through convolutional layers and pooling layers, and extracts the proposed region from the feature map, and inputs the proposed region and the original feature map into the candidate region pooling layer together, and classifies the feature regions of different sizes into the same size through ROIPooling operation, and finally performs further classification and border regression through the classification head and the border regression head to obtain the output detection frame result, and the detection frame result is the classification and positioning result of the top face frame and the front face frame; For the detection frame result obtained by the stacking object detection model, convert the detection frame result into a counting result using a counting statistics algorithm based on density clustering; as well as Using the stacking target detection model and the counting statistics algorithm to perform online stacking object counting on warehouse stacking data; The density clustering algorithm is a clustering algorithm based on DBSACN, where the distance between detection box samples is expressed as follows: Distance1(bbox 1 ,bbox 2 )=|and 1min -and 2min |+|and 1max -and 2max |, Distance2(bbox 1 ,bbox 2 )=1 / |y 1min -and 2max |+1 / |and 1max -and 2min |, Distance(bbox 1 ,bbox 2 )=Distance1(bbox 1 ,bbox 2 )+λDistance2(bbox 1 ,bbox 2 ), Among them, Distance1(bbox 1 ,bbox 2 ) is the sum of the distances between the upper and lower edges of the two boxes, Distance2(bbox 1 ,bbox 2 )The second distance is the penalty term for the distance between the upper and lower boxes, Distance(bbox 1 ,bbox 2 ) is Distance1(bbox 1 ,bbox 2 ) and Distance2(bbox 1 ,bbox 2 ) is the weighted sum of the two distances, where λ is 1; And, the counting statistics algorithm is: N=(N layer -1)*N cargo-perlayer +N top , Among them, N cargo-perlayer is the number of boxes stacked in each layer, N layer is the total number of front layers obtained by the clustering algorithm, and the total number of front layers is obtained by the density clustering algorithm for the detection result of the front frame, N top is the top box obtained by the target detection model, and the counting result of the top box is the sum of the detection results of the top surface box.
2. The method according to claim 1, It is characterized in that The training set and the validation set of the stacked images are divided into batches of a predetermined size and preprocessed, including: Scaling the stacked image to the predetermined size according to the same aspect ratio by image scaling; Use image horizontal flipping to flip the image horizontally with a probability of 0.5; The histogram equalization algorithm is used to perform histogram equalization on the brightness V component in the HSV space of the whole image.
3. The method according to claim 1, It is characterized in that The multi-task loss includes cross entropy classification loss and smoothL1 loss of bounding box regression, wherein the aspect ratio of the Anchor in the region proposal network layer is {1:2, 1:1, 2:1} and its size is {8, 16, 32}.
4. The method according to claim 1, It is characterized in that The update end condition includes: the loss value is less than a preset threshold or the number of updates reaches a preset number of updates.
5. A warehouse stacked object counting device based on machine vision, used to implement the warehouse stacked object counting method based on machine vision as claimed in any one of claims 1 to 4, It is characterized in that include: A model building module, used to build a target detection model for classifying and labeling the front and top surfaces of the stack, wherein the target detection model includes a feature extraction network and a detection / classification network; A data preprocessing module, used for dividing the training set and the validation set of the stacked images into batches of a predetermined size and performing preprocessing; A model training module is used to select any Batch in the preprocessed training set and input it into the target detection model for forward propagation, calculate the multi-task loss of the output value of the target detection model and the classification label, and update the weight of the target detection model based on the loss value and the preset optimizer back propagation, and obtain the stacked target detection model through multiple updates until the update end condition is met; A conversion module, for converting the detection frame result obtained by the stacking target detection model into a counting result by using a counting statistics algorithm based on density clustering; as well as A counting module, used for performing online stacking object counting on warehouse stacking data by using the stacking target detection model and the counting statistics algorithm; The density clustering algorithm is a clustering algorithm based on DBSACN, where the distance between detection box samples is expressed as follows: Distance1(bbox 1 ,bbox 2 )=|and 1min -and 2min |+|and 1max -and 2max |, Distance2(bbox 1 ,bbox 2 )=1 / |y 1min -and 2max |+1 / |and 1max -and 2min |, Distance(bbox 1 ,bbox 2 )=Distance1(bbox 1 ,bbox 2 )+λDistance2(bbox 1 ,bbox 2 ), Among them, Distance1(bbox 1 ,bbox 2 ) is the sum of the distances between the upper and lower edges of the two boxes, Distance2(bbox 1 ,bbox 2 )The second distance is the penalty term for the distance between the upper and lower boxes, Distance(bbox 1 ,bbox 2 ) is Distance1(bbox 1 ,bbox 2 ) and Distance2(bbox 1 ,bbox 2 ) is the weighted sum of these two distances; And, the counting statistics algorithm is: N=(N layer -1)*N cargo-perlayer +N top , Among them, N cargo-perlayer is the number of boxes stacked in each layer, N layer is the total number of positive layers obtained by the clustering algorithm, N top is the top-level box obtained by the object detection model.
6. The device according to claim 5, It is characterized in that The target detection model is a model structure based on Faster R-CNN, and the feature extraction network based on the model structure based on Faster R-CNN is a VGG16 network, a ResNet network or a ResNeXt network.
7. The device according to claim 5, It is characterized in that The data preprocessing module is specifically used to: Scaling the stacked image to the predetermined size according to the same aspect ratio by image scaling; Use image horizontal flipping to flip the image horizontally with a probability of 0.5; The histogram equalization algorithm is used to perform histogram equalization on the brightness V component in the HSV space of the whole image.
8. The device according to claim 5, It is characterized in that The update end condition includes: the loss value is less than a preset threshold or the number of updates reaches a preset number of updates.
Citation Information
Patent Citations
Crayfish grading method based on machine learning
CN111666986A