Video Stream-Based Classification Method and Device for Hanging Garments in Logistics Centers

Through deep learning and feature fusion technology based on video streams, the problem of low efficiency in the classification of hang-mounted clothing in the logistics center is solved, and efficient and accurate clothing classification is achieved.

CN114663803BActive Publication Date: 2025-07-29BAOKAI (SHANGHAI) INTELLIGENT LOGISTICS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210193004.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-07-29
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately classify garments in logistics centers, especially in the case of high-speed movement and high-light reflection. The existing methods are inefficient and insufficiently automated.

Method used

A deep learning model based on video stream is used to identify the clothing bounding box, and multi-scale image features are extracted and feature fusion is performed through the bounding box numbering and scale weighting mechanism, and a convolutional neural network is used for classification.

Benefits of technology

Real-time classification of hanging clothing is realized, processing efficiency and classification accuracy are improved, and clothing categories can be effectively identified in high-speed motion and high-light reflection environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663803B_ABST
    Figure CN114663803B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for classifying hanging clothes in a logistics center based on a video stream. The steps of the method include receiving the input video stream, identifying the clothes in the video stream based on a preset deep learning model, and marking the bounding boxes of the clothes in the image data of the video stream; sequentially numbering the bounding boxes of the clothes based on the appearance order of the clothes in the image data of the video stream; dividing the image data of the video stream into multiple first image frames, and cropping the first image frames based on the bounding boxes to obtain boundary images; for the boundary images with the same bounding box number, extracting the image features in the boundary images, assigning different weights to the image features in the boundary images of different scales based on the scale of the boundary images, and performing feature fusion on the boundary images with the same bounding box number based on the attention mechanism to obtain a fused image; inputting the fused image into a preset convolutional neural network classifier to obtain the category of the fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of logistics, and in particular, to a method and device for classifying hanging clothes in a logistics center based on a video stream. Background Art

[0002] Generally, a logistics center transports multiple types of hanging clothes at the same time, such as suits, overcoats, down jackets, etc. These clothes may come from multiple orders mixed together. Therefore, in the sorting process, the clothes need to be classified by category to facilitate subsequent packing of the clothes and then transporting them to the corresponding clothing manufacturers.

[0003] Currently, the related technologies for clothing classification can be roughly divided into identifying each one by using barcodes, or using image recognition for identification. If barcode recognition is used, it is usually carried out manually or by using a camera to read barcodes, with low efficiency; if image recognition is used, usually only a single image can be recognized. However, in a real logistics scenario, hanging clothes move at high speed on the conveying equipment, making it difficult to capture clear barcodes or single images, resulting in poor effects of the above methods in practical applications. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method for classifying hanging clothes in a logistics center based on a video stream to eliminate or improve one or more defects existing in the prior art.

[0005] One aspect of the present invention provides a method for classifying hanging clothes in a logistics center based on a video stream. The steps of the method include:

[0006] Receiving the input video stream, identifying clothes in the image data of the video stream based on a preset deep learning model, and marking the bounding boxes of the clothes in the image data of the video stream;

[0007] Numbering the bounding boxes of the clothes in sequence based on the appearance order of the clothes in the image data of the video stream;

[0008] Dividing the image data of the video stream into multiple first image frames, and cropping the first image frames based on the bounding boxes to obtain boundary images;

[0009] For boundary images with the same bounding box number, extracting image features in the boundary images, assigning different weights to the image features in boundary images of different scales based on the scale of the boundary images, and performing feature fusion on the boundary images with the same bounding box number based on the attention mechanism to obtain fused images;

[0010] Inputting the fused images into a preset convolutional neural network classifier to obtain the categories of the fused images.

[0011] With the above solution, this solution can receive video streams in real time. Since the same piece of clothing in a video stream usually follows the rule of moving from far to near or from near to far, this solution can obtain multiple images of the same piece of clothing at different scales. And the representativeness of the same feature in images at different scales is different. Therefore, this solution assigns different weights and obtains a fused image. By classifying the fused image, the category of the clothing with the same bounding box number in the original image data can be obtained. On the one hand, this solution can identify the clothing category in real time and improve the processing efficiency. Moreover, by assigning different weights to features, the accuracy of classification can be improved.

[0012] In some embodiments of the present invention, in the step of marking the bounding box of the clothing in the image data of the video stream, the bounding box is generated according to the scale of the clothing in the image data, and the image of the clothing is within the range defined by the bounding box;

[0013] The bounding box changes as the scale of the clothing it encloses becomes larger or smaller in the image data of the video stream.

[0014] In some embodiments of the present invention, in the step of marking the bounding box of the clothing in the image data of the video stream, based on a preset bounding box threshold, the size of the current bounding box is compared with the bounding box threshold in real time. If the size of the current bounding box is not within the range of the bounding box threshold, the bounding box is not displayed.

[0015] In some embodiments of the present invention, the step of dividing the image data of the video stream into multiple first image frames includes dividing the image data of the video stream into multiple initial image frames according to the frames of the image data;

[0016] Convert the initial image frames into grayscale images, and calculate the gray centroid of each grayscale image based on the gray values of each pixel point in the grayscale images;

[0017] Calculate the average gray centroid based on the gray centroids of each grayscale image, and calculate the distances between the gray centroids of each grayscale image and the average gray centroid respectively;

[0018] Based on the distances between the gray centroids and the average gray centroid, select the first preset number of grayscale images with relatively close distances from all grayscale images as the first image frames.

[0019] In some embodiments of the present invention, the gray centroid of each grayscale image is calculated based on the gray values of each pixel point in the grayscale image according to the following formula:

[0020]

[0021]

[0022] Combine x according to the above formula c , yc Obtain the coordinates (x c , y c ) of the gray centroid. In the formula, x c represents the abscissa of the gray centroid, y c represents the ordinate of the gray centroid, x ij is the pixel gray value of the pixel at the i-th row and j-th column of the grayscale image with size M*N. M is the total number of pixel rows of the grayscale image, and N is the total number of pixel columns of the grayscale image.

[0023] In some embodiments of the present invention, calculate the average values of the abscissa and ordinate of the gray centroids of all grayscale images respectively to obtain the average gray centroid.

[0024] In some embodiments of the present invention, calculate the distance between the gray centroid of each grayscale image and the average gray centroid according to the following formula:

[0025]

[0026] d i represents the distance between the gray centroid of the grayscale image i and the average gray centroid. represents the abscissa of the gray centroid of the grayscale image i. represents the ordinate of the gray centroid of the grayscale image i. represents the abscissa of the average gray centroid. represents the ordinate of the average gray centroid.

[0027] In some embodiments of the present invention, calculate the distance between the gray centroid of each grayscale image and the average gray centroid respectively, and screen out the first preset number of grayscale images with relatively close distances as the first image frame.

[0028] In some embodiments of the present invention, divide into multiple scale ranges. Each scale range corresponds to a preset feature-weight group. In the feature-weight group, weight parameters are set for each image feature. The step of assigning different weights to the features in the boundary images of different scales based on the scale of the boundary image includes:

[0029] Determine the scale range corresponding to the boundary image based on the scale of the boundary image, match the corresponding feature-weight group for the boundary image according to the scale range, and assign the corresponding weight parameters to each image feature in the boundary image.

[0030] In some embodiments of the present invention, a plurality of scale thresholds are divided, and the scale thresholds are arranged in order based on the numerical magnitudes of the plurality of scale thresholds. Each of the scale thresholds is correspondingly provided with a feature-weight group, and weight parameters are set for each image feature in the feature-weight group. The step of assigning different weights to features in boundary images of different scales based on the scale of the boundary image includes;

[0031] Sort the boundary images with the same bounding box number based on their order in the video stream. Based on the sorting order of the boundary images, compare the scale of the first boundary image in the sorted boundary images with the first scale threshold in the scale sorting;

[0032] If the scale of the first boundary image is greater than the first scale threshold, continue to compare it with the next scale threshold until the boundary image is less than or equal to the nth scale threshold. Match the feature-weight group corresponding to the nth scale threshold for the first boundary image, and assign the corresponding weight parameters to each image feature in this boundary image;

[0033] If the scale of the first boundary image is less than or equal to the first scale threshold, match the feature-weight group corresponding to the first scale threshold for the first boundary image, and assign the corresponding weight parameters to each image feature in this boundary image;

[0034] Compare the scale of the a-th boundary image in the sorted boundary images with the b-th scale threshold matched by the (a - 1)-th boundary image;

[0035] If the scale of the a-th boundary image is greater than the b-th scale threshold, continue to compare it with the (b + 1)-th scale threshold until the boundary image is less than or equal to the m-th scale threshold. Match the feature-weight group corresponding to the m-th scale threshold for the a-th boundary image, and assign the corresponding weight parameters to each image feature in this boundary image;

[0036] If the scale of the a-th boundary image is less than or equal to the b-th scale threshold, match the feature-weight group corresponding to the b-th scale threshold for the a-th boundary image, and assign the corresponding weight parameters to each image feature in this boundary image.

[0037] In some embodiments of the present invention, the convolutional neural network classifier is trained based on the following loss function formula:

[0038]

[0039] L represents the loss function value, f represents the f-th fused image, F represents the total number of fused images, is the feature vector x corresponding to the image features of the f-th fused image fThe value obtained by normalization, e represents the Euler number, g represents the g-th image feature, and s g represents the eigenvalue of the g-th image feature, G represents the total number of categories of image features, is the value obtained by normalizing the image feature s in the fused images of the same category corresponding to the f-th fused image g γ represents the weight of the feature distance function, and the feature distance function is represents the center point of the feature vectors of the fused images of the same category corresponding to the fused image f.

[0040] In some embodiments of the present invention, the step of dividing the image data of the video stream into multiple first image frames further includes performing a clarification process on the divided first image frames. The clarification process includes using Wiener filtering to remove the first image frames containing noise among the multiple first image frames.

[0041] In some embodiments of the present invention, the step of cropping the first image frame based on the bounding box to obtain a boundary image includes performing a highlight removal process on the boundary image. The highlight removal process includes

[0042] obtaining a transformation matrix corresponding to multiple boundary images with the same bounding box number based on the SURF algorithm;

[0043] dividing every α of the multiple boundary images with the same bounding box number into a fusion group, and performing an alignment process on the boundary images in the same fusion group based on the transformation matrix;

[0044] fusing the boundary images in the same fusion group into the same boundary image by combining taking the minimum pixel gray value, taking the gray average value, Gaussian difference, and median.

[0045] The additional advantages, objectives, and features of the present invention will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be pointed out and obtained specifically in the description and the drawings.

[0046] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.

[0048] Figure 1 This is a schematic diagram of an implementation method of the clothing classification method for hanging clothes in a logistics center based on a video stream according to the present invention;

[0049] Figure 2 This is a schematic diagram of the use of bounding box numbers;

[0050] Figure 3 This is a schematic diagram of the processing structure of a convolutional neural network;

[0051] Figure 4 This is a schematic diagram of another implementation method of the clothing classification method for hanging clothes in a logistics center based on a video stream according to the present invention. Detailed implementation method

[0052] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the implementation methods and drawings. Here, the schematic implementation methods and descriptions of the present invention are used to explain the present invention, but do not limit the present invention.

[0053] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0054] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0055] Here, it should also be noted that if not otherwise specified, the term "connection" in this article can not only refer to direct connection, but also represent indirect connection with intermediaries.

[0056] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0057] Introduction to the prior art:

[0058] (1) Classification by scanning codes with a barcode scanner

[0059] The most direct method is for manual classification by using a barcode scanner to read codes. This method has an almost 100% accuracy rate, but is very slow in terms of efficiency. Hanging clothes take up a lot of space, so it is very difficult to scan codes for them, and it cannot be used in a logistics center with a large order volume.

[0060] Furthermore, if manual barcode scanning and code reading are used for classification, first the clothes need to be taken off the production line, and then the barcode scanner is held for identification, which is inefficient and results in serious waste of human resources.

[0061] (2) The camera reads the barcodes for classification

[0062] This method uses a camera to read barcodes for classification. As long as the label of the clothing appears within the shooting range of the camera, its barcode can be obtained, and then the clothing is sent to a cross-belt sorter and transported to the corresponding stalls. This method is faster than manual classification, but it is not sufficient to meet the challenges of multiple orders and multiple types in the logistics center. If the label is upside down, the barcode content cannot be obtained, and thus classification cannot be carried out.

[0063] Furthermore, if a camera barcode reader is used, manual assistance is required to place the clothing label within the shooting range of the camera for barcode reading and recognition classification. Situations such as the label being upside down or partially blocked will result in unrecognizability, and the efficiency is also low. The degree of automation needs to be improved.

[0064] (3) Clothing classification method based on zero-shot recognition

[0065] This method labels the clothing features as attribute vectors, extracts the feature vectors of the clothing image, learns the mapping from the attribute vectors of the training set to the feature vectors, and then inputs the feature vectors of the test set into the learned mapping to obtain the corresponding attribute vectors of the test set and find the clothing category closest to them. The disadvantage of this method is that it does not fundamentally use clothing features for classification, and the predicted categories are still relatively few. If the number of categories is very large, classification may not be possible.

[0066] (4) Clothing classification method based on feature enhancement

[0067] This method extracts the texture features and shape features of the clothing image, then combines them into attribute features and inputs them into a discriminator, and the discriminator predicts the category of the clothing. This method has high requirements for image clarity and cannot meet the recognition requirements for blurred images during movement. Among them, texture features are relatively difficult to obtain, increasing the difficulty of integration.

[0068] (5) Clothing classification method based on deep learning

[0069] This method uses an attention mechanism to amplify the key vectors and weights of the clothing image features, uses a spatial transformation network to transform the receptive field of the image features, and then inputs the image features into a capsule network to extract spatial correlation information, and classifies the clothing according to the high-level information. Its disadvantage is that it cannot achieve real-time clothing classification, and in terms of network design, it still uses an ordinary convolutional neural network model, and the classification effect for clothes with strong edge features such as windbreakers and cheongsams is not good.

[0070] Furthermore, the clothing classification technology based on a single image has a higher degree of automation than the above methods, but has higher requirements for the resolution of the camera. Otherwise, it will lead to classification failure due to the inability to accurately extract image features, and cannot meet the needs of real-time classification of hanging clothing in the actual logistics center scenario.

[0071] The above existing methods are all based on barcodes or single images for clothing classification. However, in the real logistics scenario, hanging clothing moves at high speed on the conveying equipment, making it difficult to capture clear barcodes or single images, resulting in poor performance of the above methods in practical applications. Currently, there is relatively little research on clothing classification based on video streams. Therefore, the goal of this patent is to achieve the classification of hanging clothing in the logistics center based on video streams, and there are the following difficulties to be solved: 1. Due to the vibration generated by the movement of high-speed conveying equipment, it is difficult to obtain clear frames in the video stream; 2. Since the hanging clothing in the logistics center has plastic packaging, it is prone to high-light reflection, seriously affecting the image quality and increasing the recognition difficulty.

[0072] To solve the above problems, the present invention proposes a method for classifying hanging clothing in a logistics center based on a video stream. This method divides the obtained video stream into multiple images, first removes the high-light reflection in the images, and then performs feature fusion on multiple images of different scales through an attention mechanism to obtain image features with greater discrimination, and finally realizes the accurate classification of hanging clothing.

[0073] As Figure 1 、 4 shown, one aspect of the present invention provides a method for classifying hanging clothing in a logistics center based on a video stream. The steps of the method include,

[0074] Step S100, receiving the input video stream, identifying the clothing in the image data of the video stream based on a preset deep learning model, and marking the bounding box of the clothing in the image data of the video stream;

[0075] In some embodiments of the present invention, recording starts after the hanging clothing to be processed enters the shooting range of the camera. First, the clothing in the video is detected, marked with a bounding box and numbered. The initial number is 1. If new clothing appears in subsequent frames, the bounding box numbers are sequentially incremented to achieve synchronous tracking of multiple pieces of clothing;

[0076] Step S200, sequentially numbering the bounding boxes of the clothing based on the appearance order of the clothing in the image data of the video stream;

[0077] Step S300, dividing the image data of the video stream into multiple first image frames, and cropping the first image frames based on the bounding boxes to obtain boundary images;

[0078] In some embodiments of the present invention, according to the bounding box numbers, the same piece of clothing with the same bounding box numbers that co - appear in multiple images is cropped to obtain images of the clothing at different scales.

[0079] Step S400: For the boundary images with the same bounding box numbers, extract the image features in the boundary images, assign different weights to the image features in the boundary images of different scales based on the scale of the boundary images, and perform feature fusion on the boundary images with the same bounding box numbers based on the attention mechanism to obtain a fused image;

[0080] The image features include, but are not limited to, contour features and texture features.

[0081] Step S500: Input the fused image into a preset convolutional neural network classifier to obtain the category of the fused image.

[0082] Adopting the above - mentioned solution, this solution can receive the video stream in real - time. Since the same piece of clothing in the video stream usually follows the rule of from far to near or from near to far, this solution can obtain multiple images of the same piece of clothing at different scales. And the representativeness of the same feature in images of different scales is different. Therefore, this solution assigns different weights and obtains a fused image, getting a clear picture frame. By classifying the fused image, the category of the clothing with the same bounding box numbers in the original image data can be obtained. On the one hand, this solution can identify the clothing category in real - time and improve the processing efficiency, and by assigning different weights to the features, it can improve the classification accuracy.

[0083] In some embodiments of the present invention, the multi - scale images obtained are subjected to feature fusion to obtain a fused image with higher discrimination, and then sent to a classifier. According to the classification result, the clothing is sent to the corresponding category stalls through a cross - belt sorter to complete the real - time classification of the clothing.

[0084] In some embodiments of the present invention, if the clothing classification is completed, the corresponding bounding box number can be used again.

[0085] In some embodiments of the present invention, the sequence of the bounding box numbers can be a recyclable sequence. Since the clothing that appears first in the video will be classified first, the bounding box number of this clothing can return to the number pool. When the clothing that appears later uses this number again, they will not affect each other, and it can avoid the serial numbers being too large for a long time and increase the calculation difficulty.

[0086] Such as Figure 2As shown, in some embodiments of the present invention, a numbered circular queue can be set for the bounding box numbers. Initially, the numbers range from 1 to 100, and two pointers, front and rear, are set, pointing to the next number to be used and the number to be stored respectively. After using a number or storing a number, the pointer moves one position backward.

[0087] Numbers are taken from the head of the circular queue to number the bounding boxes in multiple images. The bounding boxes at the same position have the same number. If new clothing appears in subsequent frames, the numbering order in the previous images is used as the standard, and the numbers of the new clothing are accumulated in sequence. When the classification of the bounding boxes with the same number is completed, the numbers are put back into the number queue for the convenience of subsequent use of the bounding boxes.

[0088] In some embodiments of the present invention, in the step of marking the bounding box of the clothing in the image data of the video stream, a bounding box is generated according to the scale of the clothing in the image data, and the image of the clothing is within the range framed by the bounding box;

[0089] The bounding box becomes larger or smaller as the scale of the clothing it frames changes in the image data of the video stream.

[0090] In some embodiments of the present invention, in the step of marking the bounding box of the clothing in the image data of the video stream, based on a preset bounding box threshold, the size of the current bounding box is compared with the bounding box threshold in real time. If the size of the current bounding box is not within the range of the bounding box threshold, the bounding box is not displayed.

[0091] As shown in the following formula:

[0092] threshold low ≤S bounding box ≤threshold high .

[0093] With the above scheme, the upper and lower thresholds of the bounding box size are set, and it is judged whether the size of the bounding box in the video frame is between the upper and lower thresholds. Only the video frames within this range are retained. If not, it means the image is too large or too small and difficult to identify features, so it is directly discarded to reduce the processing burden.

[0094] In some embodiments of the present invention, the step of dividing the image data of the video stream into multiple first image frames includes dividing the image data of the video stream into multiple initial image frames according to the frames of the image data;

[0095] The initial image frames are converted into grayscale images, and the gray centroid of each grayscale image is calculated based on the gray values of the individual pixel points in the grayscale images;

[0096] Calculate the average gray centroid based on the gray centroids of each grayscale image, and calculate the distances between the gray centroids of each grayscale image and the average gray centroid respectively;

[0097] Based on the distances between the gray centroids and the average gray centroid, select the first preset number of grayscale images with relatively close distances from all grayscale images as the first image frame.

[0098] In some embodiments of the present invention, calculate the gray centroid of each grayscale image based on the gray values of each pixel point in the grayscale image according to the following formula:

[0099]

[0100]

[0101] Combine x according to the above formula c , y c to obtain the coordinates (x c , y c ) of the gray centroid. In the formula, x c represents the abscissa of the gray centroid, y c represents the ordinate of the gray centroid, x ij is the pixel gray value of the pixel point in the i-th row and j-th column of the grayscale image with the size of M*N. M is the total number of pixel rows of the grayscale image, and N is the total number of pixel columns of the grayscale image.

[0102] In some embodiments of the present invention, calculate the average values of the abscissas and ordinates of the gray centroids of all grayscale images respectively to obtain the average gray centroid

[0103] In some embodiments of the present invention, calculate the distances between the gray centroids of each grayscale image and the average gray centroid according to the following formula:

[0104]

[0105] d z represents the distance between the gray centroid of the grayscale image z and the average gray centroid, represents the abscissa of the gray centroid of the grayscale image z, represents the ordinate of the gray centroid of the grayscale image z, represents the abscissa of the average gray centroid, represents the ordinate of the average gray centroid.

[0106] In some embodiments of the present invention, calculate the distances between the gray centroids of each grayscale image and the average gray centroid respectively, and select the first preset number of grayscale images with relatively close distances as the first image frame.

[0107] In some embodiments of the present invention, a plurality of scale ranges are divided, each scale range corresponding to a preset feature - weight group. In the feature - weight group, weight parameters are set for each image feature. The step of assigning different weights to features in boundary images of different scales based on the scale of the boundary image includes,

[0108] Determine the scale range corresponding to the boundary image based on the scale of the boundary image, match the corresponding feature - weight group for the boundary image according to the scale range, and assign the corresponding weight parameters to each image feature in the boundary image.

[0109] In some embodiments of the present invention, a plurality of scale thresholds are divided, and the scale thresholds are arranged in order based on the numerical magnitudes of the plurality of scale thresholds. Each of the scale thresholds is correspondingly provided with a feature - weight group. In the feature - weight group, weight parameters are set for each image feature. The step of assigning different weights to features in boundary images of different scales based on the scale of the boundary image includes;

[0110] Sort the boundary images with the same bounding box number based on their order in the video stream. Based on the sorting order of the boundary images, compare the scale of the first boundary image in the boundary image sorting with the first scale threshold in the scale sorting;

[0111] If the scale of the first boundary image is greater than the first scale threshold, continue to compare with the next scale threshold until the boundary image is less than or equal to the nth scale threshold. Match the feature - weight group corresponding to the nth scale threshold for the first boundary image, and assign the corresponding weight parameters to each image feature in the boundary image;

[0112] If the scale of the first boundary image is less than or equal to the first scale threshold, match the feature - weight group corresponding to the first scale threshold for the first boundary image, and assign the corresponding weight parameters to each image feature in the boundary image;

[0113] Compare the scale of the ath boundary image in the boundary image sorting with the bth scale threshold matched by the (a - 1)th boundary image;

[0114] In some embodiments of the present invention, a is a positive integer greater than 1, b is a positive integer, and if a equals 2, b may equal 1.

[0115] If the scale of the ath boundary image is greater than the bth scale threshold, continue to compare with the (b + 1)th scale threshold until the boundary image is less than or equal to the mth scale threshold. Match the feature - weight group corresponding to the mth scale threshold for the ath boundary image, and assign the corresponding weight parameters to each image feature in the boundary image;

[0116] In some embodiments of the present invention, m ≥ b + 1.

[0117] If the scale of the a-th boundary image is less than or equal to the b-th scale threshold, then match the a-th boundary image with the feature-weight group corresponding to the b-th scale threshold, and assign the corresponding weight parameter to each image feature in this boundary image.

[0118] With the above solution, since the movement of the clothing on the production line is gradually approaching or gradually moving away from the camera, the clothing captured by the camera is gradually getting larger or smaller. Since the boundary images with the same bounding box number are sorted based on their order in the video stream, and the multiple boundary images in sequential order are in an order of gradually getting larger or gradually getting smaller, therefore, the scale of the a-th boundary image in the boundary image sorting does not need to be compared with the scale threshold before the b-th scale threshold matched by the (a - 1)-th boundary image. Compared with the method of comparing all intervals for each match, the processing efficiency can be significantly improved. Especially in the application scenario of the present invention, real-time processing needs to be ensured, and the processing efficiency needs to be guaranteed.

[0119] In some embodiments of the present invention, if the clothing in the scene is approaching the camera from far to near, the sorting of the boundary images is such that the smaller the scale, the smaller the number, and the number increases as the scale threshold value increases.

[0120] Assign different weights to different image features. Since the scales of the images at different times are different, the contour features of the smaller-scale images are more obvious, while the texture features of the larger-scale images are more obvious. Therefore, it is necessary to extract multi-scale information to increase the receptive field of the model. However, if the images are directly scaled to the same scale, the images may become blurred or even serrated, resulting in the inapplicability of the image features. Use time information to perform weighting in terms of scale.

[0121] Existing network models all generate high-resolution details at the spatial local points of low-resolution feature maps, while the method proposed by the present invention can generate details from all features and can determine whether two images with obvious differences have consistent highly refined features. The attention mechanism proposed by the present invention extracts and fuses features, which is to perform weighted combination of the extracted image features according to time information and image information. The key lies in the calculation of each information weight value, considering different weight parameters for each input element, so as to pay more attention to the parts similar to the input elements and suppress other useless information. The greatest advantage is that it can consider global and local connections in one step and can perform parallel computing, and can also perform real-time hanging clothing classification under the pressure of a huge amount of data in the logistics center.

[0122] The attention mechanism is applied to each convolutional layer to weight the feature information in the image, and then the output of the original layer and the weighted feature information are fused to form new features with more obvious discrimination. Another key point is that the extracted information needs to be scaled. Since the images are at different scales, the extracted information should also be normalized and scaled to a unified scale, and then the normalized information is input into the fusion module, which can fuse the feature information of each image to obtain more discriminative features.

[0123] The overall network structure is as follows Figure 3 As shown, the input is the frame sequence image after image preprocessing, which encodes both the image features and the time information, and then the encoded image features are weighted through the attention mechanism. The weighted features are fused and then sent to the final convolutional neural network classifier.

[0124] In some embodiments of the present invention, the convolutional neural network classifier is trained based on the following loss function formula:

[0125]

[0126] L represents the loss function value, f represents the f-th fused image, F represents the total number of fused images, is the value obtained by normalizing the feature vector x corresponding to the image features of the f-th fused image f where e represents the Euler number, g represents the g-th image feature, s g represents the eigenvalue of the g-th image feature, and G represents the total number of categories of image features. is the value obtained by normalizing the image feature s in the fused images of the same category corresponding to the f-th fused image g where γ represents the weight of the feature distance function, and the feature distance function is represents the center point of the feature vectors of the fused images of the same category corresponding to the fused image f.

[0127] If the category of the fused image f is pants, the fused images of the same category corresponding to the fused image f are images of the same pants category. There may be differences in their image features, but they belong to the same pants category.

[0128] In this solution, the convolutional neural network classifier can be trained using the above loss function, or it can be trained using a general training method.

[0129] Based on the loss function, update the parameter values in the convolutional neural network classifier to complete the training.

[0130] In some embodiments of the present invention, as the γ value increases, the distance between the same type of features and the center point becomes smaller, and the distribution of the same type of samples becomes more compact. At the same time, the distance between the features of different categories becomes larger, and the distinguishability of different category samples becomes higher, thereby improving the classification accuracy. Finally, the clothing is conveyed to the appropriate stalls according to the categories output by the classifier to complete the classification process.

[0131] In some embodiments of the present invention, each fused image includes a plurality of image features, and a feature vector of the fused image is constituted by the plurality of image features.

[0132] In some embodiments of the present invention, the step of dividing the image data of the video stream into a plurality of first image frames further includes denoising the divided first image frames, and the denoising step is to use Wiener filtering to remove the first image frames with noise in the plurality of first image frames.

[0133] In some embodiments of the present invention, the step of cropping the first image frame based on the bounding box to obtain a boundary image includes de-highlighting the boundary image, and the de-highlighting step includes,

[0134] Obtaining a transformation matrix corresponding to a plurality of boundary images with the same bounding box number based on the SURF algorithm;

[0135] Dividing every α of the plurality of boundary images with the same bounding box number into a fusion group, and performing alignment processing on the boundary images in the same fusion group based on the transformation matrix;

[0136] Sequentially adopting the methods of taking the minimum pixel gray value, taking the gray average value, Gaussian difference and median combination for the boundary images in the same fusion group to fuse the boundary images in the same fusion group into the same boundary image.

[0137] Adopting the above solution, since the clothing on the production line is usually provided with plastic packaging, which is prone to high-light reflection, this solution greatly guarantees the features of the image through de-highlighting by the SURF algorithm, and can eliminate the high light by combining the minimum pixel gray value, taking the gray average value, Gaussian difference and median, and can process in real time to improve the accuracy of clothing classification.

[0138] In some embodiments of the present invention, α can be 1, 2 or 3, etc.

[0139] The present invention classifies the hanging clothing with plastic packaging, and reduces the high-light reflection caused by the plastic packaging to the lowest level through image preprocessing to improve the accuracy of subsequent clothing classification;

[0140] During the transportation of hanging clothes, vibrations are generated due to the high-speed movement of the conveying equipment, resulting in blurred images. The present invention extracts multiple images of different scales from the video stream, calculates their features respectively, and uses the attention mechanism to fuse the extracted features into features with higher discrimination, reducing the negative impact of excessive equipment speed and vibrations on image recognition.

[0141] The present invention can simultaneously track multiple pieces of clothing on the conveying equipment, perform positioning, numbering differentiation, and identification respectively, effectively preventing the situation where clothes cannot be classified in time due to too fast conveying speed or too small distance between clothes.

[0142] Aiming at the problems of low automation, slow processing speed, and serious waste of human resources existing in the prior art, the present invention uses a camera to record the video stream during the transportation of hanging clothes, extracts multiple images from the video stream, locates, numbers the bounding boxes of multiple pieces of clothing in the images, differentiates and identifies them, extracts features of clothes with the same number, and uses the attention mechanism to fuse multiple images of different scales to obtain features with stronger discrimination for classification, thereby improving the classification speed and accuracy of hanging clothes in the logistics center.

[0143] An embodiment of the present invention also provides a device for classifying hanging clothes in a logistics center based on a video stream. The device includes a computer device, the computer device includes a processor and a memory, computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method described above.

[0144] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for classifying hanging clothes in a logistics center based on a video stream described above are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0145] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link.

[0146] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0147] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0148] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and variations can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A classification method for hanging clothes in a logistics center based on video streams, characterized in that, Receive the input video stream, identify the clothing in the image data of the video stream based on a preset deep learning model, and mark the bounding box of the clothing in the image data of the video stream; Number the bounding boxes of the clothing sequentially based on the order of appearance of the clothing in the image data of the video stream; Divide the image data of the video stream into multiple first image frames, and crop the first image frames based on the bounding boxes to obtain boundary images; For the boundary images with the same bounding box number, extract the image features in the boundary images. Based on the scale of the boundary images, assign different weights to the image features in the boundary images of different scales, divide multiple scale thresholds, arrange the scale thresholds in order according to the numerical size of the multiple scale thresholds, and each scale threshold is correspondingly set with a feature-weight group. In the feature-weight group, weight parameters are set for each image feature. Sort the boundary images with the same bounding box number according to their order in the video stream; Based on the sorting order of the boundary images, compare the scale of the first boundary image in the boundary image sorting with the first scale threshold in the scale sorting; If the scale of the first boundary image is greater than the first scale threshold, continue to compare with the next scale threshold until the boundary image is less than or equal to the nth scale threshold, match the feature-weight group corresponding to the nth scale threshold for the first boundary image, and assign the corresponding weight parameters to each image feature in this boundary image; If the scale of the first boundary image is less than or equal to the first scale threshold, match the feature-weight group corresponding to the first scale threshold for the first boundary image, and assign the corresponding weight parameters to each image feature in this boundary image; Compare the scale of the a-th boundary image in the boundary image sorting with the b-th scale threshold matched by the (a - 1)-th boundary image; If the scale of the a-th boundary image is greater than the b-th scale threshold, continue to compare with the (b + 1)-th scale threshold until the boundary image is less than or equal to the mth scale threshold, match the feature-weight group corresponding to the mth scale threshold for the a-th boundary image, and assign the corresponding weight parameters to each image feature in this boundary image; If the scale of the a-th boundary image is less than or equal to the b-th scale threshold, match the feature-weight group corresponding to the b-th scale threshold for the a-th boundary image, and assign the corresponding weight parameters to each image feature in this boundary image; Perform feature fusion on the boundary images with the same bounding box number based on the attention mechanism to obtain a fused image; Input the fused image into a preset convolutional neural network classifier to obtain the category of the fused image.

2. The method for classifying hanging clothes in a logistics center based on a video stream according to claim 1, wherein In the step of marking the bounding box of the clothing in the image data of the video stream, Generate a bounding box according to the scale of the clothing in the image data, and the image of the clothing is within the range framed by the bounding box; The bounding box becomes larger or smaller as the scale of the clothing it frames changes in the image data of the video stream.

3. The method for classifying hanging clothes in a logistics center based on a video stream according to claim 1 or 2, characterized in that, In the step of marking the bounding box of the clothing in the image data of the video stream, based on a preset bounding box threshold, compare the size of the current bounding box with the bounding box threshold in real time. If the size of the current bounding box is not within the range of the bounding box threshold, the bounding box is not displayed.

4. The method for classifying hanging clothes in a logistics center based on a video stream according to claim 1, wherein The step of dividing the image data of the video stream into multiple first image frames includes dividing the image data of the video stream into multiple initial image frames according to the frames of the image data; converting the initial image frames into grayscale images, and calculating the grayscale centroid of each grayscale image based on the grayscale values of each pixel point in the grayscale images; calculating the average grayscale centroid based on the grayscale centroids of each grayscale image, and respectively calculating the distances between the grayscale centroids of each grayscale image and the average grayscale centroid; screening out the first preset number of grayscale images with relatively close distances from all the grayscale images based on the distances between the grayscale centroids and the average grayscale centroid as the first image frames.

5. The method for classifying hanging garments in a logistics center based on a video stream according to claim 4, wherein Calculating the grayscale centroid of each grayscale image based on the grayscale values of each pixel point in the grayscale image according to the following formula: Combine x according to the above formula c , y c to obtain the coordinates (x c , y c ) of the gray centroid. In the formula, x c represents the abscissa of the gray centroid, and y c represents the ordinate of the gray centroid. x ij is the pixel gray value of the pixel at the i-th row and j-th column of the grayscale image with size M*N. M is the total number of pixel rows of the grayscale image, and N is the total number of pixel columns of the grayscale image.

6. The method for classifying hanging clothes in a logistics center based on a video stream according to claim 1, wherein Dividing into multiple scale ranges, each scale range corresponding to a preset feature-weight group, and setting weight parameters for each image feature in the feature-weight group. The step of assigning different weights to the features in the boundary images of different scales based on the scale of the boundary image includes determining the scale range corresponding to the boundary image based on the scale of the boundary image, matching the corresponding feature-weight group for the boundary image according to the scale range, and assigning the corresponding weight parameters to each image feature in the boundary image.

7. The method for classifying hanging clothes in a logistics center based on a video stream according to claim 1, wherein The step of cropping the first image frame based on the boundary box to obtain the boundary image includes performing highlight removal processing on the boundary image. The steps of the highlight removal processing include obtaining the transformation matrices corresponding to multiple boundary images with the same boundary box number based on the SURF algorithm; dividing every α of the multiple boundary images with the same boundary box number into a fusion group, and performing alignment processing on the boundary images in the same fusion group based on the transformation matrices; sequentially adopting the combination of taking the minimum pixel grayscale value, taking the grayscale average value, Gaussian difference, and median to fuse the boundary images in the same fusion group into the same boundary image.

8. The method for classifying hanging garments in a logistics center based on a video stream according to claim 1, wherein, Training the convolutional neural network classifier based on the following loss function formula: Let \(L\) denote the loss function value, \(f\) denote the \(f\)-th fused image, and \(F\) denote the total number of fused images. is the value obtained by normalizing the feature vector \(x\) corresponding to the image features of the \(f\)-th fused image. f Here, \(e\) represents the Euler's number, \(g\) represents the \(g\)-th image feature, and g \(s\) represents the eigenvalue of the \(g\)-th image feature of the fused images of the same class corresponding to the \(f\)-th fused image. \(G\) represents the total number of classes of image features. is the value obtained by normalizing the eigenvalue of the \(g\)-th image feature of the fused images of the same class corresponding to the \(f\)-th fused image. \(\gamma\) represents the weight of the feature distance function, and the feature distance function is represents the center point of the feature vectors of the fused images of the same class corresponding to the fused image \(f\).

9. A clothing classification device for hanging and mounting in a logistics center based on a video stream, characterized in that, The device includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Fine-grained image detection method and system based on improved RA-CNN

    CN112052876A