Method, device and storage medium for sorting items
By adopting sparse network structure and bounding box prediction module in the item sorting network, the problem of low sorting efficiency when processing complex item images in the prior art is solved, and more efficient item type identification and sorting is achieved.
Patent Information
- Application Number
- CN202111514338.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The prior art has low sorting efficiency when processing objects that contain images of high similarity, mutual occlusion, too far from the camera, or high speed rotation.
A sorting network with a sparse network structure is adopted to extract and predict the item images through the bounding box prediction module and the candidate area network to improve the accuracy and speed of object type recognition.
Through the use of sparse network structure, the sparseness and high computing performance of the network are maintained, and the speed of object type identification and sorting is improved.
Smart Images

Figure CN114359624B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent identification, and in particular to an object sorting method, device and storage medium. Background Art
[0002] Sorting of items is usually done manually, which has the problems of high labor costs and low efficiency. At present, there is also a method of using a neural network to identify the image of the items to be sorted, and controlling the execution robot to sort the items according to the classification and recognition results; but the efficiency of this method is mainly determined by the performance of the neural network. Some neural networks do not perform well when processing images of multiple items with high similarity, multiple items that occlude each other, items that are too far from the camera, and items that rotate at high speed, which greatly affects the sorting efficiency. Summary of the invention
[0003] The purpose of the present invention is to solve at least one of the technical problems existing in the prior art and to provide an item sorting method, device and storage medium.
[0004] The technical solution adopted by the present invention to solve the problem is:
[0005] A first aspect of the present invention provides an article sorting method, comprising:
[0006] Acquire a training object image and object annotation information corresponding to the training object image;
[0007] Inputting the training object image and the object annotation information into a sorting network for training until the classification accuracy of the sorting network reaches a preset target value, the sorting network comprising a bounding box prediction module, the bounding box prediction module comprising a plurality of sparse network structures, the sparse network structures being used to simultaneously perform convolution and re-aggregation on input signals at a plurality of scales;
[0008] Obtain an image of an object to be identified;
[0009] The image of the object to be identified is input into the trained sorting network for identification to obtain the identification result.
[0010] According to the first aspect of the present invention, for each of the training object images, the object labeling information is a name labeling of the object in the training object image.
[0011] According to the first aspect of the present invention, the inputting the training object image and the object labeling information into the sorting network for training until the classification accuracy of the sorting network reaches a preset target value includes:
[0012] Repeat the following steps until the classification accuracy of the sorting network reaches the preset target value;
[0013] Extracting features from the training object image to obtain a first feature;
[0014] Inputting the first feature into a candidate region network to generate a priori frame, and determining positive samples and negative samples from the priori frame;
[0015] Fusing the positive sample and the negative sample with the first feature to obtain a fused feature;
[0016] Inputting the fused features into the bounding box prediction module to obtain a bounding box and an item category within the bounding box;
[0017] A loss is calculated based on the bounding box, and parameters of the sorting network are adjusted based on the loss.
[0018] According to the first aspect of the present invention, the classification accuracy of the sorting network is a matching rate between the object category within the bounding box and the object labeling information.
[0019] According to the first aspect of the present invention, the determining of positive samples and negative samples from the priori frame includes:
[0020] Calculate the intersection-over-union ratio of the prior frame and the real frame;
[0021] The priori box corresponding to the IoU ratio that is greater than or equal to a first preset threshold is used as the positive sample, and the priori box corresponding to the IoU ratio that is less than a second preset threshold is used as the negative sample.
[0022] According to the first aspect of the present invention, before the step of inputting the fusion feature into the bounding box prediction module, the method includes:
[0023] The fused features are input into the region of interest pooling module for pooling processing.
[0024] According to a first aspect of the present invention, the sparse network structure includes a first sparse network structure; the first sparse network structure includes:
[0025] A first branch, wherein the first branch includes a plurality of first convolutional layers and a first activation function layer;
[0026] A second branch, wherein the second branch includes a plurality of second convolutional layers for dimensionality reduction, a second activation function layer, and a plurality of third convolutional layers, and the second activation function layer is connected between the plurality of the second convolutional layers and the plurality of the third convolutional layers;
[0027] A third branch, wherein the third branch comprises a plurality of fourth convolutional layers for dimensionality reduction, a third activation function layer and a plurality of fifth convolutional layers, wherein the third activation function layer is connected between the plurality of fourth convolutional layers and the plurality of fifth convolutional layers;
[0028] The fourth branch includes a first pooling layer and multiple sixth convolutional layers.
[0029] According to the first aspect of the present invention, the sparse network structure includes a second sparse network structure; the second sparse network structure includes:
[0030] A fifth branch, wherein the fifth branch includes a plurality of seventh convolutional layers and a fourth activation function layer;
[0031] A sixth branch, the sixth branch comprising a plurality of eighth convolutional layers for dimensionality reduction, a fifth activation function layer and a plurality of ninth convolutional layers, the fifth activation function layer being connected between the plurality of the eighth convolutional layers and the plurality of the ninth convolutional layers;
[0032] a seventh branch, the seventh branch comprising a plurality of tenth convolutional layers for dimensionality reduction, a sixth activation function layer and a plurality of eleventh convolutional layers, the sixth activation function layer being connected between the plurality of the tenth convolutional layers and the plurality of the eleventh convolutional layers;
[0033] An eighth branch, wherein the eighth branch includes a second pooling layer and multiple twelfth convolutional layers.
[0034] According to a second aspect of the present invention, an article sorting device comprises: a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the article sorting method as described in the first aspect of the present invention is implemented.
[0035] A third aspect of the present invention is a storage medium storing a computer program, wherein the computer program is used to execute the article sorting method as described in the first aspect of the present invention.
[0036] The above scheme has at least the following beneficial effects: item type recognition is performed based on item images through a sorting network with a sparse network structure, and the sparse network structure can simultaneously perform convolution and re-aggregation on the input signal at multiple sizes, which can maintain the sparsity of the network structure, making the size of the sorting network small and maintaining high computing performance; improving the speed of item type recognition and the speed of item sorting.
[0037] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The present invention will be further described below in conjunction with the accompanying drawings and examples.
[0039] Figure 1 is a flow chart of an article sorting method according to an embodiment of the present invention;
[0040] Figure 2 This is a flowchart of step S200. DETAILED DESCRIPTION
[0041] This section will describe in detail the specific embodiments of the present invention. The preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the accompanying drawings is to supplement the description of the text part of the specification with graphics, so that people can intuitively and vividly understand each technical feature and the overall technical solution of the present invention, but it cannot be understood as a limitation on the scope of protection of the present invention.
[0042] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0043] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0044] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0045] Reference Figure 1 , an embodiment of the first aspect of the present invention provides an item sorting method.
[0046] Item sorting methods include:
[0047] Step S100: acquiring a training object image and object annotation information corresponding to the training object image;
[0048] Step S200: inputting the training object image and the object annotation information into the sorting network for training until the classification accuracy of the sorting network reaches a preset target value, the sorting network includes a bounding box prediction module, the bounding box prediction module includes a plurality of sparse network structures, and the sparse network structures are used to simultaneously perform convolution and re-aggregation on the input signal at multiple scales;
[0049] Step S300, obtaining an image of an object to be identified;
[0050] Step S400: input the image of the object to be identified into the trained sorting network for identification to obtain the identification result.
[0051] In this embodiment, the type of item is identified based on the image of the item through a sorting network with a sparse network structure. The sparse network structure can simultaneously perform convolution and re-aggregation on the input signal at multiple sizes, which can maintain the sparsity of the network structure, making the size of the sorting network small while maintaining high computing performance, thereby improving the speed of item type identification and the speed of item sorting.
[0052] It should be noted that the overall sorting network is a fast RCNN network structure.
[0053] In certain embodiments of the first aspect of the present invention, for step S100, a training object image and object annotation information corresponding to the training object image are obtained. The training object image may be image data from a training database. These image data have been manually annotated; for each training object image, the object annotation information is the name of the object in the training object image.
[0054] Reference Figure 2 In some embodiments of the first aspect of the present invention, for step S200, inputting the training object image and the object annotation information into the sorting network for training until the classification accuracy of the sorting network reaches a preset target value includes:
[0055] Repeat steps S210 to S250 until the classification accuracy of the sorting network reaches a preset target value;
[0056] Step S210: extracting features from the training object image to obtain a first feature;
[0057] Step S220: input the first feature into the candidate region network to generate a priori box, and determine positive samples and negative samples from the priori box;
[0058] Step S230: Fusing the positive sample and the negative sample with the first feature to obtain a fused feature;
[0059] Step S240: input the fused features into a bounding box prediction module to obtain a bounding box and the category of the object in the bounding box;
[0060] Step S250: Calculate the loss according to the bounding box, and adjust the parameters of the sorting network according to the loss.
[0061] It should be noted that the classification accuracy of the sorting network is the matching rate of the object category in the bounding box and the object labeling information for all training object images input to the sorting network.
[0062] In step S210, feature extraction is performed on the training object image through the feature extraction network to obtain the first feature. Of course, before feature extraction, the training object image can be denoised, which is conducive to improving the recognition rate. Specifically, in this embodiment, the feature extraction network is a convolutional neural network. Of course, in other embodiments, the feature extraction network can be a commonly used graphic feature extraction network, such as a feature extraction network structure using a boundary feature method.
[0063] For step S220, the first feature is input into a region candidate network (RPN) to generate a priori box, and positive samples and negative samples are determined from the priori box.
[0064] In addition, positive samples and negative samples are determined from the prior frame, specifically including the following steps:
[0065] Calculate the intersection and union ratio of the prior frame and the real frame;
[0066] The priori boxes corresponding to the IoU ratio greater than or equal to the first preset threshold are taken as positive samples, and the priori boxes corresponding to the IoU ratio less than the second preset threshold are taken as negative samples.
[0067] It should be noted that Intersection over Union (IoU) is a standard for measuring the accuracy of detecting corresponding objects in a specific data set.
[0068] In addition, in this embodiment, the first preset threshold is 0.7, and the second preset threshold is 0.3.
[0069] For step 230, the positive sample and the negative sample are respectively fused with the first feature to obtain a fused feature. That is, the fused feature includes two types of fused features, namely, a fused feature obtained by fusion of the positive sample and the first feature, and a fused feature obtained by fusion of the negative sample and the first feature. The fused feature is input into the region of interest pooling module for pooling processing.
[0070] In step S240, the fused features processed by pooling are input into a bounding box prediction module to obtain a bounding box and an object category within the bounding box.
[0071] Specifically, the structure of the bounding box prediction module is as follows:
[0072] The first layer is a convolutional layer, a ReLU activation function layer, a max pooling layer, and a ReLU activation function layer connected in sequence. The convolutional layer of the first layer has a 7x7 convolution kernel, a stride of 2, and a padding of 3. The size of the max pooling layer of the first layer is 3x3, with a stride of 2.
[0073] The second layer is a convolutional layer, a ReLU activation function layer, a maximum pooling layer, and a ReLU activation function layer connected in sequence. The convolutional layer of the second layer has a 3x3 convolution kernel, a stride of 1, a padding of 1, and 192 channels. The maximum pooling layer of the second layer has a size of 3x3 and a stride of 2.
[0074] The third layer is a sparse network structure, which includes the first sparse network structure; the first sparse network structure includes:
[0075] The first branch includes 64 first convolution layers with 1x1 convolution kernels and a first activation function layer;
[0076] The second branch includes 96 second convolutional layers with 1x1 convolutional kernels for dimensionality reduction, a second activation function layer, and 128 third convolutional layers with 3x3 convolutional kernels, and the padding of the third convolutional layer is 1; the second activation function layer is connected between the 96 second convolutional layers and the 128 third convolutional layers;
[0077] The third branch includes 16 fourth convolutional layers with 1x1 convolutional kernels for dimensionality reduction, a third activation function layer, and 32 fifth convolutional layers with 5x5 convolutional kernels. The padding of the fifth convolutional layer is 2; the third activation function layer is connected between the 16 fourth convolutional layers and the 32 fifth convolutional layers;
[0078] The fourth branch includes the first pooling layer and 32 sixth convolution layers; the first pooling layer has a 3x3 convolution kernel, and the padding of the first pooling layer is 1; the sixth convolution layer has a 1x1 convolution kernel.
[0079] It should be noted that the first branch, the second branch, the third branch and the fourth branch of the first sparse network structure perform convolutions in parallel on multiple sizes at the same time, and then the outputs of the first branch, the second branch, the third branch and the fourth branch are aggregated.
[0080] In addition, the sparse network structure of the third layer includes the second sparse network structure; the second sparse network structure includes:
[0081] The fifth branch includes 128 seventh convolution layers with 1x1 convolution kernels and a fourth activation function layer;
[0082] The sixth branch includes 128 eighth convolutional layers with 1x1 convolutional kernels for dimensionality reduction, a fifth activation function layer, and 192 ninth convolutional layers with 3x3 convolutional kernels. The padding of the ninth convolutional layer is 1, and the fifth activation function layer is connected between the 128 eighth convolutional layers and the 192 ninth convolutional layers.
[0083] The seventh branch includes 32 tenth convolutional layers with 1x1 convolutional kernels for dimensionality reduction, a sixth activation function layer, and 96 eleventh convolutional layers with 5x5 convolutional kernels. The padding of the eleventh convolutional layer is 2, and the sixth activation function layer is connected between the 32 tenth convolutional layers and the 96 eleventh convolutional layers.
[0084] The eighth branch includes a second pooling layer and multiple twelfth convolutional layers. The second pooling layer has a 3x3 convolution kernel, and the padding of the second pooling layer is 1; the twelfth convolutional layer has a 1x1 convolution kernel, and the size of the output feature map is 28x28x64.
[0085] Similarly, the fifth branch, the sixth branch, the seventh branch, and the eighth branch of the second sparse network structure simultaneously perform convolutions in parallel on multiple sizes, and then the outputs of the fifth branch, the sixth branch, the seventh branch, and the eighth branch are aggregated.
[0086] This sparse network structure performs dimensionality reduction through a convolutional layer with a 1x1 convolution kernel in the early stages of multiple branches, which improves the training speed.
[0087] It should be noted that the bounding box prediction module also includes a fourth layer and a fifth layer, which are also sparse network structures, and the sparse network structures of the fourth layer and the fifth layer are structurally the same as the sparse network structure of the third layer. Of course, in other embodiments, the bounding box prediction module can also be set with a larger number of sparse network structures.
[0088] The bounding box prediction module performs bounding box prediction based on the fused features to obtain the bounding box and the category of the object in the bounding box.
[0089] For step S250, the loss is calculated based on the bounding box and the true box, the loss is back-propagated, and the parameters of the sorting network are adjusted based on the loss.
[0090] In step S300, an image of the object to be identified is captured in real time by a camera.
[0091] For step S400, the image of the object to be identified is input into the trained sorting network for identification to obtain the identification result. According to the identification result, the sorting mechanical device can be controlled to sort the objects corresponding to the image of the object to be identified. Before obtaining the final identification result, the target detection frame of the object can be corrected by using bounding box regression to improve the final detection accuracy.
[0092] In addition, an item sorting device includes a camera, a data processing module and an execution robot. The camera takes a real-time picture of the item to be sorted to obtain an image of the item to be identified, and feeds it back to the data processing module. The data processing module includes a trained sorting network, which recognizes the image of the item to be identified through the sorting network and outputs the identification result of the item classification. The execution robot is controlled to sort the item to be sorted according to the identification result.
[0093] An embodiment of the second aspect of the present invention provides an article sorting device.
[0094] The article sorting device comprises: a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the article sorting method according to the first aspect of the present invention is implemented.
[0095] The processor and the memory may be connected via a bus or other means.
[0096] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0097] The node embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] The third aspect of the present invention provides a storage medium storing a computer program for executing the article sorting method according to the first aspect of the present invention.
[0099] It will be appreciated by those skilled in the art that all or some of the steps and systems in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0100] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above implementation mode. As long as the technical effect of the present invention is achieved by the same means, it should belong to the protection scope of the present invention.
Claims
1. A method for sorting items, characterized in that: include: Acquire a training object image and object annotation information corresponding to the training object image; Inputting the training object image and the object annotation information into a sorting network for training until the classification accuracy of the sorting network reaches a preset target value, the sorting network comprising a bounding box prediction module, the bounding box prediction module comprising a plurality of sparse network structures, the sparse network structures being used to simultaneously perform convolution and re-aggregation on input signals at a plurality of scales; Obtain an image of an object to be identified; Inputting the image of the object to be identified into the trained sorting network for identification to obtain an identification result; The sparse network structure includes a first sparse network structure; The first sparse network structure includes: A first branch, wherein the first branch includes a plurality of first convolutional layers and a first activation function layer; A second branch, wherein the second branch includes a plurality of second convolutional layers for dimensionality reduction, a second activation function layer, and a plurality of third convolutional layers, and the second activation function layer is connected between the plurality of the second convolutional layers and the plurality of the third convolutional layers; A third branch, wherein the third branch comprises a plurality of fourth convolutional layers for dimensionality reduction, a third activation function layer and a plurality of fifth convolutional layers, wherein the third activation function layer is connected between the plurality of fourth convolutional layers and the plurality of fifth convolutional layers; A fourth branch, wherein the fourth branch includes a first pooling layer and a plurality of sixth convolutional layers; The sparse network structure includes a second sparse network structure; the second sparse network structure includes: A fifth branch, wherein the fifth branch includes a plurality of seventh convolutional layers and a fourth activation function layer; A sixth branch, the sixth branch comprising a plurality of eighth convolutional layers for dimensionality reduction, a fifth activation function layer and a plurality of ninth convolutional layers, the fifth activation function layer being connected between the plurality of the eighth convolutional layers and the plurality of the ninth convolutional layers; a seventh branch, the seventh branch comprising a plurality of tenth convolutional layers for dimensionality reduction, a sixth activation function layer and a plurality of eleventh convolutional layers, the sixth activation function layer being connected between the plurality of the tenth convolutional layers and the plurality of the eleventh convolutional layers; An eighth branch, wherein the eighth branch includes a second pooling layer and multiple twelfth convolutional layers.
2. The method for sorting items according to claim 1, characterized in that: For each of the training object images, the object labeling information is the name labeling of the object in the training object image.
3. The article sorting method according to claim 1, characterized in that: The training object image and the object labeling information are input into a sorting network for training, Until the classification accuracy of the sorting network reaches a preset target value, including: Repeat the following steps until the classification accuracy of the sorting network reaches the preset target value; Extracting features from the training object image to obtain a first feature; Inputting the first feature into a candidate region network to generate a priori frame, and determining positive samples and negative samples from the priori frame; Fusing the positive sample and the negative sample with the first feature to obtain a fused feature; Inputting the fused features into the bounding box prediction module to obtain a bounding box and an item category within the bounding box; A loss is calculated based on the bounding box, and parameters of the sorting network are adjusted based on the loss.
4. The article sorting method according to claim 3, characterized in that: The classification accuracy of the sorting network is a matching rate between the object category in the bounding box and the object labeling information.
5. The article sorting method according to claim 3, characterized in that: The determining of positive samples and negative samples from the prior frame includes: Calculate the intersection-over-union ratio of the prior frame and the real frame; The priori box corresponding to the IoU ratio that is greater than or equal to a first preset threshold is used as the positive sample, and the priori box corresponding to the IoU ratio that is less than a second preset threshold is used as the negative sample.
6. The article sorting method according to claim 3, characterized in that: Before the step of inputting the fusion feature into the bounding box prediction module, the method further comprises: The fused features are input into the region of interest pooling module for pooling processing.
7. An article sorting device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the article sorting method according to any one of claims 1 to 6 when executing the computer program.
8. A storage medium, characterized in that: A computer program is stored, and the computer program is used to execute the object sorting method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Article identification method and device and storage medium
CN113496241A
Method and system for detection and classification of license plates
US20170262723A1