Logistics sorting algorithm based on RCNN

Through the logistics sorting algorithm based on RCNN, the problems of high complexity and insufficient adaptability of the logistics sorting algorithm in the prior art are solved, and fast and accurate goods identification and classification are achieved, labor costs and error rates are reduced, and logistics sorting efficiency is improved.

CN120198734AActive Publication Date: 2025-06-24NANJING ZSPLAT TECH
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510335826.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing logistics sorting algorithm is complex and lacks adaptability, making it difficult to replace manual sorting, resulting in low efficiency and high error, increasing the labor cost of the logistics industry.

Method used

Using the logistics sorting algorithm based on RCNN, we collect and integrate large-scale parcel image data, build CNN and ResNet models, and combine the pyramid pooling layer to achieve the extraction and classification of cargo size, packaging materials and outer packaging integrity.

Benefits of technology

This algorithm can accurately identify and classify goods in a short time, improve processing speed, reduce labor costs, reduce error rates of manual sorting, and improve the efficiency and accuracy of logistics sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198734A_ABST
    Figure CN120198734A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of logistics sorting and deep learning, and particularly discloses an RCNN-based logistics sorting algorithm, and a model is formed by orderly stacking a plurality of convolutional layers, three residual blocks, a pyramid pooling layer and a Softmax layer. The problems of gradient disappearance and gradient explosion occurring in the deep convolutional neural network training process are solved by introducing ResNet, so that the network can learn data features more easily in the training process, and the performance and stability of the network are improved. According to the algorithm, an RCNN model is trained by collecting and integrating a large-scale wrapped image data set, model parameters are systematically adjusted in the process, and the model is iteratively optimized, so that the model parameters with optimal performance are obtained. The parcels to be sorted can be effectively classified according to the parcel size, the packaging material and the outer package integrity, so that the purposes of improving the sorting efficiency and optimizing the overall logistics operation effect are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a logistics sorting algorithm based on RCNN, which combines deep learning with logistics sorting and belongs to the technical field of deep learning applications. Background Art

[0002] Logistics sorting is widely used in multiple fields such as e-commerce, express delivery, and manufacturing. It significantly improves the efficiency and accuracy of logistics processing by quickly and accurately classifying and sorting goods. It uses advanced automation technologies and equipment to sort the items to be sorted to the corresponding destinations according to certain rules, so as to achieve efficient, accurate, and fast transmission of logistics.

[0003] Among them, the logistics sorting algorithm is a key technology in the logistics industry for optimizing the sorting process and improving the sorting efficiency. It is implemented through computer programs and is a series of rules and calculation methods for determining the sorting order, sorting path, and sorting strategy of goods, aiming to improve the sorting efficiency, reduce sorting errors, and optimize the overall operation effect of logistics.

[0004] Therefore, it is necessary to invent an algorithm with low complexity and strong adaptability that can replace manual labor and continuously, efficiently, and accurately complete sorting tasks, reducing labor costs. With the continuous progress of technology and the continuous expansion of application scenarios, the logistics sorting algorithm will continue to play an important role in promoting the continuous development and upgrading of the logistics industry. Summary of the Invention

[0005] The purpose of the present invention is to invent an algorithm that can optimize the sorting process, enabling goods to be accurately identified and classified in a short time, thereby improving the processing speed. This algorithm can extract and classify features such as the size of goods, packaging materials, and the integrity of the outer packaging, and can replace manual sorting work, reducing the labor cost of the logistics industry and at the same time reducing the error rate of manual sorting.

[0006] The technical idea of the present invention is as follows: First, collect and integrate a large-scale sample data set containing original package images, and at the same time make the image matrix sizes consistent through central cropping, and then perform normalization processing on the image data matrix. Then, build a training model based on CNN and ResNet to extract image features. After feature extraction, connect a pyramid pooling layer to complete the construction of the RCNN model. Finally, train the RCNN model with a large-scale data set and save the model with the best performance.

[0007] The present invention provides a logistics sorting algorithm based on RCNN (Residual Convolutional Neural Network), including the following steps:

[0008] Step 1: Image data collection and preprocessing,

[0009] The dataset consists of a large number of samples containing original parcel images, which show a high degree of diversity in terms of size, packaging materials, and the integrity of the outer packaging. For images with inconsistent sizes, central cropping is used to make the size of the image matrix consistent, and then the image data matrix is normalized so that the features between different dimensions are numerically comparable to prevent the impact of excessive numerical differences on the training model.

[0010] Step 2: Building the RCNN model structure;

[0011] First, a feature extraction module is built based on the convolutional neural network structure, which is used to extract features from the image data matrix. This module is composed of multiple convolutional layers, pooling layers, batch normalization layers, and Sigmoid layers stacked in an orderly manner. And "ResNet" is introduced in the feature extraction module to solve the problems of gradient disappearance and gradient explosion during the training process of the deep convolutional neural network, making it easier for the network to learn the features of the data during the training process, improving the performance and stability of the network. Then, a pyramid pooling layer is built after the feature extraction module to output image features and perform classification through the Softmax layer.

[0012] Step 3: Training the RCNN model;

[0013] The RCNN model is trained by collecting and integrating a large-scale dataset. During the training process, the model parameter configuration is systematically adjusted, and the model is optimized through iteration, and finally the model version with the best performance is saved.

[0014] As a further technical solution of the present invention, in the central cropping method in Step 1, the position and size of the cropping area are determined according to the size of the original image and the target cropping size, and the formula is as follows:

[0015]

[0016] where (left, top) represents the coordinates of the upper left corner of the cropping area, (right, bottom) represents the coordinates of the lower right corner of the cropping area, and (width, height) and (new_width, new_height) represent the image sizes before and after cropping respectively.

[0017] Furthermore, in Step 1, each pixel value is linearly transformed through normalization processing so that the original data is mapped to the interval [a, b] (usually [0, 1] or [-1, 1]), and the preprocessed data is limited within a certain range to eliminate the adverse effects caused by singular sample data. The present invention adopts the maximum-minimum normalization method, and the formula is as follows:

[0018]

[0019] Among them, x i represents the elements of the image data matrix, and max(x) and min(x) represent the maximum and minimum values of the data matrix elements respectively.

[0020] Furthermore, in the second step, the convolutional neural network extracts features layer by layer by stacking a series of layers. As the number of network layers increases, it brings problems of gradient disappearance and gradient explosion.

[0021] By introducing a residual block, that is, adding a skip connection between the input and output, to address this problem.

[0022] The input feature is represented as x, and the output feature of one or more layers is represented as F(x). The formula of F(x)+x can be implemented by a feedforward neural network with a shortcut connection.

[0023] Furthermore, in the second step, the pyramid pooling layer maps the local features extracted by the upper layer to the label space of the sample to form a global feature representation for subsequent classification. It includes multi-scale max pooling, Flatten, Dropout, and Linear.

[0024] Max pooling significantly reduces the size of the feature map by extracting the maximum data in the specified window. Its formula is as follows:

[0025]

[0026] Among them, the input data form of max pooling is (C, H in , W in ), and the output data form is (C, H out , W out ). C represents the number of channels, padding is for filling, dilatin is for dilation, kernel_size is the window size, and stride is the step size.

[0027] Linear performs a linear transformation on the input feature through the weight matrix A. Its formula is as follows:

[0028] y = xA T +b

[0029] Among them, x represents the input feature vector, A represents the weight matrix, b represents the bias vector, and y represents the output feature vector.

[0030] Further, in the third step, the model is first pre-trained on a large-scale dataset to obtain initial weights and feature extraction capabilities. Then, the pre-processed dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. The training set is used to train the model, and the model parameters are continuously adjusted through the data samples in the training set, enabling the model to better learn the features of the data and minimizing the training error. The validation set is used to select the optimal hyperparameters of the model, such as the number of training epochs, the learning rate, and the BatchSize, etc., and to preliminarily evaluate the model performance. The test set is used to finally evaluate the generalization ability of the model to verify the performance of the model on unseen data.

[0031] Cross-entropy loss is used as the loss function for model training and testing, Sigmoid is used as the activation function, SGD is used as the optimizer, the number of training epochs is 100, the learning rate is 0.001, and the BatchSize is 64.

[0032] The above parameters interact with each other and jointly affect the performance of the model. The optimal parameter configuration is found through experiments and verification.

[0033] After iterative training, the model version with the best performance is finally saved.

[0034] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the RCNN-based logistics sorting algorithm described above is implemented.

[0035] A computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the RCNN-based logistics sorting algorithm described above is implemented.

[0036] Compared with the prior art, the above technical solutions adopted by the present invention have the following advantages:

[0037] 1. The present invention processes images with inconsistent sizes through central cropping to make the image matrix sizes consistent, avoiding incorrect processing caused by inconsistent image sizes and enhancing the applicability of the algorithm.

[0038] 2. The present invention performs normalization processing on the image data matrix, enabling the features between different dimensions to have a certain comparability in terms of numerical values, and preventing the impact of overly large or small numerical values on the subsequent training model.

[0039] 3. The present invention introduces "ResNet" in the deep convolutional neural network to solve the problems of gradient disappearance and gradient explosion during the training process, enabling the network to more easily learn the features of the data during the training process and improving the performance and stability of the network.

[0040] 4. The present invention uses a pyramid pooling layer to perform pooling operations on the feature map at different scales, obtain multi-scale information, and increase the receptive field of neurons. For packages of different sizes that may have different feature scales, the pyramid pooling layer can also effectively capture these features of different scales, enabling the model to have good representation capabilities for targets of different sizes. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 FIG. is a flowchart of an RCNN-based logistics sorting algorithm disclosed by the present invention applied to an actual logistics sorting scenario.

[0042] Figure 2 FIG. is a flowchart of an RCNN-based logistics sorting algorithm disclosed by the present invention.

[0043] Figure 3 FIG. is a mathematical model of "ResNet" introduced in an RCNN-based logistics sorting algorithm disclosed by the present invention.

[0044] Figure 4 FIG. is a model structure diagram of an RCNN-based logistics sorting algorithm disclosed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To deepen the understanding of the present invention, the following detailed description is given in conjunction with the accompanying drawings for this embodiment.

[0046] Embodiment: An RCNN (Residual Convolutional Neural Network)-based logistics sorting algorithm includes the following steps:

[0047] Step 1: Image data collection and preprocessing.

[0048] The data set consists of a large number of samples containing original package images, and these images show a high degree of diversity in terms of size, packaging material, and outer package integrity. For images with inconsistent sizes, central cropping is used to make the image matrix sizes the same, and then the image data matrix is normalized to make the features between different dimensions comparable numerically and prevent the impact of excessive numerical differences on the training model.

[0049] Step 2: Construction of the RCNN model structure.

[0050] First, a feature extraction module is built based on the convolutional neural network structure, which is used to extract features from the image data matrix. This module is composed of multiple convolutional layers, pooling layers, batch normalization layers, and Sigmoid layers stacked in an orderly manner. And "ResNet" is introduced in the feature extraction module to solve the problems of gradient disappearance and gradient explosion that occur during the training process of the deep convolutional neural network, making it easier for the network to learn the features of the data during the training process, improving the performance and stability of the network. Then, a pyramid pooling layer is built after the feature extraction module to output image features and classify them through the Softmax layer.

[0051] Step 3: Training of the RCNN model;

[0052] The RCNN model is trained by collecting and integrating a large-scale dataset. During the training process, the model parameter configuration is systematically adjusted, and the model is optimized through iteration. Finally, the model version with the best performance is saved.

[0053] As a further technical solution of the present invention, in the central cropping method in Step 1, the position and size of the cropping area are determined according to the size of the original image and the target cropping size. The formula is as follows:

[0054]

[0055] Among them, (left, top) represents the coordinates of the upper left corner of the cropping area, (right, bottom) represents the coordinates of the lower right corner of the cropping area, and (width, height) and (new_width, new_height) represent the image sizes before and after cropping respectively.

[0056] Furthermore, in Step 1, each pixel value is linearly transformed through normalization processing, so that the original data is mapped to the interval [a, b] (usually [0, 1] or [-1, 1]), and the preprocessed data is limited within a certain range, thereby eliminating the adverse effects caused by singular sample data. The present invention adopts the maximum-minimum normalization method, and the formula is as follows:

[0057]

[0058] where x i represents the element of the image data matrix, and max(x) and min(x) represent the maximum and minimum values of the data matrix elements respectively.

[0059] Furthermore, in Step 2, the convolutional neural network extracts features layer by layer by stacking a series of layers. As the number of network layers increases, the problems of gradient disappearance and gradient explosion will occur.

[0060] This problem is addressed by introducing a residual block, i.e., adding a skip connection between the input and the output.

[0061] The input feature is denoted as x, and the output feature of one or more layers is denoted as F(x). The formula F(x) + x can be implemented by a feed-forward neural network with a shortcut connection.

[0062] Furthermore, in the second step, the pyramid pooling layer maps the local features extracted by the upper layer to the label space of the sample to form a global feature representation for subsequent classification. It includes multi-scale max pooling, Flatten, Dropout, and Linear.

[0063] Max pooling significantly reduces the size of the feature map by extracting the maximum data within a specified window. Its formula is as follows:

[0064]

[0065] Among them, the input data form of max pooling is (C, H in , W in ), and the output data form is (C, H out , W out ). C represents the number of channels, padding is for padding, dilation is for dilation, kernel_size is the window size, and stride is the step size.

[0066] Linear performs a linear transformation on the input feature using the weight matrix A. Its formula is as follows:

[0067] y = xA T + b

[0068] Where x represents the input feature vector, A represents the weight matrix, b represents the bias vector, and y represents the output feature vector.

[0069] Furthermore, in the third step, first pre-train the model on a large-scale dataset to obtain initial weights and feature extraction capabilities. Then divide the pre-processed dataset into a training set, a validation set, and a test set in a ratio of 8:1:1. Use the training set to train the model, continuously adjust the model parameters through the data samples in the training set so that the model can better learn the features of the data to minimize the training error. Use the validation set to select the best hyperparameters of the model, such as the number of training times, learning rate, and BatchSize, etc., and preliminarily evaluate the model performance. The test set is used to finally evaluate the generalization ability of the model to verify the performance of the model on unseen data.

[0070] Cross-entropy loss is used as the loss function for model training and testing, Sigmoid is used as the activation function, SGD is used as the optimizer, the number of training times is 100, the learning rate is 0.001, and the BatchSize is 64.

[0071] The above parameters interact with each other and jointly affect the performance of the model. The best parameter configuration is found through experiments and verification.

[0072] After iterative training, the model version with the best performance is finally saved.

[0073] A logistics sorting algorithm based on RCNN disclosed by the present invention is applied to the actual logistics sorting process as follows: First, the images of the packages to be sorted are collected by peripheral devices, and then the integrity, material, and size features of the package outer packaging are extracted and classified through the algorithm. Packages with extremely damaged outer packaging are transported to sorting port 1 and withdrawn from the sorting process, and then enter the sorting process again after manual processing of the outer packaging. For packages with damaged outer packaging that do not affect transportation, soft-packaged packages are transported to sorting port 2, and the remaining packages are transported to sorting ports 3 and 4 respectively according to their sizes, so as to achieve the purpose of quickly sorting packages.

[0074] As Figure 2 shown, a logistics sorting algorithm based on RCNN disclosed by the present invention mainly includes three steps: image data collection, preprocessing, RCNN model structure construction, and RCNN model training, and finally saves the model with the best performance.

[0075] As Figure 3 shown, a logistics sorting algorithm based on RCNN disclosed by the present invention introduces "ResNet" to address the problems of gradient disappearance and gradient explosion. The input feature is represented as x, and the output feature of one or more layers is represented as F(x). The formula of F(x)+x is realized through a feedforward neural network with shortcut connections, that is, the input and output are directly added.

[0076] As Figure 4 shown, a logistics sorting algorithm based on RCNN disclosed by the present invention has a model composed of multiple convolutional layers, three residual blocks, a pyramid pooling layer, and a Softmax layer. Each residual block contains three convolutional layers, three batch normalization layers, and a Sigmoid activation function.

[0077] As shown in Table 1, a logistics sorting algorithm based on RCNN disclosed by the present invention obtains the optimal model parameter configuration through systematic adjustment of parameters during the training process and iterative optimization.

[0078] Table 1: RCNN model parameter configuration and output size of each module,

[0079]

[0080] It should be noted that the above embodiments are not intended to limit the protection scope of the present invention, and equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. The logistics sorting algorithm based on RCNN is characterized by , including the following steps: Step 1: Image data collection and preprocessing, For images of inconsistent sizes, use center cropping to make the image matrix size consistent, and then normalize the image data matrix to make the features between different dimensions numerically comparable to a certain extent, to prevent the large numerical differences from affecting the training model. Step 2: RCNN model structure construction, First, a feature extraction module is built based on the convolutional neural network structure, which is used to extract features from the image data matrix. The feature extraction module is composed of multiple convolutional layers, pooling layers, batch normalization layers and Sigmoid layers stacked in an orderly manner. "ResNet" is introduced into the feature extraction module to solve the gradient vanishing and gradient exploding problems that occur during the training process of deep convolutional neural networks. Then, a pyramid pooling layer is built after the feature extraction module to output image features and classify them through the Softmax layer. Step 3: RCNN model training, The RCNN model is trained by collecting and integrating large-scale data sets, systematically adjusting the model parameters during the training process, optimizing the model through iteration, and finally saving the model parameters with the best performance.

2. The RCNN-based logistics sorting algorithm according to claim 1 is characterized in that: In step 1, the center cropping method, the position and size of the cropping area are determined according to the size of the original image and the target cropping size. The formula is as follows: Among them, (left, top) represents the coordinates of the upper left corner of the cropped area, (right, bottom) represents the coordinates of the lower right corner of the cropped area, (width, height) and (new_width, new_height) represent the image sizes before and after cropping, respectively.

3. The RCNN-based logistics sorting algorithm according to claim 1 is characterized in that: In step 1, each pixel value is linearly transformed through normalization processing, so that the original data is mapped to the [a, b] interval, and the preprocessed data is limited to a certain range, thereby eliminating the adverse effects caused by singular sample data. The maximum-minimum normalization method is used, and the formula is as follows: Among them, x i Represents the image data matrix element, max(x) and min(x) represent the maximum and minimum values ​​of the data matrix element respectively.

4. The RCNN-based logistics sorting algorithm according to claim 1 is characterized in that: In step 2, Introducing the residual block, that is, adding a skip connection between the input and output, The input feature is represented as x, and the output feature of one or more layers is represented as F(x). The formula of F(x)+x is implemented by a feedforward neural network with shortcut connections. In step 2, the pyramid pooling layer maps the local features extracted by the upper layer to the sample's label space to form a global feature representation, which includes multi-scale maximum pooling, Flatten, Dropout, and Linear. The maximum pooling extracts the maximum data of the specified window, and its formula is as follows: Among them, the input data form of the maximum pooling is (C,H in ,W in ), the output data format is (C,H out ,W out ), C represents the number of channels, padding represents filling, dilation represents expansion, kernel_size represents window size, and stride represents step size. Linear transforms the input features linearly through the weight matrix A. The formula is as follows: y=xA T +b Among them, x represents the input feature vector, A represents the weight matrix, b represents the bias vector, and y represents the output feature vector.

5. The RCNN-based logistics sorting algorithm according to claim 1 is characterized in that: In step three, the model is first pre-trained on a large-scale data set to obtain preliminary weights and feature extraction capabilities. Then, the pre-processed data set is divided into a training set, a validation set, and a test set in a ratio of 8:1:1 to prevent the model from overfitting. The model is trained using the training set, and the model parameters are continuously adjusted through the data samples of the training set to enable the model to better learn the characteristics of the data to minimize the training error. The validation set is used to select the best hyperparameters of the model and preliminarily evaluate the model performance. The test set is used to finally evaluate the generalization ability of the model to verify the performance of the model on unseen data. Model training and testing use cross entropy loss as the loss function, Sigmoid as the activation function, SGD as the optimizer, 100 training times, a learning rate of 0.001, and a BatchSize of 64.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the RCNN-based logistics sorting algorithm as described in any one of claims 1 to 5 above is implemented.

7. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, the RCNN-based logistics sorting algorithm as described in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Remote traffic sign detection and recognition method based on F-RCNN

    CN110163187A

  • Attention mechanism CNN-based 5-day and 9-day incubated egg embryo image classification method

    CN110309880A

  • Non-uniform texture small defect detection method based on improved Faster R-CNN model

    CN111598861A

  • Chip defect image classification method based on ResNet network

    CN113076989A

  • Airport storehouse luggage retrieval method based on improved convolutional network

    CN114048345A