Sorting robot multi-target identification system and method for storage environment

Through the multi-channel fusion CNN network structure model combined with RFID, barcode and machine vision recognition module, the diversity and real-time problems of product sorting in the warehousing environment are solved, and efficient product classification and sorting control are achieved.

CN120268664APending Publication Date: 2025-07-08HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510336181.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When sorting goods in the warehousing environment, the classification category range is small, and it is unable to adapt to diverse goods, the identification model is slow to respond, the real-time performance is not high, the degree of automation is insufficient, and the need for rapid logistics sorting is not met.

Method used

A multi-channel fusion CNN network structure model is adopted to combine RFID, barcode and machine vision recognition modules, and through data expansion and genetic algorithm optimization, a multi-objective recognition system is built, and a variety of sensor information fusion is used to classify and capture control.

Benefits of technology

It improves the accuracy and sorting efficiency of product classification, enhances the generalization performance and positioning accuracy of the model, and realizes efficient product sorting control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120268664A_ABST
    Figure CN120268664A_ABST
Patent Text Reader

Abstract

The invention discloses a sorting robot multi-target recognition system and method for a storage environment. The system comprises a commodity category recognition module and a robot sorting control module. The commodity category identification module comprises an RFID induction module, a bar code scanning module and a machine vision identification module; the RFID sensing module comprises an RFID sensor and is used for reading a signal of an article pasted with an RFID tag; the bar code scanning module comprises an industrial code scanner and is used for scanning articles with bar codes due to the fact that RFID tags cannot be pasted, the visual identification module is used for guiding a mechanical arm to grab the articles read and scanned by the RFID induction module and the bar code scanning module, and for some articles without the RFID tags and the bar codes, the visual identification module is used for guiding the mechanical arm to grab the articles read and scanned by the RFID induction module and the bar code scanning module. The visual module can also be used for identifying and sorting; the robot sorting control module comprises an object position positioning and following grabbing algorithm module, a sorting robot and an infrared object position induction sensor on an assembly line. According to the method, the accuracy of target classification and recognition in the storage environment is enhanced, the sorting efficiency is improved, the target can be closer to the real environment, and the generalization performance of the model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a multi-target recognition system and method for a sorting robot in a warehousing environment, belonging to the technical field of multi-target recognition. Background Art

[0002] With the continuous development of the logistics industry, the number of warehoused goods has gradually increased, and the categories of goods have become increasingly complex. Manual sorting of goods has become increasingly difficult. Warehouse sorting robots can not only greatly save labor, but also monitor the sorting and transportation of goods in real time to count the quantity and transportation direction of goods categories. It can greatly reduce the operation and maintenance costs of enterprises, increase the efficiency of warehouse sorting, and improve economic benefits.

[0003] Chinese Patent Application with Patent No. CN107563432A discloses a multi-target recognition system for robots based on a vision model. This method obtains the lengths of various feature vectors of parts by scanning them for classifying parts. The categories of goods classified by this method are relatively limited, and small parts blocked or loaded in containers such as boxes cannot be classified, so the range of classified goods is relatively limited.

[0004] Chinese Patent Application with Patent No. CN110653168A applied for a sorting system for multi-target recognition. This patent mainly uses a mechanical trolley equipped with a CCD camera to scan two-dimensional codes, identify goods through the two-dimensional codes and sort them. For some damaged two-dimensional codes or the two-dimensional code labels of some items that cannot be scanned due to their uneven surfaces, there are relatively strict requirements for the shape and size of commodity goods and it cannot adapt

[0005] to the diversity of goods, and its classification of diverse goods is relatively limited.

[0006] The range of commodity classification categories of the above patents is small, unable to handle frequent pipeline sorting situations, the recognition models are relatively single, and their model response is slow, with low real-time performance and low automation, unable to meet the current rapid logistics sorting situation. Summary of the Invention

[0007] To meet the requirements of the current industry for commodity sorting, the present invention proposes a multi-target recognition system and method for a sorting robot in a warehousing environment, which can recognize many commodity categories, has good real-time performance, and high detection efficiency.

[0008] The technical solution of the present invention is as follows:

[0009] A multi-target recognition system for a sorting robot in a warehousing environment includes a commodity category recognition module and a robot sorting control module;

[0010] The commodity category recognition module includes an RFID induction module, a barcode scanning module, and a machine vision recognition module;

[0011] The RFID induction module includes an RFID sensor for reading the signals of items with RFID tags pasted;

[0012] The barcode scanning module includes an industrial barcode scanner for scanning items with barcodes set because RFID tags cannot be pasted,

[0013] The vision recognition module: is used to guide the robotic arm to grab the items read and scanned by the RFID induction module and the barcode scanning module. For some items without RFID tags and barcodes pasted, they can also be identified and sorted through the vision module;

[0014] The robotic sorting control module includes an object position positioning and following grasping algorithm module, a sorting robot, and an infrared object position induction sensor on the production line.

[0015] A multi-object recognition method for a sorting robot in a warehousing environment, using the above system, includes the following steps:

[0016] Step 1: Install the RFID sensor, industrial barcode scanner, sorting robot, and infrared object position induction sensor beside the commodity transportation production line, adjust the conveying speed of the conveyor belt, and complete the communication and preliminary debugging between components;

[0017] Adjust the running speed of the production line, install the infrared object position induction sensor, build the production line according to the sorting categories of warehousing commodities, conduct preliminary debugging on the sorting robot, conduct preliminary calibration on the industrial barcode scanner and the vision system of the robot, and conduct debugging on the RFID sensor. Transmit the induction signal of the RFID, the scanning signal of the industrial barcode scanner, and the signal of the infrared object position induction sensor to the control system of the sorting robot through wired transportation.

[0018] Step 2: Collect the image information of commodities in the warehousing environment, classify and label these images, and then perform augmented dataset processing, and train the machine vision module with the augmented dataset;

[0019] Denoise the images, rotate the angles, and store the processed images in various categories to expand the commodity recognition dataset;

[0020] Gaussian filtering is used for denoising:

[0021]

[0022] where (x,y) are the coordinates of the original image; δ is a specified constant;

[0023] Image rotation uses:

[0024] x′ = x - cx

[0025] y′ = y - cy

[0026] x” = x'cos(θ) - y'sin(θ)

[0027] y” = x'sin(θ) + y'cos(θ)

[0028] x2 = x″ + cx

[0029] y2 = y″ + cy

[0030] (cx, cy) are the coordinates of the rotation center, (x2, y2) are the coordinates after rotation, and θ is the angle of image rotation. The preprocessed image information is obtained, and manual image annotation classification is performed. Then, it is divided into a training set and a validation set according to a ratio of 8:2.

[0031] Step 3: Build a visual image classification model based on a multi-channel fusion CNN module, and the multi-channel fusion CNN module is built by oneself;

[0032] Step 4: Input the augmented dataset into the multi-channel fusion CNN module for training. After training, implant the network into the visual recognition module for real-time detection;

[0033] Step 41: Divide the processed dataset into a training set, a validation set, and a test set;

[0034] Step 42: Input the image training dataset into the multi-channel fusion CNN network model obtained in Step 3 for training. Pass the image data forward through the network model to obtain the output value of the network model. According to the error value between the output value and the corresponding true value in the dataset, calculate the contribution of each network parameter to the error value through the backpropagation algorithm SGD, and update the parameters of the network model according to the gradient obtained by backpropagation, and continuously adjust the model parameters until the preset number of training times is reached or the convergence state is reached, and the training ends;

[0035] Step 43: Use the test set to test the trained model, and evaluate performance indicators such as the accuracy, recall rate, precision rate, and classification error of the network model. y i is the actual value, P i is the predicted value, m is the total number of data points, and C is the total number of categories.

[0036]

[0037] Among them, TP represents true positive, that is, predicting the true category as the true category; TN represents true negative, that is, the number of predicting the wrong category as the wrong category; FP represents false positive, that is, the number of predicting the wrong category as the true category; FN represents false negative, that is, the number of predicting the true category as the wrong category.

[0038] Step 44: Use the genetic algorithm to adjust the hyperparameters for the retraining of the model to improve the accuracy and robustness of the model. First, limit the value range of each hyperparameter, and then randomly generate a batch of data within the limited range as the initial data. For the commodity image classification part, define the fitness function as its classification cross-entropy, set the mutation probability of the population, perform the selection operation, then perform crossover and mutation, calculate the fitness of the new population, and repeat this series of operations until the maximum number of iterations is reached or a certain convergence threshold is reached, so that the optimal training model and the corresponding hyperparameters can be obtained by the genetic algorithm.

[0039] Step 45: After implanting the multi-channel fusion CNN network structure model into the sorting robot, its image is extracted by the visual recognition module, and the priority of the triple detection system is set as RFID sensor > industrial barcode scanner > visual recognition module. When the commodity category has been determined by the previous level, the signal of the next level will be skipped and directly transferred to the sorting control system of the industrial robot to guide the sorting.

[0040] Step 5: When the sorting robot sorts, it locates the object position and follows the grasping algorithm module to locate the object position and follow the grasping.

[0041] When the commodity object moves to the position where the infrared positioning device is triggered, the object will block the infrared light, resulting in signal triggering. At this time, the sorting robot locks the object. The color of the conveyor belt is selected to have a distinct color difference from the commodity color. Using the HSV color threshold segmentation method, after the obtained contour is filtered by opening operation and the internal small holes are closed by closing operation, it can be separated from the background. Calculate the minimum circumscribed rectangle with its border parallel to the conveyor belt for the separated contour to obtain its rectangular border and center. The visual recognition module follows the position of the center and transmits this position to the sorting robot control system in real time. After the control system calculates the time interval, it moves the gripper at the end of the mechanical arm of the sorting robot to directly above it for grasping. The movement of the mechanical arm adopts PID control and reaches the target position within the limited time and overshoot range.

[0042] The beneficial effects of the present invention are:

[0043] 1. The present invention is a multi-target recognition system based on the multi-channel fusion CNN network structure model combined with RFID and barcodes, which strengthens the accuracy of target classification and recognition in the warehouse environment and improves the efficiency during sorting.

[0044] 2. The present invention adjusts the training set by means of data augmentation, which can make it closer to the real environment and ensure the generalization performance of the model.

[0045] 3. The present invention adopts a strategy of multi-sensor information fusion for robot sorting and grasping control, with high positioning accuracy, stable robot control movement and good timeliness. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is an installation diagram of industrial assembly line components;

[0047] Figure 2 It is a flowchart of a triple target recognition system;

[0048] Figure 3 a is a CNN network structure diagram, and b is a composition diagram of a CONV module;

[0049] Figure 4 It is a schematic diagram of camera calibration coordinate conversion;

[0050] Figure 5 It is a schematic diagram of robotic arm hand-eye calibration;

[0051] Figure 6 It is a flowchart of a genetic algorithm;

[0052] Figure 7 (a) is a graph of the change of training loss of a multi-channel fusion CNN model, and (b) is a graph of the change of its accuracy;

[0053] Figure 8 It is a diagram of conveyor belt object positioning tracking and category recognition;

[0054] Figure 9 It is a response diagram of PID control during robotic arm grasping and sorting under disturbance. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0056] Embodiment:

[0057] As Figure 1 shown, the present invention provides a multi-target recognition system for a sorting robot used in a warehouse environment, including the following steps:

[0058] Step 1: Install components such as RFID sensors, industrial barcode scanners, sorting robots, and infrared object position sensors beside the commodity transportation assembly line. Adjust the conveyor belt speed to 0.1 m / s and the spacing between items to 0.5 m. Complete the communication and preliminary debugging between components, specifically as follows:

[0059] Step 11: Select an RFID printer and an RFID sensor. The requirements for the sensor are a reading distance of ≥ 1.5 m, a wired transmission function, and a reading rate of > 500 sheets / second.

[0060] Step 12: Select a barcode printer and an industrial camera. The requirements are a resolution of ≥ 720×1280, a reading distance of ≥ 1 m, an image frame rate of ≥ 10 fps, a field of view width of ≥ 700 mm, and a wired transmission function.

[0061] Step 13: Select an infrared position sensor with a response time of ≤ 0.01 s and a wired transmission function. Select a 6-axis robot with a movable radius of ≥ 0.5 m and equipped with a gripper at the end that can withstand a weight of ≥ 30 KG. Install one position for one type of commodity on the assembly line and only sort this type of commodity. The position spacing between sorting areas for different types of commodities is ≥ 3 m.

[0062] Step 14: Calibrate the industrial camera (including barcode scanner and machine vision camera). As Figure 4 shown, PW = (XW, YW, ZW) is the world coordinate system, PC = (XC, YC, WC) is the camera coordinate system, P = (X, Y, 1) is the image coordinate system, and (U, V) is the pixel coordinate system. The calibration can be obtained as follows: (f is the camera calibration, dX, dY are the pixel unit sizes, u0, v0 are the image centers, and R, T are the transformation matrices from the world coordinate system to the camera coordinate system)

[0063]

[0064] Perform hand-eye calibration on the camera on the robotic arm. As Figure 5 shown, it is equivalent to adding the matrix transformation from the camera to the end of the robotic arm and the matrix transformation from the end of the robotic arm to the base coordinate system of the robotic arm on the basis of camera calibration. These two matrices can be obtained by the DH method.

[0065] Step 2: Obtain a common commodity dataset for warehousing logistics and perform data augmentation processing: Convert the image to 640×640×3. Assume the size of the original image is M×N and it needs to be scaled to an image of size P×Q:

[0066]

[0067] For the scaled pixel points (x', y') corresponding to the original image (x, y), they can be obtained through bilinear interpolation:

[0068] I'(x', y') = (1 - a)(1 - b)I(x, y) + a(1 - b)I(x + 1, y) + (1 - a)bI(x, y + 1) + abI(x + 1, y + 1)

[0069] I’ is the transformed image, I is the original image, and a and b are the relative distances between the four pixels closest to the interpolation point.

[0070] Gaussian filtering is used for denoising:

[0071]

[0072] where (x, y) are the coordinates of the original image.

[0073] Image rotation is performed using:

[0074] x′ = x - cx

[0075] y′ = y - cy

[0076] x” = x'cos(θ) - y'sin(θ)

[0077] y” = x'sin(θ) + y'cos(θ)

[0078] x2 = x″ + cx

[0079] y2 = y″ + cy

[0080] (x, y) are the coordinates of the pixel points in the original image, (cx, cy) are the coordinates of the rotation center, (x2, y2) are the coordinates after rotation, and the preprocessed image is labeled and classified and divided into a training set and a validation set according to a ratio of 8:2.

[0081] Step 3: Based on the multi-channel fusion CNN network structure model, use as Figure 4The structure shown is as follows. If the size of the input image is 640×640×3, after the CONV operation in the first column (the first channel), a feature map with a size of 320×320×48 is obtained. After the CONV operation in the second column (the second channel), a feature map with a size of 160×160×48 is obtained. After the CONV operation in the third column (the third channel), a feature map with a size of 80×80×48 is obtained. After the fourth column of the feature map, it is a feature map with a size of 40×40×48. Among them, the part passing through the first channel is continued to be sent to the pooling layer for dimensionality reduction, obtaining a feature map with a size of 160×160×48, and then adding it to the feature map of the second channel to obtain a feature image integrating the information of the first and second channels, with a size of 160×160×48. The third and fourth columns are also operated in the same way as the first and second columns to obtain a feature map with a size of 40×40×48. The first path of the integrated feature map of the first and second columns passes through the CONV operation to obtain a feature map with a size of 80×80×48, and the second path passes through the pooling operation to obtain a feature map with a size of 80×80×48. Then, adding the first path and the second path together obtains a fused feature map with a size of 80×80×48. The feature map of the first path just now passes through the max pooling to obtain a feature map with a size of 40×40×48. The previously obtained integrated information feature map passes through the CONV to obtain a feature map with a size of 40×40×96, and then adding it to the feature map with a size of 40×40×48 obtained through the CONV operation before and sending it into the CONV operation to obtain a feature map with a size of 20×20×96. Concatenating this feature map with the integrated feature map of the previous third and fourth columns and the feature map obtained by the first column alone through the max pooling before obtains a feature map with a size of 40×40×192. The feature map with a size of 40×40×192 passes through the average pooling to obtain a feature map with a size of 20×20×192. The integrated feature map with a size of 40×40×48 of the previous third and fourth columns is also divided into two paths. One path passes through the CONV, and the second path passes through the max pooling. Finally, adding the two paths together obtains a feature map with a size of 20×20×48, and then concatenating this feature map with the feature map with a size of 20×20×192 to obtain a feature map with a size of 20×20×240. This feature map passes through the CONV to obtain a feature map with a size of 20×20×24, and finally sending this feature map into the average pooling to obtain a feature map with a size of 10×10×24. Flattening this feature map and sending it into the FL (fully connected layer), after passing through the fully connected layer and sending it into the activation function RELU, a feature vector with a size of 100×1 is obtained. Sending it into the fully connected - activation function layer again obtains a feature vector with a size of 6×1. Finally, the SOFTMAX layer performs class prediction. The output of the fully connected network is the number of classified commodities, and the output of each neuron is (F is the activation function, commonly TANH or RELU):

[0082] The output of the fully connected network is the number of classified commodities, and the output of each neuron is (F is the activation function, commonly TANH or RELU, B is the bias, Wi is the weight, Z is the output result, and Y i is the i-th input):

[0083]

[0084] The output of the fully connected layer is finally the probabilities of each category, and the category with the highest probability is used as the output.

[0085] Step 4: Input the augmented dataset into the multi-channel fusion CNN module for network training. After training, implant the network into the vision software for real-time detection. Specifically:

[0086] Step 41: Divide the processed dataset into a training set, a validation set, and a test set;

[0087] Step 42: Input the image training (commodity classification) dataset into the multi-channel fusion CNN network model obtained in Step 4 for training. Pass the image data forward through the network model to obtain the output value of the network model. According to the error value between the output value and the corresponding true value in the dataset, calculate the contribution of each network parameter to the error value through the backpropagation algorithm SGD, and update the parameters of the network model according to the gradient obtained by backpropagation, and continuously adjust the model parameters until the preset number of training times is reached or the convergence state is achieved, and the training ends;

[0088] Step 43: Use the test set to test the trained model, and evaluate performance indicators such as the accuracy, recall rate, precision rate, and classification error of the network model. yi is the actual value, Pi is the predicted value, m is the total number of data points, and C is the total number of categories.

[0089]

[0090] Among them, TP represents the true positive example, that is, predicting the true category as the true category; TN represents the true negative example, that is, the number of predicting the wrong category as the wrong category; FP represents the false positive example, that is, the number of predicting the wrong category as the true category; FN represents the false negative example, that is, the number of predicting the true category as the wrong category;

[0091] Step 44: Use the genetic algorithm to adjust the hyperparameters for retraining the model to improve the accuracy and robustness of the model. The flowchart of the genetic algorithm is as Figure 6As shown in the figure. First, limit the value ranges of each hyperparameter. After the limitation, randomly generate a batch of data within the value range as the initial data. For the commodity image classification part, define the fitness function as its classification cross-entropy. Set the mutation probability of the population. After the selection operation, perform crossover and then mutation. Calculate the fitness of the new population and repeat this series of operations until the maximum number of iterations is reached or a certain convergence threshold is achieved. Thus, the optimal training model and the corresponding hyperparameters can be obtained by the genetic algorithm. The final training result is as Figure 7 shown, and it can be seen that the verification accuracy rate reaches 82%.

[0092] Step 45: After implanting the multi-channel fusion CNN network structure model into the robotic arm, its image is extracted by the camera on the robotic arm. Set the priority of the triple detection system as RFID sensor > industrial barcode scanner > machine vision. When the commodity category has been determined by the previous level, the signal of the next level will be skipped and directly transferred to the sorting control system of the industrial robot to guide the sorting. The specific flowchart is as Figure 2 shown.

[0093] Step 5: The object position positioning and following grasping algorithm during robot sorting is as follows: When it is 1 meter away from the conveyor belt, since the object blocks the infrared locator and triggers the signal, at this time, the robotic arm camera will perform positioning. After obtaining the minimum circumscribed rectangle whose border is parallel to the conveyor belt, as Figure 8 shown.

[0094] Step 51: Since the time t from its position to the position directly opposite the robotic arm is 10 s, establish the origin of the grasping coordinate system at (0, 0, 0) directly below the robot, and the grasping waiting position is (0, 0, 30). Assume the current coordinates of its end gripper are (x1, y1, z1), and it needs to move to (0, 0, 30). From the DH matrix, it can be known that there is such a relationship between the rotation amounts qi of each joint of the robot and the end gripper:

[0095]

[0096] The coordinate transformation matrix T can be obtained by the DH method. Therefore, to solve the change amounts of each joint qi from the known position, the inverse matrix of T needs to be calculated. After obtaining the change amounts of each qi, based on time, it is required to reach the specified position before 8 s. PID control can be performed on it to meet the target requirements. From Figure 9It can be seen that assuming the target position is given by a step signal, through establishing a mathematical model for simulation, the simulation results show that the PID control reaches the specified position in 4 s under disturbance, and the movement is smooth under position fluctuation, indicating its good rapidity and stability. After reaching the specified position, the gripper can be opened according to its minimum circumscribed rectangle, wait for the object to arrive and then close the gripper, and then plan from the current position (0, 0, 30) to the object sorting position (x2, y2, z2). The operation is the same as before, but since the spacing between objects is 0.5 m, only 5 s are required to return to the starting position and wait for the operation.

[0097] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present invention, several improvements and deformations can still be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A multi-target recognition system for sorting robots in a warehousing environment, characterized in that It includes a commodity category recognition module and a robot sorting control module; The commodity category recognition module includes an RFID induction module, a barcode scanning module, and a machine vision recognition module; The RFID induction module includes an RFID sensor for reading the signals of items pasted with RFID tags; The barcode scanning module includes an industrial barcode scanner for scanning items with barcodes set because they cannot be pasted with RFID tags. The vision recognition module: is used to guide the robotic arm to grasp the items read and scanned by the RFID induction module and the barcode scanning module. For some items without RFID tags and barcodes pasted, they can also be identified and sorted through the vision module; The robot sorting control module includes an object position positioning and following grasping algorithm module, a sorting robot, and an infrared object position induction sensor on the assembly line.

2. A multi-target recognition method for sorting robots in a warehousing environment, characterized in that Using the system described in claim 1, it includes the following steps: Step 1: Install the RFID sensor, industrial barcode scanner, sorting robot, and infrared object position induction sensor beside the commodity transportation assembly line, adjust the conveying speed of the conveyor belt, and complete the communication and preliminary debugging between each component; Step 2: Collect the image information of commodities in the warehousing environment, classify and label these images, then perform augmented dataset processing, and train the machine vision module with the augmented dataset; Step 3: Build a visual image classification model based on a multi-channel fusion CNN module, and the multi-channel fusion CNN module is built by oneself for the network; Step 4: Input the augmented dataset into the multi-channel fusion CNN module for training. After training, implant the network into the vision recognition module for real-time detection; Step 5: When the sorting robot sorts, it performs object position positioning and following grasping through the object position positioning and following grasping algorithm module.

3. The multi-object recognition method for a sorting robot in a warehousing environment according to claim 2, characterized in that The specific steps of Step 1 are as follows: Adjust the running speed of the assembly line, install the infrared object position induction sensor, build the assembly line according to the sorting categories of warehousing commodities, conduct preliminary debugging on the sorting robot, conduct preliminary calibration on the industrial barcode scanner and the vision system of the robot, and conduct debugging on the RFID sensor. Transmit the induction signal of the RFID, the scanning signal of the industrial barcode scanner, and the signal of the infrared object position induction sensor to the control system of the sorting robot through wired transportation.

4. A multi-object recognition method for sorting robots in a warehousing environment according to claim 2, characterized in that The specific steps of Step 2 are as follows: Perform denoising processing and angle rotation on the images, and store the processed images into various categories to expand the commodity recognition dataset; When denoising, use Gaussian filtering processing: where (x, y) are the coordinates of the original image; δ is a specified constant; The image rotation adopts: X′ = X - Cx y′ = y - cy X″ = X′cos(θ) - y′sin(θ) y″ = x′sin(θ) + y′cos(θ) x2 = x″ + cx y2 = y″ + cy (cx, cy) are the coordinates of the rotation center, (x2, y2) are the coordinates after rotation, and θ is the rotation angle of the image to obtain the preprocessed image information, conduct manual image annotation classification, and then divide it into a training set and a validation set according to a ratio of 8:

2.

5. The multi-object recognition method for a sorting robot in a warehousing environment according to claim 2, wherein The specific steps of Step 4 are as follows: Step 41: Divide the processed data set into a training set, a validation set, and a test set; Step 42: Input the image training data set into the multi-channel fusion CNN network model obtained in Step 3 for training. Forward-propagate the image data through the network model to obtain the output value of the network model. According to the error value between the output value and the corresponding true value in the data set, calculate the contribution of each network parameter to the error value through the backpropagation algorithm SGD, and update the parameters of the network model according to the gradient obtained by backpropagation. Continuously adjust the model parameters until the preset number of training times is reached or the convergence state is achieved, and the training ends; Step 43: Use the test set to test the trained model, and evaluate performance metrics such as the accuracy, recall, precision, and classification error of the network model, where y i is the actual value, P i is the predicted value, m is the total number of data points, and C is the total number of categories. Among them, TP represents the true positive example, that is, predicting the true category as the true category; TN represents the true negative example, that is, the number of predicting the wrong category as the wrong category; FP represents the false positive example, that is, the number of predicting the wrong category as the true category; FN represents the false negative example, that is, the number of predicting the true category as the wrong category; Step 44: Use the genetic algorithm to adjust the hyperparameters for retraining the model to improve the accuracy and robustness of the model. First, limit the value range of each hyperparameter. After the limitation, randomly generate a batch of data within the value range as the initial data. Define the fitness function for the commodity image classification part as its classification cross-entropy, set the mutation probability of the population, perform the selection operation, then perform the crossover and swap, and then perform the mutation. Calculate the fitness of the new population, and repeat this series of operations until the maximum number of iterations is reached or a certain convergence threshold is reached, so that the optimal training model and the corresponding hyperparameters can be obtained by the genetic algorithm; Step 45: After implanting the multi-channel fusion CNN network structure model into the sorting robot, its image is extracted by the visual recognition module, and the priority of the triple detection system is set as RFID sensor > industrial barcode scanner > visual recognition module. When the commodity category has been determined by the upper level, the signal of the lower level will be skipped and directly transferred to the sorting control system of the industrial robot to guide the sorting.

6. The multi-object recognition method of the sorting robot for the warehousing environment according to claim 2, characterized in that The specific steps of Step 5 are as follows: When the commodity object travels to the position where the infrared positioning device is triggered, the object will block the infrared light, resulting in signal triggering. At this time, the sorting robot locks the object. The color of the conveyor belt is selected as a color that is significantly different from the color of the commodity. Using the HSV color threshold segmentation method, after the obtained contour is filtered by the opening operation and the internal small holes are closed by the closing operation, it can be separated from the background. Calculate the minimum circumscribed rectangle whose border is parallel to the conveyor belt for the separated contour to obtain its rectangular border and center. The sub-visual recognition module follows the position of the center and transmits this position to the sorting robot control system in real time. After the control system calculates the time interval, it moves the gripper at the end of the mechanical arm of the sorting robot to directly above it for grasping. The movement of the mechanical arm adopts PID control and reaches the target position within the limited time and overshoot range.

Citation Information

Patent Citations

  • Visual shape model-based robot multi-target recognition method

    CN107563432A

  • Multi-object identification based intelligent sorting system

    CN110653168A