An intelligent storage object recognition device based on lightweight network
By adopting a lightweight network-based intelligent storage object recognition device in the warehousing identification system, using multi-view image acquisition and Densenet to improve the model, the existing system's low recognition efficiency, low accuracy and large data volume are solved, and the recognition effect of high-precision and low data processing volume is achieved, which is suitable for the identification of irregular parts.
Patent Information
- Application Number
- CN202110359606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-04-02
AI Technical Summary
The existing warehouse identification system has low recognition efficiency and recognition accuracy, excessive data processing and poor applicability.
Using an intelligent storage object recognition device based on lightweight network, multi-angle image data is obtained using a turntable and multi-view image acquisition device, and the Densenet improved model is used for identification. This model improves recognition accuracy and reduces data processing through technologies such as Channel Split, Deep Separation Convolutional Layer and Channel Shuffle.
It realizes high-precision and low data processing volume of storage objects, which is suitable for the identification of irregular parts, greatly improving the applicability and identification efficiency of the system.
Smart Images

Figure CN113505629B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of equipment identification, and relates to an intelligent storage object identification device based on a lightweight network. Background Art
[0002] Since the 20th century, machine vision has experienced nearly 60 years of development and has seen many applications. It is also one of the hot research directions of scholars from all walks of life. There are countless cases of applying machine vision to detection, transportation, medical imaging, and monitoring systems. Applying machine vision to industrial warehousing can help realize industrial intelligence faster. Since machine vision systems can quickly obtain a large amount of information, are easy to process automatically, and are easy to integrate with design information and processing control information, in modern automated production processes, people use machine vision systems widely in fields such as working condition monitoring, finished product inspection, and quality control. However, machine vision technology is relatively complex, and the biggest difficulty is that the human visual mechanism is still unclear. Therefore, there is still a certain degree of complexity in establishing a machine vision system. In the process of studying machine vision, it is crucial to obtain a high-quality usable image. Therefore, in order to ensure image quality, the entire structure has high requirements for the light source system. Contrast, brightness, robustness (position sensitivity), etc. are all issues that need to be carefully considered during the design process. It can be foreseen that with the maturity and development of machine vision technology itself, it will be more and more widely used in modern and future manufacturing industries.
[0003] With the rapid development of the Internet of Things, all walks of life are developing intelligent warehousing systems, especially in the logistics industry. Establishing an intelligent warehousing system in an industrial environment will greatly improve search efficiency. Modern industrial warehousing systems not only have complex items, but also have different shapes and performances. Realizing dynamic access and intelligent query can provide a good guarantee for industrial intelligence.
[0004] In traditional industrial production or warehousing scenarios, the identification and tracking of production materials (devices) or goods are usually achieved by labeling (barcodes, QR codes). The labeled goods (such as QR codes) are stored in a fixed location in the warehouse by manpower. When picking up the goods, the location of the goods in the warehouse is obtained by scanning the labels to achieve the picking and placing of the goods, or the scanning system set on the conveyor belt scans the passing materials to achieve the tracking of the materials. However, for non-standard goods (small quantity) or special-shaped goods (which cannot be labeled), the labeling method is inefficient or simply cannot be achieved. For such goods, the current warehousing system mainly collects parts in a specific state at a specific angle, and then completes the part identification. In order to complete the part identification, it is often necessary to use a material sorting machine in conjunction with an image acquisition device. Not only is the collection very cumbersome and difficult to manage, but the material sorting machine often needs to be customized, which is expensive. At the same time, it is difficult to apply to irregular and complex material sorting.
[0005] In addition, target detection is one of the important research directions in the field of computer vision. The traditional target detection method is to extract features by constructing feature descriptors and then use classifiers to classify features to achieve target detection, such as Histogram of Oriented Gradient (HOG) and Support Vector Machine (SVM). With the excellent performance of deep learning in the field of image classification, convolutional neural networks have begun to be widely used in various fields of computer vision. Using deep learning to achieve target detection in the field of target detection has become a new direction. Traditional neural networks use fully connected layers to connect between layers, while the weight sharing network of convolutional neural networks greatly reduces the amount of calculation and the complexity of the network model. The translation invariance of convolutional neural networks can better process the features of images. A large number of convolutional neural networks have emerged based on image recognition methods, from the initial LeNet, AlexNet, ZFNext, VGGNet, Inception series, ResNet series to lightweight neural networks. These algorithms are modified based on Backbone to obtain high accuracy. These technologies have been gradually used in other applications such as face recognition and intelligent warehousing. Despite this, the technology still has the following defects: Due to the emergence of residual connections, it can stack deeper networks while ensuring accuracy, but this also increases the complexity and computational complexity of the network model. Some commercial computers can run these networks with superior GPUs and memory, but for some embedded devices with limited performance, the amount of data they process is too large to meet the requirements of timeliness and accuracy.
[0006] Therefore, it is of great practical significance to develop a storage object recognition device with high accuracy, small data processing volume and good applicability. Summary of the invention
[0007] The purpose of the present invention is to overcome the defects of the existing warehouse identification system, such as low recognition efficiency and recognition accuracy, excessive processing data volume and poor applicability, and to provide a warehouse object identification device with high accuracy, small data processing volume and good applicability.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] An intelligent storage object identification device based on a lightweight network includes a turntable for placing an object to be identified, wherein the turntable is connected to a turntable driving device and can rotate horizontally under the drive of the turntable driving device;
[0010] There are more than two image acquisition devices arranged around the turntable, with the centers of the fields of view aligned with the center of the turntable, and the image acquisition devices are arranged at different positions of the turntable and have different height differences with the turntable;
[0011] The image acquisition device and the turntable driving device are respectively connected to the central processing unit, and the central processing unit obtains the picture of the object to be identified through the image acquisition device, and then inputs it into the trained Densenet improved model, and the trained Densenet improved model outputs the target category;
[0012] The improvement of the Densenet improved model compared to the Densenet model is that the Back Bone in the net block (Net Block) includes Channel Split, a first convolutional layer (first Conv), a second convolutional layer (second Conv), a third convolutional layer (third Conv), a first depth separation convolutional layer (first DWConv), a second depth separation convolutional layer (second DWConv), Concat, and Channel Shuffle, wherein Channel Split, the first convolutional layer (first Conv), the first depth separation convolutional layer (first DWConv), the second convolutional layer (second Conv), Concat, and Channel Shuffle are connected in sequence, the second depth separation convolutional layer (second DWConv), the third convolutional layer (third Conv) are connected in parallel with the first convolutional layer (first Conv), the first depth separation convolutional layer (first DWConv), and the second convolutional layer (second Conv), the second depth separation convolutional layer (second DWConv) is connected to Channel Split, and the third convolutional layer (third Conv) is connected to Concat. That is, after the output channels of Back Bone are distributed by Channel Split, part of them pass through the first convolution layer (first Conv), the first depth separation convolution layer (first DWConv), and the second convolution layer (second Conv), and the other part passes through the second depth separation convolution layer (second DWConv), the third convolution layer (third Conv), and then they are merged through Concat and then shuffled through Channel Shuffle and output to the next layer as the input of the next layer.
[0013] The turntable in the lightweight network-based intelligent warehouse object recognition device of the present invention can adjust the orientation of the parts, increase data richness and reliability, and compared with the traditional machine vision that only uses top-view images for recognition, multi-perspective images (it is equipped with multiple image acquisition devices in different orientations and different heights that can obtain the external full view of the part to be recognized during the image acquisition process) can provide more part information, which is conducive to the deep learning model to learn more complete part information, and can prevent overfitting and improve the generalization ability of the model. It can not only improve the recognition efficiency, but also be suitable for the recognition of irregular parts, greatly improving the applicability of the system; in addition, the present invention also proposes for the first time to combine the convolutional neural network with the characteristics of the ShuffleNet channel intention order, the residual network in ResNet, the channel segmentation in Inception, the multiple residual networks in DenseNet and the DWConv in MobileNet to obtain the Densenet improved model, which further improves the accuracy and reduces the structure of the model compared with the traditional network and the lightweight network, and the application of the Densenet improved model to complete the target category recognition not only has a small amount of data processing, but also has high recognition accuracy, and has great application prospects.
[0014] As the preferred technical solution:
[0015] In the above-mentioned intelligent storage object recognition device based on lightweight network, the central processor runs the following program to perform data enhancement:
[0016] (1) After the object to be identified is placed at the center of the turntable, the central processing unit turns on the turntable drive device and all image acquisition devices. The turntable drive device drives the turntable to rotate horizontally. The image acquisition device regularly acquires images of the object to be identified. While the image acquisition device is acquiring images, the central processing unit records the observation angle information of the image acquisition device and maps the acquired image with the observation angle information data to form the data set. The observation angle information G of the nth image acquisition device at the tth moment is G. t n The specific expression is as follows:
[0017]
[0018]
[0019] Among them, r n is the distance between the nth image acquisition device and the center of the turntable, θ n is the angle between the nth image acquisition device and the horizontal plane where the turntable is located, is the angle between the horizontal projection of the nth image acquisition device at the tth moment and the x-axis, is the angle between the horizontal plane projection of the nth image acquisition device and the x-axis at the initial moment, the x-axis is the x-axis in the coordinate system established with the center of the turntable as the origin, the horizontal plane where the turntable is located as the x-axis, and the plane where the y-axis is located, and V represents the rotation speed of the turntable;
[0020] (2) Enhance the observation angle information data of the image acquisition device in the data set, and the enhanced observation angle information H of the nth image acquisition device at the tth moment is t n The specific expression is as follows:
[0021]
[0022]
[0023] in, is the rotation angle of the enhanced data compared to the observation angle information of the nth image acquisition device at the tth moment, [d x ,d y ] respectively represent the translation components of the enhanced data in the x-axis and y-axis directions compared to the observation angle information of the n-th image acquisition device at the t-th moment.
[0024] During the data acquisition process, the device of the present invention establishes a Cartesian coordinate system with the center of the turntable as the origin. When the object is placed at the center of the turntable and starts to acquire images, the real-time spatial polar coordinate description of the camera's viewing angle relative to the object is calculated (used to describe the spatial posture of the camera unit at different angles), thereby marking the observation viewing angle information (including recognition angle and distance parameters) based on the spatial polar coordinate representation in the acquired object image data, thereby increasing the efficiency and accuracy of the subsequent recognition process.
[0025] As described above, in an intelligent warehouse object recognition device based on a lightweight network, the state corresponding to the enhanced data is translated left and right or up and down by less than or equal to 20% compared to the state corresponding to the enhanced object, or the random rotation angle clockwise or counterclockwise is less than or equal to 30°. Only a feasible technical solution is given here, and technicians in this field can generate enhanced data through translation and rotation operations according to actual needs.
[0026] In the above-mentioned intelligent storage object recognition device based on a lightweight network, the turntable is arranged in a frame, the image acquisition device is fixed on the frame, and a light source is also fixed on the frame;
[0027] A black back plate is arranged under the turntable; the black back plate, the light source and the soft light cover together constitute the light source system of the device, which provides the conditions required for diffuse reflection. The lighting conditions are based on diffuse reflection lighting, which is used to provide a stable and suitable lighting environment for the device, avoid image acquisition from being disturbed and affected by external noise, and improve the recognition accuracy.
[0028] The turntable driving device is a driving motor.
[0029] In the above-mentioned intelligent storage object recognition device based on a lightweight network, the light source is arranged above the turntable, and the outer cover of the frame is provided with a soft light cover;
[0030] The frame is a square frame;
[0031] There are five image acquisition devices, which are arranged around and on the top of the frame, and the height differences between the five image acquisition devices and the turntable are different; the five image acquisition devices can realize parallel acquisition, and five image data of the same part from different perspectives can be obtained after one acquisition, which greatly improves the acquisition speed and efficiency;
[0032] The frame has rounded corners to prevent excessive local pressure from damaging the equipment. The frame is made of multiple aluminum alloy square tubes fixed and spliced together. The frame provides stable structural support for the image acquisition device of the intelligent visual image acquisition system. Except for the parts to be identified, the rest of the device (frame, image acquisition device, turntable, turntable drive device, light source, soft light cover and black backplane) are fixed as one to prevent installation errors or collision damage during movement.
[0033] As described above, the intelligent storage object recognition device based on a lightweight network, the Densenet improved model includes a main convolution layer (main Conv), a feature extraction layer (Pooling), a first network block, a transition layer (TransitionLayer), a second network block, and a classification layer (Classification Layer) connected in sequence; the first network block and the second network block have the same structure;
[0034] The ratio of the number of training sets to the number of test sets of the Densenet improved model is 4:1. The end point of model training is to reach the preset number of training times. The protection scope of the present invention is not limited to this, and the number of training sets and test sets can be set by those skilled in the art according to actual conditions.
[0035] An intelligent storage object identification device based on a lightweight network as described above, wherein the network block comprises a first Back Bone, a second Back Bone, a third Back Bone and a fourth Back Bone connected in sequence, the output of the first Back Bone is simultaneously the input of the second Back Bone, the third Back Bone and the fourth Back Bone, and the output of the second Back Bone is simultaneously the input of the third Back Bone and the fourth Back Bone;
[0036] The transition layer (Transition Layer) includes a feature extraction layer (Pooling);
[0037] The classification layer includes a global average pool and a Softmax classifier.
[0038] As described above, in the intelligent storage object recognition device based on lightweight network, there are five Densenet improved models, each of which corresponds to an image acquisition device one by one (i.e., the training and use of the Densenet improved model corresponds to the image acquisition device), and the training process is a process of continuously adjusting the model parameters using the image of the known category of objects acquired by the corresponding image acquisition device as input and the corresponding category probability of the object as theoretical output, and the termination condition of the training is reaching the upper limit of the number of training times. A set of data groups used to train the Densenet improved model includes images of the known category of objects acquired by the corresponding image acquisition device and their corresponding categories.
[0039] The above Softmax classifier is used to calculate the classification probability of each sample, as follows:
[0040]
[0041] In the formula, s i represents the output value of the ith neuron of the Softmax classifier, s i =F·η, F is the image feature vector of a training sample, η is the corresponding weight, and n is the number of categories to be classified;
[0042] Then according to the probability y i The training error is calculated:
[0043]
[0044] When i = k, θ ik =1, i represents the i-th category. When the original input belongs to category i, y k * =1;
[0045] The training error is used to backpropagate from the last layer of the convolutional neural network, and the cross entropy is used as the loss function. The Adam adaptive gradient optimizer is optimized, the initial learning rate is set to 0.01, the training rounds are 100 times, and the TensorFlow2.0 framework is used for experiments. The trained model is saved.
[0046] As described above, an intelligent storage object recognition device based on a lightweight network obtains the output results of each Densenet improved model, and then confirms the final result (i.e., the target category) according to the election rule or according to the weight coefficient of each Densenet improved model. The weight adjustment algorithm is used. For example, 5 Densenet improved models have a weight coefficient of 0.2. If one of the Densenet improved models predicts wrong n times multiple times, it will be in [0,2 n ], the new weight z=0.2-x*0.01. When z is 0, the system will issue an alarm and remove the Densenet improved model. Suppose the predicted results and weights of the 5 Densenet improved models are y1, y2, y3, y4, y5, z1, z2, z3, z4, z5, respectively. If y1, y2, y3 are the same, y4, y5 are the same, z1+z2+z3>z4+z5, the final result is y1. According to the different performance of the Densenet improved model, different weight value ratios can be set during initialization.
[0047] In the above-mentioned intelligent storage object recognition device based on lightweight network, the image collected by the image acquisition device needs to be preprocessed as follows before application:
[0048] (1) Grayscale;
[0049] (2) Remove image noise;
[0050] (3) Use the canny operator for edge detection;
[0051] (4) Find the smallest circumscribed square based on the edge and intercept the entire circumscribed square;
[0052] (5) Use bilinear interpolation to scale the square image to an appropriate size.
[0053] The above image preprocessing method is used to perform image standardization to solve problems such as too large resolution of the original image data, too much invalid information, and mixed image orientation (differences in size and shape of different parts).
[0054] That is, first use Gaussian blur to smooth the stains and scratches in the image; then use the canny edge detection algorithm to detect the part contour; then use the adaptive parameter loop detection method to solve the parameter problem of the above method to achieve the positioning of the part, and finally crop the image. The steps are as follows:
[0055] i) Select the initial Gaussian blur radius r. If the part edge information can be detected at this time, go directly to step iii), otherwise go to ii);
[0056] ii) Subtract 2 from the Gaussian blur radius r, reduce the blur radius and then return to step (1), and repeat this step to detect edge information;
[0057] iii) After the above two steps, the Gaussian blur radius parameter r′ that can accurately detect the edge information of the part is obtained, and the appropriate double threshold is calculated to crop and scale the image;
[0058] The selected Gaussian filter formula is as follows:
[0059]
[0060] where is the blur radius and is the standard deviation of the normal distribution.
[0061] Canny edge detection can crop the part image based on the positioning information of the part, retain the main area of the part, and remove the edge information of stains and scratches. The result of Canny edge detection mainly depends on the selection of three parameters, namely the Gaussian blur radius and the two thresholds of the gradient amplitude.
[0062] Based on the above problems, an adaptive parameter adjustment method based on parameter cycle detection is proposed to complete the adaptive selection of parameters, thereby achieving accurate detection of part edges. The adaptive selection of parameters is based on two basic assumptions:
[0063] a) The volume of noise information such as stains and scratches is much smaller than the volume of the part itself. As the Gaussian blur radius increases, the information of stains and scratches will inevitably be eliminated before the information of the part itself;
[0064] b) As long as the radius of the Gaussian blur is large enough, all the gradient information of the entire image can be smoothed out, making it impossible for the Canny operator to extract valid information.
[0065] Beneficial effects:
[0066] (1) The lightweight network-based intelligent storage object recognition device of the present invention provides a stable lighting environment, improves the adaptability to different external lighting and the recognition accuracy, and multiple cameras with different viewing angles arranged on it can synchronously and parallelly capture images, greatly improving the acquisition efficiency and accuracy. The accurate distance and angle relationship between the camera and the recognized object is established through the spatial coordinate system, and is fused and stored with the image information recognized by the corresponding camera. It is the first time to propose combining multi-observation viewing angle information with image information to improve image acquisition and subsequent recognition accuracy. The data enhancement operation is applied to increase the data volume;
[0067] (2) The intelligent warehouse object recognition device based on a lightweight network of the present invention proposes for the first time to combine the convolutional neural network with the characteristics of the channel intention order of ShuffleNet, the residual network in ResNet, the channel segmentation in Inception, the multiple residual networks in DenseNet and the DWConv in MobileNet to obtain the Densenet improved model, which further improves the accuracy and reduces the structure of the model compared with the traditional network and the lightweight network. The application of the Densenet improved model to complete the target category recognition not only has a small amount of data processing, but also has high recognition accuracy, and has great application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a schematic diagram of the overall structure of the intelligent storage object identification device based on a lightweight network of the present invention;
[0069] Figure 2 It is a schematic diagram of spatial three-dimensional coordinate representation of the intelligent storage object recognition device based on lightweight network of the present invention;
[0070] Figure 3 FIG. 1 is a method for geometric feature adaptive object image correction;
[0071] Figure 4 : is a basic structural diagram of the depthwise separable convolution of the present invention (where a is a structural diagram of the prior art, and b is a structural diagram after improvement of the present invention);
[0072] Figure 5 It is a structure diagram of the decremental feature reuse of the present invention;
[0073] Figure 6 A flow chart of the identification operation of the intelligent storage object identification device based on a lightweight network of the present invention;
[0074] Among them, 1-frame, 2-black backboard, 3-image acquisition device, 4-light source, 5-driving motor. DETAILED DESCRIPTION
[0075] The specific implementation modes of the present invention are further described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0076] In the description of the present invention, it is necessary to understand that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0077] An intelligent storage object recognition device based on a lightweight network, such as Figure 1 As shown, it includes a turntable for placing the to-be-identified parts, the turntable is connected to a turntable driving device (driving motor 5), and can rotate horizontally under the drive of the turntable driving device. The turntable is arranged in a frame 1 (a square frame, which is fixedly spliced by a plurality of aluminum alloy square tubes, and the corners of the frame are rounded transition), a black back plate 2 is arranged under the turntable, and a soft light cover is arranged on the outer cover of the frame 1;
[0078] Five image acquisition devices 3 are arranged around the turntable, and the centers of the fields of view are aligned with the center of the turntable. The image acquisition devices 3 are located above the turntable and fixed on the frame 1. The five image acquisition devices are arranged around and on the top of the frame, and the height differences between the five image acquisition devices and the turntable are different. A light source 4 is also fixed on the frame 1, and the light source 4 is arranged above the turntable.
[0079] The image acquisition device and the turntable driving device are respectively connected to the central processing unit;
[0080] The operation process of using the above device is as follows:
[0081] (a) The above device needs to be used to collect training and test samples. Various sample pieces are placed on the turntable, and the turntable and image acquisition device are turned on to collect images of various sample pieces. At the same time, the central processor runs the following program to perform data enhancement:
[0082] (a1) After the sample piece is placed at the center of the turntable, the central processing unit turns on the turntable drive device and all image acquisition devices. The turntable drive device drives the turntable to rotate horizontally. The image acquisition device regularly acquires images of the piece to be identified. While the image acquisition device is acquiring images, the central processing unit records the observation angle information of the image acquisition device and maps the acquired image with the observation angle information data to form the data set. The observation angle information G of the nth image acquisition device at the tth moment is G. t n The specific expression is as follows:
[0083]
[0084]
[0085] Among them, r n is the distance between the nth image acquisition device and the center of the turntable, θ n is the angle between the nth image acquisition device and the horizontal plane where the turntable is located, is the angle between the horizontal projection of the nth image acquisition device at the tth moment and the x-axis, is the angle between the horizontal projection of the nth image acquisition device and the x-axis at the initial moment, and the x-axis is the coordinate system established with the center of the turntable as the origin, the horizontal plane where the turntable is located as the x-axis, and the plane where the y-axis is located (the schematic diagram of the spatial coordinate representation is as follows Figure 2 ), V represents the rotation speed of the turntable;
[0086] (a2) Enhance the observation angle information data of the image acquisition device in the data set, and the enhanced observation angle information H of the nth image acquisition device at the tth moment is t n The specific expression is as follows:
[0087]
[0088]
[0089] in, is the rotation angle of the enhanced data (the state corresponding to the enhanced data is less than or equal to 20% left or right or up and down translation or the random rotation angle of clockwise or counterclockwise is less than or equal to 30° compared to the state corresponding to the enhanced object) compared to the observation angle information of the nth image acquisition device at the tth moment, [d x ,d y ] respectively represent the translation components of the enhanced data in the x-axis and y-axis directions compared to the observation angle information of the n-th image acquisition device at the t-th moment;
[0090] In addition, the image needs to be preprocessed as follows before application:
[0091] Grayscale; remove image noise; use the Canny operator for edge detection; find the smallest circumscribed square based on the edge and intercept the entire circumscribed square; scale the square image to an appropriate size using bilinear interpolation.
[0092] The above image acquisition and preprocessing are as follows:
[0093] S1-1 uses multi-threading in the UI interface to open cameras at multiple angles, places industrial objects in a multi-view part data acquisition system based on diffuse reflection lighting, and collects images from multiple angles;
[0094] S1-2 uses a geometric feature adaptive object image correction method for images collected from multiple perspectives. The specific algorithm is as follows ( Figure 3 shown):
[0095] (1) Select the initial Gaussian blur radius r. Under this radius, the edge information of most stains and scratches is erased and will not be detected. If the edge information of the part can be detected at this time, it means that under this radius, the Canny operator can detect the edge of the part without detecting noise, and directly enter step (3). At this time, r is large, and for some parts, their own edge information is also smoothed, resulting in the Canny operator being unable to detect the edge of the part body, and then entering step (2).
[0096] (2) Subtract 2 from the Gaussian blur radius r, reduce the blur radius and return to step (1). When the canny operator cannot detect edge information, steps (1) and (2) are repeated, so it is called a parameter loop detection algorithm. In the acquired part image, the edge information of the part has the largest amplitude and the highest continuity. Therefore, in the loop detection process, it can be guaranteed that the first detected edge information must be located at the edge of the part, rather than stains and scratches.
[0097] (3) After the above two steps, the Gaussian blur radius parameter r′ that can accurately detect the edge information of the part is obtained. Under this radius, the canny operator is used to realize edge detection, and a black-and-white image containing only the edge information of the part is obtained. The value of each pixel is 0 or 255, 0 is black and represents the background, and 255 is white and represents the edge of the object. By traversing the image, the upper and lower limit coordinates of its edge information on the x-axis and y-axis can be obtained, thereby obtaining the main area of the part. Let be the axis coordinate of the leftmost white pixel, be the axis coordinate of the rightmost white pixel, be the axis coordinate of the topmost white pixel, be the axis coordinate of the bottommost white pixel, then the edge coordinate area of the contour information can be obtained as between and of the axis, and between and of the axis. This paper uses square images as the standard form of data. In order to obtain complete image information and facilitate operations such as translation and rotation in data enhancement, the intercepted area is expanded 30 pixels around, that is, the upper left corner coordinates are used as the side length to determine the final intercepted area. The image intercepted in this way will make the detected part located in the center of the image and occupy the main part of the image. At this point, the positioning and cutting of parts are completed.
[0098] After many experiments, the initialization parameter r=9 of Gaussian blur radius was finally selected. When the double threshold was fixed to , the algorithm can achieve the best effect and realize error-free positioning of all part data sets. According to the positioning information of the part, the part image can be cropped, the main part area is retained, and the blank information area is cropped. However, the pixel size of the image obtained at this time is related to the volume of the part, and the pictures of parts of different volumes have different sizes. In order to obtain a more standard data format, the cropped image needs to be scaled. In order to reduce the information loss caused by scaling, the bilinear interpolation method is used to complete the scaling of the image, and all images are scaled to a uniform size of 80*80 pixels to complete the preliminary construction of the data set. After actual observation, the image under this pixel can take into account various parts of different volumes, control the image size while retaining as much detail information as possible. For small-volume parts, cropping and scaling can make the details of the main part of the part more prominent, and the number of noise points and scratches in the whole picture is also much less.
[0099] S1-3 performs operations such as translation, flipping, rotation, and contrast change on the image obtained in step S1-2, so that the image changes slightly without affecting the object features, and uses left-right translation and up-down translation transformation within 20%, random rotation within 30° clockwise and counterclockwise, and horizontal / vertical flipping and contrast change to enhance the data set;
[0100] S1-4 uses the UI interface (written in PyQt5) software to annotate the collected and preprocessed images and place them in a folder with a specified path. At the same time, the software puts part of the data into the training set and the test set in an 8:2 ratio. The specific paths are as follows:
[0101] New_data / train / camera0 / 0 / ...jpg
[0102] New_data / train / camera1 / 0 / ...jpg New_data is the name of the dataset, camera is the shooting angle, and the individual numbers are the labels of the parts.
[0103] New_data / test / camera0 / 0 / ...jpg Extract a part from the train set as the test set
[0104] New_data can set other path names by itself.
[0105] S1-5 determines whether the number of images of the same type collected in the training set exceeds 200. If so, continue with the following steps. If not, enter step S1-3 again.
[0106] (b) After obtaining a sufficient amount of training test samples through the above processing, the training test samples corresponding to the image acquisition devices are used to train the Densenet improved model (i.e., the training test samples collected by image acquisition devices I, II, III, IV and V are used to train the Densenet improved models I, II, III, IV and V, i.e., five Densenet improved models are obtained through training). The training process is a process of continuously adjusting the model parameters using the images of objects of known categories collected by the corresponding image acquisition devices as input and the corresponding category probabilities of the objects as theoretical output. The termination condition of the training is that the upper limit of the number of training times is reached (specifically, the training is to use the training error to backpropagate from the last layer of the convolutional neural network in sequence, use cross entropy as the loss function, and optimize with the Adam adaptive gradient optimizer. The initial learning rate is set to 0.01, the number of training rounds is 100, and the TensorFlow2.0 framework is used for experiments. The trained model is saved, and the ratio of the number of training sets to the number of test sets is 4:1);
[0107] Densenet improved model (specifically, it is a deep separable convolution-decreasing feature reuse structure, in which the basic structure of deep separable convolution is as follows Figure 4 As shown, the decreasing feature reuse structure (Back Bone structure) is as follows Figure 5 As shown), it includes a main convolutional layer (main Conv), a feature extraction layer, a first network block, a transition layer, a second network block and a classification layer connected in sequence;
[0108] The first network block and the second network block have the same structure. The network block includes a first Back Bone, a second Back Bone, a third Back Bone and a fourth Back Bone which are connected in sequence. The output of the first Back Bone is also the input of the second Back Bone, the third Back Bone and the fourth Back Bone, and the output of the second Back Bone is also the input of the third Back Bone and the fourth Back Bone.
[0109] The Back Bone in the network block includes Channel Split, the first convolution layer, the second convolution layer, the third convolution layer, the first depth separation convolution layer, the second depth separation convolution layer, Concat, and Channel Shuffle, wherein Channel Split, the first convolution layer, the first depth separation convolution layer, the second convolution layer, Concat, and Channel Shuffle are connected in sequence, the second depth separation convolution layer and the third convolution layer are connected in parallel with the first convolution layer, the first depth separation convolution layer, and the second convolution layer, the second depth separation convolution layer is connected to Channel Split, and the third convolution layer is connected to Concat;
[0110] The transition layer includes a feature extraction layer;
[0111] The classification layer includes a sequentially connected average pooling layer and a Softmax classifier;
[0112] The specific differences between the network structure of the present invention and the Densenet model are shown in Table 1 below:
[0113] Table 1
[0114]
[0115] (c) Complete the identification of the part to be identified (such as Figure 6 shown):
[0116] Place the part to be identified on the turntable, turn on the turntable and five image acquisition devices, and the five image acquisition devices respectively collect pictures of the part to be identified. The central processor obtains the pictures and pre-processes them, and then inputs the pictures of the part to be identified corresponding to each image acquisition device into the trained Densenet improved model corresponding to the image acquisition device. Each Densenet improved model outputs the target category, and the final result is confirmed according to the election rules.
[0117] In order to verify the accuracy of the method of the present invention, this embodiment is experimented on a private data set, and the experimental results are shown in Table 2 (wherein Densenet, SqueezeNet and MobileNet refer to replacing the Densenet improved model in this embodiment with a Densenet model, a SqueezeNet model and a MobileNet model respectively):
[0118] Table 2
[0119] Model Densenet The present invention SqueezeNet MobileNet Accuracy 96.27% 98.71% 94.92% 95.93%
[0120] From the above results, it can be found that the present invention can improve the recognition accuracy while greatly reducing network parameters, and has great application prospects.
[0121] It has been verified that the intelligent warehouse object recognition device based on a lightweight network of the present invention provides a stable lighting environment, improves the adaptability to different external lighting and the recognition accuracy, and multiple cameras with different viewing angles arranged thereon can synchronously and parallelly capture images, greatly improving the acquisition efficiency and accuracy, and establishing an accurate distance and angle relationship between the camera and the recognized object through a spatial coordinate system, and merging and storing them with the image information recognized by the corresponding camera respectively. It is the first time to propose combining multi-observation viewing angle information with image information to improve the accuracy of image acquisition and subsequent recognition, and the application of data enhancement operations increases the amount of data; it is the first time to propose combining the characteristics of the ShuffleNet channel intention order, the residual network in ResNet, the channel segmentation in Inception, the multiple residual networks in DenseNet, and the DWConv in MobileNet in the convolutional neural network to obtain the Densenet improved model, which further improves the accuracy and reduces the structure of the model compared with the traditional network and the lightweight network, and the application of the Densenet improved model to complete the target category recognition not only has a small amount of data processing, but also has high recognition accuracy, and has great application prospects.
[0122] Although specific embodiments of the present invention are described above, those skilled in the art should understand that these are merely examples and that various changes or modifications may be made to these embodiments without violating the principles and essence of the present invention.
Claims
1. An intelligent storage object recognition device based on a lightweight network, characterized in that: It includes a turntable for placing the to-be-identified parts, the turntable is connected to the turntable driving device and can rotate horizontally under the drive of the turntable driving device; There are more than two image acquisition devices arranged around the turntable, with the centers of the fields of view aligned with the center of the turntable, and the image acquisition devices are arranged at different positions of the turntable and have different height differences with the turntable; The image acquisition device and the turntable driving device are respectively connected to the central processing unit, and the central processing unit obtains the picture of the object to be identified through the image acquisition device, and then inputs it into the trained Densenet improved model, and the trained Densenet improved model outputs the target category; The improvement of the Densenet improved model compared to the Densenet model is that the Back Bone in the network block includes Channel Split, a first convolutional layer, a second convolutional layer, a third convolutional layer, a first depth separation convolutional layer, a second depth separation convolutional layer, Concat, and Channel Shuffle, wherein Channel Split, the first convolutional layer, the first depth separation convolutional layer, the second convolutional layer, Concat, and Channel Shuffle are connected in sequence, the second depth separation convolutional layer and the third convolutional layer are connected in parallel with the first convolutional layer, the first depth separation convolutional layer, and the second convolutional layer, the second depth separation convolutional layer is connected to ChannelSplit, and the third convolutional layer is connected to Concat; During the data collection process, the device establishes a Cartesian coordinate system with the center of the turntable as the origin. When the object is placed at the center of the turntable and starts to collect images, the real-time spatial polar coordinate description of the camera's view relative to the object is calculated to describe the spatial posture of the camera unit at different angles, thereby marking the observation view information based on the spatial polar coordinate representation in the acquired object image data, including the recognition angle and distance parameters, to increase the efficiency and accuracy of the subsequent recognition process; The central processing unit runs the following program to perform data enhancement: (1) After the object to be identified is placed at the center of the turntable, the central processing unit turns on the turntable drive device and all image acquisition devices. The turntable drive device drives the turntable to rotate horizontally. The image acquisition device regularly acquires images of the object to be identified. While the image acquisition device is acquiring images, the central processing unit records the observation angle information of the image acquisition device and maps the acquired image with the observation angle information data to form a data set. The observation angle information G of the nth image acquisition device at the tth moment is G. t n The specific expression is as follows: Among them, r n is the distance between the nth image acquisition device and the center of the turntable, θ n is the angle between the nth image acquisition device and the horizontal plane where the turntable is located, is the angle between the horizontal projection of the nth image acquisition device at the tth moment and the x-axis, is the angle between the horizontal plane projection of the nth image acquisition device and the x-axis at the initial moment, the x-axis is the x-axis in the coordinate system established with the center of the turntable as the origin, the horizontal plane where the turntable is located as the x-axis, and the plane where the y-axis is located, and V represents the rotation speed of the turntable; (2) Enhance the observation angle information data of the image acquisition device in the data set, and the enhanced observation angle information H of the nth image acquisition device at the tth moment is t n The specific expression is as follows: in, is the rotation angle of the enhanced data compared to the observation angle information of the nth image acquisition device at the tth moment, [d x ,d y ] respectively represent the translation components of the enhanced data in the x-axis and y-axis directions compared to the observation angle information of the n-th image acquisition device at the t-th moment.
2. According to claim 1, a light-weight network-based intelligent storage object identification device is characterized in that: The rotating disk is arranged in a frame, the image acquisition device is fixed on the frame, and a light source is also fixed on the frame; A black back plate is arranged under the turntable; The turntable driving device is a driving motor.
3. The intelligent storage object identification device based on lightweight network according to claim 2 is characterized in that: The light source is arranged above the turntable, and the frame is covered with a soft light cover; The frame is a square frame; There are five image acquisition devices, which are arranged around and on the top of the frame respectively, and the height differences between the five image acquisition devices and the turntable are all different; The corners of the frame are rounded; the frame is formed by fixedly splicing a plurality of aluminum alloy square tubes.
4. The intelligent storage object identification device based on lightweight network according to claim 1 is characterized in that: The Densenet improved model includes a main convolutional layer, a feature extraction layer, a first network block, a transition layer, a second network block, and a classification layer connected in sequence; the first network block and the second network block have the same structure; The ratio of the number of training sets to the number of test sets of the Densenet improved model is 4:
1.
5. The intelligent storage object identification device based on lightweight network according to claim 4 is characterized in that: The network block includes a first Back Bone, a second Back Bone, a third Back Bone and a fourth Back Bone connected in sequence, the output of the first Back Bone is simultaneously the input of the second Back Bone, the third Back Bone and the fourth Back Bone, and the output of the second Back Bone is simultaneously the input of the third Back Bone and the fourth Back Bone; The transition layer includes a feature extraction layer; The classification layer includes an average pooling layer and a Softmax classifier.
6. The intelligent storage object identification device based on lightweight network according to claim 3 is characterized in that: There are five Densenet improved models in total, each of which corresponds to an image acquisition device one by one. The training process is a process of continuously adjusting the model parameters using the images of objects of known categories acquired by the corresponding image acquisition device as input and the corresponding category probabilities of the objects as theoretical output. The termination condition of the training is reaching the upper limit of the number of training times.
7. The intelligent storage object identification device based on lightweight network according to claim 6 is characterized in that: After obtaining the output results of each Densenet improved model, the final result is confirmed according to the election rules.
Citation Information
Patent Citations
High-precision image navigation positioning method based on local neighborhood map
CN112129281A
Three-dimensional acquisition equipment
CN211178345U