Multi-target detection statistical method and system for unmanned warehouse goods
By adopting the improved FP-YOLO algorithm in the multi-object detection statistics of unmanned warehouse goods, the problems of slow detection speed and inaccurate results are solved, and more efficient multi-object detection and statistics are achieved.
Patent Information
- Application Number
- CN202510101559.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, in the multi-objective detection statistics of unmanned warehouse goods, there are problems such as slow detection speed and inaccurate detection results.
The FP-YOLO algorithm based on YOLOv3 algorithm is adopted to improve the feature extraction network structure, introduce loss rejection terms, and use DenseNet optimization to improve the detection performance of the model.
The problems existing in multi-object detection in and out of the database have been improved, the overall mAP, training loss reduction speed and model detection speed have been improved, and the detection effect has been improved.
Smart Images

Figure CN119942081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection, and more specifically, to a multi-target detection statistical method and system for unmanned warehouse cargo. Background Art
[0002] Computer vision technology is an artificial intelligence-based technology that analyzes and processes images and videos to achieve target counting and statistical detection and statistics. It has a wide range of applications in various fields, such as traffic management, retail, healthcare, etc. How to use computer vision technology for multi-target detection, counting and statistics, and explore the current status and future development of its application is a trend.
[0003] Object counting and statistical detection and statistics refer to the counting and statistical detection and statistical analysis of objects in images or videos. In the past, people usually used manual methods to count objects, which was time-consuming, labor-intensive and prone to errors. However, using computer vision technology for object counting and statistical detection and statistics can achieve fast, accurate and automatic results. The following will introduce some commonly used computer vision technology methods.
[0004] First of all, object detection is one of the important tasks in computer vision. Its goal is to find objects of interest in images or videos. Common object detection methods include feature-based methods, deep learning-based methods, and convolutional neural network-based methods. These methods can achieve accurate detection and positioning of targets in different scenarios, providing a basis for subsequent target counting and statistical detection and statistics.
[0005] Secondly, target tracking refers to the continuous tracking of a target in continuous images or video frames. Target tracking methods are often used for real-time tracking and statistical analysis of targets. Common target tracking methods include color histogram-based methods, correlation filter-based methods, and deep learning-based methods. These methods can track the trajectory of targets in continuous images or videos and provide a basis for target number statistics.
[0006] Then, target counting refers to counting the number of targets in a scene. Target counting methods can be classified according to the number of targets, such as single target counting and multi-target counting. Common target counting methods include sensor-based methods, density estimation-based methods, and deep learning-based methods. These methods can achieve accurate counting and statistical detection and statistics of the number of targets according to different scenarios and needs.
[0007] Finally, target statistics refers to the statistical analysis of the target's attributes. Target statistics can include the target's size, speed, dynamic trajectory, etc. Common target statistics methods include feature extraction-based methods, region-based methods, and deep learning-based methods. These methods can achieve accurate statistics and analysis of target attributes and provide data support for subsequent applications.
[0008] The central problem of computer vision is to solve how to parse the information that the computer wants from the image. In machine vision, image processing mainly consists of three tasks: image classification, target detection, and image semantic segmentation. Image classification is to solve the problem of the existence of the target in the image, which belongs to the category. The computer extracts the target features (shape, size, number, brightness, color, etc.) of the entire image through artificially designed features or feature learning methods, and then uses the classifier to distinguish the different object categories. Therefore, how to extract the features of the image is crucial in image classification. In traditional machine learning, a function bag (BoF) method is used to express the local features of the image through vector quantization, and the features of the entire image are expressed through histograms. After that, the feature expression in image classification includes: Fisher r vector and local aggregation descriptor vector. The Fisher vector expresses the information of the feature vector in a richer way. VLAD aims to reduce the amount of memory used to express the feature quantity. The image classification method based on deep learning can perform a deep hierarchical feature description of the image in a supervised or unsupervised manner, replacing the work of artificial feature design or selection of image features. With the introduction of convolutional neural networks (CNNs), deep learning has far surpassed the feature design plus classifier design method in terms of target recognition accuracy, and has gradually become the preferred solution for the current 1,000-category image classification task.
[0009] Computer vision technology is widely used and has great development potential. At present, computer vision technology has been applied to many fields. For example, in traffic management, computer vision technology can realize the counting and statistical detection and statistics of vehicles and pedestrians, thereby providing accurate traffic flow information and traffic congestion prediction. In the retail industry, computer vision technology can count and statistically detect and count customers, and then analyze and predict shopping behavior and consumption trends. In healthcare, computer vision technology can be applied to medical image analysis and disease diagnosis, thereby improving medical efficiency and accuracy. However, computer vision technology also faces some challenges and problems. First, the quality of image and video data has a great impact on the accuracy of computer vision results. Problems such as noise, blur, illumination changes and occlusion in images or videos may lead to errors in recognition and counting. Second, large-scale data collection and processing require powerful computing and storage resources, which is a challenge for some applications. In addition, the algorithms and models of computer vision technology need to be continuously optimized and improved to adapt to changes in different scenarios and needs.
[0010] Object detection is to solve the problem of how to accurately mark the location of a specific object in an image. Different from the classification algorithm, object detection also needs to solve the problem of determining the location of the object in the image. Object detection is also a key link and task at present, and it is also an important branch of computer vision research in recent years. It is mainly divided into classification object detection based on traditional features and convolutional neural network object detection based on deep learning.
[0011] In traditional target detection methods, the target feature extraction is first completed by fixing the sliding window, and then the classifier is used to determine the target category and output its location to achieve target detection. For example, Haar feature operator and adaptive boosting (Adaboost) are often used to solve practical face detection and pedestrian detection problems; the combination of gradient directional histogram (HOG) features and support vector machine (SVM) is widely used for pedestrian detection.
[0012] In recent years, with the high efficiency of convolutional neural networks in feature extraction and classification, fast and effective target detection based on deep learning framework has gradually become a new direction for target recognition in computer vision. Deep learning is a learning method that uses deep neural networks. This deep neural network replaces the traditional manual feature extraction and can automatically extract features of samples with sufficient sample learning data. With the development of image recognition research based on deep learning, image recognition technology has gradually entered the practical stage of daily life. Compared with other image classification algorithms, deep convolutional neural networks (CNNs) are different from the manually designed filters in traditional algorithms, use relatively less image preprocessing, and have the advantage of being independent of prior knowledge and artificial feature design. In 2012, Alex Krizhevsky proposed the AlexNet convolutional neural network, which opened the prelude to the widespread application of deep learning.
[0013] In the field of deep learning-based object detection, algorithms are mainly divided into two-step object detection networks based on candidate boxes and end-to-end object detection based on regression methods. Among them, the object detection network based on candidate boxes is the most famous R-CNN series. In 2014, Ross Girshick et al. borrowed the sliding window idea in traditional object detection and proposed the R-CNN algorithm based on selective search (SS). The SS algorithm merges similar pixels in the image based on the color, texture, area, position, etc. of the image, and repeatedly groups regions with similar colors and textures at different thresholds through the classification network to find out the location information where the target may exist, thereby generating 2000 candidate regions where the target may exist for subsequent network processing. In R-CNN, it is necessary to perform CNN network feature extraction on each candidate box, and learn classification for each category separately through multiple classifiers (such as SVM), so as to select the candidate box through regression estimation method to achieve classification and positioning of the target. Although the R-CNN network effectively converts the detection problem into a classification problem. However, when extracting features for each candidate box, it is necessary to repeatedly use multiple feedforward networks to integrate repeated regions, which will generate a large amount of redundant calculations and make the real-time detection speed of the overall network very slow.
[0014] In 2015, Ross proposed the FastR-CNN network to improve the defects in the R-CNN network. The network uses the region of interest convergence method to sample multiple pre-selected regions generated by the region search method on the original image, and then projects them onto the convolutional features to achieve partial sharing of the extraction. The FastR-CNN network replaces the serial feature extraction method of R-CNN and directly uses a CNN to extract features from the input image. Although FastR-CNN shares most of the calculations in network extraction, the real-time detection speed of the network has been greatly improved. Since the use of the SS algorithm will still generate a large number of redundant candidate frames, affecting the overall detection speed of the network, in the same year, Shaoqing Ren et al. proposed the FasterR-CNN network that uses a fully convolutional network regional prop network (RPN) to replace the SS algorithm to generate candidate regions. RPN introduces the prior frame method, which performs regression processing by fixing prior frames of different scales and aspect ratios around the pooled area; at the same time, RPN will generate multiple types of target region proposals with different aspect ratios in feature extraction to complete the network output of object similarity scores and detected object coordinates on the image. Object detection based on regression methods is an end-to-end fast object detection method. In applications, the YOLO series and SSD series are the most common. In 2016, Joseph Redmon et al. proposed the YOLO algorithm for end-to-end detection. The detection network directly divides the input image into S*S grids for feature extraction and then uses CNN to directly output the regression results, treating the object detection problem as a regression problem. Since this type of method only needs to perform a feedforward calculation on image extraction, the network's calculation speed is usually very fast and can meet the real-time detection requirements in many application scenarios. Different from the two-stage object detection algorithm of the R-CNN series, end-to-end object detection directly extracts features of the target in a single model and then outputs the model detection results, without preprocessing the image or combining multiple models. The end-to-end series of algorithms draws on the advantages of the R-CNN series and directly uses regression to complete object detection. Looking at the development of target detection algorithms in recent years, although the R-CNN series of algorithms have high detection accuracy in machine vision, their detection speed cannot meet the high real-time detection requirements in autonomous driving. The YOLO series and SSD are typical first-order end-to-end detectors, which reduce the overall network calculation amount and therefore have significant effects in detection speed. In the target detection algorithm based on the candidate box, the network needs to generate a certain amount of target candidate areas first, then perform feature extraction through the convolutional network, and then perform target classification and frame position regression. The YOLO series directly performs target detection based on the entire image. When extracting features, YOLO will output all target information including categories and positions at one time, and directly regress the target.
[0015] In modern logistics supply chain management, warehouse in and out management is one of the key links. Accurate and efficient warehouse in and out control can greatly improve the operating efficiency of the entire logistics system and reduce losses and waste. However, traditional warehouse in and out management usually relies on manual inspection and record keeping, which is inefficient and prone to errors. In addition, a large amount of manpower costs are required, and it is difficult to achieve all-weather automatic monitoring. Nowadays, with the rapid development of computer vision technology, it is possible to use these technical means to achieve automated warehouse in and out management. By deploying visual sensors such as cameras in warehouses or distribution centers, the entry and exit of goods can be monitored in real time. Combined with advanced image analysis algorithms, it can accurately detect and identify various types of items, automatically count the number, and provide timely feedback. This can not only greatly improve the efficiency and accuracy of warehouse in and out management, but also achieve all-weather automatic monitoring and reduce labor costs.
[0016] In the detection and statistics of multi-target materials entering and leaving the warehouse, the recognition and discrimination of target objects in complex environments is a difficult challenge and one of the key tasks that need to be solved. Object detection is not only an important branch of computer vision research, but also a key link and task in the detection and statistics of multi-target materials entering and leaving the warehouse. Due to the complexity of the actual situation, it is difficult to significantly improve the detection and statistics of multi-target materials entering and leaving the warehouse based on traditional object detection.
[0017] In the target detection application of multi-target materials entering and leaving the warehouse, the detection effect is easily affected by lighting, shooting angle, occlusion, etc. The current R-CNN series of algorithms are the most commonly used deep learning algorithms in the field of target detection. Although this series of algorithms has high detection accuracy in experiments, due to its complexity, it will produce high latency in real-time detection, making it difficult to promote this series of algorithms. The end-to-end YOLO series of algorithms based on regression methods not only reduces the complexity of the convolutional network, but also meets the requirements of real-time detection. Based on the latest YOLOv3 algorithm, this paper proposes a FP-YOLO algorithm suitable for multi-target material detection in and out of the warehouse.
[0018] In summary, the use of computer vision technology for target counting and statistical detection and statistics can provide fast, accurate and automated results, and it has been widely used in many fields. In the future, with the continuous advancement and development of artificial intelligence technology, computer vision technology will play an important role in more fields, providing us with more convenient and accurate information analysis and decision support.
[0019] However, the above technology has problems such as slow detection speed and inaccurate detection results in practical applications.
[0020] In view of the above problems, there is an urgent need for a multi-target detection statistical method and system for unmanned warehouse goods. Summary of the invention
[0021] In order to solve the deficiencies in the prior art, the present invention provides a multi-target detection statistical method and system for unmanned warehouse cargo, which improves the detection accuracy through a target detection method.
[0022] The present invention adopts the following technical solution.
[0023] The first aspect of the present invention relates to a multi-target detection statistical method for unmanned warehouse goods, and the method comprises the following steps: S1, obtaining storage in-and-out multi-target material data from the data space of a storage management enterprise; S2, performing sample expansion operation on the storage in-and-out multi-target material data to obtain the storage in-and-out multi-target material data; S3, performing data preprocessing operation on the expanded storage in-and-out multi-target material data to obtain the storage in-and-out multi-target material data; S4, constructing an FP-YOLO algorithm based on the YOLOv3 algorithm; S5, inputting the storage in-and-out multi-target material data obtained after preprocessing into the constructed FP-YOLO algorithm for feature extraction to obtain the result after feature extraction; S6, calculating and screening the first confidence and the second confidence for the result after feature extraction, performing feature fusion and DenseNet optimization on the calculated and screened data, and detecting and outputting the optimized data to obtain the target detection statistical result.
[0024] Preferably, in step 1, the warehousing in-and-out multi-objective material data is recorded in Pycharm, a software for data processing operations based on the Python programming language.
[0025] Preferably, in step 2, samples under different lighting conditions are expanded by processing such as sharpening, oversaturation and image enhancement to expand the training samples under different lighting conditions; in scenes with jittery shooting, the image is blurred.
[0026] Preferably, in step 3, a K-means algorithm is used to perform centroid clustering analysis on the KITTI dataset and the expanded dataset to determine the clustering range of the target box size in the algorithm's dataset; the IOU of the prior box and the sample's true bounding box is used as a distance metric, and based on the FP-YOLO system, a total of four feature maps of different scales are output in the YOLO layer, and each feature map generates three prior boxes of different sizes.
[0027] Preferably, in step 4, the feature extraction part of the overall architecture of the FP-YOLO system is obtained by adjusting the image size, building Darknet-62 and YOLO layer feature extraction output; after waiting for all borders to be traversed, the candidate bounding boxes are screened and calculated; the size of the feature map is adjusted to the same resolution, and feature fusion is performed on the adjusted feature map; and DenseNet is used to obtain the spliced feature output using feature information.
[0028] Preferably, in the YOLOv3 algorithm model training, the loss calculation is divided into position loss, recognition loss, and confidence loss.
[0029] Preferably, the position loss is:
[0030]
[0031] In the formula, when there is no target in the grid, λ obj =0, 1 when there is a target,
[0032] truth w , truth h is the height and width of the real target box, predict w , predict h is the height and width of the predicted bounding box output by the model,
[0033] i is the feature number, I is the number of features, x, y, w, h are the horizontal and vertical coordinates, width, and height of the center point of the predicted bounding box, and b is the current position of the predicted bounding box.
[0034] Preferably, the recognition loss is:
[0035]
[0036] c==truth class ? 1:0 is the category identified by the prior box and the actual target category truth class Is the function consistent?
[0037] c is the category number, and C is the total number of categories;
[0038] predict class is the predicted label classification.
[0039] Preferably, the confidence loss is:
[0040] loss3=(truth conf -predict conf ) 2
[0041] In the formula, truth conf is the confidence of the true bounding box, predict conf Confidence of the predicted bounding box.
[0042] The second aspect of the present invention relates to a multi-target detection and statistical system for unmanned warehouse goods using the method of the first aspect of the present invention; the system comprises an acquisition module, an expansion module, a preprocessing module, a construction module, an extraction module, and an output module; the acquisition module is used to acquire storage in-and-out multi-target material data from the data space of a storage management enterprise; the expansion module is used to perform sample expansion operations on the storage in-and-out multi-target material data to obtain the storage in-and-out multi-target material data; the preprocessing module is used to perform data preprocessing operations on the expanded storage in-and-out multi-target material data , obtaining multi-target material data of warehousing in and out; the construction module is used to construct an FP-YOLO algorithm based on the YOLOv3 algorithm; the extraction module is used to input the multi-target material data of warehousing in and out obtained after preprocessing into the constructed FP-YOLO algorithm for feature extraction to obtain the result after feature extraction; the output module is used to calculate and screen the first confidence level and the second confidence level for the result after feature extraction, perform feature fusion and DenseNet optimization on the calculated and screened data, and output the optimized data detection to obtain the target detection statistical result.
[0043] A third aspect of the present invention relates to a terminal, comprising a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described in the first aspect of the present invention.
[0044] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect of the present invention.
[0045] The beneficial effect of the present invention is that, compared with the prior art, the statistical method and system for multi-target detection of goods in unmanned warehouses in the present invention proposes an improved method for multi-target detection of goods entering and leaving the warehouse in terms of the feature extraction network structure based on the YOLOv3 algorithm, the overall algorithm architecture, and the loss calculation during model training.
[0046] The beneficial effects of the present invention also include:
[0047] 1. Aiming at the problem of YOLOv3 missing large-size targets in target detection, the YOLOv3 feature extraction network is analyzed, and based on the YOLOv3 feature extraction backbone network Darknet-52, an improved feature extraction backbone network Darknet-62 is proposed. This network structure enables the algorithm to output feature maps with a larger field of view than YOLOv3 in the YOLO layer to improve the problem of missing large-size targets in image detection.
[0048] 2. FP-YOLO proposes an algorithm architecture for parallel output of feature maps to solve the problem of network bloat and reduced detection speed caused by the deepening of network layers and the increase of output feature tensors; at the same time, in order to reduce the waiting problem in the calculation of the first confidence, second confidence and IOU of the bounding box during the YOLO detection process, a cache queue method is proposed to reduce space waste and time waiting in the calculation.
[0049] 3. It has been verified that the FP-YOLO algorithm has improved the problems in the multi-target detection of inbound and outbound storage proposed in this paper. Finally, the performance of the model before and after the improvement is evaluated by using the general evaluation criteria for image detection. According to the evaluation indicators, the training model under the FP-YOLO algorithm has improved in terms of overall mAP, training loss reduction speed, and model detection speed.
[0050] In summary, the above method should be applicable in constrained and limited multi-target scenarios and have better detection effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of a multi-target detection statistical method for unmanned warehouse goods according to the present invention. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solution and advantages of the present invention clearer and more accurate, the technical solution of the present invention is described in detail below through multiple specific implementation methods. The embodiments adopted by the present invention are only used to explain the present invention and are not used to limit the content of the present invention.
[0053] Figure 1 Schematic diagram of a multi-target detection statistical method for unmanned warehouse cargo in the present invention. Figure 1 As shown, the first aspect of the present invention relates to a multi-target detection statistical method for unmanned warehouse goods, and the method includes the following steps: S1, obtaining storage in-and-out multi-target material data from the data space of the warehousing management enterprise; S2, performing sample expansion operation on the storage in-and-out multi-target material data to obtain storage in-and-out multi-target material data; S3, performing data preprocessing operation on the expanded storage in-and-out multi-target material data to obtain storage in-and-out multi-target material data; S4, constructing an FP-YOLO algorithm based on the YOLOv3 algorithm; S5, inputting the storage in-and-out multi-target material data obtained after preprocessing into the constructed FP-YOLO algorithm for feature extraction to obtain the result after feature extraction; S6, calculating and screening the first confidence and the second confidence for the result after feature extraction, performing feature fusion and DenseNet optimization on the calculated and screened data, and outputting the optimized data detection to obtain the target detection statistical result.
[0054] In order to solve the existing technical problems and achieve the above-mentioned purpose, the technical solution provided by the present invention is a computer vision-based multi-target material detection and statistics algorithm for in-and-out warehouses. The algorithm is a new FP-YOLO algorithm based on the YOLOv3 algorithm. Its core idea is to improve the feature extraction network structure, introduce loss exclusion terms, and add some techniques that hardly increase the inference time to improve the overall performance of the model. Specifically, based on the YOLOv3 algorithm, the optimization of the FP-YOLO algorithm is proposed from the following three aspects to address the problems of in-and-out warehouse multi-target material detection and statistics: (1) In order to solve the problem of missed detection of large-size targets by the YOLOv3 algorithm, the FP-YOLO algorithm improves the feature extraction network structure in YOLOv3 to obtain feature maps of more sizes. (2) In order to solve the low recall rate and missed detection problems caused by mutual occlusion of materials, the FP-YOLO algorithm introduces two loss exclusion terms in the calculation of model training loss. (3) In order to improve the detection performance, the FP-YOLO algorithm uses DenseNet to optimize the low-resolution feature layer.
[0055] In the experimental section, we conduct extensive experiments on several real-world databases. Experimental results show that our method outperforms existing state-of-the-art object detection models.
[0056] In step S1, the warehouse storage and warehousing multi-objective material data is recorded in Pycharm, a software for data processing operations based on the Python programming language.
[0057] In order to solve the problems existing in the detection of multi-target materials in and out of the warehouse, this paper proposes an improved FP-YOLO system based on YOLOv3. The KITTI dataset is a widely used computer vision algorithm evaluation dataset; the 2D part of the KITTI dataset is mainly used in the training of this model, with a total of 2,000 images. The dataset contains real images collected from multiple scenes, each image can contain up to 45 targets, and there are different degrees of target occlusion and truncation. In addition, 2,000 actual images are added for joint training. In training, the ratio of the dataset to the test set is 8:2. Using a labeled detection dataset can improve the precise positioning of the model, while using a classification dataset enhances the model's robustness to categories.
[0058] Due to factors such as lighting and shooting angle, the detection effect may decrease, resulting in an increased false positive rate. To improve the generalization ability of the model, the FP-YOLO system refers to the data enhancement strategy of YOLOv3, and enhances the contours and details of the target in the image by adjusting the exposure, hue, and saturation. Based on the official KITTI dataset, 2,000 actual images were added, and the number of samples under different lighting conditions was expanded using common image processing methods.
[0059] In step 2, sample expansion is divided into two aspects:
[0060] First, samples under different lighting conditions are expanded. Through sharpening, oversaturation, and image enhancement, training samples under different lighting conditions are expanded. Existing scene images can extract more samples with clear features, especially in dark environments, where lack of light will lead to the loss of target features. Affected by lighting, the material features in the image may not be clear enough. Therefore, in sample production, white balance and exposure processing are used to enhance image details. The processed image provides richer material feature information than the original image, which is helpful for sample expansion. In addition, due to the difference in light intensity inside and outside the warehouse, the material information outside the warehouse may not be clear enough. Using the method of reducing exposure can obtain more sample information.
[0061] Secondly, in scenes with jittery shots, the camera may have a short focus time difference when the material is moving, resulting in blurred images. In the case of a small number of samples, blurring the image can provide more useful sample information for the dataset, thereby effectively expanding the training samples.
[0062] By introducing this artificially blurred data augmentation technology, we can not only simulate the image quality problems in the real world, but also enhance the robustness of the model to different imaging conditions. Specifically, blurring can improve the generalization ability of the model, reduce the risk of overfitting, and enhance the adaptability of the model.
[0063] In step S3, the size of the prior box needs to be modified. In the native YOLOv3 model, the default number of clusters K=9 is used to constrain the bounding box sizes corresponding to the feature maps of three different fields of view. Although this method facilitates YOLOv3 to determine the position of the predicted bounding box relative to the image and the size of the target to be detected in the target location, the size of the bounding box determined by this method is based on the PascalVOC dataset, so the prediction under the PascalVOC dataset can well contain the target features, but it is not suitable for the target detection of different types of materials in the dataset KITTI. If the YOLOv3 fixed prior box size is used, the converted predicted bounding box will produce a certain amount of drift with the actual target bounding box, so that the predicted bounding box cannot completely contain the characteristics of the target object and the precision of the target is low. Because there are not many types of targets in the training dataset, and the sizes of the targets are different from other datasets.
[0064] Therefore, the present invention uses the K-means algorithm to perform centroid cluster analysis on the KITTI dataset and the expanded dataset before training the FP-YOLO model to determine the clustering range of the target frame size in the algorithm's dataset. This method can help the model converge faster during the training process, and can also accurately predict the frame offset and target coordinates. At the same time, a more accurate prior frame size helps reduce the background impact in the predicted bounding box when predicting the frame output.
[0065] The K-means algorithm is a simple, fast and easy-to-implement data clustering analysis method. The main idea is: first divide the data into K data clusters, then calculate the average value of the point data within the cluster, and finally classify and adjust the points near the cluster to achieve effective data separation. Usually, the position of the bounding box of an object in an image is mainly determined by the coordinates of the left and right corners. Among them is the coordinate of the lower left corner of the bounding box, and is the coordinate of the upper right corner of the bounding box. However, in K-means clustering analysis, the correct metric must be used as input. Therefore, it is necessary to convert the coordinates corresponding to the target bounding box in the image into the normalized bounding box width wk and height wh required for K-means input according to the following formula.
[0066]
[0067] In YOLOv3, when using K-means clustering to perform cluster analysis of target bounding boxes, if the standard Euclidean distance is used as the distance metric, large-sized bounding boxes will introduce more errors than small-sized bounding boxes. Since the IOU value between the prior box and the true bounding box of the adjacent target is usually high, YOLOv3 proposes to use the IOU of the prior box and the true bounding box of the sample as the distance metric to reduce the impact of size differences on distance calculation. This paper uses the K-means algorithm to analyze the KITTI dataset and the extended dataset. Based on the fact that the FP-YOLO system outputs four feature maps of different scales in the YOLO layer, each feature map generates three different sizes of prior boxes. Therefore, when calculating the size of the prior box of the FP-YOLO system, the number of clusters K in K-means is set to 12.
[0068] The feature map size output by the improved YOLO layer is smaller than that of YOLO, which makes the size of the prior frame generated by FP-YOLO smaller than that of the prior frame in the YOLO algorithm, thereby improving the model's detection accuracy for small targets in the image. The three prior frames of different sizes set by YOLOv3 for each feature map have a large gap. Although the object detection similarity is high, the target frame may contain a lot of background and the object features may not be fully covered. In order to solve the problem of missing large objects, the FP-YOLO system uses feature maps of sizes 16×16 and 8×8 to jointly predict large-sized targets in the image. With a moderate prior frame gap, the coordinate formula can be used to convert the predicted frame that fits the target boundary better.
[0069] The feature extraction part of the overall architecture of the FP-YOLO system mainly consists of three parts: image resizing, building Darknet-62, and YOLO layer feature extraction output. In the multi-scale feature output process of the YOLO layer of FP-YOLO, eight downsamplings are required. Therefore, the height and width of the image fed into the FP-YOLO network should be integer multiples of 128. If the height and width of the image are integer multiples of 128, the image is directly scaled to the size of 896*896 required by the extraction network; if the height and width of the image are not integer multiples of 128, the shortest side of the image is first calculated to 896 scaling ratio, and the longest side of the original image is adjusted to 896 according to the scaling ratio; secondly, the shortest side of the image is gray-filled to 896.
[0070] In order to reduce the amount of network calculation, the image pixel values should also be normalized before feature extraction. The amount of calculation in network training will affect the running detection speed performance of the entire model. In YOLOv3, the YOLO layer output tensor calculation formula is as follows, where N anchor The number of prior boxes generated for each grid, N Fi The number of 1*1 grids generated for each feature map i, N class is the number of categories in the YOLO detection part. j is the total number of feature maps.
[0071]
[0072] The present invention mainly studies the multi-target detection of goods by the FP-YOLO system in an unmanned warehouse, and the number of detection categories Nclass is 6. In the output part of the YOLO layer of the algorithm, a grid with a side length of 1*1 will be generated in each feature map, and six corresponding prior frames of different sizes will be generated in each grid (three prior frames in a feature map have fixed sizes), and each prior frame contains 5 information (center coordinates x, y, height and width w, h, and a confidence level to reflect whether there is a target and a bounding box in the grid). The single-threaded structure used in the YOLOv3 algorithm has a large amount of calculation, and high-configuration GPU training is required for training with more image samples. Therefore, in order to improve the overall detection speed of the algorithm model and reduce the amount of calculation, the FP-YOLO system proposes to use multiple low-configuration GPUs or CPUs to perform parallel detection on the four feature map outputs of the YOLO layer.
[0073] During the YOLO detection process of YOLOv3, it is necessary to wait until all borders are completely traversed before the next step of candidate bounding box screening and calculation is performed. Therefore, the network will waste waiting time in tensor calculation. Based on the YOLO detection mechanism in YOLOv3, the present invention proposes to use a cache area queue method to reduce the calculation waiting time for each candidate border calculation and screening wait during the detection process of the FP-YOLO system. According to the different stages of YOLO detection, the queue in the cache area is divided into four parts: the first confidence queue generated in the first confidence screening stage, the second confidence queue generated in the second confidence calculation stage, the bounding box queue generated in the second confidence screening stage, and the IOU queue generated in the IOU calculation stage of the candidate bounding box and the real target bounding box.
[0074] In the above figure, let the candidate bounding box generated by the feature map be Xn. In the first confidence part, first divide the feature map into grid units according to size, then determine the grid where the target is located according to the target center, generate three sizes of prior boxes for each grid and calculate the position of each bounding box, and count the number of all current candidate bounding boxes. Arrange the candidate boxes according to size and select one of the candidate bounding boxes in turn. If the candidate box is not the last one, calculate the confidence score of the target in each bounding box; set the bounding box confidence score threshold to select the qualified candidate bounding box. If the confidence score of the bounding box is less than or equal to the set threshold, it means that the probability of the target existing in this bounding box is low and should be discarded; if the confidence score is greater than the set threshold, it is stored in the first confidence queue of the cache from top to bottom. The bounding boxes in the queue do not need to wait for all nx to be screened and directly calculate the second confidence in turn, and then store them in the second confidence queue from top to bottom. When the second confidence calculation of all bounding boxes in the first confidence queue is completed (assuming that the bounding boxes in the queue at this time are x1, x2, x3...x j), select the bounding box with the second highest confidence from the second confidence queue and record it as A. Traverse all bounding boxes in the second confidence queue (excluding bounding box A), and stop if the traversal is completed. If the traversal is not completed, calculate the intersection over union (IOU) of the bounding box and A at this time; set an IOU threshold, when the IOU of the bounding box and A is greater than the set threshold, it is determined that the overlap rate with A is too high and should be discarded; when the IOU of the bounding box and A is less than or equal to the set threshold, it is retained in the bounding box queue in the cache area. Assuming that the number of all categories is m, then a bounding box contains the center coordinates t of the bounding box coordinates. x , t y , the width and height of the bounding box t w , t h , the first confidence level P obj , all second confidence levels P l .
[0075] In the second confidence calculation screening, there is no need to wait for all bounding boxes to be screened, and you can select the bounding boxes from the cache area's bounding queue (assuming that the bounding boxes in the queue are x2, x3, x4...x k ) Take out the candidate bounding boxes one by one and complete the IOU calculation of the overlap rate with the actual border. After the calculation, the candidate bounding boxes will be sent to the IOU queue one by one (assuming that the bounding boxes in the queue are x2, x3, x4...x j ) After waiting for the IOU calculation of all categories of the four feature maps to be completed, the NMS (non-maximum suppression) method is used to sort the possible target bounding boxes according to the category probability, and the bounding box with the highest probability is selected. The IOU of other bounding boxes and the bounding box with the highest probability is calculated respectively. After using the set threshold to filter out the bounding boxes with a high overlap rate with the bounding box with the highest probability, the remaining bounding boxes are sorted and the above process is repeated until the best target bounding box under the category is selected as the predicted bounding box of the object with a size of 832*832. Finally, the scale t is scaled according to the predicted bounding box w ,t h , map the predicted bounding box size under the size of 832*832 to the original image, and finally use the translation scale t x ,t y ,The predicted bounding box position is output in the original image, and finally the object detection in the image is realized.
[0076] Compared with traditional handcrafted feature-based algorithms, deep learning-based algorithms usually obtain low-level and high-level image features through convolutional neural networks and other feature extractors. These features have different resolutions, so how to effectively process and fuse these multi-scale features has a crucial impact on the network that uses them for reasoning. Feature Pyramid Network (FPN) has done pioneering work by combining multi-scale features using a top-down approach. Path Aggregation Network (PANet) further adds a bottom-up path to FPN. The specific method is to first resize the feature maps to the same resolution and then add them together. In this method, features of different scales are treated equally.
[0077] The same method is used to fuse features in the neck of the network. Assuming the input image size is 640×640, the PANet structure uses the features extracted by the backbone network as input, which are 80×80, 40×40, and 20×20 features at P3 to P5 levels. FPN uses a top-down approach to fuse deep features with underlying features through upsampling to obtain predicted feature maps. This operation passes strong semantic features from the upper layer to the lower layer, enhancing the model's ability to learn image features, but may lose some positioning features. Therefore, PAN is added after FPN, which complements FPN and passes strong localized features from bottom to top. This comprehensively improves the robustness and learning performance of the model.
[0078] The aggregation process of multi-scale features can be expressed as:
[0079]
[0080] Where x and y are different layers on the feature pyramid, y is one or more lower layers, and x is one or more upper layers.
[0081] is the output feature of a certain layer, is the intermediate feature map on this layer.
[0082] Resize(·) is an upsampling or downsampling alignment function, which is used to align the number of features in the x-1 layer to the x layer.
[0083] ω1, ω2 and ω3 are the weights of features of different scales.
[0084] During the training process of the neural network, the feature map is reduced due to convolution and downsampling, and the feature information is lost during the transmission process. DenseNet is proposed to make more efficient use of feature information. It connects each layer in a feedforward manner, so the first layer receives all the feature maps x0, x1, ..., x l-1 as input.
[0085] x l =H l [x0,x1,…,x l-1 ] (7)
[0086] Where [x0,x1,…,x l-1 ] is the concatenation of feature maps of each layer x0,x1,…,x l-1 ,H l is a function for processing concatenated feature maps. This allows DenseNet to alleviate gradient vanishing, enhance feature propagation, promote feature reuse, and greatly reduce the number of parameters.
[0087] In the multi-target detection of goods in unmanned warehouses, dense target detection has always been a difficult problem. Since the original input image data needs to be resized according to the shortest side in the YOLOv3 algorithm, this method not only reduces the image resolution but also reduces the distance of objects that were originally close and occluded in the image. Therefore, the close distance between targets will cause the target features to be occluded, which poses a great challenge to the detection algorithm. At the same time, due to the diversity of target objects, background complexity, robustness of light intensity, mutual occlusion between targets, etc., the robustness of such objects is very high, resulting in the low recall rate of YOLOv3 target detection in dense situations, and prone to problems such as missed detection of targets. Based on the YOLOv3 loss calculation formula, the FP-YOLO system introduces two mutually exclusive loss calculations to optimize the missed detection and false detection problems caused by target occlusion in dense scenes.
[0088] The loss function is an important basis for penalizing misdetected samples in deep neural networks. The quality of its design will directly affect the convergence effect of model training. Designing a more suitable loss function for the model in the target scenario to obtain better prediction results has gradually become an important direction of model optimization.
[0089] In the training of the YOLOv3 algorithm model, the loss calculation is divided into position loss, recognition loss, and confidence loss. The loss function, also known as the cost function, maps the value of a random event or a random variable related to it to a non-negative real number to represent the "risk" or "loss" of the random event. Neural networks generally use the method of minimizing the loss function to train the network so that it has good reasoning ability. The following three parts of the loss function are used to optimize the network.
[0090] Position loss, also known as localization loss, is the error between the predicted bounding box and the true bounding box. When calculating position loss in YOLOv1, in order to make the deviation between the predicted bounding box and the true bounding box of the sample larger, YOLOv1 adopted the method of taking square root of both width and height, but this method did not have a significant penalty for the error between small-sized bounding boxes. Therefore, YOLOv3 improved the loss function and used truth w and truth h To increase the penalty for large errors in small-sized bounding boxes and increase the accuracy of generating predicted bounding boxes. In YOLOv3, the Sigmoid function is first used as the activation function to compress the translation scale of the predicted bounding box to the interval [0, 1]. This method can effectively ensure that the center of the target is in the grid cell where the prediction is performed, and reduce the excessive offset between the predicted bounding box and the actual target bounding box. Then, the position loss is calculated by calculating the mean square error between the real bounding box of the sample and the predicted bounding box. The position loss calculation formula is:
[0091]
[0092] In the formula, when there is no target in the grid, λ obj =0, 1 when there is a target.
[0093] truth w , truth h is the height and width of the real target box, predict w , predict h The height and width of the predicted bounding box output by the model.
[0094] i is the feature number, k is the number of features, x, y, w, h are the horizontal and vertical coordinates, width, and height of the center point of the predicted bounding box, and b is the current position of the predicted bounding box.
[0095] In YOLOv3, an independent Logic classifier is used instead of Softmax to complete multi-label classification under the same category. Therefore, the binary cross entropy function is used to complete the recognition loss calculation. K is the output tensor under three different feature map sizes; c is the category number, and C is the total number of categories; when the category identified by the prior box is consistent with the actual target category, c==truth class ? 1:0=0.
[0096]
[0097] predict class is the predicted label classification.
[0098] In the YOLO detection process, when there is a target in the grid of the feature map, the grid generates three prior boxes to predict the target. When one of the prior boxes has the highest IOU overlap with the real bounding box of the sample, the bounding box generated by this prior box can more accurately predict the target, so it is necessary to calculate the center coordinate error, width and height error, confidence error and classification error to update the network weight. The remaining two prior boxes only need to calculate the confidence loss. The confidence loss is:
[0099] loss3=(truth conf -predict conf ) 2
[0100] Among them, truth conf is the confidence of the true bounding box, predict conf Confidence of the predicted bounding box.
[0101] In the occlusion problem of dense target detection, due to the similar features between the target objects, the predicted bounding box positioning of the target object is prone to inaccurate. In multi-target detection, the object occlusion problem is mainly divided into three categories: occlusion between targets of the same type, occlusion between targets of different types, and occlusion between non-targets and targets;
[0102] In one embodiment, a new loss function covering loss is introduced in the model training of the FP-YOLO system. The loss function is calculated as follows:
[0103] Calculate the traditional loss between any predicted bounding box p and any real bounding box t during the algorithm iteration process, such as IOU. Sort by multiple IOU values to obtain the worst real bounding box t with the largest IOU value. mass .
[0104] Find the worst truth bounding box t mass The correlation between the regression box and any predicted bounding box p is used to obtain the first coverage loss, the second coverage loss and the third coverage loss.
[0105] Among them, the first coverage loss is:
[0106]
[0107] P + is the set of predicted bounding boxes p, K is the total number of predicted bounding boxes,
[0108] B p is the set of regression boxes for any predicted bounding box p, where B p1 With B p2 For B p Two complementary subsets of B.p1 The distance between the position of all regression boxes in and the position of the worst ground-truth bounding box is less than ρ, B p2 On the contrary. mass is the worst ground-truth bounding box. ρ is a reference quantity.
[0109] The second coverage loss is:
[0110]
[0111] The third coverage loss is:
[0112]
[0113] In the formula, ε is a small number to prevent the denominator from being 0, p i is a predicted bounding box and i is the number of the box.
[0114] ∏(·) is a mapping function, and the mapping method is the identity transformation in the whole domain.
[0115] In the above formula, pr is the predicted coordinate system, and tr is the real coordinate system. The network can be optimized by setting the weights of the confidence loss functions of different detection layers.
[0116] In the embodiment, the image data set used for model training is 3200 in total, and the test data image data set is 800. The Kerask framework is used in the Python language environment to implement the algorithm part of the FP-YOLO system, and the hardware is a GeForceRTX2080Ti model GPU with 11GB video memory for model training.
[0117] In order to improve the problem of missed detection and inaccurate positioning of large-scale targets in YOLOv3, the FP-YOLO algorithm proposes to increase the number of downsampling times in the Darknet feature extraction network to output a larger feature map for the YOLO layer to solve the problem of missed detection of large-scale targets in the image; at the same time, a priori boxes with smaller size differences are used in model training to solve the problem that the target bounding box has high similarity but incomplete feature coverage. The feature coverage and similarity of objects in the predicted bounding box before and after the algorithm improvement. Table 1 is a table of the detection effects of 7 types of targets.
[0118] Table 1
[0119]
[0120] Before the improvement, the YOLOv3 algorithm detected 4, although the target category similarity reached 0.96. However, the output target prediction bounding box contained most of the features of other targets, which also caused the problem of missed detection of 7. However, the improved FP-YOLO system not only improved the feature coverage of the output prediction bounding box of the object, but also detected 7. In Table 1, although the recognition similarity of 3 objects is the same before and after the improvement, the prediction bounding box generated by the YOLOv3 algorithm model for 3 does not fully cover the effective features of the target. Although YOLOv3 has achieved a good recognition similarity, there are problems with the size and position of the target box. Compared with the YOLOv3 algorithm, the FP-YOLO system has improved the overall feature coverage of the target object in the prediction box, but there is still a problem of more background in the bounding box.
[0121] In order to solve the problem of missed detection of large-sized targets, FP-YOLO proposed to output a feature map of 7*7, which is larger than YOLOv3 when fusion outputting YOLO features; at the same time, although the feature map with an output size of 14*14 is smaller than the feature map with a size of 13*13, it can help the model detect medium and large-sized objects; FP-YOLO also adjusted the input size of the original image to obtain richer detail information of the objects in the image, helping the model to achieve better separation of the target and background during the detection process, and improving the robustness of the model.
[0122] FP-YOLO draws on the multi-scale feature fusion method in YOLOv3, adds convolution modules and upsampling to combine and fuse different early network features to obtain a feature map of 58*58 with a smaller field of view than YOOv3. This feature map is used to extract more fine-grained information about small and medium-sized targets in the early stage of the feature extraction network, and the accuracy of small-sized target detection in images is also improved.
[0123] In summary, the FP-YOLO system adjusts the image input size of the feature extraction network and increases the number of downsampling times of the backbone network, so that the YOLO layer can output a feature map with a larger field of view suitable for large target sizes in the image. At the same time, it uses FPN-like fine-grained fusion in the early network to output a 58*58 feature map with a narrower field of view than the 56*56 feature map. After experimental comparison and improvement, the detection accuracy of the FP-YOLO system for large and small targets in the image is improved compared with YOLOv3.
[0124] Occlusion target detection is divided into mutual occlusion between target objects and mutual occlusion between non-target objects. For the occlusion problem between target objects, the FP-YOLO system introduces two repulsive terms of the repulsive force function in the loss calculation during model training; the occlusion between targets and non-targets is solved by expanding the number of samples during model training. The improved FP-YOLO system adds coverage loss, and the detection accuracy of occlusion between targets and non-targets is improved compared with the YOLOv3 algorithm.
[0125] In the target detection process, the intersection-over-union (IOU) value of the predicted bounding box of the target generated in the image detection and the actual target box is first calculated, and then the IOU threshold is set. When the IOU is greater than the set threshold, it is determined that the predicted bounding box can detect the target and is retained. Secondly, the predicted bounding box is classified. Therefore, in target detection, the size of mAP when IOU=0.5 is often used to measure the overall target detection effect of the model. When IOU=0.5, the higher the mAP (average precision), the better the model detection performance.
[0126] After the model training is completed, the model is evaluated using the test training set. In the test set, there are 400 targets in material 1, 470 targets in material 2, and 45 targets in material 3. The model before and after improvement is evaluated in terms of mAP, loss convergence during model training, and model running speed.
[0127] In image classification tasks, the target category that the predicted bounding box is to be predicted is often set as the positive category, and other categories are set as negative categories. If the predicted target is a positive sample and the actual target is a positive sample, it is recorded as a true positive (TP), and the predicted target is a positive sample and the actual target is a negative sample, it is recorded as a true negative (FP). The FP and TP of each category in the model before and after the improvement are shown in Table 2.
[0128] Table 2
[0129]
[0130] As shown in Table 2, although the YOLOv3 model detects more total number of objects 1 and 2 than the FP-YOLO system model, the number of positive samples detected by YOLOv3 in the detection of object 2 is significantly less than that of FP-YOLO, and the number of negative samples detected in this category is more than that of FP-YOLO; the detection effect of the FP-YOLO model in the detection of object 3 is significantly improved. In multi-category object detection, the average AP of all categories is used to measure the overall classification effect of the model, and the size of mAP when IOU=0.5 is used to measure the overall object detection effect of the model. When IOU=0.5, the higher the mAP, the better the model detection performance. The algorithms before and after the improvement are trained using the same training data set validation test set, where the number of pictures sent to the network for learning each time is 16 (batch-size=16). The FP-YOLO system model has a faster and more gentle loss decline rate than YOLOv3 during training, and can find the optimal solution and output the optimal parameters of the model faster than YOLOv3.
[0131] In actual engineering applications, target detection technology has requirements for the detection accuracy and speed of the model. If we only focus on improving the detection accuracy of the algorithm model, a large amount of computational redundancy will be generated and more memory will be occupied. When the YOLOv3 algorithm model is trained, a high-performance GPU is required to complete the learning of multi-threaded samples. When there are many model learning samples, the model training needs to occupy a lot of memory for calculation; but although the improved FP-YOLO system deepens the number of feature extraction network layers and increases the amount of calculation in network model training, the algorithm network architecture adopts parallel computing to use four CPUs to parallelly calculate the YOLO layer output and use the cache queue form to reduce the calculation waiting time during the detection process, so as to achieve the overall detection speed of the network. In the training test set, the detection speed of the YOLOv3 model is 16.72Fps, while that of the FP-YOLO model is 19.43Fps. The detection speed of the FP-YOLO model is higher than that of YOLOv3.
[0132] The second aspect of the present invention relates to a multi-target detection and statistical system for unmanned warehouse goods using the method of the first aspect of the present invention; the system comprises an acquisition module, an expansion module, a preprocessing module, a construction module, an extraction module, and an output module; the acquisition module is used to acquire storage in-and-out multi-target material data from the data space of a storage management enterprise; the expansion module is used to perform sample expansion operations on the storage in-and-out multi-target material data to obtain the storage in-and-out multi-target material data; the preprocessing module is used to perform data preprocessing operations on the expanded storage in-and-out multi-target material data. The multi-target material data of warehousing in and out is obtained; the construction module is used to construct an FP-YOLO algorithm based on the YOLOv3 algorithm; the extraction module is used to input the multi-target material data of warehousing in and out obtained after preprocessing into the constructed FP-YOLO algorithm for feature extraction to obtain the result after feature extraction; the output module is used to calculate and screen the first confidence and the second confidence for the result after feature extraction, perform feature fusion and DenseNet optimization on the calculated and screened data, and output the optimized data detection to obtain the target detection statistical result.
[0133] A third aspect of the present invention relates to a terminal, comprising a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described in the first aspect of the present invention.
[0134] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect of the present invention.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention is described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention still include contents that can modify or replace the specific embodiments of the present invention. Any modification or replacement that does not depart from the spirit and scope of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A multi-target detection statistical method for unmanned warehouse cargo, characterized in that: The method comprises the following steps: S1. Obtaining multi-objective material data of storage in and out from the data space of the storage management enterprise; S2. Perform sample expansion operation on the warehousing and warehouse multi-objective material data to obtain the warehousing and warehouse multi-objective material data; S3, performing data preprocessing operations on the expanded warehousing in-and-out multi-objective material data to obtain the warehousing in-and-out multi-objective material data; S4. Build the FP-YOLO algorithm based on the YOLOv3 algorithm; S5, the multi-target material data of storage in and out obtained after preprocessing is put into the constructed FP-YOLO algorithm for feature extraction to obtain the result after feature extraction; S6. Calculate and screen the first confidence level and the second confidence level for the result after feature extraction, perform feature fusion and DenseNet optimization on the calculated and screened data, and output the optimized data for target detection statistics.
2. The multi-target detection statistical method for unmanned warehouse cargo according to claim 1 is characterized by: In the step 1, The warehouse storage and warehousing multi-objective material data are recorded in Pycharm, a software for data processing operations based on the Python programming language.
3. The multi-target detection statistical method for unmanned warehouse cargo according to claim 2 is characterized by: In the step 2, Expand the samples under different lighting conditions by processing such as sharpening, oversaturation and image enhancement to expand the training samples under different lighting conditions; In scenes with shaky shots, blur the image.
4. The multi-target detection statistical method for unmanned warehouse cargo according to claim 3 is characterized by: In step 3, Use the K-means algorithm to perform centroid clustering analysis on the KITTI dataset and the expanded dataset to determine the clustering range of the target box size in the algorithm's dataset; The IOU between the prior box and the true bounding box of the sample is used as the distance metric. Based on the FP-YOLO system, four feature maps of different scales are output in the YOLO layer, and each feature map generates three prior boxes of different sizes.
5. The multi-target detection statistical method for unmanned warehouse cargo according to claim 4 is characterized by: In step 4, The feature extraction part of the overall architecture of the FP-YOLO system is obtained by adjusting the image size, building Darknet-62, and YOLO layer feature extraction output; After waiting for all bounding boxes to be traversed, candidate bounding boxes are screened and calculated; Resize the feature maps to the same resolution and perform feature fusion on the resized feature maps; DenseNet is used to use feature information to obtain the concatenated feature output.
6. The multi-target detection statistical method for unmanned warehouse cargo according to claim 5 is characterized by: In the YOLOv3 algorithm model training, the loss calculation is divided into position loss, recognition loss, and confidence loss.
7. The multi-target detection statistical method for unmanned warehouse cargo according to claim 6 is characterized by: The position loss is: In the formula, when there is no target in the grid, λ obj =0, 1 when there is a target, truth w , truth h is the height and width of the real target box, predict w ,predict h is the height and width of the predicted bounding box output by the model, i is the feature number, I is the number of features, x, y, w, h are the horizontal and vertical coordinates, width, and height of the center point of the predicted bounding box, and b is the current position of the predicted bounding box.
8. The multi-target detection statistical method for unmanned warehouse cargo according to claim 7 is characterized by: The recognition loss is: c==truth class ? 1:0 is the type identified by the prior box and the actual target type truth class Is the function consistent? c is the category number, and C is the total number of categories; predict class is the predicted label classification.
9. The multi-target detection statistical method for unmanned warehouse cargo according to claim 8, characterized in that: The confidence loss is: loss3=(truth conf -predict conf ) 2 In the formula, truth conf is the confidence of the true bounding box, predict conf is the confidence of the predicted bounding box.
10. A multi-target detection and statistics system for unmanned warehouse cargo using the method described in any one of claims 1 to 9; characterized in that: The system acquisition module, expansion module, preprocessing module, construction module, extraction module, and output module; The acquisition module is used to acquire the multi-objective material data of storage in and out from the data space of the storage management enterprise; The expansion module is used to perform sample expansion operations on the warehousing and warehouse multi-target material data to obtain the warehousing and warehouse multi-target material data; The preprocessing module is used to perform data preprocessing operations on the expanded warehousing and outgoing multi-objective material data to obtain the warehousing and outgoing multi-objective material data; The building module is used to build an FP-YOLO algorithm based on the YOLOv3 algorithm; The extraction module is used to input the pre-processed storage in-and-out multi-target material data into the constructed FP-YOLO algorithm for feature extraction to obtain the result after feature extraction; The output module is used to calculate and screen the first confidence level and the second confidence level for the result after feature extraction, perform feature fusion and DenseNet optimization on the calculated and screened data, and output the optimized data for detection to obtain target detection statistical results.
11. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Material checking machine automatic identification and counting method based on computer vision
CN121582602A