Ton bag tallying method based on monocular vision and Repulsion loss enhancement
Through the ton bag collection method based on monocular vision and Repulsion loss enhancement, the door machine hook number and ton bag number are calculated in real time, solving the problem of low efficiency and inability to obtain the ton bag number in real time in the existing technology, and achieving high-precision and high-efficiency ton bag collection.
Patent Information
- Application Number
- CN202510029960.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-23
AI Technical Summary
The existing ton bag collection technology relies on the combination of manual recording and measurement and third-party data, which is inefficient and cannot obtain the number of ton bag operations in real time.
The ton bag collection method based on monocular vision and Repulsion loss enhancement is adopted. The video stream is obtained through the camera, the number of door hooks and ton bags is calculated in real time, and the door machine status is judged by the inter-frame difference method and feature matching method. The ton bag number is detected by the Repulsion loss function introduced in the yolov5 algorithm.
Real-time calculation of the number of tons of bags is achieved, the transparency and controllability of the production process is improved, the algorithm robustness and detection accuracy are optimized, and the accuracy is reached of 96%.
Smart Images

Figure CN120031804A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of ton bag tallying, and in particular to a ton bag tallying method based on monocular vision and repulsion loss enhancement. Background Art
[0002] Smart ports are an important trend in the future development of the port sector. By applying technologies such as big data, cloud computing, and artificial intelligence to collect, analyze, and transmit data, intelligent management of key links such as port transportation, loading and unloading, storage, and packaging is carried out, which improves logistics efficiency and thus improves port operation efficiency, which is of great significance to improving port benefits.
[0003] At present, there is no corresponding technical research on intelligent tallying without ton bags in the port. Ton bag tallying adopts a working method that combines manual record measurement and third-party data. The manual measurement method is to weigh the weight of ten bags of bulk cargo, obtain the average, and then count the number of ton bags to calculate the total weight. At the same time, the daily inventory and operation plan are arranged and tracked by the tally clerk. The ton bag loading and unloading quantity of each gantry crane is the responsibility of each gantry crane. The overall efficiency is low, and the ton bag operation quantity cannot be obtained in real time;
[0004] In view of this, in-depth research was conducted on the above problems, and a ton bag tallying method based on monocular vision and repulsion loss enhancement was proposed to establish a complete set of automated ton bag tallying algorithm processes to realize real-time ton bag operation quantity calculation. Summary of the invention
[0005] The purpose of the present invention is to provide a ton bag tallying method based on monocular vision and repulsion loss enhancement, so as to solve the problem proposed in the above background technology that the existing ton bag tallying adopts a working mode combining manual recording measurement and third-party data, the overall efficiency is low, and the number of ton bag operations cannot be obtained in real time.
[0006] To achieve the above object, the present invention provides the following technical solution: a ton bag tallying method based on monocular vision and repulsion loss enhancement, comprising the following steps:
[0007] Step A: Start the door machine;
[0008] Step B: Use the camera to obtain the video stream;
[0009] Step C: Determine the state of the door machine to determine whether the door machine is stationary or moving;
[0010] Step D: Calculate the number of door crane hooks in real time;
[0011] Step E: Detect the number of ton bags when judging the movement of the gantry crane;
[0012] Step F: Combine the hook count calculation results and the ton bag quantity detection results for analysis and send them to the production integration system;
[0013] Step G: End;
[0014] In step C, two algorithms, frame difference method and feature matching method, are used to judge the state of the door crane;
[0015] In the step E, the Repulsion loss function is introduced into the yolov5 algorithm to detect the number of ton bags.
[0016] As a preferred technical solution of the present invention, the inter-frame difference method is to obtain the images of the current frame and the previous frame, obtain the absolute value of the brightness difference between the two frames by difference, and judge the degree of difference between the two frames of images according to the sum of the absolute values. The larger the sum of the absolute values, the greater the image difference, and the smaller the sum of the absolute values, the smaller the image difference.
[0017] As the preferred technical solution of the present invention, the specific process of the inter-frame difference method is as follows: first, the previous and next two frames of images are obtained, the image is downsampled to (80, 45, 3), and converted into a grayscale image, and a Gaussian blur of size (21, 21) is set to further reduce the environmental noise. The image is kept at a level where the texture can be clearly seen. Subsequently, the two images are subtracted and binarized. The pixel difference greater than 25 is black, and the others are white. Finally, a morphological corrosion operation is performed on the image to obtain the final changed area, and the degree of environmental change is judged by the obtained image.
[0018] As a preferred technical solution of the present invention, the feature matching method extracts feature points on two images as conjugate entities by selecting one or more feature descriptors, and uses a certain matching method to find the common points of the two images.
[0019] As a preferred technical solution of the present invention, the matching method is one of a brute force matching method, a cross matching method, and a KNN matching method.
[0020] As the preferred technical solution of the present invention, the specific process of the feature matching method is: obtaining the images of the current frame and the previous frame, converting them into grayscale images, obtaining a scale pyramid, performing FAST corner point detection feature extraction on the scale pyramid image to obtain feature points, performing BRIEF and Streer BRIEF feature description on the feature points, obtaining descriptors, and finally using a certain matching method to perform feature matching, and judging the similarity of the two images by comparing the number of successful pairings.
[0021] As a preferred technical solution of the present invention, in step E, a time close to the middle of the movement of the door crane is selected to capture the image, and the image is captured three times for recognition.
[0022] As the preferred technical solution of the present invention, the yolov5 algorithm consists of four parts: an input part, a backbone network, a neck, and a detection head. The input part is the beginning of the network. After acquiring the image, it is resized to a size of (640, 640, 3) and sent to the backbone network. The backbone network is the main feature extraction part of the network. By building a multi-layer Conv module, a C3 module, and an SPPF module, the image is feature extracted to extract image features at different levels, and three feature maps of different sizes (80, 80, 256), (40, 40, 512), and (20, 20, 1024) are output. The neck splices the feature maps of the backbone network input in series to construct a model pyramid, combining the deep features and shallow features of the model to improve the detection accuracy of the model. The detection head is the output part of the model, which is composed of three 1*1 convolutions, and converts the feature map output by the neck into object position information, category information, and confidence of the corresponding grid size.
[0023] As a preferred technical solution of the present invention, the repulsion loss loss function calculation formula is:
[0024] L=L Attr +a*L RepGT +β*L RepBox
[0025] L Attr In order to make the predicted box closer to the real box, L Rep The purpose is to make the predicted box away from the surrounding real boxes. Parameters α and β are used to balance the weights of the two. Let P(lP,tP,wP,hP) be the candidate box, G(lG,tG,wG,hG) be the real box, P+ be the set of positive candidate boxes, and the positive candidate box is the one whose IoU with at least one of the real boxes is greater than a certain threshold, here 0.5g={G} is the set of real boxes;
[0026] L Attr The calculation formula is:
[0027]
[0028] P∈P+ (all positive samples) is the set of detection boxes P divided according to the set IoU threshold; G p Attr Match a real target box with the maximum IoU value for each detection box P; B P is the predicted box obtained after regression offset from the detection box P;
[0029] L RepGT The calculation formula is:
[0030]
[0031] In order to increase the intersection of the predicted box and the real box instead of the union, the IoG calculation method is adopted here, and smoothLn is used as the smoothing function, which not only retains the robustness of L1 but also absorbs the fast convergence of L2;
[0032] L RepBox The calculation formula is:
[0033]
[0034] The denominator ‖ is the identity function, that is, y=x.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows: the entire algorithm process of the ton bag tallying method based on monocular vision and repulsion loss enhancement uses a single camera, which reduces the cost of new hardware for the gantry crane, achieves real-time display of production data, ensures the transparency of the production process, improves the controllability of the production process, optimizes the performance parameters of the two algorithms, the inter-frame difference method and the feature matching method, improves the robustness of the algorithm, can adapt to the needs under various complex working conditions, and has an accuracy of 96%. At the same time, the loss function of the yolov5 algorithm is optimized to improve the detection accuracy of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a flow chart of ton bag tallying by a gantry crane of the present invention;
[0037] Figure 2 This is a flow chart of the inter-frame difference method of the present invention;
[0038] Figure 3 This is a flow chart of the feature matching method of the present invention;
[0039] Figure 4 This is a data diagram of the door machine state judgment experiment of the present invention;
[0040] Figure 5 This is a diagram showing the results of a door machine state judgment experiment of the present invention;
[0041] Figure 6 This is a data diagram of the ton bag quantity detection experiment of the present invention;
[0042] Figure 7 This is a diagram showing the experimental results of the ton bag quantity detection of the present invention. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] See also Figure 1 - Figure 7 The technical solution of the present invention is a method for tallying ton bags based on monocular vision and repulsion loss enhancement, comprising the following steps:
[0045] Step A: Start the door machine;
[0046] Step B: Use the camera to obtain the video stream;
[0047] Step C: Determine the state of the door machine to determine whether the door machine is stationary or moving;
[0048] Step D: Calculate the number of door crane hooks in real time;
[0049] Step E: Detect the number of ton bags when judging the movement of the gantry crane;
[0050] Step F: Combine the hook count calculation results and the ton bag quantity detection results for analysis and send them to the production integration system;
[0051] Step G: End;
[0052] In step C, two algorithms, frame difference method and feature matching method, are used to judge the state of the door crane;
[0053] In step E, the Repulsion loss function is introduced into the yolov5 algorithm to detect the number of ton bags;
[0054] The inter-frame difference method is to obtain the image of the current frame and the previous frame, and obtain the absolute value of the brightness difference between the two frames by doing the difference. The difference between the two frames is judged according to the sum of the absolute values. The larger the sum of the absolute values, the greater the image difference, and the smaller the sum of the absolute values, the smaller the image difference.
[0055] The specific process of the inter-frame difference method is as follows: first, obtain the previous and next two frames of images, downsample the images to (80, 45, 3), and convert them into grayscale images. Set a Gaussian blur of size (21, 21) to further reduce environmental noise. The image is kept at a level where the texture can be clearly seen. Then, subtract the two images and use binarization processing. The pixel difference greater than 25 is black, and the others are white. Finally, perform a morphological corrosion operation on the image to obtain the final changed area. The degree of environmental change can be judged by the obtained image.
[0056] The feature matching method selects one or more feature descriptors, extracts the feature points on the two images as conjugate entities, and uses a certain matching method to find the common points of the two images;
[0057] The matching method is one of brute force matching, cross matching, and KNN matching;
[0058] The specific process of the feature matching method is as follows: obtain the images of the current frame and the previous frame, convert them into grayscale images, obtain the scale pyramid, perform FAST corner detection feature extraction on the scale pyramid image to obtain feature points, perform BRIEF and Streer BRIEF feature description on the feature points, obtain descriptors, and finally use a certain matching method to perform feature matching, and judge the similarity of the two images by comparing the number of successful pairings;
[0059] In step E, a time close to the middle of the door crane movement is selected to capture the image, and the image is captured three times for recognition;
[0060] The yolov5 algorithm consists of four parts: input part, backbone network, neck, and detection head. The input part is the beginning of the network. After obtaining the image, it is resized to (640, 640, 3) size and sent to the backbone network. The backbone network is the main feature extraction part of the network. By building a multi-layer Conv module, C3 module and SPPF module, the image is feature extracted to extract image features at different levels, and output feature maps of three different sizes (80, 80, 256), (40, 40, 512), and (20, 20, 1024). The neck splices the feature maps of the backbone network input in series to build a model pyramid, combining the deep and shallow features of the model to improve the detection accuracy of the model. The detection head is the output part of the model, which is composed of three 1*1 convolutions. The feature map output by the neck is converted into object position information, category information and confidence of the corresponding grid size;
[0061] The calculation formula of Repulsion loss function is:
[0062] L=L Attr +a*L RepGT +β*L RepBox L Attr In order to make the predicted box closer to the real box, L Rep The purpose is to make the predicted box away from the surrounding real boxes. Parameters α and β are used to balance the weights of the two. Let P(lP,tP,wP,hP) be the candidate box, G(lG,tG,wG,hG) be the real box, P+ be the set of positive candidate boxes, and the positive candidate box is the one whose IoU with at least one of the real boxes is greater than a certain threshold, here 0.5g={G} is the set of real boxes;
[0063] L Attr The calculation formula is:
[0064]
[0065] P∈P+ (all positive samples) is the set of detection boxes P divided according to the set IoU threshold; G pAttr Match a real target box with the maximum IoU value for each detection box P; B P is the predicted box obtained after regression offset from the detection box P;
[0066] L RepGT The calculation formula is:
[0067]
[0068] In order to increase the intersection of the predicted box and the real box instead of the union, the IoG calculation method is adopted here, and smoothLn is used as the smoothing function, which not only retains the robustness of L1 but also absorbs the fast convergence of L2;
[0069] L RepBox The calculation formula is:
[0070]
[0071] The denominator ‖ is the identity function, that is, y=x.
[0072] experiment
[0073] A port gantry crane was used as the experimental site, and video stream data in working state within three days was collected as a data set, on which the accuracy of the algorithm was verified. All experiments were implemented in Python, using the PyTorch framework, the GPU was NVIDIA GeForce 3090, the initial learning rate of the model training was set to 0.001, and it was reduced to 0.1 every 500 steps, and the BatchSize was set to 64;
[0074] Door crane state judgment: In order to verify the effect of inter-frame difference method and feature point matching method on door crane state judgment, 500 sets of continuous frame images were collected and a suitable threshold was set. When the result is less than the threshold, it indicates that the two frames are similar and the door crane is in a stationary state. When the result is greater than the threshold, it indicates that there is a gap between the two frames and the door crane is in a moving state. The accuracy and time consumption of the two algorithms are statistically analyzed. The experimental results are shown in the figure. Figure 4-5 As shown in the figure, from the experimental results, we can see that the frame difference method is better than the feature matching algorithm as a whole. The feature matching method has a high number of false positives, especially when there are scenes with large fluctuations such as sea areas in the scene, there are many unstable factors;
[0075] Ton bag quantity detection: A data set was created based on the actual working scene of the gantry crane to verify the accuracy of the yolov5s-Repulsion and other target detection algorithms proposed in this paper, and the mAP and detection time of the statistical algorithm were calculated. The experimental results are shown in the figure below. Figure 6-7As shown, the comparison of experimental results shows that, thanks to the Repulsion loss function, the mAP of the YOLOv5s-Repulsion algorithm proposed in this paper is higher than that of other comparison algorithms, and the mAP is improved by 2.12% compared with the original yolov5s algorithm, especially in the detection of ton bags under dense occlusion, which increases the Euclidean distance between the target box and other boxes and reduces the occurrence of missed detection. Because the algorithm is optimized in terms of the loss function and the network structure is consistent with yolov5s, the time consumption is the same as that of the yolov5 algorithm, which can meet actual production needs.
[0076] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.
[0077] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for tallying ton bags based on monocular vision and repulsion loss enhancement, characterized in that: The following steps are involved: Step A: Start the door machine; Step B: Use the camera to obtain the video stream; Step C: Determine the state of the door machine to determine whether the door machine is stationary or moving; Step D: Calculate the number of door crane hooks in real time; Step E: Detect the number of ton bags when judging the movement of the gantry crane; Step F: Combine the hook count calculation results and the ton bag quantity detection results for analysis and send them to the production integration system; Step G: End; In step C, two algorithms, frame difference method and feature matching method, are used to judge the state of the door crane; In the step E, the Repulsion loss function is introduced into the yolov5 algorithm to detect the number of ton bags.
2. According to claim 1, a method for tallying ton bags based on monocular vision and repulsion loss enhancement is characterized in that: The inter-frame difference method is to obtain the images of the current frame and the previous frame, obtain the absolute value of the brightness difference between the two frames by difference, and judge the degree of difference between the two frames of images according to the sum of the absolute values. The larger the sum of the absolute values, the greater the image difference, and the smaller the sum of the absolute values, the smaller the image difference.
3. The method for tallying ton bags based on monocular vision and repulsion loss enhancement according to claim 2 is characterized in that: The specific process of the inter-frame difference method is as follows: first, obtain the previous and next two frames of images, downsample the images to (80, 45, 3), and convert them into grayscale images. Set a Gaussian blur of size (21, 21) to further reduce environmental noise. The image is kept at a level where the texture can be clearly seen. Then, subtract the two images and use binarization processing. The pixel difference greater than 25 is black, and the others are white. Finally, perform a morphological corrosion operation on the image to obtain the final changed area, and judge the degree of environmental change through the obtained image.
4. According to claim 1, a method for tallying ton bags based on monocular vision and repulsion loss enhancement is characterized in that: The feature matching method selects one or more feature descriptors, extracts feature points on two images as conjugate entities, and uses a certain matching method to find common points between the two images.
5. The method for tallying ton bags based on monocular vision and repulsion loss enhancement according to claim 4 is characterized in that: The matching method is one of a brute force matching method, a cross matching method, and a KNN matching method.
6. The method for tallying ton bags based on monocular vision and repulsion loss enhancement according to claim 4 is characterized in that: The specific process of the feature matching method is as follows: obtain the images of the current frame and the previous frame, convert them into grayscale images, obtain the scale pyramid, perform FAST corner point detection feature extraction on the scale pyramid image to obtain feature points, perform BRIEF and Streer BRIEF feature description on the feature points, obtain descriptors, and finally use a certain matching method to perform feature matching, and judge the similarity of the two images by comparing the number of successful pairings.
7. The method for tallying ton bags based on monocular vision and repulsion loss enhancement according to claim 1 is characterized in that: In the step E, a time close to the middle of the movement of the door crane is selected to capture the image, and the image is captured three times for recognition.
8. The method for tallying ton bags based on monocular vision and repulsion loss enhancement according to claim 1 is characterized in that: The yolov5 algorithm consists of four parts: input part, backbone network, neck and detection head. The input part is the beginning of the network. After acquiring the image, it is resized to (640, 640, 3) size and sent to the backbone network. The backbone network is the main feature extraction part of the network. By building a multi-layer Conv module, C3 module and SPPF module, the image is feature extracted to extract image features at different levels, and three feature maps of different sizes (80, 80, 256), (40, 40, 512) and (20, 20, 1024) are output. The neck splices the feature maps of the backbone network input in series to build a model pyramid, combining the deep features and shallow features of the model to improve the detection accuracy of the model. The detection head is the output part of the model, which is composed of three 1*1 convolutions. The feature map output by the neck is converted into object position information, category information and confidence of the corresponding grid size.
9. The method for tallying ton bags based on monocular vision and repulsion loss enhancement according to claim 1 is characterized in that: The Repulsion loss function calculation formula is: L=L Attr +a*L RepGT +β*L RepBox L Attr In order to make the predicted box closer to the real box, L Rep The purpose is to make the predicted box away from the surrounding real boxes. Parameters α and β are used to balance the weights of the two. Let P(lP,tP,wP,hP) be the candidate box, G(lG,tG,wG,hG) be the real box, P+ be the set of positive candidate boxes, and the positive candidate box is the one whose IoU with at least one of the real boxes is greater than a certain threshold, here 0.5g={G} is the set of real boxes; L Attr The calculation formula is: P∈P+ (all positive samples) is the set of detection boxes P divided according to the set IoU threshold; G p Attr Match a real target box with the maximum IoU value for each detection box P; B P is the predicted box obtained after regression offset from the detection box P; L RepGT The calculation formula is: In order to increase the intersection of the predicted box and the real box instead of the union, the IoG calculation method is used here, and smoothLn is used as the smoothing function, which not only retains the robustness of L1 but also absorbs the fast convergence of L2; L RepBox The calculation formula is: The denominator ‖ is the identity function, that is, y=x.