Container positioning detection method, device and equipment and storage medium
By improving the YOLOv10 model, a dynamic positive and negative sample weight is constructed and the occlusion attention module is introduced, the accuracy of cargo container positioning detection in logistics scenarios is solved, and more efficient and accurate cargo container positioning detection is achieved.
Patent Information
- Application Number
- CN202411988180.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
When the existing target detection model deals with occlusion problems in logistics scenarios, it is difficult to accurately identify and locate partially obscured cargo boxes, resulting in reduced cargo management efficiency and accuracy.
By improving the YOLOv10 model, we construct positive sample weights determined by the linear combination of classification scores, prediction box and real box, as well as negative sample weights determined by the classification confidence and position mismatch of negative samples. Combining spatial prior values, classification scores, prediction box and real box interchange ratios, balanced classification task hyperparameters and regression task hyperparameters, and the consistency matching metric formula is determined. At the same time, multiple occlusion attention modules are introduced to enhance the feature capture capability of unoccluded areas.
This method significantly improves the accuracy of cargo container positioning detection in the processing logistics environment, and can handle occlusion situations more flexibly, improving the efficiency and accuracy of cargo management.
Smart Images

Figure CN119942054A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of logistics technology, and in particular to a cargo box positioning detection method, device, equipment and storage medium. Background Art
[0002] With the rapid development of the global logistics industry, accurate positioning of cargo boxes has become particularly important in warehouse management, cargo sorting and transportation scheduling. In practical applications, cargo box stacking, complex transportation environments and cargo box occlusions frequently occur. When dealing with occlusion problems in logistics scenarios, existing object detection models have difficulty accurately identifying and locating partially occluded cargo boxes. This deficiency affects the efficiency and accuracy of cargo management and increases the complexity of operations.
[0003] Although the application of existing deep learning-based object detection methods in cargo box positioning has made significant progress, there are still some shortcomings. Existing methods such as YOLOv10 use a novel label assignment method, which assigns multiple positive samples to each real object during training and uses a single prediction result during inference. This strategy combines the advantages of two matching strategies and achieves efficient inference without NMS. However, in complex actual scenes, mutual occlusion between cargo boxes is a common problem. The above methods perform poorly when dealing with occluded targets, and the accuracy of positioning detection is not high. Summary of the invention
[0004] The present invention provides a cargo box positioning detection method, device, equipment and storage medium, which are used to solve the technical problem of low accuracy of cargo box positioning detection.
[0005] In order to solve the above technical problems, in a first aspect, the present invention provides a method for detecting cargo box positioning, the method comprising:
[0006] Obtain the image of the cargo box to be inspected;
[0007] The cargo box image to be detected is input into the trained improved YOLOv10 model, and the predicted cargo box positioning result in the image is output; wherein, when performing consistency matching metric calculation, the YOLOv10 model constructs a positive sample weight determined by a linear combination of the classification score and the intersection-over-union ratio of the predicted box and the true box, a negative sample weight determined by the classification confidence and the position mismatch degree of the negative sample, and a consistency matching metric formula determined by the spatial prior value, the classification score, the intersection-over-union ratio of the predicted box and the true box, and a balanced classification task hyperparameter and a regression task hyperparameter.
[0008] Optionally, the introduced positive sample weight and negative sample weight include:
[0009] The calculation formula of the positive sample weight is: pos=γ 1 p+γ 2 IoU, where w pos is the positive sample weight, γ 1 and γ 2 They are parameter adjustment coefficients, which are used to adjust the impact of classification score and IoU on the weight of positive samples respectively;
[0010] The calculation formula of the negative sample weight is: neg =δ 1 ·1-p)+δ 2 (1-IoU), where w neg is the negative sample weight, δ 1 and δ 2 They are parameter adjustment coefficients, which are used to adjust the impact of classification error and mismatch on the weight of negative samples;
[0011] In the positive sample weight and negative sample weight calculation formula, p is the classification score, and IoU is the intersection over union ratio of the predicted box and the true box.
[0012] Optionally, the consistency matching metric formula includes:
[0013]
[0014] in, is the consistency matching measurement result, w pos is the positive sample weight, w neg is the negative sample weight, s is the spatial prior value, p is the classification score, For the prediction box The intersection-over-union ratio with the true box b, α is the hyperparameter for the classification task, and β is the hyperparameter for the regression task.
[0015] Optionally, the YOLOv10 model is constructed with multiple occlusion attention modules after the path aggregation network; the multiple occlusion attention modules process the multi-scale feature map obtained by the path aggregation network to obtain a feature map in which the features of the unoccluded area in the cargo box image are enhanced.
[0016] Optionally, the plurality of occlusion attention modules process the multi-scale feature map obtained by the path aggregation network, including:
[0017] Each of the occlusion degree attention modules performs a separate convolution operation on each channel of the feature map at one scale, and uses a 1×1 point-by-point convolution to merge the convolution results of each channel.
[0018] Optionally, the occlusion degree attention module includes a depth convolution unit, a first Gaussian error linear unit, a first batch normalization unit, a point-by-point convolution unit, a second Gaussian error linear unit, and a second batch normalization unit;
[0019] The deep convolution unit performs a convolution operation on the input feature map, and the convolution result is processed by the first Gaussian error linear unit and the first batch normalization unit in sequence to obtain a first feature map; the input feature map and the first feature map are added through a residual connection to obtain a second feature map; the second feature map is convolved by using the point-by-point convolution unit, and the convolution result is processed by the second Gaussian error linear unit and the second batch normalization unit in sequence to form an output.
[0020] Optionally, the predicted cargo box positioning result in the output image includes position coordinate information of each cargo box area in the cargo box image to be detected.
[0021] In a second aspect, the present invention provides a cargo box positioning detection device, including a data acquisition module and a detection module;
[0022] The data acquisition module is used to acquire the image of the cargo box to be detected;
[0023] The detection module is used to input the cargo box image to be detected into the trained improved YOLOv10 model, and output the predicted cargo box positioning result in the image; wherein, when performing consistency matching metric calculation, the YOLOv10 model constructs a positive sample weight determined by a linear combination of the classification score and the intersection-over-union ratio of the predicted box and the true box, a negative sample weight determined by the classification confidence and position mismatch degree of the negative sample, and a consistency matching metric formula determined by the spatial prior value, the classification score, the intersection-over-union ratio of the predicted box and the true box, and the hyperparameters of the balanced classification task and the hyperparameters of the regression task.
[0024] In a third aspect, the present invention provides a cargo box positioning detection device, including a memory and a processor, wherein:
[0025] The memory is used to store computer programs;
[0026] The processor is used to read the program in the memory and execute the steps of the cargo box positioning detection method provided in the first aspect above.
[0027] In a fourth aspect, the present invention provides a computer-readable storage medium having a readable computer program stored thereon, which, when executed by a processor, implements the steps of the cargo box positioning detection method provided in the first aspect above.
[0028] Compared with the prior art, the cargo box positioning detection method, device, equipment and storage medium provided by the present invention have the following beneficial effects:
[0029] Obtain a cargo box image to be detected; input the cargo box image to be detected into the trained improved YOLOv10 model, and output the predicted cargo box positioning result in the image; wherein, when the YOLOv10 model performs consistency matching metric calculation, it constructs a positive sample weight determined by a linear combination of the classification score and the intersection-over-union ratio of the predicted box and the true box, a negative sample weight determined by the classification confidence and position mismatch degree of the negative sample, and a consistency matching metric formula determined by the spatial prior value, the classification score, the intersection-over-union ratio of the predicted box and the true box, and the hyperparameters of the balanced classification task and the hyperparameters of the regression task. This method assigns dynamic weights to positive and negative samples when calculating the consistency matching metric, so that the model can flexibly handle occlusion situations, and can effectively improve the accuracy of cargo box positioning detection in a logistics environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only part of the embodiments of the present invention, rather than all of the embodiments. For ordinary technicians in this field, without paying creative work, other drawings obtained based on these drawings all fall within the scope of protection of this application.
[0031] Figure 1 It is a flow chart of a cargo box positioning detection method provided by an embodiment of the present invention;
[0032] Figure 2 It is a schematic diagram of the structure of a traditional YOLOv10 model provided in an embodiment of the present invention;
[0033] Figure 3 It is a schematic diagram of the structural relationship between the occlusion attention module and the path aggregation network provided by an embodiment of the present invention;
[0034] Figure 4 is a schematic diagram of the internal structure of the occlusion degree attention module provided by an embodiment of the present invention;
[0035] Figure 5 It is a structural schematic diagram of a cargo box positioning detection device provided by an embodiment of the present invention;
[0036] Figure 6 It is a structural schematic diagram of a cargo box positioning detection device provided by an embodiment of the present invention;
[0037] Figure 7It is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0039] In order to make the description of the disclosed content more detailed and complete, the following is an illustrative description of the implementation mode and specific examples of the present invention; however, this is not the only form of implementing or applying the specific embodiments of the present invention. The implementation mode covers the features of multiple specific embodiments and the method steps and their sequence for constructing and operating these specific embodiments. However, other specific embodiments can also be used to achieve the same or equal functions and step sequences. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.
[0041] In the description of the embodiments of the present invention, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two, and other quantifiers are similar. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention, and the embodiments of the present application and the features in the embodiments may be combined with each other without conflict.
[0042] Example 1
[0043] like Figure 1 As shown, it is a flow chart of a cargo box positioning detection method provided by an embodiment of the present invention. The cargo box positioning detection method includes the following steps.
[0044] Step S101, obtaining an image of a cargo box to be inspected;
[0045] To obtain the cargo box image, it is necessary to identify and locate the cargo box in the image.
[0046] Step S102, input the cargo box image to be detected into the trained improved YOLOv10 model, and output the predicted cargo box positioning result in the image; wherein, when performing consistency matching metric calculation, the YOLOv10 model constructs a positive sample weight determined by a linear combination of the classification score and the intersection-over-union ratio of the predicted box and the true box, a negative sample weight determined by the classification confidence and the position mismatch degree of the negative sample, and a consistency matching metric formula determined by the spatial prior value, the classification score, the intersection-over-union ratio of the predicted box and the true box, and a balanced classification task hyperparameter and a regression task hyperparameter.
[0047] The YOLOv10 model is a real-time target detection model that can be used to implement target detection and positioning, multi-category classification, feature extraction and representation, and can adapt to different scenarios and data. The YOLOv10 model mainly includes Backbone (backbone network), PAN (Path Aggregation Network), detection head Head and consistency matching metric module.
[0048] like Figure 2 As shown in the figure, Input is the input, and the image to be detected is input into the Backbone of the model. Backbone extracts the basic features of the image and passes these features to PAN; PAN combines feature maps of different scales to help the network better understand the multi-scale information in the image. Specifically, PAN further processes and aggregates the feature maps passed by the backbone network to generate multi-scale feature representations, namely multi-scale feature maps.
[0049] In the figure, Dual Label Assignments represents dual label assignment and prediction. Dual label assignment includes two branches: One-to-many Head and One-to-one Head. One-to-many Head includes two subtasks: regression and classification. The regression subtask is used to predict the position and size of the bounding box, and the classification subtask is used to predict the category of the target. One-to-one Head also includes two subtasks: regression and classification. Unlike the one-to-many Head, the regression and classification subtasks of the one-to-one Head adopt a one-to-one label assignment strategy. During model training, the one-to-many Head is used to assign multiple positive samples to each real object, while the one-to-one Head is used to obtain a single prediction result during inference.
[0050] In the figure, ConsistentMatch Metric represents the calculation of the consistency match metric. The consistency match metric module is used to calculate the consistency match metric to ensure the consistency of label assignment during training. In object detection, the consistency match metric module provides a basis for label assignment in the one-to-one assignment head and one-to-many assignment head branches by measuring the degree of match between the prediction results and the real instances.
[0051] During training, the YOLOv10 model outputs the predicted bounding box information and the consistency matching metric results, thereby performing joint optimization of two heads (one-to-many and one-to-one). When making predictions after training, the YOLOv10 model discards the one-to-many allocation head and only uses the one-to-one allocation head for prediction, outputting the predicted bounding box information.
[0052] The solution of the present invention optimizes the traditional YOLOv10 model architecture and processing, including reconstructing the consistency matching metric calculation formula when the YOLOv10 model performs consistency matching metric calculation, constructing a positive sample weight determined by a linear combination of a classification score and an intersection-over-union ratio of a predicted box and a true box, a negative sample weight determined by the classification confidence and position mismatch degree of the negative sample, and a new consistency matching metric formula determined by a spatial prior value, a classification score, an intersection-over-union ratio of a predicted box and a true box, and a hyperparameter of a balanced classification task and a hyperparameter of a regression task.
[0053] As an optional implementation manner, the positive sample weight and the negative sample weight introduced include:
[0054] The calculation formula of the positive sample weight is: pos =γ 1 p+γ 2 IoU, where w pos is the positive sample weight, γ 1 and γ 2 They are parameter adjustment coefficients, which are used to adjust the impact of classification score and IoU on the weight of positive samples respectively;
[0055] The calculation formula of the negative sample weight is: neg =δ 1 ·)1-p)+δ 2 (1-IoU), where w neg is the negative sample weight, δ 1 and δ 2 They are parameter adjustment coefficients, which are used to adjust the impact of classification error and mismatch on the weight of negative samples;
[0056] In the positive sample weight and negative sample weight calculation formula, p is the classification score, and IoU is the intersection over union ratio of the predicted box and the true box.
[0057] When calculating the consistency matching metric, different weights are introduced for positive samples and negative samples, which enhances the flexibility of label assignment. The positive sample weight w introduced is pos It reflects the importance of the model to the positive sample matching, and the negative sample weight w introduced neg Aims to reduce the interference of negative samples on the training process. The calculation of positive and negative sample weights is based on the classification confidence and matching metric of the sample. Among them, the positive sample weight is determined by the linear combination of the classification score p and IoU, which is used to measure the importance of positive samples in label assignment; the negative sample weight w neg It is determined based on the classification confidence and position mismatch degree of negative samples.
[0058] In the embodiment of the present invention, the positive sample weights and negative sample weights are introduced in the consistency matching metric calculation, which can more accurately match the supervision signal of the metric optimization model, adjust the contribution of positive samples and negative samples in label allocation, enable the model to flexibly handle occlusion situations, and effectively improve the accuracy of cargo box positioning detection in a logistics environment.
[0059] As an optional implementation, the consistency matching metric formula includes:
[0060]
[0061] in, is the consistency matching measurement result, w pos is the positive sample weight, w neg is the negative sample weight, s is the spatial prior value, p is the classification score, For the prediction box The intersection-over-union ratio with the true box b, α is the hyperparameter for the classification task, and β is the hyperparameter for the regression task.
[0062] In some embodiments, the spatial prior value in the above consistency matching metric formula is used to indicate whether the predicted anchor point is located in the real target instance. If s=1, it corresponds to a positive sample, otherwise s=0, it corresponds to a negative sample. Based on the above positive sample weights and negative sample weights, the consistency matching metric formula is obtained, which can accurately reflect the contribution of positive and negative samples in the matching process.
[0063] In the embodiment of the present invention, the weight w of the positive sample is adjusted by the consistency matching metric formula. pos Reflects the importance of the model to matching positive samples, making the negative sample weight w neg Reducing the interference of negative samples on the training process enhances the flexibility of label allocation and more accurately reflects the contribution of positive and negative samples in the matching process, which can effectively improve the accuracy of cargo box positioning detection in logistics environments.
[0064] As an optional implementation, the YOLOv10 model is constructed with multiple occlusion attention modules after the path aggregation network; the multiple occlusion attention modules process the multi-scale feature map obtained by the path aggregation network to obtain a feature map in which the features of the unoccluded area in the cargo box image are enhanced.
[0065] The solution of the present invention optimizes the traditional YOLOv10 model architecture and processing, and also includes adding an occlusion attention module after the path aggregation network processing, which is used to process the multi-scale feature map obtained by the path aggregation network processing and then input it to the label head for label allocation and prediction. Figure 3 As shown in the figure, taking the multi-scale feature map obtained by PAN processing as three feature maps of different sizes as an example, after PAN, three OAA (Occlusion-Aware Attention) modules are introduced to process the multi-scale feature map obtained by PAN processing. The occlusion attention module enhances the feature capture capability of the unoccluded area, thereby improving the feature saliency of the unaffected area, supplementing the information of those areas that are difficult to identify due to occlusion, thereby improving the overall detection performance of the model.
[0066] In an embodiment of the present invention, multiple occlusion attention modules are constructed after the path aggregation network. The multiple occlusion attention modules process the multi-scale feature map obtained by the path aggregation network, and enhance the feature capture capability of the unobstructed area to improve the feature significance of the unaffected area, so as to supplement the information of the area that is difficult to identify due to occlusion, thereby improving the overall detection performance of the model and improving the accuracy of cargo box positioning detection.
[0067] As an optional implementation, the plurality of occlusion attention modules process the multi-scale feature map obtained by the path aggregation network, including:
[0068] Each of the occlusion degree attention modules performs a separate convolution operation on each channel of the feature map at one scale, and uses a 1×1 point-by-point convolution to merge the convolution results of each channel.
[0069] After the path aggregation network processes the feature maps of different scales, the feature maps of each scale are introduced into an occlusion attention module respectively, so that each occlusion attention module processes a feature map of one scale. The occlusion attention module performs a separate convolution operation on each channel of the feature map of each scale, and then integrates the information of these channels through point-by-point convolution to enhance the features of the unoccluded areas. Considering the relationship between channels, a 1×1 point-by-point convolution can be used to merge the convolution results of each channel, thereby enhancing the ability of feature representation while keeping the number of parameters reduced.
[0070] The embodiment of the present invention performs a separate convolution operation on each channel of the feature map of each size and uses a 1×1 point-by-point convolution to merge the convolution results of each channel. While keeping the number of parameters reduced, the feature representation capability is enhanced, and the features of the unobstructed areas are enhanced, thereby further improving the accuracy of cargo box positioning detection.
[0071] As an optional implementation, the occlusion degree attention module includes a depth convolution unit, a first Gaussian error linear unit, a first batch normalization unit, a point-by-point convolution unit, a second Gaussian error linear unit, and a second batch normalization unit;
[0072] The deep convolution unit performs a convolution operation on the input feature map, and the convolution result is processed by the first Gaussian error linear unit and the first batch normalization unit in sequence to obtain a first feature map; the input feature map and the first feature map are added through a residual connection to obtain a second feature map; the second feature map is convolved by using the point-by-point convolution unit, and the convolution result is processed by the second Gaussian error linear unit and the second batch normalization unit in sequence to form an output.
[0073] In some embodiments, the internal structure diagram of the occlusion degree attention module is as follows: Figure 4 As shown in the figure, each input channel of the depth convolution unit Depthwise Conv can apply the convolution kernel independently instead of sharing the convolution kernel across channels. This convolution method can reduce the number of parameters and the amount of calculation while maintaining the expressiveness of the model; Gaussian error linear unit GELU, activation function, is used to increase nonlinear characteristics and help the model learn more complex features; batch normalization unit BN (BatchNormalization, batch normalization) is used to normalize each batch of data so that the internal covariate offset of the network will not accumulate as the network layer deepens; pointwise convolution unit Pointwise Conv performs 1x1 convolution operation. Specifically, it can be used to mix feature channels in the depth direction, which can change the number of channels of the feature map without changing its spatial dimension. Figure 4 middle Residual Connection represents the residual connection, which enables the network to learn the residual mapping and prevent the gradient from disappearing in the deep network.
[0074] The depthwise convolution unit Depthwise Conv performs a depth convolution operation on the input feature map, and convolves each channel independently to extract spatial features to obtain a convolution result. The convolution result is processed by the first Gaussian error linear unit GELU and the first batch normalization unit BN in turn to obtain a first feature map; the input feature map and the first feature map are added through a residual connection to obtain a second feature map; the second feature map is subjected to a 1x1 convolution operation using a point-by-point convolution unit to adjust the number of channels of the feature map, and the convolution result is processed by the second Gaussian error linear unit GELU and the second batch normalization unit BN in turn to form an output.
[0075] In an embodiment of the present invention, a convolution operation is performed on an input feature map through a deep convolution unit, and the convolution result is processed by a first Gaussian error linear unit and a first batch normalization unit in sequence to obtain a first feature map; the input feature map and the first feature map are added through a residual connection to obtain a second feature map; a convolution operation is performed on the second feature map using a point-by-point convolution unit, and the convolution result is processed by a second Gaussian error linear unit and the second batch normalization unit in sequence to form an output; the number of parameters and the amount of calculation can be reduced while maintaining the expressive power of the model, increasing nonlinear characteristics, helping the model to learn more complex features, so that the internal covariate offset of the network will not accumulate as the network layer deepens, and the network can learn residual mapping, which can prevent the gradient in the deep network from disappearing, thereby improving the accuracy and efficiency of cargo box positioning detection.
[0076] Based on the traditional YOLOv10 model architecture and processing, the model is optimized and improved by introducing positive sample weights and negative sample weights during label assignment processing, and introducing an occlusion attention module before label assignment processing. This enables the model to flexibly handle occlusion situations, which can effectively improve the accuracy of cargo box recognition and positioning in logistics environments.
[0077] When training the improved YOLOv10 model, first, cargo box data in various logistics scenarios is collected to build a data set, and the cargo box targets in the data set are annotated with detection boxes; then, the above annotated data set is used as a training set to train the model.
[0078] After the improved YOLOv10 model is trained, the image of the cargo box to be detected is input into it, and the predicted cargo box positioning result in the image is output.
[0079] As an optional implementation, the predicted cargo box positioning result in the output image includes position coordinate information of each cargo box area in the cargo box image to be detected.
[0080] In some embodiments, the cargo box image to be detected is input into the trained improved YOLOv10 model, and the predicted cargo box positioning result in the image is output. The predicted cargo box positioning result may include the position coordinate information of each cargo box area in the cargo box image to be detected.
[0081] In the embodiment of the present invention, by making the predicted container positioning result in the output image include the position coordinate information of each container area in the container image to be detected, clearer container positioning can be achieved.
[0082] In a specific embodiment, the trained improved YOLOv10 model is tested using a 2000 cargo box image test data set. The positioning detection accuracy of the trained improved YOLOv10 model is 96.2%. At the same time, in order to compare the effect of the trained improved YOLOv10 model, the traditional YOLOv10 model is trained using the same data, and its accuracy is 94.9%. This shows that the trained improved YOLOv10 model provided by the embodiment of the present invention has higher accuracy for cargo box positioning detection.
[0083] The cargo box positioning detection method provided by the embodiment of the present invention combines the introduction of positive sample weights and negative sample weights in the label assignment process, and the introduction of the occlusion attention module before the label assignment process, optimizes the YOLOv10 model, enables the model to flexibly handle occlusion situations, and can effectively improve the accuracy of cargo box positioning detection in a logistics environment.
[0084] Example 2
[0085] Based on the above cargo box positioning detection method, an embodiment of the present invention provides a cargo box positioning detection device, and its structural schematic diagram is as follows: Figure 5 As shown, the cargo box positioning detection device 50 includes a data acquisition module 51 and a detection module 52;
[0086] The data acquisition module 51 is used to acquire the image of the cargo box to be detected;
[0087] The detection module 52 is used to input the cargo box image to be detected into the trained improved YOLOv10 model, and output the predicted cargo box positioning result in the image; wherein, when performing consistency matching metric calculation, the YOLOv10 model constructs a positive sample weight determined by a linear combination of the classification score and the intersection-over-union ratio of the predicted box and the true box, a negative sample weight determined by the classification confidence and position mismatch degree of the negative sample, and a consistency matching metric formula determined by the spatial prior value, the classification score, the intersection-over-union ratio of the predicted box and the true box, and the hyperparameters of the balanced classification task and the hyperparameters of the regression task.
[0088] For other details about how the modules in the cargo box positioning detection device 50 implement the above technical solution, please refer to the description of the cargo box positioning detection method provided in the above invention embodiment, which will not be repeated here.
[0089] Example 3
[0090] Based on the above cargo box positioning detection method, the embodiment of the present invention further provides a cargo box positioning detection device, the structural diagram of which is as follows: Figure 6 As shown, the device 60 includes a processor 61 and a memory 62 coupled to the processor 61. The memory 62 stores a computer program, and when the computer program is executed by the processor 61, the processor 61 executes the steps of the cargo box positioning detection method in the above embodiment.
[0091] For other details about how the processor 61 in the cargo box positioning detection device 60 implements the above technical solution, please refer to the description of the cargo box positioning detection method provided in the above invention embodiment, which will not be repeated here.
[0092] Among them, the processor 61 can also be called a CPU (Central Processing Unit), and the processor 61 may be an integrated circuit chip with signal processing capabilities; the processor 61 can also be a general-purpose processor, DSP (Digital Signal Process), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, among which the general-purpose processor can be a microprocessor or the processor 61 can also be any conventional processor, etc.
[0093] Example 4
[0094] The embodiment of the present invention further provides a computer-readable storage medium, a schematic diagram of which is shown in FIG. Figure 7As shown, a readable computer program 71 is stored on the storage medium 70; wherein, the computer program 71 can be stored in the above storage medium 70 in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium 70 includes: U disk, mobile hard disk, magnetic disk or optical disk, ROM (Read-Only Memory), RAM (Random Access Memory) and other media that can store program codes, or terminal devices such as computers, servers, mobile phones, and tablets.
[0095] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0096] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0097] In addition, each functional module in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium.
[0098] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0099] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
[0100] The technical solution provided by the present application is introduced in detail above. The principles and implementation methods of the present application are explained by using specific examples in the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
[0101] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0102] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0103] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0105] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for detecting cargo box positioning, characterized in that: include: Obtain the image of the cargo box to be inspected; The cargo box image to be detected is input into the trained improved YOLOv10 model, and the predicted cargo box positioning result in the image is output; wherein, when performing consistency matching metric calculation, the YOLOv10 model constructs a positive sample weight determined by a linear combination of the classification score and the intersection-over-union ratio of the predicted box and the true box, a negative sample weight determined by the classification confidence and the position mismatch degree of the negative sample, and a consistency matching metric formula determined by the spatial prior value, the classification score, the intersection-over-union ratio of the predicted box and the true box, and a balanced classification task hyperparameter and a regression task hyperparameter.
2. The method for detecting cargo box positioning according to claim 1, characterized in that: The positive sample weight and the negative sample weight introduced include: The calculation formula of the positive sample weight is: pos =γ1·p+γ2·IoU, where w pos is the positive sample weight, γ1 and γ2 are adjustment coefficients, which are used to adjust the impact of classification score and IoU on the positive sample weight respectively; The calculation formula of the negative sample weight is: neg =δ1·(1-p)+δ2·(1-IoU), where w neg is the negative sample weight, δ1 and δ2 are adjustment coefficients, which are used to adjust the influence of classification error and mismatch on the negative sample weight respectively; In the positive sample weight and negative sample weight calculation formula, p is the classification score, and IoU is the intersection over union ratio of the predicted box and the true box.
3. The method for detecting cargo box positioning according to claim 2, characterized in that: The consistency matching measurement formula includes: in, is the consistency matching measurement result, w pos is the positive sample weight, w neg is the negative sample weight, s is the spatial prior value, p is the classification score, For the prediction box The intersection-over-union ratio with the true box b, α is the hyperparameter for the classification task, and β is the hyperparameter for the regression task.
4. The method for detecting cargo box positioning according to claim 1, characterized in that: The YOLOv10 model is constructed with multiple occlusion attention modules after the path aggregation network; The plurality of occlusion attention modules process the multi-scale feature map obtained by the path aggregation network to obtain a feature map in which the features of the unoccluded areas in the cargo box image are enhanced.
5. The method for detecting cargo box positioning according to claim 4, characterized in that: The plurality of occlusion attention modules process the multi-scale feature map obtained by the path aggregation network, including: Each of the occlusion degree attention modules performs a separate convolution operation on each channel of the feature map at one scale, and uses a 1×1 point-by-point convolution to merge the convolution results of each channel.
6. The method for detecting cargo box positioning according to claim 5, characterized in that: The occlusion degree attention module includes a depth convolution unit, a first Gaussian error linear unit, a first batch normalization unit, a point-by-point convolution unit, a second Gaussian error linear unit, and a second batch normalization unit; The deep convolution unit performs a convolution operation on the input feature map, and the convolution result is processed by the first Gaussian error linear unit and the first batch normalization unit in sequence to obtain a first feature map; the input feature map and the first feature map are added through a residual connection to obtain a second feature map; the second feature map is convolved using the point-by-point convolution unit, and the convolution result is processed by the second Gaussian error linear unit and the second batch normalization unit in sequence to form an output.
7. The method for detecting cargo box positioning according to claim 1, characterized in that: The predicted cargo box positioning result in the output image includes the position coordinate information of each cargo box area in the cargo box image to be detected.
8. A cargo box positioning detection device, characterized in that: It includes a data acquisition module and a detection module; The data acquisition module is used to acquire the image of the cargo box to be detected; The detection module is used to input the cargo box image to be detected into the trained improved YOLOv10 model, and output the predicted cargo box positioning result in the image; wherein, when performing consistency matching metric calculation, the YOLOv10 model constructs a positive sample weight determined by a linear combination of the classification score and the intersection-over-union ratio of the predicted box and the true box, a negative sample weight determined by the classification confidence and position mismatch degree of the negative sample, and a consistency matching metric formula determined by the spatial prior value, the classification score, the intersection-over-union ratio of the predicted box and the true box, and the hyperparameters of the balanced classification task and the hyperparameters of the regression task.
9. A cargo box positioning detection device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store computer programs; The processor is used to read the computer program in the memory and execute the steps of the cargo box positioning detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A readable computer program is stored thereon, and when the program is executed by a processor, the steps of the cargo box positioning detection method as described in any one of claims 1 to 7 are implemented.