A method of detecting a liquid container in a package
Patent Information
- Application Number
- CN202510658891.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-30
- Filing Date
- 2025-05-21
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-05-21
AI Technical Summary
[0005]鉴于上述的分析,本发明实施例旨在提供一种包裹中液体容器的检测方法,用以解决现有技术中对液体进行检查时效率低的问题
[0042]与现有技术相比,本发明至少可实现如下有益效果之一:
Smart Images

Figure CN120543516B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more particularly to a method for detecting liquid containers in packages. Background Technology
[0002] With the increasing demand for public safety, security inspection technologies in densely populated places such as airports and subways are constantly developing.
[0003] Currently, the market standard for liquid detection involves removing the liquid container from the package and then placing it into a designated liquid security screening device for scanning to determine the liquid's type. This method is complex and cumbersome. In high-traffic areas such as airports and subways, this traditional method often leads to congestion at security checkpoints, inconveniencing passengers and impacting the overall operational efficiency of the facility.
[0004] To detect liquids in packages, it is first necessary to detect the container holding the liquid. Therefore, there is an urgent need for a method that can quickly detect liquid containers in packages. Summary of the Invention
[0005] Based on the above analysis, the present invention aims to provide a method for detecting liquid containers in packages, thereby solving the problem of low efficiency in the prior art when inspecting liquids.
[0006] This invention provides a method for detecting a liquid container in a package, the method comprising:
[0007] The package was scanned using a CT scanner to obtain X-ray images, which were then preprocessed to obtain the image to be inspected.
[0008] The image to be detected is input into the liquid container recognition model to obtain the prediction results of each target in the image to be detected; wherein, the prediction results include the prediction category and the prediction bounding box information; the liquid container recognition model is trained based on an improved YOLOv8 neural network structure, in which a first depthwise separable convolutional module is used to replace the residual convolutional module in the backbone network of the YOLOv8 neural network structure, and a second depthwise separable convolutional module is used to replace the residual convolutional module in the neck network of the YOLOv8 neural network structure;
[0009] For each target, targets predicted as liquid containers are extracted from the prediction results, and the predicted bounding boxes of the extracted targets are displayed on the visualization screen.
[0010] Based on further improvements to the above detection method, the improved YOLOv8 neural network structure includes a backbone network, a neck network, and a detection head network;
[0011] The backbone network is used to extract features from the input image to be detected, and the obtained third, fourth and fifth feature maps are input into the neck network;
[0012] The neck network is used to fuse the third, fourth, and fifth feature maps to obtain the second, third, and fourth fused feature maps, which are then output to the detection head network.
[0013] The detection head network is used to predict the predicted category and predicted bounding box information of each target in the image to be detected based on the second, third and fourth fusion feature maps; the predicted bounding box information includes the center point coordinates, width and height.
[0014] Based on the further improvement of the above detection method, the backbone network includes a first convolutional module and first to fourth feature extraction layers connected in sequence, and the second to fourth feature extraction layers respectively output the third to fifth feature maps;
[0015] The first to third feature extraction layers have the same structure, each including a first convolutional module and a first depthwise separable convolutional module connected in sequence;
[0016] The fourth feature extraction layer consists of a first convolutional module, a first depthwise separable convolutional module, and an accelerated spatial pyramid pooling layer (SPPF) connected in sequence.
[0017] Based on further improvements to the above detection method, the neck network includes a first sampling fusion layer, a second sampling fusion layer, a first convolutional fusion layer, and a second convolutional fusion layer;
[0018] The first sampling fusion layer is used to concatenate and fuse the fourth feature map and the fifth feature map, and to transmit the fused first fused feature map to the first convolutional fusion layer and the second sampling fusion layer.
[0019] The second sampling fusion layer is used to concatenate and fuse the first fusion feature map and the third feature map, and outputs the fused second fusion feature map to the detection head network and the first convolutional fusion layer.
[0020] The first convolutional fusion layer is used to concatenate and fuse the first fusion feature map and the second fusion feature map, and output the fused third fusion feature map to the detection head network and the second convolutional fusion layer.
[0021] The second convolutional fusion layer is used to concatenate and fuse the fifth feature map and the third fusion feature map, and outputs the resulting fourth fusion feature map to the detection head network.
[0022] Based on the further improvement of the above detection method, the first sampling fusion layer and the second sampling fusion layer have the same structure, both including an upsampling layer, a splicing layer and a second depthwise separable convolutional module;
[0023] The first and second convolutional fusion layers have the same structure, both including a second convolutional module, a splicing layer, and a second depthwise separable convolutional module.
[0024] Based on the further improvement of the above detection method, in the first sampling fusion layer, the fifth feature map obtained by the fifth feature map through the upsampling layer is transmitted to the stitching layer. The fifth fusion feature map is stitched and fused with the fourth feature map in the stitching layer to obtain the sixth fusion feature map, which is transmitted to the second depthwise separable convolution module. The sixth fusion feature map is then processed by the second depthwise separable convolution module to obtain the first fusion feature map.
[0025] In the second sampling fusion layer, the first fusion feature map is passed through the upsampling layer to obtain the seventh fusion feature map, which is then transmitted to the splicing layer. The seventh fusion feature map is spliced and fused with the third feature map in the splicing layer to obtain the eighth fusion feature map, which is then transmitted to the second depthwise separable convolution module. The eighth fusion feature map is then passed through the second depthwise separable convolution module to obtain the second fusion feature map.
[0026] In the first convolutional fusion layer, the second fusion feature map is passed through the second convolutional module to obtain the ninth fusion feature map, which is then transmitted to the splicing layer. The ninth fusion feature map is spliced and fused with the first fusion feature map in the splicing layer to obtain the tenth fusion feature map, which is then transmitted to the second depthwise separable convolutional module. The tenth fusion feature map is then passed through the second depthwise separable convolutional module to obtain the third fusion feature map.
[0027] In the second convolutional fusion layer, the eleventh fusion feature map obtained by the third fusion feature map through the second convolutional module is transmitted to the splicing layer. The eleventh fusion feature map is spliced and fused with the fifth feature map in the splicing layer to obtain the twelfth fusion feature map, which is transmitted to the second depthwise separable convolutional module. The twelfth fusion feature map is then processed by the second depthwise separable convolutional module to obtain the fourth fusion feature map.
[0028] Based on further improvements to the above detection method, the first depthwise separable convolutional module and the second depthwise separable convolutional module have the same structure, both including a first convolutional layer, a separation layer, multiple depthwise separable convolutional layers, a merging layer, and a second convolutional layer.
[0029] In the first or second depthwise separable convolutional module, the input feature map is passed through the first convolutional layer to obtain the sixth feature map and then transmitted to the separation layer and the merging layer.
[0030] The sixth feature map is divided into the seventh and eighth feature maps in the separation layer and then transmitted to the merging layer and multiple depthwise separable convolutional layers, respectively.
[0031] The eighth feature map is split convolutionally performed in the first depthwise separable convolutional layer, and the resulting ninth feature map is passed to the merging layer and the next depthwise separable convolutional layer.
[0032] Except for the first depthwise separable convolutional layer, all other depthwise separable convolutional layers receive the output feature map of the previous depthwise separable convolutional layer, perform a separable convolution operation to obtain the feature map, and then transmit it to the merging layer and the next depthwise separable convolutional layer.
[0033] In the merging layer, the feature maps obtained from the sixth, seventh, and ninth feature maps and all other depth-separable convolutional layers are merged, and the merged tenth feature map is then passed to the second convolutional layer.
[0034] The tenth feature map is passed through the second convolutional layer to obtain the output feature map.
[0035] Based on further improvements to the above detection method, the depth-separable convolutional layer includes a depth convolutional layer, a first normalized layer, a first activation function layer, a third convolutional layer, a second normalized layer, and a second activation function layer connected in sequence.
[0036] Based on further improvements to the above detection method, the detection head network includes a first detection head, a second detection head, a third detection head, and a post-processing layer;
[0037] The first detection head, the second detection head, and the third detection head are used to predict the prediction results of the corresponding feature maps based on the second fusion feature map, the third fusion feature map, and the fourth fusion feature map, respectively.
[0038] The post-processing layer is used to filter the prediction results of the second, third, and fourth fusion feature maps to obtain the prediction results of each target in the image to be detected.
[0039] Based on further improvements to the above detection method, the first detection head, the second detection head, or the third detection head have the same structure, all including a third convolutional module and a fourth convolutional layer connected in sequence;
[0040] The third convolutional module is used to compress the feature dimensions of the input feature map and integrate local information.
[0041] The fourth convolutional layer is used to adjust the number of channels in the input feature map to obtain a preset number of prediction results.
[0042] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0043] 1. A liquid container recognition model is obtained by training an improved YOLOv8 neural network structure. In the improved YOLOv8 neural network structure, a first depthwise separable convolutional module replaces the residual convolutional module in the backbone network of the YOLOv8 neural network structure, and a second depthwise separable convolutional module replaces the residual convolutional module in the neck network of the YOLOv8 neural network structure. The liquid container recognition model is used to predict the target image corresponding to the package, and quickly identify the target that is a liquid container in the package. Combined with a preset gradient descent training method, the liquid container recognition model can be trained quickly to obtain a trained liquid container recognition model, further improving the application range of the detection method.
[0044] 2. By progressively extracting features from the input feature map through the first and second depthwise separable convolution modules, the computational load of using traditional residual convolution modules is reduced, thereby enabling feature map extraction and fusion and rapid prediction of the image to be detected.
[0045] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0046] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0047] Figure 1 This is a flowchart illustrating a method for detecting a liquid container in a package, as provided in an embodiment of the present invention.
[0048] Figure 2 A schematic diagram of the improved YOLOv8 neural network structure provided in an embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the structure of a first depthwise separable convolutional module or a second depthwise separable convolutional module provided in an embodiment of the present invention. Detailed Implementation
[0050] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0051] One specific embodiment of the present invention discloses a method for detecting liquid containers in a package, such as... Figure 1 As shown, the detection method includes:
[0052] Step S1: Use a CT scanner to scan the package to obtain X-ray images, and preprocess the X-ray images to obtain the image to be inspected;
[0053] Step S2: Input the image to be detected into the liquid container recognition model to obtain the prediction results of each target in the image to be detected; wherein, the prediction results include the prediction category and the prediction bounding box information; the liquid container recognition model is trained based on an improved YOLOv8 neural network structure, in which a first depthwise separable convolutional module is used to replace the residual convolutional module in the backbone network of the YOLOv8 neural network structure, and a second depthwise separable convolutional module is used to replace the residual convolutional module in the neck network of the YOLOv8 neural network structure;
[0054] Step S3: Extract targets with the predicted category of liquid container from the prediction results of each target, and display the predicted bounding box of the extracted target on the visualization screen.
[0055] Specifically, such as Figure 1 As shown, in step S1, the package is scanned using a CT scanner to obtain a corresponding X-ray image. It is understood that detailed data such as average density, average atomic number, density root mean square deviation, and atomic number root mean square deviation at any location can be analyzed from the X-ray image of the package. After determining the location of the liquid container, detailed data of the liquid container is extracted and compared with detailed data of hazardous liquids stored in the database to identify the type of liquid stored in each container, enabling rapid liquid inspection.
[0056] Specifically, in step S1, the ray image corresponding to the package is preprocessed to obtain the image to be detected in a set format.
[0057] Preferably, the pretreatment includes one or more of the following operations:
[0058] Grayscale conversion;
[0059] Binarization;
[0060] Normalization;
[0061] Scaling;
[0062] Cutting;
[0063] Sharpen;
[0064] Filtering.
[0065] Specifically, grayscale conversion is the process of converting a color image to a grayscale image, helping to reduce the complexity of image processing and storage requirements. Binarization converts an image into one containing only two pixel values, simplifying image data and making it easier to perform further analysis and processing. Normalization scales the pixel values of an image proportionally to a specific range, such as scaling to the range of 0 to 1, which can adjust the scale of image data to improve the efficiency and accuracy of analysis and modeling. Scaling and cropping adjust the size and content of an image; scaling changes the image's resolution and pixel density, while cropping changes the image's composition and focus, helping to adjust the image to meet analysis and processing needs. Sharpening enhances the edges and contours in an image, making the image appear clearer and more vivid. Filtering processes an image in the frequency or spatial domain, adjusting the frequency components of the image to improve image quality or extract useful information, enhancing or suppressing certain features in the image.
[0066] Specifically, after determining the image to be detected corresponding to the package in step S1, in step S2, the image to be detected is input into the liquid container recognition model. The liquid container recognition model recognizes the image to be detected and obtains the predicted category and predicted bounding box information of each target in the image to be detected.
[0067] Specifically, the liquid container recognition model is obtained by training an improved YOLOv8 neural network structure using a preset gradient descent training method. Specifically, both the backbone and neck networks of the YOLOv8 neural network structure include residual convolutional modules. In the improved YOLOv8 neural network structure of this invention, a first depthwise separable convolutional module replaces the residual convolutional module in the backbone network, and a second depthwise separable convolutional module replaces the residual convolutional module in the neck network. Specifically, by using the first and second depthwise separable convolutional modules to progressively extract features from the input feature map, the computational load of using traditional residual convolutional modules is effectively reduced, enabling feature map extraction and fusion, and rapid prediction of the image to be detected.
[0068] Preferably, the preset gradient descent training method is any one of the following:
[0069] Stochastic gradient descent;
[0070] Adam gradient descent algorithm;
[0071] Batch gradient descent.
[0072] Specifically, stochastic gradient descent (SGD) is computationally efficient and highly adaptable, capable of escaping local minima, which helps rotated image detection models escape local minima and find optimal solutions. Furthermore, it has a fast single-iteration speed and low memory consumption, and can quickly determine the final liquid container recognition model when training it.
[0073] Specifically, the Adam gradient descent algorithm combines the features of momentum gradient descent and RMSProp. By updating model parameters through adaptive learning rate and second-order moment estimation, it can accelerate the convergence speed of the rotated image detection model in the gradient descent process and avoid getting trapped in local optima.
[0074] Specifically, batch gradient descent uses all the data directly, eliminating the hassle of mini-batch selection, and has stable convergence, allowing for faster acquisition of stable rotated image detection models.
[0075] Preferably, the improved YOLOv8 neural network structure includes a backbone network, a neck network, and a detection head network;
[0076] The backbone network is used to extract features from the input image to be detected, and the obtained third, fourth and fifth feature maps are input into the neck network;
[0077] The neck network is used to fuse the third, fourth, and fifth feature maps to obtain the second, third, and fourth fused feature maps, which are then output to the detection head network.
[0078] The detection head network is used to predict the predicted category and predicted bounding box information of each target in the image to be detected based on the second, third and fourth fusion feature maps; the predicted bounding box information includes the center point coordinates, width and height.
[0079] Specifically, such as Figure 2 As shown, the backbone network receives the image to be detected in a set format and performs feature extraction on the image to be detected, which can obtain the first feature map P1, the second feature map P2, the third feature map P3, the fourth feature map P4 and the fifth feature map P5. The third to fifth feature maps are output from the backbone network to the neck network.
[0080] Specifically, such as Figure 2 As shown, the third to fifth feature maps are fused in the neck network to obtain the first fused feature map S1, the second fused feature map S2, the third fused feature map S3 and the fourth fused feature map S4. The second to fourth fused feature maps are then output to the detection head network.
[0081] Specifically, such as Figure 2 As shown, in the neck network, prediction results are obtained by performing predictions based on the second to fourth fused feature maps respectively. The prediction results are the predicted categories and predicted bounding box information of each target in the image to be detected. The predicted bounding box information of each target includes the center point coordinates, width and height of each target.
[0082] Preferably, such as Figure 2 As shown, the backbone network includes a first convolutional module and first to fourth feature extraction layers connected in sequence, and the second to fourth feature extraction layers respectively output the third to fifth feature maps;
[0083] The first to third feature extraction layers have the same structure, each including a first convolutional module and a first depthwise separable convolutional module connected in sequence;
[0084] The fourth feature extraction layer consists of a first convolutional module, a first depthwise separable convolutional module, and an accelerated spatial pyramid pooling layer (SPPF) connected in sequence.
[0085] Specifically, such as Figure 2 As shown, the first convolutional module is CBS1, and the first depthwise separable convolutional module is C2DSC1.
[0086] Specifically, after the image to be detected enters the backbone network, it passes through the first convolutional module, the first feature lifting layer, the second feature lifting layer, the third feature lifting layer and the fourth feature lifting layer in sequence to obtain the first feature map P1, the second feature map P2, the third feature map P3, the fourth feature map P4 and the fifth feature map P5, respectively.
[0087] Specifically, such as Figure 2 As shown, the first feature extraction layer, the second feature extraction layer, and the third feature extraction layer have the same structure, each including a first convolutional module CBS1 and a first depthwise separable convolutional module C2DSC1 connected in sequence.
[0088] Specifically, such as Figure 2 As shown, the fourth feature extraction layer comprises a first convolutional module CBS1, a first depthwise separable convolutional module C2DSC1, and an accelerated spatial pyramid pooling layer SPPF, connected in sequence. Specifically, the SPPF module improves the efficiency of feature extraction and thus the accuracy of detection without sacrificing accuracy, through multi-scale pooling, simplified computation, and increased receptive field.
[0089] Preferably, the neck network includes a first sampling fusion layer, a second sampling fusion layer, a first convolutional fusion layer, and a second convolutional fusion layer;
[0090] The first sampling fusion layer is used to concatenate and fuse the fourth feature map and the fifth feature map, and to transmit the fused first fused feature map to the first convolutional fusion layer and the second sampling fusion layer.
[0091] The second sampling fusion layer is used to concatenate and fuse the first fused feature map and the third feature map, and outputs the fused second fused feature map to the detection head network and the first convolutional fusion layer;
[0092] The first convolutional fusion layer is used to concatenate and fuse the first fusion feature map and the second fusion feature map, and output the fused third fusion feature map to the detection head network and the second convolutional fusion layer.
[0093] The second convolutional fusion layer is used to concatenate and fuse the fifth feature map and the third fusion feature map, and outputs the resulting fourth fusion feature map to the detection head network.
[0094] Specifically, such as Figure 2 As shown, the third feature map P3 in the backbone network is input to the second sampling fusion layer, and the fourth feature map P4 and the fifth feature map P5 in the backbone network are both input to the first sampling fusion layer. Specifically, the first sampling fusion layer outputs the first fused feature map S1 to the first convolutional fusion layer and the second sampling fusion layer; the second sampling fusion layer outputs the second fused feature map S2 to the detection head network and the first convolutional fusion layer; the first convolutional fusion layer outputs the third fused feature map S3 to the detection head network and the second convolutional fusion layer; and the second convolutional fusion layer outputs the fourth fused feature map S4 to the detection head network.
[0095] Preferably, such as Figure 2 As shown, the first sampling fusion layer and the second sampling fusion layer have the same structure, both including an upsampling layer, a splicing layer, and a second depthwise separable convolutional module;
[0096] The first and second convolutional fusion layers have the same structure, both including a second convolutional module, a splicing layer, and a second depthwise separable convolutional module.
[0097] Specifically, such as Figure 2 As shown, the first sampling fusion layer or the second sampling fusion layer includes an upsampling layer (upsample), a concatenation layer (concat), and a second depthwise separable convolutional module (C2DSC2); the first convolutional fusion layer or the second convolutional fusion layer includes a second convolutional module (CBS2), a concatenation layer (concat), and a second depthwise separable convolutional module (C2DSC2).
[0098] Specifically, such as Figure 2As shown, in the first sampling fusion layer, the fifth feature map P5 is upsampled to obtain the fifth fused feature map S5, which is then transmitted to the concatenation layer concat. The fifth fused feature map S5 is concatenated and fused with the fourth feature map P4 in the concatenation layer concat to obtain the sixth fused feature map S6, which is then transmitted to the second depthwise separable convolutional module C2DSC2. The sixth fused feature map S6 is then processed by the second depthwise separable convolutional module C2DSC2 to obtain the first fused feature map S1.
[0099] In the second sampling fusion layer, the first fusion feature map S1 is passed through the upsampling layer to obtain the seventh fusion feature map S7, which is then transmitted to the concat layer. The seventh fusion feature map S7 is concatenated and fused with the third feature map P3 in the concat layer to obtain the eighth fusion feature map S8, which is then transmitted to the second depthwise separable convolutional module C2DSC2. The eighth fusion feature map S8 is then passed through the second depthwise separable convolutional module C2DSC2 to obtain the second fusion feature map S2.
[0100] In the first convolutional fusion layer, the second fusion feature map S2 is passed through the second convolutional module CBS2 to obtain the ninth fusion feature map S9, which is then transmitted to the concatenation layer concat. The ninth fusion feature map S9 is concatenated and fused with the first fusion feature map S1 in the concatenation layer concat to obtain the tenth fusion feature map S10, which is then transmitted to the second depthwise separable convolutional module C2DSC2. The tenth fusion feature map S10 is passed through the second depthwise separable convolutional module C2DSC2 to obtain the third fusion feature map S3.
[0101] In the second convolutional fusion layer, the eleventh fusion feature map S11 obtained by the third fusion feature map S3 through the second convolutional module CBS2 is transmitted to the concatenation layer concat. The eleventh fusion feature map S11 is concatenated and fused with the fifth feature map P5 in the concatenation layer concat to obtain the twelfth fusion feature map S12, which is transmitted to the second depthwise separable convolutional module C2DSC2. The twelfth fusion feature map S12 is transmitted to the second depthwise separable convolutional module C2DSC2 to obtain the fourth fusion feature map S4.
[0102] Specifically, the first convolutional module CBS1 and the second convolutional module CBS2 have the same structure, both including a convolutional layer Conv, a batch normalization layer BN, and a linear unit activation function layer connected in sequence. The convolutional layer Conv can learn simple edges and textures, which helps the deeper network structure to capture more complex patterns and shapes.
[0103] Preferably, the first depthwise separable convolutional module and the second depthwise separable convolutional module have the same structure, both including a first convolutional layer, a separation layer, multiple depthwise separable convolutional layers, a merging layer, and a second convolutional layer;
[0104] In the first or second depthwise separable convolutional module, the input feature map is passed through the first convolutional layer to obtain the sixth feature map and then transmitted to the separation layer and the merging layer.
[0105] The sixth feature map is divided into the seventh and eighth feature maps in the separation layer and then transmitted to the merging layer and multiple depthwise separable convolutional layers, respectively.
[0106] The eighth feature map is split convolutionally performed in the first depthwise separable convolutional layer, and the resulting ninth feature map is passed to the merging layer and the next depthwise separable convolutional layer.
[0107] Except for the first depthwise separable convolutional layer, all other depthwise separable convolutional layers receive the output feature map of the previous depthwise separable convolutional layer, perform a separable convolution operation to obtain the feature map, and then transmit it to the merging layer and the next depthwise separable convolutional layer.
[0108] In the merging layer, the feature maps obtained from the sixth, seventh, and ninth feature maps and all other depth-separable convolutional layers are merged, and the merged tenth feature map is then passed to the second convolutional layer.
[0109] The tenth feature map is passed through the second convolutional layer to obtain the output feature map.
[0110] Preferably, the depth-separable convolutional layer comprises a depthwise convolutional layer, a first normalized layer, a first activation function layer, a third convolutional layer, a second normalized layer, and a second activation function layer connected in sequence.
[0111] Specifically, such as Figure 3 As shown, in the first or second depthwise separable convolutional module, the input feature map is passed through the first convolutional layer to obtain the sixth feature map P6. The sixth feature map P6 is transmitted to the separation layer and the merging layer. The sixth feature map P6 is divided into the seventh feature map P7 and the eighth feature map P8 in the separation layer. The seventh feature map P7 is transmitted to the merging layer, and the eighth feature map P8 is transmitted to the first depthwise separable convolutional layer.
[0112] Specifically, in the first depthwise separable convolutional layer, the eighth feature map P8 sequentially passes through a depthwise convolutional layer, a first normalization layer, a first activation function layer, a third convolutional layer, a second normalization layer, and a second activation function layer to obtain the ninth feature map P9. The depthwise convolutional layer has a 3x3 kernel (depthwise convolution), while the third convolutional layer has a 1x1 kernel (pointwise convolution). The ninth feature map P9 is then passed to the next depthwise separable convolutional layer and the merging layer. Except for the first depthwise separable convolutional layer, all other depthwise separable convolutional layers receive the output feature map from the previous depthwise separable convolutional layer, perform separable convolution operations to obtain feature maps, and then pass them to the merging layer and the next depthwise separable convolutional layer. Understandably, the feature map output from the last depthwise separable convolutional layer is directly input into the merging layer for merging.
[0113] Specifically, such as Figure 3 As shown, in the merging layer, the received sixth feature map P6, seventh feature map P7, ninth feature map P9 and all other feature maps obtained from the separable convolutional layers are merged, and the resulting tenth feature map P10 is transmitted to the second convolutional layer; the feature map obtained by passing the tenth feature map P10 through the second convolutional layer is used as the output feature map.
[0114] It is worth noting that using depthwise separable convolutional layers can effectively reduce the number of nodes and computational load in the network compared to using traditional standard convolutional layers.
[0115] Specifically, during the convolution process, D p λ represents the size of the input feature map. i D represents the number of channels in the input feature map. k λ represents the size of the convolution kernel. o This indicates the number of channels in the output feature map.
[0116] Assume the input feature map is D p ×D p ×λ i The kernel size is D k ×D k The output feature map is D p ×D p ×λ o The parameter ratio K between depthwise separable convolution and standard convolution is... 参数量 And computational complexity K 计算量 All are:
[0117]
[0118] The derivation process is as follows:
[0119] Assuming the input to the network in a standard convolutional structure is D p ×D p ×λ iIt needs to be done through D k ×D k ×λ i ×λ o The convolution operation is performed on the convolution kernel, and the resulting output layer is D. p ×D p ×λ o The size of the feature map, and the number of parameters P required for a standard convolution operation. 标准卷积 The calculation formula is:
[0120] P 标准卷积 =D k ×D k ×λ i ×λ o ;
[0121] The formula for calculating the computational cost C of standard convolution is:
[0122] C 标准卷积 =D k ×D k ×D p ×D p ×λ i ×λ o ;
[0123] Depthwise separable convolution employs both depthwise convolution and pointwise convolution operations, where depthwise convolution assumes D... p ×D p ×λ i The input layer uses D k ×D k ×λ i A convolution operation is performed on the convolution kernel to obtain a result of size D. p ×D p ×λ i Feature map, parameter quantity P during calculation 逐深度卷积 The calculation formula is:
[0124] P 逐深度卷积 =D k ×D k ×λ i ;
[0125] Computational complexity C 逐深度卷积 The calculation formula is:
[0126] C 逐深度卷积 =D k ×D k ×D p ×D p ×λ i ;
[0127] In pointwise convolution, 1×1×λ is used. i×λ o convolution kernel for D p ×D p ×λ i Perform a convolution operation on the feature map, assuming we obtain D. p ×D p ×λ o Size feature map. Parameters P in the calculation process. 逐点卷积 The calculation formula is:
[0128] P 逐点卷积 =1×1×λ i ×λ o ;
[0129] The computational workload C during the calculation process 逐点卷积 The calculation formula is:
[0130] C 逐点卷积 =1×1×D p ×D p ×λ i ×λ o ;
[0131] By comparing the number of parameters and computational cost of depthwise separable convolution with that of standard convolution, we can visually observe the difference between the two. Parameter ratio K 参数量 The calculation formula is:
[0132]
[0133] The ratio of computational quantity K 计算量 The calculation formula is:
[0134]
[0135] Therefore, when the output layer size obtained by using standard convolution operations and depthwise separable convolution operations is the same, the reduction in parameters and computational cost of depthwise separable convolution is only related to the kernel size D. k Number of output channels λ o Relevant. For convolutional structures, the number of parameters and computational cost of the normalization and activation function layers in the network are negligible. Theoretically, optimizing a 3x3 convolution operation using depthwise separable convolution can improve speed by approximately 9 times.
[0136] Preferably, the detection head network includes a first detection head, a second detection head, a third detection head, and a post-processing layer;
[0137] The first detection head, the second detection head, and the third detection head are used to predict the prediction results of the corresponding feature maps based on the second fusion feature map, the third fusion feature map, and the fourth fusion feature map, respectively.
[0138] The post-processing layer is used to filter the prediction results of the second, third, and fourth fusion feature maps to obtain the prediction results of each target in the image to be detected.
[0139] Preferably, the first detection head, the second detection head, or the third detection head have the same structure, each including a third convolutional module and a fourth convolutional layer connected in sequence;
[0140] The third convolutional module is used to compress the feature dimensions of the input feature map and integrate local information.
[0141] The fourth convolutional layer is used to adjust the number of channels in the input feature map to obtain a preset number of prediction results.
[0142] Specifically, such as Figure 3 As shown, the first detection head YOLO Head1, the second detection head YOLO Head2, and the third detection head YOLO Head3 predict the second fusion feature map S2, the third fusion feature map S3, and the fourth fusion feature map, respectively, to obtain the prediction results.
[0143] Specifically, in the post-processing layer, the prediction results corresponding to the second fusion feature map S2, the third fusion feature map S3, and the fourth fusion feature map are filtered to obtain the prediction results of each target in the image to be detected.
[0144] For example, the first detection head YOLO Head1, the second detection head YOLO Head2, and the third detection head YOLOHead3 output prediction boxes of 52*52, 26*26, and 13*13 corresponding to the second fusion feature map S2, the third fusion feature map S3, and the fourth fusion feature map, respectively. Based on the feature maps of three different sizes, predictions are made for targets of different sizes, resulting in a total of 3549 prediction boxes. Each prediction box contains the following information: the confidence level of the target, the confidence level that the target is a liquid container, the center point coordinates of the prediction bounding box, and the width and height.
[0145] The confidence level of the target in each predicted bounding box is compared with a preset target confidence threshold. Predicted bounding boxes with confidence levels below the preset threshold are removed from the 3549 predicted bounding boxes, resulting in the remaining predicted bounding boxes. The intersection-union ratio (IUR) of any two predicted bounding boxes is calculated based on their center point coordinates, width, and height. This IUR is then compared with a preset IUR threshold. If the IUR is higher than the preset threshold, the two predicted bounding boxes are considered identical. The predicted bounding box with the lower target confidence level is removed, while the predicted bounding box with the higher target confidence level is retained. Predicted bounding boxes whose target is a liquid container and whose confidence level is higher than the preset liquid container confidence level are classified as liquid containers, while those whose confidence level is not higher than the preset liquid container confidence level are classified as non-liquid containers. This process yields the various targets included in the image to be detected. The targets obtained at this point may be either liquid containers or non-liquid containers.
[0146] Specifically, such as Figure 1 As shown, in step S3, the prediction results obtained in step S2 are extracted and analyzed. It can be seen that the predicted category of the target in the image to be detected is a liquid container. The predicted bounding box information of the category of liquid container is displayed on the visualization screen for easy viewing by the user.
[0147] Compared with existing technologies, the liquid container detection method in this invention provides a liquid container recognition model trained on an improved YOLOv8 neural network structure. In the improved YOLOv8 neural network structure, a first depthwise separable convolutional module replaces the residual convolutional module in the backbone network, and a second depthwise separable convolutional module replaces the residual convolutional module in the neck network. The liquid container recognition model is used to predict the target image corresponding to the package, quickly identifying the liquid container within the package. By combining this with a preset gradient descent training method, the liquid container recognition model can be trained quickly, further expanding the application range of the detection method. Simultaneously, the first and second depthwise separable convolutional modules progressively extract features from the input feature map, reducing the computational load compared to using traditional residual convolutional modules, achieving feature map extraction and fusion, and enabling rapid prediction of the target image.
[0148] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0149] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for detecting a liquid container in a package, characterized in that, The detection method includes: The package was scanned using a CT scanner to obtain X-ray images, which were then preprocessed to obtain the image to be inspected. The image to be detected is input into the liquid container recognition model to obtain the prediction results of each target in the image to be detected; wherein, the prediction results include the prediction category and the prediction bounding box information; the liquid container recognition model is trained based on an improved YOLOv8 neural network structure, in which a first depthwise separable convolutional module is used to replace the residual convolutional module in the backbone network of the YOLOv8 neural network structure, and a second depthwise separable convolutional module is used to replace the residual convolutional module in the neck network of the YOLOv8 neural network structure; Extract targets whose predicted category is liquid container from the prediction results of each target, and display the predicted bounding box of the extracted target on the visualization screen; The first depthwise separable convolutional module and the second depthwise separable convolutional module have the same structure, both including a first convolutional layer, a separation layer, multiple depthwise separable convolutional layers, a merging layer and a second convolutional layer; In the first or second depthwise separable convolutional module, the input feature map is passed through the first convolutional layer to obtain the sixth feature map and then transmitted to the separation layer and the merging layer. The sixth feature map is divided into the seventh and eighth feature maps in the separation layer and then transmitted to the merging layer and multiple depthwise separable convolutional layers, respectively. The eighth feature map is split convolutionally performed in the first depthwise separable convolutional layer, and the resulting ninth feature map is passed to the merging layer and the next depthwise separable convolutional layer. Except for the first depthwise separable convolutional layer, all other depthwise separable convolutional layers receive the output feature map of the previous depthwise separable convolutional layer, perform a separable convolution operation to obtain the feature map, and then transmit it to the merging layer and the next depthwise separable convolutional layer. In the merging layer, the feature maps obtained from the sixth, seventh, and ninth feature maps and all other depth-separable convolutional layers are merged, and the merged tenth feature map is then passed to the second convolutional layer. The tenth feature map is passed through the second convolutional layer to obtain the output feature map.
2. The detection method according to claim 1, characterized in that, The improved YOLOv8 neural network structure includes a backbone network, a neck network, and a detection head network; The backbone network is used to extract features from the input image to be detected, and the obtained third, fourth and fifth feature maps are input into the neck network; The neck network is used to fuse the third, fourth, and fifth feature maps to obtain the second, third, and fourth fused feature maps, which are then output to the detection head network. The detection head network is used to predict the predicted category and predicted bounding box information of each target in the image to be detected based on the second, third and fourth fusion feature maps; the predicted bounding box information includes the center point coordinates, width and height.
3. The detection method according to claim 2, characterized in that, The backbone network includes a first convolutional module and first to fourth feature extraction layers connected in sequence, and the second to fourth feature extraction layers respectively output the third to fifth feature maps; The first to third feature extraction layers have the same structure, each including a first convolutional module and a first depthwise separable convolutional module connected in sequence; The fourth feature extraction layer consists of a first convolutional module, a first depthwise separable convolutional module, and an accelerated spatial pyramid pooling layer (SPPF) connected in sequence.
4. The detection method according to claim 2, characterized in that, The neck network includes a first sampling fusion layer, a second sampling fusion layer, a first convolutional fusion layer, and a second convolutional fusion layer; The first sampling fusion layer is used to concatenate and fuse the fourth feature map and the fifth feature map, and to transmit the fused first fused feature map to the first convolutional fusion layer and the second sampling fusion layer. The second sampling fusion layer is used to concatenate and fuse the first fused feature map and the third feature map, and outputs the fused second fused feature map to the detection head network and the first convolutional fusion layer; The first convolutional fusion layer is used to concatenate and fuse the first fusion feature map and the second fusion feature map, and output the fused third fusion feature map to the detection head network and the second convolutional fusion layer. The second convolutional fusion layer is used to concatenate and fuse the fifth feature map and the third fusion feature map, and outputs the resulting fourth fusion feature map to the detection head network.
5. The detection method according to claim 4, characterized in that, The first sampling fusion layer and the second sampling fusion layer have the same structure, both including an upsampling layer, a splicing layer, and a second depthwise separable convolutional module; The first and second convolutional fusion layers have the same structure, both including a second convolutional module, a splicing layer, and a second depthwise separable convolutional module.
6. The detection method according to claim 5, characterized in that, In the first sampling fusion layer, the fifth feature map is passed through the upsampling layer to obtain the fifth fused feature map, which is then transmitted to the splicing layer. The fifth fused feature map is spliced and fused with the fourth feature map in the splicing layer to obtain the sixth fused feature map, which is then transmitted to the second depthwise separable convolution module. The sixth fused feature map is then passed through the second depthwise separable convolution module to obtain the first fused feature map. In the second sampling fusion layer, the first fusion feature map is passed through the upsampling layer to obtain the seventh fusion feature map, which is then transmitted to the splicing layer. The seventh fusion feature map is spliced and fused with the third feature map in the splicing layer to obtain the eighth fusion feature map, which is then transmitted to the second depthwise separable convolution module. The eighth fusion feature map is then passed through the second depthwise separable convolution module to obtain the second fusion feature map. In the first convolutional fusion layer, the second fusion feature map is passed through the second convolutional module to obtain the ninth fusion feature map, which is then transmitted to the splicing layer. The ninth fusion feature map is spliced and fused with the first fusion feature map in the splicing layer to obtain the tenth fusion feature map, which is then transmitted to the second depthwise separable convolutional module. The tenth fusion feature map is then passed through the second depthwise separable convolutional module to obtain the third fusion feature map. In the second convolutional fusion layer, the eleventh fusion feature map obtained by the third fusion feature map through the second convolutional module is transmitted to the splicing layer. The eleventh fusion feature map is spliced and fused with the fifth feature map in the splicing layer to obtain the twelfth fusion feature map, which is transmitted to the second depthwise separable convolutional module. The twelfth fusion feature map is then processed by the second depthwise separable convolutional module to obtain the fourth fusion feature map.
7. The detection method according to claim 1, characterized in that, The depthwise separable convolutional layer comprises a depthwise convolutional layer, a first normalized layer, a first activation function layer, a third convolutional layer, a second normalized layer, and a second activation function layer connected in sequence.
8. The detection method according to claim 2, characterized in that, The detection head network includes a first detection head, a second detection head, a third detection head, and a post-processing layer; The first detection head, the second detection head, and the third detection head are used to predict the prediction results of the corresponding feature maps based on the second fusion feature map, the third fusion feature map, and the fourth fusion feature map, respectively. The post-processing layer is used to filter the prediction results of the second, third, and fourth fusion feature maps to obtain the prediction results of each target in the image to be detected.
9. The detection method according to claim 8, characterized in that, The first detection head, the second detection head, or the third detection head have the same structure, all including a third convolutional module and a fourth convolutional layer connected in sequence; The third convolutional module is used to compress the feature dimensions of the input feature map and integrate local information. The fourth convolutional layer is used to adjust the number of channels in the input feature map to obtain a preset number of prediction results.
Citation Information
Patent Citations
Underwater whale target detection method based on lightweight YOLOv4
CN114418930A
Human body target identification method and system
CN114913454A