Cigarette box appearance defect detection method and device, electronic equipment, medium and product
By using a dual-model approach to detect defects in cigarette boxes, identifying defects from both structural and textural feature dimensions and fusing the results, the problem of high false detection and high false negative rates in existing technologies is solved, achieving high-precision automated quality inspection of cigarette box appearance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA TOBACCO ZHEJIANG IND CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-02
AI Technical Summary
Existing cigarette box defect detection technologies are easily affected by external environmental interference, resulting in high false detection and false negative rates, and cannot meet the needs of high-precision quality inspection.
Two appearance defect detection models with different structures were used to detect defects in cigarette box images. Structural and local texture defects were identified from different feature dimensions, and the detection results were fused.
It effectively reduced the false detection rate and missed detection rate of cigarette box appearance defects, achieved high-precision automated quality inspection, and improved quality inspection efficiency and quality control level.
Smart Images

Figure CN122134665A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of cigarette box inspection technology, and in particular to a method, apparatus, electronic device, medium and product for detecting appearance defects in cigarette boxes. Background Technology
[0002] In the cigarette production process, the appearance quality of cigarette packaging is of paramount importance. Cigarette packaging is an essential component of cigarette products, and its appearance quality directly affects the first impression of the cigarette and the brand image.
[0003] In related technologies, traditional cigarette box defect detection mainly relies on machine vision combined with traditional image processing technology with dynamic thresholding. This involves deploying industrial cameras on the production line to collect images, and then using basic algorithms such as edge detection and Hough transform to analyze defects.
[0004] However, this defect detection method is highly susceptible to external environmental interference. For example, dust, changes in light source brightness, and mechanical vibration can all cause detection errors, resulting in a high false detection rate and a high missed detection rate. At the same time, it can only identify simple defect types, and its detection effect on complex defects is poor, which cannot meet the needs of high-precision quality inspection of cigarette box appearance. Summary of the Invention
[0005] This invention provides a method, device, electronic device, medium, and product for detecting appearance defects in cigarette boxes, so as to achieve the effect of accurately detecting appearance defects in cigarette boxes based on images using two appearance defect detection models with different structures.
[0006] According to one aspect of the present invention, a method for detecting defects in the appearance of cigarette boxes is provided, the method comprising: Acquire an image to be processed, wherein the image to be processed includes a cigarette box to be detected; Based on the first appearance defect detection model and the image to be processed, a first defect detection result corresponding to the cigarette box to be detected is determined, and based on the second appearance defect detection model and the image to be processed, a second defect detection result corresponding to the cigarette box to be detected is determined; wherein, the model structures of the first appearance defect detection model and the second appearance defect detection model are different; Based on the first defect detection result and the second defect detection result, a target defect detection result corresponding to the cigarette box to be tested is determined, and the cigarette box to be tested is processed according to the target defect detection result.
[0007] According to another aspect of the present invention, a device for detecting defects in the appearance of cigarette boxes is provided, the device comprising: An image acquisition module is used to acquire an image to be processed, wherein the image to be processed includes a cigarette box to be detected; A cigarette box defect detection module is used to determine a first defect detection result corresponding to the cigarette box to be detected based on a first appearance defect detection model and the image to be processed, and to determine a second defect detection result corresponding to the cigarette box to be detected based on a second appearance defect detection model and the image to be processed; wherein the model structures of the first appearance defect detection model and the second appearance defect detection model are different. The target detection result determination module is used to determine the target defect detection result corresponding to the cigarette box to be detected based on the first defect detection result and the second defect detection result, and to process the cigarette box to be detected based on the target defect detection result.
[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement a cigarette box appearance defect detection method as described in any of the embodiments of this disclosure.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement any of the cigarette box appearance defect detection methods in the embodiments of the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, which, when executed by a processor, implements a cigarette box appearance defect detection method as described in any of the embodiments of the present disclosure.
[0011] The technical solution of this disclosure, by acquiring an image to be processed, including a cigarette box to be inspected, provides a unified input data source for subsequent dual-model defect detection, ensuring the completeness and quantifiability of the visual information of the cigarette box to be inspected, which is the fundamental premise of the entire defect detection process. Furthermore, by determining a first defect detection result corresponding to the cigarette box to be inspected based on a first appearance defect detection model and the image to be processed, and by determining a second defect detection result corresponding to the cigarette box to be inspected based on a second appearance defect detection model and the image to be processed, two appearance defect detection models with different structures are used to detect cigarette box defects from different feature dimensions, compensating for the detection blind spots of a single model and improving the comprehensiveness and accuracy of defect identification. Furthermore, by determining a target defect detection result corresponding to the cigarette box to be inspected based on the first and second defect detection results, and processing the cigarette box to be inspected according to the target defect detection result, the dual-model detection results are fused to output a unified and accurate target defect detection result. Based on this, quality inspection processing is performed on the cigarette box, ensuring the reliability of defect judgment and achieving efficient and accurate handling of cigarette box defects. The technical solution of this disclosure solves the problems of high false detection rate and false negative rate in the prior art, as well as the inability to meet the requirements of high-precision quality inspection of cigarette box appearance. By using dual models with different structures to detect cigarette box defects from different feature dimensions and fusing the detection results, the false detection rate and false negative rate of cigarette box appearance defects are effectively reduced, realizing high-precision automated quality inspection of cigarette box appearance, meeting the requirements of cigarette production for high-precision quality inspection of cigarette box appearance, and improving the quality inspection efficiency and quality control level of the production line.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart illustrating a method for detecting appearance defects in cigarette boxes according to an embodiment of the present disclosure. Figure 2 A flowchart illustrating another method for detecting appearance defects in cigarette boxes provided in this embodiment of the present disclosure; Figure 3 A flowchart illustrating another method for detecting appearance defects in cigarette boxes provided in this embodiment of the present disclosure; Figure 4This is a schematic diagram of a cigarette box appearance defect detection device provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0017] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0018] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0019] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0020] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0021] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0022] Figure 1 This is a flowchart illustrating a method for detecting appearance defects in cigarette boxes according to an embodiment of this disclosure. This embodiment is applicable to detecting appearance defects in cigarette boxes. The method can be executed by a cigarette box appearance defect detection device, which can be implemented in hardware and / or software and can be configured in electronic devices such as computers or servers. Figure 1 As shown, the method in this embodiment includes: S110. Obtain the image to be processed, wherein the image to be processed includes the cigarette box to be detected.
[0023] The image to be processed can be a digital image containing at least one cigarette box to be inspected. The image to be processed can include images captured by image acquisition devices such as industrial cameras at designated locations on the cigarette production line, or images captured directly by image acquisition devices such as cameras. Generally, to meet the input requirements of subsequent defect detection models and improve the accuracy of appearance defect detection, the image to be processed can be obtained after preprocessing operations such as size normalization, contrast enhancement, and noise reduction. The image content of the image to be processed can cover the complete appearance of the cigarette box to be inspected, and can present various potential defects such as thread misalignment, packaging wrinkles, surface stains, and blurred printing. The cigarette box to be inspected can be a cigarette box to be inspected for appearance defects. Typically, appearance defect inspection can be performed on cigarette boxes on the cigarette production line, and these cigarette boxes that need appearance defect inspection can be used as the cigarette boxes to be inspected. The cigarette box to be inspected can be the core inspection object in the image to be processed, and its appearance quality directly affects the brand image and factory standards of the cigarette product. The cigarette box to be inspected can be any type of cigarette box, optionally including at least one of the following: cigarette pack boxes and cigarette carton boxes (a carton of cigarettes includes multiple cigarette packs). Potential appearance defects in the cigarette box to be inspected are divided into two categories: structural defects and localized texture defects.
[0024] In this embodiment, the acquisition method to be processed may include at least one of the following: real-time acquisition based on image acquisition devices such as industrial cameras pre-installed on the cigarette production line; direct acquisition by taking pictures of the cigarette box to be inspected using image acquisition devices such as cameras; acquisition from the target storage space in response to the user's selection operation; received from external devices; or a preset template image.
[0025] In one embodiment, an industrial camera can be pre-installed on the cigarette production line. Furthermore, during cigarette production, when the cigarette box to be inspected moves along the production line to a preset position for image acquisition, the pre-installed industrial camera can acquire an image of the cigarette box to obtain a raw image. Further, the raw image is pre-processed, and the pre-processed raw image is determined as the image to be processed.
[0026] S120. Based on the first appearance defect detection model and the image to be processed, determine the first defect detection result corresponding to the cigarette box to be detected, and based on the second appearance defect detection model and the image to be processed, determine the second defect detection result corresponding to the cigarette box to be detected.
[0027] The first appearance defect detection model can be a deep learning model that takes an image containing the cigarette box to be detected as input and performs appearance defect detection based on the image. The first appearance defect detection model can be a deep learning model improved from the YOLO model as a baseline architecture. In this embodiment, a multi-scale fusion module with dilated convolution can be introduced into the neck network of the YOLOv5 model to obtain the first appearance defect detection model. This multi-scale fusion module integrates modules such as multi-scale dilated convolution, adaptive average pooling, channel attention, spatial channel attention, and residual connections. The multi-scale fusion module can be used to expand the receptive field. The first appearance defect detection model can be used to detect any type of appearance defect on the cigarette box to be detected. Preferably, the first appearance defect detection model can be used to detect structural defects on the cigarette box to be detected; that is, the first appearance defect detection model has higher detection accuracy for structural defects on the cigarette box to be detected. However, the first appearance defect detection model can also detect appearance defects on the cigarette box to be detected other than structural defects. It can be understood that structural defects can be types of defects caused by problems in the cigarette box packaging molding process, resulting in abnormal spatial structures in the appearance of the cigarette box. Structural defects are characterized by deformation or displacement of the cigarette box's physical form, typically including thread misalignment and packaging wrinkles. Structural defects have obvious spatial location features, and detection accuracy can be improved by expanding the model's receptive field. The first defect detection result can be understood as the structured data set output by the first appearance defect detection model after detecting the input image. The first defect detection result can include at least three core components: defect bounding box information, defect category label, and confidence score.
[0028] The second appearance defect detection model can be a deep learning model that takes an image containing the cigarette box to be inspected as input and performs appearance defect detection based on the image. The second appearance defect detection model can be a deep learning model using the YOLOv8n architecture. The model structures of the first and second appearance defect detection models are different. The second appearance defect detection model can be used to detect any type of appearance defect on the cigarette box to be inspected. Preferably, the second appearance defect detection model has high sensitivity to changes in texture features in the image being processed and can be used to detect local texture defects on the cigarette box to be inspected. That is, the second appearance defect detection model has higher detection accuracy for local texture defects on the cigarette box to be inspected. However, the second appearance defect detection model can also detect appearance defects on the cigarette box to be inspected other than local texture defects. It can be understood that local texture defects can be a type of defect on the surface of the cigarette box caused by stains, printing process errors, or abnormal texture features in a local area. Local texture defects do not involve overall structural deformation of the cigarette box; typical examples include surface dirt and blurred printing. Local texture defects are characterized by variations in local pixel texture. Detection sensitivity can be improved through anchorless bounding box detection and precise label assignment strategies in the second appearance defect detection model. The second defect detection result can be understood as a structured dataset output by the second appearance defect detection model after detecting the input image. The second defect detection result can include at least three core components: defect bounding box information, defect category label, and confidence score.
[0029] In this embodiment, the first appearance defect detection model may include an input network, a backbone network, a neck network, and an output network. After obtaining a processing image including the cigarette box to be detected, the processing image can be input into the first appearance defect detection model. Further, the processing image can be processed sequentially by the input network, backbone network, neck network, and output network in the first appearance defect detection model to obtain a first defect detection result corresponding to the cigarette box to be detected. Additionally, the processing image can be input into a second appearance defect detection model, which performs appearance defect detection on the cigarette box in the processing image to obtain a second defect detection result corresponding to the cigarette box to be detected.
[0030] In this embodiment, both the first appearance defect detection model and the second appearance defect detection model can be trained on a pre-built deep learning model based on sample images of sample cigarette boxes and actual appearance defect information corresponding to the sample cigarette boxes. The training process is explained in detail below using the first appearance defect detection model as an example: Multiple training samples are acquired, including sample images of sample cigarette boxes and actual appearance defect information corresponding to the sample cigarette boxes; for each training sample, the sample images are input into the pre-built first deep learning model to obtain appearance defect detection information; based on the actual appearance defect information and appearance defect detection information in the training samples, a loss value is determined, and the model parameters of the first deep learning model are adjusted according to the loss value to obtain the trained first appearance defect detection model when the preset training acceptance conditions are met.
[0031] S130. Based on the first defect detection result and the second defect detection result, determine the target defect detection result corresponding to the cigarette box to be tested, and process the cigarette box to be tested according to the target defect detection result.
[0032] The target defect detection result can be understood as a final defect detection result with high confidence, generated after verifying and fusing the first and second defect detection results. Optionally, the target defect detection result may include no defects, a set of target defect bounding boxes, and / or target defect categories. The target defect detection result can be directly converted into a standard annotation file as a model training sample; or, the target defect detection result can be used to guide the subsequent processing flow of cigarette box quality inspection. The processing of the cigarette boxes to be inspected can be a subsequent quality inspection action performed based on the target defect detection result. This quality inspection action can include at least two situations: the first is to perform a release operation for cigarette boxes to be inspected that are determined to be defect-free or have defects within the acceptable range, allowing them to enter the next production process; the second is to perform a sorting and rejection operation or a manual re-inspection operation for cigarette boxes to be inspected that are determined to have unacceptable defects or have failed verification and require manual re-inspection.
[0033] In practical applications, when using neural network models to detect appearance defects in cigarette boxes, a single defect detection model is typically used to process the image and obtain the detection result against the target defect of the cigarette box. However, using a single defect detection model to detect appearance defects in cigarette boxes is prone to creating detection blind spots due to the limitations of the model's own feature extraction, and the lack of result cross-validation leads to a high rate of false positives and false negatives, making it difficult to meet the actual production needs of high-precision quality inspection of cigarette box appearance.
[0034] To address the above issues, this embodiment employs a first appearance defect detection model and a second appearance defect detection model to process the image to be processed, obtaining first and second defect detection results respectively. Furthermore, the first and second defect detection results are integrated and verified to obtain the target defect detection result corresponding to the cigarette box to be inspected. By using two models for separate detection and integrating and verifying the detection results, the detection blind spots of a single model are effectively compensated for. Cross-validation of the results significantly reduces the false detection rate and false negative rate of cigarette box appearance defects, improving the accuracy and reliability of defect detection and meeting the actual production needs for high-precision quality inspection of cigarette box appearance.
[0035] In this embodiment, the determination of the target defect detection result may include at least one of the following: processing the first defect detection result and the second defect detection result using a detection result fusion model to obtain the output target defect detection result; or processing the first defect detection result and the second defect detection result using a detection result verification fusion method to obtain the target defect detection result. The detection result fusion model may include an intelligent agent, a conversational language model, and / or a pre-trained deep learning model. The detection result verification fusion method may include semantic consensus verification, spatial consistency verification, and / or verification result fusion.
[0036] Optionally, based on the first defect detection result and the second defect detection result, the target defect detection result corresponding to the cigarette box to be detected is determined, including: inputting the first defect detection result and the second defect detection result into a pre-trained detection result fusion model, and fusing the first defect detection result and the second defect detection result through the detection result fusion model to obtain the output target defect detection result.
[0037] Optionally, the first defect detection result includes first defect category information and first defect bounding box information; the second defect detection result includes second defect category information and second defect bounding box information; based on the first defect detection result and the second defect detection result, the target defect detection result corresponding to the cigarette box to be detected is determined, including: based on the first defect detection result and the second defect detection result, the target defect detection result corresponding to the cigarette box to be detected is determined, including: using a semantic verification method to perform semantic consensus verification on the first defect category information and the second defect category information to obtain semantic consensus verification information; wherein, the semantic consensus verification information includes the semantic consensus verification result and the target defect category; if the semantic consensus verification result is passed, using a spatial matching method to perform spatial consistency verification on the first defect bounding box information and the second defect bounding box information to obtain spatial consistency verification information; the spatial consistency verification information includes the spatial consistency verification result and the target defect bounding box set; if the spatial consistency verification result is passed, the target defect detection result corresponding to the cigarette box to be detected is determined based on the target defect bounding box set and the target defect category.
[0038] In one implementation, after obtaining the first defect detection result and the second defect detection result, the first defect category information from the first defect detection result and the second defect category information from the second defect detection result are extracted. A semantic verification method is used to count the defect type sets and corresponding instance counts of both results. Only when the defect type sets are completely identical and the number of instances of the same type is completely identical is the semantic consensus verification deemed successful, and the target defect category is determined, forming semantic consensus verification information containing the verification result and the target defect category. Further, if the semantic consensus verification result is successful, the first defect bounding box information from the first defect detection result and the second defect bounding box information from the second defect detection result are extracted. A spatial matching method is used to perform spatial consistency verification on the defect bounding boxes of the same target defect category, filtering out target defect bounding boxes that meet preset conditions, forming spatial consistency verification information containing the spatial consistency verification result and the target defect bounding box set. Further, if the spatial consistency verification result is successful, the target defect bounding box set is associated and matched with the corresponding target defect category, ultimately determining the target defect detection result corresponding to the cigarette box to be detected.
[0039] The technical solution of this disclosure, by acquiring an image to be processed, including a cigarette box to be inspected, provides a unified input data source for subsequent dual-model defect detection, ensuring the completeness and quantifiability of the visual information of the cigarette box to be inspected, which is the fundamental premise of the entire defect detection process. Furthermore, by determining a first defect detection result corresponding to the cigarette box to be inspected based on a first appearance defect detection model and the image to be processed, and by determining a second defect detection result corresponding to the cigarette box to be inspected based on a second appearance defect detection model and the image to be processed, two appearance defect detection models with different structures are used to detect cigarette box defects from different feature dimensions, compensating for the detection blind spots of a single model and improving the comprehensiveness and accuracy of defect identification. Furthermore, by determining a target defect detection result corresponding to the cigarette box to be inspected based on the first and second defect detection results, and processing the cigarette box to be inspected according to the target defect detection result, the dual-model detection results are fused to output a unified and accurate target defect detection result. Based on this, quality inspection processing is performed on the cigarette box, ensuring the reliability of defect judgment and achieving efficient and accurate handling of cigarette box defects. The technical solution of this disclosure solves the problems of high false detection rate and false negative rate in the prior art, as well as the inability to meet the requirements of high-precision quality inspection of cigarette box appearance. By using dual models with different structures to detect cigarette box defects from different feature dimensions and fusing the detection results, the false detection rate and false negative rate of cigarette box appearance defects are effectively reduced, realizing high-precision automated quality inspection of cigarette box appearance, meeting the requirements of cigarette production for high-precision quality inspection of cigarette box appearance, and improving the quality inspection efficiency and quality control level of the production line.
[0040] Figure 2 This is a schematic flowchart illustrating another method for detecting appearance defects in cigarette boxes provided in this embodiment. The technical solution of this embodiment can be combined with other embodiments; for the same or related parts, descriptions of other embodiments can be used, and will not be repeated here. Figure 2 As shown, the method in this embodiment may specifically include: S210. Obtain the image to be processed, wherein the image to be processed includes the cigarette box to be detected.
[0041] S220. Receive the image to be processed through the input network and perform data augmentation on the image to be processed to obtain the image tensor corresponding to the image to be processed.
[0042] The first appearance defect detection model may include an input network, a backbone network, a neck network, and an output network. The input network, which is the first functional module of the first appearance defect detection model, is responsible for receiving the image to be processed and performing data augmentation operations on it. Data augmentation operations may include mosaic data augmentation and / or adaptive anchor box calculation. It is understood that mosaic data augmentation is a commonly used and efficient data augmentation method in YOLO series models, which can significantly improve the model's generalization ability. The tensor of the image to be processed may be a multidimensional array form transformed from the image to be processed after data augmentation by the input network. The tensor of the image to be processed may be a digital data format that the deep learning model can directly compute. The tensor dimensions corresponding to a typical image to be processed include at least one of the following: batch size, number of channels, height, and width. The tensor of the image to be processed can fully carry the pixel color and spatial location information of the image to be processed.
[0043] In this embodiment, upon obtaining the image to be processed, it can be input into the first appearance defect detection model. Further, the image to be processed can be received via an input network, and data augmentation can be performed on the image to obtain an augmented image. Further, the augmented image can be converted into a tensor format that the model can directly compute, thereby obtaining the image tensor.
[0044] S230. Receive the tensor of the image to be processed through the backbone network, and downsample and enhance the tensor of the image to be processed to obtain the basic feature map of the image to be processed at multiple scales.
[0045] The backbone network can be the core feature extraction module of the first appearance defect detection model. The backbone network can adopt the core backbone network architecture of the YOLOv5 model, such as the Cross-Stage Partial-DarkNet (CSP-DarkNet) structure, which includes a focusing layer, multiple sets of Cross-Stage Partial (CSP) modules, and convolutional layers. The focusing layer can be understood as the first feature processing module of the YOLOv5 backbone network. Through the focusing layer, the dimensionality of the image tensor to be processed can be reduced, doubling the number of image channels and halving the size without losing image details, quickly extracting shallow features of the image, adapting to the detail capture requirements of cigarette box structural defects. Multiple sets of CSP modules can be used to split the feature map output by the focusing layer into two parts. One part undergoes convolution and activation operations to enhance features, while the other part skips complex calculations. Finally, these two parts of features are fused. This design improves the feature extraction capability of cigarette box defects while reducing model computation and avoiding overfitting during training. Convolutional layers, using pre-defined kernels, slide across feature maps output by multiple CSP modules to capture local features such as edges, textures, and shapes. They are the core unit for extracting spatial features of structural defects in cigarette boxes. Working in conjunction with multiple CSP modules, convolutional layers progressively refine key features of structural defects such as thread offsets and packaging wrinkles, laying the foundation for multi-scale fusion in the neck network. Base feature maps can be feature representations at different scales output by the backbone network. Multiple scale base feature maps can be understood as multiple base feature maps, each corresponding to a different scale. Optionally, these multiple base feature maps can correspond to three scales: large, medium, and small, representing shallow detail features and deep semantic features of the image to be processed, respectively. This multi-scale design allows the model to simultaneously capture structural defect targets of different sizes.
[0046] In one implementation, the backbone network may include a focusing layer, multiple sets of cross-stage partial modules, and convolutional layers. Further, when the input network outputs a tensor of the image to be processed, this tensor can be input into the backbone network. Further, the focusing layer in the backbone network can downsample the tensor to be processed to obtain a first output feature map, which is then input into the multiple sets of cross-stage partial modules. Further, the multiple sets of cross-stage partial modules can process the first output feature map to obtain a second output feature map, which is then input into the convolutional layers. Further, the convolutional layers perform feature extraction on the second output feature map to obtain base feature maps of the output image to be processed at multiple scales, i.e., multiple base maps are obtained.
[0047] S240. Receive multiple basic feature maps through the neck network, and perform feature fusion and enhancement on the multiple basic feature maps according to the neck network to obtain a multi-scale fused feature map.
[0048] The neck network is a core improvement module in the first appearance defect detection model, its main function being to expand the model's receptive field. Based on the YOLOv5 model structure, the neck network embeds a dilated convolutional multi-scale fusion module. This module can fuse and enhance features from multiple basic feature maps at different scales, outputting a more powerful multi-scale fused feature map. Specifically, the neck network expands the receptive field of the first appearance defect detection model. The receptive field can be understood as the region of the original input image corresponding to a single pixel on the output feature map in a neural network. Expanding the receptive field allows the model to capture broader image context information without sacrificing feature resolution, thereby improving the detection sensitivity for structural defects such as thread misalignment and packaging wrinkles. The multi-scale fused feature map is the feature representation output by the neck network after fusing and enhancing basic feature maps at multiple scales. The multi-scale fused feature map integrates defect features at different scales and strengthens key defect information through an attention mechanism, resulting in a significantly better representation of structural defects in cigarette boxes than the basic feature maps.
[0049] In this embodiment, the neck network may include a multi-scale dilated convolution module, an average pooling module, a first attention module, a second attention module, and a residual module. Furthermore, when multiple basic feature maps are received through the neck network, the multi-scale dilated convolution module, the average pooling module, the first attention module, the second attention module, and the residual module can sequentially process the multiple basic feature maps to obtain a multi-scale fused feature map.
[0050] Optionally, the neck network includes a multi-scale dilated convolution module, an average pooling module, a first attention module, a second attention module, and a residual module. The neck network performs feature fusion and enhancement on multiple base feature maps to obtain a multi-scale fused feature map, including: performing parallel convolution on each base feature map using the multi-scale dilated convolution module to obtain multiple first feature maps; downsampling each first feature map using the average pooling module to obtain multiple second feature maps; enhancing defect features on each second feature map using the first attention module to obtain multiple third feature maps; locating defect features on each third feature map using the second attention module to obtain multiple fourth feature maps; and performing residual connections between the multiple fourth feature maps and multiple base feature maps using the residual module to obtain the multi-scale fused feature map.
[0051] The multi-scale spatial convolution module is a core sub-module in the neck network used to expand the model's receptive field. This module incorporates multiple sets of convolutional kernels with different dilation rates, enabling parallel convolution operations on base feature maps at different scales. This allows for the capture of multi-scale spatial features of structural defects on the cigarette box without reducing the resolution of the base feature maps, adapting to the detection needs of defects of different sizes and locations. The first feature map is the feature representation output by the multi-scale dilated convolution module after performing parallel convolution on a single base feature map. The first feature map retains the size information of the base feature map, and due to the dilated convolution, its receptive field is significantly expanded, covering a wider image area and providing stronger global spatial feature representation of structural defects in the cigarette box. The average pooling module is a sub-module in the neck network used for feature aggregation and size alignment. This sub-module employs an adaptive average pooling strategy. The average pooling module receives the first feature map and performs downsampling by calculating the average pixel value of a local region of the first feature map. This reduces the size difference of the first feature map at different scales, achieving the aggregation of feature spatial information and laying the foundation for unified processing by the subsequent attention module. The second feature map can be the feature representation output by the average pooling module after downsampling a single first feature map. The second feature map undergoes size normalization and alignment, ensuring a unified size specification across different scales. This allows it to be directly input into the attention module for processing, avoiding feature fusion errors caused by size differences. The first attention module can be a channel attention submodule built on a multilayer perceptron in the neck network. This module dynamically strengthens channel feature signals related to cigarette box structural defects by calculating the weight values of each feature channel, while weakening interference from irrelevant channels such as background and noise, thus filtering key defect features from a channel perspective. The third feature map can be the feature representation output by the first attention module after performing channel attention weighting on a single second feature map. The channel-dimensional features of the third feature map are reweighted, amplifying the feature signals of defect-related channels and suppressing the feature signals of background interference channels, significantly improving the recognizability of defect features. The second attention module can be a dual-channel attention submodule built on a convolutional block attention module in the neck network. The second attention module first filters key features again in the channel dimension, and then locates the specific area of the cigarette box defect in the spatial dimension. This allows for precise localization of structural defects from both channel and spatial dimensions, further highlighting the feature information of the defect area. The fourth feature map can be the feature representation output by the second attention module after completing channel and spatial attention filtering on a single third feature map. The fourth feature map not only strengthens the defect-related channel features but also accurately locates the spatial position of the defect, achieving optimal representation accuracy for structural defects such as thread offset and packaging wrinkles. The residual module can be a sub-module in the neck network used for feature fusion and gradient preservation.This module uses skip connections to perform pixel-by-pixel superposition calculations between the fourth feature map and the original input base feature map. Its core function is to avoid feature loss or gradient vanishing issues during model training, improving model training stability while preserving the original information of the base feature map. The multi-scale fused feature map can be the final feature representation output by the residual module after completing residual connections. The multi-scale fused feature map integrates the original information of the base feature map, as well as the enhanced features optimized by multi-scale dilated convolution, pooling, and dual attention, achieving optimal representation of structural defects in cigarette boxes. It can be directly input into the output network for defect detection.
[0052] In one implementation, when multiple basic feature maps are input into the neck network, a multi-scale dilated convolution module first performs parallel convolution operations on each basic feature map output from the backbone network to expand the receptive field and capture the multi-scale spatial features of structural defects in cigarette boxes, resulting in multiple corresponding first feature maps. Further, an average pooling module performs adaptive downsampling on each first feature map, aggregating spatial information and aligning the sizes of feature maps at different scales, resulting in multiple corresponding second feature maps. Next, a first attention module performs channel-dimensional weight allocation and feature enhancement on each second feature map, amplifying defect-related channel feature signals and weakening irrelevant background interference, resulting in multiple corresponding third feature maps. Subsequently, a second attention module performs dual-dimensional feature filtering (channel and spatial) and precise defect location localization on each third feature map, further highlighting the spatial regional features of structural defects, resulting in multiple corresponding fourth feature maps. Finally, the residual module is used to perform pixel-by-pixel residual connection and fusion of multiple fourth feature maps with multiple original basic feature maps, preserving the original feature information to avoid feature loss, and finally obtaining a multi-scale fused feature map with the best representation ability for cigarette box structural defects.
[0053] S250. Receive a multi-scale fusion image through the output network and perform detection on the first multi-scale fusion image to obtain a first defect detection result corresponding to the cigarette box to be detected. Also, determine a second defect detection result corresponding to the cigarette box to be detected based on the second appearance defect detection model and the image to be processed.
[0054] The output network can be the output module of the first appearance defect detection model, using the native detection head structure of YOLOv5. The output network can receive the multi-scale fused feature map output by the neck network, and output the first defect detection result through operations such as predicted box regression and class confidence calculation. This result includes defect bounding box information, defect class label, and confidence score.
[0055] In this embodiment, after obtaining a multi-scale fused feature map, the multi-scale fused feature map can be input into the output network. Furthermore, the output network can perform appearance defect detection based on the multi-scale fused feature map to obtain a first defect detection result corresponding to the cigarette box to be detected. Moreover, when the image to be processed is input into the first appearance defect detection model, it can also be input into a second appearance defect detection model. The second appearance defect detection model processes the image to be processed to obtain a second defect detection result corresponding to the cigarette box to be detected.
[0056] S260. Based on the first defect detection result and the second defect detection result, determine the target defect detection result corresponding to the cigarette box to be tested, and process the cigarette box to be tested according to the target defect detection result.
[0057] The technical solution of this embodiment receives an image to be processed through an input network and performs data augmentation on the image to be processed to obtain an image tensor corresponding to the image to be processed. Further, it receives the image tensor to be processed through a backbone network and performs downsampling and feature enhancement on the image tensor to obtain basic feature maps of the image to be processed at multiple scales. Further, it receives multiple basic feature maps through a neck network and performs feature fusion and enhancement on the multiple basic feature maps according to the neck network to obtain a multi-scale fused feature map. Further, it receives the multi-scale fused map through an output network and detects the first multi-scale fused map to obtain the output first defect detection result corresponding to the cigarette box to be detected. By sequentially performing data augmentation, multi-scale feature extraction, and fusion enhancement on the cigarette box image, it finally outputs an accurate first defect detection result, effectively improving the extraction accuracy and detection accuracy of cigarette box appearance defect features, and laying a high-quality single-model detection foundation for subsequent dual-model result fusion.
[0058] Figure 3 This is a schematic flowchart illustrating another method for detecting appearance defects in cigarette boxes provided in this embodiment. The technical solution of this embodiment can be combined with other embodiments; for the same or related parts, descriptions of other embodiments can be used, and will not be repeated here. Figure 3 As shown, the method in this embodiment may specifically include: S310. Obtain the image to be processed, wherein the image to be processed includes the cigarette box to be detected.
[0059] S320. Based on the first appearance defect detection model and the image to be processed, determine the first defect detection result corresponding to the cigarette box to be detected, and based on the second appearance defect detection model and the image to be processed, determine the second defect detection result corresponding to the cigarette box to be detected; the first defect detection result includes first defect category information and first defect bounding box information; the second defect detection result includes second defect category information and second defect bounding box information.
[0060] The first defect category information can be relevant data from the first defect detection results, used to characterize the type of appearance defect on the cigarette box to be detected. This includes the specific category name and / or code of the detected first appearance defect, the number of instances of each category, and is qualitative and quantitative statistical information on the appearance defects of the cigarette box. The first defect category information includes a set of first defect categories and a set of first defect quantities corresponding to the first defect category set. The first defect bounding box information can be relevant data from the first defect detection results, used to characterize the spatial location and extent of the first appearance defect on the cigarette box to be detected. This includes the bounding box information and / or confidence score corresponding to each detected appearance defect, and is quantitative information on the spatial location of the appearance defect of the cigarette box. The first defect bounding box information can include the coordinates of at least one first bounding box of a first appearance defect. The second defect category information can be relevant data from the second defect detection results, used to characterize the type of appearance defect on the cigarette box to be detected. This includes the specific category name and / or code of the detected second appearance defect, the number of instances of each category, and is qualitative and quantitative statistical information on the appearance defects of the cigarette box. The second defect category information includes a set of second defect categories and a set of first defect quantities corresponding to the second defect category set. The second defect bounding box information can be relevant data from the second defect detection results, used to characterize the spatial location and extent of the second appearance defect on the cigarette box to be detected. This includes the bounding box information and / or confidence score corresponding to each detected appearance defect, and is a quantitative information on the spatial localization of the cigarette box appearance defect. The second defect bounding box information can include the coordinates of the first bounding box of at least one second appearance defect.
[0061] S330. Use semantic verification to perform semantic consensus verification on the first defect category information and the second defect category information to obtain semantic consensus verification information; wherein, the semantic consensus verification information includes semantic consensus verification results and target defect category information.
[0062] The semantic verification method can be understood as the specific judgment method used to perform semantic consensus verification on the first defect category information and the second defect category information. The core of the semantic verification method can be to statistically compare the defect type sets and corresponding instance counts of the two models. The semantic consensus verification is considered successful only when the defect type sets are completely identical and the number of defect instances of the same type is completely matched. This is a means of verifying the consistency of the defect category judgment results of the two models. The semantic consensus verification information can be the comprehensive result data obtained after performing semantic consensus verification. It includes two core parts: the semantic consensus verification result and the target defect category. It is the output data of the semantic verification stage, providing the judgment basis and core detection object for the subsequent spatial consistency verification. The semantic consensus verification result can be the core component of the semantic consensus verification information, representing the judgment conclusion of the semantic verification stage. It includes two results: verification passed and verification failed, used to determine whether the conditions for performing subsequent spatial consistency verification are met. The target defect category information can be the comprehensive representation information of the defect categories and quantities generated after the semantic consensus verification is successful, with the two models reaching complete consensus. It is composed of the target defect category set and the target defect quantity set. The target defect category information can be the core target output of the semantic consensus verification process. It integrates the qualitative (category) and quantitative (quantity) defect information after consensus, providing a unified category and quantity basis for subsequent defect detection operations throughout the entire process. The target defect category set can be the collective term for defect categories determined after the semantic consensus verification is passed and the two models reach a complete consensus. Specifically, it is an unordered, non-repeating set directly adopted from the first defect category set. The target defect category set can serve as the basis for defect category determination in subsequent spatial consistency verification and target defect detection result generation stages, inheriting all elements from the first defect category set. The target defect quantity set can be the collective term for the number of defects determined after the semantic consensus verification is passed and the two models reach a complete consensus. It has a strict one-to-one correspondence with the target defect category set, specifically a set of values directly adopted from the first defect quantity set. This set accurately represents the actual number of detected instances corresponding to each defect category in the target defect category set, serving as a quantitative reference for subsequent defect bounding box matching, inheriting all values from the first defect quantity set.
[0063] Optionally, the first defect category information includes a first defect category set and a first defect quantity set corresponding to the first defect category set; the second defect category information includes a second defect category set and a second defect quantity set corresponding to the second defect category set; semantic consensus verification is performed on the first defect category information and the second defect category information using a semantic verification method to obtain semantic consensus verification information, including: determining whether the first defect category set and the second defect category set are consistent to obtain a first verification result, and determining whether the second defect quantity set is consistent to obtain a second verification result; if both the first verification result and the second verification result are consistent, the semantic consensus verification result is determined to be verified as passed, the first defect category set is taken as the target defect category set, the first defect quantity set is taken as the target defect quantity set, and the target defect category set and the target defect quantity set are determined as the target defect category information.
[0064] The first defect category set is a core component of the first defect category information. It refers to the set of specific category names and / or codes of all first appearance defects detected by the model from the image of the cigarette box to be inspected. It is an unordered and non-repeating defect category representation, used to reflect the range of defect types detected by the first appearance defect detection model. For example, if the first appearance defect detection model detects thread misalignment, packaging wrinkles, surface contamination, and printing blur, then the first defect category set is {thread misalignment, packaging wrinkles, surface contamination, printing blur}. The first defect quantity set is also a core component of the first defect category information. It corresponds one-to-one with the first defect category set and is the set of the actual number of instances detected for each defect category in the first defect category set. It is used to quantitatively represent the specific number of each type of first appearance defect detected by the first appearance defect detection model. For example, if the first defect category set is {thread misalignment, packaging wrinkles, surface contamination, printing blur}, and one instance of each of the four types of defects is detected, then the first defect quantity set is {1, 1, 1, 1}. The first appearance defect can be a single appearance defect on the cigarette box to be inspected, detected by the first appearance defect detection model. For example, if the first appearance defect detection model detects a pull thread misalignment defect, this pull thread misalignment is an independent first appearance defect. The second defect category set can be a core component of the second defect category information, referring to the set of specific category names and / or codes of all second appearance defects detected by the model from the image of the cigarette box to be inspected. It is an unordered and non-repeating defect category representation, used to reflect the range of defect types detected by the second appearance defect detection model. For example, if the second appearance defect detection model detects surface stains, blurred printing, pull thread misalignment, and packaging wrinkles, then the second defect category set is {surface stains, blurred printing, pull thread misalignment, packaging wrinkles}. The second defect quantity set can be a core component of the second defect category information, corresponding one-to-one with the second defect category set. It is the set of the actual number of instances detected for each defect category in the second defect category set, used to quantitatively represent the specific number of each type of second appearance defect detected by the second appearance defect detection model. For example, if the second defect category set is {surface contamination, blurred printing, thread misalignment, packaging wrinkles}, and one instance of each of the four types of defects is detected, then the first defect quantity set is {1, 1, 1, 1}. A second appearance defect can be a single appearance defect on the cigarette box to be inspected, detected by the second appearance defect detection model. For example, if the second appearance defect detection model detects one surface contamination defect, that surface contamination is an independent second appearance defect. The first verification result can be the comparison judgment conclusion of the category dimension in the semantic consensus verification. This refers to the judgment result obtained after comparing the element composition of the first defect category set and the second defect category plan through semantic verification, including both consistent and inconsistent verification results, used to verify whether the two models have the same range of defect type judgments for the cigarette box.The second verification result can be the comparison and judgment conclusion of the quantity dimension in the semantic consensus verification. It refers to the judgment result obtained after comparing the element composition of the first defect quantity set and the second defect quantity set through semantic verification. It includes two results: consistent and inconsistent. It is used to verify whether the two models judge the number of instances of various defects of cigarette boxes the same. The effective comparison of the second verification result requires the first verification result to be consistent as a premise, to ensure that the quantity comparison corresponds one-to-one with the category.
[0065] In one implementation, the first defect category set in the first defect category information and the second defect category set in the second defect category information are first compared element by element to determine whether their defect type composition is completely consistent, and a first verification result is obtained. Then, the first defect quantity set corresponding to the first defect category set in the first defect category information and the second defect quantity set corresponding to the second defect category set in the second defect category information are compared numerically to determine whether the number of various defect instances is completely consistent, and a second verification result is obtained. Further, if both the first and second verification results are consistent, the semantic consensus verification result is determined to be successful. The first defect category set is taken as the target defect category set, and the first defect quantity set corresponding to the first defect category set is taken as the target defect quantity set. The target defect category set and the target defect quantity set are then determined as the target defect category information.
[0066] It should be noted that if either the first or second verification result is inconsistent, the semantic consensus verification result can be determined as failing. The first and second defect detection results corresponding to the cigarette box to be inspected will then be marked as pending review and will not be included in the subsequent spatial consistency verification process. Subsequently, these cigarette boxes to be inspected can be collected into a manual review queue, where quality inspectors will manually determine the actual defect type, quantity, and whether there are any detection errors by combining the image to be processed with the first and second defect detection results. Based on the manual review conclusions, subsequent defect labeling or quality inspection processing will be completed. Simultaneously, these samples that fail verification can be included in the model optimization dataset for subsequent iterative training and accuracy optimization of the first and second appearance defect detection models.
[0067] S340. If the semantic consensus verification result is successful, spatial matching is used to perform spatial consistency verification on the first defect bounding box information and the second defect bounding box information to obtain spatial consistency verification information. The spatial consistency verification information includes the spatial consistency verification result and the target defect bounding box set.
[0068] The spatial matching method can be a specific judgment method used to verify the spatial consistency of the first and second defect bounding box information. Its core is to determine the intersection-union ratio (IUR) of bounding boxes of the same type under the target defect category, filter out defect bounding box pairs whose IUR meets preset filtering conditions, and perform confidence-weighted fusion. This is a means of verifying and optimizing the consistency of the dual-model defect spatial localization results. The spatial consistency verification information can be the comprehensive result data obtained after performing spatial consistency verification. It includes two core parts: the spatial consistency verification result and the target defect bounding box set. It is the output data of the spatial verification stage, providing spatial localization basis for determining the final target defect detection result. The spatial consistency verification result can be a core component of the spatial consistency verification information, representing the judgment conclusion of the spatial verification stage. It includes two results: verification passed and verification failed, used to determine whether the final target defect detection result can be determined based on this verification result. The target defect bounding box set can be a core component of the spatial consistency verification information, referring to the set of optimal defect bounding boxes obtained after confidence-weighted fusion of successfully matched defect bounding box pairs under the target defect category, assuming the spatial consistency verification result is passed.
[0069] Optionally, the first defect bounding box information includes the coordinates of the first bounding box of at least one first appearance defect; the second defect bounding box information includes the coordinates of the second bounding box of at least one second appearance defect; spatial consistency verification is performed on the first and second defect bounding box information using a spatial matching method to obtain spatial consistency verification information, including: for at least one first appearance defect, if at least one third appearance defect with the same defect category as the first appearance defect is identified from at least one second appearance defect, the spatial consistency verification result is determined to be verified as passed; for at least one third appearance defect, the intersection-union ratio (IUGR) information between the first and third appearance defects is determined based on the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the third appearance defect; based on the IUGR information corresponding to at least one third appearance defect, a fourth appearance defect matching the first appearance defect is determined, and a target defect bounding box set is determined based on the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the fourth appearance defect.
[0070] The first bounding box coordinates can represent the spatial location and extent of a single first appearance defect in the image to be processed, and are the basic building blocks of the first defect bounding box information. The first bounding box coordinates can include the vertex coordinates and / or center point coordinates of the first appearance defect's bounding box; for example, the vertex coordinates are typically used as the first bounding box coordinates. The second bounding box coordinates can represent the spatial location and extent of a single second appearance defect in the image to be processed, and are the basic building blocks of the second defect bounding box information. The second bounding box coordinates can include the vertex coordinates and / or center point coordinates of the second appearance defect's bounding box; for example, the vertex coordinates are typically used as the second bounding box coordinates. The third appearance defect can be one or more second appearance defects selected from all second appearance defects that match the defect category of a certain first appearance defect. It is essentially still a second appearance defect, but because it matches the category of the first appearance defect, it becomes a spatial matching candidate. If there is no second appearance defect with a matching category, there is no corresponding third appearance defect; it is the initial screening object for spatial matching of the first appearance defects. The intersection-over-union (IoU) ratio can be a quantitative judgment data characterizing the degree of overlap between the bounding boxes of a first appearance defect and a corresponding third appearance defect. Its core is the IoU value calculated based on the coordinates of the first and second bounding boxes of both defects, serving as the core quantitative basis for judging the spatial matching degree of the defect bounding boxes. A fourth appearance defect can be a single third appearance defect whose IoU information meets preset screening conditions from all third appearance defects corresponding to a first appearance defect; its essence remains a second appearance defect. Optionally, the preset screening conditions include at least one of the following: the IoU information is greater than a preset IoU threshold, which may include 0.5, 0.4, or 0.6, etc.; or at least one maximum value among the IoU information.
[0071] In one implementation, for at least one first appearance defect in the first defect bounding box information, at least one third appearance defect whose defect category matches that of the first appearance defect is first selected from all second appearance defects in the second defect bounding box information. If the category matching is completed, the spatial consistency verification result of the first appearance defect is determined to be successful. Further, for the selected at least one third appearance defect, the intersection-union ratio (IUR) information, representing the degree of spatial overlap, is calculated and determined based on the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the third appearance defect. Further, from all third appearance defects corresponding to the first appearance defect, the third appearance defect with the highest IUR information is selected as the fourth appearance defect that matches the first appearance defect. Further, the first bounding box coordinates of at least one first appearance defect and the second bounding box coordinates of the fourth appearance defects corresponding to each first appearance defect can be integrated into a target defect bounding box set.
[0072] S350. If the spatial consistency verification result is successful, determine the target defect detection result corresponding to the cigarette box to be detected based on the target defect bounding box set and target defect category information.
[0073] The target defect bounding box set includes the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the fourth appearance defect that matches the first appearance defect; the first appearance defect belongs to the first defect detection result; and the fourth appearance defect belongs to the second defect detection result.
[0074] Optionally, based on the target defect bounding box set and the target defect category, the target defect detection result corresponding to the cigarette box to be detected is determined, including: determining the target bounding box information corresponding to the first appearance defect based on the first bounding box coordinates of the first appearance defect, the first confidence level corresponding to the first appearance defect, the second bounding box coordinates of the fourth appearance defect, and the second confidence level corresponding to the fourth appearance defect; and determining the target defect detection result corresponding to the cigarette box to be detected based on the target bounding box information corresponding to the first appearance defect and the target defect category.
[0075] The first confidence score can be the confidence rating of the first appearance defect detection model for a single detected first appearance defect, a numerical value of 0-1. This confidence score corresponds one-to-one with the first bounding box coordinates of the first appearance defect. The closer the value is to 1, the higher the confidence of the model in locating the bounding box and classifying the first appearance defect. It is one of the core weight bases for bounding box fusion calculation. The second confidence score can be the confidence rating of the second appearance defect detection model for a single detected fourth appearance defect, a numerical value of 0-1. Its numerical definition and representation rules are completely consistent with the confidence score of the first appearance defect, and it corresponds one-to-one with the second bounding box coordinates of the fourth appearance defect. It is another core weight base for bounding box fusion calculation. The target bounding box information can be the optimal defect bounding box representation information obtained after fusion calculation for a single first appearance defect, combining its first bounding box coordinates and corresponding first confidence score, and the second bounding box coordinates and corresponding second confidence score of the successfully matched fourth appearance defect. The target bounding box information can be the fused comprehensive quantitative data, which is the final spatial location result of a single first appearance defect in the image to be processed.
[0076] In this embodiment, the determination of the target bounding box information may include at least one of the following: processing the first bounding box coordinates of the first appearance defect, the first confidence level corresponding to the first appearance defect, the second bounding box coordinates of the fourth appearance defect, and the second confidence level corresponding to the fourth appearance defect through a weighted average operation to obtain the target bounding box information; processing the first bounding box coordinates of the first appearance defect, the first confidence level corresponding to the first appearance defect, the second bounding box coordinates of the fourth appearance defect, and the second confidence level corresponding to the fourth appearance defect through an average operation to obtain the target bounding box information; and processing the first bounding box coordinates of the first appearance defect, the first confidence level corresponding to the first appearance defect, the second bounding box coordinates of the fourth appearance defect, and the second confidence level corresponding to the fourth appearance defect through an information fusion model to obtain the output target bounding box information.
[0077] Optionally, based on the first bounding box coordinates of the first appearance defect, the first confidence level corresponding to the first appearance defect, the second bounding box coordinates of the fourth appearance defect, and the second confidence level corresponding to the fourth appearance defect, the target bounding box information corresponding to the first appearance defect is determined, including: determining the product between the first bounding box coordinates and the first confidence level to obtain a first value; and determining the product between the second bounding box coordinates and the second confidence level to obtain a second value; adding the first value and the second value to obtain a third value; adding the first confidence level and the second confidence level to obtain a total confidence level; and determining the ratio between the third value and the total confidence level to obtain the target bounding box information corresponding to the first appearance defect.
[0078] In one implementation, for at least one first appearance defect in the target defect bounding box set, a first value is obtained by multiplying the coordinates of the first bounding box of the first appearance defect with a first confidence level; a second value is obtained by multiplying the coordinates of the second bounding box of a fourth appearance defect matching the first appearance defect with a second confidence level; the first value and the second value are added together to obtain a third value; the first confidence level and the second confidence level are added together to obtain a total confidence level; the ratio between the third value and the total confidence level is determined, and the obtained ratio is used as the target bounding box information corresponding to the first appearance defect. Further, the target bounding box information corresponding to at least one first appearance defect can be used to determine target defect bounding box information, and the target defect bounding box information and target defect category information can be used as the target defect detection result corresponding to the cigarette box to be detected.
[0079] For example, suppose the successfully matched appearance defect pair is ( The target bounding box information can be determined based on the following formula: ;in, Represents any first appearance defect in the target defect bounding box set; Indicates the first appearance defect A matching fourth appearance defect; Indicates the first appearance defect The corresponding first confidence level; Indicates the fourth appearance defect The corresponding second confidence level; Indicates the first appearance defect The corresponding target bounding box information; Indicates multiplication.
[0080] The technical solution provided in this disclosure uses a semantic verification method to perform semantic consensus verification on the first defect category information and the second defect category information to obtain semantic consensus verification information. The semantic consensus verification information includes the semantic consensus verification result and the target defect category information. Further, if the semantic consensus verification result is successful, a spatial matching method is used to perform spatial consistency verification on the first defect bounding box information and the second defect bounding box information to obtain spatial consistency verification information. The spatial consistency verification information includes the spatial consistency verification result and the target defect bounding box set. Further, if the spatial consistency verification result is successful, the target defect detection result corresponding to the cigarette box to be detected is determined based on the target defect bounding box set and the target defect category information. A two-layer verification mechanism, first semantic consensus verification and then spatial consistency verification, is used to verify and fuse the defect category and bounding box information output by the two models layer by layer, accurately determining the target defect detection result of the cigarette box. This significantly improves the accuracy and reliability of defect judgment and effectively reduces the false detection rate and false negative rate of cigarette box appearance defects.
[0081] This disclosure provides an optional embodiment of a method for detecting appearance defects in cigarette boxes, the specific implementation of which can be found in the following embodiments. Technical features that are the same as or similar to those in the above embodiments will not be repeated here. The following describes the method for detecting appearance defects in cigarette boxes.
[0082] First, data acquisition and preprocessing: Images of cigarette packs containing various defect types were collected to construct training samples; the images underwent preprocessing operations such as size normalization, contrast enhancement, and noise reduction. Next, the main detection model was constructed and trained: the main detection model is an improved YOLOv5 model. Its core improvement lies in introducing a multi-scale fusion module with dilated convolutions in the model's neck network. In this improved module, multi-scale dilated convolutions can capture targets of different sizes, improving multi-scale target detection capabilities; adaptive average pooling reduces feature size differences by aggregating spatial information; MLP combined with channel attention can dynamically adjust the importance of each channel, enhancing key features; CBAM (channel + spatial attention) further strengthens salient regions and suppresses background interference; residual connections preserve original features and improve training stability. Training was performed using a manually labeled dataset containing only the overall bounding boxes of the cigarette packs. This design significantly expands the model's receptive field without sacrificing feature map resolution, making it more sensitive to structural defects such as "thread offset" and "packaging wrinkles." Next, an auxiliary detection model is constructed and trained: the auxiliary detection model adopts the YOLOv8n architecture. This model's native anchorless box design and Task-Aligned Assigner label allocation strategy give it higher detection sensitivity for local texture defects such as "surface contamination" and "printing blur." The essential differences in architecture and optimization objectives between the main and auxiliary models are key to ensuring their decision independence and complementarity. Training is performed using a manually labeled dataset containing only the overall bounding boxes of cigarette boxes. Then, parallel inference is performed between the two models: the image to be processed is input into both the trained main and auxiliary detection models, and the two models perform inference independently, outputting their respective detection result sets. Each result includes bounding box coordinates, defect category labels, and confidence scores. Next, two-level verification and high-confidence label generation: this step is the core of the invention, selecting high-confidence consensus results from the independent outputs of the two models and generating the final labels: First level: semantic consensus verification; the defect type sets output by the main and auxiliary models and the number of instances of each type are counted respectively. The image passes verification only if both images have identical defect type sets and the number of instances of all the same category is also identical. Otherwise, the image is marked as "pending review" and excluded from this round of automatic annotation. Second level: Spatial consistency verification and fusion; for defect categories that pass the first level verification, precise spatial location matching is performed. For the defect bounding boxes reported by the main model... Find the box with the largest intersection-union ratio among similar results of the auxiliary model. The conditions for a successful match are: (in, (A preset threshold, for example, 0.5). For all successfully matched defect pairs ( The final bounding boxes are generated using a confidence-weighted fusion strategy. ;in, This represents the confidence score of the main model for the bounding box; This represents the confidence score of the auxiliary model for the bounding box; This represents the final coordinates of the bounding box. This strategy combines the localization advantages of both models. The final category label adopts the consensus-reached category; This indicates multiplication. Afterwards, the annotation file is output: the fused target defect detection results are written into an annotation file in a standard format (such as YOLO format).
[0083] This technical solution applies "dual-model collaborative verification" to the "automatic annotation generation" scenario. Through differentiated model architecture design and a refined two-level verification fusion algorithm, it solves the fundamental problem of unreliable automatic annotation using a single model. Furthermore, this invention provides a complete and feasible automated annotation workflow. Experimental verification shows that the accuracy of automatically annotated data generated by this method is significantly improved, greatly reducing the proportion of images requiring manual review. This high-quality annotated data can be directly used for model training, and the performance of the new model trained using it is comparable to that of a model using all manually annotated data, greatly saving time and economic costs. This invention achieves a leap from "manual annotation" or "unreliable automatic annotation" to "highly reliable automatic annotation," effectively freeing manual labor from heavy and repetitive annotation work, allowing them to focus on handling a few difficult cases, and providing strong data support for the rapid deployment and iteration of industrial vision inspection systems.
[0084] Figure 4 This is a schematic diagram of a device for detecting defects in the appearance of cigarette boxes, provided in an embodiment of this disclosure. Figure 4 As shown, the cigarette box appearance defect detection device includes: an image acquisition module 410, a cigarette box defect detection module 420, and a target detection result determination module 430. The image acquisition module 410 is used to acquire an image to be processed, which includes the cigarette box to be detected. The cigarette box defect detection module 420 is used to determine a first defect detection result corresponding to the cigarette box to be detected based on a first appearance defect detection model and the image to be processed, and to determine a second defect detection result corresponding to the cigarette box to be detected based on a second appearance defect detection model and the image to be processed; wherein the model structures of the first appearance defect detection model and the second appearance defect detection model are different. The target detection result determination module 430 is used to determine a target defect detection result corresponding to the cigarette box to be detected based on the first defect detection result and the second defect detection result, and to process the cigarette box to be detected according to the target defect detection result.
[0085] The technical solution of this disclosure, by acquiring an image to be processed, including a cigarette box to be inspected, provides a unified input data source for subsequent dual-model defect detection, ensuring the completeness and quantifiability of the visual information of the cigarette box to be inspected, which is the fundamental premise of the entire defect detection process. Furthermore, by determining a first defect detection result corresponding to the cigarette box to be inspected based on a first appearance defect detection model and the image to be processed, and by determining a second defect detection result corresponding to the cigarette box to be inspected based on a second appearance defect detection model and the image to be processed, two appearance defect detection models with different structures are used to detect cigarette box defects from different feature dimensions, compensating for the detection blind spots of a single model and improving the comprehensiveness and accuracy of defect identification. Furthermore, by determining a target defect detection result corresponding to the cigarette box to be inspected based on the first and second defect detection results, and processing the cigarette box to be inspected according to the target defect detection result, the dual-model detection results are fused to output a unified and accurate target defect detection result. Based on this, quality inspection processing is performed on the cigarette box, ensuring the reliability of defect judgment and achieving efficient and accurate handling of cigarette box defects. The technical solution of this disclosure solves the problems of high false detection rate and false negative rate in the prior art, as well as the inability to meet the requirements of high-precision quality inspection of cigarette box appearance. By using dual models with different structures to detect cigarette box defects from different feature dimensions and fusing the detection results, the false detection rate and false negative rate of cigarette box appearance defects are effectively reduced, realizing high-precision automated quality inspection of cigarette box appearance, meeting the requirements of cigarette production for high-precision quality inspection of cigarette box appearance, and improving the quality inspection efficiency and quality control level of the production line.
[0086] In some embodiments of this disclosure, optionally, the first appearance defect detection model includes an input network, a backbone network, a neck network, and an output network; the cigarette box defect detection module 420 includes: an image tensor determination submodule, a basic feature map determination submodule, a fusion feature map determination submodule, and a first detection result determination submodule. The image tensor determination submodule is used to receive the image to be processed through the input network and perform data enhancement on the image to be processed to obtain the image tensor corresponding to the image to be processed; the basic feature map determination submodule is used to receive the image tensor to be processed through the backbone network and perform downsampling and feature enhancement on the image tensor to be processed to obtain basic feature maps of the image to be processed at multiple scales; the fusion feature map determination submodule is used to receive multiple basic feature maps through the neck network, perform feature fusion and enhancement on the multiple basic feature maps according to the neck network, and obtain a multi-scale fusion feature map; the first detection result determination submodule is used to receive the multi-scale fusion map through the output network, and perform detection on the first multi-scale fusion map according to the output network to obtain the output first defect detection result corresponding to the cigarette box to be detected.
[0087] In some embodiments of this disclosure, optionally, the neck network includes a multi-scale dilated convolution module, an average pooling module, a first attention module, a second attention module, and a residual module; the fusion feature map determination submodule includes: a first feature map determination unit, a second feature map determination unit, a third feature map determination unit, a fourth feature map determination unit, and a fusion feature map determination unit. Specifically, the first feature map determination unit is used to perform parallel convolution on each basic feature map according to the multi-scale dilated convolution module to obtain multiple first feature maps; the second feature map determination unit is used to downsample each first feature map according to the average pooling module to obtain multiple second feature maps; the third feature map determination unit is used to perform defect feature enhancement on each second feature map according to the first attention module to obtain multiple third feature maps; the fourth feature map determination unit is used to perform defect feature localization on each third feature map according to the second attention module to obtain multiple fourth feature maps; and the fusion feature map determination unit is used to perform residual concatenation on the multiple fourth feature maps and multiple basic feature maps according to the residual module to obtain a multi-scale fusion feature map.
[0088] In some embodiments of this disclosure, optionally, the first defect detection result includes first defect category information and first defect bounding box information; the second defect detection result includes second defect category information and second defect bounding box information; the target detection result determination module 430 includes: a semantic consensus verification unit, a spatial consistency verification unit, and a defect detection result determination unit. The semantic consensus verification unit is used to perform semantic consensus verification on the first defect category information and the second defect category information using a semantic verification method to obtain semantic consensus verification information; wherein, the semantic consensus verification information includes a semantic consensus verification result and target defect category information; the spatial consistency verification unit is used to perform spatial consistency verification on the first defect bounding box information and the second defect bounding box information using a spatial matching method when the semantic consensus verification result is passed, to obtain spatial consistency verification information; the spatial consistency verification information includes a spatial consistency verification result and a set of target defect bounding boxes; the defect detection result determination unit is used to determine the target defect detection result corresponding to the cigarette box to be detected based on the set of target defect bounding boxes and the target defect category information when the spatial consistency verification result is passed.
[0089] In some embodiments of this disclosure, optionally, the first defect category information includes a first defect category set and a first defect quantity set corresponding to the first defect category set; the second defect category information includes a second defect category set and a second defect quantity set corresponding to the second defect category set; the semantic consensus verification unit is specifically used to determine whether the first defect category set and the second defect category set are consistent, to obtain a first verification result, and to determine whether the second defect quantity set is consistent with the second defect quantity set, to obtain a second verification result; if both the first verification result and the second verification result are consistent, the semantic consensus verification result is determined to be verified as passed, the first defect category set is taken as the target defect category set, the first defect quantity set is taken as the target defect quantity set, and the target defect category set and the target defect quantity set are determined as target defect category information.
[0090] In some embodiments of this disclosure, optionally, the first defect bounding box information includes the coordinates of the first bounding box of at least one first appearance defect; the second defect bounding box information includes the coordinates of the second bounding box of at least one second appearance defect; the spatial consistency verification unit is specifically used to determine that the spatial consistency verification result is passed when at least one third appearance defect is determined from at least one second appearance defect that is consistent with the defect category of the first appearance defect; for at least one third appearance defect, the cross-union ratio information between the first appearance defect and the third appearance defect is determined based on the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the third appearance defect; based on the cross-union ratio information corresponding to at least one third appearance defect, a fourth appearance defect that matches the first appearance defect is determined, and a target defect bounding box set is determined based on the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the fourth appearance defect.
[0091] In some embodiments of this disclosure, optionally, the target defect bounding box set includes the first bounding box coordinates of a first appearance defect and the second bounding box coordinates of a fourth appearance defect that matches the first appearance defect; the first appearance defect belongs to the first defect detection result; the fourth appearance defect belongs to the second defect detection result; the defect detection result determination unit is specifically used to determine the target bounding box information corresponding to the first appearance defect based on the first bounding box coordinates of the first appearance defect, the confidence level corresponding to the first appearance defect, the second bounding box coordinates of the fourth appearance defect, and the confidence level corresponding to the fourth appearance defect; determine the target bounding box information corresponding to the first appearance defect as the target defect bounding box information, and determine the target defect bounding box information and the target defect category information as the target defect detection result corresponding to the cigarette box to be detected.
[0092] The cigarette box appearance defect detection device provided in this embodiment can execute the cigarette box appearance defect detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0093] It is worth noting that the various units and modules included in the above-mentioned cigarette box appearance defect detection device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0094] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0095] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0096] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0097] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the cigarette box appearance defect detection method.
[0098] In some embodiments, the cigarette box appearance defect detection method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the cigarette box appearance defect detection method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the cigarette box appearance defect detection method by any other suitable means (e.g., by means of firmware).
[0099] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0100] Computer programs used to implement the cigarette box appearance defect detection method of this disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs can be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0101] This disclosure provides a computer-readable storage medium storing computer instructions for causing a processor to execute a method for detecting defects in the appearance of a cigarette box. The method includes: acquiring an image to be processed, wherein the image to be processed includes a cigarette box to be detected; determining a first defect detection result corresponding to the cigarette box to be detected based on a first appearance defect detection model and the image to be processed; and determining a second defect detection result corresponding to the cigarette box to be detected based on a second appearance defect detection model and the image to be processed. The first appearance defect detection model is used to detect structural defects on the cigarette box to be detected; the second appearance defect detection model is used to detect local texture defects on the cigarette box to be detected; and determining a target defect detection result corresponding to the cigarette box to be detected based on the first defect detection result and the second defect detection result, and processing the cigarette box to be detected based on the target defect detection result.
[0102] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0104] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0105] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0106] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of embodiments of this disclosure.
[0107] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements a media resource retrieval method according to any embodiment of this disclosure.
[0108] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0109] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0110] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for detecting appearance defects in cigarette boxes, characterized in that, include: Acquire an image to be processed, wherein the image to be processed includes a cigarette box to be detected; Based on the first appearance defect detection model and the image to be processed, a first defect detection result corresponding to the cigarette box to be detected is determined, and based on the second appearance defect detection model and the image to be processed, a second defect detection result corresponding to the cigarette box to be detected is determined; wherein, the model structures of the first appearance defect detection model and the second appearance defect detection model are different; Based on the first defect detection result and the second defect detection result, a target defect detection result corresponding to the cigarette box to be tested is determined, and the cigarette box to be tested is processed according to the target defect detection result.
2. The method for detecting appearance defects in cigarette boxes according to claim 1, characterized in that, The first appearance defect detection model includes an input network, a backbone network, a neck network, and an output network; the step of determining the first defect detection result corresponding to the cigarette box to be detected based on the first appearance defect detection model and the image to be processed includes: The image to be processed is received through the input network, and data augmentation is performed on the image to be processed to obtain the image tensor corresponding to the image to be processed. The backbone network receives the tensor of the image to be processed, and downsamples and enhances the tensor of the image to be processed to obtain the basic feature map of the image to be processed at multiple scales. Multiple basic feature maps are received through the neck network, and feature fusion and enhancement are performed on the multiple basic feature maps according to the neck network to obtain a multi-scale fused feature map; The multi-scale fusion image is received through the output network, and the multi-scale fusion image is detected to obtain the first defect detection result corresponding to the cigarette box to be detected.
3. The method for detecting appearance defects in cigarette boxes according to claim 2, characterized in that, The neck network includes a multi-scale dilated convolution module, an average pooling module, a first attention module, a second attention module, and a residual module; the step of fusing and enhancing multiple base feature maps based on the neck network to obtain a multi-scale fused feature map includes: The multi-scale dilated convolution module performs parallel convolution on each of the basic feature maps to obtain multiple first feature maps. The average pooling module downsamples each of the first feature maps to obtain multiple second feature maps. Each second feature map is enhanced with defect features based on the first attention module to obtain multiple third feature maps; Based on the second attention module, defect features are located for each of the third feature maps to obtain multiple fourth feature maps; The residual module performs residual connections on multiple fourth feature maps and multiple basic feature maps to obtain a multi-scale fused feature map.
4. The method for detecting appearance defects in cigarette boxes according to claim 1, characterized in that, The first defect detection result includes first defect category information and first defect bounding box information; The second defect detection result includes second defect category information and second defect bounding box information; The step of determining the target defect detection result corresponding to the cigarette box to be inspected based on the first defect detection result and the second defect detection result includes: The first defect category information and the second defect category information are semantically validated using a semantic validation method to obtain semantic consensus validation information; wherein, the semantic consensus validation information includes semantic consensus validation result and target defect category information; If the semantic consensus verification result is passed, spatial matching is used to perform spatial consistency verification on the first defect bounding box information and the second defect bounding box information to obtain spatial consistency verification information; the spatial consistency verification information includes spatial consistency verification result and target defect bounding box set. If the spatial consistency verification result is successful, the target defect detection result corresponding to the cigarette box to be detected is determined based on the target defect bounding box set and the target defect category information.
5. The method for detecting appearance defects in cigarette boxes according to claim 4, characterized in that, The first defect category information includes a first defect category set and a first defect quantity set corresponding to the first defect category set; The second defect category information includes a second defect category set and a second defect quantity set corresponding to the second defect category set; The semantic consensus verification of the first defect category information and the second defect category information is performed using a semantic verification method to obtain semantic consensus verification information, including: Determine whether the first defect category set is consistent with the second defect category set to obtain a first verification result; and determine whether the second defect quantity set is consistent with the second defect quantity set to obtain a second verification result. If the first verification result and the second verification result are consistent, the semantic consensus verification result is determined to be verified as passed. The first defect category set is taken as the target defect category set, the first defect quantity set is taken as the target defect quantity set, and the target defect category set and the target defect quantity set are determined as target defect category information.
6. The method for detecting appearance defects in cigarette boxes according to claim 4, characterized in that, The first defect bounding box information includes the coordinates of the first bounding box of at least one first appearance defect; the second defect bounding box information includes the coordinates of the second bounding box of at least one second appearance defect. The spatial consistency verification of the first defect bounding box information and the second defect bounding box information is performed using a spatial matching method to obtain spatial consistency verification information, including: In the case where at least one of the first appearance defects is identified from at least one of the second appearance defects and the defect category of the first appearance defect is consistent, the spatial consistency verification result is determined to be a successful verification. For at least one of the third appearance defects, the intersection-over-union ratio (IoU) between the first appearance defect and the third appearance defect is determined based on the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the third appearance defect. Based on the intersection-union ratio information corresponding to at least one of the third appearance defects, a fourth appearance defect that matches the first appearance defect is determined, and a set of target defect bounding boxes is determined based on the first bounding box coordinates of the first appearance defect and the second bounding box coordinates of the fourth appearance defect.
7. The method for detecting appearance defects in cigarette boxes according to claim 4, characterized in that, The target defect bounding box set includes the first bounding box coordinates of a first appearance defect and the second bounding box coordinates of a fourth appearance defect that matches the first appearance defect; the first appearance defect belongs to the first defect detection result; the fourth appearance defect belongs to the second defect detection result; The step of determining the target defect detection result corresponding to the cigarette box to be detected based on the target defect bounding box set and the target defect category information includes: Based on the first bounding box coordinates of the first appearance defect, the first confidence level corresponding to the first appearance defect, the second bounding box coordinates of the fourth appearance defect, and the second confidence level corresponding to the fourth appearance defect, the target bounding box information corresponding to the first appearance defect is determined. The target bounding box information corresponding to the first appearance defect is determined as the target defect bounding box information, and the target defect bounding box information and the target defect category information are determined as the target defect detection result corresponding to the cigarette box to be detected.
8. A device for detecting defects in the appearance of cigarette boxes, characterized in that, include: An image acquisition module is used to acquire an image to be processed, wherein the image to be processed includes a cigarette box to be detected; A cigarette box defect detection module is used to determine a first defect detection result corresponding to the cigarette box to be detected based on a first appearance defect detection model and the image to be processed, and to determine a second defect detection result corresponding to the cigarette box to be detected based on a second appearance defect detection model and the image to be processed; wherein the model structures of the first appearance defect detection model and the second appearance defect detection model are different. The target detection result determination module is used to determine the target defect detection result corresponding to the cigarette box to be detected based on the first defect detection result and the second defect detection result, and to process the cigarette box to be detected based on the target defect detection result.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the cigarette box appearance defect detection method as described in any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method for detecting defects in the appearance of cigarette boxes according to any one of claims 1-7.