Defect detection method based on lightweight convolutional neural network and continuous learning
Through lightweight convolutional neural network and continuous learning framework, the automation problem of carton printing defect detection is solved, efficient and accurate defect identification and adaptive detection are achieved, and the quality and efficiency of carton production are improved.
Patent Information
- Application Number
- CN202510223294.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional carton printing defect detection relies on manual inspection, which has inconsistent inspection results, is expensive and difficult to meet the needs of modern production of high speed and high precision. The surface defect detection of complex carton printing has not yet been fully automated.
The lightweight convolutional neural network and a continuous learning framework are adopted, combining deep separable convolution, reverse residual blocks, coordinate attention mechanisms, and bidirectional weighted feature pyramid networks for feature extraction and fusion. The continuous learning module dynamically updates the classifier to adapt to new defect types.
It improves the efficiency and accuracy of defect detection in carton production, reduces manual intervention, reduces costs, ensures system adaptability and stability, and supports the seamless operation of the production line and the consistency of product quality.
Smart Images

Figure CN120339163A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of defect detection, and particularly relates to a defect detection method based on a lightweight convolutional neural network and continuous learning. Background Art
[0002] As an important part of the modern packaging industry, cartons play a crucial role in protecting products, facilitating transportation, and enhancing brand image. Whether it is daily consumer goods, electronic products, or food and drugs, carton packaging can provide effective physical protection to prevent damage to products caused by external environmental factors such as impact, moisture, and light. At the same time, the printing quality of cartons not only directly affects the aesthetics of the packaging and the market image of the product, but also relates to the functionality of the product during the logistics and sales processes, such as barcode recognition and product information reading. Therefore, the quality of carton printing not only affects the safety and integrity of the product, but also has an important impact on the customer's brand perception and purchasing decisions.
[0003] Traditional carton printing defect detection mainly relies on manual inspection. Although manual inspection is relatively flexible, due to subjective factors, it is prone to inconsistent detection results. In addition, with the continuous improvement of production efficiency, relying on manual inspection is not only costly, but also difficult to meet the requirements of high-speed and high-precision modern production. The defects of manual inspection also include dependence on environmental factors such as light and worker fatigue, which may directly affect the accuracy and efficiency of detection.
[0004] In the industrial field, significant progress has been made in automated defect detection technology. Especially surface defect detection technology based on machine vision and deep learning can efficiently and accurately analyze and detect the product surface. However, in the field of carton printing, although surface defect detection technology is widely used in materials such as metals and plastics, the automated detection of complex carton printing surfaces has not been fully realized. The defects of carton printing include printing deviation, ghosting, missing printing, blurring, etc. These defects often have high requirements for the accuracy of color and patterns, and the defect size is small and the shape is complex, which poses a huge challenge to automated detection.
[0005] To address these problems, combining advanced convolutional neural network (CNN) technology, especially lightweight deep learning networks and continuous learning frameworks, can provide an efficient automated defect detection solution. Lightweight CNNs perform excellently in feature extraction and pattern recognition, and can significantly reduce the demand for computing resources while maintaining high accuracy. In addition, continuous learning technology enables the system to self-optimize and adjust in the face of continuously updated data and newly emerging defect types, ensuring that it can adapt to different types of defects during the detection process without the need to frequently retrain the entire model.
[0006] Therefore, the automated carton printing defect detection technology that combines lightweight CNN and continuous learning can not only significantly improve the detection efficiency and accuracy, reduce the subjective errors of manual detection, but also lower the production cost and enhance the overall quality control ability of the carton printing industry. The introduction of this technology is of great significance for promoting the process of intelligent manufacturing industry and provides a more intelligent, convenient and efficient solution for carton printing quality detection. Summary of the Invention
[0007] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a defect detection method based on lightweight convolutional neural network and continuous learning, and an automatic detection system based on continuous learning, which greatly improves the defect detection efficiency in the production of corrugated cartons.
[0008] To achieve the above purpose, the present invention adopts the following technical solutions:
[0009] In the first aspect, the present invention provides a defect detection method based on lightweight convolutional neural network and continuous learning, including the following steps:
[0010] Collect corrugated carton images and perform preprocessing;
[0011] Construct a corrugated carton defect detection model and perform model training. The corrugated carton defect detection model includes a lightweight convolutional neural network module, a feature fusion module and a continuous learning module; the lightweight convolutional neural network module extracts features based on depthwise separable convolution, inverted residual block and coordinate attention mechanism; the feature fusion module performs multi-scale extraction and cross-layer fusion on feature maps from different levels through a new type of bidirectional weighted feature pyramid network; the continuous learning module detects new defect types and automatically marks new samples through the input fused feature maps, and dynamically updates the classifier;
[0012] Perform defect detection on the corrugated carton image to be detected based on the trained corrugated carton defect detection model.
[0013] As a preferred technical solution, the step of collecting corrugated carton images and performing preprocessing is specifically:
[0014] Collect corrugated carton images through image acquisition devices installed on the production line.
[0015] As a preferred technical solution, the lightweight convolutional neural network module includes a depthwise separable convolution module, an inverted residual block module and a coordinate attention mechanism module, specifically:
[0016] The lightweight CNN extracts the basic features of the image through the depthwise separable convolution module to obtain the feature map I1 of each channel;
[0017] Pointwise convolution uses a convolution kernel of a set size to integrate feature maps of different channels to form a new feature map I2; the integration is to perform operations on each pixel only with all channels at the position of that pixel;
[0018] The new feature map is input into the inverted residual block module for processing. Specifically, the channels are expanded through pointwise convolution to obtain the expanded feature map I3. The feature map I3 is independently calculated within each channel through depthwise convolution operation. Finally, the expanded number of channels is compressed back to the original scale through pointwise convolution to obtain the feature map I4;
[0019] The feature map I4 is input into the coordinate attention mechanism for processing. The global information of the feature map I4 in the horizontal and vertical directions is obtained through pooling operations to obtain the feature vector in the horizontal direction and the feature vector in the vertical direction;
[0020] The feature vector in the horizontal direction and the feature vector in the vertical direction are subjected to feature encoding, feature fusion, and feature enhancement processing to obtain feature maps I5 of different scales that maintain the original size.
[0021] As a preferred technical solution, the feature encoding, feature fusion, and feature enhancement processing of the feature vector in the horizontal direction and the feature vector in the vertical direction to obtain feature maps I5 of different scales that maintain the original size are specifically as follows:
[0022] The coordinate attention mechanism module is used to perform one-dimensional global pooling on the feature map I4 in the horizontal and vertical directions respectively. Global pooling is performed on each column in the vertical direction to obtain a first feature vector representing the information of the entire column, and global pooling is performed on each row in the horizontal direction to obtain a second feature vector representing the information of the entire row;
[0023] The direction information in the first feature vector and the second feature vector is subjected to feature encoding through a shared convolutional layer to generate a first intermediate feature and a second intermediate feature with direction sensitivity;
[0024] The first intermediate feature and the second intermediate feature are processed through a convolutional layer to respectively generate a first attention map in the horizontal direction and a second attention map in the vertical direction;
[0025] The first attention map and the second attention map are multiplied element-wise with the feature map I4 to respectively weight the information of the feature map I4 in the horizontal and vertical directions, so as to enhance the features in the horizontal direction and the vertical direction, making the defects in a specific area or a specific area receive higher attention.
[0026] As a preferred technical solution, in the feature fusion module, different scale feature maps can be weighted and fused through a two-way path by a new type of bidirectional weighted feature pyramid network. Specifically:
[0027] Extract the feature map I6 with low resolution and high semantic information and the feature map I7 with high resolution and low semantic information from the feature map I5;
[0028] Bottom-up feature fusion: Upsample the feature map I6 to obtain the feature map I8, adjust the number of channels of the feature map I8 using pointwise convolution to obtain the feature map I9. In the weighted feature fusion process, control the contribution of the feature map I9 in the fusion by assigning a learnable weight, and use the weighted formula to obtain the fused feature map I 10 ;
[0029] Bottom-up feature fusion: Downsample the feature map I7 to obtain the feature map I 11 and adjust the number of channels of the feature map I8 using pointwise convolution to obtain the feature map I 12 and use the weighted formula to obtain the fused feature map I 13 ;
[0030] After the top-down and bottom-up two-way weighted fusion, the output feature map maintains the input spatial size. The output fused feature map not only retains the high-level semantic information but also combines the low-level detailed information, enabling the network to effectively detect defects of different scales on the corrugated cardboard box.
[0031] As a preferred technical solution, in the continuous learning module, the fused feature map I 10 and the feature map I 13 pass through the classifier to output the probability that each sample belongs to different defect classifications, and determine the defect category to which the sample belongs by maximizing the classification probability.
[0032] To detect whether the detection model encounters a new type of defect, use the Mahalanobis distance to measure the difference between the new input sample and the known defect categories, and judge whether the sample belongs to the existing category or a new category by calculating the distance between the sample and the center of each category;
[0033] After detecting a new defect and manually annotating it, the continuous learning mechanism will adaptively update the model. By updating the classifier parameters of the model and the parameters of the new defect classifier, gradually learn the features of the new defect, so that the model can not only handle new types of defects but also not forget the existing defect categories.
[0034] As a preferred technical solution, the continuous learning process uses a joint loss function that includes the cross-entropy loss and the new defect detection loss. The goal of the joint loss function is to maximize the probability of correct classification and at the same time enable the model to distinguish new defects from old defects.
[0035] In a second aspect, the present invention provides a defect detection system based on a lightweight convolutional neural network and continual learning, which is applied to the defect detection method based on a lightweight convolutional neural network and continual learning, and includes an image acquisition module, a defect detection model construction module, and a defect detection module;
[0036] The image acquisition module is used to acquire corrugated box images and perform preprocessing;
[0037] The defect detection model construction module is used to construct a corrugated box defect detection model and perform model training. The corrugated box defect detection model includes a lightweight convolutional neural network module, a feature fusion module, and a continual learning module; the lightweight convolutional neural network module performs feature extraction based on depthwise separable convolution, inverted residual blocks, and coordinate attention mechanisms; the feature fusion module performs multi-scale extraction and cross-layer fusion on feature maps from different levels through a novel bidirectional weighted feature pyramid network; the continual learning module detects new defect types and automatically labels new samples through the input fused feature maps, and dynamically updates the classifier;
[0038] The defect detection module is used to perform defect detection on the corrugated box image to be detected based on the trained corrugated box defect detection model.
[0039] In a third aspect, the present invention provides an electronic device, which includes:
[0040] At least one processor; and,
[0041] A memory communicatively connected to the at least one processor; wherein,
[0042] The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the defect detection method based on a lightweight convolutional neural network and continual learning.
[0043] In a fourth aspect, the present invention provides a computer-readable storage medium storing a program, which when executed by a processor, implements the defect detection method based on a lightweight convolutional neural network and continual learning.
[0044] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0045] The present invention proposes an innovative automatic real-time defect detection framework based on a continual learning mechanism, which is particularly suitable for the carton production environment with a large amount of data. This framework not only solves the challenges of traditional methods in terms of carton defect classification efficiency and accuracy, but also further improves the quality and efficiency of the entire production through its unique automation and intelligence characteristics.
[0046] First, the present invention realizes the efficient and accurate automatic identification of various defects occurring in the carton production process, greatly improving the detection speed and accuracy. By adopting an advanced continuous learning algorithm, the system can continuously learn and optimize its detection ability from newly generated data without frequently relying on manual intervention or historical data updates to adjust the classification rules. This self-evolving ability ensures that the system can maintain a high degree of adaptability and stability even in the face of rapidly changing production conditions.
[0047] Secondly, the present invention significantly reduces the dependence on a large amount of historical data and the need for manual operations, which not only reduces labor costs but also reduces quality problems caused by human errors. In addition, the real-time processing ability of the system ensures that any potential defects can be immediately detected and processed, thus avoiding the risk of defective products flowing into the market and enhancing the market competitiveness of the products.
[0048] Finally, by integrating this intelligent detection framework, the carton manufacturing process becomes more intelligent and efficient. It supports the seamless operation of the production line while ensuring the consistency and reliability of the final product quality. Therefore, the present invention not only promotes the improvement of production efficiency but also lays a foundation for realizing a more refined and intelligent production management mode.
[0049] In summary, the present invention provides a complete solution aiming to improve the efficiency and accuracy of defect detection in the carton production process, thereby enhancing the overall quality of the products, reducing resource waste, and promoting the industry to develop towards higher quality standards. Through the above technical solutions, the present invention has brought an innovative progress to the carton manufacturing industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 is a flowchart of the defect detection method based on a lightweight convolutional neural network and continuous learning according to an embodiment of the present invention;
[0052] Figure 2 is a schematic diagram of a bidirectional weighted pyramid network (BWFPN) in an embodiment of the present invention;
[0053] Figure 3 is a schematic diagram of an inverted residual block in an embodiment of the present invention;
[0054] Figure 4Schematic diagram of the coordinate attention mechanism (CA) in the embodiments of the present invention;
[0055] Figure 5 Framework for automatically detecting new defect types in the embodiments of the present invention;
[0056] Figure 6 Block diagram of the defect detection system based on the lightweight convolutional neural network and continuous learning in the embodiments of the present invention;
[0057] Figure 7 Structural diagram of the electronic device in the embodiments of the present invention. Detailed implementation manners
[0058] In order to enable those skilled in the art of this technology to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of this application.
[0059] Referring to "embodiments" in this application means that the specific features, structures or characteristics described in conjunction with the embodiments may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.
[0060] As Figure 1 shown, the defect detection method based on the lightweight convolutional neural network and continuous learning in this embodiment includes the following steps:
[0061] S1. Collect corrugated box images and perform preprocessing;
[0062] As an indispensable part of the modern logistics and packaging industries, the quality monitoring of corrugated boxes is particularly important. In order to ensure that the quality of corrugated boxes meets the standards, image acquisition technology is widely used in the production process. This process mainly relies on industrial cameras or laser scanning devices to complete. These precision devices are usually installed at key positions on the production line to perform real-time and non-destructive detection on each passing corrugated box.
[0063] Industrial cameras stand out for their high resolution and high frame rate, enabling them to capture extremely subtle defects on the surface of corrugated cardboard boxes even when moving rapidly on the production line. Such cameras can not only identify obvious flaws like scratches and cracks but also detect more concealed problems such as minor color deviations or inconsistent material thickness. Additionally, laser scanning devices construct a three-dimensional model of the object's surface by emitting laser beams and receiving reflected light, having unique advantages in detecting defects on three-dimensional or uneven surfaces.
[0064] Once the images of corrugated cardboard boxes are captured, the next step is to analyze them using advanced image processing techniques. Currently, lightweight convolutional neural networks (CNNs) have become the preferred tool for handling such problems. These networks are designed to be "lightweight" enough to meet the real-time requirements on the production line while maintaining accuracy. The corrugated cardboard box images input into the CNN may contain various potential surface defects, such as indentations, blurred printing, ghosting, missing prints, and color blurring. Through learning from a large number of labeled samples, the CNN can learn to automatically identify these defects and make accurate judgments on new corrugated cardboard box images.
[0065] The application of this technology has greatly improved the quality and efficiency of corrugated cardboard box production, reducing the omissions and errors that may occur in manual inspections. Moreover, it provides data support for production enterprises, helping them better understand the weak links in the production process, thereby taking measures for improvement and ultimately achieving the optimization of the production process and cost control. With the continuous progress of artificial intelligence and machine vision technologies, we can expect more efficient and accurate methods to emerge in the future, further promoting the innovative development of the packaging industry.
[0066] S2. Construct a corrugated cardboard box defect detection model and conduct model training. The corrugated cardboard box defect detection model includes a lightweight convolutional neural network module 101, a feature fusion module 102, and a continuous learning module 103.
[0067] In the corrugated cardboard box defect detection task, please refer to Figure 1 again. The above three modules cooperate closely to jointly improve the detection efficiency and adaptability.
[0068] First, the lightweight convolutional neural network module 101 achieves efficient feature extraction through the combination of depthwise separable convolution, inverted residual blocks, and coordinate attention mechanisms. The core goal of the lightweight convolutional neural network module is to reduce the number of network parameters and computational complexity, ensuring that defect features of corrugated cardboard boxes can be extracted quickly and accurately in scenarios with high real-time requirements. The multi-scale feature maps extracted by this module provide the basis for subsequent feature fusion.
[0069] Next, the feature fusion module 102 performs multi-scale extraction and cross-layer fusion on the feature maps from different levels through a novel bidirectional weighted feature pyramid network (BWFPN). This step is crucial for detecting various defects of different sizes, from fine scratches to large-area tears. The feature fusion module can capture and represent features of different scales, ensuring that defects are comprehensively identified. The fused feature maps further enhance the network's expressive ability, enabling it to handle complex and diverse defects.
[0070] Finally, the continuous learning module 103 detects new defect types and automatically labels new samples through the input fused feature maps, dynamically updating the classifier. This module receives the enhanced feature maps from the feature fusion module, analyzes and adapts to newly emerging defects in real time, ensuring that the network's classifier is continuously updated without the need to retrain all the data. Through continuous learning, the model can adapt to new defects, reduce storage and computational overhead, and achieve long-term efficient detection capabilities.
[0071] In the present invention, the three modules cooperate with each other. The lightweight convolutional neural network provides preliminary feature extraction, the feature fusion module further optimizes the feature expression, and the continuous learning module ensures that the model can adapt to new defect types at any time, forming a complete closed-loop for defect detection.
[0072] S2.1. The lightweight convolutional neural network module includes a depthwise separable convolution module, an inverted residual block module, and a coordinate attention mechanism module. Such a design significantly reduces the computational complexity of the model and maintains strong feature extraction capabilities, as follows:
[0073] S2.1.1. The lightweight CNN extracts the basic features of the image through the depthwise separable convolution module; in standard convolution, the convolution kernel calculates between all channels of the image. However, in depth convolution, each convolution kernel acts only on a single channel of the input image and does not span channels. Assume the size of the input corrugated cardboard image I0 is H×W×C in , where H is the height of the image, W is the width, and C in is the number of input channels (usually 3 channels for RGB). In depth convolution, a separate convolution kernel (such as a 3×3 convolution kernel) is used on each channel to extract features, which means that each channel will generate a new feature map separately
[0074] After depth convolution, the image still maintains the dimension of H×W×C in , but the features within each channel change, extracting more detailed features in that channel, such as lines, textures, dents, etc. on the cardboard box.
[0075] Although the feature maps of each channel have been extracted, there is no fusion between these feature maps yet. Next, pointwise convolution is responsible for integrating information between channels.
[0076] Pointwise convolution uses C out convolution kernels of size 1×1 (i.e., each pixel only operates with all channels at that position) to integrate information from different channels. This means that the network can combine the features from each channel to form a new feature map where C out is the number of output channels.
[0077] Assume the convolution kernel size is k×k, the computational complexity of standard convolution is H×W×C in ×k 2 ×C out . Through the decomposition of depthwise separable convolution: the computational complexity of depthwise convolution is H×W×C in ×k 2 , and the computational complexity of pointwise convolution is H×W×C in ×C out . The computational complexity of this decomposition is much lower than that of standard convolution, especially when the convolution kernel is large or the number of channels is large, the advantage is more obvious. Despite the reduction in computational complexity, depthwise separable convolution still performs well in feature extraction ability. Depthwise convolution ensures local feature extraction within each channel, while pointwise convolution maintains the integration and learning ability of global information through cross-channel information fusion. This enables the network to reduce the computational overhead without significantly sacrificing the performance of the model. After this process, the input corrugated cardboard image will be transformed into a feature map containing more detailed information, helping the network to better identify subtle defects on the surface.
[0078] S2.1.2. After depthwise separable convolution, we obtain a feature map Then it will continue to enter the inverted residual block. As Figure 2 shown, compared with the traditional residual block, the biggest difference in the design of the second inverted residual block module is that its processing order is "reversed". Specifically, the number of channels is expanded by using 1×1 pointwise convolution. This operation increases the representational ability of the network, enabling subsequent convolutions to extract information from more channels. The number of input channels is C out , and the expanded number of channels can be t×C out (where t is the expansion ratio coefficient). The expanded feature map will undergo a depthwise convolution operation. Depthwise convolution only performs calculations independently within each channel and does not operate across channels. Since it does not operate across channels, this greatly reduces the computational complexity but still can well extract local features. Finally, through 1×1 pointwise convolution, the expanded number of channels is compressed back to the original scale feature map This part can further improve the network's ability to extract detailed features in the cardboard box image while maintaining computational efficiency.
[0079] S2.1.3. After the previous convolutional and inverted residual blocks, the resulting feature map is To further enhance the accuracy of feature extraction, a third Coordinate Attention (CA) mechanism module is introduced in the present invention.
[0080] The working principle of the coordinate attention mechanism is as follows:
[0081] (1) Input feature map: The input feature map is where H and W are the height and width of the feature map respectively, and C out is the number of channels. The input feature map is usually the intermediate feature output by the convolutional layer.
[0082] (2) One-dimensional global pooling: The first step of the CA attention mechanism is to perform one-dimensional global pooling on the feature map in the horizontal and vertical directions respectively. The purpose of the pooling operation is to obtain the global information of the feature map in these two directions.
[0083] a. Pooling in the vertical direction: In the vertical direction (height direction), global pooling is performed on each column to obtain a feature vector representing the information of the entire column. The pooling operation is similar to compressing the values of the entire column and taking their global average or maximum value. The size of the vertical pooling vector is 1×W×C out .
[0084] b. Pooling in the horizontal direction: In the horizontal direction (width direction), global pooling is performed on each row to obtain a feature vector representing the information of the entire row. This pooling operation also compresses the values of the entire row, and average pooling or max pooling can also be used. The size of the resulting feature vector is H×1×C out .
[0085] (3) Feature encoding: The two pooled feature vectors are encoded through a shared convolutional layer to generate attention weights in the two directions. This process helps the network encode the directional information in the features to generate an attention map with directional sensitivity. For the feature vector after vertical pooling, an intermediate feature of size 1×W×C out / r is generated (where r is the scaling factor used to reduce the number of parameters). For the feature vector after horizontal pooling, an intermediate feature of size H×1×C out / r is also generated.
[0086] (4) Feature fusion: The above two intermediate features in the two directions are further processed through a convolutional layer to generate two attention maps, corresponding to the horizontal and vertical directions respectively. The sizes of the attention maps are 1×W×C outand H×1×C out These two attention maps respectively represent the dependency relationships of the feature map in two directions.
[0087] (5) Feature enhancement: Next, the two generated attentions Figure 1 ×W×C out and H×1×C out are multiplied element-wise with the original input feature map to weight the information of the feature map in the horizontal and vertical directions respectively. This makes the network assign higher weights to important regions (such as indentations or tears on the corrugated cardboard box), and lower weights to irrelevant regions.
[0088] a. The attention map in the vertical direction is multiplied with the original feature map channel by channel to enhance the features in the vertical direction. The size of the feature map enhanced in the vertical direction remains H×W×C out , but the feature values of each column are weighted according to the attention map, enhancing the regions important in the vertical direction (such as the vertical indentations in the corrugated cardboard box).
[0089] b. The attention map in the horizontal direction is multiplied with the original feature map channel by channel to enhance the features in the horizontal direction.
[0090] The size of the feature map enhanced in the horizontal direction remains H×W×C out , but the feature values of each row are weighted according to the attention map in the horizontal direction, enhancing the regions important in the horizontal direction (such as the horizontal tears in the corrugated cardboard box).
[0091] After the enhancement in the vertical and horizontal directions, the network weights each position of the feature map, so that specific regions (such as defect regions) receive higher attention. The results of these two enhancement steps are fused together, and the final output feature map maintains the original size
[0092] Through this mechanism, the network can automatically focus on the regions with defects on the corrugated cardboard box, such as vertical indentations and horizontal tears, improving the accuracy and efficiency of detection.
[0093] S2.2. The feature fusion module performs multi-scale extraction and cross-layer fusion on the feature maps from different levels through a novel bidirectional weighted feature pyramid network.
[0094] To effectively fuse these features at different scales, this embodiment proposes a Bidirectional Weighted Feature Pyramid Network (BWFPN). This feature pyramid network can perform top-down and bottom-up feature transfer in the feature pyramid through bidirectional information flow. The design of BWFPN enables feature maps at different scales to be weighted and fused through bidirectional paths, thus ensuring that information flow can be efficiently transferred and fused in the network.
[0095] As Figure 4 shown, different layers of the convolutional neural network extract increasingly abstract features from the input feature map through layer-by-layer convolutional operations such as feature map (low resolution, high semantic information), feature map (high resolution, low semantic information).
[0096] S2.2.1, Top-down feature fusion:
[0097] First, upsample the highest-level feature map to obtain the feature map The upsampling method can be bilinear interpolation or nearest-neighbor interpolation. For example, nearest-neighbor interpolation is the simplest upsampling method. It directly assigns the value of the nearest pixel to the pixel to be inserted. Then, use pointwise convolution to adjust the number of channels of the feature map to C out to obtain the feature map In the weighted feature fusion process, BWFPN uses a learnable weight w i to control the contribution ratio of each feature map in the fusion. For each input feature X i ∈I9, the network assigns a learnable weight w i to control its contribution to the output feature, and uses the weighted formula to obtain the fused feature map
[0098] S2.2.2, Bottom-up feature fusion:
[0099] Downsample the lowest-level feature map to obtain the feature map Common downsampling methods such as pooling or convolution operations can reduce the spatial size of the feature map. Adjust the number of channels of the downsampled low-level feature map to C out through pointwise convolution for fusion with the upper-level feature map, and obtain the feature map and use the weighted formula to obtain the fused feature map
[0100] The calculation formula for the weighted fusion F is:
[0101]
[0102] Among them, w i represents the learnable weight of the input feature, and X i is the input feature map, η is a small constant used to prevent numerical instability, and F is the weighted output feature.
[0103] After the top-down and bottom-up bidirectional weighted fusion, the output feature map maintains the input spatial dimensions H×W×C out , but the fused feature map not only contains multi-level information but also dynamically adjusts the contribution of each layer of features through the weighted fusion mechanism, making the model more effective in multi-scale detection. The size of the obtained fused feature map is still H×W×C out . The output feature map not only retains the high-level semantic information but also combines the low-level detailed information, enabling the network to effectively detect defects of different scales on the corrugated cardboard box, such as large-area creases, corner damages, partial tears, etc.
[0104] S2.3, Continuous learning and new defect detection;
[0105] As Figure 5 shown, the model receives the feature map x∈I 10 from the fused features, and I 13 . These feature maps x usually pass through a classifier, which assigns probabilities to each defect category based on the feature map. The classifier outputs the probability q sample that each sample x j belongs to different defect categories label, denoted as q sample (x j |θ(t)), where θ(t) represents the model parameters of the classifier, and q sample is the probability of classifying the sample x j . The model determines the category of the sample by maximizing q
[0106]
[0107] To detect whether the model encounters a new type of defect label new , the system uses the Mahalanobis distance to measure the difference between the new input sample x sample and the known defect category label old . The Mahalanobis distance can measure the distance between a sample and the center of the Gaussian distribution of the existing category. By calculating the distance between the sample and the center of each category, the system can determine whether the sample belongs to an existing category or a new category.
[0108] First, define the mean μ j of each defect category j and the covariance matrix ∑j :
[0109]
[0110] where N j is the number of samples of defect type j.
[0111] The calculation formula of Mahalanobis distance is as follows:
[0112]
[0113] If the Mahalanobis distance exceeds a preset threshold, then the sample is determined to be a new defect label new . The new defect is sent to the manual inspection station for annotation to ensure that the new defect category is accurately defined and annotated.
[0114] After detecting a new defect and performing manual annotation, the continuous learning mechanism adaptively updates the model. By updating the classifier parameters of the model and the parameters of the new defect detector, the model gradually learns the characteristics of the new defect, ensuring that it can not only handle new types of defects but also not forget the existing categories.
[0115] To achieve this, the continuous learning process uses a joint loss function that includes cross-entropy loss and new defect detection loss. The goal of the loss function is to maximize the probability of correct classification while enabling the model to distinguish new defects from old defects. When updating the model, a joint loss function L(θ,φ,τ (t-1) ) is used, where τ (t-1) represents the score threshold of the previous batch of data. The loss function formula is as follows:
[0116]
[0117] This formula ensures that the model Model can adaptively update the parameter θ(t) through continuous learning when processing the new defect x new . The updated model can output more accurate classification results and has the ability to identify new defects. The output of the feature map includes not only the classification results of the existing categories but also the detection results of the new categories.
[0118] S3. Perform defect detection on the corrugated cardboard box image to be detected based on the trained corrugated cardboard box defect detection model.
[0119] Defect detection based on the trained corrugated cardboard defect detection model not only greatly improves the efficiency and accuracy of product quality control, but also provides strong data analysis support for manufacturing enterprises, helping them achieve higher quality standards while reducing operating costs. With the development of artificial intelligence technology, this intelligent detection method will become more popular in the future and will be extended to more fields, bringing revolutionary changes to various industries.
[0120] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0121] Based on the same idea as the lightweight convolutional neural network and continuous learning defect detection method in the above embodiments, the present invention also provides a lightweight convolutional neural network and continuous learning defect detection system, which can be used to execute the above lightweight convolutional neural network and continuous learning defect detection method. For the sake of convenience of description, in the structural schematic diagram of the lightweight convolutional neural network and continuous learning defect detection system embodiment, only the parts related to the embodiments of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and it may include more or fewer components than those illustrated, or combine certain components, or have different component arrangements.
[0122] Please refer to Figure 6 , in another embodiment of the present application, a lightweight convolutional neural network and continuous learning defect detection system 200 is provided, which includes an image acquisition module 201, a defect detection model construction module, and a defect detection module 203;
[0123] The image acquisition module 201 is configured to acquire a corrugated cardboard image and perform preprocessing;
[0124] The defect detection model construction module 202 is configured to construct a corrugated cardboard defect detection model and perform model training. The corrugated cardboard defect detection model includes a lightweight convolutional neural network module, a feature fusion module, and a continuous learning module; the lightweight convolutional neural network module performs feature extraction based on depthwise separable convolution, inverted residual blocks, and coordinate attention mechanisms; the feature fusion module performs multi-scale extraction and cross-layer fusion on feature maps from different levels through a novel bidirectional weighted feature pyramid network; the continuous learning module detects new defect types and automatically labels new samples through the input fused feature maps, and dynamically updates the classifier;
[0125] The defect detection module 203 is configured to perform defect detection on the corrugated cardboard image to be detected based on the trained corrugated cardboard defect detection model.
[0126] It should be noted that the lightweight convolutional neural network and continuous learning-based defect detection system of the present invention corresponds one-to-one with the lightweight convolutional neural network and continuous learning-based defect detection method of the present invention. The technical features and their beneficial effects described in the embodiments of the above-mentioned lightweight convolutional neural network and continuous learning-based defect detection method are applicable to the embodiments of the lightweight convolutional neural network and continuous learning-based defect detection. For specific content, reference can be made to the description in the method embodiments of the present invention, and details will not be repeated here. This is hereby declared.
[0127] In addition, in the implementation manner of the lightweight convolutional neural network and continuous learning-based defect detection system in the above embodiments, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the lightweight convolutional neural network and continuous learning-based defect detection system is divided into different program modules to complete all or part of the functions described above.
[0128] Please refer to Figure 7 , in one embodiment, an electronic device for implementing a lightweight convolutional neural network and continuous learning-based defect detection method is provided. The electronic device 300 may include a first processor 301, a first memory 302, and a bus, and may further include a computer program stored in the first memory 302 and executable on the first processor 301, such as a lightweight convolutional neural network and continuous learning-based defect detection program 303.
[0129] Among them, the first memory 302 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the first memory 302 may be an internal storage unit of the electronic device 300, such as the mobile hard disk of the electronic device 300. In some other embodiments, the first memory 302 may also be an external storage device of the electronic device 300, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 300. Further, the first memory 302 may also include both an internal storage unit and an external storage device of the electronic device 300. The first memory 302 can be used not only to store application software installed in the electronic device 300 and various types of data, such as the code of the lightweight convolutional neural network and continuous learning defect detection program 303, etc., but also to temporarily store data that has been output or will be output.
[0130] In some embodiments, the first processor 301 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The first processor 301 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the first memory 302, and calling data stored in the first memory 302, to perform various functions of the electronic device 300 and process data.
[0131] Figure 7 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 7 the shown structure does not constitute a limitation on the electronic device 300, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0132] The lightweight convolutional neural network and continuous learning defect detection program 303 stored in the first memory 302 in the electronic device 300 is a combination of multiple instructions. When running in the first processor 301, it can achieve:
[0133] Collect corrugated cardboard box images and perform preprocessing;
[0134] Build a corrugated cardboard box defect detection model and perform model training. The corrugated cardboard box defect detection model includes a lightweight convolutional neural network module, a feature fusion module, and a continuous learning module. The lightweight convolutional neural network module performs feature extraction based on depthwise separable convolution, inverted residual blocks, and coordinate attention mechanism. The feature fusion module performs multi-scale extraction and cross-layer fusion on feature maps from different levels through a novel bidirectional weighted feature pyramid network. The continuous learning module detects new defect types and automatically marks new samples through the input fused feature maps, and dynamically updates the classifier.
[0135] Perform defect detection on the corrugated cardboard box image to be detected based on the trained corrugated cardboard box defect detection model.
[0136] Furthermore, if the modules / units integrated in the electronic device 300 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).
[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media used in the various embodiments provided in the present application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0138] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0139] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A lightweight convolutional neural network-based and continuous learning defect detection method, characterized in that, Including the following steps: Collect corrugated box images and perform preprocessing; Construct a corrugated box defect detection model and conduct model training. The corrugated box defect detection model includes a lightweight convolutional neural network module, a feature fusion module, and a continuous learning module. The lightweight convolutional neural network module performs feature extraction based on depthwise separable convolution, inverted residual blocks, and coordinate attention mechanism. The feature fusion module conducts multi-scale extraction and cross-layer fusion on feature maps from different levels through a novel bidirectional weighted feature pyramid network. The continuous learning module detects new defect types and automatically labels new samples through the input fused feature maps, and dynamically updates the classifier; Perform defect detection on the corrugated box images to be detected based on the trained corrugated box defect detection model.
2. The defect detection method based on a lightweight convolutional neural network and continuous learning according to claim 1, wherein The step of collecting corrugated box images and performing preprocessing is specifically as follows: Collect corrugated box images through image acquisition devices installed on the production line.
3. The lightweight convolutional neural network and continuous learning-based defect detection method according to claim 1, characterized in that The lightweight convolutional neural network module includes a depthwise separable convolution module, an inverted residual block module, and a coordinate attention mechanism module, specifically: The lightweight CNN extracts the basic features of the image through the depthwise separable convolution module to obtain the feature map I1 of each channel; Pointwise convolution uses a convolution kernel of a set size to integrate the feature maps of different channels to form a new feature map I2. The integration is to perform operations on each pixel only with all channels at the position of this pixel; Input the new feature map into the inverted residual block module for processing, specifically: expand the channels through pointwise convolution to obtain the expanded feature map I3, perform depth convolution operations on the feature map I3 independently within each channel, and finally compress the expanded number of channels back to the original scale through pointwise convolution to obtain the feature map I4; Input the feature map I4 into the coordinate attention mechanism for processing, and obtain the global information of the feature map I4 in the horizontal and vertical directions through pooling operations to obtain the horizontal feature vector and the vertical feature vector; Perform feature encoding, feature fusion, and feature enhancement processing on the horizontal feature vector and the vertical feature vector to obtain feature maps I5 of different scales that maintain the original size.
4. The method for defect detection based on a lightweight convolutional neural network and continuous learning according to claim 3, characterized in that, The step of performing feature encoding, feature fusion, and feature enhancement processing on the horizontal feature vector and the vertical feature vector to obtain feature maps I5 of different scales that maintain the original size is specifically as follows: Use the coordinate attention mechanism module to perform one-dimensional global pooling on the feature map I4 in the horizontal and vertical directions respectively. Perform global pooling on each column in the vertical direction to obtain a first feature vector representing the information of the entire column, and perform global pooling on each row in the horizontal direction to obtain a second feature vector representing the information of the entire row; Encode the direction information in the first feature vector and the second feature vector through a shared convolutional layer to generate a first intermediate feature and a second intermediate feature with direction sensitivity; Process the first intermediate feature and the second intermediate feature through a convolutional layer to generate a first attention map in the horizontal direction and a second attention map in the vertical direction respectively; Multiply the first attention map and the second attention map element - by - element with the feature map I4, respectively weighting the information of the feature map I4 in the horizontal and vertical directions to enhance the features in the horizontal and vertical directions, so that a specific area or the defects in a specific area receive higher attention.
5. The method for defect detection based on a lightweight convolutional neural network and continuous learning according to claim 1, characterized in that In the feature fusion module, different - scale feature maps can be weighted and fused through a two - way path by a new type of bidirectional weighted feature pyramid network. Specifically: Extract a feature map I6 with low - resolution and high - semantic information and a feature map I7 with high - resolution and low - semantic information from the feature map I5; Bottom-up feature fusion: Upsample the feature map I6 to obtain the feature map I8, use pointwise convolution to adjust the number of channels of the feature map I8 to obtain the feature map I9. In the weighted feature fusion process, a learnable weight is assigned to control the contribution of the feature map I9 in the fusion, and the weighted formula is used to obtain the fused feature map I 10 ; Bottom-up feature fusion: Downsample the feature map I7 to obtain the feature map I 11 , and use pointwise convolution to adjust the number of channels of the feature map I8 to obtain the feature map I 12 , and use the weighted formula to obtain the fused feature map I 13 ; After two - way weighted fusion through top - down and bottom - up directions, the output feature map maintains the input spatial size. The output fused feature map not only retains the high - level semantic information but also combines the low - level detailed information, enabling the network to effectively detect defects of different scales on the corrugated cardboard box.
6. The method for defect detection based on a lightweight convolutional neural network and continuous learning according to claim 5, wherein In the continuous learning module, the fused feature map I 10 and the feature map I 13 pass through a classifier to output the probability of each sample belonging to different defect classifications, and the defect category to which the sample belongs is determined by maximizing the classification probability. To detect whether the detection model encounters a new type of defect, the Mahalanobis distance is used to measure the difference between the newly input sample and the known defect categories. By calculating the distance between the sample and the center of each category, it is determined whether the sample belongs to an existing category or a new category. After detecting a new defect and manually annotating it, the continuous learning mechanism adaptively updates the model. By updating the parameters of the model's classifier and the parameters of the new - defect classifier, the features of the new defect are gradually learned, enabling the model to not only handle new types of defects but also not forget the existing defect categories.
7. The method for defect detection based on a lightweight convolutional neural network and continuous learning according to claim 6, wherein The continuous learning process uses a joint loss function that includes the cross - entropy loss and the new - defect detection loss. The objective of the joint loss function is to maximize the probability of correct classification while enabling the model to distinguish new defects from old defects.
8. A lightweight convolutional neural network and continuous learning-based defect detection system, characterized in that, Applied to the lightweight convolutional neural network - based and continuous - learning defect detection method according to any one of claims 1 - 7, it includes an image acquisition module, a defect detection model construction module, and a defect detection module; The image acquisition module is used to acquire a corrugated cardboard box image and perform pre - processing; The defect detection model construction module is used to construct a corrugated cardboard box defect detection model and perform model training. The corrugated cardboard box defect detection model includes a lightweight convolutional neural network module, a feature fusion module, and a continuous learning module. The lightweight convolutional neural network module performs feature extraction based on depth - separable convolution, inverted residual blocks, and coordinate attention mechanism. The feature fusion module performs multi - scale extraction and cross - layer fusion on feature maps from different levels through a new type of bidirectional weighted feature pyramid network. The continuous learning module detects new defect types and automatically marks new samples through the input fused feature map, and dynamically updates the classifier; The defect detection module is used to perform defect detection on the corrugated cardboard box image to be detected based on the trained corrugated cardboard box defect detection model.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to execute the defect detection method based on a lightweight convolutional neural network and continuous learning as described in any one of claims 1-7.
10. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, it implements the defect detection method based on a lightweight convolutional neural network and continuous learning as described in any one of claims 1-7.
Citation Information
Cited By
Cloth surface flaw detection method and device based on texture perception and anomaly detection
CN121169883A