Image Anomaly Detection Method and System Based on Lightweight Object Detection Model
By combining GhostConv and DWConv on the YOLOv8 model, a GH-HGV2Net backbone network is built, and feature fusion modules and lightweight shared group convolution detection heads are designed, the existing object detection methods are solved in the detection of hidden target anomalies and high computing resource consumption in complex industrial environments, achieving high precision and real-time lightweight object detection.
Patent Information
- Application Number
- CN202510357025.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-25
AI Technical Summary
When the existing object detection methods process images in complex industrial environments, it is difficult to effectively detect hidden object abnormalities, and the computing resource consumption is high, making it difficult to meet the deployment requirements of real-time and resource-constrained devices.
The lightweight object detection model based on the YOLOv8 model is built by combining GhostConv and DWConv to build a new backbone network GH-HGV2Net, and a feature fusion module and a lightweight shared group convolution detection head are designed to reduce the model parameter amount and calculation complexity while maintaining high detection accuracy.
It realizes the significant reduction in model parameters and calculation complexity while maintaining high detection accuracy, improves the real-time and applicability of the method, and can effectively respond to the high-precision and real-time requirements of industrial production and other practical application scenarios.
Smart Images

Figure CN119904624B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of object detection, relates to a lightweight object detection network model, and specifically relates to image anomaly detection, which can be applied to lightweight neural network image anomaly detection in various scenarios. Background Art
[0002] Object detection algorithms play a crucial role in industrial production and quality management, and can effectively improve production efficiency and ensure product quality. However, traditional manual detection methods have significant deficiencies, such as slow detection speed, high missed detection rate, etc., and it is difficult to meet the fast-paced and high-precision manufacturing requirements of modern industry. With the progress of technology, especially the development of deep learning methods, object detection algorithms based on convolutional neural networks (CNNs) have become the mainstream solutions. Currently, the mainstream object detection methods are mainly divided into two categories: region proposal-based methods, such as R-CNN, Faster R-CNN, etc. These methods use a region proposal network (RPN) to generate candidate regions, and then classify and regress each candidate region. Although these methods perform well in detection accuracy, their computational cost is high and the detection speed is slow, making it difficult to meet real-time requirements. Regression-based methods, such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector), transform the object detection problem into a regression problem and directly predict the location and category of the object through a single forward propagation. These methods are favored for their faster detection speed and are more suitable for real-time application scenarios. Although the above methods have made significant progress, they still face many challenges in practical applications. The target images in industrial environments usually contain complex local and global features, and their anomalies or defects may be embedded in normal structures in extremely concealed ways, thus increasing the difficulty of detection. In addition, in order to achieve large-scale deployment, reducing the computational cost and storage requirements of deep learning models remains an important issue to be solved.
[0003] Therefore, developing a lightweight object detection method with both high accuracy and low computational resource consumption can not only meet the detection requirements in complex scenarios, but also meet the real-time and deployment requirements of resource-constrained devices, which is of great significance for promoting the wide application of object detection technology in industrial production and other fields. Summary of the Invention
[0004] The present invention aims to propose solutions to two core problems, that is, on the basis of the existing YOLOv8 model, further optimize the network structure to significantly reduce the consumption of computing resources, while ensuring that the lightweight design of the model does not come at the cost of detection accuracy. Through this improvement, the present invention can effectively meet the high-precision and real-time requirements for object detection in industrial production, quality inspection and other practical application scenarios, providing a reliable solution for efficient deployment in resource-constrained environments.
[0005] To solve the above problems, the present invention provides an image anomaly detection method based on a lightweight object detection model, which combines an innovative network architecture and optimization techniques to ensure that while maintaining high detection accuracy, significantly reduce the number of model parameters and computational complexity, thereby improving the real-time performance and applicability of the method. The specific technical solutions include the following:
[0006] Step 1, construct a lightweight object detection model improved based on the YOLOv8 model, including a backbone network, a feature fusion module, and a lightweight shared group convolution detection head;
[0007] Combine the phantom convolution GhostConv and the depthwise convolution DWConv, and integrate them with the lightweight backbone network HGNetV2 to construct a new backbone network GH-HGV2Net to extract features from the input image. The entire new backbone network generates two output feature layers at the intermediate stage;
[0008] The processing process of the feature fusion module is as follows: the final output feature of the new backbone network passes through the SPPF module to generate a third output feature layer; combine the dynamic convolution DynaMicConv, the C2f module, and the phantom convolution GhostConv to construct the phantom dynamic convolution layer C2fGhostDynaMic; perform multi-level fusion on the three output feature layers through the phantom dynamic convolution layer, the upsampling layer, the phantom convolution, and the splicing operation, and finally output three feature maps of different sizes;
[0009] Combine grouped convolution and convolution operations with shared parameters to construct a lightweight shared group convolution detection head LSCD for object prediction;
[0010] Step 2, select the images containing the anomalies to be detected, label the positions of the anomalies to be detected in the images, and record the coordinate information of the anomalies to form an anomaly detection dataset;
[0011] Step 3, use the anomaly detection dataset to train the constructed object detection model, and output the trained object detection model to realize image anomaly detection.
[0012] Furthermore, the new backbone network GH-HGV2Net includes a preliminary extraction module HGStem and 4 stages;
[0013] The 4 stages are stage1 to stage4 respectively. Among them, stage1 contains a GH-HGBlock, stage2 contains a depthwise convolution DWConv and a GH-HGBlock, stage3 contains a DWConv and three consecutive GH-HGBlocks, and stage4 is the same as stage2.
[0014] Furthermore, the preliminary extraction module HGStem includes multiple ordinary convolution and pooling layers.
[0015] Furthermore, the processing flow of the GH-HGBlock is as follows: First, the input features go through n ghost convolutional GhostConv operations to generate a series of feature maps; then, these feature maps are concatenated through the Concatenate operation to connect the feature maps after each GhostConv operation; next, the concatenated feature maps go through the squeeze convolution SqueezeConv and excitation convolution ExcitationConv operations to extract the channel information and feature responses of the feature maps; finally, according to the condition that the number of channels of the original feature map is the same as that of the feature map after the intermediate operations, the original input and the feature maps generated by the GhostConv and Concatenate operations are added together to generate the final feature map.
[0016] Furthermore, the phantom dynamic convolution layer C2fGhostDynaMic includes a dynamic convolution DynamicConv and a phantom bottleneck layer GhostBottleNeck; the processing process of C2fGhostDynaMic is as follows: The input feature map first goes through the dynamic convolution DynamicConv, and then the output result is split into two parts. The first part of the result goes through two identical GhostBottleNecks, and then the second part of the split feature map is concatenated with the feature maps after the first GhostBottleNeck and the second GhostBottleNeck operations, and finally goes through another DynamicConv;
[0017] Among them, the processing process of the GhostBottleNeck is as follows: The input first goes through the phantom convolution GhostConv, then through the batch normalization layer and the activation layer, then through GhostConv, and finally added to the original feature map to output the final feature map.
[0018] Furthermore, the specific processing process of the feature fusion module is as follows:
[0019] The features extracted by the novel backbone network are first processed by the SPPF module, followed by an upsampling operation. Meanwhile, the upsampled feature map is concatenated with the feature map provided by the intermediate stage 3. Then, it passes through a C2fGhostDynaMic module with 13 layers, and then another upsampling operation is performed. The upsampled feature map is concatenated with the feature map provided by the intermediate stage 2. At this time, it passes through another C2fGhostDynaMic module with 16 layers, and the generated feature map is used to provide to the detection head for detecting small targets. The 16-layer feature map then undergoes a GhostConv upsampling operation. Then, the upsampled feature map is concatenated again with the feature map of the first concatenation, that is, the 13-layer feature map. After that, it passes through a C2fGhostDynaMic module with 19 layers and provides it to the detection head for detecting medium targets. In addition, the 19-layer feature map passes through another GhostConv operation and is then concatenated with the feature map passing through the SPPF. Finally, the concatenated feature map passes through a C2fGhostDynaMic module with 22 layers and provides it to the detection head for detecting large targets.
[0020] Furthermore, the processing process of the lightweight shared group convolution detection head LSCD is as follows:
[0021] First, feature maps of different sizes are respectively processed by the grouped convolution Conv_GN1x1 to reduce the computational load. Then, the features processed by Conv_GN1x1 are input into two grouped convolutions with shared parameters of Conv_GN3x3 to further extract feature information. Then, after the grouped convolution with shared parameters, the features are divided into three branches for detecting targets of different sizes, including small targets, medium targets, and large targets. In each branch, a regression convolution operation Conv_Reg and a classification convolution operation Conv_Cls with shared parameters are used for the regression and classification tasks of the target respectively. At the same time, in each Conv_Reg operation, a trainable scalar parameter scale is introduced to adjust the scale of target regression, so as to adapt to abnormal regions of industrial products of different sizes. The parameter scale is initialized to the positive value 1.0 and is then gradually optimized through gradient descent.
[0022] Furthermore, step 1 also includes using the LAMP pruning technique to further compress and speed up the lightweight object detection model;
[0023] In LAMP, the optimization objective of the current layer is to minimize the impact of the pruning operation on the model performance by optimizing the mask matrix ,
[0024] (1)
[0025] Among them, is the binary mask matrix of the current layer. Each element of this matrix is either 0 or 1, and is used to indicate whether the weight at that position is retained or pruned. During the training process it will automatically adjust the learning; is the 0-norm of the mask matrix, representing the number of non-zero weights allowed in the current layer, is the pruning sparsity threshold of the layer, used to constrain the upper limit of non-zero elements in the mask, represents the normalized representation of the input data, is the 2-norm of the input data, used to constrain the scale of the input data, is to find the maximum value of the vector under the condition of satisfying is the weight matrix representing the layer, is element-wise multiplication;
[0026] The optimization of the entire lightweight object detection model is expressed as:
[0027] (2)
[0028] Among them, is a constant, used to limit the number of non-zero weights in the model, is the lightweight object detection model when using the weight parameter the output for the input , is the weight after sparsification the output for the input , represents the set of all weight matrices from the first layer to the layer, is the total number of layers of the lightweight object detection model;
[0029] Under the greedy strategy, the connection with the lowest importance score is removed in each iteration. The above optimization is relaxed through Equation (3):
[0030] (3)
[0031] Among them, is the sparsified weight tensor, is the F-norm. The weight tensor is flattened into a one-dimensional vector, and the one-dimensional vector is sorted so that for any constant When it satisfies , respectively represent the -th and -th vectors after sorting; Since the product term is a constant value, the importance score of the -th index in the weight tensor is defined as, that is, the LAMP score of the -th index in the weight tensor is defined as follows:
[0032] (4)
[0033] Finally, prune the connection with the smallest global pruning LAMP score until the required global sparsity constraint is satisfied.
[0034] Furthermore, step 1 also includes enhancing the lightweight object detection model using knowledge distillation; The lightweight object detection model is divided into three models of n, s, and m according to the number of network layers, the number of parameters, and the computational complexity. The n-scale model has the smallest number of network layers, the number of parameters, and the computational complexity, the m-scale model is the largest, and the s-scale model is medium. The s-scale model is used as the teacher network, and the n-scale model is used as the student network;
[0035] Perform knowledge distillation on specific channels in the feature fusion module. First, convert the activation map of the channel into a probability distribution, use the probability distance metric to measure the difference, and represent the teacher network and the student network as and , and The activation maps of are respectively represented as and , and the general form of the channel distillation loss is expressed as:
[0036] (5)
[0037] Among them, is the KL divergence, which is used to calculate the probability distance metric between two inputs, is the softmax function, which is used to convert the activation map into a probability distribution; and are the activation maps of the teacher network and the student network respectively, and are the activation maps of specific channels in the teacher network and the student network respectively;
[0038] Use to convert the activation value into a probability distribution, as shown in Equation (6):
[0039] (6)
[0040] Among them , represents the index channel; represents the spatial position of the index channel, and are the width and height of the channel feature map, is a hyperparameter.
[0041] The present invention also provides an image anomaly detection system based on a lightweight object detection model, including:
[0042] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the image anomaly detection method based on the lightweight object detection model as described in the above technical solution.
[0043] The present invention proposes a backbone network called GH-HGV2Net for object detection tasks. The hierarchical structure can well handle the characteristics of the local and global structures of objects. Such a design can effectively extract features at different levels, so as to better capture the local fine structures and global overall structures in the image. The network adopts an HG Stage composed of GH-HGBlocks, and the convolution operation in each GH-HGBlock uses the GhostConv technology. The GhostConv convolution extracts feature information from the input feature map using a small number of ordinary convolutional kernels, and then generates the final feature map through a more inexpensive linear transformation operation. This special convolution operation can reduce the network calculation amount and parameter amount while maintaining the original feature map size and number of channels, thereby improving the network calculation efficiency and performance. In addition, the backbone network also includes an HGStem module, which is composed of ordinary convolution and pooling layers, and is used to reduce the calculation amount and extract preliminary feature representations. The role of the HGStem module is to convert the input data image from high resolution to low resolution suitable for network processing and extract preliminary feature representations. The entire backbone network consists of an HGStem module and 4 HG Stages, aiming to improve the accuracy and robustness of object detection through hierarchical feature representation and multi-scale feature extraction. Finally, the output features of the backbone network will be sent to an SPPF (Spatial Pyramid Pooling Fusion) module to extract multi-scale feature information and perform fusion. Through hierarchical feature representation, accurate capture of the local and global structures in the image and sensitive detection of the lower targets are realized, so as to further enhance the detection ability and robustness of the network.
[0044] In the object detection task, the detection of tiny objects is often highly challenging, especially when the feature size of the object is very small and easily confused with the background texture. This scenario requires that the lightweight network pay special attention to the extraction and expression ability of detailed features during design. To this end, the present invention proposes a feature fusion module C2fGHDY-PAFPN (Neck) based on dynamic convolution to improve the performance of the lightweight network in the detection of tiny objects. Dynamic convolution is introduced as part of feature extraction, and its powerful feature expression ability can better distinguish the subtle differences between object features and background textures, thereby improving the detection accuracy. Although dynamic convolution slightly increases the amount of computation compared to traditional convolution, its significant advantage in feature extraction accuracy can make up for this, resulting in a substantial improvement in detection performance. The designed new module C2fGHDY-PAFPN (Neck) is obtained by replacing all the ordinary convolution blocks in the Bottelneck of the C2f module with GhostConv and dynamic convolution. This structure uses the cross-stage feature fusion strategy and the truncated gradient flow technique to enhance the variability of the learned features between different network layers, thereby reducing the influence of redundant gradient information and enhancing the learning ability of the network. Due to the introduction of GhostConv and dynamic convolution, the C2fGHDY-PAFPN module reduces a large number of 3×3 ordinary convolutions in the original structure, greatly compressing the model size of the network, reducing the number of parameters and the amount of computation.
[0045] To address the issues of high model computational complexity and diverse object sizes and shapes in object detection tasks, the present invention designs a lightweight grouped shared-weight detection head (LSCD). This detection head can effectively reduce the number of parameters and improve computational efficiency, making it particularly suitable for object detection tasks in resource-constrained environments. Specifically, the same set of convolutional layers is used to predict detection bounding boxes for the multi-scale feature maps (Feature Map) output from the feature fusion part. Then, a learnable Scale value is used as a coefficient for each layer to scale the predicted bounding boxes. The Group Normalization (GroupNorm) layer has been proven to enhance the localization and classification performance of the detection head. By using shared convolutions, the number of parameters can be significantly reduced, making the model more lightweight, especially on resource-constrained devices. When using shared convolutions, all the normalization layers within the convolutions are replaced with GroupNorm. The parameters of the classification loss can be shared because whether it is small, medium, or large objects, the information learned in the detection layer is about the same object and is not affected. However, when considering that the regression (Reg) loss function can be shared, the scale problem needs to be taken into account. Therefore, a Scale layer is added. To handle the inconsistent object scales detected by each detection head while using shared convolutions, the Scale layer is used to scale the features. To further compress the size of the object detection model and improve the inference speed, the present invention introduces the LAMP (Layer-wise Automatic Magnitude Pruning) pruning technique. This technique is specifically designed for lightweight object detection models. Through an efficient pruning strategy, it can significantly reduce the computational resource requirements while maximizing the retention of model performance, and is applicable to lightweight optimization scenarios for various object detection tasks. The core advantage of the LAMP pruning technique is that it does not rely on hyperparameters and does not require knowledge of a specific model, making it highly versatile. Its principle is to consider each layer of the neural network as an independent operator. By analyzing the impact of pruning on "model-level" distortion, the priority of sparsity for each layer is determined. Specifically, LAMP calculates the LAMP score to re-adjust the network weights to approximately measure the performance loss caused by pruning, and automatically determines the sparsity of each layer based on the score, thereby achieving a globally optimized pruning process. This method retains all the advantages of magnitude pruning (MP) during the pruning process. By globally removing low-priority weights, the LAMP pruning technique can significantly reduce the number of model parameters and computational complexity while maintaining the detection accuracy and robustness of the model.
[0046] The pruned network model will lose some accuracy. To improve the detection accuracy, channel-based knowledge distillation technology can be used. Taking the s version of the improved model as the teacher model and the n version as the student model, knowledge distillation is performed on the three feature layers (16, 19, 22) generated by the feature fusion module, and the learning ability and detection accuracy of the student model for features are improved through channel-level distillation.
[0047] In summary, through in-depth research, the inventor team proposed a lightweight object detection model GH_HGV2-LSCD-YOLOX based on YOLOX. This model has high flexibility and adaptability and can be widely applied to a variety of industrial scenarios, including actual needs such as fabric defect detection, power equipment defect detection, and drug appearance abnormality detection. Brief Description of the Drawings
[0048] Figure 1 is a flowchart provided by an embodiment of the present invention.
[0049] Figure 2 is a model framework diagram of the improved lightweight object detection model provided by an embodiment of the present invention.
[0050] Figure 3 is a model diagram of GhostConv provided by an embodiment of the present invention.
[0051] Figure 4 is a model diagram of GH-HGBlock in the backbone network provided by an embodiment of the present invention.
[0052] Figure 5 is a structural diagram of channel-based knowledge distillation provided by an embodiment of the present invention. Detailed Description of the Invention
[0053] The present invention is optimized and improved based on the YOLOv8 model, and a lightweight object detection model GH_HGV2-LSCD-YOLOX is proposed. As Figure 1 shown, this model achieves lightweight and efficient detection through the following key steps:
[0054] First, by combining GhostConv and DWConv and integrating them with the lightweight backbone network HGNetV2, a new backbone network GH-HGV2Net is constructed, thus significantly reducing the number of model parameters. In the feature fusion stage, DynaMicConv, C2f module with strong feature expression ability and GhostConv are combined to design a dual-stream PANet structure to achieve feature fusion. Finally, through the lightweight shared group convolution detection head (LSCD), the number of parameters is further reduced while the accuracy and inference speed of the model are improved.
[0055] GH_HGV2-LSCD-YOLOX consists of three core modules: the backbone network, the feature fusion module, and the lightweight YOLO detection head. Among them, the backbone network is the improved GH-HGV2Net, and the feature fusion module is C2fGHDY-PAFPN. The overall network structure is as follows: First, during the model training and testing process, the backbone network first extracts features from the input image. After passing through the HGStem (Hierarchical Group Stem) module, 4 stages (HGStage), and the SPPF module, three output feature layers are generated: (80, 80, 512), (40, 40, 1024), and (20, 20, 1024). Secondly, these feature layers are then fed into the C2fGHDY-PAFPN feature fusion network. The feature fusion network, through a bidirectional fusion mechanism, generates three feature maps (corresponding to small, medium, and large targets) after two downsampling and two upsampling operations, with 16, 19, and 22 layers respectively, and finally feeds them into the LSCD detection head for target prediction.
[0056] To further optimize the lightweight performance of the model, the present invention introduces the LAMP (Layer-wise Automatic Magnitude Pruning) pruning technique. Compared with traditional pruning methods, the LAMP pruning technique does not require hyperparameter adjustment and does not rely on specific model knowledge. It can globally evaluate the importance of weights, thereby achieving automated and efficient pruning operations, significantly compressing the model size and improving the inference speed.
[0057] In addition, to enhance the detection accuracy of the pruned model, the present invention also adopts the channel knowledge distillation technique. Specifically, the s version of the YOLOv8 model is used as the teacher model, and the n version is used as the student model. Knowledge distillation is performed on the three feature layers (16, 19, 22) generated by the feature fusion module to improve the student model's learning ability and detection accuracy of features through channel-level distillation.
[0058] The GH_HGV2-LSCD-YOLOX model proposed by the present invention exhibits excellent performance in the n-scale network structure, can maintain a high level of detection accuracy while significantly reducing the consumption of computing resources, and provides an efficient and reliable solution for industrial production and other real-time target detection scenarios.
[0059] To more clearly elaborate on the specific implementation method of the present invention, the following will be described in detail from three key aspects: the network model structure, the pruning technique, and the knowledge distillation technique. Through the organic combination of these three core modules, the present invention realizes the lightweight, high-efficiency, and high-precision of the target detection model, providing an efficient and reliable solution for the industrial field.
[0060] I. Model Structure
[0061] GH_HGV2-LSCD-YOLOX consists of three core modules: the backbone network, the feature fusion module, and the lightweight YOLO detection head. Among them, the backbone network is the improved GH-HGV2Net, and the feature fusion module is C2fGHDY-PAFPN. The overall network structure is as follows Figure 2 shown. First, during the model training and testing process, the backbone network first extracts features from the input image. After passing through the preliminary extraction module HGStem module, as Figure 3 shown, the HGStem module consists of multiple ordinary convolutions, max pooling layers, and connection operations, which are used to reduce the computational complexity and extract preliminary feature representations. The role of the HGStem module is to convert the input industrial product image from high resolution to low resolution suitable for network processing and extract the preliminary feature representations.
[0062] Subsequently, it passes through 4 stages (HGStage). Stage 1 contains one GH-HGBlock, stage 2 contains one depthwise convolution DWConv (Depthwise Convolution) and one GH-HGBlock, stage 3 contains one DWConv and three consecutive GH-HGBlocks, and stage 4 is the same as stage 2. The aim is to improve the accuracy and robustness of image anomaly detection through hierarchical feature representation and multi-scale feature extraction.
[0063] Finally, the output features of the backbone network are fed into the SPPF module to extract and fuse multi-scale feature information. The entire backbone network generates three output feature layers after stage 2, stage 3, and the SPPF module, which are: (80, 80, 512), (40, 40, 1024), and (20, 20, 1024). Secondly, these feature layers are then fed into the C2fGHDY-PAFPN feature fusion network. The feature fusion network generates three feature maps (corresponding to small targets, medium targets, and large targets) through a bidirectional fusion mechanism, after two downsampling and two upsampling operations, and finally feeds them into the LSCD detection head for target prediction.
[0064] Among them, as Figure 3 shown, the working process of the Ghost convolution GhostConv is divided into two stages. In the first stage, a small number of convolutions (for example, normally using 128 convolutional kernels, here using 64, thus reducing the computational complexity by half). In the second stage, CheapOperation, using Figure 3The Φ in it indicates that Φ is a convolution such as 3*3 or 5*5, and it is a depth-wise convolution (convolution operation is a complete combination of convolution-batch normalization BN-nonlinear activation, and the so-called linear transformation or cheap operation refers to ordinary convolution, excluding batch normalization and nonlinear activation). The final feature map is obtained by merging the feature maps output in the first stage and the feature maps output in the second stage to get the same number of feature maps as that of ordinary convolution (the advantage is that the computational amount is reduced).
[0065] As Figure 4 shown, the GH-HGBlock is the basic building block of the backbone network, and its internal process is as follows: First, the input features go through n GhostConv convolution operations to generate a series of feature maps. Then, these feature maps are concatenated through the Concatenate operation to connect the feature maps after each GhostConv operation. Next, the concatenated feature maps go through the SqueezeConv and ExcitationConv operations for extracting the channel information and feature responses of the feature maps. Finally, according to the condition that the number of channels of the original feature map is the same as that of the feature map after intermediate operations, the original input and the feature maps generated by the GhostConv and Concat operations are added together to generate the final feature map. Such a design can make full use of the advantages of the GhostConv convolution operation, reduce the network computational amount and the number of parameters, and at the same time, through the Concatenate and addition operations, fuse the feature information at different stages, improving the network's feature expression ability and detection performance.
[0066] As Figure 2As shown, the Phantom Dynamic Convolution Layer C2fGhostDynaMic is a newly designed model that replaces the ordinary convolution in the original Bottleneck model with GhostConv and DynamicConv. The C2fGhostDynaMic module is mainly composed of DynamicConv and GhostBottleNeck. Among them, GhostBottleNeck first passes through GhostConv, then through the batch normalization layer and the activation layer, then through GhostConv again, and finally adds with the original feature map to output the final feature map. The input feature map of the C2fGhostDynaMic module first passes through DynamicConv, and then the output result is split into two parts. The first part of the result passes through two identical GhostBottleNeck modules. Then, the second part of the split feature map is concatenated with the feature maps after the operations of the first GhostBottleNeck module and the second GhostBottleNeck module. Finally, it passes through another DynamicConv module. This structure uses the cross-stage feature fusion strategy and the truncated gradient flow technology to enhance the variability of the learned features between different network layers, thereby reducing the impact of redundant gradient information and enhancing the learning ability of the network. Due to the introduction of GhostConv and DynamicConv, the C2fGhostDynaMic module reduces a large number of 3×3 ordinary convolutions in the original structure, greatly compresses the model size of the network, reduces the number of parameters and the amount of computation, enabling the model to be deployed on mobile devices and making it easier to implement edge computing detection of abnormal regions.
[0067] According to Figure 2As shown, the features with a size of 20x20x1024 extracted by the backbone network are first processed by the SPPF (Spatial Pyramid Pooling Fusion) module. Subsequently, an upsampling operation is performed, and the upsampled feature map is concatenated with the 40x40x1024 feature map provided by the third stage. Then, it passes through a C2fGhostDynaMic module (labeled as 13 layers), and then another upsampling operation is performed, and the upsampled feature map is concatenated with the 80x80x512 feature map provided by the second stage. At this time, it passes through another C2fGhostDynaMic module (labeled as 16 layers), and the generated feature map is used to provide to the detection head for detecting small targets. The 16-layer feature map then undergoes a GhostConv upsampling operation, and then, the upsampled feature map is concatenated again with the feature map of the first concatenation (13-layer feature map), and then passes through a C2fGhostDynaMic module (labeled as 19 layers), which is provided to the detection head for detecting medium targets. In addition, the 19-layer feature map undergoes another GhostConv operation and is then concatenated with the 20x20x1024 feature map that has passed through SPPF. Finally, the concatenated feature map passes through a C2fGhostDynaMic module (labeled as 22 layers) and is provided to the detection head for detecting large targets. In this way, the process of the feature map circulation of the feature fusion network is completed.
[0068] According to Figure 2As shown in the figure, in order to reduce the computational complexity of the model and the problem of diverse sizes and shapes of abnormal regions, this paper designs a detection head LSCD with grouped shared weights, which operates based on the features extracted from the 16th, 19th, and 22nd layers of the network. Specifically, first, these features are respectively processed by Conv_GN1x1 (grouped convolution) to reduce the computational complexity. Then, the features processed by Conv_GN1x1 are input into two grouped convolutions with shared parameters of Conv_GN3x3 to further extract feature information. Next, after the grouped convolution with shared parameters, we divide the features into three branches for detecting targets of different sizes, including small targets, medium targets, and large targets. In each branch, we again adopt the convolution operations Conv_Reg and Conv_Cls with shared parameters for the regression and classification tasks of the targets respectively. It should be noted that in each Conv_Reg operation, we introduce a trainable scalar parameter scale to adjust the scale of target regression, so as to adapt to abnormal regions of different sizes. The scale parameter is initialized to the positive value 1.0 and then gradually optimized through gradient descent to minimize the loss function and improve the model performance (where the loss function used is the existing loss function of the YOLOv8 model). The trained scale parameter in the inference stage is directly applied to adjust the regression value of the prediction box. The introduction of the scale parameter enables our model to more flexibly adapt to the sizes and shapes of various industrial product abnormalities. By continuously optimizing the scale parameter during the training process, we can achieve effective detection and localization of targets of different sizes, improving the accuracy and robustness of the detection.
[0069] II. Pruning Technique
[0070] This paper adopts the LAMP (Layer-wise Automatic Magnitude Pruning) pruning technique to further compress and speed up the improved lightweight network. Compared with other pruning methods, LAMP pruning does not require hyperparameters and does not rely on any model-specific knowledge. Its principle is to regard each neural network layer as an operator, and determine the priority of the hierarchical sparsity of the magnitude pruning method MP (Magnitude Pruning) based on the weight size by examining the "model-level" distortion brought by the pruning layer, that is, to re-adjust the weight size through the LAMP score, approximate the model-level distortion caused by pruning, and automatically determine the hierarchical sparsity, so as to achieve further compression and speed up of the lightweight network. Global pruning with the LAMP score is equivalent to MP with automatically determined hierarchical sparsity, and using LAMP pruning can completely retain the advantages of MP.
[0071] In LAMP, the optimization objective of the current layer is to optimize the mask matrix under the constraint conditions , to minimize the impact of pruning operations on model performance:
[0072] (1)
[0073] Among them, is the binary mask matrix of the current layer. Each element of this matrix is either 0 or 1, indicating whether the weight at that position is retained (1) or pruned (0). During the training process the learning will be automatically adjusted. is the 0-norm of the mask matrix, representing the number of non-zero weights allowed in the current layer. is the pruning sparsity threshold of the layer, used to constrain the upper limit of non-zero elements in the mask. represents the normalized representation of the input data, is the 2-norm of the input data, used to constrain the scale of the input data. is to find the maximum value of the vector under the condition of satisfying . is the weight matrix of the representation layer , is element-wise multiplication.
[0074] The optimization of the entire network can be expressed as:
[0075] (2)
[0076] Among them, is a constant, used to limit the number of non-zero weights in the model. is the network (GH_HGV2-LSCD-YOLOX model) when using the weight parameter the output for the input , is the output of the input after being sparsified by the weight , represents the set of all weight matrices from the first layer to the layer.
[0077] Under the greedy strategy, the connection with the lowest importance score is removed in each iteration. Through the spectral norm that is, the largest singular value of matrix A and the Frobenius norm that is, the square root of the sum of the squares of the matrix elements. According to the inequality relationship conditions, the above optimization can be relaxed by using the following upper bound of the model output distortion:
[0078] (3)
[0079] Among them, is the sparsified weight tensor, is the F-norm; the weight tensor is flattened into a one-dimensional vector, and the one-dimensional vector is sorted so that for any when, it satisfies , respectively represent the th and th vectors after sorting. Since the product term is a constant value, the importance score of the th index in the weight tensor is defined as, that is, the LAMP score of the th index in the weight tensor is defined as follows:
[0080] (4)
[0081] Briefly speaking, the LAMP score (Equation 4) measures the relative importance of the target connection among all surviving connections belonging to the same layer. In the same layer, connections with smaller weight magnitudes (in the same layer) have been pruned. Therefore, two connections with the same weight magnitude have different LAMP scores used. Once the LAMP scores are calculated, the connections with the smallest global pruning LAMP scores will be removed until the required global sparsity constraint is satisfied; this process is equivalent to automatically selecting hierarchical sparsity to perform MP.
[0082] III. Knowledge Distillation
[0083] In order to enhance the pruned model. Knowledge distillation is an effective method to improve the detection accuracy of the algorithm. Existing knowledge distillation methods usually use pointwise alignment or structured information alignment between spatial positions, but the channels contain a large amount of knowledge that is ignored during the distillation process. Channel-based knowledge distillation CWD can better utilize the knowledge in each channel, and the activations of the corresponding channels between the teacher and student networks should be small-step adjustments. For this purpose, first, the activations of the channels (the outputs corresponding to layers 16, 19, and 22 in this embodiment) are converted into a probability distribution, so that we can use probability distance metrics, such as KL divergence, to measure the differences. Represent the teacher and student networks as and , and 's activation maps are respectively represented as and , and the general form of the channel distillation loss is expressed as:
[0084] (5)
[0085] Among them, is the KL divergence, which is used to calculate the probability distance metric between two inputs. is the softmax function, which is used to convert the activation map into a probability distribution. and are the activation maps of the teacher network and the student network respectively. and are the activation maps of specific channels in the teacher network and the student network respectively.
[0086] Using To convert the activation values into a probability distribution, as shown in Equation (6):
[0087] (6)
[0088] Where , represents the index channel; represents the spatial position of the index channel. and are the width and height of the channel feature map. is a hyperparameter (temperature). If we use a larger , the probability will become softer, which means we will focus on a wider spatial region of each channel. By applying softmax normalization, we eliminate the scale effect between large networks and compact networks (if the number of channels between the teacher and the student does not match, a 1×1 convolutional layer is used to upsample the number of channels in the student network). The lightweight network framework GH_HGV2-LSCD-YOLOx proposed by the present invention has three scales of networks, namely n, s, and m, according to the number of network layers, the number of parameters, and the computational complexity. The n scale is the smallest, the m scale is the largest, and the s scale is medium. In the embodiment of the present invention, the s model is used as the teacher model and the n model is used as the student model, and the above three-scale division method is well known in the art. Specifically, knowledge distillation is performed on the three feature layers (small, medium, large) (16, 19, 22) extracted from the feature fusion module, as Figure 5 shown.
[0089] IV. Model Training and Detection
[0090] In this embodiment, a total of 2,800 public industrial product anomaly datasets are used, including 4 types of anomalies (hole, Knot, Line, Stain), and they are randomly selected and classified according to the ratio of training set:test set = 4:1. The input size of the images is 640x640. The experiment is conducted on the Ubuntu 20.04.4 operating system, the CPU model is Intel i9-9900k@3.60GHz, the GPU model is NVIDIA GeForce GTX3060Ti graphics card, the video memory is 8G, and the deep learning framework is selected as python-3.8.0, cuda-11.7, pytorch-2.0.1.
[0091] As shown in Table 1, the lightweight model is compared with the original model in terms of model size, inference speed, number of parameters, and mAP value. It can be seen that while the lightweight model improves the Map, although the inference speed decreases, its number of parameters and model size decrease significantly, reducing the computational cost.
[0092] Table 1 Comparison results between the lightweight model and the original model
[0093]
[0094] Another invention, an image anomaly detection system based on a lightweight object detection model is also provided in an embodiment of the present invention, including:
[0095] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the image anomaly detection method based on the lightweight object detection model as described in the above technical solution.
[0096] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. An image anomaly detection method based on a lightweight target detection model, characterized in that: The steps include: Step 1: Build a lightweight target detection model based on the improved YOLOv8 model, including three parts: backbone network, feature fusion module and lightweight shared group convolution detection head; Combining ghost convolution GhostConv and deep convolution DWConv, and integrating with the lightweight backbone network HGNetV2, a new backbone network GH-HGV2Net is constructed to extract features from the input image. The entire new backbone network generates two output feature layers in the middle stage. The processing process of the feature fusion module is as follows: the final output features of the new backbone network are passed through the SPPF module to generate the third output feature layer; the phantom dynamic convolution layer C2fGhostDynaMic is constructed by combining the dynamic convolution DynaMicConv, C2f module and ghost convolution GhostConv; the three output feature layers are multi-level fused through the ghost dynamic convolution layer, upsampling layer, ghost convolution and splicing operations, and finally three feature maps of different sizes are output; Combining group convolution and convolution operations with shared parameters to build a lightweight shared group convolution detection head LSCD for target prediction; Step 2: Select an image containing an anomaly to be detected, mark the position of the anomaly to be detected in the image, and record the coordinate information of the anomaly to form an anomaly detection data set; Step 3: Use the anomaly detection dataset to train the constructed target detection model, and output the trained target detection model to realize image anomaly detection.
2. The image anomaly detection method based on a lightweight target detection model according to claim 1, characterized in that: The new backbone network GH-HGV2Net includes a preliminary extraction module HGStem and 4 stages; The four stages are stage1 to stage4, where stage1 contains a GH-HGBlock, stage2 contains a deep convolution DWConv and a GH-HGBlock, stage3 contains a DWConv and three consecutive GH-HGBlocks, and stage4 is the same as stage2.
3. The image anomaly detection method based on a lightweight target detection model according to claim 2, characterized in that: The preliminary extraction module HGStem includes multiple ordinary convolution and pooling layers.
4. The image anomaly detection method based on a lightweight target detection model according to claim 2, characterized in that: The processing flow of GH-HGBlock is as follows: first, the input features are subjected to n ghost convolution GhostConv operations to generate a series of feature maps; then, these feature maps are spliced through the Concatenate operation to connect the feature maps after each GhostConv operation; next, the spliced feature maps are subjected to the SqueezeConv and ExcitationConv operations to extract the channel information and feature response of the feature maps; finally, according to the condition that the number of channels of the original feature map is the same as the number of channels of the feature map after the intermediate operation, the original input and the feature map generated by the GhostConv and Concatenate operations are added to generate the final feature map.
5. The image anomaly detection method based on a lightweight target detection model according to claim 1, characterized in that: The ghost dynamic convolution layer C2fGhostDynaMic includes dynamic convolution DynamicConv and ghost bottleneck layer GhostBottleNeck; the processing process of C2fGhostDynaMic is as follows: the input feature map first passes through dynamic convolution DynamicConv, and then the output result is divided into two parts. The first part of the result passes through two identical GhostBottleNecks, and then the second part of the feature map after the segmentation is spliced with the feature map after the first GhostBottleNeck and the second GhostBottleNeck operation, and finally passes through a DynamicConv; The processing process of GhostBottleNeck is as follows: the input first passes through the ghost convolution GhostConv, then passes through a batch normalization layer and an activation layer, then passes through GhostConv, and finally adds it to the original feature map to output the final feature map.
6. The image anomaly detection method based on a lightweight target detection model according to claim 1, characterized in that: The specific processing process of the feature fusion module is as follows: The features extracted by the new backbone network are first processed by the SPPF module, followed by an upsample operation, and the upsampled feature map is spliced with the feature map provided by the intermediate stage stage3; then, it passes through a C2fGhostDynaMic module marked as 13 layers, and then upsamples again, and splices the upsampled feature map with the feature map provided by the intermediate stage stage2; at this time, it passes through another C2fGhostDynaMic module marked as 16 layers, and the generated feature map is used to provide to the detection head for detecting small targets; the 16-layer feature map then passes through a GhostConv upsampling operation, and then the upsampled feature map is spliced again with the first spliced feature map, that is, the 13-layer feature map, and then passes through a C2fGhostDynaMic module marked as 19 layers, and provides it to the detection head of the target under detection; in addition, the 19-layer feature map passes through another GhostConv operation, and then spliced with the feature map after SPPF; finally, the spliced feature map passes through a C2fGhostDynaMic module marked as 22 layers and is provided to the detection head for detecting large targets.
7. The image anomaly detection method based on a lightweight target detection model according to claim 1, characterized in that: The processing process of the lightweight shared group convolution detection head LSCD is as follows: First, feature maps of different sizes are processed by group convolution Conv_GN1x1 to reduce the amount of calculation; then, the features processed by Conv_GN1x1 are input into the group convolution of two Conv_GN3x3 with shared parameters to further extract feature information. Then, after the group convolution with shared parameters, the features are divided into three branches for detecting targets of different sizes, including small targets, medium targets and large targets; in each branch, the regression convolution operation Conv_Reg and the classification convolution operation Conv_Cls with shared parameters are used for target regression and classification tasks respectively; at the same time, in each Conv_Reg operation, a trainable scalar parameter scale is introduced to adjust the scale of target regression to adapt to defects of different sizes; The parameter scale is initialized to a positive value of 1.0 and then gradually optimized by gradient descent.
8. The image anomaly detection method based on a lightweight target detection model according to claim 1, characterized in that: Step 1 also includes using LAMP pruning technology to further compress and speed up the lightweight object detection model; In LAMP, the optimization goal of the current layer is to optimize the mask matrix under the constraints. , to minimize the impact of pruning operations on model performance: (1); in, For the current The binary mask matrix of the layer. Each element of this matrix is 0 or 1, which is used to indicate whether the weight at that position is retained or pruned. Will automatically adjust learning; is the zero norm of the mask matrix, representing the number of non-zero weights allowed in the current layer, for The pruning sparsity threshold of the layer is used to constrain the upper limit of non-zero elements in the mask. represents the normalized representation of the input data, is the 2-norm of the input data, used to constrain the scale of the input data, To satisfy Find vectors under conditions The maximum value of Presentation Layer The weight matrix of is element-wise multiplication; The optimization of the entire lightweight target detection model is expressed as: (2); in, is a constant that limits the number of non-zero weights in the model, Lightweight target detection model Using the weight parameter Time input The output, is the weight after sparse For input The output, Indicates the first layer to the The set of all weight matrices of the layer, is the total number of layers of the lightweight object detection model; Under the greedy strategy, the connection with the lowest importance score is removed in each iteration, and the above optimization is relaxed by equation (3): (3); in, is the sparse weight tensor, is the F norm, and the weight tensor Flatten to a one-dimensional vector, sort the one-dimensional vector so that for any constant When, meet , Respectively represent the sorted and vectors; due to the product term is a constant value, so the weight tensor Middle The importance score of an index is defined as the weight tensor Middle The LAMP score for an index is defined as follows: (4); Finally, the connections with the smallest LAMP scores are pruned globally until the required global sparsity constraint is met.
9. The image anomaly detection method based on a lightweight target detection model according to claim 1, characterized in that: Step 1 also includes using knowledge distillation to enhance the lightweight target detection model; dividing the lightweight target detection model into three scales of n, s, and m according to the number of network layers, parameter amount, and computational complexity. The n scale has the smallest number of network layers, parameter amount, and computational complexity, the m scale has the largest scale, and the s scale is medium. The s scale is used as the teacher network and the n scale is used as the student network. To perform knowledge distillation on a specific channel in the feature fusion module, we first convert the activation map of the channel into a probability distribution and use the probability distance metric to measure the difference. The teacher network and the student network are represented as and , and The activation maps are expressed as and , the channel distillation loss is expressed as: (5); in, is the KL divergence, which is used to calculate the probability distance metric between two inputs. is the softmax function, which is used to convert the activation map into a probability distribution; and are the activation maps of the teacher network and the student network, respectively. and They are the activation maps of specific channels in the teacher network and the student network, respectively; use To convert the activation value into a probability distribution, as shown in formula (6): (6); in , Represents the index channel; represents the spatial position of the index channel, and is the width and height of the channel feature map, is a hyperparameter.
10. An image anomaly detection system based on a lightweight target detection model, characterized in that: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the image anomaly detection method based on a lightweight target detection model as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Lightweight fruit maturity detection method and system suitable for embedded equipment
CN119295989A
Method for detecting image target in smart home environment
WO2021244079A1