Method and device for detecting cosmetic defects of pharmaceutical glass tubes
By using multi-scale feature extraction and deep learning algorithms to detect appearance defects in pharmaceutical glass tubes, the problems of low efficiency and false positives/false negatives in manual inspection have been solved, achieving efficient and accurate automated inspection results.
Patent Information
- Application Number
- CN202510755288.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-06-06
AI Technical Summary
In the current technology, the detection of appearance defects in pharmaceutical glass tubes relies on manual visual quality inspection, which is labor-intensive, inefficient, and highly subjective, increasing the probability of false detection and missed detection.
We employ deep learning algorithms based on computer vision to detect appearance defects in pharmaceutical glass tubes through multi-scale feature extraction, channel feature enhancement, and deep feature extraction, combined with a target detection model. We optimize feature correlation by utilizing multi-scale feature maps and self-attention mechanisms, and achieve accurate detection by combining cross-scale feature fusion and target loss functions.
It enables efficient and accurate detection of appearance defects in pharmaceutical glass tubes, reduces labor and time costs, improves the ability to identify minute defects, and ensures the quality stability of automated and intelligent testing.
Smart Images

Figure CN120594545B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of pharmaceutical equipment and quality testing technology, and more specifically, it relates to a method and device for detecting appearance defects in pharmaceutical glass tubes. Background Technology
[0002] Pharmaceutical glass tubing, used as packaging for medicines, directly affects the safety and efficacy of the drugs. However, due to various reasons during the production process, appearance defects in pharmaceutical glass tubing occur frequently, posing certain risks to the use of the medicines. Traditional manual visual quality inspection methods are labor-intensive, inefficient, and highly subjective, increasing the probability of false positives and false negatives. Summary of the Invention
[0003] The purpose of this application is to provide a method and apparatus for detecting appearance defects in pharmaceutical glass tubes, so as to improve the efficiency and accuracy of detecting appearance defects in pharmaceutical glass tubes.
[0004] A first aspect of this application provides a method for detecting appearance defects in pharmaceutical glass tubes, comprising:
[0005] Acquire the target image to be detected, which is an image of a medicine package glass tube;
[0006] The target image to be detected is processed by the target detection model to obtain the target detection result. The target detection result is used to characterize whether there are defects in the pharmaceutical glass tube and the type of defect.
[0007] The target image to be detected is processed by a target detection model to obtain the target detection results, including:
[0008] Multi-scale feature extraction is performed on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales.
[0009] For each scale of feature map, channel feature enhancement processing is performed on the feature map of that scale to obtain the first feature map;
[0010] For each scale of feature map, deep features are extracted from the feature map of that scale from different receptive fields to obtain a second feature map;
[0011] Based on the fusion processing of each first feature map and each second feature map, the fused features corresponding to each scale feature map are obtained. Each first feature map is the first feature map corresponding to each scale feature map, and each second feature map is the second feature map corresponding to each scale feature map.
[0012] Target detection is performed based on the fusion features to obtain the target detection results.
[0013] A second aspect of this application provides a device for detecting appearance defects in pharmaceutical glass tubes, comprising:
[0014] The image acquisition module acquires the target image to be detected, which is an image of the medicine package glass tube.
[0015] The target detection module is used to perform target detection processing on the target image to be detected through a target detection model to obtain target detection results. The target detection results are used to characterize whether there are defects in the pharmaceutical glass tube and the type of defects.
[0016] The target image to be detected is processed by a target detection model to obtain the target detection results, including:
[0017] Multi-scale feature extraction is performed on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales.
[0018] For each scale of feature map, channel feature enhancement processing is performed on the feature map of that scale to obtain the first feature map;
[0019] For each scale of feature map, deep features are extracted from the feature map of that scale from different receptive fields to obtain a second feature map;
[0020] Based on the fusion processing of each first feature map and each second feature map, the fused features corresponding to each scale feature map are obtained. Each first feature map is the first feature map corresponding to each scale feature map, and each second feature map is the second feature map corresponding to each scale feature map.
[0021] The target loss function is used to perform target detection processing on each fused feature to obtain the target detection result; the target loss function is the loss function of the target detection model.
[0022] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for detecting appearance defects in pharmaceutical glass tubes.
[0023] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for detecting appearance defects in pharmaceutical glass tubes.
[0024] The beneficial effects of the method and apparatus for detecting appearance defects in pharmaceutical glass tubes provided in this application are as follows:
[0025] The applicant has discovered that existing technologies struggle to achieve automated and accurate detection of small defects, such as those on pharmaceutical glass tubes, due to their reliance on simplistic, manual visual inspection methods. These methods are labor-intensive, inefficient, and subjective, increasing the probability of false positives and false negatives. Based on this discovery, and unlike existing methods that rely on manual visual inspection, this application processes the target image to obtain different feature maps. These feature maps are then fused, and the combined fused features are used for target detection. This improves the ability to identify small defects in pharmaceutical glass tubes, accurately capturing subtle cracks, tiny bubbles, and other imperceptible surface flaws. While achieving high-precision detection, this application effectively reduces labor and time costs, providing a new technological path for the automated and intelligent detection of surface defects in pharmaceutical glass tubes. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 A flowchart illustrating a method for detecting appearance defects in pharmaceutical glass tubes according to an embodiment of this application;
[0028] Figure 2 Three network structures are provided for a pharmaceutical glass tube appearance defect detection model in one embodiment of this application;
[0029] Figure 3 The network structure of the CBA module provided in one embodiment of this application;
[0030] Figure 4 This is a diagram of the Resnet block structure provided in an embodiment of this application;
[0031] Figure 5 The network structure of the MSDW module provided in one embodiment of this application;
[0032] Figure 6(a) shows the effect of the first RT-DETR detection of pharmaceutical glass tube breakage according to an embodiment of this application;
[0033] Figure 6(b) shows the effect of the first detection model designed in this case, according to an embodiment of this application, on detecting damage to pharmaceutical glass tubes.
[0034] Figure 6(c) shows the effect of the second RT-DETR detection of pharmaceutical glass tube breakage according to an embodiment of this application;
[0035] Figure 6(d) shows the effect of the second detection model designed in this case on detecting damage to pharmaceutical glass tubes according to an embodiment of this application;
[0036] Figure 6(e) shows the effect of the third RT-DETR detection of pharmaceutical glass tube breakage according to an embodiment of this application;
[0037] Figure 6(f) shows the effect of the third detection model designed in this case on detecting damage to pharmaceutical glass tubes according to an embodiment of this application;
[0038] Figure 6(g) is an image showing the effect of the first RT-DETR method for detecting damage and bubbles in pharmaceutical glass tubes according to an embodiment of this application;
[0039] Figure 6(h) shows the effect of the first detection model designed in this case, according to an embodiment of this application, in detecting damage and bubbles in pharmaceutical glass tubes;
[0040] Figure 6(i) shows the effect of a second RT-DETR method for detecting damage and bubbles in pharmaceutical glass tubes according to an embodiment of this application;
[0041] Figure 6(j) shows the detection effect of the second detection model designed in this case according to an embodiment of this application on the breakage and bubbles of pharmaceutical glass tubes;
[0042] Figure 6(k) is a diagram showing the effect of RT-DETR detection of a pharmaceutical glass tube gas line according to an embodiment of this application;
[0043] Figure 6(l) is a diagram showing the detection effect of the detection model designed in this case according to an embodiment of this application on the gas line of the pharmaceutical glass tube;
[0044] Figure 7 A structural block diagram of a device for detecting appearance defects in pharmaceutical glass tubes provided in one embodiment of this application;
[0045] Figure 8 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0046] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0048] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a method for detecting appearance defects in pharmaceutical glass tubes according to an embodiment of this application. The method includes steps S101 and S102.
[0049] S101: Acquire the target image to be detected. The target image to be detected is the image of the medicine package glass tube.
[0050] In this embodiment, the medicine packaged glass tube can be a pre-filled syringe. The system can employ a high-precision area array camera (resolution ≥1392×1040) with an 8mm low-distortion industrial lens to achieve continuous shooting capability of 30 frames per second, capturing images of the medicine packaged glass tube to be inspected on the production line. At least three sets of visual imaging structures can be set on both sides of the production line, evenly distributed around the center of the glass tube in a 360° radius, ensuring circumferential full-coverage imaging of the medicine packaged glass tube. For example, each imaging unit may include:
[0051] Camera module: Can be equipped with an adjustable gimbal, supporting ±15° tilt angle adjustment, and adaptable to glass tubes of different diameters. For example, the diameter of the pre-filled syringe is 50-120mm.
[0052] Backlight system: A red diffuse high-brightness light source can be used, and stray light interference can be eliminated by using a dark-colored aluminum base. The light source is arranged perpendicular to the lens axis to enhance the contrast of defects such as bubbles and scratches on the glass tube surface.
[0053] S102: The target image to be detected is processed by the target detection model to obtain the target detection result. The target detection result is used to characterize whether there are defects in the pre-filled syringe and the type of defect.
[0054] In this embodiment, the target detection model can be a deep learning algorithm based on computer vision, used to automatically identify and locate specific target objects in an image. In the pre-filled syringe detection scenario, the model uses a convolutional neural network to extract and classify multi-scale features of key parts of the target image. Its core functions may include: determining whether there are defects in the image and marking the location of defects using bounding boxes, etc.
[0055] In this embodiment, the defect types can be air bubble defects caused by incomplete vacuum plugging, needle breakage caused by improper vacuum plugging height or center positioning plate offset, and the presence of stones and foreign objects caused by peristaltic pump failure or nitrogen pressure fluctuations.
[0056] Specifically, S102 may include: S1021-S1025.
[0057] S1021: Perform multi-scale feature extraction on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales.
[0058] The target image acquired by the camera is input into the backbone network to extract multi-scale features, resulting in a multi-scale feature map. The backbone network can be a ResNet-18 network, a lightweight member of the residual network family, consisting of 18 layers, including 17 convolutional layers and 1 fully connected layer. The multi-scale features are obtained through residual blocks S3-S5, yielding primary feature S3, intermediate feature S4, and high-level feature S5. The residual block is the core component of the residual network; it directly superimposes the input to the output through skip connections, allowing the network to learn residuals and solving the gradient vanishing and degradation problems of deep networks.
[0059] In this embodiment, the Attention-based Intra-scale Feature Interaction (AIFI) module can be selected to perform intra-scale interaction on the high-level feature S5. AIFI is a core component in the Real-Time Detection Transformer (RT-DETR) model used to enhance multi-scale feature interaction. Its core design concept is to optimize feature correlation within a single scale through a self-attention mechanism, while working in conjunction with the cross-scale feature fusion module. The result F5 is obtained by performing intra-scale interaction on the high-level feature S5 using the following formula.
[0060]
[0061]
[0062] Flatten flattens the spatial dimensions of S5 into a sequence form, which meets the input requirements of Transformer for processing sequence data. , , These are the query, key, and value vectors, respectively, all with the same initial value, reflecting the characteristics of the self-attention mechanism. Attn represents multi-head attention, Reshape represents reshaping the data to the same shape as S5, and CCDFM represents cross-scale deep fusion of features.
[0063] S1022: For each scale of feature map, perform channel feature enhancement processing on the feature map of that scale to obtain the first feature map.
[0064] In this embodiment, the input multi-scale feature map can be processed using 1×1 convolution and dynamic adaptive normalization to enhance channel features, resulting in a first feature map. The normalization process can employ an adaptive hyperbolic tangent function. This adaptive hyperbolic tangent function is a standard... An improved version of the function achieves dynamic adjustment of nonlinear activation by introducing learnable parameters. The final output first feature map possesses both high semantic meaning and detail fidelity, providing high-quality input for subsequent object detection heads.
[0065] S1023: For each scale of feature map, perform deep feature extraction from different receptive fields to obtain a second feature map.
[0066] In this embodiment, the input multi-scale feature map is used to extract deep features in different receptive fields. The steps may include:
[0067] First, residual blocks containing 3×3, 5×5, and 7×7 convolutional kernels can be used to extract features from the differentiated receptive field. Second, average pooling can be used to preserve overall statistical information across multiple output channels. Finally, a nonlinear Gaussian error function can be employed. Activation yields the second feature map.
[0068] S1024: Based on each first feature map and each second feature map, perform fusion processing to obtain the fused features corresponding to each scale feature map respectively. Each first feature map is the first feature map corresponding to each scale feature map respectively, and each second feature map is the second feature map corresponding to each scale feature map respectively.
[0069] In this embodiment, spatial attention modeling can be used to allocate attention weights element-wise along the spatial dimension, thereby strengthening the response intensity of key regions. This spatial attention modeling can be achieved by performing Hadamard product on each first feature map and each second feature map. The Hadamard product is an element-wise multiplication operation in matrix operations, widely used in deep learning and image processing.
[0070] In this embodiment, 1×1 convolution can be used to integrate cross-channel information, and residual connections can be used to additively fuse the fused features with the original input.
[0071] S1025: Perform target detection processing on each fused feature based on the target loss function to obtain the target detection result; the target loss function is the loss function of the target detection model.
[0072] In this embodiment, based on the fusion features obtained in the above steps, the Intersection over Union (IoU) perceptual query feature selection module and the transform decoder with an auxiliary prediction detection head can be used to perform target detection processing to obtain the target detection result.
[0073] In this embodiment, the IoU-aware query feature selection is primarily responsible for resolving the inconsistency between classification and localization. Its core function is to select candidate features from various fused features that simultaneously possess high classification confidence and a high intersection-over-union ratio (IoU), serving as the initial object query for the decoder. Traditional DETR models use randomly initialized learnable parameters for object queries, leading to a mismatch between classification scores and localization accuracy. For example, "false positive" boxes may have high classification scores but low IoU. The IoU-aware query, by constraining the training objective, forces the model to assign high classification scores to features with high IoU, thereby selecting a more reliable initial query.
[0074] In this embodiment, the transform decoder with an auxiliary prediction detection head is mainly responsible for accelerating convergence and optimizing iterations. This decoder refines the bounding boxes and class predictions of object queries layer by layer through a multi-layer iterative structure. Each layer of the decoder receives the output of the previous layer and fuses the global context through a self-attention mechanism.
[0075] For example, auxiliary prediction heads accelerate training by adding them after each decoder layer, using intermediate supervision signals to speed up model convergence. These heads provide gradient feedback during training and can be selectively removed during inference to reduce latency.
[0076] As can be concluded from the above, the method for detecting appearance defects in pharmaceutical packaging glass tubes proposed in this invention achieves accurate identification and classification of appearance defects in pre-filled syringes. Compared with traditional manual inspection methods, this method greatly improves the efficiency of appearance quality inspection and product quality stability, effectively meeting the needs of large-scale production. At the same time, it overcomes the problems of strong subjectivity and fatigue leading to missed or false detections in manual inspection, ensuring the stability of product quality. This is of great significance for promoting the high-quality development of the pharmaceutical packaging industry.
[0077] In one embodiment of this application, for each scale of feature map, the method further includes:
[0078] The feature map at this scale is reconstructed across channels to obtain the first reconstructed feature.
[0079] Based on the features after the first recombination, the hyperbolic tangent processing result is determined using the hyperbolic tangent function;
[0080] Based on the hyperbolic tangent processing results and the features after the first recombination, the feature map after normalization is determined;
[0081] The feature map at this scale undergoes channel feature enhancement processing, including:
[0082] Channel feature enhancement processing is performed on the normalized feature map corresponding to the feature map at this scale;
[0083] This includes deep feature extraction from feature maps of this scale using different receptive fields, including:
[0084] Deep features are extracted from the normalized feature maps corresponding to the feature maps at this scale using different receptive fields.
[0085] In this embodiment, cross-channel feature recombination of the feature map at this scale can be performed using 1×1 convolution to combine features from different input channels, achieving cross-channel information interaction and thus obtaining the first recombined feature. This first recombined feature maintains the same spatial size as the original image, only changing the channel dimension. 1×1 convolution can also reduce the computational complexity of normalization.
[0086] For example, if the input is RGB three channels, a 1×1 convolution can generate new channels, combining features such as color and texture.
[0087] In this embodiment, an adaptive hyperbolic tangent function can be used for normalization to obtain a normalized feature map. This normalized feature map compresses extreme values non-linearly and transforms the central part of the input almost linearly, preserving the core effect of normalization. This allows the network to focus on learning high-order feature associations rather than low-order statistical biases, automatically weakening the influence of outliers, preventing feature degradation of small target features due to gradient vanishing during sequence processing, and improving the performance of the neural network in complex data processing.
[0088] In this embodiment, the normalized feature map is obtained using the following formula.
[0089]
[0090]
[0091] Here, Tanh represents the hyperbolic tangent function. It is a learnable scalar parameter used to dynamically adjust the scaling range of the input and control the nonlinearity. and These are learnable vector parameters that allow the output to be scaled back to an appropriate range.
[0092] For example, in practical implementation, parameters It is usually initialized to 0.5, parameter Initialize as a vector of all 1s, parameters Initialize as a vector consisting entirely of zeros.
[0093] In this embodiment, the first feature enhancement path is a cross-channel feature recombination path, which performs channel feature enhancement processing on the normalized feature map.
[0094] Specifically, the normalized feature map can be recombined across channels using 1×1 convolution. Secondly, channel attention maps can be generated through activation functions, thus achieving channel feature enhancement processing.
[0095] In this embodiment, the second feature enhancement path is a multi-scale deep feature enhancement path, which extracts deep features from the normalized feature maps corresponding to the feature maps of that scale from different receptive fields.
[0096] Specifically, firstly, three sets of depthwise separable convolutions can be deployed in parallel to perform convolution operations on the normalized feature map. Here, the kernel sizes of the multi-scale convolutions are 3×3, 5×5, and 7×7, respectively. Secondly, average pooling (Avg) is performed on the multi-scale output channels to generate fused features. Then, the normalized feature map and the fused feature map are... Residual connections are performed to enhance the stability of gradient flow. Then, the feature data after residual connection is recombined across channels through 1×1 convolution to complete cross-channel information interaction. Finally, the feature data is activated by the nonlinear Gaussian error function GeLU to complete the deep feature extraction of the normalized feature map.
[0097] As can be seen from the above, cross-channel feature recombination is achieved through 1×1 convolution, and the normalization parameters are dynamically adjusted by the adaptive hyperbolic tangent function, which effectively suppresses outliers and preserves the linear response of the feature center region, thereby enhancing the gradient stability of small targets. At the same time, a dual-path collaborative enhancement strategy is adopted, which dynamically allocates weights in the channel dimension through the attention mechanism and captures local details and global semantics in the spatial dimension through multi-scale depthwise separable convolution. Combined with residual connections and GeLU activation function to optimize gradient flow, efficient and accurate detection of appearance defects in pharmaceutical glass tubes can be achieved.
[0098] In one embodiment of this application, channel feature enhancement processing is performed on the feature map at this scale, including:
[0099] Cross-channel feature recombination is performed on the feature map at this scale to obtain the second recombined feature map; based on the second recombined feature map, channel feature enhancement processing is performed through activation function to generate a channel attention map, which serves as the first feature map corresponding to the feature map at this scale.
[0100] In this embodiment, the feature map at this scale can be reconstructed across channels using a cross-channel feature reconstruction path to obtain a second reconstructed feature. The feature map at this scale is the feature map after feature normalization processing.
[0101] Specifically, a 1×1 convolution can be used to reconstruct cross-channel features from the feature map at this scale. Secondly, an activation function is used to generate channel attention maps. Specifically, the activation function here... The Sigmoid-Weighted LinearUnit (SiLU) activation function is chosen for nonlinear mapping to enhance channel features and compensate for potential loss of content due to the order constraints of the State Space Model (SSM). The order constraints of the SSM are essentially the result of both its mathematical model and hardware implementation; mathematically, they must adhere to recursive dynamics and frequency domain alignment, while hardware implementation is limited by memory access and parallel strategies. Secondly, the activation function... The core idea is to use the product of the input value and its Sigmoid activation value as the output, mathematically expressed as: ,in First feature map Depend on Received, among which This is the feature map at this scale.
[0102] From the above, we can conclude that, firstly, 1×1 convolution does not change the height and width of the feature map, only adjusting the channel dimension, thus avoiding spatial information loss and making it suitable for multi-scale feature fusion. Secondly, at each spatial location, 1×1 convolution is equivalent to a fully connected layer performing a linear transformation on the channels, but significantly reducing the number of parameters through a weight sharing mechanism. Finally, compared to other activation functions, SiLU provides smooth gradients when features are close to zero, which is beneficial for improving the model's performance and generalization ability, and its continuous differentiability is suitable for channel recombination tasks.
[0103] In one embodiment of this application, deep feature extraction is performed on the feature map at this scale from different receptive fields to obtain various second feature maps, including:
[0104] The feature maps at this scale are subjected to depthwise separable convolutions using different convolution kernels to obtain their respective depthwise convolution features.
[0105] The corresponding convolutional features are subjected to channel-level pooling to obtain fused features;
[0106] The fused features and the feature map at this scale are subjected to residual connection processing to obtain the residual-processed features;
[0107] The features after residual processing are recombined across channels to obtain the third recombined features;
[0108] The features after the third reorganization are activated by a nonlinear Gaussian error function to obtain multi-scale deep features.
[0109] In this embodiment, different receptive fields can be used to perform depthwise separable convolution by using different convolution kernels for deep feature extraction of feature maps at this scale.
[0110] Specifically, multiple sets of depthwise separable convolutions are deployed in parallel, with multi-scale convolution kernel sizes of 3×3, 5×5, and 7×7. The 3×3 depthwise convolution kernel extracts local texture features (such as scratches or bubbles in pharmaceutical glass tubes), the 5×5 depthwise convolution kernel captures medium-range contextual information (such as deformation of pharmaceutical glass tubes), and the 7×7 depthwise convolution kernel perceives global structural features (such as air lines in pharmaceutical glass tubes). The depthwise separable convolution decomposes spatial filtering and channel projection, reducing computational complexity.
[0111] In this embodiment, Avg can be used to fuse the multi-scale outputs of multiple depthwise separable convolutions to obtain fused features. .
[0112] Specifically, Avg compresses features extracted from feature maps of different scales, such as those from 3×3, 5×5, and 7×7 convolutional kernels, into global statistics along the channel dimension, eliminating local noise interference and highlighting common features across different channels. For example, in the detection of defects in pharmaceutical glass tubes, multi-scale features may contain local details of cracks (small scale) and overall morphology (large scale), and Avg can fuse these complementary pieces of information.
[0113] In this embodiment, By constructing residual connections with the feature maps at this scale, we obtain the residual connection feature data. This operation can alleviate gradient vanishing and thus enhance the stability of gradient flow.
[0114] In this embodiment, the feature data after residual concatenation can be recombined across channels through 1×1 convolution to complete cross-channel information interaction and obtain the third recombined feature. Finally, the third recombined feature can be activated by the nonlinear Gaussian error function GeLU to obtain the second feature map. The specific process is expressed by the following formula: ,in, This is the second feature map. These are the weights of a depthwise separable convolution. It is the input feature dimension. It is the output feature dimension. This is the kernel size, where i = 1, 2, 3; The sizes are set to 3×3, 5×5, and 7×7 respectively.
[0115] As can be seen from the above, this embodiment significantly improves the model's sensitivity to minor defects and its robustness to industrial scenarios by recombining cross-channel features, while reducing computational complexity.
[0116] In one embodiment of this application, for each scale of feature map, a fusion process is performed based on the first feature map and the second feature map corresponding to that scale to obtain the fused feature corresponding to that scale feature map, including:
[0117] Determine the Hadamard product of the first feature map corresponding to the feature map at this scale and the second feature map corresponding to the feature map at this scale;
[0118] Based on the Hadamard product, cross-channel feature recombination is performed to obtain the fourth recombined feature.
[0119] The fused features are obtained by fusing the features after the fourth reorganization with the feature map at this scale.
[0120] In this embodiment, the first feature map and the second feature map can be fused through dual-path parallel feature interaction to obtain the fused feature corresponding to the feature map at this scale.
[0121] Specifically, the dual-path parallel feature interaction consists of two steps: spatial attention modeling and channel information fusion. Spatial attention modeling can be achieved by performing a Hadamard product on the first and second feature maps, as shown in the formula: This spatial attention modeling can assign element-wise attention weights across spatial dimensions, enhancing the response intensity of key regions. Channel information fusion utilizes 1×1 convolutions to integrate cross-channel information, and residual connections are used to additively fuse the fused features with the original input. The formula is as follows: .in, This is a feature of fusion.
[0122] As can be seen from the above, parallel use of depthwise separable convolutions with different dilation rates can capture local details, small targets, and multi-scale spatial features. By using residual connections of depthwise separable convolution kernels, the ability to extract local spatial information is enhanced while reducing computational costs, improving the model's ability to detect targets at different scales, and enhancing the model's robustness to scale changes.
[0123] In one embodiment of this application, target detection processing is performed based on various fusion features to obtain target detection results, including:
[0124] The fused features are then concatenated to obtain the concatenated features.
[0125] The target detection results are obtained by performing target detection processing based on the features after connection processing.
[0126] In this embodiment, different defects correspond to different features after concatenation processing. Specifically, when dealing with bubble defects in medicine packaging glass tubes, since bubbles usually appear as circular or elliptical areas with blurred edges and low grayscale values in images, the features after concatenation processing help the target detection module quickly extract suspected bubble areas. Combined with morphological processing methods, noise interference is removed, and finally, the bubble defect is accurately located.
[0127] For crack defects, because they are characterized by being thin and long, with sharp edges and significant changes in grayscale values in the image, the processed features can help the target detection module quickly identify the direction and length of the crack, thereby completing the detection of crack defects.
[0128] When dealing with scratch defects, since they appear as linear features in images, by analyzing parameters such as the angle and length of the lines, the connected features can help the target detection module to accurately identify and classify scratch defects, ultimately obtaining comprehensive and accurate target detection results.
[0129] As can be seen from the above, the features after connection processing can be used to target different defects: bubbles are blurry, low-grayscale circles, and features help to quickly extract suspected areas and combine them with morphological localization; cracks are thin, sharp, and have obvious grayscale changes, and features help to identify the direction and length; scratches are linear, and features help to analyze parameters to achieve accurate classification, significantly improving the efficiency and accuracy of defect detection.
[0130] In one embodiment of this application, a sample image, a target bounding box corresponding to each target in the sample image, and the target probability that each target in the sample image belongs to a preset target are obtained;
[0131] Based on the sample image, the target bounding box corresponding to each target in the sample image, and the target probability that each target in the sample image belongs to the preset target, the initial model is trained to obtain the target detection model;
[0132] The methods for training the initial model include:
[0133] The sample image is used to perform target detection through the initial model to obtain the predicted detection result, which includes the prediction box corresponding to each target and the predicted probability that each target belongs to the preset target.
[0134] The first loss is determined based on the intersection-union ratio (IUU) between the predicted bounding box and the labeled bounding box corresponding to each target.
[0135] The second loss is determined based on the difference between the predicted probability that each target belongs to a preset target and the target probability that the target belongs to the preset target;
[0136] Based on the first loss and the second loss, the initial model is trained.
[0137] In this embodiment, the sample image can be an image of a historical medicinal glass tube collected by a camera, and can be image data of each part, including images under different angles and different lighting conditions.
[0138] The target annotation box is a rectangular box used to mark the defect position in the image, and is usually represented by a bounding box. Taking the bubble defect as an example, the annotator will outline the edge contour of the bubble in the sample image. The target probability can be the confidence that the target belongs to a preset defect category, and the value range is [0,1].
[0139] In this embodiment, a large amount of labeled pre-filled syringe image data is used to train the constructed model by constructing a loss function. For the positioning loss of the predicted box coordinates, the WIoUv3 loss function of the RT-DETR framework can be used, and a matching quality-aware loss is designed as the classification loss of the predicted category, and its expression is , which is the target loss function. Among them, represents the true label category, represents the predicted probability of the foreground category, represents the IoU between the predicted bounding box and the target box, and the parameter adjusts the sensitivity of the matching quality to the loss weight allocation, and is used to control the balance between high-quality matching samples and low-quality matching samples.
[0140] Specifically, a parameter is introduced into the IoU, so that samples with a higher IoU obtain exponentially larger weights, strengthen the gradient signal of high-quality predictions, and accelerate the learning of high-quality predictions;
[0141] The positive sample gradient is proportional to , and compared with the original VFL loss, the model pays more attention to samples with accurate positioning; the negative sample gradient is inversely proportional to , and automatically focuses on difficult samples.
[0142] During the training process, a progressive parameter tuning strategy is adopted for the parameter : initially set base_gammma = 2.0; in the initial stage of training (current_epoch < max_epoch / 3), the parameter is set to base_gammma × 0.7; in the middle stage (current_epoch < max_epoch*2 / 3), the parameter is set to base_gammma × 0.8, and in the later stage, the parameter Set to base_gammma to enhance high-quality samples. Here, current_epoch is the current training epoch, and max_epoch is the maximum training epoch.
[0143] As can be seen from the above, during the training process, continuously adjusting the model parameters, including the learning rate, training batch size, number of training iterations, and optimizing the loss function, can improve the model's detection accuracy and generalization ability.
[0144] In one embodiment of this application, a network structure for a defect detection model of pharmaceutical glass tubes is provided.
[0145] like Figure 2 (a) The acquired image data is processed by the backbone network to obtain primary features S3, intermediate features S4, and high-level features S5. These three different features are then input into the CCDFM module for deep feature extraction. Finally, an Intersection over Union (IoU) perceptual query selection feature module and a transform decoder with an auxiliary prediction detection head are used for target detection processing to obtain the target detection result. Figure 2 (b) The intermediate steps are omitted; only the main structure of the network structure of the pharmaceutical glass tube appearance defect detection model is described. Compared to... Figure 2 (b) Figure 2 (c) The AIFI module is highlighted to demonstrate its importance in performing intrascale interactions on high-level feature S5.
[0146] Figure 3 The network structure of the CBA module performs 3×3 convolution, batch normalization, and activation functions on the input data one at a time, and then gradually extracts semantic features from low to high level. Figure 4 It is a residual network, which consists of pooling layers, 1×1 convolutional layers and parallel CBA modules. The overall design aims to achieve feature compression, cross-channel interaction and multi-branch feature fusion. Figure 5 This describes the network structure of the MSDW module. It illustrates a dual-path parallel feature interaction architecture. Path 1 first uses 1×1 convolutions to achieve cross-channel feature recombination, followed by the generation of channel attention maps through an activation function. Path 2 first deploys multiple sets of depthwise separable convolutions in parallel, with multi-scale convolution kernel sizes of 3×3, 5×5, and 7×7. Then, channel-level average pooling is performed on the multi-scale outputs to generate fused features F. fusion Then, the fused features are used to construct residual connections with the initial input features. Finally, cross-channel feature recombination is achieved through 1×1 convolution to complete cross-channel information interaction, and the feature is activated by the nonlinear Gaussian error function GeLU.
[0147] In one embodiment of this application, an experimental verification is provided. The experimental environment is based on the Ubuntu 18.04 operating system, and the graphics card is an NVIDIA V100. The experiment basically adopts the parameter settings recommended by RT-DETR, and uses data augmentation and other strategies for model training. Each training iteration uses 4 images as input and is repeated 300 times.
[0148] 5352 images of pre-filled syringes from different angles and under different lighting conditions were collected and preprocessed. The appearance defects of the pre-filled syringes were classified into six categories: air lines, stones, stains, scratches, damage, and bubbles, represented by the numbers 0, 1, 2, 3, 4, and 5, respectively. The dataset was divided into training, validation, and test sets according to a 7:2:1 random sampling principle. The constructed model was then trained. After training, the test set images were inspected, and the evaluation metrics were average accuracy (mAP) and frame rate (FPS, frames per second).
[0149] Table 1 presents the test results of different detection algorithms on the pre-filled packaging dataset. Compared with the benchmark RT-DETR algorithm, the method of this invention shows a significant improvement in accuracy, with a 3.3% increase in mAP and a frame rate (FPS) of 92 frames per second. Although the running speed is slightly slower than the benchmark model, it can still meet the speed requirements of the pre-filled packaging production line.
[0150]
[0151] Figures 6(a), 6(c), 6(e), 6(g), 6(i), and 6(k) show the original RT-DETR detection results, while Figures 6(b), 6(d), 6(f), 6(h), 6(j), and 6(l) show the detection results of the detection model designed in this case. It can be observed that both methods can correctly detect larger defects, such as the stain defects in Figures 6(a), 6(b), 6(g), and 6(h), and the breakage and gas line defects in Figures 6(i), 6(j), 6(k), and 6(l). However, for smaller or less noticeable defects, the baseline model did not detect them correctly, such as the small stain at the bottle mouth in Figure 6(a) and the scratch at the bottom of the bottle in Figure 6(e). Specifically, firstly, the improved RT-DETR in this case detected small targets or defective targets with indistinct features that were missed by the RT-DETR method; secondly, the improved scheme significantly improved the confidence level of the detection results; and finally, the improved scheme was more accurate in locating the targets.
[0152] The detection algorithm and system for detecting appearance defects in pre-filled syringes provided by this invention have the advantages of high detection efficiency, high accuracy, and high level of automation. They can significantly improve the efficiency and quality of appearance defect detection in pre-filled syringes and have broad application prospects.
[0153] A method for detecting appearance defects in pharmaceutical glass tubes, corresponding to the above embodiment. Figure 7 This is a structural block diagram of a device for detecting appearance defects in pharmaceutical glass tubes according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 7 The pharmaceutical glass tube appearance defect detection device 20 includes: an image acquisition module 21 and a target detection module 22.
[0154] Image acquisition module 21 is used to acquire the target image to be detected, which is the image of the medicine package glass tube;
[0155] The target detection module 22 is used to perform target detection processing on the target image to be detected through the target detection model to obtain the target detection result. The target detection result is used to characterize whether there are defects in the pharmaceutical glass tube and the type of defect.
[0156] The target image to be detected is processed by a target detection model to obtain target detection results, including:
[0157] Multi-scale feature extraction is performed on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales.
[0158] For each scale of feature map, channel feature enhancement processing is performed on the feature map of that scale to obtain the first feature map;
[0159] For each scale of feature map, deep features are extracted from the feature map of that scale from different receptive fields to obtain a second feature map;
[0160] Based on the fusion processing of each first feature map and each second feature map, the fused features corresponding to each scale feature map are obtained. Each first feature map is the first feature map corresponding to each scale feature map, and each second feature map is the second feature map corresponding to each scale feature map.
[0161] Target detection is performed based on the fusion features to obtain the target detection results.
[0162] In one embodiment of this application, the target detection module 22 is further configured to:
[0163] For each scale of feature map, cross-channel feature recombination is performed on the feature map of that scale to obtain the first recombined feature;
[0164] Based on the features after the first recombination, the hyperbolic tangent processing result is determined using the hyperbolic tangent function;
[0165] Based on the hyperbolic tangent processing results and the features after the first recombination, the feature map after normalization is determined;
[0166] The feature map at this scale undergoes channel feature enhancement processing, including:
[0167] Channel feature enhancement processing is performed on the normalized feature map corresponding to the feature map at this scale;
[0168] This includes deep feature extraction from feature maps of this scale using different receptive fields, including:
[0169] Deep features are extracted from the normalized feature maps corresponding to the feature maps at this scale using different receptive fields.
[0170] In one embodiment of this application, the target detection module 22 is specifically used for:
[0171] Cross-channel feature recombination is performed on the feature map at this scale to obtain the second recombined feature map; based on the second recombined feature map, channel feature enhancement processing is performed through activation function to generate a channel attention map, which serves as the first feature map corresponding to the feature map at this scale.
[0172] In one embodiment of this application, the target detection module 22 is specifically used for:
[0173] The feature maps at this scale are subjected to depthwise separable convolutions using different convolution kernels to obtain their respective depthwise convolution features.
[0174] The corresponding convolutional features are subjected to channel-level pooling to obtain fused features;
[0175] The fused features and the feature map at this scale are subjected to residual connection processing to obtain the residual-processed features;
[0176] The features after residual processing are recombined across channels to obtain the third recombined features;
[0177] The features after the third reorganization are activated by a nonlinear Gaussian error function to obtain multi-scale deep features.
[0178] In one embodiment of this application, the target detection module 22 is specifically used for:
[0179] Determine the Hadamard product of the first feature map corresponding to the feature map at this scale and the second feature map corresponding to the feature map at this scale;
[0180] Based on the Hadamard product, cross-channel feature recombination is performed to obtain the fourth recombined feature.
[0181] The fused features are obtained by fusing the features after the fourth reorganization with the feature map at this scale.
[0182] In one embodiment of this application, the target detection module 22 is specifically used for:
[0183] Each fused feature is connected to obtain the connected features.
[0184] The target detection results are obtained by performing target detection processing based on the features after connection processing.
[0185] In one embodiment of this application, the target detection module 21 is specifically used for:
[0186] Obtain a sample image, the bounding box corresponding to each target in the sample image, and the probability that each target in the sample image belongs to a preset target;
[0187] Based on the sample image, the target bounding box corresponding to each target in the sample image, and the target probability that each target in the sample image belongs to the preset target, the initial model is trained to obtain the target detection model;
[0188] The methods for training the initial model include:
[0189] The sample image is used to perform target detection through the initial model to obtain the predicted detection result, which includes the prediction box corresponding to each target and the predicted probability that each target belongs to the preset target.
[0190] The first loss is determined based on the intersection-union ratio (IUU) between the predicted bounding box and the labeled bounding box corresponding to each target.
[0191] The second loss is determined based on the difference between the predicted probability that each target belongs to a preset target and the target probability that the target belongs to the preset target;
[0192] The initial model is trained based on the first loss and the second loss.
[0193] See Figure 8 , Figure 8 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 8The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above-described device embodiments, for example... Figure 7 The functions of the image acquisition module 21 and the target detection module 22 shown are illustrated.
[0194] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0195] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0196] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.
[0197] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the first and second embodiments of the method for detecting appearance defects in pharmaceutical glass tubes provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.
[0198] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0199] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0200] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0201] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0202] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0203] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0204] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0205] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for detecting appearance defects in pharmaceutical glass tubes, characterized in that, include: Acquire a target image to be detected, wherein the target image to be detected is a medicine package glass tube image; The target image to be detected is processed by a target detection model to obtain target detection results, which are used to characterize whether the pharmaceutical glass tube has defects and the type of defects. The step of performing target detection processing on the target image to be detected using a target detection model to obtain target detection results includes: Multi-scale feature extraction is performed on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales. For each scale of feature map, channel feature enhancement processing is performed on the feature map of that scale to obtain the first feature map; For each scale of feature map, deep features are extracted from the feature map of that scale from different receptive fields to obtain a second feature map; Based on the fusion processing of each first feature map and each second feature map, the fused features corresponding to each scale feature map are obtained. The first feature map is the first feature map corresponding to each scale feature map, and the second feature map is the second feature map corresponding to each scale feature map. The target detection result is obtained by performing target detection processing on each fused feature based on the target loss function; the target loss function is the loss function of the target detection model. For each scale of feature map, the method further includes: The feature map at this scale is reconstructed across channels to obtain the first reconstructed feature. Based on the features of the first recombination, the hyperbolic tangent processing result is determined using the hyperbolic tangent function; Based on the hyperbolic tangent processing result and the first recombined features, the normalized feature map is determined; The process of enhancing the channel features of the feature map at this scale includes: Channel feature enhancement processing is performed on the normalized feature map corresponding to the feature map at this scale; The step of extracting deep features from the feature map of this scale from different receptive fields includes: Deep feature extraction is performed on the normalized feature maps corresponding to the feature maps of the specified scale from different receptive fields.
2. The method according to claim 1, characterized in that, The channel feature enhancement processing of the feature map at this scale includes: Cross-channel feature recombination is performed on the feature map at this scale to obtain a second recombined feature map; based on the second recombined feature map, channel feature enhancement processing is performed through activation function to generate a channel attention map, which serves as the first feature map corresponding to the feature map at this scale.
3. The method according to claim 1, characterized in that, The process of extracting deep features from feature maps of this scale from different receptive fields to obtain various second feature maps includes: The feature maps at this scale are subjected to depthwise separable convolutions using different convolution kernels to obtain their respective depthwise convolution features. The corresponding convolutional features are then subjected to channel-level pooling to obtain fused features. The fused features and the feature map at this scale are subjected to residual connection processing to obtain the residual-processed features; The features after residual processing are recombined across channels to obtain the third recombined features; The third recombined features are activated using a nonlinear Gaussian error function to obtain multi-scale depth features.
4. The method according to claim 1, characterized in that, For each scale of feature map, a fusion process is performed based on the first and second feature maps corresponding to that scale to obtain the fused features corresponding to that scale feature map, including: Determine the Hadamard product of the first feature map corresponding to the feature map at this scale and the second feature map corresponding to the feature map at this scale; Based on the Hadamard product, cross-channel feature recombination is performed to obtain the fourth recombined feature; The features after the fourth recombination and the feature map at the same scale are fused together to obtain the fused features corresponding to the feature map at that scale.
5. The method according to any one of claims 1-4, characterized in that, The target detection processing based on each fusion feature to obtain the target detection result includes: The fused features are then concatenated to obtain the concatenated features. The target detection result is obtained by performing target detection processing based on the features after the connection processing.
6. The method according to claim 1, characterized in that, The process of constructing the target loss function includes: Obtain a sample image, the bounding box corresponding to each target in the sample image, and the probability that each target in the sample image belongs to a preset target; Based on the sample image, the target bounding box corresponding to each target in the sample image, and the target probability that each target in the sample image belongs to a preset target, the initial model is trained to obtain the target detection model; The methods for training the initial model include: The sample image is subjected to target detection through the initial model to obtain a predicted detection result, which includes a prediction box corresponding to each target and a predicted probability that each target belongs to a preset target. The first loss is determined based on the intersection-union ratio (IUU) between the predicted bounding box and the labeled bounding box corresponding to each target; The second loss is determined based on the difference between the predicted probability that each target belongs to a preset target and the target probability that the target belongs to the preset target; The initial model is trained based on the first loss and the second loss.
7. A device for detecting appearance defects in pharmaceutical glass tubes, characterized in that, include: The image acquisition module is used to acquire the target image to be detected, wherein the target image to be detected is an image of a medicine package glass tube; The target detection module is used to perform target detection processing on the target image to be detected through a target detection model to obtain target detection results. The target detection results are used to characterize whether there are defects in the pharmaceutical glass tube and the type of defects. The step of performing target detection processing on the target image to be detected using a target detection model to obtain target detection results includes: Multi-scale feature extraction is performed on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales. For each scale of feature map, channel feature enhancement processing is performed on the feature map of that scale to obtain the first feature map; For each scale of feature map, deep features are extracted from the feature map of that scale from different receptive fields to obtain a second feature map; Based on the fusion processing of each first feature map and each second feature map, the fused features corresponding to each scale feature map are obtained. The first feature map is the first feature map corresponding to each scale feature map, and the second feature map is the second feature map corresponding to each scale feature map. Target detection results are obtained by performing target detection processing based on various fusion features; For each scale of feature map, the object detection module is also used for: The feature map at this scale is reconstructed across channels to obtain the first reconstructed feature. Based on the features of the first recombination, the hyperbolic tangent processing result is determined using the hyperbolic tangent function; Based on the hyperbolic tangent processing result and the first recombined features, the normalized feature map is determined; The process of enhancing the channel features of the feature map at this scale includes: Channel feature enhancement processing is performed on the normalized feature map corresponding to the feature map at this scale; The step of extracting deep features from the feature map of this scale from different receptive fields includes: Deep feature extraction is performed on the normalized feature maps corresponding to the feature maps of the specified scale from different receptive fields.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Semiconductor wafer defect detection light source configuration method, device and equipment
CN119445046A
Steel sheet scraper defect detection method, storage medium, electronic equipment and device
CN119693339A