Medicinal glass tube appearance defect detection method and device

Through multi-scale feature extraction and deep learning algorithms, the problem of low manual detection efficiency in the appearance defect detection of pharmaceutical glass tubes is solved, and efficient and accurate automated detection is achieved.

CN120594545AActive Publication Date: 2025-09-05INST OF APPLIED MATHEMATICS HEBEI ACADEMY OF SCI

Patent Information

Application Number
CN202510755288.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-05
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing technology of traditional Chinese medicine glass tube appearance defect detection relies on artificial visual quality detection, which is labor-intensive, low-efficiency and strong subjectivity, resulting in high probability of missed detection and missed detection.

Method used

Using a deep learning algorithm based on computer vision, the multi-scale feature extraction, channel feature enhancement and deep feature extraction, combined with fusion processing, can realize the automatic detection of appearance defects of pharmaceutical glass tubes.

Benefits of technology

It improves the efficiency and accuracy of the appearance defect detection of medicinal glass tubes, reduces labor and time costs, and realizes the automated and intelligent detection of appearance defects of medicinal glass tubes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120594545A_ABST
    Figure CN120594545A_ABST
Patent Text Reader

Abstract

The invention provides a medicinal glass tube appearance defect detection method and device, and belongs to the technical field of pharmaceutical equipment and quality detection.The method comprises the steps that a to-be-detected target image is obtained, and the to-be-detected target image is a medicine package glass tube image; target detection processing is carried out on the target image to be detected through a target detection model, a target detection result is obtained, and the target detection result is used for representing whether the medicinal glass tube has defects or not and the defect type; according to the invention, the detection rate of defective products, especially small defective products, is improved, the false detection rate and the omission ratio are reduced, the production efficiency and the product quality are improved, and the construction of an intelligent production line is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of pharmaceutical equipment and quality inspection technology, and more specifically, relates to a method and device for detecting appearance defects of pharmaceutical glass tubes. Background Art

[0002] As pharmaceutical packaging, the quality of pharmaceutical glass tubing is directly related to the safety and effectiveness of the drug. However, due to various reasons during the production process, cosmetic defects often occur in pharmaceutical glass tubing, posing certain risks to drug use. Traditional manual visual quality inspection methods are labor-intensive, inefficient, and subjective, increasing the probability of false positives and missed detections. Summary of the Invention

[0003] The purpose of this application is to provide a method and device for detecting appearance defects of pharmaceutical glass tubes, so as to improve the efficiency and accuracy of detecting appearance defects of pharmaceutical glass tubes.

[0004] A first aspect of an embodiment of the present application provides a method for detecting appearance defects in pharmaceutical glass tubes, comprising: Acquire a target image to be detected, where the target image to be detected is an image of a medicine package glass tube; The target image to be detected is processed by the target detection model to obtain the target detection result. The target detection result is used to characterize whether there are defects in the pharmaceutical glass tube and the type of defects; The target image to be detected is processed by the target detection model to obtain the target detection result, including: Perform multi-scale feature extraction on the target image to be detected to obtain a multi-scale feature map, which includes feature maps of multiple scales; For each scale feature map, perform channel feature enhancement processing on the feature map of the scale to obtain a first feature map; For each scale feature map, perform deep feature extraction on the scale feature map from different receptive fields to obtain a second feature map; Based on each first feature map and each second feature map, a fusion process is performed to obtain fusion features corresponding to each scale feature map, each first feature map is a first feature map corresponding to each scale feature map, and each second feature map is a second feature map corresponding to each scale feature map; Target detection processing is performed based on each fusion feature to obtain the target detection result.

[0005] A second aspect of the embodiments of the present application provides a device for detecting appearance defects in pharmaceutical glass tubes, comprising: An image acquisition module is used to obtain an image of a target to be detected, where the image of the target to be detected is an image of a glass tube of a medicine package; A target detection module is used to perform target detection processing on the target image to be detected through a target detection model to obtain a target detection result, and the target detection result is used to characterize whether there is a defect in the pharmaceutical glass tube and the type of defect; The target image to be detected is processed by the target detection model to obtain the target detection result, including: Perform multi-scale feature extraction on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales; For each scale feature map, perform channel feature enhancement processing on the feature map of the scale to obtain a first feature map; For each scale feature map, deep feature extraction is performed on the feature map of that scale from different receptive fields to obtain a second feature map; Based on each first feature map and each second feature map, a fusion process is performed to obtain fusion features corresponding to each scale feature map, each first feature map is a first feature map corresponding to each scale feature map, and each second feature map is a second feature map corresponding to each scale feature map; Target detection processing is performed on each fusion feature based on the target loss function to obtain a target detection result; the target loss function is the loss function of the target detection model.

[0006] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned method for detecting appearance defects of pharmaceutical glass tubes are implemented.

[0007] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for detecting appearance defects of pharmaceutical glass tubes are implemented.

[0008] The beneficial effects of the method and device for detecting appearance defects of pharmaceutical glass tubes provided in the embodiments of the present application are: The applicant has discovered that it is difficult to achieve automated and precise detection of small targets such as appearance defects in pharmaceutical glass tubes in the existing technology because the existing technology solutions are too simple and mainly rely on manual visual quality detection methods, which are labor-intensive, inefficient, and highly subjective, increasing the probability of false detection and missed detection. Based on the above findings, and in contrast to the existing technology that relies on manual visual quality detection methods, the present application processes the target image to be detected to obtain different feature maps, then fuses the different feature maps, and performs target detection based on the fused features, thereby improving the recognition ability of small target defects in pharmaceutical glass tubes and accurately capturing subtle cracks, fine bubbles, and other imperceptible appearance defects. While achieving high-precision detection, the present application effectively reduces labor and time costs, providing a new technical path for the automated and intelligent detection of appearance defects in pharmaceutical glass tubes. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0010] Figure 1 A schematic flow chart of a method for detecting appearance defects in pharmaceutical glass tubes provided in one embodiment of the present application; Figure 2 Three network structures of the pharmaceutical glass tube appearance defect detection model provided in one embodiment of the present application; Figure 3 The network structure of the CBA module provided in one embodiment of the present application; Figure 4 A diagram of the Resnet block structure provided in one embodiment of the present application; Figure 5 The network structure of the MSDW module provided in one embodiment of the present application; FIG6 (a) is a first RT-DETR effect diagram of detecting a broken medicinal glass tube according to an embodiment of the present application; FIG6 (b) is a diagram showing the effect of detecting a broken medicinal glass tube using the first detection model designed in this case provided in one embodiment of the present application; FIG6 (c) is a second RT-DETR effect diagram of detecting a broken medicinal glass tube provided in an embodiment of the present application; FIG6 (d) is a diagram showing the effect of detecting a broken medicinal glass tube using a second detection model designed in this case provided in an embodiment of the present application; FIG6 (e) is a diagram showing the effect of detecting a broken medicinal glass tube using a third RT-DETR method according to an embodiment of the present application; FIG6( f ) is a diagram showing the effect of detecting a broken medicinal glass tube using the third detection model designed in this case provided in an embodiment of the present application; FIG6 (g) is a diagram showing the effect of detecting damage and bubbles in a pharmaceutical glass tube using the first RT-DETR method provided in one embodiment of the present application; FIG6 (h) is a diagram showing the effect of detecting damage and bubbles in a pharmaceutical glass tube using the first detection model designed in this case according to an embodiment of the present application; FIG6 (i) is a diagram showing the effect of detecting damage and bubbles in a pharmaceutical glass tube using a second RT-DETR method provided in one embodiment of the present application; FIG6( j ) is a diagram showing the effect of detecting damage and bubbles in a pharmaceutical glass tube using a second detection model designed in this case provided in an embodiment of the present application; FIG6 (k) is a diagram showing the effect of RT-DETR detection of gas lines in a pharmaceutical glass tube according to an embodiment of the present application; FIG6 (l) is a diagram showing the effect of the detection model designed in this case for detecting the gas line of a medicinal glass tube provided in one embodiment of the present application; Figure 7 This is a structural block diagram of a device for detecting appearance defects in pharmaceutical glass tubes provided in one embodiment of the present application; Figure 8 A schematic block diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0011] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0012] In order to make the purpose, technical solutions and advantages of this application clearer, specific embodiments will be described below with reference to the accompanying drawings.

[0013] Please refer to Figure 1 , Figure 1 A flowchart of a method for detecting appearance defects of a pharmaceutical glass tube provided in one embodiment of the present application is provided. The method includes S101 and S102.

[0014] S101: Acquire a target image to be detected, where the target image to be detected is a medicine package glass tube image.

[0015] In this embodiment, the drug package glass tube can be a pre-filled syringe. The system can use a high-precision area array camera (resolution ≥1392×1040) with an 8mm low-distortion industrial lens to achieve continuous shooting at 30 frames per second, capturing images of the drug package glass tube to be inspected along the production line. At least three sets of visual imaging structures can be set up on both sides of the production line, evenly distributed 360° around the center of the glass tube to ensure full circumferential coverage of the drug package glass tube. For example, each imaging unit can include: Camera module: It can be equipped with an adjustable gimbal, supporting ±15° pitch adjustment, to accommodate glass tubes of different diameters. For example, the diameter of pre-filled syringes ranges from 50 to 120 mm.

[0016] Backlight system: A red diffuse reflection high-brightness light source can be used, and the stray light interference can be eliminated by darkening the aluminum base. The light source is arranged perpendicular to the lens axis to enhance the contrast of defects such as bubbles and scratches on the glass tube surface.

[0017] S102: Performing target detection processing on the target image to be detected by using a target detection model to obtain a target detection result. The target detection result is used to characterize whether the prefilled syringe has defects and the type of defects.

[0018] In this embodiment, the object detection model can be a deep learning algorithm based on computer vision, used to automatically identify and locate specific target objects in images. In the prefilled syringe inspection scenario, the model uses a convolutional neural network to extract and classify multi-scale features of key areas in the target image. Its core functions may include determining the presence of defects in the image and locating the defect using bounding boxes.

[0019] In this embodiment, the defect type can be a bubble defect caused by incomplete vacuum plugging, a needle tube breakage caused by improper vacuum plug height or center positioning plate offset, and the presence of stones and foreign matter due to peristaltic pump failure or nitrogen pressure fluctuation.

[0020] Specifically, S102 may include: S1021-S1025.

[0021] S1021: Perform multi-scale feature extraction on the target image to be detected to obtain a multi-scale feature map, which includes feature maps of multiple scales.

[0022] The target image captured by the camera is input into the backbone network to extract multi-scale features of the image, resulting in a multi-scale feature map. The backbone network can be ResNet-18, a lightweight member of the residual network family with 18 layers, including 17 convolutional layers and 1 fully connected layer. Multi-scale features are extracted through residual blocks (ResNet Blocks S3-S5) to obtain primary features S3, intermediate features S4, and high-level features S5. Residual blocks are the core components of residual networks. They use skip connections to directly superimpose inputs onto outputs, enabling the network to learn residuals and address the vanishing gradient and degradation issues of deep networks.

[0023] In this embodiment, the Attention-based Intra-scale Feature Interaction (AIFI) module can be selected to perform intra-scale interaction on high-level features S5. AIFI is a core component in the Real-Time Detection Transformer (RT-DETR) model for enhancing multi-scale feature interaction. Its core design concept is to optimize feature correlation within a single scale through a self-attention mechanism, while working in conjunction with the cross-scale feature fusion module. The following formula is used to perform intra-scale interaction on high-level features S5 to obtain the result F5.

[0024]

[0025]

[0026]

[0027] Among them, Flatten flattens the spatial dimension of S5 into a sequence form, which meets the input requirements of Transformer for processing sequence data. , , The query, key, and value vectors, respectively, all have the same initial value, reflecting the characteristics of the self-attention mechanism. Attn represents multi-head attention, Reshape reshapes the data to the same shape as S5, and CCDFM represents cross-scale deep fusion of features.

[0028] S1022: For each scale feature map, perform channel feature enhancement processing on the feature map of the scale to obtain a first feature map.

[0029] In this embodiment, the input multi-scale feature map can be subjected to channel feature enhancement processing using 1×1 convolution and dynamic adaptive normalization to obtain a first feature map. Among them, the normalization processing can use an adaptive hyperbolic tangent function. The adaptive hyperbolic tangent function is a standard This improved version of the function introduces learnable parameters to dynamically adjust nonlinear activations. The final output first feature map has high semantics and detail fidelity, providing high-quality input for the subsequent object detection head.

[0030] S1023: For each scale feature map, perform deep feature extraction on the scale feature map from different receptive fields to obtain a second feature map.

[0031] In this embodiment, deep feature extraction is performed on the input multi-scale feature map in different receptive fields, and the steps may include: First, a residual block containing 3×3 convolution kernels, 5×5 convolution kernels, and 7×7 convolution kernels can be used to extract features of differentiated receptive fields. Secondly, average pooling can be used to retain overall statistical information on multi-scale output channels. Finally, a nonlinear Gaussian error function can be used. Activate to get the second feature map.

[0032] S1024: Perform fusion processing based on each first feature map and each second feature map to obtain fusion features corresponding to each scale feature map, where each first feature map is a first feature map corresponding to each scale feature map, and each second feature map is a second feature map corresponding to each scale feature map.

[0033] In this embodiment, spatial attention modeling can be used to assign attention weights element-by-element in the spatial dimension, strengthening the response strength of key areas. Spatial attention modeling can be performed by performing a Hadamard product on each first feature map and each second feature map. The Hadamard product is an element-by-element multiplication operation in matrix operations and is widely used in deep learning and image processing.

[0034] In this embodiment, 1×1 convolution can be used to integrate cross-channel information, and the fused features can be additively fused with the original input through residual connection.

[0035] S1025: Performing target detection processing on each fusion feature based on a target loss function to obtain a target detection result; the target loss function is a loss function of the target detection model.

[0036] In this embodiment, based on the fusion features obtained in the above steps, an intersection over union (IoU) perception query selection feature module and a transform decoder with an auxiliary prediction detection head can be used to perform target detection processing to obtain target detection results.

[0037] In this embodiment, the IoU-aware query selection feature is primarily responsible for resolving the inconsistency between classification and positioning. Its core function is to select candidate features with both high classification confidence and high intersection-over-union (IoU) from each fused feature as the decoder's initial object query. The object query in the traditional DETR model is a randomly initialized learnable parameter, resulting in a mismatch between classification scores and positioning accuracy. For example, a "false positive" box may have a high classification score but a low IoU. By constraining the training objective, the IoU-aware query forces the model to assign high classification scores to features with high IoU, thereby selecting a more reliable initial query.

[0038] In this embodiment, the transform decoder with an auxiliary prediction detection head is primarily responsible for accelerating convergence and optimizing iterations. Through a multi-layered iterative structure, the decoder refines the bounding box and category predictions for object queries layer by layer. Each decoder layer receives the output of the previous layer and incorporates the global context through a self-attention mechanism.

[0039] For example, auxiliary prediction heads accelerate training. Adding auxiliary prediction heads after each decoder layer accelerates model convergence through intermediate supervision signals. These heads provide gradient feedback during training and can be selectively removed during inference to reduce latency.

[0040] As can be seen from the above, the method for detecting appearance defects in pharmaceutical packaging glass tubes proposed in this invention enables accurate identification and classification of appearance defects in prefilled syringes. Compared to traditional manual inspection methods, this method significantly improves the efficiency of appearance quality inspection and product quality stability, effectively meeting the needs of large-scale production. It also overcomes the subjectivity of manual inspection and the risk of missed or false detections due to fatigue, thereby ensuring the stability of product quality. This is of great significance for promoting the high-quality development of the pharmaceutical packaging industry.

[0041] In one embodiment of the present application, for each scale feature map, the method further includes: Perform cross-channel feature reorganization on the feature map of this scale to obtain the first reorganized feature; Based on the first reorganized features, and through the hyperbolic tangent function, determining a hyperbolic tangent processing result; Determine a normalized feature map based on the hyperbolic tangent result and the features after the first reorganization; The channel feature enhancement processing is performed on the feature map of this scale, including: Perform channel feature enhancement based on the normalized feature map corresponding to the feature map of the scale; Among them, deep feature extraction is performed on the feature map of this scale from different receptive fields, including: Deep feature extraction is performed on the normalized feature maps corresponding to the feature maps of this scale from different receptive fields.

[0042] In this embodiment, cross-channel feature reorganization of the feature map at this scale can be performed using a 1×1 convolution to combine features from different input channels, enabling cross-channel information exchange. This results in a first reorganized feature map that maintains the same spatial size as the original image, changing only the channel dimension. The 1×1 convolution can also reduce the computational complexity of normalization.

[0043] For example, if the input is RGB three channels, 1×1 convolution can generate new channels and integrate color, texture and other features.

[0044] In this embodiment, an adaptive hyperbolic tangent function can be used for normalization to produce a normalized feature map. This nonlinearly compresses extreme values ​​and almost linearly transforms the central portion of the input, preserving the core effect of normalization. This allows the network to focus on learning high-order feature associations rather than low-order statistical deviations, automatically weakening the influence of outliers, preventing feature degradation caused by vanishing gradients in small target features during sequential processing, and improving the performance of neural networks when processing complex data.

[0045] In this embodiment, the normalized feature map is obtained by the following formula.

[0046]

[0047]

[0048] Here, Tanh represents the hyperbolic tangent function, is a learnable scalar parameter used to dynamically adjust the input scaling range and control the nonlinear strength. and is a learnable vector parameter that enables the output to be scaled back to the appropriate range.

[0049] For example, in the specific implementation, the parameter Usually initialized to 0.5, the parameter Initialized to all 1 vectors, parameters Initialize to a vector of all 0s.

[0050] In this embodiment, the first feature enhancement path is a cross-channel feature recombination path, which performs channel feature enhancement processing on the normalized feature map.

[0051] Specifically, the normalized feature map can be used to implement cross-channel feature reorganization using 1×1 convolution. Secondly, the channel attention map can be generated through the activation function, thereby achieving channel feature enhancement processing.

[0052] In this embodiment, the second feature enhancement path is a multi-scale deep feature enhancement path, which performs deep feature extraction on the normalized feature maps corresponding to the feature maps of the scale from different receptive fields.

[0053] Specifically, first, three sets of depth-wise separable convolutions can be deployed in parallel to perform convolution operations on the normalized feature maps. Here, the multi-scale convolution kernel sizes are 3×3, 5×5, and 7×7 respectively. Secondly, average pooling (Average Pooling, Avg) is performed on the multi-scale output execution channel to generate fused features. Then, the normalized feature map and fusion feature Residual connection is performed to enhance the stability of gradient flow, and then the feature data after residual connection is reorganized across channels through 1×1 convolution to complete cross-channel information interaction. Finally, it is activated by the nonlinear Gaussian error function GeLU to complete the deep feature extraction of the normalized feature map.

[0054] From the above, it can be concluded that cross-channel feature reorganization is achieved through 1×1 convolution, and the normalization parameters are dynamically adjusted in combination with the adaptive hyperbolic tangent function, which effectively suppresses outliers and retains the linear response of the feature center area, thereby enhancing the gradient stability of small targets; at the same time, a dual-path collaborative enhancement strategy is adopted, and weights are dynamically allocated through the attention mechanism in the channel dimension. In the spatial dimension, local details and global semantics are captured through multi-scale depth-separable convolution. The residual connection and GeLU activation function are used to optimize the gradient flow, which can achieve efficient and accurate detection of appearance defects in pharmaceutical glass tubes.

[0055] In one embodiment of the present application, performing channel feature enhancement processing on the feature map of the scale includes: The feature map of this scale is reorganized across channels to obtain the second reorganized features. Based on the second reorganized features, the channel features are enhanced through the activation function to generate a channel attention map as the first feature map corresponding to the feature map of this scale.

[0056] In this embodiment, the feature map at this scale may be reorganized across channels using a cross-channel feature reorganization path to obtain a second reorganized feature. The feature map at this scale is a feature map after feature normalization.

[0057] Specifically, 1×1 convolution can be used to reorganize the cross-channel features of the feature map of this scale. Secondly, the channel attention map is generated by the activation function. Specifically, the activation function here is The Sigmoid-Weighted LinearUnit (SiLU) activation function is selected for nonlinear mapping to enhance channel features and compensate for the content that may be lost due to the order constraint of the State Space Model (SSM). The order constraint of SSM is essentially the result of the dual effects of its mathematical model and hardware implementation. Mathematically, it must follow recursive dynamics and frequency domain alignment, while the hardware is limited by memory access and parallel strategies. Secondly, the activation function The core idea is to take the product of the input value and its Sigmoid activation value as the output. The mathematical expression is: ,in The first feature map Depend on Get, among them is the feature map of this scale.

[0058] From the above, we can conclude that, first, 1×1 convolution does not change the height and width of the feature map, only adjusting the channel dimension, avoiding spatial information loss and making it suitable for multi-scale feature fusion. Second, at each spatial location, 1×1 convolution is equivalent to a fully connected layer performing a linear transformation on the channel, but significantly reduces the number of parameters through weight sharing. Finally, compared to other activation functions, SiLU provides smooth gradients when features are close to zero, which helps improve model performance and generalization. Its continuous differentiability is well-suited for channel reorganization tasks.

[0059] In one embodiment of the present application, deep feature extraction is performed on the feature map of the scale from different receptive fields to obtain each second feature map, including: The feature maps of this scale are subjected to depth-separable convolution through different convolution kernels to obtain the corresponding depth-wise convolution features; Perform channel-level pooling on the corresponding convolution features to obtain fusion features; Perform residual connection processing on the fusion features and the feature map of the scale to obtain the residual processed features; The features after residual processing are reorganized across channels to obtain the third reorganized features; The third reorganized features are activated by a nonlinear Gaussian error function to obtain multi-scale depth features.

[0060] In this embodiment, performing deep feature extraction on feature maps of the scales using different receptive fields may be performing depthwise separable convolution using different convolution kernels.

[0061] Specifically, multiple sets of depthwise separable convolutions are deployed in parallel, with multi-scale convolution kernels of 3×3, 5×5, and 7×7 sizes. The 3×3 depthwise convolution kernel extracts local texture features (such as scratches or bubbles in pharmaceutical glass tubing), the 5×5 depthwise convolution kernel captures mid-range contextual information (such as deformation in pharmaceutical glass tubing), and the 7×7 depthwise convolution kernel perceives global structural features (such as air lines in pharmaceutical glass tubing). The depthwise separable convolution kernel decomposes spatial filtering and channel projection, reducing computational complexity.

[0062] In this embodiment, Avg can be used to fuse the multi-scale outputs of multiple groups of depth-wise separable convolutions to obtain the fusion feature .

[0063] Specifically, Avg compresses features extracted from feature maps of different scales, such as those extracted by 3×3, 5×5, and 7×7 convolution kernels, into global statistics along the channel dimension, eliminating local noise interference and highlighting common features across different channels. For example, in defect detection for pharmaceutical glass tubing, multi-scale features may include both local details (small-scale) and the overall morphology (large-scale) of a crack. Avg can fuse this complementary information.

[0064] In this embodiment, A residual connection is constructed with the feature map of this scale to obtain the feature data after the residual connection. This operation can alleviate the gradient disappearance and thus enhance the stability of the gradient flow.

[0065] In this embodiment, the feature data after residual connection can be reorganized across channels through 1×1 convolution to complete cross-channel information interaction and obtain the third reorganized features. Finally, the third reorganized features can be activated by the nonlinear Gaussian error function GeLU to obtain the second feature map. The specific process formula is expressed as follows: ,in, is the second feature map, are the weights of the depthwise separable convolution, is the feature dimension of the input, is the feature dimension of the output, is the convolution kernel size, here i=1,2,3; The sizes are set to 3×3, 5×5, and 7×7 respectively.

[0066] From the above, it can be concluded that this embodiment significantly improves the model's sensitivity to minor defects and robustness in industrial scenarios while reducing computational complexity through cross-channel feature reorganization.

[0067] In one embodiment of the present application, for each scale feature map, a fusion process is performed based on the first feature map and the second feature map corresponding to the scale feature map to obtain a fusion feature corresponding to the scale feature map, including: Determine a Hadamard product of a first feature map corresponding to the feature map of the scale and a second feature map corresponding to the feature map of the scale; Perform cross-channel feature reorganization based on the Hadamard product to obtain the fourth reorganized feature; The fourth reorganized feature is fused with the feature map of the scale to obtain the fused feature corresponding to the feature map of the scale.

[0068] In this embodiment, the first feature map and the second feature map may be fused through dual-path parallel feature interaction to obtain a fused feature corresponding to the scale feature map.

[0069] Specifically, the dual-path parallel feature interaction is divided into two steps: spatial attention modeling and channel information fusion. The spatial attention modeling can be obtained by performing the Hadamard product on the first feature map and the second feature map, and the formula is: The spatial attention modeling can distribute the attention weights element by element in the spatial dimension and strengthen the response strength of the key areas. Channel information fusion can use 1×1 convolution to integrate cross-channel information and add the fused features to the original input through residual connection. The formula is as follows .in, For fusion features.

[0070] From the above, it can be concluded that the parallel use of depthwise separable convolutions with different dilation rates can capture local details, small targets and multi-scale spatial features. Through the residual connection of the depthwise separable convolution kernel, the ability to extract local spatial information is enhanced while reducing the computational cost, improving the model's ability to detect targets of different scales and enhancing the model's robustness to scale changes.

[0071] In one embodiment of the present application, target detection processing is performed based on each fusion feature to obtain a target detection result, including: Connect each fusion feature to obtain the connected features; Target detection is performed based on the features after connection processing to obtain target detection results.

[0072] In this embodiment, different defects correspond to different features after concatenation processing. Specifically, when dealing with bubble defects in drug packaging glass tubes, since bubbles typically appear as circular or elliptical areas with blurred edges and low grayscale values ​​in images, the concatenated features help the target detection module quickly extract suspected bubble areas. Morphological processing methods are then combined to remove noise interference and ultimately accurately locate the bubble defect. For crack defects, because they are slender, have sharp edges and obvious grayscale value changes in the image, the features after connection processing can help the target detection module quickly identify the direction and length of the crack, thereby completing the detection of crack defects.

[0073] When it comes to scratch defects, since they appear as linear features in the image, by analyzing parameters such as the angle and length of the lines, the connected processed features can help the target detection module to accurately identify and classify scratch defects, and ultimately obtain comprehensive and accurate target detection results. From the above, it can be concluded that the features after connection processing can be used to deal with different defects in a targeted manner: bubbles are fuzzy, low-grayscale circles, and the features help to quickly extract suspected areas and combine them with morphological positioning; cracks are slender and sharp, with obvious grayscale changes, and the features help to identify the direction and length; scratches are linear, and the features help analyze parameters to achieve accurate classification, significantly improving the efficiency and accuracy of defect detection.

[0074] In one embodiment of the present application, a sample image, a target annotation box corresponding to each target in the sample image, and a target probability that each target in the sample image belongs to a preset target are obtained; Based on the sample image, the target annotation box corresponding to each target in the sample image, and the target probability that each target in the sample image belongs to the preset target, the initial model is trained to obtain the target detection model; The methods for training the initial model include: The sample image is tested by the initial model to obtain the predicted detection results, which include the prediction box corresponding to each target and the predicted probability of each target belonging to the preset target; Based on the intersection-over-union ratio between the prediction box corresponding to each target and the target annotation box corresponding to the target, a first loss is determined based on the intersection-over-union ratio; determining a second loss based on a difference between a predicted probability that each target belongs to the preset target and a target probability that the target belongs to the preset target; Based on the first loss and the second loss, the initial model is trained.

[0075] In this embodiment, the sample image may be a historical image of a medicinal glass tube captured by a camera, or may be image data of various parts, including images at different angles and under different lighting conditions.

[0076] The target annotation box is a rectangular box used to mark the location of defects in an image, typically represented by a bounding box. For example, for a bubble defect, the annotator would select the edge of the bubble in the sample image. The target probability is the confidence level that the target belongs to the predefined defect category, and its value range is [0, 1].

[0077] In this embodiment, the constructed model is trained by constructing a loss function using a large amount of labeled pre-filled syringe image data. For the positioning loss of the predicted bounding box coordinates, the WIoUv3 loss function of the RT-DETR framework can be adopted, and a matching quality-aware loss is designed as the classification loss of the predicted category, and its expression is , which is the target loss function. Among them, represents the true label category, represents the predicted probability of the foreground category, represents the IoU between the predicted bounding box and the target box, and the parameter adjusts the sensitivity of the matching quality to the loss weight allocation, and is used to control the balance between high-quality matching samples and low-quality matching samples.

[0078] Specifically, a parameter is introduced into the IoU, so that samples with a higher IoU obtain exponentially larger weights, strengthen the gradient signal of high-quality predictions, and accelerate the learning of high-quality predictions; The positive sample gradient is proportional to , compared with the original VFL loss, the model pays more attention to samples with accurate positioning; the negative sample gradient is inversely proportional to , automatically focusing on difficult samples.

[0079] During the training process, a progressive parameter tuning strategy is adopted for the parameter : initially set base_gammma = 2.0; in the initial stage of training (current_epoch < max_epoch / 3), the parameter is set to base_gammma × 0.7; in the middle stage (current_epoch < max_epoch*2 / 3), the parameter is set to base_gammma × 0.8, and in the later stage, the parameter is set to base_gammma to strengthen high-quality samples. Among them, current_epoch is the current training round, and max_epoch is the maximum training round.

[0080] It can be concluded from the above that during the training process, the model parameters are continuously adjusted, including the learning rate, the training batch size, the number of training times, and the optimization of the loss function, to improve the detection accuracy and generalization ability of the model.

[0081] In an embodiment of the present application, the network structure of the medicinal glass tube defect detection model is given.

[0082] Such as Figure 2The collected image data shown in (a) is processed by the backbone network to obtain primary features S3, intermediate features S4 and high-level features S5, and the three different features are input into the CCDFM module for deep feature extraction. Then, the intersection over union (IoU) perception query selection feature module and the transform decoder with auxiliary prediction detection head are used to perform target detection processing to obtain the target detection results. Figure 2 (b) Omit the intermediate steps and only describe the main structure of the network structure of the pharmaceutical glass tube appearance defect detection model. Figure 2 (b), Figure 2 (c) Highlighting of the AIFI module, demonstrating its importance in performing intrascale interactions on high-level feature S5.

[0083] Figure 3 This is the network structure of the CBA module, which performs 3×3 convolution, batch normalization and activation function on the input data once and gradually extracts semantic features from low-level to high-level. Figure 4 It is a residual network, which consists of a pooling layer, a 1×1 convolutional layer and a parallel CBA module. The overall design aims to achieve feature compression, cross-channel interaction and multi-branch feature fusion. Figure 5 The network structure of the MSDW module is shown. It shows a dual-path parallel feature interaction architecture. Path 1 first uses 1×1 convolution to achieve cross-channel feature reorganization, and then generates a channel attention map through the activation function. Path 2 first deploys multiple sets of depthwise separable convolutions in parallel, where the multi-scale convolution kernel sizes are 3×3, 5×5, and 7×7 respectively. Secondly, channel-level average pooling is performed on the multi-scale output to generate the fused feature F. fusion Then, the fusion features are combined with the initial input features to form a residual connection. Finally, cross-channel feature reorganization is achieved through 1×1 convolution, completing cross-channel information interaction, and finally activated by the nonlinear Gaussian error function GeLU.

[0084] In one embodiment of the present application, an experimental verification is provided. The experimental environment is based on the Ubuntu 18.04 operating system and an NVIDIA V100 graphics card. The experiment basically adopts the parameter settings officially recommended by RT-DETR, and uses data augmentation and other strategies for model training. Each training session inputs four images and iterates 300 times.

[0085] 5,352 images of prefilled syringes were collected from various angles and lighting conditions and preprocessed. The appearance defects of prefilled syringes were classified into six categories: air lines, stones, stains, scratches, breakage, and bubbles, with the defect types represented by the numbers 0, 1, 2, 3, 4, and 5, respectively. The dataset was divided into training, validation, and test sets using a random sampling ratio of 7:2:1. The constructed model was used for training. After training, the test set images were inspected, with mean average precision (mAP) and frame rate (FPS) used as evaluation metrics.

[0086] Table 1 shows the test results of different detection algorithms on the pre-filled dataset. Compared with the baseline RT-DETR algorithm, the proposed method shows significant improvements in accuracy, with a 3.3% increase in mean average prediction accuracy (MAP) and a frame rate of 92 frames per second. Although running slower than the baseline model, it can still meet the speed requirements of the pre-filled production line.

[0087]

[0088] Figures 6(a), 6(c), 6(e), 6(g), 6(i), and 6(k) show the original RT-DETR detection results, while Figures 6(b), 6(d), 6(f), 6(h), 6(j), and 6(l) show the detection results of the detection model designed in this case. It can be observed that both methods can correctly detect larger defects, such as the stain defects in Figures 6(a), 6(b), 6(g), and 6(h), and the breakage defects and air line defects in Figures 6(i), 6(j), 6(k), and 6(l). However, the baseline model does not correctly detect smaller defects or less obvious defects, such as the small stain on the bottle mouth in Figure 6(a) and the scratch on the bottle bottom in Figure 6(e). Specifically, first, the improved RT-DETR in this case detected small targets or defective targets with unclear features that were missed by the RT-DETR method. Second, the improved scheme significantly increased the confidence level in the detection results. Finally, the improved scheme also more accurately positioned the target.

[0089] The method and system for detecting appearance defects of prefilled syringes based on the detection algorithm provided by the present invention have the advantages of high detection efficiency, high precision, and high level of automation. They can greatly improve the efficiency and quality of appearance defect detection of prefilled syringes and have broad application prospects.

[0090] Corresponding to the above embodiment, a method for detecting appearance defects of a pharmaceutical glass tube is provided. Figure 7 This is a structural block diagram of a device for detecting appearance defects of a pharmaceutical glass tube provided in one embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 7The pharmaceutical glass tube appearance defect detection device 20 includes: an image acquisition module 21 and a target detection module 22.

[0091] The image acquisition module 21 is used to obtain the target image to be detected, which is the image of the medicine package glass tube; The target detection module 22 is used to perform target detection processing on the target image to be detected through the target detection model to obtain a target detection result. The target detection result is used to characterize whether there is a defect in the pharmaceutical glass tube and the type of defect; The target image to be detected is processed by the target detection model to obtain the target detection result, including: Perform multi-scale feature extraction on the target image to be detected to obtain a multi-scale feature map, which contains feature maps of multiple scales; For each scale feature map, perform channel feature enhancement processing on the feature map of the scale to obtain a first feature map; For each scale feature map, deep feature extraction is performed on the feature map of that scale from different receptive fields to obtain a second feature map; Based on each first feature map and each second feature map, a fusion process is performed to obtain fusion features corresponding to each scale feature map, each first feature map is a first feature map corresponding to each scale feature map, and each second feature map is a second feature map corresponding to each scale feature map; Target detection processing is performed based on each fusion feature to obtain the target detection result.

[0092] In one embodiment of the present application, the target detection module 22 is further configured to: For the feature map of each scale, perform cross-channel feature reorganization on the feature map of that scale to obtain the first reorganized feature; Based on the first reorganized features, and through the hyperbolic tangent function, determining a hyperbolic tangent processing result; Determine a normalized feature map based on the hyperbolic tangent result and the features after the first reorganization; The channel feature enhancement processing is performed on the feature map of this scale, including: Perform channel feature enhancement based on the normalized feature map corresponding to the feature map of the scale; Among them, deep feature extraction is performed on the feature map of this scale from different receptive fields, including: Deep feature extraction is performed on the normalized feature maps corresponding to the feature maps of this scale from different receptive fields.

[0093] In one embodiment of the present application, the target detection module 22 is specifically configured to: The feature map of this scale is reorganized across channels to obtain the second reorganized features. Based on the second reorganized features, the channel features are enhanced through the activation function to generate a channel attention map as the first feature map corresponding to the feature map of this scale.

[0094] In one embodiment of the present application, the target detection module 22 is specifically configured to: The feature maps of this scale are subjected to depth-separable convolution through different convolution kernels to obtain the corresponding depth-wise convolution features; Perform channel-level pooling on the corresponding convolution features to obtain fusion features; Perform residual connection processing on the fusion features and the feature map of the scale to obtain the residual processed features; The features after residual processing are reorganized across channels to obtain the third reorganized features; The third reorganized features are activated by a nonlinear Gaussian error function to obtain multi-scale depth features.

[0095] In one embodiment of the present application, the target detection module 22 is specifically configured to: Determine a Hadamard product of a first feature map corresponding to the feature map of the scale and a second feature map corresponding to the feature map of the scale; Perform cross-channel feature reorganization based on the Hadamard product to obtain the fourth reorganized feature; The fourth reorganized feature is fused with the feature map of the scale to obtain the fused feature corresponding to the feature map of the scale.

[0096] In one embodiment of the present application, the target detection module 22 is specifically configured to: Each fusion feature is connected to obtain the connected features; Target detection is performed based on the features after connection processing to obtain target detection results.

[0097] In one embodiment of the present application, the target detection module 21 is specifically configured to: Obtain a sample image, a target annotation box corresponding to each target in the sample image, and a target probability that each target in the sample image belongs to a preset target; Based on the sample image, the target annotation box corresponding to each target in the sample image, and the target probability that each target in the sample image belongs to the preset target, the initial model is trained to obtain the target detection model; The methods for training the initial model include: The sample image is tested by the initial model to obtain the predicted detection results, which include the prediction box corresponding to each target and the predicted probability of each target belonging to the preset target; Based on the intersection-over-union ratio between the prediction box corresponding to each target and the target annotation box corresponding to the target, a first loss is determined based on the intersection-over-union ratio; determining a second loss based on a difference between a predicted probability that each target belongs to the preset target and a target probability that the target belongs to the preset target; Based on the first loss and the second loss, the initial model is trained.

[0098] See also Figure 8 , Figure 8 This is a schematic block diagram of an electronic device provided in one embodiment of the present application. Figure 8 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules / units in the above-mentioned device embodiments, such as Figure 7 The functions of the image acquisition module 21 and the target detection module 22 are shown.

[0099] It should be understood that in the embodiment of the present application, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0100] The input device 302 may include a touchpad, a fingerprint collection sensor (for collecting user fingerprint information and fingerprint direction information), a microphone, etc. The output device 303 may include a display (LCD, etc.), a speaker, etc.

[0101] The memory 304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store device type information.

[0102] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present application can execute the implementation methods described in the first and second embodiments of a method for detecting appearance defects of a medicinal glass tube provided in the embodiments of the present application, and can also execute the implementation methods of the electronic device described in the embodiments of the present application, which will not be repeated here.

[0103] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.

[0104] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.

[0105] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0106] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0107] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or can be an electrical, mechanical or other form of connection.

[0108] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0109] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0110] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for detecting appearance defects of pharmaceutical glass tubes, characterized in that: include: Acquire a target image to be detected, wherein the target image to be detected is an image of a medicine package glass tube; Performing target detection processing on the target image to be detected using a target detection model to obtain a target detection result, wherein the target detection result is used to characterize whether the pharmaceutical glass tube has defects and the type of defects; The target detection process is performed on the target image to be detected by using a target detection model to obtain a target detection result, including: Performing multi-scale feature extraction on the target image to be detected to obtain a multi-scale feature map, wherein the multi-scale feature map includes feature maps of multiple scales; For each scale feature map, perform channel feature enhancement processing on the feature map of the scale to obtain a first feature map; For each scale feature map, deep feature extraction is performed on the feature map of that scale from different receptive fields to obtain a second feature map; Performing fusion processing based on each first feature map and each second feature map to obtain fusion features corresponding to each scale feature map, wherein each first feature map is a first feature map corresponding to each scale feature map, and each second feature map is a second feature map corresponding to each scale feature map; Target detection processing is performed on each fusion feature based on the target loss function to obtain a target detection result; the target loss function is the loss function of the target detection model.

2. The method according to claim 1, characterized in that For each scale feature map, the method further includes: Perform cross-channel feature reorganization on the feature map of this scale to obtain the first reorganized feature; Based on the first reorganized features, and using a hyperbolic tangent function, determining a hyperbolic tangent processing result; Determining a normalized feature map based on the hyperbolic tangent result and the first reorganized features; The performing channel feature enhancement processing on the feature map of the scale includes: Perform channel feature enhancement based on the normalized feature map corresponding to the feature map of the scale; The step of performing deep feature extraction on the feature map of the scale from different receptive fields includes: Deep feature extraction is performed on the normalized feature map corresponding to the feature map of the scale from different receptive fields.

3. The method according to claim 1, characterized in that The performing channel feature enhancement processing on the feature map of the scale includes: The feature map of this scale is reorganized across channels to obtain a second reorganized feature. Based on the second reorganized feature, channel feature enhancement processing is performed through an activation function to generate a channel attention map as the first feature map corresponding to the feature map of this scale.

4. The method according to claim 1, wherein The deep feature extraction is performed on the feature map of the scale from different receptive fields to obtain each second feature map, including: The feature maps of this scale are subjected to depth-separable convolution through different convolution kernels to obtain the corresponding depth-wise convolution features; Perform channel-level pooling on the corresponding convolutional features to obtain fusion features; Performing residual connection processing on the fusion feature and the feature map of the scale to obtain residual processed features; Performing cross-channel feature reorganization on the residual processed features to obtain third reorganized features; The third reorganized feature is activated by a nonlinear Gaussian error function to obtain the multi-scale depth feature.

5. The method according to claim 1, wherein For each scale feature map, a fusion process is performed based on the first feature map and the second feature map corresponding to the scale feature map to obtain a fusion feature corresponding to the scale feature map, including: Determine a Hadamard product of a first feature map corresponding to the feature map of the scale and a second feature map corresponding to the feature map of the scale; Performing cross-channel feature reorganization based on the Hadamard product to obtain fourth reorganized features; Based on the fourth reorganized features and the feature map of the scale, a fusion feature corresponding to the feature map of the scale is obtained.

6. The method according to any one of claims 1 to 5, characterized in that The target detection processing is performed based on each fusion feature to obtain the target detection result, including: Connecting the fused features to obtain connected features; Target detection processing is performed based on the features after the connection processing to obtain the target detection result.

7. The method according to claim 1, characterized in that The target loss function construction process includes: Obtain a sample image, a target annotation box corresponding to each target in the sample image, and a target probability that each target in the sample image belongs to a preset target; Based on the sample image, the target annotation box corresponding to each target in the sample image, and the target probability that each target in the sample image belongs to the preset target, the initial model is trained to obtain the target detection model; The methods for training the initial model include: Performing target detection on the sample image through the initial model to obtain a predicted detection result, wherein the predicted detection result includes a prediction box corresponding to each target and a predicted probability that each target belongs to a preset target; Based on the intersection-over-union ratio between the prediction box corresponding to each target and the target annotation box corresponding to the target, and determining the first loss based on the intersection-over-union ratio; determining a second loss based on a difference between the predicted probability that each target belongs to the preset target and the target probability that the target belongs to the preset target; The initial model is trained based on the first loss and the second loss.

8. A device for detecting appearance defects of pharmaceutical glass tubes, characterized in that: include: An image acquisition module is used to obtain a target image to be detected, wherein the target image to be detected is an image of a medicine package glass tube; A target monitoring module is used to perform target detection processing on the target image to be detected through a target detection model to obtain a target detection result, and the target detection result is used to characterize whether there is a defect in the pharmaceutical glass tube and the type of defect; The target detection process is performed on the target image to be detected by using a target detection model to obtain a target detection result, including: Performing multi-scale feature extraction on the target image to be detected to obtain a multi-scale feature map, wherein the multi-scale feature map includes feature maps of multiple scales; For each scale feature map, perform channel feature enhancement processing on the feature map of the scale to obtain a first feature map; For each scale feature map, deep feature extraction is performed on the feature map of that scale from different receptive fields to obtain a second feature map; Performing fusion processing based on each first feature map and each second feature map to obtain fusion features corresponding to each scale feature map, wherein each first feature map is a first feature map corresponding to each scale feature map, and each second feature map is a second feature map corresponding to each scale feature map; Target detection processing is performed based on each fusion feature to obtain the target detection result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Layout analysis method and device based on feature fusion and storage medium

    CN116935419A

  • Multi-scale receptive field low-illumination target detection method under adaptive enhancement

    CN119399448A

  • Semiconductor wafer defect detection light source configuration method, device and equipment

    CN119445046A

  • Steel sheet scraper defect detection method, storage medium, electronic equipment and device

    CN119693339A

Cited By

  • Efficient two-step multi-scale feature extraction method

    CN121214196A