Image detection method, device and equipment for injection molding workpiece and storage medium

Through the combination of the dual attention extraction network and the feature analysis network, efficient and accurate detection of injection molded workpiece defects is achieved, the problems of inconsistency and inefficiency of manual inspection are solved, and the accuracy and robustness of inspection are improved.

CN120431099AInactive Publication Date: 2025-08-05SHENZHEN JINWEIFA PLASTIC&METALPRODUCE CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510933620.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The quality inspection of existing injection molded parts relies on artificial visual inspection, and there are problems such as inconsistent inspection standards, inefficient efficiency and misjudgment of inspection, which is difficult to meet the needs of modern large-scale production.

Method used

The dual attention extraction network and feature analysis network are used to detect defects of injection molded workpieces through multi-scale pooling, attention channel fusion and deformable convolution.

Benefits of technology

It improves the accuracy and robustness of defect detection of injection molded workpieces, can effectively deal with defects of various types and shapes, and improves the overall performance of inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431099A_ABST
    Figure CN120431099A_ABST
Patent Text Reader

Abstract

The invention provides an image detection method and device for an injection molding workpiece, equipment and a storage medium, and the method comprises the steps: inputting an enhanced image set of the injection molding workpiece into a preset double-attention extraction network, and obtaining an affinity feature map; inputting the affinity feature map into a preset feature analysis network to obtain a main branch attention feature and an auxiliary branch attention feature; performing multi-scale pooling on the main branch attention features to obtain main branch features, performing convolution on the auxiliary branch attention features to obtain auxiliary branch features, and splicing the main branch features and the auxiliary branch features to obtain enhanced features; fusing the enhanced features by using an attention channel to obtain a fused feature map; and carrying out deformable convolution on the fused feature map to obtain a defect detection result of the injection molding workpiece.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a method, device, equipment and storage medium for image detection of injection molded workpieces. Background Art

[0002] As critical components in industrial production, the quality inspection of injection molded parts is directly related to the overall performance and reliability of the product. Currently, quality inspection of injection molded parts primarily relies on manual visual inspection, where inspectors visually determine whether the product contains defects such as bubbles, cracks, and color variations. This traditional manual inspection method has numerous limitations: First, test results are significantly influenced by the inspector's experience, physical strength, and subjective judgment, making it difficult to ensure consistent testing standards; second, manual inspection is inefficient and cannot meet the demands of modern large-scale production; third, prolonged visual fatigue can easily lead to missed inspections and misjudgments, compromising the accuracy of product quality control. Summary of the Invention

[0003] The present application provides an image detection method, device, equipment and storage medium for injection molded workpieces, which are used to improve the detection speed and detection accuracy of injection molded parts.

[0004] In a first aspect, an embodiment of the present application provides an image detection method for an injection molded workpiece, the method comprising: The enhanced image set of the injection molded workpiece is input into the preset dual attention extraction network to obtain the affinity feature map; Inputting the affinity feature map into a preset feature analysis network to obtain main branch attention features and auxiliary branch attention features; Performing multi-scale pooling on the main branch attention feature to obtain the main branch feature, performing convolution on the auxiliary branch attention feature to obtain the auxiliary branch feature, and splicing the main branch feature and the auxiliary branch feature to obtain the enhanced feature; Using the attention channel to fuse the enhanced features, a fused feature map is obtained; Performing deformable convolution on the fused feature map to obtain a defect detection result of the injection molded workpiece.

[0005] In a second aspect, an embodiment of the present application provides an image detection device for an injection molded workpiece, wherein the image detection device for an injection molded workpiece is configured to perform an image detection method for an injection molded workpiece as described in any one of the embodiments of the present application, the device comprising: A feature extraction module is used to input the enhanced image set of the injection molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map; A feature analysis module, configured to input the affinity feature map into a preset feature analysis network to obtain a main branch attention feature and an auxiliary branch attention feature; A feature enhancement module is used to perform multi-scale pooling on the main branch attention feature to obtain the main branch feature, perform convolution on the auxiliary branch attention feature to obtain the auxiliary branch feature, and splice the main branch feature with the auxiliary branch feature to obtain the enhanced feature; A feature fusion module, configured to fuse the enhanced features using an attention channel to obtain a fused feature map; The result output module is used to perform deformable convolution on the fused feature map to obtain the defect detection result of the injection molded workpiece.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device including a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the image detection method for injection molded workpieces as described in any one of the embodiments of the present application when executing the computer program.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the image detection method for injection molded workpieces as described in any one of the embodiments of the present application.

[0008] An embodiment of the present application provides an image detection method, apparatus, device and storage medium for injection-molded workpieces. The method includes: inputting an enhanced image set of the injection-molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map; inputting the affinity feature map into a preset feature analysis network to obtain a main branch attention feature and an auxiliary branch attention feature; performing multi-scale pooling on the main branch attention feature to obtain a main branch feature, performing convolution on the auxiliary branch attention feature to obtain an auxiliary branch feature, and splicing the main branch feature with the auxiliary branch feature to obtain an enhanced feature; using attention channels to fuse the enhanced features to obtain a fused feature map; performing deformable convolution on the fused feature map to obtain a defect detection result of the injection-molded workpiece. In the above method, comprehensive feature perception capability is provided through the dual attention mechanism, the diversity and integrity of features are guaranteed through the dual-branch attention network, and the key features are highlighted in the attention channel fusion to obtain a fused feature map. After the fused feature map is processed by deformable convolution, the adaptability to various defects is improved, and accurate defect detection results of injection molded workpieces are output, which not only ensures the accuracy of detection, but also improves the robustness of the algorithm. It can effectively deal with defects of various types, shapes and sizes on the surface of injection molded workpieces, and significantly improves the overall performance of defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A schematic flow chart of an image detection method for an injection molded workpiece provided in an embodiment of the present application; Figure 2 A schematic block diagram of an image detection device for an injection molded workpiece provided in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0012] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0013] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0014] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0015] See also Figure 1 , Figure 1 This is a schematic flow chart of an image detection method for an injection molded workpiece provided in an embodiment of the present application. Figure 1 As shown, the specific steps of the image detection method for injection molded workpieces include: S101-S105.

[0016] S101: Input the enhanced image set of the injection molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map.

[0017] In some embodiments, before inputting the enhanced image set of the injection-molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map, the method further includes: capturing an image of the injection-molded workpiece to be inspected; performing Gaussian filtering and median filtering on the image to be inspected to obtain a denoised image; performing adaptive histogram equalization on the denoised image to obtain a light-balanced image; and performing data augmentation on the light-balanced image in a preset manner to obtain an enhanced image set. The preset data augmentation includes random rotation between -15° and 15°, random horizontal flipping, random vertical flipping, random cropping, and ±20% brightness and contrast adjustment.

[0018] By performing image enhancement on the initial image to be detected, the detection accuracy can be improved. After the network enhancement is implemented, it can be input into the preset dual attention extraction network for processing.

[0019] For example, when an enhanced image set of an injection molded workpiece is input into a pre-configured dual-attention extraction network, the network comprises two cross-attention modules with shared parameters. The enhanced image set undergoes dimensionality reduction via a 1×1 convolutional layer to generate a query feature map Q and a key feature map K. The number of output channels in the convolutional layer is set to 1 / 8 of the number of input channels to reduce computational complexity. An affinity operation is used to calculate the correlation between each pixel in the feature map and pixels in the same row and column. This correlation calculation is performed using matrix multiplication to ensure that long-range dependencies between pixels are captured. The correlation values are normalized via a softmax layer to obtain horizontal and vertical correlation features. The two cross-attention modules are connected in series. The first module's receptive field covers a local area, focusing on capturing local structural features and texture information on the injection molded workpiece surface. The second module expands the receptive field through a cumulative effect, acquiring global contextual information, enabling the model to understand the correlation between defective areas and surrounding normal areas. Residual connections are used to transfer features between the two modules of the dual-attention extraction network, preventing the vanishing gradient problem in deep networks. The affinity feature map output by this network structure not only contains the structural difference information between the surface defects and normal areas of the injection molded workpiece, but also retains the detailed features such as the shape and texture of the defective area, providing a comprehensive feature representation for subsequent feature analysis.

[0020] S102: Input the affinity feature map into a preset feature analysis network to obtain the main branch attention feature and the auxiliary branch attention feature.

[0021] For example, when an affinity feature map is input into a pre-set feature analysis network, the network employs an improved CSP architecture, splitting the input affinity feature map into two parallel processing paths: a main branch and an auxiliary branch. In the feature analysis network, the main branch adjusts the number of channels using a 1×1 convolutional layer that uses the LeakyReLU activation function to enhance the network's nonlinear representation capabilities. A 3×3 convolutional layer is then used to extract deep features, with the number of kernels being twice the number of input channels to enhance feature extraction. The auxiliary branch reduces computational complexity using a 1×1 convolutional layer, retaining some of the original feature information and complementing the main branch. Both branches utilize a combination of channel-wise and spatial-wise attention mechanisms. The channel-wise attention mechanism learns inter-channel dependencies through a squeeze-excitation network structure, while the spatial-wise attention mechanism emphasizes important spatial locations by generating a two-dimensional attention map. The main branch's attention features focus on salient features of the defect area, including its edges, shape, and internal structure. The auxiliary branch's attention features retain background texture information to understand the contrast between the defect and the background. This dual-branch structural design of the feature analysis network not only balances computational efficiency and feature extraction capabilities, but also improves the discriminability of features through the introduction of the attention mechanism, making the model better adapted to the detection needs of various types of defects on the surface of injection molded workpieces, such as bubbles, sink marks, burn marks, silver wire, and other defect features of different forms.

[0022] S103. Perform multi-scale pooling on the main branch attention feature to obtain the main branch feature, perform convolution on the auxiliary branch attention feature to obtain the auxiliary branch feature, and splice the main branch feature and the auxiliary branch feature to obtain the enhanced feature.

[0023] For example, when performing multi-scale pooling on the main branch attention features, an improved SoftPool pooling operation is used instead of the traditional MaxPool pooling. SoftPool pooling accumulates activation values in an exponentially weighted manner, with the weight coefficients dynamically adjusted via learnable parameters, enabling the pooling operation to adaptively retain important features. Feature extraction is performed for three different receptive fields (2×2, 3×3, and 5×5). The pooled results for each receptive field are passed through a 1×1 convolutional layer with the number of channels adjusted to generate multi-scale main branch features. The pooling stride is set to 2 to ensure progressive downsampling of the feature map size. The auxiliary branch attention features are processed through a 3×3 convolutional layer using a dilated convolution design with a dilation rate of 2, expanding the receptive field while maintaining the original spatial resolution, generating auxiliary branch features. The main and auxiliary branch features are concatenated along the channel dimension, and the number of channels is adjusted through a 1×1 convolutional layer to form enhanced features. This SoftPool-based multi-scale feature extraction method is particularly suitable for processing defects of varying sizes on the surface of injection molded workpieces and can better preserve the characteristic information of less noticeable defects. Compared with traditional MaxPool, SoftPool can avoid excessive loss of feature information, maintain the integrity and continuity of the defect area, and improve detection accuracy.

[0024] S104: Use attention channels to fuse enhanced features and obtain a fused feature map.

[0025] For example, when utilizing attention channel fusion to enhance features, an improved channel attention mechanism is employed to adaptively weight features from different channels. Channel statistics are extracted through two parallel branches: global average pooling and global maximum pooling. These two pooling operations capture different statistical properties of the feature map. The pooled results are then passed through a shared multi-layer perceptron network to learn inter-channel dependencies. This network consists of two fully connected layers, with the number of neurons in the middle layer being 1 / 16 of the number of input channels. A ReLU activation function is used to enhance nonlinear representation. The output features of the two branches are additively fused and normalized using a sigmoid function to generate a channel weight map. The channel weight map is then multiplied by the enhanced features to achieve adaptive feature recalibration, highlighting channel features that contribute significantly to defect detection. This attention-based feature fusion approach not only improves the model's ability to recognize different types of defects in injection molded workpieces, but also suppresses background noise interference, improving the discriminability of feature representation. The introduction of the channel attention mechanism enables the model to automatically learn the importance of different feature channels, adapting to the needs of diverse detection scenarios.

[0026] S105. Perform deformable convolution on the fused feature map to obtain defect detection results of the injection molded workpiece.

[0027] For example, a modified deformable convolutional network (DCN) replaces the standard convolution operation when performing deformable convolution on the fused feature map. Deformable convolution uses an additional offset learning branch, consisting of a standard convolution layer with 2N output channels (N is the number of sampling points). The offset parameter enables adaptive adjustment of the convolution kernel sampling point positions to accommodate irregularly shaped defects on the surface of injection-molded workpieces. In the detection head network, deformable convolution is combined with a focal loss function, with a modulation factor γ set to 2 to reduce the loss weight of easily distinguishable samples and focus on learning difficult ones. The detection head employs a multi-scale prediction strategy, setting prior bounds of varying sizes at different feature levels to accommodate defects of varying sizes. Detection results include defect location coordinates (center point coordinates, width and height), confidence scores, and category information. Overlapping detection boxes are filtered using an improved non-maximum suppression (Soft-NMS) algorithm. Soft-NMS uses a Gaussian penalty function to smoothly reduce the confidence of overlapping boxes to avoid missing adjacent defects. The injection molded workpiece defect detection results output by the model include both precise defect location information and defect type classification results, supporting subsequent quality analysis and process optimization.

[0028] An embodiment of the present application provides an image detection method, apparatus, device and storage medium for injection-molded workpieces. The method includes: inputting an enhanced image set of the injection-molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map; inputting the affinity feature map into a preset feature analysis network to obtain a main branch attention feature and an auxiliary branch attention feature; performing multi-scale pooling on the main branch attention feature to obtain a main branch feature, performing convolution on the auxiliary branch attention feature to obtain an auxiliary branch feature, and splicing the main branch feature with the auxiliary branch feature to obtain an enhanced feature; using attention channels to fuse the enhanced features to obtain a fused feature map; performing deformable convolution on the fused feature map to obtain a defect detection result of the injection-molded workpiece. In the above method, comprehensive feature perception capability is provided through the dual attention mechanism, the diversity and integrity of features are guaranteed through the dual-branch attention network, and the key features are highlighted in the attention channel fusion to obtain a fused feature map. After the fused feature map is processed by deformable convolution, the adaptability to various defects is improved, and accurate defect detection results of injection molded workpieces are output, which not only ensures the accuracy of detection, but also improves the robustness of the algorithm. It can effectively deal with defects of various types, shapes and sizes on the surface of injection molded workpieces, and significantly improves the overall performance of defect detection.

[0029] In order to more clearly introduce the technical solution of the present application, the technical solution of the present application will be introduced through specific embodiments below. It should be noted that the specific embodiments are used to expand the technical solution of the present application, but are not intended to limit the present application.

[0030] In some embodiments, the enhanced image set of the injection molded workpiece is input into a preset dual-attention extraction network to obtain an affinity feature map, including: S1011-S1015.

[0031] S1011. Input the enhanced image set into a preset dual-attention extraction network, wherein the dual-attention extraction network includes: a first cross-attention unit, a second cross-attention unit and a feature mapping layer, and the first cross-attention unit and the second cross-attention unit share network parameters.

[0032] For example, the dual-attention extraction network uses an improved YOLOv5 backbone network architecture, embedding a dual cross-attention module in the deep feature extraction layer. The network input image has a resolution of 640×640 pixels. After five stages of feature extraction, feature maps at three different scales, 1 / 8, 1 / 16, and 1 / 32, are obtained. The first and second cross-attention units use the same network structure and parameters, and the shared parameter mechanism can reduce computational overhead by 50%. The feature mapping layer uses a 1×1 convolutional structure with 256 output channels to integrate multi-scale feature information.

[0033] S1012. Perform feature dimensionality reduction processing on the enhanced image set through two independent dimensionality reduction convolutional layers to obtain a query feature map Q and a key feature map K.

[0034] For example, both dimensionality reduction convolutional layers use a 1×1 convolutional structure, with the number of convolution kernels being 1 / 8 the number of input feature channels, effectively reducing computational complexity. For the input enhanced image features, the first dimensionality reduction convolutional layer generates a query feature map Q, and the second dimensionality reduction convolutional layer generates a key feature map K. The spatial resolution of the feature maps remains unchanged. Group Normalization is used during the dimensionality reduction process to standardize features, with the number of groups set to 8 to ensure the stability of the feature distribution. The reduced feature maps Q and K have the same number of channels, facilitating the subsequent calculation of attention weights.

[0035] S1013. Use the first cross-attention unit to receive the query feature map Q and the key feature map K, calculate the correlation between each pixel point and its horizontal and vertical pixels, and obtain a first directional correlation feature.

[0036] Exemplarily, the first cross-attention unit performs matrix multiplication on the feature maps Q and K to calculate the attention weight. During the calculation process, the feature map is divided into two branches, horizontal and vertical, and the correlation between the pixel point and the pixels in the same row and the pixels in the same column is calculated respectively. The correlation calculation uses a dot product operation and is normalized using a scaling factor of 1 / sqrt(d), where d is the feature dimension. The calculated correlation score is normalized by the Softmax function to generate a weight coefficient matrix, which reflects the strength of the dependency relationship between pixels at different positions.

[0037] S1014. Input the first directional correlation feature into the second cross-attention unit, further extract the correlation features between the pixels, and obtain the second directional correlation feature.

[0038] Exemplarily, the second cross-attention unit uses the same parameter configuration as the first unit to perform a secondary extraction of the first directional correlation features. By cascading two cross-attention units, the dependency between a pixel and its diagonal pixels can be captured. The second attention calculation uses a residual connection mechanism to add the input features to the attention-weighted features, effectively preventing the vanishing gradient problem. The dropout mechanism is used in the calculation of the attention weights, with a dropout rate set to 0.1, to enhance the generalization ability of the model.

[0039] S1015 , inputting the first direction correlation feature and the second direction correlation feature into a feature mapping layer, normalizing the features through a normalization function, and generating an affinity feature map.

[0040] For example, the feature mapping layer employs an adaptive weighted fusion strategy, fusing correlation features from two directions using learnable weight parameters. Layer Normalization is used during the fusion process for feature standardization, and the parameters are initialized using the He initialization method. This normalization ensures that the feature values are distributed within a reasonable range, which promotes the stability of model training. The resulting affinity feature map preserves the structural differences between surface defects and background areas of the injection molded workpiece. Each channel of the feature map corresponds to a different type of semantic feature, providing an important basis for subsequent defect detection.

[0041] In some embodiments, the affinity feature map is input into a preset feature analysis network to obtain the main branch attention feature and the auxiliary branch attention feature, including: S1021-S1023.

[0042] S1021. Input the affinity feature map into a preset feature analysis network, wherein the feature analysis network includes a main feature processing channel and an auxiliary feature processing channel. The main feature processing channel includes a plurality of exponentially weighted pooling units connected in series, and the auxiliary feature processing channel includes a convolution processing unit.

[0043] For example, the feature analysis network adopts a dual-channel parallel architecture design. The main feature processing channel is connected in series with four exponentially weighted pooling units. The receptive fields of each pooling unit are 5×5, 7×7, 9×9, and 11×11 pixels, respectively. The pooling stride is set to 2 to ensure multi-scale feature extraction. The convolution processing unit in the auxiliary feature processing channel uses a standardized convolution operation with 256 channels, a LeakyReLU activation function with an activation slope of 0.1, batch normalization parameters momentum set to 0.9, and epsilon set to 1e-5, effectively improving the stability and generalization of feature extraction.

[0044] S1022. Based on the exponentially weighted pooling unit, perform feature weighted accumulation on the pixel contribution of the affinity feature map to obtain the main branch attention feature.

[0045] Exemplarily, the exponentially weighted pooling unit assigns different weights to each pixel in the input feature map. The weight calculation is normalized using the softmax function, and the temperature parameter is set to 0.5. During the pooling operation, the feature map is divided into several overlapping regions with an overlap ratio of 0.5. The pixel values in each region are multiplied by the corresponding weights and accumulated. The weight assignment strategy is based on the significance of the pixel values. The more significant the pixel, the greater the weight. The weight decay function uses the exponential form exp(-d² / σ²), where d is the distance between pixels and σ is an adjustable parameter set to 1.5 to ensure the effective retention of surface defect features of the injection molded workpiece.

[0046] S1023. Perform feature extraction on the affinity feature map through the convolution processing unit in the auxiliary feature processing channel to obtain auxiliary branch attention features. The convolution processing unit uses a 1×1 convolution kernel to perform feature dimensionality reduction.

[0047] For example, the weight initialization of the 1×1 convolution kernel in the convolution processing unit adopts the He initialization method, and the standard deviation is set to 0.02. To enhance the robustness of feature extraction, a dropout mechanism is introduced in the convolution layer, and the dropout rate is set to 0.1. The convolution operation adopts the valid mode to ensure the accuracy of the boundary information. During the feature dimensionality reduction process, the number of channels is reduced from the original C to C / 2, where C represents the number of channels of the input feature map. The feature map after dimensionality reduction is processed by spatial pyramid pooling, and the pooling scale includes three levels: 1×1, 2×2, and 4×4. The pooling results are spliced to form a multi-scale feature representation, enhancing the model's adaptability to defects of different sizes.

[0048] In some embodiments, based on an exponentially weighted pooling unit, feature weighted accumulation is performed on the pixel contributions of the affinity feature map to obtain the main branch attention features, including: S221-S225.

[0049] S221. Adaptively block the affinity feature map to divide the feature map into multiple feature blocks of equal size, where the size of each feature block matches the pooling window size of the exponentially weighted pooling unit.

[0050] S222. Perform an exponential transformation operation on the pixel values in each feature block to obtain a first weight coefficient. The exponential transformation operation uses a learnable exponential base for adaptively adjusting the weight distribution.

[0051] S223. Based on the first weight coefficient, calculate the relative importance score of each pixel in the feature block, and generate a position weighting coefficient in combination with the spatial position information of the pixel.

[0052] S224 . Adaptively fuse the first weight coefficient with the position weight coefficient to obtain a second weight coefficient.

[0053] S225. Perform softmax normalization on the second weight coefficient, perform weighted sum operation on the normalized second weight coefficient and the corresponding pixel value to obtain the weighted feature value of the feature block, and cascade the weighted feature values of all feature blocks to generate the main branch attention feature.

[0054] For example, when an affinity feature map is processed by an exponentially weighted pooling unit, the feature map is divided into several n×n feature blocks, where n is the pooling window size. For each pixel value xij within a feature block, a learnable exponential base α (α>0) is introduced to calculate the exponential transformation exp(α·xij) to obtain the initial weight. This exponential base can be dynamically adjusted during training, enabling the model to adaptively assign weights based on different types of features. The first weight coefficient wi,j calculated based on the initial weights reflects the basic importance of the pixel in the feature block. Taking into account the spatial distribution of features, a two-dimensional Gaussian function G(i,j) is introduced as a position weight template. This function takes the center of the feature block as its origin and generates position weight coefficients pi,j based on the distance from the pixel to the center, giving pixels closer to the center a larger weight. The first weight coefficient is adaptively combined with the position weight coefficient using a learnable fusion parameter λ (0≤λ≤1) to obtain the second weight coefficient. The second weight coefficient is softmax-normalized to ensure that the sum of the weights is 1. The normalized weights are multiplied and summed with the original pixel values to obtain the weighted eigenvalues of the feature block. The weighted eigenvalues of all feature blocks are rearranged and combined according to their original spatial positional relationships to generate main branch attention features with the same spatial dimensions as the input feature map. This process preserves important feature information while taking into account spatial positional relationships, effectively extracting salient features of the target area, suppressing background interference, and improving feature expression capabilities. Through learnable exponential cardinality and fusion parameters, the model can adaptively adjust the feature extraction strategy according to different task requirements, enhancing the algorithm's flexibility and generalization capabilities.

[0055] In some embodiments, multi-scale pooling is performed on the main branch attention features to obtain the main branch features, convolution is performed on the auxiliary branch attention features to obtain the auxiliary branch features, and the main branch features and the auxiliary branch features are spliced to obtain enhanced features, including: S1031-S1036.

[0056] S1031. Adaptively group the main branch attention features into multiple independent feature groups.

[0057] Exemplarily, a dynamic channel grouping algorithm is used to process the main branch attention features. By calculating the correlation matrix between the channels of the feature map, a correlation threshold of 0.6 is set, and channels with correlations above the threshold are automatically clustered into one group. An adaptive grouping strategy based on spectral clustering is used in the grouping process. Channels between 64 and 512 are divided into four groups, and those above 512 are divided into eight groups. The number of channels in each group is dynamically adjusted to maintain a balanced distribution of information. An additional attention gating mechanism is introduced for each channel group, with a gating threshold set to 0.3 to suppress the responses of invalid feature channels and improve the discriminability of feature expression.

[0058] S1032. Apply pooling operations of different scales to each independent feature group to generate a multi-scale feature representation, wherein the original feature information is retained through skip connections during the pooling operation.

[0059] For example, SoftPool pooling operations with three different receptive fields of 5×5, 7×7, and 9×9 are performed on each independent feature group, with a pooling step size of 2 and a reflect padding mode. During the pooling process, the original features are reduced to the same number of channels as the pooled features through 1×1 convolution, and skip connection pathways are established. The skip connection uses an additive fusion method, and the fusion weights are automatically adjusted using learnable parameters. The weights are initialized using the Xavier method. For the pooling results at each scale, spatially adaptive instance normalization is used for feature normalization. The normalization parameters include the scaling factor γ and the bias term β, and the initial values of the parameters are set to 1.0 and 0, respectively.

[0060] S1033. Use the channel attention mechanism to adaptively weight the multi-scale feature representation, calculate the importance weight of each channel, and generate weighted multi-scale main branch features.

[0061] For example, a two-layer channel attention module is constructed. The first layer uses global average pooling and global maximum pooling to process feature maps in parallel. The pooling results are transformed by a shared multi-layer perceptron. The number of neurons in the middle layer of the perceptron is 1 / 16 of the number of input channels. The second layer uses inter-channel relationship modeling and uses matrix multiplication to calculate the similarity score between channels. The similarity matrix is normalized by Softmax to generate attention weights. The results of the two layers of attention are fused using weighted averaging. The weight coefficient is automatically learned through trainable parameters and initialized to 0.5. The attention-weighted features are channel-recalibrated, and the recalibration parameter γ is initialized to 0.1.

[0062] S1034. Apply 1×1 convolution to the auxiliary branch attention feature to perform channel dimensionality reduction, and then use 3×3 depth-separable convolution to extract spatial features to obtain the initial auxiliary branch feature.

[0063] For example, a 1×1 convolution is used to compress the auxiliary branch attention features, with a compression ratio of 2 and kernel parameters initialized using the He method. The compressed features are then fed into a 3×3 depthwise separable convolution layer with a depth multiplier of 2. The number of grouped convolutions equals the number of input channels, and each group of independent convolution kernels extracts spatial features. A dilated convolution mechanism is introduced into the depthwise separable convolution, with a dilation ratio of 2 to expand the receptive field. The convolution layer is followed by a BatchNorm normalization layer with a momentum parameter of 0.9 and an epsilon of 1e-5. A PReLU activation function is used, with the negative semi-axis slope initialized to 0.25.

[0064] S1035. Input the initial auxiliary branch features into the residual connection module, extract features of different receptive fields through parallel 1×1 convolution and 3×3 convolution branches, and add the results to obtain the optimized auxiliary branch features.

[0065] For example, a dual-branch residual architecture is designed. The main branch uses a three-layer structure consisting of 1×1 convolution for dimensionality reduction, 3×3 convolution for feature extraction, and 1×1 convolution for dimensionality increase. The channel compression ratio is 4, and the convolution kernel weights are initialized using a normal distribution with a standard deviation of 0.02. The shortcut branch uses 1×1 convolution for channel matching and introduces a group normalization mechanism, with the number of feature groups set to 8. The outputs of the two branches are combined using an additive fusion method. Before fusion, the features are L2-normalized with a normalization scale factor of 20. The residual connection module also includes a channel permutation operation with a permutation unit size of 4 to enhance the expressiveness of features.

[0066] S1036. Align the weighted multi-scale main branch features and the optimized auxiliary branch features in the spatial dimension and channel dimension, then concatenate them, and perform feature fusion and dimensionality reduction through 1×1 convolution to obtain enhanced features.

[0067] For example, the main branch features and auxiliary branch features are aligned, and the spatial dimensions are resized to the same size using bilinear interpolation, with bilinear as the interpolation mode. 1×1 convolution is used to resize the channel dimensions to the same number of channels, with the number of feature channels for both the main and auxiliary branches adjusted to 256. A channel importance evaluation module is introduced before feature concatenation, and the SE attention mechanism is used to calculate channel weights. The compression ratio is set to 16. The concatenated features are fused using 1×1 convolution, with the number of output channels set to 512. The convolution kernel parameters are initialized using the Kaiming uniform distribution. The fused features are recalibrated, and the recalibration parameter learning rate is set to 0.01.

[0068] In some embodiments, deformable convolution is performed on the fused feature map to obtain defect detection results of the injection molded workpiece, including: S1051-S1056.

[0069] S1051. Input the fused feature map into a preset deformable convolutional network, wherein the preset deformable convolutional network includes: an offset prediction layer and a feature sampling layer, the offset prediction layer is used to generate a spatial offset of the sampling point, and the feature sampling layer is used to extract features according to the offset sampling position.

[0070] For example, the deformable convolutional network uses a 3×3 convolution kernel structure, with each kernel containing 9 sampling points. The offset prediction layer consists of two consecutive standard convolutional layers with 256 and 18 channels, respectively. The 18 channels correspond to the x and y offsets of the 9 sampling points. The feature sampling layer uses a deformable convolution operator with 256 input channels, 512 output channels, a stride of 1, and a padding of SAME. The network parameters are initialized using the Xavier method to ensure the stability of network training. The deformable convolutional network can adaptively adjust the shape and size of the receptive field based on the morphological characteristics of surface defects in injection molded workpieces.

[0071] S1052. Perform a convolution operation on the fused feature map through the offset prediction layer to predict the two-dimensional spatial offset of each feature point to obtain an offset parameter map, which represents the position offset of the convolution kernel sampling point relative to the regular grid.

[0072] For example, the offset prediction layer processes the 256-channel fused feature map to generate an 18-channel offset parameter map. The offset prediction range is limited to [-1, 1] and is implemented using the tanh activation function. For each location (x, y) on the feature map, the offset △pn=(△x, △y) of 9 sampling points is calculated, where n∈[1, 9]. The attention mechanism is used in the offset prediction process to assign different weights to different sampling points. The weight values are automatically adjusted based on the local response strength of the feature map. This design enables the network to dynamically adjust the sampling position according to the specific shape of the injection molded workpiece defect, enhancing its adaptability to irregular defects.

[0073] S1053. Dynamically adjust the sampling position of the original convolution kernel based on the offset parameter map, calculate the new sampling coordinates, and use the bilinear interpolation method to resample the features of the non-integer coordinate positions to obtain an adaptive feature map.

[0074] For example, for each position p0 on the feature map, the new coordinates of nine offset sampling points are calculated: pn = p0 + pn + △pn, where n∈[1,9]. Because the calculated sampling coordinates are typically non-integer values, bilinear interpolation is used to perform a weighted average of the features at the four nearest integer coordinate positions, with the interpolation weight being inversely proportional to the distance from the sampling point to the integer coordinate. The interpolation calculation is accelerated using CUDA, and the effective range of the sampling points is expanded by zero padding. This adaptive sampling strategy enables the network to flexibly adjust the shape of the receptive field based on the morphological characteristics of the injection molded workpiece defects, improving the detection accuracy of various defects.

[0075] S1054: Input the adaptive feature map into a multi-scale detection head network, where the multi-scale detection head network includes a classification branch and a regression branch, which are used to predict the category probability and bounding box coordinates of the target respectively.

[0076] For example, the multi-scale detection head network makes predictions on feature maps at three different scales, with feature map sizes of 1 / 8, 1 / 16, and 1 / 32 of the input image, respectively. Both the classification and regression branches utilize 3×3 convolutional layers. The classification branch outputs 3 channels (3 prior bounding boxes at each scale), while the regression branch outputs 12 channels (4 coordinate parameters per prior bounding box). The detection head network utilizes a shared weight design, using the same classification and regression parameters for the feature maps at the three scales to reduce network parameter requirements. This multi-scale detection strategy can adapt to defects of varying sizes on injection molded parts and improve scale robustness.

[0077] S1055. Apply the sigmoid function to the output feature map of the classification branch to obtain a category probability map, and perform coordinate decoding on the output feature map of the regression branch to obtain the prediction box coordinates.

[0078] For example, the output of the classification branch is converted to the interval [0, 1] by the sigmoid function, indicating the probability that each prior box is predicted to be each defect category. The output of the regression branch uses relative coordinate encoding, and the predicted value includes the offset (tx, ty) of the center point of the bounding box relative to the prior box and the scale factor (tw, th). The coordinate decoding process uses an exponential function to process the scale factor to ensure that the width and height of the predicted box are positive values. The actual coordinates of the predicted box are calculated using the prior box size and the predicted offset. The calculation process uses vectorized operations to improve efficiency. This encoding and decoding mechanism improves the stability and accuracy of bounding box regression.

[0079] S1056. Post-process the prediction frames using a non-maximum suppression algorithm to filter out redundant detection frames with high overlap, and obtain defect detection results for the injection molded workpiece, including defect category, location coordinates, and confidence score.

[0080] For example, the non-maximum suppression algorithm sets the IoU threshold to 0.5 and the confidence threshold to 0.3. For each defect category, all predicted boxes are sorted in descending order by confidence. The predicted box with the highest confidence is retained, and the IoU value between this box and other predicted boxes is calculated. Redundant boxes with an IoU greater than the threshold are removed. The confidence score of the predicted box is determined by the product of the category probability and the objectness score. The detection results are output in JSON format, containing the category label, bounding box coordinates [x1, y1, x2, y2], and confidence score for each defect. This post-processing strategy effectively removes duplicate detections and provides accurate defect detection results.

[0081] See also Figure 2 , Figure 21 is a schematic block diagram of an image detection device for an injection molded workpiece according to an embodiment of the present application. The image detection device 200 for an injection molded workpiece is used to perform the aforementioned image detection method for an injection molded workpiece. The image detection device 200 for an injection molded workpiece can be configured in a server.

[0082] Among them, the server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0083] like Figure 2 As shown, the image detection device 200 for injection molded workpieces includes: a feature extraction module 201 , a feature analysis module 202 , a feature enhancement module 203 , a feature fusion module 204 and a result output module 205 .

[0084] The feature extraction module 201 is used to input the enhanced image set of the injection molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map.

[0085] The feature analysis module 202 is used to input the affinity feature map into a preset feature analysis network to obtain the main branch attention feature and the auxiliary branch attention feature.

[0086] The feature enhancement module 203 is used to perform multi-scale pooling on the main branch attention feature to obtain the main branch feature, perform convolution on the auxiliary branch attention feature to obtain the auxiliary branch feature, and splice the main branch feature with the auxiliary branch feature to obtain the enhanced feature.

[0087] The feature fusion module 204 is used to fuse the enhanced features using the attention channel to obtain a fused feature map.

[0088] The result output module 205 is used to perform deformable convolution on the fused feature map to obtain the defect detection result of the injection molded workpiece.

[0089] An embodiment of the present application provides an electronic device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement an image detection method for an injection molded workpiece as described in any one of the embodiments of the present application when executing the computer program.

[0090] An embodiment of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor implements an image detection method for an injection molded workpiece as described in any one of the embodiments of the present application.

[0091] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for image detection of an injection molded workpiece, characterized in that: The method comprises: The enhanced image set of the injection molded workpiece is input into the preset dual attention extraction network to obtain the affinity feature map; Inputting the affinity feature map into a preset feature analysis network to obtain main branch attention features and auxiliary branch attention features; Performing multi-scale pooling on the main branch attention feature to obtain the main branch feature, performing convolution on the auxiliary branch attention feature to obtain the auxiliary branch feature, and splicing the main branch feature and the auxiliary branch feature to obtain the enhanced feature; Using the attention channel to fuse the enhanced features, a fused feature map is obtained; Performing deformable convolution on the fused feature map to obtain a defect detection result of the injection molded workpiece.

2. The image detection method for injection molded workpiece according to claim 1, wherein: Before inputting the enhanced image set of the injection molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map, the method further includes: Taking an image of the injection molded workpiece to be inspected; Performing Gaussian filtering and median filtering on the image to be detected to obtain a denoised image; Performing adaptive histogram equalization processing on the denoised image to obtain a light-equalized image; Data enhancement is performed on the illumination-balanced image in a preset manner to obtain an enhanced image set.

3. The image detection method for injection molded workpiece according to claim 1, wherein: The enhanced image set of the injection molded workpiece is input into a preset dual attention extraction network to obtain an affinity feature map, including: Inputting the enhanced image set into a preset dual-attention extraction network, wherein the dual-attention extraction network includes: a first cross-attention unit, a second cross-attention unit, and a feature mapping layer, and the first cross-attention unit and the second cross-attention unit share network parameters; Performing feature dimensionality reduction processing on the enhanced image set through two independent dimensionality reduction convolutional layers to obtain a query feature map Q and a key feature map K; Utilizing the first cross attention unit to receive the query feature map Q and the key feature map K, calculating the correlation between each pixel and its horizontal and vertical pixels to obtain a first directional correlation feature; Inputting the first directional correlation feature into the second cross attention unit to further extract correlation features between pixels to obtain a second directional correlation feature; The first directional correlation feature and the second directional correlation feature are input into the feature mapping layer, and the features are normalized by a normalization function to generate the affinity feature map.

4. The image detection method for injection molded workpiece according to claim 1, wherein: The step of inputting the affinity feature map into a preset feature analysis network to obtain the main branch attention feature and the auxiliary branch attention feature comprises: Inputting the affinity feature map into a preset feature analysis network, wherein the feature analysis network includes a main feature processing channel and an auxiliary feature processing channel, the main feature processing channel includes a plurality of serially connected exponentially weighted pooling units, and the auxiliary feature processing channel includes: a convolution processing unit; Based on the exponentially weighted pooling unit, the pixel contributions of the affinity feature map are subjected to feature weighted accumulation to obtain the main branch attention feature; The affinity feature map is subjected to feature extraction by a convolution processing unit in the auxiliary feature processing channel to obtain auxiliary branch attention features, and the convolution processing unit uses a 1×1 convolution kernel to perform feature dimensionality reduction.

5. The image detection method for injection molded workpiece according to claim 4, wherein: The exponentially weighted pooling unit is used to perform feature weighted accumulation on the pixel contributions of the affinity feature map to obtain the main branch attention feature, including: Adaptively partitioning the affinity feature map into multiple feature blocks of equal size, wherein the size of each feature block matches the pooling window size of the exponentially weighted pooling unit; Performing an exponential transformation operation on the pixel values in each feature block to obtain a first weight coefficient, wherein the exponential transformation operation uses a learnable exponential base for adaptively adjusting the weight distribution; Based on the first weight coefficient, the relative importance score of each pixel in the feature block is calculated, and the position weight coefficient is generated in combination with the spatial position information of the pixel; Adaptively fusing the first weight coefficient with the position weight coefficient to obtain a second weight coefficient; The second weight coefficient is subjected to softmax normalization processing, and the normalized second weight coefficient is weighted and summed with the corresponding pixel value to obtain the weighted eigenvalue of the feature block, and the weighted eigenvalues of all feature blocks are cascaded and spliced to generate the main branch attention feature.

6. The image detection method for injection molded workpiece according to claim 1, wherein: The method of performing multi-scale pooling on the main branch attention feature to obtain the main branch feature, performing convolution on the auxiliary branch attention feature to obtain the auxiliary branch feature, and splicing the main branch feature with the auxiliary branch feature to obtain the enhanced feature includes: Adaptively grouping the main branch attention features into multiple independent feature groups; Applying pooling operations of different scales to each of the independent feature groups to generate multi-scale feature representations, wherein original feature information is retained through skip connections during the pooling operation; Adaptively weighting the multi-scale feature representation using a channel attention mechanism, calculating the importance weight of each channel, and generating weighted multi-scale main branch features; Apply 1×1 convolution to the auxiliary branch attention feature to perform channel dimensionality reduction, and then use 3×3 depth-separable convolution to extract spatial features to obtain the initial auxiliary branch feature; Input the initial auxiliary branch features into the residual connection module, extract features of different receptive fields through parallel 1×1 convolution and 3×3 convolution branches, and add the results to obtain optimized auxiliary branch features; The weighted multi-scale main branch features and the optimized auxiliary branch features are aligned in the spatial dimension and the channel dimension, and then spliced, and feature fusion and dimensionality reduction are performed through 1×1 convolution to obtain enhanced features.

7. The image detection method for injection molded workpiece according to claim 1, wherein: The step of performing deformable convolution on the fused feature map to obtain a defect detection result of the injection molded workpiece includes: Inputting the fused feature map into a preset deformable convolutional network, wherein the preset deformable convolutional network includes: an offset prediction layer and a feature sampling layer, wherein the offset prediction layer is used to generate a spatial offset of a sampling point, and the feature sampling layer is used to extract features according to the offset sampling position; Performing a convolution operation on the fused feature map through the offset prediction layer to predict the two-dimensional spatial offset of each feature point to obtain an offset parameter map, wherein the offset parameter map represents the position offset of the convolution kernel sampling point relative to the regular grid; Dynamically adjust the sampling position of the original convolution kernel based on the offset parameter map, calculate new sampling coordinates, and use a bilinear interpolation method to resample features at non-integer coordinate positions to obtain an adaptive feature map; Inputting the adaptive feature map into a multi-scale detection head network, wherein the multi-scale detection head network includes a classification branch and a regression branch, which are respectively used to predict the category probability and bounding box coordinates of the target; Applying a sigmoid function to the output feature map of the classification branch to obtain a category probability map, and performing coordinate decoding on the output feature map of the regression branch to obtain prediction box coordinates; The non-maximum suppression algorithm is used to post-process the prediction frames to filter out redundant detection frames with high overlap, and the defect detection results of the injection molded workpiece are obtained, including defect category, location coordinates and confidence score.

8. An image detection device for injection molded workpieces, characterized in that: The image detection device for an injection molded workpiece is used to perform the image detection method for an injection molded workpiece according to any one of claims 1 to 7, and the image detection device for an injection molded workpiece comprises: A feature extraction module is used to input the enhanced image set of the injection molded workpiece into a preset dual-attention extraction network to obtain an affinity feature map; A feature analysis module, configured to input the affinity feature map into a preset feature analysis network to obtain a main branch attention feature and an auxiliary branch attention feature; A feature enhancement module is used to perform multi-scale pooling on the main branch attention feature to obtain the main branch feature, perform convolution on the auxiliary branch attention feature to obtain the auxiliary branch feature, and splice the main branch feature with the auxiliary branch feature to obtain the enhanced feature; A feature fusion module, configured to fuse the enhanced features using an attention channel to obtain a fused feature map; The result output module is used to perform deformable convolution on the fused feature map to obtain the defect detection result of the injection molded workpiece.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the image detection method for injection molded workpieces according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the image detection method for injection-molded workpieces according to any one of claims 1 to 7.

Citation Information

Cited By

  • Deformable convolution-based convolutional neural network firmware vulnerability detection method

    CN121328631A

  • Target detection method and device

    CN121685938A

  • A target detection method and apparatus

    CN121685938B