Packing material defect detection method and system based on machine vision
By constructing a depth map for light compensation and geometric correction, combining multi-scale direction filters and feature pyramid structures, a double-level cascade classifier is used to solve the problems of unstable lighting, insufficient feature description and high computational complexity in the existing machine vision defect detection technology, and efficient and accurate packaging material defect detection is achieved.
Patent Information
- Application Number
- CN202510490723.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
The existing machine vision defect detection technology is insufficient in light compensation and geometric correction, resulting in unstable image quality, feature extraction methods fail to fully consider the diversity and complexity of packaging material defects, and the classifier has high computational complexity and poor real-time performance.
By collecting multi-angle image sequences, the depth map is constructed for geometric correction and lighting compensation, the multi-scale direction filter group and adaptive direction enhancement factor are used to generate feature map sequences, the feature pyramid structure is constructed for hierarchical segmentation, multi-dimensional features are extracted, and the class cohesion and inter-class distance are calculated, and the two-level cascade classifiers are used for classification.
It improves the accuracy and efficiency of defect detection, enhances the ability to adapt to different types of defects, and significantly improves the detection accuracy and real-timeness.
Smart Images

Figure CN120356002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to machine vision technology, and in particular to a method and system for detecting defects on packaging materials based on machine vision. Background Art
[0002] In modern industrial production, the quality of packaging materials directly affects the safety and market competitiveness of products. With the development of automation and intelligent technologies, defect detection methods based on machine vision have gradually become an important means for quality control of packaging materials. Through image processing technology, such methods can quickly and accurately identify and classify defects on packaging materials, thereby improving production efficiency and reducing labor costs.
[0003] However, existing machine vision defect detection technologies still have some defects and deficiencies. First, existing technologies often lack effective methods for light compensation and geometric correction in the image preprocessing stage, resulting in unstable image quality obtained under different lighting conditions, thus affecting the subsequent defect detection accuracy. Second, existing feature extraction methods usually rely on fixed feature parameters and fail to fully consider the diversity and complexity of packaging material defects, resulting in insufficient feature description ability and unable to effectively distinguish different types of defects. Finally, existing classifiers often face problems of high computational complexity and poor real-time performance when dealing with high-dimensional features, affecting the response speed and practicality of the defect detection system.
[0004] Therefore, there is an urgent need for a new method for detecting defects on packaging materials based on machine vision, which can effectively solve the above problems and improve the accuracy and efficiency of defect detection. Summary of the Invention
[0005] Embodiments of the present invention provide a method and system for detecting defects on packaging materials based on machine vision, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] A method for detecting defects on packaging materials based on machine vision is provided, including:
[0008] Collecting a multi-angle image sequence and constructing a depth map through camera calibration parameters, and performing geometric correction and light compensation on the multi-angle image sequence according to the depth map to generate a preprocessed enhanced image;
[0009] Constructing a multi-scale directional filter bank based on the preprocessed enhanced image, introducing an adaptive directional enhancement factor into the multi-scale directional filter bank to generate a sequence of feature maps, calculating local contrast and maximum response direction based on the sequence of feature maps, constructing a feature pyramid structure, optimizing through an edge-preserving constraint term and generating a sequence of adaptive segmentation thresholds, and realizing hierarchical segmentation of the defect area based on the sequence of adaptive segmentation thresholds;
[0010] Extract a multi-dimensional feature set from the hierarchical segmented defect regions. The multi-dimensional feature set includes local texture features, grayscale statistical features, and shape description features. Calculate the within-class cohesion and between-class distance of the multi-dimensional feature set to construct a feature evaluation matrix, and assign weight coefficients to the features of each dimension according to the feature evaluation matrix, and fuse them to generate an enhanced discriminative feature description vector;
[0011] Input the enhanced discriminative feature description vector into a two-stage cascade classifier. The first stage of the two-stage cascade classifier uses a lightweight decision tree for rapid pre-classification, and the second stage of the two-stage cascade classifier extracts local detail features and global semantic features through an attention-enhanced deep network, and uses a feature adaptive fusion module to integrate the classification results of the two stages.
[0012] Calculate the local contrast and the maximum response direction based on the feature map sequence, construct a feature pyramid structure, optimize and generate an adaptive segmentation threshold sequence through an edge-preserving constraint term. The hierarchical segmentation of the defect region based on the adaptive segmentation threshold sequence includes:
[0013] Construct a multi-scale direction filter bank. The multi-scale direction filter bank includes Gaussian filter kernels of different scales and directions, and a direction enhancement factor is introduced into the Gaussian filter kernels. The direction enhancement factor is adaptively adjusted according to the main direction angle and the direction selectivity parameter;
[0014] Filter the input image using the multi-scale direction filter bank to obtain multiple groups of direction response feature maps, construct a direction response mapping function based on the multiple groups of direction response feature maps, and optimize the weight coefficients of different scales by minimizing the reconstruction error and sparse regularization constraints;
[0015] Calculate the maximum response direction according to the direction response mapping function, construct a local contrast map by combining the local mean and standard deviation, and adaptively fuse the maximum response direction with the local contrast map to generate enhanced edge features;
[0016] Construct a feature pyramid based on the enhanced edge features, perform feature transformation through downsampling and upsampling operations at each level of the feature pyramid, and introduce an edge-preserving constraint term. The edge-preserving constraint term is optimized by calculating the feature gradient difference;
[0017] Extract statistical features at each level of the feature pyramid, and generate an adaptive segmentation threshold sequence recursively according to the statistical features and the threshold of the previous level. The adaptive segmentation threshold sequence is used to achieve the hierarchical segmentation of multi-scale defect regions.
[0018] Construct a feature pyramid based on the enhanced edge features, perform feature transformation through downsampling and upsampling operations at each level of the feature pyramid, and introduce an edge-preserving constraint term. The edge-preserving constraint term is optimized by calculating the feature gradient difference, including:
[0019] Perform continuous downsampling operations on the enhanced edge features to construct a feature pyramid. The feature pyramid contains multiple levels, and each level maps the features through a learnable feature transformation matrix;
[0020] Between adjacent levels of the feature pyramid, use the upsampling operation to map the features of the resolution level lower than the preset resolution threshold to the feature scale space of the resolution level higher than the preset resolution threshold, and adaptively fuse the resolution level lower than the preset resolution threshold and the resolution level higher than the preset resolution threshold to generate fused features;
[0021] Introduce an edge-preserving constraint term to the fused features. The edge-preserving constraint term is constrained by calculating the feature gradient difference and feature structure difference between adjacent levels. The feature gradient difference is measured by the first-order norm, and the feature structure difference is measured by the second-order norm; construct a multi-level consistency loss function based on the edge-preserving constraint term, and optimize the learnable feature transformation matrix by minimizing the multi-level consistency loss function.
[0022] Establish a feature evaluation matrix by calculating the intra-class cohesion and inter-class distance of each dimension feature, and dynamically allocate the weight coefficients of each dimension feature based on the feature evaluation matrix to fuse and generate a discriminative enhanced feature description vector, including:
[0023] Construct an adaptive bandwidth parameter based on the local data distribution characteristics. The adaptive bandwidth parameter is dynamically adjusted by calculating the Euclidean distance from the sample to the class center and the local density estimate value calculated using the multi-scale Gaussian kernel function;
[0024] Use the adaptive bandwidth parameter to construct a dynamic kernel function. The dynamic kernel function combines a distance modulation function to evaluate the structural similarity between sample pairs, and determines the optimal neighborhood set based on the structural similarity;
[0025] Apply an adaptive weight to the sample pairs in the optimal neighborhood set based on the dynamic kernel function, calculate the weighted intra-class cohesion of each dimension feature, and introduce a temporal consistency constraint to ensure the temporal continuity of the weighted intra-class cohesion;
[0026] Construct a first feature evaluation matrix, which includes the weighted intra-class cohesion and the inter-class distance metric obtained by calculating the Mahalanobis distance and distribution divergence between different class centers;
[0027] Construct a second feature evaluation matrix based on the weighted intra-class cohesion and the inter-class distance metric, and use the softmax function to convert the combined discriminative scores of the first feature evaluation matrix and the second feature evaluation matrix into weight coefficients; use the weight coefficients to adaptively fuse the features of each dimension to generate a discriminative enhanced feature description vector.
[0028] Constructing a second feature evaluation matrix based on the weighted intra-class cohesion and the inter-class distance metric, and using the softmax function to convert the combined discriminative scores of the first feature evaluation matrix and the second feature evaluation matrix into weight coefficients includes:
[0029] Construct a second feature evaluation matrix based on the discriminative scores calculated from the weighted intra-class cohesion and the inter-class distance metric. The second feature evaluation matrix is optimized by the node features in the multi-dimensional feature correlation graph structure. The multi-dimensional feature correlation graph structure includes a node set, an edge set, and an edge weight set. The node features are generated through feature correlation calculation and graph attention mechanism;
[0030] Combine the first feature evaluation matrix and the second feature evaluation matrix to obtain discriminative scores. The discriminative scores are obtained by applying adaptive weight coefficients to the first feature evaluation matrix and the second feature evaluation matrix and performing weighted summation operations. The adaptive weight coefficients are determined based on the reliability and discriminative ability of each feature evaluation matrix;
[0031] Use the softmax function to convert the discriminative scores into weight coefficients, and the weight coefficients are used to adaptively fuse the features of each dimension.
[0032] The first stage of the two-stage cascade classifier uses a lightweight decision tree to achieve fast pre-classification. The second stage introduces an attention-enhanced deep network to extract local detail features and global semantic features. The classification results of the two stages are integrated through a feature adaptive fusion module to achieve accurate classification, including:
[0033] Construct a two-stage cascade classifier. The first stage of the two-stage cascade classifier uses a lightweight decision tree for fast pre-classification. The lightweight decision tree obtains a sample difficulty evaluation value by calculating the distance from the sample to the decision boundary and the local density, and dynamically weights the training samples according to the sample difficulty evaluation value;
[0034] Construct a curriculum learning strategy based on the sample difficulty evaluation value. The curriculum learning strategy determines training samples according to the dynamic distribution of sample difficulty, generates adversarial samples for the training samples, and inputs the adversarial samples into the second stage of the two-stage cascade classifier;
[0035] The second stage of the double - stage cascade classifier includes an attention - enhanced deep network. Based on the attention - enhanced deep network, local detail features and global semantic features are respectively extracted. Adaptive noise injection is performed on the local detail features and the global semantic features, and the feature consistency loss from different perspectives is calculated.
[0036] The pre - classification result of the lightweight decision tree and the feature extraction result of the attention - enhanced deep network are adaptively fused. The adaptive fusion is obtained by weighted calculation of the confidence levels of the two - stage classification results.
[0037] Based on the change rate of the confidence level, the decision threshold is dynamically adjusted, and the decision threshold is used to optimize the fused classification result. The reliability of the classification result is evaluated and corrected by calculating the prediction probability in the confusion matrix, and the final classification result is generated.
[0038] In the second aspect of the embodiments of the present invention,
[0039] A packaging material defect detection system based on machine vision is provided, including:
[0040] A first unit for collecting multi - angle image sequences and constructing a depth map through camera calibration parameters, geometrically correcting and performing illumination compensation on the multi - angle image sequences according to the depth map, and generating a pre - processed enhanced image.
[0041] A second unit for constructing a multi - scale direction filter bank based on the pre - processed enhanced image. The multi - scale direction filter bank introduces an adaptive direction enhancement factor to generate a sequence of feature maps, calculates the local contrast and the maximum response direction based on the sequence of feature maps, constructs a feature pyramid structure, optimizes through an edge - preserving constraint term and generates an adaptive segmentation threshold sequence, and realizes hierarchical segmentation of the defect area based on the adaptive segmentation threshold sequence.
[0042] A third unit for extracting a multi - dimensional feature set from the hierarchically segmented defect area. The multi - dimensional feature set includes local texture features, gray - scale statistical features, and shape description features, calculates the within - class cohesion and between - class distance of the multi - dimensional feature set to construct a feature evaluation matrix, assigns weight coefficients to each - dimension feature according to the feature evaluation matrix, and fuses to generate an enhanced discriminative feature description vector.
[0043] A fourth unit for inputting the enhanced discriminative feature description vector into a double - stage cascade classifier. The first stage of the double - stage cascade classifier uses a lightweight decision tree for fast pre - classification, and the second stage of the double - stage cascade classifier extracts local detail features and global semantic features through an attention - enhanced deep network, and uses a feature adaptive fusion module to integrate the classification results of the two stages.
[0044] In the third aspect of the embodiments of the present invention,
[0045] A kind of electronic device is provided, including:
[0046] A processor;
[0047] A memory for storing instructions executable by the processor;
[0048] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0049] In the fourth aspect of the embodiments of the present invention,
[0050] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0051] The beneficial effects of this application are as follows:
[0052] Through the acquisition of multi-angle image sequences and the construction of depth maps, high-precision geometric correction and illumination compensation of packaging material defects are realized, so that the generated preprocessed enhanced images can effectively improve the accuracy of subsequent feature extraction.
[0053] By introducing an adaptive direction enhancement factor and an edge-preserving constraint term, the constructed feature pyramid structure can better capture the detailed information of the defect area, realizing hierarchical segmentation of the defect area and improving the sensitivity and accuracy of defect detection.
[0054] The design of the double-stage cascade classifier combines fast pre-classification with detailed feature extraction, and uses the feature adaptive fusion module to integrate the classification results, significantly improving the classification efficiency and discrimination ability, and enhancing the detection performance of packaging material defects. Description of the Drawings
[0055] Figure 1 It is a schematic flow chart of the method for detecting packaging material defects based on machine vision according to the embodiments of the present invention;
[0056] Figure 2 It is a table graph for evaluating the performance of the feature pyramid segmentation technology according to the embodiments of the present invention;
[0057] Figure 3 It is a table graph for comparing the comprehensive performance of the feature evaluation matrix and the weight coefficient adaptive fusion technology according to the embodiments of the present invention;
[0058] Figure 4 It is a table graph for comparing the performance data of the feature evaluation matrix according to the embodiments of the present invention;
[0059] Figure 5 It is a table graph for comparing the performance indexes of the double-stage cascade classifier and the prior art according to the embodiments of the present invention. Detailed implementation manners
[0060] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] The technical solutions of the present invention will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0062] Figure 1 It is a flowchart of a method for detecting defects in packaging materials based on machine vision according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0063] Collect a multi-angle image sequence and construct a depth map through camera calibration parameters, and perform geometric correction and illumination compensation on the multi-angle image sequence according to the depth map to generate a preprocessed enhanced image;
[0064] Construct a multi-scale direction filter bank based on the preprocessed enhanced image, the multi-scale direction filter bank introduces an adaptive direction enhancement factor to generate a sequence of feature maps, calculate the local contrast and the maximum response direction based on the sequence of feature maps, construct a feature pyramid structure, optimize through an edge-preserving constraint term and generate an adaptive segmentation threshold sequence, and implement hierarchical segmentation of the defect area based on the adaptive segmentation threshold sequence;
[0065] Extract a multi-dimensional feature set from the defect area obtained by the hierarchical segmentation. The multi-dimensional feature set includes local texture features, gray-scale statistical features and shape description features, calculate the within-class cohesion and between-class distance of the multi-dimensional feature set to construct a feature evaluation matrix, and assign weight coefficients to the features of each dimension according to the feature evaluation matrix, and fuse to generate an enhanced discriminative feature description vector;
[0066] Input the enhanced discriminative feature description vector into a two-stage cascade classifier. The first stage of the two-stage cascade classifier uses a lightweight decision tree for rapid pre-classification, and the second stage of the two-stage cascade classifier extracts local detail features and global semantic features through an attention-enhanced deep network, and uses a feature adaptive fusion module to integrate the classification results of the two stages.
[0067] In an alternative embodiment, based on the sequence of feature maps, calculate the local contrast and the maximum response direction, construct a feature pyramid structure, optimize and generate an adaptive segmentation threshold sequence through an edge-preserving constraint term, and implement hierarchical segmentation of the defect area based on the adaptive segmentation threshold sequence, including:
[0068] Construct a multi-scale directional filter bank, where the multi-scale directional filter bank includes Gaussian filter kernels of different scales and directions, and introduce a direction enhancement factor into the Gaussian filter kernels, and the direction enhancement factor is adaptively adjusted according to the main direction angle and the direction selectivity parameter;
[0069] Filter the input image using the multi-scale directional filter bank to obtain multiple groups of directional response feature maps, construct a directional response mapping function based on the multiple groups of directional response feature maps, and optimize the weight coefficients of different scales by minimizing the reconstruction error and the sparse regularization constraint;
[0070] Calculate the maximum response direction according to the directional response mapping function, construct a local contrast map by combining the local mean and standard deviation, and adaptively fuse the maximum response direction with the local contrast map to generate enhanced edge features;
[0071] Construct a feature pyramid based on the enhanced edge features, perform feature transformation through downsampling and upsampling operations at each level of the feature pyramid, and introduce an edge-preserving constraint term, and the edge-preserving constraint term is optimized by calculating the feature gradient difference;
[0072] Extract statistical features at each level of the feature pyramid, and generate an adaptive segmentation threshold sequence in a recursive manner according to the statistical features and the threshold of the previous level, and the adaptive segmentation threshold sequence is used to implement hierarchical segmentation of multi-scale defect areas.
[0073] Construct a multi-scale directional filter bank, which contains Gaussian filter kernels of different scales and directions. In practical applications, 5 scale parameters can be set, namely filter kernels with pixel sizes of 3×3, 5×5, 7×7, 9×9, and 11×11 respectively, and 8 directions can be set, namely 0°, 45°, 90°, 135°, 180°, 225°, 270°, and 315°. A direction enhancement factor is introduced into each Gaussian filter kernel, and this factor is adaptively adjusted according to the main direction angle and the direction selectivity parameter. Specifically, for the filter kernel in the direction of θ, the value at its center point (x, y) can be adjusted by multiplying the value calculated by the standard Gaussian function by the direction enhancement factor. The value of the direction enhancement factor is related to the angle between the current pixel point and the main direction of the filter. The smaller the angle, the larger the enhancement factor. For example, when the direction selectivity parameter is set to 1.5, the enhancement factor of the pixel points in the main direction can reach 2.3, while the enhancement factor of the pixel points perpendicular to the main direction is about 0.7.
[0074] Use the constructed multi-scale directional filter bank to filter the input image to obtain multiple groups of directional response feature maps. For an input image with a resolution of 1024×768, after being processed by filters of 5 scales and 8 directions, 40 groups of directional response feature maps can be obtained. Based on these feature maps, construct a directional response mapping function, and optimize the weight coefficients of different scales by minimizing the reconstruction error and sparse regularization constraints. In practical applications, the sparse regularization parameter can be set to 0.01, and the weight coefficients are solved by an iterative optimization method. After 50 iterations, the weight coefficients of each scale are 0.15, 0.25, 0.35, 0.15, and 0.1 respectively, indicating that the filter of the medium scale (7×7) makes the greatest contribution to defect detection.
[0075] Calculate the maximum response direction according to the directional response mapping function. The specific method is to find the direction with the largest response value among the 8 directions for each pixel point as the maximum response direction of this point. At the same time, construct a local contrast map by combining the local mean and standard deviation. The local contrast is calculated using an 11×11 sliding window. For the center point of the window, its contrast value is equal to the gray value of this point minus the mean value of the pixels in the window, and then divided by the standard deviation of the pixels in the window plus a small constant (set to 0.001). Adaptively fuse the maximum response direction with the local contrast map to generate enhanced edge features. The fusion weight is dynamically adjusted according to the size of the local contrast. Larger weights are assigned to the direction features in areas with high contrast, and vice versa in areas with low contrast. For example, when the local contrast value is greater than 0.5, the weight of the direction feature can be set to 0.7; when the contrast value is less than 0.2, the weight of the direction feature is reduced to 0.3.
[0076] Construct a feature pyramid based on the enhanced edge features, which contains 4 levels. The first level is the original resolution (1024×768), and the subsequent levels are obtained by 2-fold downsampling, which are 512×384, 256×192, and 128×96 respectively. Feature transformation is performed on each level of the feature pyramid through downsampling and upsampling operations. Max-pooling operation is used for downsampling, and bilinear interpolation method is used for upsampling. At the same time, an edge-preserving constraint term is introduced, and this constraint term is optimized by calculating the feature gradient difference. In specific implementation, the gradient direction difference between the upsampled feature map and the original feature map can be calculated. When the difference exceeds the threshold (such as 30°), the gradient information of the original feature map is retained; otherwise, the gradient information of the two is fused by weighted average, and the weight ratio is set to 7:3.
[0077] Extract statistical features on each level of the feature pyramid, including grayscale histogram, local entropy value, and gradient magnitude distribution. According to these statistical features and the threshold of the previous level, an adaptive segmentation threshold sequence is generated recursively. For the top layer of the pyramid (128×96), the initial threshold is set to 1.2 times the average grayscale value of the image; for other levels, the threshold is determined by the weighted combination of the threshold of the previous level and the statistical features of the current level. The weight coefficient is adaptively adjusted according to the level, with the weight of the top layer being 0.4, increasing by 0.1 layer by layer, and reaching 0.7 at the bottom layer. For example, if the threshold of the top layer is 120 and the peak value of the local entropy value distribution of the second layer is 135, then the threshold of the second layer can be calculated as 120×0.4 + 135×0.6 = 129.
[0078] Implement hierarchical segmentation of multi-scale defect regions based on the generated adaptive segmentation threshold sequence. For each layer in the pyramid, compare the enhanced edge feature map with the corresponding adaptive threshold, and the region greater than the threshold is marked as a defect candidate area. Then optimize the candidate area through morphological operations (opening operation and closing operation), and the size of the structuring element is set to 3×3 pixels. Starting from the top layer of the pyramid, map the detected defect regions to the next level and fuse them with the detection results of the next level. The fusion strategy uses the "or" operation, that is, as long as a region is detected as a defect at any level, the region is retained in the final result. Through experimental verification, the average detection accuracy of this method reaches 94.7% in the defect detection of materials such as metal surfaces, textiles, and glass, which is 8.3 percentage points higher than the traditional single-scale method.
[0079] Figure 2 This is the performance evaluation table graph of the feature pyramid segmentation technology in the embodiment of the present invention:
[0080] This figure shows the detailed evaluation results of the performance metrics of the feature pyramid at different pyramid levels. From the perspective of the adaptive threshold, as the level deepens (from Level 1 to Level 5), the value shows an obvious upward trend, gradually increasing from 0.152 to 0.312, with an average value of 0.232. The recall rate shows a slightly decreasing trend as the level deepens, dropping from 0.952 at Level 1 to 0.843 at Level 5, but remains at a relatively high level overall, with an average value of 0.906. The precision shows a characteristic of increasing as the level deepens, rising from 0.874 at Level 1 to 0.937 at Level 5, with an average value of 0.920. The F1 score reaches the highest at the middle levels (Level 2 and Level 3), being 0.924 and 0.925 respectively, and is slightly lower at both ends, with an average value maintained at an excellent level of 0.912. The edge preservation coefficient gradually decreases from 0.981 at Level 1 to 0.872 at Level 5, reflecting that as the resolution decreases, the ability to preserve edge details weakens, but the overall average value still reaches a high level of 0.929. These data comprehensively reflect the performance characteristics of the feature pyramid at different levels, demonstrating the superiority of this method in feature extraction and expression.
[0081] The technical solution of this application stems from the improvement of the existing defect detection technology. The existing technology mainly relies on traditional image processing techniques such as edge detection methods like the Canny operator and Sobel operator, or uses a single deep learning model for feature extraction and classification. These methods have problems such as being sensitive to noise, having high computational complexity and being difficult to meet the real-time detection requirements, and losing detailed features due to not considering the multi-scale feature hierarchical relationship during the feature extraction process. To solve these problems, this application proposes a number of improvement measures: introducing an adaptive direction enhancement factor into the Gaussian filter kernel to enhance the directional expression of edge features; optimizing the weight allocation of different scale features through minimizing the reconstruction error and sparse regularization constraints; designing an edge preservation constraint term to maintain edge information during the feature pyramid transformation process; and generating an adaptive segmentation threshold sequence in a recursive manner to improve the segmentation accuracy. Through experimental verification, the solution of this application improves the detection accuracy by 15%-20%, reduces the false alarm rate by more than 30%, and increases the detection speed by 40%. At the same time, it significantly enhances the adaptability and detection stability to different types of defects, effectively solves the problems existing in the prior art, and realizes the high efficiency, accuracy and stability of packaging material defect detection.
[0082] In an optional implementation manner, a feature pyramid is constructed based on the enhanced edge features. Feature transformation is performed through downsampling and upsampling operations at each level of the feature pyramid, and an edge preservation constraint term is introduced. The edge preservation constraint term is optimized by calculating the feature gradient difference, including:
[0083] Perform consecutive downsampling operations on the enhanced edge features to construct a feature pyramid, where the feature pyramid contains multiple levels, and each level maps the features through a learnable feature transformation matrix;
[0084] Between adjacent levels of the feature pyramid, use the upsampling operation to map the features of the resolution level lower than the preset resolution threshold to the feature scale space of the resolution level higher than the preset resolution threshold, and adaptively fuse the resolution level lower than the preset resolution threshold and the resolution level higher than the preset resolution threshold to generate fused features;
[0085] Introduce an edge-preserving constraint term for the fused features. The edge-preserving constraint term is constrained by calculating the feature gradient difference and feature structure difference between adjacent levels. The feature gradient difference is measured by the first-order norm, and the feature structure difference is measured by the second-order norm; construct a multi-level consistency loss function based on the edge-preserving constraint term, and optimize the learnable feature transformation matrix by minimizing the multi-level consistency loss function.
[0086] The enhanced edge features refer to the image features after edge enhancement processing, which contain rich edge information. The process of constructing the feature pyramid is as follows: perform consecutive downsampling operations on the enhanced edge features to generate multiple feature levels with different resolutions. Specifically, assume the initial feature resolution is 512×512, and through consecutive downsampling operations, feature levels of 256×256, 128×128, 64×64, and 32×32 are generated in sequence, forming a five-layer feature pyramid structure.
[0087] The initial value of the feature transformation matrix can be set as the identity matrix, with the size of the number of feature channels × the number of feature channels. For example, if the number of feature channels is 64, the feature transformation matrix is initialized as a 64×64 identity matrix. The role of the feature transformation matrix is to adaptively adjust the features of each level and enhance the feature expression ability.
[0088] Set the preset resolution threshold to 64×64. For the resolution level lower than this threshold (such as 32×32), use the upsampling operation to map its features to the feature scale space of the resolution level higher than the threshold (such as 64×64). The upsampling operation uses the bilinear interpolation method to maintain the spatial continuity of the features.
[0089] Let the low-resolution feature be F_low, the upsampled feature be F_up, and the high-resolution feature be F_high. Then the calculation method of the fused feature F_fusion is as follows: First, calculate the adaptive weight α. The value range of α is between 0 and 1, and the initial value is set to 0.5. Then, F_fusion = α × F_up + (1 - α) × F_high. The adaptive weight α is obtained through network learning and can automatically adjust the fusion ratio according to different image contents. For example, for regions with rich edges, the value of α may be closer to 0.4, making the high-resolution feature account for a larger proportion; while for regions with smooth textures, the value of α may be closer to 0.6, making the upsampled feature account for a larger proportion.
[0090] To maintain the consistency of edge information among different levels of the feature pyramid, an edge-preserving constraint term is introduced to the fused feature. This constraint term consists of two parts: feature gradient difference and feature structure difference.
[0091] The feature gradient difference is measured using the first-order norm, which calculates the absolute difference of feature gradients between adjacent levels. Specifically, when implementing, first calculate the horizontal gradient Gx and vertical gradient Gy for the features of each level, and then calculate the absolute difference of gradients between adjacent levels. For example, for the two levels of 64×64 and 32×32 (upsampled to 64×64), calculate the sum of |Gx_64 - Gx_32_up| and |Gy_64 - Gy_32_up| as the gradient difference. In practical applications, if the number of feature channels is 64, the gradient difference is calculated separately for each channel and then the average value is obtained.
[0092] The feature structure difference is measured using the second-order norm, which calculates the Euclidean distance of feature structures between adjacent levels. Specifically, when implementing, calculate the Laplacian operator response of adjacent-level features, and then calculate the Euclidean distance of the response values. For example, for the features of the two levels of 64×64 and 32×32 (upsampled to 64×64), calculate the Laplacian responses L_64 and L_32_up respectively, and then calculate ||L_64 - L_32_up||2 as the structure difference.
[0093] Based on the above edge-preserving constraint term, a multi-level consistency loss function is constructed. Assuming that the feature pyramid has n levels, the multi-level consistency loss function is the sum of the edge-preserving constraint terms for all adjacent-level pairs. Specifically, for a five-level pyramid structure, calculate the edge-preserving constraint terms for the four pairs of adjacent levels of 512×512 and 256×256, 256×256 and 128×128, 128×128 and 64×64, 64×64 and 32×32, and assign different weights. Usually, the weights for higher-resolution level pairs are set larger. For example, the weights can be set to 0.4, 0.3, 0.2, and 0.1 respectively.
[0094] The learnable feature transformation matrix is optimized by minimizing the multi-level consistency loss function. The optimization uses the stochastic gradient descent algorithm, with the initial learning rate set to 0.001 and decaying to 0.9 times the original value every 10 training epochs. During training, the batch size is set to 16 and the number of training epochs is 100.
[0095] This method can effectively maintain the consistency of edge information among different levels of the feature pyramid. For example, for building images with complex edge structures, traditional methods may lose detailed edges at low-resolution levels. However, with the edge-preserving constraint term in this method, the features after upsampling at the 32×32 level are highly consistent with the features at the 64×64 level in terms of edge positions, with the gradient difference reduced by about 40% and the structural difference reduced by about 35%, thereby improving the expressive power of the feature pyramid and the performance of subsequent tasks.
[0096] In an alternative embodiment, a feature evaluation matrix is established by calculating the within-class cohesion and between-class distance of each dimension feature, and the weight coefficients of each dimension feature are dynamically allocated based on the feature evaluation matrix. The discriminative enhanced feature description vector is generated by fusion, including:
[0097] An adaptive bandwidth parameter is constructed based on the local data distribution characteristics, and the adaptive bandwidth parameter is dynamically adjusted by calculating the Euclidean distance from the sample to the class center and the local density estimate value calculated using the multi-scale Gaussian kernel function;
[0098] A dynamic kernel function is constructed using the adaptive bandwidth parameter. The dynamic kernel function combines a distance modulation function to evaluate the structural similarity between sample pairs, and determines the optimal neighborhood set based on the structural similarity;
[0099] An adaptive weight is applied to the sample pairs in the optimal neighborhood set based on the dynamic kernel function, the weighted within-class cohesion of each dimension feature is calculated, and a temporal consistency constraint is introduced to ensure the temporal continuity of the weighted within-class cohesion;
[0100] A first feature evaluation matrix is constructed, and the first feature evaluation matrix includes the weighted within-class cohesion and the between-class distance metric obtained by calculating the Mahalanobis distance and distribution divergence between different class centers;
[0101] A second feature evaluation matrix is constructed based on the weighted within-class cohesion and the between-class distance metric, and the softmax function is used to convert the combined discriminative scores of the first feature evaluation matrix and the second feature evaluation matrix into weight coefficients; the weight coefficients are used to adaptively fuse each dimension feature to generate a discriminative enhanced feature description vector.
[0102] Construct adaptive bandwidth parameters based on local data distribution characteristics. For a given data set, it contains samples of multiple categories, each of which has multi-dimensional features. Calculate the Euclidean distance of each sample to the center of its category. For example, for sample x_i in category A, calculate its Euclidean distance d_i to the center point of category A. At the same time, a multi-scale Gaussian kernel function is used to calculate the local density estimate of the sample. Specifically, different scale parameters σ_1=0.1, σ_2=0.5, σ_3=1.0 are selected, and the local density ρ_i 1, ρ_i2, ρ_i 3 of sample x_i at these scales are calculated respectively, and the weighted average is taken as the final local density estimate ρ_i.
[0103] The adaptive bandwidth parameter h_i is dynamically adjusted according to the formula h_i=h_0·(1+α·d_i) / (1+β·ρ_i), where h_0 is the basic bandwidth value (such as 0.5), and α and β are adjustment parameters (such as α=0.3, β=0.2). In this way, samples far from the class center and with low local density will obtain a larger bandwidth value, and vice versa.
[0104] For any two samples x_i and x_j, the dynamic kernel function value K(x_i,x_j) is calculated by combining the distance modulation function. The distance modulation function φ(x_i,x_j) takes into account the structural similarity between sample pairs. When two samples belong to the same category and have similar local distribution characteristics, the value of φ(x_i,x_j) is larger; otherwise, it is smaller. For example, if samples x_i and x_j belong to the same category A and their local density difference is less than the threshold 0.2, then φ(x_i,x_j) = 0.9; if they belong to different categories, then φ(x_i,x_j) = 0.1.
[0105] The dynamic kernel function is finally expressed as K(x_i,x_j)=exp(-||x_i-x_j||2 / (h_i·h_j))·φ(x_i,x_j). Based on this dynamic kernel function, the optimal neighborhood set N(x_i) is determined for each sample, which contains k samples (such as k=5) with high structural similarity to sample x_i.
[0106] Apply adaptive weights to the sample pairs in the optimal neighborhood set based on a dynamic kernel function. For the neighborhood sample \(x_j\in N(x_i)\) of sample \(x_i\), its weight \(w_{ij} = K(x_i,x_j) / \sum_k K(x_i,x_k)\). Using these weights, calculate the weighted within-class cohesion of each dimensional feature. Taking the \(m\)-th dimensional feature as an example, its weighted within-class cohesion \(C_m\) is obtained by calculating the reciprocal of the weighted variance of the same-class samples on this dimension. For class \(c\), its weighted within-class cohesion \(C_m^c\) of the \(m\)-th dimensional feature is \(1 / (\sum_i\sum_jw_{ij}\cdot(x_{im} - x_{jm})^2)\), where \(x_{im}\) represents the value of the \(m\)-th dimensional feature of sample \(x_i\). To ensure the temporal continuity of the weighted within-class cohesion, introduce a temporal consistency constraint, that is, \(C_m(t)=\gamma\cdot C_m(t - 1)+(1-\gamma)\cdot C_m(t)\), where \(\gamma\) is a smoothing factor (such as \(\gamma = 0.3\)), and \(C_m(t)\) and \(C_m(t - 1)\) represent the within-class cohesion at the current moment and the previous moment respectively.
[0107] Construct the first feature evaluation matrix \(M1\). This matrix contains the weighted within-class cohesion and the between-class distance metric. The between-class distance metric \(D_m\) is obtained by calculating the Mahalanobis distance and the distribution divergence between different class centers. For example, for class \(A\) and class \(B\), their between-class distance \(D_m^{AB}\) on the \(m\)-th dimensional feature considers the differences between the two class centers on this dimension and their respective covariances. If the mean of the \(m\)-th dimensional feature of class \(A\) is \(5.2\), the mean of class \(B\) is \(8.7\), and the covariances of the two classes on this dimension are \(0.8\) and \(1.2\) respectively, then \(D_m^{AB}\) can be calculated as \(3.5 / \sqrt{(0.8 + 1.2) / 2}\approx3.5 / \sqrt{1}=3.5\). The elements of the first feature evaluation matrix \(M1\) are combinations of the within-class cohesion and the between-class distance of each dimensional feature, such as \(M1_m = C_m\cdot D_m\).
[0108] Based on the weighted within-class cohesion and the between-class distance metric, construct the second feature evaluation matrix \(M2\). \(M2\) considers the complementarity and redundancy between each dimensional feature, and its element \(M2_mn\) represents the degree of complementarity between the \(m\)-th and \(n\)-th dimensional features. When the patterns of the within-class cohesion and the between-class distance of two dimensions are different, they have a high degree of complementarity; otherwise, there may be redundancy. For example, if the within-class cohesion of the 1st and 2nd dimensional features are \(0.8\) and \(0.3\) respectively, and the between-class distances are \(2.5\) and \(4.0\) respectively, then their degree of complementarity \(M2_{12}\) may be high (such as \(0.7\)); if these indicators of the 1st and 3rd dimensions are very close, then \(M2_{13}\) may be low (such as \(0.2\)).
[0109] The softmax function is used to convert the combined discriminative scores of the first feature evaluation matrix and the second feature evaluation matrix into weight coefficients. The combined discriminative score \(S_m=\lambda\cdot M1_m+(1 - \lambda)\cdot\sum_n M2_mn\), where \(\lambda\) is a balancing factor (e.g., \(\lambda = 0.6\)). The weight coefficient \(w_m=\frac{\exp(S_m)}{\sum_k\exp(S_k)}\). For example, if there are three-dimensional features with combined discriminative scores \(S_1 = 1.2\), \(S_2 = 0.8\), and \(S_3 = 0.5\) respectively, the weight coefficients calculated by the softmax function are \(w_1\approx0.5\), \(w_2\approx0.3\), and \(w_3\approx0.2\) respectively. These weight coefficients are used to adaptively fuse the features of each dimension to generate a discriminative enhanced feature description vector \(F = [w_1\cdot f_1, w_2\cdot f_2,\cdots, w_M\cdot f_M]\), where \(f_m\) represents the \(m\)-th component of the original feature.
[0110] Figure 3 This is a table graph comparing the comprehensive performance of the feature evaluation matrix and weight coefficient adaptive fusion technology in the embodiments of the present invention:
[0111] This figure shows the performance comparison results of different methods on the packaging material defect dataset and the industrial surface defect dataset. The proposed technical solution achieves the optimal performance on the packaging material defect dataset, with an accuracy of 96.5%, an intra-class cohesion of 0.916, and an inter-class distance of 0.835; it also performs excellently on the industrial surface defect dataset, with an accuracy of 94.2%, an intra-class cohesion of 0.892, and an inter-class distance of 0.815. In contrast, the accuracies of the fixed bandwidth kernel function method are 91.2% and 89.7% respectively, with an average improvement of 4.9%; the accuracies of the traditional distance measure method are 87.6% and 86.2% respectively, with an average improvement of 9.5%; the accuracies of the method without temporal constraints are 93.4% and 90.8% respectively, with an average improvement of 3.3%; the accuracies of the equal weight method are 89.8% and 87.3% respectively, with an average improvement of 6.8%. In terms of the intra-class cohesion index, the proposed scheme has an average improvement range of 6.1% - 14.0% compared with other methods, and the average improvement amplitude of the inter-class distance index is between 6.9% - 14.1%. These data fully demonstrate that the proposed technical solution has significant advantages in three key indicators: accuracy, intra-class cohesion, and inter-class distance, especially showing the most obvious performance improvement compared with traditional methods.
[0112] The improvement of the technical solution of this application stems from the in-depth research and optimization of traditional feature fusion methods. The commonly used feature fusion methods in the prior art mainly adopt fixed weights or simple weighted average strategies, such as linear discriminant analysis (LDA) and principal component analysis (PCA) based on Fisher discriminant criteria. These methods have problems such as fixed weight allocation, inability to adaptively adjust, and ignoring the local structural relationship between features when dealing with high-dimensional feature fusion, resulting in insufficient discriminability of the fused features. To solve these problems, this application proposes an adaptive feature fusion scheme based on a dynamic kernel function and a dual feature evaluation matrix: constructing a dynamic kernel function through an adaptive bandwidth parameter to achieve an accurate evaluation of the sample structural similarity; introducing a temporal consistency constraint to ensure the stability of feature evaluation; combining the Mahalanobis distance and distribution divergence to construct a dual feature evaluation matrix, and realizing the adaptive allocation of weights through the softmax function. Experimental results show that compared with traditional methods, the discriminability of features in this scheme is improved by 25%, the classification accuracy is increased by 18%, the time overhead of feature fusion is reduced by 35%, and at the same time, the robustness and generalization ability of feature expression are significantly enhanced, effectively solving the problem of poor feature fusion effect in the prior art and providing a new solution for the effective fusion of high-dimensional features.
[0113] In an alternative embodiment, constructing a second feature evaluation matrix based on the weighted within-class cohesion and the inter-class distance metric, and using the softmax function to convert the combined discriminative scores of the first feature evaluation matrix and the second feature evaluation matrix into weight coefficients includes:
[0114] Constructing a second feature evaluation matrix based on the discriminative scores calculated from the weighted within-class cohesion and the inter-class distance metric, the second feature evaluation matrix is optimized by the node features in the multi-dimensional feature association graph structure, the multi-dimensional feature association graph structure includes a node set, an edge set and an edge weight set, and the node features are generated through feature correlation calculation and graph attention mechanism;
[0115] Combining the first feature evaluation matrix and the second feature evaluation matrix to obtain discriminative scores, the discriminative scores are obtained by applying adaptive weight coefficients to the first feature evaluation matrix and the second feature evaluation matrix and performing a weighted summation operation, and the adaptive weight coefficients are determined based on the reliability and discriminative ability of each feature evaluation matrix;
[0116] Using the softmax function to convert the discriminative scores into weight coefficients, and the weight coefficients are used for adaptive fusion of each dimension feature.
[0117] Construct a second feature evaluation matrix based on weighted intra-class cohesion and inter-class distance metrics. The system constructs the second feature evaluation matrix by calculating the discriminative ability of each feature dimension. Specifically, for each feature dimension in the dataset, calculate the intra-class cohesion and inter-class distance of samples in each category under this dimension. Assume the dataset contains 10 categories, each category has 100 samples, and each sample has 50 feature dimensions. For the j-th feature dimension, calculate the intra-class cohesion value of samples in the i-th category under this dimension, that is, calculate the variance of all samples in this category on this feature dimension. The smaller the intra-class cohesion, the more concentrated the distribution of samples in this category on this feature dimension. For example, for the 3rd feature dimension, the intra-class cohesion value of samples in the 1st category is 0.023, and the intra-class cohesion value of samples in the 2nd category is 0.045.
[0118] Calculate the weighted intra-class cohesion, that is, perform a weighted average on the intra-class cohesion values of each category, and the weights can be determined based on the number of samples in each category. For example, if the 1st category has 150 samples and the 2nd category has 100 samples, then the weight of the 1st category is 0.6, the weight of the 2nd category is 0.4, and the weighted intra-class cohesion of the 3rd feature dimension is 0.023×0.6 + 0.045×0.4 = 0.0318.
[0119] Calculate the inter-class distance metric, that is, the central distance between samples of different categories on a specific feature dimension. For example, for the 3rd feature dimension, the distance between the centers of samples in the 1st and 2nd categories is 0.75, and the distance between the centers of samples in the 1st and 3rd categories is 0.82. The larger the inter-class distance, the higher the discrimination degree between different categories on this feature dimension.
[0120] Based on the weighted intra-class cohesion and inter-class distance metrics, calculate the discriminative score of each feature dimension. The discriminative score can be obtained by the ratio of the inter-class distance to the weighted intra-class cohesion. The larger the ratio, the stronger the discriminative ability of this feature dimension. For example, the discriminative score of the 3rd feature dimension is (0.75 + 0.82) / 0.0318 = 49.37.
[0121] Form the discriminative scores of all feature dimensions into a second feature evaluation matrix. For 50 feature dimensions, construct a 50×1 matrix, where each element represents the discriminative score of the corresponding feature dimension.
[0122] To further optimize the second feature evaluation matrix, introduce a multi-dimensional feature association graph structure. This graph structure includes a node set, an edge set, and an edge weight set. The node set corresponds to each feature dimension, the edge set represents the connection relationship between feature dimensions, and the edge weight set represents the correlation strength between feature dimensions.
[0123] Determine the edge weights through feature correlation calculation. For example, calculate the Pearson correlation coefficient between the i-th feature dimension and the j-th feature dimension. If the absolute value of the correlation coefficient is greater than a preset threshold (such as 0.5), establish a connection between the two nodes, and set the edge weight to the absolute value of the correlation coefficient. For example, if the correlation coefficient between the 1st feature dimension and the 3rd feature dimension is 0.72, establish a connection between these two nodes, and the edge weight is 0.72.
[0124] Apply the graph attention mechanism to generate node features. For each node, aggregate the information of its neighbor nodes, and determine the importance of the neighbor nodes according to the edge weights. For example, the neighbor nodes of the 1st feature dimension include the 3rd, 7th, and 12th feature dimensions, and the edge weights are 0.72, 0.65, and 0.58 respectively. Then the contribution ratios of these neighbor nodes to the 1st feature dimension are 0.37, 0.33, and 0.30 respectively.
[0125] Optimize the second feature evaluation matrix through the node features. Fuse the original discriminative scores with the node features generated by the graph attention mechanism to obtain the optimized discriminative scores. For example, the original discriminative score of the 3rd feature dimension is 49.37, and the score optimized by the graph attention mechanism is 52.15.
[0126] Combine the first feature evaluation matrix and the second feature evaluation matrix. The first feature evaluation matrix is constructed based on feature importance evaluation methods (such as information gain, Gini coefficient, etc.), and is also a 50×1 matrix. For example, the score of the 3rd feature dimension in the first feature evaluation matrix is 0.85.
[0127] Determine the adaptive weight coefficients for combining the two feature evaluation matrices. The adaptive weight coefficients are determined based on the reliability and discriminative ability of each feature evaluation matrix. For example, evaluate the performance of the two matrices through cross-validation. If the average accuracy of the first feature evaluation matrix is 0.82 and the average accuracy of the second feature evaluation matrix is 0.78, then the weight coefficient of the first feature evaluation matrix is 0.82 / (0.82 + 0.78) = 0.51, and the weight coefficient of the second feature evaluation matrix is 0.78 / (0.82 + 0.78) = 0.49.
[0128] Obtain the combined discriminative scores through weighted summation operations. For example, the combined discriminative score of the 3rd feature dimension is 0.85×0.51 + 52.15×0.49 = 26.0.
[0129] The softmax function is used to convert the combined discriminative scores into weight coefficients. The softmax function converts the discriminative scores of each feature dimension into values between 0 and 1, and the sum of all values is 1. For example, for the combined discriminative scores of 50 feature dimensions, after applying the softmax function, the weight coefficient of the 3rd feature dimension is 0.072.
[0130] These weight coefficients are used to adaptively fuse the features of each dimension, realize the effective selection and combination of features, and improve the performance and robustness of the model. For example, in the feature fusion stage, the contribution of the 3rd feature dimension is the original feature value multiplied by the corresponding weight coefficient 0.072.
[0131] Figure 4 This is the performance comparison data table graph of the feature evaluation matrix for the embodiments of the present invention:
[0132] This figure shows the performance comparison results of each method under different evaluation metrics. In terms of feature discriminative scores, this technical solution reaches 0.924, 0.882, and 0.905 respectively on three packaging material symptom datasets, showing a significant improvement compared with traditional feature evaluation and single-distance evaluation methods, with the improvement ratios being 29.8%, 28.7%, and 29.8% respectively. For the intra-class cohesion index, this solution achieves performances of 0.836, 0.812, and 0.825 on the three datasets, with a 34.2%, 37.2%, and 35.9% improvement compared with traditional methods. In the inter-class distance index evaluation, this solution reaches levels of 0.785, 0.768, and 0.772 respectively, with a 43.2%, 36.9%, and 39.3% improvement compared with the baseline method. In terms of the stability of weight coefficients, this solution achieves stability scores of 0.896, 0.883, and 0.891 on the three datasets respectively, with an improvement amplitude of 41.1%, 44.3%, and 43.0%. It is particularly worth noting that in terms of computational efficiency, this solution is significantly better than other methods, and the processing time is reduced to 23.8ms, 26.4ms, and 25.3ms respectively, with an efficiency improvement of about 59% compared with traditional methods, demonstrating that this technical solution has significant computational efficiency advantages while ensuring high performance.
[0133] In an alternative embodiment, the first stage of the two-stage cascade classifier is implemented by a lightweight decision tree for fast pre-classification, and the second stage introduces an attention-enhanced deep network to extract local detail features and global semantic features. The classification results of the two stages are integrated through a feature adaptive fusion module to achieve accurate classification, including:
[0134] Construct a two - stage cascade classifier. The first stage of the two - stage cascade classifier uses a lightweight decision tree for fast pre - classification. The lightweight decision tree obtains a sample difficulty evaluation value by calculating the distance from the sample to the decision boundary and the local density, and dynamically weights the training samples according to the sample difficulty evaluation value;
[0135] Construct a curriculum learning strategy based on the sample difficulty evaluation value. The curriculum learning strategy determines training samples according to the dynamic distribution of sample difficulty, generates adversarial samples for the training samples, and inputs the adversarial samples into the second stage of the two - stage cascade classifier;
[0136] The second stage of the two - stage cascade classifier includes an attention - enhanced deep network. Based on the attention - enhanced deep network, local detail features and global semantic features are respectively extracted, adaptive noise injection is performed on the local detail features and the global semantic features, and the feature consistency loss from different perspectives is calculated;
[0137] Perform adaptive fusion on the pre - classification result of the lightweight decision tree and the feature extraction result of the attention - enhanced deep network. The adaptive fusion is obtained by weighted calculation of the confidence levels of the two - stage classification results;
[0138] Dynamically adjust the decision threshold based on the change rate of the confidence level, and use the decision threshold to optimize the fused classification result. Evaluate and correct the reliability of the classification result by calculating the prediction probability in the confusion matrix to generate the final classification result.
[0139] Construct a two - stage cascade classifier, where the first stage uses a lightweight decision tree for fast pre - classification. The lightweight decision tree evaluates the sample difficulty by calculating the distance from the sample to the decision boundary and the local density. Specifically, for an input sample x, calculate its Euclidean distance d(x) to the nearest decision boundary, and at the same time count the number of similar samples in the neighborhood of the sample x to obtain the local density ρ(x). The sample difficulty evaluation value s(x) is obtained by combining these two factors, that is, s(x)=α·d(x)+(1 - α)·(1 / ρ(x)), where α is a balance factor with a value of 0.6. For samples with a higher difficulty evaluation value (such as s(x)>0.75), a higher weight w(x)=1 + β·s(x) is assigned during training, where β is a weight adjustment coefficient set to 0.5. For example, for a sample with a relatively short distance to the decision boundary (d(x)=0.3) and a low local density (ρ(x)=5), its difficulty evaluation value s(x)=0.6×0.3 + 0.4×0.2 = 0.26, and the corresponding weight w(x)=1.13.
[0140] Construct a curriculum learning strategy based on the sample difficulty evaluation value. In the initial stage of training (such as the first 10 epochs), select samples with a lower difficulty evaluation value (s(x) < 0.4) for training; as the training progresses (such as 10 - 20 epochs), gradually introduce medium-difficulty samples (0.4 ≤ s(x) < 0.7); in the later stage of training (such as after 20 epochs), introduce all samples. Generate adversarial samples for the selected training samples. The specific method is to add a perturbation δ to the original sample x to generate x' = x + δ, where the magnitude of δ is limited by ε = 0.1 to ensure that the visual difference between the adversarial sample and the original sample is not obvious. For example, for the original image pixel values in the range [0, 255], the maximum added perturbation is ±25.5. These adversarial samples will be input into the second stage of the two-stage cascade classifier.
[0141] The second stage of the two-stage cascade classifier contains a depth network with enhanced attention. This network first extracts basic features through a convolutional layer and then divides into two branches: a local detail feature extraction branch and a global semantic feature extraction branch. The local detail feature extraction branch adopts a spatial attention mechanism to strengthen the features of key regions by calculating the importance weights A_spatial of each position in the feature map. For example, for a 512×7×7 feature map, a 7×7 attention weight map with a weight value range of [0, 1] is generated through a 1×1 convolution and a Sigmoid activation function. The global semantic feature extraction branch adopts a channel attention mechanism to calculate the importance weights A_channel of each channel through global average pooling and a fully connected layer. For example, for a 512-channel feature, a 512-dimensional attention weight vector with a weight value range of [0, 1] is generated.
[0142] To enhance the robustness of the model, adaptive noise injection is performed on the extracted features. Specifically, the noise intensity is adaptively adjusted according to the importance of the features. For regions or channels with high importance, less noise is applied; for regions or channels with low importance, more noise is applied. The noise intensity σ(i) = σ_max·(1 - A(i)), where σ_max is the maximum noise intensity, set to 0.2, and A(i) is the attention weight of the corresponding position or channel. For example, for a feature region with an attention weight of 0.8, the noise intensity is 0.04; while for a region with an attention weight of 0.3, the noise intensity is 0.14.
[0143] Enhance the generalization ability of the model by calculating the consistency loss of features from different perspectives. Specifically, compare the feature representations of the same sample under different data augmentation conditions (such as rotation, scaling, cropping, etc.), calculate the cosine similarity between them, and use it as part of the consistency loss. For example, for the original perspective feature f_1 and the perspective feature f_2 after rotating 45 degrees, calculate their cosine similarity cos(f_1, f_2), and the goal is to make this similarity close to 1.
[0144] Adaptive fusion of the pre-classification results of the lightweight decision tree and the feature extraction results of the attention-enhanced deep network. First, calculate the confidence of the two classifiers. The confidence C_dt of the decision tree is based on the distance from the sample to the decision boundary, and the confidence C_nn of the deep network is based on the maximum probability value of the softmax output. The confidence of the final classification result is calculated by weighting: C_final = γ·C_dt + (1 - γ)·C_nn, where γ is the adaptive weight and is dynamically adjusted according to the historical accuracy of the two classifiers. For example, when the accuracy of the decision tree on the validation set is 85% and the accuracy of the deep network is 92%, γ = 0.85 / (0.85 + 0.92) ≈ 0.48.
[0145] Dynamically adjust the decision threshold based on the change rate of confidence. Specifically, track the average confidence change of consecutive batches of samples. When the confidence is stable (the change rate is less than 0.05), increase the decision threshold (such as from 0.7 to 0.75), and when the confidence fluctuates greatly, decrease the decision threshold. Evaluate and correct the classification results by calculating the prediction probabilities in the confusion matrix. For example, for a sample classified as class A, if its historical probability of being misclassified as class B is high (such as greater than 0.3), then reduce its confidence of being classified as A. If the corrected confidence is lower than the decision threshold, mark this sample as "pending" and wait for more evidence or manual review.
[0146] Figure 5 Table chart for comparing the performance indicators of the two-stage cascade classifier in the embodiments of the present invention with the prior art:
[0147] This figure shows the comparison results of different classification methods in multiple performance metrics. This method has significant advantages compared with three control methods: conventional CNN, single decision tree, and ResNet-50. In terms of accuracy, this method reaches 96.7%, far higher than 85.3% of conventional CNN, 79.8% of single decision tree, and 92.1% of ResNet-50. The precision reaches 95.2%, with improvements of 11.6%, 18.8%, and 4.7% compared with other methods respectively. The recall rate reaches 94.8%, with obvious improvements compared with other methods. The F1 score reaches 0.950, reflecting excellent comprehensive performance. The false positive rate is only 3.1%, much lower than 12.5%, 15.7%, and 8.3% of other methods. In terms of computational efficiency, the inference time of this method is 29.5ms. Although it is slightly slower than 12.6ms of single decision tree, it is far better than 76.8ms of ResNet-50, reflecting good real-time performance. The storage overhead is 52.3MB, nearly half less than 97.8MB of ResNet-50, and also significantly better than 45.7MB of conventional CNN. Considering all indicators, this method achieves a low computational overhead and storage requirements while ensuring high accuracy, reaching the best balance between performance and efficiency.
[0148] The technical solution of this application stems from in-depth analysis and improvement research on existing classification technologies. Existing classification methods mainly include single classifiers based on traditional machine learning (such as support vector machines, random forests, etc.) and end-to-end models based on deep learning (such as convolutional neural networks, transformers, etc.). These methods generally have problems such as insufficient ability to handle difficult-to-classify samples, lack of adaptability in the feature extraction process, and imperfect classification result reliability evaluation mechanisms. To address these problems, this application proposes an innovative improvement plan: designing a two-stage cascaded classification architecture, organically combining lightweight decision trees and attention-enhanced deep networks; introducing a curriculum learning strategy and adversarial training mechanism based on sample difficulty; adopting adaptive noise injection and multi-perspective consistency constraints; constructing a confidence-based dynamic decision mechanism and reliability evaluation system. Experimental results show that this plan improves the classification accuracy by more than 20%, increases the computational efficiency by 40%, improves the recognition accuracy of difficult-to-classify samples by 25%, significantly enhances the model robustness, increases the anti-interference ability by 30%, and the reliability evaluation accuracy of classification results reaches more than 95%. This application effectively solves the problems in existing classification technologies such as the difficulty in balancing efficiency and accuracy and insufficient ability to handle difficult-to-classify samples, providing a new technical solution for efficient and accurate classification in complex scenarios.
[0149] In the second aspect of the embodiments of the present invention,
[0150] A packaging material defect detection system based on machine vision is provided, including:
[0151] The first unit is configured to collect multi-angle image sequences, construct a depth map based on camera calibration parameters, perform geometric correction and illumination compensation on the multi-angle image sequences according to the depth map, and generate preprocessed enhanced images;
[0152] The second unit is configured to construct a multi-scale directional filter bank based on the preprocessed enhanced images. The multi-scale directional filter bank introduces an adaptive directional enhancement factor to generate a sequence of feature maps, calculates local contrast and maximum response direction based on the sequence of feature maps, constructs a feature pyramid structure, optimizes through an edge-preserving constraint term and generates a sequence of adaptive segmentation thresholds, and realizes hierarchical segmentation of the defect regions based on the sequence of adaptive segmentation thresholds;
[0153] The third unit is configured to extract a multi-dimensional feature set from the defect regions obtained by the hierarchical segmentation. The multi-dimensional feature set includes local texture features, gray-scale statistical features, and shape description features, calculates the within-class cohesion and between-class distance of the multi-dimensional feature set to construct a feature evaluation matrix, assigns weight coefficients to the features in each dimension according to the feature evaluation matrix, and fuses to generate an enhanced discriminative feature description vector;
[0154] The fourth unit is configured to input the enhanced discriminative feature description vector into a two-stage cascade classifier. The first stage of the two-stage cascade classifier uses a lightweight decision tree for fast pre-classification, and the second stage of the two-stage cascade classifier extracts local detail features and global semantic features through an attention-enhanced deep network, and integrates the classification results of the two stages using a feature adaptive fusion module.
[0155] In the third aspect of the embodiments of the present invention,
[0156] There is provided an electronic device, including:
[0157] A processor;
[0158] A memory for storing instructions executable by the processor;
[0159] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0160] In the fourth aspect of the embodiments of the present invention,
[0161] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0162] The present invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting defects in packaging materials based on machine vision, characterized in that, Including: Collecting multi - angle image sequences and constructing a depth map through camera calibration parameters, geometrically correcting and light compensating the multi - angle image sequences according to the depth map to generate a pre - processed enhanced image; Constructing a multi - scale directional filter bank based on the pre - processed enhanced image, introducing an adaptive direction enhancement factor into the multi - scale directional filter bank to generate a sequence of feature maps, calculating local contrast and maximum response direction based on the sequence of feature maps, constructing a feature pyramid structure, optimizing through an edge - preserving constraint term and generating an adaptive segmentation threshold sequence, and realizing hierarchical segmentation of defect regions based on the adaptive segmentation threshold sequence; Extracting a multi - dimensional feature set for the defect regions obtained by the hierarchical segmentation, where the multi - dimensional feature set includes local texture features, gray - scale statistical features, and shape description features, calculating the within - class cohesion and between - class distance of the multi - dimensional feature set to construct a feature evaluation matrix, and assigning weight coefficients to each - dimensional feature according to the feature evaluation matrix to fuse and generate an enhanced discriminative feature description vector; Inputting the enhanced discriminative feature description vector into a two - stage cascade classifier, where the first stage of the two - stage cascade classifier uses a lightweight decision tree for fast pre - classification, and the second stage of the two - stage cascade classifier extracts local detail features and global semantic features through an attention - enhanced deep network, and uses a feature adaptive fusion module to integrate the classification results of the two stages.
2. The method according to claim 1, characterized in that Calculating local contrast and maximum response direction based on the sequence of feature maps, constructing a feature pyramid structure, optimizing through an edge - preserving constraint term and generating an adaptive segmentation threshold sequence, and realizing hierarchical segmentation of defect regions based on the adaptive segmentation threshold sequence includes: Constructing a multi - scale directional filter bank, where the multi - scale directional filter bank includes Gaussian filter kernels of different scales and directions, and introducing a direction enhancement factor into the Gaussian filter kernels, and the direction enhancement factor is adaptively adjusted according to the main direction angle and direction selectivity parameters; Filtering the input image using the multi - scale directional filter bank to obtain multiple groups of direction response feature maps, constructing a direction response mapping function based on the multiple groups of direction response feature maps, and optimizing the weight coefficients of different scales by minimizing the reconstruction error and sparse regularization constraints; Calculating the maximum response direction according to the direction response mapping function, constructing a local contrast map by combining the local mean and standard deviation, and adaptively fusing the maximum response direction with the local contrast map to generate enhanced edge features; Constructing a feature pyramid based on the enhanced edge features, performing feature transformation through down - sampling and up - sampling operations at each level of the feature pyramid, and introducing an edge - preserving constraint term, and the edge - preserving constraint term is optimized by calculating the feature gradient difference; Extracting statistical features at each level of the feature pyramid, and generating an adaptive segmentation threshold sequence in a recursive manner according to the statistical features and the threshold of the previous level, and the adaptive segmentation threshold sequence is used to realize hierarchical segmentation of multi - scale defect regions.
3. The method according to claim 2, characterized in that, Construct a feature pyramid based on the enhanced edge features. Perform feature transformation through downsampling and upsampling operations at each level of the feature pyramid, and introduce an edge-preserving constraint term. The edge-preserving constraint term is optimized by calculating the feature gradient difference, including: Perform successive downsampling operations on the enhanced edge features to construct a feature pyramid. The feature pyramid contains multiple levels, and each level maps the features through a learnable feature transformation matrix; Between adjacent levels of the feature pyramid, use the upsampling operation to map the features of the resolution level lower than the preset resolution threshold to the feature scale space of the resolution level higher than the preset resolution threshold, and adaptively fuse the resolution level lower than the preset resolution threshold and the resolution level higher than the preset resolution threshold to generate a fused feature; Introduce an edge-preserving constraint term for the fused feature. The edge-preserving constraint term is constrained by calculating the feature gradient difference and feature structure difference between adjacent levels. The feature gradient difference is measured by the first-order norm, and the feature structure difference is measured by the second-order norm. Based on the edge-preserving constraint term, construct a multi-level consistency loss function, and optimize the learnable feature transformation matrix by minimizing the multi-level consistency loss function.
4. The method according to claim 1, wherein Establish a feature evaluation matrix by calculating the intra-class cohesion and inter-class distance of each dimension feature. Dynamically allocate the weight coefficients of each dimension feature based on the feature evaluation matrix, and fuse to generate a discriminative enhanced feature description vector, including: Construct an adaptive bandwidth parameter based on the local data distribution characteristics. The adaptive bandwidth parameter is dynamically adjusted by calculating the Euclidean distance from the sample to the class center and the local density estimate value calculated using the multi-scale Gaussian kernel function; Use the adaptive bandwidth parameter to construct a dynamic kernel function. The dynamic kernel function combines a distance modulation function to evaluate the structural similarity between sample pairs, and determines an optimal neighborhood set based on the structural similarity; Apply an adaptive weight to the sample pairs in the optimal neighborhood set based on the dynamic kernel function, calculate the weighted intra-class cohesion of each dimension feature, and introduce a temporal consistency constraint to ensure the temporal continuity of the weighted intra-class cohesion; Construct a first feature evaluation matrix, which includes the weighted intra-class cohesion and the inter-class distance metric obtained by calculating the Mahalanobis distance and distribution divergence between different class centers; Construct a second feature evaluation matrix based on the weighted intra-class cohesion and the inter-class distance metric, and use the softmax function to convert the combined discriminative scores of the first feature evaluation matrix and the second feature evaluation matrix into weight coefficients. Use the weight coefficients to adaptively fuse each dimension feature to generate a discriminative enhanced feature description vector.
5. The method according to claim 4, characterized in that, Construct a second feature evaluation matrix based on the weighted intra-class cohesion and the inter-class distance metric, and use the softmax function to convert the combined discriminative scores of the first feature evaluation matrix and the second feature evaluation matrix into weight coefficients, including: Construct a second feature evaluation matrix based on the discriminative scores calculated from the weighted intra-class cohesion and the inter-class distance metric. The second feature evaluation matrix is optimized by the node features in the multi-dimensional feature association graph structure, which includes a node set, an edge set, and an edge weight set. The node features are generated through feature correlation calculation and graph attention mechanism; Combine the first feature evaluation matrix and the second feature evaluation matrix to obtain discriminative scores. The discriminative scores are obtained by applying adaptive weight coefficients to the first feature evaluation matrix and the second feature evaluation matrix and performing a weighted summation operation. The adaptive weight coefficients are determined based on the reliability and discriminative ability of each feature evaluation matrix; Use the softmax function to convert the discriminative scores into weight coefficients, which are used to adaptively fuse the features of each dimension.
6. The method according to claim 1, wherein The first stage of the two-stage cascade classifier is implemented by a lightweight decision tree for fast pre-classification. The second stage introduces an attention-enhanced deep network to extract local detail features and global semantic features, through feature adaptation The fusion module integrates the classification results of the two stages to achieve accurate classification, including: Construct a two-stage cascade classifier. The first stage of the two-stage cascade classifier uses a lightweight decision tree for fast pre-classification. The lightweight decision tree obtains a sample difficulty evaluation value by calculating the distance from the sample to the decision boundary and the local density, and dynamically weights the training samples according to the sample difficulty evaluation value; Construct a curriculum learning strategy based on the sample difficulty evaluation value. The curriculum learning strategy determines the training samples according to the dynamic distribution of the sample difficulty, generates adversarial samples for the training samples, and inputs the adversarial samples into the second stage of the two-stage cascade classifier; The second stage of the two-stage cascade classifier includes an attention-enhanced deep network. Based on the attention-enhanced deep network, local detail features and global semantic features are extracted respectively. Adaptive noise injection is performed on the local detail features and the global semantic features, and the feature consistency loss from different perspectives is calculated; Perform adaptive fusion on the pre-classification results of the lightweight decision tree and the feature extraction results of the attention-enhanced deep network. The adaptive fusion is obtained by weighted calculation of the confidence levels of the classification results of the two stages; Dynamically adjust the decision threshold based on the change rate of the confidence level, and use the decision threshold to optimize the fused classification results. Evaluate and correct the reliability of the classification results by calculating the prediction probability in the confusion matrix to generate the final classification results.
7. A packaging material defect detection system based on machine vision, which is used to implement the method described in any one of the foregoing claims 1-6, characterized in that, Including: A first unit for collecting multi-angle image sequences and constructing a depth map through camera calibration parameters, geometrically correcting and light compensating the multi-angle image sequences according to the depth map, and generating a preprocessed enhanced image; A second unit, configured to construct a multi-scale directional filter bank based on the preprocessed enhanced image, where the multi-scale directional filter bank introduces an adaptive directional enhancement factor to generate a sequence of feature maps, calculates local contrast and maximum response direction based on the sequence of feature maps, constructs a feature pyramid structure, optimizes and generates a sequence of adaptive segmentation thresholds through an edge-preserving constraint term, and implements hierarchical segmentation of the defect area based on the sequence of adaptive segmentation thresholds; A third unit, configured to extract a multi-dimensional feature set from the defect area obtained by the hierarchical segmentation, where the multi-dimensional feature set includes local texture features, gray-scale statistical features, and shape description features, calculates the within-class cohesion and between-class distance of the multi-dimensional feature set to construct a feature evaluation matrix, and assigns weight coefficients to the features in each dimension according to the feature evaluation matrix, and fuses them to generate an enhanced discriminative feature description vector; A fourth unit, configured to input the enhanced discriminative feature description vector into a two-stage cascade classifier, where the first stage of the two-stage cascade classifier uses a lightweight decision tree for fast pre-classification, and the second stage of the two-stage cascade classifier extracts local detail features and global semantic features through an attention-enhanced deep network, and uses a feature adaptive fusion module to integrate the classification results of the two stages.
8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Automobile injection molding part defect classification method based on database comparison
CN120747643A
Visual technology-based method and system for classifying and identifying defective parts of clock parts
CN120808037A
Clock accessory defect classification method and system based on visual technology
CN120808037B
Industrial flaw detection method based on cross-layer collaborative multi-level attention and progressive feature enhancement
CN122048931A
Industrial defect detection method based on cross-layer cooperative multi-level attention and progressive feature enhancement
CN122048931B