Softwood defect detection and grading method based on instance segmentation and visual language model

CN122841309APending Publication Date: 2026-09-29SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610997469.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0007]针对现有软木人工检测效率低、检测结果不稳定,以及现有目标检测和实例分割方法对复杂纹理、细小缺陷和多等级分类适应性不足的问题,本发明提供一种基于实例分割和视觉语言模型的软木缺陷检测及分级方法基于

Benefits of technology

(1)本发明通过将浅层卷积特征和深层卷积特征进行多尺度融合,使网络同时获得软木表面细小缺陷的边缘纹理细节和高层语义信息,能够提高红点、黑点、孔隙和蜂窝状缺陷等小目标缺陷的检测能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841309A_ABST
    Figure CN122841309A_ABST
Patent Text Reader

Abstract

This invention discloses a method for cork defect detection and grading based on instance segmentation and a visual language model. The method includes: preprocessing multi-view side images of the cork to be detected; extracting the side images of the cork to be detected through discrete points of the level set and narrow band regions; constructing a multi-scale CorkMS-RCNN detection and segmentation network, performing scale matching, L2 normalization, and splicing fusion on shallow convolutional feature maps and deep convolutional feature maps; generating candidate defect regions, aligning the candidate region features using RoI Align, outputting the defect category and bounding box through the detection branch, and outputting the segmentation mask through the mask branch; inputting the cork side images and defect detection and segmentation results into a visual language model, matching them with preset cork grade text labels to obtain the cork grade. This invention can improve the accuracy of detection, segmentation, and grading of small defects on the surface of cork with complex textures, meeting the application requirements of rapid detection and automatic sorting in cork industrial production lines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of machine vision inspection, deep learning target detection, instance segmentation, visual language model and industrial automatic sorting technology, specifically involving a method for cork defect detection and grading based on instance segmentation and visual language model. Background Technology

[0002] Cork is lightweight, elastic, wear-resistant, and highly workable, making it suitable for products such as bottle stoppers, badminton shuttlecock heads, and racket handles. During cork production, it is typically graded based on surface color, texture, and defects to ensure consistency in subsequent splicing, assembly, and finished product quality.

[0003] Current cork screening and grading processes primarily rely on manual visual inspection, which suffers from high labor intensity, low efficiency, strong subjectivity, and a high rate of misclassification. Common surface defects in cork include mold, porosity, red spots, black spots, honeycomb defects, old bark residue, missing pores, red bark, black bark, T-shaped defects, incomplete defects, and stains. Because cork is a natural material, the texture, color, porosity distribution, and defect morphology of different types of cork vary significantly, making cork grading highly challenging.

[0004] Traditional machine vision methods typically rely on color, texture, or manually designed features for classification, which lacks stability when faced with complex textures and multi-level classification tasks. Conventional convolutional neural networks can automatically learn image features, but the local receptive field of the convolutional kernel limits the ability to model global relationships in complex textures; while one-stage object detection networks are fast, they lack the ability to locate and finely segment small-scale defects; although the traditional Mask R-CNN can perform object detection and instance segmentation simultaneously, its segmentation branch relies excessively on deep convolutional features, which have a large receptive field and easily mask small defects such as red dots, black dots, and honeycomb defects.

[0005] Furthermore, cork surfaces have complex textures, irregular defect boundaries, and low local contrast, making the output rating results of detection networks easily affected by the size of the training samples and the distribution of categories. Visual language pre-trained models possess cross-modal representation capabilities and strong transfer capabilities. Matching cork images with rating text labels based on similarity can improve the adaptability of multi-grade cork classification.

[0006] Therefore, there is an urgent need for a cork defect detection and grading technology that can simultaneously achieve cork side image acquisition, accurate extraction of side regions, multi-scale detection of small defects, pixel-level segmentation, and visual language level matching, in order to meet the needs of rapid detection and automatic sorting in industrial production lines. Summary of the Invention

[0007] To address the issues of low efficiency and unstable results in existing manual cork inspection methods, as well as the insufficient adaptability of existing target detection and instance segmentation methods to complex textures, small defects, and multi-level classification, this invention provides a cork defect detection and grading method based on instance segmentation and a visual language model. This method collaboratively designs multi-view side image acquisition of cork, side region extraction based on templates and level sets, a multi-scale instance segmentation network, and visual language model level matching to achieve cork defect detection, pixel-level segmentation, grade determination, and sorting control.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: A method for cork defect detection and grading based on instance segmentation and visual language models includes a model training stage, an image detection and segmentation stage, and a cork grading and classification stage. Step S1: Obtain cork image training samples and preprocess the cork image training samples to construct training samples corresponding to cork side images, defect annotations, mask annotations and cork grade labels. Step S2: Construct a multi-scale instance segmentation network, which includes a backbone network, a feature fusion module, a region candidate network, a RoI Align module, a detection branch, and a mask branch. Step S3: Input the cork side image into the backbone network, extract multi-level convolutional features from the backbone network, and fuse the shallow convolutional features with the deep convolutional features at multiple scales to obtain a fused feature map. Step S4: Input the fused feature map into the region candidate network to generate candidate defect regions, and perform foreground / background discrimination and bounding box regression on the candidate defect regions; Step S5: Align the candidate defect regions using the RoI Align module to obtain candidate region features of a fixed size; Step S6: Input the candidate region features into the detection branch and the mask branch respectively. Output the defect category and defect bounding box through the detection branch, and output the pixel-level segmentation mask of the defect region through the mask branch. Step S7: Train a multi-scale instance segmentation network based on detection loss and mask segmentation loss; The image detection and segmentation stage includes the following steps; Step S8: Obtain the side image of the cork to be detected, and preprocess and extract the side region of the side image of the cork to be detected to obtain the side image of the cork to be detected. Step S9: Input the cork side image to be detected into the trained multi-scale instance segmentation network, and output the defect category, defect bounding box and pixel-level segmentation mask; Step S10: Input the cork side image to be detected, defect category, defect bounding box and / or pixel-level segmentation mask into the visual language model, and perform similarity matching with the preset cork grade text label; Step S11: Determine the cork grade based on the cork grade text label with the highest similarity, and output the cork defect detection and grading results.

[0009] Furthermore, the lateral region extraction includes: determining the cork image center based on the cylindrical contour features of the cork, establishing multiple radial search paths starting from the cork image center, and setting horizontal and / or vertical search templates on the radial search paths to extract the cork lateral region; the search direction of the search template is approximately perpendicular to the edge direction of the cork image to shorten the contour iteration search path and increase the edge gradient response; the lateral region extraction also includes determining the cork contour boundary based on the level set evolution method, converting the level set curve into multiple discrete contour points, and constructing a narrow band region based on the discrete contour points, so that the search template scans along the narrow band region; the level set evolution method uses the finite difference method for numerical solution, where the spatial domain partial derivatives are calculated using the central difference method and the time domain partial derivatives are calculated using the forward difference method.

[0010] Furthermore, the backbone network comprises five convolutional layers, with shallow convolutional features including feature maps from the third and fourth convolutional layers, and deep convolutional features including feature maps from the fifth convolutional layer.

[0011] Furthermore, multi-scale fusion includes: downsampling the feature maps of the third and fourth convolutional layers to a spatial resolution that matches the feature map of the fifth convolutional layer, performing L2 normalization on the feature maps of different scales, and concatenating or cascading the normalized feature maps to obtain a fused feature map.

[0012] Furthermore, the region candidate network generates multiple anchor boxes of different scales and aspect ratios on the fused feature map through a sliding window, and performs target confidence judgment and bounding box coordinate regression on each anchor box to obtain candidate defect regions.

[0013] Furthermore, the RoI Align module extracts fixed-size feature tensors from the fused feature map based on the bounding box coordinates of the candidate defect regions, and normalizes and concatenates the fixed-size feature tensors into unified candidate region features.

[0014] Furthermore, the detection branch includes a classification layer and a bounding box regression layer. The classification layer is used to output the probability that a candidate defect region belongs to each defect category, and the bounding box regression layer is used to output the bounding box coordinate correction value of the candidate defect region.

[0015] Furthermore, the mask branch includes a convolutional layer and an upsampling layer. The convolutional layer is used to generate a low-resolution binary mask for the candidate defect region, and the upsampling layer is used to restore the low-resolution binary mask to the original size corresponding to the candidate defect region to obtain a pixel-level segmentation mask.

[0016] Furthermore, the detection loss includes classification loss and bounding box regression loss. The classification loss uses cross-entropy loss, the bounding box regression loss uses Smooth L1 loss, and the mask segmentation loss uses pixel-level binary classification cross-entropy loss. The total loss of the multi-scale instance segmentation network is the weighted sum of the detection loss and the mask segmentation loss.

[0017] Furthermore, the visual language model includes an image encoder and a text encoder. The image encoder is used to encode the cork side image to be detected and / or the defect detection segmentation result into an image embedding vector. The text encoder is used to encode the cork grade text label into a text embedding vector, and the cork grade is determined based on the similarity between the image embedding vector and the text embedding vector. Compared with the prior art, the present invention has at least the following beneficial effects: (1) This invention integrates shallow convolution features and deep convolution features at multiple scales, enabling the network to simultaneously obtain edge texture details and high-level semantic information of small defects on the cork surface, thereby improving the detection capability of small target defects such as red dots, black dots, pores and honeycomb defects.

[0018] (2) This invention balances the influence of features at different scales on the detection results by L2 normalization, avoiding shallow detail features or deep semantic features from dominating the detection results alone, and improving the detection stability in complex texture scenes.

[0019] (3) The present invention uses RoI Align to align the features of the candidate defect region and generates a pixel-level segmentation mask through mask branching, which can more accurately locate the cork defect region and improve the problem that the traditional detection method only outputs the bounding box and cannot finely describe the defect morphology.

[0020] (4) The present invention uses a visual language model to match the similarity between cork images and grade text labels, which can enhance the adaptability of cork multi-grade classification and reduce the dependence on specific classification heads.

[0021] (5) This invention can be combined with an image acquisition device and a sorting execution mechanism to realize a closed-loop production line from cork feeding, image acquisition, defect detection, grade determination to automatic sorting, thereby improving industrial production efficiency and grading consistency. Attached Figure Description

[0022] Figure 1 This is a schematic diagram illustrating the cork grading system of the present invention; Figure 2This is a schematic diagram of the cork detection and sorting equipment of the present invention; Figure 3 This is a schematic diagram of the cork image acquisition device and the side imaging principle of the present invention; Figure 4 This is a schematic diagram of cork side image acquisition according to the present invention; Figure 5 This is a schematic diagram of the template search process for the cork side area of ​​the present invention; Figure 6 This is a schematic diagram of the multi-scale CorkMS-RCNN detection and segmentation network structure of the present invention. Figure 7 This is a schematic diagram of the cork grading structure based on a visual language model according to the present invention. Detailed Implementation

[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments, but the scope of protection of the present invention is not limited to the following embodiments. Unless otherwise specified, the technical features of the various embodiments of the present invention can be combined with each other.

[0024] Example 1 A method for cork defect detection and grading based on instance segmentation and visual language models includes a model training stage, an image detection and segmentation stage, and a cork grading and classification stage. Step S1: Obtain cork image training samples and preprocess the cork image training samples to construct training samples corresponding to cork side images, defect annotations, mask annotations and cork grade labels. Step S2: Construct a multi-scale instance segmentation network, which includes a backbone network, a feature fusion module, a region candidate network, a RoI Align module, a detection branch, and a mask branch. Step S3: Input the cork side image into the backbone network, extract multi-level convolutional features from the backbone network, and fuse the shallow convolutional features with the deep convolutional features at multiple scales to obtain a fused feature map. Step S4: Input the fused feature map into the region candidate network to generate candidate defect regions, and perform foreground / background discrimination and bounding box regression on the candidate defect regions; Step S5: Align the candidate defect regions using the RoI Align module to obtain candidate region features of a fixed size; Step S6: Input the candidate region features into the detection branch and the mask branch respectively. Output the defect category and defect bounding box through the detection branch, and output the pixel-level segmentation mask of the defect region through the mask branch. Step S7: Train a multi-scale instance segmentation network based on detection loss and mask segmentation loss; The image detection and segmentation stage includes the following steps; Step S8: Obtain the side image of the cork to be detected, and preprocess and extract the side region of the side image of the cork to be detected to obtain the side image of the cork to be detected. Step S9: Input the cork side image to be detected into the trained multi-scale instance segmentation network, and output the defect category, defect bounding box and pixel-level segmentation mask; Step S10: Input the cork side image to be detected, defect category, defect bounding box and / or pixel-level segmentation mask into the visual language model, and perform similarity matching with the preset cork grade text label; Step S11: Determine the cork grade based on the cork grade text label with the highest similarity, and output the cork defect detection and grading results.

[0025] Specifically, the lateral region extraction includes: determining the cork image center based on the cylindrical contour features of the cork; establishing multiple radial search paths starting from the cork image center; and setting horizontal and / or vertical search templates on the radial search paths to extract the cork lateral region; the search direction of the search template is approximately perpendicular to the edge direction of the cork image to shorten the contour iteration search path and increase the edge gradient response; the lateral region extraction also includes determining the cork contour boundary based on the level set evolution method, converting the level set curve into multiple discrete contour points, and constructing a narrow band region based on the discrete contour points, so that the search template scans along the narrow band region; the level set evolution method uses the finite difference method for numerical solution, where the spatial domain partial derivatives are calculated using the central difference method and the time domain partial derivatives are calculated using the forward difference method.

[0026] Specifically, the backbone network consists of five convolutional layers. The shallow convolutional features include the feature maps of the third and fourth convolutional layers, while the deep convolutional features include the feature map of the fifth convolutional layer.

[0027] Specifically, multi-scale fusion includes: downsampling the feature maps of the third and fourth convolutional layers to a spatial resolution that matches the feature map of the fifth convolutional layer, performing L2 normalization on the feature maps of different scales, and concatenating or cascading the normalized feature maps to obtain a fused feature map.

[0028] Specifically, the region candidate network generates multiple anchor boxes of different scales and aspect ratios on the fused feature map through a sliding window, and performs target confidence judgment and bounding box coordinate regression on each anchor box to obtain candidate defect regions.

[0029] Specifically, the RoI Align module extracts a fixed-size feature tensor from the fused feature map based on the bounding box coordinates of the candidate defect region, and normalizes and concatenates the fixed-size feature tensor into a unified candidate region feature.

[0030] Specifically, the detection branch includes a classification layer and a bounding box regression layer. The classification layer is used to output the probability that a candidate defect region belongs to each defect category, and the bounding box regression layer is used to output the bounding box coordinate correction value of the candidate defect region.

[0031] Specifically, the mask branch includes a convolutional layer and an upsampling layer. The convolutional layer is used to generate a low-resolution binary mask for the candidate defect region, and the upsampling layer is used to restore the low-resolution binary mask to the original size corresponding to the candidate defect region to obtain a pixel-level segmentation mask.

[0032] Specifically, the detection loss includes classification loss and bounding box regression loss. The classification loss uses cross-entropy loss, the bounding box regression loss uses Smooth L1 loss, and the mask segmentation loss uses pixel-level binary classification cross-entropy loss. The total loss of the multi-scale instance segmentation network is the weighted sum of the detection loss and the mask segmentation loss.

[0033] Specifically, the visual language model includes an image encoder and a text encoder. The image encoder is used to encode the cork side image to be detected and / or the defect detection segmentation result into an image embedding vector. The text encoder is used to encode the cork grade text label into a text embedding vector. The cork grade is determined based on the similarity between the image embedding vector and the text embedding vector. Step S1: Obtain cork image training samples and preprocess the cork image training samples to construct training samples corresponding to the cork side image, defect annotation, mask annotation and cork grade label.

[0034] Specifically, the cork image acquisition device includes a conveying mechanism, a camera, a lens, a ring light source, and at least one reflector, which is used to reflect the image of the cork side onto the camera imaging area. Specifically, the cork image acquisition device also includes a vibrating feeding mechanism, a sorting tray, a detection sensor, a nozzle, a collection pipe, and a storage container. The vibrating feeding mechanism is used to transport cork to the conveying mechanism, and the nozzle, collection pipe, and storage container are used to perform automatic sorting according to the cork grade.

[0035] Specifically, the cork image acquisition device includes four reflectors arranged around the cork to be detected, so as to acquire cork side images from multiple directions when the cork passes through the imaging area; the size of the cork side images is uniformly 256×256 pixels.

[0036] Specifically, the search direction of the search template is approximately perpendicular to the edge direction of the cork image to shorten the contour iteration search path and increase the edge gradient response.

[0037] Specifically, the lateral region extraction also includes determining the cork contour boundary based on the level set evolution method, converting the level set curve into multiple discrete contour points, and constructing a narrow band region based on the discrete contour points, so that the search template scans along the narrow band region.

[0038] Specifically, the level set evolution method uses the finite difference method for numerical solution, where the spatial domain partial derivatives are calculated using the central difference method and the time domain partial derivatives are calculated using the forward difference method.

[0039] Specifically, the visual language model is a CLIP model, or a visual language pre-trained model with image encoder, text encoder and cross-modal similarity matching capabilities; the preset cork grade text labels include one or more grade descriptions of 4A, 3A, 2A, A, B, C, and / or include one or more industrial grade descriptions of A, B, C, D, E.

[0040] Specifically, the defect categories include one or more of the following: mold, pores, red spots, black spots, honeycomb defects, old bark residue, missing holes, red bark, black bark, T-shaped defects, incomplete defects, and stains.

[0041] Specifically, the cork grading stage also includes: generating sorting control signals based on cork grading and / or defect detection and segmentation results, and controlling nozzles, collection pipes and / or storage containers to classify and collect cork of different grades.

[0042] A cork defect detection and grading system based on multi-scale Mask R-CNN and a visual language model includes: an image acquisition module for acquiring cork side images; a side region extraction module for extracting cork side regions based on radial search paths, search templates, and discrete points of the level set; a multi-scale feature fusion module for performing scale matching, normalization, and concatenation of shallow and deep convolutional features; a candidate region generation module for generating candidate defect regions based on a region candidate network; a detection and segmentation module for outputting defect categories, defect bounding boxes, and pixel-level segmentation masks based on RoI Align, detection branches, and mask branches; a visual language classification module for performing similarity matching between cork images and / or defect detection and segmentation results and cork grade text labels; and a grading output module for outputting cork defect detection and grading results.

[0043] An electronic device includes a processor and a memory, the memory storing a computer program that, when executed by the processor, implements a method for cork defect detection and grading based on instance segmentation and a visual language model, according to the present invention.

[0044] A computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for detecting and classifying cork defects based on instance segmentation and a visual language model.

[0045] Example 2 Cork image acquisition and preprocessing, such as Figure 2and Figure 3 As shown, the cork to be tested is fed into the distribution tray by a vibrating feeding mechanism and then sequentially passes through the image acquisition area under the drive of the conveyor mechanism. The image acquisition area is equipped with a camera, lens, ring light source, and a set of reflectors. The set of reflectors is arranged around the cork to be tested and is used to reflect side images of the cork from different directions to the camera imaging area.

[0046] In one implementation, the camera can simultaneously acquire side images of the cork from four directions and a top-view image. After acquisition, the cork images are uniformly adjusted to 256×256 pixels to meet the image resolution requirements for identifying small defects. The lens focal length can be determined based on the imaging size of the smallest defect area; when the defect size is below a preset pixel threshold and does not affect the quality of the cork product, it may not be marked as a target defect.

[0047] The preprocessing includes image cropping, resizing, grayscale or color normalization, low-quality image removal, defect annotation, and mask annotation. During the training phase, training samples are constructed based on defect categories, defect bounding boxes, pixel-level masks, and cork grade labels. The detection phase applies the same resizing and normalization processing to the cork images to be detected to obtain image data suitable for input to the trained model.

[0048] Example 3 Extraction of the side area of ​​cork, such as Figure 4 and Figure 5 As shown, since cork typically has an approximately cylindrical structure, this embodiment determines the cork center based on the cork image contour features and establishes multiple radial search paths outward from the cork center. For cork edges in different directions, horizontal and vertical search templates are set so that the search path directions are approximately perpendicular to the cork image edge directions.

[0049] The above search path design minimizes the contour iteration path and maximizes edge gradient response in the search direction. The search template scans the cork image along a radial path to extract the cork side region and form the cork side detection image.

[0050] In one implementation, a level set evolution method is used to determine the cork contour boundary. The intersection of the level set curve with each radial search path yields multiple discrete contour points, thus transforming the continuous curve-form level set into multiple discrete points. Furthermore, a narrow band region is constructed centered on these discrete contour points, allowing the template to scan along this narrow band region, thereby reducing computational load and improving the efficiency of cork side region extraction.

[0051] In the evolution of the level set, the finite difference method can be used for numerical solutions. The spatial domain partial derivatives are calculated using the central difference method, and the time domain partial derivatives are calculated using the forward difference method. Let the level set function be... Its curvature term satisfies: ,in , Let represent the first-order partial derivatives of the horizontal set function in the horizontal and vertical directions, respectively. , , Let represent the second-order partial derivatives of the level set function. By constraining the evolution of the level set using the curvature term described above, the side profile of the cork can be stably obtained, providing accurate region input for subsequent defect detection.

[0052] Example 4 Multi-scale CorkMS-RCNN detection and segmentation networks, such as Figure 6 As shown, the multi-scale CorkMS-RCNN detection and segmentation network includes a backbone network, a feature fusion module, a region candidate network, an RoI Align module, a detection branch, and a mask branch. The original cork side image is first input into the backbone network, which consists of five convolutional layers used to extract hierarchical features from shallow texture details to deep semantic information layer by layer.

[0053] Shallow feature maps have high spatial resolution, preserving detailed information such as red and black dots, pore boundaries, and honeycomb defects on the cork surface. Deep feature maps have stronger semantic expressive power, capable of characterizing the overall structure and category features of the defect region. To balance detail and semantics, this embodiment retains the feature maps of the third, fourth, and fifth convolutional layers. The feature maps of the third and fourth convolutional layers are downsampled to a spatial resolution matching that of the feature map of the fifth convolutional layer, and then L2 normalization and splicing are performed to obtain a unified fused feature map. Specifically, let the output feature maps of the third, fourth, and fifth convolutional layers be respectively... The fused feature map is input into the region candidate network. The region candidate network generates multiple anchor boxes on the fused feature map using a 3×3 convolutional sliding window. These anchor boxes have different scales and aspect ratios to accommodate variations in the size and shape of cork defects. The region candidate network performs target confidence discrimination on each anchor box to determine whether it belongs to the foreground defect region or the background region. For foreground anchor boxes, bounding box regression is further performed to obtain candidate defect regions that better fit the defect area.

[0054] Candidate defect regions are input into the RoI Align module. RoI Align extracts fixed-size candidate region features from the fused feature map based on the bounding box coordinates of the candidate defect regions, avoiding localization errors caused by candidate region coordinate quantization. The candidate region features processed by RoI Align are then input into two parallel branches: one is a detection branch, which outputs the defect category probability and bounding box coordinate correction values; the other is a mask branch, which generates a low-resolution binary mask of the candidate defect regions and upsamples it to restore the original size of the candidate regions to obtain a pixel-level segmentation mask.

[0055] The detection branch and the masking branch share the front-end multi-scale feature extraction module, but they correspond to different training objectives. The loss of the detection branch includes classification loss and bounding box regression loss, where the classification loss uses cross-entropy loss and the bounding box regression loss uses Smooth L1 loss; the loss of the masking branch uses pixel-level binary classification cross-entropy loss. The total loss function is a weighted sum of the detection loss and the masking segmentation loss. The object detection loss is expressed as: ,in Indicates the target detection loss. Represents classification loss, This represents the bounding box regression loss. These are the weighting coefficients. The classification loss uses cross-entropy loss: ,in This indicates that the candidate defect region is predicted as the first... The probability of a class The bounding box regression loss is represented by the one-hot encoding of the true class label, where n represents the number of defect classes, and the additional class represents the background class. ,in, Indicates the predicted bounding box parameters. Represents the actual bounding box parameters. This represents the center coordinates, width, and height parameters of the bounding box. The Smooth L1 loss function is expressed as: The mask segmentation loss uses pixel-level binary classification cross-entropy loss: ,in This indicates the total number of pixels in the mask area. Indicates the first The real label of each pixel The model predicts the first... The probability that a pixel belongs to a defect region. The total loss function of the multi-scale object detection and segmentation network is expressed as: ,in and The weights are used to balance the impact of object detection and mask segmentation tasks on network training. Through joint optimization of the above loss functions, the network can simultaneously learn cork defect category discrimination, defect bounding box localization, and pixel-level segmentation of defect regions.

[0056] Example 5 Visual language model for cork grading, such as Figure 7 As shown, the visual language model includes an image encoder and a text encoder. During training or inference, the image encoder converts cork side images, defect detection results, and / or mask segmentation results into fixed-dimensional image embedding vectors; the text encoder converts preset cork grade text labels into text embedding vectors corresponding to the dimensions of the image embedding vectors. The image embedding vectors and text embedding vectors reside in the same multimodal feature space, enabling the cork image features to be matched with the semantics of the cork grade text.

[0057] like Figure 1 As shown, in one embodiment, the preset cork grade text label includes descriptions such as "This cork is classified as 4A," "This cork is classified as 3A," "This cork is classified as 2A," "This cork is classified as A," "This cork is classified as B," and "This cork is classified as C." Among these, 4A, 3A, 2A, and A can correspond to higher quality cork grades, B can correspond to ordinary quality cork grades, and C can correspond to substandard or low-quality cork grades. The above grade classification is only an example; in actual applications, other grade systems such as A, B, C, D, and E can be set according to the quality standards of different enterprises or production lines, and corresponding text descriptions can be configured for each grade.

[0058] The side image of the cork to be detected is input into the image encoder; or candidate defect regions are cropped from the cork side image based on the defect bounding box, and a defect highlight image is generated by combining a pixel-level segmentation mask, and then the candidate defect region image and / or defect highlight image are input into the image encoder. Subsequently, the similarity between the image embedding features and each text embedding feature is calculated to form an image-text similarity matrix, and the text label with the highest similarity level is determined as the predicted level of the cork to be detected.

[0059] In one implementation, the visual language model can be a CLIP model; in other implementations, other pre-trained visual language models with image encoders, text encoders, and cross-modal similarity matching capabilities can also be used. By introducing graded text labels into the classification, the system can simultaneously reference cork appearance, defect distribution, and graded semantic information, improving the adaptability of multi-grade classification for complex textured cork.

[0060] Example 6 For training, detection, and industrial sorting applications, preprocessed and labeled cork side images are input into the multi-scale CorkMS-RCNN detection and segmentation network during the model training phase. Training samples include cork side images, defect category labels, defect bounding box labels, pixel-level mask labels, and cork grade labels. The network is jointly optimized using classification loss, bounding box regression loss, and pixel-level mask segmentation loss to learn the category features, spatial location features, and pixel-level morphological features of cork defects. After training, the trained multi-scale CorkMS-RCNN detection and segmentation network is deployed in the computing unit of industrial inspection equipment.

[0061] In one specific implementation, the training samples are derived from real cork images captured by a cork image acquisition device. The acquired images are uniformly adjusted to 256×256 pixels and divided into training, validation, and test sets according to a preset ratio. The Adam optimizer can be used during model training, combined with a learning rate adjustment strategy to improve training stability. The image size, training sample division method, optimizer type, and learning rate adjustment strategy described above are merely examples and should not be construed as limiting the scope of protection of this invention.

[0062] During the detection phase, the cork to be tested is conveyed to the image acquisition area via a vibrating feeding mechanism, a distribution tray, and a conveyor mechanism. The system uses a camera, lens, ring light source, and a set of reflectors to acquire multi-view side images of the cork. Subsequently, the multi-view side images are sized, normalized, and the side regions are extracted to obtain the side image of the cork to be tested.

[0063] The image of the cork side surface to be detected is input into the trained multi-scale CorkMS-RCNN detection and segmentation network to obtain the category, bounding box location, and pixel-level segmentation mask of the cork surface defects. The defects may include one or more of the following: mold, porosity, red spots, black spots, honeycomb defects, old bark residue, missing holes, red bark, black bark, T-shaped defects, incomplete defects, and stains.

[0064] Subsequently, the cork side image to be inspected and / or the aforementioned defect detection and segmentation results are input into the visual language model and matched with the preset cork grade text labels for similarity matching. For production scenarios using a six-level classification system of 4A, 3A, 2A, A, B, and C, the system can use "This cork is classified as 4A," "This cork is classified as 3A," "This cork is classified as 2A," "This cork is classified as A," "This cork is classified as B," and "This cork is classified as C" as candidate grade text labels, and determine the cork grade based on the matching results between image features and each candidate text label. For production scenarios using other enterprise standards, the corresponding grade name and grade description can be replaced with A, B, C, D, E, or other custom grade text. For multiple side images of the same cork to be inspected, defect detection, pixel-level segmentation, and grade prediction are performed separately, and a comprehensive judgment is made based on the defect category, defect quantity, defect area ratio, and / or predicted grade corresponding to each side image. When the prediction results from multiple perspectives are inconsistent, the grade corresponding to the side image with the highest degree of defect can be taken as the final grade of the cork; or the detection results and grade prediction results from multiple perspectives can be weighted and fused according to preset weights to determine the final grade of the cork.

[0065] like Figure 2 As shown, the system generates sorting control signals based on cork grade and defect detection results, controlling nozzles, collection pipes, and storage containers to classify and collect cork of different grades. For example, when cork is determined to be grade 4A, 3A, 2A, or A, it can be conveyed to the corresponding high-grade storage container; when cork is determined to be grade B, it can be conveyed to the ordinary-grade storage container; when cork is determined to be grade C or has preset serious defects, it can be conveyed to the non-conforming product storage container. Specific sorting rules can be adjusted according to production line quality standards, product uses, and storage container configuration.

[0066] Through the above process, the present invention can realize closed-loop processing of cork from image acquisition, side area extraction, defect detection, pixel-level segmentation, grade classification to automatic sorting, thereby improving the detection efficiency, grading consistency and automation level of cork industrial production lines.

[0067] Example 7 In addition to systems, electronic devices, and storage media, this invention also provides a cork defect detection and grading system based on multi-scale Mask R-CNN and a visual language model, including an image acquisition module, a lateral region extraction module, a multi-scale feature fusion module, a candidate region generation module, a detection and segmentation module, a visual language classification module, and a grading output module. Each module can be implemented by software programs, hardware circuits, or a combination of both.

[0068] The present invention also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the method described in any embodiment of the present invention. The electronic device may be an industrial computer, an edge computing device, a server, an embedded controller, or an image processing workstation.

[0069] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method described in any embodiment of the present invention. The computer-readable storage medium includes, but is not limited to, a read-only memory, a random access memory, flash memory, a solid-state drive, a magnetic disk, or an optical disk.

Claims

1. A method for cork defect detection and grading based on instance segmentation and visual language model, characterized in that, This includes the model training phase, the image detection and segmentation phase, and the cork grading and classification phase. Step S1: Obtain cork image training samples and preprocess the cork image training samples to construct training samples corresponding to cork side images, defect annotations, mask annotations and cork grade labels. Step S2: Construct a multi-scale instance segmentation network, which includes a backbone network, a feature fusion module, a region candidate network, a RoI Align module, a detection branch, and a mask branch. Step S3: Input the cork side image into the backbone network, extract multi-level convolutional features from the backbone network, and fuse the shallow convolutional features with the deep convolutional features at multiple scales to obtain a fused feature map. Step S4: Input the fused feature map into the region candidate network to generate candidate defect regions, and perform foreground / background discrimination and bounding box regression on the candidate defect regions; Step S5: Align the candidate defect regions using the RoI Align module to obtain candidate region features of a fixed size; Step S6: Input the candidate region features into the detection branch and the mask branch respectively. Output the defect category and defect bounding box through the detection branch, and output the pixel-level segmentation mask of the defect region through the mask branch. Step S7: Train a multi-scale instance segmentation network based on detection loss and mask segmentation loss; The image detection and segmentation stage includes the following steps; Step S8: Obtain the side image of the cork to be detected, and preprocess and extract the side region of the side image of the cork to be detected to obtain the side image of the cork to be detected. Step S9: Input the cork side image to be detected into the trained multi-scale instance segmentation network, and output the defect category, defect bounding box and pixel-level segmentation mask; Step S10: Input the cork side image to be detected, defect category, defect bounding box and pixel-level segmentation mask into the visual language model, and perform similarity matching with the preset cork grade text label; Step S11: Determine the cork grade based on the cork grade text label with the highest similarity, and output the cork defect detection and grading results.

2. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The lateral region extraction includes: determining the cork image center based on the cylindrical contour features of the cork; establishing multiple radial search paths starting from the cork image center; and setting horizontal and vertical search templates on the radial search paths to extract the cork lateral region. The search direction of the search template is approximately perpendicular to the edge direction of the cork image to shorten the contour iteration search path and increase the edge gradient response. The lateral region extraction also includes determining the cork contour boundary based on the level set evolution method, converting the level set curve into multiple discrete contour points, and constructing a narrow band region based on the discrete contour points, so that the search template scans along the narrow band region. The level set evolution method uses the finite difference method for numerical solution, where the spatial domain partial derivatives are calculated using the central difference method and the time domain partial derivatives are calculated using the forward difference method.

3. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The backbone network consists of five convolutional layers. Shallow convolutional features include feature maps from the third and fourth convolutional layers, while deep convolutional features include feature maps from the fifth convolutional layer.

4. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 3, characterized in that, Multi-scale fusion includes: downsampling the feature maps of the third and fourth convolutional layers to a spatial resolution that matches the feature map of the fifth convolutional layer, performing L2 normalization on the feature maps of different scales, and concatenating or cascading the normalized feature maps to obtain a fused feature map.

5. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The region candidate network generates multiple anchor boxes of different scales and aspect ratios on the fused feature map through a sliding window, and performs target confidence judgment and bounding box coordinate regression on each anchor box to obtain candidate defect regions.

6. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The RoI Align module extracts fixed-size feature tensors from the fused feature map based on the bounding box coordinates of the candidate defect regions, and normalizes and concatenates these fixed-size feature tensors into unified candidate region features.

7. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The detection branch includes a classification layer and a bounding box regression layer. The classification layer outputs the probability that a candidate defect region belongs to each defect category, while the bounding box regression layer outputs the bounding box coordinate correction value of the candidate defect region.

8. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The mask branch includes convolutional layers and upsampling layers. The convolutional layers are used to generate low-resolution binary masks for candidate defect regions, and the upsampling layers are used to restore the low-resolution binary masks to the original size corresponding to the candidate defect regions to obtain pixel-level segmentation masks.

9. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The detection loss includes classification loss and bounding box regression loss. The classification loss uses cross-entropy loss, the bounding box regression loss uses Smooth L1 loss, and the mask segmentation loss uses pixel-level binary classification cross-entropy loss. The total loss of the multi-scale instance segmentation network is the weighted sum of the detection loss and the mask segmentation loss.

10. The method for cork defect detection and grading based on instance segmentation and visual language model according to claim 1, characterized in that, The visual language model includes an image encoder and a text encoder. The image encoder encodes the cork side image to be detected and the defect detection segmentation results into image embedding vectors. The text encoder encodes the cork grade text labels into text embedding vectors and determines the cork grade based on the similarity between the image embedding vectors and the text embedding vectors.