Dual-Branch Surface Defect Segmentation Method and Medium with Semantic Guidance and Texture Prior
Through the two-branch surface defect segmentation method of semantic guidance and texture prior, the two-branch feature extraction network and feature fusion network of semantic and texture are used to solve the problem of low segmentation accuracy of complex defects and weak texture defects in existing methods, achieving higher segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202510480124.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing deep learning-based surface defect segmentation method is poor when dealing with defects with complex shapes and large size differences and weak texture defects, resulting in the problem of low segmentation accuracy.
The double-branch surface defect segmentation method of semantic guidance and texture prior is adopted. The double-branch feature extraction network and feature fusion network of semantic guidance and texture prior subnet is used to improve the feature extraction capability of defect segmentation models, especially the mining of edge texture information and the segmentation of edge contours.
The segmentation accuracy and robustness of the defect segmentation model for complex defects and weak textures is improved, and the edge contour of the defect target can be more accurately segmented.
Smart Images

Figure CN119991713B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of surface defect segmentation, and in particular, to a dual-branch surface defect segmentation method and medium with semantic guidance and texture prior. Background Art
[0002] Surface defect segmentation technology plays a crucial role in the field of industrial quality inspection and is one of the important technologies to ensure product quality and maintain stable production. In the traditional production process, the detection of product surface defects often relies on manual visual inspection, which is time-consuming and labor-intensive. More importantly, due to the interference of human subjectivity and external factors such as the environment, the detection results often have certain errors and instabilities. In recent years, with the rapid development of deep learning technology, the deep learning-based surface defect segmentation method has significantly improved the accuracy and robustness of the surface defect detection task, providing an efficient and feasible solution for industrial surface defect detection, and has been applied to many fields such as weld bubble detection, metal surface detection, and circuit board quality detection, which not only proves the practicality and effectiveness of the deep learning-based surface defect segmentation technology in actual industrial production, but also demonstrates the great potential and broad application prospects of this technology.
[0003] The common types of surface defects in industrial scenarios are dents, stains, scratches, and cracks. The deep learning-based surface defect segmentation method is to perform pixel-level classification on these surface defect targets, so as to accurately segment the surface defect targets from the whole image and obtain more refined detection results, showing higher accuracy in the surface defect detection task. Although the deep learning-based surface defect segmentation method has performed well in recent years, due to the influence of industrial imaging environment and noise, the segmentation of defects with complex shapes and large size differences and the segmentation of weak texture defects with unclear defect features still pose great challenges to most existing methods. Summary of the Invention
[0004] The purpose of the present application is to provide a dual-branch surface defect segmentation method and medium with semantic guidance and texture prior, which can solve the problems of poor segmentation and low segmentation accuracy caused by complex surface defect shapes, large size differences, and weak textures.
[0005] To achieve the above object, the present application provides the following solutions:
[0006] In the first aspect, the present application provides a dual-branch surface defect segmentation method with semantic guidance and texture prior, including:
[0007] Obtain an image of the object to be detected;
[0008] Input the image of the object to be detected into the defect segmentation model to obtain the surface defect segmentation result of the object to be detected; the defect segmentation model includes a dual-branch feature extraction network for semantics and texture, a feature fusion network for semantic guidance and texture prior, and a decoder connected in sequence; the dual-branch feature extraction network for semantics and texture is used to extract semantic information, defect texture and edge features in the image of the object to be detected, and obtain a semantic feature map and a texture feature map; the feature fusion network for semantic guidance and texture prior is used to perform feature fusion on the semantic feature map and the texture feature map; the decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map;
[0009] The feature fusion network for semantic guidance and texture prior includes a semantic guidance sub-network and a texture prior sub-network;
[0010] The semantic guidance sub-network includes: a plurality of sequentially connected semantic guidance enhancement modules; the input end of each semantic guidance enhancement module is also connected to the output ends of the texture branch and the semantic branch in the dual-branch feature extraction network for semantics and texture; the semantic guidance enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information;
[0011] The texture prior sub-network includes: a plurality of boundary guidance aggregation modules; the input end of each boundary guidance aggregation module is connected to the output end of the last semantic guidance enhancement module and the output end of the semantic branch; the output end of each boundary guidance aggregation module is connected to the input end of the decoder; the boundary guidance aggregation module is used to provide edge texture prior guidance for the output of the semantic branch and segment the edge contour of the defect target.
[0012] In a second aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned dual-branch surface defect segmentation method for semantic guidance and texture prior is implemented.
[0013] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0014] The present application provides a dual-branch surface defect segmentation method and medium based on semantic guidance and texture prior. By using a dual-branch feature extraction network for semantics and texture, the feature extraction ability of the defect segmentation model for various complex defects is improved. The feature fusion network based on semantic guidance and texture prior includes a semantic guidance sub-network and a texture prior sub-network. The semantic guidance sub-network includes multiple sequentially connected semantic guidance enhancement modules. The texture prior sub-network includes multiple boundary guidance aggregation modules. The semantic guidance enhancement module is used to guide the texture branch to mine edge texture information by using the semantic information in the semantic branch. The boundary guidance aggregation module is used to provide edge texture prior guidance for the output of the semantic branch and segment the edge contour of the defect target. By using the feature fusion network based on semantic guidance and texture prior, the defect segmentation model can better segment the edge contour of the defect target, so as to improve the expression ability of the defect segmentation model and the ability to extract the unobvious defect features in the case of weak texture, and finally effectively improve the segmentation accuracy of the defect segmentation model for surface defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is an application environment diagram of a dual-branch surface defect segmentation method based on semantic guidance and texture prior in an embodiment of the present application;
[0017] Figure 2 It is a schematic flowchart of a dual-branch surface defect segmentation method based on semantic guidance and texture prior provided in an embodiment of the present application;
[0018] Figure 3 It is a schematic diagram of an image in an industrial surface defect dataset provided in an embodiment of the present application;
[0019] Figure 4 It is a schematic diagram of the structure of a defect segmentation model provided in an embodiment of the present application;
[0020] Figure 5 It is a schematic diagram of the structure of a semantic branch provided in an embodiment of the present application;
[0021] Figure 6 It is a schematic diagram of the structure of a self-attention-based semantic module provided in an embodiment of the present application;
[0022] Figure 7 It is a schematic diagram of the structure of an edge-aware enhancement module provided in an embodiment of the present application;
[0023] Figure 8 Structural schematic diagram of the detail enhancement module provided by an embodiment of the present application;
[0024] Figure 9 Structural schematic diagram of the semantic guidance enhancement module provided by an embodiment of the present application;
[0025] Figure 10 Structural schematic diagram of the boundary guidance aggregation module provided by an embodiment of the present application;
[0026] Figure 11 Industrial surface defect target segmentation result diagram provided by an embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0029] The dual-branch surface defect segmentation method with semantic guidance and texture prior provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the image of the object to be detected to the server. After the server receives the image of the object to be detected, the server inputs the image of the object to be detected into the defect segmentation model to obtain the surface defect segmentation result of the object to be detected; the defect segmentation model includes a dual-branch feature extraction network for semantics and texture, a feature fusion network for semantic guidance and texture prior, and a decoder connected in sequence; the dual-branch feature extraction network for semantics and texture is used to extract semantic information, defect texture, and edge features in the image of the object to be detected to obtain a semantic feature map and a texture feature map; the feature fusion network for semantic guidance and texture prior is used to perform feature fusion on the semantic feature map and the texture feature map; the decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map. The server can feedback the obtained surface defect segmentation result of the object to be detected to the terminal. In addition, in some embodiments, the dual-branch surface defect segmentation method for semantic guidance and texture prior can also be implemented separately by the server or the terminal. For example, the terminal can directly perform dual-branch surface defect segmentation based on semantic guidance and texture prior on the image of the object to be detected, or the server can obtain the image of the object to be detected from the data storage system and perform dual-branch surface defect segmentation based on semantic guidance and texture prior.
[0030] Among them, the terminal can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablets, Internet of Things devices, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0031] In an exemplary embodiment, as Figure 2 shown, a dual-branch surface defect segmentation method for semantic guidance and texture prior is provided. This method is executed by a computer device, and can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server in it as an example for illustration, it includes the following steps 101 to step 102.
[0032] Step 101, obtain the image of the object to be detected. For example, obtain an image from an industrial surface defect dataset, as Figure 3 shown.
[0033] Step 102, input the image of the object to be detected into the defect segmentation model to obtain the surface defect segmentation result of the object to be detected; as Figure 4As shown, the defect segmentation model includes a dual-branch feature extraction network for semantics and texture, a feature fusion network for semantic guidance and texture prior, and a decoder, which are connected in sequence. The dual-branch feature extraction network for semantics and texture is used to extract semantic information, defect texture, and edge features from the image of the object to be detected, and obtain a semantic feature map and a texture feature map. The feature fusion network for semantic guidance and texture prior is used to fuse the semantic feature map and the texture feature map. The decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map.
[0034] As Figure 4 shown, the feature fusion network for semantic guidance and texture prior includes a semantic guidance sub-network and a texture prior sub-network.
[0035] The semantic guidance sub-network includes: a plurality of sequentially connected semantic guidance enhancement modules; the input end of each semantic guidance enhancement module is also connected to the output ends of the texture branch and the semantic branch in the dual-branch feature extraction network for semantics and texture; the semantic guidance enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information.
[0036] The texture prior sub-network includes: a plurality of boundary guidance aggregation modules; the input end of each boundary guidance aggregation module is connected to the output end of the last semantic guidance enhancement module and the output end of the semantic branch; the output end of each boundary guidance aggregation module is connected to the input end of the decoder; the boundary guidance aggregation module is used to provide edge texture prior guidance for the output of the semantic branch and segment the edge contour of the defect target.
[0037] By implementing the above steps 101 to 102, aiming at the problems that the existing methods have poor segmentation and low segmentation accuracy due to the complex surface defect shape, large size difference, and weak texture, the present application uses the dual-branch feature extraction network for semantics and texture and the feature fusion network for semantic guidance and texture prior to improve the ability of the defect segmentation model to extract various complex defects and the non-obvious defect features in the case of weak texture, and improve the segmentation accuracy and robustness of the defect segmentation model for surface defects.
[0038] In another exemplary embodiment of the present application, as Figure 4 shown, in step 102, the dual-branch feature extraction network for semantics and texture includes: a semantic branch and a texture branch.
[0039] The semantic branch adopts a semantic branch based on a self-attention module. The advantage of the self-attention module is that it can better capture long-range dependencies, achieve a larger receptive field, and build a global context relationship in the network, so as to obtain more accurate and richer high-level semantic information. The output of the self-attention module is shown in formulas (1) and (2):
[0040] (1);
[0041] (2);
[0042] Wherein, Q, K, and V represent matrices obtained by linearly transforming the input matrix f; 、 represent learnable parameter matrices; is the matrix and the matrix is the number of matrix columns, preventing the inner product values of each row of the matrix and the matrix from being too large, which affects semantic relevance; represents the transpose matrix of the matrix ; Softmax represents a non-linear activation function; LayerNorm represents a linear normalization layer; SAEN represents a self-attention module; represents the intermediate calculation amount.
[0043] Such as Figure 5 shown, the semantic branch includes a block encoding module and multiple self-attention-based semantic modules connected in sequence; the image feature resolutions of the first self-attention-based semantic module to the last self-attention-based semantic module decrease in sequence; the first self-attention-based semantic module is the self-attention-based semantic module connected to the block encoding module; the self-attention-based semantic module is used to extract semantic information in the input image. The output end of each self-attention-based semantic module is also connected to the input end of the feature fusion network of semantic guidance and texture prior. Figure 5 The H and W shown in
[0044] Figure 5 represent the height and width of the image. Figure 6
[0045] Figure 5 Such as Figure 5 shown, the feature extraction process of the semantic branch is as follows: the input surface defect image (i.e., the image of the object to be detected) first passes through the block encoding module to block and encode the image into a vector representation, and the calculation formula is as shown in formula (3):
[0046] (3);
[0047] In Equation (3), PatchEmbed represents the patch encoding module; X represents the input surface defect image; represents the vector representation result after patching and encoding.
[0048] Then, it is gradually input into the semantic modules of each layer of the semantic branch for feature extraction. The resolution of the image features also gradually decreases from 1 / 4 to 1 / 32 layer by layer. Therefore, the semantic branch can output multi-level features with rich context representations. Among them, the feed-forward neural network consists of two fully connected layers and the GELU activation function. Therefore, the output of the semantic modules of each layer of the semantic branch is shown in Equation (4):
[0049] (4);
[0050] In Equation (4), SAEN represents the self-attention module; represents the learnable parameter; and represents the bias term; and respectively represent two fully connected layers; GELU represents the non-linear activation function; LayerNorm represents the linear normalization layer; Conv represents downsampling through convolution operations, and respectively represent the outputs of the semantic modules of the nth layer and the (n + 1)th layer of the semantic branch; when n = 0, represents the feature map of the input surface defect image after passing through the patch encoding module; represents the output result of the feed-forward neural network.
[0051] As Figure 4 shown, the texture branch adopts a texture branch based on an edge-aware enhancement module. The texture branch includes a plurality of edge-aware enhancement modules connected in sequence (corresponding to Figure 4 EPEM in ); the output end of the last edge-aware enhancement module is connected to the input end of the feature fusion network of semantic guidance and texture prior; the edge-aware enhancement module is used to extract the defect texture and edge features in the input image. Figure 4 shows an example. In the texture branch, 3 layers of edge-aware enhancement modules are adopted, and the resolutions of the image features are 1 / 4, 1 / 8, and 1 / 8 in sequence. In order to preserve more detailed texture information in the defect, the resolution of the downsampled feature map of the texture branch is at least 1 / 8 of the input image resolution, and no further downsampling operation will be performed subsequently. The output of each layer of the edge-aware enhancement module of the texture branch is shown in Equation (5):
[0052] (5);
[0053] In Equation (5), and respectively represent the outputs of the h-th layer and the (h + 1)-th layer of the texture branch; when h = 0, represents the input surface defect image; MaxPooling represents downsampling through the max pooling operation; EPEM represents the edge perception enhancement module.
[0054] As Figure 7 shown, the edge perception enhancement module includes: a first convolutional layer (corresponding to Figure 7 the first 3×3 convolution from the input direction in Figure 7 ), a first normalization and activation layer, a second convolutional layer (corresponding to Figure 7 the 1×1 convolution in
[0055] ), a first activation layer, a detail enhancement convolutional layer, a second activation layer, a first addition layer, a third convolutional layer (corresponding to
[0056] the second 3×3 convolution from the input direction in
[0057] (6);
[0058] In formula (6), represents the feature map input to the edge perception enhancement module; represents the convolution operation with a 3×3 convolution kernel; BachNorm represents the batch normalization operation; Relu represents the non-linear activation function; represents the convolution operation with a 1×1 convolution kernel; DEConv represents the detail enhancement convolution; EPEM represents the edge perception enhancement module. , , represent intermediate computational quantities.
[0059] In the texture branch, the edge perception enhancement module is based on a convolutional neural network, which can effectively capture local features such as edges and textures, and combines the detail enhancement convolution, which can improve the representation and generalization ability of ordinary convolutions, and further improve the model's ability to extract defect texture and edge features. The detail enhancement convolution draws on the idea of differential convolution, and deploys four differential convolutions (central differential convolution, corner differential convolution, horizontal differential convolution, vertical differential convolution) and an ordinary convolution in parallel. Among them, the four different differential convolutions are used to enhance the gradient-level information, and the ordinary convolution is used to obtain the intensity-level information. Using the reparameterization technique, the detail enhancement convolution is equivalently converted into a normal convolution without additional parameters and computational costs.
[0060] AsFigure 8 As shown in the figure, the detail enhancement convolutional layer includes a third addition layer and a fourth convolutional layer, a central difference convolutional layer, a corner difference convolutional layer, a horizontal difference convolutional layer, and a vertical difference convolutional layer that are connected in parallel.
[0061] The input ends of the fourth convolutional layer, the central difference convolutional layer, the corner difference convolutional layer, the horizontal difference convolutional layer, and the vertical difference convolutional layer are all connected to the output end of the first activation layer. The output ends of the fourth convolutional layer, the central difference convolutional layer, the corner difference convolutional layer, the horizontal difference convolutional layer, and the vertical difference convolutional layer are all connected to the input end of the third addition layer. The output end of the third addition layer is connected to the input end of the second activation layer.
[0062] The calculation formula of the detail enhancement convolutional module is shown in Equation (7):
[0063] (7);
[0064] In Equation (7), DEConv represents the detail enhancement convolution operation, x represents the input feature map, respectively represent the kernels of the central difference convolution, the corner difference convolution, the horizontal difference convolution, the vertical difference convolution, and the ordinary convolution, represents the convolution operation, represents the transformed kernel, which combines the parallel convolutions together.
[0065] In the dual-branch feature extraction network of semantics and texture, the semantic branch based on the self-attention module can construct the global context relationship of the network, so as to obtain more accurate and rich high-level semantic information. The texture branch based on the edge-aware enhancement module can effectively capture the local features of defects and improve the network's ability to extract defect texture and edge features. The semantic feature map and the texture feature map obtained by the dual-branch feature extraction network of semantics and texture are fused bidirectionally. The semantic branch and the texture branch complement each other, improving the feature extraction ability of the feature extraction network for various complex defects.
[0066] In another exemplary embodiment of the present application, in step 102, the feature fusion network of semantic guidance and texture prior includes a semantic guidance sub-network (corresponding to Figure 4 the semantic guidance part in Figure 4 and a texture prior sub-network (corresponding to
[0067] such as Figure 4 shown, the purpose of the semantic guidance sub-network is to improve the network's ability to extract the unobvious defect features in the case of weak texture, so as to better segment the edge contour of the defect target. The semantic guidance sub-network includes: a plurality of sequentially connected semantic guidance enhancement modules (corresponding to Figure 4SGEM in the texture branch); the input end of the first semantic-guided enhancement module is respectively connected to the output end of the last edge-aware enhancement module in the texture branch and the output end of the N-M+1th self-attention-based semantic module in the semantic branch; the input end of the mth semantic-guided enhancement module is also connected to the output end of the N-M+mth self-attention-based semantic module; N is the number of self-attention-based semantic modules; M is the number of semantic-guided enhancement modules; m>1; N≥M. The semantic-guided enhancement module is used to use the rich and accurate semantic information in the semantic branch to guide the output of the texture branch, mine accurate edge texture information, and suppress the noise of low-level features in the texture branch. Figure 4 An example is shown in FIG. 1 , in which two semantically guided enhancement modules are used in the semantically guided subnetwork.
[0068] like Figure 9 As shown, the semantic guidance enhancement module includes: a first convolution and normalization layer (corresponding to Figure 9 In the example, the 3×3 convolution + normalization layer corresponding to the texture feature map), the second convolution and normalization layer (corresponding to Figure 9 In the figure, there are 3×3 convolutions + normalization corresponding to the semantic feature map), upsampling layer, element-by-element multiplication layer, third activation layer (using S-type activation function), first matrix multiplication layer, second matrix multiplication layer and fourth addition layer.
[0069] The output end of the first convolution and normalization layer is connected to the input end of the element-by-element multiplication layer and the first matrix multiplication layer respectively; the output end of the element-by-element multiplication layer is connected to the input end of the third activation layer, and the output end of the third activation layer is connected to the input end of the first matrix multiplication layer and the input end of the second matrix multiplication layer respectively.
[0070] The output end of the second convolution and normalization layer is connected to the input end of the upsampling layer, and the output end of the upsampling layer is connected to the input ends of the element-by-element multiplication layer and the second matrix multiplication layer respectively.
[0071] The output end of the first matrix multiplication and the output end of the second matrix multiplication are connected to the input end of the fourth addition layer, and the output of the fourth addition layer is the output of the semantic guidance enhancement module to which it belongs.
[0072] Combination Figure 4 The invention further comprises a plurality of semantically guided enhancement modules, wherein the input end of the first convolution and normalization layer in the first semantically guided enhancement module is connected to the output end of the last edge-aware enhancement module in the texture branch.
[0073] The input end of the second convolution and normalization layer in the first semantic guided enhancement module is connected to the output end of the N-M+1th self-attention based semantic module in the semantic branch.
[0074] The input end of the first convolution and normalization layer in the m-th semantic guidance enhancement module is connected to the output end of the (m - 1)-th semantic guidance enhancement module.
[0075] The input end of the second convolution and normalization layer in the m-th semantic guidance enhancement module is connected to the output end of the (N - M + m)-th self-attention-based semantic module.
[0076] The semantic guidance enhancement module draws on the dynamic selection and fusion mechanism and combines the pixel attention mechanism. By calculating the similarity between two pixel points, the network can autonomously select useful texture information and semantic information and fuse them together, ensuring that more emphasis is placed on retaining and utilizing key semantic information during the fusion process, thereby achieving the effect of guiding texture information with semantic information and preventing texture information from being suppressed by excessive background or global context information. The output of the semantic guidance enhancement module is shown in Equation (8):
[0077] (8);
[0078] In Equation (8), represents the texture feature map input to the semantic guidance enhancement module, represents the semantic feature map input to the semantic guidance enhancement module, represents a convolution operation with a 3×3 convolution kernel, BachNorm represents batch normalization operation, up represents bilinear upsampling operation, represents an element-wise product operation, Sigmoid represents the S-shaped activation function, which is used to map the output to the range [0, 1] to generate the weight feature map , represents a matrix product operation, and SGEM represents the semantic guidance enhancement module.
[0079] The texture prior sub-network includes: multiple boundary guidance aggregation modules (corresponding to Figure 4 BGAM in); the input end of the n-th boundary guidance aggregation module is respectively connected to the output end of the last semantic guidance enhancement module and the output end of the (N - n + 1)-th self-attention-based semantic module; the output end of each boundary guidance aggregation module is connected to the input end of the decoder; the boundary guidance aggregation module is used to provide valuable edge texture prior guidance for the outputs of each layer of the semantic branch, enabling the network to better segment the edge contour of the defective target and improving the network's expression ability and the ability to extract non-obvious defect features in the case of weak textures. Figure 4 shows an example in which 4 boundary guidance aggregation modules are used in the texture prior sub-network.
[0080] As Figure 10As shown in the figure, the boundary-guided aggregation module includes: a fourth activation layer (using the sigmoid activation function), a third matrix multiplication layer, a fifth addition layer, a sixth addition layer, a fifth convolution layer (3×3 convolution), a global average pooling layer, a sixth convolution layer (one-dimensional convolution), a fifth activation layer (using the sigmoid activation function), a fourth matrix multiplication layer, and a seventh convolution layer (1×1 convolution) connected in sequence.
[0081] The fourth activation layer is connected to the output end of the last semantic-guided enhancement module; the input end of the sixth addition layer is also connected to the input end of the fourth activation layer; the input end of the fourth matrix multiplication layer is also connected to the input end of the global average pooling layer; the output of the seventh convolution layer is the output of the boundary-guided aggregation module to which it belongs.
[0082] The input ends of the third matrix multiplication layer and the fifth addition layer in the nth boundary-guided aggregation module are also connected to the output end of the (N - n + 1)th self-attention-based semantic module.
[0083] The boundary-guided aggregation module can act on the representation learning of semantic information with texture information, provide valuable edge texture prior guidance for the output of each semantic module in the semantic branch, and enhance the feature representation of weak texture defect semantics by fusing features and combining the channel attention mechanism. The output of the boundary-guided aggregation module is shown in Equation (9):
[0084] (9);
[0085] In Equation (9), represents the texture feature map input to the boundary-guided aggregation module; represents the semantic feature map input to the boundary-guided aggregation module; Sigmoid represents the sigmoid activation function, which is used to map the output to the range [0, 1] to generate the weight feature map W or ; represents the convolution operation with a convolution kernel of 3×3; represents the matrix product operation; gap represents the global average pooling operation; represents the one-dimensional convolution with a kernel size of k, where k is the odd number closest to , B represents the number of channels in the fifth convolution layer; p represents the result of the 3×3 convolution operation; represents the convolution operation with a convolution kernel of 1×1, and BGAM represents the boundary-guided aggregation module.
[0086] Taking the design of 4 self-attention-based semantic modules, 2 semantic-guided enhancement modules, and 4 boundary-guided aggregation modules as an example, the specific process of the semantic guidance and texture prior feature fusion network is described as follows:
[0087] First, through two semantic guidance enhancement modules, the semantic information in the feature maps output by the third-layer semantic module and the fourth-layer semantic module in the semantic branch is used to recursively guide the output of the texture branch, mining accurate edge texture information and suppressing the noise of the low-level features in the texture branch. The corresponding relationship is shown in Equation (10):
[0088] (10);
[0089] In Equation (10), and respectively represent the feature map outputs after passing through the m-th and the (m - 1)-th semantic guidance enhancement modules; SGEM represents the semantic guidance enhancement module; when m = 1, represents the final feature map output by the texture branch; represents the feature map in the semantic branch with a resolution of .
[0090] Secondly, the texture information guided by the two semantic guidance enhancement modules is size-registered with the feature maps output by each layer of the semantic module in the semantic branch, such as the upsampling and downsampling operations in Figure 4 , and then input into four boundary guidance aggregation modules to provide valuable edge texture prior guidance for the feature maps output by each layer of the semantic module in the semantic branch, aiming to inject edge clues related to the texture into the representation learning of the semantic branch to enhance the feature representation of weak texture defects. The corresponding relationship is shown in Equation (11):
[0091] (11);
[0092] In Equation (11), represents the texture feature map guided by the two semantic guidance enhancement modules; represents the feature map in the semantic branch with a resolution of ; resize represents the registration operation of according to the size of the feature map of ; represents the feature map output by the boundary guidance aggregation module after the output with a resolution of in the semantic branch; BGAM represents the boundary guidance aggregation module. n refers to the n-th semantic module.
[0093] In another exemplary embodiment of the present application, in step 102, the decoder is a decoder based on a feature pyramid structure. The feature pyramid structure is a progressive upsampling structure, and combines multi-scale features through lateral connections, enhancing the model's perception ability of features at different scales, increasing the robustness and adaptability of the model, and being able to perform more fine-grained feature extraction, ultimately improving the segmentation accuracy of the model for surface defects. AsFigure 4 As shown, the decoder includes a plurality of upsampling modules connected in sequence; a splicing module is connected after each upsampling module; the input end of the i-th splicing module is further connected to the output end of the (i + 1)-th boundary-guided aggregation module in the texture prior sub-network; the last upsampling module outputs the surface defect segmentation result of the object to be detected, such as Figure 11 shown in Figure 11 In , the white area is the surface defect area. Figure 4 An example is shown in , where 3 upsampling modules and 3 splicing modules are used.
[0094] The feature map after the feature fusion network with semantic guidance and texture prior is input into the decoder based on the feature pyramid structure, and the resolution of the feature map is sequentially restored from 1 / 32 upwards to the original image resolution to output the final surface defect segmentation map, and its corresponding relationship is shown in Equation (12):
[0095] (12);
[0096] In Equation (12), and respectively represent the feature maps with resolutions of and after passing through the boundary-guided aggregation module; up represents the bilinear upsampling operation; Concat is the connection operation; represents the convolution operation with a 3×3 convolution kernel; represents the feature map with a resolution of after passing through the boundary-guided aggregation module; represents the segmentation map finally output by the model.
[0097] In another exemplary embodiment of the present application, in step 2, before inputting the image of the object to be detected into the defect segmentation model, the defect segmentation model needs to be a trained model. In the present application, the cross-entropy loss function is used to calculate the cross-entropy loss between the surface defect segmentation result obtained by the model and the true annotation information of the surface defect, and the cross-entropy loss calculation formula is shown in Equation (13):
[0098] (13);
[0099] In Equation (13), G represents the total number of pixels in the image, C represents the total number of categories (or labels), represents a binary indicator variable, which is 1 if the g-th pixel belongs to the j-th category, otherwise it is 0, represents the probability that the g-th pixel predicted by the model belongs to the j-th category.
[0100] Calculate the cross - entropy loss based on the target segmentation result obtained from the defect segmentation model and the target true annotation information, and adjust the model parameters through backpropagation based on the cross - entropy loss to obtain the trained defect segmentation model.
[0101] In this application, a dual - branch feature extraction network of semantics and texture is used to extract the features of various complex surface defects. The feature fusion network of semantic guidance and texture prior enables the model to better segment the edge contours of defect targets, so as to improve the expression ability of the model and the ability to extract the features of non - obvious defects in the case of weak texture, and finally effectively improve the segmentation accuracy and robustness of the model for surface defects.
[0102] This application also provides an application scenario, which applies the above - mentioned dual - branch surface defect segmentation method based on semantic guidance and texture prior. Specifically: The dual - branch surface defect segmentation method based on semantic guidance and texture prior provided in this embodiment can be applied to the surface defect segmentation scenario. This scenario includes an image acquisition link and a surface defect segmentation link; the image acquisition link is used to acquire the image of the industrial object to be detected; the surface defect segmentation link is used to obtain the image of the industrial object to be detected, input the image into the defect segmentation model, and obtain the surface defect segmentation result of the industrial object to be detected. The dual - branch surface defect segmentation method based on semantic guidance and texture prior provided in this embodiment belongs to the surface defect segmentation link.
[0103] In an exemplary embodiment, a computer - readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the above - mentioned dual - branch surface defect segmentation method based on semantic guidance and texture prior.
[0104] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0105] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0106] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A dual-branch surface defect segmentation method based on semantic guidance and texture prior, characterized in that Including: Obtain the image of the object to be detected; Input the image of the object to be detected into the defect segmentation model to obtain the surface defect segmentation result of the object to be detected; The defect segmentation model includes a dual-branch feature extraction network for semantics and texture, a feature fusion network for semantic guidance and texture prior, and a decoder connected in sequence; the dual-branch feature extraction network for semantics and texture is used to extract semantic information, defect texture and edge features in the image of the object to be detected, and obtain a semantic feature map and a texture feature map; the feature fusion network for semantic guidance and texture prior is used to perform feature fusion on the semantic feature map and the texture feature map; the decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map; The feature fusion network for semantic guidance and texture prior includes a semantic guidance sub-network and a texture prior sub-network; The semantic guidance sub-network includes: a plurality of sequentially connected semantic guidance enhancement modules; the input end of each semantic guidance enhancement module is also connected to the output ends of the texture branch and the semantic branch in the dual-branch feature extraction network for semantics and texture; the semantic guidance enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information; The texture prior sub-network includes: a plurality of boundary guidance aggregation modules; the input end of each boundary guidance aggregation module is connected to the output end of the last semantic guidance enhancement module and the output end of the semantic branch; the output end of each boundary guidance aggregation module is connected to the input end of the decoder; the boundary guidance aggregation module is used to provide edge texture prior guidance for the output of the semantic branch and segment the edge contour of the defect target.
2. The semantic guidance and texture prior-based dual-branch surface defect segmentation method according to claim 1, wherein The dual-branch feature extraction network for semantics and texture includes: a semantic branch and a texture branch; The semantic branch includes a block coding module and a plurality of self-attention-based semantic modules connected in sequence; the image feature resolution of the first self-attention-based semantic module to the last self-attention-based semantic module decreases in sequence; the first self-attention-based semantic module is the semantic module connected to the block coding module; the self-attention-based semantic module is used to extract semantic information in the input image; The output end of each self-attention-based semantic module is also connected to the input end of the feature fusion network for semantic guidance and texture prior; The texture branch includes a plurality of sequentially connected edge perception enhancement modules; the output end of the last edge perception enhancement module is connected to the input end of the feature fusion network for semantic guidance and texture prior; the edge perception enhancement module is used to extract defect texture and edge features in the input image.
3. The dual-branch surface defect segmentation method based on semantic guidance and texture prior according to claim 2, characterized in that The self-attention-based semantic module includes: a self-attention module, a feed-forward neural network, a linear normalization layer, and a downsampling layer connected in sequence; the input end of the linear normalization layer is also connected to the input end of the self-attention module.
4. The dual-branch surface defect segmentation method based on semantic guidance and texture prior according to claim 2, wherein The edge perception enhancement module includes: a first convolutional layer, a first normalization and activation layer, a second convolutional layer, a first activation layer, a detail enhancement convolutional layer, a second activation layer, a first addition layer, a third convolutional layer, a second normalization and activation layer, and a second addition layer connected in sequence; The input end of the first addition layer is also connected to the output end of the first activation layer; the input end of the second addition layer is also connected to the input end of the first convolutional layer.
5. The semantic guidance and texture prior-based dual-branch surface defect segmentation method according to claim 4, characterized in that The detail enhancement convolution layer includes a third addition layer and a fourth convolution layer, a center difference convolution layer, an angle difference convolution layer, a horizontal difference convolution layer, and a vertical difference convolution layer connected in parallel; The input ends of the fourth convolution layer, the center difference convolution layer, the angle difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer are all connected to the output end of the first activation layer; The output ends of the fourth convolution layer, the center difference convolution layer, the angle difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer are all connected to the input end of the third addition layer; The output terminal of the third addition layer is connected to the input terminal of the second activation layer.
6. The semantic-guided and texture prior-based double-branch surface defect segmentation method according to claim 2, wherein The semantic guidance and texture prior feature fusion network includes a semantic guidance sub-network and a texture prior sub-network; The semantic guidance sub-network includes: a plurality of semantic guidance enhancement modules connected in sequence; the input end of the first semantic guidance enhancement module is respectively connected to the output end of the last edge perception enhancement module in the texture branch and the output end of the N-M+1 th semantic module based on self-attention in the semantic branch; the input end of the m th semantic guidance enhancement module is also connected to the output end of the N-M+m th semantic module based on self-attention; N is the number of semantic modules based on self-attention; M is the number of semantic guidance enhancement modules; m>1; N≥M; the semantic guidance enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information; The texture prior subnetwork includes: multiple boundary-guided aggregation modules; the input end of the nth boundary-guided aggregation module is respectively connected to the output end of the last semantic-guided enhancement module and the output end of the N-n+1th self-attention-based semantic module; the output end of each boundary-guided aggregation module is connected to the input end of the decoder; the boundary-guided aggregation module is used to provide edge texture prior guidance for the output of each layer of the semantic branch, and segment the edge contour of the defect target.
7. The dual-branch surface defect segmentation method based on semantic guidance and texture prior according to claim 6, wherein The semantic guidance enhancement module includes: a first convolution and normalization layer, a second convolution and normalization layer, an upsampling layer, an element-by-element multiplication layer, a third activation layer, a first matrix multiplication layer, a second matrix multiplication layer and a fourth addition layer; The output end of the first convolution and normalization layer is connected to the input end of the element-by-element multiplication layer and the first matrix multiplication layer respectively; the output end of the element-by-element multiplication layer is connected to the input end of the third activation layer, and the output end of the third activation layer is connected to the input end of the first matrix multiplication layer and the input end of the second matrix multiplication layer respectively; The output end of the second convolution and normalization layer is connected to the input end of the upsampling layer, and the output end of the upsampling layer is connected to the input ends of the element-by-element multiplication layer and the second matrix multiplication layer respectively; The output end of the first matrix multiplication and the output end of the second matrix multiplication are connected to the input end of the fourth addition layer, and the output of the fourth addition layer is the output of the semantic guidance enhancement module to which it belongs; The input end of the first convolution and normalization layer in the first semantic-guided enhancement module is connected to the output end of the last edge-aware enhancement module in the texture branch; The input end of the second convolution and normalization layer in the first semantic guided enhancement module is connected to the output end of the N-M+1 th self-attention based semantic module in the semantic branch; The input end of the first convolution and normalization layer in the mth semantic guidance enhancement module is connected to the output end of the m-1th semantic guidance enhancement module; The input of the second convolution and normalization layer in the mth semantic-guided enhancement module is connected to the output of the N-M+mth self-attention-based semantic module.
8. The dual-branch surface defect segmentation method based on semantic guidance and texture prior according to claim 6, characterized in that The boundary guided aggregation module includes: a fourth activation layer, a third matrix multiplication layer, a fifth addition layer, a sixth addition layer, a fifth convolution layer, a global average pooling layer, a sixth convolution layer, a fifth activation layer, a fourth matrix multiplication layer and a seventh convolution layer connected in sequence; The fourth activation layer is connected to the output of the last semantic guidance enhancement module; the input of the sixth addition layer is also connected to the input of the fourth activation layer; the input of the fourth matrix multiplication layer is also connected to the input of the global average pooling layer; the output of the seventh convolutional layer is the output of the boundary guidance aggregation module to which it belongs; The input of the third matrix multiplication layer and the input of the fifth addition layer in the nth boundary-guided aggregation module are also connected to the output of the N-n+1th self-attention-based semantic module.
9. The dual-branch surface defect segmentation method based on semantic guidance and texture prior according to claim 6, characterized in that The decoder is a decoder based on a feature pyramid structure; the decoder includes a plurality of upsampling modules connected in sequence; a splicing module is also provided between two adjacent upsampling modules; the input end of the i-th splicing module is also connected to the output end of the i+1-th boundary-guided aggregation module in the texture prior subnetwork; and the last upsampling module outputs the surface defect segmentation result of the object to be detected.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the semantically guided and texture-prior dual-branch surface defect segmentation method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Connected double-attention multi-scale fusion semantic segmentation network
CN116630626A
Display screen defect detection method and system based on enhanced feature extraction network
CN118570212A