Semantic guidance and texture prior double-branch surface defect segmentation method and medium

Through the double-branch surface defect segmentation method of semantic guidance and texture priori, the problem of poor segmentation of complex defects and weak texture defects in the prior art is solved, and the segmentation accuracy and robustness are improved.

CN119991713AActive Publication Date: 2025-05-13NANCHANG HANGKONG UNIVERSITY

Patent Information

Application Number
CN202510480124.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing surface defect segmentation method is poor in the processing of defects with complex shapes, large size differences, and weak texture defects, and low accuracy.

Method used

The two-branch surface defect segmentation method of semantic guidance and texture prior is adopted, and the two-branch feature extraction network of semantic guidance and texture prior feature fusion network is improved to improve the defect segmentation model's extraction ability of complex defects and weak texture features.

Benefits of technology

The segmentation accuracy and robustness of the defect segmentation model for surface defects is improved, and the edge profiles of complex defects and weak texture defects can be better segmented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991713A_ABST
    Figure CN119991713A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic guidance and texture prior double-branch surface defect segmentation method and a medium, and relates to the field of surface defect segmentation, and the method comprises the steps: obtaining a to-be-detected object image; inputting the image of the object to be detected into the defect segmentation model to obtain a surface defect segmentation result of the object to be detected; the defect segmentation model comprises a semantic and texture double-branch feature extraction network, a semantic guidance and texture priori feature fusion network and a decoder which are connected in sequence; the double-branch feature extraction network is used for extracting semantic information, defect texture and edge features in the to-be-detected object image to obtain a semantic feature map and a texture feature map; the feature fusion network is used for performing feature fusion on the semantic feature map and the texture feature map; and the decoder is used for outputting a surface defect segmentation result of the object to be detected according to the fused feature map. According to the method, the problems of poor segmentation and low segmentation precision caused by complex surface defect shape, large size difference and weak texture are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of surface defect segmentation, and in particular to a semantically guided and texture prior dual-branch surface defect segmentation method and medium. Background Art

[0002] Surface defect segmentation technology plays a vital role in the field of industrial quality inspection and is one of the important technologies to ensure product quality and maintain stable production. In the traditional production process, the detection of product surface defects often relies on manual visual inspection, which is time-consuming and labor-intensive. More importantly, due to human subjectivity and interference from external factors such as the environment, the detection results often have certain errors and instability. In recent years, with the rapid development of deep learning technology, the surface defect segmentation method based on deep learning has significantly improved the accuracy and robustness of surface defect detection tasks, providing an efficient and feasible solution for industrial surface defect detection, and has been applied to many fields such as weld bubble detection, metal surface detection, and circuit board quality inspection. It not only proves the practicality and effectiveness of surface defect segmentation technology based on deep learning in actual industrial production, but also demonstrates the huge potential and broad application prospects of this technology.

[0003] Common types of surface defects in industrial scenarios are dents, stains, scratches, and cracks. The surface defect segmentation method based on deep learning is to classify these surface defect targets at the pixel level, so that the surface defect targets can be accurately segmented from the entire image, and more detailed detection results can be obtained, showing higher accuracy in surface defect detection tasks. Although the surface defect segmentation method based on deep learning has performed well in recent years, due to the influence of industrial imaging environment and noise, the segmentation of defects with complex shapes and large size differences, as well as the segmentation of weak texture defects with unclear defect features, still poses a huge challenge to most existing methods. Summary of the invention

[0004] The purpose of this application is to provide a semantically guided and texture prior dual-branch surface defect segmentation method and medium, which can solve the problems of poor segmentation and low segmentation accuracy caused by complex surface defect shapes, large size differences and weak textures.

[0005] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides a semantically guided and texture prior dual-branch surface defect segmentation method comprising: Acquire an image of the object to be detected; The image of the object to be detected is input into the defect segmentation model to obtain the surface defect segmentation result of the object to be detected; the defect segmentation model includes a semantic and texture dual-branch feature extraction network, a semantic guidance and texture prior feature fusion network and a decoder connected in sequence; the semantic and texture dual-branch feature extraction network is used to extract semantic information, defect texture and edge features in the image of the object to be detected, and obtain a semantic feature map and a texture feature map; the semantic guidance and texture prior feature fusion network is used to perform feature fusion on the semantic feature map and the texture feature map; the decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map; The semantic guidance and texture prior feature fusion network includes a semantic guidance sub-network and a texture prior sub-network; The semantic guidance sub-network comprises: a plurality of sequentially connected semantic guidance enhancement modules; the input end of each semantic guidance enhancement module is also connected to the output end of the texture branch and the semantic branch in the dual-branch feature extraction network of semantics and texture; the semantic guidance enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information; The texture prior subnetwork includes: multiple boundary-guided aggregation modules; the input end of each boundary-guided aggregation module is connected to the output end of the last semantic-guided enhancement module and the output end of the semantic branch; the output end of each boundary-guided aggregation module is connected to the input end of the decoder; the boundary-guided aggregation module is used to provide edge texture prior guidance for the output of the semantic branch and segment the edge contour of the defect target.

[0006] In a second aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned semantically guided and texture prior dual-branch surface defect segmentation method.

[0007] According to the specific embodiments provided in this application, this application discloses the following technical effects: The present application provides a semantically guided and texture prior dual-branch surface defect segmentation method and medium, which utilizes a semantically and texture dual-branch feature extraction network to improve the feature extraction capability of the defect segmentation model for various complex defects. The semantically guided and texture prior feature fusion network includes a semantically guided subnetwork and a texture prior subnetwork; the semantically guided subnetwork includes a plurality of sequentially connected semantically guided enhancement modules; the texture prior subnetwork includes a plurality of boundary guided aggregation modules; the semantically guided enhancement module is used to utilize the semantic information in the semantic branch to guide the texture branch to mine edge texture information; the boundary guided aggregation module is used to provide edge texture prior guidance for the output of the semantic branch to segment the edge contour of the defect target. The feature fusion network of semantically guided and texture prior is utilized to enable the defect segmentation model to better segment the edge contour of the defect target, so as to improve the expression capability of the defect segmentation model and the ability to extract inconspicuous defect features in weak texture conditions, and ultimately effectively improve the segmentation accuracy of the defect segmentation model for surface defects. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0009] Figure 1 This is an application environment diagram of a semantically guided and texture prior dual-branch surface defect segmentation method in one embodiment of the present application; Figure 2 A schematic flow chart of a semantically guided and texture prior dual-branch surface defect segmentation method provided in one embodiment of the present application; Figure 3 A schematic diagram of an image in an industrial surface defect dataset provided by an embodiment of the present application; Figure 4 A schematic diagram of the structure of a defect segmentation model provided in one embodiment of the present application; Figure 5 A schematic diagram of the structure of the semantic branch provided in one embodiment of the present application; Figure 6 A schematic diagram of the structure of a semantic module based on self-attention provided in one embodiment of the present application; Figure 7 A schematic diagram of the structure of an edge perception enhancement module provided in one embodiment of the present application; Figure 8 A schematic diagram of the structure of a detail enhancement module provided in one embodiment of the present application; Fig. 9A schematic diagram of the structure of a semantic guidance enhancement module provided in one embodiment of the present application; Fig.10 A schematic diagram of the structure of a boundary guidance aggregation module provided in an embodiment of the present application; Fig.11 This is a diagram of the industrial surface defect target segmentation results provided in one embodiment of the present application. DETAILED DESCRIPTION

[0010] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0011] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0012] The semantic guidance and texture prior dual-branch surface defect segmentation method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the image of the object to be detected to the server. After the server receives the image of the object to be detected, the server inputs the image of the object to be detected into the defect segmentation model to obtain the surface defect segmentation result of the object to be detected; the defect segmentation model includes a double-branch feature extraction network of semantics and texture, a feature fusion network of semantic guidance and texture prior, and a decoder connected in sequence; the double-branch feature extraction network of semantics and texture is used to extract semantic information, defect texture and edge features in the image of the object to be detected, and obtain a semantic feature map and a texture feature map; the feature fusion network of semantic guidance and texture prior is used to perform feature fusion on the semantic feature map and the texture feature map; the decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map. The server can feedback the obtained surface defect segmentation result of the object to be detected to the terminal. In addition, in some embodiments, the dual-branch surface defect segmentation method of semantic guidance and texture prior can also be implemented independently by a server or a terminal. For example, the terminal can directly perform dual-branch surface defect segmentation based on semantic guidance and texture prior on the image of the object to be inspected, or the server can obtain the image of the object to be inspected from the data storage system and perform dual-branch surface defect segmentation based on semantic guidance and texture prior.

[0013] The terminal may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.

[0014] In an exemplary embodiment, Figure 2 As shown, a semantically guided and texture prior dual-branch surface defect segmentation method is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server in is taken as an example to illustrate, including the following steps 101 to 102.

[0015] Step 101, obtain an image of the object to be detected. For example, obtain an image in an industrial surface defect dataset, such as Figure 3 shown.

[0016] Step 102, input the image of the object to be detected into the defect segmentation model to obtain the surface defect segmentation result of the object to be detected; Figure 4 As shown, the defect segmentation model includes a semantic and texture dual-branch feature extraction network, a semantic-guided and texture-prior feature fusion network and a decoder connected in sequence; the semantic and texture dual-branch feature extraction network is used to extract semantic information, defect texture and edge features in the image of the object to be detected, and obtain a semantic feature map and a texture feature map; the semantic-guided and texture-prior feature fusion network is used to perform feature fusion on the semantic feature map and the texture feature map; the decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map.

[0017] like Figure 4 As shown, the semantic-guided and texture-prior feature fusion network includes a semantic-guided subnetwork and a texture-prior subnetwork.

[0018] The semantic-guided subnetwork includes: a plurality of sequentially connected semantic-guided enhancement modules; the input end of each semantic-guided enhancement module is also connected to the output ends of the texture branch and the semantic branch in the dual-branch feature extraction network of semantics and texture; the semantic-guided enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information.

[0019] The texture prior subnetwork includes: multiple boundary-guided aggregation modules; the input end of each boundary-guided aggregation module is connected to the output end of the last semantic-guided enhancement module and the output end of the semantic branch; the output end of each boundary-guided aggregation module is connected to the input end of the decoder; the boundary-guided aggregation module is used to provide edge texture prior guidance for the output of the semantic branch and segment the edge contour of the defect target.

[0020] Implementing the above steps 101 to 102, in order to address the problem that surface defects have complex shapes, large size differences, and weak textures, which lead to poor segmentation and low segmentation accuracy in existing methods, the present application utilizes a dual-branch feature extraction network of semantics and texture and a feature fusion network of semantic guidance and texture prior to improve the defect segmentation model's ability to extract features of various complex defects and unclear defects in weak texture conditions, thereby improving the defect segmentation model's segmentation accuracy and robustness for surface defects.

[0021] In another exemplary embodiment of the present application, Figure 4 As shown, in step 102, the dual-branch feature extraction network of semantics and texture includes: a semantic branch and a texture branch.

[0022] The semantic branch adopts a semantic branch based on the self-attention module. The advantage of the self-attention module is that it can better capture long-range dependencies, achieve a larger receptive field, and build global contextual relationships in the network, thereby obtaining more accurate and richer high-level semantic information. The output of the self-attention module is shown in formula (1) and formula (2): (1); (2); In the formula, Q, K, V represent matrices obtained by linear transformation of the input matrix f; , represents the learnable parameter matrix; is a matrix and matrix The number of matrix columns prevents the matrix and matrix The inner product of each row is too large, which affects the semantic relevance; Representation Matrix The transposed matrix of ; Softmax represents a nonlinear activation function; LayerNorm represents a linear normalization layer; SAEN represents a self-attention module; Indicates the amount of intermediate calculation.

[0023] like Figure 5 As shown, the semantic branch includes a block encoding module and multiple self-attention-based semantic modules connected in sequence; the image feature resolution from the first self-attention-based semantic module to the last self-attention-based semantic module decreases in sequence; the first self-attention-based semantic module is a self-attention-based semantic module connected to the block encoding module; the self-attention-based semantic module is used to extract semantic information from the input image. The output end of each self-attention-based semantic module is also connected to the input end of the semantic guidance and texture prior feature fusion network. Figure 5H and W shown in represent the height and width of the image.

[0024] Figure 5 An example is shown, where four semantic modules based on self-attention are used in the semantic branch, and the image feature resolutions are 1 / 4, 1 / 8, 1 / 16, and 1 / 32, respectively. Figure 6 As shown, the self-attention-based semantic module of each layer includes: a self-attention module, a feedforward neural network, a linear normalization layer and a downsampling layer (downsampling operation is performed through convolution) connected in sequence; the input end of the linear normalization layer is also connected to the input end of the self-attention module.

[0025] like Figure 5 As shown in Figure 2, the feature extraction process of the semantic branch is as follows: the input surface defect image (i.e., the image of the object to be detected) is first divided into blocks and encoded into vector representation by the block encoding module. The calculation formula is shown in formula (3): (3); In formula (3), PatchEmbed represents the block encoding module; X represents the input surface defect image; Represents the vector representation result after block division and encoding.

[0026] Then it is gradually input into the semantic module of each layer of the semantic branch for feature extraction. The image feature resolution is also reduced from 1 / 4 to 1 / 32 layer by layer. Therefore, the semantic branch can output multi-level features with rich context representation. Among them, the feedforward neural network consists of two fully connected layers and GELU activation function. Therefore, the output of the semantic module of each layer of the semantic branch is shown in formula (4): (4); In formula (4), SAEN represents the self-attention module; represents a learnable parameter; and represents the bias term; and They represent two fully connected layers respectively; GELU represents a nonlinear activation function; LayerNorm represents a linear normalization layer; Conv represents downsampling through a convolution operation, and They represent the outputs of the nth layer and the n+1th layer semantic modules of the semantic branch respectively; when n=0, Represents the feature map of the input surface defect image after passing through the block encoding module; Represents the output of the feedforward neural network.

[0027] like Figure 4As shown, the texture branch adopts a texture branch based on an edge-aware enhancement module, and the texture branch includes a plurality of edge-aware enhancement modules connected in sequence (corresponding to Figure 4 The output of the last edge-aware enhancement module is connected to the input of the semantic-guided and texture-prior feature fusion network; the edge-aware enhancement module is used to extract defective texture and edge features in the input image. Figure 4 An example is shown in FIG. 3 , where a three-layer edge-aware enhancement module is used in the texture branch, and the image feature resolutions are 1 / 4, 1 / 8, and 1 / 8, respectively. In order to preserve more detailed texture information in the defect, the resolution of the feature map downsampled by the texture branch is at least 1 / 8 of the input image resolution, and no downsampling operation will be performed subsequently. The output of each layer of the edge-aware enhancement module of the texture branch is shown in formula (5): (5); In formula (5), and They represent the outputs of the hth layer and the h+1th layer of the texture branch respectively; when h=0, represents the input surface defect image; MaxPooling represents downsampling through the maximum pooling operation; EPEM represents the edge-aware enhancement module.

[0028] like Figure 7 As shown, the edge perception enhancement module includes: a first convolutional layer (corresponding to Figure 7 The first 3×3 convolution from the input direction), the first normalization and activation layer, the second convolution layer (corresponding to Figure 7 1×1 convolution in ), first activation layer, detail enhancement convolution layer, second activation layer, first addition layer, third convolution layer (corresponding to Figure 7 The second 3×3 convolution from the input direction), the second normalization and activation layer, and the second addition layer.

[0029] The input end of the first addition layer is also connected to the output end of the first activation layer; the input end of the second addition layer is also connected to the input end of the first convolutional layer.

[0030] The edge-aware enhancement module is shown in formula (6): (6); In formula (6), Feature map representing the input edge-aware enhancement module; Indicates a convolution operation with a convolution kernel of 3×3; BachNorm indicates a batch normalization operation; Relu indicates a nonlinear activation function; represents a convolution operation with a convolution kernel of 1×1; DEConv represents detail enhancement convolution; EPEM represents edge-aware enhancement module. , , Indicates the amount of intermediate calculation.

[0031] In the texture branch, the edge-aware enhancement module is based on a convolutional neural network, which can effectively capture local features such as edges and textures, and is combined with detail enhancement convolution, which can improve the representation and generalization capabilities of ordinary convolutions, and further improve the model's ability to extract defective textures and edge features. The detail enhancement convolution draws on the idea of ​​differential convolution and deploys four difference convolutions (center differential convolution, angle differential convolution, horizontal differential convolution, vertical differential convolution) and an ordinary convolution in parallel. Four different differential convolutions are used to enhance gradient-level information, and ordinary convolution is used to obtain intensity-level information. The reparameterization technique is used to equivalently convert the detail enhancement convolution into a normal convolution without additional parameters and computational cost.

[0032] like Figure 8 As shown, the detail enhancement convolution layer includes a third addition layer and a fourth convolution layer, a center difference convolution layer, an angle difference convolution layer, a horizontal difference convolution layer, and a vertical difference convolution layer connected in parallel.

[0033] The input ends of the fourth convolution layer, the center difference convolution layer, the angle difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer are all connected to the output end of the first activation layer. The output ends of the fourth convolution layer, the center difference convolution layer, the angle difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer are all connected to the input end of the third addition layer. The output end of the third addition layer is connected to the input end of the second activation layer.

[0034] The calculation formula of the detail enhancement convolution module is shown in formula (7): (7); In formula (7), DEConv represents the detail enhancement convolution operation, x represents the input feature map, They represent the kernels of center difference convolution, angle difference convolution, horizontal difference convolution, vertical difference convolution and ordinary convolution respectively. represents the convolution operation, represents the transformed kernel, which combines parallel convolutions together.

[0035] In the dual-branch feature extraction network of semantics and texture, the semantic branch based on the self-attention module can build the global contextual relationship of the network, thereby obtaining more accurate and rich high-level semantic information. The texture branch based on the edge-aware enhancement module can effectively capture the local features of defects and improve the network's ability to extract defect texture and edge features. The semantic feature map and texture feature map obtained by the dual-branch feature extraction network of semantics and texture are bidirectionally fused. The semantic branch and the texture branch complement each other, improving the feature extraction network's ability to extract features for various complex defects.

[0036] In another exemplary embodiment of the present application, in step 102, the semantic guidance and texture prior feature fusion network includes a semantic guidance sub-network (corresponding to Figure 4 The semantic guidance part in the text) and the texture prior sub-network (corresponding to Figure 4 The texture prior part in ).

[0037] like Figure 4 As shown in the figure, the purpose of the semantic guidance sub-network is to improve the network's ability to extract the features of subtle defects in the case of weak texture, so as to better segment the edge contours of defective targets. The semantic guidance sub-network includes: a plurality of sequentially connected semantic guidance enhancement modules (corresponding to Figure 4 SGEM in the texture branch); the input end of the first semantic-guided enhancement module is respectively connected to the output end of the last edge-aware enhancement module in the texture branch and the output end of the N-M+1th self-attention-based semantic module in the semantic branch; the input end of the mth semantic-guided enhancement module is also connected to the output end of the N-M+mth self-attention-based semantic module; N is the number of self-attention-based semantic modules; M is the number of semantic-guided enhancement modules; m>1; N≥M. The semantic-guided enhancement module is used to use the rich and accurate semantic information in the semantic branch to guide the output of the texture branch, mine accurate edge texture information, and suppress the noise of low-level features in the texture branch. Figure 4 An example is shown in FIG. 1 , in which two semantically guided enhancement modules are used in the semantically guided subnetwork.

[0038] like Fig. 9 As shown, the semantic guidance enhancement module includes: a first convolution and normalization layer (corresponding to Fig. 9 In the example, the 3×3 convolution + normalization layer corresponding to the texture feature map), the second convolution and normalization layer (corresponding to Fig. 9 In the figure, there are 3×3 convolutions + normalization corresponding to the semantic feature map), upsampling layer, element-by-element multiplication layer, third activation layer (using S-type activation function), first matrix multiplication layer, second matrix multiplication layer and fourth addition layer.

[0039] The output end of the first convolution and normalization layer is connected to the input end of the element-by-element multiplication layer and the first matrix multiplication layer respectively; the output end of the element-by-element multiplication layer is connected to the input end of the third activation layer, and the output end of the third activation layer is connected to the input end of the first matrix multiplication layer and the input end of the second matrix multiplication layer respectively.

[0040] The output end of the second convolution and normalization layer is connected to the input end of the upsampling layer, and the output end of the upsampling layer is connected to the input ends of the element-by-element multiplication layer and the second matrix multiplication layer respectively.

[0041] The output end of the first matrix multiplication and the output end of the second matrix multiplication are connected to the input end of the fourth addition layer, and the output of the fourth addition layer is the output of the semantic guidance enhancement module to which it belongs.

[0042] Combination Figure 4 The invention further comprises a plurality of semantically guided enhancement modules, wherein the input end of the first convolution and normalization layer in the first semantically guided enhancement module is connected to the output end of the last edge-aware enhancement module in the texture branch.

[0043] The input end of the second convolution and normalization layer in the first semantic guided enhancement module is connected to the output end of the N-M+1th self-attention based semantic module in the semantic branch.

[0044] The input end of the first convolution and normalization layer in the mth semantically guided enhancement module is connected to the output end of the m-1th semantically guided enhancement module.

[0045] The input of the second convolution and normalization layer in the mth semantic-guided enhancement module is connected to the output of the N-M+mth self-attention-based semantic module.

[0046] The semantic guidance enhancement module draws on the dynamic selection fusion mechanism and combines it with the pixel attention mechanism. By calculating the similarity between two pixels, the network can autonomously select useful texture information and semantic information and fuse them together, ensuring that more emphasis can be placed on retaining and utilizing key semantic information during the fusion process, thereby achieving the effect of using semantic information to guide texture information and preventing texture information from being suppressed by too much background or global context information. The output of the semantic guidance enhancement module is shown in formula (8): (8); In formula (8), represents the texture feature map input to the semantic-guided enhancement module, represents the semantic feature map of the semantic guided enhancement module input, represents a convolution operation with a convolution kernel of 3×3, BachNorm represents a batch normalization operation, and up represents a bilinear upsampling operation. Represents the element-by-element product operation, Sigmoid represents the S-type activation function, which is used to map the output to the range of [0, 1] and generate a weight feature map , represents the matrix product operation and SGEM represents the semantically guided enhancement module.

[0047] The texture prior sub-network includes: a plurality of boundary-guided aggregation modules (corresponding to Figure 4 The BGAM in the image is used for segmenting the edge contour of the defect target, improving the network's expression ability and the ability to extract inconspicuous defect features in the case of weak texture. Figure 4 An example is shown in FIG. 4 , where four boundary-guided aggregation modules are used in the texture prior subnetwork.

[0048] like Fig.10 As shown, the boundary-guided aggregation module includes: a fourth activation layer (using an S-type activation function), a third matrix multiplication layer, a fifth addition layer, a sixth addition layer, a fifth convolution layer (3×3 convolution), a global average pooling layer, a sixth convolution layer (one-dimensional convolution), a fifth activation layer (using an S-type activation function), a fourth matrix multiplication layer and a seventh convolution layer (1×1 convolution), which are connected in sequence.

[0049] The fourth activation layer is connected to the output of the last semantic-guided enhancement module; the input of the sixth addition layer is also connected to the input of the fourth activation layer; the input of the fourth matrix multiplication layer is also connected to the input of the global average pooling layer; the output of the seventh convolutional layer is the output of the boundary-guided aggregation module to which it belongs.

[0050] The input of the third matrix multiplication layer and the input of the fifth addition layer in the nth boundary-guided aggregation module are also connected to the output of the N-n+1th self-attention-based semantic module.

[0051] The boundary-guided aggregation module can feed back texture information into the representation learning of semantic information, and provide valuable edge texture prior guidance for the output of each layer of the semantic module of the semantic branch. By fusing features and combining the channel attention mechanism, it can highlight key channels, suppress redundant channels or noise, and enhance the feature representation of weak texture defect semantics. The output of the boundary-guided aggregation module is shown in formula (9): (9); In formula (9), Texture feature map representing the input of boundary-guided aggregation module; Represents the semantic feature map of the boundary-guided aggregation module input; Sigmoid represents the S-type activation function, which is used to map the output to the range of [0, 1] to generate the weight feature map W or ; Indicates a convolution operation with a convolution kernel of 3×3; Represents a matrix product operation; gap represents a global average pooling operation; Represents a one-dimensional convolution with a kernel size of k, where k is the closest An odd number, B represents the number of channels in the fifth convolutional layer; p represents the result of the 3×3 convolution operation; represents a convolution operation with a convolution kernel of 1×1, and BGAM represents a boundary guided aggregation module.

[0052] The following takes the design of 4 self-attention-based semantic modules, 2 semantic-guided enhancement modules, and 4 boundary-guided aggregation modules as an example to illustrate the specific process of the semantic-guided and texture-prior feature fusion network: First, through two semantic guidance enhancement modules, the feature map output by the third-layer semantic module in the semantic branch and the semantic information in the feature map output by the fourth-layer semantic module are used to recursively guide the output of the texture branch, mine accurate edge texture information, and suppress the noise of low-level features in the texture branch. The corresponding relationship is shown in formula (10): (10); In formula (10), and They represent the feature map output after passing through the mth and m-1th semantic guidance enhancement modules respectively; SGEM represents the semantic guidance enhancement module; when m=1, Represents the final feature map output by the texture branch; Indicates that the resolution in the semantic branch is feature map.

[0053] Secondly, the texture information guided by the two semantic guidance enhancement modules is aligned with the feature maps output by the semantic modules of each layer of the semantic branch. Figure 4 The upsampling and downsampling operations in are then input into four boundary-guided aggregation modules to provide valuable edge texture prior guidance for the feature maps output by each layer of the semantic module of the semantic branch. The purpose is to inject texture-related edge cues into the representation learning of the semantic branch to enhance the feature representation of weak texture defect semantics. The corresponding relationship is shown in formula (11): (11); In formula (11), Represents the texture feature map guided by two semantic-guided enhancement modules; Indicates that the resolution in the semantic branch is The feature map of according to The feature map size is used for registration operation; Indicates that the resolution in the semantic branch is The output of is passed through the feature map output by the boundary guided aggregation module; BGAM stands for boundary guided aggregation module. n refers to the nth semantic module.

[0054] In another exemplary embodiment of the present application, in step 102, the decoder is a decoder based on a feature pyramid structure. The feature pyramid structure is a progressive upsampling structure, and combines multi-scale features through lateral connections, which enhances the model's perception of features of different scales, increases the model's robustness and adaptability, and enables more fine-grained feature extraction, ultimately improving the model's segmentation accuracy for surface defects. Figure 4 As shown, the decoder includes a plurality of upsampling modules connected in sequence; each upsampling module is connected to a splicing module; the input end of the i-th splicing module is also connected to the output end of the i+1-th boundary-guided aggregation module in the texture prior subnetwork; the last upsampling module outputs the surface defect segmentation result of the object to be detected, such as Fig.11 As shown, Fig.11 In the figure, the white area is the surface defect area. Figure 4 An example is shown in , using 3 upsampling modules and 3 splicing modules.

[0055] The feature map after the semantic guidance and texture prior feature fusion network is input into the decoder based on the feature pyramid structure. The resolution of the feature map is restored from 1 / 32 upwards to the original image resolution to output the final surface defect segmentation map. The corresponding relationship is shown in formula (12): (12); In formula (12), and They represent the resolution after the boundary-guided aggregation module is and feature map; up represents the bilinear upsampling operation; Concat represents the concatenation operation; Indicates a convolution operation with a convolution kernel of 3×3; It means that the resolution after the boundary-guided aggregation module is The feature map of A segmentation map representing the final output of the model.

[0056] In another exemplary embodiment of the present application, in step 2, before the image of the object to be detected is input into the defect segmentation model, the defect segmentation model needs to be a trained model. In the present application, a cross entropy loss function is used to calculate the cross entropy loss between the surface defect segmentation result obtained by the model and the surface defect real annotation information. The cross entropy loss calculation formula is shown in formula (13): (13); In formula (13), G represents the total number of pixels in the image, C represents the total number of categories (or labels), represents a binary indicator variable, which is 1 if the g-th pixel belongs to the j-th class and 0 otherwise. It represents the probability that the g-th pixel belongs to the j-th class predicted by the model.

[0057] The cross entropy loss is calculated based on the target segmentation result obtained by the defect segmentation model and the target true annotation information, and the model parameters are adjusted through back propagation based on the cross entropy loss to obtain the trained defect segmentation model.

[0058] In this application, a dual-branch feature extraction network of semantics and texture is used to extract the features of various complex surface defects. The feature fusion network of semantic guidance and texture prior enables the model to better segment the edge contours of defective targets, so as to improve the expressive ability of the model and the ability to extract unclear defect features in weak texture conditions, and ultimately effectively improve the model's segmentation accuracy and robustness for surface defects.

[0059] The present application also provides an application scenario, which applies the above-mentioned dual-branch surface defect segmentation method based on semantic guidance and texture prior. Specifically: the dual-branch surface defect segmentation method with semantic guidance and texture prior provided in this embodiment can be applied in the surface defect segmentation scenario. The scenario includes an image acquisition link and a surface defect segmentation link; the image acquisition link is used to acquire the image of the industrial object to be inspected; the surface defect segmentation link is used to obtain the image of the industrial object to be inspected, input the image into the defect segmentation model, and obtain the surface defect segmentation result of the industrial object to be inspected. The dual-branch surface defect segmentation method with semantic guidance and texture prior provided in this embodiment belongs to the surface defect segmentation link.

[0060] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the dual-branch surface defect segmentation method based on semantic guidance and texture prior is implemented.

[0061] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0062] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0063] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A semantically guided and texture prior dual-branch surface defect segmentation method, characterized in that: include: Acquire an image of the object to be detected; Input the image of the object to be inspected into the defect segmentation model to obtain the surface defect segmentation result of the object to be inspected; The defect segmentation model includes a semantic and texture dual-branch feature extraction network, a semantic guidance and texture prior feature fusion network and a decoder connected in sequence; the semantic and texture dual-branch feature extraction network is used to extract semantic information, defect texture and edge features in the image of the object to be detected, and obtain a semantic feature map and a texture feature map; the semantic guidance and texture prior feature fusion network is used to perform feature fusion on the semantic feature map and the texture feature map; the decoder is used to output the surface defect segmentation result of the object to be detected according to the fused feature map; The semantic guidance and texture prior feature fusion network includes a semantic guidance sub-network and a texture prior sub-network; The semantic guidance sub-network comprises: a plurality of sequentially connected semantic guidance enhancement modules; the input end of each semantic guidance enhancement module is also connected to the output end of the texture branch and the semantic branch in the dual-branch feature extraction network of semantics and texture; the semantic guidance enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information; The texture prior subnetwork includes: multiple boundary-guided aggregation modules; the input end of each boundary-guided aggregation module is connected to the output end of the last semantic-guided enhancement module and the output end of the semantic branch; the output end of each boundary-guided aggregation module is connected to the input end of the decoder; the boundary-guided aggregation module is used to provide edge texture prior guidance for the output of the semantic branch and segment the edge contour of the defect target.

2. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 1, characterized in that: The dual-branch feature extraction network of semantics and texture includes: a semantic branch and a texture branch; The semantic branch includes a block coding module and a plurality of self-attention-based semantic modules connected in sequence; the image feature resolutions from the first self-attention-based semantic module to the last self-attention-based semantic module decrease in sequence; the first self-attention-based semantic module is a semantic module connected to the block coding module; the self-attention-based semantic module is used to extract semantic information from an input image; The output of each self-attention-based semantic module is also connected to the input of the semantic-guided and texture-prior feature fusion network; The texture branch includes a plurality of edge-aware enhancement modules connected in sequence; the output end of the last edge-aware enhancement module is connected to the input end of the semantic guidance and texture prior feature fusion network; the edge-aware enhancement module is used to extract defective texture and edge features in the input image.

3. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 2, characterized in that: The self-attention-based semantic module includes: a self-attention module, a feedforward neural network, a linear normalization layer and a downsampling layer connected in sequence; the input end of the linear normalization layer is also connected to the input end of the self-attention module.

4. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 2, characterized in that: The edge perception enhancement module includes: a first convolution layer, a first normalization and activation layer, a second convolution layer, a first activation layer, a detail enhancement convolution layer, a second activation layer, a first addition layer, a third convolution layer, a second normalization and activation layer, and a second addition layer connected in sequence; The input end of the first addition layer is also connected to the output end of the first activation layer; the input end of the second addition layer is also connected to the input end of the first convolutional layer.

5. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 4, characterized in that: The detail enhancement convolution layer includes a third addition layer and a fourth convolution layer, a center difference convolution layer, an angle difference convolution layer, a horizontal difference convolution layer, and a vertical difference convolution layer connected in parallel; The input ends of the fourth convolution layer, the center difference convolution layer, the angle difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer are all connected to the output end of the first activation layer; The output ends of the fourth convolution layer, the center difference convolution layer, the angle difference convolution layer, the horizontal difference convolution layer, and the vertical difference convolution layer are all connected to the input end of the third addition layer; The output terminal of the third addition layer is connected to the input terminal of the second activation layer.

6. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 2, characterized in that: The semantic guidance and texture prior feature fusion network includes a semantic guidance sub-network and a texture prior sub-network; The semantic guidance sub-network includes: a plurality of semantic guidance enhancement modules connected in sequence; the input end of the first semantic guidance enhancement module is respectively connected to the output end of the last edge perception enhancement module in the texture branch and the output end of the N-M+1 th semantic module based on self-attention in the semantic branch; the input end of the m th semantic guidance enhancement module is also connected to the output end of the N-M+m th semantic module based on self-attention; N is the number of semantic modules based on self-attention; M is the number of semantic guidance enhancement modules; m>1; N≥M; the semantic guidance enhancement module is used to use the semantic information in the semantic branch to guide the texture branch to mine edge texture information; The texture prior subnetwork includes: multiple boundary-guided aggregation modules; the input end of the nth boundary-guided aggregation module is respectively connected to the output end of the last semantic-guided enhancement module and the output end of the N-n+1th self-attention-based semantic module; the output end of each boundary-guided aggregation module is connected to the input end of the decoder; the boundary-guided aggregation module is used to provide edge texture prior guidance for the output of each layer of the semantic branch, and segment the edge contour of the defect target.

7. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 6, characterized in that: The semantic guidance enhancement module includes: a first convolution and normalization layer, a second convolution and normalization layer, an upsampling layer, an element-by-element multiplication layer, a third activation layer, a first matrix multiplication layer, a second matrix multiplication layer and a fourth addition layer; The output end of the first convolution and normalization layer is connected to the input end of the element-by-element multiplication layer and the first matrix multiplication layer respectively; the output end of the element-by-element multiplication layer is connected to the input end of the third activation layer, and the output end of the third activation layer is connected to the input end of the first matrix multiplication layer and the input end of the second matrix multiplication layer respectively; The output end of the second convolution and normalization layer is connected to the input end of the upsampling layer, and the output end of the upsampling layer is connected to the input ends of the element-by-element multiplication layer and the second matrix multiplication layer respectively; The output end of the first matrix multiplication and the output end of the second matrix multiplication are connected to the input end of the fourth addition layer, and the output of the fourth addition layer is the output of the semantic guidance enhancement module to which it belongs; The input end of the first convolution and normalization layer in the first semantic-guided enhancement module is connected to the output end of the last edge-aware enhancement module in the texture branch; The input end of the second convolution and normalization layer in the first semantic guided enhancement module is connected to the output end of the N-M+1 th self-attention based semantic module in the semantic branch; The input end of the first convolution and normalization layer in the mth semantic guidance enhancement module is connected to the output end of the m-1th semantic guidance enhancement module; The input of the second convolution and normalization layer in the mth semantic-guided enhancement module is connected to the output of the N-M+mth self-attention-based semantic module.

8. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 6, characterized in that: The boundary guided aggregation module includes: a fourth activation layer, a third matrix multiplication layer, a fifth addition layer, a sixth addition layer, a fifth convolution layer, a global average pooling layer, a sixth convolution layer, a fifth activation layer, a fourth matrix multiplication layer and a seventh convolution layer connected in sequence; The fourth activation layer is connected to the output of the last semantic guidance enhancement module; the input of the sixth addition layer is also connected to the input of the fourth activation layer; the input of the fourth matrix multiplication layer is also connected to the input of the global average pooling layer; the output of the seventh convolutional layer is the output of the boundary guidance aggregation module to which it belongs; The input of the third matrix multiplication layer and the input of the fifth addition layer in the nth boundary-guided aggregation module are also connected to the output of the N-n+1th self-attention-based semantic module.

9. The semantic-guided and texture-prior dual-branch surface defect segmentation method according to claim 6, characterized in that: The decoder is a decoder based on a feature pyramid structure; the decoder includes a plurality of upsampling modules connected in sequence; a splicing module is also provided between two adjacent upsampling modules; the input end of the i-th splicing module is also connected to the output end of the i+1-th boundary-guided aggregation module in the texture prior subnetwork; and the last upsampling module outputs the surface defect segmentation result of the object to be detected.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the semantically guided and texture-prior dual-branch surface defect segmentation method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Connected double-attention multi-scale fusion semantic segmentation network

    CN116630626A

  • Surface defect detection method and system based on external semantics and high-frequency information

    CN118396976A

  • Display screen defect detection method and system based on enhanced feature extraction network

    CN118570212A

  • Defect detection method, system and equipment for power transmission line and storage medium

    CN119649031A

  • Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment

    WO2024230038A1

Cited By

  • Concrete construction robot ground leveling quality detection method

    CN120747888A

  • A method for inspecting the quality of ground leveling by a concrete construction robot

    CN120747888B

  • An industrial defect detection method and system based on bidirectional semantic flow and boundary perception refinement

    CN122597264A