Image detection method and device, equipment and storage medium
Through semantic feature extraction and edge feature extraction of image semantic segmentation model, the shortcomings of existing image detection algorithms in edge extraction and computing efficiency are solved, and more efficient and accurate image detection is achieved.
Patent Information
- Application Number
- CN202510637708.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-29
AI Technical Summary
The existing image detection algorithms have shortcomings in edge extraction and computing efficiency, resulting in low detection efficiency, waste of computing resources and low detection accuracy.
The image semantic segmentation model is used, and the backbone network of the pooling module and attention mechanism is combined for semantic feature extraction, and the edge feature extraction is extracted by multiple hollow convolution units, which enhances the feature extraction capability and reduces the calculation amount, while retaining image detail information.
It improves the efficiency and accuracy of image detection, reduces the waste of computing resources, enhances the utilization of image feature information, and improves the detection accuracy.
Smart Images

Figure CN120563802A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of image processing and quality detection, and more specifically, to an image detection method, apparatus, device and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, image detection algorithms have gradually replaced visual inspection and are widely used in appearance recognition and quality inspection of products in various fields.
[0003] For example, electronic products can be inspected based on image detection algorithms to promptly detect defects and deficiencies. However, related image detection algorithms still have deficiencies in edge extraction and computational efficiency. Summary of the Invention
[0004] In view of the above problems, the present application provides an image detection method, apparatus, device, medium and program product.
[0005] According to the first aspect of the present application, an image detection method is provided, comprising: obtaining a target image to be detected; performing feature extraction on the target image based on a semantic feature extraction branch of an image semantic segmentation model to obtain a semantic feature image and a context feature image group, wherein the semantic feature extraction branch comprises a pooling module and a backbone network based on an attention mechanism, the semantic feature image is obtained by processing the target image using the backbone network, and the context feature image group is obtained by processing the semantic feature image using the pooling module; performing feature extraction on the target image based on an edge feature extraction branch of the image semantic segmentation model to obtain an edge feature image, the edge feature extraction branch comprising a plurality of atrous convolution units; determining an image detection result based on a reference semantic segmentation image and the semantic feature image, the context feature image group and the edge feature image, and the image detection result is used to locate the position of the abnormality in the target image.
[0006] Another aspect of the present application provides an image detection device, including: an acquisition module, used to acquire a target image to be detected; a first processing module, used to perform feature extraction on the target image based on a semantic feature extraction branch of an image semantic segmentation model, to obtain a semantic feature image and a context feature image group, wherein the semantic feature extraction branch includes a pooling module and a backbone network based on an attention mechanism, the semantic feature image is obtained by processing the target image using the backbone network, and the context feature image group is obtained by processing the semantic feature image using the pooling module; a second processing module, used to perform feature extraction on the target image based on an edge feature extraction branch of an image semantic segmentation model, to obtain an edge feature image, the edge feature extraction branch includes a cascade of multiple hole convolution units; a determination module, used to determine an image detection result based on a reference semantic segmentation image, the semantic feature image, the context feature image group and the edge feature image, and the image detection result is used to locate the position of the abnormality in the target image.
[0007] Another aspect of the present application provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0008] Another aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0009] Another aspect of the present application further provides a computer program product, including a computer program or instructions, which implement the steps of the above method when the computer program or instructions are executed by a processor.
[0010] According to an embodiment of the present application, the image semantic segmentation model includes a semantic feature extraction branch and an edge feature extraction branch. The semantic feature extraction branch may include a pooling module and a backbone network based on an attention mechanism. The edge feature extraction branch may include multiple hole convolution units. A lightweight convolutional network embedded with an attention module can be used to extract features from a target image to obtain a semantic feature image. The attention module can focus on more important key features and suppress the influence of irrelevant features, thereby enhancing the feature extraction capability of the model and reducing unnecessary computational complexity. In addition, the lightweight backbone network has a small number of parameters, which can effectively save computing resources and improve detection efficiency. The semantic feature image can be processed using a pooling module to obtain a context feature map group. The context feature map group aggregates context information of different regions in the semantic feature image, which can provide more effective and comprehensive global prior information for subsequent feature extraction. Based on the edge feature extraction branch, the target image can be feature extracted to obtain an edge feature image. The edge feature extraction branch can increase the receptive field while retaining the original detail information in the target image to compensate for the edge feature loss defect caused by the pooling module. Therefore, the image semantic segmentation model provided by the embodiment of the present application can at least partially overcome the technical problems existing in the related technologies, such as low detection efficiency, waste of computing resources, poor edge feature extraction capability, insufficient utilization of image feature information, resulting in low detection accuracy, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0012] Figure 1 A diagram schematically illustrates an application scenario of the image detection method, apparatus, device, medium, and program product according to an embodiment of the present application;
[0013] Figure 2 The following schematically shows a flow chart of an image detection method according to an embodiment of the present application;
[0014] Figure 3 A schematic diagram of the network structure of a backbone network according to an embodiment of the present application is shown;
[0015] Figure 4 Schematically illustrates a schematic diagram of the principle of depthwise separable convolution according to an embodiment of the present application;
[0016] Figure 5 The network structure diagram of the convolutional attention module according to an embodiment of the present application is schematically shown;
[0017] Figure 6 The schematic diagram of the principle of the dilated convolution according to the embodiment of the present application is schematically shown;
[0018] Figure 7 A schematic diagram of the network structure of the edge feature extraction branch according to an embodiment of the present application is shown;
[0019] Figure 8 The following schematically shows a network structure diagram of an image semantic segmentation model according to an embodiment of the present application;
[0020] Figure 9 A schematic block diagram of the structure of an image detection device according to an embodiment of the present application is shown; and
[0021] Figure 10 A block diagram of an electronic device suitable for implementing the image detection method according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0022] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0023] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0025] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0026] The accompanying drawings show some block diagrams and / or flow charts. It should be understood that some blocks in the block diagrams and / or flow charts, or combinations thereof, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when these instructions are executed by the processor, they can create a device for implementing the functions / operations described in the block diagrams and / or flow charts.
[0027] Therefore, the techniques of this application can be implemented in hardware and / or software (including firmware, microcode, etc.). Furthermore, the techniques of this application can take the form of a computer program product on a computer-readable medium having stored thereon instructions, which can be used by or in conjunction with an instruction execution system. In the context of this application, a computer-readable medium can be any medium that can contain, store, convey, propagate, or transmit instructions. For example, a computer-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium. Specific examples of computer-readable media include: magnetic storage devices, such as magnetic tape or hard disk drives (HDDs); optical storage devices, such as compact disks (CD-ROMs); memory, such as random access memory (RAM) or flash memory; and / or wired or wireless communication links.
[0028] Appearance inspection can be understood as the detection of foreign matter, flaws, and defects on the surface of components and products. With the rapid development of artificial intelligence technology, image detection algorithms have gradually replaced visual inspection and are widely used in appearance recognition and quality inspection of products in various fields.
[0029] For example, electronic products can be inspected based on image detection algorithms to promptly detect defects and deficiencies. However, related image detection algorithms still have deficiencies in edge extraction and computational efficiency.
[0030] In a related embodiment, product appearance inspection can be performed based on an image semantic segmentation model. Image semantic segmentation classifies each pixel in an image into different categories. Compared to traditional image recognition, which only performs "presence or absence" judgments, semantic segmentation can more finely classify each part of the image. Therefore, it is suitable for automatic image inspection in delicate areas, such as the inspection of electronic product appearance images.
[0031] While implementing the inventive concept of this application, the inventors discovered that related image detection methods based on semantic segmentation models suffer from the following problems: The models' unclear focus leads to low detection efficiency and waste of computational resources. Furthermore, the models' edge feature extraction capabilities are poor, and they don't fully utilize image feature information, resulting in low detection accuracy.
[0032] Embodiments of the present application provide an image detection method, apparatus, device, medium, and program product, the image detection method comprising: obtaining a target image to be detected; performing feature extraction on the target image based on a semantic feature extraction branch of an image semantic segmentation model to obtain a semantic feature image and a context feature image group, wherein the semantic feature extraction branch comprises a pooling module and a backbone network based on an attention mechanism, the semantic feature image is obtained by processing the target image using the backbone network, and the context feature image group is obtained by processing the semantic feature image using the pooling module; performing feature extraction on the target image based on an edge feature extraction branch of an image semantic segmentation model to obtain an edge feature image, the edge feature extraction branch comprising multiple hole convolution units; determining an image detection result based on a reference semantic segmentation image and the semantic feature image, the context feature image group, and the edge feature image, and the image detection result is used to locate the position of the abnormality in the target image.
[0033] Figure 1 The following schematically illustrates an application scenario of the image detection method according to an embodiment of the present application.
[0034] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0035] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0036] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0037] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0038] It should be noted that the image detection method provided in the embodiment of the present application can generally be performed by the server 105. Accordingly, the image detection device provided in the embodiment of the present application can generally be set in the server 105. The image detection method provided in the embodiment of the present application can also be performed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the image detection device provided in the embodiment of the present application can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0039] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0040] It should be noted that the image detection method and device provided in the embodiments of the present application can be used for appearance inspection of electronic products (such as servers, circuit boards, components, etc.), and can also be used for appearance inspection of other products such as mechanical parts and medical devices. The application fields of the image detection method and device provided in the embodiments of the present application are not limited.
[0041] Figure 2 The flowchart of the image detection method according to an embodiment of the present application is schematically shown.
[0042] like Figure 2 As shown, the method 200 includes operations S210 to S240.
[0043] In operation S210 , an image of a target to be detected is acquired.
[0044] In operation S220, based on the semantic feature extraction branch of the image semantic segmentation model, features are extracted from the target image to obtain a semantic feature image and a context feature image group. The semantic feature extraction branch includes a pooling module and a backbone network based on the attention mechanism. The semantic feature image is obtained by processing the target image using the backbone network, and the context feature image group is obtained by processing the semantic feature image using the pooling module.
[0045] In operation S230 , based on the edge feature extraction branch of the image semantic segmentation model, feature extraction is performed on the target image to obtain an edge feature image, where the edge feature extraction branch includes a plurality of dilated convolution units.
[0046] In operation S240 , an image detection result is determined according to the reference semantic segmentation image based on the semantic feature image, the context feature image group, and the edge feature image. The image detection result is used to locate the position of an abnormality in the target image.
[0047] The target image may be an image of the appearance of a product. For example, a high-resolution camera may be used to take a picture of the appearance of the product to be inspected to obtain the target image to be inspected.
[0048] An image semantic segmentation model can be used to detect the target image and determine the image detection results. For example, the image semantic segmentation model can be trained based on sample target images and labeled images. For example, a high-resolution camera can be used to capture a large number of sample product images to collect sample target images. The pixels in the sample target images can then be classified and labeled according to specific needs to obtain corresponding labeled images. A candidate image semantic segmentation model can be trained based on the sample target images and labeled images to obtain an image semantic segmentation model.
[0049] In one embodiment, the image semantic segmentation model may include a semantic feature extraction branch and an edge feature extraction branch.
[0050] The semantic feature extraction branch can be used to extract features from the target image, generating a semantic feature image and a contextual feature image group. The semantic feature extraction branch can include a pooling module and a backbone network based on an attention mechanism. The semantic feature image is generated by processing the target image using the backbone network, while the contextual feature image group is generated by processing the semantic feature image using the pooling module.
[0051] In one embodiment, the aforementioned backbone network can be selected as a convolutional network embedded with an attention module, aiming to obtain more important key image features and suppress the influence of irrelevant features without destroying the original convolutional network structure. The convolutional network is preferably a lightweight network model (such as the MobileNet model) to effectively reduce the number of parameters and save computing resources. The attention module can dynamically allocate more computing resources and attention to important feature parts to enhance the model's feature extraction capabilities and reduce unnecessary calculations. The feature learning of the attention module can also effectively help the feature information of the image flow and diffuse.
[0052] For example, the backbone network can extract features from the input target image to obtain a semantic feature image. The semantic feature image can represent features such as the category, shape, and texture of the product in the target image.
[0053] In one embodiment, the aforementioned pooling module can be a pyramid pooling module (PPM). The pyramid pooling module can generate multiple feature maps of different sizes by performing pooling operations of different scales on the original feature map, and then concatenate these feature maps in the channel dimension to form a composite feature map containing information of multiple scales.
[0054] For example, the pooling module can perform pooling operations at different scales on the semantic feature image to generate multiple context feature maps, thereby obtaining a context feature map group. The context feature map group aggregates context information from different regions in the semantic feature image, providing more effective and comprehensive global prior information for subsequent feature extraction.
[0055] Because pooling generally reduces the spatial dimensions of feature maps, the pooling module loses much detail in the target image, especially edge details, while expanding the receptive field. To address this, the edge feature extraction branch of the semantic segmentation model can preserve the original details in the target image while increasing the receptive field, compensating for the loss of edge features caused by the pooling module.
[0056] In one embodiment, the edge feature extraction branch may include multiple dilated convolution units. Dilated convolution, also known as expanded convolution, can expand the receptive field by introducing gaps (holes) in the convolution kernel without adding additional parameters or computational effort. Based on the edge feature extraction branch, features can be extracted from the target image to generate an edge feature image. This edge feature image can characterize features such as the outline and boundaries of the product in the target image, thereby helping the model accurately segment object boundaries.
[0057] It is understandable that the semantic feature extraction branch based on the embedded attention mechanism can focus on extracting key image features in the target image, thereby enhancing the model's feature extraction capabilities and reducing unnecessary computation. The edge feature extraction branch based on multiple dilated convolutional units can retain the original image features of the target image to a greater extent, compensating for the edge feature loss defects of the pooling module, and can help the model more accurately restore the classification content of the image, thereby effectively improving the accuracy of semantic segmentation.
[0058] The reference semantic segmentation image can be understood as a standard semantic segmentation image of a product. For example, semantic segmentation can be performed on an image of the appearance of a qualified product to obtain a reference semantic segmentation image. The qualified product has a normal appearance, so the reference semantic segmentation image can be used as the standard semantic segmentation image of the product.
[0059] Image detection results can be determined based on the semantic feature image, context feature image group, and edge feature image, using the reference semantic segmentation image. The image detection results can be used to locate anomalies in the target image, thereby determining whether the product appearance corresponding to the target image has defects or deficiencies based on the image detection results.
[0060] According to an embodiment of the present application, the image semantic segmentation model includes a semantic feature extraction branch and an edge feature extraction branch. The semantic feature extraction branch may include a pooling module and a backbone network based on an attention mechanism. The edge feature extraction branch may include multiple hole convolution units. A lightweight convolutional network embedded with an attention module can be used to extract features from a target image to obtain a semantic feature image. The attention module can focus on more important key features and suppress the influence of irrelevant features, thereby enhancing the feature extraction capability of the model and reducing unnecessary computational complexity. In addition, the lightweight backbone network has a small number of parameters, which can effectively save computing resources and improve detection efficiency. The semantic feature image can be processed using a pooling module to obtain a context feature map group. The context feature map group aggregates context information of different regions in the semantic feature image, which can provide more effective and comprehensive global prior information for subsequent feature extraction. Feature extraction can be performed on the target image based on the edge feature extraction branch to obtain an edge feature image. The edge feature extraction branch can increase the receptive field while retaining the original detail information in the target image to compensate for the edge feature loss defect caused by the pooling module. Therefore, the image semantic segmentation model provided by the embodiment of the present application can at least partially overcome technical problems such as low detection efficiency, waste of computing resources, poor edge feature extraction capability, and insufficient utilization of image feature information, resulting in low detection accuracy.
[0061] According to an embodiment of the present application, based on the semantic feature extraction branch of the image semantic segmentation model, feature extraction is performed on the target image to obtain a semantic feature image and a context feature image group, including: feature extraction is performed on the target image based on the backbone network to obtain a semantic feature image; based on the pooling module, pooling operations are performed on multiple regions of the semantic feature image respectively to obtain a context feature image group, wherein the sizes of the multiple regions are different.
[0062] The target image can be input into the backbone network to perform feature extraction on the target image based on the backbone network and output a semantic feature image. The semantic feature image can be input into the pooling module, which can divide the semantic feature image into multiple regions of different sizes and perform pooling operations of different scales on each region to obtain multiple context feature images of different scales. The multiple context feature images can be used as pooled representations of different positions of the semantic feature image, and a context feature image group can be obtained based on the multiple context feature images. Exemplarily, the aforementioned multiple regions of different sizes can include, for example, a global region and multiple sub-regions.
[0063] It can be understood that the pooling module constructs different deep features from the input semantic feature image through pooling operations based on multiple scales. It can aggregate the contextual information of different regions of the semantic feature image to provide more effective and comprehensive global prior information for subsequent feature extraction. As a result, the pooling module can fully utilize global contextual information and sub-region contextual information, while reducing the loss of contextual information between different sub-regions.
[0064] According to an embodiment of the present application, according to a reference semantic segmentation image, based on a semantic feature image, a context feature image group and an edge feature image, determining an image detection result includes: fusing the semantic feature image, the context feature image group and the edge feature image to obtain a fused feature image; decoding the fused feature image to obtain a semantic segmentation image corresponding to the target image; comparing the semantic segmentation image with the reference semantic segmentation image to determine the image detection result.
[0065] Feature fusion can comprehensively utilize multiple image features to complement each other's strengths. The integrated feature representation can more comprehensively describe the data, thereby improving model performance and accuracy. Feature fusion methods include feature addition, feature multiplication, feature concatenation, and attention fusion.
[0066] For example, a semantic feature image, a contextual feature image group, and an edge feature image can be concatenated (concat fusion) to produce a fused feature image. Compared to feature addition fusion (add fusion), concat fusion results in less information loss and can better leverage the complementary advantages of individual image features.
[0067] It can be understood that the semantic feature image provides semantic category information of the product in the target image, the context feature image group provides global and local multi-scale context information, and the edge feature image provides information such as the boundary and contour of the product in the target image. By fusing the first semantic feature image, the context feature image and the edge feature image, more original details and context information in the target image can be captured, thereby effectively enhancing the feature expression ability of the fused feature image, which in turn can help improve the robustness and accuracy of image semantic segmentation.
[0068] In one embodiment, the semantic segmentation image may further include a decoding module, which may include a convolutional layer and an activation function layer. The fused feature image may be input into the convolutional layer for feature extraction, and the feature image output from the convolutional layer may be input into the activation function layer to obtain the final output semantic segmentation image. Exemplarily, the activation function may be selected from, for example, a Softmax function, a ReLU function, etc., which is not specifically limited herein.
[0069] In one embodiment, the semantically segmented image can be compared with a reference semantic image to obtain an image detection result. The image detection result may include, for example, locations where the semantically segmented image and the reference semantic image are identical and different. The locations of anomalies in the target image can be located based on the different locations in the image detection result. Furthermore, defects or deficiencies in the product appearance corresponding to the target image can be determined based on the image detection result.
[0070] According to an embodiment of the present application, the backbone network includes N first convolutional units, N convolutional attention modules, and at least one second convolutional unit, where N is an integer greater than 1; performing feature extraction on the target image based on the backbone network to obtain a semantic feature image includes: using the first first convolutional unit to extract features of the target image to obtain a first initial semantic feature image; using the first convolutional attention module to extract features of the first initial semantic feature image to obtain a first intermediate semantic feature image; using the nth first convolutional unit to extract features of the n-1th intermediate semantic feature image to obtain an nth initial semantic feature image, n=2,…,N; using the nth convolutional attention module to extract features of the nth initial semantic feature image to obtain an nth intermediate semantic feature image; and performing feature extraction on the Nth intermediate semantic feature image based on at least one second convolutional unit to obtain a semantic feature image.
[0071] In one embodiment, the backbone network can be a MobileNet model embedded with a convolutional attention module. Compared to other backbone networks, the MobileNet model is more lightweight while maintaining good segmentation results. It strikes a better balance between accuracy and efficiency, effectively saving computing resources and improving the overall detection efficiency of the model.
[0072] As an example, a convolutional attention module can be added to the first four layers of the MobileNet model. This aims to leverage the attention mechanism early in feature extraction, preserving more of the original image features of the target image, enhancing the model's feature extraction capabilities, and reducing unnecessary computation. For example, the convolutional attention module can be a CBAM (Convolutional Block Attention Module).
[0073] Figure 3 The diagram schematically shows the network structure of the backbone network according to an embodiment of the present application.
[0074] like Figure 3 As shown, the backbone network may include 4 first convolutional units 310M, 4 convolutional attention modules 320M, and 2 second convolutional units 330M.
[0075] For example, the first convolution unit 310M can be used to extract features from the input target image 301 to obtain a first initial semantic feature image. The first convolution attention module 320M can be used to extract features from the first initial semantic feature image to obtain a first intermediate semantic feature image 302. The number of channels of the first intermediate semantic feature image 302 can be, for example, 32.
[0076] For example, the second first convolution unit 310M can be used to extract features from the first intermediate semantic feature image 302 to obtain a second initial semantic feature image. The second convolution attention module 320M can be used to extract features from the second initial semantic feature image to obtain a second intermediate semantic feature image 303. The number of channels of the second intermediate semantic feature image 303 can be, for example, 64.
[0077] For example, the third first convolution unit 310M can be used to extract features from the second intermediate semantic feature image 303 to obtain a third initial semantic feature image. The third convolutional attention module 320M can be used to extract features from the third initial semantic feature image to obtain a third intermediate semantic feature image 304. The number of channels of the third intermediate semantic feature image 304 can be, for example, 128.
[0078] For example, the fourth first convolution unit 310M can be used to extract features from the third intermediate semantic feature image 304 to obtain a fourth initial semantic feature image. The fourth convolutional attention module 320M can be used to extract features from the fourth initial semantic feature image to obtain a fourth intermediate semantic feature image 305. The number of channels of the fourth intermediate semantic feature image 305 can be, for example, 256.
[0079] For example, the first second convolution unit 330M can be used to extract features from the fourth intermediate semantic feature image 305 to obtain the fifth initial semantic feature image 306. The second second convolution unit 330M can be used to extract features from the fifth initial semantic feature image 306 to obtain the semantic feature image 307.
[0080] In one example, both the first convolution unit 310M and the second convolution unit 330M can be selected as depthwise separable convolution units. Depthwise separable convolution is an efficient convolution operation that aims to reduce the amount of computation and the number of parameters while maintaining model performance.
[0081] Depthwise convolution can be composed of channel-by-channel convolution (Depthwise Convolution) and pointwise convolution (Pointwise Convolution), which can decouple the depth information and spatial information of the image to gradually obtain refined segmentation results. Depthwise separable convolution can separate the depth information and spatial information, and by fusing the underlying feature information of different scales, it achieves the decoupling of spatial information and depth information. Compared with ordinary convolution, depthwise separable convolution uses a different convolution kernel for each channel and learns the features of each channel. The obtained feature information is more comprehensive and retains more important edge feature information, effectively reducing information loss during the upsampling process and improving the accuracy of segmentation prediction results. In addition, compared with ordinary convolution, the number of parameters of depthwise separable convolution is also greatly reduced, further reducing redundant calculations and improving the computing speed.
[0082] Figure 4 The schematic diagram of the principle of depth-wise separable convolution according to an embodiment of the present application is shown schematically. For example, Figure 4 It is the process of performing depth-wise separable convolution on a 3-channel input image.
[0083] like Figure 4 As shown in the figure, depthwise separable convolution includes channel-by-channel convolution and point-by-point convolution. Channel-by-channel convolution is a convolution operation performed on each input channel separately. A convolution kernel of channel-by-channel convolution is only responsible for one channel, and a channel is only convolved by one convolution kernel. Channel-by-channel convolution is performed entirely in a two-dimensional plane, and the number of convolution kernels is the same as the number of channels in the previous layer. The operation of point-by-point convolution is similar to that of conventional convolution. Its convolution kernel size is 1×1×M (M is the number of channels in the previous layer). Point-by-point convolution is a weighted combination of multiple feature maps output by channel-by-channel convolution in the depth direction to generate a new feature map. Figure 5 Where N is the number of 1×1 convolution kernels, which is also the number of channels of the obtained feature map.
[0084] According to an embodiment of the present application, the convolutional attention module includes a channel attention submodule and a spatial attention submodule, and the nth convolutional attention module is used to perform feature extraction on the nth initial semantic feature image to obtain the nth intermediate semantic feature image, including: inputting the nth initial semantic feature image into the channel attention submodule to obtain the channel attention weight; inputting the nth initial semantic feature image into the spatial attention submodule to obtain the spatial attention weight; using the channel attention weight to weight the nth initial semantic feature image to obtain the attention feature image; using the spatial attention weight to weight the attention feature image to obtain the nth intermediate semantic feature image.
[0085] In one embodiment, the convolutional attention module can be selected as a CBAM module. The CBAM module can include two sub-modules: a channel attention module and a spatial attention module, which perform channel and spatial attention calculations, respectively. This not only saves parameters and computing power, but also ensures that the CBAM module can be integrated into the existing network architecture as a plug-and-play module. The channel attention module (Channel Attention Mechanism) and the spatial attention module (SpatialAttention Mechanism) are two main forms of attention mechanisms. They respectively weight the feature maps of the channel dimension and the spatial dimension, thereby allowing the network to pay more attention to important features. The CBAM module combines these two attention mechanisms to effectively extract key channel features while preserving spatial information, improving the network's performance in processing complex image tasks.
[0086] In relevant deep convolutional neural networks, different convolutional layers focus on different image feature information. For example, some convolutional layers pay more attention to the channel features of the image, but overly complex spatial features will cause the network to generate a lot of non-pixel information; some convolutional layers pay more attention to the spatial features of the image, but inappropriate introduction of channel features will easily lead to overfitting.
[0087] In view of this, the embodiment of the present application forms the backbone network of the image semantic segmentation model by introducing the CBAM module of the hybrid attention mechanism into the convolutional network. The addition of the CBAM module can introduce two analysis dimensions, spatial attention and channel attention, starting from the two scopes of channel and space, to realize the sequential attention structure from channel to space. Spatial attention allows the backbone network to pay more attention to the pixel areas in the target image that play a decisive role in classification and ignore irrelevant areas. Channel attention is used to process the distribution relationship of feature map channels and distribute attention to the two dimensions at the same time, thereby enhancing the effect of the attention mechanism on improving model performance.
[0088] Taking the second convolutional attention module in the backbone network as an example, the second initial semantic feature image can be input into the channel attention submodule to obtain the channel attention weight. The second initial semantic feature image can be input into the spatial attention submodule to obtain the spatial attention weight. The second initial semantic feature image can be weighted using the channel attention weight to obtain the attention feature image. The attention feature image can be weighted using the spatial attention weight to obtain the second intermediate semantic feature image.
[0089] Figure 5 A schematic diagram of the network structure of a convolutional attention module according to an embodiment of the present application is shown schematically.
[0090] like Figure 5 As shown, the convolutional attention module may include a channel attention submodule 510M and a spatial attention submodule 520M.
[0091] For example, Figure 5 As shown, the channel attention submodule 510M can focus on the input feature image (B, C, H, W, Figure 5 The input feature is shown in the figure. Max pooling and average pooling are performed on each channel at the same time to compress the information of each channel in the spatial dimension, resulting in two vectors (B, C, 1, 1). At this time, for each channel, its spatial information is compressed into a single value by the max pooling or average pooling operation, thereby achieving the compression and extraction of global spatial information. The two vectors are then passed through a multi-layer perceptron network ( Figure 5 The MLP is combined with the superposition operation, and the output feature vector is finally generated into a channel attention weight vector through the Sigmoid function. . Figure 5 in represents the addition operation, Represents the activation function.
[0092] (1)
[0093] In formula (1), Represents the input feature image, for example, the input of the second convolutional attention module is the second initial semantic feature image. represents the maximum pooling operation, represents the average pooling operation, represents a multilayer perceptron network, Represents the activation function.
[0094] For example, Figure 5As shown, the spatial attention submodule 520M can perform average pooling and maximum pooling operations on the input feature image along the channel axis, compressing the input channel dimension (C) to 1, thereby integrating the global channel information into a single channel feature map, and obtaining two (B, 1, H, W) feature maps. In this process, for each sample B, the spatial attention submodule 520M will perform average pooling on the channel features of the sample, thereby achieving compression and merging of the global channel information. This operation helps to reduce the number of parameters and computational complexity while retaining important channel feature information. The two feature maps obtained can be spliced and reduced in dimension to return the number of channels to 1 again, so that all the obtained feature information is distributed on one channel, and then the spatial attention weight vector is generated through the Sigmoid function. .
[0095] (2)
[0096] In formula (2), Represents a convolution operation.
[0097] For example, Figure 5 As shown, the channel attention weight vector can be With the input feature image This step is used to adjust the eigenvalues of each channel, emphasize the information of important channels, and suppress the information of unimportant channels. The spatial attention weight vector can be Multiply it by the attention feature image obtained by the channel attention submodule to get the final output of the convolutional attention module ( Figure 5 (shown as Output Feature). Figure 5 in Represents matrix multiplication.
[0098] (3)
[0099] In formula (3), Represents the output result of the convolutional attention module. For example, the output of the second convolutional attention module is the second intermediate semantic feature image.
[0100] By adopting serial operations to multiply the spatial attention weights with the attention feature image to obtain the attention feature image, the original image features in the target image can be retained as much as possible and the amount of calculation can be reduced, which is conducive to improving the subsequent semantic segmentation accuracy and improving computational efficiency.
[0101] According to an embodiment of the present application, the edge feature extraction branch includes M cascaded hole convolution units, M third convolution units, and a fourth convolution unit, where M is an integer greater than 1; based on the edge feature extraction branch of the image semantic segmentation model, feature extraction is performed on the target image to obtain an edge feature image, including: using the first hole convolution unit to extract features of the target image to obtain the first initial edge feature image; using the mth hole convolution unit to perform feature extraction on the m-1th initial edge feature image to obtain the mth initial edge feature image, m=2,...,M; based on the M third convolution units, the first initial edge feature image to the Mth initial edge feature image are processed respectively to obtain the first initial fused edge feature image to the Mth initial fused edge feature image; the first initial fused edge feature image is fused to the Mth initial fused edge feature image to obtain a fused edge feature image; and the fused edge feature image is processed using the fourth convolution unit to obtain an edge feature image.
[0102] The edge feature extraction branch extracts visually significant edges and object boundaries from the target image to obtain boundary information related to the product in the target image, thereby achieving more accurate semantic segmentation of the target image. Based on the edge feature extraction branch, semantic feature extraction and edge feature extraction tasks can be performed simultaneously on the target image.
[0103] The edge feature extraction branch can include multiple dilated convolution units. Dilated convolution can be simply understood as adding spaces (with weights of 0) between the convolution kernel elements to expand the convolution kernel. Assuming a variable a is used to measure the dilation coefficient of the dilated convolution, the relationship between the actual convolution kernel size after adding the dilation and the original kernel size is: K = k + (k-1)(a-1). Here, k is the original kernel size, a is the dilation rate, and K is the actual kernel size after dilation. Otherwise, dilated convolution operates in the same manner as conventional convolution. The primary function of dilated convolution is to increase the kernel's receptive field by increasing the kernel dilation rate. This increases the receptive field without requiring pooling, preserving the image's feature information as much as possible. Therefore, it can be used in place of pooling layers. Furthermore, dilated convolution does not increase the number of network parameters, so while increasing the receptive field, it does not increase the computational load for training. Therefore, the edge feature extraction branch can increase the receptive field while retaining the original detail information in the target image to make up for the edge feature loss defect caused by the pooling module.
[0104] Figure 6 The figure schematically shows the principle of the dilated convolution according to an embodiment of the present application.
[0105] Figure 6 (a) in the figure corresponds to a 3×3 ordinary convolution. Figure 6(b) in the figure represents a 3×3 hole convolution with a dilation rate of 2. Figure 6 (c) in the figure represents a 3×3 dilated convolution with a dilation rate of 4. The 3×3 convolution kernel has only 9 weight values, just like a normal convolution, and the weights of the dilated parts are filled with 0. Figure 6 For (b), if the expansion rate a=2, then the actual convolution kernel size after adding the hole is 3+(3-1)*(2-1)=5, and the corresponding receptive field increases to (2^(a+2))-1=7. Figure 6 For (c), if the dilation rate a=4, then the actual convolution kernel size after adding the hole is 3+(3-1)*(4-1)=5, and the corresponding receptive field increases to (2^(a+2))-1=15. That is to say Figure 6 (b) and (c) in the figure can increase the receptive field to 7×7 and 15×15 respectively through the above operations. Therefore, the dilated convolution can increase the receptive field of the output unit without increasing the number of parameters.
[0106] Figure 7 The network structure diagram of the edge feature extraction branch according to an embodiment of the present application is schematically shown.
[0107] like Figure 7 As shown, the edge feature extraction branch may include cascaded five hole convolution units 710M, five third convolution units ( Figure 7 The figure is marked as 1x1Conv), and the fourth convolution unit ( Figure 7 The third convolution unit and the fourth convolution unit can both be selected as 1×1 convolution layers.
[0108] like Figure 7 As shown, the first atrous convolution unit 710M can be used to extract features from the target image 701 to obtain a first initial edge feature image 702. The second atrous convolution unit 710M can be used to extract features from the first initial edge feature image 702 to obtain a second initial edge feature image 703. The third atrous convolution unit 710M can be used to extract features from the second initial edge feature image 703 to obtain a third initial edge feature image 704. The fourth atrous convolution unit 710M can be used to extract features from the third initial edge feature image 704 to obtain a fourth initial edge feature image 705. The fifth atrous convolution unit 710M can be used to extract features from the fourth initial edge feature image 705 to obtain a fifth initial edge feature image 706.
[0109] like Figure 7As shown, the first initial edge feature image 702 can be processed by the first third convolution unit to obtain a first initial fused edge feature image. The second initial edge feature image 703 can be processed by the second third convolution unit to obtain a second initial fused edge feature image. The third initial edge feature image 704 can be processed by the third third convolution unit to obtain a third initial fused edge feature image. The fourth initial edge feature image 705 can be processed by the fourth third convolution unit to obtain a fourth initial fused edge feature image. The fifth initial edge feature image 706 can be processed by the fifth third convolution unit to obtain a fifth initial fused edge feature image. The first initial fused edge feature image 702 to the fifth initial fused edge feature image 706 can be superimposed to obtain a fused edge feature image. The fused edge feature image can be processed by the fourth convolution unit to obtain an edge feature image 707. The purpose of this is to fully utilize the information output by each dilated convolution unit to obtain more edge feature information in the target image.
[0110] According to an embodiment of the present application, the pooling module includes P parallel pooling units, P fifth convolution units, P upsampling units, and a feature fusion unit, where P is a positive integer; based on the pooling module, pooling operations are performed on multiple sub-regions of the semantic feature image respectively to obtain a context feature image group including: based on the P pooling units, pooling processing is performed on the semantic feature images respectively to obtain P pooling feature images; based on the pth fifth convolution unit, down-channel processing is performed on the pth pooling feature image output by the pth pooling unit to obtain the pth reduced dimensionality feature image, p=1,…,P; based on the pth upsampling unit, upsampling processing is performed on the pth reduced dimensionality feature image output by the pth fifth convolution unit to obtain the pth context feature image; based on the feature fusion unit, the 1st context feature image is fused to the Pth context feature image to obtain a context feature image group.
[0111] In one embodiment, the pooling module may include four parallel pooling units, four fifth convolution units, four upsampling units, and a feature fusion unit. Among them, the four fifth convolution units correspond one-to-one to the four pooling units, and the four upsampling units correspond one-to-one to the four pooling units. Exemplarily, among the aforementioned four pooling units, the first pooling unit can be used to perform a global pooling operation on the semantic feature image, and the second to fourth pooling units can correspond to three sub-regions of different sizes in the semantic feature image, and are used to perform pooling operations on the three sub-regions respectively.
[0112] As just an example, the scales of the four pooling units can be different. The semantic feature images can be pooled at different scales based on the four pooling units to obtain four pooled feature images of different sizes.
[0113] For example, the fifth convolution unit can be selected as a convolution layer with a convolution kernel size of 1×1 and a filter number of 1. Based on the first fifth convolution unit, the first pooled feature image output by the first pooling unit can be subjected to channel reduction processing to obtain a first reduced-dimensionality feature image. Based on the second fifth convolution unit, the second pooled feature image output by the second pooling unit can be subjected to channel reduction processing to obtain a second reduced-dimensionality feature image. Based on the third fifth convolution unit, the third pooled feature image output by the third pooling unit can be subjected to channel reduction processing to obtain a third reduced-dimensionality feature image. Based on the fourth fifth convolution unit, the fourth pooled feature image output by the fourth pooling unit can be subjected to channel reduction processing to obtain a fourth reduced-dimensionality feature image. The number of channels of the aforementioned first to fourth reduced-dimensionality feature images is 1, and their sizes are 1×1, 2×2, 3×3, and 6×6, respectively. The purpose of this step is to ensure that the global feature map has a high proportion in the final output feature map.
[0114] Based on the first upsampling unit, the first reduced dimensionality feature image output by the first fifth convolution unit can be upsampled to obtain the first context feature image. Based on the second upsampling unit, the second reduced dimensionality feature image output by the second fifth convolution unit can be upsampled to obtain the second context feature image. Based on the third upsampling unit, the third reduced dimensionality feature image output by the third fifth convolution unit can be upsampled to obtain the third context feature image. Based on the fourth upsampling unit, the fourth reduced dimensionality feature image output by the fourth fifth convolution unit can be upsampled to obtain the fourth context feature image. The sizes of the first to fourth context feature images are restored to be the same as the semantic feature image. Based on the feature fusion unit, the first to fourth context feature images can be fused to obtain a context feature image group.
[0115] In an optional embodiment, the upsampling unit can be a sub-pixel convolution upsampling unit. Sub-pixels can be understood as smaller units between two physical pixels. The basic principle of sub-pixel convolution upsampling is to increase the number of channels of the input feature map through convolution operations, and then reorganize these channels into spatial dimensions, thereby achieving an increase in image resolution.
[0116] Exemplarily, sub-pixel convolution upsampling may include the following steps:
[0117] 1. Convolution operation: Convolve the input feature map to generate a channel number R 2 ×C feature map, where R is the upsampling factor and C is the number of channels of the input feature map. This step extracts high-level features of the input feature map through the convolution kernel and increases the number of channels to prepare for subsequent spatial reorganization.
[0118] 2. Channel reorganization: R 2 The pixels at the same position of the feature maps are arranged into an R×R feature map. Specifically, each R×R channel block is rearranged into an R×R spatial block, thereby reducing the size of the feature map from H×W×(R 2 ×C) is converted to R×H×R×W×C.
[0119] 3. Sub-pixel distribution: The reorganized feature map forms an R×R sub-pixel distribution at each position, which together constitute the pixel at that position in the output layer. In this way, each input pixel is expanded into R×R sub-pixels, achieving R-fold upsampling.
[0120] Sub-pixel convolution upsampling can effectively improve the resolution and detail of the image without significantly increasing the amount of computation. Compared with upsampling methods such as deconvolution and interpolation, sub-pixel convolution upsampling has stronger image reconstruction capabilities, better local detail restoration, and higher output image quality.
[0121] Figure 8 A network diagram of an image semantic segmentation model according to an embodiment of the present application is schematically shown.
[0122] like Figure 8 As shown, the image semantic segmentation model can include a semantic feature extraction branch and an edge feature extraction branch. The semantic feature extraction branch can include a pooling module and a backbone network based on an attention mechanism, and the edge feature extraction branch can include multiple hole convolution units.
[0123] like Figure 8 As shown, the semantic feature extraction branch and the edge feature extraction branch can process the input target image synchronously. For example, the target image 801 can be feature extracted based on the MobileNet model embedded with the CBAM module to obtain a semantic feature image ( Figure 8 For example, a multi-scale pooling operation can be performed on the semantic feature image based on the pyramid pooling module (PPM module), and multiple context feature images ( Figure 8 Sub-pixel convolution upsampling and feature fusion operations can be performed on multiple context feature images to obtain a context feature image group ( Figure 8 For example, feature extraction can be performed on the target image 801 based on the edge feature extraction branch to obtain the first pooled feature image to the fifth pooled feature image ( Figure 8 The 1st to 5th pooled feature images can be convolved with 1×1 layers respectively and then concatenated, and then convolved with 1×1 layers again to obtain edge feature images ( Figure 8 Shown as ⑨).
[0124] like Figure 8 As shown, the semantic feature image, the context feature image group, and the edge feature image can be concatenated and fused (concat fusion) to obtain a fused feature image. The fused feature image can be input into a convolutional layer for feature extraction, and the feature image output by the convolutional layer is input into an activation function layer to obtain the final output semantic segmentation image 802.
[0125] Figure 9 The structural block diagram of the image detection device according to an embodiment of the present application is schematically shown.
[0126] like Figure 9 As shown, the apparatus 900 includes an acquisition module 910 , a first processing module 920 , a second processing module 930 and a determination module 940 .
[0127] The acquisition module 910 is used to acquire the target image to be detected. The first processing module 920 is used to perform feature extraction on the target image based on the semantic feature extraction branch of the image semantic segmentation model to obtain a semantic feature image and a context feature image group, wherein the semantic feature extraction branch includes a pooling module and a backbone network based on the attention mechanism, the semantic feature image is obtained by processing the target image using the backbone network, and the context feature image group is obtained by processing the semantic feature image using the pooling module. The second processing module 930 is used to perform feature extraction on the target image based on the edge feature extraction branch of the image semantic segmentation model to obtain an edge feature image, and the edge feature extraction branch includes a plurality of cascaded hole convolution units. The determination module 940 is used to determine the image detection result based on the reference semantic segmentation image, the semantic feature image, the context feature image group and the edge feature image, and the image detection result is used to locate the position of the abnormality in the target image.
[0128] According to an embodiment of the present application, the first processing module 920 may include a first processing sub-module and a second processing sub-module.
[0129] The first processing submodule is used to extract features from the target image based on the backbone network to obtain a semantic feature image. The second processing submodule is used to perform pooling operations on multiple regions of the semantic feature image based on the pooling module to obtain a context feature image group, wherein the sizes of the multiple regions are different.
[0130] According to an embodiment of the present application, the backbone network includes N first convolution units, N convolution attention modules, and at least one second convolution unit, where N is an integer greater than 1; the first processing submodule may include a first processing unit, a second processing unit, a third processing unit, a fourth processing unit, and a fifth processing unit.
[0131] The first processing unit is configured to extract features from the target image using the first first convolution unit to obtain a first initial semantic feature image. The second processing unit is configured to extract features from the first initial semantic feature image using the first convolution attention module to obtain a first intermediate semantic feature image. The third processing unit is configured to extract features from the n-1th intermediate semantic feature image using the nth first convolution unit to obtain an nth initial semantic feature image, where n=2,…,N. The fourth processing unit is configured to extract features from the nth initial semantic feature image using the nth convolution attention module to obtain an nth intermediate semantic feature image. The fifth processing unit is configured to extract features from the Nth intermediate semantic feature image based on at least one second convolution unit to obtain a semantic feature image.
[0132] According to an embodiment of the present application, the convolutional attention module includes a channel attention submodule and a spatial attention submodule, and the fourth processing unit may include a first processing subunit, a second processing subunit, a third processing subunit and a fourth processing subunit.
[0133] The first processing sub-unit is configured to input the nth initial semantic feature image into the channel attention sub-module to obtain the channel attention weight. The second processing sub-unit is configured to input the nth initial semantic feature image into the spatial attention sub-module to obtain the spatial attention weight. The third processing sub-unit is configured to weight the nth initial semantic feature image using the channel attention weight to obtain the attention feature image. The fourth processing sub-unit is configured to weight the attention feature image using the spatial attention weight to obtain the nth intermediate semantic feature image.
[0134] According to an embodiment of the present application, the edge feature extraction branch includes M cascaded hole convolution units, M third convolution units, and a fourth convolution unit, where M is an integer greater than 1; the second processing module 930 may include a third processing sub-module, a fourth processing sub-module, a fifth processing sub-module, a sixth processing sub-module and a seventh processing sub-module.
[0135] The third processing submodule is configured to extract features from the target image using the first dilated convolution unit to obtain a first initial edge feature image. The fourth processing submodule is configured to extract features from the m-1th initial edge feature image using the mth dilated convolution unit to obtain an mth initial edge feature image, where m = 2, ..., M. The fifth processing submodule is configured to process the first to the Mth initial edge feature images based on the M third convolution units to obtain first to Mth initial fused edge feature images. The sixth processing submodule is configured to fuse the first to the Mth initial fused edge feature images to obtain a fused edge feature image. The seventh processing submodule is configured to process the fused edge feature image using the fourth convolution unit to obtain an edge feature image.
[0136] According to an embodiment of the present application, the pooling module includes P parallel pooling units, P fifth convolution units, P upsampling units, and a feature fusion unit, where P is a positive integer; the second processing submodule may include a sixth processing unit, a seventh processing unit, an eighth processing unit, and a ninth processing unit.
[0137] The sixth processing unit is configured to perform pooling processing on the semantic feature images based on the P pooling units, respectively, to obtain P pooled feature images. The seventh processing unit is configured to perform down-channel processing on the p-th pooled feature image output by the p-th pooling unit, based on the p-th fifth convolution unit, to obtain the p-th reduced-dimensionality feature image, where p = 1, ..., P. The eighth processing unit is configured to perform upsampling processing on the p-th reduced-dimensionality feature image output by the p-th fifth convolution unit, based on the p-th upsampling unit, to obtain the p-th contextual feature image. The ninth processing unit is configured to fuse the first contextual feature image with the p-th contextual feature image based on the feature fusion unit, to obtain a contextual feature image group.
[0138] According to an embodiment of the present application, the determination module 940 may include a fusion submodule, a decoding submodule, and a comparison submodule.
[0139] The fusion submodule is used to fuse the semantic feature image, the context feature image group, and the edge feature image to obtain a fused feature image. The decoding submodule is used to decode the fused feature image to obtain a semantic segmentation image corresponding to the target image. The comparison submodule is used to compare the semantic segmentation image with the reference semantic segmentation image to determine the image detection result.
[0140] According to embodiments of the present application, any multiple modules among the acquisition module 910, the first processing module 920, the second processing module 930, and the determination module 940 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the acquisition module 910, the first processing module 920, the second processing module 930, and the determination module 940 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the acquisition module 910 , the first processing module 920 , the second processing module 930 and the determination module 940 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0141] Figure 10 A block diagram of an electronic device suitable for implementing the image detection method according to an embodiment of the present application is schematically shown.
[0142] like Figure 10 As shown, the electronic device 1000 according to an embodiment of the present application includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage unit 1008 into a random access memory (RAM) 1003. The processor 1001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.
[0143] Various programs and data required for the operation of the electronic device 1000 are stored in the RAM 1003. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The processor 1001 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 1002 and / or the RAM 1003. It should be noted that the programs may also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in the one or more memories.
[0144] According to an embodiment of the present application, electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to bus 1004. Electronic device 1000 may also include one or more of the following components connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or modem. Communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 1010 as needed, so that computer programs read from the removable media can be installed into storage section 1008 as needed.
[0145] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0146] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 1002 and / or RAM 1003 described above and / or one or more memories other than ROM 1002 and RAM 1003.
[0147] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the image detection method provided in the embodiments of the present application.
[0148] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the computer program is executed by the processor 1001. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0149] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 1009, and / or installed from the removable medium 1011. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0150] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0151] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
[0152] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. An image detection method, characterized in that: The method comprises: Acquire the target image to be detected; Based on the semantic feature extraction branch of the image semantic segmentation model, feature extraction is performed on the target image to obtain a semantic feature image and a context feature image group, wherein the semantic feature extraction branch includes a pooling module and a backbone network based on an attention mechanism, the semantic feature image is obtained by processing the target image using the backbone network, and the context feature image group is obtained by processing the semantic feature image using the pooling module; Based on the edge feature extraction branch of the image semantic segmentation model, feature extraction is performed on the target image to obtain an edge feature image, wherein the edge feature extraction branch includes multiple hole convolution units; According to the reference semantic segmentation image, an image detection result is determined based on the semantic feature image, the context feature image group and the edge feature image, and the image detection result is used to locate the position of the abnormality in the target image.
2. The method according to claim 1, characterized in that The semantic feature extraction branch based on the image semantic segmentation model performs feature extraction on the target image to obtain a semantic feature image and a context feature image group, including: Performing feature extraction on the target image based on the backbone network to obtain the semantic feature image; Based on the pooling module, a pooling operation is performed on multiple regions of the semantic feature image respectively to obtain the context feature image group, wherein the sizes of the multiple regions are different.
3. The method according to claim 2, characterized in that The backbone network includes N first convolutional units, N convolutional attention modules, and at least one second convolutional unit, where N is an integer greater than 1; extracting features of the target image based on the backbone network to obtain the semantic feature image includes: Using the first convolution unit to extract features from the target image to obtain a first initial semantic feature image; Using a first convolutional attention module to perform feature extraction on the first initial semantic feature image to obtain a first intermediate semantic feature image; Use the nth first convolution unit to extract features from the n-1th intermediate semantic feature image to obtain the nth initial semantic feature image, where n=2,…,N; Using the nth convolutional attention module to extract features from the nth initial semantic feature image to obtain the nth intermediate semantic feature image; and Feature extraction is performed on the Nth intermediate semantic feature image based on the at least one second convolution unit to obtain the semantic feature image.
4. The method according to claim 3, characterized in that The convolutional attention module includes a channel attention submodule and a spatial attention submodule. The nth convolutional attention module is used to extract features from the nth initial semantic feature image to obtain the nth intermediate semantic feature image. Inputting the nth initial semantic feature image into the channel attention submodule to obtain a channel attention weight; Inputting the nth initial semantic feature map into the spatial attention submodule to obtain a spatial attention weight; weighting the nth initial semantic feature image using the channel attention weight to obtain an attention feature image; The attention feature image is weighted using the spatial attention weight to obtain the nth intermediate semantic feature image.
5. The method according to claim 1, wherein The edge feature extraction branch includes M cascaded hole convolution units, M third convolution units, and a fourth convolution unit, where M is an integer greater than 1; the edge feature extraction branch based on the image semantic segmentation model extracts features from the target image to obtain an edge feature image including: Using the first dilated convolution unit to extract features from the target image to obtain a first initial edge feature image; Use the m-th hole convolution unit to extract the features of the m-1-th initial edge feature image to obtain the m-th initial edge feature image, where m=2,…,M; Based on the M third convolution units, processing the first initial edge feature image to the M-th initial edge feature image respectively to obtain the first initial fused edge feature image to the M-th initial fused edge feature image; fusing the first initial fused edge feature image with the Mth initial fused edge feature image to obtain a fused edge feature image; The fused edge feature image is processed by using the fourth convolution unit to obtain the edge feature image.
6. The method according to claim 2, characterized in that The pooling module includes P parallel pooling units, P fifth convolution units, P upsampling units, and feature fusion units, where P is a positive integer; the pooling operation is performed on multiple regions of the semantic feature image based on the pooling module to obtain the context feature image group, which includes: Performing pooling processing on the semantic feature images based on the P pooling units to obtain P pooled feature images; Based on the p-th fifth convolution unit, the p-th pooled feature image output by the p-th pooling unit is subjected to channel reduction processing to obtain the p-th reduced dimension feature image, where p = 1, ..., P; Based on the p-th upsampling unit, upsampling the p-th reduced-dimensionality feature image output by the p-th fifth convolution unit is performed to obtain the p-th context feature image; Based on the feature fusion unit, the first context feature image to the Pth context feature image are fused to obtain the context feature image group.
7. The method according to any one of claims 1 to 6, characterized in that The determining of the image detection result based on the reference semantic segmentation image, the semantic feature image, the context feature image group, and the edge feature image comprises: fusing the semantic feature image, the context feature image group, and the edge feature image to obtain a fused feature image; Decoding the fused feature image to obtain a semantic segmentation image corresponding to the target image; The semantic segmentation image is compared with the reference semantic segmentation image to determine an image detection result.
8. An image detection device, characterized in that: The device comprises: An acquisition module, used for acquiring a target image to be detected; a first processing module, configured to perform feature extraction on the target image based on a semantic feature extraction branch of the image semantic segmentation model to obtain a semantic feature image and a context feature image group, wherein the semantic feature extraction branch includes a pooling module and a backbone network based on an attention mechanism, the semantic feature image is obtained by processing the target image using the backbone network, and the context feature image group is obtained by processing the semantic feature image using the pooling module; A second processing module is configured to perform feature extraction on the target image based on an edge feature extraction branch of the image semantic segmentation model to obtain an edge feature image, wherein the edge feature extraction branch includes a plurality of cascaded dilated convolution units; A determination module is used to determine an image detection result based on a reference semantic segmentation image, the semantic feature image, the context feature image group and the edge feature image, wherein the image detection result is used to locate the position of the abnormality in the target image.
9. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.