Image Salient Object Detection Method Based on Edge Completeness and Clarity and Dual-Stream Network
By building a dual-stream deep network based on VGG16 model, combining position feature extraction and edge enhancement modules, the problem of insufficient edge clarity and integrity in image significant object detection is solved, and efficient and accurate significant object detection is achieved.
Patent Information
- Application Number
- CN202310533685.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-05-11
AI Technical Summary
In the prior art, image significant object detection has problems with low clarity and incompleteness of the target edge, and the real-time detection efficiency is low and the detection accuracy is insufficient.
A dual-stream deep network based on the VGG16 model is built, including the first feature extraction subnet and the second feature extraction subnet. The position feature extraction module and edge enhancement module are used to combine the traditional algorithms HC and RC significance graphs to train the network through custom loss functions to improve edge clarity and integrity.
The accuracy and speed of significant object detection are improved, the problems of edge blur and incompleteness in the prior art are overcome, and fast and accurate image significant object detection is achieved.
Smart Images

Figure CN116758302B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a method for detecting salient objects in images based on edge integrity and clarity and a two-stream network in the field of image target detection. The present invention is used to detect salient objects in a single image and ensure the clarity and integrity of the edges of the detected objects, and can be applied to scenarios such as object detection, identity recognition, and image quality evaluation. Background Art
[0002] The purpose of image salient object detection is to identify the prominent objects in an image scene that are most likely to attract human visual attention, and it is mainly applied in fields such as image compression, image quality evaluation, image segmentation, target detection, and recognition. The development of the image salient object detection task can be divided into two eras: traditional methods and deep learning methods. Traditional saliency detection methods extract various handcrafted features such as color, brightness, texture, etc. They retain the edges of salient objects well, but the integrity of the objects is not high; deep learning saliency detection models use their powerful feature extraction capabilities to locate salient objects, improving the integrity of salient objects, but the object boundaries are relatively rough.
[0003] South China University of Technology discloses a method for detecting salient objects in an image in its patent document "A Method for Image Saliency Detection Based on a Parallel Convolutional Neural Network" (application number CN 201710253255.2, publication number CN 107169954 A). This method first designs a parallel convolutional neural network structure, including a global angle detection module and a local angle detection module, and the two modules are parallelized through a fully connected layer. At the same time, two network input graphs are designed: a global padding graph and a local cropping graph, and a superpixel-based label is defined for the input, and both padding graphs are centered on superpixels. Among them, the global padding graph contains all the information of the original image, representing global features, and is used as the input of the global angle detection module. The local padding graph is a cropping graph containing the detailed information of the superpixel neighborhood and is used as the input of the local angle detection module. This method detects saliency from both global and local perspectives and can effectively detect the internal semantics and background differences of the salient object body. However, the deficiencies of this method are still: the preprocessing steps of the network input are relatively cumbersome, affecting the detection efficiency; the network does not consider edge information features, resulting in blurred and incomplete object edges after detection.
[0004] Xi'an Jiaotong University discloses an image salient object detection method based on boundary enhancement in its patent document "A Salient Object Detection Method and System Based on Boundary Enhancement" (application number: CN 202210467623.4, publication number: CN 114821059 A). This method designs a convolutional neural network model, extracts abstract feature maps with different resolutions from the training set images, performs multi-level fusion on the abstract feature maps to obtain multi-level fusion feature maps, processes the multi-level fusion feature maps to obtain a feature map containing multi-scale information; performs information transformation on the feature map containing multi-scale information and then splices and fuses it to obtain boundary information features. At the same time, a boundary detection result is obtained using the features after each level of transformation, and then further fused to obtain a fused boundary detection result; the feature map containing multi-scale information is subjected to multi-scale information extraction and then spliced with the feature containing boundary information to obtain a salient object detection result. However, the deficiencies of this method are still as follows: Using ResNet-50 as the network backbone, too many model parameters will lead to an increase in training time and a decrease in detection rate; although the network enhances the clarity of edge pixels, it does not guarantee the integrity of the edges. Summary of the Invention
[0005] The object of the present invention is to address the deficiencies of the above-mentioned prior art and propose an image salient object detection method based on edge integrity and clarity and a dual-stream network, aiming to solve the problems of low clarity and incompleteness of the object edges in the prior art.
[0006] The idea for achieving the object of the present invention is as follows: The present invention constructs a dual-stream deep network. The first feature extraction sub-network is an encoder-decoder structure with the VGG16 model as the backbone. A position feature extraction module is added at the end of the encoder to extract multi-scale features and position information of the salient object, and the output features of the third, fourth, and fifth layers at the encoder end are used as the training branches for salient object extraction, and the output feature of the second layer is used as the training branch for salient edge extraction. Joint supervision is used for model training. Since this network separates the salient edge extraction branch and integrates the salient object, the clarity of the object edge is improved. The second feature extraction sub-network in the dual-stream deep network constructed by the present invention is a multi-convolutional pooling layer network structure with grayscale images and traditional algorithms HC (Histogram-based Contrast) and RC (Region-based Contrast) for calculating the saliency map as the input. Utilizing the characteristic of the traditional algorithm saliency map to retain complete edges, multi-scale complete edge features are extracted and integrated into the training branch for salient object extraction, ensuring the integrity of the object edge.
[0007] Since the dual-stream deep network constructed in the present invention uses the VGG16 model as the backbone, compared with the ResNet-18 and ResNet-50 models used in the prior art, it has fewer network parameters and a faster rate of extracting significant targets in images, overcoming the problem of low real-time detection efficiency in the prior art. At the same time, it can detect significant targets in images without additional input processing. When training the dual-stream deep network, data augmentation technology is used. By randomly flipping, rotating, and centralizing the input images, the diversity of training samples is increased, the generalization ability of the model is improved, and the problem of low detection accuracy for images containing different features in the prior art is overcome.
[0008] The specific steps to implement the present invention are as follows:
[0009] Step 1, construct the first feature extraction sub-network:
[0010] Step 1.1, construct a position feature fusion module, the structure of which includes a first dimensionality reduction convolutional block, a second dimensionality reduction convolutional block, and a third dimensionality reduction convolutional block; among them, the structures of the first to third dimensionality reduction convolutional blocks are the same, and each dimensionality reduction convolutional block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and an upsampling layer connected in series; set the convolutional kernel sizes of the first to third convolutional layers in the first to third dimensionality reduction convolutional blocks to 3×3, the padding numbers to 1, the resolutions of the upsampling layers to 32×32, 64×64, and 128×128 respectively, and the activation layers are all implemented using the ReLU activation function; set the numbers and channel numbers of the convolutional kernels of the first to third convolutional layers in the first dimensionality reduction convolutional block to 512; set the channel numbers of the convolutional kernels of the first to third convolutional layers in the second dimensionality reduction convolutional block to 512, 256, and 256 respectively, and the numbers of convolutional kernels to 256; set the channel numbers of the convolutional kernels of the first to third convolutional layers in the third dimensionality reduction convolutional block to 256, 128, and 128 respectively, and the numbers of convolutional kernels to 128;
[0011] Step 1.2, construct an inter-layer feature aggregation module, the structure of which includes a first aggregation convolutional block and a second aggregation convolutional block; among them, the structures of the first and second aggregation convolutional blocks are the same, and each aggregation convolutional block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and an upsampling layer connected in series; set the convolutional kernel sizes of the first to third convolutional layers in the first and second aggregation convolutional blocks to 3×3, the padding numbers to 1, the resolutions of the upsampling layers to 32×32 and 64×64 respectively, and the activation layers are all implemented using the ReLU activation function; set the numbers and channel numbers of the convolutional kernels of the first to third convolutional layers in the first aggregation convolutional block to 512; set the channel numbers of the convolutional kernels of the first to third convolutional layers in the second aggregation convolutional block to 512, 256, and 256 respectively, and the numbers of convolutional kernels to 256;
[0012] Step 1.3: Construct an edge enhancement module, the structure of which includes an edge convolution block, a first upsampling convolution block, a second upsampling convolution block, and a third upsampling convolution block; among them, the edge convolution block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, and a third activation layer connected in series; the first to third upsampling convolution blocks have the same structure, and each upsampling convolution block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and a max pooling layer connected in series; set the size of the convolutional kernels of the first to third convolutional layers in the edge convolution block to 3×3, the padding number to 1, the number of channels and the number of convolutional kernels to 128, and the activation layers are all implemented using the ReLU function; set the size of the convolutional kernels of the first to third convolutional layers in the first to third upsampling convolution blocks to 3×3, the padding number to 1, the size of the pooling kernel of the pooling layer to 3×3, the stride to 2, the padding number to 1, and the activation layers are all implemented using the ReLU function; set the number of channels of the convolutional kernels of the first to third convolutional layers in the first upsampling convolution block to 128, 256, and 256 respectively, and the number of convolutional kernels to 256; set the number of channels of the convolutional kernels of the first to third convolutional layers in the second upsampling convolution block to 256, 512, and 512 respectively, and the number of convolutional kernels to 256; in the third upsampling convolution block, the number of channels and the number of convolutional kernels of the first to third convolutional layers are both set to 512;
[0013] Step 1.4, construct a feature fusion module, the structure of which includes an edge fusion convolutional block, a first fusion convolutional block, a second fusion convolutional block, and a third fusion convolutional block; among them, the edge fusion convolutional block is composed of convolutional layers; the first to third fusion convolutional blocks have the same structure, and each fusion convolutional block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, a fourth convolutional layer, and an upsampling layer; set the convolutional kernel size of the first convolutional layer in the edge convolutional block to 1×1, the padding number to 1, the number of channels of the convolutional kernel to 128, and the number of convolutional kernels to 1; for each of the first to fourth convolutional layers in the first to third fusion convolutional blocks, the convolutional kernel sizes of the first to fourth convolutional layers in the third to fifth fusion convolutional blocks are all set to 3×3, the padding number is set to 1, the resolution size of the upsampling layer is all set to 256×256, and the activation layers all use the ReLU function; set the number of channels of the convolutional kernels of the first to fourth convolutional layers in the first fusion convolutional block to 256, and the numbers of convolutional kernels are 256, 256, 256, and 1 respectively; set the number of channels of the convolutional kernels of the first to fourth convolutional layers in the second fusion convolutional block to 512, and the numbers of convolutional kernels are 512, 512, 512, and 1 respectively; set the number of channels of the convolutional kernels of the first to fourth convolutional layers in the third fusion convolutional block to 512, and the numbers of convolutional kernels are 512, 512, 512, and 1 respectively;
[0014] Step 1.5, connect in series the first feature extraction module with VGG16 model as the backbone and the position feature extraction module composed of ASPP (Atrous Spatial Pyramid Pooling); the features output by the position feature module are integrated into the output branches of the second to fourth convolutional blocks in the first feature extraction module, connect in series the output features of the third to fourth convolutional blocks with the inter-layer aggregation module, connect in series the output features of the second convolutional block, the output features of the inter-layer aggregation module, and the edge enhancement fusion module, and after connecting the edge enhancement module and the feature fusion module in series, obtain the first feature extraction sub-network;
[0015] Step 2, construct a second feature extraction sub-network:
[0016] The second feature extraction sub-network is composed of a first feature convolutional block, a second feature convolutional block, a third feature convolutional block, a fourth feature convolutional block, and a fifth feature convolutional block connected in series in sequence; keep the parameters of the first to fifth feature convolutional blocks consistent with the parameters of the first feature extraction module constructed in Step 1.1;
[0017] Step 3, construct a two-stream deep network:
[0018] Step 3.1, connect the second feature extraction sub-network in the edge enhancement module of the first feature extraction sub-network; add and fuse the features output by the third to fifth convolutional blocks of the second feature extraction sub-network through the edge enhancement module;
[0019] Step 3.2, output a significant edge detection prediction map and three significant object detection prediction maps from the output of the edge enhancement module through the feature fusion module; add and fuse the significant object detection prediction maps and activate them to obtain the final image significant object map, completing the construction of the dual-stream deep network for image significant object detection;
[0020] Step 4, generate a training set:
[0021] Step 4.1, form samples by combining 1 color image with its corresponding 1 significant object ground truth map, and select at least 10553 samples to form a sample set;
[0022] Step 4.2, generate the corresponding significant edge ground truth map, grayscale map, HC significant map, and RC significant map for each sample respectively;
[0023] Step 4.3, preprocess the sample set and the corresponding significant edge ground truth map, grayscale map, HC significant map, and RC significant map for each sample to form a training set;
[0024] Step 5, train the dual-stream deep network:
[0025] Input the training set into the dual-stream deep network, calculate the total loss of the significant edge loss and the significant object loss using the total loss function, and output the overall loss value of the network; through backpropagation of the overall loss value, use the Adaptive Moment Estimation (Adam) optimizer to iteratively update the weight parameters of the network nodes until the overall loss value converges to obtain the trained dual-stream deep network;
[0026] Step 6, detect the significant object in the image:
[0027] Input the color image to be detected and its generated grayscale map, HC significant map, and RC significant map into the trained dual-stream deep network for image significant object, and output a single-channel significant object grayscale image.
[0028] Compared with the prior art, the present invention has the following advantages:
[0029] First, since the present invention separates the significant object edge training branch network to assist in the training of the supervision model and integrates the extracted significant edge features into the significant object, it overcomes the deficiency of the blurred significant object edge pixels extracted by the prior art, making the object edge extracted by the present invention have higher clarity and improving the accuracy of significant object detection.
[0030] Second, since the present invention designs a multi-layer convolutional pooling network with the grayscale image of the original image and the saliency maps of traditional algorithms HC and RC as inputs, extracts the multi-scale complete edge features in the saliency maps of traditional algorithms, and integrates them into the salient objects, overcoming the defect that the edges of the salient objects extracted by the prior art are incomplete, making the edges of the objects extracted by the present invention more complete and improving the accuracy of salient object detection.
[0031] Third, since the dual-stream deep network constructed by the present invention uses the VGG16 model as the backbone, has few network parameters, and extracts the salient objects of images at a fast rate, overcoming the deficiency of low real-time detection efficiency in the prior art, and at the same time does not require excessive additional input processing, enabling the present invention to quickly detect images with different resolutions and having a wider application scenario.
[0032] Fourth, since the present invention uses data augmentation technology for the training set when training the dual-stream deep network, overcoming the defect of low detection accuracy for pictures containing different features in the prior art, making the present invention have the advantages of strong robustness and strong generalization. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic flow chart of the present invention;
[0034] Figure 2 is a schematic structural diagram of the dual-stream deep network designed by the present invention;
[0035] Figure 3 is a schematic structural diagram of the position feature extraction module in the first feature sub-network designed by the present invention.
[0036] Figure 4 is a schematic diagram of the content of the generated image dataset, where Figure 4 (a) is a color input image, Figure 4 (b) is a ground truth map of the salient object, Figure 4 (c) is a ground truth map of the extracted salient edge, Figure 4 (d) is the extracted grayscale image, Figure 4 (e) is the extracted HC saliency map, Figure 4 (f) is the extracted RC saliency map;
[0037] Figure 5 is a schematic diagram of the extraction result of the present invention, Figure 5 (a) is the detected salient object map, Figure 5 (b) is the detected salient edge map; DETAILED DESCRIPTION OF THE INVENTION
[0038] The following further describes the present invention with reference to the drawings and embodiments.
[0039] Refer to Figure 1, the implementation steps of the embodiments of the present invention are further described.
[0040] Step 1, construct the first feature extraction sub-network.
[0041] Refer to Figure 2 the first feature extraction sub-network in the dotted part in for a further detailed description of the network constructed by the present invention.
[0042] The first feature extraction sub-network includes a first feature extraction module, a position feature extraction module, a position feature fusion module, an inter-layer feature aggregation module, an edge enhancement module, and a feature fusion module.
[0043] Step 1.1, construct the first feature extraction module, and its structure is connected in series in turn as: a first feature convolution block, a second feature convolution block, a third feature convolution block, a fourth feature convolution block, a fifth feature convolution block, and an atrous convolution pooling block. Among them, the first feature convolution block is composed of a first convolution layer, a first activation layer, a second convolution layer, and a second activation layer connected in series; the second feature convolution block is composed of a max pooling layer, a first convolution layer, a first activation layer, a second convolution layer, and a second activation layer connected in series; the third to fifth feature convolution blocks have the same structure, and each feature convolution block is composed of a max pooling layer, a first convolution layer, a first activation layer, a second convolution layer, a second activation layer, a third convolution layer, and a third activation layer connected in series.
[0044] Set the convolution kernel sizes of the first and second convolution layers in the first feature convolution block to 3×3, the padding numbers to 1, the number of channels of the convolution kernels to 3 and 64 respectively, and the number of convolution kernels to 64 and 64 respectively; both the first and second activation layers are implemented using the ReLU activation function.
[0045] Set the pooling kernel size of the max pooling layer in the second feature convolution block to 3×3, the stride to 2, and the padding number to 1; set the convolution kernel sizes of the first and second convolution layers to 3×3, the padding numbers to 1, the number of channels of the convolution kernels to 64 and 128 respectively, and the number of convolution kernels to 128 and 128 respectively; both the first and second activation layers are implemented using the ReLU activation function.
[0046] Set the pooling kernel size of the max-pooling layers in the third to fifth feature convolution blocks to 3×3, the stride to 2, and the padding number to 1; set the convolution kernel size of the first to third convolutional layers to 3×3 and the padding number to 1; the first to third activation layers are all implemented using the ReLU activation function. In the third feature convolution block, the convolution kernel channels of the first to third convolutional layers are set to 128, 256, and 256 respectively, and the number of convolution kernels is set to 256, 256, and 256 respectively; in the fourth feature convolution block, the convolution kernel channels of the first to third convolutional layers are set to 256, 512, and 512 respectively, and the number of convolution kernels is set to 512, 512, and 512 respectively; in the fifth feature convolution block, the convolution kernel channels of the first to third convolutional layers are all set to 512, and the number of convolution kernels is all set to 512.
[0047] Step 1.2: Construct a position feature extraction module, the structure of which is successively connected in series as a max-pooling layer, a parallel dilated convolution block, and a dimensionality reduction convolution block. Among them, the parallel dilated convolution block is composed of a first convolution unit, a second convolution unit, a third convolution unit, a fourth convolution unit, and a fifth convolution unit connected in parallel. The structures of the first to fourth convolution units are the same, and each is composed of a convolutional layer, a batch normalization layer, and an activation layer connected in series. The fifth convolution unit is composed of an adaptive average pooling layer, a convolutional layer, a batch normalization layer, and an activation layer connected in series; the dimensionality reduction convolution block is composed of a convolutional layer, a batch normalization layer, and an activation layer connected in series.
[0048] The following combines Figure 3 to further describe the position feature extraction module.
[0049] For the max-pooling layer, set its pooling kernel size to 3×3, the stride to 1, and the padding number to 1.
[0050] In the parallel dilated convolution block, the convolution kernel sizes of the convolutional layers in the first to fourth convolution units are set to 1×1, 3×3, 3×3, and 3×3 respectively, the padding numbers are set to 0, 6, 12, and 18 respectively, the dilation rates are set to 1, 6, 12, and 18 respectively, the convolution kernel channels are set to 512, 512, 512, and 512 respectively, and the number of convolution kernels is set to 512; the channels of the batch normalization layers are all set to 512; the activation layers are all implemented using the ReLU activation function. In the fifth convolution unit, the pooling kernel size of the adaptive average pooling layer is set to 1×1, the convolution kernel size of the convolutional layer is set to 1×1, the convolution kernel channels are set to 512, and the number of convolution kernels is set to 512; the channels of the batch normalization layer are set to 512; the activation layer is implemented using the ReLU activation function.
[0051] In the dimensionality reduction convolution block, the size of the convolution kernel in the convolutional layer is set to 1×1, the number of channels of the convolution kernel is set to 2560, and the number of convolution kernels is set to 512; the number of channels of the batch normalization layer is set to 512; the activation layer is implemented using the ReLU activation function.
[0052] Step 1.3: Construct a position feature fusion module, whose structure includes a first dimensionality reduction convolution block, a second dimensionality reduction convolution block, and a third dimensionality reduction convolution block. Among them, the structures of the first to third dimensionality reduction convolution blocks are the same, and each dimensionality reduction convolution block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and an upsampling layer connected in series.
[0053] Set the size of the convolution kernels in the first to third convolutional layers in the first to third dimensionality reduction convolution blocks to 3×3, the padding number to 1, the resolutions of the upsampling layers to 32×32, 64×64, and 128×128 respectively, and the activation layers are all implemented using the ReLU activation function. In the first dimensionality reduction convolution block, the number of channels and the number of convolution kernels in the first to third convolutional layers are both set to 512; in the second dimensionality reduction convolution block, the number of channels of the convolution kernels in the first to third convolutional layers are set to 512, 256, and 256 respectively, and the number of convolution kernels is set to 256; in the third dimensionality reduction convolution block, the number of channels of the convolution kernels in the first to third convolutional layers are set to 256, 128, and 128 respectively, and the number of convolution kernels is set to 128.
[0054] Step 1.4: Construct an inter-layer feature aggregation module, whose structure includes a first aggregation convolution block and a second aggregation convolution block. Among them, the structures of the first and second aggregation convolution blocks are the same, and each aggregation convolution block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and an upsampling layer connected in series.
[0055] Set the size of the convolution kernels in the first to third convolutional layers in the first and second aggregation convolution blocks to 3×3, the padding number to 1, the resolutions of the upsampling layers to 32×32 and 64×64 respectively, and the activation layers are all implemented using the ReLU activation function. In the first aggregation convolution block, the number of channels and the number of convolution kernels in the first to third convolutional layers are both set to 512; in the second aggregation convolution block, the number of channels of the convolution kernels in the first to third convolutional layers are set to 512, 256, and 256 respectively, and the number of convolution kernels is set to 256.
[0056] Step 1.5: Construct an edge enhancement module, the structure of which includes an edge convolution block, a first upsampling convolution block, a second upsampling convolution block, and a third upsampling convolution block. Among them, the edge convolution block is composed of a first convolution layer, a first activation layer, a second convolution layer, a second activation layer, a third convolution layer, and a third activation layer connected in series; the first to third upsampling convolution blocks have the same structure, and each upsampling convolution block is composed of a first convolution layer, a first activation layer, a second convolution layer, a second activation layer, a third convolution layer, a third activation layer, and a max pooling layer connected in series.
[0057] Set the size of the convolution kernels of the first to third convolution layers in the edge convolution block to 3×3, the padding number to 1, the number of channels and the number of convolution kernels to 128, and the activation layers are all implemented using the ReLU function.
[0058] Set the size of the convolution kernels of the first to third convolution layers in the first to third upsampling convolution blocks to 3×3, the padding number to 1, the size of the pooling kernel of the pooling layer to 3×3, the stride to 2, and the padding number to 1. The activation layers are all implemented using the ReLU function. In the first upsampling convolution block, the number of channels of the convolution kernels of the first to third convolution layers are set to 128, 256, and 256 respectively, and the number of convolution kernels is set to 256; in the second upsampling convolution block, the number of channels of the convolution kernels of the first to third convolution layers are set to 256, 512, and 512 respectively, and the number of convolution kernels is set to 256; in the third upsampling convolution block, the number of channels and the number of convolution kernels of the first to third convolution layers are both set to 512.
[0059] Step 1.6: Construct a feature fusion module, the structure of which includes an edge fusion convolution block, a first fusion convolution block, a second fusion convolution block, and a third fusion convolution block. Among them, the edge fusion convolution block is composed of a convolution layer; the first to third fusion convolution blocks have the same structure, and each fusion convolution block is composed of a first convolution layer, a first activation layer, a second convolution layer, a second activation layer, a third convolution layer, a third activation layer, a fourth convolution layer, and an upsampling layer.
[0060] Set the size of the convolution kernel of the first convolution layer in the edge convolution block to 1×1, the padding number to 1, the number of channels of the convolution kernel to 128, and the number of convolution kernels to 1.
[0061] In the first to third fusion convolutional blocks, for the first to fourth convolutional layers, the kernel size of the first to fourth convolutional layers in each of the third to fifth fusion convolutional blocks is set to 3×3, the padding is set to 1, the resolution size of the upsampling layer is set to 256×256, and the activation layer is implemented using the ReLU function. In the first fusion convolutional block, the number of channels of the kernels of the first to fourth convolutional layers is set to 256, and the number of kernels is 256, 256, 256, 1 respectively; in the second fusion convolutional block, the number of channels of the kernels of the first to fourth convolutional layers is set to 512, and the number of kernels is 512, 512, 512, 1 respectively; in the third fusion convolutional block, the number of channels of the kernels of the first to fourth convolutional layers is set to 512, and the number of kernels is 512, 512, 512, 1 respectively.
[0062] Step 2, construct the second feature extraction sub-network.
[0063] Refer to Figure 2 The second feature extraction sub-network in the dashed part in [reference] further describes the network constructed in the present invention.
[0064] The second feature extraction sub-network is composed of second feature extraction modules, and its structure is successively connected in series as the first feature convolutional block, the second feature convolutional block, the third feature convolutional block, the fourth feature convolutional block, and the fifth feature convolutional block.
[0065] The parameters of the first to fifth feature convolutional blocks are the same as those of the first to fifth feature convolutional blocks constructed in step 1.1.
[0066] Step 3, construct a dual-stream deep network for image salient object detection.
[0067] For the convenience of describing the connection method of the first and second feature sub-networks, the present invention denotes the output features after the second to fifth convolutional blocks in the first feature extraction module of the input information flow as C2, C3, C4, C5, and the output features after the third to fifth convolutional blocks in the second feature extraction module of the information flow as S2, S3, S4.
[0068] Next, refer to Figure 2 to further describe in detail the connection method of the first and second feature sub-networks of the present invention.
[0069] Step 3.1, in the first feature extraction sub-network, connect the first feature extraction module in series with the position feature extraction module, and denote the feature output after C5 passes through the position feature extraction module as L.
[0070] Step 3.2, in the inter-layer feature aggregation module of the first feature extraction sub-network, the output feature L is divided into four paths for fusion. Among them, the first path of feature L is added to and fused with feature C5, and the output feature F5 is obtained; the second to fourth paths of feature L pass through the first to third dimensionality reduction convolutional blocks in the position feature extraction module respectively, and then are added to and fused with C4, C3, and C2 respectively. The fused output features are denoted as F4, F3, and F2 respectively.
[0071] Step 3.3, in the inter-layer feature aggregation module of the first feature extraction sub-network, feature F5 passes through the first aggregation convolutional block and is added to and fused with F4. The fused output feature is denoted as E4; feature E4 passes through the second aggregation convolutional block and is added to and fused with F3. The fused output feature is denoted as E3.
[0072] Step 3.4, in the edge enhancement module of the first feature extraction sub-network, the feature obtained after the feature F2 passes through the edge convolutional block is denoted as E2, and the features obtained after E2 passes through the first to third dimensionality increase convolutional blocks are denoted as I3, I4, and I5.
[0073] Step 3.5, in the edge enhancement module of the first feature extraction sub-network, connect the second feature extraction sub-network. Add and fuse features I3, E3, and S3, and denote the output feature as Z3; add and fuse features I4, E4, and S4, and denote the output feature as Z4; add and fuse features I5, E5, and S5, and denote the output feature as Z5.
[0074] Step 3.6, in the feature fusion module of the first feature extraction sub-network, the feature E2 passes through the edge fusion convolutional block to obtain the image significant edge map D; features Z3, Z4, and Z5 pass through the first to third fusion convolutional blocks and output the significant object prediction maps Y3, Y4, and Y5 respectively.
[0075] Step 3.7, Y3, Y4, and Y5 are added and fused to obtain the final image significant object map Y, completing the construction of the image significant object detection two-stream deep network.
[0076] Step 4, generate the training set.
[0077] The present invention selects all 10,553 color images and the corresponding significant object ground truth maps in the publicly available dataset DUTS-TR to form a sample set.
[0078] The embodiments of the present invention select a color image with a resolution of 300×400 and a binary image with a resolution of 300×400, as shown in Figure 4 (a), Figure 4 (b) respectively.
[0079] Since in the dual-stream deep network constructed in the present invention, the first feature extraction sub-network requires a significant edge ground truth map for edge loss calculation, and the second feature extraction sub-network requires the grayscale image, HC saliency map, and RC saliency map of the color image as inputs, the following further expansion of the sample set is required.
[0080] Step 4.1, generate the corresponding significant edge ground truth map according to the following formula, using Figure 4 (b) in the embodiment of the present invention as the input, and the extraction result is as shown in Figure 4 (c).
[0081]
[0082] Among them, represents the significant edge ground truth map, as shown in Figure 4 (c), x represents the significant object ground truth map, as shown in Figure 4 (b), Θ represents the binary morphological erosion operator, k represents the erosion structuring element, and its value is taken as a matrix of size 3×3:
[0083]
[0084] Step 4.2, generate a single-channel grayscale image, using Figure 4 (a) in the embodiment of the present invention as the input, and the extraction result is as shown in Figure 4 (d).
[0085] Step 4.3, generate a single-channel HC saliency image, using Figure 4 (a) in the embodiment of the present invention as the input, and the extraction result is as shown in Figure 4 (e).
[0086] Step 4.4, generate a single-channel RC saliency image, using Figure 4 (a) in the embodiment of the present invention as the input, and the extraction result is as shown in Figure 4 (f).
[0087] Step 4.5, reset the resolutions of the generated significant edge ground truth map, grayscale image, HC saliency map, and RC saliency map to 256×256 pixels, and perform random image flipping processing.
[0088] Step 4.6, subtract the mean value (104.00699, 115.66877, 122.67892) from the color images in the sample set.
[0089] Step 4.7, form a training set from the sample set, the significant edge ground truth map corresponding to each sample in the sample set, the grayscale image corresponding to each sample, the HC saliency map corresponding to each sample, and the RC saliency map corresponding to each sample.
[0090] Step 5, training the dual-stream deep network for salient object detection in images:
[0091] Input the training set into the dual-stream deep network, output a salient edge prediction map and three salient object prediction maps. Calculate the edge loss value between the salient edge prediction map and the corresponding ground truth salient edge map in the training set, calculate the object loss value between the three salient object prediction maps and the corresponding ground truth salient object maps in the training set, and use the total loss function to calculate the total loss of the salient edge loss and the salient object loss, and output the overall loss value of the network. Through backpropagation of the overall loss value, use the Adaptive Moment Estimation (Adam) optimizer to update the weight parameters of each network node until the overall loss value converges, and obtain the trained dual-stream deep network.
[0092] The total loss function is as follows:
[0093]
[0094] where L total represents the joint loss function, and L D represents the binary cross-entropy edge loss function.
[0095] represents the binary cross-entropy object loss function.
[0096] The binary cross-entropy edge loss function is as follows:
[0097] L D = -(1 - D') log(1 - D) - D' log(D)
[0098] where D' represents the ground truth salient edge map, and D represents the salient edge prediction map.
[0099] The binary cross-entropy object loss function is as follows:
[0100]
[0101] where Y' represents the ground truth salient object map, and Y3, Y3, Y3 are the three salient object prediction maps generated in Step 3.6.
[0102] Step 6, detecting the salient objects in the image.
[0103] Input the color image to be detected, its generated grayscale image, HC salient map, and RC salient map into the trained dual-stream deep network for salient object detection in images, and output a single-channel grayscale image of the salient object.
[0104] The effects of the present invention are further illustrated by the following simulation experiments.
[0105] 1. Simulation experiment conditions:
[0106] The hardware platform of the simulation experiment of the present invention is: the processor is Intel (R) Core i7-9700K CPU, the main frequency is 3.6GHz, the memory is 32GB, the graphics card is NVIDIA GeForce RTX 2080SUPER, the display
[0107] The storage is 8GB.
[0108] The software platform for the simulation experiment of the present invention is: 64-bit Windows 10 operating system and Python 3.7.
[0109] The input images used in the simulation experiments of the present invention are three-channel input color images with a size of 300×400 and a jpg format from the ECSSD public dataset. This dataset is the dataset published by Yan et al. in 2013 (The learning of "Hierarchical Image Saliency Detection on Extended CSSD" (J. Shi, Q. Yan, Li Xu, and J. Jia, IEEE Transactions on Pattern Analysis and Machine Intelligence, April 2016, pp. 717-729).
[0110] 2. Simulation experiment content and result analysis:
[0111] The simulation experiment of the present invention adopts the method of the present invention to simulate a picture in the ECSSD data set. Figure 4 The color image shown in (a) is used for salient object detection, and the results are as follows: Figure 5 (a) shows a grayscale image of a single-channel salient target prediction result with a size of 300×400 and a format of png, and Figure 5 (b) shows a grayscale image of a single-channel salient edge prediction result with a size of 300×400 and a png format.
[0112] The simulation results of the present invention show that: Figure 4 As shown in (a), the edge jaggedness of the salient target in the image detected by the present invention is smaller than that of the salient target ground truth map, and there is no obvious ghost around the target. The overlap rate with the salient edge ground truth map is high, indicating that the method of the present invention has high clarity in detecting salient targets and more complete edges.
Claims
1. An image salient object detection method based on edge integrity and clarity and a two-stream network, characterized in that Construct a deep learning network with a dual-branch composed of a first feature extraction sub-network and a second feature extraction sub-network, and generate a total loss function composed of a binary cross-entropy edge loss function and an object loss function; the steps of this detection method are as follows: Step 1, construct the first feature extraction sub-network: Step 1.1, construct a position feature fusion module, the structure of which includes a first dimensionality reduction convolutional block, a second dimensionality reduction convolutional block, and a third dimensionality reduction convolutional block; among them, the structures of the first to third dimensionality reduction convolutional blocks are the same, and each dimensionality reduction convolutional block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and an upsampling layer in series; set the convolutional kernel sizes of the first to third convolutional layers in the first to third dimensionality reduction convolutional blocks to 3×3, the padding numbers to 1, the resolutions of the upsampling layers to 32×32, 64×64, and 128×128 respectively, and the activation layers are all implemented using the ReLU activation function; set the number of channels and the number of convolutional kernels of the first to third convolutional layers in the first dimensionality reduction convolutional block to 512; set the number of channels of the convolutional kernels of the first to third convolutional layers in the second dimensionality reduction convolutional block to 512, 256, and 256 respectively, and the number of convolutional kernels to 256; set the number of channels of the convolutional kernels of the first to third convolutional layers in the third dimensionality reduction convolutional block to 256, 128, and 128 respectively, and the number of convolutional kernels to 128; Step 1.2, construct an inter-layer feature aggregation module, the structure of which includes a first aggregation convolutional block and a second aggregation convolutional block; among them, the structures of the first and second aggregation convolutional blocks are the same, and each aggregation convolutional block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and an upsampling layer in series; set the convolutional kernel sizes of the first to third convolutional layers in the first and second aggregation convolutional blocks to 3×3, the padding numbers to 1, the resolutions of the upsampling layers to 32×32 and 64×64 respectively, and the activation layers are all implemented using the ReLU activation function; set the number of channels and the number of convolutional kernels of the first to third convolutional layers in the first aggregation convolutional block to 512; set the number of channels of the convolutional kernels of the first to third convolutional layers in the second aggregation convolutional block to 512, 256, and 256 respectively, and the number of convolutional kernels to 256; Step 1.3, construct an edge enhancement module, the structure of which includes an edge convolution block, a first upsampling convolution block, a second upsampling convolution block, and a third upsampling convolution block; among them, the edge convolution block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, and a third activation layer in series; the first to third upsampling convolution blocks have the same structure, and each upsampling convolution block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, and a max pooling layer in series; set the size of the convolutional kernels of the first to third convolutional layers in the edge convolution block to 3×3, the padding number to 1, the number of channels and the number of convolutional kernels to 128, and the activation layer is implemented using the ReLU function; set the size of the convolutional kernels of the first to third convolutional layers in the first to third upsampling convolution blocks to 3×3, the padding number to 1, the size of the pooling kernel of the pooling layer to 3×3, the stride to 2, the padding number to 1, and the activation layer is implemented using the ReLU function; set the number of channels of the convolutional kernels of the first to third convolutional layers in the first upsampling convolution block to 128, 256, 256 respectively, and the number of convolutional kernels to 256; set the number of channels of the convolutional kernels of the first to third convolutional layers in the second upsampling convolution block to 256, 512, 512 respectively, and the number of convolutional kernels to 256; in the third upsampling convolution block, the number of channels and the number of convolutional kernels of the first to third convolutional layers are both set to 512; Step 1.4, construct a feature fusion module, the structure of which includes an edge fusion convolution block, a first fusion convolution block, a second fusion convolution block, and a third fusion convolution block; among them, the edge fusion convolution block is composed of a convolutional layer; the first to third fusion convolution blocks have the same structure, and each fusion convolution block is composed of a first convolutional layer, a first activation layer, a second convolutional layer, a second activation layer, a third convolutional layer, a third activation layer, a fourth convolutional layer, and an upsampling layer in series; set the size of the convolutional kernel of the first convolutional layer in the edge convolution block to 1×1, the padding number to 1, the number of channels of the convolutional kernel to 128, and the number of convolutional kernels to 1; for each of the first to fourth convolutional layers in the first to third fusion convolution blocks, set the size of the convolutional kernels of the first to fourth convolutional layers in the third to fifth fusion convolution blocks to 3×3, the padding number to 1, the resolution size of the upsampling layer to 256×256, and the activation layer is implemented using the ReLU function; set the number of channels of the convolutional kernels of the first to fourth convolutional layers in the first fusion convolution block to 256, and the number of convolutional kernels to 256, 256, 256, 1 respectively; set the number of channels of the convolutional kernels of the first to fourth convolutional layers in the second fusion convolution block to 512, and the number of convolutional kernels to 512, 512, 512, 1 respectively; set the number of channels of the convolutional kernels of the first to fourth convolutional layers in the third fusion convolution block to 512, and the number of convolutional kernels to 512, 512, 512, 1 respectively; Step 1.5, cascade the first feature extraction module with the VGG16 model as the backbone and the position feature extraction module composed of the ASPP (Atrous Spatial Pyramid Pooling) module; integrate the features output by the position feature extraction module into the output branches of the second to fourth convolutional blocks in the first feature extraction module, cascade the output features of the third to fourth convolutional blocks with the inter-layer aggregation module, cascade the output features of the second convolutional block, the output features of the inter-layer aggregation module, and the edge enhancement fusion module, and after cascading the edge enhancement module with the feature fusion module, obtain the first feature extraction sub-network; Step 2, construct the second feature extraction sub-network: The second feature extraction sub-network is successively cascaded by the first feature convolutional block, the second feature convolutional block, the third feature convolutional block, the fourth feature convolutional block, and the fifth feature convolutional block; keep the parameters of the first to fifth feature convolutional blocks consistent with the parameters of the first feature extraction module constructed in Step 1.1; Step 3, construct the dual-stream deep network: Step 3.1, connect the second feature extraction sub-network in the edge enhancement module of the first feature extraction sub-network; add and fuse the features output by the third to fifth convolutional blocks of the second feature extraction sub-network through the edge enhancement module; Step 3.2, output a significant edge detection prediction map and three significant object detection prediction maps through the feature fusion module for the output of the edge enhancement module; add and fuse the significant object detection prediction maps and activate them to obtain the final image significant object map, completing the construction of the dual-stream deep network for image significant object detection; Step 4, generate the training set: Step 4.1, form samples by combining 1 color image with its corresponding 1 significant object ground truth map, and select at least 10553 samples to form a sample set; Step 4.2, respectively generate the corresponding significant edge ground truth map, grayscale map, HC significant map, and RC significant map for each sample; Step 4.3, preprocess the sample set and the corresponding significant edge ground truth map, grayscale map, HC significant map, and RC significant map for each sample to form the training set; Step 5, train the dual-stream deep network: Input the training set into the dual-stream deep network, calculate the total loss of the significant edge loss and the significant object loss using the total loss function, and output the overall loss value of the network; through backpropagation of the overall loss value, use the Adam (Adaptive Moment Estimation) optimizer to iteratively update the weight parameters of the network nodes until the overall loss value converges, obtaining the trained dual-stream deep network; Step 6, detect the significant objects in the image: Input the color image to be detected, its generated grayscale map, HC significant map, and RC significant map into the trained dual-stream deep network for image significant objects, and output a single-channel significant object grayscale image.
2. The method for detecting salient objects in images based on edge integrity and clarity and a two-stream network according to claim 1, characterized in that, The structure of the first feature extraction module described in Step 1.5 is as follows in sequence: the first feature convolution block, the second feature convolution block, the third feature convolution block, the fourth feature convolution block, the fifth feature convolution block, and the dilated convolution pooling block; the first feature convolution block consists of a first convolution layer, a first activation layer, a second convolution layer, and a second activation layer; among them, the second feature convolution block is composed of a max pooling layer, a first convolution layer, a first activation layer, a second convolution layer, and a second activation layer connected in series; the third to fifth feature convolution blocks have the same structure, and each feature convolution block is composed of a max pooling layer, a first convolution layer, a first activation layer, a second convolution layer, a second activation layer, a third convolution layer, and a third activation layer connected in series; the convolution kernel sizes of the first and second convolution layers in the first feature convolution block are both set to 3×3, the padding numbers are both set to 1, the numbers of channels of the convolution kernels are respectively set to 3 and 64, and the numbers of convolution kernels are respectively set to 64 and 64; the first and second activation layers are both implemented using the ReLU activation function; the pooling kernel size of the max pooling layer in the second feature convolution block is set to 3×3, the stride is set to 2, and the padding number is set to 1; the convolution kernel sizes of the first and second convolution layers are both set to 3×3, the padding numbers are both set to 1, the numbers of channels of the convolution kernels are respectively set to 64 and 128, and the numbers of convolution kernels are respectively set to 128 and 128; the first and second activation layers are both implemented using the ReLU activation function; the pooling kernel sizes of the max pooling layers in the third to fifth feature convolution blocks are all set to 3×3, the strides are all set to 2, and the padding numbers are all set to 1; the convolution kernel sizes of the first to third convolution layers are all set to 3×3, and the padding numbers are all set to 1; the first to third activation layers are all implemented using the ReLU activation function; the numbers of channels of the first to third convolution layers in the third feature convolution block are respectively set to 128, 256, and 256, and the numbers of convolution kernels are respectively set to 256, 256, and 256; the numbers of channels of the first to third convolution layers in the fourth feature convolution block are respectively set to 256, 512, and 512, and the numbers of convolution kernels are respectively set to 512, 512, and 512; the numbers of channels of the first to third convolution layers in the fifth feature convolution block are all set to 512, and the numbers of convolution kernels are all set to 512.
3. The method for detecting significant objects in an image based on edge integrity and clarity and a dual-stream network according to claim 1, wherein The structure of the position feature extraction module described in Step 1.5 is successively: a max pooling layer, a parallel dilated convolution block, and a dimensionality reduction convolution block; wherein, the parallel dilated convolution block is composed of a first convolution unit, a second convolution unit, a third convolution unit, a fourth convolution unit, and a fifth convolution unit connected in parallel; the first to fourth convolution units have the same structure and are all composed of a convolution layer, a batch normalization layer, and an activation layer connected in series; the fifth convolution unit is composed of an adaptive average pooling layer, a convolution layer, a batch normalization layer, and an activation layer connected in series; the dimensionality reduction convolution block is composed of a convolution layer, a batch normalization layer, and an activation layer connected in series; the pooling kernel size of the max pooling layer is set to 3×3, the stride is set to 1, and the padding number is set to 1; the convolution kernel sizes of the convolution layers in the first to fourth convolution units in the parallel dilated convolution block are set to 1×1, 3×3, 3×3, 3×3 respectively, the padding numbers are set to 0, 6, 12, 18 respectively, the dilation rates are set to 1, 6, 12, 18 respectively, the number of channels of the convolution kernels are set to 512, 512, 512, 512 respectively, and the number of convolution kernels is set to 512; the number of channels of the batch normalization layers are all set to 512; the activation layers are all implemented using the ReLU activation function; the pooling kernel size of the adaptive average pooling layer of the fifth convolution unit is set to 1×1, the convolution kernel size of the convolution layer is set to 1×1, the number of channels of the convolution kernel is set to 512, and the number of convolution kernels is set to 512; the number of channels of the batch normalization layer is set to 512; the activation layer is implemented using the ReLU activation function; the convolution kernel size of the convolution layer in the dimensionality reduction convolution block is set to 1×1, the number of channels of the convolution kernel is set to 2560, and the number of convolution kernels is set to 512; the number of channels of the batch normalization layer is set to 512; the activation layer is implemented using the ReLU activation function.
4. The method for detecting significant objects in an image based on edge integrity and clarity and a two-stream network according to claim 1, wherein The significant edge ground truth map described in Step 4.2 is calculated using the following formula: Among them, represents the significant edge truth map, x represents the significant object truth map, Θ represents the binary morphological erosion operation, k represents the erosion structure element, and its size is a 3×3 three-dimensional unit matrix.
5. The method for detecting significant objects in an image based on edge integrity and clarity and a dual-stream network according to claim 1, wherein The preprocessing described in Step 4.3 refers to subtracting the mean value from all the color images in the training set for data centering, where the mean value is {104.00699, 115.66877, 122.67892}; resetting the resolution of all the images in the training set to 256×256 pixels and performing a random flipping operation.
6. The method for detecting significant objects in an image based on edge integrity and clarity and a two-stream network according to claim 1, characterized in that, The total loss function described in Step 5 is as follows: Among them, L total represents the combined loss function, and L D represents the binary cross-entropy margin loss function, represents the binary cross-entropy target loss function; The binary cross-entropy edge loss function is as follows: L D = -(1 - D') log(1 - D) - D' log(D) where D' represents the significant edge ground truth map, and D represents the generated significant edge prediction map; The binary cross-entropy object loss function is as follows: where Y' represents the significant object ground truth map, and Y3, Y3, Y3 are three significant object prediction maps generated by the network.
Citation Information
Patent Citations
Image significance detection method based on parallel convolution neural network
CN107169954A
An Image Saliency Detection Method Based on Parallel Convolutional Neural Networks
CN107169954B
Significant target detection method and system based on boundary enhancement
CN114821059A
High-resolution remote sensing image-oriented boundary enhanced semantic segmentation method
CN115049936A
Significant target detection method based on semantic guidance and attention mechanism
CN115205663A