A saliency object detection method for hierarchical saliency modeling using generative parameters
Through the hierarchical significant modeling method of generative parameters, using deep neural networks and hierarchical signal generation modules, the problem that traditional machine learning is difficult to adapt to the significance detection of different images is solved, and better significance object detection effect is achieved.
Patent Information
- Application Number
- CN202210087655.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-01-25
AI Technical Summary
Traditional machine learning is difficult to adaptively master the significant learning modes in different images and cannot adapt well to the requirements of significant object detection models in different scenarios.
The hierarchical significant modeling method of generative parameters is adopted, and the image features are extracted through the backbone deep neural network. The hierarchical signal generation module generates adaptive hierarchical signals, and cascaded connections are made through multiple hierarchical modules to output the significance target segmentation map.
Effective modeling and detection of significance levels in different images is realized, and the model's adaptability and segmentation accuracy are improved.
Smart Images

Figure CN114463614B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a saliency object detection method using hierarchical saliency modeling with generative parameters. Background Art
[0002] In human perception, for a given image, an observer has different feelings about the saliency levels of different regions. Therefore, humans can quickly and effectively extract useful regions in the scene according to the saliency hierarchy in the image. However, for machine learning, it is very difficult to directly learn a function that maps regions with different saliency levels to the same pixel values in the true labels. Traditional machine learning is difficult to adaptively master the learning patterns of saliency in different images and cannot well meet the requirements of saliency object detection models in different scenarios. Summary of the Invention
[0003] To solve the above problems, the present invention provides a saliency object detection method using hierarchical saliency modeling with generative parameters. The specific technical solutions adopted by the present invention are as follows:
[0004] A saliency object detection method using hierarchical saliency modeling with generative parameters, comprising the following steps:
[0005] S1. Obtain a color image dataset for training a saliency object detection network and partition its gradient response map;
[0006] S2. Based on a backbone deep neural network, a hierarchical signal generation module, and multiple saliency level modules, construct a saliency object detection network, wherein the backbone deep neural network is used to extract image features of the input RGB color image, the hierarchical signal generation module is used to generate hierarchical signals that make the saliency hierarchical modeling strategy more adaptable to the input color image according to the image features, and the multiple saliency level modules are cascaded and connected to perform hierarchical saliency modeling on the input color image by combining the image features and the hierarchical signals, so as to finally output a saliency object segmentation map;
[0007] S3. Train the constructed saliency object detection network based on the color image dataset, and use the finally trained saliency object detection network to perform saliency object detection on the color image to be detected.
[0008] Preferably, the specific implementation steps of S1 include:
[0009] S11. Obtain a color image dataset as the training data of the saliency object detection network, and each training sample therein includes a single-frame color image I trainand the corresponding manually annotated salient object segmentation map P train ;
[0010] S12. For each frame of color image I train , input it into the ResNet-50 model pre-trained on ImageNet to obtain its corresponding gradient response map G sal . According to a preset threshold, divide G sal into N non-overlapping parts {p 1 , p 2 , …, p N}, where N is the number of saliency levels of the color image I train .
[0011] Preferably, in S2, the backbone deep neural network for extracting image features is composed of K convolutional blocks in cascade. The convolutional block adopts ResNet-50 or VGG-16. The output of the k-th convolutional block is encoded by an encoding layer to obtain the image feature F k . The image features corresponding to all K convolutional blocks form {F 1 , F 2 , …, F K}.
[0012] Preferably, in S2, the specific process in the hierarchical signal generation module is as follows:
[0013] S211. In the hierarchical signal generation module, first use a Transformer decoder to generate hierarchical signals. The Transformer decoder contains L Transformer decoding layers. Each layer of the Transformer decoding layer calculates the similarity between the input image feature F K and the learnable query variable Q 0 in sequence. The calculation process in any l-th Transformer decoding layer is:
[0014] Q l = MLP(MCA(MSA(Q l-1 ), F K ))), l = 1, 2, …, L
[0015] where: Q l-1 , Q l are the calculation results output by the (l - 1)-th and l-th Transformer decoding layers respectively. MSA(·), MCA(·), and MLP(·) represent the multi-head self-attention module, the multi-layer mutual attention module, and the multi-layer perceptron module respectively;
[0016] S212. After obtaining the output Q of the last Transformer decoding layerL After that, an MLP layer shared by all significance levels is used to map it into hierarchical signals:
[0017]
[0018] where s n is the significance signal of the n-th significance level, is the n-th item of Q L ; finally, the significance signals of all significance levels are combined to form the hierarchical signal as {s 1 , s 2 , …, s N}.
[0019] Preferably, in S2, the significance target detection network altogether includes K significance hierarchical modules, each of the significance hierarchical modules includes N branches, corresponding to N significance levels; the K significance hierarchical modules are numbered in reverse order according to the cascading order, the K-th significance hierarchical module is located at the forefront, and the 1st significance hierarchical module is located at the rearmost end; for any k-th significance hierarchical module, the specific process therein is as follows:
[0020] S221. In the significance hierarchical module, first use a classifier to generate a secondary semantic mask for the input feature:
[0021]
[0022] where, H k is the input feature of the k-th significance hierarchical module, wherein the significance hierarchical module cascaded at the forefront uses the image feature F k as the input feature, and the remaining significance hierarchical modules use the output H k-1 of the previous significance hierarchical module as the input feature; is the secondary semantic mask, softmax(·) is the softmax calculation on the channel dimension, and Conv 3x3 (·) is a learnable 3×3 convolutional layer;
[0023] Then expand into N secondary semantic masks corresponding to different semantic levels Each mask represents a different semantic level of the input image; use the secondary semantic mask to divide H k into N parts where:
[0024]
[0025] where, represents element-wise multiplication, Represents the features corresponding to the nth semantic level;
[0026] S222. Based on the features obtained in S221 and the hierarchical signals {s 1 , s 2 , …, s N} obtained in S212, process the corresponding nth semantic level with each saliency signal s n respectively, by converting the signal into the convolutional kernel of the network and calculating with the features:
[0027]
[0028] where * is a 2D convolution operation, is the convolutional kernel generated by using the conversion layer n for the saliency signal s , is the calculated feature;
[0029] S223. Aggregate the features F k-1 output by the backbone deep neural network and the features obtained in S222 together:
[0030]
[0031] where H k-1 represents the final output of the kth saliency hierarchical module, Concat(·) represents the concatenation operation, and when k = 1, F 0 is an empty matrix; the final output H 1 of the first saliency hierarchical module outputs the saliency object segmentation map of the input image
[0032] Preferably, in S3, the specific method for training the model of the constructed saliency object detection network based on the color image dataset is as follows:
[0033] S31. For each training sample, based on the saliency object segmentation map train of the color image I
[0034]
[0035] predicted in S223 and use and the manually annotated saliency object segmentation map P train to calculate the first loss function L ppa :
[0036]
[0037] where l is an index for measuring the difference between the two segmented images;
[0038] S32. For each training sample, based on the secondary semantic mask obtained in S221 and {p 1 , p 2 , …, p N} obtained in S12, calculate the second loss function
[0039]
[0040] where y pos is the set of coordinate points within the range of p n ;
[0041] S53. For each training sample, calculate each final loss function as:
[0042]
[0043] where ρ is a hyperparameter that controls the weights of the two loss functions; use the Adam optimization method and the backpropagation algorithm to train the entire saliency object detection network under the loss function L total until the network converges.
[0044] Preferably, the index for measuring the difference between the two segmented images is the mean squared error.
[0045] Preferably, K is set to 5.
[0046] Preferably, N is set to 3.
[0047] Preferably, L is set to 6.
[0048] Preferably, ρ is set to 0.1.
[0049] This method is based on a deep neural network, explores the saliency differences in RGB images, establishes the saliency hierarchy in the images, and uses deep learning techniques to adaptively master the learning patterns of saliency in different images, providing the saliency hierarchy information as a prior for the model, and can better meet the requirements of the saliency object detection model in different scenarios. Compared with the methods in the prior art, the present invention has the following beneficial effects:
[0050] First, the present invention proposes a method of converting the true labels with the same pixel values in each region used in saliency detection into a series of secondary semantic labels according to the saliency differences, so as to provide hierarchical guidance for the model.
[0051] Secondly, the present invention uses the Transformer technology to explore the significant differences in RGB images and generate network parameters for extracting the features of different significant regions. By improving this part, the adaptability of the model to the significant levels between different samples can be greatly improved, and the robustness of the model can be enhanced.
[0052] Finally, the present invention explicitly models the hierarchical differences of significant objects in samples, processes different significant regions with different parameters, decomposes the features into multiple sub-semantic masks, provides prior knowledge guidance for model prediction, and obtains a better significant object detection model.
[0053] This method can effectively improve the segmentation accuracy and regional similarity of significant objects in the scene in the significant object detection task, and has good application value. For example, it can quickly identify the significant parts containing useful information in a natural image, provide a more refined object segmentation pattern for subsequent tasks such as image retrieval, visual tracking, and pedestrian re-identification, and lay a good foundation. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a schematic diagram of the basic steps of the method of the present invention;
[0055] Figure 2 is a schematic diagram of the significant object detection network structure of the present invention;
[0056] Figure 3 is a partial experimental effect diagram shown in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made within the spirit and scope of the present invention defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.
[0059] Reference Figure 1 , in a preferred embodiment of the present invention, a significant object detection method using hierarchical significant modeling of generative parameters is provided for performing pixel-level fine-grained segmentation of significant objects in a given color image. The method specifically includes the following steps:
[0060] S1. Obtain a color image dataset for training a saliency object detection network, and partition its gradient response map.
[0061] In this embodiment, the specific implementation steps of the above step S1 include:
[0062] S11. Obtain a color image dataset as the training data of the saliency object detection network, where each training sample includes a single-frame color image I train and the corresponding manually annotated saliency object segmentation map P train ;
[0063] S12. For each frame of color image I train , input it into the ResNet-50 model pre-trained on ImageNet to obtain its corresponding gradient response map G sal . According to a preset threshold, equally divide the value range of G sal into N intervals, and then divide the gradient response map G sal into N non-overlapping parts {p 1 , p 2 , …, p N} according to these N intervals, where N is the number of saliency levels of the color image I train .
[0064] S2. Based on a backbone deep neural network, a hierarchical signal generation module, and multiple saliency level modules, construct a saliency object detection network, where the backbone deep neural network is used to extract the image features of the input RGB color image, the hierarchical signal generation module is used to generate hierarchical signals that make the saliency level modeling strategy more adaptable to the input color image, and the multiple saliency level modules are connected in cascade and used to perform saliency level modeling on the input color image by combining the image features and the hierarchical signals, so as to finally output a saliency object segmentation map.
[0065] In this embodiment, in the above step S2, the structure of the saliency object detection network is as Figure 2 shown, and the structures of the backbone deep neural network, the hierarchical signal generation module, and the saliency level module included therein and the specific data processing flow inside are as follows:
[0066] In this embodiment, for the backbone deep neural network, the backbone deep neural network for extracting image features is composed of K convolutional blocks connected in cascade. The convolutional block can adopt ResNet-50 or VGG-16. The output of the kth convolutional block is encoded by an encoding layer to obtain the image feature F k , and the image features corresponding to all K convolutional blocks form {F 1 , F2 , …, F K}。
[0067] In this embodiment, for the hierarchical signal generation module, the specific process in this module is as follows:
[0068] S211. First, a Transformer decoder is used in the hierarchical signal generation module to generate hierarchical signals. The Transformer decoder includes L Transformer decoding layers. A single Transformer decoding layer contains cascaded multi-head self-attention (MSA) modules, multi-head cross-attention (MCA) modules, and multilayer perceptron modules; each layer of the Transformer decoding layer calculates the input image feature F K and the similarity with the learnable query variable Q 0 in sequence. The calculation process in any l-th layer of the Transformer decoding layer is:
[0069] Q l = MLP(MCA(MSA(Q l-1 ), F K ))), l = 1, 2, …, L
[0070] where: Q l-1 , Q l are the calculation results output by the (L - 1)-th layer and the l-th layer of the Transformer decoding layer respectively. MSA(·), MCA(·), and MLP(·) represent the multi-head self-attention module, the multi-head cross-attention module, and the multilayer perceptron module respectively;
[0071] S212. After obtaining the output Q L of the last layer of the Transformer decoding layer, use an MLP layer shared by all significance levels to map it into a hierarchical signal:
[0072]
[0073] where s n is the significance signal of the n-th significance level, is the n-th item of Q L ; finally, combine the significance signals of all significance levels to form a hierarchical signal as {s 1 , s 2 , …, s N}.
[0074] In this embodiment, for the saliency level module, there are a total of K saliency level modules in the entire saliency object detection network. Each saliency level module contains N branches, corresponding to N saliency levels; the K saliency level modules are numbered in reverse order according to the cascade sequence. The K-th saliency level module is at the front end, and the (K - 1)-th saliency level module is downstream of the K-th saliency level module, and so on. The 1st saliency level module is at the end. For any k-th saliency level module, k = 1, 2, …, K, the specific process is as follows:
[0075] S221. First, use a classifier in the saliency level module to generate a secondary semantic mask for the input feature:
[0076]
[0077] where H k is the input feature of the k-th saliency level module. The saliency level module cascaded at the front end uses the image feature F k as the input feature, and the remaining saliency level modules use the output H k-1 of the previous saliency level module as the input feature; is the secondary semantic mask, softmax(·) is the softmax calculation on the channel dimension, and Conv 3x3 (·) is a learnable 3×3 convolutional layer;
[0078] Then, expand into N secondary semantic masks corresponding to different semantic levels Each mask represents a different semantic level of the input image; use the secondary semantic mask to divide H k into N parts where:
[0079]
[0080] where, represents element-wise multiplication, represents the feature corresponding to the n-th semantic level;
[0081] S222. Based on the features obtained in S221 and the level signals {s 1 , s 2 , …, s N} obtained in S212, use each saliency signal s n to process the corresponding n-th semantic level, and convert the signal into the convolutional kernel of the network and calculate it with the feature:
[0082]
[0083] where * is a 2D convolution operation, for the saliency signal s n using a conversion layer to generate a convolution kernel, is the calculated feature;
[0084] S223. Aggregate the feature F output by the backbone deep neural network k-1 with the feature obtained in S222 together:
[0085]
[0086] where H k-1 represents the final output of the k-th saliency level module, Concat(·) represents the concatenation operation, and when k = 1, F 0 is an empty matrix; the final output H of the first saliency level module 1 after passing through a 3×3 convolutional layer, outputs the saliency object segmentation map of the input image
[0087] It should be noted that the above K, N, and L in the present invention can be adjusted according to actual needs. In this embodiment, K is set to 5, N is set to 3, and L is set to 6. Therefore, as Figure 2 shown, in the entire saliency object detection network, there is a backbone deep neural network composed of 5 cascaded convolutional blocks, a hierarchical signal generation module with 6 transformer decoding layers, and 5 saliency level modules. The encoded features output after encoding the features output by each of the 5 convolutional blocks in the backbone deep neural network are used as the inputs of different hierarchical signal generation modules. At the same time, the features output by the last convolutional block are also used as the inputs of the hierarchical signal generation module to generate hierarchical signals. The hierarchical signals and the encoded features are simultaneously used as the inputs of the 5 saliency level modules. The features output by the last saliency level module are convolved by 3×3 in the output layer to obtain the saliency object segmentation map.
[0088] S3. Train the constructed saliency object detection network based on the color image dataset, and use the finally trained saliency object detection network to perform saliency object detection on the color image to be detected.
[0089] In this embodiment, the specific method for training the constructed saliency object detection network based on the color image dataset in the above step S3 is as follows:
[0090] S31. For each training sample, based on the color image I predicted in S223 trainSalient object segmentation map
[0091]
[0092] and use and the manually annotated salient object segmentation map P train to calculate the first loss function L pp a:
[0093]
[0094] where l is an index to measure the difference between the two segmentation maps, and mean squared error MSE can be used in this embodiment;
[0095] S32. For each training sample, based on the secondary semantic mask obtained in S221 and {p 1 , p 2 , …, p N} obtained in S12 to calculate the second loss function
[0096]
[0097] where y pos is the set of coordinate points within the range of p n ;
[0098] S53. For each training sample, calculate each final loss function as:
[0099]
[0100] where ρ is a hyperparameter to control the weights of the two loss functions, and can be set to 0.1 in this embodiment; use the Adam optimization method and backpropagation algorithm to train the entire salient object detection network under the loss function L total until the network converges.
[0101] The salient object detection network that converges after the above training can be used to detect salient objects in actual RGB color images. During application, just input the RGB color image to be detected into the salient object detection network, and the salient object segmentation map can be output. The method described in S1 - S3 above will be applied to specific embodiments below so that those skilled in the art can better understand the effects of the present invention.
[0102] Embodiment
[0103] The implementation method of this embodiment is as described in the previous S1 - S3, and the specific steps will not be elaborated in detail. Below, only the effects will be demonstrated for the case data. The present invention is implemented on five datasets with ground truth annotations, namely:
[0104] DUTS dataset: This dataset contains 15,572 images and their saliency labels.
[0105] ECSSD dataset: This dataset contains 1,000 images and their saliency labels.
[0106] HKU - IS dataset: This dataset contains 4,447 images and their saliency labels.
[0107] DUT - OMRON dataset: This dataset contains 5,168 images and their saliency labels.
[0108] PASCAL dataset: This dataset contains 850 images and their saliency labels.
[0109] In this example, 10,553 image - label pairs are selected from the DUTS dataset as the training set, and the others are used as the test set. A deep learning model is established and trained through the aforementioned method.
[0110] As Figure 3 shown, in the figure, GT represents the ground truth annotation of the salient object segmentation map label. The salient object segmentation map obtained by the method of the present invention is basically consistent with the ground truth salient object segmentation map.
[0111] The detection accuracy of the detection results of this embodiment is shown in Table 1 below. The prediction accuracy of various methods is mainly compared using two metrics: the average F - measure and M. Among them, the average F - measure metric is used to measure the regional similarity between the predicted salient segmentation map and the ground truth salient segmentation map. The larger the value, the more similar the predicted result is to the ground truth result; M is the result difference of each pixel point in the predicted salient segmentation map. The smaller the value, the closer the predicted result is to the ground truth segmentation map. As shown in Table 1, compared with other methods, the method of the present invention (denoted as Our network) has obvious advantages in both the average F - measure and M metrics.
[0112] Table 1
[0113]
[0114] In the above embodiments, the RGB saliency object detection method of the present invention first converts the true label into a series of secondary semantic labels. On this basis, the transformer technology is used to explore the saliency differences in the RGB image and generate network parameters for extracting the features of different saliency regions. Finally, different parameters are used to process different saliency regions, and the features are deconstructed into multiple secondary semantic masks, providing prior knowledge guidance for model prediction and obtaining a better saliency object detection model.
[0115] Through the above technical solutions, the embodiments of the present invention develop a saliency object detection method using hierarchical saliency modeling with generative parameters based on deep learning technology. The present invention can model the saliency difference hierarchy of RGB samples, use the saliency difference as prior knowledge to guide the learning of the deep model, and can better adapt to the saliency object detection tasks in different complex scenarios.
[0116] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A saliency object detection method using hierarchical saliency modeling with generative parameters, characterized in that it includes the following steps: S1. Obtain a color image dataset for training a saliency object detection network, and partition its gradient response map; S2. Based on a backbone deep neural network, a hierarchical signal generation module, and multiple saliency hierarchy modules, construct a saliency object detection network, where the backbone deep neural network is used to extract image features of the input RGB color image, the hierarchical signal generation module is used to generate hierarchical signals that make the saliency hierarchy modeling strategy more adaptable to the input color image, and the multiple saliency hierarchy modules are connected in cascade and used to perform saliency hierarchy modeling on the input color image by combining the image features and the hierarchical signals, so as to finally output a saliency object segmentation map; S3. Based on the color image dataset, perform model training on the constructed saliency object detection network, and use the finally trained saliency object detection network to perform saliency object detection on the color image to be detected; The specific implementation steps of S1 include: S11. Obtain a color image dataset as the training data for the saliency object detection network, where each training sample includes a single-frame color image I train and the corresponding manually annotated saliency object segmentation map P train ; S12. For each frame of color image I train , input it into the ResNet-50 model pre-trained on ImageNet to obtain its corresponding gradient response map G sal . According to a preset threshold, divide G sal into N non-overlapping parts {p 1 , p 2 , …, p N}, where N is the number of significance levels of the color image I train ; In the above S2, the backbone deep neural network for extracting image features is composed of K cascaded convolutional blocks. The convolutional block adopts ResNet-50 or VGG-16. The output of the k-th convolutional block is encoded by an encoding layer to obtain the image feature F k , and the image features corresponding to all K convolutional blocks form {F 1 , F 2 , …, F K}. In S2, the specific process in the hierarchical signal generation module is as follows: S211. In the hierarchical signal generation module, a Transformer decoder is first used to generate hierarchical signals. The Transformer decoder contains L Transformer decoding layers, and each Transformer decoding layer sequentially calculates the input image feature F K and the learnable query variable Q 0 for similarity. The calculation process in any l-th Transformer decoding layer is as follows: Q l = MLP(MCA(MSA(Q l-1 ), F K )),l = 1, 2, …, L Where: Q l-1 and Q l are the calculation results output by the (l-1)-th and l-th layer Transformer decoding layers respectively. MSA(·), MCA(·), and MLP(·) represent the multi-head self-attention module, the multi-layer cross-attention module, and the multi-layer perceptron module respectively; S212. After obtaining the output Q of the last layer of the Transformer decoder, use an MLP layer shared by all significance levels to map it into hierarchical signals: L After obtaining the output Q of the last layer of the Transformer decoder, use an MLP layer shared by all significance levels to map it into hierarchical signals: where s n is the significance signal of the n-th layer of significance levels, is the n-th item of Q L ; finally, the significance signals of all significance levels are combined to form a hierarchical signal {s 1 , s 2 , …, s N}; In S2, the saliency object detection network altogether includes K saliency hierarchy modules, each saliency hierarchy module includes N branches, corresponding to N saliency levels; the K saliency hierarchy modules are numbered in reverse order according to the cascade order, the one at the forefront is the Kth saliency hierarchy module, and the one at the end is the 1st saliency hierarchy module; for any kth saliency hierarchy module, the process therein is specifically as follows: S221. In the saliency hierarchy module, first use a classifier to generate a secondary semantic mask for the input features: Among them, H k is the input feature of the k-th saliency level module, where the saliency level module cascaded at the very front takes the image feature F k as the input feature, and the remaining saliency level modules take the output H k-1 of the previous saliency level module as the input feature; is the secondary semantic mask, softmax(·) is the softmax calculation on the channel dimension, and Conv 3x3 (·) is a learnable 3×3 convolutional layer; Then is expanded into N sub-semantic masks corresponding to different semantic levels Each mask represents a different semantic level of the input image; the H is divided into N parts by using the sub-semantic masks k where Among them: Among them, represents element-wise multiplication, represents the feature corresponding to the nth semantic level; S222. Based on the features obtained in S221 and the hierarchical signals {s 1 , s 2 , …, s N} obtained in S212, each significant signal s n is used to process the corresponding nth semantic level by converting the signal into the convolutional kernel of the network and calculating it with the features: where * is a 2D convolution operation, for the saliency signal s n using the conversion layer to generate the convolution kernel, is the calculated feature; S223. Aggregate the feature F output by the backbone deep neural network k-1 with the feature obtained in S222 together: Among them, H k-1 represents the final output of the k-th saliency level module, Concat(·) represents the concatenation operation, and F is an empty matrix when k = 1; the final output H 0 of the first saliency level module 1 outputs the saliency object segmentation map of the input image after passing through a 3×3 convolutional layer 2. The saliency object detection method using hierarchical saliency modeling with generative parameters according to claim 1, characterized in that in S3, the specific method for performing model training on the constructed saliency object detection network based on the color image dataset is as follows: S31. For each training sample, based on the saliency object segmentation map of the color image I predicted in S223 train and use and the manually annotated significant object segmentation map P train calculate the first loss function L ppa : where l is an index for measuring the difference between two segmentation maps; S32. For each training sample, based on the secondary semantic mask obtained in S221 and {p 1 , p 2 , …, p N} obtained in S12, calculate the second loss function where y pos is a set of coordinate points within p n range; S53. For each training sample, calculate each final loss function as: Among them, ρ is a hyperparameter that controls the weights of the two loss functions; the entire saliency object detection network is trained under the loss function L using the Adam optimization method and the backpropagation algorithm until the network converges. total 3. The saliency object detection method using hierarchical saliency modeling with generative parameters according to claim 2, characterized in that the index for measuring the difference between two segmentation maps is the mean square error.
4. The saliency object detection method using hierarchical saliency modeling with generative parameters according to claim 2, characterized in that K is set to 5 and N is set to 3.
5. The saliency object detection method using hierarchical saliency modeling with generative parameters according to claim 2, characterized in that L is set to 6.
6. The saliency object detection method using hierarchical saliency modeling with generative parameters according to claim 2, characterized in that ρ is set to 0.1.
Citation Information
Patent Citations
RGBD image saliency detection method and system based on multi-level fusion
CN111723822A
RGB-D saliency target detection method based on depth perception and multi-mode automatic fusion
CN112651406A