A saliency prediction method for panoramic images based on self-supervised learning

By adopting a self-supervised learning method in panoramic image significance prediction, targeted training of the encoder and fusion of global and local information is solved, the problems of geometric distortion and lack of labels in panoramic image significance prediction are achieved, and high-quality significance prediction is achieved.

CN115631121BActive Publication Date: 2025-05-13SHENZHEN HUAYIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211344155.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-05-13
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

The prior art has problems with geometric distortion and lack of significance labels in the significance prediction of panoramic images, resulting in poor performance of the model and degradation of accuracy when directly applying the 2D model to the panoramic image.

Method used

A self-supervised learning method is used to train the encoder in a targeted manner using a large number of unlabeled panoramic images, fusion of global and local information, and learning the characteristics of different images in the field of view, thereby alleviating the problem of lack of significance labels.

Benefits of technology

Through this method, the significance characteristics of the panoramic image can be learned during the encoder training process, ensuring that high-quality prediction results can be obtained using only one ERP image in the prediction stage, and the accuracy and efficiency of the significance prediction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631121B_ABST
    Figure CN115631121B_ABST
Patent Text Reader

Abstract

The present invention discloses a panoramic image saliency prediction method based on self-supervised learning, comprising the following steps: S1. Training an encoder using an unlabeled ERP image set, including the following sub-steps: S11. Projecting the ERP image onto a spherical surface to obtain an image group C i and a label P i ; S12. Randomly shuffling C i ; S13. Conducting encoder training, constructing a global feature extraction network and a local feature extraction network, and learning the features of both through feature fusion to update the model parameters of the global feature extraction network; S2. Conducting decoder training; S3. Inputting the panoramic image to be recognized into the trained encoder for feature extraction, and then inputting the extracted features into the decoder to obtain the final saliency prediction. The present invention uses a large number of unlabeled panoramic images to specifically train the encoder in the saliency model, alleviating the phenomenon of poor model performance caused by the lack of saliency labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a panoramic image saliency prediction method based on self-supervised learning. Background Art

[0002] The development of the Metaverse industry has driven the production and consumption of panoramic images. Compared with traditional 2D images, panoramic images can provide users with a full field of view and bring an immersive experience. However, due to the limited field of vision of humans, only a small part of the transmitted panoramic information can be actually used, resulting in a waste of bit rate. The bright spots in the saliency image represent the area that the user may watch, so we can allocate bit rate according to the saliency image to achieve the purpose of saving bit rate. Before deep learning networks were applied to saliency prediction, researchers segmented the images and simulated the human visual attention mechanism based on manually designed features. The deep learning network learns from the label image and selects features that are more suitable for saliency prediction, obtaining more accurate and robust results.

[0003] At present, most prediction methods focus on 2D images. The long research cycle makes the 2D saliency prediction model and data set more perfect. However, due to the geometric distortion caused by projecting the panoramic image onto the plane, the effect of directly applying the 2D model to the panoramic image is not ideal. And due to the short development time, the saliency labels of panoramic images are extremely limited. Without the assistance of 2D models and data sets, it is difficult for the model to select features suitable for panoramic saliency prediction to obtain better results. Therefore, some methods project the panoramic image into a small field of view image with less distortion, predict the result through the 2D model, and then perform multi-view fusion to obtain the final prediction result. Although this type of method has high accuracy, it makes real-time prediction difficult due to the large number of predicted faces and the long time required for projection.

[0004] The patent application with the publication number CN14998310A discloses a saliency detection method and system based on image processing. First, a filter image and an HSV image corresponding to a preprocessed image are obtained; based on each channel component image of the filter image and the HSV image, multiple superpixel blocks of each channel component image and the channel level distribution of each superpixel block are obtained, and the target feature index is obtained by the difference in the channel level distribution between the superpixel blocks and the center point distance between the superpixel blocks; a saliency index model is established by the superpixel block and the target feature index to obtain a first saliency index value, and the first saliency index value of each superpixel block is corrected to obtain a second saliency index value; and the second saliency index value of each channel component image is fused to obtain the target saliency index value of the superpixel block. Enhancement processing is achieved by calculating the saliency index value of each region of each channel component image, and the detection and extraction of saliency regions in the preprocessed image are completed, thereby improving the detection accuracy and efficiency. The method designs two different manual features to simulate the human attention mechanism. Due to our limited understanding of the mechanism, the detection of saliency regions only by manual features will have a more obvious decrease in accuracy in some scenes, and the geometric distortion of the panoramic image will cause the feature failure for the plane design.

[0005] The patent application with publication number CN107274419A discloses a deep learning saliency detection method based on global prior and local context. First, the color image and the depth image are segmented into superpixels. Based on the mid-level features such as compactness, uniqueness and background of each superpixel, the global prior feature map of each superpixel is obtained, and further through the deep learning model, the global prior saliency map is obtained; then, the global prior saliency map and the local context information in the color image and the depth image are combined, and the initial saliency map is obtained through the deep learning model; finally, the initial saliency map is optimized according to spatial consistency and appearance similarity to obtain the final saliency map. The application of the present invention solves the problem that the traditional saliency detection method cannot effectively detect salient objects in complex background images, and also solves the problem that the existing saliency detection method based on deep learning causes false detection due to the presence of noise in the extracted high-level features. This solution uses convolutional neural networks to extract significant features based on traditional methods, and adopts different inputs to ensure the integrity of the extracted features. Although it effectively improves the robustness of the prediction, its structure of superimposing multiple models is bound to lead to the accumulation of errors, resulting in a decrease in accuracy. In addition, the pixel lengths of the panoramic image projection at different latitudes are not the same, resulting in the failure of the design for superpixels.

[0006] The patent application with the publication number CN107346436A discloses a visual saliency detection method for fusion image classification, including: using a visual saliency detection model including an image coding network, an image decoding network and an image recognition network, using a multi-scale image as the input of the image coding network, extracting the features of the image at multiple resolutions as the coding feature vector F; fixing the weights of the image coding network except the last two layers, training the network parameters, and obtaining the visual saliency map of the original image; using F as the input of the image decoding network, normalizing the saliency map corresponding to the original image; inputting F into the image decoding network, and finally obtaining the generated visual saliency map through an upsampling layer and a nonlinear sigmoid layer; using the image recognition network with the visual saliency map of the original image and the generated visual saliency map as input, using a convolution layer with a small convolution kernel to extract features and perform pooling processing, and finally using three fully connected layers to output the probability distribution of the generated map and the probability distribution of the classification label. The purpose of quickly and effectively analyzing and judging the image is achieved, and good effects such as saving manpower and material costs and significantly improving accuracy are obtained in the practice of image annotation, supervision and behavior prediction. This scheme takes multi-scale images as input, uses convolutional neural networks to extract multi-scale features, and uses convolutional neural networks to decode the obtained features, so that the model can be learned end-to-end. However, the model is not specifically designed based on saliency, resulting in the need to improve the model accuracy. Directly applying the planar model to panoramic images will result in a decrease in accuracy. Summary of the invention

[0007] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for predicting panoramic image saliency based on self-supervised learning, which utilizes a large number of unlabeled panoramic images to perform targeted training on the encoder in the saliency model, alleviates the phenomenon of poor model performance caused by the lack of saliency labels, and fuses global and local information during the encoder training process, so that the encoder can learn the features of images with different fields of view.

[0008] The objective of the present invention is achieved through the following technical solution: A panoramic image saliency prediction method based on self-supervised learning comprises the following steps:

[0009] S1. Train the encoder using the unlabeled ERP image set, including the following sub-steps:

[0010] S11. Format conversion: Project the ERP image onto a spherical surface to obtain a CMP image group C i and label P i , i=1,…,6;

[0011] S12, for C i Randomly shuffle to get c i , and according to c iThe original position of P i Update to get the agent task label

[0012] S13. Train the encoder and build a global feature extraction network With local feature extraction network The global features and local features are taken as input, and the features of the two are learned through feature fusion, and the model parameters of the global feature extraction network are updated;

[0013] S2. Perform decoder training: decoder g θ : Constructed to predict the final significant result

[0014] S3. The panoramic image to be identified is input into the trained encoder for feature extraction, and then the extracted features are input into the decoder to obtain the final saliency prediction.

[0015] In step S13, the global feature extraction network With local feature extraction network They are:

[0016] E->F E

[0017]

[0018] where F E is a global feature, is a local feature, E represents the ERP image; -> represents the reasoning process of the feature extraction network, the global feature extraction network With local feature extraction network All models use VGG16 with the last 5 layers removed;

[0019] Then the obtained global feature F E and local features Jointly input into the feature fusion network;

[0020] The feature fusion network includes two parts: feature transformation and dot multiplication operation: First, F E and After two fully connected layers with unshared weights, we get r E and Then transform it by the following equation:

[0021] Q E =r E W Q

[0022]

[0023]

[0024] Where W Q , W V and W K is the weight that the three types of features do not share, Q E , and Represents Query, Value and Key respectively;

[0025] Then use the dot product operation to fuse the obtained features:

[0026]

[0027] Among them, CA i is the result of feature fusion, ReLU is the activation function, Represents a function nesting operator;

[0028] Obtained CA i is used for the final position prediction:

[0029]

[0030] Training is performed using the following loss function:

[0031]

[0032] The loss function calculates the difference between the predicted value and the label value, and then performs a gradient backpropagation based on the difference and updates the gradient based on the gradient. The parameters in the model are stopped after traversing the unlabeled ERP image set 100 times to obtain the global feature extraction network

[0033] Obtaining salient images: The head and eye movement recording files are used as the training set for the decoder. First, a zero matrix with the same size as the image in the training set is established. Different viewpoint positions are recorded in the head and eye recording files. If a point is recorded in the file, it is marked as 1 in the matrix. According to the recorded position, the zero matrix is ​​updated using the following method:

[0034]

[0035] S ij It is the viewpoint graph; and the viewpoint graph is difficult to train due to its sparse matrix characteristics, so the following processing is performed:

[0036]

[0037] where G is a Gaussian kernel with a dilation angle of 5°, S E Indicated by S ij The matrix formed;

[0038] Update as follows:

[0039]

[0040] Among them, T represents the process of transformation from ERP to CMP. back Represents the process of transformation from CMP to ERP;

[0041] Loss function: According to the characteristic that most areas of the saliency image are 0, the following loss function is selected to train the decoder model:

[0042]

[0043] is the predictive distribution, is the true distribution, ε is a constant set to prevent the predicted value from being too close to 0 and causing the loss to tend to infinity, W E , H E are the width and height of the image respectively;

[0044] The loss function calculates the difference between the predicted value and the label value, and then performs gradient backpropagation based on the difference and updates the parameters in the model based on the gradient. After traversing the training set one hundred times, the decoder g is obtained. θ .

[0045] The beneficial effects of the present invention are: using a large number of unlabeled panoramic images, the encoder in the saliency model is trained in a targeted manner, thereby alleviating the phenomenon of poor model performance caused by the lack of saliency labels. In addition, during the encoder training process, global and local information are fused so that the encoder can learn the features of images with different fields of view. In the prediction stage, high-quality prediction results can be obtained using only one ERP image. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 A flowchart of panoramic image saliency prediction based on self-supervised learning of the present invention;

[0047] Figure 2 This is a viewpoint conversion diagram based on CMP and ERP of the present invention. DETAILED DESCRIPTION

[0048] Definitions of Abbreviations and Key Terms:

[0049] ERP (Equi-Rectangular Projection): Equirectangular projection, a projection method that maps spherical information to a single plane.

[0050] CMP (Cube Map Projection): Cube mapping projection, a projection method that places the spherical surface in a cube and maps it to six independent faces.

[0051] ROC (Receiver Operating characteristic Curve): Receiver operating characteristic curve, which maps the classification result into a point on the plane and judges the quality of the classifier by the position of the point.

[0052] NSS (Normalized Scanpath Saliency): Normalized scan path saliency, used to measure the difference between the saliency image and the viewpoint map.

[0053] KLD (Kullback-Leibler Divergence): KL divergence measures the difference between two probability distributions.

[0054] SIM (Similarity): Similarity, which measures the similarity between two distributions.

[0055] CC (Linear Correlation): Pearson correlation coefficient, used to measure the degree of linear correlation between images.

[0056] AUC-J (Area Under ROC Curve-Judd): A variant of the area under the ROC curve. It plots points on the ROC by giving different thresholds to obtain true positive and false positive values, and calculates the area under the surface to measure the accuracy of the classifier.

[0057] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0058] like Figure 1 As shown, a panoramic image saliency prediction method based on self-supervised learning of the present invention comprises the following steps:

[0059] S1. Train the encoder using the unlabeled ERP image set, including the following sub-steps:

[0060] S11. Format conversion: Project the ERP image onto a spherical surface. Without rotating the spherical surface, the position of the CMP surface mapped by the spherical surface on the ERP is fixed. Using this characteristic, the ERP (E) image is format converted to obtain the CMP image group C. i and label Pi , i = 1, ..., 6; at input C i In the case of ERP, the model can predict C i Location information P i to conduct training;

[0061] S12. In order to prevent the model from stopping learning due to shortcuts in the task, i Randomly shuffle to get c i , and according to c i The original position of P i Update to get the agent task label

[0062] S13. Train the encoder and build a global feature extraction network With local feature extraction network The global features and local features are taken as input, and the features of the two are learned through feature fusion, and the model parameters of the global feature extraction network are updated;

[0063] The specific training process is as follows: the encoder takes global and local information as input and learns the features of both through feature fusion; in order to extract two types of images with different fields of view and better learn the proxy task, it is necessary to set up a global feature extraction network and a local feature extraction network; although the local network can only accept a single CMP face at a time, in order to narrow the information gap between it and the global network, we choose to input all CMP faces in one training process to help build implicit global information and help the global encoder better learn local features during feature fusion.

[0064] The global feature extraction network With local feature extraction network They are:

[0065] E->F E

[0066]

[0067] where F E is a global feature, is a local feature, E represents the ERP image; -> represents the reasoning process of the feature extraction network, the global feature extraction network With local feature extraction network The model used is the VGG16 model with the last 5 layers removed. The VGG16 network is a commonly used network in this field. For its specific structure, please refer to Simonyan K, Zisserman A. Very Deep Convolutional Networks for Large-Scale Image Recognition [J]. arXiv e-prints, 2014.

[0068] Then the obtained global feature F E and local features The two are input together into the feature fusion network; the feature fusion network includes two parts: feature transformation and dot multiplication operation:

[0069] First, F E and After two fully connected layers with unshared weights, we get r E and Then transform it by the following equation:

[0070] Q E =r E W Q

[0071]

[0072]

[0073] Where W Q 、V V and W K is the weight that the three types of features do not share, Q E , and Represents Query, Value and Key respectively;

[0074] Then use the dot product operation to fuse the obtained features:

[0075]

[0076] Among them, CA i is the result of feature fusion, ReLU is the activation function, represents a function nesting operator; it can be seen that in the above process, the present invention selects ReLU instead of the original softmax as the activation function. Using softmax means that Q E Need and each This will cause the model to pay too much attention to the The features obtained are consistent with the ones we want to train The purpose is not met.

[0077] Obtained CA i is used for the final position prediction:

[0078]

[0079] in This is the prediction result of the model. Training is performed using the following loss function:

[0080]

[0081] The loss function calculates the difference between the predicted value and the label value, and then performs a gradient backpropagation based on the difference and updates the gradient based on the gradient. The parameters in the model are stopped after traversing the unlabeled ERP image set 100 times to obtain the global feature extraction network

[0082] The unlabeled ERP image set used in this embodiment comes from the unlabeled ERP images in "Djilali Y, Krishna T, McGuinnessK, et al. Rethinking 360deg Image Visual Attention Modelling With Unsupervised Learning. [C] / / International Conference on Computer Vision. 2021".

[0083] S2. Decoder training: After completing the encoder After training, the decoder g θ : Constructed to predict the final significant result g θ Based on the structure in "Pan J, Canton C, McGuinness K, et al. SalGAN: Visual Saliency Prediction with Generative Adversarial Networks[J]. 2017", a U-shaped network is used to help the decoder have a larger receptive field, so as to simulate the field of vision of humans when viewing images and make better judgments on salient areas.

[0084] Obtaining salient images: Use the head and eye movement recording files provided in the dataset "Xu Y, Dong Y, Wu J, et al. Gaze Prediction in Dynamic 360° Immersive Videos[C] / / 2018IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE, 2018" as the decoder training set. First, establish a zero matrix with the same size as the image in the training set. Different viewpoint positions will be recorded in the head and eye recording files. If a point is recorded in the file, it will be marked as 1 in the matrix; according to the recorded position, the zero matrix is ​​updated using the following method:

[0085]

[0086] S ij It is the viewpoint graph; and the viewpoint graph is difficult to train due to its sparse matrix characteristics, so the following processing is performed:

[0087]

[0088] where G is a Gaussian kernel with a dilation angle of 5°, S E Indicated by S ij The matrix is ​​composed of ; using this method means that each point on the ERP is treated equally, and the ERP has different pixel densities at different latitudes, which leads to a large gap between the saliency image obtained in this way and the real image. In order to alleviate this problem, the method is updated as follows:

[0089]

[0090] Among them, T represents the process of transformation from ERP to CMP. back represents the process of converting from CMP to ERP; CMP uses more faces in projection, resulting in less distortion and closer to the image in the real field of view. We use CMP for conversion while keeping the Gaussian kernel unchanged, and get the following: Figure 2 The results in the second column ( Figure 2 In the figure, the square image on the right side of each image is an enlarged image of the image in the small square on the left side. Figure 2 In the comparison between (b) and (c), the CMP-based transformation method tends to ignore viewpoints in higher latitude regions. Figure 2Different from the ERP format images in the image, in reality, the viewpoints in high latitudes are more concentrated. In the saliency map based on ERP conversion, it is obvious that these points with close distances are transformed into more scattered viewpoints near the equator.

[0091] Loss function: According to the characteristic that most areas of the saliency image are 0, the following loss function is selected to train the decoder model:

[0092]

[0093] is the predictive distribution, is the true distribution, ε (ε = 1e-50) is a constant set to prevent the predicted value from being too close to 0 and causing the loss to tend to infinity, W E , H E are the width and height of the image respectively.

[0094] The loss function calculates the difference between the predicted value and the label value, and then performs gradient backpropagation based on the difference and updates the parameters in the model based on the gradient. After traversing the training set one hundred times, the decoder g is obtained. θ ; Using KLD loss can help the model focus on areas where the gap between the predicted value and the true value is large, which is more in line with the use requirements of significance.

[0095] S3. The panoramic image to be identified is input into the trained encoder for feature extraction, and then the extracted features are input into the decoder to obtain the final saliency prediction.

[0096] Experimental test results: AUC-J, NSS, CC, SIM and KLD are used to evaluate the performance of the network of the present invention, compared with UNISAL (Droste R, Jiao J, Noble J A. Unified Image and Video Saliency Modeling [J]. 2020), SalGAN (Pan J, Canton C, McGuinness K, et al. SalGAN: Visual Saliency Prediction with Generative Adversarial Networks [J]. 2017), SaltiNet (Marc A, Xavier G, Kevin MG, et al. Scanpath and saliency prediction on 360 degree images [J]. Signal Processing: Image Communication, 2018, 69: 8-14), MV-SalGAN360 (Chao FY, Zhang L, Hamidouche W, et al. A Multi-FoV Viewport-Based Visual Saliency Model Using Adaptive Weighting Losses for 360$^circ$Images[J].Institute ofElectrical and Electronics Engineers(IEEE),2021》),ATSAL(《Dahou Y,Tliba M,Mcguinness K,et al.ATSal:An Attention Based Architecture for SaliencyPrediction in 360Videos[J].2020》),Rethink(《Djilali Y,Krishna T,Mcguinness K,et al.Rethinking 360deg Image Visual Attention Modelling With UnsupervisedLearning.[C] / / International Conference on Computer Vision.2021》) were evaluated on indicators and compared with the inference speed of five methods.

[0097] Table 1 Performance comparison on Salient360!

[0098]

[0099] Table 1 shows the average performance of 25 images in the Salient360! dataset (Gutierrez J, David EJ, Coutrot A, et al. Introducing UN Salient360! Benchmark: A platform for evaluating visual attention models for 360° contents [C] / / 2018 Tenth International Conference on Quality of Multimedia Experience (QoMEX). 2018) under five indicators. In the table, ↑ indicates that the larger the indicator data, the better the effect, and ↓ indicates that the smaller the indicator data, the better the effect. It can be seen that the model of the present invention outperforms other models in SIM and KLD. These two indicators measure the similarity between distributions, while other methods will misclassify non-significant areas, resulting in poor performance. And in comparison with the self-supervised model, the model of the present invention has improved in all indicators. In comparison with the direct prediction model, except for NSS (which tends to ignore false positives), our models have achieved similar or higher indicators. MV-SalGAN360 achieves the best performance in the first three indicators due to its multi-view fusion method, but this method takes more time in the inference phase. ATSAL does not have outstanding performance in images because it is a video model.

[0100] Table 2. The time required for various methods to predict a single image using Salient360! as input.

[0101]

[0102]

[0103] Table 2 shows the difference in inference speed between different models. For fair comparison, all calculations are performed on an i5-9400 CPU in a Windows environment. Due to the more complex decoder, the model of our invention takes a little longer than Rethink (Rethink+ uses the same decoder as our method and is not better in terms of indicators). MV-SalGAN360 and ATSAL require projection and multi-view fusion, which leads to longer inference time.

[0104] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A panoramic image saliency prediction method based on self-supervised learning, characterized in that: The following steps are involved: S1. Train the encoder using the unlabeled ERP image set, including the following sub-steps: S11. Format conversion: Project the ERP image onto a spherical surface to obtain a CMP image group C i and label P i , i=1,…,6; S12, for C i Randomly shuffle to get c i , and according to c i The original position of P i Update to get the agent task label S13. Train the encoder and build a global feature extraction network With local feature extraction network The global features and local features are taken as input, and the features of the two are learned through feature fusion, and the model parameters of the global feature extraction network are updated; the global feature extraction network With local feature extraction network They are: where F E is a global feature, It is a local feature, E represents the ERP image; -> represents the reasoning process of the feature extraction network, the global feature extraction network With local feature extraction network All models use VGG16 with the last 5 layers removed; Then the obtained global feature F E and local features Jointly input into the feature fusion network; The feature fusion network includes two parts: feature transformation and dot multiplication operation: First, F E and After two fully connected layers with unshared weights, we get r E and Then transform it by the following equation: Q E =r E W Q Where W Q , W V and W K is the weight that the three types of features do not share, Q E , and Represents Query, Value and Key respectively; Then use the dot product operation to fuse the obtained features: Among them, CA i is the result of feature fusion, ReLU is the activation function, Represents a function nesting operator; Obtained CA i is used for the final position prediction: Training is performed using the following loss function: The loss function calculates the difference between the predicted value and the label value, and then performs gradient backpropagation based on the difference and updates the gradient based on the gradient. The parameters in the model are stopped after traversing the unlabeled ERP image set 100 times to obtain the global feature extraction network S2. Decoder training: Decoder Constructed to predict the final significant result S3. The panoramic image to be identified is input into the trained encoder for feature extraction, and then the extracted features are input into the decoder to obtain the final saliency prediction.

2. The method for predicting panoramic image saliency based on self-supervised learning according to claim 1, characterized in that: The specific implementation method of step S2 is: Obtaining salient images: The head and eye movement recording files are used as the training set for the decoder. First, a zero matrix with the same size as the image in the training set is established. Different viewpoint positions are recorded in the head and eye recording files. If a point is recorded in the file, it is marked as 1 in the matrix. According to the recorded position, the zero matrix is ​​updated using the following method: S ij It is the viewpoint graph; and the viewpoint graph is difficult to train due to its sparse matrix characteristics, so the following processing is performed: where G is a Gaussian kernel with a dilation angle of 5°, S E Indicated by S ij The matrix formed; Update as follows: Among them, T represents the process of transformation from ERP to CMP. back Represents the process of transformation from CMP to ERP; Loss function: According to the characteristic that most areas of the saliency image are 0, the following loss function is selected to train the decoder model: is the predictive distribution, is the true distribution, ε is a constant set to prevent the predicted value from being too close to 0 and causing the loss to tend to infinity, W E , H E are the width and height of the image respectively; The loss function calculates the difference between the predicted value and the label value, and then performs gradient backpropagation based on the difference and updates the parameters in the model based on the gradient. After traversing the training set one hundred times, the decoder g is obtained. θ .

Citation Information

Patent Citations

  • Deep learning saliency detection method based on global a priori and local context

    CN107274419A

  • Visual saliency detection method combined with image classification

    CN107346436A