Weakly Supervised RGBD Image Salience Detection Method Based on Edge Detection Assistance
Through edge detection assisted methods, a multi-level, multi-tasking RGBD image significance detection network is designed, and combined with color images and depth map features, cross-modal conflict and edge positioning problems in weakly supervised RGBD image significance detection are solved, improving detection performance.
Patent Information
- Application Number
- CN202211575959.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-12-08
AI Technical Summary
In weakly supervised RGBD image significance detection, cross-modal conflicts between color images and depth maps, rough edges and noise problems of depth maps lead to increased difficulty in detection of significance targets, especially in complex scenarios.
Through edge detection assisted methods, a multi-level, multi-tasking weakly supervised RGBD image significance detection network is designed, combined with color image and depth map features, and a cross-modal feature fusion module and edge refinement module are used to optimize the loss function to improve detection performance.
Effectively utilize the advantages of color images and depth maps to avoid the problems brought by depth maps, and achieve better performance weakly supervised RGBD images significance detection.
Smart Images

Figure CN116524207B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and computer vision, and particularly to a weakly supervised RGBD image saliency detection method assisted by edge detection. Background Art
[0002] Saliency object detection is an important research content in the field of computer vision. Its goal is to simulate the human visual perception system to find the most attention-grabbing objects in an image and perform pixel-level segmentation on them. As a basic image processing problem, it plays a key role in tasks such as object detection, semantic segmentation, video tracking, and image understanding.
[0003] With the development of convolutional neural networks, many deep learning-based image saliency detection methods have been proposed. Compared with traditional methods, these methods have greatly improved performance. However, deep learning requires a large amount of training data as support, and the cost of obtaining pixel-by-pixel annotation labels required by strongly supervised saliency detection models is extremely high. Therefore, weakly supervised image saliency detection has now become a research direction actively explored by many scholars.
[0004] Weakly supervised image saliency detection models incomplete weak-level annotations and then relies on the strong generalization ability of the model to infer complete saliency objects. Commonly used weak-level annotations include noise labels, image-level labels, bounding boxes, and scribble labels, etc. Compared with pixel-by-pixel annotation labels, these low-cost labels cannot provide complete structural details of saliency objects, which poses a greater challenge for the saliency detection network model to restore the delicate edge structure of saliency objects. Currently, most methods choose to introduce traditional unsupervised saliency detection methods, image classification tasks, or edge detection tasks, etc. as assistance, and use them to help determine the position and edges of saliency objects. However, in some complex scenarios, relying solely on the color and texture features provided by color images, the edge localization problem that is difficult to solve by strongly supervised saliency detection will become even more difficult in the weakly supervised case. Weakly supervised RGBD image saliency detection can improve the saliency object detection ability in complex scenarios by introducing a depth map and using the rich structural information and position information contained in the depth map as a supplement. However, when introducing the depth map, it also brings new problems, such as cross-modal conflicts between color images and depth maps, rough edges of depth maps, and noise problems caused by low-quality depth maps, etc. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a weakly supervised RGBD image saliency detection method assisted by edge detection, which can achieve better performance in weakly supervised RGBD image saliency detection.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A weakly supervised RGBD image saliency detection method based on edge detection assistance, comprising the following steps:
[0007] Step S1: Establish a weakly supervised RGBD image saliency detection training set containing scribble annotation maps and perform data augmentation;
[0008] Step S2: Design a multi-level and multi-task weakly supervised RGBD image saliency detection network, and use this network to obtain a saliency prediction result with multi-scale edge refinement;
[0009] Step S3: Design a fusion module, and use this module to fuse the saliency prediction results with multi-scale edge refinement to obtain the final saliency prediction result;
[0010] Step S4: Design a weakly supervised RGBD image saliency detection network based on edge detection assistance, and design a loss function to optimize the network parameters to obtain a trained weakly supervised RGBD image saliency detection model based on edge detection assistance;
[0011] Step S5: Input the to-be-tested RGBD image into the trained weakly supervised RGBD image saliency detection model based on edge detection assistance to obtain a saliency detection result.
[0012] In a preferred embodiment, the specific content of step S1 is as follows:
[0013] Step S11: Divide the data set, and divide it into a training set and a test set according to a certain ratio;
[0014] Step S12: For the training set, use the brush tool in the "Adobe Photoshop 2020" software to perform scribble annotation on each group of RGBD images. Specifically, use black scribbles to annotate some salient foreground regions, use white scribbles to annotate the background regions, and use gray for the unannotated regions;
[0015] Step S13: Perform data augmentation on the images in the training set. The specific operations include adding noise, randomly cropping, flipping the images, and normalizing the color images and depth maps of each group of RGBD images in the training set and the test set to highlight the foreground regions.
[0016] In a preferred embodiment, the specific content of step S2 is as follows:
[0017] Step S21: First, input the color image and the depth map into two VGG16 networks respectively. Then, use the features extracted by the 5 convolutional layers Conv1, Conv2, Conv3, Conv4, and Conv5 and the pooling layer Pool5 as multi-level color image features and multi-level depth map features
[0018] Step S22: Design an initial saliency prediction branch. At each of the 6 levels, first concatenate the color image features and the depth map features Then send the concatenated features into the cross-modal feature fusion module CFF for fusing the color image features and the depth map features. The cross-modal feature fusion module consists of a 3×3 convolutional layer, channel attention, spatial attention, and a 3×3 convolutional layer connected in series. Finally, the fused features are reduced to 1 dimension through a convolutional layer with a kernel size of 1. This process is expressed by the following formula:
[0019]
[0020] where represents the initial saliency feature of the k-th layer, and represent the color image feature and the depth map feature of the k-th layer respectively, ⊕ represents the concatenation operation, F CFF represents the cross-modal feature fusion module in the initial saliency prediction branch, and Conv 1×1 represents the convolutional layer with a kernel size of 1;
[0021] Step S23: Design an edge detection branch to obtain the edge feature E k The process is the same as that of the initial saliency prediction branch, and the formula is as follows:
[0022]
[0023] where E k represents the edge feature of the k-th layer, and represent the color image feature and the depth map feature of the k-th layer respectively, ⊕ represents the concatenation operation, F CFF ' represents the cross-modal feature fusion module in the edge detection branch, and Conv 1×1 represents the convolutional layer with a kernel size of 1.
[0024] Step S24: Design an edge-refined saliency prediction module; at each of the 6 levels, first concatenate the initial saliency feature and the edge feature E k , and then reduce the dimension of the concatenated features to 1 dimension through a convolutional layer with a kernel size of 1. The formula is as follows:
[0025]
[0026] where S k represents the edge-refined saliency feature of the k-th layer, and Ek respectively represent the initial saliency feature and the edge feature of the k-th layer, ⊕ represents the concatenation operation, and Conv 1×1 represents a convolutional layer with a convolutional kernel of 1.
[0027] In a preferred embodiment, the step S3 is specifically as follows:
[0028] Step S31: Design a fusion module; design a fusion module to integrate deep features into shallow features layer by layer. The specific process is expressed by the following formula:
[0029]
[0030] S final = σ(Conv 3×3 (H1))
[0031] where H k represents the aggregated feature of the k-th layer, S k represents the saliency feature with edge refinement of the k-th layer, F up represents upsampling, Conv 3×3 represents a convolutional layer with a convolutional kernel of 3, σ represents the Sigmoid activation function, and S final represents the final saliency prediction result.
[0032] In a preferred embodiment, the step S4 is specifically as follows:
[0033] Step S41: Combine the multi-level and multi-task weakly supervised RGBD image saliency detection network designed in step S2 and the fusion module designed in step S3 to obtain a weakly supervised RGBD image saliency detection network assisted by edge detection;
[0034] Step S42: Design the loss function of the weakly supervised RGBD image saliency detection network assisted by edge detection as follows:
[0035]
[0036] where L represents the finally trained loss function, ∑ represents summation, k ∈ {1,…6}, and are the partial cross-entropy losses acting on the k-th layer of the initial saliency prediction branch, the k-th layer of the edge-refined saliency prediction module, and the final saliency prediction result respectively, and are the smooth losses acting on the k-th layer of the initial saliency prediction branch, the k-th layer of the edge-refined saliency prediction module, and the final saliency prediction result respectively, is the cross-entropy loss acting on the k-th layer of the edge detection branch. and The specific calculation formulas are as follows:
[0037]
[0038] S k ′ = σ(S k )
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045] E k ′ = σ(E k )
[0046]
[0047] where σ represents the Sigmoid activation function, and respectively represent the initial saliency feature and the initial saliency prediction map of the k-th layer in the initial saliency prediction branch. S k and S k ′ respectively represent the edge-refined saliency feature and the edge-refined saliency prediction map of the k-th layer in the edge-refined saliency prediction module. Y represents the input scribble annotation map, U represents the scribble area in the scribble annotation map Y, (i, j) ∈ U represents the pixel located in the scribble area, log represents the log function, S final represents the final saliency prediction result map, Δ represents the derivative, ΔI[i, j], ΔG[i, j] and ΔS final [i, j] respectively represent the maps obtained by taking the derivative of the initial saliency prediction map, the edge-refined saliency prediction map, the color image, the depth map and the final saliency prediction result map of the k-th layer. |·| represents taking the absolute value, e is a constant, and α is a fixed parameter, is defined as to avoid the result being 0. E k and E k′ represent the edge features and the edge map of the k-th layer in the edge detection branch respectively, E represents the input edge map, [i, j] represents the pixels in the i-th row and j-th column of the image, Y[i, j], S final [i, j], ΔS′ k , ΔI[i, j], ΔG[i, j], E[i, j] and E k ′[i, j] represent the values of the pixels at the i-th row and j-th column of the images Y, S′ k , S final , ΔS′ k , ΔI, ΔG, E and E k ′ respectively;
[0048] Step S43: Repeat the above steps S2 to S4 in batches until the loss function value calculated in step S4 converges and stabilizes, save the network parameters, complete the training process of the weakly supervised RGBD image saliency detection network assisted by edge detection, and obtain the weakly supervised RGBD image saliency detection model assisted by edge detection.
[0049] Compared with the prior art, the present invention has the following beneficial effects: while making full use of the advantages provided by the combination of color images and depth maps, avoiding the problems brought by depth maps, and realizing better performance of weakly supervised RGBD image saliency detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is the implementation flowchart of the preferred embodiment of the present invention.
[0051] Figure 2 is an example of a set of RGBD images and their corresponding scribble annotation maps in the preferred embodiment of the present invention.
[0052] Figure 3 is the network model structure diagram in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0054] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0055] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0056] The present invention provides a weakly supervised RGBD image saliency detection method based on edge detection assistance, as Figure 1 shown, which includes the following steps:
[0057] Step S1: Establish a weakly supervised RGBD image saliency detection training set containing scribble annotation maps and perform data augmentation;
[0058] Step S2: Design a multi-level and multi-task weakly supervised RGBD image saliency detection network and use this network to obtain a saliency prediction result with multi-scale edge refinement;
[0059] Step S3: Design a fusion module and use this module to fuse the saliency prediction results with multi-scale edge refinement to obtain the final saliency prediction result;
[0060] Step S4: Design a weakly supervised RGBD image saliency detection network based on edge detection assistance and design a loss function to optimize the network parameters to obtain a trained weakly supervised RGBD image saliency detection model based on edge detection assistance;
[0061] Step S5: Input the RGBD image to be measured into the trained weakly supervised RGBD image saliency detection model based on edge detection assistance to obtain the saliency detection result.
[0062] Further, step S1 specifically includes the following steps:
[0063] Step S11: Divide the data set and divide it into a training set and a test set according to a certain ratio;
[0064] Step S12: For the training set, use the brush tool in the "Adobe Photoshop 2020" software to perform scribble annotation on each group of RGBD images. Specifically, use black scribbles to annotate some salient foreground regions, use white scribbles to annotate the background regions, and use gray for the unannotated regions;
[0065] Step S13: Perform data augmentation on the images in the training set. The specific operations include adding noise, randomly cropping, flipping the images, and normalizing the color images and depth maps of each group of RGBD images in the training set and the test set to highlight the foreground regions.
[0066] Furthermore, step S2 specifically includes the following steps:
[0067] Step S21: First, input the color image and the depth map into two VGG16 networks respectively. Then, take the features extracted from the 5 convolutional layers Conv1, Conv2, Conv3, Conv4, and Conv5 and the pooling layer Pool5 at 6 levels as multi-level color image features and multi-level depth map features
[0068] Step S22: Design an initial saliency prediction branch. At each of the 6 levels, first concatenate the color image features and the depth map features Then, send the concatenated features into the cross-modal feature fusion module CFF for fusing the color image features and the depth map features. The cross-modal feature fusion module consists of a 3×3 convolutional layer, channel attention, spatial attention, and a 3×3 convolutional layer in series. Finally, the fused features are reduced to 1 dimension through a convolutional layer with a kernel size of 1. This process is expressed by the following formula:
[0069]
[0070] where represents the initial saliency feature at the k-th level, and represent the color image feature and the depth map feature at the k-th level respectively, ⊕ represents the concatenation operation, F CFF represents the cross-modal feature fusion module in the initial saliency prediction branch, and Conv 1×1 represents the convolutional layer with a kernel size of 1;
[0071] Step S23: Design an edge detection branch to obtain the edge feature E k The process of obtaining it is the same as that of the initial saliency prediction branch, and the formula is as follows:
[0072]
[0073] where E k represents the edge feature at the k-th level, and represent the color image feature and the depth map feature at the k-th level respectively, ⊕ represents the concatenation operation, F CFF ' represents the cross-modal feature fusion module in the edge detection branch, and Conv 1×1 represents the convolutional layer with a kernel size of 1.
[0074] Step S24: Design an edge-refined saliency prediction module; at each of the 6 levels, first concatenate the initial saliency features and edge feature E k , and then reduce the dimension of the concatenated features to 1 dimension through a convolutional layer with a convolution kernel of 1. The formula is as follows:
[0075]
[0076] where S k represents the saliency feature of edge refinement in the k-th layer, and E k represent the initial saliency feature and edge feature in the k-th layer respectively. ⊕ represents the concatenation operation, and Conv 1×1 represents a convolutional layer with a convolution kernel of 1.
[0077] Furthermore, step S3 specifically includes the following steps:
[0078] Step S31: Design a fusion module; design a fusion module to integrate deep features into shallow features layer by layer. The specific process is expressed by the formula as follows:
[0079]
[0080] S final = σ(Conv 3×3 (H1))
[0081] where H k represents the aggregated feature in the k-th layer, S k represents the saliency feature of edge refinement in the k-th layer, F up represents upsampling, Conv 3×3 represents a convolutional layer with a convolution kernel of 3, σ represents the Sigmoid activation function, and S final represents the final saliency prediction result.
[0082] Furthermore, step S4 specifically includes the following steps:
[0083] Step S41: Combine the multi-level and multi-task weakly supervised RGBD image saliency detection network designed in step S2 and the fusion module designed in step S3 to obtain a weakly supervised RGBD image saliency detection network assisted by edge detection;
[0084] Step S42: Design the loss function of the weakly supervised RGBD image saliency detection network assisted by edge detection as follows:
[0085]
[0086] where L represents the finally trained loss function, ∑ represents summation, k ∈ {1, … 6}, and They are the partial cross-entropy losses acting on the k-th layer of the initial saliency prediction branch, the k-th layer of the edge-refined saliency prediction module, and the final saliency prediction result respectively. and They are the smooth losses acting on the k-th layer of the initial saliency prediction branch, the k-th layer of the edge-refined saliency prediction module, and the final saliency prediction result respectively. is the cross-entropy loss acting on the k-th layer of the edge detection branch. and The specific calculation formulas of and are as follows:
[0087]
[0088] S k ′ = σ(S k )
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095] E k ′ = σ(E k )
[0096]
[0097] where σ represents the Sigmoid activation function, and respectively represent the initial saliency feature and the initial saliency prediction map of the k-th layer in the initial saliency prediction branch, S k and S k ′ respectively represent the edge-refined saliency feature and the edge-refined saliency prediction map of the k-th layer in the edge-refined saliency prediction module, Y represents the input scribble annotation map, U represents the scribble area in the scribble annotation map Y, (i, j) ∈ U represents the pixel located in the scribble area, log represents the log function, S final represents the final saliency prediction result map, Δ represents the derivative, ΔS′ k 、ΔI[i,j], ΔG[i,j] and ΔS final[i, j] respectively represent the graphs obtained by taking the derivatives of the initial saliency prediction map of the k-th layer, the edge-refined saliency prediction map of the k-th layer, the color image, the depth map, and the final saliency prediction result map. |·| represents taking the absolute value, e is a constant, and α is a fixed parameter. is defined as to avoid the result being 0, E k and E k ′ respectively represent the edge features of the k-th layer and the edge map of the k-th layer in the edge detection branch. E represents the input edge map. [i, j] represents the pixel at the i-th row and j-th column of the image. Y[i, j], S′ k [i, j], S final [i, j], ΔS′ k , ΔI[i, j], ΔG[i, j], E[i, j], and E k ′[i, j] respectively represent the values at the pixel at the i-th row and j-th column of the images Y, S′ k , S final , ΔS′ k , ΔI, ΔG, E, and E k ′;
[0098] Step S43: Repeat the above steps S2 to S4 in batches until the loss function value calculated in step S4 converges and stabilizes. Save the network parameters to complete the training process of the weakly supervised RGBD image saliency detection network assisted by edge detection and obtain the weakly supervised RGBD image saliency detection model assisted by edge detection.
[0099] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention that do not exceed the scope of the technical solutions of the present invention in terms of the functions and effects produced belong to the protection scope of the present invention.
Claims
1. A weakly supervised RGBD image saliency detection method based on edge detection assistance, characterized in that, It includes the following steps: Step S1: Establish a weakly supervised RGBD image saliency detection training set containing scribble annotation maps and perform data augmentation; Step S2: Design a multi-level and multi-task weakly supervised RGBD image saliency detection network, and use this network to obtain a saliency prediction result with multi-scale edge refinement; Step S3: Design a fusion module, and use this module to fuse the saliency prediction results with multi-scale edge refinement to obtain the final saliency prediction result; Step S4: Design a weakly supervised RGBD image saliency detection network assisted by edge detection, and design a loss function to optimize the network parameters to obtain a trained weakly supervised RGBD image saliency detection model assisted by edge detection; Step S5: Input the to-be-tested RGBD image into the trained weakly supervised RGBD image saliency detection model assisted by edge detection to obtain the saliency detection result; The specific content of step S2 is as follows: Step S21: First, input the color image and the depth map into two VGG16 networks respectively. Then, take the features of six levels extracted by the five convolutional layers Conv1, Conv2, Conv3, Conv4, and Conv5 and the pooling layer Pool5 as multi-level color image features k = 1…6, and multi-level depth map features k = 1…6; Step S22: Design an initial saliency prediction branch. At each of the six levels, first concatenate the color image features and the depth map features Then send the concatenated features into the cross-modal feature fusion module CFF to fuse the color image features and the depth map features; The cross-modal feature fusion module consists of a 3×3 convolutional layer, channel attention, spatial attention, and a 3×3 convolutional layer connected in series; finally, the fused features are reduced to 1 dimension through a convolutional layer with a kernel of 1, and this process is expressed by the following formula: Among them represents the initial saliency feature of the k-th layer, and respectively represent the color image feature and the depth map feature of the k-th layer, represents the splicing operation, F CFF represents the cross-modal feature fusion module in the initial saliency prediction branch, Conv 1×1 represents the convolutional layer with a convolutional kernel of 1; Step S23: Design an edge detection branch to obtain edge feature E k The process is the same as that of the initial saliency prediction branch, and the formula is as follows: Among them, E k represents the edge feature of the k-th layer, and respectively represent the color image feature and the depth map feature of the k-th layer, represents the splicing operation, and F CFF ' represents the cross-modal feature fusion module in the edge detection branch, and Conv 1×1 represents the convolutional layer with a convolutional kernel of 1; Step S24: Design an edge-refined saliency prediction module; at each of the 6 levels, first concatenate the initial saliency features and the edge feature E k , and then reduce the dimension of the concatenated features to 1 dimension through a convolutional layer with a convolution kernel of 1. The formula is as follows: where S k represents the significant features of edge refinement in the k-th layer, and E k represent the initial significant features and edge features in the k-th layer respectively, represents the concatenation operation, and Conv 1×1 represents a convolutional layer with a convolutional kernel of 1; The specific content of step S3 is as follows: Step S31: Design a fusion module; design a fusion module to integrate deep features into shallow features layer by layer, and the specific process is expressed by the following formula: S final = σ(Conv 3×3 (H1)) where H k represents the aggregated feature of the k-th layer, S k represents the saliency feature of the edge refinement of the k-th layer, F up represents upsampling, Conv 3×3 represents a convolutional layer with a convolutional kernel of 3, σ represents the Sigmoid activation function, S final represents the final saliency prediction result; The specific content of step S4 is as follows: Step S41: Combine the multi-level and multi-task weakly supervised RGBD image saliency detection network designed in step S2 and the fusion module designed in step S3 to obtain a weakly supervised RGBD image saliency detection network assisted by edge detection; Step S42: Design the loss function of the weakly supervised RGBD image saliency detection network assisted by edge detection as follows: Where L represents the finally trained loss function, ∑ represents summation, k ∈ {1, … 6}, and are the partial cross-entropy losses acting on the k-th layer of the initial saliency prediction branch, the k-th layer of the edge-refined saliency prediction module, and the final saliency prediction result respectively, and are the smooth losses acting on the k-th layer of the initial saliency prediction branch, the k-th layer of the edge-refined saliency prediction module, and the final saliency prediction result respectively, is the cross-entropy loss acting on the k-th layer of the edge detection branch; and The specific calculation formulas of are as follows: S k ' = σ(S k ) E k ' = σ(E k ) where σ represents the Sigmoid activation function, and respectively represent the initial saliency feature at the k-th layer and the initial saliency prediction map at the k-th layer in the initial saliency prediction branch, S k and S k ' respectively represent the edge-refined saliency feature at the k-th layer and the edge-refined saliency prediction map at the k-th layer in the edge-refined saliency prediction module, Y represents the input scribble annotation map, U represents the scribble area in the scribble annotation map Y, (i, j) ∈ U represents the pixel located in the scribble area, log represents the log function, S final represents the final saliency prediction result map, Δ represents the derivative, ΔS' k 、ΔI[i, j], ΔG[i, j] and ΔS final [i, j] respectively represent the maps after taking the derivative of the initial saliency prediction map at the k-th layer, the edge-refined saliency prediction map at the k-th layer, the color image, the depth map and the final saliency prediction result map, |·| represents taking the absolute value, e is a constant, α is a fixed parameter, is defined as to avoid the result being 0, E k and E k ' respectively represent the edge feature at the k-th layer and the edge map at the k-th layer in the edge detection branch, E represents the input edge map, [i, j] represents the pixel at the i-th row and j-th column of the image, Y[i, j], S′ k [i, j], S final [i, j], ΔS' k 、ΔI[i, j], ΔG[i, j], E[i, j] and E k ′[i, j] respectively represent the values at the i-th row and j-th column pixels of the image Y, S′ k 、S final 、 ΔS' k 、ΔI, ΔG, E and E k '; Step S43: Repeat the above steps S2 to S4 in batches until the loss function value calculated in step S4 converges and stabilizes, save the network parameters, complete the training process of the weakly supervised RGBD image saliency detection network assisted by edge detection, and obtain a weakly supervised RGBD image saliency detection model assisted by edge detection.
2. The weakly supervised RGBD image saliency detection method based on edge detection assistance according to claim 1, wherein The specific content of step S1 is as follows: Step S11: Divide the data set into a training set and a test set; Step S12: For the training set, use the brush tool in the "Adobe Photoshop 2020" software to perform scribble annotation on each group of RGBD images. Specifically, use black scribbles to annotate some salient foreground regions, use white scribbles to annotate the background regions, and use gray for the unannotated regions; Step S13: Perform data augmentation on the images in the training set. The specific operations include adding noise, random cropping, flipping the images, and normalizing the color images and depth maps of each group of RGBD images in the training set and the test set.
Citation Information
Patent Citations
Weak supervision RGBD image saliency detection method and system based on image classification
CN112861880A
Retinal image quality assessment, error identification and automatic quality correction
US20170270653A1