A method and system for extracting significant regions of community images
By integrating image appearance features and label semantic features in the detection of significant areas of community images, and solving fusion parameters through the fusion coefficient model, the problem that detection effect in the prior art depends on area annotation and training process is solved, and more efficient and accurate significant area detection is achieved.
Patent Information
- Application Number
- CN202111245793.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-26
AI Technical Summary
In the detection of prominent areas of community images, the prior art cannot effectively integrate image appearance features and label semantic features, resulting in the detection effect dependent on area annotation and the training process is cumbersome.
The proposed method extracts appearance features and calculates appearance characteristics for the image during the training stage, and calculates object label semantic features through object label detection. Then, the two are fused through the fusion coefficient model to solve the fusion parameters. During the testing phase, the training-derived fusion parameters are used to fuse the image appearance and label semantic features to generate the final significant graph.
By integrating image appearance features and label semantic features, the accuracy and efficiency of significant area detection of community images are improved, the training process is simplified, and the dependence on area annotation is reduced.
Smart Images

Figure CN113936147B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method and system for extracting salient regions of community images. Background Art
[0002] Attention belongs to the cognitive process of human beings, is a psychological concept, and is an important part of visual perception. Salience detection by simulating the attention mechanism through a computer involves related fields such as psychology, neuroscience, biological vision, and computer vision, and is a multi-disciplinary research field. Traditional salience detection methods usually use a variety of salience cues or prior information, such as local or global contrast, boundary prior. Since these methods use low-level manually designed features and models, they cannot recognize and understand the semantic object concepts in images. Recently, deep convolutional neural networks have achieved remarkable results in visual pattern recognition methods and have been increasingly applied to the detection of salient regions. As long as sufficient training data is provided, deep convolutional neural networks can accurately identify salient objects in complex images, with performance exceeding most traditional methods based on manually designed features and achieving good detection effects.
[0003] With the rapid development of the Internet and social platforms, a large number of social images have emerged, and they carry tag information. Although the semantics of tags have been widely used in the field of image annotation, there is not much work on applying them to salient object extraction. The literature [Wen Wang, Congyan Lang, Songhe Feng. Contextualizing Tag Ranking and Saliency Detection for Social Images. Advances in Multimedia Modeling Lecture Notes in Computer Science Volume 7733, 2013, pp 428 - 435.] integrates the tag ranking task and the saliency detection task, and iteratively performs the tag ranking and saliency detection tasks. The literature [Zhu, G., Wang, Q., Yuan, Y. Tag - saliency: Combining bottom - up and top - down information for saliency detection. Computer Vision and Image Understanding, 2014, 118(1): 40 - 49.] proposes the Tag - Saliency model, which annotates multimedia data through hierarchical over - segmentation and automatic annotation techniques. The common drawback of these two pieces of literature is that the effect of saliency annotation depends on region annotation, and the tag information is processed separately from the task of extracting salient regions.
[0004] In 2021, the Journal of Intelligent Systems published the article "A Method for Detecting Salient Regions in Community Images" by Ye Liang and Jian Yu, which focuses on the problem of detecting salient regions in community images and proposes a method for detecting salient regions based on deep features. Considering the characteristics of community images with tags, in the system framework, this paper adopts two extraction lines, saliency calculation based on CNN features and semantic calculation based on tags, and the results of the two are fused. Finally, the spatial consistency of the fused saliency map is optimized through a fully - connected conditional random field model. The drawback of this method is that it does not treat the image appearance features and tag semantic features as an integrated feature, which is rather cumbersome during training. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method and system for extracting significant regions of community images. In the training stage, appearance features are extracted from training images, and saliency features based on image appearance are calculated; object detection is performed through an object detection sub corresponding to the object label carried by the image, and object label semantic features are calculated; the object label semantic features can be regarded as a kind of prior feature, which is fused with the saliency features based on appearance, and the fusion problem of saliency features is modeled to solve the fusion parameters. In the testing stage, appearance features are extracted from test images, and saliency features based on image appearance are calculated; object detection is performed through an object detection sub corresponding to the object label carried by the image, and object label semantic features are calculated; finally, the saliency features of the image appearance and the label semantic features are fused through the fusion parameters obtained by training to obtain the final saliency map.
[0006] The first object of the present invention is to provide a method for extracting significant regions of community images, including preparing a training image set, and further including the following steps:
[0007] Step 1: Perform saliency calculation based on appearance on the images in the training image set;
[0008] Step 2: Calculate the semantic features of the object labels corresponding to the object labels carried by each image in the training image set;
[0009] Step 3: Solve the fusion coefficient of the saliency features;
[0010] Step 4: Extract the saliency features of the test image and calculate the saliency map.
[0011] Preferably, the preparation of the training image set further includes preparing a standard binary annotation of the significant region corresponding to the training image and / or a set of labels.
[0012] In any of the above solutions, preferably, step 1 includes extracting the appearance features of the image and calculating the corresponding saliency features to obtain K-dimensional appearance saliency features.
[0013] In any of the above solutions, preferably, the saliency calculation method on the k-th dimensional saliency feature channel is:
[0014]
[0015] where D(v i k ,v j k ) represents the difference between pixel x i and pixel x j on the k-th dimensional feature channel, w ijRepresents the spatial distance weight, where 1 ≤ k ≤ K, i represents the i-th pixel, and j represents the j-th pixel different from the i-th pixel.
[0016] Preferably, in any of the above solutions, step 2 includes the following sub-steps:
[0017] Step 21: Define the object label semantic feature p of pixel x in the image x Initialize it as {0, 0,..., 0}, M is the number of object labels included in the training image set, q represents the q-th object label;
[0018] Step 22: Check one by one whether the image contains M object labels;
[0019] Step 23: Calculate the object label semantic feature p of pixel x x .
[0020] Preferably, in any of the above solutions, step 22 includes:
[0021] 1) When the m-th object label exists, where 1 ≤ m ≤ M, then detect the N proposed rectangular boxes corresponding to the object label m, and the possibility of each rectangular box containing the object is f m n , then the semantic feature of object m for each pixel covered by the rectangular box is f m n , where 1 ≤ n ≤ N. If the pixel is not covered by any rectangular box of the m-th object label, then the semantic feature of the m-th object label of the pixel is 0;
[0022] 2) When the m-th object label of the image does not exist, then f m n = 0, where 1 ≤ n ≤ N.
[0023] Preferably, in any of the above solutions, the calculation formula for the object label semantic feature p x is:
[0024]
[0025] where, is a flag variable used to indicate whether the n-th proposed rectangular box is detected. If it is detected, then otherwise
[0026] Preferably, based on the calculations of step 2 and step 3, a total of K + M-dimensional saliency features are obtained, including K-dimensional appearance-based saliency features and M-dimensional label semantic features.
[0027] Preferably, in any of the above solutions, step 3 includes fusing the K+M-dimensional saliency features, and the fusion coefficient is solved by the following formula:
[0028]
[0029] When a x =0, it means that pixel x is labeled as non-salient, and F t (a x ) = S′(x);
[0030] When a x =1, it means that pixel x is labeled as salient, and F t (a x ) = 1 - S′(x), where S′(x) represents the saliency value of pixel x;
[0031] Among them, S′(x) represents the saliency value of pixel x, G(a x , a x′ ) = |a x -a x′ |.d(x, x′), where d(x, x’) represents the color difference between pixels x and x′; i represents the i-th image in the training set, y represents the y-th image in the training set, 1 ≤ y ≤ Q, and Q represents the number of images in the training set.
[0032] Preferably, in any of the above solutions, the optimization parameter is
[0033] Preferably, in any of the above solutions, step 4 includes performing saliency prediction on the test image, and splicing the K-dimensional saliency features s′ i (k) based on appearance features and the M-dimensional object label semantic features p′ x to obtain the K+M-dimensional saliency feature A of the test image. The saliency map of the test image is calculated as follows
[0034]
[0035] The second object of the invention is to provide a saliency region extraction system for community images, including an acquisition module for preparing a training image set, and further including the following modules:
[0036] Calculation module: used for performing appearance-based saliency calculation on the images in the training image set, and also used for calculating the semantic features of the corresponding object labels for each image in the training image set;
[0037] Solution module: used for solving the fusion coefficient of the saliency features;
[0038] Test module: used for extracting the saliency features of the test image and calculating the saliency map.
[0039] Preferably, the preparation of the training image set further includes preparing a set of standard binary annotations and / or labels corresponding to the training images and their corresponding salient regions.
[0040] In any of the above solutions, preferably, the calculation module is configured to extract the appearance features of the image and calculate the corresponding saliency features to obtain K-dimensional appearance saliency features.
[0041] In any of the above solutions, preferably, the saliency calculation method on the k-th dimensional saliency feature channel is as follows:
[0042]
[0043] where D(v i k , v j k ) represents the difference between pixel x i and pixel x j on the k-th dimensional feature channel, w ij represents the spatial distance weight, 1 ≤ k ≤ K, i represents the i-th pixel, and j is the j-th pixel different from the i-th pixel.
[0044] In any of the above solutions, preferably, the semantic feature calculation method of the object label includes the following sub-steps:
[0045] Step 21: Define the object label semantic feature p x of pixel x in the image and initialize it to {0, 0,..., 0}, M is the number of object labels included in the training image set,, q represents the q-th object label;
[0046] Step 22: Check one by one whether the image contains M object labels;
[0047] Step 23: Calculate the object label semantic feature p x of pixel x.
[0048] In any of the above solutions, preferably, Step 22 includes:
[0049] 1) When the m-th object label exists, 1 ≤ m ≤ M, then detect N proposed rectangular boxes corresponding to the object label m, and the possibility of each rectangular box containing the object is f m n , then the semantic feature of object m of each pixel covered by the rectangular box is f m n , 1 ≤ n ≤ N. If the pixel is not covered by any rectangular box of the m-th object label, the semantic feature of the m-th object label of the pixel is 0;
[0050] 2) When the m-th object label of the image does not exist, then f m n = 0, 1 ≤ n ≤ N.
[0051] Preferably, in any of the above solutions, the semantic feature p of the object label x has the following calculation formula:
[0052]
[0053] where, is a marking variable used to indicate whether the n-th proposed rectangle is detected. If it is detected, then otherwise
[0054] Preferably, in any of the above solutions, based on the calculation of the saliency and the semantic feature of the object label, a total of K + M-dimensional saliency features are obtained, including K-dimensional appearance-based saliency features and M-dimensional label semantic features.
[0055] Preferably, in any of the above solutions, the solving module is used to fuse the K + M-dimensional saliency features, and the fusion coefficient is solved by the following formula,
[0056]
[0057] When a x = 0, it means that the pixel x is labeled as non-salient, and F t (a x ) = S′(x);
[0058] When a x = 1, it means that the pixel x is labeled as salient, and F t (a x ) = 1 - S′(x), where S′(x) represents the saliency value of the pixel x;
[0059] where, S′(x) represents the saliency value of the pixel x, and G(a x , a x′ ) = |a x - a x′ |.d(x, x′), d(x, x’) represents the color difference between pixels x and x′; i represents the i-th image in the training set, y represents the y-th image in the training set, 1 ≤ y ≤ Q, and Q represents the number of images in the training set.
[0060] Preferably, in any of the above solutions, the optimization parameter is
[0061] Preferably, in any of the above solutions, the test module is used to perform saliency prediction on a test image, and the K-dimensional saliency feature s′ based on appearance features i (k) and the M-dimensional object label semantic feature p′ x are concatenated to obtain the K+M-dimensional saliency feature A of the test image, and the saliency map of the test image is calculated as follows
[0062]
[0063] The present invention proposes a method and system for extracting salient regions of community images, which represent the appearance features and object label semantic features of images as one feature for training, reducing the training workload and being more simple and feasible. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 FIG. is a flowchart of a preferred embodiment of the method for extracting salient regions of community images according to the present invention.
[0065] Figure 2 FIG. is a module diagram of a preferred embodiment of the system for extracting salient regions of community images according to the present invention.
[0066] Figure 3 FIG. is a flowchart of an embodiment of RCNN object detection of the method for extracting salient regions of community images according to the present invention.
[0067] Figure 4 FIG. is a schematic diagram of an original image and an object detection result of a preferred embodiment of the method for extracting salient regions of community images according to the present invention.
[0068] Figure 5 FIG. is a diagram for extracting a proposed rectangular box of a label object in an embodiment of the method for extracting salient regions of community images according to the present invention.
[0069] Figure 6 FIG. is a schematic diagram of a salient region detection result of an embodiment of salient region detection of the method for extracting salient regions of community images according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0071] Embodiment 1
[0072] As Figure 1 、 2 shown, step 100 is executed, and the acquisition module 200 prepares a training image set and prepares a set of standard binary annotations and / or labels of the salient regions corresponding to the training images.
[0073] Execute step 110. The calculation module 210 performs appearance-based saliency calculation on the images in the training image set, extracts the appearance features of the images, and calculates the corresponding saliency features to obtain K-dimensional appearance saliency features. The saliency calculation method on the k-th dimensional saliency feature channel is as follows:
[0074]
[0075] where D(v i k ,v j k ) represents the difference between pixel x i and pixel x j on the k-th dimensional feature channel, w ij represents the spatial distance weight, 1 ≤ k ≤ K, i represents the i-th pixel, and j is the j-th pixel different from the i-th pixel.
[0076] Execute step 120. The calculation module 210 calculates the semantic features of the corresponding object labels for the object labels carried by each image in the training image set. In this step, execute step 121. Define the object label semantic feature p x of pixel x in the image to be initialized as {0, 0, ……, 0}, M is the number of object labels included in the training image set,, q represents the q-th object label.
[0077] Execute step 122. Check one by one whether the image contains M object labels.
[0078] 1) When the m-th object label exists, 1 ≤ m ≤ M, then detect the N proposed rectangular boxes corresponding to the object label m. The possibility of each rectangular box containing the object is f m n , then the semantic feature of object m for each pixel covered by the rectangular box is f m n , 1 ≤ n ≤ N. If the pixel is not covered by any rectangular box of the m-th object label, the semantic feature of the m-th object label of the pixel is 0;
[0079] 2) When the m-th object label of the image does not exist, then f m n = 0, 1 ≤ n ≤ N.
[0080] Execute step 123. Calculate the object label semantic feature p x of pixel x. The calculation formula of the object label semantic feature p x is as follows:
[0081]
[0082] Among them, is a marker variable used to indicate whether the nth proposed rectangle is detected. If it is detected, then otherwise
[0083] Execute step 130. The solving module 220 solves the fusion coefficient of the saliency features. Based on the calculations in step 110 and step 120, a total of K + M-dimensional saliency features are obtained, including K-dimensional appearance-based saliency features and M-dimensional label semantic features. The K + M-dimensional saliency features are fused, and the fusion coefficient is solved by the following formula
[0084]
[0085] When a x = 0 indicates that pixel x is labeled as non-salient, F t (a x ) = S′(x);
[0086] When a x = 1 indicates that pixel x is labeled as salient, F t (a x ) = 1 - S′(x), where S′(x) represents the saliency value of pixel x;
[0087] Among them, S′(x) represents the saliency value of pixel x, G(a x , a x′ ) = |a x - a x′ |.d(x, x′), where d(x, x’) represents the color difference between pixels x and x′; i represents the ith image in the training set, y represents the yth image in the training set, 1 ≤ y ≤ Q, and Q represents the number of images in the training set.
[0088] The optimization parameter is
[0089] Execute step 140. The testing module 230 extracts the saliency features of the test image and calculates the saliency map.
[0090] Perform saliency prediction on the test image. Concatenate the K-dimensional appearance feature-based saliency feature s′ i (k) and the M-dimensional object label semantic feature p′ x to obtain the K + M-dimensional saliency feature A of the test image. The saliency map of the test image is calculated as follows
[0091]
[0092] Embodiment 2
[0093] The present invention proposes a method for extracting salient objects by fusing high-level semantic tags and low-level appearance features. In the training stage, the appearance features of the training images are extracted, and the saliency features based on the image appearance are calculated; object detection is performed through the object detection sub corresponding to the object labels carried by the images, and the object label semantic features are calculated; the object label semantic features can be regarded as a kind of prior feature, which is fused with the appearance-based saliency features to model the fusion problem of the saliency features and solve the fusion parameters. In the testing stage, the appearance features of the test images are extracted, and the saliency features based on the image appearance are calculated; object detection is performed through the object detection sub corresponding to the object labels carried by the images, and the object label semantic features are calculated; finally, the saliency features of the image appearance and the label semantic features are fused through the trained fusion parameters to obtain the final saliency map. Since the label semantic information is high-level semantic information, the present invention can better improve the traditional salient object detection method.
[0094] The present invention incorporates label semantic information into the task of extracting salient regions, and models the image appearance saliency features and label semantic information at the same time. The specific process is as follows.
[0095] Step 1. Preparation of the training set
[0096] Prepare the training image set, its corresponding standard binary annotation of the salient regions, and the set of labels.
[0097] Step 2. Calculation of appearance-based saliency features
[0098] Extract the appearance features of the images and calculate the corresponding saliency features to obtain K-dimensional appearance saliency features. The saliency calculation method on the k-th dimensional saliency feature channel is as follows:
[0099]
[0100] where D(v i k ,v j k ) represents the difference between pixel x i and pixel x j on the k-th dimensional feature channel, w ij represents the spatial distance weight, 1 ≤ k ≤ K. Finally, a total of K-dimensional saliency features based on appearance features are obtained.
[0101] Step 3. Calculation of semantic features of object labels
[0102] The training set contains a total of M object labels. Perform corresponding object detection on the object labels carried by each image. The calculation method of the object label semantic features of each pixel in the image is as follows: First, the object label semantic feature p x of pixel x in the image is initialized to {0, 0,..., 0}, Check one by one whether the image contains M object labels. If the m-th label exists, where 1 ≤ m ≤ M, then detect the N proposed rectangular bounding boxes corresponding to label m. The probability that each bounding box contains an object is f m n , then the semantic feature of object m for each pixel covered by the bounding box is f m n , where 1 ≤ n ≤ N. If a pixel is not covered by any bounding box of object m, then the semantic feature of the m-th object label for the pixel is 0; if the m-th label of the image does not exist, then f m n = 0, where 1 ≤ n ≤ N. After detecting the M object labels,
[0103] the semantic feature of the object label for pixel x is
[0104]
[0105] If pixel x is covered by the n-th bounding box of the m-th object, then Otherwise
[0106] Step 4. Solving the fusion parameters of the saliency features
[0107] Based on the calculations in Step 2 and Step 3, a total of K + M-dimensional saliency features are obtained, including K-dimensional appearance-based saliency features and M-dimensional label semantic features. This step requires fusing the K + M-dimensional saliency features. The fusion coefficients are solved by minimizing the following model.
[0108]
[0109] a x = 0 indicates that pixel x is labeled as non-salient, F n (a x ) = S′(x), where S′(x) represents the saliency value of pixel x; a x = 1 indicates that pixel x is labeled as salient, F n (a x ) = 1 - S′(x), where S′(x) represents the saliency value of pixel x. G(a x , a x′ ) = |a x - a x′ |.d(x, x′), where d(x, x’) represents the color difference between pixels x and x′; i represents the i-th image in the training set, y represents the y-th image in the training set, where 1 ≤ y ≤ Q and Q represents the number of images in the training set.
[0110] The optimized parameters are
[0111] Step 5. Salience prediction
[0112] When predicting the salience of the test image, the K-dimensional salience feature s′ i (k) based on appearance features is calculated according to the method in Step 2, and the M-dimensional object label semantic feature p′ x is calculated according to the method in Step 3. The K-dimensional salience feature s′ i (k) based on appearance features and the M-dimensional object label semantic feature p′ x are concatenated to obtain the K + M-dimensional salience feature A of the test image. The optimal parameters solved in Step 4 are used to fuse the salience feature based on appearance features and the object label semantic feature, and the salience map A’ of the test image is obtained. The calculation formula is as follows:
[0113]
[0114] Example 3
[0115] As shown in Table 1, the appearance features of the image are extracted. The appearance features include color and texture features. The color feature spaces used are RGB, HSV, and L*a*b*, and the texture features used are LBP features and the response features of the LM filter bank. The salience features are calculated on each feature channel, and the appearance-based salience features with a total of 29 dimensions are obtained.
[0116]
[0117] Table 1 Example table of the basic situation of appearance features and salience features
[0118] Example 4
[0119] Twenty object labels are selected from the training set, including bear, birds, boats, buildings, cars, cat, computer, coral, cow, dog, elk, fish, flowers, fox, horses, person, plane, tiger, train, zebra; then twenty RCNN object detection subnets corresponding to the object labels are selected for object detection, and the first 200 rectangular boxes with the highest object probabilities are selected.
[0120] Example 5
[0121] The RCNN method is from the paper: Ross Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik. RCNN - Rich feature hierarchies for accurate object detection and semantic segmentation. 2014 cvpr.
[0122] The RCNN method is a popular object detection method, and the detection process is as follows Figure 3 shown.
[0123] Example Six
[0124] As Figure 4 shown, (1) is the original image; (2) is the standard binary annotation of the significant region in the image; 20 object labels are selected from the labels of the training image set, including bear, birds, boats, buildings, cars, cat, computer, coral, cow, dog, elk, fish, flowers, fox, horses, person, plane, tiger, train, zebra. The label of the original image is cat, so the class identifier of the object in the image is 6. (3) is the rectangular box annotation of the object in the image, and the position information of the rectangular box is (100, 0, 230, 400).
[0125] Example Seven
[0126] This example illustrates the calculation process of label semantic features
[0127] As Figure 5 shown, the image has 3 detected rectangular boxes, with probabilities p 1 , p 2 , p 3 . Since the red dot pixel is surrounded by three rectangular boxes, its semantic feature is p 1 + p 2 + p 3 ; the diamond pixel is surrounded by two rectangular boxes, and its semantic feature is p 1 + p 2 ; the triangle pixel is surrounded by one rectangular box, and its semantic feature is p 3 .
[0128] Example Eight
[0129] This example illustrates an example of the significant region detection result. As Figure 6 shown, (a) is the original image, and (b) is the detection result of the significant region.
[0130] For a better understanding of the present invention, the above has been described in detail in conjunction with specific embodiments of the present invention, but it is not a limitation of the present invention. Any simple modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention. Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system embodiments, since they basically correspond to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
Claims
1. A method for extracting significant regions of community images, including preparing a training image set, characterized in that, it further includes the following steps: Step 1: Perform appearance-based saliency calculation on the images in the training image set, extract the appearance features of the images and calculate the corresponding saliency features to obtain K-dimensional appearance saliency features. The saliency calculation method on the k-th dimensional saliency feature channel is: Among them, D(v i k , v j k ) represents the difference between pixel x i and pixel x j on the k-th dimensional feature channel, w ij represents the spatial distance weight, 1 ≤ k ≤ K, i represents the i-th pixel, and j is the j-th pixel different from the i-th pixel; Step 2: Calculate the semantic features of the corresponding object labels for each image in the training image set, including the following sub-steps: Step 21: Define the object label semantic feature p of pixel x in the image x Initialize it to {0, 0, ……, 0}, M is the number of object labels included in the training image set, and q represents the q-th object label; Step 22: Check one by one whether the image contains M object labels, including: 1) When the m-th object label exists, where 1 ≤ m ≤ M, then detect the N proposed rectangular bounding boxes corresponding to the object label m, and the probability that each bounding box contains the object is f m n , then the semantic feature of object m for each pixel covered by the bounding box is f m n , where 1 ≤ n ≤ N. If a pixel is not covered by any bounding box of the m-th object label, then the semantic feature of the m-th object label for the pixel is 0; 2) When the m-th object label of the image does not exist, then f m n = 0, 1 ≤ n ≤ N; Step 23: Calculate the object label semantic feature p of pixel x x , and the calculation formula is: Among them, is a marker variable; Step 3: Solve the fusion coefficient of the saliency features. Based on the calculations in Step 1 and Step 2, a total of K+M dimensional saliency features are obtained, including K-dimensional appearance-based saliency features and M-dimensional label semantic features. Fuse the K+M dimensional saliency features. The fusion coefficient is solved by the following formula: When a x = 0 indicates that the pixel x is labeled as non-significant, F t (a x ) = S′(x); When a x = 1 indicates that pixel x is labeled as significant, F t (a x ) = 1 - S′(x), where S′(x) represents the significance value of pixel x; Among them, S’(x) represents the saliency value of pixel x, G(a x , a x′ ) = |a x - a x′ |.d(x, x’), d(x, x’) represents the color difference between pixels x and x’; i represents the i-th image in the training set, y represents the y-th image in the training set, 1 ≤ y ≤ Q, and Q represents the number of images in the training set; The optimized parameters are Step 4: Extract the saliency features of the test image and calculate the saliency map.
2. A system for extracting significant regions of community images, including an acquisition module for preparing a training image set, characterized in that, it further includes the following modules: Calculation module: Used to perform appearance-based saliency calculation on the images in the training image set, extract the appearance features of the images and calculate the corresponding saliency features to obtain K-dimensional appearance saliency features. The saliency calculation method on the k-th dimensional saliency feature channel is: Among them, D(v i k , v j k ) represents the difference between pixel x i and pixel x j on the k-th dimensional feature channel, w ij represents the spatial distance weight, 1 ≤ k ≤ K, i represents the i-th pixel, and j is the j-th pixel different from the i-th pixel; It is also used to calculate the semantic features of the corresponding object labels for each image in the training image set. The semantic feature calculation method of the object labels includes the following sub-steps: Step 21: Define the object label semantic feature p of pixel x in the image x Initialize it to {0, 0, ……, 0}, M is the number of object labels included in the training image set, and q represents the q-th object label; Step 22: Check one by one whether the image contains M object labels, including: 1) When the m-th object label exists, where 1 ≤ m ≤ M, then detect the N proposed rectangular bounding boxes corresponding to the object label m, and the probability that each bounding box contains the object is f m n , then the semantic feature of object m for each pixel covered by the bounding box is f m n , where 1 ≤ n ≤ N, if a pixel is not covered by any bounding box of the m-th object label, then the semantic feature of the m-th object label for the pixel is 0; 2) When the m-th object label of the image does not exist, then f m n = 0, 1 ≤ n ≤ N; Step 23: Calculate the object label semantic feature p of pixel x x , and the calculation formula is: Among them, is a marker variable; Solution module: Used to solve the fusion coefficient of the saliency features. Based on the calculation of the calculation module, a total of K+M dimensional saliency features are obtained, including K-dimensional appearance-based saliency features and M-dimensional label semantic features. Fuse the K+M dimensional saliency features. The fusion coefficient is solved by the following formula: When a x = 0 indicates that pixel x is labeled as non-significant, and F t (a x ) = S′(x); When a x = 1 indicates that pixel x is labeled as significant, F t (a x ) = 1 - S′(x), where S′(x) represents the significance value of pixel x; Among them, S’(x) represents the saliency value of pixel x, G(a x , a x′ ) = |a x - a x′ |.d(x, x’), d(x, x’) represents the color difference between pixels x and x’; i represents the i-th image in the training set, y represents the y-th image in the training set, 1 ≤ y ≤ Q, and Q represents the number of images in the training set; The optimized parameters are Test module: Used to extract the saliency features of the test image and calculate the saliency map.
Citation Information
Patent Citations
Significant tag sorting-based image significant target detection method
CN106127197A
Salient object extraction method based on label semantic meaning
CN107967480A