Park image focus object extraction method based on visual saliency prediction
By constructing a visual significance prediction model and a SAN semantic segmentation model, combining significance and semantic segmentation technology, the problem of extracting focal objects in park images is solved, efficient and accurate extraction of visual focal objects and optimizing the configuration of park features.
Patent Information
- Application Number
- CN202510232374.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology is difficult to effectively combine the fields of visual saliency and semantic segmentation, and it is impossible to accurately extract the focal objects in park images. The traditional eye tracking technology is cumbersome and difficult to apply to complex landscape environments on a large scale.
A visual significance prediction model is constructed, combined with the SAN semantic segmentation model, the high significance areas in the park image are obtained through visual significance prediction, and the sum of the significance values of each object category is calculated to extract the visual focus object.
Efficiently and accurately extract visual focal objects in park images, save human and material resources, assist park managers in optimizing configuration, and gain an in-depth understanding of the elements of visual attractiveness.
Smart Images

Figure CN120339812A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital landscape and intelligent data analysis, and particularly relates to a method for extracting focal objects in park images based on visual saliency prediction. Background Art
[0002] With the acceleration of the urbanization process, parks, as an important part of the urban ecosystem, play an important role in aspects such as leisure and entertainment, ecological protection, and social interaction, and are of increasing significance for improving the quality of life of urban residents and promoting sustainable development.
[0003] Visual sense is one of the main ways for humans to perceive the environment. According to research, visual perception accounts for about 76% of the overall senses. Given that landscape design and park use satisfaction largely depend on visual attractiveness, and visual attractiveness has a profound impact on people's park use intention, staying time, and emotional experience, the visual assessment of parks has become a hot topic in academic research and practical applications.
[0004] Nowadays, with the rise of social media platforms such as Dianping, a large amount of image data spontaneously generated by users provides valuable information for researchers on park landscape elements, facility layouts, and their visual effects. However, existing image analysis technologies still have obvious deficiencies, especially in the precise extraction of landscape elements and the identification of visual foci. Among them, traditional eye-tracking technology can identify the fixation points and visual attractiveness of subjects on landscape images, but this technology is cumbersome to operate and difficult to be widely applied in complex landscape environments.
[0005] In the field of computer vision, visual saliency prediction is an important research direction, aiming to automatically identify the most eye-catching regions in an image. Visual saliency prediction technology can effectively simulate the human visual mechanism by learning eye-tracking data, and can help find the focal regions in an image. In park images, these focal regions usually represent the most attractive or important landscape elements. Although certain progress has been made in the fields of visual saliency and semantic segmentation, there is no method to effectively combine the fields of visual saliency and semantic segmentation to achieve the precise extraction of focal objects in park images. There is an urgent need for innovative solutions to fill this gap. Summary of the Invention
[0006] To solve the technical problems existing in the prior art, the present invention provides a method for extracting focal objects in park images based on visual saliency prediction, which can efficiently and accurately extract visual focal objects in park images by effectively combining the technologies in the fields of visual saliency and semantic segmentation.
[0007] The object of the present invention can be achieved by adopting the following technical solutions:
[0008] A method for extracting focus objects in park images based on visual saliency prediction, the method comprising:
[0009] S1. Construct a visual saliency prediction model including a backbone network, a readout network layer, and a post-processing layer, pre-train the visual saliency prediction model on the ImageNet dataset, and train and optimize the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset;
[0010] S2. Obtain the COCO-Stuff164k dataset, and use the COCO-Stuff164k dataset to train the SAN semantic segmentation model so that the SAN semantic segmentation model can identify and output various object categories in the image;
[0011] S3. Obtain the high-saliency regions in the park image through the visual saliency prediction model, combine the object category information of the high-saliency regions in the park image output by the SAN semantic segmentation model, and calculate the sum of the saliency values of each object category;
[0012] S4. Extract the visual focus objects in the park image according to the magnitude of the sum of the saliency values of each object category.
[0013] Specifically, the step S1 includes:
[0014] Construct the backbone network of the visual saliency prediction model based on the ResNeXt50 model, pre-train the ResNeXt50 model on the ImageNet dataset, and use the pre-trained ResNeXt50 model as the backbone network of the visual saliency prediction model. The backbone network is used to extract high-dimensional feature maps from the park image;
[0015] Construct the readout network layer in the visual saliency prediction model. The readout network layer is used to map the high-dimensional feature maps extracted by the backbone network into two-dimensional saliency maps;
[0016] Construct the post-processing layer in the visual saliency prediction model. The post-processing layer is used to perform size adjustment and blurring operations on the saliency map, and normalize the processed features through the Softmax function in combination with the center bias;
[0017] Train and optimize the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset, adjust and optimize the parameters of the visual saliency prediction model, and obtain the optimized visual saliency prediction model.
[0018] Specifically, the readout network layer includes: a 1×1 convolutional layer, a normalization layer, a Softplus activation function, and a channel transformation layer. The channel transformation layer includes 5 convolutional layers, and the number of channels of the 5 convolutional layers is 16, 1, 128, 16, and 1 in sequence.
[0019] Specifically, training and optimizing the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset, and adjusting and optimizing the parameters of the visual saliency prediction model to obtain the optimized visual saliency prediction model, including:
[0020] Input the images in the SALICON dataset into the visual saliency prediction model. First, keep the weights of the backbone network fixed, and let the parameters of the readout network layer and the post-processing layer participate in the training. During the training process, use the backpropagation algorithm to calculate the loss value between the saliency map predicted by the model and the true saliency annotation of the images in the SALICON dataset. By continuously adjusting the above-mentioned parameters participating in the training, the loss value is gradually reduced;
[0021] Input the images in the MIT1003 dataset into the visual saliency prediction model. First, keep the weights of the backbone network fixed, and let the parameters of the readout network layer and the post-processing layer participate in the training. During the training process, use the backpropagation algorithm to calculate the loss value between the saliency map predicted by the model and the true saliency annotation of the images in the MIT1003 dataset. By continuously adjusting the above-mentioned parameters participating in the training, the loss value is gradually reduced, and finally the optimized visual saliency prediction model is obtained.
[0022] Specifically, the formula for the saliency map is as follows:
[0023]
[0024] Among them, P(x, y) is the saliency prediction of each pixel, S(x, y) is the saliency map, and ∑ x′,y′ P(x ′ , y ′ ) represents the sum of the saliency predictions of all pixels.
[0025] Specifically, the step S2 includes:
[0026] Obtain the COCO-Stuff164k dataset, preprocess the images in the COCO-Stuff164k dataset to obtain the training set of the SAN semantic segmentation model;
[0027] Use the training set to train the SAN semantic segmentation model, and use the standard metric method of mIoU semantic segmentation to evaluate the overlap degree between the semantic segmentation prediction result of the SAN semantic segmentation model and the true annotation.
[0028] Specifically, step S3 includes:
[0029] Obtain the visual saliency prediction values of all pixel points in the park image through the visual saliency prediction model, take the values of the first several proportions of the visual saliency prediction values of all pixel points as the high-saliency region threshold, and obtain the high-saliency region in the image according to the high-saliency region threshold;
[0030] Output the object category information of each object in the high-saliency region of the image through the SAN semantic segmentation model, and calculate the sum of the saliency values of each object category in the high-saliency region.
[0031] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0032] The present invention provides a method for extracting focus objects in park images based on visual saliency prediction. By constructing and training a visual saliency prediction model and using the COCO-Stuff164k dataset to train the SAN semantic segmentation model, the technologies in the fields of visual saliency and semantic segmentation are effectively combined. The high-saliency region in the park image is obtained through the visual saliency prediction model, the object category information of the high-saliency region in the park image is output in combination with the SAN semantic segmentation model, and the sum of the saliency values of each object category is calculated. According to the magnitude of the sum of the saliency values of each object category, the visual focus objects in the park image and the sum of their saliency values can be efficiently and accurately extracted, which has strong practical application value, can save a large amount of human and material resources, and assist park managers in deeply understanding the elements with high visual attraction in the park, so as to optimize the configuration of park elements. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0034] Figure 1 It is a flowchart of the method for extracting focus objects in park images based on visual saliency prediction according to Embodiment 1 of the present invention;
[0035] Figure 2 It is an architecture diagram of the visual saliency prediction model according to Embodiment 1 of the present invention;
[0036] Figure 3 It is an architecture diagram of the method for extracting focus objects in park images by combining the visual saliency prediction model and the semantic segmentation model according to Embodiment 1 of the present invention;
[0037] Figure 4 The original of the social media picture in a certain park in Embodiment 1 of the present invention Figure 1 , the visual saliency prediction map, and the semantic segmentation map;
[0038] Figure 5 The original of the social media picture in a certain park in Embodiment 1 of the present invention Figure 2 , the visual saliency prediction map, and the semantic segmentation map;
[0039] Figure 6 The original of the social media picture in a certain park in Embodiment 1 of the present invention Figure 3 , the visual saliency prediction map, and the semantic segmentation map. Detailed implementation manners
[0040] Next, the technical solution of the present invention will be further described in detail in conjunction with the drawings and embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The implementation manners of the present invention are not limited thereto. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0041] Embodiment 1:
[0042] As Figure 1 shown, it is a flowchart of a method for extracting the focus object of a park image based on visual saliency prediction. This embodiment provides a method for extracting the focus object of a park image based on visual saliency prediction, and the method includes:
[0043] S1. Construct a visual saliency prediction model including a backbone network, a readout network layer, and a post-processing layer. Pre-train the visual saliency prediction model on the ImageNet dataset, and train and optimize the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset.
[0044] As Figure 2 shown, it is an architecture diagram of the visual saliency prediction model. The visual saliency prediction model includes a backbone network, a readout network layer, and a post-processing layer. The backbone network is used to learn general image features and extract a high-dimensional feature map from the park image. The readout network layer is used to map the high-dimensional feature space extracted by the backbone network into a two-dimensional saliency map. The readout network layer includes a 1×1 convolutional layer, a normalization layer, a Softplus activation function, and a channel transformation layer. The channel transformation layer includes multiple convolutional layers, and the number of channels of each layer is 16, 1, 128, 16, and 1 in sequence.
[0045] S11. Construct the backbone network of the visual saliency prediction model based on the ResNeXt50 model. Pre-train the ResNeXt50 model on the ImageNet dataset. The pre-trained ResNeXt50 model serves as the backbone network of the visual saliency prediction model. The backbone network is used to learn general image features and extract high-dimensional feature maps from park images.
[0046] The ImageNet dataset is a large-scale visual dataset widely used in the research and model training of the computer vision field. It contains more than 14 million annotated images, covering approximately 20,000 categories, from common daily objects to various animals, plants, natural landscapes, artificial objects, etc., providing rich visual data to support various tasks such as image classification, object detection, and image segmentation. The ImageNet dataset provides an important benchmark for the development and evaluation of deep learning models and has played a key role in the development of computer vision technology, especially in the field of image recognition and understanding.
[0047] The ResNeXt50 model is a deep convolutional neural network model based on the residual network architecture. The network structure of the ResNeXt50 model mainly includes: an initial convolutional layer for preliminary feature extraction of the input image. Multiple residual blocks composed of multiple stacked residual blocks, each residual block containing multiple convolutional layers for learning residual mappings. A pooling layer and a fully connected layer for final feature aggregation and classification. ResNeXt50 has efficient computing and powerful feature extraction capabilities, which enable ResNeXt50 to have excellent performance in image classification and other computer vision tasks. Pre-training the ResNeXt50 model on the ImageNet dataset allows the ResNeXt50 model to learn general image features. The pre-trained ResNeXt50 model is the backbone network of the visual saliency prediction model and can extract high-dimensional feature maps from park images.
[0048] S12. Construct the readout network layer in the visual saliency prediction model. The readout network layer is used to map the high-dimensional feature maps extracted by the backbone network into a two-dimensional saliency map. The saliency map can represent the probability distribution of the regions in the image attracting human visual attention.
[0049] Specifically, the readout network layer includes a 1×1 convolutional layer, a normalization layer, a Softplus activation function, and a channel transformation layer. The channel transformation layer includes multiple convolutional layers, and the number of channels in each layer is 16, 1, 128, 16, and 1 in sequence. The 1×1 convolutional layer is used to reduce the dimension and fuse features of the high-dimensional feature map to obtain the feature map after dimension reduction. In the backbone network, the number of channels of the feature map is relatively high (2048 channels), and directly using it for saliency prediction will lead to high computational complexity and overfitting. The 1×1 convolutional layer reduces the number of channels of the high-dimensional feature map to a lower dimension (16 channels) through convolution operations in the channel dimension, thereby reducing the amount of computation and the number of parameters while retaining important feature information.
[0050] The normalization layer is used to standardize the numerical distribution of the feature map after dimension reduction to obtain and output the normalized feature map, so as to improve the stability and convergence speed of model training.
[0051] The Softplus activation function is used to map the output of the normalization layer to non-negative values. The formula of the Softplus activation function is as follows:
[0052] Softplus(x) = ln(1 + e x )
[0053] where x is the input of the neuron, ln() represents the logarithmic function, and e is the natural constant. This formula maps the input value x to the non-negative value interval through the combination of the natural logarithm and the exponential function, and the logarithmic function ln() smooths this exponential growth, making the output value softer and more continuous to some extent, which can help the model better learn and extract the features of the image.
[0054] The Softplus activation function helps the model better learn and fit the data during the construction and training of the neural network. The Softplus function has the characteristics of being smooth and non-linear. When the input value x is small, the output of the Softplus function is close to 0, which enables the neuron to remain in an inactive state to a certain extent and plays a role similar to a threshold. When the input value x is large, the output of the Softplus function approaches x, enabling the neuron to be fully activated and thus better transmit information.
[0055] The channel transformation layer consists of multiple convolutional layers with the number of channels being 16, 1, 128, 16, and 1 in sequence. It is used to process and fuse feature maps, and finally generate a high-quality single-channel saliency map. Among them, the 16-channel layer is used to extract local features and reduce the dimension; the 1-channel layer is used to generate a preliminary saliency map, which helps the subsequent layers to further optimize the saliency region; the 128-channel layer is used to enable the model to fuse more context feature information and enhance the modeling ability for saliency patterns in complex scenes; the 16-channel layer serves as a transition layer to ensure that key information is still retained during the compression process; the saliency map is essentially a single-channel heat map, representing the saliency intensity of each pixel in the image. Therefore, the last 1-channel convolutional layer directly generates the prediction result of the model. The design of these numbers of channels is finally determined through multiple experiments.
[0056] S13. Construct the post-processing layer in the visual saliency prediction model. The post-processing layer is used to perform size adjustment and blurring operations on the saliency map, and normalize the processed features through the Softmax function in combination with the center bias.
[0057] In the visual saliency prediction model, saliency maps of different sizes may be obtained in different processing stages. The size adjustment operation ensures the consistency of data in subsequent processing. Through appropriate size adjustment methods, such as interpolation, downsampling, etc., while maintaining key information, the utilization of computing resources can be optimized, and the running efficiency of the model can be improved. The blurring operation is a method for smoothing the image. In visual saliency prediction, the blurring operation helps to reduce the influence of noise on the prediction result. By appropriately blurring the feature map, the detailed information in the image can be suppressed to a certain extent, highlighting the overall features. The center bias is a characteristic considering the attention mechanism of the human visual system. When the human visual system perceives a scene, it tends to pay more attention to the central region of the visual field. In the visual saliency prediction model, combining the center bias can make the prediction result more in line with the visual perception habits of humans. Specifically, when processing the saliency map, higher weights are assigned to the pixels in the central region to reflect their importance in visual attention distribution.
[0058] The processed features are normalized through the Softmax function to obtain the visual saliency prediction value of each pixel in the image. The Softmax normalization formula is as follows:
[0059]
[0060] where f(x, y) is each element in the processed feature map, and S norm(x, y) is the predicted value of visual saliency after normalization. The denominator is the exponential sum of all feature values, ensuring that the output result conforms to the probability distribution. The result conforming to the probability distribution means that the predicted saliency value of each pixel is between [0, 1], and the sum is 1.
[0061] The Softmax function is a normalization function that can convert a set of real-valued vectors into a set of probability distributions. In the visual saliency prediction model, after operations such as previous size adjustment, blurring, and combining center bias, the obtained feature map needs to be normalized to obtain the predicted value of visual saliency for each pixel. The Softmax function can ensure that the saliency values of all pixels are between 0 and 1, and the sum of the saliency values of all pixels is 1. The obtained normalized result can be used as the final visual saliency prediction result, intuitively representing the probability of each pixel in the image becoming a salient region. This can better mine the truly salient regions in the image and improve the prediction accuracy of the model.
[0062] S14. Train and optimize the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset, adjust and optimize the parameters of the visual saliency prediction model to obtain the optimized visual saliency prediction model.
[0063] S141. Input the images in the SALICON dataset into the visual saliency prediction model. First, keep the weights of the backbone network fixed and let the parameters of the readout network layer and the parameters of the post-processing layer (including the weights of each convolutional layer in the 1×1 convolutional layer, normalization layer, and channel transformation layer) participate in the training. During the training process, use the backpropagation algorithm to calculate the loss value between the predicted saliency map of the model and the true saliency annotation of the images in the SALICON dataset. By continuously adjusting the above-mentioned parameters participating in the training, the loss value gradually decreases, so that the model can better learn the features of the images in the SALICON dataset and their saliency distribution rules. The fine-tuning in this stage aims to make the model initially adapt to the visual saliency prediction task and extract features related to human attention.
[0064] After completing the fine-tuning on the SALICON dataset, the images in the MIT1003 dataset are input into the model. First, keep the weights of the backbone network fixed, and let the parameters of the readout network layer and the parameters of the post-processing layer (including the weights of each convolutional layer in the 1×1 convolutional layer, normalization layer, and channel transformation layer) participate in the training. During the training process, continue to use the backpropagation algorithm to calculate the loss value between the saliency map predicted by the model and the true saliency annotation of the images in the MIT1003 dataset. By continuously adjusting the above-mentioned parameters participating in the training, the loss value is gradually reduced, and finally an optimized visual saliency prediction model is obtained. Through this process, the model can better adapt to the image saliency prediction tasks in different scenarios, improving its generalization ability and prediction accuracy. The high-quality annotations and diverse image contents of the MIT1003 dataset provide the model with richer training samples, enabling it to more accurately capture the salient regions in complex scenes.
[0065] The optimized visual saliency prediction model can generate a saliency map. Generating a saliency map can help accurately identify the most informative key regions in the image, thereby improving the accuracy and efficiency of detection. In the saliency map, different grayscale values or color mappings represent different degrees of saliency. High-saliency regions will be represented by specific bright colors or higher grayscale values, thus highlighting the most important and attractive parts of the image.
[0066] The formula for the saliency map is as follows:
[0067]
[0068] Among them, P(x,y) is the saliency prediction of each pixel, S(x,y) is the saliency map, and ∑ x′,y′ P(x ′ ,y ′ ) represents the sum of the saliency predictions of all pixels. The saliency map formula is a mathematical expression used to measure the saliency of each region in the image, and it clearly identifies the regions in the entire image that can most attract people's attention through calculation.
[0069] S2. Use the COCO-Stuff164k dataset to train the SAN (Side Adapter Network) semantic segmentation model so that the SAN semantic segmentation model can identify and output each object category in the image.
[0070] The SAN (Side Adapter Network) semantic segmentation model is an innovative and efficient model for solving the open-vocabulary semantic segmentation problem. In the specific training and prediction processes, the semantic segmentation model fully utilizes the advantages of large-scale vision-language models and combines the unique design of the side adapter network to accurately perceive and efficiently analyze the semantic information in images. The COCO-Stuff164k dataset is a rich and comprehensive image dataset widely used in the field of computer vision. The COCO-Stuff164k dataset covers an extremely wide range of image categories, containing more than 164,000 high-quality images from various different scenes and backgrounds. In each scene, rich semantic information is carefully annotated, detailing various objects, regions, and their relationships in the images.
[0071] S21. Obtain the COCO-Stuff164k dataset, preprocess the images in the COCO-Stuff164k dataset to obtain the training set of the SAN semantic segmentation model.
[0072] Download the COCO-Stuff164k dataset from the specified channels to ensure the integrity and accuracy of the data. Preprocess the images in the COCO-Stuff164k dataset. The preprocessing includes: Resizing the images, resizing the COCO-Stuff164k dataset images to a size suitable for model input. Image normalization, perform normalization on the images, that is, subtract the mean and divide by the standard deviation to make the distribution of the image data more stable, which helps the model to converge. Annotated data processing, perform corresponding processing on the pixel-level annotation information to ensure that the annotated data corresponds accurately to the preprocessed image data. Use the preprocessed COCO-Stuff164k dataset as the training set of the SAN semantic segmentation model.
[0073] S22. Use the training set to train the SAN semantic segmentation model, and use the mIoU (mean Intersection over Union) standard metric method for semantic segmentation to evaluate the overlap between the semantic segmentation prediction results and the ground truth annotations of the SAN semantic segmentation model. First, determine the intersection and union of each category in the prediction results and the ground truth annotations respectively, then calculate the ratio of the two, and then average this ratio for each category to obtain the mIoU value. This value can comprehensively evaluate the segmentation ability of the model on multiple categories. The calculation formula of the mIoU value is as follows:
[0074]
[0075] where N is the total number of categories, A i is the predicted region of category i, B iis the true annotation area of class i, |A i ∩B i | is the intersection of the predicted area and the true annotation area of class i (i.e., the correctly predicted area), |A i ∪B i | is the union of the predicted area and the true annotation area of class i (i.e., the total area of the predicted area and the annotation area). The mIoU value can comprehensively and synthetically reflect the matching degree between the model prediction result and the true annotation for each class.
[0076] The formula for the semantic segmentation prediction result is as follows:
[0077]
[0078] where S seg (x, y) represents the semantic segmentation prediction result at the given image coordinate position (x, y), argmax represents finding the class index corresponding to the maximum probability, C i is the pixel of the i-th class in the image, P(C i |x, y) is the conditional probability of the i-th class in the image at the coordinate position (x, y).
[0079] Based on this COCO-Stuff164k dataset, the SAN semantic segmentation model is trained to enable the SAN semantic segmentation model to have good semantic segmentation ability to recognize any class in the image and can output the object category information of the high-salience regions in the image. In this embodiment, the backbone of the SAN model is ViT-B_16. During the image processing, the image cropping size is set to 640x640. The mIoU value of this model in relevant tests and evaluations is 41.93, which reflects the good performance of the model in processing tasks such as image segmentation.
[0080] S3. Obtain the high-salience regions in the park image through the visual salience prediction model, combine the SAN semantic segmentation model to output the object category information of the high-salience regions in the park image, and calculate the sum of the salience values of each object category. As Figure 3 shown, it is the architecture diagram of the park image focus object extraction method combining the visual salience prediction model and the semantic segmentation model.
[0081] S31. Obtain the visual salience prediction values of all pixel points in the park image through the visual salience prediction model, take the values of the first several proportions of the visual salience prediction values of all pixel points as the high-salience region threshold, and obtain the high-salience regions in the image according to the high-salience region threshold.
[0082] In this embodiment, the visual saliency prediction values of all pixel points in the park image are obtained through the visual saliency prediction model, and the 90% threshold of the visual saliency prediction value is set as the high-saliency region threshold. Specifically, it means that the visual saliency prediction values of all pixel points in the image are statistically analyzed, and a critical value corresponding to the larger top 90% of the values is taken as the threshold. That is to say, the pixel points in the image are sorted according to their saliency prediction values from large to small, and the lower limit value in the range of the saliency prediction values of the top 90% of the pixel points is selected as the standard for dividing the high-saliency region and the non-high-saliency region. The setting of this threshold is based on the quantitative analysis of the saliency degrees of different pixel points in the image. By determining a relatively high proportion (90%) of the saliency value range, it can more comprehensively and accurately cover most of the regions with high saliency in the image, so as to more effectively screen out the regions where the objects or background elements that can truly attract human visual attention are located, providing a more accurate basis for subsequent steps such as semantic segmentation and calculation of the sum of saliency values.
[0083] The formula for the high-saliency region is as follows:
[0084] R salient ={(x, y)|S(x, y)>θ}
[0085] Where θ is the threshold, S(x, y) is the saliency map, and R salient is the high-saliency region. (x, y) represents the coordinate position of each pixel point in the image, and θ is the set high-saliency region threshold. When the visual saliency prediction value of a certain pixel point is greater than this threshold, it can be considered that the pixel point belongs to the high-saliency region. In this way, the regions with high saliency in the image can be more accurately screened out.
[0086] S32. Output the information of each object category in the high-saliency region of the image through the SAN semantic segmentation model, and calculate the sum of the saliency values of each object category in the high-saliency region.
[0087] For the determined high-saliency region, calculate the sum of the saliency values of each object category in this specific region. The calculation formula for the sum of the saliency values is as follows:
[0088]
[0089] Where S(x, y) is the saliency map, and I(S seg (x, y)=C i ) is the indicator function, which takes the value of 1 when the pixel belongs to class C i and 0 otherwise. Through this formula, the object with the highest visual attention, that is, the visual focus object, is effectively extracted, and the corresponding sum of the saliency values is obtained.
[0090] S4. Extract the visual focus objects in the park image according to the sum of the saliency values of each object category.
[0091] Sort the sum of the saliency values of each object category from largest to smallest, that is, compare the sum of the saliency values of different visual focus objects. The object with the largest sum of saliency values is largely the most prominent visual focus in the image. It is also possible to set a threshold. When the sum of the saliency values exceeds a certain set threshold, it can be determined as a valid visual focus object.
[0092] After sorting the sum of the saliency values of each category from largest to smallest, it can be saved as a CSV file, which includes the names of each category and the corresponding sum of the saliency values. The data in this CSV file is stored in plain text form, and the data items are separated by commas ",". By viewing this CSV file, it is possible to quickly and clearly understand the relevant situation of each category, thus providing strong support and basis for further decision-making and analysis.
[0093] In this embodiment, taking the review pictures on Dianping of a park in Guangzhou as an example, the feasibility of the method for extracting the focus objects in the park image by fusing visual saliency prediction and semantic segmentation is demonstrated.
[0094] As Figure 4 shown, for the original Figure 1 , visual saliency prediction map, and semantic segmentation map of the social media pictures in a certain park, the object categories and the sum of the saliency values (sorted from largest to smallest) in the high-saliency region are:
[0095]
[0096] As Figure 5 shown, for the original Figure 2 , visual saliency prediction map, and semantic segmentation map of the social media pictures in a certain park, the object categories and the sum of the saliency values (sorted from largest to smallest) in the high-saliency region are:
[0097]
[0098] As Figure 6 shown, for the original Figure 3 , visual saliency prediction map, and semantic segmentation map of the social media pictures in a certain park, the object categories and the sum of the saliency values (sorted from largest to smallest) in the high-saliency region are:
[0099]
[0100] In summary, the present invention provides a method for extracting focus objects in park images based on visual saliency prediction. By constructing and training a visual saliency prediction model and using the COCO-Stuff164k dataset to train the SAN semantic segmentation model, the technologies in the fields of visual saliency and semantic segmentation are effectively combined. The high-saliency regions in park images are obtained through the visual saliency prediction model, and the object category information of the high-saliency regions in park images is output in combination with the SAN semantic segmentation model. The sum of the saliency values of each object category is calculated. According to the magnitude of the sum of the saliency values of each object category, the visual focus objects in park images and the sum of their saliency values can be efficiently and accurately extracted, which has strong practical application value, can save a large amount of human and material resources, and assist park managers in deeply understanding the elements with high visual attraction in the park, so as to optimize the configuration of park elements.
[0101] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for extracting focus objects in park images based on visual saliency prediction, characterized in that, The method includes: S1. Construct a visual saliency prediction model including a backbone network, a readout network layer, and a post-processing layer. Pre-train the visual saliency prediction model on the ImageNet dataset, and train and optimize the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset; S2. Obtain the COCO-Stuff164k dataset, and use the COCO-Stuff164k dataset to train the SAN semantic segmentation model so that the SAN semantic segmentation model can identify and output the object categories in the image; S3. Obtain the high-saliency regions in the park image through the visual saliency prediction model, combine the object category information of the high-saliency regions in the park image output by the SAN semantic segmentation model, and calculate the sum of the saliency values of each object category; S4. Extract the visual focus objects in the park image according to the magnitudes of the sum of the saliency values of each object category.
2. The method for extracting the focus object of a park image based on visual saliency prediction according to claim 1, wherein The step S1 includes: Construct the backbone network of the visual saliency prediction model based on the ResNeXt50 model, pre-train the ResNeXt50 model on the ImageNet dataset, and use the pre-trained ResNeXt50 model as the backbone network of the visual saliency prediction model. The backbone network is used to extract high-dimensional feature maps from the park image; Construct the readout network layer in the visual saliency prediction model. The readout network layer is used to map the high-dimensional feature maps extracted by the backbone network into a two-dimensional saliency map; Construct the post-processing layer in the visual saliency prediction model. The post-processing layer is used to perform size adjustment and blurring operations on the saliency map, and normalize the processed features through the Softmax function in combination with the center bias; Train and optimize the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset, adjust and optimize the parameters of the visual saliency prediction model, and obtain the optimized visual saliency prediction model.
3. The method for extracting the focus object of the park image based on visual saliency prediction according to claim 2, wherein, The readout network layer includes: a 1×1 convolutional layer, a normalization layer, a Softplus activation function, and a channel transformation layer. The channel transformation layer includes 5 convolutional layers, and the number of channels of the 5 convolutional layers is 16, 1, 128, 16, and 1 in sequence.
4. A method for extracting the focus object of a park image based on visual saliency prediction according to claim 2, characterized in that The training and optimization of the visual saliency prediction model on the SALICON dataset and the MIT1003 dataset, adjusting and optimizing the parameters of the visual saliency prediction model, and obtaining the optimized visual saliency prediction model includes: Input the images in the SALICON dataset into the visual saliency prediction model. First, keep the weights of the backbone network fixed, and let the parameters of the readout network layer and the post-processing layer participate in the training. During the training process, use the backpropagation algorithm to calculate the loss value between the saliency map predicted by the model and the true saliency annotation of the images in the SALICON dataset. By continuously adjusting the above-mentioned parameters participating in the training, the loss value is gradually reduced; Input the images in the MIT1003 dataset into the visual saliency prediction model. First, keep the weights of the backbone network fixed, and let the parameters of the readout network layer and the post-processing layer participate in the training. During the training process, use the backpropagation algorithm to calculate the loss value between the saliency map predicted by the model and the true saliency annotation of the images in the MIT1003 dataset. By continuously adjusting the above-mentioned parameters participating in the training, the loss value is gradually reduced, and finally an optimized visual saliency prediction model is obtained.
5. A method for extracting the focus object of a park image based on visual saliency prediction according to claim 4, characterized in that, The formula for the saliency map is as follows: Among them, P(x, y) is the saliency prediction of each pixel, S(x, y) is the saliency map, and ∑ x′,y′ P(x ′ ,y ′ ) represents the sum of saliency predictions of all pixels.
6. The method for extracting the focal object of a park image based on visual saliency prediction according to claim 2, wherein The step S2 includes: Obtain the COCO-Stuff164k dataset, preprocess the images in the COCO-Stuff164k dataset to obtain the training set of the SAN semantic segmentation model; Use the training set to train the SAN semantic segmentation model, and use the standard metric method of mIoU semantic segmentation to evaluate the overlap degree between the semantic segmentation prediction result of the SAN semantic segmentation model and the true annotation.
7. A method for extracting the focus object of a park image based on visual saliency prediction according to claim 5, characterized in that, The step S3 includes: Obtain the visual saliency prediction values of all pixel points in the park image through the visual saliency prediction model, take the values of the first several proportions of the visual saliency prediction values of all pixel points as the high-saliency region threshold, and obtain the high-saliency region in the image according to the high-saliency region threshold; Output the object category information of each high-saliency region in the image through the SAN semantic segmentation model, and calculate the sum of the saliency values of each object category in the high-saliency region.
8. A method for extracting the focal object of a park image based on visual saliency prediction according to claim 7, characterized in that, The calculation formula for the sum of the saliency values is as follows: Among them, S(x, y) is the saliency map, and I(S seg (x, y) = C i ) is the indicator function, which takes the value of 1 when the pixel belongs to class C i and 0 otherwise.