Image saliency target detection method and system

By introducing a multi-scale feature aggregation module and an SE channel attention module in the U-shaped network, the problem of difficulty in using deep and shallow semantic information at the same time in the prior art is solved, and high accuracy detection of multi-scale and complex targets is achieved.

CN119992057APending Publication Date: 2025-05-13SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093533.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, it is difficult to extract and utilize deep semantic information and shallow semantic information simultaneously in significance object detection, resulting in low accuracy of the system in multi-scale object detection and complex texture object detection.

Method used

Using a U-shaped network-based image significance object detection method, combined with a multi-scale feature aggregation module and an SE channel attention module, the final significant map is generated through multi-layer refinement and layer-by-layer fusion.

Benefits of technology

Accurate positioning of the edges of salient objects in the image is achieved, and the deep and shallow semantic information is fully utilized, which improves the accuracy of the system in multi-scale and complex object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992057A_ABST
    Figure CN119992057A_ABST
Patent Text Reader

Abstract

The invention provides an image saliency target detection method and system. The method comprises the following steps: acquiring a to-be-detected X-ray image and preprocessing the X-ray image to obtain a feature map; inputting the preprocessed image into the trained U-shaped network; performing coarse feature extraction on the salient region in the image in the U-shaped network; the extracted coarse features are refined layer by layer from a high level to a low level, and a multi-layer refined feature map with richer and richer details is obtained; and carrying out layer-by-layer fusion on the refined feature map to generate a final saliency map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image salient object detection, and in particular relates to an image salient object detection method and system. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Salient object detection is a key task in the field of computer vision. It aims to automatically extract the target area that can best attract human visual attention from an image or video. It plays an important role in many application scenarios. For example, in image retrieval, by detecting salient objects, the content of interest to users can be matched more accurately. In the field of intelligent security, it can quickly locate suspicious targets in the monitoring screen, such as people or abnormal objects that break into restricted areas. In the scenario of autonomous driving, helping vehicles identify salient objects such as pedestrians and vehicles on the road is crucial for driving safety. Early salient object detection methods were mainly based on manual features, such as color contrast and texture differences. These methods can achieve certain results in simple scenes, but in complex natural scenes, their performance is greatly limited due to factors such as the diversity of objects, changes in lighting, and background interference. For example, when the target is similar in color to the background or the internal texture of the target is complex, methods based on manual features often find it difficult to accurately extract the salient area of ​​the target. Traditional machine learning methods, such as support vector machines (SVMs), require complex feature engineering to be manually designed when used for salient object detection, and the generalization ability of the model is limited, making it difficult to adapt to various types of image data.

[0004] With the rapid development of deep learning technology, convolutional neural networks (CNN) have achieved great success in the field of computer vision. CNN can automatically learn the feature representation of images, overcoming the limitations of manual features. Many CNN-based salient object detection methods have been proposed, which have greatly improved the accuracy and robustness.

[0005] Compared with other traditional methods, salient object recognition based on convolutional neural networks has more powerful functions in extracting deep semantic information and shallow detail information, and can extract information at more scales. At the same time, the model's calculation speed and parameter tolerance are greatly improved. The general structure of a convolutional neural network is conical. Initially, a larger shallow image is input, and after layers of convolution processing, a smaller deep image is finally output. Larger shallow images have more detailed and complex detail information; smaller deep images have more semantic and location information. More and more models based on making full use of shallow and deep information have been proposed. Among these models, the U-net network is sought after by many people. Because in this network, a richer feature map can be established more conveniently; at the same time, this model can more comprehensively utilize different types of shallow and deep information.

[0006] The attention mechanism is inspired by the human visual system. When humans observe an image, they naturally focus on certain key areas and ignore other parts. In computer vision, the attention mechanism allows the model to selectively focus on important information in the image. For salient object detection, the attention mechanism can help the model better focus on the target area and suppress background noise. It can adaptively learn the importance weights of different positions and feature channels, thereby improving the accuracy of detection.

[0007] Objects in images have different scales. In natural scenes, there may be both large-scale objects (such as buildings) and small-scale objects (such as pedestrians in the distance). Features of a single scale are difficult to effectively represent objects of different sizes. Multi-scale feature fusion technology can integrate features extracted from different levels of convolutional layers, which have different receptive fields and can capture features from fine-grained local features to coarse-grained global features. By fusing these multi-scale features, the objects in the image can be described more comprehensively, improving the performance of salient object detection on objects of different scales.

[0008] The inventors found that in the detection of salient targets, the sizes of targets in the image vary. Small-scale targets such as traffic signs and birds in the distance may account for a small proportion of the image, and their feature information is relatively small, so they are easily missed by the model. Although large-scale targets are rich in feature information, they may not be fully and accurately detected due to factors such as the receptive field. Traditional single-scale methods are difficult to effectively detect targets of different sizes at the same time, and a detection mechanism that can adapt to multi-scale targets is required. Some targets have complex textures and structures inside. Taking animals as an example, the texture of the animal's fur, the structure of different parts of the body, etc. will increase the complexity of the internal features of the target. When the internal features of the target are very different, the model may incorrectly segment different parts of the target, and cannot extract the entire target as a salient area well, but instead focuses on the local details inside the target, ignoring the overall saliency of the target. Salient target detection methods based solely on appearance features may lack semantic understanding of image content. For example, in a scene containing a person and related tools, if the model cannot understand the semantic relationship between the person and the tool, it may mistakenly divide attention to each part instead of detecting the person and its related semantic whole (such as a person holding a tool) as a salient unit. This requires the model to be able to combine semantic information to better detect salient objects.

[0009] After searching, it is found that there are technical solutions in the existing patents and paper documents about using attention mechanism and feature fusion and other related technologies to realize image saliency detection. The technical problems of the above solutions are: they do not extract and utilize deep semantic information and shallow semantic information, and cannot achieve better fusion of multiple feature maps containing different semantic information at the same time, and the accuracy of the system's salient target recognition is not high. Summary of the invention

[0010] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a method for detecting salient objects in an image, which can comprehensively improve the accuracy of salient object detection.

[0011] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0012] In a first aspect, a method for detecting salient objects in an image is disclosed, comprising:

[0013] Acquire the X-ray image to be detected and perform preprocessing to obtain a feature map;

[0014] Input the preprocessed image into the trained U-shaped network;

[0015] In the U-shaped network, coarse features are extracted from the salient areas in the image;

[0016] The extracted coarse features are refined layer by layer from high level to low level to obtain a multi-layer refined feature map with increasingly rich details;

[0017] The refined feature maps are fused layer by layer to generate the final saliency map.

[0018] As a further technical solution, the U-shaped network includes: a multi-scale feature aggregation module; the multi-scale feature aggregation module includes five branches;

[0019] The feature maps are input into five branches respectively. The first branch performs down-sampling on the input feature map through multiple convolutions continuously, and then restores it to the original size through up-sampling.

[0020] The second branch does not process the input feature map;

[0021] The third branch downsamples the input feature map by 2 times;

[0022] The fourth branch downsamples the input feature map by 4 times;

[0023] The fifth branch downsamples the input feature map by 8 times;

[0024] The results of the five branches are added together and then output after convolution processing.

[0025] As a further technical solution, the U-shaped network further includes: a pyramid pooling module;

[0026] The pyramid pooling module includes an average pooling layer, an identity mapping layer, a global average pooling layer and two adaptive average pooling layers;

[0027] The feature map is processed by the average pooling layer for the first time and the output result is output;

[0028] The above output results are processed by the identity mapping layer, the global average pooling layer and the two adaptive average pooling layers respectively;

[0029] The four results are upsampled separately to make them uniform in scale;

[0030] The up-sampled results are concatenated and then output.

[0031] Among them, the identity mapping layer does not perform any operation on the input;

[0032] Global average pooling averages all elements on each channel of each input feature map;

[0033] The adaptive average pooling layer pools the input feature map to any specified output size.

[0034] As a further technical solution, the U-shaped network further includes: an SE channel attention module;

[0035] The SE channel attention module includes a global average pooling layer;

[0036] The global average pooling is used to perform a Squeeze operation on the input feature map to compress the spatial dimension and generate a global feature descriptor;

[0037] The obtained global feature descriptor is subjected to dimension reduction to obtain a feature vector after dimension reduction;

[0038] Apply activation function to the reduced feature vector to learn more complex feature representation;

[0039] The output after activation function processing is upgraded and weights are generated for each channel;

[0040] Normalize the dimensionally upgraded feature vector to generate the weight of each channel;

[0041] The generated channel weight vector is expanded to the same spatial dimension as the input feature map, and then multiplied with the input feature map to obtain the output of the SE channel attention module.

[0042] In a second aspect, a system for detecting salient objects in an image is disclosed, comprising:

[0043] Data acquisition module: configured to: acquire the X-ray image to be detected and perform preprocessing to obtain a feature map;

[0044] Data processing module: configured to: input the preprocessed image into the trained U-shaped network;

[0045] In the U-shaped network, coarse features are extracted from the salient areas in the image;

[0046] The extracted coarse features are refined layer by layer from high level to low level to obtain a multi-layer refined feature map with increasingly rich details;

[0047] The refined feature maps are fused layer by layer to generate the final saliency map.

[0048] One or more of the above technical solutions have the following beneficial effects:

[0049] The technical solution of the present invention can achieve accurate positioning of the edges of salient targets in an image without an additional edge detection module; at the same time, it can fully explore and utilize the deep and shallow semantic information in the process, and fuse these images of different scales and types, so as to achieve the accuracy of the system's salient target detection as a whole.

[0050] This technology proposes a multi-scale feature aggregation module and introduces the SE module into the overall model, which improves the system's ability to extract and utilize deep semantic information and shallow semantic information, while allowing multiple feature maps containing different semantic information to be better integrated to improve the accuracy of the system's salient target recognition.

[0051] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0053] Figure 1 Schematic diagram of the structure of the pyramid pooling module in Embodiment 1 of the present disclosure;

[0054] Figure 2 is a schematic diagram of the structure of the channel attention module in Embodiment 1 of the present disclosure;

[0055] Figure 3 is a schematic diagram of a multi-scale feature aggregation module in Embodiment 1 of the present disclosure;

[0056] Figure 4 is a schematic diagram of the overall model process in Example 1 of the present disclosure;

[0057] Figure 5 It is a schematic diagram of the feature extraction process in Example 1 of the present disclosure. DETAILED DESCRIPTION

[0058] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0059] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0060] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0061] In view of the defects of salient target recognition, the sub-technical scheme of this embodiment is based on the U-Net framework and proposes a salient target recognition method. In order to improve the overall performance of salient target recognition and fully tap the potential of channel information in the pooling layer and FPN model, a channel global guidance network based on the pooling layer is proposed. On the one hand, the proportion of the pooling layer in the model task is increased to tap the importance of the pooling layer to the salient target recognition task. On the other hand, the channel attention mechanism is introduced to explore the potential of the channel attention mechanism in the salient target recognition task. Through this method, on the one hand, the positioning accuracy of image edge detection can be improved, and on the other hand, the utilization efficiency of semantic information can be improved. The two can comprehensively improve the accuracy of salient target detection.

[0062] Among them, the FPN model is an existing model, and the technical solution of this embodiment is an improved solution based on the FPN model.

[0063] The idea of ​​the channel global guidance module model based on the pooling layer comes from:

[0064] In order to improve the overall performance of salient object recognition and fully explore the potential of channel information in the pooling layer and FPN model, a channel global guidance network based on the pooling layer is proposed. On the one hand, the proportion of the pooling layer in the model task is increased to explore the importance of the pooling layer for the salient object recognition task, and on the other hand, the channel attention mechanism is introduced to explore the potential of the channel attention mechanism in the salient object recognition task.

[0065] (1) Pyramid pooling module:

[0066] Fully collect and process the deepest semantic information in the model.

[0067] (2) SE channel attention module:

[0068] The SE (Squeeze-and-Excitation) attention mechanism is a lightweight channel attention mechanism that aims to enhance the network's representation ability by dynamically recalibrating channel feature responses so that the network can pay more attention to important channel information and suppress relatively unimportant channel information. This mechanism can be embedded in various deep neural networks to improve performance without significantly increasing the computational burden. The core of the SE attention mechanism is to generate channel-level attention weights through two key operations: Squeeze and Excitation, and apply these weights to the original feature map. It mainly uses global average pooling and a small number of fully connected layers, and the number of parameters is relatively small. It can be easily inserted into various network architectures to improve the performance of the network, especially for tasks that are easily affected by channel feature information, such as classification, target detection, etc. By adjusting the channel weights, the network's attention to different channel features is clarified, which increases the interpretability of the network to a certain extent.

[0069] Multi-scale feature aggregation module: Since the network involves multiple branches, how to fuse this information together and make full use of it is a new problem. In order to solve this problem, a multi-scale feature aggregation module is proposed. This module can not only perfectly fuse information from different sources, but also further process this information, make fuller use of this information, and improve the final system effect.

[0070] Embodiment 1

[0071] This embodiment discloses a method for detecting salient objects in an image, including:

[0072] Acquire the X-ray image to be detected and perform preprocessing;

[0073] The preprocessed image is input into the trained U-shaped network to extract coarse features of the salient areas in the image. Figure 5 Then, the extracted coarse features are refined layer by layer from high level to low level to obtain multi-layer refined feature maps with increasingly rich details, and the refined feature maps are fused layer by layer to generate the final saliency map.

[0074] The layer-by-layer fusion of refined feature maps refers to the fusion of multiple "F" modules in the direction of the arrow, and the final saliency map output is the gray module output at the end of the arrow.

[0075] The image is input into the model after preprocessing, specifically Figure 4 The first red module in the figure, and the image processed by each module can be called a feature map. The feature map output by the previous module is directly input into the next module. The output of each module is also called a feature map.

[0076] The SE module is also called SE channel attention implementation process:

[0077] Step 1: First, use the global average pooling layer to squeeze the input feature map, compress the spatial dimension to 1x1, and set the convolution kernel of the adaptive average pooling layer to 1x1 convolution; assuming that the input feature map is X∈R H×W×C , where H represents height, W represents width, and C represents the number of channels. The calculation formula for the global average pooling operation is as follows:

[0078]

[0079] Among them, x c(i, j) represents the pixel value at position (i, j) in the cth channel. After this step, the input image X is compressed from the dimensions of HxWxC to a vector z of 1x1xC. This vector is the global descriptor, which contains the global information of each channel.

[0080] Under this operation, no matter what the H and W of the input feature map are, it will pool the input feature map to a size of 1X1 to generate a global feature descriptor.

[0081] Step 2: excitation operation; this operation mainly consists of a fully connected layer and a nonlinear transformation. First, the compressed z vector is input into the first fully connected layer (FC1). In order to reduce the amount of calculation and the number of parameters, a dimensionality reduction operation is usually performed. Let the dimensionality reduction ratio be r. If the original number of channels is C, then the number of channels output after FC1 is C / r. The calculation of the fully connected layer involves matrix multiplication and bias addition. Assuming that W1 is the weight matrix of FC1 and b1 is the bias vector, the output after FC1 can be expressed as y1=δ(W1z+b1), where δ is the ReLU activation function, which is used to introduce nonlinearity so that the network can learn more complex channel relationships. Next, the input is sent to the second fully connected layer (FC2), the purpose of which is to restore the number of channels to the original C. Let W2 be the weight matrix of FC2 and b2 be the bias vector, then the output after FC2 can be expressed as y2=(W2y1+b2). ​​Finally, the Sigmoid function is applied to y2 to obtain the channel attention weight vector s, whose element s c The value range of (c=1,2,…,c) is between (0,1). (Wherein, the weight matrix and bias vector are randomly generated at the beginning, and the best set of weight matrix and bias vector are automatically determined in the subsequent overall model training process)

[0082] Step 3: Scale operation: multiply the obtained channel attention weight vector s with the original input feature map X in the channel dimension. That is, for the cth channel X of the original feature map X c , after scaling the channel, we get

[0083] Feature map after attention recalibration This enables the network to enhance or suppress channel features according to the importance of each channel. It can continue to be used as input to subsequent network layers for further processing.

[0084] In this implementation example, the obtained global feature descriptor is reduced in dimension by reducing the number of channels from C to C / r through a fully connected layer or a 1x1 convolution, where r is the reduction rate. The purpose of this step is to reduce the amount of calculation and the number of parameters.

[0085] Applying the ReLU activation function to the reduced-dimensional feature vector enables the network to learn more complex feature representations. The output after applying the ReLU activation function is dimensionally increased, and the number of channels is restored from C / r to C through a 1x1 convolution, while generating weights for each channel.

[0086] The Sigmoid function is used to normalize the results and map them to [0,1] to generate the weight of each channel. Each element represents the importance weight of the corresponding channel.

[0087] The channel weights are applied to the original feature map. In order to apply the generated channel weights to the original feature map, the original feature map is recalibrated to highlight the information of important channels. The generated channel weight vector is expanded to the same spatial dimension as the input feature map and then multiplied with the original feature map.

[0088] The SE channel attention module uses global average pooling to squeeze the input feature map to obtain a global feature descriptor in the channel dimension; then it performs an Excitation operation through a fully connected layer (or 1x1 convolution) and an activation function to obtain channel weights; finally, the channel weights are applied to the original feature map to recalibrate the channel features, so that the network pays more attention to important channels and suppresses unimportant channels, thereby improving the performance of the network.

[0089] Multi-scale feature aggregation module implementation process:

[0090] Step 1: The feature map is directly input into the A module;

[0091] Step 2: The feature map passes through five branches, namely:

[0092] The first branch: three 3x3 convolutions are performed continuously to achieve 8x downsampling, and then 8x upsampling is performed to restore to the original size;

[0093] The first branch: no treatment;

[0094] The third branch: downsampled by 2, 4, and 8 times respectively, and then downsampled and added according to the corresponding size;

[0095] The fourth branch: add together the five results after processing the first branch to the third branch;

[0096] The fifth branch: The result in the fourth branch is subjected to a 3x3 convolution.

[0097] Step 3: Output the result of the fifth branch.

[0098] In the above multi-scale feature aggregation module, the three pieces of information are sampled separately and then added together before being transmitted to the module. First, they are downsampled three times and convolved three times, and then superimposed as shown in the figure, and then upsampled three times; at the same time, the fourth branch is downsampled three times and convolved three times, and finally upsampled. After the four branches are processed, they are added together and then convolved.

[0099] Pyramid Pooling Module

[0100] Step 1: Input feature map X∈R H×W×C , where C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map.

[0101] Step 2: The feature map is processed for the first time through the average pooling layer, and the pooling scale is 1x1; the average pooling formula is as follows:

[0102]

[0103] where k i corresponds to the pooling scale of the pooling layer. i The values ​​are 2, 4, and 8 from shallow to high layers, and (m, n) refers to the position of the channel input x on channel c.

[0104] Step 3: The output of the above steps is passed through the identity mapping layer, the global average pooling layer and two adaptive average pooling layers.

[0105] No operation is performed in the identity mapping layer, i.e., y = x;

[0106] Global Average Pooling (GAP) is a pooling operation that averages all elements on each channel of each input feature map. The feature map is converted from the dimensions of [C, H, W] to the dimensions of [C, 1, 1], where C is the number of channels, H is the height of the feature map, and W is the width of the feature map. The processing principle is:

[0107]

[0108] Among them, y c is the element of the cth channel of the output feature map, x c,i,j is the element at position (i, j) of the cth channel of the input feature map.

[0109] Adaptive Average Pooling is a pooling operation that can pool the input feature map to any specified output size, independent of the size of the input feature map. This is different from the traditional average pooling layer, which requires the size and step size of the pooling kernel to be specified in advance, while adaptive average pooling only requires the size of the final output to be specified. For the output feature map X, the dimension is [C,H in ,W in ], and the desired output feature map size is [C,H out ,W out ], the calculation formula is as follows:

[0110]

[0111] Among them, y c,i,j is the element of the cth channel at position (i, j) of the output feature map, x c,m,n is the element of the cth channel at position (m,n) of the input feature map, k h and k w is the pooling kernel size automatically calculated based on the input and output sizes, s h and w are the strides, which are automatically adjusted based on the input and output sizes to ensure that the average pooling result is of the desired size.

[0112] Step 4: Upsample the four results to make them uniform in scale.

[0113] Step 5: Concatenate the up-sampled results and output them.

[0114] In the pyramid pooling module, the feature map is processed by four different scales of pooling after the input module, and then convolved, and finally connected together after upsampling. We make some adjustments to the pyramid pooling module here. The overall process is still processed by four branches and then connected together, but these four branches are the identity mapping layer, the global average pooling layer, and two adaptive average pooling layers with output sizes of 3x3 and 5x5 respectively.

[0115] The training process of the above model includes:

[0116] Step 1: Build an experimental platform, use Python to build models, and call the PyTorch open source framework to build convolution models. All experiments are conducted on a server (NAVIDIA RTX-3080 GPU, 24G video memory).

[0117] Step 2: Prepare training and test datasets. There are a series of classic datasets for salient object recognition. In order to make the experimental results more comprehensive, five datasets are selected to test our model in many aspects. The five classic datasets are ECSSD, PASCAL-S, DUT-OMRON, HKU-IS, and DUTS-E.

[0118] Step 3: Preprocess the data set, crop the photos in the data set to make them uniform in size, and subtract irrelevant parts of the photos; then use Gaussian filtering and other filtering techniques to remove noise from the photos. We enhance the data by randomly and horizontally flipping the images in advance and adjust the image size to 384x384.

[0119] Step 4: Pre-train the backbone network. The backbone network is mainly used to extract and encode image information, which is convenient for subsequent modules to follow up. This paper selects VGG-16 and RESNET-50 models as the backbone networks, and pre-trains them respectively, and saves the training path and training parameter weights. When pre-training the initialization parameters of the backbone network, we use ImageNet for unified training operations. For other parameters in the training model, we use the random package to randomly generate a series of parameters, and then update and adjust the optimal parameters through continuous training. There are 36 epochs in the training process, and the initial learning rate is 5e-5. In all experimental processes, we use the Adam optimizer to train and optimize weights, and set the weight decay to 5e-4, and the training volume is 10 at a time. For the loss function, we use the standard binary cross entropy loss function, and the loss function formula is as follows:

[0120]

[0121] Step 5: Train the overall model. Set the model parameters the same as in step 4 and train the model in the DUT-TR dataset.

[0122] Step 6: Test on the five data sets in step 2. In order to better reflect the effect of our model, we use three common indicators, F value, MAE value, and S value, to intuitively show the results of our model. The F value parameter is set to 0.3.

[0123] The above steps are the overall operation steps of the model.

[0124] Embodiment 2

[0125] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0126] Embodiment 3

[0127] The purpose of this embodiment is to provide a computer-readable storage medium.

[0128] A computer-readable storage medium stores a computer program, which executes the steps of the above method when executed by a processor.

[0129] Embodiment 4

[0130] The purpose of this embodiment is to provide an image salient object detection system, including:

[0131] Data acquisition module: configured to: acquire the X-ray image to be detected and perform preprocessing to obtain a feature map;

[0132] Data processing module: configured to: input the preprocessed image into the trained U-shaped network;

[0133] In the U-shaped network, coarse features are extracted from the salient areas in the image;

[0134] The extracted coarse features are refined layer by layer from high level to low level to obtain a multi-layer refined feature map with increasingly rich details;

[0135] The refined feature maps are fused layer by layer to generate the final saliency map.

[0136] Embodiment 5

[0137] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any of the above embodiments.

[0138] The steps involved in the apparatus of the above embodiment correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0139] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0140] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A method for detecting salient objects in an image, characterized in that: include: Acquire the X-ray image to be detected and perform preprocessing to obtain a feature map; Input the preprocessed image into the trained U-shaped network; In the U-shaped network, coarse features are extracted from the salient areas in the image; The extracted coarse features are refined layer by layer from high level to low level to obtain a multi-layer refined feature map with increasingly rich details; The refined feature maps are fused layer by layer to generate the final saliency map.

2. The method for detecting salient objects in an image according to claim 1, wherein: The U-shaped network includes: a multi-scale feature aggregation module; the multi-scale feature aggregation module includes five branches; The feature maps are input into five branches respectively. The first branch performs down-sampling on the input feature map through multiple convolutions continuously, and then restores it to the original size through up-sampling. The second branch does not process the input feature map; The third branch downsamples the input feature map by 2 times; The fourth branch downsamples the input feature map by 4 times; The fifth branch downsamples the input feature map by 8 times; The results of the five branches are added together and then output after convolution processing.

3. The method for detecting salient objects in an image according to claim 1, wherein: The U-shaped network also includes: a pyramid pooling module; The pyramid pooling module includes an average pooling layer, an identity mapping layer, a global average pooling layer and two adaptive average pooling layers; The feature map is processed by the average pooling layer for the first time and the output result is output; The above output results are processed by the identity mapping layer, the global average pooling layer and the two adaptive average pooling layers respectively; The four results are upsampled separately to make them uniform in scale; The up-sampled results are concatenated and then output.

4. The method for detecting salient objects in an image according to claim 3, wherein: No operation is performed on the input in the identity mapping layer; Global average pooling averages all elements on each channel of each input feature map; The adaptive average pooling layer pools the input feature map to any specified output size.

5. The method for detecting salient objects in an image according to claim 1, wherein: The U-shaped network also includes: an SE channel attention module; The SE channel attention module includes a global average pooling layer; The global average pooling is used to perform a Squeeze operation on the input feature map to compress the spatial dimension and generate a global feature descriptor; The obtained global feature descriptor is subjected to dimension reduction to obtain a feature vector after dimension reduction; Apply activation function to the reduced feature vector to learn more complex feature representation; The output after activation function processing is upgraded and weights are generated for each channel; Normalize the dimensionally upgraded feature vector to generate the weight of each channel; The generated channel weight vector is expanded to the same spatial dimension as the input feature map, and then multiplied with the input feature map to obtain the output of the SE channel attention module.

6. The method for detecting salient objects in an image according to claim 1, wherein: The U-network uses the standard binary cross entropy loss function during training.

7. An image salient object detection system, characterized in that: include: Data acquisition module: configured to: acquire the X-ray image to be detected and perform preprocessing to obtain a feature map; Data processing module: configured to: input the preprocessed image into the trained U-shaped network; In the U-shaped network, coarse features are extracted from the salient areas in the image; The extracted coarse features are refined layer by layer from high level to low level to obtain a multi-layer refined feature map with increasingly rich details; The refined feature maps are fused layer by layer to generate the final saliency map.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are performed.