A method and apparatus for image saliency detection
By using a multi-path parallel architecture for image saliency detection and fusing feature information from high-resolution and low-resolution paths, the problem of low detection accuracy in existing technologies is solved, and high-precision image saliency detection is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AGRICULTURAL BANK OF CHINA
- Filing Date
- 2022-09-26
- Publication Date
- 2026-06-02
Smart Images

Figure CN115690395B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image saliency detection method and apparatus. Background Technology
[0002] In the era of big data, quickly and accurately locating salient targets in massive images is an important task for intelligent detection. Image saliency detection, as an important precursor to this task, directly or indirectly determines the ultimate success or failure of the task.
[0003] However, current image saliency detection schemes suffer from low detection accuracy. Summary of the Invention
[0004] This invention provides an image saliency detection method and apparatus to achieve high-precision image saliency detection.
[0005] According to one aspect of the present invention, an image saliency detection method is provided, which may include:
[0006] The image to be detected and the trained image saliency detection model are obtained. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module.
[0007] The image to be detected is input into the first convolutional module to obtain a high-resolution feature map, and the high-resolution feature map is processed to obtain a low-resolution feature map.
[0008] The high-resolution feature map is input into the second convolutional module to obtain the high-resolution path map, and the low-resolution feature map is input into the third convolutional module to obtain the low-resolution path map.
[0009] The high-resolution path map and the low-resolution path map are fused to obtain a result map that represents the salient targets detected from the image to be detected;
[0010] Among them, the resolution of the image to be detected, the high-resolution feature map, and the high-resolution path map are consistent, and the resolution of the low-resolution feature map and the low-resolution path map are consistent.
[0011] According to another aspect of the present invention, an image saliency detection apparatus is provided, which may include:
[0012] The model acquisition module is used to acquire the image to be detected and the trained image saliency detection model. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module.
[0013] The feature map acquisition module is used to input the image to be detected into the first convolution module to obtain a high-resolution feature map, and to process the high-resolution feature map to obtain a low-resolution feature map.
[0014] The path map acquisition module is used to input high-resolution feature maps into the second convolutional module to obtain high-resolution path maps, and to input low-resolution feature maps into the third convolutional module to obtain low-resolution path maps.
[0015] The result map module is used to fuse high-resolution path maps and low-resolution path maps to obtain a result map representing salient targets detected from the image to be detected;
[0016] Among them, the resolution of the image to be detected, the high-resolution feature map, and the high-resolution path map are consistent, and the resolution of the low-resolution feature map and the low-resolution path map are consistent.
[0017] The technical solution of this invention can acquire a target image and a trained image saliency detection model. The trained image saliency detection model can achieve high-precision image saliency detection. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module. The target image is input into the first convolutional module to obtain a high-resolution feature map, and the high-resolution feature map is processed to obtain a low-resolution feature map. The high-resolution feature map is input into the second convolutional module to obtain a high-resolution path map, and the low-resolution feature map is input into the third convolutional module to obtain a low-resolution path map. The high-resolution path map and the low-resolution path map are fused, allowing the high-level semantic information in the low-resolution path map output by the low-resolution path to supplement the low-level semantic information in the high-resolution path map output by the high-resolution path, resulting in a result map representing the salient targets detected from the target image. The target image, the high-resolution feature map, and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution, enabling the model to more easily capture salient targets in the image during image fusion. The above technical solution fuses the outputs of the high-resolution path and the low-resolution path of the image saliency detection model. While maintaining the high-resolution path for detecting image saliency, it introduces feature information obtained from the low-resolution path as auxiliary information, thereby achieving high-precision image saliency detection.
[0018] It should be understood that the description in this section is not intended to identify key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of an image saliency detection method provided in Embodiment 1 of the present invention;
[0021] Figure 2 This is a flowchart of another image saliency detection method provided in Embodiment 2 of the present invention;
[0022] Figure 3This is a schematic diagram of a process for obtaining an initial fused feature map provided in Embodiment 2 of the present invention;
[0023] Figure 4 This is a flowchart of another image saliency detection method provided in Embodiment 3 of the present invention;
[0024] Figure 5 This is a flowchart of another image saliency detection method provided in Embodiment 4 of the present invention;
[0025] Figure 6 This is a flowchart of an optional example of an image saliency detection method provided in Embodiment 4 of the present invention;
[0026] Figure 7 This is an illustration of a portion of the feature maps that can be obtained on the high-resolution path of the saliency detection model in the image saliency detection method provided in Embodiment 4 of the present invention;
[0027] Figure 8 This is a structural block diagram of the image saliency detection device provided in Embodiment 5 of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The same applies to "target," "original," etc., and will not be repeated here. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1This is a flowchart of an image saliency detection method provided in Embodiment 1 of the present invention. This embodiment is applicable to image saliency detection, specifically high-precision image saliency detection. The method can be executed by the image saliency detection device provided in this embodiment of the present invention, which can be implemented in software and / or hardware.
[0032] See Figure 1 The method of this invention specifically includes the following steps:
[0033] S110. Obtain the image to be detected and the trained image saliency detection model, wherein the image saliency detection model includes a high-resolution path and a low-resolution path, the high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module.
[0034] In this invention, the image to be detected can be understood as the image for which saliency detection is required. The image saliency detection model can be understood as a trained model capable of performing saliency detection on the image to be detected. In this embodiment, the image saliency detection model is a multi-path parallel architecture model. A high-resolution path can be understood as a path in the image saliency detection model that processes high-resolution images. A low-resolution path can be understood as a path in the image saliency detection model that processes low-resolution images; the number of low-resolution paths can be one or more; optionally, the number of low-resolution paths can be related to the number of modules with convolutional layers in the high-resolution paths. A first convolutional module can be understood as a module with convolutional layers on the high-resolution path. A second convolutional module can be understood as a module with convolutional layers in the high-resolution path connected to the first convolutional module. Specifically, the connection relationship can be between the last convolutional layer in the first convolutional module and the first convolutional layer in the second convolutional module; the number of second convolutional modules can be one or more. If there are multiple second convolutional modules, the first second convolutional module is connected to the last convolutional layer of the first convolutional module, and the multiple second convolutional modules are interconnected. The third convolutional module can be understood as a module with convolutional layers on the low-resolution path; the number of third convolutional modules can be one or more; optionally, the number of third convolutional modules can be related to the number of low-resolution paths.
[0035] S120. Input the image to be detected into the first convolution module to obtain a high-resolution feature map, and process the high-resolution feature map to obtain a low-resolution feature map.
[0036] Here, a high-resolution feature map can be understood as an image with high resolution and features that can reflect whether the image is significant, obtained after the image to be detected has been processed by the first convolutional module; a high-resolution feature map can be a feature map whose resolution has not decreased compared to the image to be detected. A low-resolution feature map can be understood as an image with reduced resolution after processing the high-resolution feature map, but which still has features that can reflect whether the image is significant; optionally, the resolution of the low-resolution feature map can be half the resolution of the high-resolution feature map.
[0037] Specifically, the image to be detected is input into the first convolutional module. After processing by the first convolutional module, a high-resolution feature map can be obtained. Then, the high-resolution feature map is further processed to obtain a low-resolution feature map. The method of processing the high-resolution feature map can be downsampling through pooling layers, and there is no specific limitation on the method of processing the high-resolution feature map.
[0038] S130. Input the high-resolution feature map into the second convolutional module to obtain the high-resolution path map, and input the low-resolution feature map into the third convolutional module to obtain the low-resolution path map.
[0039] In this context, the high-resolution path map can be understood as the feature map corresponding to the high-resolution path after processing by the second convolutional module. The low-resolution path map can be understood as the feature map corresponding to the low-resolution path after processing by the third convolutional module.
[0040] It is understandable that if there are multiple second convolutional modules, the final output of the high-resolution feature map after processing by each second convolutional module is the high-resolution path map. If there are multiple third convolutional modules, the final output of the low-resolution feature map after processing by each third convolutional module is the low-resolution path map. If there are multiple low-resolution paths, the number of low-resolution path maps is related to the number of low-resolution paths; for example, a low-resolution path map can be output for each low-resolution path.
[0041] It is important to note that the resolution of the image to be detected, the high-resolution feature map, and the high-resolution path map are consistent, as are the resolution of the low-resolution feature map and the low-resolution path map; that is, each feature map on the high-resolution path has the same resolution as the image to be detected, and each feature map on the low-resolution path has the same resolution as the feature map on the same path.
[0042] S140. The high-resolution path map and the low-resolution path map are fused to obtain a result map representing the salient targets detected from the image to be detected; wherein the image to be detected, the high-resolution feature map and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution.
[0043] In this context, a salient target can be understood as a target obtained after performing salientity detection on an image; a salient target can be an object or region with salient features, such as a tourist in a desert or the area where a flock of sheep is located on a grassland. The resulting image can be understood as an image that can represent the salient targets in the image to be detected.
[0044] Specifically, the high-resolution path map and the low-resolution path map are fused to obtain a result map representing the salient targets detected from the image to be detected. Optionally, the high-resolution path map and the low-resolution path map can be fused through a convolution operation; there is no specific limitation on the method of fusion.
[0045] For example, the image saliency detection model of this embodiment of the invention can be built based on Visual Geometry Group Network-16 (VGG-16). The last max pooling layer and all subsequent fully connected layers in the deep layers of the VGG-16 network can be removed. The first module with a convolutional layer in the VGG-16 network is used as the first convolutional module, followed by a second convolutional module. Each other module with a convolutional layer except the first one is used as the first third convolutional module for each low-resolution path. The pooling layers in the VGG-16 network can be used to process the high-resolution feature map to obtain the low-resolution feature map.
[0046] Optionally, the high-resolution path map and the low-resolution path map are fused to obtain a result map representing the salient targets detected from the image to be detected, including: obtaining an upsampling rate, wherein the upsampling rate is determined based on the resolution of the high-resolution path map and the resolution of the low-resolution path map; upsampling the low-resolution path map based on the upsampling rate to obtain an upsampled path map with the same resolution as the high-resolution path map; and fusing the high-resolution path map and the upsampled path map to obtain a result map representing the salient targets detected from the image to be detected.
[0047] The upsampling rate can be understood as the rate at which a low-resolution path map is upsampled to the same resolution as a high-resolution path map. For example, if an image saliency detection model has five paths—one high-resolution path and four low-resolution paths—each with half the resolution of the previous path, then the upsampling rates for the four low-resolution paths are {2, 4, 8, 16}. The upsampled path map can be understood as a feature map obtained by upsampling the low-resolution path map, resulting in a map with the same resolution as the high-resolution path map.
[0048] The technical solution of this invention can acquire a target image and a trained image saliency detection model. The trained image saliency detection model can achieve high-precision image saliency detection. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module. The target image is input into the first convolutional module to obtain a high-resolution feature map, and the high-resolution feature map is processed to obtain a low-resolution feature map. The high-resolution feature map is input into the second convolutional module to obtain a high-resolution path map, and the low-resolution feature map is input into the third convolutional module to obtain a low-resolution path map. The high-resolution path map and the low-resolution path map are fused, allowing the high-level semantic information in the low-resolution path map output by the low-resolution path to supplement the low-level semantic information in the high-resolution path map output by the high-resolution path, resulting in a result map representing the salient targets detected from the target image. The target image, the high-resolution feature map, and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution, enabling the model to more easily capture salient targets in the image during image fusion. The above technical solution fuses the outputs of the high-resolution path and the low-resolution path of the image saliency detection model. While maintaining the high-resolution path for detecting image saliency, it introduces feature information obtained from the low-resolution path as auxiliary information, thereby achieving high-precision image saliency detection.
[0049] Example 2
[0050] Figure 2This is a flowchart of another image saliency detection method provided in Embodiment 2 of the present invention. This embodiment is an optimization based on the above-described technical solutions. In this embodiment, optionally, the second convolutional module includes a high-resolution pre-weighting module and a high-resolution post-weighting module, and the high-resolution path further includes a high-resolution weighting module connected between the high-resolution pre-weighting module and the high-resolution post-weighting module for implementing weighted connections; the third convolutional module includes a low-resolution pre-weighting module and a low-resolution post-weighting module, and the low-resolution path further includes a low-resolution weighting module connected between the low-resolution pre-weighting module and the low-resolution post-weighting module for implementing weighted connections; a high-resolution feature map is input into the high-resolution pre-weighting module, and a low-resolution feature map is input into the low-resolution pre-weighting module; the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module are input into the high-resolution weighting module, and the output result of the high-resolution weighting module is input into the high-resolution post-weighting module to obtain a high-resolution path map; the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module are input into the low-resolution weighting module, and the output result of the low-resolution weighting module is input into the low-resolution post-weighting module to obtain a low-resolution path map. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.
[0051] See Figure 2 The method in this embodiment may specifically include the following steps:
[0052] S210. Obtain the image to be detected and the trained image saliency detection model, wherein the image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module. The second convolutional module includes a high-resolution pre-weighted module and a high-resolution post-weighted module. The high-resolution path also includes a high-resolution weighted module connected between the high-resolution pre-weighted module and the high-resolution post-weighted module to implement weighted connections. The third convolutional module includes a low-resolution pre-weighted module and a low-resolution post-weighted module. The low-resolution path also includes a low-resolution weighted module connected between the low-resolution pre-weighted module and the low-resolution post-weighted module to implement weighted connections.
[0053] The high-resolution pre-weighting module can be understood as the second convolutional module on the high-resolution path, preceding the high-resolution weighting module. The high-resolution post-weighting module can be understood as the second convolutional module on the high-resolution path, following the high-resolution weighting module. The low-resolution pre-weighting module can be understood as the third convolutional module on the low-resolution path, preceding the low-resolution weighting module. The low-resolution post-weighting module can be understood as the third convolutional module on the low-resolution path, following the low-resolution weighting module. The high-resolution weighting module can be understood as a module on the high-resolution path that weights the high-resolution path map and the low-resolution path map into a high-resolution path map, and can be used to selectively fuse feature information from different resolution paths. Optionally, there can be multiple high-resolution weighting modules, each located between different second convolutional modules. The low-resolution weighting module can be understood as a module on the low-resolution path that weights the high-resolution path map and the low-resolution path map into a low-resolution path map, and can be used to selectively fuse feature information from different resolution paths. Optionally, there can be multiple low-resolution weighting modules, each located between different third convolutional modules.
[0054] It is understandable that the high-resolution weighting module before and after high-resolution weighting are different for each high-resolution weighting module. For example, if there are second convolutional modules numbered 1, 2, and 3 on the high-resolution path, and the high-resolution weighting module connected between the second convolutional modules numbered 1 and 2 is numbered 'a', then the high-resolution weighting module before 'a' is the second convolutional module numbered 1, and the high-resolution weighting module after 'a' is the second convolutional module numbered 2. If the high-resolution weighting module connected between the second convolutional modules numbered 2 and 3 is numbered 'b', then the high-resolution weighting module before 'b' is the second convolutional module numbered 2, and the high-resolution weighting module after 'b' is the second convolutional module numbered 3. Each low-resolution weighted module has a different pre-weighted and post-weighted low-resolution module. For example, if there are third convolutional modules numbered 4, 5, and 6 on the low-resolution path, and the low-resolution weighted module connected between the third convolutional modules numbered 4 and 5 is numbered c, then the pre-weighted low-resolution module of c is the third convolutional module numbered 4, and the post-weighted low-resolution module of c is the third convolutional module numbered 5. If the low-resolution weighted module connected between the third convolutional modules numbered 5 and 6 is numbered d, then the pre-weighted low-resolution module of d is the third convolutional module numbered 5, and the post-weighted low-resolution module of d is the third convolutional module numbered 6.
[0055] S220. Input the image to be detected into the first convolution module to obtain a high-resolution feature map, and process the high-resolution feature map to obtain a low-resolution feature map.
[0056] S230. Input the high-resolution feature map into the high-resolution pre-weighting module, and input the low-resolution feature map into the low-resolution pre-weighting module.
[0057] It is understood that, in the embodiments of the present invention, a high-resolution feature map can be input into a high-resolution pre-weighting module so that the high-resolution pre-weighting module processes the high-resolution feature map into an output result that can be input into the high-resolution weighting module; and a low-resolution feature map can be input into a low-resolution pre-weighting module so that the low-resolution pre-weighting module processes the low-resolution feature map into an output result that can be input into the high-resolution weighting module.
[0058] S240. Input the output results of the high-resolution weighting module and the low-resolution weighting module into the high-resolution weighting module, and input the output results of the high-resolution weighting module into the high-resolution weighted module to obtain the high-resolution path map.
[0059] Specifically, the outputs of the high-resolution pre-weighting module and the low-resolution pre-weighting module are input into the high-resolution weighting module. The high-resolution weighting module can perform weighting processing on the high-resolution feature map and the low-resolution feature map. For example, the low-resolution feature map can be upsampled first, and then the upsampled low-resolution feature map with the same resolution as the high-resolution feature map can be weighted together with the high-resolution feature map. The output of the high-resolution weighting module is then input into the high-resolution post-weighting module to obtain the high-resolution path map.
[0060] S250. Input the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module into the low-resolution weighting module, and input the output results of the low-resolution weighting module into the low-resolution post-weighting module to obtain the low-resolution path map.
[0061] Specifically, the outputs of the high-resolution pre-weighting module and the low-resolution pre-weighting module are input into the low-resolution weighting module. The low-resolution weighting module can perform weighted processing on the high-resolution feature map and the low-resolution feature map. For example, the high-resolution feature map can be downsampled first, and then the downsampled high-resolution feature map with the same resolution as the low-resolution feature map can be weighted together with the low-resolution feature map. The output of the low-resolution weighting module is then input into the low-resolution weighted module to obtain the low-resolution path map.
[0062] S260. The high-resolution path map and the low-resolution path map are fused to obtain a result map representing the salient targets detected from the image to be detected; wherein the image to be detected, the high-resolution feature map and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution.
[0063] The technical solution of this invention involves inputting a high-resolution feature map into a high-resolution pre-weighting module and a low-resolution feature map into a low-resolution pre-weighting module; inputting the outputs of both modules into a high-resolution weighting module and then into a high-resolution post-weighting module to obtain a high-resolution path map; and inputting the outputs of both modules into a low-resolution weighting module and then into a low-resolution post-weighting module to obtain a low-resolution path map. This allows for the fusion of features corresponding to low-resolution paths with features corresponding to high-resolution paths, thereby effectively fusing globally effective features and significantly improving feature utilization.
[0064] An optional technical solution involves representing the output of the high-resolution pre-weighting module through a high-resolution feature channel, and representing the output of the low-resolution pre-weighting module through a low-resolution feature channel. Inputting the outputs of both modules into the high-resolution weighting module includes: upsampling the low-resolution feature channel to obtain an upsampled feature channel, wherein the resolution of the upsampled feature channel is consistent with that of the high-resolution feature channel; inputting the high-resolution feature channel and the upsampled feature channel into the high-resolution weighting module to perform the following steps: concatenating the high-resolution feature channel and the upsampled feature channel to obtain an initial fused feature map; determining a weight vector based on the initial fused feature map, wherein the length of the weight vector is consistent with the sum of the number of high-resolution feature channels and the number of low-resolution feature channels; processing the initial fused feature map based on the weight vector to obtain a target fused feature map; and inputting the output of the high-resolution weighting module into the high-resolution post-weighting module to obtain a high-resolution path map, including: obtaining the target fused feature map output by the high-resolution weighting module and inputting the target fused feature map into the high-resolution post-weighting module to obtain a high-resolution path map.
[0065] In this context, the high-resolution feature channel can be understood as the channel-represented feature map output by the high-resolution weighting module before weighting. The low-resolution feature channel can be understood as the channel-represented feature map output by the low-resolution weighting module before weighting. The upsampled feature channel can be understood as the low-resolution feature channel upsampled to a channel-represented feature map with the same resolution as the high-resolution feature channel. The initial fused feature map can be understood as the feature map obtained by concatenating the high-resolution feature channel and the upsampled feature channel; for example, the upsampled feature channel can be concatenated with the high-resolution feature channel. The weight vector can be understood as a vector indicating the magnitude of the corresponding feature channel weights; the weight vector represents the probability that the corresponding channel detection is significant; the number of elements in the weight vector is the same as the number of channels in the initial fused feature map. The target fused feature map can be understood as the feature map output after the high-resolution weighting module has weighted the high-resolution and low-resolution feature channels.
[0066] It is understandable that since each channel of the feature map can be regarded as a feature detector, and different channels focus on different effective features, the contribution of each channel to the final saliency detection is also very different. Therefore, in this embodiment of the invention, the output of the high-resolution pre-weighting module can be represented by the high-resolution feature channel, and the output of the low-resolution pre-weighting module can be represented by the low-resolution feature channel. Weights are calculated and assigned to each channel of the feature map to measure the probability of detecting salient targets in that channel.
[0067] For example, see Figure 3 For the obtained n1 high-resolution feature channels and n2 low-resolution feature channels, the resolution of each low-resolution feature channel is upsampled to obtain an upsampled feature channel with the same resolution as the high-resolution feature channel; the n1 high-resolution feature channels and the n2 upsampled feature channels are concatenated to obtain a feature map with n1+n2 channels; the feature map with n1+n2 channels is reduced to a feature map with (n1+n2) / 4 channels, and then it is increased to an initial fused feature map with n1+n2 channels.
[0068] Furthermore, in this embodiment of the invention, the high-resolution feature channels can be downsampled to obtain downsampled feature channels, wherein the resolution of the downsampled feature channels is consistent with that of the low-resolution feature channels. The low-resolution feature channels and the downsampled feature channels are input into a low-resolution weighting module to perform the following steps: concatenating the low-resolution feature channels and the downsampled feature channels to obtain a low-resolution initial fusion feature map; determining a low-resolution weight vector based on the low-resolution initial fusion feature map, wherein the length of the low-resolution weight vector is consistent with the sum of the number of high-resolution feature channels and the number of low-resolution feature channels; processing the low-resolution initial fusion feature map based on the low-resolution weight vector to obtain a low-resolution target fusion feature map; and inputting the output of the low-resolution weighting module into a low-resolution weighted module to obtain a low-resolution path map, including: obtaining the low-resolution target fusion feature map output by the low-resolution weighting module and inputting the low-resolution target fusion feature map into the low-resolution weighted module to obtain a low-resolution path map. The downsampled feature channels can be understood as high-resolution feature channels downsampled to a channel-represented feature map with the same resolution as the low-resolution feature channels. The low-resolution initial fusion feature map can be understood as the feature map obtained by concatenating the low-resolution feature channels and the downsampled feature channels; for example, the downsampled feature channels can be concatenated with the low-resolution feature channels. The low-resolution weight vector can be understood as a vector indicating the magnitude of the corresponding channel weights; the low-resolution weight vector represents the probability that the corresponding channel detection is significant; the number of elements in the low-resolution weight vector is the same as the number of channels in the low-resolution initial fusion feature map. The low-resolution target fusion feature map can be understood as the feature map output after the low-resolution weighting module has weighted the low-resolution and high-resolution feature channels.
[0069] For example, if an image saliency detection model has one high-resolution path and three low-resolution paths, the high-resolution feature channel corresponding to the high-resolution path and the upsampled feature channel corresponding to the low-resolution path are used... This means that the images input to each path before entering the high-resolution weighting module can be expanded by channel. A high-resolution feature channel and three upsampled feature channels are combined using the formula... The initial fused feature map is obtained by performing the connection. The weight vector is W; the initial fused feature map With weight vector W according to the formula Multiply to generate a target fusion feature map. in, It is the k-th channel of the feature map of the j-th path; w j and h jThese represent the width and height of the feature map, respectively, and C is the number of channels; Concat(·) represents the channel concatenation operation.
[0070] In this embodiment of the invention, weighting multiple paths according to channels can effectively integrate channels from multiple paths. Channels that differ significantly from the saliency of the effective information are assigned smaller weights, so that they have less impact on the subsequent calculation process; channels that differ less significantly from the saliency of the effective information are assigned larger weights, so that the channels with larger weights will be the main objects of subsequent calculations and play an important role, thereby further improving the utilization rate of features.
[0071] Based on the above scheme, another optional technical solution is to determine the weight vector based on the initial fused feature map, including: performing a global max pooling operation on the initial fused feature map to obtain an intermediate vector; performing a fully connected operation with a ReLU activation function on the intermediate vector; and performing a fully connected operation with a Sigmoid activation function on the fully connected intermediate vector to obtain the weight vector.
[0072] The intermediate vector can be understood as the vector obtained by performing global max pooling on the initial fused feature map. The ReLU activation function is the Rectified Linear Unit (ReLU), a commonly used activation function in non-linear artificial neural networks. The Sigmoid activation function is a common sigmoid function in biology, often used as an activation function in neural networks. The Sigmoid activation function is a non-linear activation function; in this embodiment of the invention, the Sigmoid activation function restricts the value of each element in W to the interval [0,1].
[0073] For example, combining the channel connection operation example above, according to the formula In the initial fusion feature map The above uses global max pooling to obtain a pool of length . The intermediate vector is used to perform a full connection operation with ReLU activation on the intermediate vector, and a full connection operation with Sigmoid activation is performed on the fully connected intermediate vector to obtain the desired weight vector. Here, GMP(·) is the global max pooling operation function, and θ1 and θ2 correspond to the layer parameters of ReLU(·) and Sigmoid(·), respectively.
[0074] In this embodiment of the invention, the weight vector is obtained by performing a global max pooling operation, a full connection operation with a ReLU activation function, and a full connection operation with a Sigmoid activation function. This method can accurately and effectively obtain the weight vector and effectively prevent gradient explosion during the training process.
[0075] Example 3
[0076] Figure 4 This is a flowchart of another image saliency detection method provided in Embodiment 3 of the present invention. This embodiment is based on the above-described technical solutions and optimized. In this embodiment, optionally, the high-resolution path further includes a high-resolution enhancement module connected to the second convolution module for region enhancement; the high-resolution feature map is input into the second convolution module to obtain the feature map before enhancement, and the feature map before enhancement is input into the high-resolution enhancement module to obtain the feature map after enhancement; the feature map after enhancement is used as the high-resolution path map. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.
[0077] See Figure 4 The method in this embodiment may specifically include the following steps:
[0078] S310. Obtain the image to be detected and the trained image saliency detection model, wherein the image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and a high-resolution enhancement module connected to the second convolutional module for performing region enhancement. The low-resolution path includes a third convolutional module.
[0079] The high-resolution enhancement module can be understood as a module that can enhance the saliency representation of the input image; optionally, in the presence of at least two second convolutional modules, the high-resolution enhancement module can be located after the last second convolutional module on the high-resolution path and connected to the last second convolutional module on the high-resolution path.
[0080] S320. Input the image to be detected into the first convolution module to obtain a high-resolution feature map, and process the high-resolution feature map to obtain a low-resolution feature map.
[0081] S330. Input the high-resolution feature map into the second convolutional module to obtain the feature map before enhancement, and input the feature map before enhancement into the high-resolution enhancement module to obtain the feature map after enhancement.
[0082] The pre-enhancement feature map can be understood as the image output by the second convolutional module after the high-resolution feature map is input into it. The post-enhancement feature map can be understood as the image output by the high-resolution enhancement module after the pre-enhancement feature map is input into it, and the enhancement module improves the saliency representation capability.
[0083] S340. Use the enhanced feature map as a high-resolution path map.
[0084] Furthermore, in this embodiment of the invention, optionally, the low-resolution path further includes a low-resolution enhancement module connected to the third convolutional module for region enhancement. The low-resolution feature map is input into the third convolutional module to obtain a low-resolution path map, including: inputting the low-resolution feature map into the third convolutional module to obtain a feature map before low-resolution enhancement, and inputting the feature map before low-resolution enhancement into the low-resolution enhancement module to obtain a feature map after low-resolution enhancement; the feature map after low-resolution enhancement is used as the low-resolution path map. Optionally, when there are multiple low-resolution paths, the number of low-resolution path maps can be the same as the number of low-resolution paths, that is, a low-resolution enhancement module can be connected after the last third convolutional module of each low-resolution path.
[0085] S350. Input the low-resolution feature map into the third convolutional module to obtain the low-resolution path map.
[0086] S360. The high-resolution path map and the low-resolution path map are fused to obtain a result map representing the salient targets detected from the image to be detected; wherein the image to be detected, the high-resolution feature map and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution.
[0087] The technical solution of this invention includes a high-resolution enhancement module connected to the second convolution module for performing region enhancement in the high-resolution path; the high-resolution feature map is input into the second convolution module to obtain the feature map before enhancement, and the feature map before enhancement is input into the high-resolution enhancement module to obtain the feature map after enhancement; the feature map after enhancement is used as the high-resolution path map, thereby improving the detection capability of the image saliency detection model and making the detection results more prominent and obvious.
[0088] An optional technical solution involves inputting the pre-enhancement feature map into a high-resolution enhancement module, comprising: inputting the pre-enhancement feature map into the high-resolution enhancement module to perform the following steps through the high-resolution enhancement module: obtaining a confidence map corresponding to the pre-enhancement feature map; processing the confidence map to obtain an enhancement map, wherein a portion of all values in the enhancement map are less than or equal to 1, and another portion of values are greater than 1; and processing the pre-enhancement feature map based on the enhancement map.
[0089] The confidence map can be understood as an image representing the probability that each pixel in the original feature map is a salient target. A larger value in the confidence map indicates that the corresponding pixel is more likely to be a salient target; conversely, a smaller value indicates that the corresponding pixel is less likely to be a salient target. The enhancement map can be understood as the confidence map after enhancement. In the enhancement map, pixels that are likely to be salient targets have values greater than 1, while pixels that are likely not to be salient targets have values less than or equal to 1. When processing the original feature map based on the enhancement map, such as multiplying the enhancement map with the original feature map, the pixel values of pixels that are likely to be salient targets in the original feature map can be increased, while the pixel values of pixels that are likely not to be salient targets in the original feature map can be decreased. This increases the difference between the pixel values of the two types of pixels, making it easier to distinguish between pixels that are likely to be salient targets and those that are likely not. Therefore, the enhancement map is better than the confidence map at representing the probability that each pixel in the original feature map is a salient target.
[0090] Furthermore, the low-resolution feature map before enhancement is input into the low-resolution enhancement module, including: inputting the low-resolution feature map before enhancement into the low-resolution enhancement module to perform the following steps: obtaining a low-resolution confidence map corresponding to the low-resolution feature map before enhancement; processing the low-resolution confidence map to obtain a low-resolution enhanced map, wherein a portion of all values in the low-resolution enhanced map are less than or equal to 1, and another portion are greater than 1; processing the low-resolution feature map before enhancement based on the low-resolution enhanced map. The low-resolution confidence map can be understood as an image that characterizes the probability that each pixel in the low-resolution feature map before enhancement is a salient target. The low-resolution enhanced map can be understood as a low-resolution confidence map after enhancement of the low-resolution confidence map, and the low-resolution enhanced map better represents the probability that each pixel in the low-resolution feature map before enhancement is a salient target than the low-resolution confidence map itself.
[0091] In this embodiment of the invention, a confidence map corresponding to the feature map before enhancement is obtained; the confidence map is processed to obtain an enhancement map; and the feature map before enhancement is processed based on the enhancement map to enhance the saliency representation of the image and weaken the useless information of non-salient targets.
[0092] Optionally, obtaining a confidence map corresponding to the feature map before enhancement includes: reducing the dimensionality of the feature map before enhancement to obtain a dimensionality-reduced feature map; processing the dimensionality-reduced feature map based on a convolutional layer with a sigmoid activation function to obtain a confidence map corresponding to the feature map before enhancement, wherein all values in the confidence map are between 0 and 1; and / or processing the confidence map to obtain an enhancement map, including: copying the confidence map to obtain a copied map, and summing the confidence map and the copied map to obtain the enhancement map. Furthermore, the processing method for obtaining the confidence map corresponding to the feature map before enhancement and / or processing the confidence map to obtain the enhancement map on the low-resolution path is the same as the processing method on the high-resolution path.
[0093] In this context, the dimensionality-reduced feature map can be understood as the image obtained after reducing the dimensionality of the original feature map. The copy image can be understood as an image that is completely identical to the confidence map after it has been copied.
[0094] For example, if an image saliency detection model has one high-resolution path and four low-resolution paths, the pre-enhancement feature map uses... S indicates that n This is the pre-enhancement feature map for the nth path; expressed by the formula... For each path, the pre-enhancement feature map is reduced in dimensionality using a 16-channel convolution operation with a 3×3 kernel, resulting in a dimensionality-reduced feature map. This dimensionality-reduced feature map is then processed using a convolutional layer with a sigmoid activation function to obtain a result similar to the pre-enhancement feature map. The corresponding confidence plot; the confidence plot is copied to obtain a copy plot, and the confidence plot and the copy plot are summed to obtain the augmented plot. It should be noted that, since the values in the confidence plot are between 0 and 1, the enhancement plot... If some of the values in the set are less than or equal to 1, and another set of values are greater than 1, then... This can be viewed as an enhanced confidence map. The enhanced map better represents the probability that each pixel in the original feature map is a salient object than the confidence map, thus enhancing the features of salient objects while weakening the features of insignificant objects. (Using the formula...) Enhanced image With strong pre-feature map S n Multiply the results and then process the result using a convolutional layer with sigmoid activation to obtain the enhanced feature map. This allows us to obtain the enhanced feature map. It possesses strong features of salient targets, while also exhibiting weak features of insignificant targets. Here, Add(·) is a pixel-wise addition operation, Conv(·) is a convolution operation, and θ3, θ4, and θ5 are the layer parameters of Conv(·), Sigmoid(·), and Add(·), respectively.
[0095] Example 4
[0096] Figure 5 This is a flowchart of another image saliency detection method provided in Embodiment 4 of the present invention. This embodiment is based on the above-described technical solutions and optimized. In this embodiment, optionally, the image saliency detection model is pre-trained through the following steps: obtaining training images and target label maps of salient objects in the training images, and using the training images and target label maps as a set of training samples; obtaining the original saliency detection model, wherein the network structure of the original saliency detection model is the same as the network structure of the image saliency detection model to be trained; training the original saliency detection model based on multiple sets of training samples to obtain the image saliency detection model. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.
[0097] See Figure 5 The method in this embodiment may specifically include the following steps:
[0098] S410. Obtain the training image and the target label map of the salient object in the training image, and use the training image and the target label map as a set of training samples.
[0099] In this context, training images can be understood as images used to train the image saliency detection model, and target label maps can be understood as label maps that characterize whether pixels in the training images belong to salient targets. Training samples can be understood as samples that include training images and their corresponding target label maps, used to train the image saliency detection model.
[0100] For example, a training image contains salient targets, and a target label map exists for these salient targets. The pixel value corresponding to the salient target in the training image is 1 at the corresponding position in the target label map, and the pixel value corresponding to the non-salient target in the training image is 0 at the corresponding position in the target label map. This training image and its corresponding target label map are used as a set of training samples.
[0101] S420. Obtain the original saliency detection model, wherein the network structure of the original saliency detection model is the same as the network structure of the image saliency detection model to be trained.
[0102] The original saliency detection model can be understood as the image saliency detection model to be trained.
[0103] S430. The original saliency detection model is trained based on multiple sets of training samples to obtain the image saliency detection model.
[0104] Specifically, multiple sets of training samples are input into the original saliency detection model to train the original saliency detection model and obtain the image saliency detection model.
[0105] For example, a set consisting of multiple training samples is used It means that I i L represents the training image. i The target label map represents the salient objects in the training images, where N is the size of the multiple training samples X. If the image saliency detection model has one high-resolution path and four low-resolution paths, This represents the five paths of the image saliency detection model, where the first path, P1, is a high-resolution path, and the remaining paths, P2, P3, P4, and P5, are low-resolution paths. The input image for path P1 is I. i The resolution of each feature map in this path remains 224×224. For the remaining low-resolution feature maps input to P2 to P5, the resolution is halved in turn, and remains unchanged in this path.
[0106] For example, the set of multiple training samples can be the data-augmented DUTS-TR dataset, which is 6 times the size of the original set of multiple training samples. The validation set is randomly selected from the DUTS-TR dataset at a ratio of 0.1. During training, the size of each image is uniformly set to 224×224, and the Adaptive Moment Estimation (Adam) optimizer is used for parameter updates, with an initial learning rate of 10. -6 In this embodiment of the invention, the parameters of the first module of each path in the original saliency detection model can use the pre-trained default parameters of the VGG-16 network, while all other layer parameters are randomly initialized. The entire training process is expected to take approximately 80 hours. Furthermore, Figure 5 The second and third convolutional modules can be composed of three 16-channel convolutional layers with a kernel size of 3×3. The entire original saliency detection model is implemented based on the popular and open-source Keras framework. Hardware-wise, a six-core PC can be used, equipped with an Intel Core i7-7800x 3.50GHz CPU (16GB RAM-300B) and a GTX 1080ti GPU.
[0107] S440. Obtain the image to be detected and the trained image saliency detection model, wherein the image saliency detection model includes a high-resolution path and a low-resolution path, the high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module.
[0108] S450. Input the image to be detected into the first convolution module to obtain a high-resolution feature map, and process the high-resolution feature map to obtain a low-resolution feature map.
[0109] S460. Input the high-resolution feature map into the second convolutional module to obtain the high-resolution path map, and input the low-resolution feature map into the third convolutional module to obtain the low-resolution path map.
[0110] S470. The high-resolution path map and the low-resolution path map are fused to obtain a result map representing the salient targets detected from the image to be detected; wherein the image to be detected, the high-resolution feature map and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution.
[0111] It should be noted that, in this embodiment of the invention, S410-S430 is performed to pre-train an image saliency detection model; and S440-S470 is performed to perform image saliency detection based on the trained image saliency detection model.
[0112] The technical solution of this invention pre-trains an image saliency detection model through the following steps: acquiring a training image and target label maps of salient objects in the training image, and using the training image and target label maps as a set of training samples; acquiring an original saliency detection model; training the original saliency detection model based on multiple sets of training samples to obtain an image saliency detection model, thereby enabling the detection of image saliency.
[0113] An optional technical solution uses training images and target label maps as a set of training samples, including: extracting edges from the target label maps to obtain edge label maps, wherein the edge label maps correspond to the target edges of salient objects in the training images; using the training images, target label maps, and edge label maps as a set of training samples; training the original saliency detection model based on multiple sets of training samples to obtain an image saliency detection model, including: for each set of training samples, inputting the training images from the training samples into the original saliency detection model to obtain the edge probability map predicted by the second convolutional module and the target probability map predicted by the third convolutional module in the original saliency detection model; calculating the target loss based on the target probability map and the target label map, and calculating the edge loss based on the edge probability map and the edge label map; obtaining the final loss based on the target loss and the edge loss, and adjusting the network parameters in the original saliency detection model based on the final loss to train the image saliency detection model.
[0114] In this context, the edge label map can be understood as a label map that characterizes whether a pixel in the training image belongs to the edge of a salient target. A target edge can be understood as the edge region of a salient target in the training image; the width of the target edge can be a pre-set fixed value or determined according to the size of the salient target, without specific limitations here. The edge probability map can be understood as an image that characterizes the probability that a pixel in the training image predicted by the second convolutional module belongs to a target edge. The target probability map can be understood as an image that characterizes the probability that a pixel in the training image predicted by the third convolutional module belongs to a salient target. The target loss can be understood as the loss for salient targets. The edge loss can be understood as the loss for the edges of salient targets. The final loss can be understood as the loss used to adjust the network parameters in the original saliency detection model. The network parameters can be understood as the parameters in the original saliency detection model.
[0115] Understandably, before training the original saliency detection model based on multiple sets of training samples, the resolution of the training images, target label maps, and / or edge label maps in each training sample can be adjusted so that the training samples can be applied to the different resolutions corresponding to different paths of the original image saliency detection model.
[0116] It should be noted that a Laplacian filter can be applied to both the training image and the target label map to obtain the target edges of salient objects in the training image and their corresponding edge label maps. In this embodiment of the invention, the method for obtaining the target edges of salient objects in the training image and their corresponding edge label maps is not specifically limited.
[0117] For example, the edge labels in the edge label map are represented by L. b(x,y)∈{0,1} is used to represent the target labels in the target label graph, and L is used to represent the target labels in the target label graph. b The loss can be represented as (x,y)∈{0,1}, where (x,y) represents the coordinates of a pixel. If the original image saliency detection model has one high-resolution path and four low-resolution paths, the loss of the five paths can be expressed as... This can be expressed by the formula: Calculate the edge loss The four low-resolution paths can be expressed by the formula. The target loss for each of the four low-resolution paths was calculated. For the loss on the j-th path, α is the probability that the predicted pixel (x, y) belongs to a salient target or a target edge. α is a constant and can be set to 0.7. The final loss is obtained based on the target loss and edge loss corresponding to the five paths, and the network parameters in the original saliency detection model are adjusted based on the final loss to train the image saliency detection model.
[0118] In this embodiment of the invention, the final loss is obtained based on the target loss and the edge loss, and the network parameters in the original saliency detection model are adjusted based on the final loss to train the image saliency detection model. This can avoid the problem of low image saliency detection accuracy due to complex edges and achieve high-precision image saliency detection.
[0119] Based on the above scheme, another optional technical solution, after inputting the training images in the training samples into the original saliency detection model, further includes: obtaining a fusion probability map, wherein the fusion probability map is predicted by the original saliency detection model for salient targets in the training images; calculating the fusion loss based on the fusion probability map and the target label map; and obtaining the final loss based on the target loss and the edge loss, including obtaining the final loss based on the target loss, the edge loss and the fusion loss.
[0120] The fusion probability map can be understood as the probability map obtained by the original saliency detection model for predicting whether a pixel in the training image belongs to a salient target. The fusion loss can be understood as the loss calculated based on the fusion probability map and the target label map.
[0121] For example, after inputting the training images from the training samples into the original saliency detection model, a fusion probability map can be obtained. Combining this with the example above of obtaining the final loss based on target loss and edge loss, the fusion loss Ψ f It can be done through formula Ψ f =-∑ (x,y) [L fo(x,y)log(P fo (x,y))+(1-L fo (x,y))log(1-P fo The result is calculated from (x,y)]. Where P... fo (x, y) represents the probability value at coordinate point (x, y) on the fusion probability map. Based on the target loss, edge loss, and fusion loss, according to the formula... The final loss is Ψ.
[0122] In this embodiment of the invention, the final loss is obtained based on the target loss, edge loss and fusion loss, which can effectively guide the original saliency detection model to learn more features of salient targets.
[0123] To better understand the technical solutions of the above embodiments of the present invention, another optional example is provided here. For example, see... Figure 6In this embodiment of the invention, the last max-pooling layer and all subsequent fully connected layers in the deep layers of the VGG-16 network are removed. Based on the outflow direction of data in the VGG-16 network, the modules with convolutional layers in the VGG-16 network are named Conv1_2, Conv2_2, Conv3_3, Conv4_3, and Conv5_3, respectively. The first number in the name of each module with convolutional layers in the VGG-16 network indicates the path in which the module is located, and the second number indicates the layer in which the last convolutional layer of the module is located. Based on the modules with convolutional layers in the VGG-16 network, an image saliency detection model with one high-resolution path and four low-resolution paths is constructed. The image to be detected is input into the image saliency detection model. The high-resolution feature map of the image to be detected, after being processed by Conv1_2, flows into the first high-resolution path. The high-resolution feature map is then pooled into a low-resolution feature map with a resolution half that of the original high-resolution feature map. This low-resolution feature map is input into the Conv2_2 module, and the processed low-resolution feature map output by the Conv2_2 module flows into the second path, which serves as the low-resolution path. The processed low-resolution feature map output by the Conv2_2 module is then pooled into a low-resolution feature map with a resolution half that of the original low-resolution feature map. This low-resolution feature map with a resolution half that of the original low-resolution feature map is input into Conv3_3. This process continues until the final low-resolution feature map flowing into Conv5_3 is 1 / 16 of the resolution of the high-resolution feature map. In the image saliency detection model, the second and third convolutional modules (excluding the first module) in each path are named Block1_1, Block1_2, Block1_3, Block1_4, Block2_1, Block2_2, Block2_3, Block3_1, Block3_2, and Block4_1, respectively. The first number in the name of each second and third convolutional module indicates the path number of the module, and the second number indicates the non-original VGG-16 network module within that path. High-resolution weighted modules are added after Block1_2 and Block1_3 to weight the feature maps of multiple paths on high-resolution paths. Low-resolution weighted modules are added after Block2_1, Block2_2, Conv3_3, Block3_1, and Conv4_3 to weight the feature maps of multiple paths on low-resolution paths.The weighted feature maps of each path are input into either the high-resolution or low-resolution enhancement module. Each feature map is then subjected to convolution (Conv) to reduce its dimensionality. The reduced feature maps are then copied and processed in the same way through convolutional layers with sigmoid activation to obtain confidence maps. These two identical confidence maps are then added together to obtain the enhanced map. The enhanced map is multiplied by the weighted feature map, and the result is processed through a convolutional layer with sigmoid activation to obtain the enhanced feature map. The enhanced feature maps corresponding to each path are upsampled according to the path order at upsampling rates {1, 2, 4, 8, 16} to obtain upsampled path maps with the same resolution as the feature maps on the high-resolution paths. These upsampled path maps are then fused through convolution operations and convolutional layers with sigmoid activation to obtain the result map representing the salient targets detected in the image.
[0124] To better understand the above technical solutions, embodiments of the present invention provide partial illustrations that can be obtained along the high-resolution path of the saliency detection model. For example, see... Figure 7 , Figure 7 The top left image shows the feature map output by the first high-resolution weighted module on the high-resolution path; Figure 7 The upper middle figure in the image shows the feature map output by the second high-resolution weighted module on the high-resolution path; Figure 7 The top right image shows the feature map output by the high-resolution enhancement module on the high-resolution path;
[0125] Figure 7 The lower left image in the image is the ground truth map of the edge of a salient target. Figure 7 The lower middle figure in the image is a binarized version of the feature map of the final output of the high-resolution path; Figure 7 The lower right figure shows the feature map of the final output of the high-resolution path. The comparison results between the above technical solution and the saliency detection methods in this field are shown in Table 1 below. It can be seen that the above technical solution is superior to the saliency detection methods in this field, and the detection accuracy of the above technical solution is higher.
[0126] Table 1 Testing and Evaluation Table
[0127]
[0128]
[0129] Example 5
[0130] Figure 8This is a structural block diagram of the image saliency detection device provided in Embodiment 5 of the present invention. This device is used to execute the image saliency detection method provided in any of the above embodiments. This device and the image saliency detection methods of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the image saliency detection device can be found in the embodiments of the above image saliency detection methods. See also... Figure 8 The device may specifically include: a model acquisition module 510, a feature map acquisition module 520, a path map acquisition module 530, and a result map acquisition module 540.
[0131] The model acquisition module 510 is used to acquire the image to be detected and the trained image saliency detection model. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module.
[0132] The feature map acquisition module 520 is used to input the image to be detected into the first convolution module to obtain a high-resolution feature map, and to process the high-resolution feature map to obtain a low-resolution feature map.
[0133] The path map acquisition module 530 is used to input high-resolution feature maps into the second convolution module to obtain high-resolution path maps, and to input low-resolution feature maps into the third convolution module to obtain low-resolution path maps.
[0134] Result map module 540 is used to fuse high-resolution path maps and low-resolution path maps to obtain a result map representing salient targets detected from the image to be detected;
[0135] Among them, the resolution of the image to be detected, the high-resolution feature map, and the high-resolution path map are consistent, and the resolution of the low-resolution feature map and the low-resolution path map are consistent.
[0136] Optionally, the second convolutional module includes a high-resolution pre-weighted module and a high-resolution post-weighted module, and the high-resolution path further includes a high-resolution weighted module connected between the high-resolution pre-weighted module and the high-resolution post-weighted module to implement weighted connections; the third convolutional module includes a low-resolution pre-weighted module and a low-resolution post-weighted module, and the low-resolution path further includes a low-resolution weighted module connected between the low-resolution pre-weighted module and the low-resolution post-weighted module to implement weighted connections; based on this, the path graph yields module 530, including:
[0137] The feature map input unit is used to input high-resolution feature maps into the high-resolution pre-weighting module and to input low-resolution feature maps into the low-resolution pre-weighting module.
[0138] The high-resolution path graph acquisition unit is used to input the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module into the high-resolution weighting module, and input the output results of the high-resolution weighting module into the high-resolution post-weighting module to obtain the high-resolution path graph;
[0139] The low-resolution path graph acquisition unit is used to input the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module into the low-resolution weighting module, and input the output results of the low-resolution weighting module into the low-resolution post-weighting module to obtain the low-resolution path graph.
[0140] Based on the above scheme, optionally, the output of the high-resolution pre-weighting module is represented through a high-resolution feature channel, and the output of the low-resolution pre-weighting module is represented through a low-resolution feature channel; the high-resolution path graph obtains the unit, including:
[0141] The upsampled feature channel is used to obtain a sub-unit, which is used to upsample the low-resolution feature channel to obtain the upsampled feature channel. The resolution of the upsampled feature channel is the same as that of the high-resolution feature channel.
[0142] The target fusion feature map yields sub-units, which are used to input the high-resolution feature channels and upsampled feature channels into the high-resolution weighting module. The high-resolution weighting module then performs the following steps: concatenates the high-resolution feature channels and upsampled feature channels to obtain an initial fusion feature map; determines a weight vector based on the initial fusion feature map, wherein the length of the weight vector is consistent with the sum of the number of high-resolution feature channels and the number of low-resolution feature channels; and processes the initial fusion feature map based on the weight vector to obtain the target fusion feature map.
[0143] The high-resolution path map is obtained by sub-units, which are used to obtain the target fusion feature map output by the high-resolution weighted module. The target fusion feature map is then input into the high-resolution weighted module to obtain the high-resolution path map.
[0144] Based on the above scheme, optionally, the target fused feature map is used to obtain sub-units, specifically for: performing a global max pooling operation on the initial fused feature map to obtain an intermediate vector; performing a fully connected operation with a ReLU activation function on the intermediate vector; and performing a fully connected operation with a Sigmoid activation function on the fully connected intermediate vector to obtain a weight vector.
[0145] Optionally, the high-resolution path also includes a high-resolution enhancement module connected to the second convolutional module for region enhancement, resulting in module 530, which includes:
[0146] The enhanced feature map acquisition unit is used to input the high-resolution feature map into the second convolutional module to obtain the unenhanced feature map, and then input the unenhanced feature map into the high-resolution enhancement module to obtain the enhanced feature map;
[0147] High-resolution path maps are used as units to convert enhanced feature maps into high-resolution path maps.
[0148] Based on the above scheme, optionally, the enhanced feature map can be used to obtain units, including:
[0149] The pre-enhancement feature map processing subunit is used to input the pre-enhancement feature map into the high-resolution enhancement module, so that the high-resolution enhancement module performs the following steps: obtaining a confidence map corresponding to the pre-enhancement feature map; processing the confidence map to obtain an enhancement map, wherein a portion of all values in the enhancement map are less than or equal to 1, and another portion of values are greater than 1; and processing the pre-enhancement feature map based on the enhancement map.
[0150] Optionally, based on the above-described apparatus, the image saliency detection apparatus may further include:
[0151] The training samples are used as modules to obtain the training images and the target label maps of salient objects in the training images, and the training images and target label maps are used as a set of training samples;
[0152] The original saliency detection model acquisition module is used to acquire the original saliency detection model, wherein the network structure of the original saliency detection model is the same as the network structure of the image saliency detection model to be trained;
[0153] The saliency detection model acquisition module is used to train the original saliency detection model based on multiple sets of training samples to obtain the image saliency detection model.
[0154] Based on the above scheme, optionally, the training samples, as modules, include:
[0155] The edge label map acquisition unit is used to extract edges from the target label map to obtain the edge label map, wherein the edge label map corresponds to the target edges of salient objects in the training image;
[0156] Training samples are used as units to combine training images, target label maps, and edge label maps into a set of training samples;
[0157] The saliency detection model obtains modules, including:
[0158] The edge probability map acquisition unit is used to input the training images in the training samples into the original saliency detection model for each group of training samples, and obtain the edge probability map predicted by the second convolution module and the target probability map predicted by the third convolution module in the original saliency detection model.
[0159] The edge loss calculation unit is used to calculate the target loss based on the target probability map and the target label map, and to calculate the edge loss based on the edge probability map and the edge label map.
[0160] The image saliency detection model training unit is used to obtain the final loss based on the target loss and edge loss, and to adjust the network parameters in the original saliency detection model based on the final loss to train the image saliency detection model.
[0161] Based on the above scheme, the optional saliency detection model acquisition module may also include:
[0162] The fusion probability map is obtained by inputting the training images from the training samples into the original saliency detection model to obtain the fusion probability map, where the fusion probability map is the prediction of the salient targets in the training images by the original saliency detection model.
[0163] The fusion loss calculation unit is used to calculate the fusion loss based on the fusion probability map and the target label map.
[0164] Image saliency detection model training unit, including:
[0165] The final loss is obtained by sub-units, which are used to obtain the final loss based on the target loss, edge loss, and fusion loss.
[0166] The image saliency detection device provided in this embodiment of the invention acquires the image to be detected and a trained image saliency detection model through a model acquisition module. The trained image saliency detection model can achieve high-precision image saliency detection. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together, and the low-resolution path includes a third convolutional module. A feature map acquisition module inputs the image to be detected into the first convolutional module to obtain a high-resolution feature map, and processes the high-resolution feature map to obtain a low-resolution feature map. A path map acquisition module inputs the high-resolution feature map into the second convolutional module. In this process, a high-resolution path map is obtained, and a low-resolution feature map is input into the third convolutional module to obtain a low-resolution path map. The resulting image module then fuses the high-resolution and low-resolution path maps. This allows the high-level semantic information in the low-resolution path map to supplement the low-level semantic information in the high-resolution path map, resulting in a final image representing the salient targets detected in the image to be detected. The resolution of the image to be detected, the high-resolution feature map, and the high-resolution path map are consistent, as are the resolutions of the low-resolution feature map and the low-resolution path map. This allows the model to more easily capture salient targets in the image during image fusion. This technical solution fuses the outputs of the high-resolution and low-resolution paths of the image saliency detection model. While maintaining the high-resolution path's ability to detect image saliency, it introduces feature information obtained from the low-resolution path as auxiliary information, thereby achieving high-precision image saliency detection.
[0167] The image saliency detection device provided in the embodiments of the present invention can execute the image saliency detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0168] It is worth noting that in the embodiments of the above-mentioned image saliency detection device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0169] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0170] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for image saliency detection, the method comprising: include: The image to be detected and the trained image saliency detection model are obtained. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together. The low-resolution path includes a third convolutional module. The number of low-resolution paths is related to the number of modules with convolutional layers in the high-resolution path. The image to be detected is input into the first convolution module to obtain a high-resolution feature map, and the high-resolution feature map is processed to obtain a low-resolution feature map. The high-resolution feature map is input into the second convolutional module to obtain a high-resolution path map, and the low-resolution feature map is input into the third convolutional module to obtain a low-resolution path map. The high-resolution path map and the low-resolution path map are fused to obtain a result map representing the salient targets detected from the image to be detected; Wherein, the image to be detected, the high-resolution feature map, and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution; The second convolutional module includes a high-resolution pre-weighting module and a high-resolution post-weighting module, and the high-resolution path further includes a high-resolution weighting module for implementing weighted connections between the high-resolution pre-weighting module and the high-resolution post-weighting module; The third convolutional module includes a low-resolution pre-weighted module and a low-resolution post-weighted module. The low-resolution path also includes a low-resolution weighted module connected between the low-resolution pre-weighted module and the low-resolution post-weighted module to implement weighted connections. The high-resolution feature map is input into the second convolutional module to obtain a high-resolution path map, including: The high-resolution feature map is input into the high-resolution pre-weighting module, and the low-resolution feature map is input into the low-resolution pre-weighting module; The output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module are input into the high-resolution weighting module, and the output results of the high-resolution weighting module are input into the high-resolution post-weighting module to obtain a high-resolution path map; The low-resolution feature map is input into the third convolutional module to obtain a low-resolution path map, including: The output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module are input into the low-resolution weighting module, and the output results of the low-resolution weighting module are input into the low-resolution post-weighting module to obtain a low-resolution path map.
2. The method of claim 1, wherein, The output of the high-resolution pre-weighting module is represented by a high-resolution feature channel, and the output of the low-resolution pre-weighting module is represented by a low-resolution feature channel. The step of inputting the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module into the high-resolution weighting module includes: The low-resolution feature channel is upsampled to obtain an upsampled feature channel, wherein the resolution of the upsampled feature channel is the same as that of the high-resolution feature channel; The high-resolution feature channel and the upsampled feature channel are input into the high-resolution weighting module to perform the following steps: The high-resolution feature channels are connected to the upsampled feature channels to obtain an initial fused feature map. A weight vector is determined based on the initial fused feature map, wherein the length of the weight vector is the same as the sum of the number of high-resolution feature channels and the number of low-resolution feature channels. The initial fused feature map is processed based on the weight vector to obtain the target fused feature map; The step of inputting the output of the high-resolution weighting module into the high-resolution weighted module to obtain the high-resolution path map includes: The target fusion feature map output by the high-resolution weighting module is obtained, and the target fusion feature map is input into the high-resolution weighted module to obtain a high-resolution path map.
3. The method of claim 2, wherein, The step of determining the weight vector based on the initial fused feature map includes: Perform global max pooling on the initial fused feature map to obtain an intermediate vector; Perform a full connection operation with ReLU activation on the intermediate vector, and then perform a full connection operation with Sigmoid activation on the fully connected intermediate vector to obtain a weight vector.
4. The method of claim 1, wherein, The high-resolution path also includes a high-resolution enhancement module connected to the second convolutional module for region enhancement. The step of inputting the high-resolution feature map into the second convolutional module to obtain the high-resolution path map includes: The high-resolution feature map is input into the second convolutional module to obtain the feature map before enhancement, and the feature map before enhancement is input into the high-resolution enhancement module to obtain the feature map after enhancement; The enhanced feature map is used as a high-resolution path map.
5. The method of claim 4, wherein, The step of inputting the pre-enhancement feature map into the high-resolution enhancement module includes: The pre-enhancement feature map is input into the high-resolution enhancement module to perform the following steps: A confidence map corresponding to the pre-enhancement feature map is obtained; The confidence graph is processed to obtain an augmentation graph, wherein a portion of the values in the augmentation graph are less than or equal to 1, and another portion of the values are greater than 1; The feature map before enhancement is processed based on the enhancement map.
6. The method of claim 1, wherein, The image saliency detection model is pre-trained through the following steps: Obtain training images and target label maps of salient objects in the training images, and use the training images and target label maps as a set of training samples; Obtain the original saliency detection model, wherein the network structure of the original saliency detection model is the same as the network structure of the image saliency detection model to be trained; The original saliency detection model is trained based on multiple sets of training samples to obtain the image saliency detection model.
7. The method of claim 6, wherein, The step of using the training image and the target label image as a set of training samples includes: Edge extraction is performed on the target label map to obtain an edge label map, wherein the edge label map corresponds to the target edge of a salient target in the training image; The training image, the target label map, and the edge label map are used as a set of training samples; The step of training the original saliency detection model based on multiple sets of training samples to obtain the image saliency detection model includes: For each set of training samples, the training image in the training samples is input into the original saliency detection model to obtain the edge probability map predicted by the second convolution module and the target probability map predicted by the third convolution module in the original saliency detection model. The target loss is calculated based on the target probability map and the target label map, and the edge loss is calculated based on the edge probability map and the edge label map; The final loss is obtained based on the target loss and the edge loss, and the network parameters in the original saliency detection model are adjusted based on the final loss to train the image saliency detection model.
8. The method of claim 7, wherein, After inputting the training images from the training samples into the original saliency detection model, the method further includes: A fusion probability map is obtained, wherein the fusion probability map is predicted by the original saliency detection model for salient targets in the training image; The fusion loss is calculated based on the fusion probability map and the target label map; The process of obtaining the final loss based on the target loss and the edge loss includes: The final loss is obtained based on the target loss, the edge loss, and the fusion loss.
9. An image saliency detection apparatus, characterized by comprising: include: The model acquisition module is used to acquire the image to be detected and the trained image saliency detection model. The image saliency detection model includes a high-resolution path and a low-resolution path. The high-resolution path includes a first convolutional module and a second convolutional module connected together. The low-resolution path includes a third convolutional module. The number of low-resolution paths is related to the number of modules with convolutional layers in the high-resolution path. The feature map acquisition module is used to input the image to be detected into the first convolution module to obtain a high-resolution feature map, and to process the high-resolution feature map to obtain a low-resolution feature map. The path map acquisition module is used to input the high-resolution feature map into the second convolution module to obtain a high-resolution path map, and to input the low-resolution feature map into the third convolution module to obtain a low-resolution path map; The result map acquisition module is used to fuse the high-resolution path map and the low-resolution path map to obtain a result map representing the salient targets detected from the image to be detected; Wherein, the image to be detected, the high-resolution feature map, and the high-resolution path map have the same resolution, and the low-resolution feature map and the low-resolution path map have the same resolution; The second convolutional module includes a high-resolution pre-weighted module and a high-resolution post-weighted module. The high-resolution path further includes a high-resolution weighted module connecting the high-resolution pre-weighted module and the high-resolution post-weighted module to implement weighted connections. The third convolutional module includes a low-resolution pre-weighted module and a low-resolution post-weighted module. The low-resolution path further includes a low-resolution weighted module connecting the low-resolution pre-weighted module and the low-resolution post-weighted module to implement weighted connections. The path graph generation module includes: A feature map input unit is used to input the high-resolution feature map into the high-resolution pre-weighting module and to input the low-resolution feature map into the low-resolution pre-weighting module; The high-resolution path graph obtaining unit is used to input the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module into the high-resolution weighting module, and input the output results of the high-resolution weighting module into the high-resolution post-weighting module to obtain a high-resolution path graph; The low-resolution path graph obtaining unit is used to input the output results of the high-resolution pre-weighting module and the low-resolution pre-weighting module into the low-resolution weighting module, and input the output results of the low-resolution weighting module into the low-resolution post-weighting module to obtain the low-resolution path graph.