Salient target detection method and system, electronic equipment and storage medium

By combining a dual-branch encoder and a cross-model interactive fusion module with an edge-aware enhancement module, the problems of blur and unclear edges in complex scenes in salient object detection are solved, achieving more accurate salient object detection.

CN120833467AActive Publication Date: 2025-10-24WENZHOU ELECTRIC POWER DESIGN CO LTD PUHUA TENDERING CONSULTING BRANCH

Patent Information

Application Number
CN202510860185.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-24
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Existing salient object detection methods have difficulty accurately capturing the irregular topological structure and cluttered background of salient objects when processing complex scenes, resulting in blurred prediction results, unclear segmentation edges, and limited depth that cannot capture long-distance dependencies.

Method used

A dual-branch encoder is used to extract local detail features and global semantic features. Feature fusion and enhancement are performed through the cross-model interaction fusion module and the edge-aware enhancement module. Combined with the decoder processing, full interaction and integration of local and global information is achieved, and edge detail information is enhanced.

Benefits of technology

The accuracy and robustness of salient object detection are improved, especially in complex scenes, which can generate better results, with clearer segmentation edges and improved detection capabilities for small objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833467A_ABST
    Figure CN120833467A_ABST
Patent Text Reader

Abstract

The invention discloses a saliency target detection method and system, electronic equipment and a storage medium, and relates to the technical field of saliency target detection.In the method, a constructed saliency target detection model is trained based on a plurality of sample images, and in the constructed saliency target detection model, a saliency target detection result is obtained. A plurality of local detail features and a plurality of global semantic features of a target sample image are extracted through a double-branch encoder, and after part of the local detail features and all the local detail features are interactively fused, a plurality of enhanced local detail features and a plurality of enhanced global detail features are obtained; performing feature splicing on each enhanced local detail feature and the corresponding enhanced global detail feature, combining all the spliced features and the remaining local detail features, and performing processing through an edge perception enhancement module and a decoder to obtain a saliency target detection result; and performing accurate saliency target detection on the target image by using the trained saliency target detection model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of salient object detection, and in particular to a salient object detection method, system, electronic device and storage medium. BACKGROUND

[0002] The salient object detection (SOD) task aims to mimic human visual perception and automatically identify and segment the most eye-catching objects or regions from natural images. The "salience" of an object can be reflected in shape, size, color, and spatial location, etc. Due to its powerful function, salient object detection has wide application value in the fields of autonomous driving, image retrieval, video segmentation, image cropping, semantic segmentation, and object recognition, etc.

[0003] Early salient object detection methods use hand-crafted features, i.e. using hand-crafted color and texture to mine contrastive regions. However, these methods cannot extract high-level semantic information and perform poorly in some complex scenes. The development of deep learning has greatly broken the bottleneck of hand-crafted methods and will continue to make steady progress.

[0004] Deep learning-based salient object detection methods mainly rely on deep neural networks to extract discriminative features and divide salient regions with high contrast in the surrounding environment. Although existing methods have made significant breakthroughs, there are still challenges such as irregular topological structure of salient objects and cluttered background. Specifically, salient objects may exhibit complex geometric structures such as non-rigid deformation, internal holes, branching, or adhesion, as well as small color or texture differences at the junction between salient objects and background. This will lead to blurred prediction results and unclear segmentation edges.

[0005] Convolutional neural network-based methods can extract discriminative features through multiple layers of convolution. However, due to their limited depth, they cannot capture long-range dependencies and produce good results in complex scenes, and they also destroy the integrity of objects. Transformer-based methods can extract global context semantic information. However, due to the global self-attention mechanism, they cannot obtain sufficient local detail information, and the prediction result may be blurred when dealing with small objects. SUMMARY

[0006] The technical problem to be solved by the present application is to overcome the deficiencies of the prior art, and specifically provides a salient object detection method, system, electronic device and storage medium, as follows: 1) In a first aspect, the present application provides a salient object detection method, and the specific technical solutions are as follows: The constructed salient object detection model is trained based on a plurality of sample images to obtain a trained salient object detection model, wherein the constructed salient object detection model comprises a double-branch encoder, a cross-model interaction fusion module, an edge perception enhancement module and a decoder, a plurality of local detail features and a plurality of global semantic features of a target sample image are extracted through the double-branch encoder, a plurality of enhanced local detail features and a plurality of enhanced global detail features are obtained by respectively passing part of the local detail features and all the global semantic features through a cross-model interaction fusion module according to a preset corresponding relationship, feature splicing is performed on each enhanced local detail feature and the corresponding enhanced global detail feature to obtain a plurality of spliced features, and a salient object detection result is obtained by combining all the spliced features and the remaining part of the local detail features and processing them through the edge perception enhancement module and the decoder, wherein the target sample image is any target sample image. The trained salient object detection model is used for salient object detection on a target image.

[0007] The salient object detection method provided by the application has the following advantages: The double-branch encoder can extract local detail features and global semantic features respectively, the local detail features are helpful for capturing complex geometric structures such as non-rigid deformation and internal cavities, the global semantic features are conducive to grasping the overall scene and resisting cluttered background interference, and the cross-model interaction fusion module further enhances these features, so that the salient object detection model can more accurately identify the target and avoid fuzzy prediction results, and can effectively cope with the irregular topological structure of the salient object and the cluttered background challenge; the salient object detection model realizes sufficient interaction and integration of local and global information through the cross-model interaction fusion and subsequent processing module, breaks the depth limit, better understands the scene from a global perspective, and at the same time takes into account the local information, guarantees the integrity of the object, and can generate better results even in complex scenes, solving the problem that the limited depth cannot capture long-distance dependency relationships; the edge perception enhancement module can strengthen edge detail information, and in combination with the local detail feature extraction of the double-branch encoder, can accurately capture the edge of the salient object, even in the case of small color or texture difference at the junction of the target and the background, the segmentation edge can be made clearer, and the overall effect of salient object detection is improved, especially when processing small targets, the prediction result can be effectively avoided.

[0008] On the basis of the above-mentioned scheme, the salient object detection method of the application can be further improved as follows.

[0009] Further, the first branch of the double-branch encoder comprises N first network layers arranged in sequence, and the second branch of the double-branch encoder comprises M second network layers arranged in sequence. The plurality of local detail features and the plurality of global semantic features of the target sample image are extracted by the double-branch encoder, and part of the local detail features and all the global semantic features are interactively fused by a cross-model interactive fusion module according to a preset corresponding relationship to obtain a plurality of enhanced local detail features and a plurality of enhanced global detail features. Each enhanced local detail feature and the corresponding enhanced global detail feature are spliced to obtain a plurality of spliced features, including: After the target sample image is input into the first first network layer, the local detail features output by the first first network layer to the nth first network layer are obtained. The local detail features output by the nth-1 first network layer are input into the first second network layer to obtain the global semantic features output by the first second network layer, and the cross-model interactive fusion module is used to interactively fuse the local detail features output by the nth first network layer and the global semantic features output by the first second network layer, so that the local detail information in the local detail features output by the nth first network layer is integrated into the global semantic features output by the first second network layer to obtain the first enhanced global semantic features, and the global context information in the global semantic features output by the first second network layer is integrated into the local detail features output by the nth first network layer to obtain the first enhanced local detail features. After the first enhanced global semantic features and the first enhanced local detail features are spliced, the first spliced feature is obtained. The first enhanced global semantic features are used as the input of the second second network layer, the first enhanced local detail features are used as the input of the nth+1 first network layer, and the output of the second second network layer and the output of the nth+1 first network layer are interactively fused by the second cross-model interactive fusion module to obtain the second enhanced global semantic features and the first enhanced local detail features, and the second spliced feature is obtained by splicing. Until the Mth spliced feature is obtained, wherein n and M are positive integers, and N=M+n-1.

[0010] The beneficial effects of the above further scheme are: through the multi-layer network structure of the double-branch encoder, the local detail features and the global semantic features of the image can be systematically extracted from different levels to ensure the richness and integrity of the features; the design of the cross-model interactive fusion module realizes the dynamic interactive fusion of the local detail features and the global semantic features. This interactive mechanism not only enhances the semantic understanding ability of the local features, but also improves the detail accuracy of the global features, so that the model can more accurately capture key information when processing complex targets; the feature splicing operation organically combines the enhanced local and global features to provide more comprehensive feature representation for subsequent processing, which improves the accuracy and robustness of the saliency target detection, especially when processing targets with complex structures.

[0011] Further, after combining all the splicing features and the local detail features of the remaining part, and processing by the edge perception enhancement module and the decoder, the salient object detection result is obtained, including: Each splicing feature is processed by a convolution layer to obtain a convolution result corresponding to each splicing feature; The convolution result corresponding to the Mth splicing feature is input into an ASPP module to obtain multi-scale features, the multi-scale features and the convolution result corresponding to the Mth splicing feature are input into a first decoder, and the output of the first decoder and the convolution result corresponding to the Mth splicing feature are input into a first edge perception enhancement module to enhance the edge information of the salient object by the first edge perception enhancement module, the output of the M-1th splicing feature and the first edge perception enhancement module are input into a second decoder, and the output of the second decoder and the convolution result corresponding to the M-1th splicing feature are input into a second edge perception enhancement module, until the output of the mth edge perception enhancement module is obtained, where m is a positive integer, m=M-n+1; The output of the mth edge perception enhancement module and the local detail feature output by the n-1th first network layer are input into an m+1th decoder, the output of the m+1th decoder and the local detail feature output by the n-1th first network layer are input into an m+1th edge perception enhancement module, the output of the m+1th edge perception enhancement module and the local detail feature output by the n-2th first network layer are input into an n+2th decoder, until the output of the Nth decoder is obtained, and the output of the Nth decoder is taken as the salient object detection result.

[0012] The beneficial effects of the above further scheme are: by processing each splicing feature through a convolution layer and inputting it into an ASPP (Atrous Spatial Pyramid Pooling) module, the technology can capture multi-scale information. The ASPP module uses convolution kernels with different dilation rates to extract features in parallel, effectively solving the problem of target size variation and enhancing the detection ability of different size salient objects; the features processed by the decoder are input into the edge perception enhancement module, which significantly improves the detection accuracy of the target edge. This solves the problem of blurred and unclear edges in the prior art, making the boundary of the detection result more accurate and detailed; the cascaded structure design of the decoder enables the features to integrate information from different levels during the decoding process, and combines the local detail features of the remaining part to comprehensively improve the model's ability to grasp details, making the detection result more complete and detailed.

[0013] Further, the cross-model interaction fusion module is specifically used for: After receiving the local detail feature and the global semantic feature, the received local detail feature is sequentially processed by the first channel attention module and the first spatial attention module to obtain a first feature, and the received local detail feature and the received global semantic feature are processed by the first cross-model spatial attention module to obtain a second feature, and the first feature and the second feature are element-wise added to obtain an enhanced local detail feature; The received global detail feature is sequentially processed by the second channel attention module and the second spatial attention module to obtain a third feature, and the received local detail feature and the received global semantic feature are processed by the second cross-model spatial attention module to obtain a fourth feature, and the third feature and the fourth feature are element-wise added to obtain an enhanced global semantic feature.

[0014] The beneficial effects of the above further scheme are: through the combination of channel attention and spatial attention modules, the local detail feature is processed in multiple dimensions, further extracting and enhancing key information, and improving the discriminability of the feature; using the cross-model spatial attention mechanism, combining the local detail feature and the global semantic feature, the effective fusion of the two is realized, and the semantic understanding of the local feature and the detail accuracy of the global feature are enhanced; through the element-wise addition method, the features obtained by different paths are fused to realize feature complementation, strengthen the key features of the salient target, and improve the accuracy of detection; the enhanced local detail feature and the global semantic feature more accurately represent the salient target, improve the performance of the detection model in complex scenes, and the effect is significant especially when processing targets with complex structure and background.

[0015] Further, the edge perception enhancement module is specifically configured to: obtain a boundary map in the received convolution result or the local detail feature, calculate a Hadamard product of the boundary map and an output of the received decoder, and process the Hadamard product through a convolution layer to obtain a convolution result corresponding to the Hadamard product, and element-wise add the convolution result corresponding to the Hadamard product and the output of the received decoder to obtain an output of the edge perception enhancement module.

[0016] The beneficial effects of the above further scheme are: by extracting the boundary map in the convolution result and calculating the Hadamard product, focusing on the edge information of the target, strengthening the boundary features of the salient target, and making the edge more clear and explicit; the result of the Hadamard product after convolution processing is element-wise added to the decoder output to realize the fusion of multi-level features, optimize the feature expression, and improve the model's ability to grasp the target as a whole and details; the enhanced edge information helps the salient target detection result to more accurately fit the actual target boundary, improves the accuracy and robustness of detection, and the effect is significant especially when processing complex background and edge blurred targets.

[0017] Further, the method further comprises: in the training of the constructed salient object detection model, a first loss function is used to supervise the training of the boundary map, and a second loss function is used to supervise the training of the salient object detection result.

[0018] The beneficial effect of the above further scheme is that the first loss function is used to supervise the boundary map, and the second loss function is used to supervise the salient detection result, so that the salient object detection model is optimized in multiple aspects. The boundary precision can be improved, the salient object boundary is clearer and more accurate, the accuracy and robustness of target detection are improved, the detection capability of the salient object detection model in a complex scene is improved, and the generalization performance of the salient object detection model is improved.

[0019] 2) In a second aspect, the present application also provides a salient object detection system, and the specific technical scheme is as follows: The system comprises a model training module and a salient object detection module. The model training module is configured to train the constructed salient object detection model based on a plurality of sample images to obtain a trained salient object detection model, wherein the constructed salient object detection model comprises a double-branch encoder, a cross-model interaction fusion module, an edge perception enhancement module, and a decoder. The double-branch encoder is configured to extract a plurality of local detail features and a plurality of global semantic features of a target sample image. Part of the local detail features and all the global semantic features are respectively subjected to interaction fusion through a cross-model interaction fusion module according to a preset corresponding relationship, so as to obtain a plurality of enhanced local detail features and a plurality of enhanced global detail features. Each enhanced local detail feature and the corresponding enhanced global detail feature are subjected to feature splicing, so as to obtain a plurality of spliced features. The salient object detection result is obtained by combining all the spliced features and the remaining part of the local detail features, and processing the combined features through the edge perception enhancement module and the decoder. The salient object detection module is configured to perform salient object detection on a target image by using the trained salient object detection model.

[0020] On the basis of the above scheme, the salient object detection system of the present application can be further improved as follows.

[0021] Further, the first branch of the double-branch encoder comprises N first network layers arranged in sequence, and the second branch of the double-branch encoder comprises M second network layers arranged in sequence. The target sample image is input into the first network layer, and the local detail features output by the first network layer and the n-th network layer are obtained. The local detail feature output by the n-1th first network layer is input into the 1st second network layer to obtain global semantic features output by the 1st second network layer, and the local detail feature output by the n th first network layer and the global semantic features output by the 1st second network layer are interactively fused through the 1st cross-model interactive fusion module, so that the local detail information in the local detail feature output by the n th first network layer is integrated into the global semantic features output by the 1st second network layer to obtain the 1st enhanced global semantic features, and the global context information in the global semantic features output by the 1st second network layer is integrated into the local detail feature output by the n th first network layer to obtain the 1st enhanced local detail feature. After the 1st enhanced global semantic features and the 1st enhanced local detail feature are spliced, the 1st spliced feature is obtained. The 1st enhanced global semantic features are taken as the input of the 2nd second network layer, and the 1st enhanced local detail feature is taken as the input of the n+1th first network layer. The outputs of the 2nd second network layer and the n+1th first network layer are interactively fused through the 2nd cross-model interactive fusion module to obtain the 2nd enhanced global semantic features and the 1st enhanced local detail feature, and then the 2nd spliced feature is obtained by splicing. Until the Mth spliced feature is obtained, wherein n and M are positive integers, N=M+n-1.

[0022] Further, each spliced feature is processed through a convolution layer to obtain a convolution result corresponding to each spliced feature; the convolution result corresponding to the Mth spliced feature is input into an ASPP module to obtain multi-scale features, and the multi-scale features and the convolution result corresponding to the Mth spliced feature are input into a 1st decoder, and the output of the 1st decoder and the convolution result corresponding to the Mth spliced feature are input into a 1st edge perception enhancement module to enhance the edge information of the salient target through the 1st edge perception enhancement module. The output of the 1st edge perception enhancement module and the M-1th spliced feature are input into a 2nd decoder, and the output of the 2nd decoder and the convolution result corresponding to the M-1th spliced feature are input into a 2nd edge perception enhancement module, until the output of the mth edge perception enhancement module is obtained, wherein m is a positive integer, m=M-n+1; the output of the mth edge perception enhancement module and the local detail feature output by the n-1th first network layer are input into an m+1th decoder, the output of the m+1th decoder and the local detail feature output by the n-1th first network layer are input into an m+1th edge perception enhancement module, and the output of the m+1th edge perception enhancement module and the local detail feature output by the n-2th first network layer are input into an n+2th decoder, until the output of the Nth decoder is obtained, which is taken as the salient target detection result.

[0023] Further, the cross-model interaction fusion module is specifically configured to: After receiving the local detail feature and the global semantic feature, the received local detail feature is sequentially processed by the first channel attention module and the first spatial attention module to obtain a first feature, and the received local detail feature and the received global semantic feature are processed by the first cross-model spatial attention module to obtain a second feature, and the first feature and the second feature are element-wise added to obtain an enhanced local detail feature. The received global detail feature is sequentially processed by the second channel attention module and the second spatial attention module to obtain a third feature, and the received local detail feature and the received global semantic feature are processed by the second cross-model spatial attention module to obtain a fourth feature, and the third feature and the fourth feature are element-wise added to obtain an enhanced global semantic feature.

[0024] Further, the edge-aware enhancement module is specifically configured to: obtain a boundary map in the received convolution result or the local detail feature, calculate a Hadamard product of the boundary map and the output of the received decoder, process the Hadamard product through a convolution layer to obtain a convolution result corresponding to the Hadamard product, and element-wise add the convolution result corresponding to the Hadamard product and the output of the received decoder to obtain an output of the edge-aware enhancement module.

[0025] Further, the model training module is further configured to: in training the constructed saliency target detection model, a first loss function is used to train and supervise the boundary map, and a second loss function is used to train and supervise the saliency target detection result.

[0026] 3) In a third aspect, the present application also provides an electronic device, which comprises a processor and a memory coupled with the processor, and the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to enable the electronic device to implement any of the above saliency target detection methods.

[0027] 4) In a fourth aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement any of the above saliency target detection methods.

[0028] It should be noted that the technical solutions of the second to fourth aspects of the present application and the corresponding possible implementation manners have the beneficial effects as described above for the first aspect and the corresponding possible implementation manners, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced as follows: Figure 1 A flowchart of a salient object detection method according to an embodiment of the present application; Figure 2 A network structure diagram of a salient object detection model; Figure 3 A network structure diagram of a cross-model interaction fusion module; Figure 4 A network structure diagram of an edge perception enhancement module; Figure 5 An introduction diagram of an international standard data set for evaluating the technical effects of the present application; Figure 6 An introduction diagram of evaluation indexes for evaluating the technical effects of the present application; Figure 7 An introduction diagram of a comparative model; Figure 8 A comparative diagram of salient object detection results (1); Figure 9 A comparative diagram of salient object detection results (2); Figure 10 A structure diagram of a salient object detection system according to an embodiment of the present application; Figure 11 A structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] The principles and features of the present application are described below, and the examples are only used to explain the present application, and are not used to limit the scope of the present application.

[0031] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0032] As shown in Figure 1 A salient object detection method according to an embodiment of the present application includes the following steps: S1, based on a plurality of sample images, a constructed salient target detection model is trained to obtain a trained salient target detection model, wherein the constructed salient target detection model comprises: a double-branch encoder, a cross-model interaction fusion module, an edge perception enhancement module and a decoder, a plurality of local detail features and a plurality of global semantic features of a target sample image are extracted through the double-branch encoder, part of the local detail features and all the global semantic features are respectively subjected to interaction fusion through a cross-model interaction fusion module according to a preset corresponding relationship, a plurality of enhanced local detail features and a plurality of enhanced global detail features are obtained, each enhanced local detail feature and the corresponding enhanced global detail feature are spliced to obtain a plurality of spliced features, and the salient target detection result is obtained by combining all the spliced features and the remaining part of the local detail features through the edge perception enhancement module and the decoder, wherein the target sample image is any target sample image; S2, using the trained salient target detection model to perform salient target detection on a target image.

[0033] Wherein, the target image is an image in the field of image video compression technology, image content understanding technology, automatic driving technology, visual tracking technology or image retrieval technology, or other fields according to actual conditions, specifically: 1) In the field of image video compression technology, the target image can be an image containing natural scenery, including blue sky, green trees, grassland and distant mountains. The detected green trees, blue-green grassland and distant mountains with clear outlines are taken as the salient target detection result. These areas are usually the focus of human eyes and contain rich texture and color information. In the image video compression process, the quality of these salient target areas is preferentially guaranteed. For non-salient areas, such as the relatively single-color area of the sky, the resolution can be appropriately reduced or more significantly compressed, thereby reducing the data volume, improving the compression efficiency, saving the storage space and transmission bandwidth without affecting the viewer's perception of the main landscape.

[0034] 2) In the field of image content understanding, the target image is a city street view image containing pedestrians, vehicles, buildings, roads, and other elements. The detected pedestrians on the street and vehicles on the road are considered as salient target detection results. Pedestrians may stand out due to their bright clothing or unique movements, while vehicles become the focus due to their larger size and movement on the road. This helps to quickly understand the semantic content of the image. For example, in an intelligent security system, by detecting the saliency of pedestrians and vehicles in the street view, real-time analysis of traffic flow and pedestrian behavior patterns can be performed. Further, by combining other technologies such as pedestrian re-identification, specific individuals can be tracked, or traffic violations such as speeding can be monitored.

[0035] 3) In the field of autonomous driving technology, the target image is a road condition image in front of the vehicle, including vehicles on other lanes, traffic signs, road markings, etc. The detected vehicles on the same lane in front of the vehicle, especially those close to the vehicle, are considered as salient target detection results. At the same time, traffic signs on both sides of the road (such as speed limit signs, turning signs, etc.) are also detected. For autonomous vehicles, identifying vehicles in front is crucial for maintaining a safe distance and braking in time. Salient target detection can help vehicles quickly and accurately locate the position and state of surrounding vehicles. The detection of traffic signs can provide information on driving rules for vehicles, such as adjusting speed according to speed limit signs and preparing to turn in advance according to turning signs, thereby improving the safety and reliability of autonomous driving.

[0036] 4) In the field of visual tracking technology, the target image is a sequence of basketball game video images captured during a sports match. The images contain basketball players, basketballs, referees, and spectators. In these image sequences, the basketball itself and the player holding the ball are detected as salient target detection results. The basketball stands out due to its rapid movement and distinctive color (usually orange), while the player holding the ball becomes the focus due to the central role of their actions (such as dribbling, passing, and shooting). In video analysis of sports events, by detecting the saliency of the basketball and the player holding the ball, visual tracking can be more effective. For example, automatically tracking the trajectory of the basketball and analyzing the dribbling, passing, and shooting actions of the player can provide deeper game analysis data for coaches and spectators. In multi-target tracking scenarios, prioritizing the tracking of these salient targets can improve the accuracy and efficiency of tracking.

[0037] 5) In the field of image retrieval technology, the target image is a user-uploaded image containing flowers, including blooming flowers, flower stems, and green leaves. The detected flowers in the target image are the salient object detection results. Flowers usually have bright colors and unique shapes, which can easily attract people's attention in the target image. In the image retrieval system, by extracting the features of salient objects (such as flowers), more accurate image matching can be performed. When the user wants to find similar flower pictures, the system will search the database according to the salient features of the flowers (color, shape, etc.), improving the accuracy and relevance of the search, and providing more accurate results for online search for specific types of flower pictures or identification of flower species.

[0038] In the field of salient object detection, it is crucial to build customized datasets according to the needs of different application scenarios. For example, in medical image analysis, a large number of images labeled with lesion areas can be collected to train a model that can accurately identify lesions to assist doctors in diagnosis. In the field of autonomous driving, an image dataset containing various road conditions and obstacles can be established to train a model that focuses on detecting key targets such as pedestrians and vehicles, improving driving safety. This targeted model building approach can make salient object detection technology more accurate in serving specific fields, achieving deep integration of technology and practical application, and promoting the intelligent development of various industries.

[0039] Optionally, in the above technical solution, the first branch of the double-branch encoder includes N first network layers arranged in sequence, and the second branch of the double-branch encoder includes M second network layers arranged in sequence. Wherein, N can be 2, 3, 4 or 5, etc., which can be set according to actual conditions, and M can be 1, 2, 3 or 4, etc.

[0040] Wherein, the network structure for extracting local detail features constitutes the first branch, and the network structure for extracting global detail features constitutes the second branch. The network structure for extracting local detail features can be a convolutional neural network, a DenseNet model, or a U-Net model. The network structure for extracting global detail features can be a Transformer, a Dense Attention network, or an LSTM model. The network structure for extracting global detail features and the network structure for extracting local detail features can also be selected according to actual conditions.

[0041] When the network structure for extracting local detail features is a convolutional neural network, a VGG-16 network (the VGG-16 network belongs to the convolutional neural network) can be specifically used, and the first network layers of the VGG-16 network are all convolutional blocks (VggBlock). When the network structure for extracting global detail features can be a Transformer, a Swin-Transformer (the Swin-Transformer belongs to the Swin-Transformer) can be specifically used, and the second network layers of the Swin-Transformer are all SwinBlocks.

[0042] The plurality of local detail features and the plurality of global semantic features of the target sample image are extracted by the double-branch encoder, and part of the local detail features and all the global semantic features are respectively interacted and fused by a preset corresponding relationship through a cross-model interaction fusion module to obtain a plurality of enhanced local detail features and a plurality of enhanced global detail features. Each enhanced local detail feature and the corresponding enhanced global detail feature are spliced to obtain a plurality of spliced features, including: S10, inputting the target sample image into the first network layer to obtain the local detail features output by the first network layer to the nth network layer; S11, inputting the local detail features output by the (n-1)th network layer into the first second network layer to obtain the global semantic features output by the first second network layer, and interacting and fusing the local detail features output by the nth network layer and the global semantic features output by the first second network layer through the first cross-model interaction fusion module to make the local detail information in the local detail features output by the nth network layer into the global semantic features output by the first second network layer, obtain the first enhanced global semantic features, and make the global context information in the global semantic features output by the first second network layer into the local detail features output by the nth network layer, obtain the first enhanced local detail features, take the first enhanced global semantic features as the input of the second second network layer, take the first enhanced local detail features as the input of the (n+1)th network layer, and interact and fuse the output of the second second network layer and the output of the (n+1)th network layer through the second cross-model interaction fusion module to obtain the second enhanced global semantic features and the first enhanced local detail features, and perform feature splicing to obtain the second spliced feature, until the Mth spliced feature is obtained, wherein n and M are positive integers, and N=M+n-1.

[0043] As shown in Figure 2 Taking “N=5, M=3, n=3” as an example, the process of “obtaining a plurality of spliced features” is described as follows:

[0044]

[0045] ​​​​​​​​​​​​​​​​​​​​​​​​As the 4th first network layer VggBlock-4, the 1st enhanced global semantic feature As the input of the 2nd second network layer SwinBlock-2, the 4th first network layer VggBlock-4 outputs the 4th local detail feature , the 2nd second network layer SwinBlock-2 outputs the 2nd global semantic feature , the 4th local detail feature and the 2nd global semantic feature are input into the 2nd cross-model interaction fusion module CMIM-2 to obtain the 2nd enhanced local detail feature and the 2nd enhanced global semantic feature , the 2nd enhanced local detail feature and the 2nd enhanced global semantic feature are spliced to obtain the 2nd spliced feature, wherein the size of the 4th local detail feature is 28x28x256, and the size of the 2nd global semantic feature is 28x28x512.

[0046] ④ the 2nd enhanced local detail feature As the 5th first network layer VggBlock-5, the 2nd enhanced global semantic feature As the input of the 3rd second network layer SwinBlock-3, the 5th first network layer VggBlock-5 outputs the 5th local detail feature , the 3rd second network layer SwinBlock-3 outputs the 3rd global semantic feature , the 5th local detail feature and the 3rd global semantic feature are input into the 3rd cross-model interaction fusion module CMIM-3 to obtain the 3rd enhanced local detail feature and the 3rd enhanced global semantic feature , the 3rd enhanced local detail feature and the 3rd enhanced global semantic feature are spliced to obtain the 3rd spliced feature, wherein the size of the 5th local detail feature is 14x14x256, and the size of the 3rd global semantic feature is 14x14x1024.

[0047] Optionally, in the above technical solution, after the combination of all the splicing features and the local detail features of the remaining part is processed by the edge perception enhancement module and the decoder, a salient object detection result is obtained, including: S12, each splicing feature is processed by a convolution layer to obtain a convolution result corresponding to each splicing feature, specifically: The first splicing feature is processed by a convolution layer (conv), and the obtained convolution result is denoted as The second splicing feature is processed by a convolution layer (conv), and the obtained convolution result is denoted as The third splicing feature is processed by a convolution layer (conv), and the obtained convolution result is denoted as .

[0048] The above feature splicing process and the process of processing each splicing feature by a convolution layer can be represented by the following formula: denotes the feature splicing of and , wherein denotes the i-th enhanced local detail feature, denotes the i-th enhanced global semantic feature, i is a positive integer, and i≤M. denotes that the convolution operation is performed on , and specifically convolution operation can be performed.

[0049] S13, the convolution result corresponding to the Mth splicing feature is input into the ASPP module to obtain multi-scale features, the multi-scale features and the convolution result corresponding to the Mth splicing feature are input into the first decoder, and the output of the first decoder and the convolution result corresponding to the Mth splicing feature are input into the first edge perception enhancement module to enhance the edge information of the salient object through the first edge perception enhancement module. The output of the first edge perception enhancement module and the M-1th splicing feature are input into the second decoder, and the output of the second decoder and the convolution result corresponding to the M-1th splicing feature are input into the second edge perception enhancement module, until the output of the mth edge perception enhancement module is obtained, wherein m is a positive integer, m=M-n+1, at this time m=5-3+1=3, so: ① the convolution result corresponding to the third splicing feature is input into the ASPP module to obtain multi-scale features; the multi-scale features and the convolution result corresponding to the third splicing feature are input into the first decoder Decoder1, and the first decoder Decoder1 processes the multi-scale features and the convolution result corresponding to the third splicing feature After processing, the output of the first decoder Decoder1 is obtained The output of the first decoder Decoder1 corresponding to the third splicing feature is input into the first edge-aware enhancement module EEM1 to obtain the output of the first edge-aware enhancement module EEM1.

[0050] ②The convolution result corresponding to the second splicing feature and the output of the first edge-aware enhancement module EEM1 are input into the second decoder Decoder2 to obtain the output of the second decoder Decoder2 The convolution result corresponding to the second splicing feature and the output of the second decoder Decoder2 are input into the second edge-aware enhancement module EEM2 to obtain the output of the second edge-aware enhancement module EEM2.

[0051] ③The convolution result corresponding to the first splicing feature and the output of the second edge-aware enhancement module EEM2 are input into the third decoder Decoder3 to obtain the output of the third decoder Decoder3 The convolution result corresponding to the first splicing feature and the output of the third decoder Decoder3 are input into the third edge-aware enhancement module EEM3 to obtain the output of the third edge-aware enhancement module EEM3.

[0052] S14, the output of the mth edge-aware enhancement module and the local detail feature output by the n-1th first network layer are input into the m+1th decoder, the output of the m+1th decoder and the local detail feature output by the n-1th first network layer are input into the m+1th edge-aware enhancement module, the output of the m+1th edge-aware enhancement module and the local detail feature output by the n-2th first network layer are input into the n+2th decoder, until the output of the Nth decoder is obtained, and the output of the Nth decoder is taken as the saliency target detection result, specifically: The output of the third edge-aware enhancement module EEM3 and the local detail feature output by the second first network layer are input into the fourth decoder Decoder4 to obtain the output of the fourth decoder Decoder4 The output of the fourth decoder Decoder4 and the local detail feature output by the second first network layer ​​​The fourth edge-aware enhancement module is inputted, and the output of the fourth edge-aware enhancement module and the local detail feature output by the first network layer are inputted The fifth decoder Decoder5, and the output of the fifth decoder is taken as the saliency target detection result.

[0053] Optionally, in the above technical solution, as shown in Figure 3 The cross-model interaction fusion module is specifically used for: S31, when receiving the local detail feature and the global semantic feature, the received local detail feature is sequentially processed by the first channel attention module and the first spatial attention module to obtain a first feature, the received local detail feature and the received global semantic feature are processed by the first cross-model spatial attention module to obtain a second feature, and the first feature and the second feature are element-wise added to obtain an enhanced local detail feature, specifically: S310, when receiving the local detail feature and the global semantic feature , the received local detail feature is inputted into the first channel attention module CA-1, the first channel attention module CA-1 performs a series of pooling and convolution operations on the received local detail feature to obtain channel attention weights , and the channel information in the local detail feature is enhanced by using the channel attention weights to obtain a first intermediate feature , which can be specifically represented by the following formula: wherein, represents a function, represents a global average pooling operation on , represents a global maximum pooling operation on , and represents: a convolution operation on represents a Hadamard product symbol, , is a positive integer, .

[0054] S311, after the first intermediate feature is processed by the first spatial attention module SA-1, a first feature is obtained, specifically, the first spatial attention module SA-1 obtains​​​​ the spatial attention weight , and the spatial attention weight enhances the important spatial position information in the first intermediate feature , which can be represented by the following formula: wherein, indicates that the first intermediate feature is subjected to convolution operation.

[0055] S312, the received local detail feature and the global semantic feature are processed by the first cross-model spatial attention module SA_CM-1 to obtain a second feature , wherein the first cross-model spatial attention module SA_CM-1 obtains a global semantic related spatial attention weight , and the spatial attention weight enhances the global semantic information in the received local detail feature , which can be represented by the following formula: wherein, indicates that the first intermediate feature is subjected to convolution operation.

[0056] S313, the first feature and the second feature are element-wise added to obtain an enhanced local detail feature , which can be represented by the following formula: The branch for obtaining the enhanced local detail feature can be referred to as a SwinToVgg branch, and the SwinToVgg branch can introduce the context information in the global semantic feature in the Swin-Transformer into the VGG-16 network.

[0057] S32, the received global detail feature is processed by the second channel attention module and the second spatial attention module in sequence to obtain a third feature, the received local detail feature and the received global semantic feature are processed by the second cross-model spatial attention module to obtain a fourth feature, and the third feature and the fourth feature are element-wise added to obtain an enhanced global semantic feature, specifically: S320, the received global semantic feature After being processed by the second channel attention module CA-2, the second intermediate feature is obtained Specifically, The second channel attention module CA-2 enhances the received global semantic feature by a series of pooling and convolution operations to obtain channel attention weights , and uses the channel attention weights to enhance the channel information in the received global semantic feature to obtain the second intermediate feature , which can be specifically represented by the following formula: Wherein, represents performing a global average pooling operation on , represents performing a global maximum pooling operation on , and represents performing a convolution operation on

[0058] S321, the second intermediate feature After being processed by the second spatial attention module SA-2, the third feature is obtained , specifically, The second spatial attention module SA-2 obtains the spatial attention weights of the third feature , and uses the spatial attention weights to enhance the important spatial position information in the third feature , which can be specifically represented by the following formula: Wherein, represents performing a convolution operation on

[0059] S322, the received local detail feature and the global semantic feature After being processed by the second cross-model spatial attention module SA_CM-2, the fourth feature is obtained , specifically, the second cross-model spatial attention module SA_CM-2 obtains the spatial attention weights related to the local details, and uses the spatial attention weights to enhance the global semantic feature ​​​​​​The local detail information in can be expressed by the following formula: in, Indicates: Yes conduct Convolution operation.

[0060] S323, the third feature and the fourth characteristic Add element by element to obtain enhanced global semantic features , which can be expressed by the following formula: The branch of the enhanced global semantic features can be called the VggToSwin branch. The VggToSwin branch can introduce the local detail information in the local detail features of the VGG-16 network into the Swin-Transformer.

[0061] Optionally, in the above technical solution, the edge-aware enhancement module is specifically used to: obtain the received convolution result or the boundary map in the local detail feature, and calculate the Hadamard product of the boundary map and the received output of the decoder, and process the Hadamard product through a convolution layer to obtain the convolution result corresponding to the Hadamard product, and add the convolution result corresponding to the Hadamard product and the received output of the decoder element by element to obtain the output of the edge-aware enhancement module.

[0062] Among them, the input of the edge perception enhancement module includes and ,when hour, Indicates the The convolution result corresponding to the splicing features; when hour, Indicates: The local detail features output by the first network layer; Indicates: The output of a decoder.

[0063] like Figure 4 As shown in Figure 2, at different decoding stages, the edge information of the salient target is enhanced through the edge-aware enhancement module. The specific data processing process of the edge-aware enhancement module is as follows: S40, calculation along the channel dimension The maximum and mean values ​​of The maximum value and mean value of The concatenated features of The concatenated features are subjected to a 3×3 convolution operation, and the convolution operation result is input into function, get the calculation result, subtract the calculation result from 1, and get The corresponding boundary map , which can be expressed by the following formula: Indicates calculation along the channel dimension The mean of Calculated along the channel dimension The maximum value of Representation: Through the convolution layer The concatenated features are subjected to 3×3 convolution operation. Indicates: enter The calculation result obtained by the function.

[0064] S41. Calculate boundary graph and The Hadamard product is processed through a convolution layer to obtain the convolution result corresponding to the Hadamard product. The convolution result corresponding to the Hadamard product is summed up. Perform element-by-element addition to obtain the output of the edge-aware enhancement module , which can be expressed by the following formula: in, Denotes: the output of the qth edge-aware enhancement module.

[0065] Optionally, in the above technical solution, the following is further included: In order to ensure accurate boundary positioning, each intermediate boundary map is supervised by a first loss function, wherein the first loss function can be a binary cross entropy (BCE) loss function, a mean square error loss function, or a multi-class cross entropy loss function. This application uses the binary cross entropy loss function for illustration, which can be expressed by the following formula: in, Denotes: Binary cross entropy loss function. express: The corresponding real boundary map, Specifically, it can be obtained from the true saliency map corresponding to the preset sample image through the Canny algorithm. represents: the number of generated boundary graphs, Represents: the sum of the binary cross entropy losses corresponding to all boundary maps.

[0066] Optionally, further comprising: performing a convolution operation on the output of each decoder through a convolution layer, and utilizing a function to obtain a salient object detection result corresponding to the output of each decoder, which can be expressed by the following formula: wherein, represents: performing a convolution operation on the output of the i-th decoder through a convolution layer; represents: inputting a function,

[0067] It should be noted that in S2, the salient object detection model trained is used to perform salient object detection on the target image, and the salient object detection result obtained through the output of the last decoder is selected as the final salient object detection result, which is then provided to the user.

[0068] Optionally, in the above technical solution, the salient object detection result is supervised and trained by using a second loss function.

[0069] In order to effectively train and optimize the salient object detection model, a plurality of loss functions can be combined to form the second loss function, that is, the loss calculated by the second loss function is the sum of the losses calculated by the plurality of loss functions. In this embodiment, the first loss function, the binary cross entropy loss function and the Intersection over Union (IoU) loss function are used to construct the second loss function, that is, the loss calculated by the second loss function is the sum of the losses calculated by the first loss function, the binary cross entropy loss function and the Intersection over Union (IoU) loss function. It should be noted that the first loss function not only supervises the boundary map, but also supervises and optimizes the salient object detection model.

[0070] Based on Figures 5 to 9 , the salient object detection effect of the trained salient object detection model is described. Specifically, the salient object detection model of the present application and the comparative model (including RCSB and DPNet, etc.) are trained by using the international standard data set (DUTS-TE data set, ECSSD data set and DUT-OMRON) shown in Figure 5 , , , and ​​​​​The salient target detection effects of the trained salient target detection model and the trained comparison model were compared using evaluation indicators such as Figure 7 As shown in the figure, the salient target detection results obtained by the trained salient target detection model are compared with the salient target detection results of the trained comparison model. The comparison results are shown in the figure. Figure 8 As shown in FIG, it can be seen that the salient target detection model trained in this application has good salient target detection accuracy. The introduction of the comparison model is as follows: Figure 7 shown.

[0071] In the above embodiments, although the steps are numbered S1, S2, etc., these are only specific embodiments given by the present invention. Those skilled in the art may adjust the execution order of S1, S2, etc. according to actual conditions, which is also within the scope of protection of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0072] like Figure 10 As shown, a salient object detection system 200 according to an embodiment of the present invention includes a model training module 201 and a salient object detection module 202; The model training module 201 is used to: train the constructed salient target detection model based on multiple sample images to obtain a trained salient target detection model, wherein the constructed salient target detection model includes: a dual-branch encoder, a cross-model interactive fusion module, an edge-aware enhancement module and a decoder, extracting multiple local detail features and multiple global semantic features of the target sample image through the dual-branch encoder, interactively fusing some local detail features and all global semantic features through a cross-model interactive fusion module according to a preset correspondence, thereby obtaining multiple enhanced local detail features and multiple enhanced global detail features, performing feature splicing on each enhanced local detail feature and the corresponding enhanced global detail feature to obtain multiple spliced ​​features, combining all spliced ​​features and the remaining local detail features after processing through the edge-aware enhancement module and the decoder to obtain a salient target detection result, wherein the target sample image is any target sample image; The salient object detection module 202 is used to perform salient object detection on the target image using the trained salient object detection model.

[0073] Optionally, in the above technical solution, the first branch of the dual-branch encoder includes N first network layers arranged in sequence, and the second branch of the dual-branch encoder includes M second network layers arranged in sequence; after the target sample image is input into the first first network layer, local detail features output by the first to nth first network layers are obtained; The local detail feature output by the n-1th first network layer is input into the 1st second network layer to obtain global semantic features output by the 1st second network layer, and the local detail feature output by the nth first network layer and the global semantic features output by the 1st second network layer are interactively fused through the 1st cross-model interactive fusion module, so that the local detail information in the local detail feature output by the nth first network layer is integrated into the global semantic features output by the 1st second network layer to obtain 1st enhanced global semantic features, and the global context information in the global semantic features output by the 1st second network layer is integrated into the local detail feature output by the nth first network layer to obtain 1st enhanced local detail features. After the 1st enhanced global semantic features and the 1st enhanced local detail features are spliced, the 1st spliced features are obtained. The 1st enhanced global semantic features are taken as the input of the 2nd second network layer, and the 1st enhanced local detail features are taken as the input of the n+1th first network layer. The outputs of the 2nd second network layer and the n+1th first network layer are interactively fused through the 2nd cross-model interactive fusion module to obtain the 2nd enhanced global semantic features and the 1st enhanced local detail features, and the 2nd spliced features are obtained by splicing. Until the Mth spliced feature is obtained, wherein n and M are positive integers, N=M+n-1.

[0074] Optionally, in the above technical solution, each spliced feature is processed through a convolution layer to obtain a convolution result corresponding to each spliced feature; the convolution result corresponding to the Mth spliced feature is input into an ASPP module to obtain multi-scale features, the multi-scale features and the convolution result corresponding to the Mth spliced feature are input into a 1st decoder, and the output of the 1st decoder and the convolution result corresponding to the Mth spliced feature are input into a 1st edge perception enhancement module to enhance the edge information of the salient target through the 1st edge perception enhancement module. The output of the 1st edge perception enhancement module and the M-1th spliced feature are input into a 2nd decoder, and the output of the 2nd decoder and the convolution result corresponding to the M-1th spliced feature are input into a 2nd edge perception enhancement module, until the output of the mth edge perception enhancement module is obtained, wherein m is a positive integer, m=M-n+1; the output of the mth edge perception enhancement module and the local detail feature output by the n-1th first network layer are input into an m+1th decoder, the output of the m+1th decoder and the local detail feature output by the n-1th first network layer are input into an m+1th edge perception enhancement module, and the output of the m+1th edge perception enhancement module and the local detail feature output by the n-2th first network layer are input into an n+2th decoder, until the output of the Nth decoder is obtained, and the output of the Nth decoder is taken as the salient target detection result.

[0075] Optionally, in the above technical solution, the cross-model interaction fusion module is specifically configured to: After receiving the local detail feature and the global semantic feature, the received local detail feature is sequentially processed by the first channel attention module and the first spatial attention module to obtain a first feature, and the received local detail feature and the received global semantic feature are processed by the first cross-model spatial attention module to obtain a second feature, and the first feature and the second feature are element-wise added to obtain an enhanced local detail feature. The received global detail feature is sequentially processed by the second channel attention module and the second spatial attention module to obtain a third feature, and the received local detail feature and the received global semantic feature are processed by the second cross-model spatial attention module to obtain a fourth feature, and the third feature and the fourth feature are element-wise added to obtain an enhanced global semantic feature.

[0076] Optionally, in the above technical solution, the edge-aware enhancement module is specifically configured to: obtain a boundary map in the received convolution result or the local detail feature, calculate a Hadamard product of the boundary map and an output of the received decoder, process the Hadamard product through a convolution layer to obtain a convolution result corresponding to the Hadamard product, and element-wise add the convolution result corresponding to the Hadamard product and the output of the received decoder to obtain an output of the edge-aware enhancement module.

[0077] Optionally, in the above technical solution, the model training module is further configured to: in training the constructed salient object detection model, a first loss function is used to train and supervise the boundary map, and a second loss function is used to train and supervise the salient object detection result.

[0078] It should be noted that the beneficial effects of the salient object detection system 200 provided by the above embodiments are the same as those of the salient object detection method, which will not be repeated here. In addition, when the system provided by the above embodiments implements its functions, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the system is divided into different functional modules according to actual conditions to complete all or part of the above described functions. In addition, the system and method embodiments provided by the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0079] Among them, the salient object detection system of the application can be a computer program (including program code) running in a computer device, for example, the salient object detection system of the application is an application software, which can be used to execute the corresponding steps in the salient object detection method of the application.

[0080] In some embodiments, the salient object detection system of the present invention can be implemented using a combination of software and hardware. As an example, the salient object detection system of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the salient object detection method of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0081] The modules described in the embodiments of the present invention may be implemented in software or hardware, and the name of a module does not necessarily limit the module itself.

[0082] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the processor executes the computer program, any of the above-mentioned salient target detection methods is implemented. That is, an electronic device according to an embodiment of the present invention may include but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the salient target detection method shown in any embodiment of the present invention by calling the computer program.

[0083] In an alternative embodiment, an electronic device is provided, such as Figure 11 As shown, Figure 11 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0084] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in connection with the present disclosure. The processor 4001 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0085] The bus 4002 can include a path for transmitting information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 11 The bus 4002 is represented by only one thick line, but it does not mean that there is only one bus or only one type of bus.

[0086] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto.

[0087] The memory 4003 is used to store application code (computer program) for executing the solution of the present invention, and is controlled by the processor 4001. The processor 4001 is used to execute the application code stored in the memory 4003 to implement the content shown in the above method embodiment.

[0088] Among them, the electronic device can also be a terminal device, and the terminal device can be any device that can install applications, including at least one of a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart TV, and a smart car device.

[0089] It should be noted that Figure 11 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0090] A computer-readable storage medium according to an embodiment of the present invention stores a computer program, which, when executed by a processor, implements any of the above-mentioned salient object detection methods.

[0091] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0092] In an exemplary embodiment, a computer program product or computer program is also provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the aforementioned salient object detection methods.

[0093] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0094] It should be understood that the flowchart and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of various embodiments of the present application. In this regard, each block in the flowchart and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.

[0095] The computer readable storage medium of embodiments of the present application can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, the computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0096] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0097] The above description is merely an illustration of preferred embodiments of the present invention and the underlying technical principles. Those skilled in the art should understand that the scope of the present invention is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned concepts. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this invention.

[0098] It should be noted that the terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and to define a specific order or precedence. Where appropriate, the order used for similar objects may be interchanged, such that the embodiments of the present application described herein can be implemented in an order other than the order shown or described.

[0099] Those skilled in the art will appreciate that the present invention may be implemented as a system, method, or computer program product. Therefore, the present invention may be implemented entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present invention may be implemented as a computer program product embodied in one or more computer-readable media containing computer-readable program code.

[0100] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A salient object detection method, characterized in that, The method comprises the following steps: Based on a plurality of sample images, a constructed salient object detection model is trained to obtain a trained salient object detection model, wherein the constructed salient object detection model comprises a double-branch encoder, a cross-model interaction fusion module, an edge perception enhancement module and a decoder. The double-branch encoder is used to extract a plurality of local detail features and a plurality of global semantic features of a target sample image. Part of the local detail features and all the global semantic features are respectively subjected to interaction fusion through a cross-model interaction fusion module according to a preset corresponding relationship to obtain a plurality of enhanced local detail features and a plurality of enhanced global detail features. Each enhanced local detail feature and the corresponding enhanced global detail feature are subjected to feature splicing to obtain a plurality of spliced features. The salient object detection result is obtained by combining all the spliced features and the remaining part of the local detail features, and processing the features through the edge perception enhancement module and the decoder. The target sample image is any target sample image. The trained salient object detection model is used for salient object detection of a target image.

2. The salient object detection method of claim 1, wherein, The first branch of the double-branch encoder comprises N first network layers arranged in sequence, and the second branch of the double-branch encoder comprises M second network layers arranged in sequence. The double-branch encoder is used to extract a plurality of local detail features and a plurality of global semantic features of a target sample image. Part of the local detail features and all the global semantic features are respectively subjected to interaction fusion through a cross-model interaction fusion module according to a preset corresponding relationship to obtain a plurality of enhanced local detail features and a plurality of enhanced global detail features. Each enhanced local detail feature and the corresponding enhanced global detail feature are subjected to feature splicing to obtain a plurality of spliced features, comprising: After the target sample image is input into the first first network layer, the local detail features output by the first first network layer to the nth first network layer are obtained. The local detail feature output by the n-1th first network layer is input into the 1st second network layer to obtain global semantic features output by the 1st second network layer, and the local detail feature output by the nth first network layer is interactively fused with the global semantic features output by the 1st second network layer through the 1st cross-model interactive fusion module, so that the local detail information in the local detail feature output by the nth first network layer is integrated into the global semantic features output by the 1st second network layer to obtain 1st enhanced global semantic features, and the global context information in the global semantic features output by the 1st second network layer is integrated into the local detail feature output by the nth first network layer to obtain 1st enhanced local detail features. After the 1st enhanced global semantic features and the 1st enhanced local detail features are spliced, the 1st spliced features are obtained. The 1st enhanced global semantic features are taken as the input of the 2nd second network layer, and the 1st enhanced local detail features are taken as the input of the n+1th first network layer. The output of the 2nd second network layer and the output of the n+1th first network layer are interactively fused through the 2nd cross-model interactive fusion module to obtain the 2nd enhanced global semantic features and the 1st enhanced local detail features, and the 2nd spliced features are obtained by splicing. Until the Mth spliced feature is obtained, wherein n and M are positive integers, N=M+n-1.

3. The salient object detection method of claim 2, wherein, After all spliced features and the local detail features of the remaining part are processed by the edge perception enhancement module and the decoder, the salient object detection result is obtained, including: Each spliced feature is processed through a convolution layer to obtain a convolution result corresponding to each spliced feature; The convolution result corresponding to the Mth spliced feature is input into the ASPP module to obtain multi-scale features. The multi-scale features and the convolution result corresponding to the Mth spliced feature are input into the 1st decoder, and the output of the 1st decoder and the convolution result corresponding to the Mth spliced feature are input into the 1st edge perception enhancement module to enhance the edge information of the salient object through the 1st edge perception enhancement module. The output of the 1st edge perception enhancement module and the M-1th spliced feature are input into the 2nd decoder, and the output of the 2nd decoder and the convolution result corresponding to the M-1th spliced feature are input into the 2nd edge perception enhancement module, until the output of the mth edge perception enhancement module is obtained, m is a positive integer, m=M-n+1. The output of the mth edge perception enhancement module and the local detail feature output by the n-1th first network layer are input into the m+1th decoder, the output of the m+1th decoder and the local detail feature output by the n-1th first network layer are input into the m+1th edge perception enhancement module, the output of the m+1th edge perception enhancement module and the local detail feature output by the n-2th first network layer are input into the n+2th decoder, until the output of the Nth decoder is obtained, and the output of the Nth decoder is taken as the salient object detection result.

4. The salient object detection method of claim 3, wherein, The cross-model interaction fusion module is specifically used for: After receiving the local detail feature and the global semantic feature, the received local detail feature is sequentially processed by the first channel attention module and the first spatial attention module to obtain a first feature, and the received local detail feature and the received global semantic feature are processed by the first cross-model spatial attention module to obtain a second feature, and the first feature and the second feature are added element by element to obtain an enhanced local detail feature; The received global detail feature is sequentially processed by the second channel attention module and the second spatial attention module to obtain a third feature, and the received local detail feature and the received global semantic feature are processed by the second cross-model spatial attention module to obtain a fourth feature, and the third feature and the fourth feature are added element by element to obtain an enhanced global semantic feature.

5. The salient object detection method of claim 4, wherein, The edge-aware enhancement module is specifically configured to: obtain a boundary map in the received convolution result or the local detail feature, calculate a Hadamard product of the boundary map and an output of the received decoder, process the Hadamard product through a convolution layer to obtain a convolution result corresponding to the Hadamard product, and add the convolution result corresponding to the Hadamard product and the output of the received decoder element by element to obtain an output of the edge-aware enhancement module.

6. The salient object detection method of claim 5, wherein, Further comprising: In the training of the constructed saliency target detection model, a first loss function is used to train and supervise the boundary map, and a second loss function is used to train and supervise the saliency target detection result.

7. A salient object detection system characterized by, The model training module and the saliency target detection module are included. The model training module is configured to: based on a plurality of sample images, train a constructed saliency target detection model to obtain a trained saliency target detection model, wherein the constructed saliency target detection model includes a double-branch encoder, a cross-model interaction fusion module, an edge-aware enhancement module, and a decoder, the double-branch encoder extracts a plurality of local detail features and a plurality of global semantic features of a target sample image, part of the local detail features and all the global semantic features are respectively interacted and fused through a cross-model interaction fusion module according to a preset corresponding relationship to obtain a plurality of enhanced local detail features and a plurality of enhanced global detail features, each enhanced local detail feature and the corresponding enhanced global detail feature are spliced to obtain a plurality of spliced features, and all the spliced features and the remaining part of the local detail features are processed by the edge-aware enhancement module and the decoder to obtain a saliency target detection result, wherein the target sample image is any target sample image; The saliency target detection module is configured to: use the trained saliency target detection model to perform saliency target detection on a target image.

8. The salient object detection system of claim 7, wherein, The target sample image is input into the first first network layer, and local detail features output by the first to the nth first network layer are obtained; the local detail features output by the (n-1)th first network layer are input into the first second network layer, global semantic features output by the first second network layer are obtained, and the local detail features output by the nth first network layer and the global semantic features output by the first second network layer are interactively fused through the first cross-model interactive fusion module, so that local detail information in the local detail features output by the nth first network layer is fused into the global semantic features output by the first second network layer to obtain first enhanced global semantic features, and global context information in the global semantic features output by the first second network layer is fused into the local detail features output by the nth first network layer to obtain first enhanced local detail features; after the first enhanced global semantic features and the first enhanced local detail features are spliced, a first spliced feature is obtained; the first enhanced global semantic features are taken as input of a second second network layer, the first enhanced local detail features are taken as input of an (n+1)th first network layer, and the output of the second second network layer and the output of the (n+1)th first network layer are interactively fused through a second cross-model interactive fusion module to obtain second enhanced global semantic features and first enhanced local detail features, which are spliced to obtain a second spliced feature, and this process is repeated until an Mth spliced feature is obtained, wherein n and M are positive integers, and N=M+n-1.

9. An electronic device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the salient object detection method in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the salient object detection method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Salient target detection method and system based on local and global context fusion

    CN114581747A

  • Significance target detection method fusing mixed attention

    CN117935031A

  • Improved UNet type RGB-D saliency target detection system

    CN119785005A

  • Person re-identification method and apparatus for fusing global features with ladder-shaped local features

    WO2024021394A1

Cited By

  • Target detection method and system with hierarchical scanning perception fused with super-resolution enhancement

    CN121921499A