Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

153 results about "Salient object detection" patented technology

Salient object detection is a task based on a visual attention mechanism, in which algorithms aim to explore objects or regions more attentive than the surrounding areas on the scene or images.

Optical remote sensing image salient target detection method based on progressive attention enhancement

The invention discloses an optical remote sensing image salient target detection method based on progressive attention enhancement, and belongs to the technical field of computer vision. The method comprises the following steps: preprocessing an original data set; inputting the preprocessed image into a hierarchical progressive fusion encoder, capturing a global irregular topological structure and local fine-grained image details, and realizing cross-hierarchical feature fusion; inputting the output characteristics of the encoder into a global context enhancement module, and capturing multi-level context information by adopting a parallel multi-branch structure; and inputting the output features of the hierarchical progressive fusion encoder and the global context enhancement module into a multi-scale progressive attention enhancement decoder, carrying out hierarchical decoding on the input features by adopting a saliency-guided attention mechanism, and gradually aggregating deep semantic information and shallow detail features to realize coarse-to-fine progressive optimization, so as to improve the robustness of the multi-scale progressive attention enhancement decoder. And finally generating a saliency map. The method can effectively improve the processing performance of an irregular topological structure and a complex context relationship in the optical remote sensing image.
Owner:SHIJIAZHUANG TIEDAO UNIV

Multi-dimensional frequency domain and deformable attention fusion saliency target detection method

The invention relates to the field of saliency target detection, and particularly discloses a multi-dimensional frequency domain and deformable attention fusion saliency target detection method, which comprises the steps of S1, inputting an infrared image to be detected; s2, performing multi-scale feature extraction and fusion to obtain low-level and high-level features; s3, phase spectrum analysis is carried out to extract frequency domain primary perception features; s4, fusing the frequency domain features to obtain frequency domain saliency features; s5, a deformable space attention module extracts space enhanced perception features; and S6, fusing the features to generate a prediction map and constraining the prediction map by a loss function. According to the method, the problems of insufficient frequency domain utilization, weak global context and detail retention and poor complex deformation target detection of an existing spatial domain method are solved, and the detection precision and robustness of a multi-scale and deformation target in a complex scene are effectively improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

RGB-D salient target detection method based on semantic features and biological inspiration

The invention relates to the technical field of computer vision, in particular to an RGB-D salient target detection method based on semantic features and biological inspiration. The method comprises the following steps: performing feature extraction on an RGB-D image sample through an encoder to obtain low-level, middle-level and high-level semantic features of an RGB image and a depth image; fusing the low-level semantic features of the RGB image and the depth image through a multi-stage fusion module to obtain low-level fusion features; fusing the intermediate semantic features of the RGB image and the depth image to obtain intermediate fusion features; fusing the advanced semantic features of the RGB image and the depth image to obtain an advanced fusion feature; and decoding the low-level, middle-level and high-level fusion features through a cortex decoder to obtain a saliency prediction map. According to the method, the efficiency, the robustness and the generalization ability of the RGB-D saliency target detection model are improved.
Owner:JIANGXI NORMAL UNIV

Light field salient target detection method based on edge perception and hierarchical fusion

The invention relates to the technical field of light field image salient target detection, in particular to a light field salient target detection method based on edge perception and hierarchical fusion, and the method comprises the steps: carrying out the multi-scale feature extraction of a focus stack image and a full-focus image through a backbone network; performing edge fusion enhancement on the focus stack features of four layers of different scales through an SEPM module and an EFM module; fusing high-level multi-modal semantic information from the global and local aspects by using an LHFM module; fusing low-layer space information and refining a salient target by using an LLFM module; and aggregating multi-scale information of a high layer and a low layer, and decoding the multi-scale information into an accurate saliency prediction image by using a detection head. According to the method, edge perception and a lightweight hierarchical fusion strategy are combined, the model parameter quantity and the calculation complexity are remarkably reduced while the high detection performance is kept, and the optimal balance between the performance and the efficiency is achieved.
Owner:CHONGQING UNIV OF TECH

Three-mode saliency target detection method and system based on frequency domain decomposition and reconstruction

The invention discloses a three-mode saliency target detection method and system based on frequency domain decomposition and reconstruction. The method comprises the following steps: firstly, respectively preprocessing a training set and a test set in a three-mode saliency target detection data set; secondly, constructing a three-mode saliency target detection network based on frequency domain decomposition and reconstruction; and finally, sending the preprocessed training set image into a three-mode saliency target detection network for processing, outputting a prediction map consistent with the input image in size, completing target detection, and performing training and testing. According to the invention, through designing the interaction, fusion and enhancement network, the information complementation advantages of three modes of visible light, depth and thermal imaging are fully utilized, the synergistic interaction and global perception efficiency among multi-mode information are further enhanced, and accurate salient target detection is realized.
Owner:HANGZHOU DIANZI UNIV

RGB-D salient target detection method based on decoupling contrast learning

The invention discloses an RGB-D salient target detection method based on decoupling contrast learning, and designs a saliency detection framework integrating expression enhancement, modal collaborative perception and structural discrimination learning by combining a structural heterogeneity problem in multi-modal modeling and utilizing the frequency domain structural advantage of a deep mode and the long-distance modeling capability of Transform. By introducing wavelet convolution and Transform joint modeling, a cross-modal interaction parallel fusion mechanism and a pixel-level structure perception contrast learning strategy, high-precision, multi-scale and boundary clear detection of a salient target area in a complex scene is realized. The method can effectively solve the problems of large information difference between modes of the RGB and the depth map, difficulty in structure alignment, fuzzy boundary prediction, weak feature expression ability and the like, significantly improves semantic consistency and structural integrity of the salient region, and has good cross-modal generalization ability and robustness.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

RGB-D saliency target detection method

The invention discloses an RGB-D saliency target detection method, and relates to the technical field of computer vision. Comprising the following steps: acquiring a color image, depth information and a corresponding RGB-D saliency target annotation graph from an RGB-D saliency detection data set; inputting the color image and the depth information into a cross-modal saliency detection network to obtain an RGB-D saliency target prediction map; the cross-modal saliency detection network is trained through the RGB-D saliency target prediction map and the RGB-D saliency target annotation map, and the trained cross-modal saliency detection network is obtained; and inputting a to-be-processed color image and to-be-processed depth information into the trained cross-modal saliency detection network to obtain an RGB-D saliency target recognition graph. According to the method, the visual integrity and detail fidelity of saliency detection are remarkably improved, and the accuracy of a saliency detection result is enhanced.
Owner:NORTHWEST NORMAL UNIVERSITY

Salient target detection method based on Mama network bidirectional guidance model

The invention discloses a saliency target detection method based on a Mama network two-way guidance model. The method comprises the steps that a two-way model framework based on the Mama network is composed of an encoder branch, an edge branch, a saliency branch and a decoder branch; the encoder branch performs block division on the obtained original image to be detected, inputs the image to the Mama feature extractor and performs down-sampling to obtain five-layer features, the first three-layer low-layer features are respectively subjected to convolution processing and then are used as edge features to be input into the edge branch, the first four-layer features are used as significant features to be input into the significant branch, and the significant branch is used as edge features to be input into the edge branch; the fifth-layer high-level features are subjected to two-stage mixed attention processing, global features are obtained, and positioning guidance is provided for subsequent feature fusion; the edge branch performs receptive field expansion and spatial enhancement on the input edge features to obtain edge fusion features. According to the invention, the precision and efficiency of saliency target detection can be well improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Underwater image salient target detection method and system based on conditional diffusion model

The invention belongs to the field of salient target detection, and provides an underwater image salient target detection method and system based on a conditional diffusion model, and the method comprises the steps: carrying out the feature extraction of an RGB image and a depth image in each time step, and obtaining a plurality of RGB features and depth features with different resolutions; fourier domain perception enhancement is carried out on the RGB features and the depth features under the same resolution, and global Fourier optimization features are obtained; generating splicing features based on the RGB features and the depth features under the same resolution, and performing spatial domain perception enhancement on the splicing features to obtain spatial optimization features; fusing the global Fourier optimization features and the spatial optimization features to obtain fusion results, and fusing the fusion results under different resolutions to obtain condition features of each time step; and predicting the condition features of different time steps and the loading reference image to obtain prediction results of different time steps, and screening and aggregating based on the prediction results of different time steps to obtain an aggregation detection result.
Owner:HAINAN UNIV

Underwater salient target detection method and system based on double-flow fusion network

The invention discloses an underwater salient target detection method and system based on a double-flow fusion network, and belongs to the technical field of computer vision. The method comprises the following steps: respectively extracting multi-scale features of an RGB image and a depth image through a double-flow encoder; in the shallow layer, fusing and enhancing the edge and detail information of the bimodal features through an edge fusion module; in a deep layer, content-adaptive cross-modal semantic fusion is realized in a frequency domain through a dynamic filtering module; fusing the multi-scale features through a cross-layer aggregation decoder to generate a rough saliency map; extracting detail features from the original RGB image through a global detail purification network; and finally, fusing the rough saliency map and the detail features, and outputting an underwater saliency target prediction map. The objective of the invention is to improve the precision and boundary definition of salient target detection in an underwater complex scene.
Owner:NANKAI UNIV

Optical remote sensing image salient target detection method based on Mama dynamic clustering and bidirectional calibration

The invention discloses an optical remote sensing image salient target detection method based on Mama dynamic clustering and bidirectional calibration, and belongs to the field of computer vision. The method comprises the following steps: preprocessing an original data set; inputting the preprocessed image into a lightweight encoder, capturing multi-scale features and refining local textures and edges; the multi-scale features output by the lightweight encoder are input into a dynamic clustering module based on Mamba, and interaction enhancement of global semantic modeling and dynamic local feature capture is achieved; inputting the output features of the Mama-based dynamic clustering module into a bidirectional cross-scale calibration module to realize cross-scale feature bidirectional complementation and semantic detail enhancement; inputting the output features of the bidirectional cross-scale calibration module into an edge attention combined repair module to realize attention hole repair and boundary precision enhancement; and finally, realizing feature aggregation and spatial resolution recovery through a decoder, and finally generating a saliency map. The method is used for solving the problems of target scale inconsistency and boundary blur in the remote sensing image.
Owner:SHIJIAZHUANG TIEDAO UNIV

Multi-domain and Mama collaborative saliency target detection method for 360-degree image

The invention provides a multi-domain and Mama collaborative saliency target detection method oriented to a 360-degree image, mainly relates to a saliency region detection method oriented to image equatorial region structure modeling and global guidance enhancement, introduces PVT as a backbone network, extracts multi-scale features, inputs the multi-scale features to a frequency domain-space domain coordination module, and finally, obtains a multi-scale target detection result. The multi-scale features extracted by the PVT backbone are fully fused through frequency domain and spatial domain information, so that multi-scale edge details in the image can be effectively captured, and the significance boundary of the equatorial region is enhanced; an attention fusion Mama module is introduced, by fusing output features of a frequency domain-space domain coordination module, the Mama module can effectively improve structural guidance and semantic complementation of equator saliency information on a polar region, and finally a lightweight multi-stage feature aggregation module is designed for generating a saliency feature map. According to the detection method provided by the invention, the most advanced performance can be obtained under the condition of relatively low calculation complexity.
Owner:JIANGXI UNIV OF SCI & TECH

Multi-modal image saliency target detection method

The invention discloses a multi-modal image saliency target detection method, which comprises the following steps of: constructing a training set comprising a color visible light image, an infrared image and a depth image, and constructing a neural network which consists of a feature extraction module, a three-modal feature fusion module and a combined decoding module, the feature extraction module extracts features and scale information of three modal images, the three-modal feature fusion module integrates features through a plurality of three-modal fusion modules, and the combined decoding module outputs a saliency target image through a plurality of prediction branches; a neural network is trained based on the training set to obtain a neural network model, and the neural network model can be used for testing saliency target detection of the image pair; the method has the advantages that the saliency target detection problem of the three-mode combined input image can be effectively solved, and the saliency target detection precision is high.
Owner:NINGBO UNIV

RGB-D lightweight semantic segmentation method fusing frequency domain guidance

The invention provides an RGB-D lightweight semantic segmentation method fusing frequency domain guidance, relates to the technical field of image processing, and designs a frequency domain guidance prompt adapter to improve consistency and propagation efficiency of cross-layer semantic features. Secondly, a spectrum-guided dynamic convolution module is provided, and efficient multi-scale feature modeling is realized while spatial domain and frequency domain features are fused. And finally, constructing a multi-scale frequency domain agency attention module, and enhancing semantic interaction and global modeling capability among different scale features in a low-overhead mode. According to the method, on a plurality of RGB-D and RGB-L semantic segmentation data sets, excellent segmentation performance can be achieved with low parameter quantity, good generalization ability is shown in five data sets in an RGB-D salient target detection task, and research results show that the segmentation result of the method is more accurate in a scene with a complex structure.
Owner:LIAO NING GONG CHENG JI SHU DA XUE E ER DUO SI YAN JIU YUAN

Salient target detection methods for RGB-D images

The present invention aims to provide a salient target detection method for RGB-D images, comprising: sampling RGB image and depth map samples at a specified number to form sampling data, and preprocessing the sampling data using data augmentation techniques to obtain data to be processed; in the cross-modal fusion of each layer in the feature extraction stage, using an attention mechanism to infer salient regions to determine the saliency degree of different regions, and fusing at each layer to obtain low-level features and high-level features; performing concatenation and convolution operations on the high-level features and the low-level features to obtain joint features of RGB features and depth features; dynamically allocating the weights of RGB image and depth map features according to the joint features of RGB features and depth features to obtain a salient target detection model; and realizing the salient target detection result according to the salient target detection model. The method described in this invention can detect salient targets well and handle scenes with low depth map contrast.
Owner:GUANGDONG UNIV OF TECH

Memory-edge guided weakly supervised video salient object detection method and system

This invention relates to a weakly supervised video salient object detection method and system based on memory-edge guidance, belonging to the field of object detection technology. It includes: dividing a given long video sequence into non-overlapping sliding windows to extract several consecutive video frames; inputting these video frames into a trained memory-edge guidance network model to achieve weakly supervised video salient object detection; specifically, it includes: extracting features from consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales; using salient cues mined from historical frames to enhance the semantic representation of relevant objects in the current frame; and then inputting these features into a decoder to obtain the final salient object detection map. This invention achieves accurate localization and fine segmentation of salient objects in videos.
Owner:SHANDONG UNIV

Image processing method, apparatus and device

The present application provides an image processing method, device and equipment, which can be applied to the technical field of image processing. The image processing method comprises: pre-processing an input image to obtain input features; inputting the input features into a visual encoder and a multi-layer perception machine in a visual center decoupler respectively to obtain enhanced special features and salient object detection special features; the visual encoder aggregates local region features based on the input features to obtain the enhanced special features, and the multi-layer perception machine captures edge information based on the input features to obtain the salient object detection special features; inputting the enhanced special features into an enhancement network to obtain enhanced output features; the enhancement network takes illumination weights of different color channels and local binary pattern features of the input image as illumination constraints, and enhances the enhanced special features to obtain the enhanced output features; and inputting the salient object detection special features and the enhanced output features into a salient object detection network to detect a salient object.
Owner:TIANJIN UNIV

A fully supervised salient target detection method

The present application relates to a kind of full supervision's salient object detection method, constructs complete multi-branch feature fusion refinement network MFFRNet as salient object detection model;Again training set in data set is input to the proposed MFFRNet model training, every time completing a round will be back propagated once, to optimize MFFRNet model parameter;With data set test set, the performance of model is evaluated;Finally, the model after evaluation is used for salient object detection.The model effectively fuses the detail information of low-level feature and the semantic information of high-level feature.The module designed for low-level feature utilizes asymmetric convolution to reduce background noise and other interference factors, and a module designed for high-level feature obtains rich semantic information.Meanwhile, aliasing effects caused by frequent up-sampling are effectively handled.The method effectively captures salient objects and obtains saliency prediction map, and has strong robustness.
Owner:SHANGHAI INST OF TECH

Weak supervision salient target detection method based on progressive edge guide feature aggregation

The invention discloses a weak supervision salient target detection method based on progressive edge guidance, which is used for solving the problems of fuzzy target positioning and rough boundary in a complex scene. The method comprises a multi-stage feature aggregation module and an edge guiding module. The multi-stage feature aggregation module adopts a two-stage fusion mechanism. In the semantic guiding stage, semantic features are used for guiding the current decoding process. In the bidirectional interaction stage, maximum pooling and average pooling are fused, and high-level semantics and shallow-layer features are integrated. The edge guiding module selectively fuses key level features, combines channel weight strengthening and residual connection, progressively optimizes boundaries and generates a high-quality edge graph. And the edge graph and the multi-level features are fused layer by layer, so that the transmission of edge information from a deep layer to a shallow layer is realized. Finally, the validity of the method is verified on the S-DUTS data set.
Owner:CHANGCHUN UNIV OF SCI & TECH

An RGB-D salient object detection method based on boundary deformable convolution guidance

The application discloses an RGB-D salient object detection method based on boundary deformable convolution guidance, comprising the following steps: step one, respectively extracting features of an RGB mode and a depth map mode; step two, fusing the features of the two modes through a cross-modal attention fusion feature module to mine common and complementary features of salient objects; step three, inputting the feature map into an encoder deep layer embedded with an adjacent multi-scale feature enhancement module to obtain global context feature information; step four, generating a boundary clue map of the salient objects by constructing a boundary feature extraction module; and step five, generating a saliency map by using the generated boundary clue map and deformable convolution guidance. The application mines and strengthens the commonness of salient objects by cross-fusion of the depth map and the RGB image, effectively captures salient objects with different sizes and uncertain quantities by using adjacent level feature interaction, and solves the boundary blur problem of the saliency map by using the edge clue map to guide the model decoding.
Owner:ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE

RGB-D-based salient target detection method and system, storage medium and equipment

The invention discloses an RGB-D-based salient target detection method and system, a storage medium and equipment, and relates to the field of target detection, and the method comprises the steps: obtaining a picture: obtaining a to-be-detected RGB picture and a to-be-detected depth picture; in the stage of fusion coding, the RGB picture and the depth picture are respectively input into an RGB channel and a depth channel of a ResNet-50 convolutional neural network as a backbone network for processing; on the basis of a space channel attention mechanism, performing feature fusion on output of corresponding layers of the RGB channel and the depth channel to obtain RGB fusion output and depth fusion output of a corresponding stage; multi-stage fusion coding, wherein RGB fusion outputs of different stages are fused to obtain RGB multi-stage feature fusion output; fusing the depth fusion outputs of different stages to obtain depth multi-stage feature fusion output; and decoding: merging, compressing and fusing the RGB multi-stage feature fusion output and the depth multi-stage feature fusion output to obtain a detection result of the salient target. The purpose of improving the detection precision is achieved.
Owner:WEST CHINA HOSPITAL SICHUAN UNIV +3

Salient object detection method based on part-object relationship based on disentangled capsule routing

This invention discloses a method for detecting salient objects with a partial-object relationship based on disentangled capsule routing, comprising: extracting basic deep features from an input image to obtain basic deep features at five different scales; further extracting multi-receptive field deep features from the basic deep features; utilizing a pre-trained capsule network based on a disentangled routing algorithm to analyze the deep features at three deep scales to obtain corresponding partial-object relationship features; fusing the corresponding partial-object relationship features with the deep features to obtain a fused feature at three deep scales; and further fusing the deep features at the first and second scales and the fused feature at three deep scales to obtain a fused feature, and generating a saliency map based on the fused feature. This invention solves the problems of the existing techniques, such as the large number of network parameters and slow inference speed, by achieving better foreground and background segmentation and faster network inference speed. It can be used in image preprocessing in computer vision.
Owner:CHANGZHOU UNIV

Double-flow network salient target detection method introducing light field features

The invention discloses a double-flow network salient target detection method introducing light field features, and the method comprises the following steps: S1, creating a data set which comprises a focal sheet and an RGB image; s2, extracting features of the focal plate and the RGB image through a double-flow encoder; s3, feature fusion: S3-1, fusing the extracted focal sheet features, and fusing effective information in the focal sheet by using a focal sheet dimension attention module; s3-2, fusing the fused focal plate features obtained in the step S3-1 and the extracted RGB image features through a cross-modal feature fusion module to obtain cross-modal fusion features; and S4, carrying out step-by-step decoding on the cross-modal fusion features obtained in the step S3 through a decoding module. According to the method, the features of the target image, the features of the collaborative image and the features of the depth image can be effectively fused through the cross-modal feature fusion module. Therefore, the improvement of the traditional salient target detection based on RGB input through the input of the light field has a good effect.
Owner:张世龙

Salient object detection method, device and equipment

The invention discloses a saliency object detection method, device and equipment, and belongs to the technical field of image processing. The saliency object detection method comprises the following steps: acquiring a first image; dividing the first image into a plurality of color areas according to the pixel values of the pixel points in the first image; wherein the pixel value similarity of the pixel points in the same color area is greater than or equal to a pixel value similarity threshold; extracting at least one feature of each color area; determining a contrast of the first color region relative to the first feature; wherein the first color region is any one color region in the plurality of color regions, and the first feature is any one feature in the at least one feature; determining a histogram contrast of the first color region; determining a first comprehensive contrast ratio of the first color area according to the contrast ratio of the first color area relative to the first feature and the histogram contrast ratio of the first color area; and determining a salient object in the first image according to the first comprehensive contrast of the plurality of color regions.
Owner:VIVO MOBILE COMM CO LTD

A video salient object detection method and system based on spatio-temporal context scene relationship propagation

The application provides a video salient object detection method based on spatio-temporal context scene relationship propagation, and relates to the technical field of video salient object detection. First, scene analysis is performed on each frame of video in a video frame sequence to obtain an instance-level object corresponding to each frame of video. Then, global instance-level features, local instance-level features and intra-frame low-level features of each frame of video and the corresponding instance-level object are extracted. Then, a matching frame of any frame of video is randomly sampled from the same video frame sequence, global instance-level features of the matching frame are extracted, the global instance-level features are integrated into global instance-level features of the matching frame, and time features between video frames are obtained. The dense spatial attention mechanism is used to integrate the local instance-level features into the global instance-level features to obtain spatial features. The time features and the spatial features are spliced to obtain spatio-temporal features. The spatio-temporal features and the global instance-level features are input into a convolutional neural network based on a gated recurrent unit for updating to obtain high-level spatio-temporal features. Finally, the intra-frame low-level features and the high-level spatio-temporal features are fused and decoded to generate a video salient object mask detection result. The application utilizes rich inter-frame and intra-frame scene relationship information in the video, and improves the accuracy of video salient object detection in complex scenes.
Owner:GUANGDONG UNIV OF TECH

Progressive attention augmented optical remote sensing image salient object detection method

The application discloses a kind of based on progressive attention enhancement optical remote sensing image salient target detection method, belong to computer vision technical field.The method includes: the original data set is preprocessed;The image after pre-processing is input hierarchical progressive fusion encoder, captures global irregular topological structure and local fine-grained image details, and realizes cross-level feature fusion;The output feature of encoder is input global context enhancement module, adopts parallel multi-branch structure, captures multi-level context information;The output feature of hierarchical progressive fusion encoder and global context enhancement module is input multi-scale progressive attention enhancement decoder, adopts saliency guided attention mechanism, carries out hierarchical decoding to input feature, gradually aggregates deep semantic information and shallow detail features, realizes from coarse to fine progressive optimization, finally generates saliency map.The application can effectively improve the processing performance of irregular topological structure and complex context relationship in optical remote sensing image.
Owner:SHIJIAZHUANG TIEDAO UNIV

RGB-D salient target detection method based on potential perception and hierarchical fusion

The invention relates to the technical field of computer vision and deep learning, in particular to an RGB-D salient target detection method based on potential perception and hierarchical fusion. Respectively extracting RGB (Red, Green, Blue) and depth features by taking SMT (Surface Mount Technology) and MobileNetV2 as backbone networks; a potential label is generated by using the depth map and the true value map, and a quality factor output by a depth quality estimation module is supervised; inputting the RGB and depth features into a potential perception interaction module; the multi-scale fusion module and the position sensing fusion module are used for differentially fusing different levels of multi-modal features and mining modal complementarity; a differential decoding strategy is adopted for fusion features of a shallow layer and a deep layer through a layered refining decoder, and semantic guidance and feature refinement are combined; and finally, fusing the edge information into a final saliency map. According to the method, the relevance and complementarity among multi-level and multi-modal features are fully mined, the influence of low-quality features is dynamically inhibited, and the accuracy and robustness of salient target detection are improved.
Owner:HEBEI UNIV OF TECH

Multi-modal target detection method based on graph network information interaction

The invention provides a multi-modal target detection method based on graph network information interaction, and the method comprises the steps: collecting visible light and infrared images through a camera and an infrared thermal imaging head, and extracting multi-scale features through a feature extraction module, and the multi-modal feature interaction module performs inter-modal and intra-modal information interaction on the bimodal multi-scale features and enhances feature representation, the gated fusion module fuses the interacted features to generate multi-modal fusion features, and the multi-modal fusion feature detection head outputs a prediction result. According to the method, a multi-modal feature interaction module based on a graph network is designed, complementary information and long-range spatial dependence among multi-modal data are captured through graph reasoning, then the performance of salient target detection is improved, and the whole process is divided into two stages, namely inter-modal graph reasoning and intra-modal graph reasoning; the two stages act together, so that information between modals can be fully fused, and space structures in the modals are strengthened.
Owner:EAST CHINA JIAOTONG UNIVERSITY

A thumbnail generation method based on salient object detection and image quality evaluation

ActiveCN116433486BPattern recognitionData set
The application discloses a thumbnail generation method based on salient object detection and image quality evaluation, comprising the following steps: 1) preparing a salient object detection data set and training a model YOLO_SAL based on the data set; 2) inputting an image into the model YOLO_SAL to determine a salient core region of the image; 3) generating a to-be-screened set around the salient core region of the image through a cropping algorithm; and 4) screening out a thumbnail with the best aesthetic quality from the to-be-screened set through an image quality evaluation model SAMP_Net based on a composition rule. The application solves the problems of the existing thumbnail generation method based on deep learning, such as complex model, complicated steps, difficulty in landing and incomplete labeling target, etc. by combining YOLO_SAL and SAMP_Net, and improves the detection speed and the aesthetic quality of the thumbnail while ensuring that the predicted image core region has saliency and integrity.
Owner:SOUTH CHINA UNIV OF TECH

RGB-D image saliency detection method based on frequency decoupling mode interaction

The invention provides an RGB-D image salient target detection method based on frequency decoupling mode interaction, and the method comprises the steps: firstly carrying out the multi-stage feature extraction of RGB-D mode data, and obtaining the multi-stage feature representation of an RGB-D mode; performing frequency domain sensing cross-modal interaction on the RGB-D modal multi-stage feature representation to obtain frequency sensing cross-modal spatial domain interaction features; discriminative enhancement and cross-modal fusion are carried out on the cross-modal spatial domain interaction features of frequency sensing, and final multi-modal fusion features are obtained; and based on the final multi-modal fusion features, through multi-scale aggregation and global dependence modeling, generating a saliency target detection prediction map. According to the method, cross-modal association in a frequency domain is explicitly modeled in a feature learning process by using a frequency cross Mama fusion module, so that the problem that internal relationships among modals may be ignored when fusion is directly performed in a spatial domain in a traditional method is remarkably relieved, and the performance and generalization ability of salient target detection are improved.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS