Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

48 results about "Salient objects" patented technology

The definition of Salient Objects is an intrinsic property of the image, which can be reliably perceived among human subjects. The Definition of Salient Objects. Unlike fixation datasets, the most widely used salient object segmentation dataset is heavily biased.

Systems and methods for adjusting automatic image capture settings using a saliency-based region of interest

PCT designated stageWO2026019420A1Character and pattern recognitionSalient objectsHeat map
An example method includes generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The method also includes detecting, based on the one or more saliency heat maps, a salient object in the scene. The method additionally includes responsive to the detecting of the salient obj ect: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The method also includes adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.
Owner:GOOGLE LLC

Salient target rapid positioning method for electric power vision model

The invention discloses a salient target rapid positioning method for an electric power visual model, which belongs to the technical field of electric power system inspection, processes image data of electric power communication equipment through a deep learning algorithm, performs feature extraction and classification by adopting a convolutional neural network (CNN), and identifies abnormal conditions of salient targets in an electric power system in real time. Therefore, accurate positioning and detection of the target are realized. Through an integrated edge calculation method, the system can complete primary processing and analysis of image data on intelligent terminal equipment. Therefore, the delay of data transmission is reduced, and the response speed of target positioning and detection is improved. And the real-time requirement is met, so that the system can quickly respond in a complex electric power inspection environment, and the system is suitable for field inspection and equipment state monitoring of an electric power system. A high-quality electric power visual data set is constructed by integrating operation data of various devices such as an electric power transmission network, a transformer substation and the like, historical inspection data and sensor data.
Owner:STATE GRID SIJI FEITIAN (LANZHOU) CLOUD TECH CO LTD

Image processing method, apparatus and device

The present application provides an image processing method, device and equipment, which can be applied to the technical field of image processing. The image processing method comprises: pre-processing an input image to obtain input features; inputting the input features into a visual encoder and a multi-layer perception machine in a visual center decoupler respectively to obtain enhanced special features and salient object detection special features; the visual encoder aggregates local region features based on the input features to obtain the enhanced special features, and the multi-layer perception machine captures edge information based on the input features to obtain the salient object detection special features; inputting the enhanced special features into an enhancement network to obtain enhanced output features; the enhancement network takes illumination weights of different color channels and local binary pattern features of the input image as illumination constraints, and enhances the enhanced special features to obtain the enhanced output features; and inputting the salient object detection special features and the enhanced output features into a salient object detection network to detect a salient object.
Owner:TIANJIN UNIV

A fully supervised salient target detection method

The present application relates to a kind of full supervision's salient object detection method, constructs complete multi-branch feature fusion refinement network MFFRNet as salient object detection model;Again training set in data set is input to the proposed MFFRNet model training, every time completing a round will be back propagated once, to optimize MFFRNet model parameter;With data set test set, the performance of model is evaluated;Finally, the model after evaluation is used for salient object detection.The model effectively fuses the detail information of low-level feature and the semantic information of high-level feature.The module designed for low-level feature utilizes asymmetric convolution to reduce background noise and other interference factors, and a module designed for high-level feature obtains rich semantic information.Meanwhile, aliasing effects caused by frequent up-sampling are effectively handled.The method effectively captures salient objects and obtains saliency prediction map, and has strong robustness.
Owner:SHANGHAI INST OF TECH

Unified architecture for interactive and salient segmentation of objects in videos and images

PCT designated stageWO2026013604A1Image enhancementImage analysisPattern recognitionSalient objects
A method and an electronic apparatus for performing unified segmentation of media content are provided. The method includes: determining a guidance map for an input frame based on a salient object from a past frame output mask and user-interacted objects in the media, operating in either salient mode or selective mode. The input frame of the media is cropped based on the guidance map and the salient ROIs of the salient object. A weighted grayscale image of the cropped frame is generated from the past frame output mask. A fused spatio-color mesh grid representation of the cropped frame in YUV format is determined. The cropped image frame, along with the weighted grayscale image and the fused spatio-color mesh grid representation, is input into a segmentation model. The segmentation model generates either a salient object segmentation or a user-interacted object segmentation for the media.
Owner:SAMSUNG ELECTRONICS CO LTD

An RGB-D salient object detection method based on boundary deformable convolution guidance

The application discloses an RGB-D salient object detection method based on boundary deformable convolution guidance, comprising the following steps: step one, respectively extracting features of an RGB mode and a depth map mode; step two, fusing the features of the two modes through a cross-modal attention fusion feature module to mine common and complementary features of salient objects; step three, inputting the feature map into an encoder deep layer embedded with an adjacent multi-scale feature enhancement module to obtain global context feature information; step four, generating a boundary clue map of the salient objects by constructing a boundary feature extraction module; and step five, generating a saliency map by using the generated boundary clue map and deformable convolution guidance. The application mines and strengthens the commonness of salient objects by cross-fusion of the depth map and the RGB image, effectively captures salient objects with different sizes and uncertain quantities by using adjacent level feature interaction, and solves the boundary blur problem of the saliency map by using the edge clue map to guide the model decoding.
Owner:ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE

Salient object detection method, device and equipment

PendingCN121305118ACharacter and pattern recognitionImaging processingSalient objects
The invention discloses a saliency object detection method, device and equipment, and belongs to the technical field of image processing. The saliency object detection method comprises the following steps: acquiring a first image; dividing the first image into a plurality of color areas according to the pixel values of the pixel points in the first image; wherein the pixel value similarity of the pixel points in the same color area is greater than or equal to a pixel value similarity threshold; extracting at least one feature of each color area; determining a contrast of the first color region relative to the first feature; wherein the first color region is any one color region in the plurality of color regions, and the first feature is any one feature in the at least one feature; determining a histogram contrast of the first color region; determining a first comprehensive contrast ratio of the first color area according to the contrast ratio of the first color area relative to the first feature and the histogram contrast ratio of the first color area; and determining a salient object in the first image according to the first comprehensive contrast of the plurality of color regions.
Owner:VIVO MOBILE COMM CO LTD

Floatation froth stability estimation method based on infrared time-series salient object segmentation

The application provides a flotation froth stability estimation method based on infrared time sequence significant target segmentation. First, a U-shaped network based on a ConvNeXt network design is used, a ConvLSTM network is embedded in the encoder to realize time sequence infrared saliency information extraction, and a cross attention mechanism is added to the ConvNeXt network to realize initial positioning of the infrared saliency area. Second, a residual refinement network with a U-shaped encoder-decoder structure is constructed, the residual between the deep learning saliency map and the true value is learned to improve the edge details of the saliency area, and fine segmentation of the time sequence significant target is realized. Finally, according to the significant target segmentation result, the froth stability is calculated, the deviation and abnormal threshold of the froth stability in the time sequence under different working conditions are counted, the significant area of the froth infrared video image is detected to realize the segmentation of the merged and broken bubbles, and the froth stability is evaluated according to the segmentation result.
Owner:FUZHOU UNIV

Unified architecture for interactive and salient segmentation of objects in videos and images

PendingUS20260051065A1Image enhancementImage analysisPattern recognitionSalient objects
A method and an electronic apparatus for performing unified segmentation of media content are provided. The method includes: determining a guidance map for an input frame based on a salient object from a past frame output mask and user-interacted objects in the media, operating in either salient mode or selective mode. The input frame of the media is cropped based on the guidance map and the salient ROIs of the salient object. A weighted grayscale image of the cropped frame is generated from the past frame output mask. A fused spatio-color mesh grid representation of the cropped frame in YUV format is determined. The cropped image frame, along with the weighted grayscale image and the fused spatio-color mesh grid representation, is input into a segmentation model. The segmentation model generates either a salient object segmentation or a user-interacted object segmentation for the media.
Owner:SAMSUNG ELECTRONICS CO LTD

A method for detecting salient targets in images

This application provides a method for detecting salient objects in images. It constructs a detection model based on a lightweight Mobilenetv2 backbone network and introduces fusion side connections into this backbone network to progressively fuse features from each layer. The method predicts salient objects at multiple scales and performs supervised learning, effectively avoiding overfitting. The proposed method constructs a lightweight detection model, and by introducing fusion side connections, it fully integrates features from each layer, making the model's performance comparable to larger existing models. The lightweight and high-performance detection model constructed in this application is suitable for deployment on television terminals and can be applied to television application scenarios such as visual object tracking and intelligent picture quality settings.
Owner:HISENSE ELECTRONIC TECH (WUHAN) CO LTD

Layout extraction system for regional annotation of images

PendingUS20260141529A1Natural language translationImage enhancementSalient objectsRadiology
A system may access an input image. The system may generate a plurality of segments based on one or more segmentation models and the input image, each segment from among the plurality of segments representing a corresponding salient object. The system may generate a depth map based on a depth estimation model. The system may layer the plurality of segments, based on the depth map and border regions between pairs of segments, to generate a plurality of ordered segments. The system may execute a vision-language model to generate a text annotation of the image based on the plurality of ordered segments.
Owner:REVE AI INC

A method for cooperative salient object detection and storage medium

The application provides a kind of synergistic salient object detection method and storage medium, the method is realized by salient feature enhancement and global information guidance, constructs synergistic salient object detection model, in down-sampling network, image feature is extracted by VGG16 backbone network, and the saliency of image feature is enhanced using coordination attention module, and dynamic convolution collaborative search module is used to search common salient object feature as synergistic feature, in up-sampling network, receptive field inflation technology is used to increase receptive field, and long-distance dependence information of image is obtained by non-local module, to optimize synergistic feature, and as the input of global information guidance fusion module, to reduce non-salient background interference. Finally, the whole synergistic salient object detection model is optimized by loss function. The method is fast in operation, and the final synergistic salient object prediction result is complete in structure and accurate in target.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Iteratively applying neural networks to automatically segment objects portrayed in digital images

The present disclosure relates to systems, method, and computer readable media that iteratively apply a neural network to a digital image at a reduced resolution to automatically identify pixels of salient objects portrayed within the digital image. For example, the disclosed systems can generate a reduced-resolution digital image from an input digital image and apply a neural network to identify a region corresponding to a salient object. The disclosed systems can then iteratively apply the neural network to additional reduced-resolution digital images (based on the identified region) to generate one or more reduced-resolution segmentation maps that roughly indicate pixels of the salient object. In addition, the systems described herein can perform post-processing based on the reduced-resolution segmentation map(s) and the input digital image to accurately determine pixels that correspond to the salient object.
Owner:ADOBE INC

Video salient object detection model training method and device, electronic equipment and storage medium

This application discloses a training method, apparatus, electronic device, and computer-readable storage medium for a video salient object detection model. Addressing the issues of insufficient multimodal fusion, temporal instability, and high label dependence, a two-stage training framework is proposed: The first stage uses cross-modal unsupervised contrastive learning to mine consistency and complementarity information between RGB and deep modalities, generating and iteratively optimizing salient object pseudo-labels to improve their quality; the second stage uses the optimized salient object pseudo-labels as supervision signals, selecting historical frames and adjacent frames to construct a reference set, training the target model through temporal feature fusion, and iteratively updating network parameters, enabling the model to obtain a stable representation in the time dimension, enhancing its ability to model long-term and short-term dependencies, suppressing dynamic interference, and maintaining target continuity. This method eliminates the need for manually labeled data, effectively reducing data costs through cross-modal contrastive learning and temporal feature fusion, while improving detection accuracy and robustness in complex scenes.
Owner:KEENON ROBOTICS CO LTD

RGB-T image saliency object detection method based on Mmba feedback iterative network

PendingCN121708283ACharacter and pattern recognitionSaliency mapSalient objects
The invention discloses an RGB-T image saliency object detection method based on a Mama feedback iterative network, and belongs to the field of computer vision. According to the network, a Mama encoder with double branches is adopted to extract multiple scales of features; then, a cross-layer feature fusion module is used for integrating features of a subsequent layer and features of a current layer, so that cross-scale correlation among multi-scale features extracted by Mama is enhanced, and significance performance of significant objects under different scales is enhanced; after cross-layer feature fusion, the feature enhancement module is then used for further extracting salient object information and increasing the proportion of salient features in a feature space; after the two modals are processed through the feature enhancement module to obtain refined features, the multi-modal feature fusion module combines the corresponding features in each layer to generate fused features, so that stronger semantics and details are shown; and finally, the generated features are sent to a feedback iteration architecture for two additional iterations to generate a clearer and more complete saliency map. The method is used for solving the common problems of feature detail loss, serious noise interference, poor physical consistency and the like of the saliency object detected in the prior art.
Owner:HARBIN ENG UNIV

A method for detecting a salient object based on a multi-scale dilated convolutional neural network

This invention discloses a salient object detection method based on a multi-scale dilated convolutional neural network. The method includes: extracting multi-scale features from the input image; inputting the multi-scale features into a dilated residual convolutional module to obtain fused features including contextual information of the multi-scale features; inputting the fused features into multiple channel attention modules to obtain multiple salient features; performing dimensionality reduction activation on each salient feature to generate a saliency map; and performing deep supervised training using a hybrid loss function that combines cross-entropy and cross-union loss. The method of this invention, based on a multi-scale dilated convolutional neural network, fully captures rich global and local semantic information in the image by using a dilated residual convolutional module, solving the problem of shallow encoder depth and insufficient information extraction. Simultaneously, the designed channel attention modules enable the network to focus on the target region, effectively improving the accuracy of object detection.
Owner:HEBEI HANGUANG HEAVY IND

Salient target detection method, device and system and electronic equipment

The invention provides a saliency target detection method, device and system and electronic equipment, and belongs to the field of computer vision. The method comprises the following steps: acquiring a first visible light image and a first thermal infrared image of a to-be-detected object; inputting the first visible light image and the first thermal infrared image into a pre-trained first model to obtain a first saliency target image of the to-be-detected object; the first model is an image fusion neural network model determined based on a second visible light image and a second thermal infrared image of the training object; the first model comprises a self-adaptive enhancement module, a coding and fusion module and a three-stream differential cooperative decoder. In conclusion, the technical scheme provided by the invention can progressively solve the technical problems of low detection precision, weak anti-interference capability, poor fusion effect and the like of the existing method layer by layer from three core links of input enhancement, feature fusion and decoding collaboration, improves the precision of saliency target detection, and can adapt to various application scenes.
Owner:NINGBO PORT INFORMATION COMM CO LTD +1

Generating image object segmentations utilizing graph-cut partitioning in self-supervised object discovery

The present disclosure is directed toward systems, methods, and non-transitory computer readable media that provide self-supervised object discovery systems that combine motion and appearance information to generate segmentation masks from a digital image or digital video and delineate one or more salient objects within the digital image / digital video. The disclosed systems utilize a neural network encoder to generate a fully connected graph based on image patches from the digital input, incorporating image patch feature and optical flow patch feature similarities to produce edge weights. The disclosed systems partition the generated graph to produce a segmentation mask. Furthermore, the disclosed systems iteratively train a segmentation network based on the segmentation mask as a pseudo-ground truth via a bootstrapped, self-training process. By utilizing both motion and appearance information to generate a bi-partitioned graph, the disclosed systems produce high-quality object segmentation masks that represent a foreground and background of digital inputs.
Owner:ADOBE INC

Identifying salient regions based on multi-resolution partitioning

In implementation of techniques for generating salient regions based on multi-resolution partitioning, a computing device implements a salient object system to receive a digital image including a salient object. The salient object system generates a first mask for the salient object by partitioning the digital image into salient and non-salient regions. The salient object system also generates a second mask for the salient object that has a resolution that is different than the first mask by partitioning a resampled version of the digital image into salient and non-salient regions. Based on the first mask and the second mask, the salient object system generates an indication of a salient region of the digital image using a machine learning model. The salient object system then displays the indication of the salient region in a user interface.
Owner:ADOBE INC

A method for detecting salient objects based on semantic information guidance

The present application relates to the technical field of image processing, in particular to a salient object detection method based on semantic information guidance. The present application constructs a salient object detection model based on semantic information guidance; divides the pictures in a salient image data set into a training set, a validation set and a test set of the salient object detection model; trains the salient object detection model; inputs the test set into the trained salient object detection model to obtain four evaluation indexes; when the four evaluation indexes meet the actual application requirements, the corresponding salient object detection model is used for salient object detection of images; otherwise, the learning rate is adjusted, the salient object detection model is retrained until the four evaluation indexes of the salient object detection model meet the actual application requirements. The salient object detection model constructed by the present application realizes the purpose of completely segmenting out salient objects and keeping accurate details, and improves the overall feature extraction capability.
Owner:HENAN UNIVERSITY

Automatic extraction of salient objects in virtual environments for object modification and transmission

ActiveUS12586341B2Image analysisAnimationVirtualizationSalient objects
Automatic extraction of salient object properties in virtual environments for object modification and transmission. In some implementations, a computer-implemented method includes determining a reference avatar and obtaining properties of an object in the virtual environment, the properties including spatial, visual, and / or audio properties. Saliency factors of the object are determined, each saliency factor normalized to a numeric range and based on a different set of properties of the object, where one or more saliency factors are additionally based on a property of the reference avatar. A saliency measure of the object is determined with respect to the reference avatar based on a combination of the saliency factors. If the saliency measure is greater than a threshold saliency measure, the reference avatar or object are automatically modified in the virtual environment based on the object, and otherwise the modification is omitted.
Owner:ROBLOX CORP

Method for detecting salient objects from rgb-d images based on cross-modal interaction and correction

The present application relates to a kind of RGB-D image saliency target detection method based on cross-modal interaction and revision, comprising:1, in the encoding stage, color image encoder and depth map encoder respectively extract the features of color image modal and depth map modal, the high-level features of color image modal and depth map modal are gradually guided to be integrated unit and are cross-modal interaction and obtain RGB-D feature;2, feature revision intermediate structure is revised to the feature of color image modal, depth map modal and RGB-D modal obtained in the encoding stage in self-modal and cross-modal;3, in the decoding stage, color image modal and depth map modal are decoded respectively, and each level decoding feature is sent into importance gate fusion unit and is fused decoding, to complete the decoding of RGB-D modal, obtain the final saliency map.The present application respectively in different stages carries out interaction and revision to feature, realizes the more comprehensive fusion of two kinds of modal and the extraction of complementary information.
Owner:BEIJING JIAOTONG UNIV

Method for detecting salient objects in optical remote sensing images based on edge-guided enhancement network

The application discloses an optical remote sensing image salient object detection method based on an edge guidance enhanced network, which comprises the steps of image preprocessing, generating a predicted edge map tensor E, establishing an EGE network, and generating a saliency map using the EGE network. The application makes full use of edge information, accurately identifies the object boundary using a lightweight network, and generates a predicted edge map. The processed backbone features are gradually fused between high-level and low-level, the backbone features are concentrated on the target area, and interference sources are excluded. These fused features are integrated into deeper features, missing target information in the background is explored through a reverse attention mechanism, and missing target details in the background are better extracted.
Owner:HEBEI NORMAL UNIV

Multi-modal false news detection method combined with multi-channel causal attention mechanism

PendingCN121302114ABiological modelsPattern recognitionSalient objects
The invention particularly relates to a multi-mode false news detection method combined with a multi-channel causal attention mechanism, which comprises the following steps of: extracting noun information of a text by adopting a word segmentation tool, detecting a salient object in an image by utilizing a target detection model, and extracting key elements in two modes of the text and the image; and generating an intervention text and an intervention image by masking the text nouns and the image objects one by one. Text features of the text and the intervention text are extracted through a pre-training language model, and image features of the image and the intervention image are extracted through a pre-training image model. A cross-modal cross attention structure is introduced into the shared semantic space, and cross-modal fusion features are generated. And calculating a causal score of each key element, and feeding back the causal score to an attention mechanism to further guide attention distribution. A multi-channel attention network is constructed, and various attention information is dynamically fused through a gating mechanism. Finally, causal-guided fusion features are used as input, and a cross entropy loss function is adopted to carry out classification training and model optimization.
Owner:XIAN UNIV OF POSTS & TELECOMM

An image recognition detection method and system based on weakly supervised learning

The application relates to the technical field of object recognition, and discloses an image recognition detection method and system based on weak supervision learning, which comprises the following steps: S1, generating initial pseudo labels of a to-be-detected image through point annotation, an edge detector and a self-adaptive flood filling algorithm; S2, constructing a point supervision salient object detection model based on a visual converter; S3, performing the first round of training on the point supervision salient object detection model to obtain a final salient map; and S4, suppressing non-salient objects of the final salient map, performing the second round of training on the point supervision salient object detection model, and completing the recognition of the image. The application solves the problem that the existing image recognition detection technology can only cover part of an image, cannot ignore invalid objects, leads to low efficiency, and has the characteristics of strong supervision capability.
Owner:GUANGDONG TOBACCO SHANWEI CO LTD

Lightweight RGB-D underwater salient object detection method based on frequency domain decoupling fusion

The application discloses a lightweight RGB-D underwater salient object detection method based on frequency domain decoupling fusion, inputs an RGB image and a depth image into a frequency domain decoupling fusion network, first extracts RGB features and depth features through a double-flow encoder. Then, a frequency domain decoupling fusion method is used to obtain double-mode fusion features, and then a hierarchical weighted fusion module is used for adaptive weighted fusion to obtain multi-level semantic features. Finally, a saliency prediction map is obtained based on the multi-level semantic features. The application innovatively designs a frequency domain decoupling fusion module, which effectively improves the efficiency and quality of feature fusion by decoupling analysis and adaptive fusion of depth features and RGB features in the frequency domain space. Secondly, the application proposes a hierarchical weighted fusion mechanism, which dynamically evaluates the relative contribution of adjacent level features, realizes adaptive weighted fusion of features, and significantly enhances the detection ability of the model to multi-scale salient objects.
Owner:LISHUI RES INST OF HANGZHOU UNIV OF ELECTRONIC SCI & TECH +1

RGB-D image saliency object detection method and system based on joint prior guidance

The invention discloses an RGB-D image saliency object detection method and system based on joint prior guidance, and belongs to the technical field of image processing. Self-adaptive effective fusion of cross-level and cross-modal features is realized under the guidance of boundary priori and main body priori, salient detail information is reserved, and the expression and generalization ability of a priori region enhancement model for a salient object key boundary region is enhanced. And finally, the definition of the boundary of the saliency object detection result of the RGB-D image and the integrity of the main body are improved. The method is used for solving the problem of how to effectively carry out cross-modal feature fusion according to guidance of prior information and the problem of how to improve the expression ability of a model to a boundary by utilizing prior guidance.
Owner:HARBIN ENG UNIV

Method for visible-thermal infrared salient object detection based on correlation modeling

The application discloses a visible light-thermal infrared saliency target detection method based on correlation modeling, gradually models the correlation of two modes from the space, region and image level, and predicts accurate saliency detection results in a non-aligned visible light-thermal infrared image pair. At the space level, the correlation modeling is guided to pay attention to the saliency region and align the salient objects in the two modes by embedding an adapter containing salient object semantic information. At the region level, the partial correlation between the two modes is modeled for the aligned salient region. At the image level, the region-level correlation is extended to the whole image of the visible light mode to predict the accurate salient object region in the visible light image. The application achieves the best effect on the non-aligned, weak-aligned and aligned saliency target detection datasets.
Owner:ANHUI UNIV

A method, device and electronic equipment for detecting salient objects based on multi-modal prediction reconstruction error

The application provides a salient object detection method and device based on multi-modal prediction reconstruction error and an electronic device. The method comprises: acquiring images of at least one mode and establishing a mode existence mask vector; performing feature coding on each existing mode image to obtain multi-scale features, and performing feature fusion to obtain a multi-scale global scene representation; performing self-prediction reconstruction and cross-modal prediction reconstruction based on the multi-scale global scene representation to obtain self-reconstruction results and cross-modal reconstruction results; calculating multi-scale self-reconstruction errors and multi-scale cross-modal reconstruction errors corresponding to each mode to obtain multi-scale error maps and initial saliency maps; constructing a gaze area and cropping image blocks from each mode image to calculate initial local saliency maps, and obtaining a refined saliency map through iterative refinement; calculating a gating saliency map corresponding to a global error feature vector, and fusing the gating saliency map and the refined saliency map to obtain a final salient object detection result.
Owner:HANGZHOU DIANZI UNIV

Salient target detection method and device based on multi-modal prediction reconstruction error, and electronic equipment

The invention provides a saliency target detection method and device based on a multi-modal prediction reconstruction error and electronic equipment. The method comprises the following steps: acquiring an image of at least one modal and establishing a modal existence mask vector; performing feature coding on each existing modal image to obtain multi-scale features, and performing feature fusion to obtain multi-scale global scene representation; performing self-prediction reconstruction and cross-modal prediction reconstruction based on the multi-scale global scene representation to obtain a self-reconstruction result and a cross-modal reconstruction result; calculating a multi-scale self-reconstruction error and a multi-scale cross-modal reconstruction error corresponding to each modal to obtain a multi-scale error graph and an initial saliency graph; constructing a staring area, cutting image blocks from each modal image to calculate an initial local saliency map, and obtaining a refined saliency map through iterative refinement; and calculating a gated saliency map corresponding to the global error feature vector, and fusing the gated saliency map and the refined saliency map to obtain a final saliency target detection result.
Owner:HANGZHOU DIANZI UNIV