Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

60 results about "Significance map" patented technology

Equipment fault mode identification and diagnosis method based on deep learning

The invention relates to an equipment fault mode identification and diagnosis method based on deep learning, and aims to improve the accuracy and generalization ability of equipment fault diagnosis. A vibration signal, a temperature signal, an acoustic signal, a current signal and image data of equipment are synchronously acquired through a multi-modal data acquisition system, and weighted fusion is performed on different modal data by adopting a self-adaptive multi-head attention mechanism. And then, performing time sequence modeling by using a bidirectional long-short term memory network (BiLSTM), and finally outputting a fault type and a fault saliency map to help operation and maintenance personnel to position a fault area. And through a transfer learning technology, the adaptability and diagnosis precision of the model under different equipment and working conditions are further improved. The method can be widely applied to fault diagnosis and intelligent operation and maintenance of various devices, and the operation reliability and the maintenance efficiency of the devices are effectively improved.
Owner:BEIJING BOHUA XINZHI SCI & TECH +1

Navigation method and device based on multi-modal knowledge enhancement, equipment and medium

The invention relates to the technical field of intelligent navigation, financial science and technology and medical health, and discloses a navigation method, device and equipment based on multi-modal knowledge enhancement and a medium, and the method comprises the following steps: obtaining a navigation instruction, and analyzing the navigation instruction to obtain a target constraint, a spatial constraint and an entity constraint; generating a structured target description text based on the target constraint, the spatial constraint and the entity constraint; an input image is acquired, a semantic saliency map is generated based on the input image and the structured target description text, and the semantic saliency map comprises a plurality of candidate areas; calculating a score of each candidate region of the semantic saliency map based on a preset scene knowledge graph to obtain a fusion score; and performing path planning according to the fusion score of each candidate region to obtain a target path. Navigation can be realized without navigation data training, and meanwhile, a potential target can be focused according to semantics, so that the navigation precision is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Adaptive diffusion image editing method and system based on concept attention

The invention discloses a self-adaptive diffusion image editing method and system based on concept attention, and the method comprises the following steps: constructing a paired data set; analyzing the editing instruction, and extracting a key concept; a pre-trained T5 language model is utilized to convert the key concept into text embedding, and the text embedding is mapped to an image feature space; modifying a diffusion model based on a Transform architecture, embedding a concept attention module in an attention layer of a multi-modal diffusion converter, calculating an attention score between image features and concept embedding, and generating a concept saliency map; in the denoising process, the weight of the target area is adjusted by using the concept saliency map so as to realize accurate editing. According to the method, under the condition that the global image quality is not affected, the editing precision can be improved, interference to a non-target area is reduced, and meanwhile, reinforcement learning and real-time feedback are combined, so that the model can be adaptively optimized, and an editing result better meeting the user requirement is generated.
Owner:NANJING UNIV OF POSTS & TELECOMM

Saliency-guided time sequence adversarial sample generation method and system

The invention discloses a saliency-guided time sequence adversarial sample generation method and a saliency-guided time sequence adversarial sample generation system, and the method comprises the steps: obtaining a target time sequence classification model and a corresponding original input sample, carrying out the supervised training of the model through a time sequence training set, so as to guarantee the prediction accuracy of the model for the original input sample, and setting a disturbance iteration parameter; generating a saliency map of the input sample based on the target time sequence classification model; constructing a composite loss function fusing classification loss and saliency alignment loss for the target time sequence classification model based on the saliency map, carrying out iterative optimization on the model based on a projection gradient descent method framework and in combination with a saliency guidance strategy to update disturbance, terminating the optimization process to obtain updated disturbance when an iteration condition is reached, and carrying out the optimization of the target time sequence classification model. And outputting a final confrontation sample based on the obtained disturbance, and completing the generation of the confrontation sample of the time sequence. The attack success rate and the non-concealment of the time sequence confrontation sample are considered at the same time.
Owner:WUHAN UNIV

Method and system for judging corrosion condition of electrode foil

The invention belongs to the technical field of condition judgment, and particularly relates to an electrode foil corrosion condition judgment method and system, and the method comprises the following steps: S1, obtaining a grayscale image of a to-be-analyzed electrode foil; constructing a structure tensor based on the composite gradient reflecting the local brightness and texture information of the pixel points, determining an anisotropic diffusion coefficient according to the structure tensor, and performing iterative anisotropic diffusion filtering processing on the grayscale image by using the anisotropic diffusion coefficient to obtain a filtered image; and S2, fusing the anisotropic diffusion coefficient determined for each pixel point with the pixel intensity of the filtered image to obtain a high-contrast corrosion significance map. According to the method, the contour edge information of the corrosion area can be kept while the complex texture and noise of the background area of the electrode foil image are smoothed, the contradiction between denoising and edge protection of a traditional filtering method is solved, the result of the corrosion condition is more reliable, and the automation level of electrode foil product quality detection is improved.
Owner:HUBEI FUYIDA ELECTRONIC TECH CO LTD

Coding of significance maps and transform coefficient blocks

A higher coding efficiency for coding a significance map indicating positions of significant transform coefficients within a transform coefficient block is achieved by the scan order by which the sequentially extracted syntax elements indicating, for associated positions within the transform coefficient block, as to whether at the respective position a significant or insignificant transform coefficient is situated, are sequentially associated to the positions of the transform coefficient block, among the positions of the transform coefficient block depends on the positions of the significant transform coefficients indicated by previously associated syntax elements. Alternatively, the first-type elements may be context-adaptively entropy decoded using contexts which are individually selected for each of the syntax elements dependent on a number of significant transform coefficients in a neighborhood of the respective syntax element, indicated as being significant by any of the preceding syntax elements.
Owner:DOLBY VIDEO COMPRESSION LLC

An Interactive Salience Mining Method for RGB-D Salient Object Detection

The present invention belongs to the technical field of computer vision and image processing, and in particular relates to an interactive saliency mining method for RGB-D salient object detection, comprising the following steps: S1, using a two-stream encoder network to extract multi-level cross-modal features of RGB and depth images; S2, proposing a cross-modal interaction module to achieve the interaction and aggregation of these different features through a series of matrix operations; S3, attempting to separate the saliency perception information from the complex environment and establishing a fusion method between adjacent features under the guidance of a relatively rough saliency map; S4, further extracting context information from the saliency perception information and background information. The present invention obtains multi-layer cross-modal fusion features in the encoding stage, uses a progressive saliency mining module to decode multi-level features, gradually filters out complex background interference to make the saliency map more accurate, gradually refines the saliency region to make the saliency map more precise, and realizes the optimal decoding process.
Owner:CHANGCHUN UNIV OF SCI & TECH

Coding of significance maps and transform coefficient blocks

A higher coding efficiency for coding a significance map indicating positions of significant transform coefficients within a transform coefficient block is achieved by the scan order by which the sequentially extracted syntax elements indicating, for associated positions within the transform coefficient block, as to whether at the respective position a significant or insignificant transform coefficient is situated, are sequentially associated to the positions of the transform coefficient block, among the positions of the transform coefficient block depends on the positions of the significant transform coefficients indicated by previously associated syntax elements. Alternatively, the first-type elements may be context-adaptively entropy decoded using contexts which are individually selected for each of the syntax elements dependent on a number of significant transform coefficients in a neighborhood of the respective syntax element, indicated as being significant by any of the preceding syntax elements.
Owner:DOLBY VIDEO COMPRESSION LLC

Methods and apparatus for unified significance map coding

Methods and apparatus are provided for unified significance map coding. An apparatus includes a video encoder (400) for encoding transform coefficients for at least a portion of a picture. The transform coefficients are obtained using a plurality of transforms. One or more context sharing maps are generated for the transform coefficients based on a unified rule. The one or more context sharing maps are for providing at least one context that is shared among at least some of the transform coefficients obtained from at least two different ones of the plurality of transforms.
Owner:INTERDIGITAL MADISON PATENT HLDG

An RGB-D Salient Object Detection Method Based on a Multimodal Difference Fusion Network

The present invention provides an RGB-D salient object detection method based on a multi-modal difference fusion network, belonging to the field of image saliency detection technology. The method uses a Swin Transformer to extract RGB and Depth features containing global context information for inferring salient objects in the scene. The present invention mainly explores the differences between the RGB and Depth modalities to analyze the connections and differences of saliency in these two modalities, and designs a difference fusion network to fuse cross-modal features for capturing complete salient objects. The present invention includes the following steps: (1) using a Swin Transformer to extract cross-modal features; (2) using a bidirectional fusion method to fuse RGB and Depth features to generate a Fusion stream; (3) using a three-stream difference supervision mechanism to obtain the differences between modalities; (4) using this difference to fuse cross-modal features; (5) using a target cascade aggregation decoder to perform saliency inference and decoding on the fused cross-modal features to generate a predicted saliency map.
Owner:ANHUI UNIV OF SCI & TECH

A model pruning method of an interpretable CNN classification model

The application relates to a model pruning method of an interpretable CNN classification model, belongs to the field of image compression, and solves the problems of high operation complexity, large time and memory consumption and difficulty in deployment on terminal equipment of an existing deep CNN model, and solves the problem of lack of interpretability of an existing model pruning algorithm. The method comprises the following steps: inputting a training picture into a neural network model to be pruned, and extracting a feature map matrix of each convolution layer; upsampling the feature map matrix to the size of the input picture, and then performing a normalization operation to construct a saliency map; multiplying the saliency map and the input picture element by element to construct a weighted input picture; subtracting the input picture from the weighted picture element by element to construct an attention region occlusion map; inputting the attention occlusion map into the model to be pruned, observing the change of model accuracy as an importance score of the channel, and pruning the channel to obtain a pruned lightweight model. The application realizes high pruning rate of the model and improves the interpretability of the pruning process.
Owner:DALIAN UNIV OF TECH

A salient object detection method in panoramic images based on multi-projection representation

The present invention relates to a method for detecting salient objects in panoramic images based on multi-projection representation. An end-to-end detection network with an encoder-decoder structure is constructed, using an equirectangular projection image and four corresponding cube-expanded images as inputs to the detection network. In the encoder stage, the equirectangular projection branch and the cube-expanded branch extract features using a fifty-layer deep residual network (ResNet-50) with shared parameters. In the decoder stage, a dynamic weighted fusion module adaptively fuses the equirectangular projection features and the four cube-expanded features, while a filtering and refinement module combines the encoding and decoding features to obtain a final saliency map. In this invention, the detection network combines two panoramic image representation methods, equirectangular projection and cube-expanded, using the equirectangular projection image and the four corresponding cube-expanded images as inputs. The cube-expanded images provide supplementary information to the equirectangular projection images, ensuring target integrity.
Owner:BEIJING JIAOTONG UNIV

Multi-level significance maps for encoding and decoding

Methods of encoding and decoding for video data are described in which multi-level significance maps are used in the encoding and decoding processes. The significant-coefficient flags that form the significance map are grouped into contiguous groups, and a significant-coefficient-group flag signifies for each group whether that group contains no non-zero significant-coefficient flags. If there are no non-zero significant-coefficient flags in the group, then the significant-coefficient-group flag is set to zero. The set of significant-coefficient-group flags is encoded in the bitstream. Any significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is non-zero are encoded in the bitstream, whereas significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is zero are not encoded in the bitstream.
Owner:VELOS MEDIA LLC

Unsupervised semantic segmentation method and system based on saliency map and deep learning

The present application relates to the field of image processing technology, and in particular, to an unsupervised semantic segmentation method and system based on saliency maps and deep learning. The present application does not need to rely on strong or weak supervision implementation methods. Without using labeled data, it generates a clear and corrected foreground mask through unsupervised representation learning, combined with a pixel-level target attention classification head and a saliency map, so that the outline of the corrected foreground mask is clearer, thereby achieving unsupervised semantic segmentation. While improving the quality of pseudo-labels and the accuracy of semantic segmentation by correcting the foreground mask, it effectively reduces the system's dependence on manual annotation. At the same time, it can also be used as an automatic annotation system to automatically annotate image data sets, significantly reducing the labor cost of data annotation.
Owner:SUN YAT SEN UNIV

Multi-Level Significance Maps for Encoding and Decoding

Methods of encoding and decoding for video data are described in which multi-level significance maps are used in the encoding and decoding processes. The significant-coefficient flags that form the significance map are grouped into contiguous groups, and a significant-coefficient-group flag signifies for each group whether that group contains no non-zero significant-coefficient flags. If there are no non-zero significant-coefficient flags in the group, then the significant-coefficient-group flag is set to zero. The set of significant-coefficient-group flags is encoded in the bitstream. Any significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is non-zero are encoded in the bitstream, whereas significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is zero are not encoded in the bitstream.
Owner:VELOS MEDIA LLC

Video image bitrate allocation method, system and device, and storage medium

Disclosed in embodiments of the present application are a video image bitrate allocation method, system and device, and a storage medium. In the technical solution provided by the embodiments of the present application, a saliency map of a target image is acquired, and real-time coding quality information of a target region in the target image is determined on the basis of the saliency map, wherein the target region is a region of interest (ROI) or a non-ROI; a coding quality redundancy amplitude of the target region is determined on the basis of the real-time coding quality information and a set quality discrimination threshold, and an offset bitrate of the target region is determined on the basis of the coding quality redundancy amplitude; and an initial bitrate of each coding block in the target region is acquired, and a coding bitrate of each coding block in the target region is configured on the basis of the initial bitrate and the offset bitrate. By using the described technical means, the bitrates of the ROI and the non-RON can be adaptively configured on the basis of the video content attribute of the video image, thereby reducing bandwidth costs while ensuring the coding quality, optimizing the video coding effect.
Owner:PENINSULA INFORMATION TECH INC

Method and device for spatially resolved localization of a defect in a component

Method for spatially resolved localization of a defect in a component, the method comprising the steps: - Providing (S1) a salience map of a NOK image or a NOK fault class image of a component exhibiting a defect; - Apply (S2) at least one dimension reduction method to the provided salience map to generate a dimensionally reduced salience map; and - spatially resolved localization (S3) of the component defect based on the dimensionally reduced salience map.
Owner:ROBERT BOSCH GMBH

A learning-based saliency localization method

The present application relates to the technical field of map saliency localization, and particularly relates to a saliency localization method based on learning. The saliency localization method based on learning provided by the present application adopts a saliency prediction model to simulate the mechanism in a SLAM framework. A saliency map is predicted according to a saliency model, and the model can capture scene semantics and geometric information. The value of the saliency map is used as the weight of a feature point in a traditional bundle adjustment method. Detailed experiments conducted on the KITTI and EuRoc datasets with the most advanced algorithms show that the algorithm proposed by the present application is superior to existing algorithms in indoor and outdoor environments, and significantly improves the positioning accuracy and robustness.
Owner:CHONGQING UNIV

A texture fabric defect detection method and medium based on an unsupervised mode

The present invention discloses a texture fabric defect detection method and medium based on an unsupervised model, comprising: equally dividing an input image into blocks to obtain a plurality of sub-images of equal size; using blocks at the edges of the sub-images as a set of texture background regions to be selected, obtaining a set of feature vectors of the selected texture background regions; removing outliers in the geometry of the selected texture background regions, using the remaining regions as texture backgrounds, and calculating texture background feature descriptors; traversing all blocks of the image and generating a block weight map based on the degree of deviation from the texture background; performing bilateral filtering on the input image to obtain a multi-channel center-surround mechanism saliency map; fusing the block weight map with the saliency map to generate a defect labeling map, thereby completing texture fabric defect detection. The present invention aims to improve the defect detection rate during the fabric defect detection process, and at the same time, improve the quantity and quality of the fabric defect sample library during the detection process.
Owner:SOUTH CHINA UNIV OF TECH +1

Electrode foil corrosion condition determination method and system

The present application belongs to the technical field of situation judgment, and particularly relates to a kind of electrode foil corrosion situation judgment method and system, comprising the following steps: S1, the gray scale image of electrode foil to be analyzed is obtained;Based on the composite gradient reflecting the local brightness and texture information of pixel point, the structure tensor is constructed, and the anisotropic diffusion coefficient is determined according to the structure tensor, the gray scale image is processed by iterative anisotropic diffusion filtering using the anisotropic diffusion coefficient, and the filtered image is obtained;S2, the anisotropic diffusion coefficient determined for each pixel point is fused with the pixel intensity of the filtered image to obtain a high-contrast corrosion significance map.The present application can smooth the complex texture and noise in the background area of electrode foil image while maintaining the outline edge information of the corrosion area, solve the contradiction between denoising and edge preservation in traditional filtering method, make the corrosion situation result more reliable, and improve the automation level of electrode foil product quality detection.
Owner:HUBEI FUYIDA ELECTRONIC TECH CO LTD

Apparatus and method for evaluating saliency map determiner

Apparatuses and methods for evaluating a saliency map determiner are provided. According to various embodiments, a method of evaluating a saliency map determiner is described, the method comprising: adding a predefined pattern to a plurality of training dataset units to train identification of a data class, wherein each training dataset unit comprises a representation of the data class to be identified; training a neural network with the plurality of training dataset units comprising the predefined pattern; determining, by the saliency map determiner, a saliency map for the data class; and evaluating the saliency map determiner based on whether the determined saliency map comprises a context of the data class introduced by the addition of the predefined pattern.
Owner:ROBERT BOSCH GMBH

Coding method, compression network training method and related equipment

The invention discloses a coding method, a compression network training method and related equipment, and belongs to the field of image processing, and the coding method comprises the steps: carrying out the saliency feature extraction of a to-be-coded feature map of a to-be-coded image, and obtaining a target saliency map, the target saliency map is used for representing the importance of each region of the to-be-coded image; performing coding processing on the to-be-coded feature map based on the target saliency map to obtain a coded code stream; wherein the coding processing comprises the following steps: determining a code rate weight of the to-be-coded feature map according to the target saliency map; and performing compression coding on the to-be-coded feature map according to the code rate weight.
Owner:VIVO MOBILE COMM CO LTD

Open vocabulary image semantic segmentation method based on hierarchical aggregation and prompt optimization

The invention discloses an open vocabulary image semantic segmentation method based on hierarchical aggregation and prompt optimization. Semantic segmentation of any open category target is achieved under the condition that extra training is not needed. The method comprises the following steps: firstly, preprocessing an input image, carrying out image semantic analysis in combination with a visual language model, and constructing a text prompt matched with image content; then, cross-modal attention information is extracted, a category saliency map is generated, and a soft discarding strategy is introduced in the saliency mining process to reduce out-of-distribution noise and artifact interference; further performing quality evaluation on significance results obtained by multiple rounds of iteration, and obtaining stable and consistent significance representation by adopting a layered weighted aggregation mode; and finally, automatically generating prompt information based on an aggregation result, inputting the prompt information into the segmentation model, and outputting a semantic segmentation mask of the target.
Owner:HOHAI UNIV

A multi-modal based salient object detection method and device, and related medium

The application discloses a kind of based on multimodal salient object detection method, device and related medium, the method includes inputting the color image to be detected into visual language model and carrying out multimodal feature processing, obtains semantic alignment feature set;Multi-granularity semantic reasoning is carried out to semantic alignment feature set, and positioning enhanced feature map is obtained;Again, detail feature map extraction is carried out to the color image to be detected by image feature neural network and detail feature map and positioning enhanced feature map are fused, and finally saliency map is obtained.Such that, saliency reasoning process can explicitly introduce semantic information, and is synergized using multi-granularity semantic expert system algorithm and detail recovery fusion algorithm, and the positioning accuracy of saliency map in complex scene is improved.
Owner:ZHEJIANG DIANCHUANG INFORMATION TECH CO LTD +1

Adversarial sample generation method based on multi-level significance regional knowledge distillation

The invention relates to an adversarial sample generation method based on multi-level salient region knowledge distillation, and belongs to the technical field of adversarial attacks in the field of computer vision. According to the method, key pixels for judging a single image by a neural network are positioned by utilizing multi-level saliency region information guidance, saliency region images of all feature layers are obtained, and a multi-level saliency image set is formed; and integrating salient region images in the salient image set by using a salient image set knowledge distillation strategy so as to generate a higher-quality adversarial attack. According to the method, significant differences of different levels are comprehensively considered, and the attack success rate and the concealment are improved; according to the method, a saliency region integration strategy based on data set distillation is designed, and the generalization ability of attack resistance is improved.
Owner:CHINESE PEOPLES LIBERATION ARMY UNIT 32801

A scale-invariant based co-salient object detection method

The application discloses a kind of based on scale invariance collaborative salient target detection method, it is related to monitoring safety technical field, method includes: obtaining the image sequence converted from video stream in monitoring system, it is input to the main network of pre-training model to extract multi-scale feature;Multi-scale feature is enhanced and fusion processing is generated initial saliency map;By residual refinement structure, restore details edge, generate accurate saliency map;Scale consistency constraint and semantic feature feedback are applied to optimization, obtain target saliency map and output.The application can effectively solve the problem that model lacks semantic guidance, generalization ability is poor in prior art to different resolution images, significantly improve the accuracy and robustness of salient target detection in monitoring system, applicable to a variety of complex monitoring scenes.
Owner:WUHAN INST OF TECH

Engine part defect detection method, device and equipment based on AI vision and medium

The application relates to an AI vision-based engine part defect detection method, device, equipment and medium. The method comprises the following steps: acquiring labeled real engine part images and defect masks to form initial samples, training by using a multi-scale topological perception feature extraction network, generating a deep feature map and a topological saliency map, and based on this, performing defect data coding and feature space synthesis to construct a large-scale synthetic augmented training dataset; taking the trained feature extraction network as a backbone, integrating a dynamic domain self-adaptive module to construct a defect detection framework, obtaining a robust teacher detection model and a parameter offset library through multi-working-condition simulation training, and performing two-stage progressive knowledge distillation processing to obtain a defect detection engine. The method significantly improves the accuracy, robustness and efficiency of engine part defect detection through multi-scale topological perception feature extraction, dynamic domain self-adaptation and progressive knowledge distillation.
Owner:张贺

Sparse adversarial disturbance generation method based on dynamic dimension

The invention provides a sparse adversarial disturbance generation method based on dynamic dimensions, and the method comprises the steps: 1) inputting an image sample collected by a visible light camera into a visual deep neural network model, and outputting the confidence of the sample belonging to each category; 2) calculating a confrontation saliency map according to the confidence coefficient, and selecting K pixels with the highest saliency as modifiable pixels; 3) transmitting the optimized gradient to the input image one by one by using a back propagation algorithm; 4) disturbing the modifiable pixels selected in the step 2) according to the gradient obtained in the step 3) to obtain a sparse adversarial sample; 5) verifying whether the sparse adversarial sample can successfully change the classification result of the deep neural network, and if the classification result is successfully changed or the number of modifiable pixels reaches an upper limit, exiting the algorithm; otherwise, increasing the number K of modifiable pixels, and taking the sparse adversarial sample obtained in the step 4) as the input of the step 1).
Owner:CHINA UNIV OF MINING & TECH

Image recognition method based on class activation map algorithm and related device

Embodiments of the present application relate to the technical field of artificial intelligence, and disclose an image recognition method based on a class activation map algorithm and related devices, the method comprising: obtaining an image to be recognized; inputting the image to be recognized into a class activation map algorithm based on multi-label gradient feedback, the class activation map algorithm based on multi-label gradient feedback performing feature convolution on the image to be recognized to obtain a multi-scale feature map, performing gradient backpropagation on feature maps of each label category when performing gradient backpropagation on a current category, determining an inter-category influence factor according to a gradient value set backpropagated by all label categories, determining an activation map weight of the current category according to the inter-category influence factor and an influence factor of different pixel positions of the current category, and obtaining a recognition result according to the activation map weight and the multi-scale feature map; wherein the current category is any category in a plurality of label categories; and outputting the recognition result. Embodiments of the present application improve the generation quality of a saliency map.
Owner:SHENZHEN UNIV

Methods and apparatus for unified significance map coding

Methods and apparatus are provided for unified significance map coding. An apparatus includes a video encoder (400) for encoding transform coefficients for at least a portion of a picture. The transform coefficients are obtained using a plurality of transforms. One or more context sharing maps are generated for the transform coefficients based on a unified rule. The one or more context sharing maps are for providing at least one context that is shared among at least some of the transform coefficients obtained from at least two different ones of the plurality of transforms.
Owner:INTERDIGITAL MADISON PATENT HLDG