Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

44 results about "Significance map" patented technology

Adaptive diffusion image editing method and system based on concept attention

The invention discloses a self-adaptive diffusion image editing method and system based on concept attention, and the method comprises the following steps: constructing a paired data set; analyzing the editing instruction, and extracting a key concept; a pre-trained T5 language model is utilized to convert the key concept into text embedding, and the text embedding is mapped to an image feature space; modifying a diffusion model based on a Transform architecture, embedding a concept attention module in an attention layer of a multi-modal diffusion converter, calculating an attention score between image features and concept embedding, and generating a concept saliency map; in the denoising process, the weight of the target area is adjusted by using the concept saliency map so as to realize accurate editing. According to the method, under the condition that the global image quality is not affected, the editing precision can be improved, interference to a non-target area is reduced, and meanwhile, reinforcement learning and real-time feedback are combined, so that the model can be adaptively optimized, and an editing result better meeting the user requirement is generated.
Owner:NANJING UNIV OF POSTS & TELECOMM

Saliency-guided time sequence adversarial sample generation method and system

The invention discloses a saliency-guided time sequence adversarial sample generation method and a saliency-guided time sequence adversarial sample generation system, and the method comprises the steps: obtaining a target time sequence classification model and a corresponding original input sample, carrying out the supervised training of the model through a time sequence training set, so as to guarantee the prediction accuracy of the model for the original input sample, and setting a disturbance iteration parameter; generating a saliency map of the input sample based on the target time sequence classification model; constructing a composite loss function fusing classification loss and saliency alignment loss for the target time sequence classification model based on the saliency map, carrying out iterative optimization on the model based on a projection gradient descent method framework and in combination with a saliency guidance strategy to update disturbance, terminating the optimization process to obtain updated disturbance when an iteration condition is reached, and carrying out the optimization of the target time sequence classification model. And outputting a final confrontation sample based on the obtained disturbance, and completing the generation of the confrontation sample of the time sequence. The attack success rate and the non-concealment of the time sequence confrontation sample are considered at the same time.
Owner:WUHAN UNIV

Method and system for judging corrosion condition of electrode foil

The invention belongs to the technical field of condition judgment, and particularly relates to an electrode foil corrosion condition judgment method and system, and the method comprises the following steps: S1, obtaining a grayscale image of a to-be-analyzed electrode foil; constructing a structure tensor based on the composite gradient reflecting the local brightness and texture information of the pixel points, determining an anisotropic diffusion coefficient according to the structure tensor, and performing iterative anisotropic diffusion filtering processing on the grayscale image by using the anisotropic diffusion coefficient to obtain a filtered image; and S2, fusing the anisotropic diffusion coefficient determined for each pixel point with the pixel intensity of the filtered image to obtain a high-contrast corrosion significance map. According to the method, the contour edge information of the corrosion area can be kept while the complex texture and noise of the background area of the electrode foil image are smoothed, the contradiction between denoising and edge protection of a traditional filtering method is solved, the result of the corrosion condition is more reliable, and the automation level of electrode foil product quality detection is improved.
Owner:HUBEI FUYIDA ELECTRONIC TECH CO LTD

Coding of significance maps and transform coefficient blocks

A higher coding efficiency for coding a significance map indicating positions of significant transform coefficients within a transform coefficient block is achieved by the scan order by which the sequentially extracted syntax elements indicating, for associated positions within the transform coefficient block, as to whether at the respective position a significant or insignificant transform coefficient is situated, are sequentially associated to the positions of the transform coefficient block, among the positions of the transform coefficient block depends on the positions of the significant transform coefficients indicated by previously associated syntax elements. Alternatively, the first-type elements may be context-adaptively entropy decoded using contexts which are individually selected for each of the syntax elements dependent on a number of significant transform coefficients in a neighborhood of the respective syntax element, indicated as being significant by any of the preceding syntax elements.
Owner:DOLBY VIDEO COMPRESSION LLC

Methods and apparatus for unified significance map coding

Methods and apparatus are provided for unified significance map coding. An apparatus includes a video encoder (400) for encoding transform coefficients for at least a portion of a picture. The transform coefficients are obtained using a plurality of transforms. One or more context sharing maps are generated for the transform coefficients based on a unified rule. The one or more context sharing maps are for providing at least one context that is shared among at least some of the transform coefficients obtained from at least two different ones of the plurality of transforms.
Owner:INTERDIGITAL MADISON PATENT HLDG

A model pruning method of an interpretable CNN classification model

The application relates to a model pruning method of an interpretable CNN classification model, belongs to the field of image compression, and solves the problems of high operation complexity, large time and memory consumption and difficulty in deployment on terminal equipment of an existing deep CNN model, and solves the problem of lack of interpretability of an existing model pruning algorithm. The method comprises the following steps: inputting a training picture into a neural network model to be pruned, and extracting a feature map matrix of each convolution layer; upsampling the feature map matrix to the size of the input picture, and then performing a normalization operation to construct a saliency map; multiplying the saliency map and the input picture element by element to construct a weighted input picture; subtracting the input picture from the weighted picture element by element to construct an attention region occlusion map; inputting the attention occlusion map into the model to be pruned, observing the change of model accuracy as an importance score of the channel, and pruning the channel to obtain a pruned lightweight model. The application realizes high pruning rate of the model and improves the interpretability of the pruning process.
Owner:DALIAN UNIV OF TECH

Unsupervised semantic segmentation method and system based on saliency map and deep learning

The present application relates to the field of image processing technology, and in particular, to an unsupervised semantic segmentation method and system based on saliency maps and deep learning. The present application does not need to rely on strong or weak supervision implementation methods. Without using labeled data, it generates a clear and corrected foreground mask through unsupervised representation learning, combined with a pixel-level target attention classification head and a saliency map, so that the outline of the corrected foreground mask is clearer, thereby achieving unsupervised semantic segmentation. While improving the quality of pseudo-labels and the accuracy of semantic segmentation by correcting the foreground mask, it effectively reduces the system's dependence on manual annotation. At the same time, it can also be used as an automatic annotation system to automatically annotate image data sets, significantly reducing the labor cost of data annotation.
Owner:SUN YAT SEN UNIV

Multi-Level Significance Maps for Encoding and Decoding

Methods of encoding and decoding for video data are described in which multi-level significance maps are used in the encoding and decoding processes. The significant-coefficient flags that form the significance map are grouped into contiguous groups, and a significant-coefficient-group flag signifies for each group whether that group contains no non-zero significant-coefficient flags. If there are no non-zero significant-coefficient flags in the group, then the significant-coefficient-group flag is set to zero. The set of significant-coefficient-group flags is encoded in the bitstream. Any significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is non-zero are encoded in the bitstream, whereas significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is zero are not encoded in the bitstream.
Owner:VELOS MEDIA LLC

Video image bitrate allocation method, system and device, and storage medium

Disclosed in embodiments of the present application are a video image bitrate allocation method, system and device, and a storage medium. In the technical solution provided by the embodiments of the present application, a saliency map of a target image is acquired, and real-time coding quality information of a target region in the target image is determined on the basis of the saliency map, wherein the target region is a region of interest (ROI) or a non-ROI; a coding quality redundancy amplitude of the target region is determined on the basis of the real-time coding quality information and a set quality discrimination threshold, and an offset bitrate of the target region is determined on the basis of the coding quality redundancy amplitude; and an initial bitrate of each coding block in the target region is acquired, and a coding bitrate of each coding block in the target region is configured on the basis of the initial bitrate and the offset bitrate. By using the described technical means, the bitrates of the ROI and the non-RON can be adaptively configured on the basis of the video content attribute of the video image, thereby reducing bandwidth costs while ensuring the coding quality, optimizing the video coding effect.
Owner:PENINSULA INFORMATION TECH INC

Method and device for spatially resolved localization of a defect in a component

Method for spatially resolved localization of a defect in a component, the method comprising the steps: - Providing (S1) a salience map of a NOK image or a NOK fault class image of a component exhibiting a defect; - Apply (S2) at least one dimension reduction method to the provided salience map to generate a dimensionally reduced salience map; and - spatially resolved localization (S3) of the component defect based on the dimensionally reduced salience map.
Owner:ROBERT BOSCH GMBH

A learning-based saliency localization method

The present application relates to the technical field of map saliency localization, and particularly relates to a saliency localization method based on learning. The saliency localization method based on learning provided by the present application adopts a saliency prediction model to simulate the mechanism in a SLAM framework. A saliency map is predicted according to a saliency model, and the model can capture scene semantics and geometric information. The value of the saliency map is used as the weight of a feature point in a traditional bundle adjustment method. Detailed experiments conducted on the KITTI and EuRoc datasets with the most advanced algorithms show that the algorithm proposed by the present application is superior to existing algorithms in indoor and outdoor environments, and significantly improves the positioning accuracy and robustness.
Owner:CHONGQING UNIV

Electrode foil corrosion condition determination method and system

The present application belongs to the technical field of situation judgment, and particularly relates to a kind of electrode foil corrosion situation judgment method and system, comprising the following steps: S1, the gray scale image of electrode foil to be analyzed is obtained;Based on the composite gradient reflecting the local brightness and texture information of pixel point, the structure tensor is constructed, and the anisotropic diffusion coefficient is determined according to the structure tensor, the gray scale image is processed by iterative anisotropic diffusion filtering using the anisotropic diffusion coefficient, and the filtered image is obtained;S2, the anisotropic diffusion coefficient determined for each pixel point is fused with the pixel intensity of the filtered image to obtain a high-contrast corrosion significance map.The present application can smooth the complex texture and noise in the background area of electrode foil image while maintaining the outline edge information of the corrosion area, solve the contradiction between denoising and edge preservation in traditional filtering method, make the corrosion situation result more reliable, and improve the automation level of electrode foil product quality detection.
Owner:HUBEI FUYIDA ELECTRONIC TECH CO LTD

Apparatus and method for evaluating saliency map determiner

Apparatuses and methods for evaluating a saliency map determiner are provided. According to various embodiments, a method of evaluating a saliency map determiner is described, the method comprising: adding a predefined pattern to a plurality of training dataset units to train identification of a data class, wherein each training dataset unit comprises a representation of the data class to be identified; training a neural network with the plurality of training dataset units comprising the predefined pattern; determining, by the saliency map determiner, a saliency map for the data class; and evaluating the saliency map determiner based on whether the determined saliency map comprises a context of the data class introduced by the addition of the predefined pattern.
Owner:ROBERT BOSCH GMBH

Coding method, compression network training method and related equipment

The invention discloses a coding method, a compression network training method and related equipment, and belongs to the field of image processing, and the coding method comprises the steps: carrying out the saliency feature extraction of a to-be-coded feature map of a to-be-coded image, and obtaining a target saliency map, the target saliency map is used for representing the importance of each region of the to-be-coded image; performing coding processing on the to-be-coded feature map based on the target saliency map to obtain a coded code stream; wherein the coding processing comprises the following steps: determining a code rate weight of the to-be-coded feature map according to the target saliency map; and performing compression coding on the to-be-coded feature map according to the code rate weight.
Owner:VIVO MOBILE COMM CO LTD

Open vocabulary image semantic segmentation method based on hierarchical aggregation and prompt optimization

The invention discloses an open vocabulary image semantic segmentation method based on hierarchical aggregation and prompt optimization. Semantic segmentation of any open category target is achieved under the condition that extra training is not needed. The method comprises the following steps: firstly, preprocessing an input image, carrying out image semantic analysis in combination with a visual language model, and constructing a text prompt matched with image content; then, cross-modal attention information is extracted, a category saliency map is generated, and a soft discarding strategy is introduced in the saliency mining process to reduce out-of-distribution noise and artifact interference; further performing quality evaluation on significance results obtained by multiple rounds of iteration, and obtaining stable and consistent significance representation by adopting a layered weighted aggregation mode; and finally, automatically generating prompt information based on an aggregation result, inputting the prompt information into the segmentation model, and outputting a semantic segmentation mask of the target.
Owner:HOHAI UNIV

A multi-modal based salient object detection method and device, and related medium

The application discloses a kind of based on multimodal salient object detection method, device and related medium, the method includes inputting the color image to be detected into visual language model and carrying out multimodal feature processing, obtains semantic alignment feature set;Multi-granularity semantic reasoning is carried out to semantic alignment feature set, and positioning enhanced feature map is obtained;Again, detail feature map extraction is carried out to the color image to be detected by image feature neural network and detail feature map and positioning enhanced feature map are fused, and finally saliency map is obtained.Such that, saliency reasoning process can explicitly introduce semantic information, and is synergized using multi-granularity semantic expert system algorithm and detail recovery fusion algorithm, and the positioning accuracy of saliency map in complex scene is improved.
Owner:ZHEJIANG DIANCHUANG INFORMATION TECH CO LTD +1

Adversarial sample generation method based on multi-level significance regional knowledge distillation

The invention relates to an adversarial sample generation method based on multi-level salient region knowledge distillation, and belongs to the technical field of adversarial attacks in the field of computer vision. According to the method, key pixels for judging a single image by a neural network are positioned by utilizing multi-level saliency region information guidance, saliency region images of all feature layers are obtained, and a multi-level saliency image set is formed; and integrating salient region images in the salient image set by using a salient image set knowledge distillation strategy so as to generate a higher-quality adversarial attack. According to the method, significant differences of different levels are comprehensively considered, and the attack success rate and the concealment are improved; according to the method, a saliency region integration strategy based on data set distillation is designed, and the generalization ability of attack resistance is improved.
Owner:CHINESE PEOPLES LIBERATION ARMY UNIT 32801

A scale-invariant based co-salient object detection method

The application discloses a kind of based on scale invariance collaborative salient target detection method, it is related to monitoring safety technical field, method includes: obtaining the image sequence converted from video stream in monitoring system, it is input to the main network of pre-training model to extract multi-scale feature;Multi-scale feature is enhanced and fusion processing is generated initial saliency map;By residual refinement structure, restore details edge, generate accurate saliency map;Scale consistency constraint and semantic feature feedback are applied to optimization, obtain target saliency map and output.The application can effectively solve the problem that model lacks semantic guidance, generalization ability is poor in prior art to different resolution images, significantly improve the accuracy and robustness of salient target detection in monitoring system, applicable to a variety of complex monitoring scenes.
Owner:WUHAN INST OF TECH

Engine part defect detection method, device and equipment based on AI vision and medium

The application relates to an AI vision-based engine part defect detection method, device, equipment and medium. The method comprises the following steps: acquiring labeled real engine part images and defect masks to form initial samples, training by using a multi-scale topological perception feature extraction network, generating a deep feature map and a topological saliency map, and based on this, performing defect data coding and feature space synthesis to construct a large-scale synthetic augmented training dataset; taking the trained feature extraction network as a backbone, integrating a dynamic domain self-adaptive module to construct a defect detection framework, obtaining a robust teacher detection model and a parameter offset library through multi-working-condition simulation training, and performing two-stage progressive knowledge distillation processing to obtain a defect detection engine. The method significantly improves the accuracy, robustness and efficiency of engine part defect detection through multi-scale topological perception feature extraction, dynamic domain self-adaptation and progressive knowledge distillation.
Owner:张贺

Image recognition method based on class activation map algorithm and related device

Embodiments of the present application relate to the technical field of artificial intelligence, and disclose an image recognition method based on a class activation map algorithm and related devices, the method comprising: obtaining an image to be recognized; inputting the image to be recognized into a class activation map algorithm based on multi-label gradient feedback, the class activation map algorithm based on multi-label gradient feedback performing feature convolution on the image to be recognized to obtain a multi-scale feature map, performing gradient backpropagation on feature maps of each label category when performing gradient backpropagation on a current category, determining an inter-category influence factor according to a gradient value set backpropagated by all label categories, determining an activation map weight of the current category according to the inter-category influence factor and an influence factor of different pixel positions of the current category, and obtaining a recognition result according to the activation map weight and the multi-scale feature map; wherein the current category is any category in a plurality of label categories; and outputting the recognition result. Embodiments of the present application improve the generation quality of a saliency map.
Owner:SHENZHEN UNIV

Methods and apparatus for unified significance map coding

Methods and apparatus are provided for unified significance map coding. An apparatus includes a video encoder (400) for encoding transform coefficients for at least a portion of a picture. The transform coefficients are obtained using a plurality of transforms. One or more context sharing maps are generated for the transform coefficients based on a unified rule. The one or more context sharing maps are for providing at least one context that is shared among at least some of the transform coefficients obtained from at least two different ones of the plurality of transforms.
Owner:INTERDIGITAL MADISON PATENT HLDG

Compression for sparse data structures utilizing mode search approximation

Embodiments are generally directed to compression for compression for sparse data structures utilizing mode search approximation. An embodiment of an apparatus includes one or more processors including a graphics processor to process data; and a memory for storage of data, including compressed data. The one or more processors are to provide for compression of a data structure, including identification of a mode in the data structure, the data structure including a plurality of values and the mode being a most repeated value in a data structure, wherein identification of the mode includes application of a mode approximation operation, and encoding of an output vector to include the identified mode, a significance map to indicate locations at which the mode is present in the data structure, and remaining uncompressed data from the data structure.
Owner:INTEL CORP

Significance map and transform coefficient block encoding

The present application relates to the coding of a significance map and a block of transform coefficients. A higher coding efficiency for coding a significance map indicating positions of significant transform coefficients within a block of transform coefficients is achieved by a scan order, by means of which the syntax elements, which are sequentially extracted, indicating for an associated position within the block of transform coefficients whether a significant transform coefficient or a non-significant transform coefficient is located at the respective position, are sequentially associated to the transform coefficient block positions among the transform coefficient block positions according to the significant transform coefficient positions indicated by previously associated syntax elements. Optionally, the first type elements can be context-adaptively entropy decoded using a context, the number of positions at which a significant transform coefficient is located according to previously extracted and associated first type syntax elements, being individually selected for each of said first type syntax elements. Even optionally, the values are scanned in a sub-block manner and the context is selected based on sub-block statistics.
Owner:DOLBY VIDEO COMPRESSION LLC

Rich text document processing method based on region-level retrieval enhanced generation

The invention relates to a visual rich text document processing method based on region-level retrieval enhancement generation. The visual rich text document processing method comprises the following steps of: 1, segmenting each document image in a document library into a plurality of non-overlapped image blocks; 2, respectively extracting an embedded vector Vq of the query text q and an embedded vector set {VRi} of the image block set R by using a visual language model M; 3, calculating the similarity si between the Vq and an image block embedding vector VRi through a region retriever I, and generating a saliency map G; step 4, carrying out binarization processing on the saliency map G, and carrying out connected region analysis on the saliency blocks based on a preset neighborhood range to generate one or more candidate bounding boxes; and 5, inputting the image region corresponding to the candidate bounding box into a large language model generator, and generating a text answer corresponding to the query. According to the method, the retrieval granularity is improved from the document level to the region level, and the problems of distraction and performance reduction caused by redundant visual content in a traditional method are effectively solved.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Space-time saliency prompting method and device for open vocabulary video action recognition

The invention provides a space-time saliency prompting method and device for open vocabulary video action recognition, and relates to the technical field of computer vision. Comprising the following steps: inputting an input video into a video-level space-time saliency prompting module, obtaining a coarse-grained saliency map through cross-modal alignment attention branches, obtaining a fine-grained saliency map through instance-level self-attention branches, and fusing the coarse-grained saliency map and the fine-grained saliency map to obtain a video-level space-time saliency prompting map; performing element-by-element multiplication with an input video, and introducing residual connection to obtain an enhanced discriminative action feature; and injecting the fine-grained saliency map into an input video to generate space-time action features, and inputting the space-time action features into a Token-level space-time saliency prompt module to obtain an open vocabulary video action recognition result. The invention provides a novel space-time significance prompting method, which can dynamically focus on an instance-level discriminative action area, and can generate a plurality of levels of action significance prompts which complement each other, so that the generalization ability of a model is remarkably enhanced.
Owner:UNIV OF SCI & TECH BEIJING

Multi-Level Significance Maps for Encoding and Decoding

PendingUS20250287001A1Digital video signal modificationAlgorithmSignificance map
Methods of encoding and decoding for video data are described in which multi-level significance maps are used in the encoding and decoding processes. The significant-coefficient flags that form the significance map are grouped into contiguous groups, and a significant-coefficient-group flag signifies for each group whether that group contains no non-zero significant-coefficient flags. If there are no non-zero significant-coefficient flags in the group, then the significant-coefficient-group flag is set to zero. The set of significant-coefficient-group flags is encoded in the bitstream. Any significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is non-zero are encoded in the bitstream, whereas significant-coefficient flags that fall within a group that has a significant-coefficient-group flag that is zero are not encoded in the bitstream.
Owner:VELOS MEDIA LLC

Exposure control method and device, computer equipment and storage medium

PendingCN121665121AImage extractionSaliency map
The invention discloses an exposure control method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a current frame image; extracting multi-scale texture gradient features of the current frame image; constructing a distance sensing model based on camera internal parameters, and mapping the multi-scale texture gradient features into spatial distance weights according to the distance sensing model; generating a significance map according to the spatial distance weight; determining exposure control parameters according to the saliency map; and adjusting image exposure according to the exposure control parameters. By analyzing texture gradient features of an image and utilizing internal reference of a camera to perform distance perception mapping, exposure of an area closer to the camera and more likely to be a main body in a picture can be actively deduced and preferentially optimized under the condition that targets such as a human face are not successfully detected; the problem of deadlock of the whole exposure process caused by failure of initial detection is avoided, and the fundamental problem that exposure control cannot be started under severe illumination conditions such as backlight is solved.
Owner:SHENZHEN TVT DIGITAL TECH CO LTD

A method, device and electronic equipment for detecting salient objects based on multi-modal prediction reconstruction error

The application provides a salient object detection method and device based on multi-modal prediction reconstruction error and an electronic device. The method comprises: acquiring images of at least one mode and establishing a mode existence mask vector; performing feature coding on each existing mode image to obtain multi-scale features, and performing feature fusion to obtain a multi-scale global scene representation; performing self-prediction reconstruction and cross-modal prediction reconstruction based on the multi-scale global scene representation to obtain self-reconstruction results and cross-modal reconstruction results; calculating multi-scale self-reconstruction errors and multi-scale cross-modal reconstruction errors corresponding to each mode to obtain multi-scale error maps and initial saliency maps; constructing a gaze area and cropping image blocks from each mode image to calculate initial local saliency maps, and obtaining a refined saliency map through iterative refinement; calculating a gating saliency map corresponding to a global error feature vector, and fusing the gating saliency map and the refined saliency map to obtain a final salient object detection result.
Owner:HANGZHOU DIANZI UNIV

Salient target detection method and device based on multi-modal prediction reconstruction error, and electronic equipment

The invention provides a saliency target detection method and device based on a multi-modal prediction reconstruction error and electronic equipment. The method comprises the following steps: acquiring an image of at least one modal and establishing a modal existence mask vector; performing feature coding on each existing modal image to obtain multi-scale features, and performing feature fusion to obtain multi-scale global scene representation; performing self-prediction reconstruction and cross-modal prediction reconstruction based on the multi-scale global scene representation to obtain a self-reconstruction result and a cross-modal reconstruction result; calculating a multi-scale self-reconstruction error and a multi-scale cross-modal reconstruction error corresponding to each modal to obtain a multi-scale error graph and an initial saliency graph; constructing a staring area, cutting image blocks from each modal image to calculate an initial local saliency map, and obtaining a refined saliency map through iterative refinement; and calculating a gated saliency map corresponding to the global error feature vector, and fusing the gated saliency map and the refined saliency map to obtain a final saliency target detection result.
Owner:HANGZHOU DIANZI UNIV

Compression for sparse data structures utilizing mode search approximation

Embodiments are generally directed to compression for compression for sparse data structures utilizing mode search approximation. An embodiment of an apparatus includes one or more processors including a graphics processor to process data; and a memory for storage of data, including compressed data. The one or more processors are to provide for compression of a data structure, including identification of a mode in the data structure, the data structure including a plurality of values and the mode being a most repeated value in a data structure, wherein identification of the mode includes application of a mode approximation operation, and encoding of an output vector to include the identified mode, a significance map to indicate locations at which the mode is present in the data structure, and remaining uncompressed data from the data structure.
Owner:INTEL CORP