Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Significance map" patented technology

Method and system for judging corrosion condition of electrode foil

ActiveCN121329920AImage enhancementImage analysisSilhouette edgeLightness
The invention belongs to the technical field of condition judgment, and particularly relates to an electrode foil corrosion condition judgment method and system, and the method comprises the following steps: S1, obtaining a grayscale image of a to-be-analyzed electrode foil; constructing a structure tensor based on the composite gradient reflecting the local brightness and texture information of the pixel points, determining an anisotropic diffusion coefficient according to the structure tensor, and performing iterative anisotropic diffusion filtering processing on the grayscale image by using the anisotropic diffusion coefficient to obtain a filtered image; and S2, fusing the anisotropic diffusion coefficient determined for each pixel point with the pixel intensity of the filtered image to obtain a high-contrast corrosion significance map. According to the method, the contour edge information of the corrosion area can be kept while the complex texture and noise of the background area of the electrode foil image are smoothed, the contradiction between denoising and edge protection of a traditional filtering method is solved, the result of the corrosion condition is more reliable, and the automation level of electrode foil product quality detection is improved.
Owner:HUBEI FUYIDA ELECTRONIC TECH CO LTD

Coding of significance maps and transform coefficient blocks

A higher coding efficiency for coding a significance map indicating positions of significant transform coefficients within a transform coefficient block is achieved by the scan order by which the sequentially extracted syntax elements indicating, for associated positions within the transform coefficient block, as to whether at the respective position a significant or insignificant transform coefficient is situated, are sequentially associated to the positions of the transform coefficient block, among the positions of the transform coefficient block depends on the positions of the significant transform coefficients indicated by previously associated syntax elements. Alternatively, the first-type elements may be context-adaptively entropy decoded using contexts which are individually selected for each of the syntax elements dependent on a number of significant transform coefficients in a neighborhood of the respective syntax element, indicated as being significant by any of the preceding syntax elements.
Owner:DOLBY VIDEO COMPRESSION LLC

Methods and apparatus for unified significance map coding

Methods and apparatus are provided for unified significance map coding. An apparatus includes a video encoder (400) for encoding transform coefficients for at least a portion of a picture. The transform coefficients are obtained using a plurality of transforms. One or more context sharing maps are generated for the transform coefficients based on a unified rule. The one or more context sharing maps are for providing at least one context that is shared among at least some of the transform coefficients obtained from at least two different ones of the plurality of transforms.
Owner:INTERDIGITAL MADISON PATENT HLDG

A model pruning method of an interpretable CNN classification model

The application relates to a model pruning method of an interpretable CNN classification model, belongs to the field of image compression, and solves the problems of high operation complexity, large time and memory consumption and difficulty in deployment on terminal equipment of an existing deep CNN model, and solves the problem of lack of interpretability of an existing model pruning algorithm. The method comprises the following steps: inputting a training picture into a neural network model to be pruned, and extracting a feature map matrix of each convolution layer; upsampling the feature map matrix to the size of the input picture, and then performing a normalization operation to construct a saliency map; multiplying the saliency map and the input picture element by element to construct a weighted input picture; subtracting the input picture from the weighted picture element by element to construct an attention region occlusion map; inputting the attention occlusion map into the model to be pruned, observing the change of model accuracy as an importance score of the channel, and pruning the channel to obtain a pruned lightweight model. The application realizes high pruning rate of the model and improves the interpretability of the pruning process.
Owner:DALIAN UNIV OF TECH

Method and device for spatially resolved localization of a defect in a component

PendingDE102024207745A1Image enhancementImage analysisSpatially resolvedComputer graphics (images)
Method for spatially resolved localization of a defect in a component, the method comprising the steps: - Providing (S1) a salience map of a NOK image or a NOK fault class image of a component exhibiting a defect; - Apply (S2) at least one dimension reduction method to the provided salience map to generate a dimensionally reduced salience map; and - spatially resolved localization (S3) of the component defect based on the dimensionally reduced salience map.
Owner:ROBERT BOSCH GMBH

Electrode foil corrosion condition determination method and system

ActiveCN121329920BSilhouette edgeMaterials science
The present application belongs to the technical field of situation judgment, and particularly relates to a kind of electrode foil corrosion situation judgment method and system, comprising the following steps: S1, the gray scale image of electrode foil to be analyzed is obtained;Based on the composite gradient reflecting the local brightness and texture information of pixel point, the structure tensor is constructed, and the anisotropic diffusion coefficient is determined according to the structure tensor, the gray scale image is processed by iterative anisotropic diffusion filtering using the anisotropic diffusion coefficient, and the filtered image is obtained;S2, the anisotropic diffusion coefficient determined for each pixel point is fused with the pixel intensity of the filtered image to obtain a high-contrast corrosion significance map.The present application can smooth the complex texture and noise in the background area of electrode foil image while maintaining the outline edge information of the corrosion area, solve the contradiction between denoising and edge preservation in traditional filtering method, make the corrosion situation result more reliable, and improve the automation level of electrode foil product quality detection.
Owner:HUBEI FUYIDA ELECTRONIC TECH CO LTD

Open vocabulary image semantic segmentation method based on hierarchical aggregation and prompt optimization

The invention discloses an open vocabulary image semantic segmentation method based on hierarchical aggregation and prompt optimization. Semantic segmentation of any open category target is achieved under the condition that extra training is not needed. The method comprises the following steps: firstly, preprocessing an input image, carrying out image semantic analysis in combination with a visual language model, and constructing a text prompt matched with image content; then, cross-modal attention information is extracted, a category saliency map is generated, and a soft discarding strategy is introduced in the saliency mining process to reduce out-of-distribution noise and artifact interference; further performing quality evaluation on significance results obtained by multiple rounds of iteration, and obtaining stable and consistent significance representation by adopting a layered weighted aggregation mode; and finally, automatically generating prompt information based on an aggregation result, inputting the prompt information into the segmentation model, and outputting a semantic segmentation mask of the target.
Owner:HOHAI UNIV

A multi-modal based salient object detection method and device, and related medium

The application discloses a kind of based on multimodal salient object detection method, device and related medium, the method includes inputting the color image to be detected into visual language model and carrying out multimodal feature processing, obtains semantic alignment feature set;Multi-granularity semantic reasoning is carried out to semantic alignment feature set, and positioning enhanced feature map is obtained;Again, detail feature map extraction is carried out to the color image to be detected by image feature neural network and detail feature map and positioning enhanced feature map are fused, and finally saliency map is obtained.Such that, saliency reasoning process can explicitly introduce semantic information, and is synergized using multi-granularity semantic expert system algorithm and detail recovery fusion algorithm, and the positioning accuracy of saliency map in complex scene is improved.
Owner:ZHEJIANG DIANCHUANG INFORMATION TECH CO LTD +1

Engine part defect detection method, device and equipment based on AI vision and medium

The application relates to an AI vision-based engine part defect detection method, device, equipment and medium. The method comprises the following steps: acquiring labeled real engine part images and defect masks to form initial samples, training by using a multi-scale topological perception feature extraction network, generating a deep feature map and a topological saliency map, and based on this, performing defect data coding and feature space synthesis to construct a large-scale synthetic augmented training dataset; taking the trained feature extraction network as a backbone, integrating a dynamic domain self-adaptive module to construct a defect detection framework, obtaining a robust teacher detection model and a parameter offset library through multi-working-condition simulation training, and performing two-stage progressive knowledge distillation processing to obtain a defect detection engine. The method significantly improves the accuracy, robustness and efficiency of engine part defect detection through multi-scale topological perception feature extraction, dynamic domain self-adaptation and progressive knowledge distillation.
Owner:张贺

Image recognition method based on class activation map algorithm and related device

Embodiments of the present application relate to the technical field of artificial intelligence, and disclose an image recognition method based on a class activation map algorithm and related devices, the method comprising: obtaining an image to be recognized; inputting the image to be recognized into a class activation map algorithm based on multi-label gradient feedback, the class activation map algorithm based on multi-label gradient feedback performing feature convolution on the image to be recognized to obtain a multi-scale feature map, performing gradient backpropagation on feature maps of each label category when performing gradient backpropagation on a current category, determining an inter-category influence factor according to a gradient value set backpropagated by all label categories, determining an activation map weight of the current category according to the inter-category influence factor and an influence factor of different pixel positions of the current category, and obtaining a recognition result according to the activation map weight and the multi-scale feature map; wherein the current category is any category in a plurality of label categories; and outputting the recognition result. Embodiments of the present application improve the generation quality of a saliency map.
Owner:SHENZHEN UNIV

Methods and apparatus for unified significance map coding

Methods and apparatus are provided for unified significance map coding. An apparatus includes a video encoder (400) for encoding transform coefficients for at least a portion of a picture. The transform coefficients are obtained using a plurality of transforms. One or more context sharing maps are generated for the transform coefficients based on a unified rule. The one or more context sharing maps are for providing at least one context that is shared among at least some of the transform coefficients obtained from at least two different ones of the plurality of transforms.
Owner:INTERDIGITAL MADISON PATENT HLDG

Rich text document processing method based on region-level retrieval enhanced generation

The invention relates to a visual rich text document processing method based on region-level retrieval enhancement generation. The visual rich text document processing method comprises the following steps of: 1, segmenting each document image in a document library into a plurality of non-overlapped image blocks; 2, respectively extracting an embedded vector Vq of the query text q and an embedded vector set {VRi} of the image block set R by using a visual language model M; 3, calculating the similarity si between the Vq and an image block embedding vector VRi through a region retriever I, and generating a saliency map G; step 4, carrying out binarization processing on the saliency map G, and carrying out connected region analysis on the saliency blocks based on a preset neighborhood range to generate one or more candidate bounding boxes; and 5, inputting the image region corresponding to the candidate bounding box into a large language model generator, and generating a text answer corresponding to the query. According to the method, the retrieval granularity is improved from the document level to the region level, and the problems of distraction and performance reduction caused by redundant visual content in a traditional method are effectively solved.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Space-time saliency prompting method and device for open vocabulary video action recognition

The invention provides a space-time saliency prompting method and device for open vocabulary video action recognition, and relates to the technical field of computer vision. Comprising the following steps: inputting an input video into a video-level space-time saliency prompting module, obtaining a coarse-grained saliency map through cross-modal alignment attention branches, obtaining a fine-grained saliency map through instance-level self-attention branches, and fusing the coarse-grained saliency map and the fine-grained saliency map to obtain a video-level space-time saliency prompting map; performing element-by-element multiplication with an input video, and introducing residual connection to obtain an enhanced discriminative action feature; and injecting the fine-grained saliency map into an input video to generate space-time action features, and inputting the space-time action features into a Token-level space-time saliency prompt module to obtain an open vocabulary video action recognition result. The invention provides a novel space-time significance prompting method, which can dynamically focus on an instance-level discriminative action area, and can generate a plurality of levels of action significance prompts which complement each other, so that the generalization ability of a model is remarkably enhanced.
Owner:UNIV OF SCI & TECH BEIJING

Exposure control method and device, computer equipment and storage medium

PendingCN121665121AImage extractionSaliency map
The invention discloses an exposure control method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a current frame image; extracting multi-scale texture gradient features of the current frame image; constructing a distance sensing model based on camera internal parameters, and mapping the multi-scale texture gradient features into spatial distance weights according to the distance sensing model; generating a significance map according to the spatial distance weight; determining exposure control parameters according to the saliency map; and adjusting image exposure according to the exposure control parameters. By analyzing texture gradient features of an image and utilizing internal reference of a camera to perform distance perception mapping, exposure of an area closer to the camera and more likely to be a main body in a picture can be actively deduced and preferentially optimized under the condition that targets such as a human face are not successfully detected; the problem of deadlock of the whole exposure process caused by failure of initial detection is avoided, and the fundamental problem that exposure control cannot be started under severe illumination conditions such as backlight is solved.
Owner:SHENZHEN TVT DIGITAL TECH CO LTD

A method, device and electronic equipment for detecting salient objects based on multi-modal prediction reconstruction error

The application provides a salient object detection method and device based on multi-modal prediction reconstruction error and an electronic device. The method comprises: acquiring images of at least one mode and establishing a mode existence mask vector; performing feature coding on each existing mode image to obtain multi-scale features, and performing feature fusion to obtain a multi-scale global scene representation; performing self-prediction reconstruction and cross-modal prediction reconstruction based on the multi-scale global scene representation to obtain self-reconstruction results and cross-modal reconstruction results; calculating multi-scale self-reconstruction errors and multi-scale cross-modal reconstruction errors corresponding to each mode to obtain multi-scale error maps and initial saliency maps; constructing a gaze area and cropping image blocks from each mode image to calculate initial local saliency maps, and obtaining a refined saliency map through iterative refinement; calculating a gating saliency map corresponding to a global error feature vector, and fusing the gating saliency map and the refined saliency map to obtain a final salient object detection result.
Owner:HANGZHOU DIANZI UNIV

Salient target detection method and device based on multi-modal prediction reconstruction error, and electronic equipment

The invention provides a saliency target detection method and device based on a multi-modal prediction reconstruction error and electronic equipment. The method comprises the following steps: acquiring an image of at least one modal and establishing a modal existence mask vector; performing feature coding on each existing modal image to obtain multi-scale features, and performing feature fusion to obtain multi-scale global scene representation; performing self-prediction reconstruction and cross-modal prediction reconstruction based on the multi-scale global scene representation to obtain a self-reconstruction result and a cross-modal reconstruction result; calculating a multi-scale self-reconstruction error and a multi-scale cross-modal reconstruction error corresponding to each modal to obtain a multi-scale error graph and an initial saliency graph; constructing a staring area, cutting image blocks from each modal image to calculate an initial local saliency map, and obtaining a refined saliency map through iterative refinement; and calculating a gated saliency map corresponding to the global error feature vector, and fusing the gated saliency map and the refined saliency map to obtain a final saliency target detection result.
Owner:HANGZHOU DIANZI UNIV

Compression for sparse data structures utilizing mode search approximation

Embodiments are generally directed to compression for compression for sparse data structures utilizing mode search approximation. An embodiment of an apparatus includes one or more processors including a graphics processor to process data; and a memory for storage of data, including compressed data. The one or more processors are to provide for compression of a data structure, including identification of a mode in the data structure, the data structure including a plurality of values and the mode being a most repeated value in a data structure, wherein identification of the mode includes application of a mode approximation operation, and encoding of an output vector to include the identified mode, a significance map to indicate locations at which the mode is present in the data structure, and remaining uncompressed data from the data structure.
Owner:INTEL CORP

Method and Apparatus for the Spatially Resolved Localization of a Defect in a Component

PendingUS20260051046A1Image enhancementImage analysisPattern recognitionSpatially resolved
A method for the spatially resolved localization of a defect in a component includes (i) providing a saliency map of a NOK image or a NOK defect class image of a component exhibiting a defect, (ii) applying at least one dimension reduction method to the provided saliency map to generate a dimension-reduced saliency map, and (iii) spatially resolved localization of the defect of the component on the basis of the dimension-reduced saliency map.
Owner:ROBERT BOSCH GMBH

General adversarial perturbation attack method and system for scene-oriented text recognition model

The application provides a general adversarial perturbation attack method and system for a scene text recognition model, belongs to the field of computer vision and adversarial machine learning, and comprises the following steps: S1, inputting a scene text recognition model dataset into a trained scene text recognition model for recognition, and selecting correctly recognized pictures to construct a training dataset; S2, constructing an original adversarial saliency map of each picture in the training dataset based on a loss function; S3, performing positive-negative separation, normalization and global average processing on the original adversarial saliency map to obtain a global average positive-negative adversarial saliency map; S4, generating a general adversarial perturbation based on the global average positive-negative adversarial saliency map; and S5, inputting the general adversarial perturbation into a scene text recognition model to be detected for attack. The application generates a perturbation by using an adversarial saliency map, improves attack migration performance, reduces attack running time, and provides technical support for robustness evaluation and privacy protection of an STR model.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Towards zero-shot anomaly detection and reasoning with multimodal large language models

According to one aspect, towards zero-shot anomaly detection and reasoning with multimodal large language models (MLLMs) may include generating a set of one or more visual tokens based on an input image, generating a significance map for the set of visual tokens based on the set of visual tokens and look-twice feature matching (LTFM), and identifying one or more anomalous visual tokens associated with the input image based on the set of one or more visual tokens associated with the input image and the significance map.
Owner:HONDA MOTOR CO LTD

Concrete chiseling roughness illumination correction detection method and system

The invention belongs to the technical field of concrete construction quality detection, and particularly discloses a concrete chiseling roughness illumination correction detection method and system, and the method comprises the steps: collecting a color image and a depth map of a chiseling surface, and calculating a roughness saliency map and an exposure quality map; performing adaptive exposure correction on each image to obtain an enhanced image; taking the enhanced image, the depth map, the roughness saliency map and the exposure quality map as common input, and extracting initial texture features and geometric structure features to obtain an initial fusion feature map; respectively constructing a high-resolution texture flow and a low-resolution semantic geometric flow based on the initial fusion feature map, and fusing the high-resolution texture flow and the low-resolution semantic geometric flow to obtain a second fusion feature map; and performing frequency domain enhancement on the second fusion feature map to obtain a frequency domain enhanced feature map, and performing scabbling roughness grade classification to realize concrete scabbling roughness illumination correction detection. According to the invention, the problems of insufficient image reliability, limited model discrimination precision and poor environmental adaptability in the prior art are solved.
Owner:HOHAI UNIV

Correlation filter tracking method based on saliency perception and spatio-temporal regularization

ActiveCN115526912BImage enhancementImage analysisPattern recognitionScale estimation
The application provides a correlation filter tracking method based on saliency perception and space-time regularization, introduces a space regularization and a time constraint term into a target function, establishes a space regularization and time constraint model, converts the model into a frequency domain, decomposes the converted model into multiple sub-problems, solves the sub-problems respectively, obtains an optimized model, acquires a saliency map of a target region, fuses the saliency map of the target into initial space regularization weight coefficients to obtain new weight coefficients based on saliency perception, learns a correlation filter according to the optimized model, uses the correlation filter to locate the target and simultaneously completes scale estimation of the target, and finally updates model parameters to complete filter tracking, which is helpful to enhance time continuity and consistency of the model, effectively reduce boundary effects, effectively improve tracking performance and tracking efficiency, and enable the tracker to adapt to appearance changes and suppress background interference.
Owner:XIAN TECH UNIV

Image recognition method based on class activation graph algorithm and related device

The embodiment of the invention relates to the technical field of artificial intelligence, and discloses an image recognition method based on a class activation graph algorithm and a related device, and the method comprises the steps: obtaining a to-be-recognized image; inputting the to-be-recognized image into a class activation graph algorithm based on multi-label gradient feedback, performing feature convolution on the to-be-recognized image by the class activation graph algorithm based on the multi-label gradient feedback to obtain a multi-scale feature graph, and performing gradient return on the feature graph of each label category when performing gradient return on the current category to obtain a multi-scale feature graph; determining an inter-category influence factor according to a gradient value set returned by all label categories, determining an activation graph weight of the current category according to the inter-category influence factor and influence factors of different pixel positions of the current category, and obtaining a recognition result according to the activation graph weight and the multi-scale feature graph; wherein the current category is any category in a plurality of label categories; and outputting the identification result. According to the embodiment of the invention, the generation quality of the saliency map is improved.
Owner:SHENZHEN UNIV

Lightweight depth enhancement method for underwater salient target detection

The invention relates to a lightweight depth enhancement method for underwater salient target detection, and the method comprises the steps: employing an encoder-decoder architecture, enabling an encoder to take an OfficientFormerV2 as a backbone network, extracting the multistage semantic features of an RGB image and a depth image, and completing the fusion enhancement of the RGB features and the depth features through a constructed depth enhancement module; the decoder is of a lightweight structure, integrates a multi-level feature fusion strategy, a multi-task learning strategy and a multi-level supervision method, and achieves synchronous prediction of a saliency map and a saliency boundary map. Experimental results show that the detection performance of the method on USOD10K and USOD underwater data sets is superior to that of most existing lightweight methods, the performance of the method is similar to that of a high-performance non-lightweight method, the underwater salient target detection precision is effectively improved while the lightweight characteristic is kept, and the problems that a traditional model is large in parameter quantity, difficult to deploy and poor in underwater scene adaptability are solved.
Owner:HANGZHOU DIANZI UNIV

Code rate fine regulation and control method based on audio and video multi-mode perception

The invention relates to the technical field of video coding, and discloses a code rate fine regulation and control method based on audio and video multi-mode perception, which comprises the following steps of: 1, extracting and normalizing auditory perception features; step 2, calculating to obtain a frame-level auditory significance score Saud; 3, fusing the audiovisual saliency map with the frame-level auditory saliency score Saud to obtain a code rate adjustment factor matrix F; and step 4, obtaining a local regulatory factor value FCU corresponding to each coding unit through F, comparing the local regulatory factor value FCU with a regulatory factor reference value Fref to obtain a relative deviation, introducing the deviation value into QP offset control logic, and dynamically correcting the original QP offset of the encoder. The space-time resolution capability of the saliency-guided code rate adjustment factor matrix F is enhanced, and the method has higher flexibility and is suitable for a task scene of audio and video fusion.
Owner:HANGZHOU DIANZI UNIV

Multi-mode-based salient object detection method and device and related medium

The invention discloses a multi-modal-based salient object detection method and device and a related medium, and the method comprises the steps: inputting a to-be-detected color image into a visual language model for multi-modal feature processing, and obtaining a semantic alignment feature set; performing multi-granularity semantic reasoning on the semantic alignment feature set to obtain a positioning enhancement feature map; and carrying out detail feature map extraction on the to-be-detected color image through the image feature neural network, and fusing the detail feature map with the positioning enhancement feature map to finally obtain a saliency map. Thus, semantic information can be explicitly introduced in the saliency reasoning process, cooperation is carried out through a multi-granularity semantic expert system algorithm and a detail recovery fusion algorithm, and the positioning precision of the saliency map in a complex scene is improved.
Owner:ZHEJIANG DIANCHUANG INFORMATION TECH CO LTD +1