Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

53 results about "Saliency map" patented technology

In computer vision, a saliency map is an image that shows each pixel's unique quality. The goal of a saliency map is to simplify and/or change the representation of an image into something that is more meaningful and easier to analyze. For example, if a pixel has a high grey level or other unique color quality in a color image, that pixel's quality will show in the saliency map and in an obvious way. Saliency is a kind of image segmentation.

An optical remote sensing image processing system and training method

PendingCN122289668AFeature miningSaliency map
An optical remote sensing image processing system and training method, relating to the fields of computer vision and image processing technology, alleviates the difficulties in balancing high accuracy and high efficiency in saliency enhancement in existing optical remote sensing image processing technologies, which suffer from high computational costs, incomplete structures, and unclear boundaries. The optical remote sensing image processing system includes a basic feature extraction network module for extracting basic features from the optical remote sensing image to be enhanced; a multi-directional feature mining and aggregation module for obtaining corresponding feature map sequences based on the basic feature map sequences; a cross-scale edge information fusion module for fusing the feature map sequences through cross-scale edge scanning; and a saliency enhancement module for obtaining a refined saliency map. This invention is applicable to remote sensing and UAV image analysis. It can quickly identify salient areas on the Earth's surface, such as buildings, ships, and disaster areas.
Owner:CHANGCHUN UNIV

Target detection method and device in special railway operation scene and storage medium

This invention relates to the field of target detection technology, specifically disclosing a target detection method, device, and storage medium for dedicated railway operation scenarios. The method includes: constructing a training dataset; building a basic target detection model based on the detection accuracy requirements of dedicated railway operation scenarios; optimizing the basic target detection model, including constructing a location-aware space-frequency coordination module in the backbone network to generate a spatial high-frequency saliency map using a dual-stream architecture and local frequency domain changes; constructing a frequency-guided cascaded feature alignment module in the feature pyramid network to guide the geometric offset of the predicted feature map based on the spatial high-frequency saliency map, and deforming and aligning the semantic attention map; and performing multi-round iterative training based on a combined loss function including detection loss and auxiliary constraint loss. The target detection method for dedicated railway operation scenarios provided by this invention can improve the accuracy of detecting small targets in dedicated railway scenarios.
Owner:RIZHAO PORT GRP CO LTD +2

A research method for visual attention mechanisms based on EEG microstates

This invention discloses a method for studying visual attention mechanisms based on EEG microstates, comprising: collecting EEG signals from subjects while watching videos and preprocessing them; extracting saliency maps from the videos, and extracting statistical features from the perspectives of spatial saliency information in the local temporal domain and spatiotemporal saliency information changes in the global time series, respectively, to obtain sIQR features and tsIQR features; converting the preprocessed EEG signals into EEG topology map sequences and performing spatial clustering to extract EEG microstate templates, determining the number of templates and selecting the final microstate templates, and then backfitting them to the EEG signals to obtain microstate sequences; extracting microstate features and depth features from the microstate sequences, examining the statistical differences between microstate features and sIQR and tsIQR features, constructing decoding models based on microstate features or depth features, and using the decoding models to decode segments and videos respectively to obtain segment labels and video labels.
Owner:SHENZHEN UNIV

Metal surface drilling positioning method based on artificial intelligence

ActiveCN121259098BEdge mapsNetwork output
This invention discloses an artificial intelligence-based method for drilling and locating holes on metal surfaces, comprising the following steps: acquiring multiple frames of images of the metal surface to construct an original multi-frame image sequence; performing weighted fusion on the original multi-frame image sequence to obtain a fused image, a specular mask, and an edge map; constructing an improved HorNet backbone network to output multi-scale texture feature maps; inputting the multi-scale texture feature maps into a texture-coordinate joint head to generate a texture saliency map and initial drilling coordinate parameters; inputting the initial drilling coordinate parameters into a projection-aware coordinate transformer to obtain drilling positioning coordinates in the machine tool workpiece coordinate system; constructing a multi-task loss function and performing end-to-end training; inputting the fused image into the trained and converged improved HorNet network to output the drilling positioning coordinates. This invention can achieve high-precision drilling and locating under complex metal surface conditions and is suitable for intelligent manufacturing and CNC machining scenarios.
Owner:XIANGTAN UNIV

Electronic component package defect detection method and detection system based on image acquisition

The application discloses an electronic component packaging defect detection method and system based on image acquisition, and relates to the field of image analysis. The method comprises the following steps: collecting a multi-modal image of an electronic component packaging to be detected; performing pixel alignment on the multi-modal image, and stacking the multi-modal image after pixel alignment as R, G and B channels respectively to generate a three-channel pseudo-color image; inputting the three-channel pseudo-color image into a pre-trained convolutional autoencoder model to obtain a reconstructed three-channel image; analyzing the difference between the three-channel pseudo-color image and the reconstructed three-channel image to obtain an analysis result, and generating a three-channel residual image according to the analysis result; generating a single-channel defect saliency map based on the three-channel residual image; and when there is a connected region with a pixel value exceeding a preset pixel threshold in the single-channel defect saliency map, determining that the connected region is a real defect. The application can effectively reduce the false positive rate and improve the production efficiency.
Owner:伯芯半导体科技(湖北)有限公司

A low signal-to-noise ratio infrared dim target detection method

The application discloses a low-signal-to-noise-ratio infrared dim and weak target detection method and relates to the technical field of image processing. The method comprises the following steps: continuously acquiring an infrared dim and weak target image sequence, setting a target motion feature set, estimating the position of the target in each frame according to the target motion feature set, superimposing the obtained multiple frames of images according to the motion feature set, processing the image group obtained through superimposition to obtain a candidate target saliency map, segmenting the saliency map to extract the candidate target, assigning a uniform motion feature attribute to the obtained candidate target pixels, assigning a gray value from the corresponding image group according to the uniform feature attribute, assigning the mean value of the image group in the pixel direction to the remaining pixels, completing the accurate information enhancement of the target, and detecting and outputting the enhanced image by using a spatial domain detection algorithm. The method can realize the detection of multiple low-signal-to-noise-ratio dim and weak targets and improves the discovery capability of an infrared detection system on a long-distance target.
Owner:SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES

An optimization-based guided filter image fusion change detection method and device

ActiveCN118071772Bguaranteed edgeSuppress abnormal noiseSaliency mapPixel value difference
The application discloses a kind of based on optimization's guiding filter image fusion change detection method and device, method includes: by extreme minimum scale difference operator, pixel value difference value is to the picture of same scene different phase, obtains difference map;Using the way of guiding filter, difference map is fused on global scale, to obtain the optimal difference map;Mean filter is used to phase map I1 and I2, and the base layer difference map corresponding to each difference map is obtained, and the optimized brightness saliency map F1 and F2 are obtained by pca fusion as the image of generating saliency map, according to the principle of maximum saliency to determine weight map;Weight map is regularized for double-scale difference map reconstruction, and the final difference map is obtained;In clustering segmentation stage, by introducing soft threshold function, the final difference map is further processed, to suppress existing abnormal noise.The device includes: processor and memory.
Owner:XINJIANG UNIVERSITY

Adaptive confidence thresholding with saliency-driven refinement

A method for saliency-driven refinement of object detection proposals includes obtaining image data generated by one or more sensors of a vehicle; extracting, from the image data, a plurality of features to generate a plurality of feature maps; upsampling the plurality of feature maps to generate a plurality of upsampled feature maps, wherein the upsampling increases a spatial resolution of the plurality of feature maps; projecting the plurality of upsampled feature maps onto a plurality of Bird's Eye View (BEV) feature maps; generating, based on the plurality of BEV feature maps, a plurality of object detection proposals for the image data; generating one or more saliency maps for the image data; and applying a confidence threshold, based on the one or more saliency maps, to the plurality of object detection proposals to generate refined object detection proposals.
Owner:QUALCOMM INC

A video saliency map generation method, device and storage medium

Embodiments of the present application relate to a video saliency map generation method, device and storage medium. The method comprises: constructing a plurality of Gaussian pyramids containing different scales for each preset feature channel according to a plurality of preset feature channel information of a video frame; wherein the preset feature channels include brightness, color and edge; determining a first saliency map for each preset feature channel according to the Gaussian pyramids of a plurality of preset feature channels; obtaining simulated optical flow by calculating motion information under a preset dimension through inter-frame difference after different level image displacement according to brightness information of a plurality of video frames; extracting a second saliency map of a preset direction and a preset speed based on a center-edge difference method according to optical flow-speed information; and determining a video saliency map according to a static saliency map synthesized according to the first saliency map and a dynamic saliency map synthesized according to the second saliency map. The technical solution of the embodiments of the present application has fast saliency map calculation speed, small calculation overhead, good effect and strong interpretability.
Owner:CHINESE INST FOR BRAIN RES BEIJING

A SAR image target recognition method based on saliency map guidance

The application provides a SAR image target recognition method based on a saliency map guide, comprising: generating a corresponding target saliency map according to an original SAR image, wherein the target saliency map contains target key information; extracting deep features and shallow features in the original SAR image based on the target saliency map, and guiding the deep features and the shallow features to be fused; refining the fused image features by using a multi-layer hollow convolution to obtain target features; and performing type recognition on the target features by using a full connection layer to obtain a target recognition result. The application avoids noise interference in a background image, improves in-class consistency of target features, can efficiently and robustly realize mining of basic features of a target, and improves precision and efficiency of SAR image target recognition.
Owner:BEIJING INST OF TECH

Diamond quality detection method and system based on image recognition

PendingCN122391180APattern recognitionMaximum eigenvalue
The present application belongs to the technical field of image recognition, and particularly relates to a diamond quality detection method and system based on image recognition, comprising the following steps: obtaining a diamond image to be detected, calculating the structure tensor in the local neighborhood of each pixel of the diamond image, constructing an initial structure guide map according to the eigenvalue distribution, and recording the difference between the maximum eigenvalue and the minimum eigenvalue of the structure tensor as an anisotropy degree map; performing multi-scale shear wave transformation on the diamond image, extracting high-frequency subband coefficients at each scale to calculate local energy, and fusing the local energy at each scale to obtain a multi-scale edge saliency map; and fusing the initial structure guide map and the multi-scale edge saliency map to generate a composite guide map. The present application avoids the edge blurring phenomenon caused by traditional filtering, and improves the completeness of diamond surface defect extraction and the accuracy of overall quality detection.
Owner:SHANGQIU LIREN SUPERHARD MATERIAL PROD CO LTD

Gesture recognition and interaction method and system based on complex background

The application provides a gesture recognition and interaction method and system based on a complex background, relates to the technical field of gesture recognition, and comprises the following steps: acquiring an image sequence, counting background and skin color features, calculating a background probability distribution map, and extracting multi-level features to predict a gesture key point heat map. The background probability inverse mapping is used as an initial attention bias, a double-branch attention mechanism integrating skin color and structure saliency maps is constructed, and the key point heat map is used for morphological guidance correction to obtain morphologically constrained attention weights. After the weights are transmitted across levels, the multi-level features are weighted and fused to obtain background-suppressed gesture features, and finally gesture recognition is realized and interaction instructions are generated. The application effectively suppresses complex background interference and improves the accuracy and robustness of gesture recognition.
Owner:BEAVER INTELLIGENT MANUFACTURING (BEIJING) TECHNOLOGY CO LTD

Satellite image compression transmission method and system for reservoir monitoring station image transmission

The application relates to the technical field of image compression transmission, and discloses a satellite image transmission monitoring station image compression transmission method and system for a reservoir, which comprises the following steps: obtaining an original image, removing system errors to obtain a clear image; extracting a waveband edge intensity and a texture gradient, performing a connected domain morphological close operation after double-threshold screening, and determining a core target region boundary; expanding an adjacent pixel outward by counting a boundary pixel histogram to obtain an expansion region; calculating a saliency map based on the region, screening high saliency pixels, performing connected component marking, and generating a protection area mask; performing wavelet transformation on the expansion region, encoding high-fidelity data according to a mask low quantization step, high quantization encoding of a background region to generate compressed data, and fusing the high-fidelity data and the compressed data to obtain a hybrid compressed image; after verification, when a receiving end decodes, core data is losslessly restored, a background region is reconstructed in detail, and a complete image is output. The method can solve the problem of image distortion.
Owner:GUANGDONG WISDOM SHUIYUN TECH CO LTD

Light field image perceptual coding method based on intra coding tree unit level code rate allocation

The application discloses a light field image perceptual coding method based on intra coding tree unit level code rate allocation, which firstly selects part of sub-aperture images in a sub-aperture image array and arranges the part of sub-aperture images into a pseudo video sequence; then obtains a depth map and a saliency map of a center sub-aperture image by using a depth estimation network and a saliency detection network; then calculates a code rate allocation weight of each coding tree unit in the selected sub-aperture image by using the center sub-aperture image, the depth map and the saliency map, and performs target code rate allocation by using the code rate allocation weight; finally, synthesizes a sub-aperture image which is not selected by using a light field angle super-resolution reconstruction network, and combines the sub-aperture image with a decoded sub-aperture image to form a complete decoded light field image; the method has the advantages that perceptual redundancy existing in the light field image is effectively removed, and the visual quality and structural consistency of a salient region can be maintained at a lower code rate.
Owner:NINGBO UNIV

Surgical nursing operation compliance image intelligent detection method and system

PendingCN122313359ASaliency mapImage detection
This invention discloses an intelligent detection method and system for compliance images of surgical nursing operations, belonging to the field of image detection and processing technology. The method includes: first, acquiring operating room monitoring video streams and extracting the spatial mask of sterile areas, the outlines of nursing instruments, and key skeletal points; based on this, constructing a spatial attention matrix and an operational topology diagram of hand-instrument interaction to generate a spatiotemporal feature map to be detected; next, calling compliance benchmark reconstruction logic to reconstruct features, calculating pixel-level feature differences and performing weighted mapping to generate an initial anomaly saliency map; finally, denoising and spatially weighted aggregation are performed on the anomaly saliency map to output the detection results. This invention learns the compliance feature distribution through reconstruction logic, avoiding dependence on scarce violation samples, and enhances the feature representation of high-risk areas and interactive behaviors by utilizing spatial attention and topology, significantly improving the detection sensitivity and accuracy of images of rare violations in surgical nursing operations.
Owner:THE FIRST AFFILIATED HOSPITAL OF TIANJIN UNIV OF TRADITIONAL CHINESE MEDICINE

A multi-modal based salient object detection method and device, and related medium

The application discloses a kind of based on multimodal salient object detection method, device and related medium, the method includes inputting the color image to be detected into visual language model and carrying out multimodal feature processing, obtains semantic alignment feature set;Multi-granularity semantic reasoning is carried out to semantic alignment feature set, and positioning enhanced feature map is obtained;Again, detail feature map extraction is carried out to the color image to be detected by image feature neural network and detail feature map and positioning enhanced feature map are fused, and finally saliency map is obtained.Such that, saliency reasoning process can explicitly introduce semantic information, and is synergized using multi-granularity semantic expert system algorithm and detail recovery fusion algorithm, and the positioning accuracy of saliency map in complex scene is improved.
Owner:ZHEJIANG DIANCHUANG INFORMATION TECH CO LTD +1

An AI recognition-based internet of things device anomaly detection system

This invention discloses an anomaly detection system for IoT devices based on AI recognition, comprising: a data acquisition and sequence construction module for constructing an input sequence; a variable sampling downsampling module for feeding the input sequence into the downsampling path of an improved SCINet, introducing Gumbel-Softmax to compute a time sampling mask; a dynamic temporal interactive attention module for outputting intermediate feature maps; an upsampling reconstruction and saliency generation module for generating an anomaly saliency map; an anomaly saliency feature sparsity module for outputting a reconstructed sequence; a decision feature generation module for obtaining decision features; a Deep SVDD training and detection module for calculating anomaly scores and outputting category labels; and a decision fusion and event recording module for outputting the anomaly time location and device identifier. This invention achieves high-precision anomaly detection on multi-channel time-series data of IoT devices.
Owner:HEFEI HUIMENG CLOUD CHAIN INFORMATION TECH CO LTD

System and method for automatic tagging of images and video in an operative report

PendingUS20260203798A1Operative reportEngineering
Systems and methods for automatic tagging of images and video in surgical streams are described. A plurality of machine learning models, trained on annotated surgical data, are used to extract salient images and video clips from surgical video streams. In addition, speech transcription models process audio streams to generate transcriptions that are then associated with the tagged media. Subsequently, the system synchronizes the multimodal data and generates structured operative records. After synchronization, billing rules are applied to produce accurate billing reports. Applications of the system include improving surgical documentation, reducing administrative burden, enhancing billing accuracy, and accelerating revenue cycles in healthcare environments.
Owner:VAIM TECHNOLOGIES LLC

Engine part defect detection method, device and equipment based on AI vision and medium

The application relates to an AI vision-based engine part defect detection method, device, equipment and medium. The method comprises the following steps: acquiring labeled real engine part images and defect masks to form initial samples, training by using a multi-scale topological perception feature extraction network, generating a deep feature map and a topological saliency map, and based on this, performing defect data coding and feature space synthesis to construct a large-scale synthetic augmented training dataset; taking the trained feature extraction network as a backbone, integrating a dynamic domain self-adaptive module to construct a defect detection framework, obtaining a robust teacher detection model and a parameter offset library through multi-working-condition simulation training, and performing two-stage progressive knowledge distillation processing to obtain a defect detection engine. The method significantly improves the accuracy, robustness and efficiency of engine part defect detection through multi-scale topological perception feature extraction, dynamic domain self-adaptation and progressive knowledge distillation.
Owner:张贺

A method for enhancing spatio-temporal joint local contrast infrared dim small moving target detection

The application provides a kind of enhanced spatio-temporal joint local contrast infrared weak and small moving target detection method, comprising the following steps: S1, obtaining original infrared image sequence;S2, the local contrast of each image space is calculated, S3, the original infrared image sequence of step S1 is extracted with three layers sliding window to the time domain profile line of small target, and the time domain local contrast graph is obtained;S4, the local contrast graph of space and the time domain local contrast graph are fused to obtain target saliency map;S5, weak and small target is segmented from target saliency map by adaptive threshold segmentation, and target information is output.The method of the application utilizes two efficient methods of single frame detection and time pixel profile detection, realizes faster detection speed and stronger detection real-time, and solves the problems of high performance method, high cost and long processing time.
Owner:HANGZHOU INST FOR ADVANCED STUDY UCAS

A Method and System for Building Roof Defect Detection Based on Unmanned Aerial Vehicles

ActiveCN121329878BEnhance edge response continuityImprove discriminationImage analysisOptically investigating flaws/contaminationSaliency mapMachine vision
This invention relates to the field of machine vision, and in particular to a method and system for detecting defects on building roofs based on unmanned aerial vehicles (UAVs). The method involves acquiring grayscale images of building roofs captured by the UAV; calculating morphological gradient maps in multiple directions of the grayscale image; calculating the variance of each pixel across all gradient values ​​to generate a gradient direction consistency map; and weighted fusion of all morphological gradient maps based on the gradient direction consistency map to generate an edge saliency map. Within the neighborhood of each pixel in the grayscale image, the frequency of grayscale pairs is attenuated and weighted according to the Euclidean distance between neighboring pixels and the center pixel, and spatial location-sensitive entropy is calculated based on the weighted co-occurrence probability to generate a texture complexity map. Further, a defect response map is obtained; a local discrimination threshold corresponding to each pixel is calculated; and when the value of a pixel in the defect response map is greater than the corresponding local discrimination threshold, the pixel is determined to be a defect point.
Owner:CHINA CONSTR FIFTH ENG DIV CORP LTD

A medical image segmentation method based on information guidance and boundary perception

PendingCN122289682Aprecise positioningaccurate diagnosisSaliency mapRadiology
This invention proposes a medical image segmentation method based on information guidance and boundary awareness. The method first performs data augmentation to expand the training set; then, a Transformer encoder is used to extract multi-scale features from the image. To accurately locate lesions, an information fusion module is designed, utilizing deep feature fusion to obtain global semantic information. Addressing the challenge of boundary segmentation, a boundary awareness module is introduced, processing boundary regions through multi-scale strip-shaped depthwise separable convolutions and achieving accurate prediction guided by location information. Simultaneously, a global / local feature extraction module processes intermediate layer features, filtering out useless information; and a channel multi-scale module is used to mine deep global semantics at the channel level. Finally, a region fusion module is set at the decoder end, combining boundary information and saliency maps to strengthen uncertain regions and improve the segmentation results. Experiments demonstrate that this method outperforms existing mainstream technologies and effectively improves the problem of difficult lesion boundary segmentation in medical images.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY

Infrared Small Target Detection Method Based on Gradient Direction Difference and Thermonuclear Diffusion

This invention discloses a method for detecting small infrared targets under complex background conditions, comprising the following steps: calculating the local gradient direction of the target image using a window method; weighting the residual saliency image based on the local gradient direction of the target image; extracting segmented regions from the weighted residual saliency image based on brightness information and defining potential target seed points based on local features; enhancing the energy response of suspicious regions through thermal kernel diffusion, thereby improving the contrast between the target and the background; describing pixels considered to belong to the target region in the segmented thermal kernel diffusion energy map, amplifying the difference between the target region and the background region, and generating an enhancement map; fusing the weighted residual saliency image and the enhancement map, and achieving the final target saliency map through adaptive segmentation. This invention has the advantages of strong robustness, low false alarm rate, and high detection accuracy.
Owner:NANJING UNIV OF SCI & TECH +1

A method for counting same-color blocks based on machine vision

PendingCN122335650APattern recognitionDensity based
This invention discloses a machine vision-based method for counting blocks of the same color, comprising the following steps: image acquisition and preprocessing; color space transformation and focusing to generate a color feature saliency map; generation of a two-dimensional spatial-color distribution vector set; seed point localization based on density peak search; object region division and final counting based on adaptive region growing; and output of the counting results. This invention completely eliminates the reliance on prior information about object shape and complete contours, utilizing color as a stable feature as the core clue. It effectively solves the problems of contact and occlusion between objects, accurately separating and counting even tightly clustered objects. It avoids complex and time-consuming model training processes, achieving a lightweight and highly robust solution based on traditional image processing and intelligent algorithms. It improves the overall robustness and adaptability of the system, enabling stable operation even under slight changes in lighting and relatively complex backgrounds, and is easily deployed on low-cost hardware platforms.
Owner:HANGZHOU HUICUI INTELLIGENT TECH CO LTD

Request tracking methods, systems, apparatuses, storage media, and computing devices

This application provides a request tracing method, system, apparatus, storage medium, and computing device. The method includes: using a first intelligent agent on a microservice host to collect request processing data of various requests related to each deployed microservice component, and caching it in a local cache of a preset size on the microservice host; the request processing data includes request span metadata and system resource information; generating a sequence of indicators to be detected corresponding to various indicators to be detected based on the request processing data, and using a saliency map of the indicator sequences to be detected to perform anomaly detection on the indicator sequences; if anomalies are detected in at least one indicator sequence to be detected, sending historical request processing data stored in the local cache and newly collected request processing data to an analysis server in the microservice architecture, and sending a data upload notification to a second intelligent agent on other microservice hosts related to the microservice host.
Owner:TSINGHUA UNIVERSITY

Driver attention prediction method and system based on semantic guidance multi-layer spatio-temporal feature fusion

The application discloses a driver attention prediction method and system based on semantic guidance multi-layer spatio-temporal feature fusion. The method comprises the following steps: acquiring a continuous traffic video segment, and uniformly adjusting the continuous traffic video segment to a fixed size input into a trained driver attention prediction network; adopting a Video Swin Transformer as a backbone network to extract multi-layer spatio-temporal features, and acquiring shallow layer spatial details and deep layer semantic context; guiding deep layer semantic information to flow to the shallow layer and promoting semantic information transmission through multi-level residual connection and up-sampling operation; introducing a hierarchical feature reorganization module to adaptively reconstruct each layer feature after fusion, highlight key salient regions, suppress redundant irrelevant features, and enhance saliency expression; respectively deconvolving and decoding the hierarchical features of four independent branches to generate an intermediate saliency map, and finally splicing and fusing to generate a final prediction result. The method can accurately predict the driver attention distribution, and has important significance for the development of advanced auxiliary driving.
Owner:SHIJIAZHUANG TIEDAO UNIV +1

Video frame adjusting method and electronic device

The application discloses a video frame adjusting method and an electronic device. The method establishes time domain correlation through a dense motion field between saliency maps, thereby generating a predicted saliency map. This process simulates the continuity of visual attention in time sequence, and provides a more accurate benchmark for evaluating the difference between the actual saliency map and the predicted saliency map. The actual saliency map and the predicted saliency map of a frame image are analyzed to determine the offset of the quantization parameter of the frame image, and the deviation between the actual distribution of the current frame image attention and the prediction is quantified. Based on this, the code rate is adjusted, and the visual experience is optimized.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

A garden plant disease and pest pattern recognition method and system

This invention discloses a method and system for pattern recognition of garden plant diseases and pests, belonging to the field of pattern recognition. The method first converts the acquired garden plant images into a color space composed of opposing brightness and color components, and initially locates candidate leaf regions using a preset threshold range. Simultaneously, visual saliency detection is performed based on frequency domain analysis to generate a saliency map and mark salient regions. A logical AND operation is performed between the candidate leaf regions and the salient regions, combined with morphological closing operations to generate a binary mask for background suppression, extracting the target main body region of the leaf. Subsequently, a parallel multi-branch feature extraction network is used to extract multi-dimensional features from the leaf regions. This network includes three branches using standard convolution, a dilation rate of 2, and a dilation rate of 4 to capture disease features at different scales. This invention improves the accuracy and robustness of disease and pest identification.
Owner:SHENZHEN FUQING ECOLOGICAL TECH CO LTD

Spatio-temporal saliency cueing method and apparatus for open vocabulary video action recognition

This invention provides a spatiotemporal saliency cues method and apparatus for open-vocabulary video action recognition, relating to the field of computer vision technology. It includes: inputting an input video to a video-level spatiotemporal saliency cues module; obtaining a coarse-grained saliency map through cross-modal aligned attention branches; obtaining a fine-grained saliency map through instance-level self-attention branches; fusing the two to obtain a video-level spatiotemporal saliency cues map; performing element-wise multiplication with the input video and introducing residual connections to obtain enhanced discriminative action features; injecting the fine-grained saliency map into the input video to generate spatiotemporal action features; and inputting these features into a token-level spatiotemporal saliency cues module to obtain the open-vocabulary video action recognition result. This invention proposes a novel spatiotemporal saliency cues method that can dynamically focus on instance-level discriminative action regions and generate multiple levels of action saliency cues that complement each other, thereby significantly enhancing the model's generalization ability.
Owner:UNIV OF SCI & TECH BEIJING