Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Visual attentiveness" patented technology

Image feature enhancement method and system based on learnable unary function gating

PendingCN122367778ARadiologyImaging Feature
This invention relates to the field of image feature enhancement technology, providing an image feature enhancement method and system based on learnable unary function gating. The method includes: dividing the input image into image patches and mapping them to visual tokens; in a multi-head attention layer, calculating a visual attention aggregation score based on an attention weight matrix to quantify the degree of abnormal attention received by the image patch; inputting the normalized visual token, aggregation score, and two-dimensional position code into a gating module composed of a learnable unary function to generate a gating matrix; after each attention head completes SDPA output and before multi-head stitching, performing element-wise gating modulation using the gating matrix, and stitching and projecting to obtain the enhanced image features. This invention, through the synergy of aggregation score and learnable unary function, suppresses abnormal attention propagation from background noise, enhances the expression of key features of small targets, and improves the feature discriminativeness and task adaptability of the visual Transformer.
Owner:CCTEG BEIJING HUAYU ENG

A research method for visual attention mechanisms based on EEG microstates

This invention discloses a method for studying visual attention mechanisms based on EEG microstates, comprising: collecting EEG signals from subjects while watching videos and preprocessing them; extracting saliency maps from the videos, and extracting statistical features from the perspectives of spatial saliency information in the local temporal domain and spatiotemporal saliency information changes in the global time series, respectively, to obtain sIQR features and tsIQR features; converting the preprocessed EEG signals into EEG topology map sequences and performing spatial clustering to extract EEG microstate templates, determining the number of templates and selecting the final microstate templates, and then backfitting them to the EEG signals to obtain microstate sequences; extracting microstate features and depth features from the microstate sequences, examining the statistical differences between microstate features and sIQR and tsIQR features, constructing decoding models based on microstate features or depth features, and using the decoding models to decode segments and videos respectively to obtain segment labels and video labels.
Owner:SHENZHEN UNIV

A steel plate surface defect recognition method fusing visual attention mechanism

The application provides a steel plate surface defect recognition method fusing visual attention mechanism, and belongs to the technical field of defect detection based on computer vision; multi-source images of a steel plate, production line process and quality detection data are synchronously collected to construct a standardized tensor benchmark with aligned physical attributes and unified data modalities. A visual attention feature coding network fusing physical information constraints is built, and PINN constraint correction is used to remove the interference of ambient light brightness gain, so that the intrinsic reflection attribute features of defects are accurately extracted. Through spatiotemporal collaborative alignment and correlation modeling of defects and process heterogeneous data, the contribution weight of process parameters to defects is quantified by adopting double-track machine learning, and a process target evaluation baseline is built. Finally, the defect topology reasoning and multi-task prediction are completed by relying on the graph attention network GAT, the local visual features and global process constraints are deeply fused, and the defect position mask and category recognition result are output; the application significantly improves the steel plate quality inspection accuracy and system decision reliability.
Owner:RIZHAO YULAN NEW MATERIAL CO LTD

A traffic car customer service marketing method and system based on a large language model

ActiveCN120851956Baccurate perceptionaccurate quantitative analysisInput/output for user-computer interactionBiological modelsPersonalizationData set
The application relates to the technical field of intelligent interaction systems and automobile customer service marketing, and discloses a traffic automobile customer service marketing method and system based on a large language model, wherein the traffic automobile customer service marketing method based on the large language model comprises the following steps: collecting driver fixation point data by using an eye movement tracking algorithm to generate eye movement trajectory data sets; constructing a visual attention heat map to form a user attention distribution model; calculating multi-medium characteristic parameters to establish a medium characteristic model; using the large language model to generate and adapt personalized marketing content and output multi-medium compatible marketing information; realizing intelligent attention guidance according to the user attention distribution model, the medium characteristic model and the multi-medium compatible marketing information; and realizing accurate perception and quantitative analysis of the visual attention of the driver through the eye movement tracking technology, so that the marketing system can master the user focus points in real time, and the technical problem that a traditional interface cannot perceive actual user focus points is solved.
Owner:CHINACHEM PUHUI (CHANGCHUN) DATA SERVICE CO LTD

An interface determination method and device, a storage medium and an electronic device

PendingCN122285167AImprove adaptation efficiencyImaging processingSimulation
This application discloses an interface determination method, device, storage medium, and electronic device, relating to the technical fields of image processing, artificial intelligence, and mobile application interface adaptation. It simulates the user's natural visual scanning process in a right-to-left reading mode by constructing a visual attention energy field. During this process, a directional cognitive flow model is used to process the directional cognitive flow, obtaining the reading path and operation decision path formed under the current interface. Based on the modeling of a preset interface structure model, spatial relationships, and potential paths, the current interface is matched and determined to be consistent with preset interface habits. This allows for the determination of the actual user experience of the current interface in real-world use. Furthermore, this solution does not rely on testers with language and cultural backgrounds for manual experience and repeated verification during the matching and determination of preset language interface habits, thereby improving the efficiency of verifying Arabic interface adaptation.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Method and device for literacy development

A literacy development device includes a light source that provides a general white illumination area and an adjustable light within the illumination area, both from the same light source. The adjustable light can be changed in shape and color through a LCD screen provided within the device. The device contains a rechargeable battery and various buttons for changing functions, light parameters, and other controls. The device is preferably usable in teaching applications, especially with children and other students with learning and / or attention issues. The device is useful in drawing attention to certain areas of a surface that includes written language, including a letter, a group of letters, a word, or a sentence, through use of the adjustable light to highlight, underline, or otherwise draw visual attention to a particular area of the surface within the illumination area.
Owner:BRITE LITER INC

A digital marketing user portrait generation method and system based on deep learning

The application discloses a kind of digital marketing user portrait generation method and system based on deep learning, it is related to digital marketing technical field, including, to original fine-grained interaction event flow is standardized, and the output standardized user microcosmic behavior time sequence;Standardized user microcosmic behavior time sequence is input to deep time sequence neural network, and the output user cognitive load quantification score and visual attention focus distribution vector;Based on user cognitive load quantification score and visual attention focus distribution vector, to static long-term interest portrait is weighted and focused fusion, generates situational dynamic user portrait vector;Situational dynamic user portrait vector is input counterfactual explanation engine.The application solves the problem that the marketing communication opportunity and form adaptability caused by lack of real-time psychological state assessment of user in prior art, improves the precision of digital marketing, user experience and communication efficiency.
Owner:CHINA NAT INST OF STANDARDIZATION

Method and device for autofocusing an object by a front camera

The application provides an automatic focusing method and device for focusing on an object by a front camera, which comprises the following steps: performing line-of-sight estimation according to face information, eye information and head information collected by the front camera on a photographer to obtain a 3D gaze ray; capturing a current picture by a rear camera; calculating the intersection of the 3D gaze ray and the current picture; determining the position of the intersection in the current picture, and driving the rear camera to focus on the position in the current picture according to the intersection. The 3D gaze ray of the photographer is determined by the front camera, the intersection of the 3D gaze ray on the current picture captured by the rear camera is determined, then the visual attention of the photographer on a certain object to be photographed is determined, and finally the object to be photographed is focused. The automatic focusing shooting with the photographer's attention as the center is realized, the accuracy of matching the shooting image with the shooting intention of the photographer is improved, and the relevance of the shooting imaging is enhanced.
Owner:SHENZHEN GUANGQIZHIJING TECHNOLOGY CO LTD

Detecting visual attention during user speech

An example process includes: concurrently receiving an audio stream and a video stream; determining, based on a first portion of the audio stream received within a predetermined duration before a current time and a first portion of the video stream received within the predetermined duration before the current time, whether a visual attention of a user is directed to an electronic device while the user is speaking; and in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is speaking: identifying a second portion of the audio stream to include user speech intended for the electronic device; initiating, by a digital assistant operating on the electronic device, a task based the second portion of the audio stream; and providing an output indicative of the initiated task.
Owner:APPLE INC

Detection of visual attention during user utterances

An exemplary process includes receiving an audio stream and a video stream simultaneously, determining whether the user's visual attention is directed to the electronic device while the user is speaking, based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within a predetermined duration before the current time, identifying a second portion of the audio stream to include the user utterance addressed to the electronic device in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking, starting a task based on the second portion of the audio stream by a digital assistant operating on the electronic device, and providing an output indicating the started task.
Owner:APPLE INC

Video frame adjusting method and electronic device

The application discloses a video frame adjusting method and an electronic device. The method establishes time domain correlation through a dense motion field between saliency maps, thereby generating a predicted saliency map. This process simulates the continuity of visual attention in time sequence, and provides a more accurate benchmark for evaluating the difference between the actual saliency map and the predicted saliency map. The actual saliency map and the predicted saliency map of a frame image are analyzed to determine the offset of the quantization parameter of the frame image, and the deviation between the actual distribution of the current frame image attention and the prediction is quantified. Based on this, the code rate is adjusted, and the visual experience is optimized.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

A method and system for achieving integrated dimming

This invention provides a method and system for achieving comprehensive dimming, relating to the field of human-computer interaction technology. The method includes: generating and fixing a static basic mapping table for each display device; extracting the display content features of each display device in real time, as well as the visual attention weight of the operator in the current scene; calculating the global physiological confidence level to verify glare sources; combining the visual attention weights, performing spatial pairing and cross-referencing on each display device to calculate the causal matching degree, generating a spatially selective dynamic attenuation coefficient for each display device; and distributing dimming instructions containing the spatially selective dynamic attenuation coefficient to the corresponding display devices via a hybrid communication network. Each display device performs dynamic compensation and low-level nonlinear inverse mapping by combining its own real-time state parameters with the static basic mapping table, and outputs a pulse width modulation duty cycle for driving light emission.
Owner:SHANGHAI ZHONGCHUAN SDT-NERC CO LTD

UI world state construction interaction method and system based on human visual perception

The present application relates to the technical field of human-computer interaction, in particular to a UI world state construction interaction method and system based on human visual perception, comprising: dividing the interface perception process into discrete perception cycles, collecting multi-source perception data in parallel in each cycle and converting into candidate object descriptors with uniform format. By calculating the theoretical perception boundary combined with the viewport, scroll position and visual focus area, the candidate object set that can be perceived by the user is screened out, and correlation analysis is carried out to construct a unified world state. The audit log of the traceable decision process is generated according to the object source and confidence. The method reduces the data processing amount by simulating the visual attention mechanism, improves the state construction efficiency and accuracy, and at the same time, the process audit enhances the credibility of the state model and the maintainability of the system.
Owner:杨思嘉

Emergency training virtual reality closed-loop regulation system and method based on multi-modal perception fusion

PendingCN122337058ASimulationMulti source data
This invention discloses a multimodal perception fusion-based emergency training virtual reality closed-loop control system and method, belonging to the field of virtual reality technology. It includes: a multi-source data synchronous acquisition module, which synchronously acquires trainees' eye-tracking data, visual three-dimensional skeletal posture key point data, and wrist-worn physiological sensor data; a virtual-real space dynamic registration and multimodal feature extraction engine, which calculates mapping relationships and extracts feature vectors representing visual attention allocation patterns, kinematic features of movements, and autonomic nervous system stress response parameters; a real-time state quantification assessment module, which outputs a real-time state assessment vector containing a comprehensive risk avoidance index and a state stability coefficient; a virtual environment dynamic control module, which matches teaching intervention strategies and generates scene control instructions; and an immersive situational cognition guidance and feedback module, which, when the state stability coefficient is determined to be lower than a preset stress safety threshold, prohibits the triggering of new preset environmental disaster events and activates a parameterized flexible situational guidance mechanism.
Owner:TIANJIN YUNLAN INTERNET OF THINGS TECHNOLOGY CO LTD

Advertisement effect evaluation method and system based on eye movement data visualization

PendingCN122288796AInformation processingPupil diameter
This invention relates to the field of commercial information processing technology and discloses a method and system for evaluating advertising effectiveness based on eye-tracking data visualization. Addressing the problems of existing evaluation methods' difficulty in objectively quantifying visual attention, and the high cost, complex operation, and lack of customized advertising analysis functions of general-purpose eye-tracking analysis software, this invention configures an eye-tracking device and performs multi-point calibration to simultaneously collect data such as fixation point coordinates, duration, and pupil diameter. Through multidimensional cleaning, fixation event extraction, and coordinate normalization, a fixation heatmap is generated using Gaussian kernel density estimation, and a fixation trajectory map is generated through temporal interpolation. The invention also automatically calculates indicators such as the proportion of fixation duration in the core focal area, overlap, and the proportion of blind spots. This invention achieves full automation from data collection to report generation, objectively and accurately revealing the distribution of visual attention in advertising, transforming physiological data into intuitive design basis, and providing a standardized and repeatable quantitative tool for advertising effectiveness evaluation.
Owner:CHONGQING UNIV OF TECH

An Automatic Method for Generating Brain CT Medical Reports Based on Hierarchical Attention Based on Co-occurrence Relationships

This invention discloses an automatic generation method for brain CT medical reports based on co-occurrence relation hierarchical attention. The method preprocesses the brain CT dataset and establishes a vocabulary; constructs a feature extractor for brain CT images to extract visual features; and builds a co-occurrence relation semantic attention module to extract semantic attention features of common medical terms in brain CT images, which includes a word embedding layer and a semantic attention mechanism. A topic vector-guided visual attention module is also constructed, where topic vectors fuse semantic information from common and rare medical terms to fully express sentence-level medical terminology topics. These topics then guide the visual attention mechanism to capture important lesion region features. This method combines the co-occurrence relationships between common medical terms to infer missing semantic information, thereby extracting richer semantic attention features. This hierarchical collaboration improves the accuracy and diversity of the generated brain CT medical reports.
Owner:BEIJING UNIV OF TECH

Method for assessing the need to alert a driver to a risky situation

Method for assessing the need to alert a driver to a risky situation. Method (100) for assisting the driving of a motor vehicle implementing an assessment of the need for an alert to determine whether or not it is necessary to alert a driver to a risky situation, comprising the following steps: - Determination (101) of a line of sight (21) of the driver; - Reception (102) of an image (22) of a scene surrounding the vehicle, and determination (103) of the point of intersection (23) with the line of sight; - Calculation (104) of a two-dimensional Gaussian (24) centered on this point of intersection, the abscissas of the Gaussian corresponding to the coordinates of the points in the image of the scene; - Determination (108) of a visual attention function (28) by deformation of said Gaussian (24), using data relating to an overall alerting need, and intrinsic characteristics of elements of interest on the surrounding scene.- For each element of interest (25i), calculation (109) of a degree of attention which depends on a value of the visual attention function (28) at the location of the element of interest, and comparison (110) with a threshold (Th_1) to determine whether or not it is necessary to alert the driver to a risky situation. Figure 1.
Owner:AMPERE SAS

A method and apparatus for optimizing direct preferences in CT reports based on visual attention masks.

ActiveCN122091065Asuppress false positivesSuppress false negative hallucinationsImage analysisNatural language data processingImaging processingVision based
This application discloses a method and apparatus for direct preference optimization (DPO) of CT reports based on visual attention masking, relating to the field of medical image processing technology. The method includes: actively blocking access to visual information through a visual attention masking mechanism, causing the model to generate hallucinatory outputs based on its own linguistic prior knowledge while visual access is blocked, thus forming non-preferred reports. This effectively suppresses false positive and false negative hallucinations. Furthermore, a reference-guided generation strategy ensures a high degree of consistency in syntactic structure between preferred and non-preferred reports, allowing DPO training to focus on distinguishing factual content from hallucinatory descriptions, avoiding shortcut learning, and significantly improving hallucination suppression efficiency. This solves the problems of existing technologies that over-rely on linguistic prior knowledge leading to factual hallucinations, lack of mechanisms to construct non-preferred samples by combining visual evidence, and inability to improve clinical accuracy of the model without requiring a large amount of additional labeled data.
Owner:BEIHANG UNIV

Network access identity auditing method and system based on multi-modal data fusion

This invention provides a method and system for network access identity verification based on multimodal data fusion. The method includes: collecting multiple modal data from several known users during the network access process; extracting features from each modal data to obtain a basic feature set corresponding to each modal data; reconstructing the basic feature sets of each modality into a phase space trajectory matrix, and generating a four-dimensional feature matrix based on the phase space trajectory matrix; performing chaotic mapping expansion, slicing, and concatenation on the four-dimensional feature matrix to obtain a concatenated vector, and stacking all the concatenated vectors row-wise to obtain a two-dimensional fusion feature matrix; and training an initial multimodal fusion identity verification model based on the two-dimensional fusion feature matrix. This invention, through multimodal fusion, deeply extracts and fully amplifies the comprehensive differences between minors and adults in terms of operating habits, vocalization states, visual attention, and keystroke mechanics, significantly improving the recognition accuracy.
Owner:CHINA UNICOM (JIANGXI) IND INTERNET CO LTD

A visual effect presentation method of urban green landscape based on greenness rate

ActiveCN120821367BAccurately obtain dynamic feelingsLandscape designUrban green space
The application discloses a kind of urban green landscape visual effect presentation methods based on green feeling rate, including main steps are using active panoramic camera equipment, capture the complete panoramic image of urban green space from the perspective of pedestrian, then build interactive virtual reality environment of different types of urban green space, record the visual attention distribution data of pedestrian by eye tracking device, then calculate the correlation between y and x in visual attention distribution data using linear regression module, obtain the urban green landscape design scheme data to be presented and establish three-dimensional model, subsequently the model is imported into virtual reality engine and builds virtual reality scene, then select panoramic angle picture in the scene, then the picture is imported into picture processing software, the percentage of fixation point number in each Euclidean distance interval is calculated, and each Euclidean distance interval is labeled with different colors according to the percentage value, to realize the visualization of green feeling rate in the module.
Owner:TSINGHUA UNIVERSITY

A pluggable bionic optical neural network system with self-adaptive perception capability

This application relates to the field of biomimetic neurotechnology and proposes a pluggable biomimetic optical neural network system with adaptive sensing capabilities. The system includes a laser source, a pinhole amplifier, a digital micromirror device, a beam splitter prism unit, a photodetector, a pluggable metasurface structure, a spatial light modulator (SLM), a mirror, and a digital camera with a charge-coupled device (CCD) image sensor. This application can adjust the state of the pluggable metasurface structure according to the contrast of the input image light, and adaptively extract high-frequency information of the object based on the contrast of the surrounding environment, thereby improving the accuracy of target recognition tasks. Furthermore, the activation level of neurons in the diffraction layer can be updated in real time based on the initial classification result of the input image light, improving the dynamic adaptive adjustment capability of visual attention, and thus achieving adaptive adjustment to the perceived target.
Owner:SOUTH CHINA NORMAL UNIV

Eye movement fixation duration weighted natural image roi division method and system

This invention relates to the field of image processing technology, specifically disclosing a method and system for natural image ROI segmentation based on eye-tracking gaze duration weighting. Addressing the discrepancy between the segmentation results of traditional methods and the actual attention regions, this method extracts effective gaze points and their gaze durations through eye-tracking data preprocessing. A gaze duration-weighted heatmap is constructed using density peak clustering and a gaze duration weighting model, effectively enhancing the feature representation of high-attention regions. By fusing the eye-tracking heatmap with multimodal semantic features such as color, texture, and edge, the semantic consistency of ROI segmentation is significantly improved. Furthermore, by introducing image complexity adaptive adjustment of the segmentation threshold, flexible adaptation to the segmentation needs of images with varying complexity is achieved. Experimental results demonstrate that the performance of this method and system is significantly superior to traditional methods, thus providing solid theoretical and practical support for intelligent ROI segmentation of high-resolution natural images oriented towards visual attention.
Owner:CHONGQING UNIV OF TECH