Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Visual attention" patented technology

A basketball video highlight automatic generation method and system based on multi-modal feature fusion

PendingCN122269082AImplement depthAchieve precise integrationBiological modelsCharacter and pattern recognitionTimestampEngineering
The application relates to a basketball video highlight automatic generation method and system based on multi-modal feature fusion, and belongs to the technical field of computer vision and multimedia intelligent analysis. The method comprises the following steps: S1, extracting and aligning a visual feature sequence and an audio feature sequence; S2, using bidirectional LSTM to perform time sequence coding on the audio-visual feature, constructing a cross-modal attention mechanism taking audio as a query and vision as a key value, and obtaining multi-modal fusion features; S3, inputting the fusion features into a 1D CNN to output a frame-by-frame highlight event probability sequence; S4, performing peak value searching based on an adaptive threshold of the probability sequence to obtain a coarse-grained timestamp; S5, taking the coarse-grained timestamp as a center to intercept local features, inputting the local features into bidirectional LSTM to predict starting and ending offsets, and obtaining accurate start and end time boundaries; and S6, clipping and splicing to generate a highlight video according to the accurate boundaries. The application realizes audio-visual semantic depth alignment through an audio-guided visual attention mechanism, and improves boundary positioning accuracy through a coarse-precision two-stage positioning architecture.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A mixed reality data processing and interaction response method, device and system

PendingCN122336209APersonalizationMixed reality
This invention discloses a mixed reality data processing and interactive response method, apparatus, and system. The method includes: acquiring corresponding spatial feature points in VR and AR environments; optimizing the solution of rigid body transformation matrices to achieve sub-millimeter-level spatial alignment; continuously monitoring alignment errors during virtual-real fusion rendering, triggering a repositioning process when errors exceed a threshold; evaluating user operations through multimodal data fusion and generating real-time AR correction guidance; and dynamically adjusting rendering parameters and resource preloading based on visual attention focus to ensure end-to-end latency and interactive feedback latency are both below set thresholds. The apparatus includes modules for feature point acquisition, spatial alignment calculation, error monitoring and repositioning control, multimodal data interface, and real-time rendering control. The system includes a server and a user interaction terminal, supporting intelligent training and adaptive interactive optimization for virtual-real fusion. This application addresses problems such as low virtual-real spatial alignment accuracy, high interactive latency, and insufficient personalized adaptation.
Owner:CHINA LIFE INSURANCE CO LTD

Intelligent structured medical record generation method and system based on multi-modal doctor-patient interaction

The application provides a kind of intelligent structured medical record generation method and system based on multimodal doctor-patient interaction, which comprises: real-time acquisition of dialogue voice and transcription into text sequence, while recognizing visual attention entity by listening to mouse operation in electronic medical record system.Based on the history of visual attention entity and text sequence, a logical demonstration track is constructed, and an implicit reward function is derived from it using a reverse reinforcement learning algorithm. Use the function to calculate the action reward value of each combination of visual attention entity and dialogue text, select the highest value combination as the optimal alignment strategy to determine the timing causal relationship. Map the entity and text to the medical knowledge graph, extract the shortest semantic path as the implicit clinical reasoning chain, and generate structured electronic medical record. The application deduces the diagnosis and treatment decision logic from the doctor's multimodal behavior through reverse reinforcement learning, solving the technical problem that traditional methods cannot establish the internal causal relationship between the doctor's visual attention focus and spoken content.
Owner:WUHAN SHENGBOHUI INFORMATION TECH CO LTD +1

Anti-cant indication system for shooting devices

The present invention relates to an audio-based anti-cant indication system for shooting devices such as rifles, pistols, bows, and crossbows. The system includes an inclination sensor configured to detect angular tilt relative to gravity and a processing module configured to compare the detected tilt against a predefined deadband threshold corresponding to a plumb orientation. When the shooting device deviates from the plumb condition, the processing module activates an audio output module that generates distinct audio signals representing left-tilt and right-tilt conditions, while producing a null or muted output when the device remains within the deadband. The system may further include variable tone pitch or volume, wireless audio transmission to earbuds or headsets, calibration mechanisms, front-rear inclination detection, and idle-state shutoff. The invention provides continuous, hands-free cant feedback without requiring visual attention.
Owner:STURDIVANT CHARLES N

Systems, media, and methods for scene enhancement with respect to visual attention

PendingCN122349647AAttention modelUser input
A method can include receiving (i) a visual representation of a scene or image data and (ii) user input specifying one or more goals for modifying the visual representation. The method can also include executing a visual attention model on the visual representation and testing combinations of scene modifications. The method can also include evaluating the impact of the scene modifications on achieving the received one or more user goals and outputting one or more recommended sets of one or more visual representation modifications that satisfy the received one or more user goals upon receiving confirmation that the one or more user goals have been achieved.
Owner:3M INNOVATIVE PROPERTIES CO

Marketing video effect evaluation method combining eye tracking and emotion recognition

The application provides a marketing video effect evaluation method combining eye movement tracking and emotion recognition, and relates to the technical field of marketing effect evaluation, and is characterized in that the method comprises the following steps: multi-modal data synchronous acquisition, eye movement data processing and visual attention area extraction, facial expression recognition and emotion state analysis, multi-modal data space-time alignment and feature fusion, multi-dimensional evaluation index calculation, video segment level effect evaluation, comprehensive effect evaluation, generation of an evaluation report and optimization suggestions.The application has the following advantages: multi-dimensional evaluation indexes are constructed, the marketing video effect can be more objectively, comprehensively and accurately evaluated, multi-dimensional, fine and objective evaluation of the marketing video effect is realized, targeted and personalized optimization suggestions are given and an effect improvement value is predicted, and a scientific basis is provided for optimization and iteration of the marketing video.
Owner:BEIJING POLYTECHNIC

In-game data display methods, systems, devices, and media

This application discloses a method, system, device, and medium for displaying in-game data, relating to the technical field of games. The method includes: acquiring scene state information of a target player in the current game scene, wherein the scene state information includes first state information of a target character, field-of-view parameters of a virtual camera, and second state information of multiple virtual objects within the field of view, the target character being the character controlled by the target player; performing semantic mining on the scene state information to obtain a personalized visual attention vector for the target character, and determining a standard visual attention vector corresponding to the current game scene based on the scene state information; acquiring the target player's historical scene state information, and constructing a standard visual prediction model based on the historical scene state information; and traversing a set of target rendering parameters based on the standard visual prediction model and a genetic algorithm. This application has the effect of improving the player's gaming experience.
Owner:NEXT TECHNOLOGY (CHENGDU) CO LTD

A logistics transportation system with electronic fence and a data calibration method thereof

The application relates to a logistics transportation system with an electronic fence and a data calibration method thereof, and relates to the technical field of intelligent workshop logistics transportation calibration. The intelligent workshop logistics transportation system comprises an intelligent fence management terminal, which is used for virtually dividing a visual attention fence area according to an actual conveying line and converting the visual attention fence area into actual attention fence coordinates; a dynamic visual acquisition terminal, which is used for positioning the conveying position of an article, triggering dynamic acquisition of the article in combination with the actual attention fence coordinates, and obtaining an article video stream; a monitoring and identification execution terminal, which is used for cleaning and identifying the article video stream, obtaining an abnormal article, generating an abnormal state code, and completing abnormal removal in combination with an external execution mechanism; and a digital twin optimization terminal, which is used for creating a digital production line according to the production line layout in combination with the visual attention fence area, and simulating and optimizing the fence layout. Through real-time tracking of the position of the article and calculation of the distribution center of the article, an offset correction value is generated, and calibration data is dynamically updated, so that self-calibration and self-adaptive optimization of the electronic fence are realized.
Owner:YUANYUZHI INFORMATION TECHNOLOGY (KUNSHAN) CO LTD

A core subgraph-based training-free multi-modal large language model fine-grained positioning method

PendingCN122451086ALinguistic modelAlgorithm
The application belongs to the technical field of computers, and particularly relates to a core subgraph-based training-free multi-modal large language model fine-grained positioning method. The method comprises the following steps: 1) using a pre-trained multi-modal large language model to respectively perform feature coding on input images and texts, and respectively dividing the images and texts into a plurality of word units to construct visual and text word unit vector sequence sets; 2) constructing a visual semantic graph structure according to the semantic similarity relationship between visual word units, taking each visual word unit as a graph node, and taking the similarity between the nodes as an edge weight, so as to form a normalized adjacency matrix; 3) calculating the importance score of each visual word unit, and taking a region specified by a user in an arbitrary shape as a core subgraph; 4) performing a visual attention diffusion process in the visual semantic relationship graph based on the core subgraph, and determining a subgraph according to the importance distribution after diffusion; and 5) inputting the screened subgraph and the coded text word unit into a large language model to generate an answer through decoding.
Owner:NANKAI UNIV

Intelligent safety helmet real-time decision algorithm optimization method based on edge computing

PendingCN122454502APacket lossVideo image
The present application relates to the field of data processing, specifically to an intelligent safety helmet real-time decision algorithm optimization method based on edge computing, comprising the following steps: acquiring a video image to establish a gradient matrix, generating a visual attention anchor point, performing multi-level convolution calculation integration on the image to allocate space and generate target coordinates, mapping the polar radius and polar angle to establish an edge network, collecting environmental parameters to construct a compensation factor, performing dynamic bias updating and node determination on the tensor to output a decision result. In the present application, the feature tensor is executed on the local device end to output an anti-interference decision through activation judgment, completely eliminating the data round-trip delay and link packet loss risk caused by remote communication, significantly improving the operation decision reliability of the device in a limited network environment, completely breaking the environmental constraints by updating the network bottom layer parameters and weighting aggregation of the compensation factor, greatly enhancing the self-adaptability and anti-interference protection level in complex and variable industrial sites.
Owner:MIANYANG CITY UNIV

Human-object interaction detection based on three-dimensional position and gaze region prior

The application provides a human-object interaction detection method based on three-dimensional position and gaze area priori. First, a high-quality instance-level visual feature is extracted by DETR, then an accurate three-dimensional position priori is constructed by a depth estimation model, and 27 three-dimensional orientation labels are calculated to finely describe the spatial orientation relationship between the human body and the object, and then the human visual attention priori is obtained by combining a gaze prediction model, finally, these heterogeneous features are uniformly encoded and fused by MLP, so as to combine the multi-modal query vector and the initial visual feature for interaction recognition. The application can not only identify the interaction behavior itself, but also clearly distinguish direct interaction, indirect interaction and no interaction and other subtle scenes, so as to solve the problems of misjudgment and limitation caused by excessive dependence on apparent visual features or single spatial relationship in the prior art, and improve the practicability, robustness and interpretability of the system in an open environment.
Owner:HUNAN NORMAL UNIVERSITY

Method for multilingual learning of language models using rlhf using synthetic eye gaze trajectories

FIELD: computer technology.SUBSTANCE: method for multilingual training of language models using reinforcement learning based on human feedback using synthetic gaze trajectories, comprising the steps of feeding multilingual text to a gaze prediction model based on a multilingual eye movement corpus and a multilingual BERT model to generate a fixation sequence, computing visual attention features including first-pass regression rate, skip rate, first-pass and total fixation counts, and normalized Levenshtein distance, generating a gaze-aware reward model by projecting features into a latent space through a fully connected network, distributing rewards to tokens proportional to fixation probabilities, and optimizing the policy using PPO or GRPO algorithms with a modified advantage that takes into account the Levenshtein distance.EFFECT: multilingual training in thirteen languages without collecting real eye tracking data, achieving the accuracy of reproducing the features of visual attention.6 cl
Owner:AVTONOMNAYA NEKOMMERCHESKAYA ORGANIZATSIYA VYSSHEGO OBRAZOVANIYA UNIV INNOPOLIS

Method and device for generating first-view video based on mask diffusion model and gaze point constraint

The application discloses a first-view video generation method and device based on a mask diffusion model and a gaze point constraint, constructs an end-to-end deep learning framework for the needs of video filling, prediction and visual attention area control in first-view video generation, and realizes diversified first-view video generation conforming to visual logic. The method first divides an input video into a conditional frame set and an unknown frame set, adds random noise to the unknown frame by using a dynamic mask module I to generate a noisy frame, then generates the video by using a full 3D convolutional neural network module II, combines diffusion steps and gaze point trajectories as conditional constraints, guides the network to generate video content conforming to space-time rules, finally uses a gaze point positioning module III to perform saliency prediction, and jointly optimizes the loss of the inverse denoising generation process and the loss of the gaze point probability graph, so that the rationality and diversity of the generated video are further improved. Through the joint optimization of the mask diffusion strategy and the gaze point constraint, the space-time coherence, noise resistance and adaptability to the gaze point trajectory of the generated video can be effectively improved, and the method is especially suitable for complex multi-face expression interaction scenes.
Owner:CHINA UNIV OF MINING & TECH

A network security sensitive event joint extraction method based on a compensation path planning strategy and a hybrid neural network

The application relates to the technical field of natural language processing, in particular to a network security sensitive event joint extraction method based on a compensation path planning strategy and a hybrid neural network, which comprises the following steps: obtaining text data, audio data, picture data and video data in the network security field; preprocessing the text data, training the preprocessed text data, obtaining a word vector, and connecting context semantic information through the word vector; preprocessing the audio data and the video data by using an ECA-ATT neural network, preprocessing the picture data by using an MLP-Mixer neural network; fusing the word vector, the preprocessed audio data, the preprocessed picture data and the preprocessed video data; inputting the fused data into a hybrid neural network built based on a word-graph visual attention mechanism; and outputting network security sensitive events. The application can realize real-time monitoring and early warning of website content and effectively identify bad information.
Owner:DALIAN VOCATIONAL & TECHNICAL COLLEGE (DALIAN OPEN UNIVERSITY)

A joint test method for road traffic safety

ActiveCN121904991BDriver/operatorData set
The application discloses a kind of road traffic safety joint test method, belong to traffic safety technical field.The prior art in traditional road traffic safety monitoring method exists the problem of accident data loss and multi-subject collaborative record limitation;The application comprises the following steps: S1.according to high-risk scene related data, set accident black point, collect accident point data, combine the SUMO vehicle follow model constructed, and the accident black point scene is reproduced;S2.based on the accident black point scene reproduced in step S1, combined with the SUMO vehicle follow model, multi-body joint test is carried out, and the high-precision test results obtained are observed and recorded.The application effectively improves the data accuracy of accident reproduction in traffic safety simulation, can reproduce the continuous trajectory of all traffic subjects, visual attention and driver control details, and can be applied to obtain the precursor of accident and the background traffic flow data set.
Owner:SHENZHEN URBAN TRANSPORT PLANNING CENT CO LTD +1

Geospatial brain-like navigation route planning methods, devices, equipment and storage media

This application proposes a geospatial brain-like navigation route planning method, apparatus, device, and storage medium. The method includes: extracting typical scene features of each image data in a navigation task dataset using a brain-like scene recognition model, wherein the similarity between the feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model and human attention maps is less than a preset similarity threshold; and inputting the typical scene features of each image data into a brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset. This application's embodiment extracts typical scene features of each image data using a brain-like scene recognition model comprising multiple brain-like neurons. Because typical scene features conform to human visual attention, the number of typical scene features is small, and they are all crucial local detail features for navigation decision-making, thus ensuring the correctness of navigation decisions while improving the robustness of the model under various visual interferences.
Owner:BEIJING NORMAL UNIVERSITY

Method and system for evaluating and optimizing visual attention load of interface in main control room of nuclear power plant

The application discloses a nuclear power plant main control room interface visual attention load evaluation and optimization method and system, belongs to the technical field of human-computer interface of nuclear power plant, and the method comprises the following steps: collecting eye movement data of an operator; calculating a heat map entropy, a gaze duration weighted coefficient and a pupil change rate coefficient based on the eye movement data, and generating a comprehensive attention load index; establishing a correlation model of attention entropy and response time, identifying a high-load area and an inefficient information area, and marking a safety-critical visual bottleneck; identifying an operation condition type, and generating an interface optimization suggestion adaptive to the operation condition; reevaluating after applying the optimization suggestion, and forming a closed-loop optimization; and the application accurately quantifies visual load of the operator through a multi-dimensional weighted fusion algorithm, automatically identifies interface design defects based on an entropy-time correlation model, generates adaptive optimization schemes for different operation conditions, and continuously improves through closed-loop feedback, so that the operator's key parameter identification time is shortened, and the operation error rate is reduced.
Owner:NANJING UNIV OF SCI & TECH

An intelligent automobile grading early warning system based on driver risk perception reliability

The application relates to an intelligent automobile hierarchical early warning system based on driver risk perception reliability, which comprises a driver risk area judgment module based on eye movement information, an objective environment anisotropic risk field calculation module, an objective environment anisotropic risk field visual field conversion module, an intelligent automobile risk area judgment module based on driver intention, a human-vehicle risk perception result representation module based on risk perception accumulation effect and decay effect, a driver risk perception reliability quantification module and a hierarchical early warning module based on driver risk perception reliability; the application realizes longitudinal and lateral combined early warning from the driver risk perception level by constructing an anisotropic driving risk field, dividing a driving visual field area, capturing a driver visual attention point and performing space-time convolution operation, effectively solves problems such as inaccurate traffic environment risk description, difficult quantitative description and evaluation of driver risk perception conditions, longitudinal and lateral early warning fragmentation and early warning lag and the like, and improves the performance of the early warning system.
Owner:JILIN UNIVERSITY

An adaptive closed-loop brain-computer interface neurofeedback training system

The application is particularly a self-adaptive closed-loop brain-computer interface neural feedback training system, and relates to the technical fields of brain-computer interface and neural feedback, comprising: a multi-modal signal acquisition module; a multi-modal feature extraction module; a coefficient fusion module; and a feedback training module.In the application, high-precision crystal oscillator clock is adopted to realize time synchronization of electroencephalogram, eye movement and behavior signals, and fusion distortion caused by signal misplacement is completely eliminated.Three types of signals cover cognitive state, visual attention and motion characteristics respectively, forming complementary state evaluation dimensions; and fine quantization algorithms are designed for each modality: electroencephalogram coefficients are screened by extreme value to eliminate instantaneous artifacts, eye movement coefficients calibrate pixel diameter to physical diameter to eliminate imaging interference, and behavior coefficients take into account large motion intensity and high-frequency posture micro-variation.
Owner:HANGZHOU BRAIN MIRACLE INTELLIGENT TECHNOLOGY CO LTD