Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

136 results about "Visual attention" patented technology

Visual language model illusion suppression method based on adaptive dynamic attention intervention

The invention discloses a visual language model illusion suppression method based on self-adaptive dynamic attention intervention. The method is used for reducing the problem that a visual language model generates wrong associated information. The method comprises the steps that text input, visual input and historical response are acquired, and an attention accumulation vector is initialized; in the layer-by-layer calculation process of the language model, dynamically adjusting a non-normalized attention matrix, enhancing the weight of a visual sensitive attention head, and performing visual Token pruning in a deep network to optimize cross-modal information interaction; and finally, generating an output Token based on the adjusted attention mechanism, and carrying out loop iteration until a complete response is generated. According to the method, a method of combining text deviation correction through self-adaptive attention head modification and visual attention convergence-based Token pruning is adopted, so that the performance of the model in a multi-modal task is remarkably improved, and the illusion phenomenon caused by a language modal dominant reasoning process is effectively relieved.
Owner:ZHEJIANG UNIV OF TECH

Thermal runaway prevention and heat dissipation optimization method and system based on capacitor module

The invention provides a thermal runaway prevention and heat dissipation optimization method and system based on a capacitor module, and relates to the technical field of capacitor module thermal management, and the method comprises the steps: extracting the multi-source operation data features of a single capacitor through an improved layered visual attention network, carrying out the data complementation through combining a bidirectional probability diffusion model, and obtaining a multi-source operation data feature of the single capacitor; and discovering a causal relationship between network identification data by using a causal structure, generating a performance analysis report, constructing a spatio-temporal dynamic graph neural network to predict temperature field distribution, performing fault diagnosis and uncertainty quantification in combination with a multi-task learning framework, further constructing a layered early warning mechanism, and performing early warning based on a layered reinforcement learning framework. And a dynamic prediction map and a performance analysis report are combined, a heat dissipation strategy is optimized, cooperative control of the heat dissipation units is realized, an optimal heat dissipation scheme is finally obtained, thermal runaway is effectively prevented, and the safety and reliability of the capacitor module are improved.
Owner:BEIJING RUIHE DEBAO THERMAL TECH CO LTD

AI tax auditing method based on multi-modal data and attention mechanism

The invention discloses an AI tax auditing method based on multi-modal data and an attention mechanism, and relates to the technical field of tax auditing, and the method comprises the steps: carrying out the preprocessing and input of multi-modal tax data, and carrying out the unified representation of the data; identifying risks and capturing attention weights based on an artificial intelligence auditing model of an attention mechanism; mapping the captured attention weight back to a corresponding data unit of the multi-mode tax auditing original data, and generating a visual attention auditing evidence chain; and analyzing the attention weight generated in the analysis process of the artificial intelligence auditing model to identify the attention distribution abnormity existing in the process of processing the specific data area or the evidence chain link, and generating auditing blind spot early warning information based on the identified attention distribution abnormity. According to the method, the interpretability and credibility of AI auditing can be improved, the auditing efficiency and quality are improved, and a novel man-machine cooperation mode is constructed.
Owner:BEIJING HUACHEN HONGYI CONSULTING CO LTD

Immersive audio-video follow-up adjustment method and system

The invention is applicable to the field of intelligent audio adjustment, and provides an immersive audio-video follow-up adjustment method and system, and the method comprises the steps: constructing a multi-dimensional perception system, and collecting multi-source information in real time; carrying out fusion processing on the collected multi-source information based on a deep neural network model, mining a dynamic mapping relation between a user state and the video content through space-time correlation analysis, and identifying a user interaction intention and an emotional tone and a space scene attribute of the video content; according to a fusion processing result, calling a dynamic parameter adjustment engine, and generating an audio parameter adjustment scheme in real time; a user experience feedback closed loop is constructed, a visual attention area of a user is collected through eye movement tracking equipment, and personalized adjustment preference parameters are generated; according to the method and the device, the audio effect is accurately matched with the user state and the audio and video content, the naturalness and the adaptability of immersive experience are remarkably improved, universality and individual differences are considered, and the audio experience which is more suitable for scenes and needs of the user is brought to the user.
Owner:SHENZHEN ZIDOO TECH CO LTD

Method and device for generating first view angle video based on mask diffusion model and fixation point constraint

The invention discloses a first view angle video generation method and device based on a mask diffusion model and fixation point constraint, and the method comprises the steps: constructing an end-to-end deep learning frame for the demands of video filling, prediction and visual attention region control in the generation of a first view angle video; the diversified first-view-angle video generation conforming to the visual logic is realized. The method comprises the following steps: firstly, dividing an input video into a condition frame set and an unknown frame set, and adding random noise to an unknown frame by using a dynamic mask module I to generate a noisized frame; performing reverse de-noising generation on the video through a full 3D convolutional neural network module II, and guiding the network to generate video content conforming to a space-time law by taking a diffusion step and a fixation point track as conditional constraints; and finally, saliency prediction is carried out by using a fixation point positioning module III. By jointly optimizing the loss of the reverse denoising generation process and the loss of the fixation point probability graph, the reasonability and diversity of the generated video are further improved. According to the method, through joint optimization of the mask diffusion strategy and fixation point constraint, the space-time continuity, the anti-noise capability and the adaptability to the fixation point trajectory of the generated video can be effectively improved, and the method is particularly suitable for a complex multi-face expression interaction scene.
Owner:CHINA UNIV OF MINING & TECH

Multi-target license plate recognition method based on visual attention mechanism

The invention discloses a multi-target license plate recognition method based on a visual attention mechanism, and belongs to the technical field of image processing and mode recognition, and the method comprises the steps: obtaining an input image, carrying out the multi-scale feature extraction, and generating a multi-scale feature map; performing spatial saliency calculation and normalization processing on the multi-scale feature map to generate an attention map; carrying out region division on the input image, and adopting differentiated image preprocessing strategies for different regions to generate a preprocessed image; based on the preprocessed image and the attention map, multi-target detection is carried out through a target detection network, candidate area screening is carried out, and a candidate license plate area is generated; and performing binarization processing, character segmentation and feature recognition on the candidate license plate region, and performing verification in combination with context information to generate a license plate recognition result. The attention map is generated by adopting a visual attention mechanism, differentiated image preprocessing and multi-target detection are guided according to the attention map, and multi-target license plate recognition can be completed in a complex scene.
Owner:SHENZHEN BOTE TECH CO LTD

Eye movement tracking-based closed cabin human eye perception evaluation method

A closed cabin human eye perception evaluation method based on eye movement tracking belongs to the technical field of computer vision, and comprises the following steps: collecting eye movement data of simulation personnel in a specific simulation environment, analyzing the eye movement data, and analyzing and calculating the reaction time and the reaction accuracy of the simulation personnel to a moving target; the visual attention distribution and information processing capability of simulation personnel in a specific task situation can be revealed, so that the decision-making efficiency and the response capability of the simulation personnel in a complex dynamic environment are reflected; the continuity of the spliced screen is calculated, and the fluency degree of the connection effect at the spliced position of the screen is reflected; the eye movement focusing degree is analyzed, the attention level of simulation personnel on a target area can be quantified, whether visual attention is concentrated in a key information area or not is helped to be judged, multiple indexes are input into a regression model, and a comprehensive human eye perception score is obtained. The method has the characteristic of human eye perception evaluation for multi-screen splicing, and a comprehensive, accurate and easy-to-operate evaluation tool is provided.
Owner:JILIN UNIVERSITY +1

Power consumption optimization control system for high-refresh-rate driving chip

The invention relates to the technical field of display system driving, and discloses a high-refresh-rate driving chip power consumption optimization control system which comprises an image content complexity modeling module for extracting region sparsity features, a visual attention prediction module for generating an attention heat map, and a display module for displaying the attention heat map. The combined graph modeling and optimization module fuses the information to construct a graph structure and predicts a refresh level, the refresh strategy generation module generates an optimal refresh strategy based on the level, the soft and hard translation execution module converts the strategy into a driving instruction and issues and executes the driving instruction, and the feedback scheduling and self-adaptive correction module dynamically adjusts the graph structure according to chip feedback. The method is used for optimizing next frame refresh control. According to the method, through region division and graph structure modeling, image region refreshing demand modeling is carried out in combination with an attention mechanism and loss feedback, differential refreshing is realized, invalid refreshing of a low-change region is effectively avoided, the content perception ability of a refreshing strategy is improved, and thus the overall power consumption is reduced.
Owner:SHENZHEN HISTONE OPTOELECTRONICS TECH CO LTD

WebM protocol low-delay video and audio translation and subtitle optimization method and system

The invention discloses a WebM protocol low-delay video and audio translation and subtitle optimization method and system, and belongs to the technical field of data processing, and the method comprises the steps: analyzing a WebM audio and video stream; establishing a video track-audio track association list and constructing a dynamic causal map; generating an initial translation text based on the audio frame sequence, the video frame sequence and the dynamic causal atlas, extracting lip movement and action semantic data of video frames through a reverse generation model to generate a completion translation text, and embedding SimpleBlock elements; defining a subtitle core gene, a non-core gene and a position gene according to the code rate data, the complete translation text and a user visual attention thermodynamic diagram, dynamically cutting the non-core gene based on code rate fluctuation, and adjusting the position gene in combination with the thermodynamic diagram to generate adaptive subtitle data; by synchronously playing and collecting feedback data, the edge weight and the subtitle position gene of the dynamic causal atlas are optimized. According to the invention, the robustness of audio and video translation and the self-adaptability of subtitle display are realized.
Owner:JIANGSU ZHIMENG INTELLIGENT TECH CO LTD

AR scenic area guide method and system based on positioning

The invention relates to the technical field of AR preview, in particular to a positioning-based AR scenic area guide method and system, and the method comprises the steps: obtaining the real-time three-dimensional position of a tourist through a Google ARCore fusion positioning engine based on a preset scenic area three-dimensional model and a space reference coordinate system; carrying out visual anchor point dynamic matching according to the real-time three-dimensional position, and identifying a current scene area corresponding to the current position of the tourist; performing virtual content scheduling and space rendering according to the current scene area, and generating space rendering content; and obtaining a visual attention behavior of a tourist on the space rendering content, screening out tourist-interested guide content according to the visual attention behavior, and generating a visual guide path according to the tourist-interested guide content. According to the method, by adjusting the loading and scheduling strategy of the virtual content, the requirements of personalized content display and smooth space guidance are met, and the experience feeling of tourists is enhanced.
Owner:湖北云雷信息技术有限公司

Air-ground pedestrian re-identification method combining multi-frame information and prompt learning

The invention discloses an air-ground pedestrian re-identification method combining multi-frame information and prompt learning, and the method comprises the following steps: inputting a video sequence into a trained visual encoder model, mapping the same pedestrian at different visual angles into a consistent feature space, and achieving the cross-visual-angle pedestrian re-identification and tracking; according to the visual encoder model, a random rotation transformation strategy of structure perception is introduced in the input embedding stage of a visual encoder backbone network, pedestrian vector features of each frame are rotated, visual angle rotation distortion generated by aerial shooting is simulated, and an enhanced sequence is generated. Extracting global features of the enhanced sequence and the unenhanced sequence through a backbone network; inputting the global feature into an inter-frame information attention module for time dimension attention calculation to obtain an average feature of multi-frame fusion; and then inputting the multi-frame fusion average features into a prompting and guiding visual attention module to generate a text prompt so as to guide model discriminative character features.
Owner:SOUTH CHINA UNIV OF TECH

Intelligent structured medical record generation method and system based on multi-modal doctor-patient interaction

The invention provides an intelligent structured medical record generation method and system based on multi-modal doctor-patient interaction, and the method comprises the steps: collecting dialogue voice in real time, transcribing the dialogue voice into a text sequence, and recognizing a visual attention entity through monitoring the operation of a mouse in an electronic medical record system; a logic demonstration track is constructed based on a historical visual attention entity and a text sequence, and an implicit reward function is reversely derived from the logic demonstration track by using a reverse reinforcement learning algorithm. The function is used for calculating the action return value of each combination of the visual attention entity and the dialogue text, and the highest value combination is selected as an optimal alignment strategy to determine the time sequence causal relationship. And mapping the entity and the text to a medical knowledge graph, extracting a shortest semantic path as an implicit clinical reasoning chain, and generating a structured electronic medical record. According to the method, diagnosis and treatment decision logic is deduced from multi-modal behaviors of doctors through reverse reinforcement learning, and the technical problem that internal causal association between visual attention focuses of doctors and oral contents cannot be established in a traditional method is solved.
Owner:WUHAN SHENGBOHUI INFORMATION TECH CO LTD +1

Real-time eye identification and visual detection process control method

The invention relates to a real-time eye identification and visual detection process control method, which comprises the following steps of: applying bitamporal edge disturbance stimulation in a visual field, inducing spontaneous dominant reaction of a dominant eye of a user, collecting perceptual response difference of a non-gazing area to the disturbance stimulation, and judging the current dominant state of the dominant eye; on the basis of micro-differential pressure distribution of different quadrant areas of eye sockets, speculating a visual attention gravity center in real time, and executing spatial geometric deformation on an ROI (Region of Interest) in a visual detection process, so that a detection area dynamically adapts to a gaze offset trend of a user; the method comprises the following steps: respectively constructing two physically independent visual detection links of a feature target strong discrimination process driven by a main visual eye and an environment contour weak confirmation process driven by an auxiliary visual eye, introducing a logic thermal impedance adjustment mechanism, and setting a buffer time window to absorb processing fluctuation caused by state switching; and dynamically matching a visual detection flow path by constructing a mapping relationship between the eye dominant frequency and the user intention type on the basis of statistical information of the dominant eye identification frequency.
Owner:AIR FORCE MEDICAL CENT PLA

Video image compression algorithm based on H.266 / VVC standard

The invention relates to the technical field of video image compression, in particular to a video image compression algorithm based on an H.266 / VVC standard, which comprises the steps of generating a visual attention thermodynamic diagram through time-space domain saliency detection, adjusting a coding tree unit QP value in a partitioned manner to realize adaptive quantization, and designing an asymmetric quantization matrix to optimize a transformation quantization effect. And the generative adversarial network is used to enhance the quality of the reconstructed frame. According to the method, the blocking effect of a complex texture region can be effectively inhibited, the quantization distortion of a human eye sensitive region is reduced, the edge preserving capability is improved, the subjective visual quality and compression efficiency of the video are remarkably improved, and the high-definition video transmission requirement is met.
Owner:HANGZHOU ZHILING TECHNOLOGY CO LTD

Multi-modal large model question and answer method, device and equipment based on attention entropy and medium

The invention discloses a multi-modal large model question and answer method, device and equipment based on attention entropy and a medium, and relates to the field of artificial intelligence and machine learning. Comprising the following steps: inputting a target image and a problem text into a multi-modal large model to obtain an image marking sequence and a text lexical element sequence; connecting an rth pruning layer in front of an rth decoding layer of the decoder, calculating an attention matrix from rth vision to text and an attention matrix from rth text to vision according to an rth image mark sequence and an rth text lexical element sequence, and determining an ith information density weight corresponding to an ith image mark in the rth image mark sequence; based on the information density weight corresponding to each image mark in the rth image mark sequence, screening out an rth reserved image mark sequence; and inputting the rth reserved image mark sequence and the rth text lexical element sequence into the rth decoding layer to obtain an answer text output by the large language model, so as to effectively balance the calculation efficiency and the semantic integrity under the condition of not needing additional training.
Owner:TSINGHUA UNIVERSITY

Remote sensing image self-learning segmentation method

The invention provides a remote sensing image self-learning segmentation method, belongs to the technical field of image processing, and aims to realize automatic, accurate, sufficient and reliable high-resolution remote sensing image segmentation. Comprising the following steps: responding to an input remote sensing image, and performing spectrum correction operation on the remote sensing image to obtain a first color image; carrying out super-pixel division on the first color image by utilizing an edge recognition and super-pixel division method for simulating a non-classical receptive field to obtain super-pixel plaques under different scales; performing clustering operation on the superpixel plaques to obtain a clustering result; training a multi-feature deformation lightweight neural network based on a visual attention mechanism by using a clustering result, and generating a super-pixel region recognition model; and performing remote sensing image segmentation by using the super-pixel region recognition model. According to the method, the self-learning ability of human eye visual perception can be simulated, then image segmentation is carried out through self-adaptive analysis, self-learning identification and self-checking correction, and a segmentation result is rapidly and accurately obtained.
Owner:NORTHWEST ENGINEERING CORPORATION LIMITED

A basketball video highlight automatic generation method and system based on multi-modal feature fusion

The application relates to a basketball video highlight automatic generation method and system based on multi-modal feature fusion, and belongs to the technical field of computer vision and multimedia intelligent analysis. The method comprises the following steps: S1, extracting and aligning a visual feature sequence and an audio feature sequence; S2, using bidirectional LSTM to perform time sequence coding on the audio-visual feature, constructing a cross-modal attention mechanism taking audio as a query and vision as a key value, and obtaining multi-modal fusion features; S3, inputting the fusion features into a 1D CNN to output a frame-by-frame highlight event probability sequence; S4, performing peak value searching based on an adaptive threshold of the probability sequence to obtain a coarse-grained timestamp; S5, taking the coarse-grained timestamp as a center to intercept local features, inputting the local features into bidirectional LSTM to predict starting and ending offsets, and obtaining accurate start and end time boundaries; and S6, clipping and splicing to generate a highlight video according to the accurate boundaries. The application realizes audio-visual semantic depth alignment through an audio-guided visual attention mechanism, and improves boundary positioning accuracy through a coarse-precision two-stage positioning architecture.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Self-adaptive closed-loop brain-computer interface neural feedback training system

The invention particularly relates to a self-adaptive closed-loop brain-computer interface neural feedback training system, and relates to the technical field of brain-computer interfaces and neural feedback. A multi-modal feature extraction module; a coefficient fusion module; and a feedback training module. In the invention, a high-precision crystal oscillator clock is adopted to realize time synchronization of electroencephalogram, eye movement and behavior signals, fusion distortion caused by signal dislocation is thoroughly eliminated, and the three types of signals respectively cover cognitive states, visual attention and motion characteristics to form complementary state evaluation dimensions; a refined quantization algorithm is designed for each mode, wherein instantaneous artifacts are eliminated through extreme value screening of the electroencephalogram coefficient, the pixel diameter is calibrated into the physical diameter through the eye movement coefficient so as to eliminate imaging interference, and the large-amplitude movement intensity and high-frequency posture micro change are considered in the behavior coefficient.
Owner:HANGZHOU BRAIN MIRACLE INTELLIGENT TECHNOLOGY CO LTD

A complex scene instruction expression understanding method based on cross-modal eye movement attention perception

The application discloses a complex scene instruction expression understanding method based on cross-modal eye movement attention perception, and belongs to the fields of computer vision, machine learning and multi-modal understanding. The application simulates the eye visual attention perception area and the transfer process by designing a dynamic deformable attention mechanism of language perception, using an eye gaze spectrum as supervision information, adaptively capturing the corresponding visual area according to language features, and designing an eye movement spectrum driven Transformer decoder to infer the target area position of the language instruction by gradually fusing the visual feature representation, thereby effectively improving the complex scene instruction expression understanding precision.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A mixed reality data processing and interaction response method, device and system

PendingCN122336209APersonalizationMixed reality
This invention discloses a mixed reality data processing and interactive response method, apparatus, and system. The method includes: acquiring corresponding spatial feature points in VR and AR environments; optimizing the solution of rigid body transformation matrices to achieve sub-millimeter-level spatial alignment; continuously monitoring alignment errors during virtual-real fusion rendering, triggering a repositioning process when errors exceed a threshold; evaluating user operations through multimodal data fusion and generating real-time AR correction guidance; and dynamically adjusting rendering parameters and resource preloading based on visual attention focus to ensure end-to-end latency and interactive feedback latency are both below set thresholds. The apparatus includes modules for feature point acquisition, spatial alignment calculation, error monitoring and repositioning control, multimodal data interface, and real-time rendering control. The system includes a server and a user interaction terminal, supporting intelligent training and adaptive interactive optimization for virtual-real fusion. This application addresses problems such as low virtual-real spatial alignment accuracy, high interactive latency, and insufficient personalized adaptation.
Owner:CHINA LIFE INSURANCE CO LTD

Systems and methods for training machine-learned visual attention models

Systems and methods of the present disclosure are directed to a method for training a machine-learned visual attention model. The method can include obtaining image data that depicts a head of a person and an additional entity. The method can include processing the image data with an encoder portion of the visual attention model to obtain latent head and entity encodings. The method can include processing the latent encodings with the visual attention model to obtain a visual attention value and processing the latent encodings with a machine-learned visual location model to obtain a visual location estimation. The method can include training the models by evaluating a loss function that evaluates differences between the visual location estimation and a pseudo visual location label derived from the image data and between the visual attention value and a ground truth visual attention label.
Owner:GOOGLE LLC

Trained highlighting for human checking of marketing materials

PCT designated stageWO2026024183A1Office automationEngineeringKnowledge management
Method of training an automatic highlighting system for supporting a human checker of automatically generated visual marketing materials, comprising: presenting to the human checker a piece of automatically generated visual marketing material to be checked; obtaining from the human checker at least one label representing a result of the checking of the presented piece of automatically generated visual marketing material; for one or more areas within the presented piece of automatically generated visual marketing material, obtaining at least one human attention indication regarding the human checker's visual attention to the respective area during the checking; and training the automatic highlighting system using at least: the presented piece of automatically generated visual marketing material, the obtained at least one label, and the at least one human attention indication obtained for at least one of the one or more areas within the presented piece of automatically generated visual marketing material.
Owner:HEINEKEN SUPPLY CHAIN BV

An unmanned aerial vehicle form design method and device based on emotional driving and a medium

The application belongs to the field of product generative design, and discloses a UAV form design method and device based on emotion driving and a medium. The method comprises the following steps: collecting user comment data, and extracting and screening core emotional words by using a BERTopic model; obtaining visual attention data of users on different UAV form components through an eye movement experiment, and calculating objective importance weight of each form component by combining an entropy weight-TOPSIS method; introducing the objective importance weight as prior knowledge into a deep learning model, constructing a Transformer-BiLSTM model fused with the prior weight, establishing a mapping relationship between user emotion and key UAV form, and predicting and generating an optimal form combination sketch; generating the sketch into a high-fidelity rendering image by using a stable diffusion model, combining physiological signals collected through eye movement and electroencephalogram multi-modal experiments, objectively verifying and optimizing the generated scheme, and determining an optimal design scheme, so that intelligent conversion from abstract emotional semantics to concrete product form is realized, and product design efficiency is improved.
Owner:NANCHANG UNIV

Method, device, medium and equipment for controlling proportion of AR / VR / MR glasses function interface

The present disclosure relates to a method, device, medium and equipment for controlling the proportion of AR / VR / MR glasses function interface, which comprises the following steps: determining the visual angle range information of the corresponding eye features of a user, generating a preset initial proportion control interface in the corresponding display area according to the visual angle range information, determining the display order of the function icons corresponding to each set function in the line-of-sight range corresponding to the visual angle range information from the initial proportion control interface, determining the first visual attention function icon corresponding to the visual attention point information according to the display order and the function icon, the visual attention point information and the target function icon, determining the second visual attention function icon based on the visual attention point information, generating a display proportion locking instruction to lock the display content and display proportion of the corresponding visual attention function icon in the current preset initial proportion control interface, and generating a target proportion control interface. Thus, the operation efficiency of the multifunctional operation interface is improved.
Owner:GUANGZHOU HAOXINWEI TECH IND CO LTD

Vehicle-mounted navigation layered display method based on driving scene self-adaption

The invention relates to the technical field of intelligent cabins, and discloses a driving scene self-adaption-based vehicle-mounted navigation hierarchical display method, which comprises the following steps of: acquiring vehicle kinetic parameters, an external meteorological environment and road topological data in real time, calculating a driving scene complexity index reflecting the cognitive load of a driver, and determining the radius of a visual attention area according to the driving scene complexity index; constructing a hierarchical rendering model comprising a basic environment layer, a path guiding layer and a key warning layer; dividing the display area into a central focusing area and an edge peripheral area based on the radius of the visual attention area by taking the center of the screen as an original point; and keeping high-fidelity display in the central focusing area, and executing gradient fuzzification processing on the basic environment layer in the edge peripheral area. The visual attention area is dynamically reduced along with the increase of scene complexity, the human eye visual tunneling effect is simulated, the cognitive load of a driver is effectively reduced, and the driving safety is improved.
Owner:深圳市科乐达电子科技有限公司

Intelligent structured medical record generation method and system based on multi-modal doctor-patient interaction

The application provides a kind of intelligent structured medical record generation method and system based on multimodal doctor-patient interaction, which comprises: real-time acquisition of dialogue voice and transcription into text sequence, while recognizing visual attention entity by listening to mouse operation in electronic medical record system.Based on the history of visual attention entity and text sequence, a logical demonstration track is constructed, and an implicit reward function is derived from it using a reverse reinforcement learning algorithm. Use the function to calculate the action reward value of each combination of visual attention entity and dialogue text, select the highest value combination as the optimal alignment strategy to determine the timing causal relationship. Map the entity and text to the medical knowledge graph, extract the shortest semantic path as the implicit clinical reasoning chain, and generate structured electronic medical record. The application deduces the diagnosis and treatment decision logic from the doctor's multimodal behavior through reverse reinforcement learning, solving the technical problem that traditional methods cannot establish the internal causal relationship between the doctor's visual attention focus and spoken content.
Owner:WUHAN SHENGBOHUI INFORMATION TECH CO LTD +1

Eye movement tracking mixed reality electric power training response method and system

The invention provides an eye movement tracking mixed reality electric power training response method and system, and the method comprises the steps: collecting eye movement data with a timestamp of a trainee, and generating a continuous sight vector sequence through coordinate system conversion and data processing; recognizing a gazing event, mapping the sight line of the gazing event to a virtual object in the mixed reality scene, and determining a gazing area; constructing a finite-state machine comprising a plurality of state nodes and a transfer relationship, and based on the gazing area and the equipment simulation state, judging whether a state transfer condition is met or not by adopting the finite-state machine and updating a task state; and outputting a response signal according to the updated task state, and performing synchronous rendering display in the mixed reality interface. According to the method and the device, a task control mode taking an eye movement watching behavior as a core is realized, so that an operation behavior is synchronized with a visual attention process, and natural interaction control without a traditional handle or gesture is realized.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Anti-cant indication system for shooting devices

The present invention relates to an audio-based anti-cant indication system for shooting devices such as rifles, pistols, bows, and crossbows. The system includes an inclination sensor configured to detect angular tilt relative to gravity and a processing module configured to compare the detected tilt against a predefined deadband threshold corresponding to a plumb orientation. When the shooting device deviates from the plumb condition, the processing module activates an audio output module that generates distinct audio signals representing left-tilt and right-tilt conditions, while producing a null or muted output when the device remains within the deadband. The system may further include variable tone pitch or volume, wireless audio transmission to earbuds or headsets, calibration mechanisms, front-rear inclination detection, and idle-state shutoff. The invention provides continuous, hands-free cant feedback without requiring visual attention.
Owner:STURDIVANT CHARLES N

Wide-range remote sensing image rural residential area extraction method and system fusing target detection and visual attention mechanism

The present application belongs to the technical field of residential area extraction, and discloses a large-range remote sensing image rural street-type residential area extraction method and system fusing target detection and visual attention mechanism, which comprises the following steps: generating a residential area candidate area detection frame on a large-range remote sensing image through a YOLOv7 model detection; and performing accurate extraction of the residential area based on the candidate area detection frame by using a PgNet model to obtain a rural residential area on the large-range remote sensing image. The present application combines the candidate mechanism of the target detection technology and the visual attention mechanism method, and roughly removes the vegetation and terrain, which is conducive to the rapid and automatic extraction of the rural residential area and the buildings thereof, and improves the search efficiency. After obtaining the candidate frame, the PgNet efficient extraction mechanism is used, so that the rural street-type residential area can be rapidly and automatically extracted in remote sensing images of different sizes.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Unmanned aerial vehicle / unmanned ship relative navigation method of eagle-eye-vision-imitating strong-vision mechanism

The invention discloses an unmanned aerial vehicle / unmanned ship relative navigation method of an eagle-eye-vision-imitating strong-vision mechanism. The method comprises the following steps: step 1, designing an image recognition feature extraction backbone network; 2, designing an eagle eye visual space attention model; 3, designing an eagle eye visual channel attention model; 4, designing an eagle eye vision double-fovea structure feature pyramid network; 5, constructing an eagle-eye-vision-imitating strong-vision target intelligent detection network, and outputting a detection result; step 6, selecting a target identification result; and 7, calculating the relative distance of the unmanned aerial vehicle / unmanned ship. The method has the advantages that (1) visual attention generated by input stimulus driving from bottom to top is enhanced; 2) target features of different environments and different scales are mined, and the target detection precision is improved; and 3) accurate relative position navigation information is provided for unmanned aerial vehicle / boat water-air cooperative control.
Owner:BEIHANG UNIV