Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

505 results about "Gaze" patented technology

In critical theory, sociology, and psychoanalysis, the philosophic term the gaze (French le regard) describes the act of seeing and the act of being seen. The concept and the social applications of the gaze have been defined and explained by existentialist and phenomenologist philosophers; Jean-Paul Sartre, in Being and Nothingness (1943); Michel Foucault in Discipline and Punish: The Birth of the Prison (1975) developed the concept of the gaze to illustrate the dynamics of socio-political power relations and the social dynamics of society's mechanisms of discipline; and Jacques Derrida, in The Animal that Therefore I Am (More to Come) (1997) elaborated upon the inter-species relations that exist among animals and human beings, which are established by way of the gaze.

Lightweight display interface rendering optimization system

The invention discloses a lightweight display interface rendering optimization system, which relates to the technical field of image processing, and comprises an acquisition module, an interface analysis module, a behavior analysis module and an attention analysis and rendering module, dynamically calculating an effective view radius by utilizing a visual tunneling effect; the interface analysis module performs character string fuzzy matching in combination with the context input by the user, identifies the search intention of the user and improves the weight of a corresponding component; the attention analysis module performs multi-modal weighted fusion on the physiological fixation data and the sparse content saliency thermodynamic diagram; according to the method, by recognizing the semantic intention and the physiological fixation point of the user, on the premise that core visual experience continuity is guaranteed, GPU load and video memory bandwidth occupation are remarkably reduced, and balance of high-performance display and low-power-consumption operation is achieved.
Owner:SHENZHEN ZHILINTAI ELECTRONIC TECH CO LTD

Method and system for automatically adapting teaching atmosphere in immersive teaching environment

The invention belongs to the field of virtual reality teaching application, and provides a teaching atmosphere automatic adaptation method and system in an immersive teaching environment. The method comprises the following steps: acquiring an eye movement image; recognizing a fixation point; carrying out ROI tracking; carrying out ROI boundary fusion; adaptively optimizing the object; adjusting the brightness of the ROI; and watching object interaction. According to the method, the immersion and interactivity of a future classroom can be improved, the use experience of an immersive virtual environment is facilitated, and deep fusion of an intelligent teaching environment and self cognition of a user is promoted.
Owner:HUAZHONG NORMAL UNIV

Deep forgery attribution and detection method based on staring guide CLIP model

The invention relates to the technical field of deep faking attribution and detection methods, in particular to a deep faking attribution and detection method based on a gaze guide CLIP model, and the method specifically comprises the following steps: collecting face images, constructing a deep faking attribution and detection data set, enabling each face image in the data set to have a corresponding faking attribution label and a faking detection label, preprocessing images in the data set, and then dividing into a training set and a test set; constructing a zero-sample deep counterfeit attribution and detection model, inputting the preprocessed data set into the model, and performing model processing to obtain predicted image counterfeit detection category features and predicted image counterfeit attribution category features; and performing optimization training on the deep counterfeit attribution and detection model through the training set to obtain an optimized and trained model, and testing the optimized and trained model by using data in the test set. According to the invention, the source of the deep pseudo image is traced through the deep learning method, and the faked attribution and detection can be carried out more accurately.
Owner:SHANDONG ARTIFICIAL INTELLIGENCE INSTITUTE +1

Nursing scene data enhancement and generation method for intention recognition model training

The invention discloses a nursing scene data enhancement and generation method for intention recognition model training, and belongs to the technical field of artificial intelligence and smart old-age care. The method aims at solving the problems that in the prior art, high-quality nursing scene training data is deficient, and the obtaining cost is high. According to the core technical scheme, the method comprises the steps that firstly, a staring sequence and other context information (such as time, place and physiological signals) of a user are processed through a multi-modal information fusion model, and a structured initial situation vector is generated; secondly, inputting the vector into a scene generation model combined with a nursing knowledge base, and automatically generating a batch of basic nursing scene data with intention labels; key points are that a data enhancement module is introduced, and a plurality of innovative strategies such as situation element disturbance, virtual physiological data synthesis, gaze path variation and virtual scene deduction are adopted to deeply process basic data, so that the diversity and complexity of a data set are greatly enriched; and finally, combining the basic data with the enhanced data to construct a final comprehensive training data set. According to the method, large-scale and high-fidelity training data can be generated in a low-cost and high-efficiency manner, and the accuracy and robustness of the intention recognition model in a real nursing environment are remarkably improved.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Image processing and training method and device, equipment, head-mounted display and medium

The invention discloses an image processing and training method and device, equipment, a head-mounted display and a medium. Acquiring fixation point information, an original image and a fixation point area image; generating a first input tensor based on the gaze point information, generating a second input tensor based on the gaze point region image, and connecting the first input tensor and the second input tensor to generate a first intermediate tensor; taking the first input tensor as a guide tensor, and performing weight distribution of channel weight and space weight on the first intermediate tensor based on a channel attention mechanism and a space attention mechanism to generate a second intermediate tensor; extracting multi-scale fuzzy features of a preset fuzzy image to generate a third intermediate tensor fuzzy image, and generating the third intermediate tensor fuzzy image based on a pre-calibrated point spread function of each sub-region in the fixation point region; connecting the second intermediate tensor and the third intermediate tensor, and inputting a fourth intermediate tensor generated by the connection into an encoder and a decoder to generate a pre-corrected image after confusion spot correction; and generating a to-be-displayed image based on the pre-corrected image and the original image.
Owner:YONGJIANG LAB

Camera focusing for video pass-through systems

The invention provides a camera focusing method and device for video transparent transmission equipment. Gaze information from the gaze tracking subsystem and depth information from the depth tracking system may be utilized to determine a depth to focus. A combination of depth information and gaze information is used.
Owner:APPLE INC

User interaction intention perception method and system based on intelligent glasses

The invention relates to the field of intelligent perception, and mainly relates to a user interaction intention perception method and system based on intelligent glasses, and the method comprises the steps: obtaining the abnormal degree of an eye jump drop point according to the main eye jump and the subsequent correction eye jump times after detecting that the current eye jump drop point of a user enters an interaction region of an interactive object; acquiring a pupil diameter change sequence after the main eye jump, and acquiring a pupil transient response score based on the pupil diameter change sequence; after a certain time from the falling point time of the main eye jump, acquiring a real-time double-variable contour ellipse area based on information of a plurality of latest fixation points, and acquiring a fixation stability score based on the real-time double-variable contour ellipse area; obtaining a real-time gaze residence time according to the falling point time of the main eye jump, and obtaining a residence time evaluation score based on the real-time gaze residence time; fusing the abnormal degree, the pupil instantaneous response score, the stability score and the evaluation score to obtain a dynamic interaction intention score of the user, and determining the interaction intention of the user based on the intention score. The perception accuracy of the user interaction intention is improved.
Owner:刘海俊

Method and device for optimizing picture fluency of vehicle-mounted display screen

The invention relates to the technical field of vehicle-mounted human-computer interaction and graphic rendering, and discloses a method and a device for optimizing picture fluency of a vehicle-mounted display screen. The method comprises the following steps: collecting multi-source data and preprocessing the multi-source data to obtain a denoised sight line position and light intensity data; fusing the data and determining real-time sight focus coordinates of the driver through light correction; if the focus is located in the navigation area, preferentially calculating a rendering thread of the focus, judging the residence time and calibrating the locking precision; the navigation area frame rate is improved according to the locking state, and secondary area transition data are prepared; residual computing resources are obtained, a degradation rendering scheme is determined, and resource balanced distribution is achieved; extracting a pre-rendering frame when a total load peak value is detected, and obtaining seamless transition data; and finally, injecting the pre-rendered frame into an output buffer area and fusing the real-time coordinates to generate a smooth picture frame sequence. According to the method, dynamic resource allocation and smooth picture output can be realized, and the vehicle-mounted display fluency and the driving safety are remarkably improved.
Owner:SHENZHEN JIUYANG INTELLIGENT TECH CO LTD

Method and system for performing spatial foveation based on eye gaze

A method includes determining an eye gaze location of a user and generating a spatial foveation map based on the eye gaze location. The method also includes receiving an image, forming a spatially foveated image using the image and the spatial foveation map, and transmitting the spatially foveated image to a wearable device. The method further includes spatially defoveating the spatially foveated image to produce a spatially defoveated image and displaying the spatially defoveated image.
Owner:MAGIC LEAP INC

Devices, methods, and graphical user interfaces for displaying sets of controls in response to gaze and / or gesture inputs

In some embodiments, a computer system enables a user to invoke display of transport controls (and / or other controls associated with controlling playback of content) using gaze inputs, gesture inputs, or a combination of these. In some embodiments, in response to detecting a first user input, the computer system displays a first set of controls in a reduced-prominence state (e.g., in a manner that is not unduly distracting to the user), and in response to detecting a second user input, the computer system displays a second set of controls in an increased-prominence state (e.g., in a more visually prominent state). The second set of controls optionally includes more controls than the first set of controls.
Owner:APPLE INC

Target abnormal movement early warning method based on eye walking along with hook

The invention discloses a target abnormal movement early-warning method based on eye walking along with a hook, and the method comprises the following steps: collecting an eye movement data stream, constructing a fixation behavior data set, and generating a fixation track sequence; constructing the fixation points as graph nodes, generating graph edges according to a time sequence and spatial proximity, and forming a graph structure sequence; inputting to the improved ST-GCN model, and outputting a target prediction vector; identifying offset candidate segments; performing trajectory morphological analysis on the offset candidate segments, and judging whether formation conditions of a trajectory loopback structure are met or not; if not, calculating an access frequency domain; if the access frequency is greater than a preset frequency threshold, outputting a low early warning signal; and screening the high-weight fixation segment based on the offset candidate segment, carrying out similarity matching, and outputting formal early warning information according to a matching result. According to the invention, multi-level accurate early warning of the abnormal motion state of the target is realized, and the method has the advantages of clear structure, strong real-time performance, good adaptability and the like.
Owner:BEIJING GUOXINZHIKE TECH DEV CO LTD

Eye tracking and gaze estimation using off-axis camera

Techniques related to the computation of gaze vectors of users of wearable devices are disclosed. A neural network may be trained through first and second training steps. The neural network may include a set of feature encoding layers and a plurality of sets of task-specific layers that each operate on an output of the set of feature encoding layers. During the first training step, a first image of a first eye may be provided to the neural network, eye segmentation data may be generated using the neural network, and the set of feature encoding layers may be trained. During the second training step, a second image of a second eye may be provided to the neural network, network output data may be generated using the neural network, and the plurality of sets of task-specific layers may be trained.
Owner:MAGIC LEAP INC

Pilot ability assessment method based on dynamic time warping and hierarchical clustering

The invention belongs to the technical field of pilot ability assessment, and particularly discloses a pilot ability assessment method based on dynamic time warping and hierarchical clustering, which comprises the following steps of: acquiring eye movement data of a tested pilot in a flight simulation task process, preprocessing the eye movement data, extracting behavior indexes of each stage of a flight task, and calculating a pilot ability assessment result; generating a fixation area number sequence according to the fixation point position; in the selection target evaluation stage, a dynamic time warping algorithm is used for carrying out nonlinear alignment on the gaze sequences of all the pilots, and an eye movement difference degree matrix between the pilots is generated; based on the difference degree matrix, adopting a hierarchical clustering algorithm to group the pilots; and outputting an ability evaluation result of the pilot according to the distribution of the pilot in the difference degree matrix and the deviation information of the pilot and the teacher watching sequence. According to the method, structured comparison and capability grade evaluation of complex cognitive behaviors can be realized, and an evaluation result has relatively high objectivity and interpretability.
Owner:NAVAL AVIATION UNIV

Multi-modal digital art content generation and personalized recommendation system and method

According to the multi-modal digital art content generation and personalized recommendation system and method, a sequential dynamic knowledge graph is constructed, a pattern, color and narrative structure triple is extracted from an art database, a natural language processing technology is adopted to mine implicit relations, and the artistic content is recommended to the personalized recommendation system and method. Multi-modal feature fusion is realized by utilizing visual and text feature extraction and a gating attention mechanism, a style template is automatically loaded according to equipment performance, rendering parameters are dynamically adjusted, and a recommendation strategy is optimized through facial expression analysis, gaze point distribution calculation and emotional state inference. Rendering detail levels are adjusted based on data acquired by a touch screen sensor and a motion sensor, design contents conforming to culture specifications are generated by combining a conditional GAN correction network, real-time updating of culture symbols and related rules is supported, and the system has strong intelligent processing capability and is suitable for popularization and application. The efficient generation and personalized recommendation requirements of the digital art content in multiple scenes can be met, and the user experience and the interaction effect are remarkably improved.
Owner:南昌理工学院

Multi-task fixation point estimation method

The invention discloses a multi-task fixation point estimation method, which relates to the technical field of computer vision and deep learning, and comprises the following steps of: firstly, constructing a data set for providing associated data from tasks to labels and supporting multi-task cooperative training, and then constructing an MTHGaze model, through a semi-structured framework, deep learning, a geometric constraint fitting model and key tasks of cooperative processing are fused, a loss function is designed and trained, defects of each task are solved through targeted loss function design, and a scientific training process enables the model to be iteratively optimized. Finally, the targets of improving the estimation precision of the fixation point, ensuring the physical rationality and enhancing the environmental robustness are achieved; and finally, performing fixation point mapping. According to the estimation method, through data set construction, data consistency can be guaranteed, training noise can be reduced, a real reference can be provided for loss function calculation, error back propagation can be realized, and Euclidean distance errors can be minimized through L2 norms of real fixation point coordinates and prediction coordinates.
Owner:CHINA JILIANG UNIV

Real-time eye identification and visual detection process control method

The invention relates to a real-time eye identification and visual detection process control method, which comprises the following steps of: applying bitamporal edge disturbance stimulation in a visual field, inducing spontaneous dominant reaction of a dominant eye of a user, collecting perceptual response difference of a non-gazing area to the disturbance stimulation, and judging the current dominant state of the dominant eye; on the basis of micro-differential pressure distribution of different quadrant areas of eye sockets, speculating a visual attention gravity center in real time, and executing spatial geometric deformation on an ROI (Region of Interest) in a visual detection process, so that a detection area dynamically adapts to a gaze offset trend of a user; the method comprises the following steps: respectively constructing two physically independent visual detection links of a feature target strong discrimination process driven by a main visual eye and an environment contour weak confirmation process driven by an auxiliary visual eye, introducing a logic thermal impedance adjustment mechanism, and setting a buffer time window to absorb processing fluctuation caused by state switching; and dynamically matching a visual detection flow path by constructing a mapping relationship between the eye dominant frequency and the user intention type on the basis of statistical information of the dominant eye identification frequency.
Owner:AIR FORCE MEDICAL CENT PLA

Instant intention unlocking system based on facial recognition and voice interaction

The invention discloses an instant intention unlocking system based on facial recognition and voice interaction, and aims at solving the problems that an existing voice assistant depends on wake-up words, interaction is low in efficiency, scene perception is weak, and energy consumption is high. The system adopts a three-layer architecture including a hardware layer (a camera, a microphone, a processor and a storage module), a software layer (a voice recognition engine, a face recognition algorithm, a scene perception module and an authority control module) and a logic layer (an intention fusion judgment module and a task routing and execution module). 'gaze + voice 'is taken as core triggering, and wakeup words are avoided; face recognition is used for confirming identity and watching state, voice recognition is used for analyzing instructions, and an application scene is fused to generate an instant intention triggering model; authority hierarchical management and control are realized, and energy consumption optimization, auxiliary triggering, cross-terminal synchronization and emergency help modules are arranged. Interaction efficiency can be improved, safety and convenience are balanced, energy consumption is reduced, and the method is adaptive to multiple terminals and multiple scenes.
Owner:冯东坡

Intelligent learning planning system and method based on multi-dimensional learning data driving

The invention relates to the technical field of intelligent learning planning, and discloses an intelligent learning planning system based on multi-dimensional learning data driving. Comprising a main control module, a multi-modal data acquisition module, a data quality enhancement module, a multi-modal feature fusion module, a dynamic strategy planning module, a man-machine intention alignment module, a knowledge topology management module, an evaluation and federal optimization module and an auxiliary visualization module, and the main control module is used for coordinating operation of each module and data scheduling; the multi-modal data acquisition module acquires HRV, EDA and other physiological signals through a wearable device, records a fixation track and an alpha wave proportion through an eye tracker / electroencephalograph, and receives a user instruction in real time through a voice interaction device. According to the method, the dynamic adaptability is higher, full-process dynamic adjustment of stage division-parameter fine adjustment-load balance is achieved through hierarchical reinforcement learning and NSGA-III optimization, and a learning path can be flexibly adapted according to the real-time state (such as sudden fatigue and rapid knowledge point mastering) of a user.
Owner:SHENWEI INTELLIGENT TECHNOLOGY (KUNMING) CO LTD

Large-screen business consultation interaction method based on multi-mode identification

The invention discloses a large-screen business consultation interaction method based on multi-modal identification, and relates to the technical field of man-machine interaction, and the method comprises the steps: collecting multi-modal data, carrying out the time calibration and space registration, and generating a multi-modal input data set; semantic interaction features are extracted from the multi-modal input data set, a cross-modal feature fusion method is used for semantic alignment, and a user intention vector is generated; performing large-screen interface semantic analysis on the user intention vector, identifying function partitions, calculating similarity, and generating a business consultation identifier; retrieving a business knowledge network according to the business consultation identifier to obtain a business knowledge node set, and performing semantic association analysis and content integration to generate a business consultation result; and performing natural language generation and visual content mapping processing on the business consultation result to generate business consultation display content, and collecting voice, gestures and gaze feedback of the user to form consultation interaction feedback content. The interactive response accuracy and the intelligent level are improved.
Owner:YLZ INFORMATION TECHNOLOGY CO LTD

Intelligent cabin control method and system

The invention provides an intelligent cabin control method and system, and the method comprises the steps: obtaining the original data of a sensor through a TOF module, and extracting the point cloud data of a target region of a target user based on the original data of the sensor; acquiring hand point cloud data and face point cloud data of the target user according to the target area point cloud data, acquiring hand joint coordinates based on the hand point cloud data, and acquiring facial expressions and fixation points based on the face point cloud data; and further controlling corresponding cabin functions based on the hand joint coordinates, the facial expressions and the fixation points. According to the method and the device, the user can control the functions of the intelligent cabin by waving the fingers or sliding the corresponding amplitude by gestures, a touch screen or a physical key is not needed, and the driving safety is guaranteed; meanwhile, related information is obtained based on the point cloud data, the recognition accuracy can be guaranteed, the computing power consumption is small, and the driving safety is further guaranteed.
Owner:HUIZHOU DESAY SV AUTOMOTIVE

Eye movement trajectory analysis and classification method and system based on three-dimensional fixation estimation

The invention discloses an eye movement track analysis and classification method and system based on three-dimensional fixation estimation, and relates to the field of computer vision, and the method comprises the specific steps: obtaining a face video, estimating a head posture based on the face video, and extracting a plurality of face images; inputting the facial image into a multi-task learning model, extracting facial features and eye features in parallel, processing the fused features, and outputting an eye prediction state and a three-dimensional fixation vector estimation value; converting the three-dimensional fixation vector estimation value into a fixation point sequence under a screen coordinate system based on a relative pose relationship between the camera and the screen; clustering the fixation point sequence to obtain multi-dimensional eye movement features; and inputting the head posture and the multi-dimensional eye movement features into a fusion de-noising variational auto-encoder to obtain a classification result. According to the method, the text reading task is combined with the space-time distance function, and the extracted multi-dimensional eye movement features such as fixation, eye jump and review can accurately reflect the reading process of the subject, so that the classification accuracy is improved.
Owner:SICHUAN UNIV

Capturing and processing content for in-vehicle and extended display, curated journaling, and real-time assistance during vehicular travel

Methods and systems are described for imaging and content generation. During a road trip, a driver or passenger observes an object or scene from a vehicle. From a perspective of the driver or passenger, the object or scene flows out of view in a relatively short amount of time. Images and video of the object or scene are captured with cameras. A field of view is extended beyond what the driver or passenger can easily observe. The capture is configurable to focus on particular types of objects or scenes. Recordings of billboards and road signs are easily displayed after passing the object. Object identification, gaze determination, and interest determination are provided. Curated content is generated. Applications to extended reality environments are provided. Artificial intelligence systems, including neural networks, and models are utilized to improve the imaging and content generation. Related apparatuses, devices, techniques, and articles are also described.
Owner:ADEIA GUIDES INC

A gaze point prediction method and system for VR large space immersive tour

The application belongs to the field of artificial intelligence, computer vision and computer graphics, and discloses a gaze point prediction method and system for VR large space immersive tour, which comprises the following steps: acquiring an eye image transmitted by a near-eye camera; constructing a gaze point prediction network model; inputting the eye image into the gaze point prediction network model to output a predicted line-of-sight direction. The application provides a method for accurately predicting a gaze point in real time, and through the method of eliminating 80% irrelevant pixels in the input image and the method of outputting a prediction value by a multi-level neural network, the calculation complexity is reduced in multiple dimensions, thereby reducing the system delay and improving the rendering quality.
Owner:北京渲光科技有限公司

Data evaluation method and device based on artificial intelligence, computer equipment and medium

The invention relates to a data evaluation method and device based on artificial intelligence, computer equipment and a medium, and is suitable for smart televisions, laptops, projectors, tablet personal computers and the like, and the method comprises the following steps: collecting video streams related to a user, and obtaining environment illumination intensity; performing face key point detection on the video stream to obtain key point information; performing blink detection and frequency calculation based on the key point information to obtain a blink frequency; performing head posture estimation on the key point information, and calculating continuous fixation duration based on a posture estimation result; generating a watching distance of the user to the screen of the target equipment; integrating the environment illumination intensity, the continuous fixation duration, the blinking frequency and the watching distance to obtain index data; evaluating the face image of the user, and obtaining a target index threshold value based on the target age; and performing data evaluation on the index data based on the target index threshold to generate a health evaluation result. According to the invention, the accuracy of health assessment processing of eye using behaviors is improved.
Owner:彩迅工业(中山)有限公司

Remote eye gaze cursor control techniques

A system adapted to interact with a user by monitoring activity of a user's eye gaze, the system comprising: a user web camera configured to determine a location at which the user's eyes gaze on a user display configured to display digital content to the user; an operator display configured to view the live interaction of the eye gaze of the user from a movement of the eye gaze controlled cursor projected in the operator display; a processor having a gaze estimation module for receiving eye movement and head movement signals from the user web camera as calibration data for estimating eye gaze, the processor configured to: generate an interactive stimulus to the user display; stimulating an interaction with the interaction via a user's gaze-based interaction or blink-based interaction; and analyzing and determining a user focus based on eye tracking of the interaction from the interaction stimulus.
Owner:SQUID EYES PTE LTD

Gaze determination using one or more neural networks

Apparatuses, systems, and techniques are presented to predict gaze of an observer. In at least one embodiment, a network is trained to predict a gaze of one or more users based, at least in part, on one or more gazes corresponding to objects not always visible to the one or more users.
Owner:NVIDIA CORP

Electronic equipment and method for disabling functionality

An electronic equipment (100) communicatively connected to a gaze-tracking device (120), an input device (110) and a display device (130), the electronic equipment (100) comprising processing circuitry (404), wherein the processing circuitry (404) is configured to obtain, from the input device (110), pointer information comprising pointing area information pertaining to at least one pointing area (142) comprised in a User Interface, UI, (140) displayed by the display device (130), obtain, from the gaze-tracking device (120), gaze information comprising at least one gaze area information pertaining to at least one gaze area (146) comprised in the UI (140), and in response to the processing circuitry (404) detecting that the pointer information meets a pointer condition associated with a target area (144) and that the gaze information meets a first gaze condition associated with the target area (144): disable a functionality associated with the target area (144). The pointer condition comprises at least part of the pointer information matching a pointer pattern and the first gaze condition comprises at least part of the gaze condition matching a first gaze pattern. Together, these two patterns indicate that the user intends to disable the functionality of the target area.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Human eye fixation point prediction method based on deep convolutional network and frequency domain feature enhancement

The invention discloses a human eye fixation point prediction method based on a deep convolutional network and frequency domain feature enhancement, and the method comprises the steps: employing the deep convolutional network (namely a residual network ResNet) as a backbone network to form an encoder branch, and employing five layers of coding blocks to extract the spatial features of five layers of an input image; the five layers of spatial features extracted by the encoder are respectively sent to a frequency domain feature enhancement module for frequency domain feature enhancement based on discrete cosine transform (DCT) so as to enhance the features of each layer of human eye fixation area and reduce interference features; sending the frequency-domain-enhanced characteristics of each level into a decoding block of a corresponding layer of a decoder branch, and carrying out decoding and spatial up-sampling operation in sequence from a high layer to a low layer; and the output of the last decoding block of the decoder branch is subjected to 1 * 1 convolution and double up-sampling to obtain a human eye fixation point prediction result. According to the method, frequency domain feature enhancement and sequential decoding from a high layer to a low layer are carried out on the spatial multi-layer convolution features extracted by the backbone network, so that the human eye fixation point prediction precision is improved.
Owner:TONGDA COLLEGE OF NANJING UNIV OF POSTS & TELECOMM