The invention discloses a
human eye fixation point prediction method based on a deep convolutional network and
frequency domain feature enhancement, and the method comprises the steps: employing the deep convolutional network (namely a residual network ResNet) as a
backbone network to form an
encoder branch, and employing five
layers of coding blocks to extract the spatial features of five
layers of an input image; the five
layers of spatial features extracted by the
encoder are respectively sent to a
frequency domain feature enhancement module for
frequency domain feature enhancement based on
discrete cosine transform (DCT) so as to enhance the features of each layer of
human eye fixation area and reduce interference features; sending the frequency-domain-enhanced characteristics of each level into a decoding block of a corresponding layer of a decoder
branch, and carrying out decoding and spatial up-sampling operation in sequence from a high layer to a low layer; and the output of the last decoding block of the decoder
branch is subjected to 1 * 1
convolution and double up-sampling to obtain a
human eye fixation point prediction result. According to the method, frequency domain feature enhancement and
sequential decoding from a high layer to a low layer are carried out on the spatial multi-layer
convolution features extracted by the
backbone network, so that the human eye
fixation point prediction precision is improved.