The invention discloses a first
view angle video generation method and device based on a
mask diffusion model and
fixation point constraint, and the method comprises the steps: constructing an end-to-end
deep learning frame for the demands of video filling, prediction and
visual attention region control in the generation of a first
view angle video; the diversified first-view-angle video generation conforming to the visual logic is realized. The method comprises the following steps: firstly, dividing an input video into a condition frame set and an unknown frame set, and adding
random noise to an unknown frame by using a dynamic
mask module I to generate a noisized frame; performing reverse de-noising generation on the video through a full 3D
convolutional neural network module II, and guiding the network to generate video content conforming to a space-time law by taking a
diffusion step and a
fixation point track as conditional constraints; and finally, saliency prediction is carried out by using a
fixation point positioning module III. By jointly optimizing the loss of the reverse denoising
generation process and the loss of the fixation point probability graph, the reasonability and diversity of the generated video are further improved. According to the method, through joint optimization of the
mask diffusion strategy and fixation point constraint, the space-time continuity, the anti-
noise capability and the adaptability to the fixation
point trajectory of the generated video can be effectively improved, and the method is particularly suitable for a complex multi-face expression interaction scene.