A neural network
inference acceleration method and device are provided. The method includes: acquiring, by a memory having a computing unit, a video and a prompt word; acquiring, by the memory having the computing unit, at least one image related to the prompt word from the video; dividing, by the memory having the computing unit, each of the at least one image into a plurality of blocks; storing, by the memory having the computing unit, the plurality of blocks; storing, by the memory having the computing unit, position information of the plurality of blocks in each of the at least one image; acquiring, by a processor, each of the at least one image as the plurality of blocks and the position information; dividing, by the processor, each of the at least one image into a plurality of sub-blocks; acquiring, by the processor, from the memory having the computing unit, an embedded vector of a pre-stored sub-block corresponding to each of the plurality of sub-blocks, respectively; and performing an
inference operation using the embedded vector and a first neural network.