The invention relates to the technical field of
artificial intelligence, can be applied to business scenes of financial science and technology,
medical health and the like, and discloses a visual token generation method, device, equipment and medium based on a shared index, and the method comprises the steps: obtaining an input image, and extracting semantic features and pixel features through a semantic
encoder and a pixel
encoder; calculating the distance between each feature and a
codebook thereof, and carrying out weighted summation to determine a shared index; retrieving quantitative features from the
codebook using a shared index; respectively generating a reconstructed image and a reconstructed
semantic feature by using a pixel decoder and a semantic decoder; jointly optimizing an
encoder, a
codebook and a decoder based on a reconstruction result; and generating a unified visual token sequence for the target task image by using the optimized component. Through double-flow
feature extraction, shared mapping quantization and joint loss optimization, global
semantic information and local pixel details can be reserved in the visual token at the same time, so that the model has accurate understanding ability, high-fidelity images can be generated, and the performance of understanding and generating tasks is improved.