The invention discloses a Token-based
visual task generation method, and belongs to the technical field of intelligent task
automation, and the method comprises the steps: S1, cross-
modal alignment; s2, performing visual Token
processing; s3, constructing a task description Token sequence: constructing the task description Token sequence based on a predefined
visual task template library according to requirements of a user or a specific application scene; s4, checking task feasibility; s5, task
priority scheduling; s6, training a task generation model; s7, dynamic task allocation; and S8, optimizing the model. According to the method, the long
sequence processing capability is optimized through a hierarchical merging strategy, the compactness of feature expression is realized while space position information is reserved, linear projection and enhanced position coding are combined to form a visual Token sequence with strong representation capability, local detail features are contained, a global context relationship is kept, and the method is suitable for the visual Token sequence with high representation capability. High-information-density feature input is provided for subsequent task
processing, and the
processing precision of various visual algorithms is effectively improved.