The invention relates to a multi-agent video multi-round question and answer method based on
implicit knowledge enhancement, and the method comprises the steps: merging the video features of a real agent and the features of a
virtual agent, and generating an extended agent
feature set; constructing a multi-agent video question and answer model; training a multi-agent video question and answer model based on the extended agent
feature set to obtain a trained multi-agent video question and answer model; and inputting the video features of the plurality of real agents and the to-be-processed
question text features into the trained multi-agent video question and answer model to obtain a final video question and answer result. According to the method and the device, support features of historical rounds are expressed as
virtual agent features, and unified modeling is carried out on the
virtual agent features and real agent features of the current round, so that efficient
multiplexing of context information in multi-round
questions and answers is realized. The method has remarkable advantages in question and answer accuracy, calculation efficiency and communication overhead, and a new solution is provided for multi-round multi-
modal question and answer tasks.