One or more embodiments of the invention provide a question and answer method, and an agent training method and device. According to the method, after an
intelligent agent obtains first question and answer information, whether visual focusing
processing needs to be carried out on an inquiry image or not can be inferred based on the first question and answer information, and under the condition that visual focusing
processing needs to be carried out on the inquiry image through
inference, an image of a target area related to a first question in the inquiry image can be extracted; obtaining a regional image; and retrieving and generating a first answer corresponding to the first question by using an
image retrieval tool and taking the regional image as a retrieval condition. Under the condition that visual focusing
processing does not need to be carried out on the inquiry image through reasoning, the complete inquiry image can be used as a retrieval condition for retrieval. Afterwards,
text retrieval can be dynamically performed or a final answer can be made based on an
image retrieval result.