The application relates to the technical field of
artificial intelligence, in particular to a multi-
modal human-computer interaction
content filtering method, device and medium. The method comprises the following steps: collecting multi-
modal interaction data; obtaining a current application scene, calculating single invalidity scores of various data, obtaining joint invalidity scores in combination with multi-
modal weights corresponding to the scene, and judging whether the joint invalidity scores are lower than a preset threshold. If yes, scene-specific violation judgment weights are called, multi-modal violation features are extracted from the data, and current fusion features are obtained by combining the weights; the latest N (N>1) historical fusion features are obtained, the current fusion features are combined to generate modified fusion features, the modified fusion features are input into a pre-trained
deep learning model, and a risk confidence is obtained. When the risk confidence exceeds a preset threshold, an interaction instruction transmission link is immediately
cut off, and multi-
modal data is encrypted and stored as evidence, so that accurate filtering and high-
risk prevention and control of multi-modal interaction content are realized. The application can effectively improve the judgment accuracy.