The invention discloses an intelligent voice
interaction method and device, which are applied to the technical field of
data processing, and the method comprises the steps: obtaining multi-
modal original data of a user image, voice and text,
processing through a
specific model and a tool, detecting and
cutting a human face through RetinaFace, then extracting an image emotion feature through a ResNet-50
network model, extracting a voice MFCC feature through a torchaudio
library, and carrying out the recognition of the MFCC feature through the RetinaFace; the method comprises the following steps: extracting semantic features of a text by a Sension-BERT Chinese model, and generating three pieces of standardized single-mode
feature data; and the dimension is unified through linear projection, a cross-
modal attention mechanism and a Transform
encoder fusion feature are combined, and an emotion classifier is input to obtain an
emotion recognition result. Then, according to the Plutchik emotional wheel theory, a service scene and role constraint matching text and voice double-
response strategy is combined, multi-
modal interaction response content is generated, finally,
continuous interaction is achieved in combination with intelligent hardware and circulating monitoring, and an intelligent interaction result meeting the emotion and scene requirements of a user is output.