The invention relates to the technical field of voice
processing, can be applied to business scenes of financial science and technology,
medical health and the like, and discloses a context information-based voice
emotion recognition method, device, equipment and medium, which comprises the following steps: receiving an original voice
stream and generating an independent voice segment, recognizing a text and determining a speaker role type, and extracting an acoustic feature index; and generating a preliminary emotion
label, generating context information in combination with the historical dialogue text, and inputting the context information, the preliminary emotion
label, the speaker role type and the acoustic feature index into a multi-
modal fusion module to generate an emotion judgment result. According to the method, multi-
modal fusion is realized on the basis of context information by combining voice, text and role information, so that the emotion change of each role can be accurately recognized and understood in a complex dialogue scene, the problems of large
single sentence emotion judgment error and
neglect of the context information in a traditional method are avoided, and the accuracy and stability of
emotion recognition are effectively improved.