The application relates to the technical field of
speech recognition, in particular to a speech-based
emotion recognition method, which comprises the following steps: according to an input speech
signal, frame division and windowing are carried out. The application can obtain more profound insights into emotional expression by decomposing the input speech
signal into two physical sources of glottal excitation and
vocal tract response for independent modeling and analysis, obtaining a
vocal tract transfer function set via
linear predictive coding operation, and applying inverse filtering to the original speech
signal to reconstruct an approximate glottal
pulse sequence, effectively stripping the influence of
vocal tract resonance on the signal, so that perturbation parameters, open quotient and closed quotient and the like representing vocal cord vibration patterns can be directly calculated, at the same time, the vocal tract
transfer function set is used to identify and track the dynamic trajectory of the
formant, and the
phase difference cosine mean between adjacent frames is combined to quantify the
sound production stability, and subtle dynamic adjustment and control stability of
sound production organs such as the
oral cavity and tongue position caused by
emotional changes are captured.