The invention relates to the technical field of voice
processing, can be applied to business scenes such as financial science and technology and
medical health, and discloses a multi-speaker dialogue
voice analysis method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-analyzed multi-speaker dialogue voice, determining a naturalness
score based on an acoustic feature, and obtaining a multi-speaker dialogue
voice analysis result; determining a
semantic consistency score based on voice embedding and semantic embedding corresponding to a preset text, determining a speaker consistency
score based on embedding of a plurality of speakers of the same speaker, determining an interaction rationality score based on voice alternate overlapping duration, determining a diversity score based on a variance of voice features, and fusing the scores, the comprehensive
mass fraction is obtained. According to the invention, through quantitative evaluation of five dimensions of naturalness,
semantic consistency, speaker consistency, interaction rationality and diversity, a comprehensive quality
scoring system is established, so that the
evaluation result simultaneously reflects voice fluency, content matching degree, identity stability, interaction
rhythm rationality and feature richness.