A voice changing method, apparatus and device for a speech dialog in a live-streaming room, and a medium, which relate to the field of network
live streaming. The method comprises: in response to a speech speaking event of a speaker user in a live-streaming room, detecting and determining a speech segment in target audio data; on the basis of a preset target
pitch value, performing tone-change
processing on a segment
pitch feature of the speech segment, so as to obtain an optimized
pitch feature; on the basis of the optimized pitch feature and a target
timbre feature, performing voice-change
processing on the speech segment, so as to obtain a voice-changed segment; and replacing a corresponding speech segment in the target audio data with the voice-changed segment, so as to obtain voice-changed audio data, and sending the voice-changed audio data to a
receiver user in the live-streaming room. The present application significantly improves the performance of voice changing technology, and solves the problems in the conventional technology in the aspects of real-time performance, personalized services, pitch adjustment naturalness, real
timbre restoration, etc.