The invention discloses a
speech synthesis method suitable for an off-line scene of
the Internet of Things, and relates to the technical field of speech
data processing. The method comprises the following steps: acquiring an original human voice
syllable sample and a to-be-synthesized target
syllable sequence; calculating short-time root-mean-square energy of each frame of the original voice, and performing endpoint detection, framing windowing and FFT (
Fast Fourier Transform) to obtain a
frequency domain spectrum; extracting an amplitude higher than a preset multiple of the average amplitude in each frame of spectrum and a corresponding frequency as a characteristic wave peak, and packaging the characteristic wave peak into a compressed characteristic data packet to be stored locally; extracting a compressed data packet corresponding to the target
syllable from the local, taking the characteristic
wave crest of each frame as a node, complementing the
frequency spectrum amplitude between nodes through linear interpolation, and adjusting the amplitude according to the short-time energy of each frame to obtain a
time domain voice frame; and carrying out frame splicing and tone adjustment, and outputting high-quality human audio voice. According to the invention, while the low storage requirement of
speech synthesis is met, the requirements of instantaneity and clear tone quality are met.