The application provides a
speech synthesis method,
system,
electronic equipment and storage medium, wherein the method comprises: inputting the speech features of the source speech of a target object into an
encoder in a
speech synthesis model to obtain first encoding features and second encoding features of the source speech; inputting the first encoding features of the source speech into an age
perception module in the
speech synthesis model to obtain target age features of first age data, and obtaining target age features of second age data according to the target age features of the first age data; inputting the target age features of the second age data, the second encoding features of the source speech and
target text to be synthesized into a speech synthesis module in the speech synthesis model to obtain target speech of the target object under the second age data. Through age feature decoupling and age feature stretching, the application effectively reduces the cost and complexity of data collection while improving the high-precision synthesis of specific age speech.