This application relates to a method, apparatus, device, medium, and program product for generating
speech training data. The method includes: acquiring initial training data, the initial training data comprising at least one audio-text pair;
processing the audio and text in each audio-text pair to obtain first
timestamp information corresponding to each word in the audio-text pair; based on the audio in each audio-text pair, obtaining each sub-language event and second
timestamp information corresponding to the sub-language event; based on the first
timestamp information corresponding to each word and the second timestamp information corresponding to the sub-language event, generating text
insertion positions corresponding to each sub-language event; and based on the text
insertion positions corresponding to each sub-language event and the initial training data, obtaining target training data. This method can reduce costs.