In some implementations, a
user device may receive a set of words, for conversion to speech, including a plurality of subsets of the set of words, where each subset of the plurality of subsets is associated with one or more corresponding formatting properties. The
user device may identify, within the set of words, a first subset of the plurality of subsets as relevant based on the one or more corresponding formatting properties associated with the first subset. Additionally, the
user device may identify, within the set of words, a second subset of the plurality of subsets as not relevant based on the one or more corresponding formatting properties associated with the second subset. Accordingly, the user device may input the first subset to a text-to-speech
algorithm.