Embodiments of the present application provide a multi-sound character disambiguation method and device,
electronic equipment and storage medium, including: obtaining attribute information of a target multi-sound character including
mask information, word segmentation information, part-of-speech information and
semantic information, inputting the attribute information into a
Transformer encoder including an initial classifier, a final classifier and a tone classifier, splicing the output results to generate a first
pinyin prediction result, and determining a final
pinyin prediction result according to
pinyin weight information of the target multi-sound character and the first pinyin prediction result. In the case of insufficient data or
data imbalance, the initial classifier, the final classifier and the tone classifier are used to fully
train the initial classifier, the final classifier and the tone classifier, so as to improve the multi-sound character prediction accuracy. At the same time, by increasing the pinyin weight information, the possible multi-sound character pronunciation can be limited in advance, so that the multi-sound character disambiguation prediction result is more accurate.