The invention relates to the technical field of
natural language processing, and discloses a Yi
language speech recognition method based on self-supervision and attention
feature fusion. The method comprises a feature
encoder module, a comparative learning module, a
mask language modeling module, a joint optimization and
feature fusion module and a decoder module, the feature
encoder module adopts a
convolutional neural network structure and converts continuous waveform signals into feature representation suitable for subsequent modeling, and the comparative learning module performs
feature fusion on the continuous waveform signals through a Gumbel-Softmax technology. The method comprises the following steps that: a
mask language modeling module and a feature fusion module are integrated, deviation caused by manual definition or clustering is avoided, the
mask language modeling module obviously enhances semantic understanding and tone modeling capabilities of a model in a low-resource scene, a self-attention feature
fusion mechanism is introduced into the feature fusion module, continuous features, discrete unit representation and
semantic context representation from an acoustic level are spliced, and a self-attention feature
fusion mechanism is introduced into the self-attention feature
fusion mechanism. The decoder module adopts a decoder
structure based on
connection time sequence classification, and the corresponding relation between the voice and the text can be achieved without strict alignment labeling.