The invention discloses a universal voice representation model training method and a related device, and relates to the technical field of
artificial intelligence, and the method comprises the steps: obtaining a voice sample and a
mask voice sample corresponding to the voice sample, inputting the voice sample into a teacher model, obtaining a teacher voice representation, inputting the
mask voice sample into a student model, and obtaining a student voice representation; obtaining a first student voice representation, generating global knowledge
distillation loss and local comparison loss according to the teacher voice representation and the first student voice representation, and updating target parameters of the student model by using the global knowledge
distillation loss and the local comparison loss; and the trained student model is used as a general voice representation model. According to the method, knowledge
distillation and comparative learning are combined, so that the student model can extract multi-aspect voice information, and therefore, the method can adapt to various voice tasks, and the universality is higher. The characteristics of the
mask area can be accurately reconstructed through double knowledge migration, so that the accuracy of voice representation extracted by the student model is higher.