The application discloses an online conference translation method and
system, and belongs to the technical field of
speech recognition. The language space is constrained by the interactive side
signal, and the recognition uncertainty is reduced. In the recording stage, a target language
label generated by a sliding gesture is introduced. The
label reflects the language of the speaker to whom the speech is directed. The language
label of the speaker's mother tongue is combined, the candidate language set is optimally converged into a two-language set of "mother tongue + target language", and the search space of
speech recognition in the language dimension is significantly reduced from the "complete language set of the conference" to the "two-language set". Based on this, the misrecognition probability caused by multi-language competition is reduced, especially the probability of misrecognizing foreign language terms as mother tongue homophonic words is reduced, and it can be quickly determined which languages are involved in each
sentence of speech, so as to quickly determine which mixed
language model is used for
speech recognition, and the speech recognition efficiency of mixed
language speech is improved.