The invention relates to the technical field of
speech recognition, and provides an end-to-end
speech recognition method and
system based on a keyword attention enhancement mechanism, and the method comprises the steps: extracting entity keywords through a keyword searcher; the voice cache manager receives continuous
air traffic control audio clips, converts the continuous
air traffic control audio clips into an audio Mel
spectrogram and then converts the audio Mel
spectrogram into an audio embedded sequence; the keyword
encoder unit maps the lexical element embedding representation sequence into a keyword embedding sequence; the audio
transliteration decoder unit performs keyword attention enhancement calculation and converts splicing vectors of the audio embedding sequence, the
transliteration start mark and the lexical element embedding representation sequence into a prediction text sequence; and inputting the new lexical element embedded representation sequence into a text translator through autoregression until the audio transwriting decoder unit outputs a transwriting end mark, and outputting the predicted text sequence as a
speech recognition text. According to the invention, real-time identification of the streaming input voice is realized, and the method has the advantages of high identification precision, low
response delay, flexible deployment and the like.