The invention discloses a voice
recognition system and method based on audiovisual bimodal
perception, and relates to the technical field of voice recognition, and the method comprises the steps: carrying out the feature
recursion of an aligned bimodal sequence based on an audiovisual frame
package, obtaining an audio high-level representation and a visual high-level representation, and calculating the
modal reliability; and fusing the audio high-level representation and the visual high-level representation based on the
modal reliability, and evaluating the prefix stability and confidence state through a streaming decoder to obtain a streaming identification information packet. According to the method, the audio high-level representation and the visual high-level representation are fused based on the
modal reliability, and the streaming decoder is used for evaluating the prefix stability and the confidence state to obtain the streaming identification information packet, so that the decoding output has the confirmable stable fragment and the credibility representation, and the continuity and the consistency of the
speech identification text are improved.