This invention provides a method and
system for long-range lip-reading recognition using unmanned aerial vehicles (UAVs), belonging to the fields of
artificial intelligence and UAV technology. The
system utilizes a camera, IMU (Integrated Mutor Unit), and
microphone mounted on the UAV to acquire video streams, IMU data streams, and audio streams. Motion compensation is applied to the video
stream using IMU data to stabilize image frames. Then, image sequences of the target speaker's lips and nasal /
jaw region are extracted and input into a lip-reading recognition model to obtain initial recognition text and visual confidence scores. Simultaneously, the
signal-to-
noise ratio (SNR) of the audio
stream is calculated. When the visual
confidence score is low and the SNR is high, the
audio recognition process is triggered, generating
audio recognition text and an audio
confidence score. Finally, the visual and audio confidence scores, along with the lip movement-audio temporal matching degree, are weighted and fused to generate the final recognition text. This invention effectively solves the problems of video
instability and background interference in long-range lip-reading recognition, improving recognition accuracy in complex environments.