The present invention proposes a laryngeal
paralysis diagnosis auxiliary method combined with audio and
video processing, comprising the following steps: obtaining an original laryngoscope video with audio information and video information; predicting the audio information through a keyword recognition model to extract an effective
phonation segment, and obtaining a first laryngoscope segment corresponding to the effective
phonation segment from the original laryngoscope video segment according to the matching relationship between the audio information and the video information; detecting and identifying the
glottis area in the first laryngoscope segment, and segmenting the
glottis area from the first laryngoscope segment; marking the
glottis area and calculating the physical properties of the glottis area, the physical properties including one or more of the glottis area area, vocal cord opening and closing angle, and vocal cord width. The present invention uses a multimodal
analysis method combined with audio and video to segment the laryngoscope video, extract key video segments, and provide a variety of
objective evaluation indicators for doctors to refer to, thereby improving diagnostic efficiency and quality.