The application discloses a vocal cord white
lesion diagnosis method and
system based on a pre-training
large model and attention transfer, and relates to the technical field of medical
image processing and
artificial intelligence auxiliary diagnosis. Frames are extracted from a laryngoscope video, a standardized image dataset is constructed, and a pre-trained visual
backbone network is used to extract multi-scale features, and an exogenous probability
heat map is generated as visual prior by means of a multi-
modal medical basic model. By calibrating the prior to the attention domain and aligning it with the internal attention of the network, the model is guided to focus on the
lesion area. A two-
stage classification strategy is adopted, first screening "
cancer" and "non-
cancer", and then subdividing "non-
cancer" lesions into "special infection or
inflammation", "vocal cord white spot-low risk" and "vocal cord white spot-high risk", and improving robustness through
ensemble learning. The
system synchronously outputs an interpretable
heat map and a structured diagnosis conclusion, and supports clinical interaction. It effectively improves the consistency of diagnosis, clinical applicability and work efficiency.