The invention discloses a tumor and target region synchronous segmentation method and
system based on visual language guidance. The method comprises the following steps: constructing a multi-
modal data set fusing a medical image and a clinical text; a unified
encoder is adopted to extract shared visual features, a double-
branch decoder is adopted to predict a
tumor region (GTV) and a clinical target region (CTV) respectively, and
information sharing and task differentiation modeling are achieved; introducing a medical pre-training large
language model, and respectively extracting language features related to GTV and CTV; a visual language collaborative attention module is designed, dynamic fusion and alignment of visual languages are carried out, task related
semantics are guided and enhanced, and redundant information is inhibited; and outputting a precise synchronous segmentation result of the GTV and the CTV after fusing the multi-
modal features. Compared with an existing scheme, the problems that the boundary of a
tumor target region in a medical image is fuzzy and multi-
modal information is insufficient in utilization are effectively solved, the segmentation precision and robustness are remarkably improved, and reliable support is provided for clinical radiotherapy planning.