The invention relates to a method and a device for generating a report under the assistance of a
label, and the method comprises the following steps: 1) extracting a structured
label set from a text report of a sample based on a large
language model, wherein the structured
label set comprises a multi-classification group consisting of dichotomous labels and
mutual exclusion options; 2) aggregating the labels in batches, after a threshold value is reached, merging and de-weighting, performing specification and
mutual exclusion group merging on synonymous, near-synonymous and redundant labels, and converging into a unified label
library; 3) based on the text report and the tag
library, outputting a tag subset of each sample through a large
language model; 4) multi-
modal multi-label classification model training: extracting each visual
modal feature, performing weighted aggregation and splicing, and outputting each label group logits through a classification head to perform weighted group loss optimization; (5) carrying out joint training on the multi-
modal large
language model by using samples of'only images-reports' and'images + labels-reports', and (6) carrying out label prediction and screening on the images by using the classification model, and inputting'images + prediction labels' into the multi-modal large language model to obtain a final report.