A
radiology report automatic generation method,
system, device and medium based on
patient specific priori knowledge uses a special token to represent missing patient clinical context information, so that a text
encoder can process complete and incomplete clinical context input in a unified manner, and then robust clinical context features are extracted; the method comprises the following steps: constructing a space-time fusion network STF, integrating previous medical images of a patient, establishing a difference mapping relation between a current image and a historical image for modeling an evolutionary process of a
disease, and extracting space-time visual features with time dependence; establishing an attention-enhanced hierarchical fusion network for fusing multiple
layers of hidden states in a visual
encoder so as to extract multiple
layers of hierarchical visual features with rich
semantics; a prior
perception progressive fusion network is introduced, and
patient specific prior knowledge and hierarchical visual features are gradually fused in a coarse-to-fine mode to generate multi-
modal features facing
radiology report generation; the method comprises the following steps: designing a two-stage training strategy: in the first stage, aiming at image-
text alignment, improving the accuracy of medical image-
text retrieval; and in the second stage, the generation of the
radiology report is taken as a target, a text decoder is optimized, and the performance of the generated
radiology report in the aspects of clinical semantic accuracy and language expression quality is improved.