The invention belongs to the field of medical vision-language pre-training, and relates to a medical vision-language pre-training method based on multi-view and text combination, which comprises the following steps of: preprocessing chest X-
ray image-report data; inputting the preprocessed image into a view
encoder to obtain local features VF and VL and global representations gF and gL of a positive view and a side view, wherein the local features VF and VL of the positive view and the local features VL of the side view and the global representations gF and gL of the positive view and the side view are compared with the report sequence XT of the
mask and the report sequence # imgabs0 # of the
mask; respectively inputting the sequence XT and # imgabs1 # into a report
encoder to obtain a local feature T, a global representation gT and a
mask report representation # imgabs2 #, and inputting VF, VL, T, gF, gL, gT and # imgabs3 # into a positive side view feature integration module to obtain a mask
report generation text T'and a
prediction probability P thereof; inputting the VF, the VL and the T into a positive and lateral feature alignment module to obtain fine-grained representations F and L of positive and lateral views; calculating a
loss function value according to P, F, L, gF, gL and gT, and updating
model parameters according to the
loss function value until a pre-trained medical vision-language general model is obtained; according to the method, the
lateral view is introduced into pre-training, so that the diagnosis accuracy is improved.