The invention provides a PDF
analysis method and
system based on image-text
semantic alignment, and belongs to the technical field of
document processing. The method comprises the following steps: reading PDF or picture file
byte data, and converting pictures into PDF
byte data in a unified format; creating an image and a Markdown
directory according to the output
directory, the PDF name and an
analysis method; extracting
byte data of a specified page of the PDF according to starting and ending page numbers; selecting a back end to analyze PDF byte data to obtain a reasoning result, an image
list and the like; generating an intermediate
JSON containing PDF detailed information according to an analysis result; analyzing the positions of the text and the picture, and associating information to realize
semantic alignment of the picture and the text; and generating multiple types of
machine readable files according to the intermediate
JSON and storing the files in a specified
directory. According to the method, the PDF multi-mode content can be accurately extracted and subjected to
semantic alignment, the PDF multi-mode content is efficiently converted into a
machine readable format, the analysis accuracy and
usability are improved, and the method is suitable for PDF analysis of
multiple image-text tables such as scientific and technical literatures.