The embodiment of the application relates to a
processing method and device for mining patent information in combination with a large
language model, the method comprising: selecting a large
language model as a first model, assigning five target tasks to the first model, and constructing five task data sets to fine-tune the first model; receiving a PDF file of a
small molecule drug patent after fine-tuning is completed and performing image, text block and table extraction on the PDF file; then performing molecular image / molecular
label / irrelevant table screening and molecular activity table splicing based on the model; then constructing a
label-image corresponding relationship table based on the molecular image / molecular
label / spliced
table information sequence and the first model, and constructing a label-activity information corresponding relationship table based on the first model and the spliced
table information sequence; and merging the two corresponding relationship tables to obtain a multi-
modal comprehensive information table; finally, a corresponding mining report is fed back by the patent file and the multi-
modal comprehensive information table. Through the application, the
data mining efficiency can be improved.