A processing method and device for mining patent information in combination with a large language model

By automating the processing of small molecule drug patent information using large language models, the problem of low efficiency in manual data collection has been solved, achieving efficient data mining and improved R&D efficiency.

CN119782382BActive Publication Date: 2025-12-12BEIJING DP TECH CO LTD

Patent Information

Application Number
CN202411803496.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-12-12
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

In existing technologies, the data collection efficiency of small molecule drug patent information is low and requires manual operation, which leads to a decrease in research and development efficiency.

Method used

Using a pre-trained large language model, images, text blocks, and tables are extracted from PDF files of small molecule drug patents by specifying five target tasks and instruction text templates. A correspondence table between label and image and label and activity information is constructed to achieve automated data mining.

Benefits of technology

It improves the efficiency of small molecule drug patent data mining and reduces costs, thereby enhancing drug development efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782382B_ABST
    Figure CN119782382B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to a processing method and device for mining patent information in combination with a large language model, the method comprising: selecting a large language model as a first model, assigning five target tasks to the first model, and constructing five task data sets to fine-tune the first model; receiving a PDF file of a small molecule drug patent after fine-tuning is completed and performing image, text block and table extraction on the PDF file; then performing molecular image / molecular label / irrelevant table screening and molecular activity table splicing based on the model; then constructing a label-image corresponding relationship table based on the molecular image / molecular label / spliced table information sequence and the first model, and constructing a label-activity information corresponding relationship table based on the first model and the spliced table information sequence; and merging the two corresponding relationship tables to obtain a multi-modal comprehensive information table; finally, a corresponding mining report is fed back by the patent file and the multi-modal comprehensive information table. Through the application, the data mining efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a processing method and device for mining patent information in combination with a large language model. BACKGROUND

[0002] In the development process of small molecule drugs, collecting and analyzing patent information is a necessary step. A large amount of compound molecular structures, chemical names and biological activity data contained in patents are important information sources for drug design and optimization. However, these data usually exist in non-standardized formats and are scattered in the text, tables and images of patent files, which requires researchers to collect information through manual operation. Obviously, this will inevitably reduce the R&D efficiency. SUMMARY

[0003] The purpose of the present application is to overcome the defects of the prior art, and to provide a processing method and device for mining patent information in combination with a large language model, an electronic device and a computer readable storage medium. The present application selects a type of pre-trained multi-modal large language model as a first model, specifies five target tasks (molecular image recognition on input images, molecular label recognition on input text, classification on input tables, similarity recognition on two input data rows, and molecular label column recognition on input tables), customizes five instruction text templates, constructs five data sets, and fine-tunes the first model based on the five data sets and the corresponding five instruction text templates; after the model fine-tuning is completed, the user inputs a PDF file of a small molecule drug patent, and extracts images, text blocks and tables from the patent file based on the PDF content extraction tool; then, based on the first instruction template and the first model, the first image information sequence is screened for molecular images, based on the second instruction template and the first model, the first text block information sequence is screened for molecular labels, based on the third instruction template and the first model, the first table information sequence is filtered for other tables, and based on the fourth instruction template, the first model and the second table information sequence, the molecular activity table and the table without headers are spliced; then, based on the first molecular image information sequence, the first molecular label information sequence, the first spliced table information sequence, the fifth instruction template and the first model, a label-image correspondence table is constructed, and based on the fifth instruction template, the first model and the first spliced table information sequence, a label-activity information correspondence table is constructed; finally, based on the label-image correspondence table and the label-activity information correspondence table, a multi-modal table is merged to obtain a corresponding multi-modal comprehensive information table, which is fed back to the current user. The patent information mining mechanism provided by the present application can mine molecular activity data from small molecule drug patents, and based on the present application, the data mining efficiency can be improved, the data mining cost can be reduced, and the R&D efficiency can be improved.

[0004] To achieve the above object, the embodiment of the present application provides a processing method for mining patent information combined with a large language model, which comprises the following steps:

[0005] a plurality of pre-trained multi-modal large language models are selected as corresponding first models; five target tasks are specified for the first models, which are referred to as first, second, third, fourth and fifth tasks; instruction text templates are customized for the first, second, third, fourth and fifth tasks, which are referred to as first, second, third, fourth and fifth instruction templates; five task data sets are constructed by collecting public small molecule drug patent literature information, which are referred to as first, second, third, fourth and fifth data sets; the first model is fine-tuned based on the first, second, third, fourth and fifth data sets and the first, second, third, fourth and fifth instruction templates; the first task is to perform molecular image recognition on an input image; the second task is to perform molecular label recognition on an input text; the third task is to classify an input table; the fourth task is to perform similarity recognition on two input data rows, and the total number of data fields of the two data rows is consistent; the fifth task is to perform molecular label column recognition on an input table;

[0006] After the fine-tuning of the first model is completed, a user-input small molecule drug patent PDF file is received as a corresponding first patent file; image, text block and table extraction are performed on the first patent file based on a pre-set PDF content extraction tool to obtain a corresponding first image information sequence, a first text block information sequence and a first table information sequence;

[0007] Based on the first instruction template and the first model, molecular image screening processing is performed on the first image information sequence to obtain a corresponding first molecular image information sequence; based on the second instruction template and the first model, molecular label screening processing is performed on the first text block information sequence to obtain a corresponding first molecular label information sequence; based on the third instruction template and the first model, other table filtering processing is performed on the first table information sequence to obtain a corresponding second table information sequence; based on the fourth instruction template, the first model and the second table information sequence, molecular activity table and table without header are spliced to obtain a corresponding first spliced table information sequence;

[0008] constructing a label-image correspondence relationship table based on the first molecular image information sequence, the first molecular label information sequence, the first splicing table information sequence, the fifth instruction template, and the first model; and constructing a label-activity information correspondence relationship table based on the fifth instruction template, the first model, and the first splicing table information sequence; and performing multi-modal table merging based on the label-image correspondence relationship table and the label-activity information correspondence relationship table to obtain a corresponding multi-modal comprehensive information table;

[0009] a corresponding first mining report composed of the first patent file and the multi-modal comprehensive information table is fed back to the current user.

[0010] Preferably, the first model at least includes a GPT series model, an LLaMA series model, a Qwen series model, and a Wenxin series model.

[0011] The first, second, third, fourth, and fifth instruction templates are each a formatted instruction text module.

[0012] The first instruction template is used to take an input image as a corresponding current image; and through a piece of natural language description text, instructs the first model to identify whether the current image is a molecular structure graph to obtain a corresponding first identification result, and takes the first identification result as a corresponding first instruction feedback output; the first identification result includes yes and no.

[0013] The second instruction template is used to take an input text as a corresponding current context; and through a piece of natural language description text, instructs the first model to identify whether the current context contains molecular label information to obtain a corresponding second identification result, and when the second identification result is yes, further performs molecular label information extraction processing on the current context to obtain corresponding first molecular label extraction information, and when the second identification result is no, sets the first molecular label extraction information as empty, and takes the obtained second identification result and the first molecular label extraction information as a corresponding second instruction feedback output; the second identification result includes yes and no.

[0014] The third instruction template is used to take an input table as a corresponding current table; and through a piece of natural language description text, instructs the first model to identify which one of a molecular activity table, a table without a header, and other tables the current table belongs to to obtain a corresponding third identification result, and takes the third identification result as a corresponding third instruction feedback output; the third identification result includes a molecular activity table, a table without a header, and other tables.

[0015] The fourth instruction template is used to input two data rows as corresponding current data row A and current data row B; and through a piece of natural language description text, the first model is instructed to perform field-by-field data similarity estimation on the current data row A and B according to the data field correspondence relationship to obtain corresponding field similarity, and average similarity calculation is performed based on all the obtained field similarity, and it is identified whether all the field similarity and the average similarity are higher than a preset similarity threshold, if all are higher than the similarity threshold, the corresponding fourth identification result is set as similar, if at least one of the field similarity or the average similarity is lower than the similarity threshold, the corresponding fourth identification result is set as dissimilar, and the fourth identification result is taken as corresponding fourth instruction feedback output; the fourth identification result includes similar and dissimilar;

[0016] The fifth instruction template is used to input a table as a corresponding current table; and through a piece of natural language description text, the first model is instructed to identify whether there is a molecular label column in the current table to obtain a corresponding fifth identification result, and when the fifth identification result is yes, the current table is further subjected to molecular label column information extraction processing to obtain corresponding first column extraction information, and when the fifth identification result is no, the corresponding first column extraction information is set as empty, and the fifth identification result and the first column extraction information are combined to form corresponding fifth instruction feedback output; the fifth identification result includes yes and no; the first column extraction information is composed of one or more first column unit information when the first column extraction information is not empty; the first column unit information includes row index information and molecular label information;

[0017] The first image information sequence includes a plurality of first image information; the first image information includes a first identifier, a first page code, a first upper left corner row and column coordinate, a first upper right corner row and column coordinate, a first lower right corner row and column coordinate, a first lower left corner row and column coordinate, and a first image; the first image is a visual diagram of a molecular structure;

[0018] The first text block information sequence includes a plurality of first text block information; the first text block information includes a second identifier, a second page code, a second upper left corner row and column coordinate, a second upper right corner row and column coordinate, a second lower right corner row and column coordinate, a second lower left corner row and column coordinate, and a first block text;

[0019] The first table information sequence includes a plurality of first table information; the first table information includes a third identifier, a third page code, a third upper left corner row and column coordinate, a third upper right corner row and column coordinate, a third lower right corner row and column coordinate, a third lower left corner row and column coordinate, and a first table; the first table includes one or more first table rows; the first table row includes a plurality of first row unit texts;

[0020] The first splicing table information sequence includes a plurality of first splicing table information; the first splicing table information includes a fourth identifier, a pre-page page number, a post-page page number, a pre-page upper-left corner row-column coordinate, a pre-page upper-right corner row-column coordinate, a pre-page lower-right corner row-column coordinate, a pre-page lower-left corner row-column coordinate, a post-page upper-left corner row-column coordinate, a post-page upper-right corner row-column coordinate, a post-page lower-right corner row-column coordinate, a post-page lower-left corner row-column coordinate, and a first splicing table; the first splicing table includes one or more second table rows; and the second table row includes a plurality of second row cell texts.

[0021] The label-image correspondence table includes a plurality of first correspondence records; the first correspondence record includes a first molecular label field and a first molecular image field.

[0022] The label-activity information correspondence table includes a plurality of second correspondence records; the second correspondence record includes a second molecular label field and a plurality of first molecular activity fields.

[0023] The multi-modal comprehensive information table includes a plurality of third correspondence records; the third correspondence record includes the third molecular label field, the second molecular image field, and a plurality of second molecular activity fields.

[0024] Preferably, the first image information sequence is subjected to molecular image screening processing based on the first instruction template and the first model to obtain a corresponding first molecular image information sequence, specifically including:

[0025] All the first images of the first image information sequence are iterated once; and in this iteration process, the first image currently iterated is input into the first instruction template as a corresponding current image to assemble a corresponding current task instruction; the current task instruction is input into the first model to perform a molecular image recognition task to obtain a corresponding first instruction feedback; the first recognition result in the first instruction feedback is taken as a corresponding current recognition result; when the current recognition result is no, the first image information corresponding to the current image is deleted from the first image information sequence; and when the iteration is ended, the latest first image information sequence is taken as the corresponding first molecular image information sequence.

[0026] Preferably, the first text block information sequence is subjected to molecular label screening processing based on the second instruction template and the first model to obtain a corresponding first molecular label information sequence, specifically including:

[0027] performing a round of traversal on all the first block texts of the first text block information sequence; and in the round of traversal, taking the first block text currently traversed as a corresponding current text, inputting the second instruction template to assemble a corresponding instruction text to obtain a corresponding current task instruction; inputting the current task instruction into the first model to process a molecular label identification task to obtain a corresponding second instruction feedback; taking the second identification result of the second instruction feedback and the first molecular label extraction information as a corresponding current identification result and a current molecular label; and identifying the current identification result; if the current identification result is yes, resetting the text content of the first block text corresponding to the current text in the first text block information sequence as the corresponding current molecular label; if the current identification result is no, deleting the first text block information corresponding to the current text from the first text block information sequence; and at the end of the round of traversal, taking the latest first text block information sequence as the corresponding first molecular label information sequence.

[0028] Preferably, the other table filtering processing of the first table information sequence based on the third instruction template and the first model obtains a corresponding second table information sequence, specifically including:

[0029] performing a round of traversal on all the first tables of the first table information sequence; and in the round of traversal, taking the first table currently traversed as a corresponding current table, inputting the third instruction template to assemble a corresponding instruction text to obtain a corresponding current task instruction; inputting the current task instruction into the first model to process a table classification task to obtain a corresponding third instruction feedback; taking the third identification result of the third instruction as a corresponding current identification result; and identifying the current identification result; if the current identification result is a molecular activity table, marking the first table information corresponding to the current table as a corresponding first type table information; if the current identification result is a table without a header, marking the first table information corresponding to the current table as a corresponding second type table information; if the current identification result is other table, deleting the first table information corresponding to the current table from the first table information sequence; and at the end of the round of traversal, taking the latest first table information sequence as the corresponding second table information sequence.

[0030] Preferably, the molecular activity table and the table without a header are spliced based on the fourth instruction template, the first model and the second table information sequence to obtain a corresponding first spliced table information sequence, specifically including:

[0031] Step 61, initializing an empty sequence as the third table information sequence corresponding to the third table information sequence; and taking the first first table information of the second table information sequence as the current table information corresponding to the third table information sequence;

[0032] Step 62, identifying whether the current table information is marked as the first type of table information; if the current table information is marked as the first type of table information, identifying whether the row coordinate of the third lower left corner row-column coordinate of the current table information is the bottom row of the page; if yes, turning to step 63; if no, turning to step 64; if the current table information is not marked as the first type of table information, turning to step 65;

[0033] Step 63, add 1 to the third page code of the current table information to obtain a corresponding next page code; match each third page code in the second table information sequence with the next page code and the third upper left corner row and column coordinates, and take the second table information whose row coordinate of the third upper left corner row and column coordinates is the top row of the page as corresponding candidate table information; identify whether the candidate table information is empty; if the candidate table information is empty, go to step 64; if the candidate table information is not empty, record the first table of the candidate table information as a corresponding post table, record the first table of the current table information as a corresponding pre table, count the total number of the first row cell text of any first table row of the pre table to obtain a corresponding first cell total number, count the total number of the first row cell text of any first table row of the post table to obtain a corresponding second cell total number, and identify whether the first and second cell total numbers are equal; if the first and second cell total numbers are not equal, go to step 64; if the first and second cell total numbers are equal, input the second first table row of the pre table and the first first table row of the post table as corresponding first and second data rows into the fourth instruction template to obtain corresponding instruction text assembly to obtain a corresponding current task instruction, input the current task instruction into the first model to process a data row similarity identification task to obtain a corresponding fourth instruction feedback, take the fourth identification result in the fourth instruction feedback as a corresponding current identification result, and identify the current identification result; if the current identification result is not similar, go to step 64;If the current recognition result is similar, a first splicing table information with empty content is added in the third table information sequence as corresponding current splicing table information, a unique identification is assigned to the current splicing table information as corresponding current splicing table identification, the post-table is added to the end of the pre-table to obtain corresponding current splicing table, the fourth identification of the current splicing table information is set as the current splicing table identification, the pre-page and post-page page numbers of the current splicing table information are set as the third page numbers of the pre-table and post-table, the pre-page top-left coordinate, pre-page top-right coordinate, pre-page bottom-right coordinate and pre-page bottom-left coordinate of the current splicing table information are set as the third top-left coordinate, third top-right coordinate, third bottom-right coordinate and third bottom-left coordinate of the pre-table, and the post-page top-left coordinate, post-page top-right coordinate, post-page bottom-right coordinate and post-page bottom-left coordinate of the current splicing table information are set as the third top-left coordinate, third top-right coordinate, third bottom-right coordinate and third bottom-left coordinate of the post-table, the first splicing table of the current splicing table information is set as the current splicing table, and after the setting of the current splicing table information is completed, the process goes to step 65.

[0034] In step 64, the first table of the current table information is recorded as corresponding current table, a first splicing table information with empty content is added in the third table information sequence as corresponding current splicing table information, a unique identification is assigned to the current splicing table information as corresponding current splicing table identification, the fourth identification of the current splicing table information is set as the current splicing table identification, the pre-page page number of the current splicing table information is set as the third page number of the current table, the post-page page number of the current splicing table information is set as empty, the pre-page top-left coordinate, pre-page top-right coordinate, pre-page bottom-right coordinate and pre-page bottom-left coordinate of the current splicing table information are set as the third top-left coordinate, third top-right coordinate, third bottom-right coordinate and third bottom-left coordinate of the current table, the post-page top-left coordinate, post-page top-right coordinate, post-page bottom-right coordinate and post-page bottom-left coordinate of the current splicing table information are all set as empty, and the first splicing table of the current splicing table information is set as corresponding current table.

[0035] Step 65, identify whether the current table information is the last first table information of the second table information sequence; if not, extract the next first table information of the second table information sequence as the new current table information and return to step 62; if yes, output the latest first spliced table information sequence as the splicing processing result of this time.

[0036] Preferably, the first molecular image information sequence, the first molecular label information sequence, the first spliced table information sequence, the fifth instruction template and the first model are used to construct a label-image correspondence table, and specifically include:

[0037] Step 71, take the first image of the first molecular image information sequence as the corresponding current molecular image;

[0038] Step 72, take the first page code corresponding to the current molecular image as the corresponding current image page code, and record the area covered by the first top-left corner row and column coordinates, the first top-right corner row and column coordinates, the first bottom-right corner row and column coordinates and the first bottom-left corner row and column coordinates of the current molecular image as the corresponding current image area;

[0039] Step 73, record the first spliced table information in which the front page page code or the rear page page code of the first spliced table information sequence matches the current image page code as the corresponding first same-page table information; confirm the coverage range of the current first same-page table information on the current image page code based on the front / rear page top-left corner row and column coordinates, the front / rear page top-right corner row and column coordinates, the front / rear page bottom-right corner row and column coordinates, and the front / rear page bottom-left corner row and column coordinates of each first same-page table information to obtain the corresponding first same-page table area; when the number of first same-page table areas is not zero, confirm whether there is an intersection between the first same-page table area and the current image area to obtain the corresponding first confirmation result; identify the first confirmation result; if the first confirmation result is not exist, go to step 74; if the first confirmation result is exist, take the first same-page table information corresponding to the first same-page table area which has intersection with the current image area as the corresponding current same-page table information, and go to step 75;

[0040] Step 74, record the first text block information in which the second page number in the first molecular label information sequence matches the current image page number as corresponding first same-page text block information; and confirm the coverage range of the current first same-page text block information on the current image page based on the second upper-left corner coordinates, the second upper-right corner coordinates, the second lower-right corner coordinates and the second lower-left corner coordinates of each first same-page text block information to obtain the corresponding first same-page text block area; and take the first block text of the first text block information corresponding to the first same-page text block area closest to the current image area as the corresponding current molecular label; and go to step 76;

[0041] Step 75, input the first spliced table of the current same-page table information into the fifth instruction template to obtain corresponding instruction text assembly to obtain corresponding current task instruction; and input the current task instruction into the first model to process the molecular label column recognition task to obtain corresponding fifth instruction feedback; and extract the corresponding first column extraction information from the fifth instruction feedback; and record the full table coverage area of the single page or double page of the current same-page table information as the corresponding first table area; and record the area covered by the first upper-left corner coordinates, the first upper-right corner coordinates, the first lower-right corner coordinates and the first lower-left corner coordinates of each first molecular image in the first molecular image information sequence as the corresponding first image area; and record the first image area in which there is an intersection with the first table area as the corresponding first intra-table image area, and sort all the first intra-table image areas in the order of intersection area from front to back to form a corresponding first intra-table image area sequence; and take the sequence index of the first intra-table image area corresponding to the current image area in the first intra-table image area sequence as the corresponding current row index; and extract the molecular label information of the first column unit information in the first column extraction information that matches the current row index as the corresponding current molecular label;

[0042] Step 76, take the current molecular image as a corresponding first molecular image field; and take the current molecular label as a corresponding first molecular label field; and form a corresponding first correspondence record by the obtained first molecular label field and the first molecular image field;

[0043] Step 77, identify whether the current molecular image is the last first molecular image of the first molecular image information sequence; if not, extract the next first molecular image of the first molecular image information sequence as a new current molecular image, and return to step 72; if yes, the corresponding label-image correspondence table is composed of all the first correspondence records obtained.

[0044] Preferably, the label-activity information correspondence table is constructed based on the fifth instruction template, the first model and the first splicing table information sequence, specifically comprising:

[0045] Step 81, extract the first second table row of all the first splicing tables of the first splicing table information sequence to form a corresponding second table row set; and extract all the second row unit texts of the second table row set to form a corresponding second row unit text set; and perform unit text deduplication processing on the second row unit text set; and perform a round of iteration on all the second row unit texts of the second row unit text set after deduplication; and in this round of iteration, the second row unit text currently iterated is taken as a corresponding current field name; and based on a preset molecular activity keyword dictionary, identify whether the current field name is a type of molecular activity name to obtain a corresponding activity name recognition result; and when the activity name recognition result is yes, mark the current field name as a corresponding molecular activity field name; and at the end of this round of iteration, all the molecular activity field names obtained are composed of a corresponding molecular activity field name set;

[0046] Step 82, a round of traversal is performed on all the first splicing tables of the first splicing table information sequence; during the round of traversal, the first splicing table currently traversed is taken as a corresponding current splicing table; the current splicing table is input into the fifth instruction template to obtain a corresponding current task instruction through corresponding instruction text assembly; the current task instruction is input into the first model to obtain a corresponding fifth instruction feedback through molecular label column identification task processing; the first column extraction information corresponding to the fifth instruction feedback is extracted; the first second table row of the current splicing table is taken as a corresponding current header row; the second row cell text corresponding to the first column extraction information in the current header row is taken as a corresponding molecular label field name; all the second row cell texts in the current header row except the molecular label field name are polled once; during the polling, the second row cell text currently polled is recorded as a corresponding current field name, and it is identified whether the current field name is included in the molecular activity field name set; if yes, the second row cell text currently polled is taken as a corresponding molecular activity field name; if not, the arrangement index of the second row cell text corresponding to the current field name in the current header row is recorded as a corresponding to-be-deleted column index; when the polling is completed, the second row cell texts corresponding to all the to-be-deleted column indexes in all the second table rows of the current splicing table are deleted.

[0047] Step 83, after the round of traversal of all the first splicing tables of the first splicing table information sequence is completed, all field contents of all the first splicing tables are fully connected and combined with the molecular label field name as a connection key field name to obtain a new table with a molecular label field and all molecular activity fields as a corresponding label-activity information corresponding relationship table.

[0048] The second aspect of the embodiment of the application provides a device for implementing the processing method for mining patent information in combination with a large language model, the device comprising: a large language model fine-tuning module, a patent information receiving module, a patent information noise reduction module, a patent information mining module, and a mining report feedback module.

[0049] The large language model fine-tuning module is configured to select a plurality of pre-trained multi-modal large language models as corresponding first models; assign five target tasks to the first models, denoted as first, second, third, fourth, and fifth tasks; customize instruction text templates for the first, second, third, fourth, and fifth tasks, denoted as first, second, third, fourth, and fifth instruction templates; construct five task data sets based on collected public small molecule drug patent literature information, denoted as first, second, third, fourth, and fifth data sets; and fine-tune the first models based on the first, second, third, fourth, and fifth data sets and the first, second, third, fourth, and fifth instruction templates; the first task is to perform molecular image recognition on the input image; the second task is to perform molecular label recognition on the input text; the third task is to classify the input table; the fourth task is to perform similarity recognition on two input data rows, and the total number of data fields of the two data rows is consistent; and the fifth task is to perform molecular label column recognition on the input table;

[0050] The patent information receiving module is configured to receive a user-inputted PDF file of a small molecule drug patent as a corresponding first patent file after the fine-tuning of the first model is completed; and perform image, text block, and table extraction on the first patent file based on a pre-set PDF content extraction tool to obtain corresponding first image information sequence, first text block information sequence, and first table information sequence;

[0051] The patent information denoising module is configured to perform molecular image screening processing on the first image information sequence based on the first instruction template and the first model to obtain a corresponding first molecular image information sequence; perform molecular label screening processing on the first text block information sequence based on the second instruction template and the first model to obtain a corresponding first molecular label information sequence; perform other table filtering processing on the first table information sequence based on the third instruction template and the first model to obtain a corresponding second table information sequence; and perform molecular activity table and table without header splicing processing based on the fourth instruction template, the first model, and the second table information sequence to obtain a corresponding first spliced table information sequence;

[0052] The patent information mining module is configured to construct a label-image correspondence table based on the first molecular image information sequence, the first molecular label information sequence, the first spliced table information sequence, the fifth instruction template, and the first model; construct a label-activity information correspondence table based on the fifth instruction template, the first model, and the first spliced table information sequence; and perform multi-modal table merging based on the label-image correspondence table and the label-activity information correspondence table to obtain a corresponding multi-modal comprehensive information table;

[0053] The mining report feedback module is configured to feed back a corresponding first mining report to the current user, which is composed of the first patent file and the multi-modal comprehensive information table.

[0054] The third aspect of the embodiment of the present application provides an electronic device, comprising a memory, a processor and a transceiver.

[0055] The processor is configured to be coupled with the memory, read and execute instructions in the memory, so as to realize the method steps of the first aspect.

[0056] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transmission and reception.

[0057] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores computer instructions, when the computer instructions are executed by a computer, the computer instructions make the computer execute the instructions of the method in the first aspect.

[0058] The embodiment of the present application provides a processing method and device for mining patent information in combination with a large language model, electronic equipment and a computer readable storage medium. From the above content, it can be known that the embodiment of the present application selects a plurality of pre-trained multi-modal large language models as a first model, and specifies five target tasks (molecular image recognition on an input image, molecular label recognition on input text, classification on an input table, similarity recognition on two input data rows, and molecular label column recognition on an input table), customizes five instruction text templates, constructs five data sets, and fine-tunes the first model based on the five data sets and the corresponding five instruction text templates; after the model fine-tuning is completed, a PDF file of a small molecule drug patent input by a user is received, and image, text block and table extraction is performed on the patent file based on a PDF content extraction tool; then, the first image information sequence is screened based on the first instruction template and the first model, the first text block information sequence is screened based on the second instruction template and the first model, the first table information sequence is filtered based on the third instruction template and the first model, and the molecular activity table and the table without a header are spliced based on the fourth instruction template, the first model and the second table information sequence; then, a label-image correspondence table is constructed based on the first molecular image information sequence, the first molecular label information sequence, the first spliced table information sequence, the fifth instruction template and the first model, and a label-activity information correspondence table is constructed based on the fifth instruction template, the first model and the first spliced table information sequence; finally, a multi-modal table is merged based on the label-image correspondence table and the label-activity information correspondence table to obtain a corresponding multi-modal comprehensive information table, and the current user is fed back. The patent information mining mechanism provided by the embodiment of the present application can mine molecular activity data for small molecule drug patents, and based on the embodiment of the present application, the data mining efficiency is improved, the data mining cost is reduced, and the research and development efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 A processing method for mining patent information in combination with a large language model is provided for the embodiment one of the present application;

[0060] Figure 2 A module structure diagram of a processing device for mining patent information in combination with a large language model is provided for the embodiment two of the present application;

[0061] Figure 3 A structure schematic diagram of an electronic equipment is provided for the embodiment three of the present application. DETAILED DESCRIPTION

[0062] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the present application in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the protection scope of the present application.

[0063] The embodiment one of the present application provides a processing method for mining patent information combined with a large language model, which is shown in a schematic diagram. Figure 1 The embodiment one of the present application provides a processing method for mining patent information combined with a large language model, which is shown in a schematic diagram.

[0064] Step 1, a kind of multi-modal large language model completed pre-training is selected as the corresponding first model; and five target tasks are specified for the first model, which are the corresponding first, second, third, fourth and fifth tasks; and instruction text templates are customized for the first, second, third, fourth and fifth tasks, which are the corresponding first, second, third, fourth and fifth instruction templates; and five task data sets are constructed by collecting public small molecule drug patent literature information, which are the corresponding first, second, third, fourth and fifth data sets; and the first model is fine-tuned based on the first, second, third, fourth and fifth data sets and the corresponding first, second, third, fourth and fifth instruction templates;

[0065] Specifically, step 11, a kind of multi-modal large language model completed pre-training is selected as the corresponding first model;

[0066] Here, the first model of the embodiment of the present application at least includes GPT series model, LLaMA series model, Qwen series model, Wenxin series model;

[0067] Step 12, five target tasks are specified for the first model, which are the corresponding first, second, third, fourth and fifth tasks;

[0068] Here, the five target tasks specified for the first model of the embodiment of the present application are: the first task is to perform molecular image recognition on the input image; the second task is to perform molecular label recognition on the input text; the third task is to classify the input table; the fourth task is to perform similarity recognition on the input two data rows, and the total number of data fields of the two data rows is consistent; the fifth task is to perform molecular label column recognition on the input table;

[0069] Step 13, instruction text templates are customized for the first, second, third, fourth and fifth tasks, which are the corresponding first, second, third, fourth and fifth instruction templates;

[0070] Here, the first, second, third, fourth and fifth instruction templates given by the embodiments of the present application are each a formatted instruction text module; wherein:

[0071] 1) The first instruction template is used to take the input image as the corresponding current image; and through a piece of natural language description text, it instructs the first model to identify whether the current image is a molecular structure diagram to obtain the corresponding first identification result, and takes the first identification result as the corresponding first instruction feedback output; the first identification result includes yes and no;

[0072] 2) The second instruction template is used to take the input text as the corresponding current context; and through a piece of natural language description text, it instructs the first model to identify whether the current context contains molecular label information to obtain the corresponding second identification result, and when the second identification result is yes, it further performs molecular label information extraction processing on the current context to obtain the corresponding first molecular label extraction information, and when the second identification result is no, it sets the corresponding first molecular label extraction information to be empty, and the second identification result and the first molecular label extraction information obtained are used to form the corresponding second instruction feedback output; the second identification result includes yes and no;

[0073] 3) The third instruction template is used to take the input table as the corresponding current table; and through a piece of natural language description text, it instructs the first model to identify which one of the molecular activity table, the table without header and the other table the current table belongs to to obtain the corresponding third identification result, and takes the third identification result as the corresponding third instruction feedback output; the third identification result includes the molecular activity table, the table without header and the other table;

[0074] 4) The fourth instruction template is used to take the input two data rows as the corresponding current data row A and current data row B; and through a piece of natural language description text, it instructs the first model to perform field-by-field data similarity estimation on the current data row A and B according to the data field correspondence relationship to obtain the corresponding field similarity, and performs average similarity calculation based on all the obtained field similarities, and identifies whether all the field similarities and the average similarity are higher than a preset similarity threshold, if they are all higher than the similarity threshold, it sets the corresponding fourth identification result to be similar, if at least one of the field similarities or the average similarity is lower than the similarity threshold, it sets the corresponding fourth identification result to be dissimilar, and takes the fourth identification result as the corresponding fourth instruction feedback output; the fourth identification result includes similar and dissimilar;

[0075] 5) the fifth instruction template is used to take the input table as a corresponding current table; and through a piece of natural language description text, the first model is instructed to identify whether there is a molecular label column in the current table to obtain a corresponding fifth identification result, and when the fifth identification result is yes, the current table is further subjected to molecular label column information extraction processing to obtain corresponding first column extraction information, and when the fifth identification result is no, the corresponding first column extraction information is set to be empty, and the fifth identification result and the first column extraction information obtained are used to form a corresponding fifth instruction feedback output; the fifth identification result includes yes and no; when the first column extraction information is not empty, the first column extraction information is composed of one or more first column unit information; the first column unit information includes row index information and molecular label information;

[0076] Step 14, and five task data sets are constructed by collecting public small molecule drug patent literature information, denoted as corresponding first, second, third, fourth and fifth data sets;

[0077] Here, the data contents of the first, second, third, fourth and fifth data sets constructed by the big data collection of the embodiments of the present application are as follows:

[0078] 1) the first data set includes a plurality of first data records; the first data record includes a first training image and a first label identification result; the first training image includes a positive sample image and a negative sample image, the positive sample image is a visual image of a molecular structure, and the negative sample image is a visual image irrelevant to the molecular structure; the first label identification result includes yes and no, the first label identification result is yes when the first training image is the positive sample image, and the first label identification result is no when the first training image is the negative sample image;

[0079] 2) the second data set includes a plurality of second data records; the second data record includes a first training text, a second label identification result and a first label molecular label; the first training text includes a positive sample text and a negative sample text, the positive sample text is a text segment with molecular label information, and the negative sample text is a text segment without molecular label information; the second label identification result includes yes and no, the second label identification result is yes when the first training text is the positive sample text, the first label molecular label is the molecular label information in the current first training text, the second label identification result is no when the first training text is the negative sample text, and the first label molecular label is empty;

[0080] 3) the third data set comprises a plurality of third data records; the third data record comprises a first training table and a third label identification result; the table type of the first training table comprises a molecular activity data table type, a table header-free data table type, and any other data table type that is neither the molecular activity data table type nor the table header-free data table type; the field set of the molecular activity data table type should include a molecular label field and at least one type of molecular activity field, and the type of the molecular activity field at least includes catalytic activity, binding activity, signal transduction activity, toxicity activity, antioxidant activity, drug activity, photosensitivity, electrical conductivity, and thermal sensitivity; the data table corresponding to the table header-free data table type has no field set; the third label identification result includes a molecular activity table, a table header-free table, and other tables;

[0081] 4) the fourth data set comprises a plurality of fourth data records; the fourth data record comprises a first training data row, a second training data row, and a fourth label identification result; the number of data fields of the first and second training data rows is consistent; the fourth label identification result includes similarity and dissimilarity; when the data format of the two field data corresponding to each data field in the first and second training data rows meets the same format requirement, the fourth label identification result is similarity, and when the data format of the two field data corresponding to at least one data field in the first and second training data rows does not meet the same format requirement, the fourth label identification result is dissimilarity;

[0082] 5) the fifth data set comprises a plurality of fifth data records; the fifth data record comprises a second training table, a fifth label identification result, and first label column information; the second training table comprises a positive sample table and a negative sample table, the positive sample table is a data table with a molecular label field, and the negative sample table is a data table without a molecular label field; the fifth label identification result includes yes and no; when the second training table is the positive sample table, the fifth label identification result is yes, the first label column information is composed of one or more first label column unit information, and the first label column unit information includes row index information and molecular label information; when the second training table is the negative sample table, the fifth label identification result is no, and the first label column information is empty;

[0083] Step 15, and fine-tuning the first model based on the first, second, third, fourth, and fifth data sets and the corresponding first, second, third, fourth, and fifth instruction templates;

[0084] Specifically comprising: step 151, first task fine-tuning the first model based on the first data set and the corresponding first instruction template;

[0085] Specifically comprising: step 1511, dividing the first data set into two sub-data sets based on a preset first split ratio, denoted as a corresponding first training set and a first evaluation set;

[0086] Here, the first split ratio is a pre-set ratio parameter, for example, 8:2; the first training set and the first evaluation set are both composed of a plurality of first data records; the ratio of the total number of records of the first training set to the first evaluation set meets the first split ratio;

[0087] Step 1512, a round of traversal is performed on all first data records of the first training set; and in the round of traversal, the currently traversed first data record is taken as a corresponding current training record; and the first training image of the current training record is extracted as a corresponding current image; and the current image is input into the first instruction template to assemble a corresponding first task instruction; and the first task instruction is input into the first model to process the molecular image recognition task to obtain a corresponding first instruction feedback; and the corresponding first recognition result is extracted from the first instruction feedback; and the first recognition result and the first label recognition result of the current training record form a corresponding first prediction-label pair; and at the end of the round of traversal, all first prediction-label pairs obtained in the round of traversal form a corresponding first prediction-label pair set;

[0088] Step 1513, the first prediction-label pair set is brought into a pre-set first model loss function L M1 ; and based on a pre-set first model optimizer, the first model is fine-tuned in a direction of minimizing the first model loss function L M1 ;

[0089] Wherein, the first model loss function L M1 includes at least an L1 loss function, an L2 loss function, a log-likelihood loss function, a cross-entropy loss function and a cross-point loss function; and the first model optimizer includes at least an SGD optimizer, an Adam optimizer and an AdamW optimizer;

[0090] Step 1514, a round of traversal is performed on all first data records of the first evaluation set; and in the round of traversal, the currently traversed first data record is taken as a corresponding current evaluation record; and the first evaluation image of the current evaluation record is extracted as a corresponding current image; and the current image is input into the first instruction template to assemble a corresponding first task instruction; and the first task instruction is input into the first model to process the molecular image recognition task to obtain a corresponding first instruction feedback; and the corresponding first recognition result is extracted from the first instruction feedback; and the first recognition result and the first label recognition result of the current evaluation record form a corresponding second prediction-label pair; and at the end of the round of traversal, all second prediction-label pairs obtained in the round of traversal form a corresponding second prediction-label pair set;

[0091] Step 1515, based on the second prediction-label pair set, the accuracy, precision, recall and F1 score are calculated to obtain the corresponding first accuracy, first precision, first recall and first F1 score;

[0092] Step 1516, the first accuracy, the first precision, the first recall and the first F1 score are identified; if the first accuracy does not satisfy the preset first accuracy range or the first precision does not satisfy the preset first precision range or the first recall does not satisfy the preset first recall range or the first F1 score does not satisfy the preset first F1 score range, return to step 1512 for continue training; if the first accuracy satisfies the first accuracy range, the first precision satisfies the first precision range, the first recall satisfies the first recall range, and the first F1 score satisfies the first F1 score range, stop the current fine-tuning and confirm that the first task fine-tuning is completed;

[0093] Here, the first accuracy range, the first precision range, the first recall range and the first F1 score range are four set numerical ranges;

[0094] Step 152, after the first task fine-tuning is completed, the first model is fine-tuned for a second task based on the second data set and the corresponding second instruction template;

[0095] Specifically, it includes: step 1521, based on a preset second split ratio, the second data set is divided into two sub-data sets, denoted as a corresponding second training set and a second evaluation set;

[0096] Here, the second split ratio is a pre-set ratio parameter, for example, 8:2; the second training set and the second evaluation set are both composed of multiple second data records; the total number ratio of the second training set and the second evaluation set meets the second split ratio;

[0097] Step 1522, the first second data record of the second training set is taken as a corresponding current training record;

[0098] Step 1523, the first training text of the current training record is input into the second instruction template to assemble a corresponding instruction text to obtain a corresponding second task instruction; and the second task instruction is input into the first model to process the molecular label identification task to obtain a corresponding second instruction feedback; and the corresponding second identification result and the first molecular label extraction information are extracted from the second instruction feedback;

[0099] Step 1524, the second identification result and the second label identification result of the current training record are respectively converted into binary classification vectors to obtain corresponding prediction classification vectors y pre and label classification vectors y tag; and the first molecular label extraction information and the first label molecular label of the current training record are respectively subjected to word embedding coding processing to obtain corresponding predicted text vectors z pre and label text vectors z tag ; the predicted classification vector y pre , the label classification vector y tag , the predicted text vector z pre and the label text vector z tag are input into a preset second model loss function L M2 ; and based on a preset second model optimizer, the first model is fine-tuned in a direction of minimizing the second model loss function L M2 ;

[0100] Here, the second model loss function L M2 is:

[0101] L M2 =L M2-a (y pre ,y tag )+L M2-n (z pre ,z tag );

[0102] Wherein, the sub-loss function L M2-a includes at least a log-likelihood loss function, a cross-entropy loss function and a cross-point loss function, and the sub-loss function L M2-b includes at least an L1 loss function and an L2 loss function.

[0103] In addition, the second model optimizer includes at least an SGD optimizer, an Adam optimizer and an AdamW optimizer.

[0104] Step 1525, identify whether the current training record is the last second data record of the second training set; if yes, go to step 1526; if no, extract the next second data record of the second training set as a new current training record and return to step 1523.

[0105] Step 1526, a round of traversal is performed on all second data records of the second evaluation set; and during the round of traversal, a currently traversed second data record is taken as a corresponding current evaluation record; a first training text of the current evaluation record is input into the second instruction template to assemble a corresponding second task instruction; the second task instruction is input into the first model to perform molecular label identification task processing to obtain a corresponding second instruction feedback; a corresponding second identification result and first molecular label extraction information are extracted from the second instruction feedback; and the first molecular label extraction information and the first label molecular label of the current evaluation record form a corresponding third prediction-label pair; and when the round of traversal ends, all third prediction-label pairs obtained in the round of traversal form a corresponding third prediction-label pair set;

[0106] Step 1527, based on the third prediction-label pair set, a second accuracy, a second precision, a second recall and a second F1 score are calculated by calculating the accuracy, the precision, the recall and the F1 score;

[0107] Step 1528, the second accuracy, the second precision, the second recall and the second F1 score are identified; if the second accuracy does not satisfy a preset second accuracy range or the second precision does not satisfy a preset second precision range or the second recall does not satisfy a preset second recall range or the second F1 score does not satisfy a preset second F1 score range, return to step 1522 to continue training; if the second accuracy satisfies the second accuracy range, the second precision satisfies the second precision range, the second recall satisfies the second recall range, and the second F1 score satisfies the second F1 score range, stop the current fine-tuning and confirm that the second task fine-tuning is completed;

[0108] Here, the second accuracy range, the second precision range, the second recall range and the second F1 score range are four set numerical ranges;

[0109] Step 153, after the second task fine-tuning is completed, the first model is fine-tuned for a third task based on the third data set and a corresponding third instruction template;

[0110] Specifically, step 1531, the third data set is divided into two sub-data sets based on a preset third split ratio, denoted as a corresponding third training set and a third evaluation set;

[0111] Here, the third split ratio is a pre-set ratio parameter, for example, 8:2; the third training set and the third evaluation set are both composed of multiple third data records; the total number ratio of the third training set and the third evaluation set satisfies the third split ratio;

[0112] Step 1532, a round of traversal is performed on all third data records of the third training set; and in the round of traversal, a currently traversed third data record is taken as a corresponding current training record; and a first training table of the current training record is input into the third instruction template to assemble a corresponding third task instruction; and the third task instruction is input into the first model to perform table classification task processing to obtain a corresponding third instruction feedback; and a corresponding third recognition result is extracted from the third instruction feedback; and the third recognition result and a third label recognition result of the current training record form a corresponding fourth prediction-label pair; and when the round of traversal ends, all fourth prediction-label pairs obtained in the round of traversal form a corresponding fourth prediction-label pair set;

[0113] Step 1533, the fourth prediction-label pair set is brought into a preset third model loss function L M3 ; and based on a preset third model optimizer, the first model is fine-tuned in a direction of minimizing the third model loss function L M3 ;

[0114] The third model loss function L M3 includes at least an L1 loss function, an L2 loss function, a cross-entropy loss function and a cross-point loss function; and the third model optimizer includes at least an SGD optimizer, an Adam optimizer and an AdamW optimizer.

[0115] Step 1534, a round of traversal is performed on all third data records of the third evaluation set; and in the round of traversal, a currently traversed third data record is taken as a corresponding current evaluation record; and a first training table of the current evaluation record is input into the third instruction template to assemble a corresponding third task instruction; and the third task instruction is input into the first model to perform table classification task processing to obtain a corresponding third instruction feedback; and a corresponding third recognition result is extracted from the third instruction feedback; and the third recognition result and a third label recognition result of the current evaluation record form a corresponding fifth prediction-label pair; and when the round of traversal ends, all fifth prediction-label pairs obtained in the round of traversal form a corresponding fifth prediction-label pair set;

[0116] Step 1535, based on the fifth prediction-label pair set, accuracy, precision, recall and F1 score are calculated to obtain corresponding third accuracy, third precision, third recall and third F1 score;

[0117] Step 1536, the third accuracy, the third precision, the third recall and the third F1 score are identified; if the third accuracy does not satisfy the preset third accuracy range or the third precision does not satisfy the preset third precision range or the third recall does not satisfy the preset third recall range or the third F1 score does not satisfy the preset third F1 score range, return to step 1532 for continuous training; if the third accuracy satisfies the third accuracy range, the third precision satisfies the third precision range, the third recall satisfies the third recall range and the third F1 score satisfies the third F1 score range, stop the current fine-tuning and confirm that the third task fine-tuning is completed;

[0118] Here, the third accuracy range, the third precision range, the third recall range and the third F1 score range are four set numerical ranges;

[0119] Step 154, after the third task fine-tuning is completed, the first model is fine-tuned for a fourth task based on a fourth data set and a corresponding fourth instruction template;

[0120] Specifically, step 1541, the fourth data set is divided into two sub-data sets based on a preset fourth split ratio, denoted as a corresponding fourth training set and a fourth evaluation set;

[0121] Here, the fourth split ratio is a pre-set ratio parameter, for example, 8:2; the fourth training set and the fourth evaluation set are both composed of multiple fourth data records; the total number ratio of the records of the fourth training set and the fourth evaluation set satisfies the fourth split ratio;

[0122] Step 1542, all fourth data records in the fourth training set are iterated for one round; and in the current iteration process, the currently iterated fourth data record is taken as a corresponding current training record; and the first and second training data rows of the current training record are input into the fourth instruction template to assemble a corresponding fourth task instruction; and the fourth task instruction is input into the first model to process the data row similarity identification task to obtain a corresponding fourth instruction feedback; and the corresponding fourth identification result is extracted from the fourth instruction feedback; and the fourth identification result and the fourth label identification result of the current training record form a corresponding sixth prediction-label pair; and at the end of the current iteration, all the sixth prediction-label pairs obtained in the current iteration form a corresponding sixth prediction-label pair set;

[0123] Step 1543, the sixth prediction-label pair set is brought into a preset fourth model loss function L M4 ; and based on a preset fourth model optimizer, the first model is fine-tuned in the direction of minimizing the fourth model loss function L M4 ;

[0124] Wherein, the fourth model loss function LM4 at least including an L1 loss function, an L2 loss function, a log-likelihood loss function, a cross-entropy loss function, and a cross-point loss function; and the fourth model optimizer at least includes an SGD optimizer, an Adam optimizer, an AdamW optimizer;

[0125] Step 1544, a round of traversal is performed on all fourth data records of the fourth evaluation set; and in the round of traversal, the fourth data record currently traversed is taken as a corresponding current evaluation record; and the first and second training data rows of the current evaluation record are input into the fourth instruction template to assemble a corresponding fourth task instruction; and the fourth task instruction is input into the first model to perform data row similarity identification task processing to obtain a corresponding fourth instruction feedback; and the corresponding fourth identification result is extracted from the fourth instruction feedback; and the fourth identification result and the fourth label identification result of the current evaluation record form a corresponding seventh prediction-label pair; and at the end of the round of traversal, all the seventh prediction-label pairs obtained in the round of traversal form a corresponding seventh prediction-label pair set;

[0126] Step 1545, the seventh prediction-label pair set is used to calculate the fourth accuracy, the fourth precision, the fourth recall, and the fourth F1 score;

[0127] Step 1546, the fourth accuracy, the fourth precision, the fourth recall, and the fourth F1 score are identified; if the fourth accuracy does not satisfy the preset fourth accuracy range, or the fourth precision does not satisfy the preset fourth precision range, or the fourth recall does not satisfy the preset fourth recall range, or the fourth F1 score does not satisfy the preset fourth F1 score range, then step 1542 is returned to continue training; if the fourth accuracy satisfies the fourth accuracy range, the fourth precision satisfies the fourth precision range, the fourth recall satisfies the fourth recall range, and the fourth F1 score satisfies the fourth F1 score range, then the current fine-tuning is stopped and it is confirmed that the fourth task fine-tuning is completed;

[0128] Here, the fourth accuracy range, the fourth precision range, the fourth recall range, and the fourth F1 score range are four set numerical ranges;

[0129] Step 155, after the fourth task fine-tuning is completed, the first model is fine-tuned for a fifth task based on a fifth data set and a corresponding fifth instruction template;

[0130] Specifically, step 1551, the fifth data set is divided into two sub-data sets based on a preset fifth segmentation ratio, which are taken as a corresponding fifth training set and a fifth evaluation set;

[0131] Here, the fifth split ratio is a pre-set ratio parameter, for example, 8:2; the fifth training set and the fifth evaluation set are both composed of multiple fifth data records; the ratio of the total number of records of the fifth training set to the fifth evaluation set meets the fifth split ratio;

[0132] Step 1552, the first fifth data record of the fifth training set is taken as the corresponding current training record;

[0133] Step 1553, the second training table of the current training record is input into the fifth instruction template for corresponding instruction text assembly to obtain a corresponding fifth task instruction; and the fifth task instruction is input into the first model for molecular label column identification task processing to obtain a corresponding fifth instruction feedback; and the corresponding fifth identification result and the first column extraction information are extracted from the fifth instruction feedback;

[0134] Step 1554, the fifth identification result and the fifth label identification result of the current training record are respectively subjected to binary classification vector conversion to obtain corresponding prediction classification vectors x pre and label classification vectors x tag ;

[0135] Step 1555, the total number of first column unit information of the first column extraction information is counted to obtain a corresponding first total number, and the total number of first label column unit information of the first label column information of the current training record is counted to obtain a corresponding second total number; the maximum value of the first total number and the second total number is taken as a corresponding current maximum value N; when there is a total number less than the current maximum value N in the first total number and the second total number, the first column extraction information or the first label column information corresponding to the first total number or the second total number less than the current maximum value N is taken as a corresponding current column information, and the total number of column unit information of the current column information is supplemented to N based on the way of supplementing all empty column unit information to the current column information; the row index information of each first column unit information of the current first column extraction information is subjected to one-hot encoding to obtain a corresponding row index encoding vector y pre,i , and the molecular label information of each first column unit information is subjected to word embedding encoding to obtain a corresponding label encoding vector z pre,i , 1≤unit index i≤N; the row index information of each first label column unit information of the current first label column information is subjected to one-hot encoding to obtain a corresponding row index encoding vector y tag,i , and the molecular label information of each first label column unit information is subjected to word embedding encoding to obtain a corresponding label encoding vector z tag,i ;

[0136] Step 1556, the prediction classification vector x pre , the label classification vector x tag , the N row index encoding vectors y pre,i , the N row index encoding vectors ytag,i N label encoding vectors z pre,i N label encoding vectors z tag,i into a preset fifth model loss function L M5 ; and based on a preset fifth model optimizer, a round of fine-tuning is performed on the first model in a direction of minimizing the fifth model loss function L M5 ;

[0137] Here, the fifth model loss function L M5 is:

[0138]

[0139] Wherein, the sub-loss function L M5-a includes at least a log-likelihood loss function, a cross-entropy loss function and a cross-point loss function, the sub-loss function L M5-b includes at least an L1 loss function and an L2 loss function, and the sub-loss function L M5-c includes at least an L1 loss function and an L2 loss function.

[0140] In addition, the fifth model optimizer includes at least an SGD optimizer, an Adam optimizer and an AdamW optimizer.

[0141] Step 1557, identify whether the current training record is the last fifth data record of the fifth training set; if so, go to step 1558; if not, extract the next fifth data record of the fifth training set as a new current training record and return to step 1553;

[0142] Step 1558, a round of traversal is performed on all fifth data records of the fifth evaluation set; and in this round of traversal, the currently traversed fifth data record is taken as a corresponding current evaluation record; and the second training table of the current evaluation record is input into the fifth instruction template for corresponding instruction text assembly to obtain a corresponding fifth task instruction; and the fifth task instruction is input into the first model for molecular label column identification task processing to obtain a corresponding fifth instruction feedback; and the corresponding fifth identification result and the first column extraction information are extracted from the fifth instruction feedback; and the first column extraction information and the first label column information of the current evaluation record are combined to form a corresponding eighth prediction-label pair; and at the end of this round of traversal, all eighth prediction-label pairs obtained in this round of traversal are combined to form a corresponding eighth prediction-label pair set.

[0143] Step 1559, based on the eighth prediction-label pair set, the accuracy, precision, recall and F1 score are calculated to obtain the corresponding fifth accuracy, fifth precision, fifth recall and fifth F1 score; and the fifth accuracy, fifth precision, fifth recall and fifth F1 score are identified; if the fifth accuracy does not satisfy the preset fifth accuracy range or the fifth precision does not satisfy the preset fifth precision range or the fifth recall does not satisfy the preset fifth recall range or the fifth F1 score does not satisfy the preset fifth F1 score range, return to step 1552 for continuous training; if the fifth accuracy satisfies the fifth accuracy range, the fifth precision satisfies the fifth precision range, the fifth recall satisfies the fifth recall range and the fifth F1 score satisfies the fifth F1 score range, stop the current fine-tuning and confirm that the fifth task fine-tuning is completed;

[0144] Here, the fifth accuracy range, the fifth precision range, the fifth recall range and the fifth F1 score range are four set numerical ranges;

[0145] Step 156, after the fifth task fine-tuning is completed, it is confirmed that the first model fine-tuning is completed.

[0146] Step 2, after the first model fine-tuning is completed, a PDF file of a small molecule drug patent input by a user is received as a corresponding first patent file; and based on a preset PDF content extraction tool, image, text block and table extraction are performed on the first patent file to obtain a corresponding first image information sequence, a first text block information sequence and a first table information sequence.

[0147] Here, the PDF content extraction tool is a kind of mature PDF file analysis tool / library, and commonly used ones include PyPDF, pdfminer.six, PyMuPDF, pdfplumber, pypdfium2, Unstructured, Tabula, Camelot, etc.

[0148] The data of the first image information sequence, the first text block information sequence and the first table information sequence extracted by the PDF content extraction tool in the embodiment of the application are as follows: 1) the first image information sequence includes a plurality of first image information; the first image information includes a first identifier, a first page code, a first upper left corner row and column coordinate, a first upper right corner row and column coordinate, a first lower right corner row and column coordinate, a first lower left corner row and column coordinate and a first image; the first image is a visual diagram of a molecular structure. 2) the first text block information sequence includes a plurality of first text block information; the first text block information includes a second identifier, a second page code, a second upper left corner row and column coordinate, a second upper right corner row and column coordinate, a second lower right corner row and column coordinate, a second lower left corner row and column coordinate and a first block text. 3) the first table information sequence includes a plurality of first table information; the first table information includes a third identifier, a third page code, a third upper left corner row and column coordinate, a third upper right corner row and column coordinate, a third lower right corner row and column coordinate, a third lower left corner row and column coordinate and a first table; the first table includes one or more first table rows; the first table row includes a plurality of first row cell texts.

[0149] Step 3, performing molecular image screening processing on the first image information sequence based on the first instruction template and the first model to obtain a corresponding first molecular image information sequence; performing molecular label screening processing on the first text block information sequence based on the second instruction template and the first model to obtain a corresponding first molecular label information sequence; performing other table filtering processing on the first table information sequence based on the third instruction template and the first model to obtain a corresponding second table information sequence; and performing molecular activity table and table without header splicing processing on the fourth instruction template, the first model and the second table information sequence to obtain a corresponding first splicing table information sequence.

[0150] Specifically includes: step 31, performing molecular image screening processing on the first image information sequence based on the first instruction template and the first model to obtain a corresponding first molecular image information sequence.

[0151] Specifically includes: performing one round of iteration on all first images of the first image information sequence; and in this round of iteration, inputting the first image currently iterated as a corresponding current image into the first instruction template to assemble a corresponding current task instruction; inputting the current task instruction into the first model to perform a molecular image recognition task to obtain a corresponding first instruction feedback; taking the first recognition result in the first instruction feedback as a corresponding current recognition result; when the current recognition result is no, deleting the first image information corresponding to the current image from the first image information sequence; and when this round of iteration ends, taking the latest first image information sequence as a corresponding first molecular image information sequence.

[0152] Step 32, performing molecular label screening processing on the first text block information sequence based on the second instruction template and the first model to obtain a corresponding first molecular label information sequence;

[0153] Specifically including: performing one round of traversal on all first block texts of the first text block information sequence; and in the current round of traversal, taking the first block text currently traversed as a corresponding current text, inputting the second instruction template to obtain a corresponding current task instruction through corresponding instruction text assembly; inputting the current task instruction into the first model to obtain a corresponding second instruction feedback through molecular label identification task processing; taking the second identification result of the second instruction feedback and the first molecular label extraction information as a corresponding current identification result and a current molecular label; and identifying the current identification result; if the current identification result is yes, resetting the text content of the first block text corresponding to the current text in the first text block information sequence as the corresponding current molecular label; if the current identification result is no, deleting the first text block information corresponding to the current text from the first text block information sequence; and at the end of the current round of traversal, taking the latest first text block information sequence as the corresponding first molecular label information sequence;

[0154] Step 33, performing other table filtering processing on the first table information sequence based on the third instruction template and the first model to obtain a corresponding second table information sequence;

[0155] Specifically including: performing one round of traversal on all first tables of the first table information sequence; and in the current round of traversal, taking the first table currently traversed as a corresponding current table, inputting the third instruction template to obtain a corresponding current task instruction through corresponding instruction text assembly; inputting the current task instruction into the first model to obtain a corresponding third instruction feedback through table classification task processing; taking the third identification result of the third instruction as a corresponding current identification result; and identifying the current identification result; if the current identification result is a molecular activity table, marking the first table information corresponding to the current table as a corresponding first type of table information; if the current identification result is a table without a header, marking the first table information corresponding to the current table as a corresponding second type of table information; if the current identification result is other table, deleting the first table information corresponding to the current table from the first table information sequence; and at the end of the current round of traversal, taking the latest first table information sequence as the corresponding second table information sequence;

[0156] Step 34, performing molecular activity table and table without header splicing processing based on the fourth instruction template, the first model and the second table information sequence to obtain a corresponding first spliced table information sequence;

[0157] The first splicing table information sequence includes a plurality of first splicing table information; the first splicing table information includes a fourth identifier, a front page page number, a rear page page number, a front page upper left corner row and column coordinate, a front page upper right corner row and column coordinate, a front page lower right corner row and column coordinate, a front page lower left corner row and column coordinate, a rear page upper left corner row and column coordinate, a rear page upper right corner row and column coordinate, a rear page lower right corner row and column coordinate, a rear page lower left corner row and column coordinate, and a first splicing table; the first splicing table includes one or more second table rows; the second table row includes a plurality of second row cell texts.

[0158] Specifically, the step 341 includes: initializing an empty sequence as a corresponding third table information sequence; and taking the first first table information of the second table information sequence as a corresponding current table information.

[0159] The step 342 includes: identifying whether the current table information is marked as a first type of table information; if the current table information is marked as the first type of table information, identifying whether a row coordinate of a third lower left corner row and column coordinate of the current table information is a page bottom row; if yes, turning to the step 343; if no, turning to the step 344; if the current table information is not marked as the first type of table information, turning to the step 345.

[0160] Step 343, add 1 to the third page code of the current table information to obtain the corresponding back page code; match each third page code and back page code in the second table information sequence, and take the second table information whose row coordinate of the third upper left corner column coordinate is the top row of the page as the corresponding candidate table information; identify whether the candidate table information is empty; if the candidate table information is empty, go to step 344; if the candidate table information is not empty, take the first table of the candidate table information as the corresponding back table, and take the first table of the current table information as the corresponding front table, count the total number of the first row cell text of any first table row of the front table to obtain the corresponding first cell total number, count the total number of the first row cell text of any first table row of the back table to obtain the corresponding second cell total number, and identify whether the first and second cell total numbers are equal; if the first and second cell total numbers are not equal, go to step 344; if the first and second cell total numbers are equal, take the second first table row of the front table and the first first table row of the back table as the corresponding first and second data rows, input the fourth instruction template to obtain the corresponding current task instruction, input the current task instruction into the first model to process the data row similarity identification task to obtain the corresponding fourth instruction feedback, take the fourth identification result in the fourth instruction feedback as the corresponding current identification result, and identify the current identification result; if the current identification result is not similar, go to step 344; if the current identification result is similar, add a first spliced table information with empty content to the third table information sequence as the corresponding current spliced table information, assign a unique identification to the current spliced table information as the corresponding current spliced table identification, add the back table to the end of the front table to obtain the corresponding current spliced table, set the fourth identification of the current spliced table information as the corresponding current spliced table identification, set the front and back page page codes of the current spliced table information as the third page codes of the front and back tables, set the front page upper left corner column coordinate, front page upper right corner column coordinate, front page lower right corner column coordinate and front page lower left corner column coordinate of the current spliced table information as the third upper left corner column coordinate, third right upper corner column coordinate, third right lower corner column coordinate and third left lower corner column coordinate of the front table, set the back page upper left corner column coordinate, back page upper right corner column coordinate, back page right lower corner column coordinate and back page left lower corner column coordinate of the current spliced table information as the third upper left corner column coordinate, third right upper corner column coordinate, third right lower corner column coordinate and third left lower corner column coordinate of the back table, set the first spliced table of the current spliced table information as the corresponding current spliced table, and go to step 345 after the setting of the current spliced table information is completed.

[0161] Step 344, record the first table of the current table information as the corresponding current table; and add a first splicing table information with empty content in the third table information sequence as the corresponding current splicing table information; and assign a unique identifier to the current splicing table information as the corresponding current splicing table identifier; and set the fourth identifier of the current splicing table information as the corresponding current splicing table identifier; and set the front page page number of the current splicing table information as the third page number corresponding to the current table; and set the back page page number of the current splicing table information as empty; and set the front page top-left corner row and column coordinates, the front page top-right corner row and column coordinates, the front page bottom-right corner row and column coordinates, and the front page bottom-left corner row and column coordinates of the current splicing table information as the third top-left corner row and column coordinates, the third top-right corner row and column coordinates, the third bottom-right corner row and column coordinates, and the third bottom-left corner row and column coordinates corresponding to the current table; and set the back page top-left corner row and column coordinates, the back page top-right corner row and column coordinates, the back page bottom-right corner row and column coordinates, and the back page bottom-left corner row and column coordinates of the current splicing table information as empty; and set the first splicing table of the current splicing table information as the corresponding current table;

[0162] Step 345, identify whether the current table information is the last first table information of the second table information sequence; if not, extract the next first table information of the second table information sequence as the new current table information and return to step 342; if yes, output the latest first splicing table information sequence as the current splicing processing result.

[0163] Step 4, construct a label-image correspondence table based on the first molecular image information sequence, the first molecular label information sequence, the first splicing table information sequence, the fifth instruction template, and the first model; and construct a label-activity information correspondence table based on the fifth instruction template, the first model, and the first splicing table information sequence; and perform multi-modal table merging based on the label-image correspondence table and the label-activity information correspondence table to obtain the corresponding multi-modal comprehensive information table;

[0164] Specifically includes: step 41, constructing a label-image correspondence table based on the first molecular image information sequence, the first molecular label information sequence, the first splicing table information sequence, the fifth instruction template, and the first model;

[0165] Specifically includes: step 411, taking the first image of the first molecular image information sequence as the corresponding current molecular image;

[0166] Step 412, taking the first page number corresponding to the current molecular image as the corresponding current image page number, and recording the area covered by the first top-left corner row and column coordinates, the first top-right corner row and column coordinates, the first bottom-right corner row and column coordinates, and the first bottom-left corner row and column coordinates corresponding to the current molecular image as the corresponding current image area;

[0167] Step 413, record the first splicing table information in which the leading page number or the trailing page number matches the current image page number as the corresponding first same-page table information; and confirm the coverage range of the current first same-page table information on the current image page based on the upper left corner row and column coordinates of the leading / trailing page, the upper right corner row and column coordinates of the leading / trailing page, the lower right corner row and column coordinates of the leading / trailing page, and the lower left corner row and column coordinates of the leading / trailing page to obtain the corresponding first same-page table region; and when the number of the first same-page table regions is not zero, confirm whether there is an intersection between one of the first same-page table regions and the current image region to obtain the corresponding first confirmation result; and identify the first confirmation result; if the first confirmation result is not, go to step 414; if the first confirmation result is, take the first same-page table information corresponding to the first same-page table region which has an intersection with the current image region as the corresponding current same-page table information, and go to step 415;

[0168] Step 414, record the first text block information in which the second page number matches the current image page number in the first molecular label information sequence as the corresponding first same-page text block information; and confirm the coverage range of the current first same-page text block information on the current image page based on the second upper left corner row and column coordinates, the second upper right corner row and column coordinates, the second lower right corner row and column coordinates, and the second lower left corner row and column coordinates of each first same-page text block information to obtain the corresponding first same-page text block region; and take the first block text of the first text block information corresponding to the first same-page text block region closest to the current image region as the corresponding current molecular label; and go to step 416;

[0169] Step 415, input the first spliced table of the current same-page table information into the fifth instruction template to assemble the corresponding instruction text to obtain the corresponding current task instruction; input the current task instruction into the first model to perform the molecular label column identification task processing to obtain the corresponding fifth instruction feedback; extract the corresponding first column extraction information from the fifth instruction feedback; record the full table coverage area of the single page or double page of the current same-page table information as the corresponding first table area; record the area covered by the first top-left corner row and column coordinates, the first top-right corner row and column coordinates, the first bottom-right corner row and column coordinates and the first bottom-left corner row and column coordinates of each first molecular image in the first molecular image information sequence as the corresponding first image area; record the first image area having intersection with the first table area as the corresponding first intra-table image area, and sort all the first intra-table image areas in the order from front to back to form the corresponding first intra-table image area sequence; take the sequence index of the first intra-table image area corresponding to the current image area in the first intra-table image area sequence as the corresponding current row index; extract the molecular label information of the first column unit information having the matching row index information in the first column extraction information and the current row index as the corresponding current molecular label;

[0170] Step 416, take the current molecular image as a corresponding first molecular image field; take the current molecular label as a corresponding first molecular label field; and form a corresponding first correspondence record by the obtained first molecular label field and first molecular image field;

[0171] Step 417, identify whether the current molecular image is the last first molecular image in the first molecular image information sequence; if not, extract the next first molecular image in the first molecular image information sequence as a new current molecular image, and return to step 412; if yes, form a corresponding label-image correspondence table by all the obtained first correspondence records;

[0172] Here, the label-image correspondence table of the embodiment of the application includes a plurality of first correspondence records; the first correspondence record includes a first molecular label field and a first molecular image field;

[0173] Step 42, and construct a label-activity information correspondence table based on the fifth instruction template, the first model and the first spliced table information sequence;

[0174] Specifically comprising: step 421, extracting the first second table row of all first splicing tables of the first splicing table information sequence to form a corresponding second table row set; and extracting all second row unit texts of the second table row set to form a corresponding second row unit text set; and performing unit text deduplication processing on the second row unit text set; and performing a round of iteration on all second row unit texts of the second row unit text set after deduplication; and in this round of iteration, taking the second row unit text currently iterated as a corresponding current field name; and identifying whether the current field name is a type of molecular activity name based on a preset molecular activity keyword dictionary to obtain a corresponding activity name identification result; and when the activity name identification result is yes, marking the current field name as a corresponding molecular activity field name; and at the end of this round of iteration, forming a corresponding molecular activity field name set from all obtained molecular activity field names;

[0175] Step 422, performing a round of iteration on all first splicing tables of the first splicing table information sequence; and in this round of iteration, taking the first splicing table currently iterated as a corresponding current splicing table; and inputting the current splicing table into the fifth instruction template to assemble a corresponding current task instruction; and inputting the current task instruction into the first model to perform a molecular label column identification task to obtain a corresponding fifth instruction feedback; and extracting corresponding first column extraction information from the fifth instruction feedback; and taking the first second table row of the current splicing table as a corresponding current header row; and taking the second row unit text in the current header row corresponding to the first column extraction information as a corresponding molecular label field name; and performing a round of polling on all second row unit texts in the current header row except the molecular label field name; and in this round of polling, taking the second row unit text currently polled as a corresponding current field name, and identifying whether the current field name is included in the molecular activity field name set, if yes, taking the second row unit text currently polled as a corresponding molecular activity field name, if not, taking the arrangement index of the second row unit text corresponding to the current field name in the current header row as a corresponding to-be-deleted column index; and at the end of this round of polling, deleting the second row unit text corresponding to all to-be-deleted column indexes in all second table rows of the current splicing table;

[0176] Step 423, after the end of the round of iteration on all first splicing tables of the first splicing table information sequence, performing full connection and merging on all field contents of all first splicing tables with the molecular label field name as the connection key field name to obtain a new table with a molecular label field and all molecular activity fields as a corresponding label-activity information corresponding relationship table;

[0177] Here, the data structure of the label-activity information corresponding table obtained by the embodiment of the application is: the label-activity information corresponding table comprises a plurality of second corresponding records; each second corresponding record comprises a second molecular label field and a plurality of first molecular activity fields;

[0178] Step 43, and based on the label-image corresponding table and the label-activity information corresponding table, multi-modal table merging is performed to obtain a corresponding multi-modal comprehensive information table;

[0179] Specifically, the label-image corresponding table and the label-activity information corresponding table are fully connected and merged with the molecular label field as the connection key field to obtain a new table with the molecular label field, the molecular image field and all the molecular activity fields as the corresponding multi-modal comprehensive information table.

[0180] Here, the data structure of the multi-modal comprehensive information table obtained by the embodiment of the application is: the multi-modal comprehensive information table comprises a plurality of third corresponding records; each third corresponding record comprises a third molecular label field, a second molecular image field and a plurality of second molecular activity fields.

[0181] Step 5, the corresponding first mining report is composed of the first patent file and the multi-modal comprehensive information table, and is fed back to the current user.

[0182] Figure 2 A module structure diagram of a processing device for mining patent information combined with a large language model is provided for the second embodiment of the application. The device is a terminal device or a server for implementing the foregoing method embodiments, and can also be a device capable of enabling the foregoing terminal device or server to implement the foregoing method embodiments, for example, the device can be a device or a chip system of the foregoing terminal device or server. As shown in the figure, the device comprises: a large language model fine-tuning module 201, a patent information receiving module 202, a patent information noise reduction module 203, a patent information mining module 204 and a mining report feedback module 205. Figure 2

[0183] ​The large language model fine-tuning module 201 is configured to select a plurality of pre-trained multi-modal large language models as corresponding first models; assign five target tasks to the first models, denoted as first, second, third, fourth and fifth tasks; customize instruction text templates for the first, second, third, fourth and fifth tasks, denoted as first, second, third, fourth and fifth instruction templates; construct five task data sets based on the collected public small molecule drug patent literature information, denoted as first, second, third, fourth and fifth data sets; and fine-tune the first models based on the first, second, third, fourth and fifth data sets and the corresponding first, second, third, fourth and fifth instruction templates; the first task is to perform molecular image recognition on the input image; the second task is to perform molecular label recognition on the input text; the third task is to classify the input table; the fourth task is to perform similarity recognition on the input two data rows, and the total number of data fields of the two data rows is consistent; and the fifth task is to perform molecular label column recognition on the input table.

[0184] The patent information receiving module 202 is configured to receive a PDF file of a small molecule drug patent input by a user as a corresponding first patent file after the fine-tuning of the first model is completed; and perform image, text block and table extraction on the first patent file based on a pre-set PDF content extraction tool to obtain corresponding first image information sequence, first text block information sequence and first table information sequence.

[0185] The patent information denoising module 203 is configured to perform molecular image screening processing on the first image information sequence based on the first instruction template and the first model to obtain a corresponding first molecular image information sequence; perform molecular label screening processing on the first text block information sequence based on the second instruction template and the first model to obtain a corresponding first molecular label information sequence; perform other table filtering processing on the first table information sequence based on the third instruction template and the first model to obtain a corresponding second table information sequence; and perform molecular activity table and no header table splicing processing based on the fourth instruction template, the first model and the second table information sequence to obtain a corresponding first spliced table information sequence.

[0186] The patent information mining module 204 is configured to construct a label-image correspondence table based on the first molecular image information sequence, the first molecular label information sequence, the first spliced table information sequence, the fifth instruction template and the first model; construct a label-activity information correspondence table based on the fifth instruction template, the first model and the first spliced table information sequence; and perform multi-modal table merging based on the label-image correspondence table and the label-activity information correspondence table to obtain a corresponding multi-modal comprehensive information table.

[0187] The mining report feedback module 205 is configured to feed back a corresponding first mining report composed of the first patent file and the multi-modal comprehensive information table to the current user.

[0188] The processing device for mining patent information in combination with a large language model provided by the embodiment of the present application can execute the method steps in the method embodiments described above, and has similar implementation principles and technical effects, which will not be described here again.

[0189] It should be noted that the division of each module of the above device is only a logical division of functions, and all or part of the modules can be integrated into one physical entity or physically separated when actually implemented. Moreover, all of the modules can be implemented in the form of software invoked by a processing element; all of the modules can also be implemented in the form of hardware; some modules can be implemented in the form of software invoked by a processing element, and some modules can be implemented in the form of hardware. For example, the large language model fine-tuning module can be a separately established processing element, or can be integrated into a certain chip of the above device, in addition, the large language model fine-tuning module can also be stored in the form of program code in the memory of the above device, and the function of the above determination module can be invoked and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of the modules can be integrated together or independently implemented. The processing element described herein can be an integrated circuit having a signal processing capability. In the implementation process, each step of the above method or each module can be completed by the integrated logic circuit of hardware or the instruction in the form of software in the processing element.

[0190] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of program code invoked by a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of invoking program code. For another example, the modules can be integrated together to implement in the form of a system on a chip (SOC).

[0191] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed by a computer, all or part of the computer program instructions generate the processes or functions described in the foregoing method embodiments. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (solid state disk, SSD)) and the like.

[0192] Figure 3 A structural schematic diagram of an electronic device is provided for the third embodiment of the present application. The electronic device can be a terminal device or a server for implementing the method of the foregoing embodiments, or a terminal device or a server connected with the foregoing terminal device or server for implementing the method of the foregoing embodiments. As shown in the figure, the electronic device can include a processor 301 (such as a CPU), a memory 302, a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiving action of the transceiver 303. The memory 302 can store various instructions for completing various processing functions and implementing the processing steps described in the foregoing embodiment methods. Preferably, the electronic device related to the embodiments of the present application further includes a power supply 304, a system bus 305 and a communication port 306. The system bus 305 is used to realize the communication connection between elements. The communication port 306 is used for the connection and communication between the electronic device and other external devices. Figure 3

[0193] In Figure 3 ​The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as a client, a read-write library and a read-only library). The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory, such as at least one disk memory.

[0194] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; can also be a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0195] It should be noted that the embodiments of the present application also provide a computer readable storage medium, which stores instructions, and when the instructions run on a computer, the computer executes the method and process provided in the above embodiments.

[0196] The embodiment of the present application provides a processing method and device for mining patent information combined with a large language model, electronic equipment and a computer readable storage medium. From the above content, it can be known that the embodiment of the present application selects a kind of multi-modal large language model completed pre-training as a first model, and specifies five target tasks (molecular image recognition on input image, molecular label recognition on input text, classification on input table, similarity recognition on two input data rows, and molecular label column recognition on input table), customizes five instruction text templates, constructs five data sets, and fine-tunes the first model based on the five data sets and the corresponding five instruction text templates; after the model fine-tuning is completed, a PDF file of a small molecule drug patent input by a user is received, and image, text block and table extraction is performed on the patent file based on a PDF content extraction tool; then, the first image information sequence is screened based on the first instruction template and the first model, the first text block information sequence is screened based on the second instruction template and the first model, other table filtering is performed on the first table information sequence based on the third instruction template and the first model, and the molecular activity table and the table without header are spliced based on the fourth instruction template, the first model and the second table information sequence; then, a label-image corresponding relationship table is constructed based on the first molecular image information sequence, the first molecular label information sequence, the first spliced table information sequence, the fifth instruction template and the first model, and a label-activity information corresponding relationship table is constructed based on the fifth instruction template, the first model and the first spliced table information sequence; finally, the multi-modal table is merged based on the label-image corresponding relationship table and the label-activity information corresponding relationship table to obtain a corresponding multi-modal comprehensive information table, and the current user is fed back. The patent information mining mechanism given by the embodiment of the present application can mine molecular activity data for small molecule drug patents, and based on the embodiment of the present application, the data mining efficiency is improved, the data mining cost is reduced, and the research and development efficiency is improved.

[0197] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0198] The above specific embodiments further illustrate the purpose, technical solutions and advantages of the present application, and it should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application, and any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A processing method for mining patent information in combination with a large language model, characterized in that, The method comprises: selecting a plurality of pre-trained multi-modal large language models as corresponding first models; and assigning five target tasks to the first model, denoted as the corresponding first, second, third, fourth and fifth tasks; and customizing instruction text templates for the first, second, third, fourth and fifth tasks, denoted as the corresponding first, second, third, fourth and fifth instruction templates; and constructing five task data sets by collecting public small molecule drug patent literature information, denoted as the corresponding first, second, third, fourth and fifth data sets; and fine-tuning the first model based on the first, second, third, fourth and fifth data sets and the corresponding first, second, third, fourth and fifth instruction templates; the first task is to perform molecular image recognition on the input image; the second task is to perform molecular label recognition on the input text; the third task is to classify the input table; the fourth task is to perform similarity recognition on the input two data rows, and the total number of data fields of the two data rows is consistent; the fifth task is to perform molecular label column recognition on the input table; After the first model fine-tuning is completed, a user input small molecule drug patent PDF file is received as a corresponding first patent file; and based on a pre-set PDF content extraction tool, image, text block and table extraction is performed on the first patent file to obtain a corresponding first image information sequence, a first text block information sequence and a first table information sequence; based on the first instruction template and the first model, the first image information sequence is subjected to molecular image screening processing to obtain a corresponding first molecular image information sequence; and based on the second instruction template and the first model, the first text block information sequence is subjected to molecular label screening processing to obtain a corresponding first molecular label information sequence; and based on the third instruction template and the first model, the first table information sequence is subjected to other table filtering processing to obtain a corresponding second table information sequence; and based on the fourth instruction template, the first model and the second table information sequence, molecular activity table and table without header splicing processing is performed to obtain a corresponding first splicing table information sequence; wherein the fourth instruction template and the first model are used to perform data row similarity recognition on the second row of the molecular activity table and the first row of the table without header when the total number of row unit texts of the molecular activity table at the bottom of the front page is equal to the total number of row unit texts of the table without header at the top of the back page; based on the first molecular image information sequence, the first molecular label information sequence, the first splicing table information sequence, the fifth instruction template and the first model, a label-image correspondence table is constructed; and based on the fifth instruction template, the first model and the first splicing table information sequence, a label-activity information correspondence table is constructed; and based on the label-image correspondence table and the label-activity information correspondence table, multi-modal table merging is performed to obtain a corresponding multi-modal comprehensive information table; the first patent file and the multi-modal comprehensive information table constitute a corresponding first mining report to feed back to the current user.

2. The processing method for mining patent information by combining a large language model according to claim 1, characterized in that, the first model at least comprises a GPT series model, an LLaMA series model, a Qwen series model, and a Wenxin series model; the first, second, third, fourth, and fifth instruction templates are each a formatted instruction text module; the first instruction template is used to input an image as a corresponding current image; and through a piece of natural language description text, the first model is instructed to identify whether the current image is a molecular structure diagram to obtain a corresponding first identification result, and the first identification result is fed back as a corresponding first instruction output; the first identification result includes yes and no; the second instruction template is used to input text as a corresponding current context; and through a piece of natural language description text, the first model is instructed to identify whether the current context contains molecular label information to obtain a corresponding second identification result, and when the second identification result is yes, the current context is further processed to extract molecular label information to obtain corresponding first molecular label extraction information, and when the second identification result is no, the first molecular label extraction information is set to empty, and the second identification result and the first molecular label extraction information are combined to form a corresponding second instruction feedback output; the second identification result includes yes and no; the third instruction template is used to input a table as a corresponding current table; and through a piece of natural language description text, the first model is instructed to identify which of the molecular activity table, the table without headers, and other tables the current table belongs to to obtain a corresponding third identification result, and the third identification result is fed back as a corresponding third instruction output; the third identification result includes a molecular activity table, a table without headers, and other tables; the fourth instruction template is used to input two data rows as corresponding current data row A and current data row B; and through a piece of natural language description text, the first model is instructed to estimate the data similarity of the current data row A and B field by field according to the data field correspondence relationship to obtain a corresponding field similarity, and based on the obtained all field similarities, an average similarity is calculated, and whether all the field similarities and the average similarity are higher than a preset similarity threshold is identified, if all are higher than the similarity threshold, the fourth identification result is set to similar, if at least one of the field similarities or the average similarity is lower than the similarity threshold, the fourth identification result is set to dissimilar, and the fourth identification result is fed back as a corresponding fourth instruction output; the fourth identification result includes similar and dissimilar; the fifth instruction template is used to input a table as a corresponding current table; And through a piece of natural language description text indicates the first model for the current table whether there is a molecular label column identification obtained corresponding fifth identification result, and in the fifth identification result is yes, further to the current table for molecular label column information extraction processing to obtain the corresponding first column extraction information, and in the fifth identification result is no, set the corresponding first column extraction information is empty, and by the fifth identification result and the first column extraction information obtained corresponding fifth instruction feedback output; The fifth identification result includes yes and no; The first column extraction information is not empty by one or more first column unit information; The first column unit information includes row index information and molecular label information; The first image information sequence includes a plurality of first image information; The first image information includes a first identifier, a first page code, a first upper left corner row and column coordinates, a first upper right corner row and column coordinates, a first lower right corner row and column coordinates, a first lower left corner row and column coordinates and a first image; The first image is a visual diagram of a molecular structure; The first text block information sequence includes a plurality of first text block information; The first text block information includes a second identifier, a second page code, a second upper left corner row and column coordinates, a second upper right corner row and column coordinates, a second lower right corner row and column coordinates, a second lower left corner row and column coordinates and a first block text; The first table information sequence includes a plurality of first table information; The first table information includes a third identifier, a third page code, a third upper left corner row and column coordinates, a third upper right corner row and column coordinates, a third lower right corner row and column coordinates, a third lower left corner row and column coordinates and a first table; The first table includes one or more first table rows; The first table row includes a plurality of first row unit texts; The first spliced table information sequence includes a plurality of first spliced table information; The first spliced table information includes a fourth identifier, a front page page code, a rear page page code, a front page upper left corner row and column coordinates, a front page upper right corner row and column coordinates, a front page lower right corner row and column coordinates, a front page lower left corner row and column coordinates, a rear page upper left corner row and column coordinates, a rear page upper right corner row and column coordinates, a rear page lower right corner row and column coordinates, a rear page lower left corner row and column coordinates and a first spliced table; The first spliced table includes one or more second table rows; The second table row includes a plurality of second row unit texts; The label-image correspondence table includes a plurality of first correspondence records; The first correspondence record includes a first molecular label field and a first molecular image field; The label-activity information correspondence table includes a plurality of second correspondence records; The second correspondence record includes a second molecular label field and a plurality of first molecular activity fields; The multi-modal comprehensive information table includes a plurality of third correspondence records; The third correspondence record includes a third molecular label field, a second molecular image field and a plurality of second molecular activity fields.

3. The processing method for mining patent information by combining a large language model according to claim 2, characterized in that, The first image information sequence is screened based on the first instruction template and the first model to obtain a corresponding first molecular image information sequence, specifically including: performing a round of traversal on all the first images of the first image information sequence; and during the round of traversal, inputting the first image currently traversed as a corresponding current image into the first instruction template to assemble a corresponding instruction text to obtain a corresponding current task instruction; inputting the current task instruction into the first model to process a molecular image recognition task to obtain a corresponding first instruction feedback; inputting the first recognition result in the first instruction feedback as a corresponding current recognition result; when the current recognition result is no, deleting the first image information corresponding to the current image from the first image information sequence; and when the round of traversal ends, inputting the latest first image information sequence as a corresponding first molecular image information sequence.

4. The processing method for mining patent information by combining a large language model according to claim 2, characterized in that, The molecular label screening processing of the first text block information sequence based on the second instruction template and the first model obtains a corresponding first molecular label information sequence, and specifically includes: performing a round of traversal on all the first images of the first image information sequence; and during the round of traversal, inputting the first image currently traversed as a corresponding current image into the first instruction template to assemble a corresponding instruction text to obtain a corresponding current task instruction; inputting the current task instruction into the first model to process a molecular image recognition task to obtain a corresponding first instruction feedback; inputting the first recognition result in the first instruction feedback as a corresponding current recognition result; when the current recognition result is no, deleting the first image information corresponding to the current image from the first image information sequence; and when the round of traversal ends, inputting the latest first image information sequence as a corresponding first molecular image information sequence.

5. The processing method for mining patent information by combining a large language model according to claim 2, characterized in that, The other table filtering processing of the first table information sequence based on the third instruction template and the first model obtains a corresponding second table information sequence, and specifically includes: The other table filtering processing of the first table information sequence based on the third instruction template and the first model obtains a corresponding second table information sequence, and specifically includes: performing a round of traversal on all the first tables of the first table information sequence; and during the round of traversal, inputting the first table being currently traversed as a corresponding current table into the third instruction template to obtain a corresponding current task instruction through corresponding instruction text assembly; inputting the current task instruction into the first model to obtain a corresponding third instruction feedback through table classification task processing; taking the third recognition result of the third instruction as a corresponding current recognition result; and performing recognition on the current recognition result; if the current recognition result is a molecule activity table, marking the first table information corresponding to the current table as a corresponding first type of table information; if the current recognition result is a table without a header, marking the first table information corresponding to the current table as a corresponding second type of table information; if the current recognition result is other table, deleting the first table information corresponding to the current table from the first table information sequence; and at the end of the round of traversal, taking the latest first table information sequence as a corresponding second table information sequence.

6. The processing method for mining patent information by combining a large language model according to claim 5, characterized in that, The molecule activity table and the table without a header are spliced based on the fourth instruction template, the first model and the second table information sequence to obtain a corresponding first spliced table information sequence, and the method specifically comprises: Step 61: initializing an empty sequence as a corresponding third table information sequence; and taking the first table information of the second table information sequence as a corresponding current table information; Step 62: identifying whether the current table information is marked as the first type of table information; if the current table information is marked as the first type of table information, identifying whether the row coordinate of the third lower left corner row-column coordinate of the current table information is the bottom row of the page, if yes, turning to step 63, and if no, turning to step 64; if the current table information is not marked as the first type of table information, turning to step 65; Step 63: taking the third left bottom corner row-column coordinate of the current table information as a corresponding first left bottom corner row-column coordinate; and taking the third right bottom corner row-column coordinate of the current table information as a corresponding second left bottom corner row-column coordinate; and turning to step 66; Step 63, add 1 to the third page code of the current table information to obtain a corresponding next page code; match each third page code in the second table information sequence with the next page code and the third upper left corner row and column coordinates, and take the second table information whose row coordinate of the third upper left corner row and column coordinates is the top row of the page as corresponding candidate table information; identify whether the candidate table information is empty; if the candidate table information is empty, go to step 64; if the candidate table information is not empty, record the first table of the candidate table information as a corresponding post table, record the first table of the current table information as a corresponding pre table, count the total number of the first row cell text of any first table row of the pre table to obtain a corresponding first cell total number, count the total number of the first row cell text of any first table row of the post table to obtain a corresponding second cell total number, and identify whether the first and second cell total numbers are equal; if the first and second cell total numbers are not equal, go to step 64; if the first and second cell total numbers are equal, input the second first table row of the pre table and the first first table row of the post table as corresponding first and second data rows into the fourth instruction template to obtain corresponding instruction text assembly to obtain a corresponding current task instruction, input the current task instruction into the first model to process a data row similarity identification task to obtain a corresponding fourth instruction feedback, take the fourth identification result in the fourth instruction feedback as a corresponding current identification result, and identify the current identification result; if the current identification result is not similar, go to step 64;If the current recognition result is similar, a first spliced table information with empty content is added in the third table information sequence as corresponding current spliced table information, a unique identification is allocated to the current spliced table information as corresponding current spliced table identification, the postpositioned table is added to the end of the prepositioned table to obtain corresponding current spliced table, the fourth identification of the current spliced table information is set as the current spliced table identification, the prepositioned page and postpositioned page page numbers of the current spliced table information are set as the third page numbers of the prepositioned table and postpositioned table, the prepositioned page upper left corner row and column coordinates, prepositioned page upper right corner row and column coordinates, prepositioned page lower right corner row and column coordinates and prepositioned page lower left corner row and column coordinates of the current spliced table information are set as the third upper left corner row and column coordinates, third upper right corner row and column coordinates, third lower right corner row and column coordinates and third lower left corner row and column coordinates of the prepositioned table, and the postpositioned page upper left corner row and column coordinates, postpositioned page upper right corner row and column coordinates, postpositioned page lower right corner row and column coordinates and postpositioned page lower left corner row and column coordinates of the current spliced table information are set as the third upper left corner row and column coordinates, third upper right corner row and column coordinates, third lower right corner row and column coordinates and third lower left corner row and column coordinates of the postpositioned table, the first spliced table of the current spliced table information is set as the current spliced table, and after the setting of the current spliced table information is finished, step 65 is performed. Step 64, record the first table of the current table information as the corresponding current table; and add a first splicing table information with empty content in the third table information sequence as the corresponding current splicing table information; and assign a unique identification to the current splicing table information as the corresponding current splicing table identification; and set the fourth identification of the current splicing table information as the corresponding current splicing table identification; and set the pre-page page number of the current splicing table information as the third page number corresponding to the current table; and set the post-page page number of the current splicing table information as empty; and set the top-left corner row and column coordinates, the top-right corner row and column coordinates, the bottom-right corner row and column coordinates and the bottom-left corner row and column coordinates of the pre-page of the current splicing table information as the third top-left corner row and column coordinates, the third top-right corner row and column coordinates, the third bottom-right corner row and column coordinates and the third bottom-left corner row and column coordinates corresponding to the current table; and set the top-left corner row and column coordinates, the top-right corner row and column coordinates, the bottom-right corner row and column coordinates and the bottom-left corner row and column coordinates of the post-page of the current splicing table information as empty; and set the first splicing table of the current splicing table information as the corresponding current table; Step 65, identify whether the current table information is the last first table information of the second table information sequence; if not, extract the next first table information of the second table information sequence as the new current table information and return to step 62; if yes, output the latest first splicing table information sequence as the current splicing processing result.

7. The processing method for mining patent information by combining a large language model according to claim 2, characterized in that, The first molecular image information sequence, the first molecular label information sequence, the first splicing table information sequence, the fifth instruction template and the first model are used to construct a label-image correspondence relationship table, and the label-image correspondence relationship table specifically includes: Step 71, take the first image of the first molecular image information sequence as the corresponding current molecular image; Step 72, take the first page number corresponding to the current molecular image as the corresponding current image page number, and record the area covered by the first top-left corner row and column coordinates, the first top-right corner row and column coordinates, the first bottom-right corner row and column coordinates and the first bottom-left corner row and column coordinates corresponding to the current molecular image as the corresponding current image area; Step 73, record the first splicing table information in which the leading page number or the trailing page number matches the current image page number as corresponding first same-page table information; and based on the upper left corner row and column coordinates of the leading / trailing page, the upper right corner row and column coordinates of the leading / trailing page, the lower right corner row and column coordinates of the leading / trailing page, and the lower left corner row and column coordinates of the leading / trailing page of each first same-page table information, confirm the coverage range of the current first same-page table information on the current image page to obtain corresponding first same-page table area; and when the number of the first same-page table areas is not zero, confirm whether there is an intersection between one of the first same-page table areas and the current image area to obtain corresponding first confirmation result; and identify the first confirmation result; if the first confirmation result is non-existent, go to step 74; if the first confirmation result is existent, take the first same-page table information corresponding to the first same-page table area which has intersection with the current image area as corresponding current same-page table information, and go to step 75; Step 74, record the first text block information in which the second page number matches the current image page number in the first molecular label information sequence as corresponding first same-page text block information; and based on the second upper left corner row and column coordinates, the second upper right corner row and column coordinates, the second lower right corner row and column coordinates, and the second lower left corner row and column coordinates of each first same-page text block information, confirm the coverage range of the current first same-page text block information on the current image page to obtain corresponding first same-page text block area; and take the first block text of the first text block information corresponding to the first same-page text block area closest to the current image area as corresponding current molecular label; and go to step 76; Step 75, input the first spliced table of the current same-page table information into the fifth instruction template to assemble corresponding instruction text to obtain a corresponding current task instruction; input the current task instruction into the first model to process the molecular label column identification task to obtain a corresponding fifth instruction feedback; extract the corresponding first column extraction information from the fifth instruction feedback; record the full table coverage area of the single page or double page of the current same-page table information as a corresponding first table area; record the first top-left corner row and column coordinates, the first top-right corner row and column coordinates, the first bottom-right corner row and column coordinates, and the first bottom-left corner row and column coordinates of each first molecular image in the first molecular image information sequence as a corresponding first image area; record the first image area that has an intersection with the first table area as a corresponding first intra-table image area, and sort all the first intra-table image areas in the intersection area from front to back to form a corresponding first intra-table image area sequence; take the sequence index of the first intra-table image area corresponding to the current image area in the first intra-table image area sequence as a corresponding current row index; extract the molecular label information of the first column unit information in the first column extraction information that matches the current row index as a corresponding current molecular label; Step 76, take the current molecular image as a corresponding first molecular image field; take the current molecular label as a corresponding first molecular label field; and form a corresponding first correspondence record by the obtained first molecular label field and first molecular image field; Step 77, identify whether the current molecular image is the last first molecular image in the first molecular image information sequence; if not, extract the next first molecular image in the first molecular image information sequence as a new current molecular image, and return to step 72; if yes, form a corresponding label-image correspondence table by all the first correspondence records obtained.

8. The processing method for mining patent information by combining a large language model according to claim 2, characterized in that, The label-activity information correspondence table is constructed based on the fifth instruction template, the first model, and the first spliced table information sequence, and specifically includes: Step 81, extract the first second table row of all the first splicing tables of the first splicing table information sequence to form a corresponding second table row set; and extract all the second row unit texts of the second table row set to form a corresponding second row unit text set; and perform unit text deduplication processing on the second row unit text set; and perform a round of traversal on all the second row unit texts of the second row unit text set after deduplication; and in this round of traversal, the second row unit text currently traversed is taken as the corresponding current field name; and whether the current field name is a type of molecular activity name is identified based on a preset molecular activity keyword dictionary to obtain a corresponding activity name identification result; and when the activity name identification result is yes, the current field name is marked as a corresponding molecular activity field name; and at the end of this round of traversal, all the molecular activity field names obtained form a corresponding molecular activity field name set; Step 82, perform a round of traversal on all the first splicing tables of the first splicing table information sequence; and in this round of traversal, the first splicing table currently traversed is taken as the corresponding current splicing table; and the current splicing table is input into the fifth instruction template to assemble a corresponding instruction text to obtain a corresponding current task instruction; and the current task instruction is input into the first model to perform a molecular label column identification task to obtain a corresponding fifth instruction feedback; and the first column extraction information is extracted from the fifth instruction feedback; and the first second table row of the current splicing table is taken as a corresponding current header row; and the second row unit text corresponding to the first column extraction information in the current header row is taken as a corresponding molecular label field name; and all the second row unit texts in the current header row except the molecular label field name are polled once; and in this polling process, the second row unit text currently polled is recorded as a corresponding current field name, and whether the current field name is included in the molecular activity field name set is identified, if yes, the second row unit text currently polled is taken as a corresponding molecular activity field name, if not, the arrangement index of the second row unit text corresponding to the current field name in the current header row is recorded as a corresponding to-be-deleted column index; and at the end of this polling, the second row unit texts corresponding to all the to-be-deleted column indexes in all the second table rows of the current splicing table are deleted; Step 83, after the end of the round of traversal on all the first splicing tables of the first splicing table information sequence, all field contents of all the first splicing tables are fully connected and merged with the molecular label field name as the connection key field name to obtain a new table with a molecular label field and all molecular activity fields as a corresponding label-activity information corresponding relationship table.

9. An apparatus for performing the processing method of mining patent information in combination with a large language model according to any one of claims 1-8, characterized in that, The device comprises a large language model fine-tuning module, a patent information receiving module, a patent information denoising module, a patent information mining module, and a mining report feedback module. The large language model fine-tuning module is configured to select a plurality of pre-trained multi-modal large language models as corresponding first models; assign five target tasks to the first models, denoted as first, second, third, fourth, and fifth tasks; customize instruction text templates for the first, second, third, fourth, and fifth tasks, denoted as first, second, third, fourth, and fifth instruction templates; construct five task data sets based on the first, second, third, fourth, and fifth data sets, denoted as first, second, third, fourth, and fifth data sets; and fine-tune the first models based on the first, second, third, fourth, and fifth data sets and the corresponding first, second, third, fourth, and fifth instruction templates; the first task is to perform molecular image recognition on the input image; the second task is to perform molecular label recognition on the input text; the third task is to classify the input table; the fourth task is to perform similarity recognition on the input two data rows, and the total number of data fields of the two data rows is consistent; and the fifth task is to perform molecular label column recognition on the input table; The patent information receiving module is configured to receive a user inputted small molecule drug patent PDF file as a corresponding first patent file after the first model fine-tuning is completed; and perform image, text block, and table extraction on the first patent file based on a pre-set PDF content extraction tool to obtain corresponding first image information sequence, first text block information sequence, and first table information sequence; The patent information denoising module is configured to perform molecular image screening processing on the first image information sequence based on the first instruction template and the first model to obtain a corresponding first molecular image information sequence; perform molecular label screening processing on the first text block information sequence based on the second instruction template and the first model to obtain a corresponding first molecular label information sequence; perform other table filtering processing on the first table information sequence based on the third instruction template and the first model to obtain a corresponding second table information sequence; and perform molecular activity table and no header table splicing processing based on the fourth instruction template, the first model, and the second table information sequence to obtain a corresponding first spliced table information sequence; wherein the fourth instruction template and the first model are configured to perform data row similarity recognition on the second row of the molecular activity table and the first row of the no header table when the total number of row unit texts at the bottom of the front page and the top of the back page of the molecular activity table and the no header table is equal. The patent information mining module is configured to construct a label-image correspondence table based on the first molecular image information sequence, the first molecular label information sequence, the first splicing table information sequence, the fifth instruction template, and the first model; construct a label-activity information correspondence table based on the fifth instruction template, the first model, and the first splicing table information sequence; and perform multi-modal table merging based on the label-image correspondence table and the label-activity information correspondence table to obtain a corresponding multi-modal comprehensive information table. The mining report feedback module is configured to feed back a corresponding first mining report composed of the first patent file and the multi-modal comprehensive information table to a current user.

10. An electronic device, comprising: The system comprises: a memory, a processor, and a transceiver; the processor is configured to be coupled with the memory, read and execute instructions in the memory to implement the method in any one of claims 1-8; the transceiver is coupled with the processor and controlled by the processor to perform message transmission and reception.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, when the computer instructions are executed by a computer, the computer executes the method in any one of claims 1-8. The computer readable storage medium stores computer instructions, when the computer instructions are executed by a computer, the computer executes the method in any one of claims 1-8.

Citation Information

Patent Citations

  • Multi-source heterogeneous customer service knowledge fusion method and application

    CN118569365A

  • Text information generation method and device, model training method and device and electronic equipment

    CN118587729A

Cited By

  • A method and system for iterative optimization of catalyst literature hints based on multi-modal information

    CN122417216A