A biomedical relation extraction method

By constructing a general instruction template and integrating a large-scale dataset, and fine-tuning a general large language model, the biomedical relation extraction task is transformed into a text generation task. This solves the problems of small dataset size and lack of domain knowledge in existing models for biomedical relation extraction, and achieves more efficient biomedical relation extraction.

CN119311777BActive Publication Date: 2025-10-24SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411343496.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-10-24
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

Existing biomedical pre-trained models and generative large language models suffer from problems such as small dataset size, lack of domain-specific knowledge, difficulty in understanding complex semantic differences, and lack of reasoning ability in biomedical relation extraction tasks, resulting in insufficient extraction accuracy.

Method used

We construct a general instruction template, integrate a large-scale biomedical relation extraction dataset, and transform the general large language model into a biomedical extraction task-specific model that does not require manual annotation by fine-tuning it. We then use instruction optimization methods to transform the biomedical relation extraction task into a text generation task, align it with the pre-training task of the generative large language model, and update the model parameters.

Benefits of technology

It significantly improves the performance and generalization ability of biomedical relationship extraction, enabling more accurate understanding and analysis of biomedical text content, and improving the efficiency and accuracy of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119311777B_ABST
    Figure CN119311777B_ABST
Patent Text Reader

Abstract

The application provides a biomedical relationship extraction method, relates to the technical field of data processing, and comprises the following steps: acquiring an existing biomedical relationship extraction dataset and integrating the biomedical relationship extraction dataset into a biomedical relationship extraction large-scale dataset; constructing a general instruction template; according to the biomedical relationship extraction large-scale dataset, the general instruction template is used to reorganize a large-scale instruction dataset; the general large language model is fine-tuned by using the large-scale instruction data, and the parameters of the general large language model are updated; and the biomedical relationship is extracted by using the fine-tuned general large language model. The application can accurately extract the relationship between biomedical entities.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of data processing, and particularly relates to a biomedical relation extraction method. BACKGROUND

[0002] Extracting relationships between biomedical entities from a large number of literature is crucial for building knowledge graphs, identifying biomarkers, drug discovery and reuse, and clinical decision-making. With the development of big data technology, the quantity and diversity of biomedical literature have increased significantly. In the face of such a huge amount of information, it is impractical to rely solely on human efforts to extract entity relationships from literature. Therefore, it is necessary to promote deep learning methods to realize automated and intelligent information extraction and analysis, thereby promoting the development of artificial intelligence in the field of science.

[0003] Biomedical relation extraction (BioRE) aims to mine the relationships between biomedical concepts, including genes (proteins), diseases, chemicals (drugs), variations (mutations), etc. The core function of this technology is to identify and extract the semantic relationships of common biomedical entity pairs, such as <chemical, disease>, <chemical, gene>, and <disease, gene>, and then convert the unstructured information scattered in the text into structured knowledge. The biomedical relation extraction task can be roughly divided into sentence-level and document-level. The former requires the model to extract the relationship of a single entity pair from a sentence, however, usually multiple sentences are needed to describe a relationship. Therefore, it is necessary to extract the relationship of one or more entity pairs in a document composed of several sentences.

[0004] In recent years, biomedical pre-trained language models (PLMs) have shown strong capabilities in various biomedical tasks, which are pre-trained on the BERT architecture or trained from scratch. The main method for the biomedical relation extraction (BioRE) task is to fine-tune biomedical pre-trained language models (PLMs) on the target dataset. In order to make up for the gap between pre-training and fine-tuning, some studies turn to design appropriate prompts, redefining the biomedical relation extraction task as a form of pre-training task, i.e. mask prediction. Some other works integrate external knowledge into the model. The following disadvantages exist:

[0005] The model is generally small in size, limiting the model's ability to capture complex language patterns and contextual information; mainly using biomedical literature (such as PubMed [6] , PMC [7]) The training model leads to a narrow range of linguistic phenomena and knowledge; it is optimized for specific biomedical tasks and lacks the flexibility of multi-task processing.

[0006] Large language models (LLM) have broken the limits of natural language processing and shown excellent problem-solving capabilities. They are a special class of pre-trained language models (PLMs) obtained by scaling model size, pre-training corpus, and computation. However, the source code of most existing models is not publicly available, also known as closed-source models, and their training data and training process are still unknown.

[0007] In recent years, generative general large language models have shown breakthrough performance in translation, reasoning, code generation, and other natural language processing tasks. Unlike previous pre-trained models, generative large language models do not need to update parameters, but only need to use in-context learning (ICL) to adapt to completely different tasks. Due to their extraordinary ability in text understanding, generation, and generalization, generative large language models have been used to solve relationship extraction tasks in general fields and have achieved satisfactory performance. However, general open-source large language models perform poorly in biomedical relationship extraction tasks. They have the following shortcomings:

[0008] Lack of domain-specific knowledge required to explain and extract relationships in complex biomedical text; biomedical text rich in professional terms and abbreviations is not fully trained; the model has difficulty understanding and distinguishing subtle semantic differences, resulting in insufficient accuracy in handling fine-grained relationships; lack of complex reasoning ability required to infer implicit relationships through complex indirect statements. SUMMARY

[0009] In view of the above shortcomings in the prior art, the present application provides a method for extracting biomedical relationships.

[0010] To achieve the above purpose, the technical scheme adopted by the present application is as follows: A method for extracting biomedical relationships, comprising the following steps:

[0011] S1, obtaining existing biomedical relationship extraction datasets and integrating them into a biomedical relationship extraction large dataset;

[0012] S2, constructing a general instruction template;

[0013] S3, using the general instruction template to reorganize the biomedical relationship extraction large dataset into a large instruction dataset;

[0014] S4, fine-tuning the general large language model by using a large-scale instruction dataset, updating parameters of the general large language model;

[0015] S5, extracting biomedical relations by using the fine-tuned general large language model.

[0016] The present application has the beneficial effects: the present application proposes a general large language model constructed by instruction optimization, which can convert the general large language model into a biomedical extraction task specific model without manual annotation, and significantly improve the performance in various biomedical extraction tasks. The general large language model constructed by instruction optimization in the present application solves the problem that the existing biomedical relation extraction task dataset is small in size and only targets specific tasks (such as sentence-level or document-level biomedical relation extraction); solves the problem that the existing biomedical pre-training model is limited by the dataset and the model size, and cannot fully learn the context semantic relationship; and solves the problem that although the large language model can capture more accurate context semantic relationship, it performs poorly in biomedical relation extraction tasks due to the lack of medical knowledge.

[0017] Further, the S1 is specifically:

[0018] An existing biomedical relation extraction dataset is obtained, and the biomedical relation extraction dataset is filtered, wherein the filtered biomedical relation extraction dataset includes sentence-level and document-level biomedical extraction tasks, and includes different biomedical extraction tasks;

[0019] Biological relations describing entities are obtained from the filtered biomedical relation extraction dataset, wherein the entities include genes, variations, diseases, and chemicals;

[0020] All obtained biological relations describing entities are integrated into a large-scale biomedical relation extraction dataset.

[0021] The beneficial effects of the above further scheme are: the present application integrates the existing biomedical relation extraction task related dataset to form a large-scale biomedical relation extraction task dataset, and integrates the large-scale dataset into instruction-response pair dataset, increases the scale of biomedical relation extraction field data, improves the diversity and standardization of biomedical relation extraction data, lays a foundation for subsequent research, and has reference value in the field of biomedical relation extraction. And the dataset is carefully selected from the existing biomedical extraction dataset in the literature, which widely covers various biomedical relations between entities usually discussed. This set includes biomedical extraction tasks of different complexities, including sentence-level and document-level.

[0022] Further, the general instruction template in S2 includes a task description design unit, a parameter design unit, and a response design unit.

[0023] The task description design unit is used for describing the biomedical relation extraction task.

[0024] The parameter design unit is used for marking entities in the input text.

[0025] The response design unit is used for responding to the answer, wherein the response type includes direct response and thought chain response.

[0026] The above further scheme has the beneficial effects that: the present application designs two types of general instruction templates, which are suitable for various types (sentence level, document level) of biomedical relation extraction tasks, aligns the biomedical relation extraction downstream task with the large language model generation task, and can explicitly guide the large language model to generate accurate responses. Using the two templates, the existing biomedical extraction data set can be converted into a biomedical extraction instruction without manual annotation.

[0027] Further, the task description design unit includes a first task description design subunit and a second task description design subunit.

[0028] The description content of the first task description design subunit is: in response to a piece of text, two entities in the biomedical literature are annotated, and options describing the relationship between the two entities are obtained.

[0029] The description content of the second task description design subunit is: in response to a piece of text in the biomedical literature, the specific entity is replaced by an entity type, and options describing the relationship between the two entities are obtained.

[0030] The parameter design unit includes a first input text form and a second input text form.

[0031] The input of the first input text form is to provide entity names, wherein the entity names are contained in entity type markers.

[0032] The input of the second input text form is to hide the entity names and only include the entity types.

[0033] The beneficial effects of the above further scheme are: the application designs two task description units and their corresponding text input formats, the core difference of which is that one way hides the entity name and only keeps the entity category, while the other way keeps the entity name and refers to it through an identifier. First of all, both formats have wide applicability, whether it is for sentence-level or document-level biomedical relation extraction tasks, they can be converted into a unified biomedical relation extraction processing flow through the two task descriptions and text input forms. Secondly, the first text input form hides the entity name and only uses the entity type, which can guide the large model to pay more attention to the context rather than the specific entity name; the second text input form keeps the entity name and refers to it through an identifier, which can enhance the model's attention to the relationship between specific entities and clearly indicate that the model needs to extract the relationship between these entities referenced by the special identifier.

[0034] Further, the expression of the answer of the thought chain response is as follows:

[0035]

[0036]

[0037] Among them, represents the answer of the thought chain response, represents the general large language model integrating the reasoning process of each step of the thought chain the answer of the thought chain response generated , represents the biomedical relation extraction task instruction, represents the general large language model parameter, represents the reasoning process of the thought chain at the t step, represents the reasoning process of the thought chain, T step, represents the process of the general large language model generating the inference result of the next step of the thought chain according to the reasoning result of the last step of the thought chain , represents the reasoning process of the thought chain at the t -1 step.

[0038] The beneficial effects of the above further scheme are: the application adds the thought chain response in the instruction, which enhances the generalization ability of the model to unknown biomedical extraction tasks.

[0039] Further, the S4 is specifically:

[0040] Expressing the relation extraction problem as a conditional text generation task using the proposed general instruction template;

[0041] align the biomedical relation extraction task with the pre-training task of the generative large language model, update the parameters of the general large language model, and complete fine-tuning, wherein the pre-training task of the generative large language model is a conditional text generation task.

[0042] The beneficial effect of the above further scheme is that the present application reorganizes the large-scale biomedical relation extraction dataset into a large-scale instruction dataset using a general instruction template. In this instruction-type dataset, the biomedical relation extraction task is transformed into a text generation task, and the purpose is to let the large model generate subsequent text to supplement the relationship between entities. This transformation process is consistent with the pre-training task target of the generative large language model, because both are essentially text generation tasks. This consistency of goals ensures a high degree of unity of the objective function from the pre-training stage to the fine-tuning stage, which helps to make full use of the knowledge accumulation of the pre-training model, thereby improving the performance and generalization ability of the model on the biomedical relation extraction task.

[0043] Further, the aligning of the biomedical relation extraction task with the pre-training task of the generative large language model is specifically:

[0044] inputting the instruction-response pair text into an open-source pre-trained large language model;

[0045] taking the biomedical relation extraction task as a downstream task;

[0046] redefining the downstream as a text generation task using the reorganized large-scale instruction dataset, and aligning the text generation task with the pre-training task of the generative large language model to complete fine-tuning.

[0047] The beneficial effect of the above further scheme is that the present application takes the biomedical relation extraction task as a downstream task, and transforms it into a text generation task with the help of a general template, thereby realizing the consistency of the biomedical relation extraction and the pre-training task target of the generative large language model. During model training, not all parameters of the entire large language model are updated during backpropagation, but the parameters of the main structure of the model are frozen, and only a small number of output layer parameters are updated. This method makes the network better adapt to the biomedical relation extraction task while saving computing resources.

[0048] Further, the S5 is specifically:

[0049] organize the biomedical text into the form of task description and input text using a general instruction template;

[0050] input the reorganized task description and input text into the fine-tuned general large language model;

[0051] The content received by the fine-tuned general large language model is taken as an instruction, and a probability distribution of the next word is predicted;

[0052] According to the probability distribution, the fine-tuned general large language model is used to add the optimal word to the end of the prefix input to continue generating words until all responses are completed, wherein the response process is as follows:

[0053]

[0054] wherein, represents the generated probability, represents all generated characters, represents the instruction received by the general large language model and the biomedical text input data, represents the i th character, represents the m th character, represents the character generated at the th position, represents the probability that the character generated at the I th position is under the condition of a given input, represents the joint probability of generating the response under the condition of a given input; I R

[0055] All generated words are organized to form a sentence, and the sentence is used to indicate the relationship between biomedical entities to complete the extraction of biomedical relationships.

[0056] The beneficial effects of the above further scheme are: the biomedical text is embedded into the general template designed in the application, and then input into the fine-tuned large language model. The model generates the character with the highest probability to form the final relationship between biological entities and outputs, completing the relationship extraction task of the biomedical text. Thanks to the general template, the large-scale instruction data set constructed in the application, and the fine-tuned large language model, the biomedical text content can be more accurately understood and analyzed, and the efficiency and accuracy of information extraction are significantly improved. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 The flowchart of the method of the application.

[0058] Figure 2 The specific example diagram of the two templates in the embodiment.

[0059] Figure 3 The example diagram of the direct response and the thinking chain response of the biomedical relationship extraction instruction in the embodiment. ​​

[0060] Figure 4 An alignment diagram of the biomedical relation extraction task in this embodiment and the pre-training task of the generative large language model. DETAILED DESCRIPTION

[0061] The specific embodiments of the present application are described below to facilitate the understanding of the present application for those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, it is obvious that various changes are within the spirit and scope of the present application defined and determined by the appended claims, and all the inventions utilizing the concept of the present application are within the scope of protection.

[0062] Before explaining the present solution, the following terms are explained:

[0063] Biomedical relation extraction (BioRE): Biomedical relation extraction is a natural language processing technique aimed at automatically identifying and extracting relationships between entities from biomedical literature.

[0064] Protein-protein interaction (PPI): Protein-protein interaction refers to the direct physical contact between two or more proteins within a cell, and these interactions are crucial for many biological processes such as signal transduction, metabolic pathways, etc.

[0065] Pre-trained language models (PLMs): Pre-trained language models are a type of machine learning model that undergoes unsupervised learning on a large amount of text data, learning the statistical patterns of language, and thus being able to generate and understand natural language.

[0066] Large language models (LLM): Large language models are a class of pre-trained language models with a large number of parameters, capable of exhibiting excellent performance on a wide range of natural language processing tasks.

[0067] Prompt: In natural language processing, a prompt is a piece of text or input sequence used to guide the model to generate a specific output.

[0068] In-context learning (ICL): In-context learning is a method that enables the model to generate correct answers in a given context without explicit fine-tuning. By providing a series of examples, the model can understand the task without additional training.

[0069] F-score: F-score is a metric used to evaluate the performance of a classification or information retrieval system, which is a weighted average of precision and recall.

[0070] Chain of Thouts (CoT): Chain of Thouts is a technique that allows the model to output the thinking process or reasoning steps before generating the final answer.

[0071] Supervised Fine-Tuning (SFT): Supervised Fine-Tuning is the process of further training a pre-trained model using a labeled specific task dataset.

[0072] As shown in Figure 1 , the present application provides a method for extracting biomedical relationships, which is implemented as follows:

[0073] S1, obtain existing biomedical relationship extraction datasets and integrate them into a biomedical relationship extraction large-scale dataset, which is specifically:

[0074] Obtain existing biomedical relationship extraction datasets and filter the biomedical relationship extraction datasets, wherein the filtered biomedical relationship extraction datasets include sentence-level and document-level biomedical extraction tasks, and include different biomedical extraction tasks;

[0075] Obtain the biological relationship of the entity from the filtered biomedical relationship extraction dataset, wherein the entity includes gene, variation, disease, and chemical;

[0076] Integrate all the obtained biological relationships of the entity into a biomedical relationship extraction large-scale dataset.

[0077] In this embodiment, the biomedical relationship extraction related dataset is sorted out, and the data with the following main characteristics is selected: (1) widely covers different biomedical extraction tasks and complexity; (2) contains sentence-level and document-level relationships. Based on the above criteria, BioRED, AIMed, DrugProt, BC5CDR, ChemProt, HPRD50, DDI and DisGeNet datasets are finally selected, and the selected datasets are integrated to form a large-scale dataset. The biomedical relationships between the four commonly used entities (i.e. gene, variation, disease and chemical) in the dataset are selected. The biomedical relationships selected in different datasets are different, and the final organization form of the dataset is shown in Table 1. Table 1 is a summary table of the constructed instruction dataset.

[0078] Table 1

[0079]

[0080] S2, construct a general instruction template, wherein the general instruction template comprises a task description design unit, a parameter design unit and a response design unit;

[0081] The task description design unit is configured to describe a biomedical relation extraction task.

[0082] The parameter design unit is configured to mark entities in input text.

[0083] The response design unit is configured to respond to an answer, wherein the response type comprises a direct response and a thinking chain response.

[0084] The task description design unit comprises a first task description design subunit and a second task description design subunit.

[0085] The first task description design subunit is configured to, in response to a piece of text, annotate two entities in biomedical literature, and obtain options describing the relationship between the two entities.

[0086] The second task description design subunit is configured to, in response to a piece of text in biomedical literature, replace specific entities with entity types, and obtain options describing the relationship between the two entities.

[0087] The parameter design unit comprises a first input text form and a second input text form.

[0088] The first input text form is configured to provide entity names, wherein the entity names are contained in entity type markers.

[0089] The second input text form is configured to hide entity names and only include entity types.

[0090] In this embodiment, the specific entity is a professional term of a gene, a disease or a drug in a piece of biomedical text, such as "sodium iodine transport protein" and "congenital thyroiditis". The biomedical relation extraction task is to extract the relationship between these terms.

[0091] In this embodiment, the design of the general instruction template includes task description design, parameter design and response design. Specifically as follows:

[0092] Task description design: The biomedical relation extraction task is described, and two description templates are designed to make it easy for large language models to understand. Template 1 describes the content as follows: The following is a piece of text that annotates two entities in biomedical literature. Please choose one from “none, association, binding, comparison, conversion, co-processing, drug interaction, negative correlation, positive correlation” that correctly describes the relationship between the two entities. Template 2 describes the content as follows: The following is a piece of text in biomedical literature, where specific entities are replaced by entity types. Please choose one from “no association, association” that correctly describes the relationship between the two entities.

[0093] Parameter design: Two types of input text are designed, and special markers are used to annotate entities in the input text, Figure 2 are specific examples of the two templates. The figure shows the input of the biomedical relation extraction based on the BERT pre-trained language model (left), the biomedical relation extraction instruction template 1 proposed for the given entity name sample (middle), and the biomedical relation extraction instruction template 2 proposed for the sample including only entity types (right), such as the sample in the GAD dataset. The template is instantiated using parameters (dark part), including relationship type, input text, and entity information. The input of template 1 provides entity names, which are contained in entity type markers, such as “due to the presence of a new deletion in the @gene target $ sodium / iodine symporter @ / gene target $, the @disease source $ congenital hypothyroidism @ / disease source $ occurs.” The input of template 2 hides the entity name and only contains the entity type, such as “the C1772T polymorphism of the @gene $ has no association with the progression or metastasis of the @disease $.”

[0094] In this embodiment, the response design provides two types of response (answer) templates. The first type is a direct response (short for response), which provides a gold answer, for example, the correct relationship type of the annotated entity pair in the given input text. This response is simple and can serve as an explicit reference for the model to learn the correct output of each input text. The second response is a chain of thought response (Chain of Thought, CoT). Its answer provides both the gold answer and the reasoning logic. The chain of thought technology decomposes the solution to a problem into a series of steps , where each step is a function of the previous step:

[0095]

[0096] The final response is composed of all steps:

[0097]

[0098] where, represents the answer of the chain of thought response, representing the reasoning process of each step of the general large language model integrated thought chain the answer of the generated thought chain response , representing the biomedical relation extraction task instruction, representing the general large language model parameters, representing the reasoning process of the thought chain at the t-th step, representing that the thought chain infers T-step reasoning, representing the process of generating the inference result of the next step of the thought chain by the general large language model according to the inference result of the previous step of the thought chain , representing the reasoning process of the thought chain at the t-1-th step.

[0099] In this embodiment, the reasoning steps of the thought chain response help the model to understand the underlying logic and context required to obtain the correct relation type. As shown in Figure 3 , an example of a direct response and a thought chain response of a biomedical relation extraction instruction is shown.

[0100] S3, according to the biomedical relation extraction large-scale data set, using the general instruction template, reorganizing into a large-scale instruction data set;

[0101] In this embodiment, the biomedical relation extraction large-scale data set is reorganized into an instruction data set according to the template as follows:

[0102] Entity annotation. Mark the entities in the original large-scale data set, and add label symbols to the entities according to the ways of template 1 and template 2, such as @gene target$ and @gene$.

[0103] Relation annotation. Add entity relation labels to each input text, and the relation labels include gold response and thought chain response.

[0104] Instruction-response pair merging. Organize the instruction description, input text, and response into a whole paragraph of text, and save it as a new large-scale instruction data set.

[0105] S4, using the large-scale instruction data set, fine-tuning the general large language model, and updating the parameters of the general large language model, which is specifically:

[0106] Express the relation extraction problem as a conditional text generation task using the proposed general instruction template;

[0107] The biomedical relationship extraction task is aligned with the pre-training task of the generative large language model, and the parameters of the general large language model are updated to complete fine-tuning. The pre-training task of the generative large language model is the conditional text generation task. The large-scale instruction dataset consists of three parts: (1) instructions arranged in the format of a general instruction template; (2) sentences of biological entity relationships to be extracted; (3) the relationships corresponding to the entities in the biomedical sentences, that is, labels. Each record in the dataset consists of an instruction, a relationship sentence to be analyzed, and a label. These labels include not only the final extracted biomedical relationship results, but also the thought chain process experienced when reaching this conclusion.

[0108] Align the biomedical relation extraction task with the pre-training task of a generative large language model, specifically:

[0109] Input the command-response text into the generative large language model;

[0110] Taking biomedical relationship extraction as a downstream task;

[0111] The downstream is redefined as a text generation task using a reorganized large-scale instruction dataset, and the text generation task is aligned with the pre-training task of the generative large language model to complete fine-tuning.

[0112] In this embodiment, a general large language model is used to perform fine-tuning on the reconstructed large-scale instruction dataset, specifically:

[0113] The biomedical relation extraction task is formulated as a conditional text generation task using the proposed instruction template. Specifically, the last relation in the input text is hidden, and the large language model is expected to automatically generate the relation.

[0114] Align the biomedical relation extraction task with the pre-training task of generative large language models (i.e., text generation), e.g. Figure 4 As shown on the right, the specific operation involves inputting command-response text into the open-source pre-trained large language model LLaMA-2, using the biomedical relationship extraction task as a downstream task. The LLaMA-2 model is then fine-tuned for instructions and the model parameters are updated. The experimental setup includes using eight NVIDIA A100 GPUs to fine-tune the model for three epochs, with a learning rate of 2e-5 and a batch size of 128. For the LLaMA-2 7B and LLaMA-2 13B models, a cosine learning rate scheduler with a 3% warm-up phase is used to optimize the parameters.

[0115] In this embodiment, the large-scale instruction data set consists of general instruction templates and corresponding relationships, namely Figure 3 It can be understood as converting large-scale data sets into Figure 3The format is organized into a whole paragraph, and is reorganized as a large-scale instruction data set, the input of the data set is the instruction + the sentence to be extracted relationship, and the label is the thinking chain response process + the biomedical relationship finally obtained.

[0116] In this embodiment, the generative large language model highlights that the downstream task expects the large model to generate biomedical relationships and complete the whole paragraph, which is consistent with the pre-training task objective of the generative large language model. Specifically, the general large language model can be divided into two categories. One is the large language model represented by Bert, whose neural network architecture is the encoder of Transformer, and the pre-training task is more like doing a fill-in-the-blank task. The other is the large language model represented by GPT and Llama, whose neural network architecture is the decoder of Transformer, and the pre-training task is more like doing a creative post-generation task. The latter is referred to as a generative large language model in the present application. In the biomedical relationship extraction of the present application, the downstream task is designed in the form of a sentence to be completed, and the large model is required to complete the relationship between two entities, for example, the general instruction template designed in the present application requires the model to complete the following sentence: “Therefore, the relationship between @disease$ and @disease$ is…” In this way, the biomedical relationship extraction task is consistent with the pre-training task of the generative large language model, both of which are to accurately supplement the post information, so it is mentioned that “align the biomedical relationship extraction task with the pre-training task of the generative large language model, complete fine-tuning, and update the parameters of the general large language model.” When the objectives of the pre-training stage and the fine-tuning stage are consistent, it is often considered that the potential of the pre-trained large model can be maximized, which helps to improve the performance of the large model on specific tasks.

[0117] In this embodiment, the pre-training process of the generative large language model uses the decoder component of the Transformer architecture to predict and generate subsequent characters step by step in an autoregressive manner, aiming to optimize the similarity between generated text and real text, thereby enhancing the text generation capability of the model. In the relationship extraction task in the biomedical field, traditional methods focus on explicit classification to determine the specific relationship between two entities in the text, and the model determines the probability of each relationship, and finally classifies the relationship between entities as the relationship with the highest probability. However, this method innovatively uses an implicit conversion strategy to reconstruct the biomedical relationship extraction into a generative task using the general template designed by the invention. The specific form is to generate a text sequence like "@Disease A and @Disease B have a relationship of…". This conversion redefines the downstream task as a text generation task, which is consistent with the pre-training task of the generative large language model. Through the design of the general instruction template, the pre-training task of the generative large language model (text generation task) and the biomedical relationship extraction task (converted generative task) can be successfully unified in form. This consistency ensures that the objective functions of the pre-training and fine-tuning stages are highly consistent, which helps to maximize the use of the knowledge accumulation of the pre-trained large model, and thus improves the performance and generalization ability of the model in specific (i.e., biomedical relationship extraction) tasks.

[0118] S5, using the fine-tuned general large language model to extract biomedical relationships, specifically:

[0119] Using the general instruction template, organize the biomedical text into the form of task description and input text;

[0120] Input the reorganized task description and input text into the fine-tuned general large language model;

[0121] The fine-tuned general large language model receives the content as an instruction to predict the probability distribution of the next word;

[0122] According to the probability distribution, the fine-tuned general large language model adds the optimal word to the end of the prefix input to continue generating words until all responses are complete;

[0123] Organize all generated words into a sentence, and use the sentence to indicate the relationship between biomedical entities to complete the extraction of biomedical relationships.

[0124] In this embodiment, the biomedical large language model optimized by the instruction is used to extract the relationship between entities in the biomedical text, specifically:

[0125] The biomedical text is organized into a task description and input text according to the instruction template. The reorganized text is input into the biomedical large language model fine-tuned, which uses the received instructions as instructions to predict the probability distribution of the next word. According to the probability distribution, the model adds the most likely word to the end of the prefix input to continue generating subsequent words until the entire response is completed. The response process of the model is described as follows: given an instruction I (composed of input text X and other user-defined content), a response The goal is to maximize the conditional probability in the autoregressive formula:

[0126]

[0127] where, P (Y | X) represents the generated probability, Y represents all generated characters, X represents the instructions received by the general large language model and the biomedical text input data, Y represents the i th character, Y represents the m th character, Y represents the character generated at the th position, P (Y | X) represents the probability that the character generated at the I th position is given the input , P (Y | X) represents the joint probability of generating the response I given the input R .

[0128] Finally, the generated subsequent words are organized together to form a sentence indicating the relationship between biomedical entities. At this point, the biomedical relationship extraction task is completed.

[0129] The biomedical relationship extraction professional large language model constructed by instruction tuning proposed in the present application converts the general large language model into a biomedical extraction task specific model without manual annotation, significantly improves the performance in various biomedical extraction tasks, and shows the most accurate results on multiple biomedical relationship extraction data sets.

[0130] The biomedical relation extraction professional large language model constructed through instruction optimization has certain robustness and generalization, and also performs well on untrained biomedical relation extraction datasets. In Table 2, the N-ary

[30] and EU-ADR

[52] two datasets are not included in the instruction dataset for training, but the biomedical relation extraction professional large language model proposed by the present application still performs well under the REaMA-13B-supervised fine-tuning experiment setting. Table 2 is the evaluation table of different models on the N-ARY and EU-ADR datasets.

[0131] Table 2

[0132]

Claims

1. A method of extracting biomedical relations, characterized by, The method comprises the following steps: S1, obtaining an existing biomedical relation extraction dataset and integrating it into a biomedical relation extraction large-scale dataset; S2, constructing a general instruction template; The general instruction template in S2 comprises a task description design unit, a parameter design unit and a response design unit; The task description design unit is used for describing the biomedical relation extraction task; the task description design unit comprises a first task description design subunit and a second task description design subunit; the description content of the first task description design subunit is: in response to a piece of text, annotating two entities in biomedical literature, obtaining options describing the relationship between the two entities; the description content of the second task description design subunit is: in response to a piece of text in biomedical literature, the specific entity is replaced by an entity type, and options describing the relationship between the two entities are obtained; The parameter design unit is used for marking entities in the input text; the parameter design unit comprises a first input text form and a second input text form; the input of the first input text form is to provide an entity name, wherein the entity name is contained in an entity type label; the input of the second input text form is to hide the entity name and only includes the entity type; The response design unit is used for responding to answers, wherein the response type comprises direct response and thought chain response; the expression of the answer of the thought chain response is as follows: wherein, represents the answer of the thinking chain response, represents the reasoning process of the general large language model integrating each step of the thinking chain the answer of the generated thinking chain response , represents the biomedical relation extraction task instruction, represents the general large language model parameters, represents the reasoning process of the thinking chain at the t step, represents that the thinking chain infers the T step reasoning, represents the process of the general large language model generating the inference result of the next step of the thinking chain according to the reasoning result of the previous step of the thinking chain , represents the reasoning process of the thinking chain at the t -1 step. S3, according to the biomedical relation extraction large-scale dataset, using the general instruction template, re-grouping into a large-scale instruction dataset; S4, using the large-scale instruction dataset, fine-tuning the general large language model, updating the parameters of the general large language model; S5, using the fine-tuned general large language model, extracting biomedical relations.

2. The biomedical relation extraction method of claim 1, wherein, The S1 is specifically: Obtaining an existing biomedical relation extraction dataset and screening the biomedical relation extraction dataset, wherein the screened biomedical relation extraction dataset comprises sentence-level and document-level biomedical relation extraction tasks, and comprises different biomedical relation extraction tasks; Obtaining the biological relationship of the entity from the screened biomedical relation extraction dataset, wherein the entity comprises gene, variation, disease and chemical; Integrating all the obtained biological relationships of the entity into a biomedical relation extraction large-scale dataset.

3. The biomedical relation extraction method of claim 1, wherein, The S4 is specifically: Expressing the relation extraction question as a conditional text generation task using the proposed general instruction template; Aligning the biomedical relation extraction task with the pre-training task of the generative large language model, updating the parameters of the general large language model, and completing the fine-tuning, wherein the pre-training task of the generative large language model is a conditional text generation task.

4. The biomedical relation extraction method of claim 3, wherein, The alignment of the biomedical relation extraction task with the pre-training task of the generative large language model is specifically: Inputting the instruction-response pair text into an open-source pre-trained large language model; Taking the biomedical relation extraction task as a downstream task; The downstream is redefined as a text generation task by using the reorganized large-scale instruction dataset, and the text generation task is aligned with the pre-training task of the generative large language model, and fine tuning is completed.

5. The biomedical relation extraction method of claim 4, wherein, The S5 is specifically: The biomedical text is organized into the form of task description and input text by using a general instruction template; The reorganized task description and input text are input into the fine-tuned general large language model; The content received by the fine-tuned general large language model is taken as an instruction to predict the probability distribution of the next word; According to the probability distribution, the fine-tuned general large language model is used to add the optimal word to the end of the prefix input to continue generating words until all responses are completed, wherein the response process is as follows: wherein, denotes the probability of generation, denotes all characters generated, denotes instructions received by the general large language model and biomedical text input data, denotes the i character, denotes the m characters generated, denotes the character generated at the position, denotes the probability that the character generated at the position is given the input I , denotes the joint probability of generating the response R given the input I ;​​ All generated words are organized to form a sentence, and the sentence is used to indicate the relationship between biomedical entities to complete the extraction of biomedical relationships.

Citation Information

Patent Citations

  • Inverse argument generation model, model training and reasoning method and evaluation standard based on large model

    CN117407589A

  • Unsupervised Template Extraction

    US20190034408A1