Automatic cue word generation and evaluation method for relation extraction in biomedical field
By establishing a prompt word scene generation template and a relationship extraction prompt word template in the field of biomedical science, generating and evaluating prompt words, the time-consuming and labor-intensive generation and evaluation of effective prompt words in the field of biomedical science, is solved, and efficient and accurate automatic generation and evaluation of prompt words are achieved.
Patent Information
- Application Number
- CN202510567442.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art generates effective prompt words and performs accurate evaluation in the field of biomedical science, which has time-consuming and labor-intensive and data quality-dependent problems, resulting in inaccurate output results of large language models.
Establish multiple prompt word scene generation templates and relationship extraction prompt word templates, combine the relationship name set, article and large language model in the field of biomedical science to generate and evaluate the initial relationship extraction prompt words, and determine the optimal prompt words through F1 score.
It reduces the need for a large amount of labeled data, reduces the cost and time of data preparation, improves the accuracy and reliability of relationship extraction, adapts to different document types and fields, and quickly generates high-quality prompt words.
Smart Images

Figure CN120449868A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field. Background Art
[0002] With the rapid development of artificial intelligence (AI), large language models (LLMs) are increasingly being used in the biomedical field, bringing new opportunities for biomedical research, medical diagnosis, and drug development. For example, models such as OpenBioLLM-Llama3 have demonstrated outstanding performance in biomedical text understanding and generation, enabling them to handle a wide range of biomedical tasks. However, in practical applications, generating effective prompt words and accurately evaluating them for diverse biomedical scenarios remains a pressing challenge.
[0003] Currently, the biomedical field is highly specialized, and LLM applications in this field often rely on training and fine-tuning with manually annotated data. This approach is not only time-consuming and labor-intensive, but also relies on data quality. Furthermore, due to the limited capabilities of LLM, manually designed prompts are often misinterpreted, resulting in inaccurate model output or incompatibility. For example, in medical question-answering systems, while large language models are used to generate responses, the design and evaluation of prompts remain deficient and require further optimization. Summary of the Invention
[0004] The present invention discloses a method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field, so as to overcome the above technical problems.
[0005] In order to achieve the above object, the technical solution of the present invention is:
[0006] A method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field includes the following steps:
[0007] S1: Based on the prompt word generation principle, multiple prompt word scene generation templates and relationship extraction prompt word templates are established;
[0008] S2: Generate templates and large language models based on a preset set of relationship names, biomedical articles, and prompt word scenarios in the biomedical field to obtain a prompt word scenario group;
[0009] S3: Obtain initial relation extraction prompt words based on entity combinations, prompt word scene groups, relation extraction prompt word templates, and large language models in biomedical articles;
[0010] The initial relationship extraction prompt words include a plurality of prompt word units corresponding to the kth relationship name in the updated synonym phrase set of the i-th relationship name; wherein the prompt word unit corresponding to the kth relationship name in the updated synonym phrase set of the i-th relationship name includes: the kth relationship name in the updated synonym phrase set of the i-th relationship name, and the error-prone name in the final error-prone group of the kth relationship name in the updated synonym phrase set of the i-th relationship name; wherein i is the index of the relationship name in the relationship name set in the biomedical field; and k is the index of the relationship name in the updated synonym phrase set.
[0011] S4: obtaining a relationship extraction prompt phrase according to the initial relationship extraction prompt word;
[0012] S5: Obtain the score values of the relationship extraction prompt words in the relationship extraction prompt word group to determine the optimal relationship extraction prompt word in the relationship extraction prompt word group, and complete the automatic generation and evaluation of the prompt words.
[0013] Furthermore, the method for obtaining the prompt word scene group is as follows:
[0014] S21: Generate a template and a large language model based on the i-th relationship name in a preset relationship name set within the biomedical field, a biomedical article, and a first prompt word scenario, obtain synonyms or synonym phrases of the i-th relationship name that appear in the biomedical article, and obtain a synonym phrase set for the i-th relationship name; wherein i is the index number of the relationship name in the relationship name set within the biomedical field, i=1,…,I; and I is the total number of relationship names in the relationship name set within the biomedical field;
[0015] S22: Generate a template and a large language model based on the synonym phrase set of the relationship name, the biomedical article, and the second prompt word scenario, obtain multiple expressions of the synonym phrases in the synonym phrase set of the i-th relationship name in the biomedical article, and add them to the synonym phrase set of the i-th relationship name; and update the synonym phrase set of the i-th relationship name;
[0016] S23: generating a template and a large language model based on the updated synonym phrase set of the i-th relationship name, the biomedical article, and the third prompt word scenario, obtaining antonym phrases of the synonym phrases in the synonym phrase set of the i-th relationship name in the biomedical article to form an error-prone group;
[0017] S24: Based on a preset entity type set in the biomedical field, the head entity type and the tail entity type of the kth relationship name in the updated synonym phrase set of the ith relationship name are replaced with entity types in the entity type set, respectively, to obtain a first error-prone relationship name of the kth relationship name in the updated synonym phrase set, so as to form a first error-prone name group of the kth relationship name in the updated synonym phrase set, and a template and a large language model are generated according to the biomedical article and the fourth prompt word scenario, to determine whether the first error-prone relationship name appears in the biomedical article;
[0018] If the first error-prone relationship name appears in the biomedical article, the first error-prone relationship name is added to the error-prone group of the kth relationship name in the updated synonym phrase set, and the error-prone group of the kth relationship name in the updated synonym phrase set is updated;
[0019] S25: Based on a preset entity type set in the biomedical field, swapping the head entity type and the tail entity type of the k-th relationship name in the updated synonym phrase set of the relationship name, obtaining a second easily-mistaken relationship name of the k-th relationship name in the updated synonym phrase set, so as to form a second easily-mistaken name group of the k-th relationship name in the updated synonym phrase set, and generating a template and a large language model based on the biomedical article and the fourth prompt word scenario, to determine whether the second easily-mistaken relationship name appears in the biomedical article;
[0020] If a second error-prone relationship name appears in the biomedical article, the second error-prone relationship name is added to the error-prone group of the kth relationship name in the updated synonym phrase set, and the error-prone group of the kth relationship name in the updated synonym phrase set is updated for the second time;
[0021] S26: Based on a preset set of relationship names in the biomedical field, obtaining a third error-prone relationship name having the same entity type as the kth relationship name in the updated synonym phrase set of the relationship name, and generating a template and a large language model based on the biomedical article and the fourth prompt word scenario, to determine whether the third error-prone relationship name appears in the biomedical article;
[0022] If a third error-prone relationship name appears in the biomedical article, the third error-prone relationship name is added to the error-prone group of the k-th relationship name in the updated synonym phrase set, and the error-prone group of the k-th relationship name in the updated synonym phrase set is updated for the third time to obtain a final error-prone group of the k-th relationship name in the updated synonym phrase set for the i-th relationship name;
[0023] S27: Obtain the prompt word scenario group of the i-th relationship name based on the preset relationship name set in the biomedical field, the updated synonym phrase set of the i-th relationship name, and the final error-prone group of the k-th relationship name in the updated synonym phrase set of the i-th relationship name.
[0024] Furthermore, the method for obtaining the relationship extraction prompt phrase is as follows:
[0025] S41: using the initial relationship extraction prompt word as the first relationship extraction prompt word in the relationship extraction prompt word group;
[0026] S42: Delete the prompt word units corresponding to the first to kth relationship names in the updated synonym phrase set of the i-th relationship name from the initial relationship extraction prompt words to obtain the k+1th relationship extraction prompt word; and obtain the relationship extraction prompt phrase; wherein k=1, ..., K-1;
[0027] Wherein, k is the index of the relationship name in the updated synonym phrase set; K is the total number of relationship names in the updated synonym phrase set.
[0028] Furthermore, the method for determining the optimal relationship extraction prompt word in the relationship extraction prompt word group is as follows:
[0029] S51: Set the coefficients of the first relationship name and the error-prone name of the first relationship name in the initial relationship extraction prompt word to 1, and obtain the F1 score F1 of the initial relationship extraction prompt word baseline ;
[0030] S52: Set the coefficients of the first relationship name and the error-prone name of the first relationship name in the initial relationship extraction prompt word to 0, obtain a new prompt word, and obtain the F1 score F1 of the new prompt word new ; Get the contribution V of the first relation extraction prompt word unit in the initial relation extraction prompt word i , the formula used is as follows:
[0031] V i =F1 baseline -F1 new ;
[0032] S53: If V i <0, remove the first relation extraction prompt word unit in the initial relation extraction prompt word;
[0033] If V i≥0, then the coefficient of restoring the first relation name in the initial relation extraction prompt word is 1, and the coefficient of restoring the non-jth error-prone name corresponding to the first relation name is 1, j = 1, ..., J; get the second new prompt word pt j ′, and obtain the second new prompt word pt j ′’s F1 score newj , and then obtain the contribution V of the j-th error-prone name corresponding to the first relation name in the initial relation extraction prompt word j :The formula used is as follows:
[0034] V j =F1 baseline -F1 newj ;
[0035] S54: If V j <0, remove the j-th error-prone name corresponding to the first relationship name in the initial relationship extraction prompt; otherwise, restore the coefficient of the j-th error-prone name of the first relationship name in the initial relationship extraction prompt to 1; update the initial relationship extraction prompt to obtain the updated relationship extraction prompt; j is the index of the error-prone name; J is the total number of error-prone names;
[0036] S55: Obtain the updated F1 score of the relation extraction prompt word;
[0037] If the F1 score of the updated relation extraction prompt word is greater than the F1 score of the initial relation extraction prompt word baseline The F1 score in ; then execute S56,
[0038] S56: Based on the updated relation extraction prompt words, re-execute S51-S55 until the F1 score of the updated relation extraction prompt phrase no longer decreases. At this time, the updated relation extraction prompt word is the optimal relation extraction prompt phrase.
[0039] Furthermore, the relationship extraction prompt word template is established as follows:
[0040] {Biomedical Articles}
[0041] You are a biomedical expert. Based on the biomedical article above, please judge whether "{relationship name}" is true according to the following prompt words and scenarios. Please give your reasons according to the scenarios:
[0042] {Prompt word scene group}
[0043] 99. None of the above, please answer: 99
[0044] Do not output any other information.
[0045] 6. Furthermore, the plurality of prompt word scene generation templates include: a first prompt word scene generation template, a second prompt word scene generation template, a third prompt word scene generation template, and a fourth prompt word scene generation template;
[0046] The first prompt word scene generation template is established as follows:
[0047] {Biomedical Articles}
[0048] You are a biomedical expert. Give the synonyms and synonym groups of {relationship name} that appear in the above biomedical articles.
[0049] Output format: Synonym 1 | Synonym 2 | Synonym 3 | ...
[0050] No other information is output;
[0051] The second prompt word scene generation template is established as follows:
[0052] {Biomedical Articles}
[0053] You are a biomedical expert. Give multiple forms of expression of {relationship name} in the above biomedical article, including passive voice and nominal forms. For example, "{gene} induces {disease}" can be expressed as: "{disease} is induced by {gene}" or "{gene} is the inducing factor of {disease}";
[0054] Output format: expression1|expression2|expression3|...
[0055] Do not output any other information;
[0056] The third prompt word scene generation template is established as follows:
[0057] {Biomedical Articles}
[0058] You are a biomedical expert. Give all the antonyms of {relationship name} that appear in the above biomedical article.
[0059] The fourth prompt word scene generation template is established as follows:
[0060] {Biomedical Articles}
[0061] You are a biomedical expert. Determine whether {relationship name} appears in the above biomedical article. If so, reply 1. If not, reply 0.
[0062] No other information is output.
[0063] Beneficial Effects: The present invention provides a method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field. Based on prompt word generation principles, multiple prompt word scenario generation templates and relationship extraction prompt word templates are established. A prompt word scenario group is obtained based on a preset set of relationship names, biomedical articles, and a large language model within the biomedical field. Initial relationship extraction prompt words are then obtained based on entity combinations in the biomedical articles in conjunction with the large language model. Finally, the optimal relationship extraction prompt word is determined based on the score of the relationship extraction prompt word, completing the automatic generation and evaluation of prompt words. The present invention can generate effective prompt words for different biomedical scenarios and accurately evaluate them. By establishing prompt word scenario generation templates and relationship extraction prompt word templates, the need for large amounts of annotated data is effectively reduced, lowering the cost and time of data preparation, making relationship extraction technology more readily applicable to data-scarce fields. The relationship extraction process can be performed more quickly, eliminating the need to train / fine-tune a large language model, thereby accelerating the development cycle. Furthermore, the present invention is adaptable to different document types and fields. Improving the accuracy of relationship extraction can automatically and vividly generate high-quality prompt words, and accurately evaluate the prompt words through a scientific evaluation system, thereby improving the effectiveness and reliability of relationship extraction in the biomedical field using large language models, and providing stronger support for biomedical research and practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0065] Figure 1 This is a flow chart of the method for automatically generating and evaluating prompt words for biomedical field relationship extraction according to the present invention. DETAILED DESCRIPTION
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0067] This embodiment introduces a method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field, including the following steps: Figure 1 As shown:
[0068] S1: Based on the prompt word generation principle, establish the prompt word scene generation template and the relationship extraction prompt word template;
[0069] Specifically, the prompt word generation principle of the present embodiment is formulated based on biomedical knowledge and the characteristics of large language models, combined with experience, and is as follows:
[0070] Principle 1. Designing prompts for large language models as single-choice questions rather than multiple-choice or open-ended questions can reduce ambiguity and improve answer accuracy.
[0071] Principle 2. Require large language models to answer numbers (such as 1 for "yes" and 0 for "no") rather than answering yes or no.
[0072] Principle 3. Expand synonyms and synonym groups for relationship names. For example, for the relationship name "{gene} is a marker for {disease}," it can be expanded to include synonyms or phrases such as "{gene} mutations are associated with {disease}," "{gene} induces {disease}," and "{gene} treats {disease}."
[0073] Principle 4. Expand relationship names to multiple forms, including passive voice and nominalization. For example, for the relationship name "{gene} induces {disease}", use the form: "{disease} is induced by {gene}" or "{gene} is the inducing factor of {disease}".
[0074] Principle 5. Use the antonym of the relation name (if applicable) to construct negative examples. For example, for the relation name "{chemical} enhances {gene} expression," LLM may mistakenly consider "{chemical} inhibits {gene} expression" as a positive example. Therefore, the latter should be considered a negative example.
[0075] Principle 6. Replace correct entities with potentially incorrect entity types to create negative examples. Large language models (LLMs) can misidentify subjects or objects in complex sentences. For example, given the relation name "{gene} is a marker for {disease}," the LLM might mistakenly identify "Glutathione (chemical, non-gene) is effective in treating liver toxicity (disease)" as a positive example. To avoid this, add "{chemical} is effective in treating {disease}" as a negative example.
[0076] Principle 7. Swap entity types to construct negative examples. For example, LLM may believe that "{chemical} affects {gene}" and "{gene} affects {chemical}" have the same meaning, so they need to be swapped to construct negative examples.
[0077] Among them, genes, chemicals, and diseases are entity types.
[0078] Principle 8. Replace relation names with easily confused terms to construct negative examples. For example, for the relation name "{chemical} enhances {gene} expression", LLM may confuse it with "{chemical} enhances {gene} activity", "{chemical} enhances {gene} transport", or "{chemical} enhances {gene} metabolism", so these confusing terms should be used to construct negative examples.
[0079] Preferably, according to Principle 3, the first prompt word scene generation template is established as:
[0080] {Biomedical Articles}
[0081] You are a biomedical expert. Give the synonyms and synonym groups of {relationship name} that appear in the above biomedical articles.
[0082] For example, for a relationship named "{gene} is a marker of {disease}", it can be expanded to synonyms or phrases such as "{gene} mutation is associated with {disease}", "{gene} induces {disease}", and "{gene} treats {disease}".
[0083] Output format: Synonym 1 | Synonym 2 | Synonym 3 | ...
[0084] No other information is output.
[0085] Preferably, according to Principle 4, the second prompt word scene generation template is established as:
[0086] {Biomedical Articles}
[0087] You are a biomedical expert. Give multiple ways to express {relationship name} in the biomedical article, including passive voice and nominal forms. For example, "{gene} induces {disease}" can be expressed as: "{disease} is induced by {gene}" or "{gene} is the inducing factor of {disease}."
[0088] Output format: expression1|expression2|expression3|...
[0089] No other information is output.
[0090] Preferably, according to Principle 5, the third prompt word scene generation template is established as:
[0091] {Biomedical Articles}
[0092] You are a biomedical expert. Give all the antonyms for {relationship name} that appear in the above biomedical article.
[0093] For example, the antonym of "{chemical} enhances {gene} expression" is "{chemical} inhibits {gene} expression"
[0094] Output format: antonym 1 | antonym 2 | antonym 3 | ...
[0095] No other information is output.
[0096] Preferably, according to principles 6, 7, and 8, the fourth prompt word scene generation template is established as follows:
[0097] {Biomedical Articles}
[0098] You are a biomedical expert. Determine whether {relationship name} appears in the above biomedical article. If it does, reply 1; if it does not, reply 0. Do not output any other information.
[0099] Preferably, according to Principle 1 and Principle 2, the relationship extraction prompt word template is established as follows:
[0100] {Biomedical Articles}
[0101] You are a biomedical expert. Based on the biomedical article above, please judge whether "{relationship name}" is true according to the following prompt words and scenarios. Please give your reasons according to the scenarios:
[0102] {Prompt word scene group}
[0103] 99. None of the above, please answer: 99
[0104] Do not output any other information.
[0105] S2: Generate a prompt word scene group based on a preset set of relationship names in the biomedical field, biomedical articles, entity types, relationship names, prompt word scene generation templates and a large language model.
[0106] Specifically, in one embodiment of the present invention, the relationship name is "{gene} is a marker for {disease}", where gene and disease are entity types. The first entity in each relationship name is the head entity, and the last entity is the tail entity.
[0107] S21: According to the i-th relationship name r in the preset relationship name set R in the biomedical field i , biomedical articles, first prompt word scene generation template and large language model, obtain the i-th relationship name r i Synonyms or synonym phrases appearing in biomedical articles to obtain the relationship between the i-th relation name r i and the i-th relation name r iThe synonym phrases appearing in the biomedical article together constitute a synonym phrase set of the i-th relationship name; wherein i is the index number of the relationship name in the relationship name set in the biomedical field, i=1,…,I; I is the total number of relationship names in the relationship name set in the biomedical field;
[0108] Specifically, loop through all the relationship names r in the relationship name set R i , i = 1, ..., I, and the known biomedical articles D; and input into the large language model together with the prompt word scene generation template, return the i-th relationship name r i Synonymous phrases in biomedical articles i1 ,s i2 ,s i3,…}, and the i-th relationship name r in the relationship name set i Join the synonym phrase S of the i-th relation name together i In the example, the synonym phrases of the relationship name are updated. Where i represents the index number of the relationship name in the relationship name set in the biomedical field; s i1 The first synonym phrase in the synonym phrase set representing the i-th relation name;
[0109] For example, the first relationship name r1 in the preset relationship name set R in the biomedical field is "{gene} is a marker of {disease}", and the LLM return value is "{gene} mutation is related to {disease}", "{gene} induces {disease}", "{gene} treats {disease}", which are incorporated into the synonym group S1 of the first relationship name r1 together with the first relationship name r1, S1 = {"{gene} mutation is related to {disease}", "{gene} induces {disease}", "{gene} treats {disease}", "{gene} is a marker of {disease}", ...}.
[0110] S22: Generate a template and a large language model based on the synonym phrase set of the relationship name, the biomedical article, and the second prompt word scenario, obtain multiple expressions of the synonym phrases in the synonym phrase set of the i-th relationship name in the biomedical article, and add them to the synonym phrase set of the i-th relationship name; and update the synonym phrase set of the i-th relationship name;
[0111] In this embodiment, all the relationship names s in the synonym phrase set of the relationship name are traversed in a loop. ik , the relationship name s in the synonym phrase set of the relationship name ik, a known biomedical article D, a second prompt word scenario generation template, and the LLM are input together to return multiple expressions in biomedical articles and add them to the synonym phrase set of the relationship name, and update the synonym phrase set of the relationship name; k represents the index number of the synonym phrase of the i-th relationship name; s ik represents the kth synonym phrase in the set of synonym phrases for the i-th relation name;
[0112] For example, a synonym phrase for the relationship name "{gene} is a marker of {disease}" is "{gene} mutation is related to {disease}"; after it is input into the LLM together with the second prompt word scene generation template, the LLM return value is another expression of the relationship name: "{disease} is related to {gene} mutation", which is incorporated into the synonym phrase set of the relationship name, and the synonym phrase set of the i-th relationship name is updated, S = {"{gene} is a marker of {disease}", "{disease} is related to {gene} mutation,...", "{gene} mutation is related to {disease}"}.
[0113] S23: Based on the updated synonym phrase set of the i-th relationship name, biomedical articles, the third prompt word scenario generation template and the large language model, the updated synonym phrase set of the i-th relationship name, the biomedical articles, and the third prompt word scenario generation template are input into the LLM, and the antonym phrases of the synonym phrases in the synonym phrase set of the i-th relationship name in the biomedical articles are obtained. The antonym phrases are returned and incorporated into the error-prone group TE to form an error-prone group.
[0114] For example, the first synonym phrase s in the synonym phrase set of the first relation name 11 = "{gene} induces {disease}", LLM return values are "{gene} does not induce {disease}", "{gene} inhibits {disease}", then TE1 = {"{gene} does not induce {disease}", "{gene} inhibits {disease}", ...}, TE1 represents the antonym phrase set of the synonym phrases in the synonym phrase set of the first relation name in biomedical articles;
[0115] S24: Based on a preset entity type set in the biomedical field, the head entity type and the tail entity type of the kth relationship name in the updated synonym phrase set of the ith relationship name are replaced with entity types in the entity type set respectively, and the first error-prone relationship name of the kth relationship name in the updated synonym phrase set is obtained to form a first error-prone name group of the kth relationship name in the updated synonym phrase set, and a template and a large language model are generated according to the fourth prompt word scenario to determine whether the first error-prone relationship name appears in the biomedical article. If the first error-prone relationship name appears in the biomedical article, the first error-prone relationship name is added to the error-prone group of the kth relationship name in the updated synonym phrase set, and the error-prone group of the kth relationship name in the updated synonym phrase set is updated;
[0116] Specifically, according to the updated synonym phrase set of the relationship name, loop through the kth relationship name s in the updated synonym phrase set of the i-th relationship name ik , will s ik The head entity type t in head Replaced with the preset entity type set T in the biomedical field is not equal to t head Other entity types, and ik The tail entity type t in tail Replaced with entity type set T not equal to t tail Other entity types are combined into a candidate error-prone name group, and the fourth prompt word scenario generates a template. Input LLM, return 1 or 0, and put the candidate error-prone name with return value = 1 into the error-prone group. For example: the relationship name s ik = "{gene} is a marker for {disease}", where {gene} is called the head entity type t head , {disease} is called the tail entity type in the relationship name t tail , the entity type set T = {"gene", "disease", "chemical"}, and the composition of the alternative error-prone name set = {"{compound} is a marker for {disease}", "{compound} is a marker for {gene}", "{compound} is a marker for {compound}", "{disease} is a marker for {compound}", "{disease} is a marker for {disease}", "{disease} is a marker for {gene}", "{gene} is a marker for {compound}"}. Input LLM, return value = {1,0,0,0,0,0,0,0}. TE i ={"{compound} is a marker for {disease}"}. TE ikrepresents the antonym phrase set of the kth synonym phrase in the updated synonym phrase set of the i-th relation name in the biomedical article;
[0117] S25: Based on a preset set of entity types in the biomedical field, the head entity type and the tail entity type of the kth relationship name in the updated synonym phrase set of the relationship name are exchanged to obtain a second error-prone relationship name of the kth relationship name in the updated synonym phrase set to form a second error-prone name group of the kth relationship name in the updated synonym phrase set, and a template and a large language model are generated according to the biomedical article and the fourth prompt word scenario to determine whether the second error-prone relationship name appears in the biomedical article. If the second error-prone relationship name appears in the biomedical article, the second error-prone relationship name is added to the error-prone group of the kth relationship name in the updated synonym phrase set, and the error-prone group of the kth relationship name in the updated synonym phrase set is updated for the second time;
[0118] Specifically, the head entity type t of the kth relationship name in the updated synonym phrase set of the relationship name is head and the tail entity type t tail Exchange to form a second error-prone name group, generate a template with the known document D and the fourth prompt word scenario, input LLM, return 1 or 0, put the alternative error-prone name with return value = 1 into the error-prone group, and update the error-prone group for the second time.
[0119] For example: Relationship name s ik = "{gene} affects {disease}", where {gene} is called the head entity type t head , {disease} is called the tail entity type in the relationship name t tail , forming a group of alternative error-prone names = {"{disease} affects {gene}"}. Input LLM, return value = {1}. Then TE ik ={“{Disease} affects {gene}”,..}.
[0120] S26: according to a preset set of relationship names in the biomedical field, obtaining a third error-prone relationship name of the same entity type as the k-th relationship name in the updated synonym phrase set of the relationship name, and generating a template and a large language model according to the biomedical article and the fourth prompt word scenario, determining whether the third error-prone relationship name appears in the biomedical article, and if the third error-prone relationship name appears in the biomedical article, adding the third error-prone relationship name to the error-prone group of the k-th relationship name in the updated synonym phrase set, and performing a third update on the error-prone group of the k-th relationship name in the updated synonym phrase set to obtain a final error-prone group of the k-th relationship name in the updated synonym phrase set of the i-th relationship name;
[0121] Specifically, all relationship names in the relationship name in the updated synonym phrase set of the relationship name are looped through, and the third error-prone relationship names with the same entity type are found to form an alternative error-prone relationship name combination, which is generated with the biomedical article and the fourth prompt word scene, and the error-prone group is updated for the third time to obtain the final error-prone group of the i-th relationship name;
[0122] For example, for the entity names "{Chemical} Enhances {Gene} Expression," "{Chemical} Enhances {Gene} Activity," "{Chemical} Enhances {Gene} Transport," and "{Chemical} Enhances {Gene} Metabolism," since their entity types are both {Chemical} and {Gene}, a potential error-prone name combination is formed: {"{Chemical} Enhances {Gene} Expression," "{Chemical} Enhances {Gene} Activity," "{Chemical} Enhances {Gene} Transport," and "{Chemical} Enhances {Gene} Metabolism"}. Inputting LLM returns a value of {1,1,1,1}. The third error-prone group then becomes {"{Chemical} Enhances {Gene} Expression," "{Chemical} Enhances {Gene} Activity," "{Chemical} Enhances {Gene} Transport," "{Chemical} Enhances {Gene} Metabolism," ...}
[0123] The first relationship name r1 = "{gene} is a marker for {disease}", the relationship between the synonym group and the error-prone group is shown in Table 1.
[0124] Table 1 Examples of the relationship between synonym group S1 and error-prone group TE1
[0125]
[0126]
[0127] S27: Obtain a prompt word scenario group according to a preset set of relationship names in the biomedical field, an updated set of synonym phrases of the relationship names, and a final error-prone group of the relationship names.
[0128] Specifically, the prompt word scene group P in this embodiment is P = {R, S, TE}, where S is a positive example, that is, a set of correct relationship names, which is a set of synonym phrases of the updated relationship names; TE is a negative example, that is, a set of incorrect relationship names, which is the final error-prone group.
[0129] S3: Based on the entity combination, prompt word scene group, relationship extraction prompt word template and large language model in the biomedical article, the initial relationship extraction prompt words are obtained; the initial relationship extraction prompt words include multiple prompt word units corresponding to the kth relationship name in the updated synonym phrase set of the i-th relationship name; wherein, the prompt word unit corresponding to the kth relationship name in the updated synonym phrase set of the i-th relationship name includes: the kth relationship name in the updated synonym phrase set of the i-th relationship name, and the error-prone name in the final error-prone group of the kth relationship name in the updated synonym phrase set of the i-th relationship name.
[0130] Specifically, the prompt word scene group, the relationship extraction prompt word template, the biomedical article D, and the entity combination EP involved in the biomedical article D are combined to form an initial relationship extraction prompt word.
[0131] For example, R, S, and TE of the prompt word scene group P are shown in Table 1. The entity combination EP is shown in Table 2.
[0132] Table 2 Example of entity combination
[0133]
[0134]
[0135] In one embodiment of the present invention, taking the entity combination of the first relationship name r1 = "{gene} is a marker of {disease}" as an example, the initial relationship extraction prompt words are obtained:
[0136] The biomedical articles are as follows:
[0137] Carvedilol regulates the expression of hypoxia-inducible factor-1α and vascular endothelial growth factor in a rat model of volume-overload heart failure.
[0138] Background: Beta-blockers have been increasingly recognized as a beneficial therapy for congestive heart failure. Hypoxia-inducible factor-1α (HIF-1α) is tightly regulated in the ventricular myocardium. However, the expression of HIF-1α in chronic heart failure due to volume overload and its changes after beta-blocker treatment remain unclear.
[0139] Methods and Results: To test the hypothesis that HIF-1α plays a role in myocardial failure caused by volume overload, we induced volume-overload heart failure in adult Sprague-Dawley rats by creating aortocaval fistula for 4 weeks. Carvedilol was then administered postoperatively at a dose of 50 mg / kg body weight / day. Compared with the sham group (heart weight / body weight ratio, 2.6 ± 0.3), the heart weight / body weight ratio increased to 3.9 ± 0.7 in the aortocaval fistula group (P < 0.001). Left ventricular end-diastolic diameter increased from 6.5 ± 0.5 mm to 8.7 ± 0.6 mm (P < 0.001). In the aortocaval fistula group, heart weight and ventricular dimensions returned to baseline following carvedilol treatment. Western blot results showed that HIF-1α, vascular endothelial growth factor (VEGF), and brain natriuretic peptide (BNP) protein expression was upregulated in the aortocaval fistula group, while nerve growth factor β (NGF-β) expression was downregulated. Real-time polymerase chain reaction (PCR) results showed that HIF-1α, VEGF, and BNP mRNA expression increased, while NGF-β mRNA expression decreased, in the aortocaval fistula group. After carvedilol treatment, protein and mRNA expression of HIF-1α, VEGF, BNP, and NGF-β returned to baseline levels. Increased immunohistochemical staining for HIF-1α, VEGF, and BNP was observed in the ventricular myocardium of the aortocaval fistula group, and carvedilol restored this staining to normal.
[0140] Conclusion: In a rat model of volume-overload heart failure, HIF-1a and VEGF mRNA and protein expressions are upregulated. Carvedilol treatment is associated with the reversal of HIF-1a and VEGF dysregulation in failing ventricular myocardium.
[0141] Extract prompt word template based on relationship:
[0142] You are a biomedical expert. Based on the above biomedical article, please judge whether "HIF-1α is a marker for heart failure" is true based on the following scenarios. Give your reasons according to the scenarios. The initial relationship extraction prompt word pt0 is as follows:
[0143] 1. If HIF-1α is a marker for heart failure, answer 1.
[0144] Otherwise, if heart failure is not associated with HIF-1α mutations, answer E1.1.
[0145] Otherwise, if the chemical is associated with HIF-1α mutation, answer E1.2.
[0146] 2. Otherwise, if HIF-1α mutations are associated with heart failure, answer 2.
[0147] Otherwise, if HIF-1α mutations are not associated with heart failure, answer E2.1.
[0148] Otherwise, if the HIF-1α mutation is chemically related, answer E2.2.
[0149] 3. Otherwise, if heart failure is induced by HIF-1α, answer 3.
[0150] Otherwise, if heart failure is not induced by HIF-1α, answer E3.1.
[0151] 4. Otherwise, if HIF-1α induces heart failure, answer 4.
[0152] Otherwise, if HIF-1α does not induce heart failure, answer E4.1.
[0153] Otherwise, if HIF-1α inhibits heart failure, answer E4.2
[0154] 99. None of the above, please answer: 99
[0155] Do not output any other information.
[0156] Among them, "HIF-1α is a marker for heart failure" is an element in the updated synonym phrase set of the relationship name, and "heart failure is not related to HIF-1α mutation" and "chemicals are related to HIF-1α mutation" are elements in the final error-prone group corresponding to the elements in the updated synonym phrase set of the relationship name.
[0157] S4: obtaining a relationship extraction prompt phrase according to the initial relationship extraction prompt word;
[0158] Preferably, the method for obtaining the relationship extraction prompt phrase is as follows:
[0159] S41: using the initial relationship extraction prompt word as the first relationship extraction prompt word in the relationship extraction prompt word group;
[0160] S42: Delete the prompt word units corresponding to the first to kth relationship names in the updated synonym phrase set of the i-th relationship name from the initial relationship extraction prompt words to obtain the k+1th relationship extraction prompt word; and obtain the relationship extraction prompt phrase; wherein k=1, ..., K-1;
[0161] Among them, k is the index of the relationship name in the updated synonym phrase set, and is also the index number of the relationship extraction prompt word unit in the initial relationship extraction prompt word; K is the total number of relationship names in the updated synonym phrase set, and is also the total number of relationship extraction prompt word units in the initial relationship extraction prompt word.
[0162] Specifically, in this embodiment, the initial relationship extraction prompt word is represented as pt0, which is the first relationship extraction prompt word in the relationship extraction prompt phrase group; pt0 includes K relationship extraction prompt word units corresponding to the relationship names in the synonym phrase set after the relationship names are updated;
[0163] Delete the prompt word units corresponding to the 1st to kth relationship names in the synonym phrase set of the updated relationship name from the initial relationship extraction prompt word. That is, delete the first relationship name in pt0 and its corresponding error-prone name. The remaining part is the second relationship extraction prompt word pt1 in the relationship extraction prompt phrase. pt1 includes K-1 relationship extraction prompt word units corresponding to the relationship names in the updated synonym phrase set of the relationship name. Using the same method, the other relationship extraction prompt words in the relationship extraction prompt phrase can be obtained. Finally, the relationship extraction prompt phrase is obtained.
[0164] For example, delete the first relationship name s1 from pt0, as well as E1.1 and E1.2 corresponding to s1's subscript 1. This means deleting the first relationship extraction prompt word unit and generating a new prompt word pt1. Similarly, a relationship extraction prompt phrase group PT1 = {pt0, pt1, pt2, ...} is generated. Here, s1 is the first relationship name in the updated synonym phrase group for the relationship name; E1.1 and E1.2 are both error-prone names in the final error-prone group corresponding to the first relationship name in the updated synonym phrase group for the relationship name; pt1 is the second relationship extraction prompt word in the relationship extraction prompt phrase group, and pt2 is the third relationship extraction prompt word in the relationship extraction prompt phrase group.
[0165] S5: Obtain the score value of the relationship extraction prompt word in the relationship extraction prompt phrase group; determine the optimal relationship extraction prompt word in the relationship extraction prompt phrase group, and complete the automatic generation and evaluation of the prompt word.
[0166] Specifically, according to the relationship extraction prompt phrases, a prompt word combination optimization algorithm based on a greedy strategy is used to evaluate the contribution of each relationship extraction prompt word scenario and the optimal prompt word in the relationship extraction prompt phrases.
[0167] Specifically, traverse the relationship extraction prompt phrase PT of the i-th relationship name iEach element pt0, pt1, pt2, ... in the LLM is input to obtain the result a, where a is the predicted relation extraction prompt word. If the first letter of the value of a is E, that is, the predicted result is a mistake-prone name in the mistake-prone group, then the prediction result is a negative example, that is, LLM determines that the entity combination does not have the relation name. Otherwise, it is a positive example, that is, LLM determines that the entity combination does have the relation name. Save the entity combination of the positive example as the prediction result
[0168] Next, the LLM prediction results With the marked results Compare. If the predicted result exists in the labeled result, TP (True Positive) is +1. If the predicted result does not exist in the labeled result, FP (False Positive) is +1. If the labeled result does not exist in the predicted result, FN (False Negative) is +1.
[0169] Based on this, the evaluation indicators of the relationship extraction prompt words in each relationship extraction prompt phrase are obtained, including: precision rate P (precision_score), recall rate Recall (recall_score), and F1 score value (f1_score) are calculated as follows:
[0170] P=TP / (TP+FP)
[0171] Recall = TP / (TP+FN)
[0172] F1=2*(P*Recall) / (P+Recall)
[0173] Specifically, the updated synonym phrase set S for relation names aims to improve recall and reduce FN. The third updated error-prone phrase TE aims to improve P and reduce FP. After deleting one or more cue word scenarios from S and TE corresponding to each cue word unit, the greater the decrease in F1, the greater the contribution of the deleted cue word scenario. This is used to evaluate the effectiveness of the cue word scenario.
[0174] Preferably, this embodiment is based on a greedy strategy combination optimization algorithm, which aims to improve the solution efficiency by gradually narrowing the search space, thereby obtaining the optimal relationship extraction prompt word pt for a certain relationship name. 最优 i .
[0175] Specifically, when a prompt word scene group P = {R, S, TE} is given, each element of S and TE has a decision coefficient C = {0, 1}, and there are n and m elements in S and TE respectively, the relation extraction prompt phrase PT is generated by P and C, and is used as the input of LLM to find the relation extraction prompt phrase pt that maximizes F1. k , pt k∈PT, that is:
[0176]
[0177] Among them, F1(pt k ) represents the kth relation extraction prompt word pt in the relation extraction prompt phrase k F1 score; R represents the set of relation names; c k represents the coefficient of the positive example in the k-th relation extraction prompt word; c j represents the coefficient of the counterexample in the kth relation extraction prompt word; C represents the decision coefficient; s k represents the positive example in the k-th relation extraction prompt word; te j represents the counterexample in the k-th relation extraction prompt word; TE represents the final error-prone group; S represents the set of synonym phrases of the updated relation name;
[0178] Specifically, the algorithm of this embodiment is designed as follows:
[0179] S51: First, initialize: set the coefficients of the first relationship name and the error-prone name of the first relationship name in the initial relationship extraction prompt word to 1, and obtain the F1 score F1 of the initial relationship extraction prompt word baseline .
[0180] S52: Set the coefficients of the first relationship name and the error-prone name of the first relationship name in the initial relationship extraction prompt word to 0, obtain a new prompt word, and obtain the F1 score F1 of the new prompt word new ; Get the contribution V of the first relation extraction prompt word unit in the initial relation extraction prompt word i , the formula used is as follows:
[0181] V i =F1 baseline -F1 new .
[0182] S53: If V i <0, remove the first relation extraction prompt word unit in the initial relation extraction prompt word, including the first relation name and the corresponding error-prone name;
[0183] If V i ≥0, then the coefficient of the first relation name in the initial relation extraction prompt word is restored to 1, and the coefficient of the non-j-th error-prone name corresponding to the first relation name is restored to 1. At this time, the coefficient of the j-th error-prone name corresponding to the first relation name is still 0; obtain the second new prompt word pt j ′, and obtain the second new prompt word pt j ′’s F1 score newj, and then obtain the contribution V of the j-th error-prone name corresponding to the first relation name in the initial relation extraction prompt word j :The formula used is as follows:
[0184] V j =F1 baseline -F1 newj .
[0185] S54: If V j <0, remove the j-th error-prone name corresponding to the first relationship name in the initial relationship extraction prompt; otherwise, restore the coefficient of the j-th error-prone name of the first relationship name in the initial relationship extraction prompt to 1; update the initial relationship extraction prompt to obtain the updated relationship extraction prompt;
[0186] S55: Obtain the updated F1 score of the relation extraction prompt word;
[0187] If the F1 score of the updated relation extraction prompt word is greater than the F1 score of the initial relation extraction prompt word baseline The F1 score in ; then execute S56,
[0188] S56: Based on the updated relation extraction prompt words, re-execute S51-S55 until the F1 score of the updated relation extraction prompt phrase no longer decreases. At this time, the updated relation extraction prompt word is the optimal relation extraction prompt phrase.
[0189] Specifically, in this embodiment, after traversing all the relationship extraction prompt word units in the first relationship extraction prompt word in the relationship extraction prompt word group, if there is still no F1 score that continues to decrease, the F1 score calculation is continued based on the second relationship extraction prompt word in the relationship extraction prompt word group until the optimal relationship extraction prompt word is obtained.
[0190] This embodiment focuses on document-level relationship extraction based on a general large language model in the biomedical field, and is particularly aimed at the automatic generation and evaluation of relationship extraction prompt words. It has the following beneficial effects:
[0191] Reduced data dependency: The need for large amounts of labeled data is reduced, which reduces the cost and time of data preparation and makes relationship extraction technology easier to apply to areas where data is scarce.
[0192] Improved efficiency: The relationship extraction process can be performed faster without the need to train / fine-tune large models, thus speeding up the development cycle.
[0193] Enhanced generalization capabilities: able to adapt to different document types and domains.
[0194] Improve accuracy: Improve the accuracy of relationship extraction.
[0195] Reduced manual intervention: Automated relationship extraction reduces the need for manual intervention, lowers labor costs, and increases processing speed.
[0196] The combined effects of these advantages make the document-level relationship extraction method based on a universal large model in this embodiment valuable in a variety of application scenarios. It automatically and vividly generates high-quality prompt words based on different biomedical scenarios, and accurately evaluates these prompt words through a scientific evaluation system. This improves the effectiveness and reliability of relationship extraction using large language models in the biomedical field, providing stronger support for biomedical research and practice.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field, characterized in that: The steps include: S1: Based on the prompt word generation principle, multiple prompt word scene generation templates and relationship extraction prompt word templates are established; S2: Generate templates and large language models based on a preset set of relationship names, biomedical articles, and prompt word scenarios in the biomedical field to obtain a prompt word scenario group; S3: Obtain initial relation extraction prompt words based on entity combinations, prompt word scene groups, relation extraction prompt word templates, and large language models in biomedical articles; The initial relationship extraction prompt words include a plurality of prompt word units corresponding to the kth relationship name in the updated synonym phrase set of the i-th relationship name; wherein the prompt word unit corresponding to the kth relationship name in the updated synonym phrase set of the i-th relationship name includes: the kth relationship name in the updated synonym phrase set of the i-th relationship name, and the error-prone name in the final error-prone group of the kth relationship name in the updated synonym phrase set of the i-th relationship name; wherein i is the index of the relationship name in the relationship name set in the biomedical field; and k is the index of the relationship name in the updated synonym phrase set. S4: obtaining a relationship extraction prompt phrase according to the initial relationship extraction prompt word; S5: Obtain the score values of the relationship extraction prompt words in the relationship extraction prompt word group to determine the optimal relationship extraction prompt word in the relationship extraction prompt word group, and complete the automatic generation and evaluation of the prompt words.
2. The method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field according to claim 1, characterized in that: The method for obtaining the prompt word scene group is as follows: S21: Generate a template and a large language model based on the i-th relationship name in a preset relationship name set within the biomedical field, a biomedical article, and a first prompt word scenario, obtain synonyms or synonym phrases of the i-th relationship name that appear in the biomedical article, and obtain a synonym phrase set for the i-th relationship name; wherein i is the index number of the relationship name in the relationship name set within the biomedical field, i=1,…,I; and I is the total number of relationship names in the relationship name set within the biomedical field; S22: Generate a template and a large language model based on the synonym phrase set of the relationship name, the biomedical article, and the second prompt word scenario, obtain multiple expressions of the synonym phrases in the synonym phrase set of the i-th relationship name in the biomedical article, and add them to the synonym phrase set of the i-th relationship name; and update the synonym phrase set of the i-th relationship name; S23: generating a template and a large language model based on the updated synonym phrase set of the i-th relationship name, the biomedical article, and the third prompt word scenario, obtaining antonym phrases of the synonym phrases in the synonym phrase set of the i-th relationship name in the biomedical article to form an error-prone group; S24: Based on a preset entity type set in the biomedical field, the head entity type and the tail entity type of the kth relationship name in the updated synonym phrase set of the ith relationship name are replaced with entity types in the entity type set, respectively, to obtain a first error-prone relationship name of the kth relationship name in the updated synonym phrase set, so as to form a first error-prone name group of the kth relationship name in the updated synonym phrase set, and a template and a large language model are generated according to the biomedical article and the fourth prompt word scenario, to determine whether the first error-prone relationship name appears in the biomedical article; If the first error-prone relationship name appears in the biomedical article, the first error-prone relationship name is added to the error-prone group of the kth relationship name in the updated synonym phrase set, and the error-prone group of the kth relationship name in the updated synonym phrase set is updated; S25: Based on a preset entity type set in the biomedical field, swapping the head entity type and the tail entity type of the k-th relationship name in the updated synonym phrase set of the relationship name, obtaining a second easily-mistaken relationship name of the k-th relationship name in the updated synonym phrase set, so as to form a second easily-mistaken name group of the k-th relationship name in the updated synonym phrase set, and generating a template and a large language model based on the biomedical article and the fourth prompt word scenario, to determine whether the second easily-mistaken relationship name appears in the biomedical article; If a second error-prone relationship name appears in the biomedical article, the second error-prone relationship name is added to the error-prone group of the kth relationship name in the updated synonym phrase set, and the error-prone group of the kth relationship name in the updated synonym phrase set is updated for the second time; S26: Based on a preset set of relationship names in the biomedical field, obtaining a third error-prone relationship name having the same entity type as the kth relationship name in the updated synonym phrase set of the relationship name, and generating a template and a large language model based on the biomedical article and the fourth prompt word scenario, to determine whether the third error-prone relationship name appears in the biomedical article; If a third error-prone relationship name appears in the biomedical article, the third error-prone relationship name is added to the error-prone group of the k-th relationship name in the updated synonym phrase set, and the error-prone group of the k-th relationship name in the updated synonym phrase set is updated for the third time to obtain a final error-prone group of the k-th relationship name in the updated synonym phrase set for the i-th relationship name; S27: Obtain the prompt word scenario group of the i-th relationship name based on the preset relationship name set in the biomedical field, the updated synonym phrase set of the i-th relationship name, and the final error-prone group of the k-th relationship name in the updated synonym phrase set of the i-th relationship name.
3. The method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field according to claim 2, characterized in that: The method for obtaining the relationship extraction prompt phrase is as follows: S41: using the initial relationship extraction prompt word as the first relationship extraction prompt word in the relationship extraction prompt word group; S42: Delete the prompt word units corresponding to the first to kth relationship names in the updated synonym phrase set of the i-th relationship name from the initial relationship extraction prompt words to obtain the k+1th relationship extraction prompt word; and obtain the relationship extraction prompt phrase; wherein k=1, ..., K-1; Wherein, k is the index of the relationship name in the updated synonym phrase set; K is the total number of relationship names in the updated synonym phrase set.
4. The method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field according to claim 3, characterized in that: The method used to determine the optimal relationship extraction prompt word in the relationship extraction prompt word group is as follows: S51: Set the coefficients of the first relationship name and the error-prone name of the first relationship name in the initial relationship extraction prompt word to 1, and obtain the F1 score F1 of the initial relationship extraction prompt word baseline ; S52: Set the coefficients of the first relationship name and the error-prone name of the first relationship name in the initial relationship extraction prompt word to 0, obtain a new prompt word, and obtain the F1 score F1 of the new prompt word new ; Get the contribution V of the first relation extraction prompt word unit in the initial relation extraction prompt word i , the formula used is as follows: V i =F1 baseline -F1 new ; S53: If V i <0, remove the first relation extraction prompt word unit in the initial relation extraction prompt word; If V i ≥0, then the coefficient of restoring the first relation name in the initial relation extraction prompt word is 1, and the coefficient of restoring the non-jth error-prone name corresponding to the first relation name is 1, j = 1, ..., J; get the second new prompt word pt j ′, and obtain the second new prompt word pt j ′’s F1 score newj , and then obtain the contribution V of the j-th error-prone name corresponding to the first relation name in the initial relation extraction prompt word j :The formula used is as follows: V j =F1 baseline -F1 newj ; S54: If V j <0, remove the j-th error-prone name corresponding to the first relationship name in the initial relationship extraction prompt; otherwise, restore the coefficient of the j-th error-prone name of the first relationship name in the initial relationship extraction prompt to 1; update the initial relationship extraction prompt to obtain the updated relationship extraction prompt; j is the index of the error-prone name; J is the total number of error-prone names; S55: Obtain the updated F1 score of the relation extraction prompt word; If the F1 score of the updated relation extraction prompt word is greater than the F1 score of the initial relation extraction prompt word baseline The F1 score in ; then execute S56, S56: Based on the updated relation extraction prompt words, re-execute S51-S55 until the F1 score of the updated relation extraction prompt phrase no longer decreases. At this time, the updated relation extraction prompt word is the optimal relation extraction prompt phrase.
5. The method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field according to claim 1, characterized in that: The relationship extraction prompt word template is established as follows: {Biomedical Articles} You are a biomedical expert. Based on the above biomedical article, please judge whether "{relationship name}" is true according to the following prompt words and scenarios. Please give your reasons according to the scenarios: {Prompt word scene group} 99. None of the above, please answer: 99 Do not output any other information.
6. The method for automatically generating and evaluating prompt words for relationship extraction in the biomedical field according to claim 1, characterized in that: The multiple prompt word scene generation templates include: a first prompt word scene generation template, a second prompt word scene generation template, a third prompt word scene generation template, and a fourth prompt word scene generation template; The first prompt word scene generation template is established as follows: {Biomedical Articles} You are a biomedical expert. Give the synonyms and synonym groups of {relationship name} that appear in the above biomedical articles. Output format: Synonym 1 | Synonym 2 | Synonym 3 | ... No other information is output; The second prompt word scene generation template is established as follows: {Biomedical Articles} You are a biomedical expert. Give multiple forms of expression of {relationship name} in the above biomedical article, including passive voice and nominal forms. For example, "{gene} induces {disease}" can be expressed in the following forms: "{disease} is induced by {gene}" or "{gene} is the inducing factor of {disease}"; Output format: expression1|expression2|expression3|... No other information is output; The third prompt word scene generation template is established as follows: {Biomedical Articles} You are a biomedical expert. Give all the antonyms of {relationship name} that appear in the above biomedical article. The fourth prompt word scene generation template is established as follows: {Biomedical Articles} You are a biomedical expert. Determine whether {relationship name} appears in the above biomedical article. If so, reply 1. If not, reply 0. No other information is output.