Method and system for automatically generating medicine safety information using generative ai
Generative AI automates drug safety reporting by extracting and classifying drug safety information, addressing human error and inconsistency, and enhancing efficiency and reliability.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-04-02
AI Technical Summary
Current drug safety information reporting in the pharmaceutical industry is prone to human errors, lacks consistency, and is time-consuming due to reliance on subjective human analysis and review of drug adverse events.
A method and system using Generative AI to automatically generate drug safety reports by extracting keywords, calculating similarities, summarizing relevant information, and classifying it into standardized formats, with self-verification to ensure accuracy and consistency.
Automated generation of drug safety reports reduces human error, ensures high accuracy and consistency, and saves time while maintaining report quality.
Smart Images

Figure KR2025014304_02042026_PF_FP_ABST
Abstract
Description
Method and System for Automatic Generation of Drug Safety Information Using Generative AI
[0001] The present invention relates to a method and system for automatically generating reports using Generative AI, and more specifically, to a method and system for automatically generating reports on adverse drug reactions using Generative AI.
[0002] A drug adverse event report is a report prepared to report information regarding adverse events (side effects) that occur after the administration of drugs, etc.
[0003] Figure 1 is a diagram provided in the description of a conventional drug side effect reporting processor.
[0004] Referring to Figure 1, in the current pharmaceutical industry, the Pharmacovigilance (PV) manager of a pharmaceutical company collects safety information one by one, reviews and analyzes it, and then directly writes standardized domestic and international drug adverse event reports.
[0005] In this process, since safety information is collected and analyzed based on the subjective standards of the person in charge or the company, there is a risk that the accuracy and consistency of the results may be compromised.
[0006] In other words, as the task of analyzing various document formats and content, identifying errors, extracting necessary information, and writing and reviewing drug side effect reports was previously performed mostly by individuals (persons in charge), problems such as the following may exist.
[0007] 1. Human Error: When humans directly review drug safety information, errors may occur due to mistakes. For example, this may include typographical errors or reports of adverse drug reactions resulting from incorrect interpretation of information.
[0008] 2. Lack of Consistency: Since the standards for reviewing and writing drug safety clue information may vary from person to person, there is a problem of inconsistency in the results.
[0009] 3. Time consumption: Reviewing and writing drug safety clue information takes a lot of time, which can reduce work efficiency.
[0010] Accordingly, to resolve the aforementioned problems, it is necessary to explore automation methods for the preparation and / or review of drug adverse event reports.
[0011] The present invention has been devised to solve the aforementioned problems, and the objective of the present invention is to provide a method and system for automatically generating drug safety information using generative AI, which can automatically generate drug side effect reports by applying various domain data regarding the safety of specific drugs to a generative AI model.
[0012] A method for automatically generating drug safety information using generative AI according to an embodiment of the present invention for achieving the above objective comprises: a step in which a system receives a user prompt containing information on the safety of a specific drug and information on the approval details of the drug; a step in which the system extracts keywords based on pre-set keyword extraction criteria using a topic modeling model; a step in which the system calculates a similarity between the user prompt and the extracted keywords; a step in which the system extracts and summarizes information on the safety of a drug having a similarity greater than or equal to a pre-set threshold; a step in which the system classifies the summarized information on the safety of the drug according to a pre-set drug side effect report format; and a step in which the system generates a drug side effect report for the drug using the classified information.
[0013] And information regarding the safety of the drug may include at least one of structured or unstructured data in text form, structured or unstructured data in image form, and structured or unstructured data in voice form regarding the safety of the drug.
[0014] In addition, the pre-established keyword extraction criteria may include at least one of the ingredients, dosage, indications (efficacy and effect), and target patients.
[0015] And in the step of calculating similarity, the similarity between the user prompt and the extracted keywords can be calculated by utilizing Latent Dirichlet Allocation (LDA) topic modeling or semantic network analysis techniques.
[0016] In addition, the summarizing step can extract and summarize information regarding the safety of drugs having a similarity level greater than a preset threshold based on a Multimodal Large Language Model.
[0017] And the classification step can be performed to match summarized information on the safety of the drug, based on a RAG (Retrieval-Augmented Generation) model linked to a multimodal language model, to each item of a pre-established drug adverse effect report form.
[0018] In addition, the pre-established drug adverse event report form may be the drug adverse event report form provided by the Korean Association of Adverse Event Reporting Systems (KAERS) or the International Council of Organizations for Medical Sciences (CIOMS).
[0019] And, according to one embodiment of the present invention, a method for automatically generating drug safety information using generative AI may further include the step of, when the classification of information regarding the safety of summarized drugs is completed, the system performing self-verification for each matching item of the classified information.
[0020] In addition, the step of performing self-verification involves performing self-verification for each item that matches the classified information, and if the verification score is below a preset threshold, the classification of information regarding the safety of the summarized drug is re-performed, provided that the preset threshold of the verification score may be set differently for each item that matches the classified information depending on whether it is essential information.
[0021] Meanwhile, an automatic drug safety information generation system utilizing generative AI according to another embodiment of the present invention comprises: an input unit for inputting a user prompt containing information on the safety of a specific drug and information on the approval details of the drug; and a processor that utilizes a topic modeling model to extract keywords based on preset keyword extraction criteria, calculates the similarity between the user prompt and the extracted keywords, extracts and summarizes information on the safety of a drug having a similarity greater than or equal to a preset threshold, classifies the summarized information on the safety of the drug according to a preset drug side effect report format, and generates a drug side effect report for the drug using the classified information.
[0022] As described above, according to the embodiments of the present invention, by automating the task of writing drug side effect reports, report writing time can be saved, high accuracy and reliability can be ensured by minimizing human error, and reports of consistent quality can be obtained.
[0023] Figure 1 is a drawing provided in the description of a conventional drug side effect reporting processor,
[0024] FIG. 2 is a drawing provided in the description of the configuration of an automatic drug safety information generation system utilizing generative AI according to an embodiment of the present invention.
[0025] FIG. 3 is a drawing provided for a more detailed configuration description of the input unit illustrated in FIG. 2.
[0026] FIG. 4 is a drawing provided for a more detailed configuration description of the processor shown in FIG. 2.
[0027] FIG. 5 is a flowchart provided for explaining a method for automatically generating drug safety information using generative AI according to an embodiment of the present invention, and
[0028] FIG. 6 is a drawing provided in the description of a drug side effect report created using generative AI according to one embodiment of the present invention.
[0029] The present invention will be described in more detail below with reference to the drawings. To clearly explain the invention, parts unrelated to the description have been omitted from the drawings, and in the drawings, the width, length, thickness, etc., of the components may be exaggerated for convenience.
[0030] FIG. 2 is a drawing provided for the configuration description of a system for automatically generating drug safety information using generative AI according to one embodiment of the present invention, and FIG. 3 is a drawing provided for the more detailed configuration description of the input unit shown in FIG. 2.
[0031] The automatic drug safety information generation system utilizing generative AI according to the present embodiment (hereinafter collectively referred to as the "system") is provided to automatically generate drug side effect reports by applying various domain data regarding the safety of a specific drug to a generative AI model.
[0032] To this end, the system may include an input unit (100), a processor (200), and a storage unit (300).
[0033] The input unit (100) receives information regarding the safety of a specific drug (e.g., voluntary report, occasional report, periodic report, literature, etc.) and information regarding the approval details of the drug, and receives a user prompt (e.g., user requirements for writing a drug side effect report, etc.) to automatically write a drug side effect report according to a pre-set drug side effect report form.
[0034] Here, information regarding the safety of the drug may include at least one of structured or unstructured data in text form, structured or unstructured data in image form, and structured or unstructured data in voice form regarding the safety of the drug. Additionally, information regarding the approved specifications of the drug may be received via a user prompt.
[0035] To this end, the input unit (100) may include a drug safety information input unit (110) equipped with a communication module connected to a network to receive information regarding the safety of a specific drug, and a user prompt input unit (120) that receives a user prompt (e.g., information regarding the approval details of the drug) for automatically creating a drug side effect report according to a preset drug side effect report form and transmits it to a processor (200).
[0036] The storage unit (300) is provided to store programs and data necessary for the operation of the processor (200).
[0037] The processor (200) can process various domain data regarding the safety of a specific drug to a generative AI model to automatically generate a drug side effect report.
[0038] To this end, the generative AI model according to the present embodiment may be composed of a multimodal language model and a topic modeling model linked to the multimodal language model, an LDA (Latent Dirichlet Allocation) topic model that calculates similarity by performing LDA topic modeling or a semantic network analysis model that calculates similarity by utilizing a semantic network analysis technique and a RAG (Retrieval-ugmented Generation) model.
[0039] Specifically, the processor (200) can extract keywords based on preset keyword extraction criteria using a topic modeling model, calculate the similarity between the user prompt and the extracted keywords, and extract and summarize information on the safety of a drug having a similarity greater than or equal to a preset threshold.
[0040] And the processor (200) can classify information on the safety of summarized drugs according to a pre-set drug side effect report form and use the classified information to generate a drug side effect report for the drug.
[0041] Figure 4 is a drawing provided for a more detailed configuration description of the processor shown in Figure 2.
[0042] Referring to FIG. 4, the processor (200) may include a pre-training unit (210), a fine-tuning unit (220), a keyword extraction unit (230), a similarity calculation unit (240), a summary information generation unit (250), a search and classification unit (260), a classification result verification unit (270), and a report generation unit (280).
[0043] The pre-training unit (210) can perform a pre-training procedure for a generative AI model composed of a multimodal language model and a topic modeling model linked to the multimodal language model, an LDA topic model or a semantic network analysis model and a RAG model, based on information regarding the safety of a specific drug and information regarding the approval details of the drug.
[0044] That is, the pre-training unit (210) performs a pre-training procedure for a generative AI model to generate a drug side effect report for the drug that meets the requirements within the user prompt, based on at least one of the data among text-based structured data or unstructured data regarding the safety of the drug, image-based structured data or unstructured data, and voice-based structured data or unstructured data, and information regarding the approval details of the drug that is input (uploaded).
[0045] The fine-tuning unit (220) can perform a fine-tuning procedure so that, once the pre-training procedure is completed, the pre-trained generative AI model is specialized for each domain of data regarding the stability of the input drug (e.g., structured data or unstructured data in text form, structured data or unstructured data in image form, and structured data or unstructured data in voice form, etc.).
[0046] At this time, the fine-tuning unit (220) can perform a fine-tuning procedure for each data domain, for each multimodal language model and the topic modeling model, LDA topic model or semantic network analysis model and RAG model linked to the multimodal language model.
[0047] That is, the multimodal language model may be composed of a first multimodal language model fine-tuned to be specialized for structured data in text form, a second multimodal language model fine-tuned to be specialized for unstructured data in text form, a third multimodal language model fine-tuned to be specialized for structured data in image form and the structured data in image form and the corresponding caption data, a fourth multimodal language model fine-tuned to be specialized for unstructured data in image form and the corresponding caption data, a fifth multimodal language model fine-tuned to be specialized for converting structured data in voice form into text form and processing it, and a sixth multimodal language model fine-tuned to be specialized for converting unstructured data in voice form into text form and processing it.
[0048] This applies equally to topic modeling models, LDA topic models, semantic network analysis models, and RAG models linked to the multimodal language model, allowing for fine-tuning procedures to be performed for each data domain.
[0049] The keyword extraction unit (230) can extract keywords based on pre-set keyword extraction criteria by utilizing a topic modeling model in which a fine-tuning procedure has been completed.
[0050] Here, the topic modeling model can be implemented using KeyBERT, etc., and the pre-set keyword extraction criteria correspond to information regarding the approved indications of the drug, and specifically may include at least one of the ingredients, dosage, indications (efficacy and effect), and target patients.
[0051] That is, the keyword extraction unit (230) can extract keywords based on at least one criterion among ingredients, dosage, indications (efficacy and effect) and target patients by utilizing a topic modeling model in which a fine-tuning procedure has been completed.
[0052] The similarity calculation unit (240) can calculate the similarity between the user prompt and the extracted keywords by utilizing an LDA topic model or a semantic network analysis model that has completed a fine-tuning procedure.
[0053] The summary information generation unit (250) can extract and summarize information regarding the safety of a drug having a similarity greater than or equal to a preset threshold when the similarity between the user prompt and the extracted keyword is calculated through the similarity calculation unit (240).
[0054] Specifically, the summary information generation unit (250) can extract and summarize information regarding the safety of a drug having a similarity greater than a preset threshold based on a multimodal language model in which a fine-tuning procedure has been completed.
[0055] The search and classification unit (260) can perform search and classification operations to match information on the safety of summarized drugs based on the RAG model with the fine-tuning procedure completed to each item of the pre-set drug side effect report form.
[0056] The established drug adverse event report form may be the Korea Food and Drug Administration’s Adverse Event Reporting System (KAERS) or the drug adverse event report form provided by the Council for International Organization of Medical Sciences (CIOMS).
[0057] The classification result verification unit (270) can perform self-verification for each item that matches the classified information.
[0058] Specifically, the classification result verification unit (270) can perform self-verification for each item that matches the classified information, and if the verification score is less than a preset threshold, the search and classification work of information on the safety of the summarized medicine can be performed again.
[0059] In this case, the classification result verification unit (270) can set the pre-set threshold of the verification score differently for each item that matches the classified information, depending on whether it corresponds to the essential information, so that the entry of essential information is necessarily executed.
[0060] Here, essential information may include information on adverse events (side effects), information on suspected drugs, patient information, reporter information, etc.
[0061] Additionally, the classification result verification unit (270) can, even after the classification work is re-performed, if the verification score is below a preset threshold, allow the fine-tuning procedure of the RAG model to be re-performed so as to be specialized for items where the verification score is below the preset threshold, and utilize the fine-tuned RAG model to be specialized for items where the verification score is below the preset threshold, thereby allowing the search and classification work of information on the safety of summarized drugs to be re-performed.
[0062] The report generation unit (280) can generate a drug side effect report for the drug by utilizing information classified based on a multimodal language model in which a fine-tuning procedure has been completed.
[0063] FIG. 5 is a flowchart provided in the description of a method for automatically generating drug safety information using generative AI according to an embodiment of the present invention, and FIG. 6 is a diagram provided in the description of a drug side effect report generated using generative AI according to an embodiment of the present invention.
[0064] The method for automatically generating drug safety information using generative AI according to the present embodiment can be executed by the system described above with reference to FIGS. 2 to 4.
[0065] Specifically, after the system performs pre-training and fine-tuning procedures for a generative AI model composed of a multimodal language model, a topic modeling model linked to the multimodal language model, an LDA topic model or a semantic network analysis model, and a RAG model, information regarding the safety of a specific drug (e.g., voluntary report, ad-hoc report, periodic report, literature, etc.) is received (uploaded) (S510) and information regarding the approval details of the drug is input (S520), the system can apply the received / input information by applying it to the generative AI model for which the pre-training and fine-tuning procedures have been completed.
[0066] Specifically, the system can extract keywords based on pre-set keyword extraction criteria using a topic modeling model with a fine-tuning procedure completed (S530), and calculate the similarity between the user prompt and the extracted keywords using an LDA topic model or a semantic network analysis model with a fine-tuning procedure completed (S540).
[0067] And when the similarity between the user prompt and the extracted keyword is calculated through the similarity calculation unit (240), the system can extract and summarize information on the safety of the drug having a similarity greater than or equal to a preset threshold (S550).
[0068] In addition, the system can perform search and classification tasks to match information on the safety of the summarized drug based on the RAG model with the fine-tuning procedure completed to each item of the pre-set drug side effect report form (S560).
[0069] In this case, the established drug adverse event report form may be the Korea Food and Drug Administration’s Adverse Event Reporting System (KAERS) or the drug adverse event report form provided by the Council for International Organization of Medical Sciences (CIOMS).
[0070] And the system can generate (write) a drug side effect report for the drug by utilizing information classified based on a multimodal language model with a fine-tuning procedure completed (S570).
[0071] In addition, the system can utilize generative AI models to translate written drug side effect reports into other languages (e.g., English to Korean, Korean to English, etc.).
[0072] Additionally, after the search and classification tasks are performed, the system can perform self-verification for each matching item of the classified information and then perform the task of generating a drug side effect report.
[0073] By automating the task of writing drug side effect reports, time spent on report preparation can be saved, high accuracy and reliability can be ensured by minimizing human error, and reports of consistent quality can be obtained.
[0074] Meanwhile, it goes without saying that the technical concept of the present invention may also be applied to a computer-readable recording medium containing a computer program that enables the device and method according to the present embodiment to perform their functions. Furthermore, the technical concept according to various embodiments of the present invention may be implemented in the form of computer-readable code recorded on a computer-readable recording medium. A computer-readable recording medium may be any data storage device that can be read by a computer and store data. For example, a computer-readable recording medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical disk, hard disk drive, etc. Additionally, computer-readable code or a program stored on a computer-readable recording medium may be transmitted through a network connected between computers.
[0075] Furthermore, although preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above. Various modifications are possible by those skilled in the art without departing from the essence of the invention as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present invention.
Claims
1. A step in which the system receives a user prompt containing information on the safety of a specific drug and information on the approved specifications of the drug; A step in which the system extracts keywords based on pre-established keyword extraction criteria using a topic modeling model; A step in which the system calculates the similarity between a user prompt and an extracted keyword; A step in which the system extracts and summarizes information on the safety of drugs having a similarity greater than or equal to a preset threshold; A step in which the system classifies information on the safety of summarized drugs according to a pre-established drug adverse event report form; and A method for automatically generating drug safety information using generative AI, comprising the step of the system generating a drug side effect report for the drug using classified information.
2. In Claim 1, Information on the safety of medicines is, A method for automatically generating drug safety information using generative AI, characterized by including at least one of structured or unstructured data in text form, structured or unstructured data in image form, and structured or unstructured data in voice form regarding the safety of the drug.
3. In Claim 1, The pre-set keyword extraction criteria are, A method for automatically generating drug safety information using generative AI, characterized by including at least one of a component, dosage, indication (efficacy and effect), and target patients.
4. In Claim 1, The step of calculating similarity is, A method for automatically generating drug safety information using generative AI, characterized by calculating the similarity between a user prompt and extracted keywords by utilizing Latent Dirichlet Allocation (LDA) topic modeling or semantic network analysis techniques.
5. In Claim 1, The summarizing step is, A method for automatically generating drug safety information using generative AI, characterized by extracting and summarizing information on the safety of drugs having a similarity level greater than a preset threshold based on a Multimodal Large Language Model.
6. In Claim 5, The classification step is, A method for automatically generating drug safety information using generative AI, characterized by classifying summarized drug safety information based on a RAG (Retrieval-Augmented Generation) model linked to a multimodal language model to match it to each item of a pre-established drug side effect report form.
7. In Claim 6, The pre-established drug side effect report form is, A method for automatically generating drug safety information using generative AI, characterized by being a drug adverse event report form provided by the Korea Advanced Drug Response and Emerging Drugs Reporting System (KAERS) or the International Association of Organizations for Medical Sciences (CIOMS).
8. In Claim 6, A method for automatically generating drug safety information using generative AI, characterized by further including the step of, once the classification of summarized drug safety information is completed, the system performing self-verification for each matching item of the classified information.
9. In Claim 8, The step of performing self-verification is, Self-verification is performed on the classified information for each matching item, and if the verification score is below a preset threshold, the classification of the summarized drug safety information is re-executed, A method for automatically generating drug safety information using generative AI, characterized in that the pre-set threshold of the verification score is set differently for each item matching classified information, depending on whether it corresponds to essential information.
10. An input section for entering a user prompt that includes information on the safety of a specific drug and information on the approved specifications of the drug; and An automatic drug safety information generation system utilizing generative AI, comprising: a processor that utilizes a topic modeling model to extract keywords based on preset keyword extraction criteria, calculates similarity between a user prompt and the extracted keywords, extracts and summarizes information on the safety of drugs having similarity greater than a preset threshold, classifies the summarized information on the safety of drugs according to a preset drug side effect report format, and generates a drug side effect report for the corresponding drug using the classified information.