Adverse drug reaction prediction system based on large language model embedded layer
By using a large language model embedded layer in the drug adverse reaction prediction system and combining it with the molecular characteristics of the drug, the limitations of drug adverse reaction detection in the prior art are solved, and more efficient and accurate prediction of adverse reactions is achieved, especially in the prediction of unknown adverse reactions.
Patent Information
- Application Number
- CN202510072816.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has limitations in the detection of adverse drug reactions, making it difficult to capture rare or long-term side effects, and traditional methods are less efficient when processing multimodal data, and lack the ability to predict potential unknown adverse reactions of drugs.
A drug adverse reaction prediction system based on the large language model embedding layer is adopted. By replacing the embedded layer of the T-TA multi-head attention mechanism unit with the first embedded layer trained by the drug molecular data set, the embedded layer of the large language model is combined with the characteristics of the drug molecular to achieve real-time and accurate prediction of adverse reactions.
It improves the accuracy and efficiency of adverse reaction prediction, enhances the prediction ability of unknown adverse reactions, optimizes performance on small-scale data sets, provides more comprehensive drug safety prediction, and helps R&D personnel to discover potential risks in a timely manner.
Smart Images

Figure CN120015362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of drug safety prediction, and in particular to a drug adverse reaction prediction system based on a large language model embedding layer. Background Art
[0002] Adverse drug reactions are an important issue in drug development and medical applications. Adverse drug reactions can cause serious harm to patients' health and even threaten their lives. In existing drug adverse reaction detection, they are mainly obtained through clinical trials and voluntary reporting systems, but these methods have obvious limitations in practical applications. Due to scale and time constraints, clinical trials are difficult to capture rare or long-term side effects. In addition, the sample population in clinical trials is usually relatively single, and it is difficult to fully cover the reactions of people of different ages, genders and genetic backgrounds. Although the voluntary reporting system provides a more diverse source of data, the reported data is usually incomplete, and it is especially easy to overlook mild or common adverse reactions. In addition, there is a time delay in reporting, which cannot provide timely risk predictions for drug developers and doctors.
[0003] In the field of computational methods, traditional prediction methods based on drug molecular structure rely on detailed chemical molecular information, and have limited prediction effects on new drugs that lack molecular data. At the same time, traditional machine learning methods, such as random forests and support vector machines, are inefficient when processing multimodal data, such as drug text information and molecular structure data, and the model performance is difficult to meet actual needs. In recent years, graph neural networks have attracted attention for their good performance in drug interaction modeling, but due to their high computational complexity, they are difficult to apply in actual large-scale data scenarios. In addition, most existing algorithms focus on querying and analyzing known adverse reactions, but lack the ability to predict potential unknown adverse reactions of drugs, which makes it difficult for R&D personnel and doctors to discover the potential risks of drugs in a timely manner, increasing the safety risks of drug use. Therefore, how to detect and predict adverse drug reactions more accurately and in real time is a technical problem that needs to be solved. Summary of the invention
[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a drug adverse reaction prediction system based on a large language model embedding layer. By replacing the embedding layer of the T-TA multi-head attention mechanism unit with the first embedding layer, the large language model embedding layer is fully combined with the molecular characteristics of the drug, so as to accurately predict the potential adverse reactions of the drug in real time.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] According to one aspect of the present invention, a drug adverse reaction prediction system based on a large language model embedding layer is provided, the system comprising: a training sample construction module and a prediction module;
[0007] Among them, the training sample module includes a first data input unit and a first data processing unit; the prediction module includes a first data reading unit, a T-TA multi-head attention mechanism unit, a decoder and an output unit; the embedding layer in the T-TA multi-head attention mechanism unit is the first embedding layer obtained after the medical Llama model is trained using a drug molecule dataset.
[0008] Furthermore, the first embedding layer is obtained by training and extracting an embedding layer replacement module.
[0009] Furthermore, the embedding layer replacement module includes a drug molecule data set construction unit, a medical Llama model training unit and an extraction unit; wherein the medical Llama model training unit includes a second data reading unit and a medical Llama model unit.
[0010] Furthermore, the drug molecule data set construction unit includes a second data input unit and a second data processing unit; the second data input unit collects the drug name and the corresponding molecular attribute data, and the molecular attributes include molecular structure, molecular size, chirality and toxicity data; the second data processing unit performs cleaning, deduplication and structuring operations to obtain the drug molecule data set.
[0011] Furthermore, the medical Llama model unit is trained using the drug molecule dataset, so that the embedding layer of the medical Llama model unit learns the molecular attribute data of the drug molecules, combines it with the original medical text data information in the embedding layer, and obtains the first embedding layer through the extraction unit.
[0012] Furthermore, in the training sample module, the first data input unit collects drug names and corresponding adverse reactions, and inputs them into the first data processing unit for preprocessing of cleaning and deduplication operations to obtain a training data set.
[0013] Furthermore, when calculating the attention score matrix, the T-TA multi-head attention mechanism unit uses a diagonal masking operation to avoid the attention interaction between the same token and itself. The expression of the attention score is:
[0014]
[0015] Among them, Q i is the query vector produced by the first embedding layer; K is the key vector that provides matching references to capture the deep relationship between drugs and adverse reactions; V is the vector that retains the feature information of the embedding layer and generates the final prediction representation value; d is the key vector dimension.
[0016] Furthermore, an input isolation mechanism is used in the T-TA multi-head attention mechanism unit to isolate the key-value input from the network flow and fix it to the sum of word embedding and position embedding. The query input is updated layer by layer only during reasoning by referring to the fixed output of the first embedding layer, and the position embedding is input into the query Q of the first encoding layer.
[0017] Furthermore, the information of the first embedding layer is encoded and converted into an embedding representation, and the embedding representation is input into the T-TA multi-head attention mechanism unit.
[0018] Furthermore, after embedding the first embedding layer, the decoder and the T-TA multi-head attention mechanism unit are jointly trained using the training dataset.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] (1) Improved the accuracy and efficiency of adverse reaction prediction: By fine-tuning the training of a large language model and extracting its embedding layer, and fusing the multimodal data of drug text information and molecular structure data, the limitations of traditional methods in dealing with complex drug-adverse reaction relationships have been overcome, and the prediction accuracy of potential adverse drug reactions has been improved. The T-TA multi-head attention mechanism that replaces the first embedding layer and the joint training of the decoder are used to optimize the inference speed, making the prediction process more efficient.
[0021] (2) Enhanced prediction capability for unknown adverse reactions: The system not only relies on known adverse drug reaction datasets for training, but also constructs a drug molecule dataset containing drug molecular structure, molecular size, and chirality information through drug molecule datasets, so that the medical Llama model can learn deeper semantic relationships, including the prediction of unknown or rare adverse reactions. Therefore, the system can provide more comprehensive safety predictions in the early stages of drug development or when evaluating new drugs, helping R&D personnel and doctors to promptly identify potential risks.
[0022] (3) Optimizing performance on small-scale datasets: By fine-tuning the large language model and embedding specific drug molecule information, the data scarcity problem is effectively addressed, and good prediction performance can still be achieved on small-scale datasets. In addition, this system does not need to fix the length of the adverse reaction sequence of the drug to be tested, and can more flexibly and comprehensively predict the potential adverse reactions of the drug, further enhancing its applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a structural block diagram of the drug adverse reaction prediction system based on the large language model embedding layer. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0025] like Figure 1 As shown, a drug adverse reaction prediction system based on a large language model embedding layer includes: a training sample construction module and a prediction module, wherein the training sample module includes a first data input unit and a first data processing unit; the prediction module includes a first data reading unit, a T-TA multi-head attention mechanism unit, a decoder and an output unit; the embedding layer in the T-TA multi-head attention mechanism unit is the first embedding layer obtained after the medical Llama model is trained using a drug molecule data set. The first embedding layer is obtained by training and extracting the embedding layer replacement module. The embedding layer replacement module includes a drug molecule data set construction unit, a medical Llama model training unit and an extraction unit; wherein the medical Llama model training unit includes a second data reading unit and a medical Llama model unit. The drug molecule data set construction unit includes a second data input unit and a second data processing unit; the second data input unit collects the drug name and the corresponding molecular attribute data, and the molecular attributes include molecular structure, molecular size, chirality and toxicity data; the second data processing unit performs cleaning, deduplication and structuring operations to obtain a drug molecule data set.
[0026] The medical Llama model unit is trained using the drug molecule data set, so that the embedding layer of the medical Llama model unit learns the drug molecule attribute data, combines it with the original medical text data information in the embedding layer, and obtains the first embedding layer through the extraction unit. In the training sample module, the first data input unit collects the drug name and the corresponding adverse reaction, and inputs it into the first data processing unit for pre-processing of cleaning and deduplication operations to obtain the training data set.
[0027] The information of the first embedding layer is encoded and converted into an embedded representation, and the embedded representation is input into the T-TA multi-head attention mechanism unit. After embedding the first embedding layer, the decoder and the T-TA multi-head attention mechanism unit are jointly trained using the training dataset.
[0028] When calculating the attention score matrix, the T-TA multi-head attention mechanism unit uses diagonal masking to avoid the attention interaction between the same token and itself. The expression of the attention score is:
[0029]
[0030] Among them, Q iis the query vector produced by the first embedding layer; K is the key vector that provides matching references to capture the deep relationship between drugs and adverse reactions; V is the vector that retains the feature information of the embedding layer and generates the final prediction representation value; d is the key vector dimension.
[0031] The T-TA multi-head attention mechanism unit uses an input isolation mechanism to isolate the key-value input from the network flow and fix it to the sum of word embedding and position embedding. The query input is only updated layer by layer during reasoning by referring to the fixed output of the first embedding layer, and the position embedding is input into the query Q of the first encoding layer.
[0032] The specific steps of the method for building a drug adverse reaction prediction system based on a large language model embedding layer include:
[0033] S1, the first data input unit in the training sample module collects the drug name and the corresponding adverse reaction, and inputs it into the first data processing unit for preprocessing to obtain a training data set; the second data input unit of the drug molecule data set construction unit collects the molecular structure data, molecular size data and chirality of the drug, and the second data processing unit performs preprocessing to obtain a drug molecule data set;
[0034] S2. Use the drug molecule dataset to train the medical Llama model unit;
[0035] S3. Use the extraction unit to extract the embedding layer of the trained medical Llama model unit as the first embedding layer, and replace the embedding layer of the T-TA multi-head attention mechanism unit with the first embedding layer;
[0036] S4, jointly train the decoder and T-TA multi-head attention mechanism unit using the training dataset.
[0037] This embodiment is widely applicable to scenarios such as drug development and medical diagnosis. By replacing the embedding layer of the traditional model with the embedding layer of the large language model trained with the drug molecule data set, the computational efficiency and model performance of the prediction are improved. At the same time, this embodiment can accurately predict the potential undetected adverse reactions of drugs, greatly improving the timeliness and accuracy of adverse reaction prediction, providing forward-looking warnings for users' drug safety, providing scientific decision-making support for drug developers and medical practitioners, and also providing guarantees for patients' drug safety.
[0038] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A drug adverse reaction prediction system based on a large language model embedding layer, characterized in that: The system comprises: a training sample construction module and a prediction module; Among them, the training sample module includes a first data input unit and a first data processing unit; the prediction module includes a first data reading unit, a T-TA multi-head attention mechanism unit, a decoder and an output unit; the embedding layer in the T-TA multi-head attention mechanism unit is the first embedding layer obtained after the medical Llama model is trained using a drug molecule dataset.
2. A drug adverse reaction prediction system based on a large language model embedding layer according to claim 1, characterized in that: The first embedding layer is obtained by training and extracting an embedding layer replacement module.
3. A drug adverse reaction prediction system based on a large language model embedding layer according to claim 2, characterized in that: The embedding layer replacement module includes a drug molecule data set construction unit, a medical Llama model training unit and an extraction unit; wherein the medical Llama model training unit includes a second data reading unit and a medical Llama model unit.
4. A drug adverse reaction prediction system based on a large language model embedding layer according to claim 3, characterized in that: The drug molecule data set construction unit includes a second data input unit and a second data processing unit; the second data input unit collects drug names and corresponding molecular attribute data, and the molecular attributes include molecular structure, molecular size, chirality and toxicity data; the second data processing unit performs cleaning, deduplication and structuring operations to obtain a drug molecule data set.
5. A drug adverse reaction prediction system based on a large language model embedding layer according to claim 4, characterized in that: The medical Llama model unit is trained using the drug molecule data set, so that the embedding layer of the medical Llama model unit learns the molecular attribute data of the drug molecules, combines it with the original medical text data information in the embedding layer, and obtains the first embedding layer through the extraction unit.
6. A drug adverse reaction prediction system based on a large language model embedding layer according to claim 1, characterized in that: In the training sample module, the first data input unit collects drug names and corresponding adverse reactions, and inputs them into the first data processing unit for preprocessing of cleaning and deduplication operations to obtain a training data set.
7. The drug adverse reaction prediction system based on a large language model embedding layer according to claim 1, characterized in that: The T-TA multi-head attention mechanism unit uses diagonal masking to avoid the attention interaction between the same token and itself when calculating the attention score matrix. The expression of the attention score is: Among them, Q i is the query vector produced by the first embedding layer; K is the key vector that provides matching references to capture the deep relationship between drugs and adverse reactions; V is the vector that retains the feature information of the embedding layer and generates the final prediction representation value; d is the key vector dimension.
8. The drug adverse reaction prediction system based on a large language model embedding layer according to claim 1, characterized in that: The T-TA multi-head attention mechanism unit uses an input isolation mechanism to isolate the key-value input from the network flow and fix it to the sum of word embedding and position embedding. The query input is updated layer by layer only during reasoning by referring to the fixed output of the first embedding layer, and the position embedding is input into the query Q of the first encoding layer.
9. The drug adverse reaction prediction system based on a large language model embedding layer according to claim 1, characterized in that: The information of the first embedding layer is encoded and converted into an embedding representation, and the embedding representation is input into the T-TA multi-head attention mechanism unit.
10. The drug adverse reaction prediction system based on a large language model embedding layer according to claim 1, characterized in that: After the first embedding layer, the decoder and the T-TA multi-head attention mechanism unit are jointly trained using the training dataset.