Rareness-aware drug recommendation method, system, device, and medium
Patent Information
- Application Number
- CN202611113748.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-27
AI Technical Summary
当前主流方法均采用单一共享骨干网络,以统一变换方式处理所有就诊数据,为不同临床负担、不同罕见度的就诊分配同等建模容量,无法适配电子健康记录的长尾分布特性与药物推荐的安全敏感需求
(一)本发明通过构建罕见度感知专家混合药物推荐模型,实现诊断、操作与药物之间关联关系的协同建模,在不显著增加单位计算量的前提下扩展了模型容量,提高了药物推荐的整体准确性与处理效率;
Smart Images

Figure CN122619260B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of electronic health record data processing and artificial intelligence technology, specifically a method, system, device, and medium for drug recommendation based on rarity perception. Background Technology
[0002] Drug recommendation based on Electronic Health Records (EHRs) is an important research direction in medical artificial intelligence. Its core is to recommend appropriate drug combinations and avoid drug interactions based on the diagnosis and operation codes of the patient. It is widely used in scenarios such as clinical decision support, rational drug use, and personalized diagnosis and treatment.
[0003] Deep learning technology has significantly driven the iteration of drug recommendation methods. Existing research has optimized task performance through various methods such as temporal coding, drug representation learning, graph networks, and pre-training fine-tuning, but core architectural limitations still exist. Current mainstream methods all use a single shared backbone network to process all patient data in a uniform transformation manner, allocating equal modeling capacity to patients with different clinical burdens and rarity levels. This approach cannot adapt to the long-tail distribution characteristics of electronic health records and the safety-sensitive requirements of drug recommendation.
[0004] The long-tail distribution of clinical data amplifies the shortcomings of this architecture: First, a few common diagnostic codes dominate model training, while a large number of rare codes lack supervisory signals; second, rare visits often have more complex clinical information, yet they share modeling parameters with simple and common visits, causing the core problem of a mismatch between modeling needs and model capacity supply.
[0005] Expert hybrid models can activate some sub-networks through conditional computation, decoupling modeling capacity from the shared backbone network and effectively alleviating the above problems. However, directly applying standard expert hybrid models to drug recommendation tasks still has shortcomings. On the one hand, traditional routing allocation methods lack clinical interpretability and cannot match the rarity and complexity of visits. On the other hand, long-tail data distribution can lead to expert collapse, and rare visits still cannot obtain dedicated modeling capacity.
[0006] Therefore, in long-tail distribution scenarios, how to improve the accuracy and safety of drug recommendations is a technical problem that urgently needs to be solved. Summary of the Invention
[0007] The technical objective of this invention is to provide a drug recommendation method, system, device, and medium based on rarity perception to address the problem of improving the accuracy and safety of drug recommendations in long-tail distribution scenarios.
[0008] The technical objective of this invention is achieved as follows: a drug recommendation method based on rarity perception, the specific method of which is as follows: Construct a drug recommendation dataset: Collect an electronic health record dataset including diagnostic codes, operation codes, and prescription drug sets. Construct a consultation lexical sequence and a drug recommendation dataset based on the electronic health record dataset. Divide the drug recommendation dataset into a training set, a validation set, and a test set. Calculate the smoothed clinical evidence matrix: Statistically count the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs on the training set. Obtain a smoothed clinical evidence matrix that characterizes the strength of association based on the co-occurrence counts and marginal counts, and then obtain the significance of the evidence for the current medical visit. Construct a rareness-aware expert hybrid drug recommendation model: For each consultation word sequence, activate only the top K expert routers in the sparse expert hybrid block to obtain the consultation representation for this consultation. Map the rareness score, number of diagnostic codes, number of operation codes, and evidence significance into a routing context vector through a lightweight multilayer perceptron. Based on the rareness score of each consultation and the corresponding expert routing probability distribution, dynamically construct expert preference information related to rare samples, map the consultation representation into a drug recommendation probability vector, and obtain the recommended drug set. Training and optimizing the drug recommendation model: The drug recommendation model is trained based on the drug recommendation dataset so that it can output the drug recommendation results corresponding to the target medical visit.
[0009] As a preferred option, the drug recommendation dataset is constructed as follows: Constructing a sequence of medical visit lexical terms: For each medical visit, a [CLS] term is placed before the diagnosis code and a [SEP] term is placed before the operation code, forming a unified sequence of medical visit lexical terms, in the form of: ;in, Indicates the patient The sequence of medical terminology; Indicates patient index, =1,…,N; For patients The total number of medical visits, i.e., the patient's total number of visits. The length of the term sequence for medical consultation =1,…, ; Constructing a consultation representation: The output hidden state of the [CLS] token is used as the consultation representation, in the form of... ; This indicates the patient's communication style during each medical visit; Indicates the time sequence number of the medical visit within the corresponding patient's medical visit lexical sequence; Represents a set of diagnostic codes; Represents the set of operation codes; Indicates a collection of prescription drugs; Indicates the degree of rarity at the time of medical visit. ; Represents the minimum-maximum normalization on the training set; | | Represents the set of diagnostic codes The number of elements (i.e., the number of diagnostic codes); Represents the set of diagnostic codes The first in One diagnostic code; express Total number of visits in the training set The frequency of the training set; Constructing a drug recommendation dataset: Based on the sequence of medical visit terms and the medical visit representation, the drug recommendation dataset is represented as follows: Where N represents the total number of patients.
[0010] More preferably, the calculation of the smoothed clinical evidence matrix is as follows: On the training set, the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs are statistically analyzed. Based on the co-occurrence counts and marginal counts, the conditional log-likelihood ratio with additive smoothing is estimated to obtain a smoothed clinical evidence matrix representing the strength of association. Based on the diagnosis and procedure codes of the current visit, the corresponding non-negative association evidence is extracted from the smoothed clinical evidence matrix. The corresponding row vectors are summed and aggregated to obtain a positive clinical evidence vector. The ratio of the infinite norm to the L1 norm of the positive clinical evidence vector is calculated, and this ratio is used as the evidence significance of the current visit. The conditional log-likelihood of additive smoothing is estimated as follows: ; ; ; in, Indicating in diagnostic coding Drugs under the conditions of occurrence The smoothed conditional probability used; Indicates diagnostic code Not found; Indicating in diagnostic coding Drugs in the absence of conditions The smoothed conditional probability used; Indicates diagnostic code With drugs Co-occurrence count in the number of visits to the training set; and These represent diagnostic codes. With drugs Count the respective edges in the training set; Indicates the additive smoothing coefficient; Represents positive numbers; Indicates diagnostic code With drugs The corresponding smoothed clinical evidence matrix entries, with positive values representing drugs. In the corresponding diagnostic code When a value appears, it is more likely to be used than when it is missing; conversely, negative values are more likely to be used. The numerical stability range is truncated, and entries with co-occurrence counts below the minimum threshold are set to zero to suppress spurious associations caused by rare encodings. The non-negative part is denoted as... ; Operations—Drug Entries The calculation method is as follows: ; ; ; in, Operation code With drugs Co-occurrence count in the number of visits to the training set; and They represent operation codes respectively. With drugs Count the edges in the training set; Operation code With drugs The corresponding smoothed clinical evidence matrix entries have the same positive and negative meanings as the diagnosis-drug entries. Their absolute values are also truncated within a stable range, and entries with co-occurrence counts below the minimum threshold are set to zero.
[0011] More preferably, the sparse expert hybrid block includes a multi-head self-attention layer, residual connections and layer normalization, the top-K expert routers, and an expert feedforward network group, which includes... A parallel set of expert feedforward networks Each expert feedforward network is an independent feedforward neural network used to perform nonlinear mapping of input feature representations. Each expert feedforward network has the same network structure but uses independent learnable parameters. Each expert feedforward network is used to perform independent feature transformation on each word in the consultation word sequence, including diagnostic coding words, operation coding words, and special words [CLS] and [SEP]. The structure of the expert feedforward network includes a first linear transformation layer, a nonlinear activation layer, and a second linear transformation layer. The first linear transformation layer is used to map each word in the consultation word sequence in the original feature space to the intermediate feature space. The nonlinear activation layer is used to enhance the feature representation ability of each word in the consultation word sequence. The second linear transformation layer is used to map the intermediate features back to the original feature space.
[0012] More preferably, a lightweight multilayer perceptron is used to map the four-dimensional interpretable patient-level features—rarity score, number of diagnostic codes, number of operational codes, and evidential significance—to a routing context vector that matches the expert scoring dimension. This allows the patient-level information of rareness score and evidential significance to be injected into the expert feedforward network for selection. The lightweight multilayer perceptron is a two-layer fully connected network. The first fully connected layer linearly maps the four-dimensional interpretable patient-level features—rarity score, number of diagnostic codes, number of operational codes, and evidential significance—to the intermediate hidden dimension and applies a non-linear activation function. The second fully connected layer then linearly maps the output to a routing context vector. .
[0013] More optimally, the rareness-aware expert hybrid drug recommendation model is constructed as follows: Sparse Expert Hybrid Transform: The Top-K expert router calculates the matching probability between the input [CLS] term and each expert feedforward network based on the consultation representation and consultation-level context information, and selects the top K expert feedforward networks with the highest probabilities to participate in the calculation. That is, for each input [CLS] term, only the selected expert feedforward networks are activated, and the unselected expert feedforward networks do not participate in the calculation process of the current term. The formula is as follows: ;in, Indicates seeking medical treatment The Middle The hidden state of each word element after the multi-head self-attention layer and before the expert feedforward network; This refers to a lexical router, which generates scores for each expert feedforward network based on the hidden state of the lexical itself, thereby selecting the expert to be activated for the corresponding lexical. A learnable linear projection matrix of shape equal to the number of expert feedforward networks multiplied by the hidden dimension is used to linearly map the word hidden state to a scoring vector of length equal to the number of expert feedforward networks. Represents the routing context vector; This indicates that the patient-level routing context will be mapped to expert scoring; The learnable routing bias scale is used to control the strength of the patient-level context-induced bias. The output representations generated by each activated expert feedforward network are then weighted and fused according to their corresponding routing weights to obtain the final representation after sparse expert hybrid transformation, as shown in the formula: ;in, This represents the probabilities of the top K experts that have been renormalized on the selected expert feedforward network. This indicates the index of the selected expert feedforward network, and its value is the TopK set of indices of the first K selected expert feedforward networks; Indicates the first A first expert feedforward network, that is, the first expert feedforward network applied to the hidden state of the input lexical unit. A two-layer fully connected feedforward transform; Show the first The output of an expert feedforward network for the corresponding word element; Rarity-aware routing: This approach encodes three types of routing signals using four interpretable visit-level features: rareness score, number of diagnostic codes, number of operational codes, and significance of evidence. The formula is as follows: ;in, Indicates the rarity score; and These represent the number of diagnostic codes and the number of operation codes for this visit (i.e., the number of elements in the diagnostic code set and the operation code set), respectively, which together characterize the clinical burden. Indicates the significance of the evidence; Indicates vector transpose; This represents a column vector composed of four interpretable visit-level features: rareness score, number of diagnostic codes, number of operation codes, and significance of evidence. Positive clinical evidence is then aggregated for the current visit, using the following formula: ;in, The two summations in the equation are respectively... , For the summation index, the lower bound is always 1, and the upper bound is always the number of elements in the diagnostic code set. The number of elements in the encoding set of operations ; , These represent the first [number] of this medical visit. The diagnostic code and the first Each operation code; , This indicates that the non-negative part of the smoothed clinical evidence matrix contains... , The corresponding row vectors; and the significance is defined as: High significance indicates that the current code sharply points to a small number of drugs, while low significance indicates diffuse or weak evidence; finally, Projected as routing context vector and contribute by Scaled route bias; compared to routes that rely solely on lexical features, this context binds expert choices to clinically interpretable heterogeneity; Load balancing: through loss function Promote balanced utilization of experts among those in the expert feedforward network group; among which... Indicated by expert feedforward network The percentage of terms used for the preferred route; Indicating expert feedforward network Average routing probability on a batch; pre-factor Used to calibrate the loss to its optimal value for uniform routing; Rarity Monotonicity: By statistically analyzing the expert routing distribution of historical patient visit samples, expert preference information related to rare samples is dynamically constructed. Specifically, for historical patient visit samples, the rarity score corresponding to the historical patient visit sample and the average expert probability distribution after passing through the first K expert routers are recorded. Based on the set of historical patient visit samples, the expert selection tendency of samples with different rarity levels is weighted statistically to obtain the rare preference expert representation. ;in, The expert distribution vector representing the rare sample preference corresponds to an expert feedforward network, which is used to measure the adaptation tendency of the corresponding expert feedforward network to rare medical samples. Let represent any medical visit sample in the historical medical visit sample set, where the total number of samples in the historical medical visit sample set is denoted as . ; Indicates the first Rarity score of a historical medical visit sample; Indicates the first The average expert probability distribution obtained by calculating a routing network from a sample of historical medical visits; This represents the normalized weights calculated based on the rarity score; This represents any index variable in the historical medical visit sample set, i.e., all historical medical visit sample numbers traversed during the summation process; since samples with higher rarity have greater weight, therefore... This can more fully reflect the expert selection preferences corresponding to rare patient visits; subsequently, based on the expert distribution vector of rare sample preferences... The ranking results of the feedforward network weights of each expert are used to select the top expert with the highest weight. An expert feedforward network is used as a rare preference expert, and an expert preference mask is constructed. Among them, expert preference masking Positions with a value of 1 indicate that the corresponding expert is more inclined to handle rare samples; this is the expert preference mask. A value of 0 indicates that the corresponding expert was not classified as a rare preference expert; further, let... The average routing distribution of the current consultation layer and word units, for the current consultation Compared with historical medical records ,definition and ;in, and The current medical visit Compared with historical medical records The aggregation weights of the route distribution on rare preference experts, i.e., the inner product of the route distribution vector and the mask, characterize the degree of dependence of medical visits on rare preference experts; the monotonic alignment loss is: ;in, , The margin is a hyperparameter that controls the strength of the constraint. This is the rarity threshold, applied only to the current patient visit. Compared with historical medical records The absolute value of the difference between the rarity scores is greater than Only then are the two included in the constraint set B to participate in the loss calculation, so as to avoid introducing noise between medical visits with similar rarity; The number of patient visit pairs in the patient visit code set B; The sign function (takes values of +1, 0, or -1, used to indicate the relative magnitude and direction of two rare cases); subscript " " indicates the positive part operation This means that a penalty is only applied when the value within the parentheses is positive; if the current patient is receiving treatment... Compared to history, the sample is the key. Even rarer, monotonic alignment loss encourages current medical visits. Greater attention should be paid to specialists with rare preferences; if the current medical visit is... More common, then the opposite; rare preference expert mask is re-evaluated online, so no expert is permanently bound to a specific disease or rarity group.
[0014] More preferably, the drug recommendation model is trained as follows: Masked Encoding Reconstruction: 15% of the diagnostic and operational codes are replaced with masked [MASK] terms, and the original multi-label clinical codes are reconstructed by the prediction head using binary cross-entropy. ;in, This represents the binary cross-entropy loss in masked encoding reconstruction; The sum of the indices of the diagnostic coding vocabulary and the operational coding vocabulary is given. The index is set to a lower bound of 1 and an upper bound of the diagnostic vocabulary size. With the size of the operation vocabulary The sum (i.e. This allows for the traversal of all codes in the diagnostic coding dictionary and the operational coding dictionary; Indicates the first The real multi-label of the encoded data (whether the encoded data appeared in this visit before being masked, with a value of 0 or 1); Indicates the prediction head for the first The predicted probability of each encoded output; Clinical Evidence Reconstruction (CER): Clinical evidence reconstruction is the core pre-training objective for evidence perception. For each visit, a priority-weighted evidence score is calculated for each drug. ;in, Indicates current medical visit For the first Priority-weighted evidence score for each drug code; Indicates the rate of decay in the priority of clinical evidence reconstruction (index term) The weight of the code decreases monotonically as its position in the patient visit increases, thus assigning higher weight to earlier codes. and These represent the positional order of the diagnostic code and the operation code within this medical visit, respectively. This indicates that during this medical visit, the first One diagnostic code; Indicates the first Each operation code; This represents the non-negative part of the smoothed clinical evidence matrix (i.e., the matrix obtained after truncating negative values to 0). and These represent diagnostic codes. Operation codes Drug coding Non-negative clinical evidence items between; and The codes are arranged according to their source records within the medical visit, thus assigning higher weight to earlier codes to serve as proxies for the primary diagnosis priority; the soft objective of clinical evidence reconstruction is: When no positive evidence is available, a uniform distribution is used as a fallback; a dedicated clinical evidence reconstruction prediction head provides... ;in, This indicates that the clinical evidence reconstruction predicts the impact of the current medical visit. The output, unnormalized evidence prediction score vector; Indicates a visit to the doctor. and Let represent the learnable weight matrix and bias vector of the clinical evidence reconstruction prediction head, respectively; the loss for clinical evidence reconstruction pre-training is: ;in, This indicates the loss in pre-training for reconstructing clinical evidence; Represents the relationship between two probability distributions Divergence (relative entropy) is used to measure the difference between the predicted distribution and the soft target distribution; This indicates that clinical evidence reconstructs soft objectives; Indicates distillation temperature, used for smoothing. Hyperparameters of the distribution =1.0; Perceptual fine-tuning: The drug prediction head maps visit representations to drug probabilities. ; This represents the unnormalized prediction score vector output by the drug prediction head, with a length equal to the size of the drug vocabulary. and These represent the learnable weight matrix and bias vector of the drug prediction head, respectively. The Sigmoid activation function maps each score element-wise to the [0,1] interval; Here are the predicted probability vectors for each drug; the fine-tuning objective is: ;in, This represents the multi-label prediction loss; Indicates multiple-labeled items; Indicates punishment Common prescriptions for labeled drug pairs, This indicates that the data was extracted from the TWOSIDES database (a publicly available database of drug interactions derived from adverse drug event reports using statistical methods, recording a large number of known adverse drug interactions). Adjacency matrix metric; This indicates the first item in the drug glossary. The drug and the first There are known harmful interactions between these drugs (of which) , (For the index of drugs in the thesaurus) This indicates that there are no known harmful interactions between the two. Drug recommendation inference: Input the visit statement into the drug prediction head to generate a prediction score. Then, the recommendation probability of each candidate drug is obtained by mapping, and drugs with probabilities higher than a set threshold are included in the recommendation set. This allows us to obtain the final medication recommendation for the target patient.
[0015] A drug recommendation system based on rarity awareness, which implements the drug recommendation method based on rarity awareness as described above; the system includes: The drug recommendation dataset construction unit is used to collect electronic health record datasets including diagnostic codes, operation codes, and prescription drug sets. For each visit in the electronic health record dataset, a unified visit lexical sequence is constructed based on the corresponding diagnostic and operation codes. The visit lexical sequence is subjected to unified representation processing of lexical embedding, position embedding, and segment embedding to obtain the visit representation and calculate the visit-level rarity score. The drug recommendation dataset is constructed based on the visit lexical sequence and visit representation and is divided into training set, validation set, and test set. The smoothed clinical evidence matrix calculation unit is used to statistically count the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs on the training set. Based on the co-occurrence counts and marginal counts, it estimates the conditional log-likelihood ratio with additive smoothing to obtain a smoothed clinical evidence matrix representing the strength of association. Based on the diagnosis and procedure codes of the current visit, it extracts the corresponding non-negative association evidence from the smoothed clinical evidence matrix, sums and aggregates the corresponding row vectors to obtain a positive clinical evidence vector, and calculates the ratio of the infinite norm to the L1 norm of the positive clinical evidence vector. The ratio of the infinite norm to the L1 norm of the positive clinical evidence vector is used as the evidence significance of the current visit. The rareness-aware expert hybrid drug recommendation model construction unit is used to activate only the top K experts in the sparse expert hybrid block for each consultation term sequence to obtain the consultation representation for this consultation. Then, based on the consultation-level routing context composed of consultation-level rareness score, diagnostic code, operation code, and evidence significance, it is projected into a context vector through a lightweight multilayer perceptron, and then scaled by a learnable routing bias scale to generate a routing bias to guide expert selection. Furthermore, through load balancing and monotonicity constraints related to rareness, expert allocation can be adaptively adjusted according to changes in the rarity of the consultation. Corresponding expert preference constraints are constructed by dynamically evaluating the adaptability of different experts to rare samples based on expert usage. Finally, the consultation representation is mapped to a drug recommendation probability vector through a drug prediction head and thresholded to obtain the recommended drug set. The drug recommendation model training optimization unit is used to perform pre-training based on the drug recommendation dataset through masking, clinical evidence reconstruction, and... A multi-stage training framework for perceptual fine-tuning is used to train the drug recommendation model, and the parameters of the drug recommendation model are optimized by constructing a joint loss function, so that the drug recommendation model can output the drug recommendation result corresponding to the target medical visit.
[0016] An electronic device includes: a memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the drug recommendation method based on rarity perception as described above.
[0017] A computer-readable storage medium storing a computer program that can be executed by a processor to implement the rareness-based drug recommendation method described above.
[0018] The drug recommendation method, system, device, and medium based on rarity perception of the present invention have the following advantages: (i) This invention constructs a rareness-aware expert hybrid drug recommendation model to achieve collaborative modeling of the relationship between diagnosis, operation and drug, expands the model capacity without significantly increasing the unit computational load, and improves the overall accuracy and processing efficiency of drug recommendation; (ii) This invention decouples the modeling capacity from the single shared backbone network by replacing the dense Transformer feedforward layer with a sparsely activated expert network, so that patients with different rarity and complexity are handled by different expert subsets, thus alleviating the mismatch between unified capacity allocation and modeling requirements. (iii) This invention improves the consistency between expert allocation and clinical heterogeneity by using a rareness-aware routing module to condition expert selection based on interpretable signals such as rareness, clinical burden and evidence significance. (iv) This invention enables rare preference specialization to emerge spontaneously without the need for manual specification of expert specialization direction through a monotonic rarity alignment module, and combines load balancing to suppress expert collapse under long-tail distribution, thereby improving the balance and stability of expert utilization. (v) This invention reconstructs pre-training from clinical evidence, distills group-level diagnostic / operation-drug association evidence into the patient encoder, alleviates the problem of insufficient supervision of rare visit labels, enhances the model's ability to model complex clinical semantics, and improves the accuracy of recommendation results; (vi) This invention is achieved through The combined optimization of perceptual fine-tuning and multiple losses improves recommendation accuracy while controlling the risk of harmful drug interactions, thus balancing recommendation effectiveness and safety. (vii) This invention, through the synergistic effect of rareness-aware capacity allocation and evidence-guided expert specialization, enables the model to maintain consistent performance advantages in rare disease grouping, thereby improving the model's generalization ability and fairness in different rare disease groups. Attached Figure Description
[0019] The invention will be further described below with reference to the accompanying drawings.
[0020] Appendix Figure 1 This is a flowchart of a drug recommendation method based on rarity awareness. Appendix Figure 2 A schematic diagram of the structure of a sparse expert hybrid block; Appendix Figure 3 A flowchart illustrating the process of training a drug recommendation model; Appendix Figure 4 The structural block diagram of the building unit for the rareness perception expert hybrid drug recommendation model. Detailed Implementation
[0021] The drug recommendation method, system, device, and medium based on rarity perception of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Example 1:
[0023] As attached Figure 1 As shown in the figure, this embodiment provides a drug recommendation method based on rarity perception, which is as follows: S1. Constructing a Drug Recommendation Dataset: Collect an electronic health record dataset including diagnostic codes, operation codes, and prescription drug sets. For each visit in the electronic health record dataset, construct a unified visit term sequence based on the corresponding diagnostic and operation codes. Perform unified representation processing on the visit term sequence using term embedding, position embedding, and segment embedding to obtain the visit representation and calculate the visit-level rarity score. Construct a drug recommendation dataset based on the visit term sequence and visit representation. Divide the drug recommendation dataset into a training set, a validation set, and a test set (in a ratio of 4:1:1). The training set, validation set, and test set are used for model training, validation, and testing, respectively. S2. Calculate the smoothed clinical evidence matrix: On the training set, count the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs. Estimate the conditional log-likelihood ratio with additive smoothing based on the co-occurrence counts and marginal counts to obtain the smoothed clinical evidence matrix representing the strength of association. Based on the diagnosis and procedure codes of the current visit, extract the corresponding non-negative association evidence from the smoothed clinical evidence matrix, sum and aggregate the corresponding row vectors to obtain the positive clinical evidence vector, and calculate the ratio of the infinite norm to the L1 norm of the positive clinical evidence vector. Use this ratio as the significance of the evidence for the current visit. S3. Construct a rareness-aware expert hybrid drug recommendation model: For each consultation term sequence, only the top K experts in the sparse expert hybrid block are activated to obtain the consultation representation for this consultation; and the consultation-level routing context, composed of consultation-level rareness score, diagnosis code, operation code, and evidence significance, is projected into a context vector through a lightweight multilayer perceptron, and then the routing bias is generated after learningable routing bias scaling to guide expert selection; furthermore, through load balancing and monotonicity constraints related to rareness, the expert allocation can be adaptively adjusted according to the rarity of the consultation; by dynamically evaluating the adaptability of different experts to rare samples based on expert usage, corresponding expert preference constraints are constructed; finally, the consultation representation is mapped into a drug recommendation probability vector through a drug prediction head and thresholded to obtain the recommended drug set; S4. Training and Optimizing the Drug Recommendation Model: Based on the drug recommendation dataset, the drug recommendation model is trained through a multi-stage training framework of mask pre-training, clinical evidence reconstruction pre-training, and DDI-aware fine-tuning. The parameters of the drug recommendation model are optimized by constructing a joint loss function, so that the drug recommendation model can output the drug recommendation results corresponding to the target medical visit.
[0024] The specific steps for constructing the drug recommendation dataset in step S1 of this embodiment are as follows: S101. Constructing a Medical Visit Terminology Sequence: For each medical visit, place a [CLS] term before the diagnosis code and a [SEP] term before the operation code to form a unified medical visit terminology sequence, in the form of: ;in, Indicates the patient The sequence of medical terminology; Indicates patient index, =1,…,N; For patients The total number of medical visits, i.e., the patient's total number of visits. The length of the term sequence for medical consultation =1,…, ; S102. Constructing a Medical Visit Representation: The output hidden state of the [CLS] lexical element is used as the medical visit representation, in the form of... ; This indicates the patient's communication style during each medical visit; Indicates the time sequence number of the medical visit within the corresponding patient's medical visit lexical sequence; Represents a set of diagnostic codes; Represents the set of operation codes; Indicates a collection of prescription drugs; Indicates the degree of rarity at the time of medical visit. ; Represents the minimum-maximum normalization on the training set; | | Represents the set of diagnostic codes The number of elements (i.e., the number of diagnostic codes); Represents the set of diagnostic codes The first in One diagnostic code; express Total number of visits in the training set The training set frequency; mean-based aggregation is more stable than minimum-based aggregation when a single rare comorbidity co-occurs with several frequent diagnoses. This score is also used as a routing signal and as a basis for dividing the evaluation into five rareness quantile groups. The basis; S103. Constructing a Drug Recommendation Dataset: Based on the medical visit terminology sequence and medical visit representation, the drug recommendation dataset is represented as follows: Where N represents the total number of patients.
[0025] In this embodiment, step S2, which estimates the conditional log-likelihood ratio with additive smoothing based on the co-occurrence count and their respective edge counts, is as follows: For diagnostic-drug entries The conditional log-likelihood ratio with additive smoothing is estimated as follows: ; ; ; in, Indicating in diagnostic coding The smoothed conditional probability that drug A is used given the given conditions; Indicates diagnostic code Not found; Indicating in diagnostic coding Drugs in the absence of conditions The smoothed conditional probability used; Indicates diagnostic code With drugs Co-occurrence count in the number of visits to the training set; and These represent diagnostic codes. With drugs Each edge in the training set is counted (total number of occurrences). This represents the additive smoothing coefficient, used to avoid zero probability estimates or excessive fluctuations when the co-occurrence count is small. =0.5; It represents a very small positive constant, used to ensure that the denominator is non-zero and to ensure the numerical stability of logarithmic operations; Indicates diagnostic code With drugs The corresponding smoothed clinical evidence matrix entries, with positive values representing drugs. In the corresponding diagnostic code When a value appears, it is more likely to be used than when it is missing; conversely, negative values are more likely to be used. The numerical stability range is truncated, and entries with co-occurrence counts below the minimum threshold are set to zero to suppress spurious associations caused by rare encodings. The non-negative part is denoted as... , for appendix Figure 4 The patient encoder in the middle provides significant evidence and provides for the attached Figure 3The reconstruction of clinical evidence provides soft targets; drug safety is extracted from the TWOSIDES database (a publicly available database of drug-drug interactions mined from adverse drug event reports using statistical methods, recording a large number of known adverse drug interactions). Adjacency matrix measure, This indicates the first item in the drug glossary. The drug and the first There are known harmful interactions between these drugs (of which) , (This is the index of the drug in the dictionary), and a value of 0 indicates that there is no known harmful interaction between the two.
[0026] For Operations - Drug Entries The conditional log-likelihood ratio with additive smoothing is estimated as follows: ; ; ; in, Operation code With drugs Co-occurrence count in the number of visits to the training set; and They represent operation codes respectively. With drugs Count the edges in the training set; Operation code With drugs The corresponding smoothed clinical evidence matrix entries have the same positive and negative meanings as the diagnosis-drug entries. Their absolute values are also truncated within a stable range, and entries with co-occurrence counts below the minimum threshold are set to zero.
[0027] As attached Figure 2 As shown, the sparse expert hybrid block in step S3 of this embodiment includes a multi-head self-attention layer, residual connections and layer normalization, the top-K expert routers, and an expert feedforward network group. The expert feedforward network group includes... A parallel set of expert feedforward networks Each expert feedforward network is an independent feedforward neural network used to perform nonlinear mapping of input feature representations. Each expert feedforward network has the same network structure but uses independent learnable parameters. Each expert feedforward network is used to perform independent feature transformation on each word in the consultation word sequence, including diagnostic coding words, operation coding words, and special words [CLS] and [SEP]. The structure of the expert feedforward network includes a first linear transformation layer, a nonlinear activation layer, and a second linear transformation layer. The first linear transformation layer is used to map each word in the consultation word sequence in the original feature space to the intermediate feature space. The nonlinear activation layer is used to enhance the feature representation ability of each word in the consultation word sequence. The second linear transformation layer is used to map the intermediate features back to the original feature space.
[0028] In step S3 of this embodiment, the lightweight multilayer perceptron is used to map the four-dimensional interpretable patient-level features (rarity score, number of diagnostic codes, number of operational codes, and evidential significance) into a routing context vector that matches the expert scoring dimension. This injects the patient-level information of rareness score and evidential significance into the expert feedforward network for selection. The lightweight multilayer perceptron is a two-layer fully connected network. The first fully connected layer linearly maps the four-dimensional interpretable patient-level features (rarity score, number of diagnostic codes, number of operational codes, and evidential significance) to the intermediate hidden dimension and applies a non-linear activation function. The second fully connected layer then linearly maps the output to a routing context vector. .
[0029] The specific steps in step S3 of this embodiment for constructing the rareness-aware expert hybrid drug recommendation model are as follows: S301, Sparse Expert Hybrid Transform: A Top-K expert router calculates the matching probability between the input [CLS] term and each expert feedforward network based on the consultation representation and consultation-level context information, and selects the top K expert feedforward networks with the highest probabilities to participate in the calculation. That is, for each input [CLS] term, only the selected expert feedforward networks are activated, and the unselected expert feedforward networks do not participate in the calculation process of the current term. The formula is: ;in, Indicates seeking medical treatment The Middle The hidden state of each word element after the multi-head self-attention layer and before the expert feedforward network; This refers to a lexical router, which generates scores for each expert feedforward network based on the hidden state of the lexical itself, thereby selecting the expert to be activated for the corresponding lexical. A learnable linear projection matrix of shape equal to the number of expert feedforward networks multiplied by the hidden dimension is used to linearly map the word hidden state to a scoring vector of length equal to the number of expert feedforward networks. Represents the routing context vector; This indicates that the patient-level routing context will be mapped to expert scoring; The learnable routing bias scale is used to control the strength of the patient-level context-induced bias. The output representations generated by each activated expert feedforward network are then weighted and fused according to their corresponding routing weights to obtain the final representation after sparse expert hybrid transformation, as shown in the formula: ;in, This represents the probabilities of the top K experts that have been renormalized on the selected expert feedforward network; k represents the index of the selected expert feedforward network, which is the set of indices of the selected top K expert feedforward networks, TopK. This represents the k-th expert feedforward network, which is the k-th two-layer fully connected feedforward transformation applied to the hidden state of the input token; This represents the output of the k-th expert feedforward network to the corresponding word; S302, Rareness-Aware Routing: Three types of routing signals are encoded using four interpretable visit-level features: rareness score, number of diagnostic codes, number of operational codes, and significance of evidence. The formula is as follows: ;in, Indicates the rarity score; and These represent the number of diagnostic codes and the number of operation codes for this visit (i.e., the number of elements in the diagnostic code set and the operation code set), respectively, which together characterize the clinical burden. Indicates the significance of the evidence; Indicates vector transpose; This represents a column vector composed of four interpretable visit-level features: rareness score, number of diagnostic codes, number of operation codes, and significance of evidence. Positive clinical evidence is then aggregated for the current visit, using the following formula: ;in, The two summations in the equation are respectively... , For the summation index, the lower bound is always 1, and the upper bound is always the number of elements in the diagnostic code set. The number of elements in the encoding set of operations ; , These represent the first [number] of this medical visit. The diagnostic code and the first Each operation code; , This indicates that the non-negative part of the smoothed clinical evidence matrix contains... , The corresponding row vectors; and the significance is defined as: High significance indicates that the current code sharply points to a small number of drugs, while low significance indicates diffuse or weak evidence; finally, Projected as routing context vector and contribute by Scaled route bias; compared to routes that rely solely on lexical features, this context binds expert choices to clinically interpretable heterogeneity; S303, Load Balancing: via Encourage balanced utilization of expert feedforward networks; among which, Indicated by expert feedforward network The percentage of terms used for the preferred route; Indicating expert feedforward network Average routing probability on a batch; pre-factor Used to calibrate the loss to its optimal value for uniform routing; S304. Rarity Monotonicity: By statistically analyzing the expert routing distribution of historical medical visit samples, expert preference information related to rare samples is dynamically constructed. Specifically, for historical medical visit samples, the rarity score corresponding to the historical medical visit sample and the average expert probability distribution after passing through the first K expert routers are recorded. Based on the set of historical medical visit samples, the expert selection tendency of samples with different rarity levels is weighted statistically to obtain the rare preference expert representation. ;in, The expert distribution vector representing the rare sample preference corresponds to an expert feedforward network, which is used to measure the adaptation tendency of the corresponding expert feedforward network to rare medical samples. Let represent any medical visit sample in the historical medical visit sample set, where the total number of samples in the historical medical visit sample set is denoted as . ; Indicates the first Rarity score of a historical medical visit sample; Indicates the first The average expert probability distribution obtained by calculating a routing network from a sample of historical medical visits; This represents the normalized weights calculated based on the rarity score; This represents any index variable in the historical medical visit sample set, i.e., all historical medical visit sample numbers traversed during the summation process; since samples with higher rarity have greater weight, therefore... This can more fully reflect the expert selection preferences corresponding to rare patient visits; subsequently, based on the expert distribution vector of rare sample preferences... The ranking results of the feedforward network weights of each expert are used to select the top expert with the highest weight. An expert feedforward network is used as a rare preference expert, and an expert preference mask is constructed. Among them, expert preference masking Positions with a value of 1 indicate that the corresponding expert is more inclined to handle rare samples; this is the expert preference mask. A value of 0 indicates that the corresponding expert was not classified as a rare preference expert; further, let... The average routing distribution of the current consultation layer and word units, for the current consultation Compared with historical medical records ,definition and ;in, and The current medical visit Compared with historical medical records The aggregation weights of the route distribution on rare preference experts, i.e., the inner product of the route distribution vector and the mask, characterize the degree of dependence of medical visits on rare preference experts; the monotonic alignment loss is: ;in, , The margin is a hyperparameter that controls the strength of the constraint. This is the rarity threshold, applied only to the current patient visit. Compared with historical medical records The absolute value of the difference between the rarity scores is greater than Only then are the two included in the constraint set B to participate in the loss calculation, so as to avoid introducing noise between medical visits with similar rarity; The number of patient visit pairs in the patient visit code set B; The sign function (takes values of +1, 0, or -1, used to indicate the relative magnitude and direction of two rare cases); the subscript "+" indicates the positive part operation. That is, a penalty is applied only when the value in parentheses is positive; if the current visit t is rarer than the historical visit in sample b, the monotonic alignment loss encourages the current visit t to pay more attention to rare-preference experts; if the current visit t is more common, the opposite is true; the rare-preference expert mask is re-evaluated online, so no expert is permanently bound to a specific disease or rarity group.
[0030] As attached Figure 3 As shown, the specific training of the drug recommendation model in step S4 of this embodiment is as follows: S401, Mask Encoding Reconstruction: Replace 15% of the diagnostic and operational codes with masked [MASK] terms, and reconstruct the original multi-label clinical codes from the prediction head using binary cross-entropy. ;in, This represents the binary cross-entropy loss in masked encoding reconstruction; The sum of the indices of the diagnostic coding vocabulary and the operational coding vocabulary is given. The index is set to a lower bound of 1 and an upper bound of the diagnostic vocabulary size. With the size of the operation vocabulary The sum (i.e. This allows for the traversal of all codes in the diagnostic coding dictionary and the operational coding dictionary; Indicates the first The real multi-label of the encoded data (whether the encoded data appeared in this visit before being masked, with a value of 0 or 1); Indicates the prediction head for the first The predicted probability of each encoded output; S402, Clinical Evidence Reconstruction (CER): Clinical evidence reconstruction is the core pre-training objective for evidence perception. For each medical visit, a priority-weighted evidence score is calculated for each drug. ;in, Indicates current medical visit For the first Priority-weighted evidence score for each drug code; Indicates the rate of decay in the priority of clinical evidence reconstruction (index term) The weight of the code decreases monotonically as its position in the patient visit increases, thus assigning higher weight to earlier codes. and These represent the positional order of the diagnostic code and the operation code within this medical visit, respectively. This indicates that during this medical visit, the first One diagnostic code; Indicates the first Each operation code; This represents the non-negative part of the smoothed clinical evidence matrix (i.e., the matrix obtained after truncating negative values to 0). and These represent diagnostic codes. Operation codes Drug coding Non-negative clinical evidence items between; and The codes are arranged according to their source records within the medical visit, thus assigning higher weight to earlier codes to serve as proxies for the primary diagnosis priority; the soft objective of clinical evidence reconstruction is: When no positive evidence is available, a uniform distribution is used as a fallback; a dedicated clinical evidence reconstruction prediction head provides... ;in, This indicates that the clinical evidence reconstruction predicts the impact of the current medical visit. The output, unnormalized evidence prediction score vector; Indicates a visit to the doctor. and Let represent the learnable weight matrix and bias vector of the clinical evidence reconstruction prediction head, respectively; the loss for clinical evidence reconstruction pre-training is: ;in, This indicates the loss in pre-training for reconstructing clinical evidence; Represents the relationship between two probability distributions Divergence (relative entropy) is used to measure the difference between the predicted distribution and the soft target distribution; This indicates that clinical evidence reconstructs soft objectives; Indicates distillation temperature, used for smoothing. The hyperparameter of the distribution is τ = 1.0; S403 Perceptual fine-tuning: The drug prediction head maps visit representations to drug probabilities. ; This represents the unnormalized prediction score vector output by the drug prediction head, with a length equal to the size of the drug vocabulary. and These represent the learnable weight matrix and bias vector of the drug prediction head, respectively. The Sigmoid activation function maps each score element-wise to the [0,1] interval; Here are the predicted probability vectors for each drug; the fine-tuning objective is: ;in, This represents the multi-label prediction loss; Indicates multiple-labeled items; Indicates punishment Common prescriptions for labeled drug pairs, This indicates that the data was extracted from the TWOSIDES database (a publicly available database of drug interactions derived from adverse drug event reports using statistical methods, recording a large number of known adverse drug interactions). Adjacency matrix metric; This indicates the first item in the drug glossary. The drug and the first There are known harmful interactions between these drugs (of which) , (For the index of drugs in the thesaurus) This indicates that there are no known harmful interactions between the two. S404, Drug Recommendation Reasoning: Input the visit information into the drug prediction head to generate a prediction score. Then, the recommendation probability of each candidate drug is obtained by mapping, and drugs with probabilities higher than a set threshold are included in the recommendation set. This allows for the acquisition of the final drug recommendation result for the target medical visit; specifically, in this embodiment, a threshold δ = 0.5 is set, and drugs with a probability higher than this threshold are included in the recommendation set. This allows us to obtain the final medication recommendation for the target patient.
[0031] [Experimental Verification] This embodiment was evaluated on two publicly available intensive care electronic health record datasets, MIMIC-III and MIMIC-IV. The data were divided into training, validation, and test sets in a 4:1:1 ratio according to the patient dimension. Key statistical information of the processed datasets is shown in Table 1.
[0032]
[0033] This embodiment employs a 3-layer Transformer encoder, 4 attention heads, and a 512-dimensional hidden representation; the expert hybrid module is configured with N... E =4 expert feedforward networks, using routing with the number of experts K=2 for the first K experts, the routers are conditionalized to the rareness-aware context by the β scaling bias term in equation (5); the smoothed clinical evidence matrix is estimated only on the training set, the smoothing coefficient α=0.5, the minimum co-occurrence threshold is 1; the monotonic alignment interval η=0.1; loss weights =0.03、 =0.85、 =0.001、 =0.01; all parameters were optimized using AdamW (learning rate). Weight decay of 0.1, random deactivation ( The accuracy rate was 0.3%. Evaluation metrics included: Jaccard coefficient (primary indicator) for accuracy, PRAUC and F1 score; and safety metrics. Rate and average number of recommended drugs ( ); and the Jaccard grouping (G1 to G5) based on rarity quantiles.
[0034] The model proposed in this embodiment achieves better results than other methods on the MIMIC-III dataset. The comparison of experimental results is shown in Table 2. Higher Jaccard, PRAUC, and F1 scores are better. The lower the better.
[0035]
[0036] The drug recommendation model in this embodiment also achieved better results than other methods on the MIMIC-IV dataset. The comparison of experimental results is shown in Table 3.
[0037]
[0038] Example 2: This embodiment provides a drug recommendation system based on rarity awareness, which is used to implement the drug recommendation method based on rarity awareness in Embodiment 1; the system includes: The drug recommendation dataset construction unit is used to collect electronic health record datasets including diagnostic codes, operation codes, and prescription drug sets. For each visit in the electronic health record dataset, a unified visit lexical sequence is constructed based on the corresponding diagnostic and operation codes. The visit lexical sequence is subjected to unified representation processing of lexical embedding, position embedding, and segment embedding to obtain the visit representation and calculate the visit-level rarity score. The drug recommendation dataset is constructed based on the visit lexical sequence and visit representation and is divided into training set, validation set, and test set. The smoothed clinical evidence matrix calculation unit is used to statistically count the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs on the training set. Based on the co-occurrence counts and marginal counts, it estimates the conditional log-likelihood ratio with additive smoothing to obtain a smoothed clinical evidence matrix representing the strength of association. Based on the diagnosis and procedure codes of the current visit, it extracts the corresponding non-negative association evidence from the smoothed clinical evidence matrix, sums and aggregates the corresponding row vectors to obtain a positive clinical evidence vector, and calculates the ratio of the infinite norm to the L1 norm of the positive clinical evidence vector. The ratio of the infinite norm to the L1 norm of the positive clinical evidence vector is used as the evidence significance of the current visit. The rareness-aware expert hybrid drug recommendation model construction unit is used to activate only the top K experts in the sparse expert hybrid block for each consultation term sequence to obtain the consultation representation for this consultation. Then, based on the consultation-level routing context composed of consultation-level rareness score, diagnostic code, operation code, and evidence significance, it is projected into a context vector through a lightweight multilayer perceptron, and then scaled by a learnable routing bias scale to generate a routing bias to guide expert selection. Furthermore, through load balancing and monotonicity constraints related to rareness, expert allocation can be adaptively adjusted according to changes in the rarity of the consultation. Corresponding expert preference constraints are constructed by dynamically evaluating the adaptability of different experts to rare samples based on expert usage. Finally, the consultation representation is mapped to a drug recommendation probability vector through a drug prediction head and thresholded to obtain the recommended drug set. The drug recommendation model training and optimization unit is used to train the drug recommendation model based on the drug recommendation dataset through a multi-stage training framework of mask pre-training, clinical evidence reconstruction pre-training, and DDI-aware fine-tuning. It also optimizes the parameters of the drug recommendation model by constructing a joint loss function, so that the drug recommendation model can output the drug recommendation result corresponding to the target visit.
[0039] As attached Figure 4 As shown, the rareness-aware expert hybrid drug recommendation model construction unit in this embodiment includes a clinical evidence matrix construction module, a patient encoder module, a rareness-aware routing module, a monotonic rareness alignment module, and a drug recommendation module.
[0040] The clinical evidence matrix construction module takes the training set of the drug recommendation dataset as input, performs statistical analysis on the co-occurrence relationships between diagnosis and drug, and operation and drug, and outputs a smoothed clinical evidence matrix. and its non-negative part and will Simultaneously, the patient encoder module is fed in as a significant evidence signal, and multi-stage training is fed in as a soft target for pre-training clinical evidence reconstruction. Specifically, for each pair of diagnostic codes and drugs (or operation codes and drugs), their co-occurrence counts and their respective marginal counts are counted, and the conditional log-likelihood ratio with additive smoothing is estimated to obtain a smoothed clinical evidence matrix characterizing the strength of the association. Positive values indicate that the drug is more likely to be used when the corresponding code appears, while negative values have the opposite effect. The clinical evidence matrix construction module regards the smoothed clinical evidence matrix as an empirical structural prior, providing population-level drug use evidence for routing and pre-training.
[0041] The patient encoder module takes the sequence of medical terminology consisting of diagnostic and operational codes as input, performs embedding representation on each terminology, and processes it through multiple layers of sparse expert hybrid blocks to output the medical terminology for that visit. Specifically, for each visit, a classification [CLS] terminology is placed before the diagnostic code, and a separation [SEP] terminology is placed before the operational code to form a unified clinical sequence; the representation of each terminology is the sum of terminology embedding, segment embedding used to distinguish between diagnosis and operation, and learnable position embedding. Subsequently, the sequence passes through several layers of sparse expert hybrid blocks: each block contains multi-head self-attention, residual and layer normalization, top-K expert routers, and a routing expert group consisting of several expert feedforward networks; unlike applying the same feedforward network to all terms, this module only activates the top-K experts for each term, thereby expanding the representation capacity without activating all expert parameters; the output hidden state of the [CLS] terminology is used as the medical terminology and sent to the drug recommendation module.
[0042] The rarity-aware routing module takes the patient-level routing context as input and outputs a routing probability distribution. Specifically, the module constructs a routing context using four interpretable patient-level features—patient rarity score, number of diagnoses, number of operations, and evidence significance—projected onto a lightweight multilayer perceptron as a context vector. This vector is then scaled using a learnable routing bias and added to the lexical-level routing score to jointly determine expert selection. Evidence significance characterizes whether the current patient-level code sharply points to a small number of drugs: high significance indicates concentrated evidence, while low significance indicates diffuse or weak evidence. This module directly links expert selection to clinically interpretable patient heterogeneity.
[0043] The monotonic rarity alignment module takes the rarity score of each visit and its corresponding expert routing probability distribution as input to generate an alignment loss that constrains expert selection behavior, and incorporates it into the overall training objective. Specifically, based on the traditional MoE load balancing constraint, this module further introduces a monotonicity constraint related to rarity, enabling expert assignment to adaptively adjust as the rarity of visits changes. The module dynamically evaluates the adaptability of different experts to rare samples based on expert usage during model training and constructs corresponding expert preference constraints. When the rarity of one visit sample is higher than that of another, the model is guided to increase its preference for selecting experts suitable for rare samples; conversely, it maintains the corresponding adjustment direction. Since this expert preference relationship is continuously updated during training, experts are not fixedly assigned to a certain type of disease or a specific rarity region, thus avoiding expert degradation and load imbalance while promoting the formation of adaptive specialization capabilities among experts related to sample characteristics.
[0044] The drug recommendation module takes the visit representation as input, processes it through a drug prediction head, a sigmoid activation function, and thresholding, and outputs a set of recommended drugs for the target visit. Specifically, for a longitudinal visit record, the patient encoder first recalculates the routing context and activates the top K experts for each term to obtain the final [CLS] visit representation. Then, the drug prediction head maps it to the predicted probability of each drug, processes it through a sigmoid activation function, and binarizes it with a threshold. Drugs with probabilities higher than the threshold are included in the recommendation set, and the final output is the recommended drug combination for this visit.
[0045] This embodiment models drug recommendation as a rareness-aware conditional capacity allocation problem, replacing the dense Transformer feedforward layer with a sparsely activated expert network, and conditionalizing expert routing to three interpretable visit-level signals: diagnostic rareness, clinical burden, and evidence significance derived from the Smoothed Clinical Evidence (SCE) matrix. To combat expert collapse under long-tail training, this embodiment introduces a monotonic rareness alignment constraint, encouraging rarer visits to rely more heavily on rare-preference experts, allowing expert specialization to emerge spontaneously without manual designation. Simultaneously, it introduces Clinical Evidence Reconstruction (CER) pre-training, distilling the SCE matrix into the patient encoder before further processing. This approach allows for the transfer of population-level drug evidence to rare visits with limited supervision; compared to existing methods, this embodiment maintains competitiveness. While significantly improving the overall accuracy of drug recommendations, this approach also extends its advantages to rare visit grouping, solving the problems of unified capacity allocation and insufficient supervision of rare visits in existing technologies. Furthermore, the key innovation of this embodiment lies in constructing a sparse expert hybrid block that co-conditions clinical evidence and rarity, and achieving interpretable capacity allocation through monotonic rarity alignment and clinical evidence reconstruction pre-training.
[0046] Example 3:
[0047] This embodiment also provides an electronic device, including: a memory and a processor; The memory stores the instructions executed by the computer. The processor executes computer execution instructions stored in the memory, causing the processor to perform the rareness-based drug recommendation method in any embodiment of the present invention.
[0048] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.
[0049] Memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0050] Example 4:
[0051] This embodiment also provides a computer-readable storage medium storing a plurality of instructions, which are loaded by a processor to cause the processor to execute the drug recommendation method based on rarity perception in any embodiment of the present invention. Specifically, a system or apparatus equipped with a storage medium storing software program code that implements the functions of any of the above embodiments can be provided, and the computer (or CPU or MPU) of the system or apparatus can read and execute the program code stored in the storage medium.
[0052] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.
[0053] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.
[0054] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0055] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A drug recommendation method based on rarity perception, characterized in that, The method is as follows: Construct a drug recommendation dataset: Collect an electronic health record dataset including diagnostic codes, operation codes, and prescription drug sets. Construct a consultation lexical sequence and a drug recommendation dataset based on the electronic health record dataset. Divide the drug recommendation dataset into a training set, a validation set, and a test set. Calculate the smoothed clinical evidence matrix: Statistically count the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs on the training set. Obtain a smoothed clinical evidence matrix that characterizes the strength of association based on the co-occurrence counts and marginal counts, and then obtain the significance of the evidence for the current medical visit. Construct a rareness-aware expert hybrid drug recommendation model: For each consultation word sequence, activate only the top K expert routers in the sparse expert hybrid block to obtain the consultation representation for this consultation. Map the rareness score, number of diagnostic codes, number of operation codes, and evidence significance into a routing context vector through a lightweight multilayer perceptron. Based on the rareness score of each consultation and the corresponding expert routing probability distribution, dynamically construct expert preference information related to rare samples, map the consultation representation into a drug recommendation probability vector, and obtain the recommended drug set. Training and optimizing the drug recommendation model: The drug recommendation model is trained based on the drug recommendation dataset so that it can output the drug recommendation results corresponding to the target medical visit.
2. The drug recommendation method based on rarity perception according to claim 1, characterized in that, The drug recommendation dataset is constructed as follows: Constructing a sequence of medical visit lexical terms: For each medical visit, a [CLS] term is placed before the diagnosis code and a [SEP] term is placed before the operation code, forming a unified sequence of medical visit lexical terms, in the form of: ;in, Indicates the patient The sequence of medical terminology; Indicates patient index, =1,…,N; For patients The total number of medical visits, i.e., the patient's total number of visits. The length of the term sequence for medical consultation =1,…, ; Constructing a consultation representation: The output hidden state of the [CLS] token is used as the consultation representation, in the form of... ; This indicates the patient's communication style during each medical visit; Indicates the time sequence number of the medical visit within the corresponding patient's medical visit lexical sequence; Represents a set of diagnostic codes; Represents the set of operation codes; Indicates a collection of prescription drugs; Indicates the degree of rarity at the time of medical visit. ; Represents the minimum-maximum normalization on the training set; | | Represents the set of diagnostic codes The number of elements; Represents the set of diagnostic codes The first in One diagnostic code; express Total number of visits in the training set The frequency of the training set; Constructing a drug recommendation dataset: Based on the sequence of medical visit terms and the medical visit representation, the drug recommendation dataset is represented as follows: Where N represents the total number of patients.
3. The drug recommendation method based on rarity perception according to claim 1 or 2, characterized in that, The calculation of the smoothed clinical evidence matrix is as follows: On the training set, the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs are statistically analyzed. Based on the co-occurrence counts and marginal counts, the conditional log-likelihood ratio with additive smoothing is estimated to obtain a smoothed clinical evidence matrix representing the strength of association. Based on the diagnosis and procedure codes of the current visit, corresponding non-negative association evidence is extracted from the smoothed clinical evidence matrix. The corresponding row vectors are summed and aggregated to obtain a positive clinical evidence vector. The ratio of the infinite norm to the L1 norm of the positive clinical evidence vector is calculated, and this ratio is used as the evidence significance of the current visit. The conditional log-likelihood of additive smoothing is estimated as follows: ; ; ; in, Indicating in diagnostic coding Drugs under the conditions of occurrence The smoothed conditional probability used; Indicates diagnostic code Not found; Indicating in diagnostic coding Drugs in the absence of conditions The smoothed conditional probability used; Indicates diagnostic code With drugs Co-occurrence count in the number of visits to the training set; and These represent diagnostic codes. With drugs Count the respective edges in the training set; Indicates the additive smoothing coefficient; Represents positive numbers; Indicates diagnostic code With drugs The corresponding smoothed clinical evidence matrix entries, with positive values representing drugs. In the corresponding diagnostic code When a value appears, it is more likely to be used than when it is missing; conversely, negative values are more likely to be used. The numerical stability range is truncated, and entries with co-occurrence counts below the minimum threshold are set to zero to suppress spurious associations caused by rare encodings. The non-negative part is denoted as... ; Operations—Drug Entries The calculation method is as follows: ; ; ; in, Operation code With drugs Co-occurrence count in the number of visits to the training set; and They represent operation codes respectively. With drugs Count the edges in the training set; Operation code With drugs The corresponding smoothed clinical evidence matrix entries have the same positive and negative meanings as the diagnosis-drug entries. Their absolute values are also truncated within a stable range, and entries with co-occurrence counts below the minimum threshold are set to zero.
4. The drug recommendation method based on rarity perception according to claim 3, characterized in that, The sparse expert hybrid block includes a multi-head self-attention layer, residual connections and layer normalization, the top K expert routers, and an expert feedforward network group. The expert feedforward network group includes... A parallel set of expert feedforward networks Each expert feedforward network is an independent feedforward neural network used to perform nonlinear mapping of input feature representations. Each expert feedforward network has the same network structure but uses independent learnable parameters. Each expert feedforward network is used to perform independent feature transformation on each word in the consultation word sequence, including diagnostic coding words, operation coding words, and special words [CLS] and [SEP]. The structure of the expert feedforward network includes a first linear transformation layer, a nonlinear activation layer, and a second linear transformation layer. The first linear transformation layer is used to map each word in the consultation word sequence in the original feature space to the intermediate feature space. The nonlinear activation layer is used to enhance the feature representation ability of each word in the medical consultation word sequence; the second linear transformation layer is used to map the intermediate features back to the original feature space.
5. The drug recommendation method based on rarity perception according to claim 4, characterized in that, A lightweight multilayer perceptron is used to map the four-dimensional interpretable patient-level features (rarity score, number of diagnostic codes, number of operational codes, and significance of evidence) to a routing context vector that matches the expert scoring dimension. This allows the patient-level information of rareness score and significance of evidence to be injected into the expert feedforward network for selection. The lightweight multilayer perceptron is a two-layer fully connected network. The first fully connected layer linearly maps the four-dimensional interpretable patient-level features (rarity score, number of diagnostic codes, number of operational codes, and significance of evidence) to the intermediate hidden dimension and applies a non-linear activation function. The second fully connected layer then linearly maps the output to a routing context vector. .
6. The drug recommendation method based on rarity perception according to claim 5, characterized in that, The specific details of constructing the rareness-aware expert hybrid drug recommendation model are as follows: Sparse expert hybrid transform: The first K expert routers calculate the matching probability between the input [CLS] term and each expert feedforward network based on the consultation representation and consultation-level context information, and select the top K expert feedforward networks with the highest probabilities to participate in the calculation. That is, for each input [CLS] term, only the selected expert feedforward networks are activated, and the unselected expert feedforward networks do not participate in the calculation process of the current term. The formula is: ;in, Indicates seeking medical treatment The Middle The hidden state of each word element after the multi-head self-attention layer and before the expert feedforward network; This refers to a lexical router, which generates scores for each expert feedforward network based on the hidden state of the lexical itself, thereby selecting the expert to be activated for the corresponding lexical. A learnable linear projection matrix of shape equal to the number of expert feedforward networks multiplied by the hidden dimension is used to linearly map the word hidden state to a scoring vector of length equal to the number of expert feedforward networks. Represents the routing context vector; This indicates that the patient-level routing context will be mapped to expert scoring; The learnable routing bias scale is used to control the strength of the patient-level context-induced bias. The output representations generated by each activated expert feedforward network are then weighted and fused according to their corresponding routing weights to obtain the final representation after sparse expert hybrid transformation, as shown in the formula: ;in, This represents the probabilities of the top K experts that have been renormalized on the selected expert feedforward network. This indicates the index of the selected expert feedforward network, and its value is the TopK set of indices of the first K selected expert feedforward networks; Indicates the first A first expert feedforward network, that is, the first expert feedforward network applied to the hidden state of the input lexical unit. A two-layer fully connected feedforward transform; Show the first The output of an expert feedforward network for the corresponding word element; Rarity-aware routing: This approach encodes three types of routing signals using four interpretable visit-level features: rareness score, number of diagnostic codes, number of operational codes, and significance of evidence. The formula is as follows: ;in, Indicates the rarity score; and These represent the number of diagnostic codes and the number of operation codes for this visit, respectively, which together characterize the clinical burden; Indicates the significance of the evidence; Indicates vector transpose; This represents a column vector composed of four interpretable visit-level features: rareness score, number of diagnostic codes, number of operation codes, and significance of evidence. Positive clinical evidence is then aggregated for the current visit, using the following formula: ;in, The two summations in the equation are respectively... , For the summation index, the lower bound is always 1, and the upper bound is always the number of elements in the diagnostic code set. The number of elements in the encoding set of operations ; , These represent the first [number] of this medical visit. The diagnostic code and the first Each operation code; , This indicates that the non-negative part of the smoothed clinical evidence matrix contains... , The corresponding row vectors; and the significance is defined as: High significance indicates that the current code sharply points to a small number of drugs, while low significance indicates diffuse or weak evidence; finally, Projected as routing context vector and contribute by Scaling of route offset; Load balancing: through loss function Promote balanced utilization of experts among those in the expert feedforward network group; among which... Indicated by expert feedforward network The percentage of terms used for the preferred route; Indicating expert feedforward network Average routing probability on a batch; pre-factor Used to calibrate the loss to its optimal value for uniform routing; Rarity Monotonicity: By statistically analyzing the expert routing distribution of historical patient visit samples, expert preference information related to rare samples is dynamically constructed. Specifically, for historical patient visit samples, the rarity score corresponding to the historical patient visit sample and the average expert probability distribution after passing through the first K expert routers are recorded. Based on the set of historical patient visit samples, the expert selection tendency of samples with different rarity levels is weighted statistically to obtain the rare preference expert representation. ;in, The expert distribution vector representing the rare sample preference corresponds to an expert feedforward network, which is used to measure the adaptation tendency of the corresponding expert feedforward network to rare medical samples. Let represent any medical visit sample in the historical medical visit sample set, where the total number of samples in the historical medical visit sample set is denoted as . ; Indicates the first Rarity score of a historical medical visit sample; Indicates the first The average expert probability distribution obtained by calculating a routing network from a sample of historical medical visits; This represents the normalized weights calculated based on the rarity score; Let $\mathbf{ ... The ranking results of the feedforward network weights of each expert are used to select the top expert with the highest weight. An expert feedforward network is used as a rare preference expert, and an expert preference mask is constructed. Among them, expert preference masking Positions with a value of 1 indicate that the corresponding expert is more inclined to handle rare samples; this is the expert preference mask. A value of 0 indicates that the corresponding expert was not classified as a rare preference expert; further, let... The average routing distribution of the current consultation layer and word units, for the current consultation Compared with historical medical records ,definition and ;in, and The current medical visit Compared with historical medical records The aggregation weights of the route distribution on rare preference experts, i.e., the inner product of the route distribution vector and the mask, characterize the degree of dependence of medical visits on rare preference experts; the monotonic alignment loss is: ;in, , For intervals; This is the rarity threshold, applied only to the current patient visit. Compared with historical medical records The absolute value of the difference between the rarity scores is greater than Only then are the two included in the constraint set B to participate in the loss calculation, so as to avoid introducing noise between medical visits with similar rarity; The number of patient visit pairs in the patient visit code set B; For symbolic functions; subscript " " indicates the positive part operation This means that a penalty is only applied when the value within the parentheses is positive; if the current patient is receiving treatment... Compared to history, the sample is the key. Even rarer, monotonic alignment loss encourages current medical visits. Greater attention should be paid to specialists with rare preferences; if the current medical visit is... More commonly, the direction is the opposite.
7. The drug recommendation method based on rarity perception according to claim 6, characterized in that, The specific steps for training the drug recommendation model are as follows: Masked Encoding Reconstruction: The diagnostic and operational codes are replaced with masked [MASK] terms, and the original multi-label clinical codes are reconstructed by the prediction head using binary cross-entropy. ;in, This represents the binary cross-entropy loss in masked encoding reconstruction; The sum of the indices of the diagnostic coding vocabulary and the operational coding vocabulary is given. The index is set to a lower bound of 1 and an upper bound of the diagnostic vocabulary size. With the size of the operation vocabulary The sum of these values is used to iterate through all codes in the diagnostic coding vocabulary and the operational coding vocabulary. Indicates the first A true multi-label coded tag; Indicates the prediction head for the first The predicted probability of each encoded output; Clinical evidence reconstruction: For each visit, calculate a priority-weighted evidence score for each drug. ;in, Indicates current medical visit For the first Priority-weighted evidence score for each drug code; Indicates the rate of decay in the priority of clinical evidence reconstruction; and These represent the positional order of the diagnostic code and the operation code within this medical visit, respectively. This indicates that during this medical visit, the first One diagnostic code; Indicates the first Each operation code; This represents the non-negative part of the smoothed clinical evidence matrix. and These represent diagnostic codes. Operation codes Drug coding Non-negative clinical evidence items between; and The codes are arranged according to their source records within the medical visit, thus assigning higher weight to earlier codes to serve as proxies for the primary diagnosis priority; the soft objective of clinical evidence reconstruction is: When no positive evidence is available, a uniform distribution is used as a fallback; a dedicated clinical evidence reconstruction prediction head provides... ;in, This indicates that the clinical evidence reconstruction predicts the impact of the current medical visit. The output, unnormalized evidence prediction score vector; Indicates a visit to the doctor. and Let represent the learnable weight matrix and bias vector of the clinical evidence reconstruction prediction head, respectively; the loss for clinical evidence reconstruction pre-training is: ;in, This indicates the loss in pre-training for reconstructing clinical evidence; Represents the relationship between two probability distributions Divergence is used to measure the difference between the predicted distribution and the soft target distribution. This indicates that clinical evidence reconstructs soft objectives; Indicates distillation temperature, used for smoothing. Hyperparameters of the distribution; Perceptual fine-tuning: The drug prediction head maps visit representations to drug probabilities. ; This represents the unnormalized prediction score vector output by the drug prediction head, with a length equal to the size of the drug vocabulary. and These represent the learnable weight matrix and bias vector of the drug prediction head, respectively. The Sigmoid activation function maps each score element-wise to the [0,1] interval; Here are the predicted probability vectors for each drug; the fine-tuning objective is: ;in, This represents the multi-label prediction loss; Indicates multiple-labeled items; Indicates punishment Common prescriptions for labeled drug pairs, This indicates that it was extracted from the TWOSIDES database. Adjacency matrix metric; This indicates the first item in the drug glossary. The drug and the first There are known harmful interactions between these drugs. This indicates that there are no known harmful interactions between the two. Drug recommendation inference: Input the visit statement into the drug prediction head to generate a prediction score. Then, the recommendation probability of each candidate drug is obtained by mapping, and drugs with probabilities higher than a set threshold are included in the recommendation set. This allows us to obtain the final medication recommendation for the target patient.
8. A drug recommendation system based on rarity perception, characterized in that, This system is used to implement the drug recommendation method based on rarity perception as described in any one of claims 1 to 7; The system includes: The drug recommendation dataset construction unit is used to collect electronic health record datasets including diagnostic codes, operation codes, and prescription drug sets. For each visit in the electronic health record dataset, a unified visit lexical sequence is constructed based on the corresponding diagnostic and operation codes. The visit lexical sequence is subjected to unified representation processing of lexical embedding, position embedding, and segment embedding to obtain the visit representation and calculate the visit-level rarity score. The drug recommendation dataset is constructed based on the visit lexical sequence and visit representation and is divided into training set, validation set, and test set. The smoothed clinical evidence matrix calculation unit is used to statistically count the co-occurrence counts and marginal counts of diagnoses, procedures, and drugs on the training set. Based on the co-occurrence counts and marginal counts, it estimates the conditional log-likelihood ratio with additive smoothing to obtain a smoothed clinical evidence matrix representing the strength of association. Based on the diagnosis and procedure codes of the current visit, it extracts the corresponding non-negative association evidence from the smoothed clinical evidence matrix, sums and aggregates the corresponding row vectors to obtain a positive clinical evidence vector, and calculates the ratio of the infinite norm to the L1 norm of the positive clinical evidence vector. The ratio of the infinite norm to the L1 norm of the positive clinical evidence vector is used as the evidence significance of the current visit. The rareness-aware expert hybrid drug recommendation model construction unit is used to activate only the top K experts in the sparse expert hybrid block for each consultation term sequence to obtain the consultation representation for this consultation. Then, based on the consultation-level routing context composed of consultation-level rareness score, diagnostic code, operation code, and evidence significance, it is projected into a context vector through a lightweight multilayer perceptron, and then scaled by a learnable routing bias scale to generate a routing bias to guide expert selection. Furthermore, through load balancing and monotonicity constraints related to rareness, expert allocation can be adaptively adjusted according to changes in the rarity of the consultation. Corresponding expert preference constraints are constructed by dynamically evaluating the adaptability of different experts to rare samples based on expert usage. Finally, the consultation representation is mapped to a drug recommendation probability vector through a drug prediction head and thresholded to obtain the recommended drug set. The drug recommendation model training and optimization unit is used to train the drug recommendation model based on the drug recommendation dataset through a multi-stage training framework of mask pre-training, clinical evidence reconstruction pre-training, and DDI-aware fine-tuning. It also optimizes the parameters of the drug recommendation model by constructing a joint loss function, so that the drug recommendation model can output the drug recommendation result corresponding to the target visit.
9. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the drug recommendation method based on rarity perception as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the drug recommendation method based on rarity perception as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal fusion Bayesian medical auxiliary diagnosis method and system
CN121862365A
Federated Distributed Computational Graph Platform for Genomic Medicine and Biological System Analysis
US20250259695A1