A training method of a traditional Chinese medicine prescription prediction model
Patent Information
- Application Number
- CN202511999337.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-12-29
AI Technical Summary
[0005]鉴于上述的分析,本发明实施例旨在提供一种中医处方预测模型的训练方法,用以解决现有处方预测不准确的问题
[0016] Compared with existing technologies, the training method of the TCM prescription prediction model provided in this embodiment of the invention constructs a sample set by acquiring time-series medical data of multiple patients, constructs a multi-task reasoning model with decoupled stage perception and representation, trains the multi-task reasoning model based on the sample set, and obtains a trained TCM prescription prediction model. This combines syndrome prediction, treatment prediction and prescription prediction tasks, improves the accuracy of the final prescription prediction task through multi-task learning, and provides doctors with reliable auxiliary decision support.
Smart Images

Figure CN122050719B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traditional Chinese medicine prescription prediction technology, and in particular to a training method for a traditional Chinese medicine prescription prediction model. Background Technology
[0002] The overall process of TCM diagnosis and treatment can be divided into four key stages: First, symptom collection, where doctors use the four diagnostic methods of observation, auscultation and olfaction, inquiry, and palpation to obtain information such as the patient's tongue appearance, pulse, complexion, and voice; then, the differentiation stage, where doctors infer the cause, location, and pathogenesis of the disease based on the collected symptoms, and establish the syndrome type; next, the treatment process, where doctors select targeted treatment methods based on the differentiation results, such as clearing heat or calming the mind; and finally, prescription formulation, where doctors combine the patient's specific condition, constitution, and historical health status to formulate a personalized prescription containing multiple Chinese herbs.
[0003] In early research on the intelligentization of Traditional Chinese Medicine (TCM), machine learning methods primarily utilized probabilistic models, statistical correlations, and topic modeling to explore the potential connections between symptoms and Chinese herbal medicines. These methods had limited ability to model high-order nonlinear relationships between concepts such as symptoms and medications. With the development of artificial intelligence technologies such as deep learning and reinforcement learning, more and more techniques are being used to uncover the deep connections between symptoms and prescribed medications.
[0004] However, existing methods still simplify TCM diagnosis and treatment into an end-to-end prediction task of "symptom-prescription", failing to fully consider the inherent connections and synergistic mechanisms among multiple sub-tasks such as syndrome identification, treatment selection and prescription generation, resulting in inaccurate prescription prediction. Summary of the Invention
[0005] Based on the above analysis, the present invention aims to provide a training method for a traditional Chinese medicine prescription prediction model to solve the problem of inaccurate prescription prediction in existing methods.
[0006] On one hand, embodiments of the present invention provide a training method for a traditional Chinese medicine prescription prediction model, comprising the following steps: Acquire time-series medical visit data of multiple patients, and construct a sample set based on the time-series medical visit data; A multi-task reasoning model based on stage-aware representation decoupling is constructed. The multi-task reasoning model is used to predict the results of multiple diagnosis and treatment stages using a stage-by-stage generation method. The multi-tasks include syndrome prediction task, treatment method prediction task, and prescription prediction task. The multi-task reasoning model is trained based on the sample set to obtain a trained TCM prescription prediction model.
[0007] Based on further improvements to the above method, a sample set is constructed using the time-series medical data in the following manner: Extract the time-series medical data for each patient at t time steps; the time-series medical data for each time step includes symptoms, syndromes, treatment methods, and prescription data; The symptom, syndrome, treatment, and prescription data from the previous t-1 time steps, and the symptom data from the t-th time step, are used as the input data for the sample. The syndrome, treatment, and prescription data from the t-th time step are used as the labels for the sample, resulting in a sample. The t-th time step is the target time step. The obtained samples constitute a sample set.
[0008] Based on further improvements to the above method, the multi-task reasoning model includes: The feature extraction module is used to extract the symptom feature representation at the target time step and the feature representation of historical time series medical data before the target time step; The stage-aware representation decoupling module is used to decouple the symptom feature representation of the target time step to obtain shared features and differentiated features; and to obtain an enhanced representation of the target time step based on the shared features and differentiated features. The inference module is used to perform stage-by-stage inference based on feature representations of historical time-series medical data, symptom feature representations of the target time step, and enhanced representations of the target time step, to obtain prediction results for multiple treatment stages.
[0009] Based on a further improvement of the above method, the multi-task reasoning module includes: The sequence construction unit is used to construct symptom feature sequences, syndrome feature sequences, treatment feature sequences, and prescription feature sequences based on feature representations and enhanced representations of target time steps from historical time-series medical data. The gated loop unit is used to dynamically capture the symptom feature sequence, syndrome feature sequence, treatment feature sequence, and prescription feature sequence to obtain the symptom enhancement time sequence representation, syndrome enhancement time sequence representation, treatment enhancement time sequence representation, and prescription enhancement time sequence representation, respectively. The syndrome prediction module is used to predict the syndrome at the target time step based on the symptom enhancement time series representation and the syndrome enhancement time series representation. The treatment prediction module is used to predict the treatment method at the target time step based on the temporal representation of symptom enhancement, temporal representation of syndrome enhancement, and temporal representation of treatment method enhancement. The prescription prediction module is used to predict prescriptions for the target time step based on symptom enhancement time-series representation, syndrome enhancement time-series representation, treatment enhancement time-series representation, and prescription enhancement time-series representation.
[0010] Based on further improvements to the above method, The feature representation of the historical time series medical visit data includes symptom feature representation, syndrome feature representation, treatment feature representation, and prescription feature representation for each historical time step; The enhanced representation of the target time step includes enhanced representation of syndrome, enhanced representation of treatment method, and enhanced representation of prescription; Based on the feature representation of historical time-series medical visit data and the enhanced representation of the target time step, the following methods are used to construct symptom feature sequences, syndrome feature sequences, treatment method feature sequences, and prescription feature sequences: The symptom feature representations of each historical time step and the target time step are concatenated to construct a symptom feature sequence. The syndrome feature sequence is constructed by concatenating the syndrome feature representation of each historical time step with the syndrome enhancement representation of the target time step. Concatenate the treatment feature representation of each historical time step and the treatment enhancement representation of the target time step to construct a treatment feature sequence; The prescription feature sequence is constructed by concatenating the prescription feature representation of each historical time step with the prescription augmentation representation of the target time step.
[0011] Based on further improvements to the above method, the multi-task inference model is trained using the following loss function: ; in, Indicates the predicted loss of prescriptions. This indicates the method of treatment and the prediction of losses. Indicates the predicted loss based on the symptoms. Indicates time difference alignment loss. Indicates the decoupling loss. , , and This represents the weight hyperparameter.
[0012] Based on further improvements to the above method, the differentiated features include syndrome differentiation features, treatment differentiation features, and prescription differentiation features; The decoupling loss is calculated using the following formula: ; in, This function represents the orthogonality measure between two vectors. Indicates shared features, Indicates the differentiated characteristics of syndromes. This indicates the differentiated characteristics of treatment methods. This demonstrates the differentiated characteristics of prescriptions.
[0013] Based on the above method, a further improvement is made to calculate the time difference alignment loss using the following method: ; in, This represents the feature difference between the i-th time step and the (i-1)-th time step in the symptom feature sequence. This represents the feature difference between the i-th time step and the (i-1)-th time step in the syndrome feature sequence. This represents the feature difference between the i-th time step and the (i-1)-th time step in the treatment feature sequence. Let represent the feature difference between the i-th time step and the (i-1)-th time step in the prescription feature sequence, and t represent the number of time steps in the symptom feature sequence. This represents the cosine similarity function.
[0014] Based on further improvements to the above method, the following formula is used to calculate the prescription prediction loss: ; in, Indicates the total number of symptoms. This indicates whether a label for the i-th symptom exists at the target time step. This indicates the prediction result of whether the i-th symptom exists at the target time step.
[0015] Based on a further improvement of the above method, the feature extraction module uses a graph neural network to extract feature representations of historical time-series medical data before the target time step.
[0016] Compared with existing technologies, the training method of the TCM prescription prediction model provided in this embodiment of the invention constructs a sample set by acquiring time-series medical data of multiple patients, constructs a multi-task reasoning model with decoupled stage perception and representation, trains the multi-task reasoning model based on the sample set, and obtains a trained TCM prescription prediction model. This combines syndrome prediction, treatment prediction and prescription prediction tasks, improves the accuracy of the final prescription prediction task through multi-task learning, and provides doctors with reliable auxiliary decision support.
[0017] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0018] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1 This is a flowchart of the training method for the medical prescription prediction model in an embodiment of the present invention; Figure 2 This is a performance comparison chart of different models in the embodiments of the present invention; Figure 3These are box plots showing the performance of different models in the embodiments of the present invention. Detailed Implementation
[0019] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0020] A specific embodiment of the present invention discloses a training method for a traditional Chinese medicine prescription prediction model, such as... Figure 1 As shown, it includes the following steps: S1. Obtain time-series medical visit data of multiple patients, and construct a sample set based on the time-series medical visit data; S2. Construct a multi-task reasoning model based on stage-aware representation decoupling. The multi-task reasoning model is used to predict the results of multiple diagnosis and treatment stages using a stage-by-stage generation method. The multi-tasks include syndrome prediction task, treatment method prediction task, and prescription prediction task. S3. The multi-task reasoning model is trained based on the sample set to obtain a trained TCM prescription prediction model.
[0021] In practice, TCM medical records typically contain information from multiple patient visits at different time periods, forming a multi-stage longitudinal data sequence. Let the complete medical record for each patient be represented as... ,in This represents the number of clinical visits by the patient at time step t, and Each consultation adhered to the TCM principle of "differentiation of syndromes and treatment based on syndrome differentiation," which can be formally defined as... ,in , , and These represent the set of symptoms recorded during each medical visit, the established syndrome type, the treatment methods used, and the final prescription (which typically consists of multiple Chinese herbal medicines). The patient's time-series medical visit data includes the symptoms, syndrome, treatment method, and prescription at each time step.
[0022] In practice, unique heat coding can be used to encode each TCM symptom to obtain a coding vector for each symptom; coding can be used to encode each TCM syndrome to obtain a coding vector for each syndrome; coding can be used to encode each TCM treatment method to obtain a coding vector for each treatment method; and coding can be used to encode each TCM drug to obtain a coding vector for each drug.
[0023] In practice, time-series medical records of multiple patients are acquired to construct a sample set. For example, for a given patient, the time-series medical records from time step 1 to t-1 and the syndrome at time step t are used as the sample input data. The syndrome, treatment method, and prescription at time step t are used as labels to construct a sample, thus building the sample set. Time step t is the target time step to be predicted.
[0024] This invention models the TCM diagnosis and treatment process as a multi-task learning problem. Specifically, given a patient's longitudinal historical medical records... And the set of symptoms observed at the time of the current medical visit. The goal of the model is to predict three key outcomes simultaneously, based on full utilization of temporal and stage information: (1) the syndrome type corresponding to the current medical visit. (2) Corresponding treatment methods (3) The final recommended personalized Chinese medicine prescription .
[0025] In implementation, this invention constructs a multi-task reasoning model to generate prediction results for each stage of diagnosis and treatment using a stage-by-stage generation approach. The multi-tasks include syndrome prediction, treatment method prediction, and prescription prediction.
[0026] Specifically, multi-task inference models include: The feature extraction module is used to extract the symptom feature representation at the target time step and the feature representation of historical time series medical data before the target time step; The stage-aware representation decoupling module is used to decouple the symptom feature representation of the target time step to obtain shared features and differentiated features; and to obtain an enhanced representation of the target time step based on the shared features and differentiated features. The inference module is used to perform stage-by-stage inference based on feature representations of historical time-series medical data, symptom feature representations of the target time step, and enhanced representations of the target time step, to obtain prediction results for multiple treatment stages.
[0027] In practice, the target time step is the time step that the multi-task inference model wants to predict. In this embodiment of the invention, the t-th time step is taken as the target time step.
[0028] The encoded vectors of all symptoms at the target time step constitute the symptom feature representation of the target time step.
[0029] During implementation, the feature extraction module uses a graph neural network to extract feature representations of historical time-series medical data prior to the target time step.
[0030] During implementation, a graph structure is constructed for the medical data at each time step. The nodes in the graph structure represent the symptoms, syndromes, treatments, and prescriptions in the medical data at that time step. Connection edges are established between symptom nodes and syndrome nodes; connection edges are established between syndrome nodes and treatment nodes; and connection edges are established between treatment nodes and prescription nodes, thus obtaining the graph structure for that time step.
[0031] In implementation, based on the clinical paradigm of TCM syndrome differentiation and treatment, this invention defines each type of clinical event (each symptom, each syndrome, each treatment plan, and each prescription) in each consultation as a node in the graph, and establishes embedding matrices for the four event types respectively: ; ; ; ; Each row in the embedding matrix corresponds to a single clinical event instance. Its initial representation is defined as follows: ; ; in, , , , These respectively represent the symptoms observed during a single medical visit, the diagnosed syndrome, the treatment method used, and the corresponding prescription. Indicates the total number of symptoms, Indicates the total number of symptoms. Indicates the total number of treatments. d represents the total number of prescriptions, and d represents the dimension of the initial feature. It represents the set of real numbers.
[0032] The initial features of each symptom node are obtained through the embedding matrix. Initial characteristics of each syndrome node Initial characteristics of each healing node and the initial features of each prescription node .
[0033] In constructing the graph, we establish connections between events in adjacent diagnostic stages based on the logical sequence of TCM diagnosis and treatment, forming edge sets. These correspond to the relationships of "symptom → syndrome", "syndrome → treatment method", and "treatment method → prescription", respectively. All edges are bidirectional, with an initial weight of 1, thus forming a four-part graph structure that is fully connected within each visit.
[0034] After constructing the graph structure, the graph neural network uses an improved graph attention mechanism to model the dependencies between nodes in the graph structure, thereby obtaining the feature representation of each node in the graph.
[0035] In implementation, after constructing the graph corresponding to each time step, this invention performs feature interaction and information propagation between nodes based on predefined edge relationships. It employs the GATv2 graph attention mechanism to model the dependencies between nodes and dynamically calculates edge weights through an adaptive attention mechanism, thereby improving the accuracy of feature fusion and semantic expressive power, resulting in a feature representation of each node in the graph. The input to the graph attention mechanism is the initial features of the nodes.
[0036] After obtaining the feature representation of each node in the graph structure, information is aggregated for nodes of the same type to form an overall representation of that type. GraphNorm is then used to normalize the aggregation results to eliminate scale differences between different event types. This yields the feature representation of the medical data at each time step, including symptom feature representation, syndrome feature representation, treatment method feature representation, and prescription feature representation.
[0037] For example, for the (t-1)th time step, the symptom characteristics are specifically calculated as follows: ; in, This represents the symptom characteristics at time step t-1. Let represent the set of symptom nodes in the graph structure corresponding to the (t-1)th time step. This represents the feature representation of the i-th symptom node in the graph structure corresponding to the (t-1)-th time step. This indicates GraphNorm normalization.
[0038] The symptom characteristics at time step (t-1) are obtained in the same way. Characteristics of governance methods and prescription feature representation .
[0039] To simultaneously acquire common information across all clinical stages and the unique differences between each stage, this invention proposes stage-aware representation decoupling. Stage-aware representation decoupling allows for the explicit description of how a physician's focus on different symptom characteristics changes throughout the same consultation, thereby enabling precise modeling of stages such as syndrome differentiation, treatment plan formulation, and prescription generation.
[0040] Specifically, the stage-aware representation decoupling module includes one shared encoder and three differential encoders. The shared encoder is used to extract global information common to all stages, while the differential encoders for each stage extract differential features related to a specific diagnosis and treatment stage.
[0041] Specifically, the encoder structure is defined as follows: ; Indicates a shared encoder. Parameters representing the shared encoder, Indicates the encoder for symptom differentiality. The parameters representing the syndrome difference encoder, Indicates the encoder for differences in treatment methods. The parameters represent the encoder for treatment method differences. Indicates a prescription difference encoder, Parameters representing the prescription difference encoder, The dimension of the initial feature.
[0042] The symptom feature representation at the target time step is input into the shared encoder and each differential encoder to decouple the features, thus obtaining the shared features. And three differentiating characteristics, namely, syndrome differentiation characteristics Differentiated characteristics of treatment methods and prescription differentiation characteristics Shared features reflect the overall information of a patient's medical process, while differentiated features depict the unique diagnostic and treatment information at each stage.
[0043] In implementation, the shared encoder and the differential encoder can adopt existing encoder structures, such as multilayer linear perceptrons, Transformers, recurrent neural networks, etc. This invention does not impose any restrictions on the encoder structure. The experimental results presented in this paper use the multilayer linear perceptron, which has the fewest parameters and the simplest structure, as an example.
[0044] By using a stage-aware representation decoupling module to distinguish between shared features and stage-specific features, key information corresponding to different stages can be extracted, providing a foundation for subsequent accurate predictions.
[0045] Then, based on the shared and differentiated features, an enhanced representation of the target time step is obtained. This enhanced representation includes syndrome enhancement, treatment method enhancement, and prescription enhancement. In implementation, the enhanced representation of the target time step is obtained in the following manner: ; in," "Indicates a splicing operation, , and Both represent fully connected layers. , and This represents the parameters of the corresponding fully connected layer. This indicates the enhancement of symptoms at the target time step. The treatment method at the target time step is enhanced representation. The prescription enhancement characterization represents the target time step.
[0046] After obtaining the enhanced representation of the target time step, multi-task prediction is performed through the inference module.
[0047] In this embodiment of the invention, to achieve dynamic prediction of the TCM diagnosis and treatment process, a hierarchical progressive training framework is proposed, which sequentially completes the prediction tasks of three stages: "syndrome identification—treatment method determination—prescription recommendation." This framework uses symptom characteristics as the input starting point and gradually transmits information in the time and stage dimensions, thereby achieving hierarchical accumulation and reasoning of clinical knowledge.
[0048] Specifically, the reasoning module includes: The sequence construction unit is used to construct symptom feature sequences, syndrome feature sequences, treatment feature sequences, and prescription feature sequences based on feature representations and enhanced representations of target time steps from historical time-series medical data. The gated loop unit is used to dynamically capture the symptom feature sequence, syndrome feature sequence, treatment feature sequence, and prescription feature sequence to obtain the symptom enhancement time sequence representation, syndrome enhancement time sequence representation, treatment enhancement time sequence representation, and prescription enhancement time sequence representation, respectively. The syndrome prediction module is used to predict the syndrome at the target time step based on the symptom enhancement time series representation and the syndrome enhancement time series representation. The treatment prediction module is used to predict the treatment method at the target time step based on the temporal representation of symptom enhancement, temporal representation of syndrome enhancement, and temporal representation of treatment method enhancement. The prescription prediction module is used to predict prescriptions for the target time step based on symptom enhancement time-series representation, syndrome enhancement time-series representation, treatment enhancement time-series representation, and prescription enhancement time-series representation.
[0049] Specifically, based on the feature representation of historical time-series medical visit data and the enhanced representation of the target time step, the following methods are used to construct symptom feature sequences, syndrome feature sequences, treatment method feature sequences, and prescription feature sequences: The symptom feature representations of each historical time step and the target time step are concatenated to construct a symptom feature sequence. The syndrome feature sequence is constructed by concatenating the syndrome feature representation of each historical time step with the syndrome enhancement representation of the target time step. Concatenate the treatment feature representation of each historical time step and the treatment enhancement representation of the target time step to construct a treatment feature sequence; The prescription feature sequence is constructed by concatenating the prescription feature representation of each historical time step with the prescription augmentation representation of the target time step.
[0050] During implementation, the symptom feature representations at time steps 1 to t-1 are concatenated with the symptom features at time step t to form a symptom feature sequence. The symptom feature representations at time steps 1 to t-1 are concatenated with the symptom enhancement representations at time step t to form a symptom feature sequence. The symptom feature representations at time steps 1 to t-1 are concatenated with the symptom enhancement representations at time step t to form a symptom feature sequence. The treatment feature representations of time steps 1 to t-1 are concatenated with the treatment enhancement representation of time step t to form the treatment feature sequence. The prescription feature representations from time steps 1 to t-1 are concatenated with the prescription enhancement representation from time step t to form the prescription feature sequence. .
[0051] Furthermore, by using a gated recurrent unit (GRU) to dynamically capture the features of each feature sequence, the temporal representations of symptom enhancement, syndrome enhancement, treatment enhancement, and prescription enhancement are obtained, represented as follows: ; ; ; ; in, , , and Indicates a GRU cell. , , and This represents the parameters of the corresponding GRU unit. This indicates the temporal progression of symptoms. Indicating the temporal representation of symptom enhancement, This indicates that the treatment method enhances the temporal representation. This indicates the timing of prescription enhancement.
[0052] In the syndrome identification stage, i.e. the syndrome prediction module, the symptom enhancement time sequence representation and the syndrome enhancement time sequence representation are concatenated and input into the syndrome identification decoder to obtain the multi-label syndrome prediction result.
[0053] In the treatment method determination stage, namely the treatment method prediction module, the temporal representation of symptom enhancement, the temporal representation of syndrome enhancement, and the temporal representation of treatment method enhancement are concatenated and input into the treatment method determination decoder, and the predicted result of the treatment method is output.
[0054] In the prescription recommendation stage, namely the prescription prediction module, the symptom enhancement time sequence representation, syndrome enhancement time sequence representation, treatment enhancement time sequence representation and prescription enhancement time sequence representation are concatenated and input into the prescription recommendation decoder, and the prescription prediction result is output.
[0055] In implementation, the structures of the syndrome identification decoder, treatment method determination decoder, and prescription recommendation decoder can adopt existing decoder structures, such as multilayer linear perceptrons, Transformers, recurrent neural networks, etc. This invention does not impose restrictions on the decoder structure. The experimental results presented in this paper use the multilayer linear perceptron, which has the smallest number of parameters and the simplest structure, as an example.
[0056] After constructing the multi-task reasoning model, the multi-task reasoning model is trained based on the constructed sample set to obtain the trained multi-task reasoning model.
[0057] Specifically, the multi-task inference model is trained based on the following loss function: ; in, Indicates the predicted loss of prescriptions. This indicates the method of treatment and the prediction of losses. Indicates the predicted loss based on the symptoms. Indicates time difference alignment loss. Indicates the decoupling loss. , , and This represents the weight hyperparameter.
[0058] In practice, prescription, treatment method, and syndrome prediction are all multi-label predictions. Therefore, the cross-entropy loss can be used for prescription prediction loss, treatment method prediction loss, and syndrome prediction loss.
[0059] For example, the following formula can be used to calculate prescription prediction loss: ; in, Indicates the total number of symptoms. This indicates whether a label for the i-th symptom exists at the target time step. This indicates the prediction result of whether the i-th symptom exists at the target time step.
[0060] Through this hierarchical and progressive prediction mechanism, the model can achieve dynamic reasoning and accurate prediction from symptoms to prescriptions while fully preserving the hierarchical structure of TCM diagnosis and treatment logic, thereby significantly improving the interpretability and clinical practical value of TCM intelligent diagnosis and treatment system.
[0061] To avoid information aliasing between shared features and stage-specific features, this invention further designs an orthogonal constraint mechanism by constructing a decoupling loss function. Constraining the correlation between different features to achieve independence of the feature space.
[0062] Specifically, the decoupling loss is calculated using the following formula: ; in, This represents the orthogonality measure between two vectors, which ensures the independence between different eigenvectors by minimizing the absolute value of their inner product. Indicates shared features, Indicates the differentiated characteristics of syndromes. This indicates the differentiated characteristics of treatment methods. This indicates the differentiating characteristics of prescriptions.
[0063] By introducing orthogonal constraints to decouple the feature space, the most diagnostically valuable symptom information in different stages is highlighted, thereby improving the accuracy of multi-task prediction.
[0064] To ensure consistency in the temporal trends across different stages of a patient's medical history, this invention introduces a cross-stage time difference alignment loss to uniformly model the dynamic evolution of symptoms, treatment plans, and prescriptions over time. The time difference alignment loss calculates the representational differences between two consecutive visits within each stage to characterize the temporal trends of the patient's condition.
[0065] Specifically, the time difference alignment loss is calculated using the following method: ; in, This represents the feature difference between the i-th time step and the (i-1)-th time step in the symptom feature sequence. This represents the feature difference between the i-th time step and the (i-1)-th time step in the syndrome feature sequence. This represents the feature difference between the i-th time step and the (i-1)-th time step in the treatment feature sequence. Let represent the feature difference between the i-th time step and the (i-1)-th time step in the prescription feature sequence, and t represent the number of time steps in the symptom feature sequence. This represents the cosine similarity function, used to measure the directional consistency of time difference vectors at different stages.
[0066] During implementation, the feature difference between the i-th time step and the (i-1)-th time step is calculated in the following way: ; in, This represents the symptom feature representation at the i-th time step in the symptom feature sequence. This represents the symptom feature representation at the (i-1)th time step in the symptom feature sequence; Let represent the syndrome feature representation at the i-th time step in the syndrome feature sequence. This represents the syndrome feature representation at the (i-1)th time step in the syndrome feature sequence; Let represent the treatment feature representation at the i-th time step in the treatment feature sequence. This represents the treatment feature representation at the (i-1)th time step in the treatment feature sequence; Let represent the prescription feature representation at the i-th time step in the prescription feature sequence. This represents the prescription feature representation at the (i-1)th time step in the prescription feature sequence.
[0067] , , and The mapping function, represented by a multilayer linear perceptron network, projects the time difference vector onto a shared low-dimensional representation space, thereby enabling comparability of change patterns at different stages. , , and This indicates the corresponding parameter.
[0068] By introducing a time difference alignment loss, the consistency of temporal changes during multiple medical visits is captured and maintained, thereby modeling the complex relationship between time series and multi-stage sequences. By minimizing this loss function, the consistent trend of the temporal evolution trajectory of each stage can be effectively ensured in the shared representation space, thus achieving a dynamic alignment expression of the disease condition on the longitudinal time axis.
[0069] During the training phase, by minimizing the aforementioned comprehensive loss function... This enables joint learning and optimization of model parameters; during the inference phase, the model uses the same inference process as the training phase to achieve end-to-end prediction of diagnosis and treatment results.
[0070] To illustrate the effectiveness of this invention, two real-world datasets were used for training and validation. Dataset 1 originated from inpatient medical records of the Department of Respiratory Medicine at the First Affiliated Hospital of Henan University of Traditional Chinese Medicine, recording symptoms, syndromes, treatment methods, and prescription information during multiple patient visits. Dataset 2 consisted of real clinical data provided by the Traditional Chinese Medicine Oncology Treatment Center of Chongqing University Cancer Hospital, part of the "Research, Development, and Application of Intelligent Auxiliary Diagnosis and Treatment Platform for TCM Syndrome Differentiation and Treatment" project, a national development and reform commission engineering project for the integration of biotechnology and information technology. After standardization and preprocessing, the two datasets were structured according to the four-step TCM diagnosis and treatment process, divided into training, validation, and test sets with a ratio of 0.7:0.1:0.2 to support effective model training and generalization performance validation.
[0071] To verify the effectiveness and superiority of the method of this invention, we selected six representative baseline models for comparative experiments, as follows: PTM: A topic model based on Traditional Chinese Medicine knowledge, used for prescription generation; SMGCN: A graph convolution model based on multi-heterogeneous graphs for symptom feature extraction; TCMPR: Prescription recommendation based on the mapping relationship between traditional Chinese medicine and symptoms; KDHR: A multi-graph convolution method for fusing attribute information of traditional Chinese medicine; PresRecST: A phased modeling method combining residual networks and knowledge graphs; SDPR: A generative model based on a four-part graph and employing multi-task and contrastive learning.
[0072]
[0073] Table 1: Performance Comparison of Different Methods on the Multi-Stage Dialectical Treatment Dataset Table 1 presents the comprehensive performance of various TCM prescription recommendation models across different datasets. Early TCM prescription recommendation methods primarily relied on low-order statistics (such as topic models) to mine co-occurrence patterns between symptoms and medications, making it difficult to capture high-order dependencies and exhibiting limited generalization ability. Although subsequent methods introduced graph structures and external knowledge to enhance expressive power, they still simplified the diagnosis and treatment process into a single "symptom-to-prescription" mapping, neglecting the hierarchical logic of syndrome differentiation and treatment. Recent multi-stage modeling methods have improved performance to some extent by guiding subsequent decisions based on the results of the previous stage. Meanwhile, the method of this invention effectively captures the temporal evolution of the disease and distinguishes between common and specific symptom features by introducing temporal difference alignment and stage feature decoupling mechanisms, thus more closely resembling the real clinical reasoning process and achieving higher prediction accuracy and generalization ability. It significantly outperforms existing technologies on all indicators, demonstrating excellent overall performance and stability.
[0074]
[0075] Table 2: Performance Comparison of the Invention in Syndrome and Treatment Methods To verify the predictive effectiveness of the method of this invention in various stages of diagnosis and treatment, a multi-stage performance evaluation experiment was conducted, and the results are shown in Table 2. On two datasets, compared with the multilayer perceptron model and the staged diagnosis and treatment modeling method, this invention shows significant advantages in all indicators of key aspects such as syndrome identification and treatment determination. Its superior performance mainly stems from two key designs: first, a stage-aware representation decoupling mechanism, which can distinguish common and specific symptom characteristics at different stages, accurately reflecting the doctor's focus during the diagnosis and treatment process; second, a cross-stage time difference alignment mechanism, which can capture the dynamic changes in the disease condition over time, improving the ability to characterize the patient's longitudinal diagnosis and treatment patterns. By jointly modeling stage features and time dependencies, this invention achieves higher accuracy and consistency throughout the entire process of diagnosis, treatment, and prescription generation.
[0076] This experiment further verifies the stability and adaptability of the method of the present invention in scenarios with sparse medical records, and the results are as follows: Figure 2 and Figure 3 As shown. Figure 2 This demonstrates the response of different models to the rate of decline in F1@15 (the F1 score is calculated within the range of the top 15 predictions ranked by confidence in the model's prediction results) as the number of visits is limited. Our TMRMed model significantly exhibits the least performance degradation, highlighting its ability to maintain high performance even under conditions of sparse visit information. Furthermore, Figure 3 The stability of each model was depicted, with our model showing a significantly smaller box plot, indicating minimal fluctuations and superior robustness to visit-related perturbations. While the performance of all models declined with decreasing available visits, our method exhibited the smallest decrease, maintaining high prediction accuracy and stability, demonstrating strong robustness even with limited information. Mechanistically, this advantage stems primarily from two design features: first, the temporal difference alignment mechanism models the dynamic relationships between consecutive visits, enabling the model to reasonably infer changes in the patient's condition with limited data; second, the stage-aware representation decoupling mechanism automatically focuses on key symptom features, maintaining diagnostic accuracy even in low-data environments. In summary, this invention demonstrates excellent robustness and practical value in data-scarce clinical scenarios.
[0077] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0078] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A training method for a traditional Chinese medicine prescription prediction model, characterized in that, Includes the following steps: Acquire time-series medical visit data of multiple patients, and construct a sample set based on the time-series medical visit data; A multi-task reasoning model based on stage-aware representation decoupling is constructed. The multi-task reasoning model is used to predict the results of multiple diagnosis and treatment stages using a stage-by-stage generation method. The multi-tasks include syndrome prediction task, treatment method prediction task, and prescription prediction task. The multi-task reasoning model is trained based on the sample set to obtain a trained TCM prescription prediction model. The multi-task reasoning model includes: The feature extraction module is used to extract the symptom feature representation at the target time step and the feature representation of historical time series medical data before the target time step; The stage-aware representation decoupling module is used to decouple the symptom feature representation of the target time step to obtain shared features and differentiated features; and to obtain an enhanced representation of the target time step based on the shared features and differentiated features. The inference module is used to perform stage-by-stage inference based on the feature representation of historical time series medical data, the symptom feature representation of the target time step, and the enhanced representation of the target time step, so as to obtain the prediction results of multiple diagnosis and treatment stages. The training loss for training the multi-task inference model includes temporal difference alignment loss; the temporal difference alignment loss is calculated as follows: ; in, This represents the feature difference between the i-th time step and the (i-1)-th time step in the symptom feature sequence. This represents the feature difference between the i-th time step and the (i-1)-th time step in the syndrome feature sequence. This represents the feature difference between the i-th time step and the (i-1)-th time step in the treatment feature sequence. Let represent the feature difference between the i-th time step and the (i-1)-th time step in the prescription feature sequence, and t represent the number of time steps in the symptom feature sequence. This represents the cosine similarity function.
2. The training method for the TCM prescription prediction model according to claim 1, characterized in that, Based on the aforementioned time-series medical visit data, a sample set was constructed using the following method: Extract the time-series medical data for each patient at t time steps; the time-series medical data for each time step includes symptoms, syndromes, treatment methods, and prescription data; The symptom, syndrome, treatment, and prescription data from the previous t-1 time steps, and the symptom data from the t-th time step, are used as the input data for the sample. The syndrome, treatment, and prescription data from the t-th time step are used as the labels for the sample, resulting in a sample. The t-th time step is the target time step. The obtained samples constitute a sample set.
3. The training method for the TCM prescription prediction model according to claim 1, characterized in that, The reasoning module includes: The sequence construction unit is used to construct symptom feature sequences, syndrome feature sequences, treatment feature sequences, and prescription feature sequences based on feature representations and enhanced representations of target time steps from historical time-series medical data. The gated loop unit is used to dynamically capture the symptom feature sequence, syndrome feature sequence, treatment feature sequence, and prescription feature sequence to obtain the symptom enhancement time sequence representation, syndrome enhancement time sequence representation, treatment enhancement time sequence representation, and prescription enhancement time sequence representation, respectively. The syndrome prediction module is used to predict the syndrome at the target time step based on the symptom enhancement time series representation and the syndrome enhancement time series representation. The treatment prediction module is used to predict the treatment method at the target time step based on the temporal representation of symptom enhancement, temporal representation of syndrome enhancement, and temporal representation of treatment method enhancement. The prescription prediction module is used to predict prescriptions for the target time step based on symptom enhancement time-series representation, syndrome enhancement time-series representation, treatment enhancement time-series representation, and prescription enhancement time-series representation.
4. The training method for the TCM prescription prediction model according to claim 3, characterized in that, The feature representation of the historical time series medical visit data includes symptom feature representation, syndrome feature representation, treatment feature representation, and prescription feature representation for each historical time step; The enhanced representation of the target time step includes enhanced representation of syndrome, enhanced representation of treatment method, and enhanced representation of prescription; Based on the feature representation of historical time-series medical visit data and the enhanced representation of the target time step, the following methods are used to construct symptom feature sequences, syndrome feature sequences, treatment method feature sequences, and prescription feature sequences: The symptom feature representations of each historical time step and the target time step are concatenated to construct a symptom feature sequence. The syndrome feature sequence is constructed by concatenating the syndrome feature representation of each historical time step with the syndrome enhancement representation of the target time step. Concatenate the treatment feature representation of each historical time step and the treatment enhancement representation of the target time step to construct a treatment feature sequence; The prescription feature sequence is constructed by concatenating the prescription feature representation of each historical time step with the prescription augmentation representation of the target time step.
5. The training method for the TCM prescription prediction model according to claim 3, characterized in that, The multi-task inference model is trained based on the following loss function: ; in, Indicates the predicted loss of prescriptions. This indicates the method of treatment and the prediction of losses. Indicates the predicted loss based on the symptoms. Indicates time difference alignment loss. Indicates the decoupling loss. , , and This represents the weight hyperparameter.
6. The training method for the TCM prescription prediction model according to claim 5, characterized in that, The differentiated features include syndrome differentiation features, treatment differentiation features, and prescription differentiation features; The decoupling loss is calculated using the following formula: ; in, This function represents the orthogonality measure between two vectors. Indicates shared features, Indicates the differentiated characteristics of syndromes. This indicates the differentiated characteristics of treatment methods. This demonstrates the differentiated characteristics of prescriptions.
7. The training method for the TCM prescription prediction model according to claim 5, characterized in that, The prescription prediction loss is calculated using the following formula: ; in, Indicates the total number of symptoms. This indicates whether a label for the i-th symptom exists at the target time step. This indicates the prediction result of whether the i-th symptom exists at the target time step.
8. The training method for the TCM prescription prediction model according to claim 1, characterized in that, The feature extraction module uses a graph neural network to extract feature representations of historical time-series medical data prior to the target time step.
Citation Information
Patent Citations
Diagnosis and treatment result prediction method fusing time sequence and traditional Chinese medicine multi-stage diagnosis and treatment
CN121439114A