Dynamic monitoring method and system of TCM disease based on big data time series analysis
By extracting disease characteristics and symptoms from TCM disease records and conducting in-depth analysis using a disease dynamic development prediction network, the problem of reliance on experience in traditional TCM diagnosis has been solved. This enables precise dynamic monitoring and prediction of TCM diseases, improving the accuracy and reliability of disease prediction.
Patent Information
- Application Number
- CN202411646991.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Traditional Chinese medicine diagnostic methods rely on doctors' experience and patients' self-reports, making it difficult to achieve objective and dynamic monitoring of the disease. Existing time series analysis methods ignore the complex correlations and temporal dependencies between disease characteristics, resulting in limited accuracy and timeliness of disease monitoring.
By extracting past case disease characteristic data, current node and predicted case patient symptoms from the patient's TCM disease record data set, and using a disease dynamic development prediction network for in-depth analysis, the network is optimized to improve the accuracy of disease prediction. This includes constructing a semantic classification framework, privacy semantic encoding and nested structure to process data, and optimizing the disease dynamic development prediction network.
It enables precise and dynamic monitoring of the progression of diseases in TCM, improves the accuracy and foresight of disease prediction, reduces uncertainty, and provides reliable data support for TCM clinical decision-making.
Smart Images

Figure CN119495390B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a method and system for dynamic monitoring of TCM disease conditions based on big data time series analysis. Background Technology
[0002] In traditional Chinese medicine (TCM) clinical practice, accurate monitoring and prediction of disease progression are crucial for developing effective treatment plans and improving patient prognosis. However, traditional TCM diagnostic methods rely primarily on the doctor's experience and the patient's self-report. While this method can reflect the patient's condition to some extent, it is often influenced by subjective factors and makes it difficult to achieve objective and dynamic monitoring of the disease.
[0003] With the rapid development of big data technology, it has become possible to dynamically monitor disease conditions using time series analysis. Time series analysis is a statistical technique that can reveal the patterns and trends of data changes over time through the mining and analysis of historical data. However, most time series analysis methods currently applied to disease monitoring in Traditional Chinese Medicine (TCM) focus only on data from a single point in time, neglecting the continuity and dynamism of disease development, resulting in limited accuracy and timeliness in disease monitoring.
[0004] Furthermore, existing disease prediction models often overlook the complex relationships and temporal dependencies between disease features when processing TCM disease data, which affects the reliability of the prediction results. Summary of the Invention
[0005] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for dynamic monitoring of TCM disease conditions based on big data time series analysis, the method comprising:
[0006] The TCM patient condition record data set is used to obtain past sample condition characteristic data set, current sample patient symptoms, and predicted sample patient symptoms. The TCM patient condition record data set includes multiple sample patient symptoms organized based on time series. The past sample condition characteristic data set includes multiple sample patient symptoms located before the target time node. The current sample patient symptoms are the sample patient symptoms corresponding to the target time node. The predicted sample patient symptoms are the sample patient symptoms after the target time node.
[0007] The set of past case disease characteristics data and the symptoms of the current node patient are loaded into the disease dynamic development prediction network to determine the first confidence level of the occurrence of the symptoms of the current node patient at the target time node;
[0008] The set of past case disease characteristics data and the predicted case patient symptoms are loaded into the disease dynamic development prediction network to determine the second confidence level of the predicted case patient symptoms appearing after the target time node;
[0009] Based on the loss between the labeled confidence level of the symptoms exhibited by the patient in this node and the first confidence level, the disease dynamic development prediction network is optimized. Furthermore, based on the loss between the labeled confidence level of the predicted symptoms exhibited by the patient in this node and the second confidence level, the disease dynamic development prediction network is optimized.
[0010] In one possible implementation of the first aspect, optimizing the disease progression prediction network based on the loss between the labeled confidence level of the predicted sample patient's symptoms and the second confidence level includes:
[0011] Obtain first training supervision data and second training supervision data of the predicted sample patient's symptoms. The proportion of positive sample labels used in the first training supervision data to characterize the predicted sample patient's symptoms is 1. The proportion of positive sample labels used in the second training supervision data to characterize the predicted sample patient's symptoms is x, where x is not less than 0 and x is not greater than 1.
[0012] Based on the first training supervision data and the second training supervision data, the labeled confidence scores of the predicted sample patient's symptoms are generated;
[0013] The disease progression prediction network is optimized based on the loss between the labeled confidence level of the predicted patient's symptoms and the second confidence level.
[0014] In one possible implementation of the first aspect, the disease progression prediction network is a taught disease progression prediction network; the acquisition of the second training supervision data of the predicted sample patient's symptom presentation includes:
[0015] The set of past case disease characteristics data and the predicted case patient symptoms are loaded into the teaching disease dynamic development prediction network to determine the third confidence level of the predicted case patient symptoms appearing after the target time node; the number of network parameters of the teaching disease dynamic development prediction network is not less than the number of network parameters of the teaching disease dynamic development prediction network.
[0016] The third confidence level is used as the second training supervision data for the predicted symptoms of the patient.
[0017] In one possible implementation of the first aspect, the patient's condition dynamic development prediction network is the patient's condition dynamic development prediction network in the current round of network parameter optimization phase, and the teaching condition dynamic development prediction network is the teaching condition dynamic development prediction network in the current round of network parameter optimization phase; the method further includes:
[0018] Based on the teaching condition dynamic development prediction network of the previous network parameter optimization stage and the teaching condition dynamic development prediction network of the current network parameter optimization stage, the teaching condition dynamic development prediction network of the current network parameter optimization stage is generated.
[0019] In one possible implementation of the first aspect, generating the teaching condition dynamic development prediction network for the current round of network parameter optimization based on the teaching condition dynamic development prediction network of the previous network parameter optimization stage and the teaching condition dynamic development prediction network of the current round of network parameter optimization stage includes:
[0020] The neuron weight information of the teaching disease dynamic development prediction network in the previous network parameter optimization stage is fused with the first training degradation factor to generate the first parameter information;
[0021] The neuron weight information of the dynamic development prediction network of the patient's condition in the current network parameter optimization stage is fused with the second training degradation factor to generate the second parameter information; the sum of the first training degradation factor and the second training degradation factor is 1.
[0022] The first parameter information and the second parameter information are fused to generate the neuron weight information of the teaching disease dynamic development prediction network in the current round of network parameter optimization stage.
[0023] In one possible implementation of the first aspect, generating the labeled confidence score of the predicted sample patient's symptom presentation based on the first training supervision data and the second training supervision data includes:
[0024] The second training supervision data is fused with the first influence factor to generate the third training supervision data;
[0025] The first training supervision data is fused with the second influence factor to generate the fourth training supervision data; the sum of the first influence factor and the second influence factor is 1.
[0026] The third and fourth training supervision data are fused together to generate the labeled confidence scores of the predicted patient symptoms.
[0027] In one possible implementation of the first aspect, the method further includes:
[0028] From multiple candidate first impact factors, a reference impact factor that matches the second training supervision data is determined; the reference impact factor is used as the first impact factor of the second training supervision data.
[0029] In one possible implementation of the first aspect, obtaining the predicted sample patient's symptoms includes:
[0030] The predicted symptoms of the patient are randomly extracted from the patient's TCM medical record data set, which includes the symptoms of patients following the symptoms of the patient in this node.
[0031] In one possible implementation of the first aspect, the method further includes:
[0032] Based on the optimized disease dynamic development prediction network, the target disease feature data set and the target patient's symptoms are analyzed to predict the target confidence level of the time node corresponding to the target patient's symptoms.
[0033] For example, in one possible implementation of the first aspect, before the steps of obtaining the set of past sample disease characteristics data, the symptoms of the current sample patient, and the predicted symptoms of the sample patient from the patient's TCM disease record data set, the method further includes:
[0034] A structural analysis was performed on the patient TCM condition record data in the aforementioned patient TCM condition record dataset to determine the different types of data elements included and the relationship structure between each data element.
[0035] Based on prior knowledge, data, and privacy protection requirements in the field of traditional Chinese medicine health management, a semantic classification framework is constructed. This framework is used to classify each data element according to its semantic category in traditional Chinese medicine theory and to define corresponding semantic tags and related attributes. The semantic tags will be used to perform privacy semantic encoding on the data elements.
[0036] According to the semantic classification framework, each data element in the patient's TCM condition record data is labeled. For each data element, the category to which the data element belongs in the semantic classification framework is found, and a corresponding semantic label is assigned to the data element. Based on the labeled semantic label, privacy weight information is added to each data element. The privacy weight information is determined based on the importance of the data element in privacy protection. In this way, the data elements labeled with semantic labels and privacy weights are recombined to form a preliminary data structure after privacy semantic encoding.
[0037] Various health data patterns are extracted from the initial data structure after privacy semantic encoding. Each health data pattern is characterized and a health data pattern library is established. The characteristic description includes the combination of data elements that constitute the health data pattern, the value range or characteristic value of each data element, and the logical relationship between each data element.
[0038] For each data subset in the initial data structure after privacy semantic encoding, it is matched with various health data patterns in the health data pattern library. The data subset is a data part divided according to different diagnostic stages or body systems.
[0039] For a successfully matched subset of data, the name or number of the health data pattern matched by the successfully matched subset is marked, and the similarity degree when the match is successful is recorded to generate a health data pattern matching result. The similarity degree is quantitatively evaluated based on the proportion of matching data elements and the matching status of key data elements. In addition, a special mark is made for a subset of data that does not match any health data pattern, indicating that the subset of data is an abnormal subset of data or a new unidentified health data pattern.
[0040] The preliminary data structure of privacy semantic encoding is integrated with the health data pattern matching results. Specifically, for each data element, the semantic label, privacy weight information, matched health data pattern and similarity information of the data element are associated. Based on the association results, a corresponding nested structure is constructed. The outer structure of the nested structure contains the semantic category information of the data element, and the inner structure contains the data element itself, privacy weight, and pattern matching result information.
[0041] Based on the nested structure, the patient's TCM medical record data is preprocessed for privacy protection.
[0042] In another aspect, embodiments of the present invention also provide a dynamic monitoring system for TCM disease conditions based on big data time series analysis, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or code. The processor is used to execute the programs, instructions, or code in the machine-readable storage medium to implement the above-mentioned method.
[0043] Based on the above, this application embodiment extracts the characteristics of past cases, the current node, and the symptoms of predicted cases from the patient's TCM disease record data set, and uses a disease dynamic development prediction network to conduct in-depth analysis of these data. This achieves precise dynamic monitoring of the development of TCM diseases. It can not only determine the first confidence level of the symptoms of the current case at the target time node, but also predict the second confidence level of the symptoms of the predicted case at future time nodes, thereby improving the accuracy and foresight of disease prediction. By continuously optimizing the disease dynamic development prediction network, the uncertainty in disease prediction is effectively reduced, providing more reliable data support for TCM clinical decision-making and helping to improve the effect of TCM diagnosis and treatment. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the execution flow of the method for dynamic monitoring of TCM disease based on big data time series analysis provided in this embodiment of the invention.
[0045] Figure 2 This is a schematic diagram of the hardware architecture of a dynamic monitoring system for TCM disease conditions based on big data time series analysis provided in an embodiment of the present invention. Detailed Implementation
[0046] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a method for dynamic monitoring of TCM disease conditions based on big data time series analysis, provided in one embodiment of the present invention. The following is a detailed description of this method for dynamic monitoring of TCM disease conditions based on big data time series analysis.
[0047] Step S110: Obtain the past sample patient condition characteristic data set, the symptoms of the current sample patient, and the predicted symptoms of the sample patient from the patient's TCM condition record data set. The patient's TCM condition record data set includes the symptoms of multiple sample patients organized based on time series. The past sample patient condition characteristic data set includes the symptoms of multiple sample patients located before the target time node. The symptoms of the current sample patient are the symptoms of the sample patient corresponding to the target time node. The predicted symptoms of the sample patient are the symptoms of the sample patient after the target time node.
[0048] In this embodiment, the server begins processing patient medical data from a traditional Chinese medicine (TCM) medical institution. This institution's patient medical record dataset contains numerous patients' TCM treatment records over many years. These records are organized chronologically, detailing each patient's various symptoms at different points in time. For example, for patient A, the record begins with their initial consultation and includes information such as pulse characteristics (wiry pulse, slippery pulse, etc.), tongue characteristics (red tongue with yellow coating, pale tongue with white coating, etc.), physical discomfort symptoms (headache, fatigue, cough, etc.), and other relevant TCM diagnostic information (such as Qi and blood status, organ function, etc.).
[0049] The server now needs to determine a target time point for a specific analysis task. Let's say this target time point is the 30th day after patient A's initial visit. The server retrieves relevant data from the patient's TCM medical record dataset. The past sample disease characteristic data set consists of all of patient A's symptom records before this 30th day. For example, between day 1 and day 29, patient A's pulse was consistently wiry, her tongue was red with a yellow coating, and she experienced intermittent headaches and fatigue. This symptom data constitutes the past sample disease characteristic data set.
[0050] The symptoms presented by the patient at this node are those corresponding to the target time point of day 30. For example, the record on day 30 shows that patient A's pulse became thready, headache worsened, and new symptoms such as irritability appeared. These are the symptoms presented by the patient at this node.
[0051] The predicted symptoms of the sample patients are those exhibited after day 30. The server randomly extracts these symptoms from subsequent records (according to one possible implementation). For example, if patient A's pulse reverts to a wiry pulse between days 31 and 35, and their headache symptoms lessen but they develop insomnia, these are used as the predicted symptoms of the sample patients. By accurately obtaining these three types of data from a large-scale dataset of patients' TCM medical records, a foundational data source is provided for predicting the dynamic development of the disease.
[0052] Step S120: Load the set of past case disease characteristics data and the symptoms of the current node case patient into the disease dynamic development prediction network, and determine the first confidence level of the occurrence of the symptoms of the current node case patient at the target time node.
[0053] In this embodiment, the server loads the previously acquired set of past case disease feature data (such as the symptom data of patient A from day 1 to day 29) and the symptoms of the current node's case patient (symptom data on day 30) into a pre-constructed disease dynamic development prediction network. This disease dynamic development prediction network is a neural network model trained on a large amount of TCM disease data, which can learn the complex relationships between different symptoms and the development pattern of the disease over time.
[0054] For example, neurons in a network may correspond to different symptom features or combinations of symptoms, and the connection weights between neurons represent the degree of correlation between these symptoms. When given a set of past sample disease feature data and the symptoms exhibited by the current node's sample patient, the network performs a series of calculations. Suppose that, based on previously learned patterns, the network analyzes the symptoms exhibited by the current node's sample patient A on day 30, such as a change in pulse to a thin pulse, worsening headache, and irritability.
[0055] For example, a comprehensive assessment can be made based on the correlation between past symptoms (such as a wiry pulse, red tongue with yellow coating, intermittent headaches, and fatigue) and current symptoms, as well as the probability of such symptom shifts in a large number of other similar cases. If many patients in the large training data show similar symptom shifts around day 30 after similar past symptoms, the network will assign a high first confidence score. Conversely, if such symptom shifts are rare in the training data, the first confidence score will be low. For instance, after calculation, if the server obtains a first confidence score of 0.7 for the symptoms of the current node's sample patient appearing on day 30, it means that the network believes there is a 70% probability that patient A will experience these symptoms on day 30 based on existing data and learned patterns.
[0056] Step S130: Load the set of past case disease characteristics data and the predicted case patient symptoms into the disease dynamic development prediction network, and determine the second confidence level of the predicted case patient symptoms appearing after the target time node.
[0057] We continued to use the previous set of past case disease characteristic data (symptom data of patient A from day 1 to day 29), and at the same time, we loaded the predicted symptoms of the case patients (patient A’s pulse changed back to a stringy pulse, headache symptoms were relieved and insomnia symptoms appeared between day 31 and day 35) into the disease dynamic development prediction network.
[0058] Similarly, the network analyzes based on its internal structure and learned knowledge. Since the network has learned the time-series correlation patterns of different symptoms, it assesses the likelihood of these symptoms appearing after day 30 based on patient A's previous symptom development (symptoms from day 1 to day 29) and the symptoms exhibited by current predicted sample patients. For example, the network might discover that in previous cases, when patients had similar symptoms before day 30 and developed specific symptoms on day 30 (such as the day 30 symptoms in step S120), some patients subsequently experienced symptoms such as a return to a wiry pulse, reduced headache, and insomnia.
[0059] Based on factors such as the matching ratio and the strength of the association between symptoms, the network calculates a second confidence level. Assuming the calculated second confidence level is 0.6, it means that, based on the existing data and the network's learning results, there is a 60% probability that patient A will exhibit the symptoms predicted in these sample patients after day 30.
[0060] Step S140: Based on the loss between the labeled confidence level of the symptoms exhibited by the patient in this node and the first confidence level, optimize the disease dynamic development prediction network; and based on the loss between the labeled confidence level of the predicted patient's symptoms and the second confidence level, optimize the disease dynamic development prediction network.
[0061] First, the labeled confidence level for the symptoms presented by the patients in this example is a definite value based on actual medical diagnosis or, more accurately, expert evaluation. For instance, if patient A develops symptoms on day 30 (pulse becomes weak, headache worsens, and irritability, etc.), after careful expert evaluation and extensive clinical experience, the labeled confidence level is determined to be 0.8, meaning that the expert believes there is an 80% probability that the patient experienced these symptoms under these circumstances.
[0062] The server calculates the loss between the labeled confidence level (0.8) and the first confidence level (0.7) of the patient's symptoms in this node example. This loss can be calculated using various loss functions, such as mean squared error (MSE). After calculating the loss value, the server uses this loss value to adjust the parameters of the disease progression prediction network. For example, if using the backpropagation algorithm, the server calculates the gradient of each neuron's weight based on the loss value, and then updates the weights according to a certain learning rate, so that the network can give results closer to the labeled confidence level when dealing with similar situations in the future.
[0063] To predict the symptoms of sample patients, it is also necessary to determine the labeled confidence level. Following the method mentioned earlier, let's assume that the labeled confidence level for predicting the symptoms of sample patients (patient A's pulse returning to a wiry pulse, headache symptoms lessening, and insomnia symptoms appearing between day 31 and day 35) is 0.75. This is based on expert experience or more accurate diagnostic methods.
[0064] The server calculates the loss between the labeled confidence level (0.75) and the second confidence level (0.6) for predicting the symptoms of sample patients, also using an appropriate loss function. This loss value is then used to optimize the disease dynamics prediction network. For example, the weights of neurons related to symptom association analysis are adjusted to make the network's confidence assessment of predicting future symptom occurrences more accurate. By continuously optimizing the network based on the loss between the labeled confidence level and the network-calculated confidence level, the disease dynamics prediction network can continuously improve its prediction accuracy and better adapt to the complex dynamic changes in TCM diseases.
[0065] Throughout the process, by processing TCM disease record data, performing calculations in the disease dynamic development prediction network, and optimizing the network based on confidence loss, the ability to predict the dynamic development of TCM diseases is continuously improved, thereby providing more accurate and valuable support for TCM clinical diagnosis and health management.
[0066] Based on the above steps, this embodiment of the application extracts the characteristics of past cases, the current node, and the symptoms of predicted cases from the patient's TCM disease record data set, and uses a disease dynamic development prediction network to conduct in-depth analysis of these data. This achieves precise dynamic monitoring of the development of TCM diseases. It can not only determine the first confidence level of the symptoms of the current case at the target time node, but also predict the second confidence level of the symptoms of the predicted case at future time nodes, thereby improving the accuracy and foresight of disease prediction. By continuously optimizing the disease dynamic development prediction network, the uncertainty in disease prediction is effectively reduced, providing more reliable data support for TCM clinical decision-making and helping to improve the effect of TCM diagnosis and treatment.
[0067] In one possible implementation, step S140 includes:
[0068] Step S141: Obtain the first training supervision data and the second training supervision data of the predicted sample patient's symptoms. The proportion of positive sample labels used in the first training supervision data to characterize the predicted sample patient's symptoms is 1. The proportion of positive sample labels used in the second training supervision data to characterize the predicted sample patient's symptoms is x, where x is not less than 0 and x is not greater than 1.
[0069] Step S142: Based on the first training supervision data and the second training supervision data, generate the labeled confidence level of the predicted sample patient's symptoms.
[0070] Step S143: Optimize the disease dynamic development prediction network based on the loss between the labeled confidence level of the predicted sample patient's symptoms and the second confidence level.
[0071] In this embodiment, the first step is to acquire first and second training supervision data for predicting the symptoms of the sample patients. In the previously mentioned example of patient A, for the predicted symptoms of patient A between days 31 and 35, such as a return to a wiry pulse, reduced headache, and the appearance of insomnia, the proportion of positive sample labels used in the first training supervision data to characterize these predicted symptoms is 1. This means that from an ideal, completely deterministic perspective, these symptoms are considered to perfectly conform to a certain expected pattern. For example, in a large amount of existing, accurately diagnosed case data considered standard examples, when a patient previously had a similar symptom progression to patient A from day 1 to day 30, the appearance of these symptoms between days 31 and 35 is clearly identified as a typical disease progression, so the proportion of positive sample labels is set to 1.
[0072] The proportion of positive sample labels used in the second training supervision data to characterize the predicted symptoms of patient cases is x (0 ≤ x ≤ 1). This is a more flexible data set that considers more complex real-world situations. Continuing with patient A as an example, the server will consider more factors to determine this x value. For instance, the server will analyze the uncertainty of whether patients with similar initial symptoms will develop these symptoms between day 31 and day 35 due to individual differences (such as age, physical condition, and living environment). Suppose that in some young, physically fit patient groups, even if the initial symptoms are similar, the proportion of patients developing these symptoms later is relatively low. After statistical analysis of this group of patients, if the proportion of these symptoms is found to be 0.6, then this 0.6 might be determined as the x value. The determination of this x value is based on a comprehensive consideration of a large amount of data from different types of patients, not only limited to the symptoms themselves, but also involving the individual characteristics of patients and various complex relationships in the overall data.
[0073] Next, based on the first and second training supervised data, a labeled confidence score is generated to predict the symptoms of the sample patients. The server performs complex calculations based on the acquired training supervised data. For example, suppose the server uses a weighted fusion method to generate the labeled confidence score. For the first training supervised data, since it represents a positive sample label ratio of 1, which is very certain information, the server will give it a relatively high weight. For the second training supervised data, since its x-value (let's say 0.6) reflects a situation with some uncertainty, it will be given a relatively low weight. The specific weight allocation may be determined based on experience or previous data analysis, such as giving the first training supervised data a weight of 0.7 and the second training supervised data a weight of 0.3. Then, the labeled confidence score for predicting the symptoms of the sample patients is calculated as: 1 × 0.7 + 0.6 × 0.3 = 0.7 + 0.18 = 0.88. This 0.88 is the labeled confidence level obtained by combining two types of training and supervision data. It represents a comprehensive assessment confidence level that patient A will exhibit the symptoms predicted by these sample patients from day 31 to day 35, from a comprehensive perspective.
[0074] Finally, the disease progression prediction network is optimized based on the loss between the labeled confidence level (0.88) and the second confidence level (0.6 calculated previously) for predicting the symptoms of sample patients. The server uses an appropriate loss function to quantify this loss, such as mean squared error (MSE). The formula for calculating the loss is: MSE = (0.88 - 0.6)² = 0.0784. This loss value reflects the difference between the labeled confidence level for predicting the symptoms of sample patients and the second confidence level calculated by the network. The server then uses this loss value to optimize the disease progression prediction network.
[0075] During the optimization process, the server employs techniques such as backpropagation. The weights of each neuron in the disease progression prediction network are adjusted based on the loss value. Assuming the disease progression prediction network is a multi-layered neural network with complex connections between neurons, each connection has a weight. After calculating the loss value, the server updates the weights accordingly. For example, for the neuron connection weights related to predicting patient A's symptoms from day 31 to day 35, the server updates the weights according to a certain learning rate (assuming a learning rate of 0.01). If a weight is w, the updated weight w' = w - 0.01 × (∂MSE / ∂w), where ∂MSE / ∂w is the partial derivative of the loss function MSE with respect to the weight w. By continuously adjusting the weights based on the loss, the disease progression prediction network can continuously learn and improve, enabling more accurate calculation of confidence levels and improving the accuracy of predicting the dynamic progression of diseases in Traditional Chinese Medicine when handling similar prediction tasks in the future. Throughout the process, the server precisely processes various data, calculates relevant values, and uses these results to optimize the network to adapt to the complex and ever-changing realities of TCM patients.
[0076] In one possible implementation, the disease progression prediction network is a taught disease progression prediction network. Acquiring second training supervision data for the predicted patient's symptoms includes: loading the set of past patient disease feature data and the predicted patient's symptoms into the teaching disease progression prediction network, and determining a third confidence level for the predicted patient's symptoms to appear after a target time point. The number of network parameters in the teaching disease progression prediction network is not less than the number of network parameters in the taught disease progression prediction network. The third confidence level is used as the second training supervision data for the predicted patient's symptoms.
[0077] In this embodiment, firstly, the server performs the following operations to obtain the second training supervision data for predicting the symptoms of the sample patients. Taking the previously mentioned patient A as an example, the server has already obtained the set of past sample disease characteristic data of patient A (such as symptom data from day 1 to day 29) and the predicted symptoms of the sample patients (pulse returning to a wiry pulse, headache symptoms lessening, and insomnia appearing between day 31 and day 35). The server loads this data into the teaching disease dynamic development prediction network. This teaching disease dynamic development prediction network is a relatively complex network with a large number of network parameters, which is no less than the number of network parameters of the learned disease dynamic development prediction network. Due to the large number of network parameters, more potential relationships and patterns can be mined from the data.
[0078] When the server loads the set of past case disease characteristics data of patient A and the predicted symptoms of the patient into the teaching disease dynamic development prediction network, the network performs calculations based on its complex internal structure and pre-learned knowledge. This network contains numerous neurons, each corresponding to different disease characteristics or combinations of characteristics, and the connection weights between neurons reflect the degree of correlation between these characteristics. The network comprehensively considers the relationship between patient A's previous symptom development (such as pulse, tongue appearance, and other symptoms from day 1 to day 29) and the predicted symptoms of the patient (symptoms from day 31 to day 35). For example, the teaching disease dynamic development prediction network may have learned the probability relationship between the subsequent reduction of headache symptoms and insomnia symptoms under certain pulse and tongue appearance patterns. Through a series of complex calculations, the network determines the third confidence level for the predicted symptoms of the patient to appear after the target time point. Assuming the calculated third confidence level is 0.72, this 0.72 means that, according to the analysis of the teaching disease dynamic development prediction network, there is a 72% probability that patient A will experience these symptoms between day 31 and day 35. The server then uses this third confidence level as the second training supervision data to predict the symptoms of the sample patients.
[0079] In one possible implementation, the patient's condition dynamic development prediction network is the patient's condition dynamic development prediction network of the current round of network parameter optimization, and the teaching condition dynamic development prediction network is the teaching condition dynamic development prediction network of the current round of network parameter optimization. The method further includes: generating the teaching condition dynamic development prediction network of the current round of network parameter optimization based on the teaching condition dynamic development prediction network of the previous network parameter optimization stage and the patient's condition dynamic development prediction network of the current round of network parameter optimization.
[0080] In one possible implementation, the process of generating the teaching condition dynamic development prediction network for the current round of network parameter optimization based on the teaching condition dynamic development prediction network from the previous network parameter optimization phase and the student condition dynamic development prediction network from the current round of network parameter optimization includes:
[0081] The neuron weight information of the teaching disease dynamic development prediction network in the previous network parameter optimization stage is fused with the first training degradation factor to generate the first parameter information.
[0082] The neuron weight information of the predictive network for the dynamic development of the patient's condition during the current round of network parameter optimization is fused with a second training degradation factor to generate second parameter information. The sum of the first training degradation factor and the second training degradation factor is 1.
[0083] Fuse the first parameter information and the second parameter information to generate the neuron weight information of the taught disease condition dynamic development prediction network in this round of network parameter optimization phase.
[0084] Next, in a possible implementation manner, the taught disease condition dynamic development prediction network is the taught disease condition dynamic development prediction network in this round of network parameter optimization phase, the taught disease condition dynamic development prediction network is the taught disease condition dynamic development prediction network in this round of network parameter optimization phase, and there is an operation of generating the taught disease condition dynamic development prediction network in this round of network parameter optimization phase based on the taught disease condition dynamic development prediction network in the previous network parameter optimization phase and the taught disease condition dynamic development prediction network in this round of network parameter optimization phase.
[0085] When the server performs this operation, it first fuses the neuron weight information of the taught disease condition dynamic development prediction network in the previous network parameter optimization phase with the first training degradation factor to generate the first parameter information. Assume that the weight of a certain neuron in the taught disease condition dynamic development prediction network in the previous network parameter optimization phase is w1, and the first training degradation factor is a (0 < a < 1), then the weight w1' of this neuron after fusion is w1 * a. Such an operation is performed on all neurons in the network to generate the first parameter information. In this process, the role of the first training degradation factor is to adjust the neuron weight information of the previous taught disease condition dynamic development prediction network to a certain extent to meet the new network optimization requirements.
[0086] At the same time, the server fuses the neuron weight information of the taught disease condition dynamic development prediction network in this round of network parameter optimization phase with the second training degradation factor to generate the second parameter information. Assume that the weight of a certain neuron in the taught disease condition dynamic development prediction network in this round of network parameter optimization phase is w2, and the second training degradation factor is b (0 < b < 1 and a + b = 1), then the weight w2' of this neuron after fusion is w2 * b. Similarly, such an operation is performed on all neurons in the network to obtain the second parameter information.
[0087] Finally, the server fuses the first and second parameter information to generate the neuron weight information for the teaching-based disease dynamic development prediction network in this round of network parameter optimization. For each neuron in the network, the weight w1' from the first parameter information and the weight w2' from the second parameter information are fused. For example, the fused neuron weight w3 = w1' + w2'. By performing this operation on all neurons in the network, the generation of the neuron weight information for the teaching-based disease dynamic development prediction network in this round of network parameter optimization is completed. This fusion method combines the advantages of the teaching-based disease dynamic development prediction network from the previous network parameter optimization stage and the teaching-based disease dynamic development prediction network in this round of network parameter optimization, enabling the newly generated teaching-based disease dynamic development prediction network in this round of network parameter optimization to better adapt to the needs of disease dynamic development prediction and improve the accuracy and reliability of prediction. Throughout the process, the server precisely processes various data, network weight information, and different factors to optimize and improve the disease dynamic development prediction network, thereby better addressing the complex and ever-changing realities of TCM diseases.
[0088] When processing these tasks, the server needs to handle a large amount of patient condition data and perform various calculations and operations with precision. From obtaining second training supervision data from the teaching disease dynamic development prediction network to generating new teaching disease dynamic development prediction networks based on networks at different stages, each step involves a deep understanding and precise operation of the data and network structure. For example, in processing patient A's condition data, the server must ensure the accuracy of data at each stage and the correctness of parameter transformation and fusion between different networks. This requires the server to have powerful computing and data processing capabilities to cope with the complexity and diversity of TCM condition data. At the same time, this complex network construction and optimization process is also aimed at improving the accuracy of disease prediction and providing more valuable support for TCM clinical diagnosis and health management. By continuously optimizing the teaching disease dynamic development prediction network, the server can better utilize existing condition data, uncover more valuable disease development patterns, and thus provide more accurate services for patient condition prediction and health management.
[0089] In one possible implementation, step S142 includes:
[0090] Step S1421: The second training supervision data is fused with the first influence factor to generate the third training supervision data.
[0091] Step S1422: The first training supervision data is fused with the second influence factor to generate the fourth training supervision data. The sum of the first influence factor and the second influence factor is 1.
[0092] Step S1423: The third training supervision data and the fourth training supervision data are fused to generate the labeled confidence scores of the predicted sample patient's symptoms.
[0093] In one possible implementation, the method further includes: determining a reference impact factor that matches the second training supervision data from a plurality of candidate first impact factors; and using the reference impact factor as the first impact factor of the second training supervision data.
[0094] In this embodiment, taking the case of the aforementioned patient A as an example, the first training supervision data (the proportion of positive sample labels used to characterize the symptoms of the predicted sample patient is 1) and the second training supervision data (assuming that the third confidence level of patient A's symptoms from day 31 to day 35 obtained in the previous teaching of the disease dynamic development prediction network is 0.72, this value is used as the second training supervision data).
[0095] First, the server determines a reference impact factor from multiple candidate first impact factors that matches the second training-supervised data, and uses this reference impact factor as the first impact factor for the second training-supervised data. The server maintains a set of candidate first impact factors, determined based on extensive prior data processing and analysis experience. Each candidate first impact factor has associated data characteristics or conditional ranges. For example, when the value of the second training-supervised data falls within a certain range, it may correspond to a specific candidate first impact factor. Suppose that when the value of the second training-supervised data is in the range of 0.7 - 0.8, the corresponding candidate first impact factor is 0.3. Since the previously obtained second training-supervised data value is 0.72, this 0.3 is determined as the reference impact factor matching the second training-supervised data, which is also the first impact factor for the second training-supervised data.
[0096] Next, the server fuses the second training-supervised data with the first influence factor to generate the third training-supervised data. In the case of patient A, the second training-supervised data is 0.72, and the first influence factor is 0.3. Therefore, the third training-supervised data = 0.72 * 0.3 = 0.216. This calculation process adjusts the second training-supervised data according to the proportion of the first influence factor to obtain a new intermediate data, namely the third training-supervised data.
[0097] Simultaneously, the server fuses the first training supervision data with the second influence factor to generate the fourth training supervision data. Since the sum of the first and second influence factors is 1, and the first influence factor is already determined to be 0.3, the second influence factor is 1 - 0.3 = 0.7. The first training supervision data is 1 (representing a proportion of positive sample labels is 1), therefore the fourth training supervision data = 1 * 0.7 = 0.7. This process also adjusts the first training supervision data according to the corresponding proportions to obtain the fourth training supervision data.
[0098] Finally, the server merges the third and fourth training supervision data to generate a labeled confidence score for predicting the symptoms of the predicted sample patients. For patient A, the third training supervision data score is 0.216, and the fourth training supervision data score is 0.7, so the labeled confidence score = 0.216 + 0.7 = 0.916. This labeled confidence score represents the overall confidence level for predicting that patient A will exhibit the symptoms of the predicted sample patients between days 31 and 35, after comprehensively considering the first and second training supervision data and their respective influencing factors.
[0099] Throughout the process, the server precisely identifies the first influencing factor that matches the second training supervision data, and adjusts and fuses the first and second training supervision data according to rigorous calculation rules to obtain the labeled confidence score for predicting the symptoms of the sample patients. This labeled confidence score will be used subsequently to optimize the disease dynamic development prediction network to improve the accuracy of the network's disease development prediction. In processing these operations, the server needs to precisely manage and calculate large amounts of data and various parameters, which relies on its powerful computing and data processing capabilities. For example, when determining the matching reference influencing factor from the candidate first influencing factors, it needs to quickly and accurately find the appropriate influencing factor among many candidates based on the values of the second training supervision data. During data fusion calculations, the accuracy of the calculations must also be ensured to guarantee that the final labeled confidence score accurately reflects the comprehensive assessment of the disease prediction. This precise data processing and calculation process is a crucial guarantee for the effective operation of the entire disease dynamic development prediction system, and helps to improve the application value of traditional Chinese medicine disease prediction in clinical practice and health management.
[0100] In one possible implementation, obtaining the predicted sample patient's symptoms includes:
[0101] The predicted symptoms of the patient are randomly extracted from the patient's TCM medical record data set, which includes the symptoms of patients following the symptoms of the patient in this node.
[0102] In one possible implementation, the method further includes:
[0103] Based on the optimized disease dynamic development prediction network, the target disease feature data set and the target patient's symptoms are analyzed to predict the target confidence level of the time node corresponding to the target patient's symptoms.
[0104] In this embodiment, when obtaining the predicted symptoms of the sample patients, the symptoms are randomly extracted from the patient's TCM medical record data set after the symptoms of the sample patient at this node. Taking the previously mentioned patient A as an example, patient A's TCM medical record data set contains numerous symptom information organized chronologically from the initial consultation. After determining the symptoms of the sample patient at this node (assuming patient A's pulse becomes thready, headache worsens, and irritability occurs on day 30), the server will view the symptoms of all sample patients after the symptoms of the sample patient at this node, that is, the symptom records from day 31 onwards. These records contain various possible changes in symptoms, such as further changes in pulse, relief or worsening of headache, the appearance of new symptoms (such as insomnia, cough, etc.), and changes in other TCM-related symptoms such as tongue appearance and qi and blood. The server randomly extracts from these subsequent symptom records to obtain the predicted symptoms of the sample patients. For example, the server might randomly select symptom records from patient A between day 31 and day 35, where the pulse changes back to a wiry pulse, headache symptoms lessen, and insomnia appears. These selected symptoms become the predictive sample patient's symptoms. This random sampling method can, to some extent, represent multiple possibilities for the development of a patient's condition, providing different sample data for predicting the subsequent dynamic development of the condition.
[0105] After completing the above operations and optimizing the disease progression prediction network based on the previous steps, the server still needs to perform an operation based on the optimized disease progression prediction network to analyze any input target disease feature data set and target patient symptoms, and predict the target confidence score for the time point corresponding to the target patient's symptoms. Suppose there is a new patient B, whose target disease feature data set includes patient B's initial symptom information, such as pulse (e.g., deep pulse), tongue appearance (e.g., pale tongue with white coating), physical discomfort symptoms (e.g., fatigue, aversion to cold), and other information related to traditional Chinese medicine diagnosis (e.g., qi and blood deficiency). The target patient's symptoms may be symptoms that appear at a specific time point (e.g., 15 days after the initial visit), such as a possible worsening of cough or changes in pulse.
[0106] The server inputs the target disease feature data set and the target patient's symptoms into the optimized disease dynamic development prediction network. This optimized network has been trained and optimized using a large amount of previous patient data (such as the data processing of patient A), and its internal neuron weights and other parameters have been adjusted to a relatively accurate state. The neurons in the network correspond to different disease features or combinations of features, and the connection weights between neurons represent the degree of correlation between these features. When the target disease feature data set and the target patient's symptoms are input into the network, it performs complex calculations based on its internal structure and learned patterns.
[0107] First, the network identifies the initial disease state pattern of patient B based on initial symptom information (such as deep pulse, pale tongue with white coating, fatigue, and aversion to cold) in the target disease feature dataset. Then, combining the target patient's symptoms (symptoms that may appear on day 15), the network searches for similar patterns in its large set of learned disease development patterns. For example, the network may discover a proportion of patients in previously processed patient data who initially had similar pulse, tongue appearance, and physical discomfort symptoms, and around day 15, developed symptoms such as worsening cough and changes in pulse.
[0108] Based on these matching results and factors such as the strength of associations between different learned symptoms, the network calculates the target confidence level for the occurrence of the target patient's symptoms on day 15. Assuming a calculated target confidence level of 0.65, this means that, based on the existing data and the network's learning results, there is a 65% probability that patient B will exhibit these target patient symptoms on day 15 after the initial visit. In this way, the server uses the optimized disease dynamic development prediction network to predict the condition of new patients, providing valuable reference for TCM clinical diagnosis and health management. Throughout the process, the server accurately acquires the predicted symptoms of sample patients and uses the optimized network to predict the condition of new patients, demonstrating the effectiveness and practicality of the entire system in TCM disease analysis and prediction.
[0109] For example, in one possible implementation, prior to step S110, the method further includes:
[0110] Step S101: Perform structural analysis on the patient TCM condition record data in the patient TCM condition record data set to determine the different types of data elements and the relationship structure between each data element.
[0111] Step S102: Based on prior knowledge and data and privacy protection requirements in the field of TCM health management, a semantic classification framework is constructed. The semantic classification framework is used to classify each data element according to its semantic category in TCM theory and to define corresponding semantic tags and related attributes. The semantic tags will be used to perform privacy semantic encoding on the data elements.
[0112] Step S103: According to the semantic classification framework, each data element in the patient's TCM condition record data is labeled. For each data element, the category to which the data element belongs in the semantic classification framework is found, and a corresponding semantic label is assigned to the data element. Based on the labeled semantic label, privacy weight information is added to each data element. The privacy weight information is determined based on the importance of the data element in privacy protection. In this way, the data elements labeled with semantic labels and privacy weights are recombined to form a preliminary data structure after privacy semantic encoding.
[0113] Step S104: Extract various health data patterns from the preliminary data structure after privacy semantic encoding, perform feature description on each health data pattern, and establish a health data pattern library. The feature description includes the combination of data elements constituting the health data pattern, the value range or feature value of each data element, and the logical relationship between each data element.
[0114] Step S105: For each data subset in the preliminary data structure after privacy semantic encoding, match it with various health data patterns in the health data pattern library. The data subset is a data part divided according to different diagnostic stages or body systems.
[0115] Step S106: For a successfully matched subset of data, mark the name or number of the health data pattern matched by the successfully matched subset, record the similarity degree when the match is successful, and generate a health data pattern matching result. The similarity degree is quantitatively evaluated based on the proportion of matching data elements and the matching status of key data elements. In addition, a special mark is made for a subset of data that does not match any health data pattern, indicating that the subset of data is an abnormal subset of data or a new unidentified health data pattern.
[0116] Step S107: Integrate the preliminary data structure of privacy semantic encoding with the health data pattern matching results. Specifically, for each data element, associate the semantic label, privacy weight information, matched health data pattern, and similarity information of the data element. Based on the association results, construct a corresponding nested structure. The outer structure of the nested structure contains the semantic category information of the data element, and the inner structure contains the data element itself, privacy weight, and pattern matching result information.
[0117] Step S108: Perform privacy protection preprocessing on the patient's TCM condition record data based on the nested structure.
[0118] In this embodiment, the server first performs structural analysis on the patient TCM medical record data in the patient TCM medical record dataset. Taking patient data from a large TCM medical institution as an example, this patient TCM medical record data contains a wealth of information. The server carefully sorts through this data to determine the different types of data elements contained therein and the relationship structure between these data elements. For example, the patient's basic information (name, age, gender, etc.), TCM diagnostic information (pulse, tongue appearance, Qi and blood status, etc.), symptom descriptions (headache, fatigue, cough, etc.), and treatment process (medication, acupuncture points, etc.) are all different types of data elements. These data elements have complex relationship structures; for example, age may be related to the incidence probability of certain diseases, pulse and tongue appearance may jointly reflect the patient's organ function status, and medication may be adjusted according to the diagnosis results and symptoms.
[0119] Next, based on prior knowledge in the field of Traditional Chinese Medicine (TCM) health management and the requirements for data and privacy protection, the server constructs a semantic classification framework. This framework classifies each data element according to semantic categories in TCM theory. For example, pulse data elements are classified according to the meanings represented by different pulse types in TCM theory. For instance, a wiry pulse might be classified as a pulse type related to liver function, with its semantic label defined as "Liver Pulse Related - Wiry Pulse," and related attributes such as pulse characteristics (straight and long, like pressing a string). For headaches in symptom descriptions, they might be classified according to the location and nature of the pain, such as "Head Yangming Meridian Headache - Frontal Pain" as the semantic label, with related attributes including the frequency of pain attacks and accompanying symptoms. These semantic labels will be used for privacy-preserving semantic encoding of data elements.
[0120] Then, following the constructed semantic classification framework, the server labels each data element in the patient's TCM medical record data. For each data element, the server looks up its category within the semantic classification framework and assigns it a corresponding semantic label. Next, based on the semantic labels, privacy weight information is added to each data element. This privacy weight is determined based on the importance of the data element in privacy protection. For example, basic information such as the patient's name and ID number has a high privacy weight because its leakage could directly expose the patient's privacy, so it may be assigned a higher value, such as 0.9; while some general symptom descriptions, such as occasional mild cough, have a relatively low privacy weight, perhaps assigned 0.3. In this way, the data elements labeled with semantic tags and privacy weights are recombine to form a preliminary data structure after privacy semantic encoding.
[0121] From the initial data structure encoded with privacy semantics, the server begins to extract various health data patterns. Taking a common ailment (such as the common cold) as an example, there might be a health data pattern whose characteristic description includes a combination of data elements such as a rapid and floating pulse, a thin white tongue coating, and symptoms including fever, headache, and cough. The rapid and floating pulse ranges from 90 to 100 beats per minute, and the thin white tongue coating is characterized by a thin, white tongue coating. The logical relationship between the data elements is that the rapid and floating pulse, the thin white tongue coating, and symptoms such as fever, headache, and cough have a causal or accompanying relationship. The server provides such a detailed characteristic description for each identified health data pattern, thereby establishing a health data pattern library.
[0122] For each data subset in the initial data structure after privacy semantic encoding (these subsets are data parts divided according to different diagnostic stages or body systems; for example, data related to respiratory symptoms, diagnosis, and treatment are divided into one subset, and data related to the digestive system are divided into another subset), the server matches it against various health data patterns in the health data pattern library. For the respiratory system-related data subset, which includes data elements such as pulse, tongue appearance, cough, and respiratory rate, the server compares these data elements with various patterns in the health data pattern library. If a data subset has a floating and rapid pulse, a thin white tongue, frequent cough, and a slightly increased respiratory rate, it matches the previously established cold health data pattern.
[0123] For successfully matched subsets of data, the server labels the health data pattern matched within that subset with its name or number (e.g., "Cold - Pattern 1") and records the similarity score at the time of the match. This similarity score is quantitatively evaluated based on the proportion of matching data elements and the matching status of key data elements. For example, if the pulse, tongue, and cough—three key data elements—in a subset of data perfectly match the cold health data pattern, while the change in respiratory rate differs slightly from the standard value in the pattern, the calculated similarity score might be 0.8. For subsets of data that do not match any health data pattern, the server specially labels them as abnormal subsets or new, unrecognized health data patterns. For example, if a subset of data contains a specific pulse and a combination of symptoms never recorded before, the server labels it as a new, unrecognized health data pattern for further research.
[0124] Finally, the server integrates the preliminary data structure with privacy semantic encoding with the health data pattern matching results. Specifically, for each data element, its semantic label, privacy weight information, matched health data pattern, and similarity score are associated, and a corresponding nested structure is constructed based on the association results. For example, for a data element (such as a pulse pattern), the outer structure contains its semantic category information (pulse pattern - related to external pathogens), and the inner structure contains the data element itself (pulse pattern), privacy weight (0.3), and pattern matching result information (matched to a cold - pattern 1, similarity score 0.8). Through this nested structure, the server can better perform privacy-preserving preprocessing on patients' TCM medical record data. For example, in subsequent data processing, if data sharing or analysis is involved, the server can decide how to process the data based on privacy weights and pattern matching results, satisfying the needs of data analysis while effectively protecting patient privacy. Simultaneously, this structure also helps to more accurately identify and analyze medical data, laying a solid foundation for subsequent operations such as obtaining relevant medical characteristic data from the patient's TCM medical record data set. Throughout the process, the server, through a rigorous data processing workflow, effectively managed, protected privacy, and prepared for data mining of TCM patient condition records.
[0125] Figure 2 This invention illustrates the hardware structure of a TCM disease dynamic monitoring system 100 based on big data time series analysis, provided by an embodiment of the present invention, for implementing the above-described method for dynamic monitoring of TCM disease based on big data time series analysis. Figure 2 As shown, the TCM disease dynamic monitoring system 100 based on big data time series analysis may include a processor 110, a machine-readable storage medium 120, a bus 130, and a communication unit 140.
[0126] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data acquired from an external terminal. In some embodiments, the machine-readable storage medium 120 may store data and / or instructions used by the Traditional Chinese Medicine Disease Dynamic Monitoring System 100 based on big data time series analysis to perform or use in order to accomplish the exemplary methods described in this invention.
[0127] In the specific implementation process, one or more processors 110 execute computer-executable instructions stored in machine-readable storage medium 120, so that processor 110 can execute the dynamic monitoring method of TCM disease based on big data time series analysis as described in the above method embodiment. Processor 110, machine-readable storage medium 120 and communication unit 140 are connected through bus 130. Processor 110 can be used to control the sending and receiving actions of communication unit 140.
[0128] The specific implementation process of processor 110 can be found in the various method embodiments executed by the above-mentioned TCM disease dynamic monitoring system 100 based on big data time series analysis. The implementation principle and technical effect are similar, and will not be repeated here.
[0129] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned method for dynamic monitoring of TCM disease based on big data time series analysis is implemented.
[0130] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A method for dynamic monitoring of TCM disease conditions based on big data time series analysis, characterized in that, The method includes: The TCM patient condition record data set is used to obtain past sample condition characteristic data set, current sample patient symptoms, and predicted sample patient symptoms. The TCM patient condition record data set includes multiple sample patient symptoms organized based on time series. The past sample condition characteristic data set includes multiple sample patient symptoms located before the target time node. The current sample patient symptoms are the sample patient symptoms corresponding to the target time node. The predicted sample patient symptoms are the sample patient symptoms after the target time node. The set of past case disease characteristics data and the symptoms of the current node patient are loaded into the disease dynamic development prediction network to determine the first confidence level of the occurrence of the symptoms of the current node patient at the target time node; The set of past case disease characteristics data and the predicted case patient symptoms are loaded into the disease dynamic development prediction network to determine the second confidence level of the predicted case patient symptoms appearing after the target time node; Based on the loss between the labeled confidence level of the symptoms exhibited by the patient in this node and the first confidence level, the disease dynamic development prediction network is optimized, and based on the loss between the labeled confidence level of the symptoms exhibited by the predicted patient and the second confidence level, the disease dynamic development prediction network is optimized. Before the steps of obtaining the set of past sample disease characteristics data, the symptoms of the current sample patient, and the predicted symptoms of the sample patient from the patient's TCM disease record data set, the method further includes: A structural analysis was performed on the patient TCM condition record data in the aforementioned patient TCM condition record dataset to determine the different types of data elements included and the relationship structure between each data element. Based on prior knowledge, data, and privacy protection requirements in the field of traditional Chinese medicine health management, a semantic classification framework is constructed. This framework is used to classify each data element according to its semantic category in traditional Chinese medicine theory and to define corresponding semantic tags and related attributes. The semantic tags will be used to perform privacy semantic encoding on the data elements. According to the semantic classification framework, each data element in the patient's TCM condition record data is labeled. For each data element, the category to which the data element belongs in the semantic classification framework is found, and a corresponding semantic label is assigned to the data element. Based on the labeled semantic label, privacy weight information is added to each data element. The privacy weight information is determined based on the importance of the data element in privacy protection. In this way, the data elements labeled with semantic labels and privacy weights are recombined to form a preliminary data structure after privacy semantic encoding. Various health data patterns are extracted from the initial data structure after privacy semantic encoding. Each health data pattern is characterized and a health data pattern library is established. The characteristic description includes the combination of data elements that constitute the health data pattern, the value range or characteristic value of each data element, and the logical relationship between each data element. For each data subset in the initial data structure after privacy semantic encoding, it is matched with various health data patterns in the health data pattern library. The data subset is a data part divided according to different diagnostic stages or body systems. For a successfully matched subset of data, the name or number of the health data pattern matched by the successfully matched subset is marked, and the similarity degree when the match is successful is recorded to generate a health data pattern matching result. The similarity degree is quantitatively evaluated based on the proportion of matching data elements and the matching status of key data elements. In addition, a special mark is made for a subset of data that does not match any health data pattern, indicating that the subset of data is an abnormal subset of data or a new unidentified health data pattern. The preliminary data structure of privacy semantic encoding is integrated with the health data pattern matching results. Specifically, for each data element, the semantic label, privacy weight information, matched health data pattern and similarity information of the data element are associated. Based on the association results, a corresponding nested structure is constructed. The outer structure of the nested structure contains the semantic category information of the data element, and the inner structure contains the data element itself, privacy weight, and pattern matching result information. Based on the nested structure, the patient's TCM medical record data is preprocessed for privacy protection.
2. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to claim 1, characterized in that, The optimization of the disease progression prediction network based on the loss between the labeled confidence level and the second confidence level of the predicted patient symptoms includes: Obtain first training supervision data and second training supervision data of the predicted sample patient's symptoms. The proportion of positive sample labels used in the first training supervision data to characterize the predicted sample patient's symptoms is 1. The proportion of positive sample labels used in the second training supervision data to characterize the predicted sample patient's symptoms is x, where x is not less than 0 and x is not greater than 1. Based on the first training supervision data and the second training supervision data, the labeled confidence scores of the predicted sample patient's symptoms are generated; The disease progression prediction network is optimized based on the loss between the labeled confidence level of the predicted patient's symptoms and the second confidence level.
3. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to claim 2, characterized in that, The disease progression prediction network is a taught disease progression prediction network; the second training supervision data for the symptoms exhibited by the predicted sample patients is obtained, including: The set of past case disease characteristics data and the predicted case patient symptoms are loaded into the teaching disease dynamic development prediction network to determine the third confidence level of the predicted case patient symptoms appearing after the target time node; the number of network parameters of the teaching disease dynamic development prediction network is not less than the number of network parameters of the teaching disease dynamic development prediction network. The third confidence level is used as the second training supervision data for the predicted symptoms of the patient.
4. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to claim 3, characterized in that, The patient's condition dynamic development prediction network is the patient's condition dynamic development prediction network in the current round of network parameter optimization phase, and the teaching condition dynamic development prediction network is the teaching condition dynamic development prediction network in the current round of network parameter optimization phase; the method further includes: Based on the teaching condition dynamic development prediction network of the previous network parameter optimization stage and the teaching condition dynamic development prediction network of the current network parameter optimization stage, the teaching condition dynamic development prediction network of the current network parameter optimization stage is generated.
5. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to claim 4, characterized in that, The method of generating the teaching condition dynamic development prediction network for the current round of network parameter optimization based on the teaching condition dynamic development prediction network of the previous network parameter optimization stage and the teaching condition dynamic development prediction network of the current round of network parameter optimization includes: The neuron weight information of the teaching disease dynamic development prediction network in the previous network parameter optimization stage is fused with the first training degradation factor to generate the first parameter information; The neuron weight information of the dynamic development prediction network of the patient's condition in the current network parameter optimization stage is fused with the second training degradation factor to generate the second parameter information; the sum of the first training degradation factor and the second training degradation factor is 1. The first parameter information and the second parameter information are fused to generate the neuron weight information of the teaching disease dynamic development prediction network in the current round of network parameter optimization stage.
6. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to claim 2, characterized in that, The step of generating labeled confidence scores for the predicted patient symptoms based on the first training supervision data and the second training supervision data includes: The second training supervision data is fused with the first influence factor to generate the third training supervision data; The first training supervision data is fused with the second influence factor to generate the fourth training supervision data; the sum of the first influence factor and the second influence factor is 1. The third and fourth training supervision data are fused together to generate the labeled confidence scores of the predicted patient symptoms.
7. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to claim 6, characterized in that, The method further includes: From multiple candidate first impact factors, a reference impact factor that matches the second training supervision data is determined; the reference impact factor is used as the first impact factor of the second training supervision data.
8. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to any one of claims 1-7, characterized in that, Obtaining the symptoms exhibited by the predicted sample patients includes: The predicted symptoms of the patient are randomly extracted from the patient's TCM medical record data set, which includes the symptoms of patients following the symptoms of the patient in this node.
9. The method for dynamic monitoring of TCM disease conditions based on big data time series analysis according to any one of claims 1-7, characterized in that, The method further includes: Based on the optimized disease dynamic development prediction network, the target disease feature data set and the target patient's symptoms are analyzed to predict the target confidence level of the time node corresponding to the target patient's symptoms.
10. A dynamic monitoring system for TCM disease conditions based on big data time series analysis, characterized in that, The TCM disease dynamic monitoring system based on big data time series analysis includes a processor and a memory, the memory and the processor are connected, the memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to realize the TCM disease dynamic monitoring method based on big data time series analysis as described in any one of claims 1-9.
Citation Information
Patent Citations
Traditional Chinese medicine auxiliary diagnosis system
CN109346185A
Deep neural network medical image diagnosis method based on confidence coefficient calibration
CN114549469A