Tuberculosis data management control system based on AI disease simulation model
By developing a tuberculosis data management and control system based on AI disease simulation model, deeply integrating multiomics data and analyzing it using AI algorithms, the shortcomings in tuberculosis data management and diagnosis and treatment innovation in the existing technology are solved, and the precise analysis of the interaction mechanism between tuberculosis pathogens and hosts are achieved and the generation of personalized diagnosis and treatment paths are achieved.
Patent Information
- Application Number
- CN202510309725.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the existing technology to efficiently integrate multi-source data to achieve a comprehensive analysis of the interaction mechanism between tuberculosis pathogens and hosts, resulting in a lack of specificity and sensitivity in the research and development of diagnostic reagents, lack of personalized considerations in the treatment plan, and insufficient application space for AI technology in the entire process innovation of tuberculosis diagnosis and treatment.
Develop a tuberculosis data management and control system based on AI disease simulation model. By deeply integrating genomics, proteomics, metabolomics and other multiomic data, using AI algorithms for in-depth integration and analysis, identifying target molecules, developing diagnostic reagents, optimizing clinical trial plans, and generating personalized diagnosis and treatment paths.
The precise analysis of the interaction mechanism between tuberculosis pathogens and hosts has been achieved, the research and development efficiency and personalization of diagnostic reagents have been improved, the clinical trial plan has been optimized, more effective personalized diagnosis and treatment paths have been generated, and the safety and integrity of the data have been ensured.
Smart Images

Figure CN120164635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data management and control, and specifically to a tuberculosis data management and control system based on an AI disease simulation model. Background Art
[0002] As an infectious disease, tuberculosis remains a major challenge in the field of global public health to this day. Its complex and variable pathogen characteristics and host response mechanisms make the accurate diagnosis and effective treatment of tuberculosis the focus and difficulty of medical research. With the continuous progress of medical technology, especially the rapid development of multi-omics technologies such as genomics, proteomics, and metabolomics, it provides an unprecedented data basis for the in-depth study of tuberculosis. However, how to efficiently integrate and utilize these multi-source data to promote the innovation of tuberculosis diagnosis and treatment technologies has become an urgent problem to be solved.
[0003] Traditional tuberculosis diagnosis and treatment technologies mainly rely on the analysis of data with limited dimensions and empirical judgments, and it is difficult to comprehensively capture the complex interaction mechanisms between tuberculosis pathogens and the host. This leads to the development of diagnostic reagents often lacking sufficient specificity and sensitivity, and it is difficult to meet the needs of precision medicine. At the same time, the design of treatment plans also lacks personalized considerations and it is difficult to accurately match the specific conditions of different patients. In addition, there is still much room for improvement in existing patents in using AI technology to deeply integrate multi-source data to drive the innovation of the entire tuberculosis diagnosis and treatment process, especially in constructing an efficient data management and control system to coordinate multi-link work, where there are obvious shortcomings.
[0004] In view of the above problems, it is necessary to optimize the existing tuberculosis data management and control system. By deeply integrating multi-omics data such as genomics, proteomics, and metabolomics through AI technology, a comprehensive analysis of tuberculosis pathogens and their interaction mechanisms with the host has been achieved. Therefore, it is of great significance to develop a tuberculosis data management and control system based on an AI disease simulation model that can comprehensively achieve the above characteristics. Summary of the Invention
[0005] The purpose of the present invention is to make up for the deficiencies of the existing technology and provide a tuberculosis data management and control system based on an AI disease simulation model. It can deeply integrate and analyze genomics, proteomics, metabolomics and other multi-omics data through the use of AI algorithms to accurately identify the molecular characteristics of tuberculosis pathogens and their interactions with the host. At the same time, the system also covers multiple key links such as the precise research and development and optimization of diagnostic reagents, the simulation and optimization of clinical trials, and the generation of personalized diagnosis and treatment paths, realizing data management and intelligent control of the entire tuberculosis diagnosis and treatment process.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: A tuberculosis data management and control system based on an AI disease simulation model, which system comprises the following components: a data collection module, a multi-omics data integration and analysis module, a diagnostic reagent research and development and optimization module, a clinical trial simulation module, a personalized path generation module, and a data management center;
[0007] The data collection module constructs a multi-source data collection interface, establishes connections with hospital information systems, research institution databases, and chemical substance libraries, collects data including Mycobacterium tuberculosis gene sequences, protein structures, patient clinical information, known reagent properties, chemical substance library data, and multi-omics data, stores them in corresponding database areas after deduplication, error correction, and classification, and establishes data indexes for retrieval and invocation;
[0008] The multi-omics data integration and analysis module integrates genomics, proteomics, and metabolomics data, and constructs an analysis model based on deep learning. Taking the multi-omics data as input, it mines the features and patterns in the data, identifies target molecules specifically related to tuberculosis and having diagnostic value, transmits the identified target molecule data to the diagnostic reagent research and development and optimization module and the data management center, and simultaneously establishes a data transmission log;
[0009] The diagnostic reagent research and development and optimization module, based on the provided target molecule information, builds an AI model for the research and development and optimization of tuberculosis diagnostic reagents, uses deep learning algorithms to perform virtual screening on substances in the chemical substance library, calculates the affinity and binding energy parameters of the interaction by constructing a simulation environment for the interaction between substances and target molecules, predicts potential diagnostic markers, generates a list of marker candidates according to the virtual screening results, determines a preliminary diagnostic reagent formulation framework in combination with the component information in the chemical substance library, stores the formulation framework data and the list of marker candidates in the diagnostic reagent data sub-library, and simultaneously develops a formulation optimization algorithm based on patient individual data. Taking the patient's physical condition, disease severity, and drug resistance test result data as input, and combining with the preliminary formulation framework, it performs personalized optimization by adjusting the proportion of components in the formulation, stores the finally optimized diagnostic reagent formulation data and the entire research and development and optimization process data in the diagnostic reagent data sub-library, and conducts data interaction with the data management center;
[0010] The clinical trial simulation module constructs a mathematical model and an algorithm framework based on the pathophysiological mechanism, transmission law, and population infection characteristics of tuberculosis. Combining with AI algorithms, it inputs the collected data of different populations into the disease simulation model. At the same time, it designs a clinical trial simulation algorithm to pre-test the detection differences of reagents in different populations. By simulating different test scenarios, sample sizes, and detection process factors, it calculates the performance indicators of the reagents, discovers the factors affecting the clinical trial results in advance, generates optimization suggestions for the clinical trial plan according to the simulation results, and stores the simulation data and plan suggestions in the clinical trial data sub-library;
[0011] The personalized path generation module constructs a personalized diagnosis and treatment path generation model based on AI algorithms. Taking the individual data of patients as input, it predicts the possible development trends of the disease according to the natural history of tuberculosis and the individual characteristics of patients. Combining with the disease progression prediction results, based on the efficacy data and applicable population data of different treatment methods, it generates a personalized tuberculosis diagnosis and treatment path for patients, stores the personalized diagnosis and treatment path data in the diagnosis and treatment path data sub-library, establishes a data interaction channel with the data management center. When the individual data of patients changes, it receives notifications from the data management center, recalculates and adjusts the diagnosis and treatment path, and feeds back the adjusted path data to the data management center;
[0012] The data management center builds a centralized data storage architecture, including relational and non-relational databases to store the data of each module, and establishes an efficient data index and query system. It designs multi-mode query interfaces and implements permission management to ensure data security. At the same time, it formulates data backup and recovery strategies, regularly performs full and incremental backups of data to remote locations or the cloud, establishes a data analysis task scheduling mechanism, regularly integrates and analyzes the data of each module, monitors the data change trends and outliers, discovers anomalies, sends notifications to the corresponding modules, and records logs.
[0013] Furthermore, the multi-omics data integration and analysis module mines the features and patterns in the data, and identifies target molecules that are specifically related to tuberculosis and have diagnostic value. Its calculation formula is: where R ij is the association degree value between the i-th genomics data feature and the j-th proteomics data feature, reflecting the strength of the synergistic association between the two in the tuberculosis mechanism. M ik is the value of the i-th genomics data feature in the k-th sample, is the average value of the i-th genomics data feature in all samples, N jk is the value of the j-th proteomics data feature in the k-th sample, is the average value of the j-th proteomics data feature in all samples, N is the total number of samples, λ is the attenuation coefficient, d ijis the distance metric between the i-th genomics data feature and the j-th proteomics data feature in the feature space.
[0014] Furthermore, the diagnostic reagent R & D and optimization module uses a deep learning algorithm to perform virtual screening on the substances in the chemical library, and its algorithm formula is: where S m is the comprehensive score of the m-th substance as a tuberculosis diagnostic marker, and the score indicates the possibility of becoming an effective diagnostic marker. L is the number of features for evaluating the interaction between the substance and the target molecule, f1(C1, T m ) is a function regarding the l-th feature, C1 is the calculated value of the l-th feature in the interaction, T m is the chemical structure or property parameter related to the m-th substance, and w1 is the weight coefficient of the l-th feature.
[0015] Furthermore, the diagnostic reagent R & D and optimization module develops a formula optimization algorithm based on the patient's individual data, and its algorithm formula is: C p = α × C g + β × C r + γ × C s , where C p is the proportion vector of the optimized diagnostic reagent formula components, C g is the proportion vector of the basic formula components based on the diagnostic marker screening results, C r is the proportion vector of the formula components adjusted according to the patient's drug resistance data, C s is the proportion vector of the formula components adjusted based on the patient's physical condition, and α, β, γ are the weight coefficients corresponding to C g , C r , C s , which are determined according to the data of different patient groups to balance the influence of various factors on the formula.
[0016] Furthermore, in the diagnostic reagent R & D and optimization module, combined with the preliminary formula framework, personalized optimization is carried out by adjusting the proportion of the components in the formula. Specifically, the policy function π(α, β, γ|s) is defined, where s represents the system state, including the patient data features and the performance indicators of the current formula. The policy function represents the probability distribution of selecting the weight coefficients α, β, γ under the given state s. The goal is to have good treatment effect and few adverse reactions, so a high reward is given. Then the cumulative reward calculation formula is: where T is the treatment cycle or observation cycle, r t is the immediate reward at time t. According to the policy gradient theorem, the update formula of the weight coefficient is: where θ is the parameter of the policy function, including α, β, γ, N is the number of patient samples, α n,t , βn,t , γ n,t is the weight coefficient of the nth patient at time t, s n,t is the corresponding system state, R n is the cumulative reward of the nth patient. By updating the weight coefficient along the policy gradient direction, the policy can better adapt to the patient's response to achieve the purpose of optimizing the formula.
[0017] Furthermore, the clinical trial simulation module designs a clinical trial simulation algorithm to pre-test the detection differences of the reagent in different populations. Its algorithm formula is: where, E t is the comprehensive effect prediction index of the clinical trial, I is the number of patient groups in the simulated clinical trial, J is the number of evaluation indexes, x ij is the simulated result value of the ith patient group on the jth evaluation index, y ij is the weight coefficient of the ith patient group on the jth evaluation index, z ij is the sample size normalization coefficient of the ith patient group on the jth evaluation index.
[0018] Furthermore, in the clinical trial simulation module, the preliminary determination of the weight coefficient y ij is based on the standard operating procedures of tuberculosis clinical trials and the key focus directions of previous studies. When the key focus direction of the clinical trial changes or new evaluation indexes are considered more important, adjust y ij value, and at the same time, as the dynamic change of the sample size of the trial groups, update the z ij value in real time.
[0019] Furthermore, the personalized path generation module constructs a personalized diagnosis and treatment path generation model based on the AI algorithm. Its model formula is: where, D represents the comprehensive evaluation index of the disease development trend, and its value range is between 0 and 1. The closer it is to 1, the more obvious the trend of the disease developing in an adverse direction, and the closer it is to 0, the more the disease tends to be stable or improved. I is the total number of feature categories considered, α i is the weight coefficient of the ith type of feature, P i represents the parameter related to the natural history of tuberculosis in the ith type of feature, f i (P i ) is a function transformation based on the natural history parameter to convert it into a quantitative value related to the disease development trend, H i represents the parameter related to the individual characteristics of the patient in the ith type of feature, g i (H i ) is a function transformation for the individual characteristic parameter to convert it into a quantitative value matching the disease development trend.
[0020] Furthermore, the personalized path generation module combines the disease progression prediction results and generates a personalized tuberculosis diagnosis and treatment path for the patient based on the efficacy data and applicable population data of different treatment methods. The formula for generating the diagnosis and treatment path is: P n = Σ k k = 1 K uk k × vk k (PD n n), where P n n is the personalized diagnosis and treatment path score of the nth patient, used to evaluate the rationality and effectiveness of the path. K is the number of factors affecting the generation of the diagnosis and treatment path. uk k is the weight coefficient of the kth factor, and vk k (PD n n) is a function based on the kth individual data factor of the nth patient.
[0021] Furthermore, in the personalized path generation module, the initial determination of the weight coefficient uk k is based on the historical diagnosis and treatment data of a large number of tuberculosis patients. When new tuberculosis treatment technologies or drugs or new patient group characteristics are included in the system for consideration, the importance of each factor is re-evaluated. By analyzing the new diagnosis and treatment data, the value of uk k is recalculated.
[0022] Compared with the prior art, the tuberculosis data management and control system based on the AI disease simulation model has the following beneficial effects:
[0023] First, the present invention deeply integrates multi-omics data such as genomics, proteomics, and metabolomics through AI algorithms, accurately identifies target molecules closely related to tuberculosis specificity, provides a core basis for the research and development of diagnostic reagents, and at the same time performs personalized optimization according to the individual data of patients to ensure that the formula of the diagnostic reagent more accurately matches the actual needs of patients. In addition, through the clinical trial simulation module, it is possible to predict the detection differences of the reagent in different populations, insight into the factors that may affect the test results in advance, and thus generate a more feasible clinical trial plan.
[0024] Second, the present invention deeply integrates multi-omics data and uses AI technology for intelligent control. It not only solves the deficiencies of traditional technologies but also realizes the data management and control of the entire tuberculosis diagnosis and treatment process. It not only improves the R & D efficiency and personalization degree of diagnostic reagents but also optimizes the clinical trial plan, generates a more effective personalized diagnosis and treatment path. In addition, the data management center, as the core hub of the system, ensures the security and integrity of data.
[0025] Other advantages, objects, and features of the present invention will be set forth in part in the following description, and in part will be obvious to those skilled in the art upon examination of the following, or may be learned from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is a flow operation diagram of a tuberculosis data management and control system based on an AI disease simulation model;
[0028] Figure 2 It is a flowchart of a tuberculosis data management and control system based on an AI disease simulation model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific embodiments, structures, features, and their effects of the present invention as follows.
[0030] Embodiment 1
[0031] For the research and development of new diagnostic reagents and the optimization of clinical trials, the data collection module collects data from major tuberculosis research institutions, hospitals, and public health databases, obtaining clinical data of patients from different regions, different age groups, different genders, and covering various tuberculosis types (including latent infection, newly diagnosed pulmonary tuberculosis, drug-resistant pulmonary tuberculosis, etc.), including detailed symptom records, disease progression processes, treatment histories, and treatment results, etc. At the same time, a large number of mycobacterium tuberculosis samples are collected for whole-genome sequencing, proteomics analysis, and metabolomics detection to obtain their gene sequence characteristics, protein expression profiles, and metabolite changes. In addition, existing tuberculosis diagnostic reagent information, including component, performance index, scope of application, and clinical feedback data, is integrated. After cleaning and classifying and sorting these data, they are stored in the corresponding database areas, and data indexes are established for quick retrieval and call.
[0032] The multi-omics data integration and analysis module first standardizes the collected genomics, proteomics, and metabolomics data to eliminate the differences caused by different experimental batches and different detection platforms. For the gene expression data in genomics data, the reads per kilobase per million mapped reads (RPKM) or transcripts per million (TPM) method is used for normalization to make the gene expression levels comparable among different samples. For proteomics data, the protein quantification data is normalized and corrected according to the internal standard protein or total protein content. For metabolomics data, the internal standard method and data standardization algorithm are used to convert the original detection signals of metabolites into relative contents or concentration values to ensure the consistency of detection results for different batches.
[0033] The Al algorithm is used to integrate and analyze the standardized multi-omics data, and the formula is used to calculate the correlation degree R between gene and protein features ij , where M ik is the value of the i-th gene feature in the k-th sample (such as the expression level of gene i in sample k), is the average value of the i-th gene feature in all samples, N jk is the value of the j-th protein feature in the k-th sample (such as the activity level of protein j in sample k), is the average value of the j-th protein feature in all samples, N is the total number of samples, λ is the attenuation coefficient, and d ij is the distance metric between the i-th gene feature and the j-th protein feature in the functional space (such as determined by the shortest path length or functional similarity score in the gene-protein interaction network, etc.). In this way, potential target molecule combinations that play a key role in the pathogenesis of tuberculosis and are different from existing diagnostic targets are identified. For example, it is found that there is a close association between a new mycobacterium tuberculosis secreted protein and a key enzyme in a specific metabolic pathway in host cells, and it shows a unique change pattern during the infection process of drug-resistant mycobacterium tuberculosis. These potential target information is transmitted to the diagnostic reagent R & D and optimization module and the data management center, and the data transmission log is detailedly recorded, including information such as transmission time, data volume, and key features of target molecules.
[0034] Based on the provided target molecule information, the diagnostic reagent R & D and optimization module constructs an AI model for the R & D and optimization of tuberculosis diagnostic reagents, and uses deep learning algorithms to perform virtual screening on the substances in the chemical substance library. Its algorithm formula is: where, S m is the comprehensive score of the m-th substance as a tuberculosis diagnostic marker, L is the number of features for evaluating the interaction between the substance and the target molecule, f l (C l , Tm ) is a function of the l-th feature (for example, for the binding affinity feature, f l (C i , T m ) can be a function transformation of the binding energy value obtained from molecular docking simulations to convert the binding energy into a score contribution value related to the effectiveness of the diagnostic marker), C l is the calculated value of the l-th feature in the interaction between the substance and the target molecule, T m is the relevant chemical structure or property parameter of the m-th substance, w l is the weight coefficient of the l-th feature (obtained by learning a large amount of data on the interaction between chemical substances and target molecules in the training set, reflecting the relative importance of different features in judging the effectiveness of diagnostic markers). Compounds with higher scores are selected as candidate markers, and a preliminary diagnostic reagent formulation framework is determined, including the types, concentration ranges of candidate markers, and necessary excipient components (such as buffers, stabilizers, preservatives, etc.). At the same time, considering the synergistic effects and stability factors between markers, the preliminary formulation framework data is stored in the diagnostic reagent data sub-library.
[0035] Optimize the preliminary formulation according to the patient's individual data. Obtain the patient's individual data from the data acquisition module, including physical condition data (such as liver and kidney function indicators, immune cell counts, nutritional status indicators, etc.), disease severity data (such as mycobacterium tuberculosis load, scope and severity of lung lesions, etc.), drug resistance test result data (drug resistance genotypes and phenotypes to various anti-tuberculosis drugs). Develop a formulation optimization algorithm based on the patient's individual data, and its algorithm formula is: C p = α × C g + β × C r + γ × C s , where C p is the optimized diagnostic reagent formulation component ratio vector, C g is the basic formulation component ratio vector based on the diagnostic marker screening results, C r is the formulation component ratio vector adjusted according to the patient's drug resistance data, C s is the formulation component ratio vector adjusted based on the patient's physical condition, α, β, γ are the weight coefficients corresponding to C g , C r , C s , determined according to different patient population data, balancing the influence of various factors on the formulation, and defining the policy function π(α, β, γ|s), where s represents the system state, including patient data characteristics and performance indicators of the current formulation. The policy function represents the probability distribution of selecting the weight coefficients α, β, γ under the given state s. The goal is to have good treatment effects and few adverse reactions, then a high reward is given, and the cumulative reward calculation formula is: Among them, T is the treatment cycle or observation cycle, and r t is the immediate reward at time t. According to the policy gradient theorem, the update formula for the weight coefficient is: where θ is the parameter of the policy function, including α, β, γ, N is the number of patient samples, α n,t , β n,t , γ n,t is the weight coefficient of the nth patient at time t, s n,t is the corresponding system state, R n is the cumulative reward of the nth patient. By updating the weight coefficient along the policy gradient direction, the policy can better adapt to the patient's response to achieve the purpose of optimizing the formula.
[0036] The clinical trial simulation module establishes a tuberculosis disease simulation model, constructs a mathematical model framework based on the epidemiological data, pathophysiological knowledge, and clinical research results of tuberculosis, combines with AI algorithms, and uses the formula to estimate the comprehensive effect E t of the clinical trial. Among them, I is the number of patient groups in the simulated clinical trial (grouped according to factors such as age, gender, disease type, drug resistance status, etc.), J is the number of evaluation indicators (such as diagnostic accuracy, specificity, sensitivity, treatment effectiveness, adverse reaction incidence rate, recurrence rate, etc.), x ij is the simulated result value of the ith patient group on the jth evaluation indicator (for example, in the diagnostic accuracy indicator, x ij can be the correct diagnosis proportion of the patients in this group using the new diagnostic reagent predicted by the model. In the treatment effectiveness indicator, x ij can be the mycobacterium tuberculosis negative conversion rate or the proportion of lung lesion shrinkage of the patients in this group after a specific treatment plan simulated by the model, etc.), y ij is the weight coefficient of the ith patient group on the jth evaluation indicator (determined according to the key focus direction of the clinical trial and the importance of different indicators. For example, if this trial focuses on the performance of the new diagnostic reagent in drug-resistant patients, the weight coefficient y ij of the diagnostic accuracy indicator for the drug-resistant patient group is relatively high. If the safety of the treatment plan is concerned, the weight coefficients of the adverse reaction incidence rate indicators for each group are relatively high), z ij is the sample size normalization coefficient of the ith patient group on the jth evaluation indicator (determined according to the proportion of the group sample size to the total sample size, used to eliminate the influence of different group sample size differences on the results. For example, if the total sample size is 1000 and the sample size of a certain group is 200, then the z ij value of this group in the calculation of each indicator is 200 / 1000 = 0.2).
[0037] Based on the simulation results, potential problems are identified, such as insufficient sensitivity of the diagnostic reagent in certain specific patient groups (such as elderly drug-resistant patients), or a relatively high incidence of adverse reactions of a certain drug combination in specific genotype patients in the combined drug treatment plan. Then, optimization suggestions for the clinical trial plan are generated, such as increasing the sample size of specific genotype patients to more accurately evaluate the performance of the diagnostic reagent, adjusting the drug combination or dosage to reduce the incidence of adverse reactions, optimizing the detection process or time points to improve diagnostic accuracy, etc. The simulation data and plan suggestions are stored and synchronized with the data management center. The data synchronization frequency is set to once a day to ensure the timeliness and consistency of the data. The data management center classifies and stores the simulation data and plan suggestions, and conducts analysis and mining to provide a reference basis for subsequent clinical trial design. At the same time, a data monitoring mechanism is established to regularly check the accuracy of the simulation data and the rationality of the plan suggestions. If any abnormalities are found, the clinical trial simulation module is notified in a timely manner for adjustment and optimization.
[0038] Example 2
[0039] For personalized diagnosis and treatment and dynamic adjustment, when a patient is first diagnosed with tuberculosis, the data collection module collects their comprehensive information, including personal basic information (age, gender, occupation, etc.), past medical history (whether there are other chronic diseases such as cardiovascular diseases, lung diseases, etc.), family medical history (whether there are tuberculosis patients in the family), living habits (exercise, diet, sleep conditions, etc.), and at the same time conducts a detailed clinical examination to obtain physical condition data (such as blood routine, liver and kidney function indicators, cardiopulmonary function test results, etc.), disease severity data (determining the scope of lung lesions through chest imaging examinations, mycobacterium tuberculosis load detection, etc.), drug resistance test results (determining the sensitivity or drug resistance to various anti-tuberculosis drugs). In addition, multi-omics data of this patient (if there have been relevant tests before) and the diagnosis and treatment data and outcome information of similar patients are obtained from the data management center. After sorting and preprocessing these data, they are stored and a data index is established.
[0040] The personalized path generation module uses machine learning algorithms to construct a disease progression prediction model, takes the patient's individual data as input to predict the disease development trend, and uses the formula to calculate the comprehensive evaluation index D of the disease development trend. Among them, I is the total number of feature categories, α i is the weight coefficient, P i is a parameter related to the natural history of tuberculosis, f i (P i ) is its functional transformation, H i is a parameter related to the patient's individual characteristics, g i (H i) is its functional transformation. For example, for a young male patient with a mild smoking history and no other chronic diseases, and the tuberculosis bacteria he is infected with are initially treated sensitive type. Through the model prediction, his disease may improve relatively quickly under standardized treatment, but if the smoking habit is not improved, there may be a certain risk of recurrence. According to the prediction results, combined with the efficacy data and applicable population data of different treatment methods (such as different drug combinations, physical therapy methods, nutritional support programs, etc.), a personalized diagnosis and treatment path is generated. For example, the standard treatment plan of first-line anti-tuberculosis drugs is adopted, and at the same time, the patient is advised to quit smoking and given nutritional supplement suggestions, and sputum smear, chest imaging examinations, and liver and kidney function monitoring are carried out regularly. The examination time interval is determined according to the disease progression prediction and the patient's individual situation. For example, it is checked once every two weeks at the initial stage of treatment, and once a month after the condition is stable, etc. The diagnosis and treatment path data is stored and interacted with the data management center.
[0041] During the treatment process, as new data is generated (such as the results of each examination, changes in the patient's living habits, etc.), the data management center promptly notifies the personalized path generation module. For example, after 3 months of treatment, there are still a small amount of tuberculosis bacteria in the sputum smear examination of the patient, and it is found that the patient has insufficient sleep and irregular diet due to high work pressure. The personalized path generation module re-analyzes the data and adjusts the diagnosis and treatment path. It may add an auxiliary drug to enhance the treatment effect, adjust the nutritional support plan, and increase the examination frequency to once every two weeks. At the same time, the adjusted path data is fed back to the data management center to update the patient's diagnosis and treatment records and data storage information.
[0042] During the adjustment process, the disease progression prediction model is used again to evaluate the impact of the adjusted diagnosis and treatment path on the disease development trend. Substitute the new data (such as the drug concentration monitoring data after adding the auxiliary drug, the change data of the physical condition indicators after adjusting the diet and sleep, etc.) into the formula Recalculate the comprehensive evaluation index D of the disease development trend, observe the change trend of the indicators, ensure that the adjusted diagnosis and treatment path develops in a direction beneficial to the patient's recovery, and at the same time feed back the adjusted path data to the data management center to update the patient's diagnosis and treatment records and data storage information. The data management center backs up and archives the updated data for subsequent analysis and research.
[0043] After a patient completes a course of treatment, the data management center comprehensively evaluates the treatment effect based on various data, comparing the chest imaging data before and after treatment (such as complete absorption, partial absorption, or no significant change in lung lesions), sputum smear and tubercle bacillus culture results (whether turning negative and the time of turning negative), and the patient's physical condition indicators (whether the weight increases, whether the symptoms are completely relieved, whether the liver and kidney functions return to normal, etc.). Combining the records of the adjustment of the diagnosis and treatment path during the treatment process (such as the types, doses, and times of drug adjustments, the increase or decrease of examination items, and the changes in time intervals), it evaluates the effectiveness of this personalized diagnosis and treatment path. For example, if the tubercle bacillus is completely cleared after treatment, the lung lesions are significantly absorbed (such as the lesion area is reduced by more than 80%), and the physical condition is good (the weight increases by more than 5%, there are no obvious discomfort symptoms, and the liver and kidney function indicators are normal), it is evaluated as a successful treatment. The evaluation results are stored in the data management center, and these data are used to optimize the disease progression prediction model and the diagnosis and treatment path generation model. For example, if it is found that a certain type of patient (such as young men, with a mild smoking history, and newly diagnosed sensitive type tubercle bacillus infection) has a greater impact on the treatment effect under specific changes in living habits (such as lack of sleep and irregular diet), the parameters in the relevant model will be adjusted by increasing the weight coefficient α of the living habit feature category in the model i , and optimizing the function transformations f i (P i ) and g i (H i ) (such as re-determining the quantitative impact relationship between lack of sleep and irregular diet on disease progression based on the new data), so as to provide more accurate diagnosis and treatment services for subsequent similar patients. At the same time, the successful case experience is sorted into a knowledge base and stored classified according to the different characteristics of the patients (age, gender, disease type, drug resistance situation, living habits, etc.), for doctors to refer to during clinical decision-making. For example, when facing a new young male tuberculosis patient with similar living habits, doctors can quickly retrieve the knowledge base and obtain relevant diagnosis and treatment experience and suggestions, such as paying more attention to lifestyle intervention and the use of early adjuvant drugs when formulating the diagnosis and treatment path, thereby continuously improving the level of personalized tuberculosis diagnosis and treatment and providing high-quality medical services for more patients.
[0044] The above is only a preferred embodiment of the present invention, and it does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to form equivalent embodiments with equivalent changes, but as long as it does not depart from the technical content of the present invention, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A tuberculosis data management and control system based on AI disease simulation model, characterized in that: The system includes the following components: data acquisition module, multi-omics data integration and analysis module, diagnostic reagent development and optimization module, clinical trial simulation module, personalized pathway generation module and data management center; The data acquisition module constructs a multi-source data collection interface, establishes connections with hospital information systems, scientific research institution databases and chemical substance libraries, collects data including tuberculosis gene sequences, protein structures, patient clinical information, known reagent performance, chemical substance library data and multi-omics data, and stores them in the corresponding database area after deduplication, error correction, classification and sorting, and establishes data indexes for retrieval and call; The multi-omics data integration and analysis module integrates genomics, proteomics and metabolomics data, and constructs an analysis model based on deep learning. It uses multi-omics data as input, mines features and patterns in the data, identifies target molecules that are specifically related to tuberculosis and have diagnostic value, transmits the identified target molecule data to the diagnostic reagent development and optimization module and the data management center, and establishes a data transmission log; The diagnostic reagent development and optimization module builds an AI model for tuberculosis diagnostic reagent development and optimization based on the target molecule information provided, uses a deep learning algorithm to virtually screen substances in the chemical substance library, and calculates the affinity and binding energy parameters of the interaction by constructing a simulated environment for the interaction between the substance and the target molecule, predicts potential diagnostic markers, generates a candidate list of markers based on the virtual screening results, and determines a preliminary diagnostic reagent formula framework in combination with the component information in the chemical substance library, stores the formula framework data and the candidate list of markers in the diagnostic reagent data sub-library, and develops a formula optimization algorithm based on individual patient data, takes the patient's physical condition, disease severity, and drug resistance test result data as input, and combines the preliminary formula framework to perform personalized optimization by adjusting the proportion of the ingredients in the formula, stores the final optimized diagnostic reagent formula data and the entire development and optimization process data in the diagnostic reagent data sub-library, and exchanges data with the data management center; The clinical trial simulation module constructs a mathematical model and algorithm framework based on the knowledge of the pathophysiological mechanism, transmission law and population infection characteristics of tuberculosis, and inputs the collected data of different populations into the disease simulation model in combination with the AI algorithm. At the same time, a clinical trial simulation algorithm is designed to predict the detection differences of reagents in different populations, calculate the performance indicators of reagents by simulating different test scenarios, sample sizes and detection process factors, discover factors affecting clinical trial results in advance, generate clinical trial program optimization suggestions based on simulation results, and store simulation data and program suggestions in the clinical trial data sub-library; The personalized pathway generation module constructs a personalized diagnosis and treatment pathway generation model based on an AI algorithm, takes individual patient data as input, predicts the possible development trend of the disease according to the natural history of tuberculosis and individual patient characteristics, combines the disease progression prediction results, and generates a personalized tuberculosis diagnosis and treatment pathway for the patient based on the efficacy data of different treatment methods and applicable population data, stores the personalized diagnosis and treatment pathway data in the diagnosis and treatment pathway data sub-library, establishes a data interaction channel with the data management center, receives notifications from the data management center when individual patient data changes, recalculates and adjusts the diagnosis and treatment pathway, and feeds back the adjusted pathway data to the data management center; The data management center builds a centralized data storage architecture, including relational and non-relational databases to store data in each module, and establishes an efficient data indexing and query system, designs multi-mode query interfaces and implements authority management to ensure data security. At the same time, it formulates data backup and recovery strategies, regularly backs up data in full and incremental amounts to a remote location or the cloud, establishes a data analysis task scheduling mechanism, regularly integrates and analyzes data from each module, monitors data change trends and abnormal values, and sends notifications to the corresponding modules and records logs when abnormalities are found.
2. According to claim 1, a tuberculosis data management and control system based on an AI disease simulation model is characterized in that: The multi-omics data integration and analysis module mines the features and patterns in the data to identify target molecules that are specifically related to tuberculosis and have diagnostic value. The calculation formula is: Among them, R ij is the correlation value between the i-th genomic data feature and the j-th proteomic data feature, reflecting the strength of the synergistic correlation between the two in the mechanism of tuberculosis. ik is the value of the i-th genomics data feature in the k-th sample, is the average value of the i-th genomics data feature in all samples, N jk is the value of the jth proteomics data feature in the kth sample, is the average value of the jth proteomics data feature in all samples, N is the total number of samples, λ is the decay coefficient, and d ij It is the distance measure between the i-th genomics data feature and the j-th proteomics data feature in the feature space.
3. According to claim 1, a tuberculosis data management and control system based on an AI disease simulation model is characterized in that: The diagnostic reagent development and optimization module uses a deep learning algorithm to perform virtual screening of substances in the chemical substance library. The algorithm formula is: Among them, S m is the comprehensive score of the mth substance as a diagnostic marker for tuberculosis. The score indicates the possibility of becoming an effective diagnostic marker. L is the number of features for evaluating the interaction between the substance and the target molecule. f l (C l ,T m ) is a function of the lth feature, C l is the calculated value of the lth feature in the interaction, T m is the chemical structure or property parameter related to the mth substance, w l is the weight coefficient of the lth feature.
4. According to claim 1, a tuberculosis data management and control system based on an AI disease simulation model is characterized in that: The diagnostic reagent development and optimization module develops a formulation optimization algorithm based on individual patient data, and the algorithm formula is: C p =α×C g +β×C r +γ×C s , where C p is the optimized diagnostic reagent formula component ratio vector, C g is the basic formula component ratio vector based on the diagnostic marker screening results, C r is the formula component ratio vector adjusted according to the patient's drug resistance data, C s is the ratio vector of the formula ingredients adjusted based on the patient's physical condition, α, β, γ are the corresponding C g , C r , C s The weight coefficient is determined based on the data of different patient groups to balance the impact of various factors on the formula.
5. According to claim 1, a tuberculosis data management and control system based on an AI disease simulation model is characterized in that: In the diagnostic reagent development and optimization module, combined with the preliminary formula framework, personalized optimization is performed by adjusting the proportion of ingredients in the formula. Specifically, a strategy function π(α, β, γ|s) is defined, where s represents the system state, including patient data characteristics and performance indicators of the current formula. The strategy function represents the probability distribution of selecting weight coefficients α, β, and γ under a given state s. The goal is to have a good treatment effect and few adverse reactions, and a high reward is given. The calculation formula for the cumulative reward is: Where T is the treatment period or observation period, r t is the immediate reward at time t. According to the policy gradient theorem, the update formula of the weight coefficient is: Among them, θ is the parameter of the strategy function, including α, β, γ, N is the number of patient samples, α n,t , β n,t , γ n,t is the weight coefficient of the nth patient at time t, s n,t is the corresponding system state, R n is the cumulative reward of the nth patient. By updating the weight coefficient along the policy gradient direction, the strategy can better adapt to the patient's response to achieve the purpose of optimizing the formula.
6. A tuberculosis data management and control system based on an AI disease simulation model according to claim 1, characterized in that: The clinical trial simulation module designs a clinical trial simulation algorithm to predict the detection differences of reagents in different populations. The algorithm formula is: Among them, E t is the comprehensive effect estimation index of the clinical trial, I is the number of patient groups in the simulated clinical trial, J is the number of evaluation indicators, and x ij is the simulation result value of the i-th patient group on the j-th evaluation index, y ij is the weight coefficient of the i-th patient group on the j-th evaluation index, z ij is the sample size normalization coefficient of the i-th patient group on the j-th evaluation indicator.
7. A tuberculosis data management and control system based on an AI disease simulation model according to claim 6, characterized in that: In the clinical trial simulation module, the weight coefficient y ij The initial determination of y is based on the standard operating procedures of tuberculosis clinical trials and the focus of previous studies. When the focus of clinical trials changes or new evaluation indicators are considered more important, adjust y ij The value of z is updated in real time as the sample size of the experimental group changes dynamically. ij The value of .
8. The tuberculosis data management and control system based on the AI disease simulation model according to claim 1 is characterized in that: The personalized pathway generation module constructs a personalized diagnosis and treatment pathway generation model based on the AI algorithm, and its model formula is: Among them, D represents the comprehensive evaluation index of the disease development trend, and its value range is between 0 and 1. The closer it is to 1, the more obvious the trend of the disease developing in an unfavorable direction is, and the closer it is to 0, the more likely the disease is to be stable or improved. I is the total number of feature categories considered, and α i is the weight coefficient of the i-th feature, P i represents the parameter related to the natural history of tuberculosis in the i-th category, f i (P i ) is a function based on the natural history parameters to convert them into quantitative values related to the development trend of the disease. i represents the parameter related to the individual characteristics of the patient in the i-th category, g i (H i ) is a functional transformation of individual characteristic parameters, converting them into quantitative values that match the development trend of the disease.
9. The tuberculosis data management and control system based on the AI disease simulation model according to claim 1, characterized in that: The personalized pathway generation module combines the disease progression prediction results, the efficacy data of different treatment methods and the applicable population data, and generates a personalized tuberculosis diagnosis and treatment pathway for the patient. The diagnosis and treatment pathway generation formula is: n =∑ k =1 K u k ×v k (PD n ), where P n is the personalized diagnosis and treatment pathway score of the nth patient, which is used to evaluate the rationality and effectiveness of the pathway. K is the number of factors that affect the generation of the diagnosis and treatment pathway. u k is the weight coefficient of the kth factor, v k (PD n ) is a function based on the kth individual data factor of the nth patient.
10. A tuberculosis data management and control system based on an AI disease simulation model according to claim 9, characterized in that: In the personalized path generation module, the weight coefficient u k The initial determination of u is based on the historical diagnosis and treatment data of a large number of tuberculosis patients. When new tuberculosis treatment technologies or drugs appear or new patient population characteristics are taken into account, the importance of each factor is re-evaluated and u is recalculated by analyzing the new diagnosis and treatment data. k The value of .