Data processing method and device based on clinical test data
By calculating the posterior probability of clinical trial data using a Bayesian theorem-based method, a probability ranking list of indications is generated, which solves the quantification and interpretability problems of existing diagnostic aids and improves the scientificity and credibility of diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MEDICAL MO (BEIJING) MEDICAL INFORMATION TECH CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to provide prospective and quantitative diagnostic assistance when utilizing clinical trial data, and the decision-making process of machine learning models lacks clear clinical reasoning logic, reducing doctors' trust and willingness to adopt them.
By acquiring data from target clinical trials, identifying sets of symptoms and indications, calculating prior probabilities and historical likelihoods using historical clinical databases, applying Bayes' theorem to calculate posterior probabilities, and generating a probability ranking list of indications, quantitative diagnostic support is provided.
It enables the extraction of reliable and interpretable probabilistic correlation knowledge from massive amounts of data, improving the objectivity, consistency, and scientific rigor of diagnosis, and providing a quantitative assessment tool that conforms to clinical logic.
Smart Images

Figure CN121880415A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a data processing method and apparatus based on clinical trial data. Background Technology
[0002] In clinical diagnostic practice, physicians' decisions heavily rely on the comprehensive interpretation of diverse and heterogeneous data, including patient clinical manifestations, laboratory tests, and imaging results. With the widespread adoption of medical informatics, massive amounts of clinical trial data and real electronic medical record data have accumulated, revealing complex statistical correlations between diseases and clinical manifestations. However, current utilization of this data largely remains at the level of descriptive statistics, retrospective analysis, or simple matching based on fixed rules, failing to fully leverage its inherent deep correlations for prospective and quantitative diagnostic assistance.
[0003] Specifically, existing technologies have the following limitations: First, diagnostic support largely relies on rigid rules or knowledge bases summarized from expert experience, making it difficult to adapt to the diversity and uncertainty of clinical cases, and insufficient in recognizing emerging or rare disease patterns. Second, most systems can only provide qualitative or binary (yes / no) judgments, unable to offer evidence-based, comparable quantitative probabilities, making it difficult for doctors to intuitively assess the relative likelihood between different diagnostic hypotheses. Third, although some studies have attempted to use machine learning models for disease prediction, these end-to-end models often resemble "black boxes," lacking clear and clinically logical interpretability in their decision-making process, thus reducing clinicians' trust and willingness to adopt them.
[0004] Therefore, how to systematically extract reliable and interpretable probabilistic correlation knowledge from massive historical clinical trial data, and construct a clinical auxiliary judgment framework that can quantify diagnostic uncertainty and has a transparent reasoning process, has become a key technical problem for improving diagnostic accuracy and the scientific nature of decision-making. This invention aims to address how to combine data-driven approaches with clinical logical reasoning to achieve an objective quantitative assessment of diagnostic probability. Summary of the Invention
[0005] In view of this, embodiments of this application provide a data processing method and apparatus based on clinical trial data to perform quantifiable evaluation of clinical trial data in order to assist in the determination of indications.
[0006] In a first aspect, embodiments of this application provide a data processing method based on clinical trial data, the data processing method comprising: After obtaining the target clinical trial data, determine the symptoms included in the target clinical trial data, and determine the set of indications that cause the symptoms; Based on historical clinical databases, determine the prior probability of each indication and the historical likelihood of each symptom under the condition of each indication. For each indication, the combined probability of all symptoms occurring simultaneously under that indication is calculated based on the historical likelihood of each symptom corresponding to that indication. Based on the prior probability of the indication and the combined probability corresponding to the indication, the posterior probability of the indication when all symptoms are present is determined according to Bayes' theorem. A probability ranking list of indications is generated based on the posterior probabilities from high to low.
[0007] Secondly, embodiments of this application provide a data processing apparatus based on clinical trial data, the data processing apparatus comprising: The acquisition unit is used to acquire target clinical trial data, determine the symptoms included in the target clinical trial data, and determine the set of indications that cause the symptoms; The first determining unit is used to determine, based on a historical clinical database, the prior probability of each indication occurring, and the historical likelihood of each symptom under the condition of each indication. The calculation unit is used to calculate the overall probability of the simultaneous occurrence of all symptoms under each indication based on the historical likelihood of each symptom corresponding to that indication. The second determining unit is used to determine the posterior probability of the indication when all symptoms are present, based on the prior probability of the indication and the comprehensive probability corresponding to the indication, according to Bayes' theorem. The generation unit is used to generate a probability sorting list of indications based on the posterior probabilities from high to low.
[0008] The technical solution provided in this application includes, but is not limited to, the following beneficial effects: In this application, the prior probability and historical likelihood of the symptoms corresponding to the target clinical trial data can be determined through historical clinical databases. The prior probability reflects the basic prevalence of the indication in the population, and the historical likelihood mines the correlation between symptoms and indications from historical cases. The two are dynamically updated through Bayesian formulas to form a reasoning chain that conforms to clinical logic. Secondly, the conditional independence assumption transforms the complex problem of multi-feature joint probability estimation into the product calculation of single-feature statistics, which not only solves the problem of computational feasibility under high-dimensional data, but also maintains the simplicity of the reasoning process. Furthermore, the probabilistic output intuitively presents the relative probability of different diagnostic hypotheses, providing quantitative evidence for assisting in the determination of indications. This application transforms data-driven statistical laws into clinically understandable auxiliary support tools, effectively improving the objectivity, consistency and scientific nature of diagnostic decisions.
[0009] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a data processing method based on clinical trial data provided in this application embodiment; Figure 2 This is a schematic diagram of a data processing device based on clinical trial data provided in an embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0013] It should be noted in advance that the symptoms in this application are clinical manifestations, such as high fever (body temperature >39℃), cough, etc., and the indications can be understood as specific diseases, such as influenza, common cold, bacterial pneumonia. For ease of understanding, examples of high fever (denoted as e1), cough (denoted as e2), influenza (also known as flu, denoted as c1), common cold (also known as cold, denoted as c2), and bacterial pneumonia (also known as pneumonia, denoted as c3) will be used in the following illustrations.
[0014] Figure 1 A flowchart illustrating a data processing method based on clinical trial data provided in this application embodiment is shown below. Figure 1 The data processing method includes the following steps: Step 101: After obtaining the target clinical trial data, determine the symptoms included in the target clinical trial data, and determine the set of indications that cause the symptoms.
[0015] Step 102: Based on the historical clinical database, determine the prior probability of each indication and the historical likelihood of each symptom under the condition of each indication.
[0016] Step 103: For each indication, calculate the overall probability of all symptoms occurring simultaneously under that indication based on the historical likelihood of each symptom corresponding to that indication.
[0017] Step 104: Based on the prior probability of the indication and the combined probability corresponding to the indication, determine the posterior probability of the indication when all symptoms are present, according to Bayes' theorem. Step 105: Generate a probability ranking list of indications based on the posterior probabilities from high to low.
[0018] In this application, the prior probability and historical likelihood of the symptoms corresponding to the target clinical trial data can be determined through historical clinical databases. The prior probability reflects the basic prevalence of the indication in the population, and the historical likelihood mines the correlation between symptoms and indications from historical cases. The two are dynamically updated through Bayesian formulas to form a reasoning chain that conforms to clinical logic. Secondly, the conditional independence assumption transforms the complex problem of multi-feature joint probability estimation into the product calculation of single-feature statistics, which not only solves the problem of computational feasibility under high-dimensional data, but also maintains the simplicity of the reasoning process. Furthermore, the probabilistic output intuitively presents the relative probability of different diagnostic hypotheses, providing quantitative evidence for assisting in the determination of indications. This application transforms data-driven statistical laws into clinically understandable auxiliary support tools, effectively improving the objectivity, consistency and scientific nature of diagnostic decisions.
[0019] In one feasible implementation, when performing the steps of determining the prior probability of each indication and the historical likelihood of each symptom under the condition of each indication based on a historical clinical database, the historical clinical database is first queried to determine a first number of occurrences of each indication and a second number of occurrences of each symptom in each indication; then, for each indication, a first ratio is calculated to the sum of the first numbers of the indication and the first numbers of all indications, and the first ratio is used as the prior probability of the indication; finally, for each indication, a first proportion of occurrence of each symptom in the indication is determined based on the first number of occurrences of the indication and the second number of occurrences of each symptom in each indication, and the first proportion is used as the historical likelihood of the corresponding symptom under the condition of the indication.
[0020] For example, the symptoms include {e1,e2}, and the set of indications C={c1,c2,c3}. At this point, the transformation from unstructured data to standardized features and classification labels that the model can process has been completed. Suppose that there are 1000 cases in the historical clinical database that simultaneously contain {e1,e2}, including 300 cases of influenza, 550 cases of the common cold, and 150 cases of bacterial pneumonia. The prior probability of influenza is P(c1)=300 / 1000=0.30; the prior probability of the common cold is P(c2)=550 / 1000=0.55; and the prior probability of pneumonia is P(c3)=150 / 1000=0.15.
[0021] For influenza (c1), out of 300 cases, if 270 cases have high fever, then the historical likelihood of high fever in influenza is P(e1|c1) = 270 / 300 = 0.90; if 210 cases have cough, then the historical likelihood of cough in influenza is P(e2|c1) = 210 / 300 = 0.70. For the common cold (c2): If 330 out of 550 cases had high fever, then the historical likelihood of high fever in the common cold is P(e1|c2) = 330 / 550 = 0.60; if 440 cases had cough, then the historical likelihood of cough in the common cold is P(e2|c2) = 440 / 550 = 0.80. For pneumonia (c3): Of the 150 cases, if 120 cases had high fever, the historical likelihood of high fever in pneumonia is P(e1|c3) = 120 / 150 = 0.80; if 135 cases had cough, the historical likelihood of cough in pneumonia is P(e2|c3) = 135 / 150 = 0.90.
[0022] In one feasible implementation, when performing the step of calculating the comprehensive probability of the simultaneous occurrence of all symptoms under each indication based on the historical likelihood of each symptom corresponding to that indication, for each indication, a first product of the historical likelihoods of each symptom under that indication is calculated, and the first product is used as the comprehensive probability of the simultaneous occurrence of all symptoms under that indication.
[0023] For example, the combined probability of high fever and cough occurring simultaneously in influenza is P(E|c1) = P(e1|c1) × P(e2|c1) = 0.90 × 0.70 = 0.630; the combined probability of high fever and cough occurring simultaneously in the common cold is P(E|c2) = P(e1|c2) × P(e2|c2) = 0.60 × 0.80 = 0.480; and the combined probability of high fever and cough occurring simultaneously in pneumonia is P(E|c3) = P(e1|c3) × P(e2|c3) = 0.80 × 0.90 = 0.720.
[0024] In one feasible implementation, when determining the posterior probability of an indication when all symptoms are present, based on the prior probability and the comprehensive probability corresponding to the indication, according to Bayes' theorem, firstly, for each indication, the second product of the historical likelihood and the comprehensive probability corresponding to the indication is calculated; then, after obtaining the second product for each indication, a second proportion of the second product of the indication in the sum of the second products of all indications is calculated, and the second proportion is used as the posterior probability of the indication.
[0025] For example, applying Bayes' theorem updates the prior probability to the posterior probability, where the posterior probability P(c i |E) is the final answer we need. After seeing the "high fever and cough" symptom in the target clinical trial data, what is the probability of "high fever and cough" occurring simultaneously for each indication? First, we need to calculate the total probability of "high fever and cough" occurring simultaneously: P(E)=P(c1)P(E|c1)+P(c2)P(E|c2)+P(c3)P(E|c3) =(0.30×0.630)+(0.55×0.480)+(0.15×0.720) = 0.189+0.264+0.108=0.561.
[0026] Then, the posterior probability that "high fever and cough" both occur in influenza is calculated as P(c1|E)=[P(c1)×P(E|c1)] / P(E)=0.189 / 0.561≈0.337; The posterior probability that "high fever and cough" both occur during a cold is calculated as P(c2|E) = [P(c2) × P(E|c2)] / P(E) = 0.264 / 0.561 ≈ 0.471; The posterior probability that "high fever and cough" both occur in pneumonia is calculated as P(c3|E) = [P(c3) × P(E|c3)] / P(E) = 0.108 / 0.561 ≈ 0.192.
[0027] The above calculation process also includes normalization.
[0028] In one feasible implementation, after obtaining the probability ranking list, the probability ranking list is displayed on the user terminal.
[0029] For example, a probability-sorted list L p={(c2: common cold, 0.471), (c1: influenza, 0.337), (c3: bacterial pneumonia, 0.192)}, this list presents the identification results to users in a clear and quantitative way: the data shown in this target clinical trial are most likely (47.1%) to be the common cold, followed by influenza (33.7%), and pneumonia is relatively less likely (19.2%). This probability ranking list is displayed on the user's mobile terminal to help the user confirm the indication.
[0030] Figure 2 A schematic diagram of a data processing device based on clinical trial data provided in this application embodiment is shown below. Figure 2 As shown, the data processing device includes: The acquisition unit 21 is used to acquire target clinical trial data, determine the symptoms included in the target clinical trial data, and determine the set of indications that cause the symptoms; The first determining unit 22 is used to determine, based on a historical clinical database, the prior probability of each indication occurring, and the historical likelihood of each symptom under the condition of each indication. The calculation unit 23 is used to calculate the comprehensive probability of the simultaneous occurrence of all symptoms under each indication based on the historical likelihood of each symptom corresponding to that indication. The second determining unit 24 is used to determine the posterior probability of the indication when all symptoms are present, based on the prior probability of the indication and the comprehensive probability corresponding to the indication, according to Bayes' theorem. The generation unit 25 is used to generate a probability sorting list of indications according to the posterior probabilities from high to low.
[0031] In one feasible implementation, the first determining unit is configured to determine, based on a historical clinical database, the prior probability of each indication occurring, and the historical likelihood of each symptom under the condition of each indication, including: The historical clinical database was queried to determine the first number of occurrences for each indication and the second number of occurrences for each symptom in each indication; For each indication, a first ratio is calculated to the sum of the first quantities of all indications, and the first ratio is used as the prior probability of that indication. For each indication, a first proportion of each symptom is determined based on a first number of occurrences of the indication and a second number of occurrences of each symptom in each indication, so as to determine the first proportion as the historical likelihood of the corresponding symptom under the condition of the indication.
[0032] In one feasible implementation, the computing unit is used to calculate the overall probability of the simultaneous occurrence of all symptoms under each indication based on the historical likelihood of each symptom corresponding to that indication, including: For each indication, the first product of the historical likelihoods of each symptom under that indication is calculated, and the first product is used as the combined probability of all symptoms occurring simultaneously under that indication.
[0033] In one feasible implementation, the second determining unit is configured to determine the posterior probability of the indication when all symptoms are present, according to Bayes' theorem, based on the prior probability of the indication and the combined probability corresponding to the indication. This includes: For each indication, calculate the second product of the historical likelihood corresponding to that indication and the overall probability corresponding to that indication; After obtaining the second product for each indication, a second proportion of the second product of that indication in the sum of the second products of all indications is calculated, and the second proportion is used as the posterior probability of that indication.
[0034] In one feasible implementation, the data processing apparatus further includes: The display unit is used to display the probability sorting list on the user terminal after obtaining the probability sorting list.
[0035] about Figure 2 For explanations of the principles behind the relevant content, please refer to... Figure 1 Detailed explanations of the relevant content will not be repeated here.
[0036] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0037] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0038] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0039] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0040] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0041] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A data processing method based on clinical trial data, characterized in that, The data processing method includes: After obtaining the target clinical trial data, determine the symptoms included in the target clinical trial data, and determine the set of indications that cause the symptoms; Based on historical clinical databases, determine the prior probability of each indication and the historical likelihood of each symptom under the condition of each indication. For each indication, the combined probability of all symptoms occurring simultaneously under that indication is calculated based on the historical likelihood of each symptom corresponding to that indication. Based on the prior probability of the indication and the combined probability corresponding to the indication, the posterior probability of the indication when all symptoms are present is determined according to Bayes' theorem. A probability ranking list of indications is generated based on the posterior probabilities from high to low.
2. The data processing method as described in claim 1, characterized in that, The determination of the prior probability of each indication based on a historical clinical database, and the historical likelihood of each symptom under the condition of each indication, includes: The historical clinical database was queried to determine the first number of occurrences for each indication and the second number of occurrences for each symptom in each indication; For each indication, a first ratio is calculated to the sum of the first quantities of all indications, and the first ratio is used as the prior probability of that indication. For each indication, a first proportion of each symptom is determined based on a first number of occurrences of the indication and a second number of occurrences of each symptom in each indication, so as to determine the first proportion as the historical likelihood of the corresponding symptom under the condition of the indication.
3. The data processing method as described in claim 1, characterized in that, For each indication, based on the historical likelihood of each symptom corresponding to that indication, the comprehensive probability of the simultaneous occurrence of all symptoms under that indication is calculated, including: For each indication, the first product of the historical likelihoods of each symptom under that indication is calculated, and the first product is used as the combined probability of all symptoms occurring simultaneously under that indication.
4. The data processing method as described in claim 1, characterized in that, The determination of the posterior probability of the indication when all symptoms are present, based on the prior probability and the combined probability corresponding to the indication, according to Bayes' theorem, includes: For each indication, calculate the second product of the historical likelihood corresponding to that indication and the overall probability corresponding to that indication; After obtaining the second product for each indication, a second proportion of the second product of that indication in the sum of the second products of all indications is calculated, and the second proportion is used as the posterior probability of that indication.
5. The data processing method as described in claim 1, characterized in that, After obtaining the probability sorting list, the probability sorting list is displayed on the user terminal.
6. A data processing device based on clinical trial data, characterized in that, The data processing device includes: The acquisition unit is used to acquire target clinical trial data, determine the symptoms included in the target clinical trial data, and determine the set of indications that cause the symptoms; The first determining unit is used to determine, based on a historical clinical database, the prior probability of each indication occurring, and the historical likelihood of each symptom under the condition of each indication. The calculation unit is used to calculate the overall probability of the simultaneous occurrence of all symptoms under each indication based on the historical likelihood of each symptom corresponding to that indication. The second determining unit is used to determine the posterior probability of the indication when all symptoms are present, based on the prior probability of the indication and the comprehensive probability corresponding to the indication, according to Bayes' theorem. The generation unit is used to generate a probability sorting list of indications based on the posterior probabilities from high to low.
7. The data processing apparatus as described in claim 6, characterized in that, The first determining unit is configured to determine, based on a historical clinical database, the prior probability of each indication occurring, and the historical likelihood of each symptom under the condition of each indication, including: The historical clinical database was queried to determine the first number of occurrences for each indication and the second number of occurrences for each symptom in each indication; For each indication, a first ratio is calculated to the sum of the first quantities of all indications, and the first ratio is used as the prior probability of that indication. For each indication, a first proportion of each symptom is determined based on a first number of occurrences of the indication and a second number of occurrences of each symptom in each indication, so as to determine the first proportion as the historical likelihood of the corresponding symptom under the condition of the indication.
8. The data processing apparatus as described in claim 6, characterized in that, The calculation unit is used to calculate the overall probability of the simultaneous occurrence of all symptoms under each indication based on the historical likelihood of each symptom corresponding to that indication, including: For each indication, the first product of the historical likelihoods of each symptom under that indication is calculated, and the first product is used as the combined probability of all symptoms occurring simultaneously under that indication.
9. The data processing apparatus as described in claim 6, characterized in that, The second determining unit is used to determine the posterior probability of the indication when all symptoms are present, according to Bayes' theorem, based on the prior probability of the indication and the comprehensive probability corresponding to the indication. This includes: For each indication, calculate the second product of the historical likelihood corresponding to that indication and the overall probability corresponding to that indication; After obtaining the second product for each indication, a second proportion of the second product of that indication in the sum of the second products of all indications is calculated, and the second proportion is used as the posterior probability of that indication.
10. The data processing apparatus as claimed in claim 6, characterized in that, The data processing device further includes: The display unit is used to display the probability sorting list on the user terminal after obtaining the probability sorting list.
Citation Information
Cited By
A data analysis method and device based on multi-dimensional feature similarity
CN122337567A