A data processing method and apparatus

By selecting a target model that matches the prediction scenario from multiple prediction models and verifying its performance using multiple validation models, the problem of low model selection efficiency in existing technologies is solved, and efficient and accurate prediction model selection and prediction index acquisition are achieved.

CN114334137BActive Publication Date: 2026-03-06BEIJING JIAHE HAISEN HEALTH TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111646407.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2026-03-06
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

How to select the most suitable prediction model from numerous prediction models to improve the efficiency and accuracy of model selection.

Method used

By acquiring the data to be processed and the information of the prediction scenario, a target prediction model that matches the information of the prediction scenario is selected from multiple pre-built prediction models. The clinical conclusion prediction performance is verified using multiple validation models. The optimal prediction model is selected based on the validation results, and the correlation between the information of the prediction scenario and the optimal prediction model is established.

Benefits of technology

It has achieved automated selection of the optimal prediction model, improved the efficiency and accuracy of model selection, reduced calculation errors, and improved the accuracy and calculation efficiency of prediction indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114334137B_ABST
    Figure CN114334137B_ABST
Patent Text Reader

Abstract

This application provides a data processing method and apparatus. This application utilizes computer technology to help simplify the workflow, collects and stores prediction models and automatically imports data, calculates prediction indicators, and then verifies the prediction indicators based on a series of verification models to quickly select the optimal prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology, and in particular to a data processing method and apparatus. Background Technology

[0002] With the development of big data medical technology, the value of medical data has received unprecedented attention. In the era of big data, the technologies for data acquisition, storage, analysis and prediction have developed rapidly. The ability to accurately predict the probability of a certain clinical outcome is also an inherent requirement of precision medicine. As a quantitative tool for risk assessment, clinical prediction models can provide doctors, patients and health administrators with more objective and accurate information for decision-making, and therefore their application is becoming more and more widespread.

[0003] After years of data accumulation, a certain number of prediction models for various application scenarios have been accumulated. However, the problem remains: how to select the prediction model that is more suitable for the application scenario from among the many prediction models. Summary of the Invention

[0004] This application provides the following technical solution:

[0005] A data processing method, comprising:

[0006] Obtain the first set of data to be processed and the predicted scenario information;

[0007] Select at least one target prediction model that matches the prediction scenario information from a plurality of pre-built prediction models, the prediction model being used to predict clinical conclusions;

[0008] The first data to be processed is input into the target prediction model to obtain the prediction index obtained by the target prediction model;

[0009] The prediction indicators are input into multiple validation models to obtain the validation results output by each validation model. The validation results characterize the clinical conclusion prediction performance of the target prediction model.

[0010] Based on the verification results output by each of the verification models, the optimal prediction model that meets the model selection criteria is selected from at least one of the target prediction models.

[0011] Establish the correlation between the predicted scenario information and the optimal prediction model, and save the correlation.

[0012] Optionally, the plurality of said verification models include:

[0013] Goodness-of-fit test model, ROC curve model, calibration curve model, and DCA curve model;

[0014] The validation results output by the goodness-of-fit test model characterize the degree of matching between the predicted index and the baseline clinical conclusion.

[0015] The validation results output by the ROC curve model characterize the relationship between the specificity and sensitivity of the prediction index.

[0016] The validation results output by the calibration curve model characterize the accuracy of the predicted index.

[0017] The validation results output by the DCA curve model characterize the practicality of the prediction index.

[0018] Optionally, the method further includes:

[0019] Obtain the second set of data to be processed and the target prediction scene information;

[0020] Based on the correlation between the predicted scenario information and the optimal prediction model, the optimal prediction model associated with the target predicted scenario information is determined.

[0021] The second data to be processed is input into the optimal prediction model to obtain the prediction index output by the optimal prediction model.

[0022] Optionally, the method further includes:

[0023] Establish the correlation between the prediction index obtained from the optimal prediction model and the first data to be processed;

[0024] Save the correlation between the prediction index obtained from the optimal prediction model and the first data to be processed.

[0025] Optionally, the method further includes:

[0026] Obtain the second set of data to be processed;

[0027] In the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed, it is determined whether there is any first data to be processed that is consistent with the second data to be processed.

[0028] If it exists, the prediction index obtained by the optimal prediction model associated with the first data to be processed is used as the prediction index to be used for the second data to be processed.

[0029] A data processing apparatus, comprising:

[0030] The first acquisition module is used to acquire the first data to be processed and the prediction scene information;

[0031] The first selection module is used to select at least one target prediction model from a plurality of pre-built prediction models that matches the prediction scenario information, the prediction model being used to predict clinical conclusions.

[0032] The first prediction module is used to input the first data to be processed into the target prediction model to obtain the prediction index obtained by the target prediction model.

[0033] The validation module is used to input the prediction index into multiple validation models respectively, and obtain the validation result output by each validation model. The validation result characterizes the clinical conclusion prediction performance of the target prediction model.

[0034] The second selection module is used to combine the verification results output by each of the verification models to select the optimal prediction model that meets the model screening conditions from at least one of the target prediction models.

[0035] The first module for establishing and saving is used to establish the correlation between the predicted scenario information and the optimal prediction model, and to save the correlation.

[0036] Optionally, the plurality of said verification models include:

[0037] Goodness-of-fit test model, ROC curve model, calibration curve model, and DCA curve model;

[0038] The validation results output by the goodness-of-fit test model characterize the degree of matching between the predicted index and the baseline clinical conclusion.

[0039] The validation results output by the ROC curve model characterize the relationship between the specificity and sensitivity of the prediction index.

[0040] The validation results output by the calibration curve model characterize the accuracy of the predicted index.

[0041] The validation results output by the DCA curve model characterize the practicality of the prediction index.

[0042] Optionally, the device further includes:

[0043] The second acquisition module is used to acquire the second data to be processed and the target prediction scene information;

[0044] The first determining module is used to determine the optimal prediction model associated with the target prediction scene information based on the correlation between the prediction scene information and the optimal prediction model.

[0045] The second prediction module is used to input the second data to be processed into the optimal prediction model to obtain the prediction index output by the optimal prediction model.

[0046] Optionally, the device further includes:

[0047] The second establishment module is used to establish the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0048] The second storage module is used to store the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0049] Optionally, the device further includes:

[0050] The third acquisition module is used to acquire the second data to be processed;

[0051] The second determining module is used to determine whether there is any first data to be processed that is consistent with the second data to be processed in the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0052] The third determining module is used to, if there is a first data to be processed that is consistent with the second data to be processed, use the prediction index obtained by the optimal prediction model associated with the first data to be processed as the prediction index to be used corresponding to the second data to be processed.

[0053] Compared with the prior art, the beneficial effects of this application are as follows:

[0054] In this application, by acquiring first data to be processed and prediction scenario information, at least one target prediction model matching the prediction scenario information is selected from a plurality of pre-constructed prediction models. The prediction model is used to predict clinical conclusions. The first data to be processed is input into the target prediction model to obtain the prediction index obtained by the target prediction model. The prediction index is input into a plurality of validation models to obtain the validation result output by each validation model. The validation result characterizes the clinical conclusion prediction performance of the target prediction model. Combining the validation results output by each validation model, the optimal prediction model that meets the model screening conditions is selected from at least one target prediction model, thereby realizing the automated screening of the optimal prediction model and ensuring the efficiency of model screening.

[0055] Furthermore, by inputting the prediction indicators into multiple validation models and obtaining the validation results output by each model, and by combining the validation results output by each model, the optimal prediction model can be selected, thus ensuring the accuracy of the selection. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a flowchart of a data processing method provided in Embodiment 1 of this application;

[0058] Figure 2 This is a flowchart of a data processing method provided in Embodiment 2 of this application;

[0059] Figure 3 This is a flowchart of a data processing method provided in Embodiment 3 of this application;

[0060] Figure 4 This is a flowchart of a data processing method provided in Embodiment 4 of this application;

[0061] Figure 5 This is a schematic diagram of the structure of a data processing device provided in this application. Detailed Implementation

[0062] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0063] To address the aforementioned problems, this application provides a data processing method, which will be described below.

[0064] Reference Figure 1 Here is a flowchart of a data processing method provided in Embodiment 1 of this application, as shown below. Figure 1 As shown, the method may include, but is not limited to, the following steps:

[0065] Step S11: Obtain the first data to be processed and the prediction scenario information.

[0066] In this embodiment, the first data to be processed can be understood as: medical-related clinical data.

[0067] Predictive scenario information can be understood as information that can characterize clinical predictive scenarios (e.g., lung cancer mortality prediction scenarios).

[0068] In this embodiment, obtaining the first data to be processed may include:

[0069] Retrieve the first piece of data to be processed from the database. This first piece of data can be stored in the database as a data table in CSV or Excel format.

[0070] In this embodiment, the electronic device can provide a prediction scene selection interface, which includes various prediction scene information. Users can select prediction scene information on the prediction scene selection interface, and the electronic device accordingly retrieves the selected prediction scene information from the interface. Users can select all prediction scene information or select only some as needed, ensuring flexibility in filtering.

[0071] Step S12: Select at least one target prediction model from a plurality of pre-built prediction models that matches the prediction scenario information, the prediction model being used to predict clinical conclusions.

[0072] In this embodiment, a correspondence between each pre-built prediction model and its applicable prediction scenario can be established in advance. Based on each pre-built prediction model and its applicable prediction scenario, at least one target prediction model that matches the prediction scenario information can be selected from the multiple pre-built prediction models.

[0073] The pre-built prediction models may include, but are not limited to, mathematical formulas, web page calculation models, and scoring models.

[0074] Mathematical formulas can be constructed in the following ways:

[0075] Mathematical formulas can be automatically retrieved from data sources (such as literature) and their parameters and variable names can be stored in a database.

[0076] Web page computing models can be constructed in the following ways:

[0077] The system retrieves websites from the internet suitable for computation, uses these websites as webpage computation models, and stores their web addresses in a database. When using the webpage computation model, it retrieves the website's web address from the database and accesses the website based on that address.

[0078] The scoring model can be constructed in the following ways:

[0079] Extract the variable names and their corresponding scores from the rating tables in the literature, and store the variable names and their corresponding scores in the database parameters and variables.

[0080] Step S13: Input the first data to be processed into the target prediction model to obtain the prediction index obtained by the target prediction model.

[0081] Step S14: Input the prediction index into multiple validation models respectively to obtain the validation result output by each validation model. The validation result characterizes the clinical conclusion prediction performance of the target prediction model.

[0082] The multiple verification models may include, but are not limited to:

[0083] Goodness-of-fit test model, ROC curve model, calibration curve model and DCA curve model.

[0084] The goodness-of-fit test model can calculate the chi-square value and p-value, calculate the predicted probability of each individual's outcome event, and regroup the data according to the order of predicted probabilities (usually into 5-10 groups) to perform the Hosmer-Lemeshow (HL) goodness-of-fit test to examine the degree of agreement between the predicted results and the actual situation.

[0085] The ROC curve model, which can be understood as a working characteristic curve, is a curve plotted based on a series of different thresholds, dividing the model into two categories. The true positive rate (sensitivity) is plotted on the ordinate, and the false positive rate (1-specificity) on the abscissa. The ROC curve graphically combines sensitivity and specificity, accurately reflecting the relationship between the specificity and sensitivity of the model's predicted values. The closer the ROC curve is to the upper left corner, and the larger the area under the curve, the greater its predictive value. It can also be used to compare different indicators. Generally, an area under the curve greater than 0.8 is considered to have higher diagnostic value; however, the specific predictive value needs to be considered in conjunction with clinical practice. The discriminative ability of the predictive model is evaluated by calculating the area under the curve (AUC), sen, spe, and accuracy.

[0086] Calibration curve models are primarily used to evaluate the accuracy of a model. A calibration curve is a scatter plot with predicted occurrences on the x-axis and actual occurrences on the y-axis. A straight line is fitted to this scatter plot; if the line is a straight line with a 45-degree slope passing through the origin, the model's accuracy is very good; the further away from a 45-degree slope line from the origin, the worse the prediction accuracy. In logistic regression analysis, the calibration curve is essentially a visualization of the Hosmer-Lemeshow goodness-of-fit test results.

[0087] The DCA curve model can be evaluated for its practicality using Decision Curve Analysis (DCA). The curve is plotted with the threshold probability on the x-axis and the net benefit (benefit minus harm) on the y-axis. The closer the curve is to the upper right corner, the better the practicality of the predictive model. In practical applications, there are two reference lines representing two extreme cases. The horizontal line indicates all samples are negative, no one receives intervention, and the net benefit is 0. The sloping line indicates all samples are positive, everyone receives intervention, and the net benefit is a negative-sloping line. A DCA curve closer to these two lines indicates poor clinical practicality.

[0088] Accordingly, the validation results output by the goodness-of-fit test model characterize the degree of matching between the predicted index and the baseline clinical conclusion;

[0089] The validation results output by the ROC curve model characterize the relationship between the specificity and sensitivity of the prediction index.

[0090] The validation results output by the calibration curve model characterize the accuracy of the predicted index.

[0091] The validation results output by the DCA curve model characterize the practicality of the prediction index.

[0092] In this embodiment, an open interface can be provided, which can receive custom scripts uploaded by developers, offering flexibility and convenience. The custom scripts are used to verify the prediction metrics according to custom verification rules.

[0093] Step S15: Combining the verification results output by each verification model, select the optimal prediction model that meets the model selection criteria from at least one target prediction model.

[0094] Model selection criteria can be set based on the types of multiple validation models. For example, model selection criteria can be, but are not limited to, the accuracy of the predictive indicator meeting a set accuracy threshold and the usefulness of the predictive indicator meeting a set usefulness threshold.

[0095] Step S16: Establish the correlation between the predicted scenario information and the optimal prediction model, and save the correlation.

[0096] In this application, by acquiring first data to be processed and prediction scenario information, at least one target prediction model matching the prediction scenario information is selected from a plurality of pre-constructed prediction models. The prediction model is used to predict clinical conclusions. The first data to be processed is input into the target prediction model to obtain the prediction index obtained by the target prediction model. The prediction index is input into a plurality of validation models to obtain the validation result output by each validation model. The validation result characterizes the clinical conclusion prediction performance of the target prediction model. Combining the validation results output by each validation model, the optimal prediction model that meets the model screening conditions is selected from at least one target prediction model, thereby realizing the automated screening of the optimal prediction model and ensuring the efficiency of model screening.

[0097] Furthermore, by inputting the prediction indicators into multiple validation models and obtaining the validation results output by each model, and by combining the validation results output by each model, the optimal prediction model can be selected, thus ensuring the accuracy of the selection.

[0098] For mathematical formulas and scoring models, users only need to select the prediction scenario information and import the data to automatically calculate all the prediction indicators and verify the prediction performance of the prediction model. This is very convenient and fast, reducing calculation errors caused by users inputting data from the analytical formulas and improving efficiency.

[0099] For web-based computational models, batch processing can be used to import parameter data all at once and calculate predictive indicators, improving the efficiency of indicator calculation. Furthermore, compared to errors caused by manually inputting data, the source can be traced, improving data accuracy.

[0100] As another optional embodiment of this application, refer to Figure 2 This is a flowchart of a data processing method provided in Embodiment 2 of this application. This embodiment is mainly an extension of the data processing method described in Embodiment 1 above, such as... Figure 2 As shown, the method may include, but is not limited to, the following steps:

[0101] Step S21: Obtain the first data to be processed and the prediction scene information.

[0102] Step S22: Select at least one target prediction model from a plurality of pre-built prediction models that matches the prediction scenario information, the prediction model being used to predict clinical conclusions.

[0103] Step S23: Input the first data to be processed into the target prediction model to obtain the prediction index obtained by the target prediction model.

[0104] Step S24: Input the prediction index into multiple validation models respectively to obtain the validation result output by each validation model. The validation result characterizes the clinical conclusion prediction performance of the target prediction model.

[0105] Step S25: Combining the verification results output by each verification model, select the optimal prediction model that meets the model selection criteria from at least one target prediction model.

[0106] Step S26: Establish the correlation between the predicted scene information and the optimal prediction model, and save the correlation.

[0107] For detailed procedures of steps S21-S26, please refer to the relevant description of steps S11-S16 in Example 1, which will not be repeated here.

[0108] Step S27: Establish the correlation between the prediction index obtained from the optimal prediction model and the first data to be processed.

[0109] Step S28: Save the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0110] In this embodiment, based on the correlation between the prediction index obtained from the optimal prediction model and the first data to be processed, the prediction index corresponding to the data to be processed that is consistent with the first data to be processed can be determined, which simplifies the prediction process and improves the efficiency of obtaining the prediction index.

[0111] As another optional embodiment of this application, refer to Figure 3 This is a flowchart of a data processing method provided in Embodiment 3 of this application. This embodiment is mainly an extension of the data processing method described in Embodiment 2 above, such as... Figure 3 As shown, the method may include, but is not limited to, the following steps:

[0112] Step S31: Obtain the first data to be processed and the prediction scene information.

[0113] Step S32: Select at least one target prediction model from a plurality of pre-built prediction models that matches the prediction scenario information, the prediction model being used to predict clinical conclusions.

[0114] Step S33: Input the first data to be processed into the target prediction model to obtain the prediction index obtained by the target prediction model.

[0115] Step S34: Input the prediction index into multiple validation models respectively to obtain the validation result output by each validation model. The validation result characterizes the clinical conclusion prediction performance of the target prediction model.

[0116] Step S35: Combining the verification results output by each verification model, select the optimal prediction model that meets the model selection criteria from at least one target prediction model.

[0117] Step S36: Establish the correlation between the predicted scene information and the optimal prediction model, and save the correlation.

[0118] Step S37: Establish the correlation between the prediction index obtained from the optimal prediction model and the first data to be processed.

[0119] Step S38: Save the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0120] For a detailed description of steps S31-S38, please refer to the relevant description of steps S21-S28 in Example 2, which will not be repeated here.

[0121] Step S39: Obtain the second data to be processed.

[0122] Based on selecting the optimal prediction model and establishing and saving the correlation between the prediction indicators obtained by the optimal prediction model and the first data to be processed, the optimal prediction model and the correlation between the prediction indicators obtained by the optimal prediction model and the first data to be processed can be applied. Specifically, the second data to be processed can be obtained.

[0123] Step S310: In the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed, determine whether there is any first data to be processed that is consistent with the second data to be processed.

[0124] If it exists, proceed to step S311; if it does not exist, the user can input the target prediction scene information, and based on the correlation between the prediction scene information and the optimal prediction model, determine the optimal prediction model associated with the target prediction scene information, input the second data to be processed into the optimal prediction model, and obtain the prediction index output by the optimal prediction model.

[0125] Step S311: Use the prediction index obtained by the optimal prediction model associated with the first data to be processed as the prediction index to be used for the second data to be processed.

[0126] In this embodiment, when the optimal prediction model and the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed can be applied, based on the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed, it is determined whether there is a first data to be processed that is consistent with the second data to be processed. If there is, the prediction index obtained by the optimal prediction model associated with the first data to be processed is used as the prediction index to be used corresponding to the second data to be processed, which simplifies the prediction process and can improve the efficiency of obtaining prediction indicators.

[0127] As another optional embodiment of this application, refer to Figure 4 This is a flowchart of a data processing method provided in Embodiment 4 of this application. This embodiment is mainly an extension of the data processing method described in Embodiment 1 above, such as... Figure 4 As shown, the method may include, but is not limited to, the following steps:

[0128] Step S41: Obtain the first data to be processed and the prediction scene information.

[0129] Step S42: Select at least one target prediction model from a plurality of pre-built prediction models that matches the prediction scenario information, the prediction model being used to predict clinical conclusions.

[0130] Step S43: Input the first data to be processed into the target prediction model to obtain the prediction index obtained by the target prediction model.

[0131] Step S44: Input the prediction index into multiple validation models respectively to obtain the validation result output by each validation model. The validation result characterizes the clinical conclusion prediction performance of the target prediction model.

[0132] Step S45: Combining the verification results output by each verification model, select the optimal prediction model that meets the model selection criteria from at least one target prediction model.

[0133] Step S46: Establish the correlation between the predicted scene information and the optimal prediction model, and save the correlation.

[0134] For detailed procedures of steps S41-S46, please refer to the relevant description of steps S11-S16 in Example 1, which will not be repeated here.

[0135] Step S47: Obtain the second data to be processed and the target prediction scene information.

[0136] Step S48: Based on the correlation between the predicted scene information and the optimal prediction model, determine the optimal prediction model associated with the target predicted scene information.

[0137] Step S49: Input the second data to be processed into the optimal prediction model to obtain the prediction index output by the optimal prediction model.

[0138] In this embodiment, based on establishing and saving the correlation between the predicted scene information and the optimal prediction model, the correlation can be applied for corresponding processing. Specifically, steps S47-S49 can be executed to directly match the optimal prediction model for predicting the second data to be processed, thereby obtaining prediction indicators and improving prediction efficiency and accuracy.

[0139] The data processing apparatus provided in this application will be described below. The data processing apparatus described below can be referred to in correspondence with the data processing method described above.

[0140] Please see Figure 5 The data processing device includes: a first acquisition module 100, a first selection module 200, a first prediction module 300, a verification module 400, a second selection module 500, and a first creation and saving module 600.

[0141] The first acquisition module 100 is used to acquire the first data to be processed and the prediction scene information;

[0142] The first selection module 200 is used to select at least one target prediction model that matches the prediction scenario information from a plurality of pre-built prediction models, the prediction model being used to predict clinical conclusions.

[0143] The first prediction module 300 is used to input the first data to be processed into the target prediction model to obtain the prediction index obtained by the target prediction model.

[0144] The verification module 400 is used to input the prediction index into multiple verification models respectively, and obtain the verification result output by each verification model. The verification result characterizes the clinical conclusion prediction performance of the target prediction model.

[0145] The second selection module 500 is used to select the optimal prediction model that meets the model screening conditions from at least one target prediction model by combining the verification results output by each of the verification models.

[0146] The first establishment and saving module 600 is used to establish the correlation between the predicted scene information and the optimal prediction model, and to save the correlation.

[0147] In this embodiment, the plurality of verification models may include:

[0148] Goodness-of-fit test model, ROC curve model, calibration curve model, and DCA curve model;

[0149] The validation results output by the goodness-of-fit test model characterize the degree of matching between the predicted index and the baseline clinical conclusion.

[0150] The validation results output by the ROC curve model characterize the relationship between the specificity and sensitivity of the prediction index.

[0151] The validation results output by the calibration curve model characterize the accuracy of the predicted index.

[0152] The validation results output by the DCA curve model characterize the practicality of the prediction index.

[0153] In this embodiment, the data processing apparatus may further include:

[0154] The second acquisition module is used to acquire the second data to be processed and the target prediction scene information;

[0155] The first determining module is used to determine the optimal prediction model associated with the target prediction scene information based on the correlation between the prediction scene information and the optimal prediction model.

[0156] The second prediction module is used to input the second data to be processed into the optimal prediction model to obtain the prediction index output by the optimal prediction model.

[0157] In this embodiment, the data processing apparatus may further include:

[0158] The second establishment module is used to establish the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0159] The second storage module is used to store the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0160] In this embodiment, the data processing apparatus may further include:

[0161] The third acquisition module is used to acquire the second data to be processed;

[0162] The second determining module is used to determine whether there is any first data to be processed that is consistent with the second data to be processed in the correlation between the prediction index obtained by the optimal prediction model and the first data to be processed.

[0163] The third determining module is used to, if there is a first data to be processed that is consistent with the second data to be processed, use the prediction index obtained by the optimal prediction model associated with the first data to be processed as the prediction index to be used corresponding to the second data to be processed.

[0164] It should be noted that each embodiment focuses on describing the differences from other embodiments, and the same or similar parts between the embodiments can be referred to accordingly. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0165] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0166] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0167] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0168] The data processing method and apparatus provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A data processing method, characterized by, The method comprises the following steps: obtaining first to-be-processed data and prediction scene information; selecting at least one target prediction model matched with the prediction scene information from a plurality of prediction models pre-constructed based on a correspondence relationship between each prediction model and the prediction scene to which the prediction model is applicable, the prediction model being used for predicting a clinical conclusion; inputting the first to-be-processed data into the target prediction model to obtain a prediction index obtained by the target prediction model; inputting the prediction index into a plurality of verification models respectively to obtain a verification result output by each verification model, the verification result representing a clinical conclusion prediction performance of the target prediction model; the plurality of verification models comprise a goodness-of-fit test model, a ROC curve model, a calibration curve model and a DCA curve model; the verification result output by the goodness-of-fit test model represents a matching degree between the prediction index and a benchmark clinical conclusion, the goodness-of-fit test model being used for evaluating an agreement degree between a prediction result of the target prediction model and an actual condition; the verification result output by the ROC curve model represents a relationship between a specificity and a sensitivity of the prediction index, the ROC curve model being used for evaluating a discrimination ability of the target prediction model; the verification result output by the calibration curve model represents an accuracy of the prediction index, the calibration curve model being used for evaluating an accuracy evaluation of the target prediction model; the verification result output by the DCA curve model represents a practicality degree of the prediction index, the DCA curve model being used for evaluating a practicality problem of the target prediction model; selecting an optimal prediction model meeting a model screening condition from the at least one target prediction model in combination with the verification result output by each verification model, the model screening condition being set according to types of the plurality of verification models; establishing an association relationship between the prediction scene information and the optimal prediction model and saving the association relationship; the method further comprises: establishing an association relationship between the prediction index obtained by the optimal prediction model and the first to-be-processed data and saving the association relationship between the prediction index obtained by the optimal prediction model and the first to-be-processed data; obtaining second to-be-processed data; determining whether there is first to-be-processed data consistent with the second to-be-processed data in the association relationship between the prediction index obtained by the optimal prediction model and the first to-be-processed data; if there is, taking the prediction index obtained by the optimal prediction model associated with the first to-be-processed data as a to-be-used prediction index corresponding to the second to-be-processed data.

2. The method of claim 1, wherein, the method further comprises: obtaining second to-be-processed data and target prediction scene information; determining an optimal prediction model associated with the target prediction scene information based on the association relationship between the prediction scene information and the optimal prediction model; inputting the second to-be-processed data into the optimal prediction model to obtain a prediction index output by the optimal prediction model.

3. A data processing apparatus, characterized by, The method comprises the following steps: a first obtaining module is configured to obtain first to-be-processed data and prediction scene information; The first selection module is configured to select at least one target prediction model matched with the prediction scene information from a plurality of pre-constructed prediction models, the prediction model being used for predicting a clinical conclusion. The first prediction module is configured to input the first to-be-processed data into the target prediction model to obtain a prediction index obtained by the target prediction model. The verification module is configured to input the prediction index into a plurality of verification models respectively to obtain a verification result output by each of the verification models, the verification result representing a clinical conclusion prediction performance of the target prediction model. The plurality of verification models include a goodness-of-fit test model, a ROC curve model, a calibration curve model and a DCA curve model; the verification result output by the goodness-of-fit test model represents a matching degree between the prediction index and a benchmark clinical conclusion, the goodness-of-fit test model being used for evaluating an agreement degree between a prediction result of the target prediction model and an actual condition; the verification result output by the ROC curve model represents a relationship between a specificity and a sensitivity of the prediction index, the ROC curve model being used for evaluating a discrimination ability of the target prediction model; the verification result output by the calibration curve model represents an accuracy of the prediction index, the calibration curve model being used for evaluating an accuracy evaluation of the target prediction model; and the verification result output by the DCA curve model represents a practicality degree of the prediction index, the DCA curve model being used for evaluating a practicality problem of the target prediction model. The second selection module is configured to select an optimal prediction model satisfying a model screening condition from the at least one target prediction model in combination with the verification result output by each of the verification models, the model screening condition being set according to types of the plurality of verification models. The first establishment and saving module is configured to establish an association relationship between the prediction scene information and the optimal prediction model, and save the association relationship. The device further includes: The second establishment module is configured to establish an association relationship between the prediction index obtained by the optimal prediction model and the first to-be-processed data. The second saving module is configured to save the association relationship between the prediction index obtained by the optimal prediction model and the first to-be-processed data. The third acquisition module is configured to acquire second to-be-processed data. The second determination module is configured to determine whether there is first to-be-processed data consistent with the second to-be-processed data in the association relationship between the prediction index obtained by the optimal prediction model and the first to-be-processed data. The third determination module is configured to, if there is first to-be-processed data consistent with the second to-be-processed data, take the prediction index obtained by the optimal prediction model associated with the first to-be-processed data as a to-be-used prediction index corresponding to the second to-be-processed data.

4. The apparatus of claim 3, wherein, The device further includes: The second acquisition module is configured to acquire second to-be-processed data and target prediction scene information. The first determination module is configured to determine an optimal prediction model associated with the target prediction scene information based on the association relationship between the prediction scene information and the optimal prediction model. A second prediction module is configured to input the second to-be-processed data into the optimal prediction model to obtain a prediction index output by the optimal prediction model.

Citation Information

Patent Citations

  • Method and device for obtaining clinical data prediction model, readable medium and electronic equipment

    CN110648764A