Disease Prediction Set Processing Method, Device, Electronic Device and Storage Medium

By identifying and adjusting the sort of disease prediction sets of error-prone samples, the problem of low top1 accuracy caused by confusing disease clusters is solved, and the overall accuracy of disease prediction sets is improved.

CN113921144BActive Publication Date: 2025-07-04TSINGHUA UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111194507.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-23
Filing Date
2021-10-13
Publication Date
2025-07-04
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

The existing disease prediction model has a low accuracy rate in output top1, which is mainly due to the similar symptoms of disease clusters or disease pairs, which makes it difficult to distinguish, which in turn affects the accuracy of disease prediction sets.

Method used

By obtaining the symptom characteristics of the disease prediction set, scoring the detection samples based on the degree of confusion, identifying the samples that are prone to errors, and adjusting the confidence order of the disease prediction set to improve the prediction accuracy.

Benefits of technology

The top1 accuracy of the disease prediction set is improved, while the top2 and top3 accuracy are maintained or improved, enhancing the overall prediction effect of the disease prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113921144B_ABST
    Figure CN113921144B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, electronic device, and storage medium for processing a disease prediction set. Among them, the method includes obtaining a disease prediction set of a sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature; scoring the sample to be detected based on the degree of confusion of at least one symptom feature to determine whether the sample to be detected is an error-prone sample; if the sample to be detected is an error-prone sample, adjust the confidence ranking in the disease prediction set of the sample to be detected. By the above method, the present application can improve the prediction accuracy of the disease prediction set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent medical technology, and particularly to a method for processing a disease prediction set, an electronic device, and a storage medium. Background Art

[0002] With the development of information technology, using information technology to replace or assist manual execution of many tasks such as production and decision-making has become one of the mainstream trends in the development of current information technology. In the medical scenario, information technology can provide prediction recommendations for medical diagnosis problems, thereby assisting doctors in medical diagnosis.

[0003] Currently, the accuracy of the top 2 or top 3, etc. in the disease prediction set output by the disease prediction model is relatively high, while the accuracy of the top 1 is relatively low. The reason is that it is affected by easily confused disease clusters or disease pairs, and the symptoms of easily confused disease clusters or disease pairs are very similar. Therefore, it is difficult to find a combination of one or more symptoms as a gold standard to distinguish easily confused diagnosis clusters or disease pairs, resulting in difficulty in improving the prediction accuracy of the disease prediction set. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a method, device, electronic device, and storage medium for processing a disease prediction set to improve the prediction accuracy of the disease prediction set.

[0005] To solve the above technical problem, in the first aspect of the embodiments of this application, a method for processing a disease prediction set is provided. The method includes: obtaining a disease prediction set of a sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature; scoring the sample to be detected based on the degree of confusion of at least one symptom feature to determine whether the sample to be detected is an error-prone sample; if the sample to be detected is an error-prone sample, adjusting the confidence ranking in the disease prediction set of the sample to be detected.

[0006] To solve the above technical problem, in the second aspect of the embodiments of this application, a device for processing a disease prediction set is provided. The device includes: an obtaining module, configured to obtain a disease prediction set of a sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature; a processing module, configured to score the sample to be detected based on the degree of confusion of at least one symptom feature to determine whether the sample to be detected is an error-prone sample; an adjustment module, configured to, if the sample to be detected is an error-prone sample, adjust the confidence ranking in the disease prediction set of the sample to be detected.

[0007] To solve the above technical problems, a third aspect of the embodiments of the present application provides an electronic device, which includes a processor and a memory connected to the processor. The memory is used to store program data, and the processor is used to execute the program data to implement the foregoing method.

[0008] To solve the above technical problems, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, in which program data is stored. When the program data is executed by a processor, it is used to implement the foregoing method.

[0009] The beneficial effect of the present application is as follows: Different from the prior art, the present application obtains a disease prediction set of a sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature. Then, based on the degree of confusion of at least one symptom feature, the sample to be detected is scored to comprehensively consider the degree of easy prediction error of each symptom feature in the sample to be detected, and the degree of easy prediction error of the sample to be detected is obtained and reflected in the score. Thus, an error-prone sample can be determined based on the score of the sample to be detected, and then by adjusting the confidence ranking in the disease prediction set of the error-prone sample, the prediction accuracy of the disease prediction set can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a schematic flowchart of an embodiment of the method for processing a disease prediction set of the present application;

[0011] Figure 2 is another schematic flowchart of an embodiment of the method for processing a disease prediction set of the present application;

[0012] Figure 3 is still another schematic flowchart of an embodiment of the method for processing a disease prediction set of the present application;

[0013] Figure 4 is a schematic flowchart of another embodiment of the method for processing a disease prediction set of the present application;

[0014] Figure 5 is Figure 4 a schematic flowchart of another implementation manner of step S24 in

[0015] Figure 6 is a schematic flowchart of an embodiment of the method for obtaining a preset disease combination of the present application;

[0016] Figure 7 is a schematic flowchart of an embodiment of the method for obtaining a confusion score table of the present application;

[0017] Figure 8 is Figure 7 a schematic flowchart of another embodiment of step S45 in

[0018] Figure 9 It is another process schematic diagram of an embodiment of the method for obtaining the confusion score table of the present application.

[0019] Figure 10 It is a process schematic diagram of an embodiment of the method for obtaining the preset score threshold of the present application.

[0020] Figure 11 It is another process schematic diagram of an embodiment of the method for obtaining the preset score threshold of the present application.

[0021] Figure 12 It is a process schematic diagram of the training and testing stages of the present application.

[0022] Figure 13 It is a framework schematic diagram of an embodiment of the disease prediction set processing device of the present application.

[0023] Figure 14 It is a framework schematic diagram of an embodiment of the electronic device of the present application.

[0024] Figure 15 It is a framework schematic diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0026] The terms "first" and "second" in the present application are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0027] The symptoms of confusing disease clusters or confusing disease pairs are very similar, or they are the causes or concomitant diseases of each other. Therefore, the disease prediction model can better distinguish them from other diseases, but it is very difficult to select the correct prediction within the confusing diseases. A typical example is bronchial asthma and pneumonia, both of which present some respiratory symptoms, such as shortness of breath, coughing, etc. In addition, pneumonia may also cause fever, but it may not be accompanied by fever. Therefore, fever cannot be used as the gold standard for differentiating the two. Therefore, it is unrealistic and inaccurate to simply find a combination of one or more symptoms as the gold standard to distinguish confusing disease clusters or confusing disease pairs.

[0028] To address the problem of difficult identification of confusing diseases, the present application provides a method for processing a disease prediction set. By obtaining the disease prediction set of a sample to be detected, where the sample to be detected includes at least one symptom feature, it is possible to score the sample to be detected based on the degree of confusion of at least one symptom feature to determine whether the sample to be detected is an error-prone sample. If it is determined that the sample to be detected is an error-prone sample, the confidence ranking in the disease prediction set of the sample to be detected can be adjusted, that is, the ranking of at least two disease types in the disease prediction set output by the model is adjusted, so that the accuracy of the finally output disease prediction set is higher.

[0029] When "embodiment" is mentioned in the present application, it means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0030] For the method provided in the embodiments of the present application, the execution subject of each step can be an electronic device. For example, the electronic device can be a mobile phone, a computer (such as a tablet computer, a notebook computer, a desktop computer, etc.), a wearable device, a medical device, etc. It should be noted that the method for processing a disease prediction set provided in the present application is a method for processing the disease prediction set output by a disease prediction model to improve the prediction accuracy of the disease prediction set, where the disease prediction set is used to assist doctors in medical diagnosis.

[0031] Please refer to Figures 1 to 3 , Figure 1 which is a schematic flowchart of an embodiment of the method for processing a disease prediction set of the present application, Figure 2 which is another schematic flowchart of an embodiment of the method for processing a disease prediction set of the present application, Figure 3 which is yet another schematic flowchart of an embodiment of the method for processing a disease prediction set of the present application.

[0032] AsFigure 1 As shown in Figure 1 , the method may include the following steps:

[0033] Step S11: Obtain a disease prediction set of a sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature.

[0034] Wherein, as Figure 2 shown in Figure 2 , the sample to be detected may include at least one symptom. One or more symptoms may correspond to one symptom feature, that is, the disease feature may be composed of one or more symptoms. A symptom generally refers to the abnormal sensations subjectively felt by a patient or certain objective pathological changes caused by a series of functional, metabolic, and morphological structure abnormalities in the body during the disease process. The sample to be detected may be a text recording at least one symptom.

[0035] Optionally, the symptom may be either described in natural language or a feature extracted from the patient's medical image information, such as the current status description of a certain organ of the patient extracted by a convolutional network, for example, whether there is a shadow in the lungs obtained from a CT image, information such as a parenchymal hypoechoic area obtained from an ultrasound instrument, etc.

[0036] In some embodiments, when selecting a combination of multiple symptoms as a symptom feature, the multiple symptoms may be sorted in a certain order (such as lexicographical order) as the unique and fixed expression of the symptom feature. For example, "fever" and "cough" are sorted in lexicographical order as "fever" -> "cough", and this is used as the unique and fixed expression of the combination of "fever" and "cough". At this time, the symptom feature does not consider the order of the symptoms. In other embodiments, "fever" -> "cough" may also be used as one symptom feature, and "cough" -> "fever" may be used as another symptom feature. At this time, the order of the symptoms is considered.

[0037] In some specific embodiments, the sample to be detected may be a medical record or a text corresponding to the patient's symptoms recorded in the medical record. A medical record is a record of the medical activities of medical staff in the process of examining, diagnosing, and treating the occurrence, development, and outcome of a patient's disease. Generally, the patient's symptoms are recorded in it.

[0038] In some embodiments, the disease prediction set can be obtained by processing a sample to be detected using a disease prediction model. Among them, the disease prediction model can be any text classification model, such as BERT, TextRCNN, DPCNN, TextGCN, Fasttext, etc. Taking the BERT model as an example, the basic principle is to use text classification to map a text describing a patient's symptoms to the predicted disease name, that is, the label of this text, so as to obtain the disease prediction set. Specifically, the medical record data can be used to train the disease prediction model. After obtaining a relatively perfect model by adjusting parameters and other methods, the sample to be detected is input into the trained disease prediction model to obtain the disease prediction set.

[0039] Among them, the methods for optimizing the disease prediction model include adjusting the model network structure. Common means include GRU (Gate Recurrent Unit) or LSTM (Long-Short Term Memory), that is, upgrading the basic unit of RNN (Recurrent Neural Networks) to GRU or LSTM structure to obtain better long-term memory ability. Secondly, some training techniques can also be adopted, such as multi-fold calculation. In addition, the hyperparameters of the model can also be adjusted, such as the learning rate, batch number, dropout ratio, etc. In this embodiment, selecting to perform disease prediction based on the model is beneficial for processing complex long texts, adapting to real-world needs, and can obtain good basic performance. On this basis, using the disease prediction set processing method provided in this embodiment to further optimize the disease prediction set can obtain better prediction effects.

[0040] In this embodiment, the disease prediction model can output a disease prediction set including at least two disease types, and the at least two disease types are sorted by confidence. As Figure 2 shown, it can be sorted in descending order of confidence, such as Disease A -> Disease B -> Disease C ->...

[0041] Step S12: Score the sample to be detected based on the confusion degree of at least one symptom feature to determine whether the sample to be detected is an error-prone sample.

[0042] Among them, the confusion degree of the symptom feature is obtained through big data analysis. For example, objective data obtained by a machine analyzing and statistically processing a large amount of medical record data. For specific details, refer to the description of the embodiment of the confusion score table later.

[0043] In some embodiments, the confusing score table records the scores of each symptom feature, and the scores of the symptom features are used to reflect the confusing degree of the symptom features. Specifically, the confusing score table of the symptom features can be used to obtain the scores of each symptom feature in the sample to be detected. The scores can reflect the confusing degree of the symptom features, that is, the degree of easy prediction error. By comprehensively considering the confusing degrees of multiple symptom features in the sample to be detected, the sample to be detected is scored, so that the degree of easy prediction error of the sample to be detected can be reflected in the scores. Furthermore, it can be determined whether the sample to be detected is an error-prone sample according to the score of the sample to be detected, and the accuracy rate is relatively high.

[0044] In some embodiments, step S12 is only executed when the highest confidence level in the disease prediction set of the sample to be detected is less than the preset confidence threshold. Here, the credibility of the top1 result with the highest confidence level greater than the preset confidence threshold is relatively high. Therefore, such samples to be detected can be directly output without reordering, and the current top1 result is directly output.

[0045] Step S13: If the sample to be detected is an error-prone sample, the confidence level ranking in the disease prediction set of the sample to be detected is adjusted.

[0046] If the sample to be detected is an error-prone sample, it means that the accuracy rate of the original confidence level ranking in the disease prediction set of the sample to be detected is relatively low. Therefore, the confidence level ranking can be adjusted.

[0047] As Figure 2 shown, in some embodiments, the sample to be detected includes symptom 1 to symptom N (N is a positive integer). Among them, at least two of symptom 1 to symptom N can form a symptom feature, such as symptom 1 and symptom 2. Among them, at least one of the different symptom features is different, or the arrangement order of the symptoms is different. It can be understood that different symptom features can include the same symptom, such as symptom feature 1 (including symptom 1 and symptom 2) and symptom feature 2 (including symptom 1 and symptom 3). The sample to be detected can obtain a disease prediction set through a disease prediction model. The disease prediction set includes multiple disease types sorted by confidence level from high to low, that is, for example, disease A -> disease B -> disease C ->... Then, the confusing score table of the symptom features can be used to determine whether the sample to be detected is an error-prone sample. If it is an error-prone sample, the confidence level in the disease prediction set of the sample to be detected is reordered. For example, disease B is adjusted to swap positions with disease A. When only the top1 result is output, disease B is output; if it is not an error-prone sample, that is, a normal sample, disease A is directly output.

[0048] Specifically, when the sample to be detected is an error-prone sample, it indicates that the top1 accuracy rate in the current disease prediction set is relatively low. Therefore, the order of the disease types with the highest confidence levels is adjusted. For example, the ranking of the disease type with the highest confidence level is adjusted backward, and the rankings of the disease types with other confidence levels are adjusted forward to swap the top1 disease type. For example, if the previous confidence ranking is Disease A -> Disease B -> Disease C, the adjusted confidence ranking can be Disease B -> Disease A -> Disease C, or Disease C -> Disease B -> Disease A.

[0049] As an example, as Figure 3 shown, the symptoms in the sample to be detected may specifically include: cough, sneezing, runny nose, no fever, no vomiting, no shortness of breath, no chest tightness and shortness of breath, no diarrhea. After the sample to be detected is processed by a disease prediction model (such as BERT), a disease prediction set can be obtained. The disease types in the disease prediction set are arranged in descending order of confidence as "pneumonia" > "bronchial asthma" -> "rhinitis" -> "enteritis" -> "dermatitis". It can be found that there are easily confused diseases in the disease prediction set directly output by the disease prediction model, that is, "pneumonia" and "bronchial asthma" are an easily confused disease pair, and their corresponding symptoms are similar, so that the model is difficult to distinguish between the two. Therefore, the top2 accuracy rate or top3 accuracy rate of the disease prediction set is relatively high, but the top1 accuracy rate is relatively low.

[0050] Among them, the confidence level of the top1 result (such as "pneumonia") is the highest, and the confidence level of the top5 result (such as "dermatitis") is the lowest. The top1 accuracy rate value refers to the accuracy rate given for the evaluation criterion of "when the top1 result is recognized as correct, it is determined to be correct". The top2 accuracy rate value refers to the accuracy rate given for the evaluation criterion of "as long as one of the top1 or top2 hits the correct result, it is determined to be correct", and so on.

[0051] In this regard, the embodiment of the present application scores the sample to be detected based on the degree of confusion of at least one symptom feature, so that the degree of easy prediction error of the sample to be detected can be reflected in the score. Thus, it can be determined whether the sample to be detected is an error-prone sample according to the score of the sample to be detected, and then the confidence ranking in the disease prediction set of the error-prone sample can be adjusted to improve the prediction accuracy rate of the disease prediction set. In the above example, the true disease type is "bronchial asthma". The embodiment of the present application can identify and process the error-prone sample, and then adjust the confidence ranking in the disease prediction set. As Figure 3 shown in the optimized top3 result output, "bronchial asthma" is adjusted to the first place, that is, swapped with "pneumonia", so that the top1 accuracy rate in the disease prediction set can be improved while ensuring that the top2 accuracy rate or top3 accuracy rate remains unchanged.

[0052] Among them, it can be understood that Figure 3 The top 3 results output in the middle are only for illustration. In other embodiments, the number of output results can be selected according to needs. For example, the top 1 result, the top 2 result, the top 5 result, etc. can be output.

[0053] In this embodiment, by obtaining the disease prediction set of the sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, the sample to be detected includes at least one symptom feature, and then scoring the sample to be detected based on the degree of confusion of at least one symptom feature, so as to comprehensively consider the degree of easy prediction error of each symptom feature in the sample to be detected, obtain the degree of easy prediction error of the sample to be detected, and reflect it in the score, so that the error-prone sample can be determined based on the score of the sample to be detected, and then by adjusting the confidence ranking in the disease prediction set of the error-prone sample, the prediction accuracy of the disease prediction set can be improved.

[0054] Please refer to Figures 4 to 5 , Figure 4 is a schematic flowchart of another embodiment of the method for processing the disease prediction set of the present application, Figure 5 is Figure 4 a schematic flowchart of another embodiment of step S24 in

[0055] As Figure 4 shown, the method may include the following steps:

[0056] Step S21: Obtain the disease prediction set of the sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature.

[0057] For the description of this step, reference can be made to step S11 in the above embodiment, which will not be elaborated here.

[0058] Step S22: Determine the disease combination to be matched of at least two disease types in the disease prediction set.

[0059] Among them, the disease combination to be matched includes at least two disease types. In some embodiments, the first several disease types sorted by confidence from high to low in the disease prediction set can be used as the disease combination to be matched. In the present application, several means two or more.

[0060] Optionally, the number of disease types in the disease combination to be matched can be selected according to actual needs, and this embodiment does not limit this. For the convenience of description, in this embodiment, the number of disease types in the disease combination to be matched is 2, that is, two disease types are used as an example for illustration.

[0061] Continue to refer toFigure 2 In the example shown, the top two disease types sorted by confidence level from high to low in the disease prediction set can be selected as the disease combination to be matched, namely "pneumonia" and "bronchial asthma".

[0062] Step S23: Determine whether the disease combination to be matched matches the preset disease combination.

[0063] If the disease combination to be matched matches the preset disease combination, then execute Step S24.

[0064] If the disease combination to be matched does not match the preset disease combination, then end and exit the process.

[0065] Among them, the preset disease combination can be composed of at least two easily confused disease types, such as the two easily confused disease types of "bronchial asthma" and "pneumonia". Among them, the preset disease combination can be obtained through big data statistics, or can also be set by the user according to experience. This application also provides a way to obtain the preset disease combination. For details, please refer to the subsequent embodiments and will not be elaborated here.

[0066] In some embodiments, there are multiple types of preset disease combinations. Therefore, whether the disease combination to be matched matches the preset disease combination can refer to whether there is a preset disease combination that matches the disease combination to be matched among multiple preset disease combinations.

[0067] In some embodiments, it can be determined whether the types of diseases in the disease combination to be matched are the same as the types of diseases in the preset disease combination. If they are the same, it is determined that the disease combination to be matched matches the preset disease combination; otherwise, it is determined that the disease combination to be matched does not match the preset disease combination. In this embodiment, the order of the disease types in the disease combination to be matched may not be used as a matching requirement. For example, the disease combination to be matched 1 (including disease B - disease A) and the disease combination to be matched 2 (including disease A - disease B) both match the preset disease combination (including disease A - disease B). In other embodiments, the order of the disease types in the disease combination to be matched can also be used as a matching requirement, that is, when the sorting of the disease types in the disease combination to be matched is the same as the order of the disease types in the preset disease combination, it is determined that the disease combination to be matched matches the preset disease combination.

[0068] In this embodiment, if the disease combination to be matched of the sample to be detected matches the preset disease combination, it can be indicated that the error-prone possibility of the sample to be detected is relatively high. If the disease combination to be matched does not match the preset disease combination, it can be indicated that the error-prone possibility of the sample to be detected is relatively low. Therefore, through steps S22 and S23, the disease prediction set of the sample to be detected can be preliminarily screened first to screen out the samples to be detected with a relatively high error-prone possibility, and then the sample to be detected is scored by using the confusion score table of symptom features. For the samples to be detected with a relatively low error-prone possibility, the process can be ended to simplify the process steps and save computing resources.

[0069] Step S24: Use the confusion score table of symptom features to score the sample to be detected to determine whether the sample to be detected is an error-prone sample.

[0070] In this embodiment, using the confusion score table of symptom features to score the sample to be detected may include steps S241 to S242:

[0071] Step S241: Find the score corresponding to each symptom feature in the sample to be detected in the confusion score table.

[0072] In this embodiment, the confusion score table is used to record the scores of all symptom features corresponding to at least one preset disease combination. Specifically, the preset disease combination that matches the combination to be matched can be determined, and then the symptom features that overlap with the symptom features corresponding to the preset disease combination among the symptom features that appear in the sample to be detected are found in the confusion score table, and then the scores of the overlapping symptom features are obtained.

[0073] In other embodiments, each preset disease combination may correspond to a confusion score table respectively, that is, a confusion score table is used to record the scores of all symptom features corresponding to one preset disease combination.

[0074] Step S242: Calculate the sum of the scores corresponding to at least one symptom feature of the sample to be detected to obtain the score of the sample to be detected.

[0075] Specifically, the score of the sample to be detected can be calculated by using formula (1) as follows:

[0076]

[0077] In formula (1), g i represents the score of the sample to be detected, s j represents the score corresponding to the jth symptom feature in the sample to be detected in the confusion score table. The sample to be detected includes a total of j overlapping symptom features in the confusion score table, where the value of j is a positive integer (N).

[0078] In this embodiment, determining whether a sample to be detected is an error-prone sample may include steps S243 to S245:

[0079] Step S243: Determine whether the score of the sample to be detected is greater than a preset score threshold.

[0080] If the score of the sample to be detected is greater than the preset score threshold, then execute step S244.

[0081] If the score of the sample to be detected is less than or equal to the preset score threshold, then execute step S245.

[0082] Among them, the preset score threshold can be set according to the actual situation. In some embodiments, the preset score threshold can be an empirical value obtained by the user based on experience. In other embodiments, the preset score threshold can be obtained through big data calculation, and the specific description can be referred to the following embodiments.

[0083] Step S244: Determine that the sample to be detected is an error-prone sample.

[0084] After determining that the sample to be detected is an error-prone sample, continue to execute step S25.

[0085] Step S245: Determine that the sample to be detected is a normal sample.

[0086] After determining that the sample to be detected is a normal sample, the process can be exited and the current disease prediction set can be directly output.

[0087] Step S25: If the sample to be detected is an error-prone sample, adjust the confidence ranking in the disease prediction set of the sample to be detected.

[0088] The description of this step can be referred to step S13 in the above embodiment, and will not be repeated here.

[0089] Please refer to Figure 6 , Figure 6 is a schematic flowchart of an embodiment of the method for obtaining a preset disease combination in the present application. As Figure 6 shown, the preset disease combination is obtained through the following steps:

[0090] Step S31: Obtain the disease training sets of multiple training samples, where the disease training set includes at least two disease types sorted by confidence, and the training samples include at least one symptom feature and the true disease type.

[0091] Among them, the training samples may include at least one symptom. One or more symptoms may correspond to one symptom feature, that is, the disease feature may be composed of one or more symptoms. The number of training samples can be selected according to the actual situation and is not limited here. For example, the number of training samples is 100 to 10,000. For the description of symptoms, reference can be made to step S11 and will not be elaborated here.

[0092] In some specific embodiments, the training samples may be medical records or the text in the medical records that correspondingly records the patient's symptoms and the true disease type. The true disease type may be the disease type given by the doctor, which can be considered as the true disease type here. Different from the above embodiments, in this embodiment, since the training samples also include the true disease type for subsequent verification of the disease training set to determine the training samples with incorrect predictions and the training samples with correct predictions, and then count the undetermined disease combinations with a relatively large number in the incorrect samples as the preset disease combinations.

[0093] In some embodiments, the disease training set may be obtained by processing the training samples through a disease prediction model. Specifically, as Figure 12 shown, a training sample set containing multiple training samples can be input into the disease prediction model, and the disease prediction model processes each of the multiple training samples to obtain multiple disease training sets.

[0094] In this embodiment, the disease prediction model can output a disease training set including at least two disease types, and these two disease types are sorted according to the confidence level. Specifically, they can be sorted in descending order of confidence level.

[0095] It can be understood that the process of obtaining the disease training set of multiple training samples in this embodiment is similar to the process of obtaining the disease prediction set of the sample to be detected in the above embodiment. Therefore, the description of this step can refer to the corresponding position in the above embodiment and will not be elaborated here.

[0096] Step S32: Determine the incorrect samples among the multiple training samples according to the matching result between the disease training set of each training sample and the true disease type.

[0097] Specifically, it can be determined whether the disease type with the highest confidence level in the disease training set of the training sample is the same as the true disease type. If not, it is determined that the training sample is an incorrect sample, that is, the prediction of the training sample is incorrect; otherwise, it is determined that the training sample is a correct sample, that is, the prediction of the training sample is correct. Thus, the incorrect samples among the multiple training samples can be determined.

[0098] Step S33: Determine the undetermined disease combination of each incorrect sample, where the undetermined disease combination is at least two disease types sequentially selected from the disease training set of the incorrect sample according to the confidence level from high to low.

[0099] The reason why the wrong samples are preset wrongly may be that they are affected by confounding diseases, making it difficult for the disease prediction model to distinguish confounding diseases.

[0100] Among them, the pending disease combination includes at least two disease types. In some embodiments, the first multiple disease types in the disease training set of the error sample ranked from high to low in confidence can be used as the pending disease combination. Optionally, the number of disease types in the pending disease combination can be selected according to actual needs, and this embodiment does not limit this.

[0101] For the convenience of explanation, this embodiment is illustrated by taking the number of disease types in the pending disease combination as 2. Specifically, the first two disease types in the pending disease combination ranked from high to low in confidence can be selected as the disease combination to be matched. Each error sample corresponds to one pending disease combination.

[0102] Step S34: Count the number of each pending disease combination.

[0103] Step S35: Sort the number of each pending disease combination from high to low.

[0104] Step S36: selecting at least one of the plurality of pending disease combinations that is ranked high as the preset disease combination.

[0105] It is understandable that the pending disease combinations that are ranked higher among the multiple pending disease combinations are relatively weak and error-prone parts of the model prediction, so they can be used as the preset disease combinations to be corrected next.

[0106] In this embodiment, the purpose is to use the pending disease combination with a larger number in multiple erroneous samples as the preset disease combination. The specific method of obtaining the preset disease combination is not limited to the method provided in the above embodiment.

[0107] See also Figures 7 to 8 , Figure 7 1 is a flow chart of an embodiment of a method for obtaining a confusing score table of the present application. Figure 8 yes Figure 7 A flow chart of another embodiment of step S45 is shown in FIG. Figure 9 This is another flowchart of an embodiment of a method for obtaining an easily confused score table of the present application.

[0108] like Figure 7 As shown, the confusing score table is obtained by the following steps:

[0109] Step S41: Obtain a disease training set of multiple training samples, wherein the disease training set includes at least two disease types sorted by confidence, and the training samples include at least one symptom feature and a real disease type.

[0110] Different from the above embodiments, in this embodiment, since the training samples further include the true disease types for subsequent verification of the disease training set to determine the training samples with prediction errors and the training samples with correct predictions, and then the symptom features of these training samples are extracted respectively and calculated comprehensively to form a confusion score table.

[0111] For the description of this step, reference can be made to step S31 in the above embodiments, which will not be elaborated here.

[0112] Step S42: From the multiple training samples corresponding to each preset disease combination, select the training samples with prediction errors and whose corresponding true disease types are included in the preset disease combination to form an error sample set, and select the training samples with correct predictions to form a correct sample set.

[0113] As Figure 9 shown, when the top1 result in the disease training set of the training sample is different from the true disease type, it is determined that the prediction is incorrect, and then the training sample is determined as an error sample. For the error sample, it can be further determined whether the true disease type of the error sample is included in the corresponding preset disease combination. If the true disease type is included, the training sample can be classified into the error sample set. In addition, the training samples with correct predictions can be directly determined as correct samples and classified into the correct sample set.

[0114] As Figure 9 shown, for example, a preset disease combination is bronchial asthma and pneumonia. Thus, multiple training samples with the pending disease combination of bronchial asthma and pneumonia can be first selected from the training sample set, and then from the selected multiple training samples, select the training samples with prediction errors but the second highest confidence disease in the disease prediction set equal to the true disease type (i.e., the true disease type is included in the preset disease combination) to form an error sample set, and select the training samples with correct predictions to form a correct sample set. It can be found that both training sample 2 and training sample 3 have incorrect predictions, but the true disease type "rhinitis" is not included in the corresponding preset disease combination ("bronchial asthma" and "pneumonia") of training sample 3. Therefore, training sample 2 is classified into the error sample set, while training sample 3 is not classified into the error sample set. In addition, training sample 1 with correct prediction is classified into the correct sample set.

[0115] Step S43: Determine all the symptom features in the error sample set and all the symptom features in the correct sample set.

[0116] In some embodiments, the manner of selecting symptom features can be determined according to specific requirements. In some embodiments, when selecting a combination of multiple symptoms as a symptom feature, the multiple symptoms can be sorted in a certain order (such as lexicographical order) as the unique and fixed expression of the symptom feature.

[0117] Among them, the symptoms of each training sample in the error sample set or the correct sample set can be permuted and combined to obtain multiple symptom features. Alternatively, the user can select a combination of one or more symptoms as a symptom feature according to actual needs.

[0118] It can be understood that the error score table and the correct score table can include the same symptom features or different symptom features. According to Figure 9 Judging from the first three symptom features shown in the table, both the error score table and the correct score table include symptom feature 1. In addition, the error score table further includes symptom feature 2 and symptom feature 3, and the correct score table further includes symptom feature 4 and symptom feature 5.

[0119] Step S44: Calculate the probability of occurrence of each symptom feature in the error sample set as the first score of the symptom feature, and calculate the probability of occurrence of each symptom feature in the correct sample set as the second score of the symptom feature.

[0120] Specifically, the number of occurrences of each symptom feature in the error samples corresponding to the error sample set can be counted and normalized (i.e., divided by the sum of the number of occurrences of all symptom features in the error sample set), and the obtained value is the first score of each symptom feature in the current preset disease combination pair. Then, each symptom feature and its first score can be put into one-to-one correspondence to form an error score table for use in generating the subsequent confusion score table.

[0121] Similarly, the number of occurrences of each symptom feature in the correct samples corresponding to the correct sample set can be counted and normalized (i.e., divided by the sum of the number of occurrences of all symptom features in the correct sample set), and the obtained value is the second score of each symptom feature in the current preset disease combination pair. Then, each symptom feature and its second score can be put into one-to-one correspondence to form a correct score table for use in generating the subsequent confusion score table.

[0122] Step S45: Calculate the comprehensive score of the symptom feature according to the first score and the second score of each symptom feature.

[0123] Specifically, when the symptom feature is included in the correct sample set, the comprehensive score of the symptom feature is equal to the difference between the first score and the second score of the symptom feature; when the symptom feature is not included in the correct sample set, the comprehensive score of the symptom feature is equal to the first score of the symptom feature.

[0124] Such asFigure 8 As shown, in some embodiments, step S45 may specifically include steps S451 to S453:

[0125] Step S451: Determine whether the symptom feature is in the correct sample set.

[0126] If so, execute step S452.

[0127] Otherwise, execute step S453.

[0128] Step S452: Calculate the difference between the product of the first score and the first weight and the product of the second score and the second weight as the comprehensive score of the symptom feature.

[0129] Step S453: Calculate the product of the first score and the first weight as the comprehensive score of the symptom feature.

[0130] Specifically, the comprehensive score of the symptom feature can be calculated using formula (2) as follows:

[0131] In formula (2), S final The comprehensive score of a symptom feature (F) in the confusion score table, which is also the final score of this symptom feature. S error Represents the first score of the symptom feature (F) in the incorrect sample set, S correct Represents the second score of the symptom feature (F) in the correct sample set, α error Represents the weight of the first score and α correct Represents the weight of the second score. Among them, α error And α correct The values of can be selected according to actual needs and are not limited here.

[0132] It can be understood that the higher the comprehensive score of the symptom feature, the greater the degree of error-proneness of this symptom feature, and the lower the comprehensive score of the symptom feature, the smaller the degree of error-proneness of this symptom feature. Therefore, the principle of making the confusion score table is that the more times a certain symptom feature appears in the incorrect sample set and the fewer times it appears in the correct sample set, the higher the comprehensive score of this symptom feature, and vice versa, the lower the comprehensive score of this symptom feature.

[0133] Step S46: Associatively record each symptom feature and its corresponding comprehensive score in the confusion score table.

[0134] Such as Figure 9 The confusion score table shown includes all symptom features in the incorrect sample set and the correct sample set, that is, symptom feature 1 to symptom feature 5, etc. The scores recorded in the confusion score table for each symptom feature refer to the comprehensive score of the symptom feature.

[0135] Please refer to Figures 10 to 11 , Figure 10 which is a schematic flowchart of an embodiment of the method for obtaining the preset score threshold in the present application, Figure 11 and which is another schematic flowchart of an embodiment of the method for obtaining the preset score threshold in the present application.

[0136] As Figure 10 shown, the preset score threshold is obtained through the following steps:

[0137] Step S51: Obtain at least one symptom feature of a plurality of test samples corresponding to a preset disease combination.

[0138] As Figure 12 shown, in this embodiment, the test samples are different from the training samples used to calculate the confusion score table in the above embodiment, and in other embodiments, they may also be the same. The test samples can be medical records, or texts in the medical records that record the patient's symptoms and the true disease type. It can be understood that the test samples are similar to the samples to be detected and the training samples, for example, all include at least one symptom feature.

[0139] In some specific embodiments, there are 100 medical records in a batch. 80 medical records can be used as training samples to calculate the confusion score table, and the other 20 medical records can be used as test samples to calculate the preset score threshold and the accuracy rate of the disease test set after test optimization.

[0140] It can be understood that each preset disease combination corresponds to a preset score threshold. Therefore, in this embodiment, by processing at least one symptom of a plurality of test samples corresponding to the preset disease combination, the preset score threshold corresponding to the preset disease combination can be obtained.

[0141] In some embodiments, the highest confidence level in the disease test set of the test samples is less than the preset confidence level threshold, that is, only the test samples with the highest confidence level less than the preset confidence level threshold are used to calculate the preset confidence level threshold. The purpose is to exclude the samples that the model predicts with great certainty and not modify them twice.

[0142] Step S52: Use the confusion score table of the symptom features to score each of the plurality of test samples respectively.

[0143] This step is similar to the step of scoring the sample to be detected using the confusion score table of the symptom features in the above embodiment. Therefore, the description of this step can be referred to the corresponding position in the above embodiment and will not be repeated here.

[0144] Step S53: Calculate the sum of the mean and standard deviation of the scores of the plurality of test samples according to the score of each test sample, and use it as the preset score threshold corresponding to the preset disease combination.

[0145] Specifically, the formula (3) can be used to calculate the preset score threshold corresponding to the preset disease combination as follows:

[0146]

[0147] In formula (3), g thread represents the preset score threshold, g i represents the score of the test sample, n represents the number of test samples, and β represents the weight of the standard deviation.

[0148] Among them, β can be selected according to actual needs. The larger β is, the larger g thread is, so the number of error-prone samples found is less, but it can ensure that the accuracy of the error-prone samples found is higher, that is, a larger proportion of the samples are indeed mispredicted. Therefore, the value of β can be selected according to actual needs.

[0149] Among them, the preset score threshold represents the normal range of the scores of multiple test samples in the preset disease combination. If it exceeds the preset score threshold, it means that the test sample is very similar to the samples that made mistakes in the training samples. Therefore, it can be determined that the test sample is an error-prone sample, and thus the confidence ranking corresponding to the test sample needs to be adjusted. For example, the disease corresponding to the top1 result is modified to another disease type in the preset disease combination.

[0150] In some embodiments, the score of the test sample can also be used to determine whether the test sample is an error-prone sample. Specifically, the score of the test sample can be compared with the corresponding preset score threshold. If the score of the test sample is greater than the preset score threshold, it is determined that the test sample is an error-prone sample; otherwise, it is determined that the test sample is a normal sample. Then, the confidence ranking in the disease test set of the determined error-prone samples is adjusted, and then the adjusted disease test set is output. Finally, the disease test set is compared with the true disease type, and thus the accuracy of the optimized disease test set can be judged. Further, the preset score threshold can also be adjusted according to the accuracy of the disease test set. For example, the value of β can be adjusted. By continuously adjusting the value of β, the value of β corresponding to the highest accuracy of the disease test set is found, and the preset score threshold is calculated using this β so that a higher accuracy can be obtained when using the preset score threshold in subsequent actual use.

[0151] Such as Figure 11As shown, the test sample set includes a total of 24 test samples. The disease combinations to be determined (such as the top 2 combination) of the 24 test samples match a preset disease combination. After preliminary screening, 15 test samples with a confidence level lower than the preset confidence threshold (such as 0.85) can be obtained. The confidence levels of the top 1 results of these test samples are relatively low. Therefore, the scores of each test sample can be calculated using the confusion score table, and the preset score threshold 0.5412 corresponding to the preset disease combination can be calculated using the scores of the 15 test samples. Then, the preset score threshold can be compared with the scores of the test samples. The test sample 1 and test sample 2 with scores greater than 0.5412 are determined as error-prone samples, and the error-prone samples need to be re-sorted. The other test samples with scores less than or equal to 0.5412 are determined as normal samples, and the normal samples can directly output the current results.

[0152] Please refer to Figure 12 , Figure 12 which is a schematic flow chart of the training and testing phases of this application.

[0153] At the beginning, multiple training samples and multiple test samples are input into the disease prediction model for model prediction to obtain multiple disease training sets and multiple disease test sets respectively. Then, based on the multiple disease training sets, big data analysis can be performed to determine which disease types are easily confused diseases, so as to determine the preset disease combination (which can refer to steps S31 - S36 in the above embodiment correspondingly). After that, the correct score table and the wrong score table corresponding to each preset disease combination are calculated, and the confusion score table is calculated based on the correct score table and the wrong score table (which can refer to steps S41 - S46 in the above embodiment correspondingly). After obtaining the confusion score table, the scores of each test sample corresponding to the preset disease combination can be obtained using the confusion score table, the preset score threshold corresponding to the preset disease combination is calculated, and then the error-prone samples are determined using the preset score threshold, and the confidence levels of the disease test sets of the error-prone samples are re-sorted, and then the disease test sets are output (which can refer to steps S51 - S53 in the above embodiment correspondingly).

[0154] When the related technology makes disease prediction based on data mining, the general idea is to find features such as symptoms that are strongly related to each disease, and then perform field matching with medical records. However, it is difficult to process complex information through data mining, and the idea is relatively simple, making it difficult to further optimize. In this embodiment, disease prediction is based on a model, which has the ability to process complex long texts, is more adaptable to the actual situation, and there are more diverse optimization schemes. On this basis, a method related to the model training result is used for further optimization. This method not only retains the performance advantages of the model itself but also makes the most use of the training data of the model.

[0155] Different from the conventional disease prediction solutions based on models, this application combines the idea of data mining to further improve performance. It separately extracts the features of a preset disease combination, breaking down the big problem of improving the prediction accuracy of the disease prediction set into small problems of distinguishing multiple diseases in the preset disease combination. This not only breaks through the limit of the prediction accuracy of the current disease prediction set and further improves the accuracy of the disease prediction set, but also has a certain degree of interpretability. In addition, this application uses the data obtained during the training of the machine learning model to solve the shortcomings of the data mining method itself, such as being single-line and having limited thinking. Therefore, it can find more in-depth and essential features in the medical records for prediction, rather than being limited to finding the combination of one or more symptoms as the gold standard for prediction like ordinary data mining methods, resulting in poor accuracy.

[0156] In addition, this application provides an adaptive error correction mechanism that can set different error-prone sample judgment conditions according to the disease test set of each preset disease combination (that is, each preset disease combination has a corresponding preset score threshold), so that each preset disease combination can obtain the optimal output result. In some embodiments, the preset score threshold can be automatically generated according to the disease test set without the need to manually check the internal situation of each preset disease combination to select the threshold, which is faster and has a certain degree of interpretability. For example, from the perspective of the statistical meaning of the standard deviation, some samples with scores higher than the preset score threshold can be determined as error samples and re-sorted, which can improve the disease prediction effect of this part of the samples.

[0157] Please refer to Figure 13 , Figure 13 which is a schematic framework diagram of an embodiment of the disease prediction set processing device of this application.

[0158] As Figure 13 shown, the disease prediction set processing device 100 may include an acquisition module 110, a processing module 120, and an adjustment module 130. Among them, the acquisition module 110 is used to acquire the disease prediction set of the sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature. The processing module 120 is used to score the sample to be detected based on the confusion degree of at least one symptom feature to determine whether the sample to be detected is an error-prone sample. The adjustment module 130 is used to adjust the confidence ranking in the disease prediction set of the sample to be detected if the sample to be detected is an error-prone sample.

[0159] In some embodiments, the processing module 120 is also used to find the corresponding score of each symptom feature in the sample to be tested in the confusion score table, wherein the confusion score table records the score of each symptom feature, and the score of the symptom feature is used to reflect the degree of confusion of the symptom feature; calculate the sum of the scores corresponding to at least one symptom feature of the sample to be tested to obtain the score of the sample to be tested.

[0160] In some embodiments, the processing module 120 is further configured to determine whether the score of the sample to be detected is greater than a preset score threshold; if the score of the sample to be detected is greater than the preset score threshold, the sample to be detected is determined to be an error-prone sample.

[0161] In some embodiments, the preset score threshold is obtained in the following manner: wherein the acquisition module 110 is also used to obtain at least one symptom feature of multiple test samples corresponding to the preset disease combination; the processing module 120 also uses the confusing score table of the symptom features to score the multiple test samples respectively, and then calculates the mean and the sum of the standard deviation of the scores of the multiple test samples based on the score of each test sample as the preset score threshold corresponding to the preset disease combination.

[0162] In some embodiments, the highest confidence score in the disease test set of the test sample is less than a preset confidence threshold.

[0163] In some embodiments, the preset disease combination is obtained in the following manner: the acquisition module 110 is also used to obtain a disease training set of multiple training samples, wherein the disease training set includes at least two disease types sorted by confidence, and the training samples include at least one symptom feature and a true disease type; the processing module 120 is also used to determine the error samples in the multiple training samples based on the matching results of the disease training set and the true disease type of each training sample, and determine the pending disease combination of each error sample, wherein the pending disease combination is at least two disease types selected from the disease training set of the error sample in descending order according to confidence, the number of each pending disease combination is counted, the number of each pending disease combination is sorted from high to low, and at least one pending disease combination with a top ranking among the multiple pending disease combinations is selected as the preset disease combination.

[0164] In some embodiments, the processing module 120 is also used to determine a disease combination to be matched of at least two disease types in the disease prediction set, and then determine whether the disease combination to be matched matches the preset disease combination; if the disease combination to be matched matches the preset disease combination, the sample to be tested is scored based on the degree of confusion of at least one symptom feature.

[0165] In some embodiments, the confusing score table is obtained in the following manner: The obtaining module 110 is further configured to obtain a disease training set of a plurality of training samples, where the disease training set includes at least two disease types sorted by confidence, and the training samples include at least one symptom feature and a true disease type; The processing module 120 is further configured to select, from the multiple training samples corresponding to each preset disease combination, the training samples with incorrect predictions and where the corresponding true disease type is included in the preset disease combination to form an incorrect sample set, and select the training samples with correct predictions to form a correct sample set, then determine all the symptom features in the incorrect sample set, and determine all the symptom features in the correct sample set, then calculate the probability of occurrence of each symptom feature in the incorrect sample set as the first score of the symptom feature, and calculate the probability of occurrence of each symptom feature in the correct sample set as the second score of the symptom feature, then calculate the comprehensive score of the symptom feature according to the first score and the second score of each symptom feature, and finally associate and record each symptom feature and the corresponding comprehensive score in the confusing score table.

[0166] In some embodiments, the processing module 120 is further configured to determine whether the symptom feature is in the correct sample set; if so, calculate the difference between the product of the first score and the first weight and the product of the second score and the second weight as the comprehensive score of the symptom feature; otherwise, calculate the product of the first score and the first weight as the comprehensive score of the symptom feature.

[0167] For the descriptions of the above steps, reference can be made to the corresponding positions in the foregoing method embodiments, and details are not described herein again. It can be understood that the above module division is only a logical function division, and there may be other division methods in actual implementation.

[0168] Please refer to Figure 14 , Figure 14 which is a schematic framework diagram of an embodiment of the electronic device of the present application.

[0169] As Figure 14 shown, the electronic device 200 includes a processor 210 and a memory 220 connected to the processor 210. The memory 220 is used to store program data, and the processor 210 is used to execute the program data to implement the steps in any of the foregoing method embodiments.

[0170] The electronic device 200 includes, but is not limited to, a television, a desktop computer, a laptop computer, a handheld computer, a wearable device, a head-mounted display, a reader device, a portable music player, a portable game console, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cell phone, a personal digital assistant (PDA), an augmented reality (AR) device, and a virtual reality (VR) device.

[0171] Specifically, the processor 210 is used to control itself and the memory 220 to implement the steps in any of the above method embodiments. The processor 210 may also be referred to as a CPU (Central Processing Unit). The processor 210 may be an integrated circuit chip with signal processing capabilities. The processor 210 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. Additionally, the processor 210 may be implemented by multiple integrated circuit chips together.

[0172] Please refer to Figure 15 , Figure 15 which is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application.

[0173] As Figure 15 shown, the computer-readable storage medium 300 stores program data 310, and when the program data 310 is executed by the processor, it is used to implement the steps in any of the above method embodiments.

[0174] The computer-readable storage medium 300 may specifically be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store computer programs, or it may also be a server storing the computer program. The server can send the stored computer program to other devices for running, or it can also run the stored computer program by itself.

[0175] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in electrical, mechanical, or other forms.

[0176] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0178] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0179] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for processing a disease prediction set, characterized in that include: Obtaining a disease prediction set of a sample to be tested, wherein the disease prediction set includes at least two disease types sorted by confidence, and the sample to be tested includes at least one symptom feature; Finding the corresponding score of each symptom feature in the sample to be detected in the confusion score table, wherein the confusion score table records the score of each symptom feature, and the score of the symptom feature is used to reflect the degree of confusion of the symptom feature; Calculating the sum of the scores corresponding to at least one symptom feature of the sample to be tested to obtain the score of the sample to be tested; Determine whether the score of the sample to be detected is greater than a preset score threshold; If the score of the sample to be detected is greater than a preset score threshold, the sample to be detected is determined to be an error-prone sample; If the sample to be detected is an error-prone sample, the confidence ranking of the sample to be detected in the disease prediction set is adjusted.

2. The method according to claim 1, wherein The preset score threshold is obtained in the following manner: Obtaining at least one symptom feature of a plurality of test samples corresponding to a preset disease combination; Using the confusing score table of symptom characteristics, scoring the plurality of test samples respectively; According to the score of each of the test samples, the sum of the mean and the standard deviation of the scores of the multiple test samples is calculated as the preset score threshold corresponding to the preset disease combination.

3. The method according to claim 2, wherein The highest confidence in the disease test set of the test sample is less than a preset confidence threshold.

4. The method according to claim 2, wherein The preset disease combination is obtained in the following manner: Acquire a disease training set of multiple training samples, wherein the disease training set includes at least two disease types sorted by confidence, and the training samples include at least one symptom feature and a real disease type; Determine an erroneous sample among the plurality of training samples according to a matching result between the disease training set and the real disease type of each training sample; Determine a pending disease combination for each of the erroneous samples; wherein the pending disease combination is at least two disease types selected from the disease training set of the erroneous samples in descending order according to confidence; Counting the number of each of the pending disease combinations; Sort the number of each of the pending disease combinations from high to low; At least one of the plurality of pending disease combinations that is ranked high is selected as the preset disease combination.

5. The method according to claim 1 or 2, characterized in that, The confusing score table is obtained by: Acquire a disease training set of multiple training samples, wherein the disease training set includes at least two disease types sorted by confidence, and the training samples include at least one symptom feature and a real disease type; From the plurality of training samples corresponding to each preset disease combination, the training samples with incorrect predictions and containing the corresponding real disease type in the preset disease combination are selected to form an error sample set, and the training samples with correct predictions are selected to form a correct sample set; Determine all of the symptom features in the error sample set, and determine all of the symptom features in the correct sample set; Calculate the probability of occurrence of each of the symptom features in the error sample set as the first score of the symptom feature, and calculate the probability of occurrence of each of the symptom features in the correct sample set as the second score of the symptom feature; Calculate the comprehensive score of the symptom feature according to the first score and the second score of each of the symptom features; Associate and record each of the symptom features and the corresponding comprehensive score in the confusion score table.

6. The method according to claim 5, characterized in that, The calculating the comprehensive score of the symptom feature according to the first score and the second score of each of the symptom features includes: Determine whether the symptom feature is in the correct sample set; If so, calculate the difference between the product of the first score and the first weight and the product of the second score and the second weight as the comprehensive score of the symptom feature; Otherwise, calculate the product of the first score and the first weight as the comprehensive score of the symptom feature.

7. The method according to claim 1, characterized in that Before scoring the sample to be detected based on the confusion degree of the at least one symptom feature, it further includes: Determine the disease combination to be matched of at least two disease types in the disease prediction set; Determine whether the disease combination to be matched matches the preset disease combination; If the disease combination to be matched matches the preset disease combination, score the sample to be detected based on the confusion degree of the at least one symptom feature.

8. A disease prediction set processing device, characterized in that, It includes: An acquisition module, configured to acquire a disease prediction set of a sample to be detected, where the disease prediction set includes at least two disease types sorted by confidence, and the sample to be detected includes at least one symptom feature; A processing module, configured to find the score corresponding to each symptom feature in the sample to be detected in the confusion score table, where the confusion score table records the scores of each of the symptom features, and the scores of the symptom features are used to reflect the confusion degree of the symptom features; calculate the sum of the scores corresponding to at least one symptom feature of the sample to be detected to obtain the score of the sample to be detected; determine whether the score of the sample to be detected is greater than a preset score threshold; if the score of the sample to be detected is greater than the preset score threshold, determine that the sample to be detected is an error-prone sample; An adjustment module, configured to adjust the confidence ranking in the disease prediction set of the sample to be detected if the sample to be detected is an error-prone sample.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory connected to the processor, The memory is used to store program data, and the processor is used to execute the program data to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The program data is stored in the computer-readable storage medium, and when the program data is executed by the processor, it is used to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Adjuvant disease diagnosis method based on patient test results

    CN107066791A